* feat(runtime): allow Bun runtimes that provide node:sqlite
* fix(process): drop execa buffer encoding under Bun spawn (Bun rejects non-spawn options)
* chore(process): cite oven-sh/bun#36049 in bun spawn workaround
* docs(install): bun with node:sqlite can run openclaw; bun install workspace caveat
* fix(process): clear execa buffer encoding under Bun without mutating read-only options
Three defects that compounded into a test suite nobody could run and nobody
was running.
The suite has been red on main since 2026-07-20. `connect_frame_matches_gateway_schema`
asserts a TLS-pinned connection advertises no capabilities, but #111933 wrote
that assertion while caps held only inline-widgets, and #111920 made
`agent-kind` unconditional the same day. No textual conflict, so both landed
and the assertion has been wrong ever since. Pinning only withdraws inline
widgets, so assert exactly that.
`cargo test` could not run on macOS at all: tauri-plugin-notifications links a
Swift static library, nothing adds an rpath for the Swift runtime, and every
test binary aborted at load with `Library not loaded:
@rpath/libswift_Concurrency.dylib`. Emit the rpath from build.rs.
Neither surfaced because linux-app.yml never ran `cargo test` - it only checked
formatting and built bundles. Run the suite on Linux, and upgrade the macOS job
from `check` to `test` so link-time breakage like the rpath is caught at all;
`check` never links, so it cannot see this class of failure.
* refactor(channels): declare thread addressing as a channel trait
* fix(tasks): require declared thread capability for direct parent-review delivery
* docs(tasks): record the no-target-parsing tradeoff at the delivery gate
* fix(linux): restore the macOS build of the Tauri companion
Tauri gates `WebviewWindowBuilder::transparent` behind its `macos-private-api`
feature on macOS, so the companion stopped compiling for macOS when Quick Chat
landed in #109947: `no method named 'transparent' found for struct
'WebviewWindowBuilder'`. Enable that feature together with the matching
`app.macOSPrivateApi` config flag, which Tauri requires to agree with it or
tauri-build fails the allowlist check.
Quick Chat is built for a transparent window - quickchat.css makes html and
body transparent behind a 16px-radius translucent card with a drop shadow - so
compiling macOS by dropping `.transparent(true)` would ship an opaque rectangle
instead of the floating composer.
Also add a per-PR macOS `cargo check`. The only job that compiled the macOS
target sits behind a manual dispatch input in linux-app-release.yml, which is
why a broken macOS build survived for nine days.
* ci(linux): build the macOS companion on macos-15
tauri-plugin-notifications' Swift sources use typed throws
(`throws(FFIResult)`), which requires Swift 6. The macos-14 runner image still
ships Swift 5.x and cannot parse them, so the plugin's build script panics with
"Swift build failed for target: arm64-apple-macosx13.0".
This hit the new per-PR check immediately, and it means the release workflow's
`build_macos` job could not have produced a macOS bundle either - it compiles
the same crate on the same image, but only runs behind a manual dispatch input
so nothing surfaced it. Move both to macos-15.
* perf(sessions): stop per-row full-store scans in ACP meta reads
Every per-session ACP meta read listed and JSON-parsed the entire session
store just to find one row, making gateway sessions.list O(rows^2). A CPU
profile on team.openclaw.ai showed 12.7s of a 78.5s window inside
parseSqliteSessionEntryJson via buildGatewaySessionRow -> readAcpSessionMeta.
Resolve the store key with exact read-only single-row probes (exact key,
then lowercased key) and keep the full scan only as the fallback for legacy
case-variant keys. Adds the missing read-only exact-probe accessor next to
loadExactSessionEntry.
* refactor(acp): split session store binding out of session-meta
max-lines pushed session-meta.ts past 700; the store-binding helpers
(resolveStoreEntryForSessionKey, resolveSessionStorePathForAcp,
readSessionEntryFromStore) are one cohesive unit and move to
session-meta-store.ts. Re-export keeps the public entrypoint stable.
* perf(sessions): case-variant fallback scans keys, not parsed entries
Review caught that a per-row miss still parsed the full store. The fallback
only needs persisted keys to find a legacy case-variant, so list keys without
touching entry_json and probe the single winner.
* feat(ui): complete the Labs roster with the remaining experimental gates
Labs shipped with two entries while three more experimental surfaces stayed
reachable only by hand-editing config. Adds Tool Search, lean local-model tools,
and message audit metadata, each with the subtext and documentation link the
page already promises.
Two of them do not fit the boolean gate the registry assumed, so the registry
now states each row's on/off values instead of hardcoding true/false:
- Tool Search is a mode, and resolveToolSearchConfig defaults an unset mode to
"code" even in object form. A bare enable would therefore select the surface
with the weakest recall rather than the bounded directory this row advertises,
so enabling writes `mode: "directory"` alongside the gate in one patch.
- Message audit is `off | direct | all`. Labs offers the conservative `direct`,
so turning it on cannot start recording group or unknown conversations the
operator never opted into, and `all` deliberately does not read as on — that
is a broader choice made elsewhere and this row must not quietly narrow it.
The audit row carries a restart hint because startGatewayEventSubscriptions
resolves the mode once and bakes it into the recorder, which outlives the reload
plan's `logging: none` rule. The other four are read per agent run.
* fix(ui): read a broader audit mode as enabled instead of narrowing it
`all` records more than the `direct` this row offers, but it is still on. Strict
equality against onValue rendered it as off, which made the switch look available
and would have silently rewritten a deliberately broader operator setting down to
`direct` on the next click.
Separate "what enabling writes" from "what counts as enabled": onValue stays the
value Labs sets, activeValues lists every value that reads as on. Turning the row
off from `all` now writes `off`, which is the only narrowing the operator asked
for.
The test that covered this asserted the opposite of what its own comment
described; corrected, plus coverage for the off-from-all write.
* fix(ui): mirror the runtime rule for Tool Search enablement
ToolSearchConfig is a union, and resolveToolSearchConfig treats an object that
configures anything besides `enabled` as already on:
`readBoolean(raw.enabled, configured)`. Reading only the `enabled` leaf showed
`{ mode: "tools" }` as off, so the row offered a switch that would have replaced
that operator's mode with `directory` — the same narrowing the audit row was
just fixed for.
Give the registry a `readEnabled` override for gates whose enablement is not a
single leaf, and point Tool Search at a copy of the runtime rule with the
resolver named so the two stay comparable. The other four rows keep the leaf
read and declare `readEnabled: null` explicitly.
* chore(ui): keep LabFeatureValue local to the registry
Only the registry names the type, and the hard-zero Knip production scan rejects
an export with no production consumer outside its own module.
* fix(scripts): run changed checks locally when Blacksmith never ran them
AGENTS.md already says trusted-source work falls back to local execution when
the remote backend is unavailable, but the tooling did not implement it: a lease,
broker, DNS, or network failure surfaced as a plain exit 1, so the lanes were
reported red without ever having been evaluated. That is worse than slow — a run
that never happened looked the same as a run that failed.
Tee the wrapper output and use its run summary as the discriminator. The summary
only appears once the command reached the box, so a failure carrying
`command-exit` is a real verdict and propagates unchanged; anything else never
produced one and re-runs locally, with a loud note so the proof summary records
which machine produced it.
Deliberately a positive test for `command-exit` rather than a blocklist of
infrastructure errors. Guessing wrong toward "infrastructure" only re-runs the
checks locally; guessing wrong toward "real failure" would block on an outage.
It must never widen to "fall back on any non-zero exit" — prompt snapshots are
Linux-only truth and would pass locally on macOS, turning a red gate green.
The sparse-checkout path keeps no fallback: it exists precisely because the
checkout cannot resolve the diff refs, so there is nothing local to run.
* fix(scripts): require positive pre-dispatch evidence before falling back
A missing run summary does not prove the remote never started: a wrapper that
crashes or loses its output transport after dispatch looks identical. Reading
that absence as "never ran" would rerun locally and could turn an unknown or
failing Linux-only lane green, which is the exact masking this guard exists to
prevent.
Require a positive pre-dispatch signature instead, and keep the command-exit
veto. Also reapply backpressure on the tee: inherited stdio got it from the OS,
piping does not, so a verbose delegated run could buffer its whole output here.
* fix(linux): draw the macOS tray icon from a template silhouette
AppKit renders menu bar template images from the alpha channel alone, so the
opaque rounded-tile 32x32.png painted a solid white square instead of the
mascot. Add a dedicated tray-template.svg and its 36px render, a silhouette
with the eyes knocked back out, and select it only on macOS. Its geometry
mirrors the native macOS app's CritterIconRenderer at rest so both clients
wear the same face; other platforms keep the full-color 32x32.png.
* chore(linux): drop the root changelog entry from the tray icon fix
CHANGELOG.md is release-owned in this repo (AGENTS.md) and `scripts/pr
prepare-run` rejects normal PRs that touch it; release generation derives
entries from merged PRs. The release-note context lives in the PR body and
the preceding commit message instead.
A non-main agent store deliberately does not persist an OAuth credential the
main store already owns at the same or newer expiry. The SQLite migration
verified its write by reloading and comparing every imported profile, so that
intentional dedup read as data loss: verification threw, the migration aborted,
and the legacy auth-profiles.json stayed on disk.
The runtime then refuses to start while any legacy credential file exists, and
the error it prints tells the operator to run openclaw doctor --fix, which is
what had just failed. That is a boot loop with circular remediation.
Treat a profile the local store intentionally deduped to main as verified
rather than missing.
* build(deps): drop stale root partial-json and dead ownership-manifest entries
partial-json is owned by packages/ai (which declares its own copy); glob and
markdown-it no longer exist as root dependencies. Audit follow-up to #114006.
* build(deps): keep root partial-json (openclaw/ai bundling contract); drop only dead manifest entries
* refactor(prompt): plain inbound context labels with a provenance marker
Replaces trust-worded inbound context labels ("(untrusted metadata)",
"(untrusted, for context)") with plain labels plus a fixed provenance
marker suffix appended to every OpenClaw-injected context header.
Detection keys on the marker, not label text, so strippers stay correct
across UI, TUI, replay, /trace segmentation, memory recall, and the Swift
chat preprocessor. Drops sanitizeInboundSystemTags in favor of the marker
boundary plus trusted system-prompt narration.
Renames the untrusted-named plugin SDK context identifiers to
channel-provenance names, keeping deprecated aliases registered for
removal after 2026-09-08.
Adds `openclaw doctor --fix` migrations that rewrite legacy inbound
labels in stored SQLite transcripts and purge legacy envelope-
contaminated LanceDB recall rows.
* fix(ci): resolve gate failures for plain inbound context labels
- doctor sqlite readers: open read-only connections via openNodeSqliteDatabase
so the Kysely connection-boundary guardrail holds; unexport the now-internal
transcript snapshot type (Knip unused-export gate).
- compat registry: split the record table into registry-records.ts and
plugin-sdk-subpath-records.ts. The new compat record pushed registry.ts past
the 700-line oxlint cap; suppressions are disallowed, so follow the existing
sibling record-module pattern. Public exports and PluginCompatCode literals
unchanged.
- acp-runtime test: assert current finalization behavior (newline normalization
only). The bracket de-fang and System: rewrite it expected were removed with
sanitizeInboundSystemTags; forged system lines are neutralized at the
system-event queue, the single chokepoint feeding the System:-per-line render.
- regenerate docs_map and the plugin SDK API baseline manifest.
* fix(prompt): harden inbound context label migration and drop in-band sanitizer
Review follow-ups on the plain-label + provenance-marker change:
- Remove src/security/system-tags.ts. Rewriting inbound text to neutralize
look-alike `System:`/`[System]` markers corrupted legitimate user text and is
not a real injection boundary; role separation plus external-content wrapping
is. Explicit product decision, recorded at the system-event queue.
- Narrow the LanceDB legacy-row purge so it cannot delete benign memories. It
now requires a complete known legacy sentinel line, a legacy label followed by
a fenced JSON body, or the complete legacy external-content header. The prior
predicates matched ordinary prose such as `Notes (untrusted metadata):`, and
deletion is irreversible.
- Make explicit-empty canonical ChannelStructuredContext win over the deprecated
alias via a present/absent result instead of collapsing `[]` to undefined.
- Keep `\r?` in the active-memory doctor rule. It is the only rule spanning the
header's line break, migrated assistant rows skip newline normalization, and
without it the marked-header replace wins and the body strips to empty. Added
a CRLF regression test.
- Fix stale comments that described removed behavior, and cover the Swift
prose-block strip path.
Claude-Session: https://claude.ai/code/session_01WNzsPddQmxy9Y7jKD4wAxH
* feat(ui): add a Memory settings page with Dreaming as a tab
Memory config was scattered across five surfaces: the memory.* schema section
lived on AI & Agents with 43 of 51 keys behind the Advanced tier, the memory
slot owner was only visible on Plugins, dreaming's knobs were JSON-only, its
status UI sat under Agents, and Memory Import was a separate route.
/settings/memory now owns that surface, following the MCP page shape (curated
rows above an embedded schema editor):
- Overview: the exclusive memory slot rendered as a segmented control over
installed memory-kind plugins, memory.backend promoted out of Advanced with
the qmd sub-config revealed only when qmd is selected, additive add-on rows,
and a Memory Import link.
- Search: the memory.search surface via the embedded editor.
- Dreaming: the global frequency/model/timezone/storage/phase knobs, which
previously required hand-editing openclaw.json, plus an agent picker feeding
the existing dream scene/diary/advanced panel for the agent-scoped reads.
Engine selection calls plugins.setEnabled so the gateway's exclusive slot
policy stays the single owner instead of being duplicated in the UI.
* fix(ui): redirect stale ai-agents memory deep links to the memory page
* fix(ui): report memory runtime defaults on the Memory page
The Dreaming tab rendered its own defaults instead of the ones
resolveMemoryDreamingConfig applies, so a config carrying only
dreaming.enabled showed all three phases off while they were running, and
an unset storage mode read as inline instead of separate. Toggle specs now
carry the runtime fallback and the storage default is stated once, both
pointing at src/memory-host-sdk/dreaming.ts.
Three more surfaces asserted things the runtime does not do:
- plugins.slots.memory "none" is the explicit-off sentinel, not an engine
id, so the segmented control selected nothing. The slot now resolves to a
closed auto/off/pinned selection with its own hint.
- memory.backend is resolved by the memory runtime the slot owner
registers, which only memory-core ships, so the row is hidden for any
other engine instead of saving a value nothing reads.
- The Dreaming tab wrote config.dreaming for whichever plugin owns the
slot even when that plugin's schema cannot hold it. It now reuses the
enablement flow's schema check (resolveDreamingConfigPathSupport, shared
with updateDreamingEnabled) and renders an unsupported state instead.
Also key the plugin-catalog sync on the connected phase: the connecting ->
connected transition keeps the same client object, so a page mounted
during the handshake never loaded the catalog and never showed the engine
picker.
The tab keeps the autosave status line and restart banner the embedded
editor renders on the other tabs; these knobs autosave, but nothing
reported it. The pure view moved to memory-dreaming.ts with the element in
memory-dreaming-page.ts, matching memory.ts/memory-page.ts.
* fix(ui): resolve the memory slot through the canonical policy
The Memory page re-derived plugins.slots.memory instead of using the rule the
runtime applies, which broke both directions of the engine control:
- An unset slot was reported as "the first enabled memory-kind plugin in the
catalog". The runtime resolves it to the slot's default owner
(DEFAULT_SLOT_BY_KEY.memory), so the page could show one engine as active
while another was loaded, reveal or hide the backend row for the wrong
plugin, and target the wrong plugin when switching memory off.
- Off called plugins.setEnabled(false), which writes enablement only. The slot
stayed pinned, so the choice did not survive a refresh and re-enabling that
plugin from the Plugins page silently switched memory back on.
resolveSlotSelection now lives next to defaultSlotIdForKey in
src/plugins/slots.ts and owns the rule once; config normalization consumes it
and the page imports it instead of restating it. Off writes the explicit "none"
sentinel through the config form, so it round-trips; picking an engine still
goes through plugins.setEnabled, which is where the exclusive slot policy
lives. The dreaming controller's own copy of the rule is gone too.
Four smaller fixes on the same surface:
- A failed engine change is reported next to the control instead of being
swallowed, so the selector no longer just snaps back.
- Dreaming's numeric inputs carry the memory-core manifest's integer/min/max
bounds and refuse out-of-range edits at the field, rather than patching a
value autosave then fails to write.
- Settings search destinations carry the Memory tab that renders the matched
child, so a memory.search hit no longer lands on Overview, whose narrowed
editor omits it.
- The Dreaming tab caches only a definitive schema-capability answer. An
offline or failed lookup now reports "unknown" and is retried on reconnect
instead of permanently suppressing the recheck.
* fix(ui): model unknown memory state instead of collapsing it
The Memory page reported unknowns as decided values. An empty catalog meant
loading, disconnected, or a failed plugins.list, yet add-on rows rendered
"Disabled"; catalog completions were keyed on client identity, which survives a
phase flip, so a stale load could repopulate a disconnected page or overwrite a
newer read; and `?tab=` was adopted once per distinct value, so a repeat
navigation to a tab the user had left was ignored.
Replace the ad-hoc nullable fields with closed shapes. MemoryCatalog is a
loading/unavailable/ready union, so absence of an entry only decides anything
inside `ready`, and MemoryAddonRow carries a four-state enablement the view
renders without ever inventing an "off". CatalogConnection is one object per
(client, connected) transition and doubles as the request generation an
in-flight load carries, so obsolete completions are dropped by identity. The tab
is no longer page state at all: the URL owns it, tab clicks navigate, and every
arrival is honored.
Settings search now resolves the engine/backend through the same
resolveMemoryBackend the page uses and matches only the `memory.*` children the
page can surface, so a `memory.qmd` hit under the built-in backend no longer
routes to an Overview whose editor omits it.
* fix(ui): surface a disabled memory owner and anchor curated backend search
The slot and plugin enablement are independent config surfaces, so
`plugins.slots.memory` can name a plugin the catalog reports as disabled.
The engine control showed that plugin as selected, and because re-picking an
already-selected radio fires no change event, there was no way back on. Add an
explicit enable row for that state and let the same-id write through when the
owner is not running; picking Off stays a no-op.
`memory.backend` is curated out of the schema editor, so the generic
`#config-section-memory` anchor scrolled past it. Fold the memory tab and hash
choice into one `memoryDestination` owner that routes a curated-only match to
the new anchor above the editor.
* fix(ui): scope the dreaming capability probe to its connection
The probe was deduplicated by plugin id alone, which cannot tell a current
answer from a stale one. A disconnect and reconnect on the same slot owner left
the token armed, so the reconnect read as "already in flight" and swallowed the
retry that an `unknown` answer requires — leaving an unsupported engine's knobs
editable until some unrelated config notification arrived. An A -> B -> A switch
had the mirror problem: the old A response was accepted for the new A probe.
Make the in-flight probe an object whose identity is the generation, drop it
whenever the owner or the connection changes, and accept only the completion
that still owns the slot. Same shape as the catalog guard on the Memory page.
* fix(ui): satisfy the lint and dead-export gates on the memory page
Exhaustive switches need a terminal `default:` to satisfy
typescript/consistent-return, matching the existing view-status.ts shape.
Seven symbols were exported with no production consumer outside their own
module, which the hard-zero Knip production scan rejects. Tests alone do not
make internals contracts, so drop the exports and reach the behavior through
each module's public surface instead: the view props type comes from
`Parameters<typeof renderMemory>`, the tab panel is found by its ARIA role, and
the dreaming number/storage helpers are proven through `renderDreamingSettings`.
Folding those helper unit tests into the render path also corrected one of them:
a `type="number"` input coerces unparseable text to empty, so the "reject
garbage" case was unreachable through the real control. Replaced with the
inclusive-bound and clear-the-field cases, which are reachable.
* refactor(ui): keep the memory schema facts out of the startup bundle
Settings pages are already lazy — the config route is `import("./config-page.ts")`
— but settings search runs from app-host at startup, and it needed the same
answers about which `memory.*` children are reachable and where a match lives.
Importing those from the view module dragged lit, hub-tabs, and settings-ui into
the startup chunk with it, blowing the Control UI startup budget.
Move the rendering-free facts (slot/backend resolution, tab and curated key
lists, schema narrowing, the anchor id) into memory-schema.ts, which imports
only record-coerce and the shared slot policy. The view keeps the templates and
now consumes the same module, so there is still one owner per fact.
* chore(ui): record the memory settings surface in the startup budget baseline
Routing settings search through memory-schema.ts instead of the view module
recovered 10,872 B of the startup chunk (334,992 -> 324,120 B), which is back
under the 324,608 B ceiling. The remaining 2,795 B over the old baseline is the
honest cost of the new surface: its i18n strings, plus the slot/backend facts
the startup search index has to read.
Measured by hosted CI (run 30189972795); this worktree cannot build locally
because pnpm wants to purge a node_modules shared with other running agents.
Request shaping is selected by version predicates in llm-core, so a Claude id
from a generation newer than anything the plugin encodes fell through to pre-4.6
shaping: manual budget_tokens thinking plus caller sampling parameters, both of
which current models reject.
Resolve such an id through the plugin's existing forward-compat path and stamp
params.canonicalModelId, the same seam Bedrock and Mantle use to map a
provider-native id onto a canonical Claude contract. Shared contracts are
untouched, so Claude served through third-party providers keeps its own
shaping.
* fix(worktrees): reclaim git worktree locks left by dead processes
A gateway that dies without releasing its worktrees leaves
"openclaw pid=<dead>" locks behind. lockState() already classifies those as
{ kind: "dead" }, and release()/remove() reclaim them, but
lockWorktreeForProcess() threw instead. Every later acquire then failed with
"fatal: ... is already locked", and retainGitLock() swallowed the error and ran
the agent with no lock at all -- so the stale lock permanently disabled the
protection it was meant to provide.
Reclaim a dead-owner lock on acquire, the way the sibling call sites do.
* docs(worktrees): enumerate every reclaim path in the scope note