* fix(gateway): bound audit and Codex backlogs
Live Gateway SQLite lock failures and process heap pressure exposed two
independent queue owners. Route best-effort audit persistence through the
canonical shared-state connection with bounded contention retries, and remove
the per-notification Codex yield so the keyed turn queue can drain directly.
Follow-up to #126033 and #126073.
* fix(gateway): annotate raw SQLite cold-open probe
* test(codex): register notification burst shard
* fix(cli): report unknown subcommands even with --help
* fix(cli): scope Commander hook lint exceptions
* fix(cli): document Commander hook assertion
* refactor(cli): delegate the Commander help hook through super
* fix(snapshot): survive cold PowerShell starts in Windows staging gates
CI run 31775262530, checks-windows-node-test-1 attempt 1, showed the fail-closed ACL probe timing out during PowerShell first-use module preparation. Centralize encoded one-shot spawning, budget 60 seconds for cold starts, and preserve the underlying probe failure as the error cause.
* fix(snapshot): sanitize PowerShell failure causes in Windows staging gates
* fix(secrets): explain the sanitized plan-file failure cause suppression
check-lint-core-2 flagged preserve-caught-error at the private plan file
catch; retaining the raw error would re-leak the -EncodedCommand argv the
sanitization contract strips, so the suppression is intentional (same
idiom as setup-inference-activate.ts).
* test(lint): register the private-plan-file suppression in the inventory
* test(infra): give the LAN-host real PowerShell spawn a cold-start budget
checks-windows-node-test-2 (run 31804325922) hit the same cold-start flake
class this PR fixes: the codepage-proof test spawns real powershell.exe
bounded at 3s, which a cold runner cannot meet. Production keeps its
fail-open 3s route-hint probe; only the test's real-spawn verification
uses the shared cold-spawn budget.
* fix(agents): race in-flight tool promises against run abort (#103905)
wrapToolWithAbortSignal forwarded the combined abort signal into tool
execute calls but awaited the returned promise without racing it, so a
tool that never observes the signal kept running after stuck-session
recovery aborted the run and delivered its result into a dead run.
Race the execute promise against the combined signal so the wrapped
call rejects with AbortError as soon as the abort fires. The tool
promise itself cannot be cancelled in JavaScript; it keeps running in
the background and its late settlement is detached so it never lands
in the aborted run. Tool results and errors pass through unchanged
while the run is alive.
Refs #103905
* fix(agents): satisfy oxlint abort-race rejection rules without losing tool errors
* fix(lint): annotate intentional non-Error pass-through rejections
The race intentionally forwards tool rejections untouched (including
non-Error values), which trips oxlint's promise-rejection rules; mark
both sites with oxlint-disable-next-line justifications.
* test(scripts): allowlist intentional abort-race rejection suppression
The raw non-Error pass-through reject() in agent-tools.abort.ts is the
documented tool-error contract; register its prefer-promise-reject-errors
suppression in the production allowlist.
* fix(agents): observe sync-aborted tool rejections
Co-authored-by: thomas.szbay <thomas.szbay@xydigit.com>
---------
Co-authored-by: thomas.szbay <thomas.szbay@example.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: thomas.szbay <thomas.szbay@xydigit.com>
A landed change removed the no-useless-assignment suppression in
extensions/reef/src/transport.ts without pruning the allowlist tail,
failing checks-node-compact-small-7 on every PR (merge-skew).
* feat(ui): animated Clawd mascot on the new-session welcome hero
Ports the native mascot to the Control UI as a canvas-backed
<openclaw-mascot> Lit component: exact 120x120 vector geometry, mood
loops (idle/curious/thinking/working/happy/celebrating/sad/sleepy/
attentive), deterministic seeded animator, DPR-aware rendering,
visibility-paused rAF, theme-reactive palette, and a reduced-motion
static pose. The new-session welcome hero swaps its static lobster
sprite for the living mascot.
* fix(ui): cancel stale mascot gestures on mood change
* fix(ui): keep MascotEffect internal to the pose module
* test(lint): allowlist the canvas fill suppression for the web mascot
The repository hard-zero sweep (bae9752c5a, #108641) left five CI lanes
red on every PR: a stale max-lines baseline entry for the shrunken
live-cache-regression-runner, session-accessor debt baseline above the
real count, two sqlite-reliability contract types un-exported while
scripts/lib still imports them, a deleted-file entry on the production
lint-suppression allowlist, and a stale wildcard-barrel assertion in the
parallels smoke model test.
User-facing name is now OpenClaw (the system speaks); internal code name is
system-agent. Gateway methods crestodian.* -> openclaw.chat/openclaw.setup.*,
agent tool -> openclaw, reserved agent ids openclaw + retired crestodian.
openclaw setup routes: onboarding flags -> onboard, -m/--yes -> system agent,
bare configured interactive -> OpenClaw chat, unconfigured -> onboarding.
Hidden crestodian CLI and /crestodian TUI aliases kept; docs moved to
docs/cli/openclaw.md with redirect stub. macOS/Android strings in lockstep.
Refs #107237
* feat: node-hosted plugins — dynamic tools, MCP servers, and skills
Nodes become declarative plugin hosts:
- node.pluginTools.update: node hosts publish plugin-registered agent tool
descriptors; gateway materializes them as agent tools executing via
node.invoke under the node command allowlist, with tools.effective
invalidation and node online/offline removal.
- Trusted paired-node descriptors: no gateway-side plugin registration
required; gateway.nodes.pluginTools.enabled off-switch (default on);
description/count caps; deterministic node-prefixed collision names.
- Declarative node-hosted MCP: nodeHost.mcp.servers (McpServerConfig shape)
starts MCP clients on the node host, publishes tools as pluginId node-mcp,
executes via built-in mcp.tools.call.v1 with per-layer timeouts, failure
isolation, and orphan-safe shutdown. No re-pairing when servers change.
- Node-hosted skills: node.skills.update publishes ~/.openclaw/skills
content (64 skills/64KB/512KB caps both sides); gateway merges them into
the skills snapshot while connected and exec host=node is available, with
node:// locators, node-prefixed collisions, disabled command dispatch,
and gateway.nodes.skills.enabled + nodeHost.skills.enabled switches.
- Security: node-supplied pluginIds cannot satisfy pluginId-scoped tool
allowlists unless gateway-registered; reserved node-mcp id requires the
core MCP descriptor shape; protocol registry kept out of public
plugin-sdk dts.
- E2E: pond harness proves publication, MCP round-trip, skills locator, and
disconnect/reconnect for all three surfaces.
* style: format node-plugin-tools test
* fix(skills): keep status loader unfiltered when eligibility is passed
skills.status started passing eligibility for the node-skill merge, which
flipped loadWorkspaceSkillEntries into filtered mode and dropped disabled
skills from status reports (QA plugin-lifecycle-hot-reload timeout). Status
now merges node skills explicitly around an unfiltered load. Also: regen
docs_map for new node docs sections; add the intentional node-host MCP
onclose suppression to the lint-suppression allowlist.
* refactor(pairing): move device pairing store to shared SQLite state DB
Device pairing, pending requests, and bootstrap tokens now live in the
device_pairing_* / device_bootstrap_tokens tables of state/openclaw.sqlite
instead of devices/{paired,pending,bootstrap}.json. Gateways import legacy
paired records once at startup (before the node-surface fold) and archive
the JSON files with a .migrated suffix; transient pending/bootstrap rows
are dropped. The unshipped node_pairing_* tables are removed from the
schema and dropped from existing DBs. Doctor now flags un-imported legacy
store files instead of corrupt-JSON reads.
* refactor(pairing): drop stale awaits now that store persistence is synchronous
* refactor(pairing): extract leaf record types to break store/domain module cycle
* feat: correlate native search outcomes in audit history
Metadata-only audit ledger for agent runs and tool actions in the shared
state DB: stable event identity, closed action/status/error vocabularies,
one-way-hashed tool-call ids, never-inferred terminal outcomes for native
web-search (explicit completed/failed only; otherwise unknown), bounded
retention, audit.list gateway RPC and openclaw audit CLI. Squashed from
the 82-commit audit stack for replay onto current main.
* feat(audit): add audit.enabled config gate (default on)
The metadata-only audit ledger records by default: an audit trail enabled
only after an incident cannot explain the incident, and the rows are
strictly less sensitive than the transcripts every install already
stores. audit.enabled=false stops new writes at the gateway subscription
seam; audit.list and openclaw audit keep serving existing records until
they expire. Documented in the configuration reference, protocol page,
and CLI reference.
* fix: repair full-matrix CI findings after rebase
- break the dynamic-tools/dynamic-tool-execution import cycle by
extracting resolveCodexToolAbortTerminalReason into a leaf module
- restore main's session-worktree protocol exports lost in the
index.ts auto-merge
- register the audit event writer worker as a knip entry point
- docs table formatting; subagent wait-cancellation test scoped to its
audit intent (outcome + timing) and advanced past main's new
lifecycle-timeout retry grace
* fix(diffs): share SSR preloads and repair language-pack hydration
Render viewer and file documents from a single @pierre/diffs SSR preload
per file (mode=both previously ran the full diff+highlight pipeline
twice; 651ms -> 303ms on an 8-file patch), apply the file-mode font bump
as a document-level override, and keep hydration payloads
variant-faithful.
Fix the language-pack runtime downgrading pack-only languages to plain
text at hydration by defining a per-target build flag and forwarding it
to payload normalization.
Also: case-insensitive language hints, identical before/after
short-circuit with details.changed, patch input failures classified as
tool input errors, canonical config values now win over deprecated
aliases, hash-pinned viewer runtime served immutable, truthful
browser-vs-render errors, timing-safe artifact token compare, unref
idle browser timer.
* docs(changelog): link diffs rendering entry to PR
* test(diffs): narrow manifest validation results before value access
* test(tooling): allowlist diffs viewer-client define suppression
* fix(agents): fail fast with attributable reason after MCP stdio session dies mid-run
Wires MCP Client onclose/onerror during bundle-mcp session creation so a
crashed/exited server flips session.connected instead of staying stale.
Next tool/resource/prompt call throws a domain-specific 'is disconnected'
error immediately instead of surfacing the SDK's generic 'Not connected'.
A disconnected reused session is retired and rebuilt fresh on the next
catalog pass rather than reused, since the SDK chains onclose/onerror
cumulatively on repeat connect() and the stdio transport never clears its
read buffer on an unexpected exit.
* fix(agents): retire a reused MCP session that dies mid-refresh, not just pre-refresh
codex review found: the catch-path retirement in getCatalog()'s per-server
task only covered two cases (fresh session that never connected this pass,
and a non-reused session that failed for any reason) - a reused session
that was healthy when this pass started but disconnects mid-refresh (child
process dies between ensureSessionConnected() returning and
listAllToolsBestEffort() finishing) hit neither branch, so onclose flipped
connected=false but the dead session object stayed in the map until the
next catalog rebuild happened to notice it. Added the missing branch plus
a regression test that kills the child process mid tools/list on a reused
session and asserts the session is purged from the map within the same
refresh (not just fail-fast on tool calls, which already worked).
* fix(agents): retire a dead reused MCP session even across overlapping catalog generations
codex round-2 review found: the mid-refresh retirement branch still
skipped when sharedWithNewerGeneration was true, which protects a
still-alive session another generation is actively using - but once
onclose flips session.connected=false, the transport is dead for every
generation sharing that object, so that guard no longer applies.
Simplified to a single !session.connected branch that always retires
(retireSessionIfCurrent already no-ops safely if a newer generation
replaced the map entry), and dropped the now-write-only
connectedForCatalog local it replaced.
* test(scripts): allow MCP SDK callback suppressions
* fix(agents): retire closed MCP sessions
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>