Propagate the receipt-acknowledgment steps from the canonical
ClawSweeper dispatch template (openclaw/clawsweeper#1080): mint a
minimal issues:write App token and post an idempotent
clawsweeper-pr-ack marker comment for non-draft opened and
ready_for_review pull requests, before review dispatch.
Use channels.telegram.groups as the shared default when an account omits groups. Explicit account maps, including {}, remain full replacements. This keeps chat admission and sender restrictions on one policy path and fixes silent multi-account authorization failures.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Abandon unexpected acceptance failures instead of leaking accepting slots. Reject non-Git or blank suggestions before side effects, and protect unseen pending suggestions ahead of accepted replay state.
* fix(doctor): migrate legacy agent databases discovered on disk and bound registry writes to the active state dir
* fix(doctor): discover configured agent databases without registry rows
* fix(doctor): preserve filesystem and configured database identity
* fix(doctor): prioritize configured agent database identity
* fix(doctor): prefer recorded agent database ownership
* fix(gateway): rebuild chat metadata when auth-profile snapshots change
chat.metadata cached a prepared generation built before the runtime
auth-profile store snapshot was published (empty-store fallback) and
generationFactsMatch never compared auth state, so the Control UI showed
"No models available" after a gateway restart until an unrelated config
edit. Capture per-agent auth snapshot revisions in the generation facts,
subscribe the metadata lifecycle to auth-store mutations, and run one
awaited revision-aware catch-up refresh after listener registration so
publications that precede registration are still observed.
* fix(gateway): bound cloud worker tunnel startup and surface dispatch failure detail
Live stress-testing the Control UI cloud flow found sessions.dispatch
hanging unbounded (observed 12+ min) when the worker SSH tunnel could
not connect: the runner's exited promise settled only on "close" (a
spawn error never settled it), the reconnect loop swallowed every
failure with no logging, the poisoned ready promise was re-handed to
every later dispatch, and dispatch errors dropped the actionable reason
recorded in worker_environments.last_error.
- settle exited on the real "exit" event, wire the owner abort signal
into spawn, bound stop()'s post-SIGKILL wait, and fail stop honestly
when termination is unconfirmed instead of fabricating an exit
- add a 60s per-attempt readiness deadline and log each failed connect
attempt (bounded, redacted); keep an unconfirmed child tracked and
wait for its real exit before retrying
- add a 3-minute epoch-fenced startTunnel deadline with a typed,
actionable error; detach its cleanup so the deadline holds
- append the bounded recorded reason to the five dispatch-visible
worker environment error messages (docs already promise these)
- remote socket setup: drop "--" from chmod (BSD/macOS chmod treats it
as a filename), which blocked every tunnel to a macOS worker host
- workspace quiescence: tolerate EPERM without crashing the protocol
while keeping unsignalable freeze targets counted as live so
quiescence fails closed
* test(gateway): type worker child kill mock
Replies, previews, and media in channel Direct Messages topics now remain in their originating topic. Bot-private and forum topic routing remains unchanged.
* perf(agents): keep turn-path model catalog reads off the full live build
First agent turns (embedded and cron) resolved thinking capability through
loadPreparedModelCatalogSnapshot without readOnly, which materialized the
full live model-runtime catalog: ambient synthetic-auth discovery fanned out
to every registered provider and loaded plugin discovery modules through
jiti source transform (3,172 TS modules, 36s event-loop block, +600MB heap,
58.7s model-selection on a cold gateway).
- add loadProviderScopedThinkingCatalog: manifest metadata first, then a
provider-scoped read-only static catalog, then scoped live discovery only
for runtime-discovery providers (preserves #116584 Ollama semantics)
- route scopedLiveProviderDiscovery through the scoped read-only loader
- scope live-mode ambient synthetic-auth refs to the requested providers
- bound the last-resort synthetic-auth sweep to discovery entry modules
- memoize per-turn plugin skill dir resolution/republish (single-slot,
lifecycle-cleared; was a full walk + symlink republish every turn)
Cold first turn 72.7s -> ~22s wall (remaining cost is provider prefill of
the ~19.5k-token default prompt); model-selection 58,726ms -> 124ms.
* test(agents): align model-catalog.runtime mocks with scoped thinking catalog seam
Explicit vi.mock factories must export every binding prod touches; the new
loadProviderScopedThinkingCatalog export is now mocked everywhere the module
is stubbed, and the live-model-switch Ollama hydration test asserts the new
provider-scoped seam instead of the retired unscoped snapshot call shape.
* test(agents): export scoped thinking catalog from every prepared-catalog mock; split synthetic-auth helpers
- add loadProviderScopedThinkingCatalog to all explicit prepared-model-catalog
and model-catalog.runtime mock factories (vi.mock factories must export every
binding prod touches)
- move synthetic-auth ref scoping/resolution into
prepared-model-runtime.synthetic-auth.ts; keeps facts under the max-lines cap
* test(agents): prove scoped thinking hydration for runtime-only models
Boundary proof for the ClawSweeper review gap: the three-tier helper stops at
manifest or scoped-static when they resolve, and runs provider-scoped live
discovery (no broad fanout) only for runtime-only models; cron selection
hydrates through the same scoped helper and skips it entirely for thinking=off.
* test(agents): accept rest args in scoped thinking catalog mocks
* fix(cloud-workers): close lifecycle ownership gaps
Own bootstrap cleanup at the operation boundary and make fallback workspace sync converge across retries. Re-establish tunnel readiness per connection, retire placements before destructive session mutation, and keep operator diagnostics lightweight and redacted. Cover destructive lifecycle paths in their original execution order.
* fix(cloud-workers): drain local claims before retirement
delete/reset drain admitted local work, re-read exact identity, retire before destructive cleanup; active-claim/race tests.
* fix(cloud-workers): bind retry cleanup to workspace owner
Attest canonical HOME and the exact managed path.
Revalidate ownership before recursive fallback cleanup.
Cover malicious paths and ownership drift with tests.
* fix(cloud-workers): fence fallback workspace receivers
* refactor(plugin-sdk): delete the heavy runtime-doctor barrel
Nothing may pull the state-db/kysely graph through a doctor barrel anymore.
The barrel's remaining heavy exports move to two narrow private-local
subpaths, each with a single purpose:
- doctor-repair-runtime: install-path diagnosis, plugin config removal, and
state-database schema detect/repair (matrix doctor, voice-call lazy import)
- plugin-state-store-runtime: the sync keyed-store factory. It stays out of
plugin-state-runtime because hot channel entrypoints import that at module
load and opening a store pulls the state-database graph.
Doctor closures also stop pulling ssrf-runtime (fetch-guard + gateway net)
for two legacy private-network helpers that live in the lighter ssrf-policy
subpath: mattermost, nextcloud-talk, tlon, matrix.
The closure guard now forbids the two new heavy subpaths instead of the
deleted barrel, so the invariant keeps being enforced where it still applies.
* perf(doctor): keep heavy graphs out of every doctor closure
Doctor enumeration cold-loads each declaring plugin's contract closure, so
one heavy import in a closure is paid by the whole sweep. Four barrels were
still dragging unrelated graphs in for trivial helpers; each is repaired at
the leaf rather than by caching downstream:
- Legacy private-network config migration moves to a config leaf. It only
reshapes records, but lived beside the SSRF runtime (DNS, proxy, logging),
costing mattermost ~2.7s. ssrf-policy re-exports it, surface unchanged.
- Streaming config readers move to a leaf. They read two config keys, but
streaming.ts also formats tool aggregates, pulling tool-display/logging/
acp-core; that cost slack ~2.3s.
- signal took the channel-secret barrel for isRecord; the canonical plugin
record guard is string-coerce-runtime (root AGENTS.md).
- llm-task took the provider-model barrel for parseModelRef, now a narrow
model-ref-parse subpath.
Full doctor enumeration of all 42 declaring plugins, built mode:
legacy config rules 6668ms -> 1265ms, state migrations 184ms -> 127ms.
No plugin remains an outlier; the slowest is now ~380ms against a ~200ms floor.
Public export surfaces of every touched SDK subpath are byte-identical
(verified by diffing built module exports before/after); the API baseline
hashes move only because re-exported declarations emit differently.
The closure guard gains rules for each repaired barrel so the invariant
holds for future closures.
* fix(release): exclude new private-local declarations from the published package
Same pack-path rule as c41da3759f: private-local subpaths ship without d.ts.
* fix(doctor): repair the closure guard violations that break main
The landed guard fails on main: three closures import heavy barrels for one
symbol each. Two more surfaced once the guard learned about the provider-model
barrel. Each gets a narrow subpath at the leaf:
- telegram sent-message-cache + state-migrations took the session-store barrel
(session accessor + state-db) for resolveStorePath -> session-store-paths
- discord thread-bindings.state took the channel-outbound barrel (reply
pipeline + channel registry) for one identity write -> outbound-echo-runtime
- discord model-picker took the provider-model barrel for normalizeProviderId,
which model-ref-parse now exposes beside parseModelRef
The guard also stops walking artifacts of plugins whose manifest declares no
doctor surface. Such a declaration gates the artifact off every enumeration
path exactly as resolvePluginDoctorContracts does, so its closure cost is never
paid; anthropic ("doctorContract": {}) was being held to a cost it cannot
incur. Absent declarations still load eagerly and stay enforced.
Side effect worth naming: discord's built doctor contract now loads again.
On main both discord and telegram fail to require in packaged builds (an
ESM-only transitive dep) and silently lose their repairs; this restores
discord and takes enumerated legacy config rules from 87 to 99. Telegram's
built artifact still pulls execa through dist chunking - a build-level defect
with a different owner, filed as follow-up.
Make timeout paths cancel or release every pending test continuation before cleanup, so failures cannot poison the remaining Swift suite. Inject the retry sleeper to preserve invalidation-before-backoff ordering without real-time polling.