* fix(auth): create a fresh install with canonical shared-auth ownership
A brand-new install was born in the retired shape. `parseSharedAuthStoreOwnership(undefined)` returns
`legacy-main`, which is the correct compat answer for an existing install whose profiles really do
live in the main agent database -- but a new install has no ownership row and no legacy data either,
so onboarding wrote its first credential into `agents/main/agent/openclaw-agent.sqlite` and the
operator's very first `openclaw doctor` told them to run a migration for state OpenClaw had created
seconds earlier. The main agent also stayed undeletable until they did.
Record `auth.sharedStore = {"location":"state-db"}` when the shared store is first written and the
legacy source provably holds nothing: no `auth_profile_store` row, no `auth_profile_state` row, and
no unfinished cleanup ledger entry. Any legacy row, or any inspection error, leaves ownership alone
so doctor keeps owning the relocation. The check is memoized per ownership generation with a WeakSet
keyed on the process-stable ownership object, so a legacy root is inspected once per process and
doctor's committed flip naturally invalidates it.
The legacy row inspection moves out of `state-migrations.shared-auth-store.ts` into the auth-profiles
owner so doctor and runtime share one contract instead of runtime importing migration code. Explicit
main-agent credential writes now follow the shared target, which is a no-op on legacy roots where
both routes already resolve to the same file.
No SQLite schema change; the ownership row is data. Existing installs take exactly the path they take
today.
* fix(auth): preserve JSON-era shared credentials
* docs(auth): explain why doctor names the main agent dir during a shared JSON import
* test(auth): assert shared-owner runtime reads
* test(doctor): read migrated catalog credentials through the shared owner
A fresh root records state-db shared ownership, so the model-catalog credential migration persists into the shared store rather than the agent file. The assertion read the agent file directly and saw an empty store while all three credentials were present and correct in state/openclaw.sqlite. Read through the owner for that state root instead of pinning storage layout; the credential contents are still asserted exactly.
* fix(cli): message send cannot address channels from npm-installed plugins
Target resolution, channel enumeration, and target-prefix inference only consulted the process-root channel registry, so message CLI actions running against a scoped registry handle could not see installed channel plugins even though selection and send execution could. Carry the selection-resolved plugin into target resolution, fall back to the registry handle in scope for resolver-owned lookups, and list runtime-visible channel plugins for channel selection and prefix inference.
* fix(cli): keep runtime-visible channel reads import-light
Importing channel-resolution from the target-prefix leaf pulled the plugin bootstrap/loader graph into every consumer and reordered module loading under distant vi.mock factories (subagent-registry.steer-restart failed in CI with a hoisting TDZ). Move the scoped-registry reads into a dedicated import-light module, share its registry matcher with channel-resolution, and drop the mock workarounds the heavier graph had required.
* chore(ui): re-baseline startup JS for the outbound scoped-registry reads
CI measured 348285 B gzip on the merge ref (baseline 347023 + 1056 tolerance). The first CI round measured 347784 B, so most of the growth is main-side drift since the 2026-08-19 baseline; the outbound changes account for roughly 60 B in a local A/B. Updated with the documented --update-baseline --startup-js-bytes flow using the CI value.
* Revert "chore(ui): re-baseline startup JS for the outbound scoped-registry reads"
This reverts commit f60bd4c45f.
* fix(cli): plan broadcast accounts from runtime-visible channel plugins
The unscoped message broadcast --account planner still enumerated only process-root plugins, so a registry-scoped installed channel could not join broadcast candidate planning. Use the runtime-visible read and cover the scoped and no-scope paths.
* fix(cli): honor scoped channel plugin precedence
* fix(outbound): preserve loaded plugin fallback order
---------
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
A reachable gateway whose health check failed either escaped as a raw thrown
error or hit healthCommand's CLI-style runtime.exit(1), killing non-interactive
onboarding before logNonInteractiveOnboardingJson could emit the --json summary
and polluting JSON stdout with human diagnostic text. Route health failures
through logNonInteractiveOnboardingFailure (structured ok:false payload in
--json mode, framed text otherwise) via healthCommandNonExiting, and capture
healthCommand's human output off stdout in --json runs so the diagnostic lands
in the payload's detail field instead.
The first `openclaw doctor` on a freshly onboarded install reported "Persisted
plugin registry is missing or stale. Repair with `openclaw doctor --fix`".
Nothing was wrong: the `installed_plugin_index` row had never existed, `plugins
list` reported 148 plugins without it, no retired `plugins.installs` records
were present, and starting the gateway once builds the row by itself -- measured
going 0 -> 1 across a single `gateway run`.
`preflightPluginRegistryInstallMigration` returned a single `action: "migrate"`
whether the index was absent or unreadable, and the health issue name
`registry-missing-or-stale` shows the conflation: missing and stale are
different states and only one is a problem.
Split that into `initialize | migrate`. A root with no persisted index, no
install records, and no retired config records is initialization the gateway
owns, so doctor stays quiet. Install records without a readable index stays a
migration, as does a config still carrying retired `plugins.installs` records,
or a caller that supplied no config to prove otherwise. `doctor --fix` still
builds the index in every case -- this changes the warning, not the repair.
Production +5 LOC.
* fix(wizard): keep embedded health-check failures visible instead of exiting mid-flow
healthCommand's reachable-gateway auth diagnostic paths call runtime.exit(1),
which with defaultRuntime hard-kills the hosting configure wizard, onboarding
finalize, or doctor daemon flow mid-render — dropping the failure framing,
docs guidance, and outro. Add healthCommandNonExiting, which traps that
CLI-style exit into ExitError so the host flow owns the outcome, and use it
at every embedded call site.
Also fix the doctor e2e harness createConfigIO mock missing configPath, which
broke doctor.runs-legacy-state-migrations on current main.
* fix(onboard): reflect a failed health check in the finalize outro
A reachable gateway whose health check failed still ended onboarding with the
plain success outro because completion gating only read the earlier
reachability probe. Record the health outcome as its own fact and end with a
dedicated outro pointing at openclaw health. (ClawSweeper P1 on #126758.)
`openclaw onboard --openai-api-key 'openclaw onboard --auth-choice ...'`
correctly refuses with "Paste the API key value, not an OpenClaw onboarding
command", exits 1, and writes no config. The identical value supplied through
`OPENAI_API_KEY` -- the form `docs/start/wizard-cli-automation.md` documents for
automation -- exited 0 with empty stderr and persisted the command string as the
credential.
`isMalformedApiKeyInput` was already imported into this file but guarded only
the `flagKey` branch; both `resolveEnvKey()` branches returned unchecked. The
operator finished onboarding believing they were configured, and nothing told
them otherwise until an agent turn failed or they happened to run
`openclaw doctor`, which classifies that exact value as `malformed_api_key` and
prints the hint they never saw.
Route every operator-supplied key -- flag, env, and secret-ref env -- through one
guard, and name the environment variable in the message so an operator with
several exported keys knows which one is wrong. The stored-profile branch stays
unguarded on purpose: a bad key already on disk is doctor's to diagnose, and
refusing there would strand someone re-onboarding to replace it.
`openclaw onboard` refuses a corrupt openclaw.json and tells the operator to run
`openclaw doctor --fix`. Doctor then answered with one sentence -- "Config could
not be parsed or recovered ... refusing to apply repairs" -- named no next step,
and exited 1. The operator was left looping between two commands that pointed at
each other.
The refuse path also wrote openclaw.json.clobbered.<timestamp> and called it
"Original preserved", but it had not clobbered anything: at that point the
snapshot is a reread of the live file, so the copy was byte-identical to the
untouched config. Three failed runs left three identical copies.
Drop the copy and say what to do instead: name the file, state that it cannot be
repaired automatically, and point at `openclaw config validate` for the exact
parse position, hand-editing, or moving the file aside and re-running
`openclaw onboard`. Commands go through formatCliCommand so profile and
container invocations stay pasteable.
`doctor-config-preflight.ts` was the only caller of the public
preserveConfigSnapshotAsClobbered wrapper, so the wrapper, its factory entry and
its barrel export go too; the genuine recovery paths keep using the core helper
and still preserve real originals. Production -16 LOC.
* fix(models): keep direct-credential gateway probes on isolated runtime generations
* fix(models): carry isolated-runtime probe mode on the loader type
Drop the call-site cast for the widened runner params: the lazy loader now
declares the isolated-read-only capable runEmbeddedAgent shape, keeping the
assertion-safety ratchet at its grandfathered baseline for this file.
* fix(config): validate config writes against the config being written
writeConfigFileFromContext passed the pre-write snapshot's plugin
metadata into strict validation. During onboarding that snapshot belongs
to an intermediate config written by agent creation, which has no plugin
entries, so its scoped manifest registry is empty. Validating the final
candidate against it made every plugin entry added by the same write look
unknown, and non-interactive onboarding warned that the openai and codex
entries it had just written were stale or uninstalled.
Drop the stale snapshot so validation resolves the manifest registry from
the candidate it is actually validating. Strict semantic validation is
unchanged, and the registry load stays lazy.
* fix(ci): raise the Control UI startup JS baseline to unblock main
main is red on the Control UI startup-JS ratchet: unrelated PR #126725 measures 348289 B and this branch measures 348351 B against a 347023 B baseline + 1056 B tolerance. Twenty-six UI commits have landed since the last bump (#126474), none individually large. Baseline moves to the CI-measured 348351 B, well under the 358400 B maintainer-approved ceiling that still guards cumulative creep.
Plugin install, replacement, and uninstall clear process memos through
registerPluginMetadataProcessMemoLifecycleClear, but four executable-
authority caches never registered, so retired plugin callbacks kept
executing after the registry moved on:
- createConfigScopedPromiseLoader (document/web-content extractor lists)
now self-registers its clear at the factory, so no caller can leak
resolved plugin callbacks past a lifecycle change.
- Provider policy surface maps (bundled + external, including cached
negative entries) clear on lifecycle changes.
- Public surface loader now drops module exports, loader closures, and
native require cache entries, not just resolved locations.
- SDK facade loader registers the same clear for facade exports and
loader state; imported-plugin history is preserved as diagnostics.
The tracked-roots + native-require eviction pattern from provider
discovery is extracted into clearPluginModuleLoaderLifecycleCache and
reused by provider discovery, doctor contracts, the public surface
loader, and the facade loader, removing two near-copies.
Regression tests fail pre-fix: replaced or uninstalled plugin callbacks
must not run after clearPluginMetadataLifecycleCaches, proven down to
on-disk artifact replacement through the native require chain.
* fix(tasks): rank terminal tasks by completion and keep Recent terminal-only
- updateTaskStateByRunId backfills lastEventAt from endedAt for terminal
finalizers (mirrors markTaskTerminalById), keeping activity monotonic
- both taskUpdatedAt projections rank terminal tasks by the maximum
available activity timestamp, healing stale rows while preserving later
delivery/terminal-outcome events recorded after completion
- Tasks page Recent fetch filters to terminal statuses so queued/running
rows cannot starve the Recent section
Related to #100911
* refactor(tasks): normalize completion at registry owner
Absorb terminal timestamp ordering into the canonical registry lifecycle boundary, remove duplicated projection and writer policy, and prove Recent remains visible behind 200 active tasks in Chromium.
Co-authored-by: SunnyShu0925 <shu.zongyu@xydigit.com>
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>