* fix(agents): remove a deleted agent's cron jobs on the offline delete path
Follow-up to #127037, which fixed the exec-approvals half of the same gap and
named this one explicitly.
`agents delete` tries the Gateway first and falls back to a local path. The
Gateway handler nests two transactional cleanups around the roster commit --
cron wrapping approvals wrapping the config write. After #127037 the offline
path did the inner one; it still skipped cron. So deleting an agent without a
Gateway left its scheduled jobs enabled:
$ openclaw agents delete cronprobe --force
Deleted agent: cronprobe <- no mention of cron
$ sqlite3 <state>/state/openclaw.sqlite "select job_id, name, agent_id, enabled from cron_jobs"
975cb750-... | cronprobe-job | cronprobe | 1
$ openclaw cron list
cronprobe-job every 1h Next: in 59m idle
To be accurate about severity: this is not silent. Each firing records
`error: "cron job agent is unavailable: cronprobe"` and `cron list` flips to
`error`. The defect is that the job keeps its schedule forever, and that
recreating an agent with the same id points it at the new agent.
The fallback had collapsed two different reasons into one `null` return, which
is what made the fix look unsafe at first: credential failures happen *before*
transport, so a live scheduler may still own the cron store, while an
unreachable Gateway means nothing else is holding it. `maybeDeleteAgentThroughGateway`
now returns a discriminated union, and only the unreachable branch mutates the
store directly. The credentials branch commits the roster, warns, and sets
`cronCleanupSkipped: true` in JSON.
The local `CronService` construction already existed inside
`local-request-context.ts`; it moves to `src/cron/local-service.ts` and both
callers share it rather than growing a second cron mutation path. That
extraction also switches the default-owner resolver from
`tryResolveLegacyCompatibilityAgentId` to `tryResolveAmbientOwnerAgentId`, which
is a superset -- it honors an explicitly configured
`agents.defaults.systemAgent.agentId` and otherwise falls back to exactly the
previous function. Live testing showed agentless memory-dreaming jobs need it to
load under explicit agent ownership.
Production +89/-62.
* test(agents): split the delete suite so the new cron coverage stays under the cap
The 40-line cron regression test added in the previous commit pushed
`src/commands/agents.delete.test.ts` to 1018 code lines, over the 1000 cap, and
`check-lint-core-3` went red. Repo policy forbids a `max-lines` suppression.
Unlike the earlier `cron/view.test.ts` split there was no describe-level seam --
23 flat tests in a single describe -- so the split follows subject instead. The
seven workspace-lifecycle tests (trashing, sharing, overlap, symlink reachability,
workspace-state cleanup) move to `agents.delete.workspace.test.ts`.
`vi.mock` and `vi.hoisted` are per-file and cannot be imported, so the mock
preamble and the shared `beforeEach` are declared in both files; the helper block
above them is unchanged in each. Each file then imports only what it uses, which
is why the import lists differ.
Trimming to a hair under the cap by moving only the new test was possible and
rejected: it would have left the file at ~978 code lines, back at the cap within
a couple of changes. This leaves 749 and 603 physical lines.
No test content changed: 27 passed before, 27 after.
Remove the premature visibility classifier and let one proof agent configure and exercise the disposable Telegram gateway. Align mock response timing with the 15-minute lane budget while preserving credential isolation through the alias-token proxy.
Preserve model-supplied filename identity for mutations while keeping existence-checked Unicode-equivalent fallback for reads.
Co-authored-by: yetval <yetvald@gmail.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
`openclaw onboard --help` advertised `claude-cli`, which the command
rejects, and hid `token`, which it accepts and its own error text
recommends. Three surfaces each rebuilt the accepted set independently:
Commander help asked for legacy aliases, the reset preflight added a
hardcoded `BUILT_IN_AUTH_CHOICES`, and the non-interactive dispatcher
added a hardcoded `GENERIC_NON_INTERACTIVE_AUTH_CHOICES`.
`formatAuthChoiceChoicesForCli` is now the single owner of that set. It
emits the generic token-provider choices (`setup-token`, `token`,
`apiKey`) that `AuthChoice` has always declared built-in, so every
surface renders and validates the same list.
Deprecated aliases stay out of it. `includeLegacyAliases: true` arrived
with a mechanical help-text refactor, never as a product decision: the
pre-refactor hardcoded help string listed no legacy alias, the reset
preflight's copy was unreachable because normalization runs first, and
the deprecation error already names the replacement. `oauth` is likewise
normalized to `setup-token` before any validator sees it, so the
dispatcher's `oauth` arm and its slot in the accepted list were dead.
Production LOC: -26.
* fix(gateway): recover credential-file accounts on secrets reload
Preserve independently discovered credential-file degradation across runtime snapshot refreshes, and re-inspect only the affected account when secrets are reloaded. Healthy sibling accounts remain running while status and doctor retain exact-owner diagnostics until recovery or teardown.
* test(gateway): prove credential-file reload recovery
* test(gateway): assert redacted reload error code
Bind scheduler-owned cron, hook, and heartbeat runs to lifecycle-fenced Gateway context so trusted built-in tools resolve after startup or reload without inheriting request client state.
Co-authored-by: Marvinthebored <marvin.assistant@lindsey.jp>
Co-authored-by: Marvinthebored <peter@lindsey.jp>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(cron): reject blank --command-cwd/--command-input before kind forge
Presence-only typeof checks counted blank cwd/input as command-specific
edits, forging {kind:"command"} patches that could convert agentTurn or
script jobs into empty command payloads. Require non-blank values first.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(cron): allow empty --command-input while rejecting blank cwd
ClawSweeper: Gateway command stdin is an unrestricted string, so empty
or whitespace --command-input must still patch through. Keep the blank
--command-cwd reject that prevents forging a command payload with no cwd.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* test(cron): focus command blank option coverage
Amp-Thread-ID: https://ampcode.com/threads/T-01a0220d-eaa0-76b4-adb9-68841f015b75
* fix(cron): reject blank cwd before edit reads
Amp-Thread-ID: https://ampcode.com/threads/T-01a0220d-eaa0-76b4-adb9-68841f015b75
---------
Co-authored-by: zyw02 <zyw02@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Amp <amp@ampcode.com>
Two defects in one lifecycle, found by following one thread.
`agents delete` tries the Gateway first and falls back to a local path when the
Gateway is unreachable -- its own header calls this "gateway delegation and local
cleanup fallback". The Gateway path wraps deletion in
`withAgentExecApprovalsRemoved`; the local fallback did not. So deleting an agent
with no Gateway running left its exec approvals allowlist behind, silently.
That residue is not inert. Verified end to end: approve a pattern for an agent,
delete it through the offline path, recreate an agent with the same id, and
`approvals get` shows the old allowlist live under the new agent. The operator
never granted it. The offline path now uses the same journal-fenced,
rollback-capable helper as the Gateway path.
Removal rather than a warning: an exec approval is a latent authority grant, and
leaving one behind contradicts the deletion contract the Gateway already
enforces. `agents delete` already trashes the workspace and agent dir, so
removing a policy entry is not out of character for it.
The second defect is in the shared helper and affects the Gateway path too, which
means it is shipped today. It matched policy keys with
normalizeAgentId(policyKey) === key
and `normalizeAgentId` falls back to `"main"` for anything it cannot represent.
`"*"` normalizes to `""` and therefore returns `"main"`, so deleting the `main`
agent also matched -- and removed -- the `"*"` wildcard allowlist. Strict
normalization now skips unrepresentable keys, so valid aliases are still removed
while the wildcard and unrelated agents survive.
Sibling sweep of `pruneAgentConfig`, which already handled bindings, subagent
allowlists, heartbeat/system-agent ownership, and Talk ownership: broadcast
targets, hook mappings, and hook agent allowlists also retained the deleted id.
Fixed, preserving `"*"` in the hook allowlist.
Follow-ups deliberately not taken here, each needing its own owner or design:
offline deletion still lacks the Gateway's transactional cron-job cleanup;
plugin-owned routing and thread bindings need an agent-deletion lifecycle hook;
remote node approvals need distributed cleanup or explicit warning semantics;
approval `agentFilter` cleanup needs a disabled-state design, since removing its
last entry would widen policy rather than narrow it.
Production +45/-18.
The generic provider-missing-auth error told operators to create a new agent:
No API key found for provider "openai". Auth store: <path> (agentDir: <path>).
Configure auth for this agent (openclaw agents add <id>) or copy only portable
static auth profiles from the main agentDir.
`openclaw agents add <id>` creates an agent. It does nothing about a missing
provider credential, so following it leaves you with a second agent that also
has no key.
Reproduced on a real build: a fresh install with only an Anthropic key, then
`openclaw memory index`. There is exactly one configured agent, `main`, and it is
correctly configured -- nothing about `agents add` applies. Doctor, diagnosing the
identical condition on the same install, gets it right, listing how to supply the
key and how to disable memory search. The command that actually failed gave worse
advice than the advisory check.
`models auth paste-api-key --provider <id>` is the suggestion rather than
`models auth login`, because `login` requires an owning provider plugin while this
generic fallback is also reached for pluginless, inline, and custom providers, and
for a plugin-owned provider that returns no specialized message. Pasting a key
resolves every "No API key found" case the fallback can produce. The `--agent`
hint is folded in for non-default agents, and `formatCliCommand` keeps the command
correct under a profile or container.
One message rather than two: distinguishing "no credential anywhere" from "this
agent lacks a profile another agent has" would need cross-agent store reads the
resolver does not do today, and the combined sentence is correct for both.
Production +1/-1.
* feat(control-ui): stream live draft previews in the typing indicator
Multi-identity sessions now show what a teammate is typing, not just that
they are typing: the composer's per-keystroke session.typing sends carry a
bounded tail of the draft (optional preview field, 400 code points max),
the gateway throttle re-emits on changed payloads at 250ms (boolean-only
stays at 1s, trailing edge keeps the latest draft), and the transcript
renders a per-actor bubble with the live text plus a blinking caret.
Actors without preview data keep the three-dot bubble.
Previews are ephemeral presence: never persisted, never part of the
session transcript or model context, excluded from aria-live regions, and
gated by the existing >=2-live-viewers, sharing-role, and incognito
checks. No new config surface.
* chore(protocol): regenerate Swift gateway models for typing preview
* fix(gateway): aggregate typing previews across same-actor connections
A boolean-only session.typing update from a second connection of the same
actor (another tab or device) erased their live draft preview, because
typing liveness aggregated per actor while the broadcast preview came only
from the latest request. Preview aggregation now lives with the connection
aggregation owner: updateTypingConnections tracks per-connection previews
and returns the newest non-empty preview among live connections, so the
broadcast keeps the active draft until its connection stops or expires.
Regression fails pre-fix (event lost its preview field).