* Move raw transcript from system to user prompt to protect provenance.
* Type fix.
* fix(voice-call): harden transcript context handling
* fix(voice-call): initialize inbound Twilio control state
* test(voice-call): align runtime coordinator fixture
---------
Co-authored-by: joshavant <830519+joshavant@users.noreply.github.com>
Keep the writer-owned idempotency key stable across transcript redaction so final Codex snapshots replay against the admitted SQLite row instead of dropping mirrored history.\n\nCloses #126244
* fix: stop Full access sessions from requesting exec approval
* fix: propagate Full access policy to compaction
* fix: source compaction permissions from session state
* fix(codex): preserve model effort capabilities
Keep public model identities separate from app-server execution routing, and retain provider-owned complete effort metadata when account discovery is partial.
Fixes#126005
* refactor(codex): avoid redundant thread rotation
* fix(codex): preserve model fallbacks without leaking wire ids
* perf(gateway): remove repeated logging and delivery scans
Exact session-delivery retries no longer scan the full queue. Logging and diagnostics reuse lifecycle-owned settings and listener interest so uninterested projections are skipped, while outbound WebSocket summaries are built only after recipient admission.
* fix(infra): break diagnostic listener import cycle
Keep event-type validation at the diagnostic dispatcher while the process-wide listener presence counter remains a leaf module.
* test(cli): use logging override owner
Exercise late one-shot JSON diagnostics through the canonical logger override setter so lifecycle-cached console settings are invalidated as they are in production.
* test(auth): use logging override owner
Configure the locked-update warning test through the canonical logger override setter so lifecycle-cached console settings are invalidated before assertion.
* test(gateway): normalize redacted media fixture
Compare durable inbound media facts against the public redaction contract so random identifiers that resemble sensitive text do not make the Gateway suite flaky.
* fix(gateway): bound audit and Codex backlogs
Live Gateway SQLite lock failures and process heap pressure exposed two
independent queue owners. Route best-effort audit persistence through the
canonical shared-state connection with bounded contention retries, and remove
the per-notification Codex yield so the keyed turn queue can drain directly.
Follow-up to #126033 and #126073.
* fix(gateway): annotate raw SQLite cold-open probe
* test(codex): register notification burst shard
* fix(copilot): add OpenClaw prompt guidance
Copilot append-mode system messages included credential safety, workspace bootstrap, and extra context but omitted OpenClaw delegation and reply-delivery policy.
Build guidance from the final policy-filtered tool surface so visible delegated work, Skill Workshop, and source replies follow the same behavior as Codex.
* fix(copilot): break prompt guidance import cycle
The CI architecture gate detected a cycle through attempt-config and prompt-guidance. Isolate raw-run mode detection in a leaf module.
Preserve completed tool work when native Codex compaction fails, close failed compaction progress, and bypass unrelated model/auth failover before isolated finalization. Fixes#125789.
The `## Delegation` guidance added in #125691 lived only in
buildAgentSystemPrompt, so Codex-runtime agents never received it: the
Codex harness builds its own developer instructions in
extensions/codex/src/app-server/thread-prompt.ts and imports nothing
from the system-prompt builders. Live A/B on gpt-5.6-luna had the native
runtime answer "spawn a visible session" while the Codex runtime
answered "spawn a hidden subagent".
Move the policy into src/agents/delegation-guidance.ts, owning both the
main-session mode resolver and the section text, and export it through
the agent-harness plugin SDK barrel that the Codex harness already uses.
The hidden-delegation vocabulary is injected by each runtime, so core
never names a plugin-owned tool: native passes `sessions_spawn`, Codex
passes native `spawn_agent`. Visible sessions stay `sessions_spawn`
with visible=true on both runtimes because Codex-native children are
never OpenClaw sessions.
Also narrows the Codex line that told the model to use `sessions_spawn`
only for OpenClaw/ACP delegation; it now scopes that to internal
legwork, so user-facing deliverables still route to a visible session.
* fix(ci): stop codex lane cold-graph hangs
The side-question domain-policy test loaded the complete agent-harness tool graph inside a one-second readiness race, making the serial non-isolated Codex shard fail or stay silent under cold imports. Build the test's web_search marker and real web_fetch tool from the narrow implementation, then synchronize on turn startup before issuing the tool call. Cap each Codex test process at 12 files so CI gets bounded time-to-first-output as defense in depth.\n\nRefs #125839
* fix(test): keep codex web fetch fixture on sdk boundary
Load the real web_fetch factory on demand through the existing local-only plugin test runtime. This preserves the narrow cold-graph fix without letting a bundled plugin test reach into core internals.
* fix(gateway): stop terminal PTYs on session archive
Bind agent terminals to the durable session incarnation, drain exact ownership during archive, and terminate every job-control process group in the PTY session.\n\nCloses #125769
* test(gateway): cover terminal cleanup on archive
* test(gateway): align terminal outcome assertions
* test(gateway): preserve session exports in invoke test
* fix(gateway): await terminal exit before archive
* test(codex): consolidate supervised instruction coverage
Move the duplicated two-attempt regression into the canonical thread lifecycle test so the exact two-worker extension shard stays bounded on low-core CI.\n\nRelated: #125783
* fix(auth): preserve WHAM classifications and failure recording
WHAM 401/403 state now drives accurate re-auth guidance, while inline hook failures are contained after persistence so recorded failures cannot escape or be masked.
* docs(plugin-sdk): define auth cooldown classifications
document the additive cooldown diagnostic contract and cover its canonical public-SDK projection.
* fix(auth): keep WHAM diagnostics source-compatible
keep cooldownReason canonical, persist exact WHAM diagnostics in optional cooldownClassification, and preserve operator guidance plus failure-hook containment.
* fix(auth): keep failover on canonical cooldown reasons
ensure optional WHAM diagnostics never drive scheduling and discard mismatched persisted reason/classification pairs.
Preserve operator approval terminal reasons through the harness, map denials and timeouts to Codex decline, and retain visible timeout evidence so the turn can continue instead of being killed.
* fix(setup): refresh Codex registry with staged install
* fix(macos): verify inference before onboarding handoff
* fix(setup): use native Codex home for subscription auth
* fix(codex): honor attempt-scoped setup config
* fix(macos): align onboarding handoff with reopen
* fix(setup): await prepared model convergence
* fix(ui): avoid false auth state for empty catalog
* fix(setup): scope catalog convergence to Codex gateway
* fix(setup): publish the committed runtime catalog
* fix(models): project configured static runtime models
* fix(codex): expose app-server model catalog
* fix(models): preserve Codex auth across reloads
* fix(ci): align Codex onboarding checks
* test(ui): stabilize dock suppression environment
* fix(codex): honor discovery config in app-server model catalog
The manifest documents discovery.enabled (bundled fallback list) and
discovery.timeoutMs (default 2500ms) for model discovery; the new catalog
path used the generic 60s request timeout and ignored the enable gate.
Also drop the test-only listModels injection seam in favor of vi.mock.
* fix(setup): refuse prepared Codex auth over an explicit remote transport
configureCodexCliPreparedAuth silently rewrote an explicitly configured
websocket/unix app-server to local stdio (keeping a dangling url), moving
the credential boundary onto this host. Fail setup with actionable
guidance instead; also surface the root cause when the prepared model
catalog refresh fails after activation.
* refactor(agents): one canonical model-catalog identity key
Three near-identical key helpers existed (models-list-result,
models-list-configured-static, harness/model-catalog). Export
resolveModelCatalogIdentityKey from the route-policy owner, collapse the
duplicate dedupe loops into dedupeByKey, make donor enrichment Map-based,
and inline the one-off harness-augment wrapper.
* fix(macos): restore custodian handoff for fresh activations
Landing every finish on the plain dashboard stranded the custodian
first-run flow (memory import, channels, permissions, hatch). Fresh
activations now hand off to custodian onboarding; live-verified
pre-existing setups reopen the normal dashboard, matching the removed
already-configured shortcut. Tests pin the destination per path.
Also isolate the post-startup Codex login test from developer machines:
ambient OPENAI_API_KEY and a real Codex login made it assert-fail.
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(codex): bound no-engine continuity projections to half the context window
A degraded or absent context engine sends fresh-thread continuity through
projectContextEngineAssemblyForCodex with the whole-window projection cap
((window - 20k) x 4 chars), so a large uncompacted transcript renders into
a single turn/start input consuming up to 90% of the model context window.
That turn fills its own native thread, the next turn's token fuse rotates
it, and the following fresh thread re-projects the transcript again -
observed as 11 near-window turn inputs on cold threads in one day
(openclaw/openclaw#125254).
Continuity projections now use a dedicated cap that reserves half the
context token budget, so the fresh thread keeps headroom for later turns
and the existing delta-resume path can actually engage. The active-engine
projection path keeps its whole-window cap unchanged.
Related: #125254
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN
* fix(codex): size continuity projections from real token cost, not the optimistic estimate
The continuity cap reserved half the context window in tokens but converted
that budget to characters with APPROX_RENDERED_CHARS_PER_TOKEN = 4, so at a
258,400-token window it permitted 516,800 chars. A live projection measured
703,134 chars for 226,146 input tokens, meaning that cap really costs about
166k tokens (~64% of the window), not the intended 129.2k.
Codex reports input tokens only after a turn and bounds turn input by
characters, so the projection cannot be sized in verified tokens before it is
sent. Convert the continuity budget at a conservative 3 chars/token instead,
which holds the reserved half in real tokens at the densest ratio observed.
Related: #125254
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN
* fix(codex): scope the continuity sizing claim to the density it was measured at
The half-window claim was stated as a guarantee, but the 3 chars/token
conversion rests on one observed projection. Input that tokenizes more densely
(CJK, base64, minified code) still exceeds the reserved half, so the constant is
renamed to CONTINUITY_EMPIRICAL_CHARS_PER_TOKEN and its comment says plainly
that it is an empirical floor rather than a bound.
The invariant test is narrowed to the measured density, and a companion test
pins the break-even ratio at 3 chars/token so the limitation is visible in the
suite instead of implied. Choosing between a guaranteed worst-case bound and
this empirical cap is a maintainer-owned tradeoff, left open on the PR.
Related: #125254
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN
* feat(codex): size continuity projections from the session's observed token density
Each completed Codex turn now records a calibration sample on the thread
binding: prompt chars actually sent vs the provider-reported input token cost
(uncached + cache read + cache write). The no-engine continuity cap converts
its half-window token budget at that observed ratio instead of a fixed
chars-per-token guess, so capChars / ratio stays at the reserved budget for
any content density - CJK, base64, and minified code included. The sample is
captured before startup rotation so a rotated-away thread's density still
sizes the fresh thread's projection; without a sample the empirical 3
chars/token default applies, and degenerate samples clamp to [0.5, 4].
Related: #125254
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN
* fix(codex): make continuity calibration monotone - samples only tighten the cap
Red-team finding: a loose sample (up to 4 chars/token) followed by denser
content could size the cap past the empirical default, and stale or
non-continuity samples persist on the binding. Clamping the calibrated ratio
at the empirical default makes every such failure mode degrade to the
uncalibrated behavior instead of past it, and the invariant test asserts
monotonicity across poisoned samples directly.
Related: #125254
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN
* fix(codex): record continuity calibration only from continuity projections
ClawSweeper P2: calibration ran after every successful turn filtered only by
prompt size, so a dense direct or active-engine prompt could persist a sample
whose density later shrinks continuity history it never measured. The
no-engine continuity appliers now mark the prompt state, finalize gates the
sample on that marker, and a cross-mode regression proves a large direct
prompt records nothing.
Related: #125254
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN
---------
Co-authored-by: Marvinthebored <262704729+Marvinthebored@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(agents): unify agent status into a durable progress_card
Replace the write-only update_plan to-do tool and the fragmented plan
rendering with one durable status artifact per session: progress_card
({plan?, markdown?}, replace-on-write, 8 KiB markdown / 50-step caps).
Cards persist in a lazy-additive session_progress_cards table in the
per-agent DB (no schema-version bump), broadcast progressCard.changed,
and render from the store with exactly one live placement per view
(session rail when visible, else the composer-adjacent bar); transcripts
collapse to one-line receipts, and the sidebar hovercard shows other
sessions' cards inline (markdown + <progress>, DOMPurify allowlist, no
iframes). The three stream-derived plan renderers and their dedup
heuristics are deleted.
Codex runs disable the native plan tool per thread
(tools.update_plan.enabled=false) and receive progress_card via the
dynamic-tool bridge; compaction restore now reinjects the card (steps +
bounded markdown). Card writes still emit the legacy plan stream event so
native apps and channels keep working until their per-platform
migrations. Policy names map update_plan -> progress_card; the shipped
tools.updatePlan=false kill switch is honored.
Net -277 production LOC; -480 test LOC.
* test(agents): regenerate Codex prompt snapshots for update_plan thread-config disable
* chore(protocol): allowlist progressCard.changed for native apps pending card migration
* fix(ci): repair progress card integration checks
* fix(codex): canonicalize native progress cards
* test(gateway): reconcile progress card method order
* test(codex): stabilize native approval fixture
Show an explicit waiting acknowledgment when sessions_yield ends an otherwise-silent interactive turn, while keeping private resume context out of channel delivery and preserving existing visible replies.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(agents): report interrupted tool outcomes honestly
Prevent restart recovery from claiming that an interrupted tool call completed when no successful matching result was recorded.
* fix(agents): keep ambiguous recovery restart-safe
Classify failed replay-unsafe tool results at the shared restart-recovery owner so interrupted side effects remain unavailable until their external state is verified.
* fix(agents): distinguish missing tool outcomes
Carry the existing missing_tool_result fact into projected Codex transcripts so restart recovery restricts only genuinely unknown outcomes, while confirmed failures remain retryable.