Files
openclaw/extensions/codex
Marvinthebored 3d016975c3 fix(codex): degraded-engine continuity no longer projects the whole context window per turn (#125324)
* fix(codex): bound no-engine continuity projections to half the context window

A degraded or absent context engine sends fresh-thread continuity through
projectContextEngineAssemblyForCodex with the whole-window projection cap
((window - 20k) x 4 chars), so a large uncompacted transcript renders into
a single turn/start input consuming up to 90% of the model context window.
That turn fills its own native thread, the next turn's token fuse rotates
it, and the following fresh thread re-projects the transcript again -
observed as 11 near-window turn inputs on cold threads in one day
(openclaw/openclaw#125254).

Continuity projections now use a dedicated cap that reserves half the
context token budget, so the fresh thread keeps headroom for later turns
and the existing delta-resume path can actually engage. The active-engine
projection path keeps its whole-window cap unchanged.

Related: #125254

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN

* fix(codex): size continuity projections from real token cost, not the optimistic estimate

The continuity cap reserved half the context window in tokens but converted
that budget to characters with APPROX_RENDERED_CHARS_PER_TOKEN = 4, so at a
258,400-token window it permitted 516,800 chars. A live projection measured
703,134 chars for 226,146 input tokens, meaning that cap really costs about
166k tokens (~64% of the window), not the intended 129.2k.

Codex reports input tokens only after a turn and bounds turn input by
characters, so the projection cannot be sized in verified tokens before it is
sent. Convert the continuity budget at a conservative 3 chars/token instead,
which holds the reserved half in real tokens at the densest ratio observed.

Related: #125254

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN

* fix(codex): scope the continuity sizing claim to the density it was measured at

The half-window claim was stated as a guarantee, but the 3 chars/token
conversion rests on one observed projection. Input that tokenizes more densely
(CJK, base64, minified code) still exceeds the reserved half, so the constant is
renamed to CONTINUITY_EMPIRICAL_CHARS_PER_TOKEN and its comment says plainly
that it is an empirical floor rather than a bound.

The invariant test is narrowed to the measured density, and a companion test
pins the break-even ratio at 3 chars/token so the limitation is visible in the
suite instead of implied. Choosing between a guaranteed worst-case bound and
this empirical cap is a maintainer-owned tradeoff, left open on the PR.

Related: #125254

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN

* feat(codex): size continuity projections from the session's observed token density

Each completed Codex turn now records a calibration sample on the thread
binding: prompt chars actually sent vs the provider-reported input token cost
(uncached + cache read + cache write). The no-engine continuity cap converts
its half-window token budget at that observed ratio instead of a fixed
chars-per-token guess, so capChars / ratio stays at the reserved budget for
any content density - CJK, base64, and minified code included. The sample is
captured before startup rotation so a rotated-away thread's density still
sizes the fresh thread's projection; without a sample the empirical 3
chars/token default applies, and degenerate samples clamp to [0.5, 4].

Related: #125254

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN

* fix(codex): make continuity calibration monotone - samples only tighten the cap

Red-team finding: a loose sample (up to 4 chars/token) followed by denser
content could size the cap past the empirical default, and stale or
non-continuity samples persist on the binding. Clamping the calibrated ratio
at the empirical default makes every such failure mode degrade to the
uncalibrated behavior instead of past it, and the invariant test asserts
monotonicity across poisoned samples directly.

Related: #125254

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN

* fix(codex): record continuity calibration only from continuity projections

ClawSweeper P2: calibration ran after every successful turn filtered only by
prompt size, so a dense direct or active-engine prompt could persist a sample
whose density later shrinks continuity history it never measured. The
no-engine continuity appliers now mark the prompt state, finalize gates the
sample on that marker, and a cross-mode regression proves a large direct
prompt records nothing.

Related: #125254

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKtoZgXWnaAH8rmLiSydpN

---------

Co-authored-by: Marvinthebored <262704729+Marvinthebored@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 17:41:45 -07:00
..

OpenClaw Codex

Official OpenClaw plugin for OpenAI Codex app-server integration. It exposes the Codex-managed GPT model catalog, the Codex runtime surfaces used by OpenClaw agents, and opt-in supervision of native Codex sessions.

Install from OpenClaw:

openclaw plugins install @openclaw/codex

Use this plugin when you want OpenClaw to run Codex-backed model turns, media understanding, and prompt overlays through the Codex app-server harness, or to browse non-archived Codex CLI, VS Code, Atlas, and ChatGPT sessions and paginated transcripts across paired computers.

Guided onboarding attempts to install and enable supervision after it detects a native Codex installation and the selected inference backend passes its live check; Codex does not need to be the primary backend. Supervision activates when that opportunistic plugin setup succeeds. App Server availability is checked when supervision connects. An explicit Codex plugin disable, plugin-policy block, or supervision.enabled: false prevents opportunistic enablement. Manual setups enable plugins.entries.codex.config.supervision.enabled. Without explicit App Server connection settings, supervision uses a managed user-home stdio connection; explicit appServer settings are honored.

The Gateway-backed operator CLI is:

openclaw codex sessions [--search <text>] [--host <id>] [--limit <count>] [--cursor <cursor>] [--json] [--url <url>] [--token <token>] [--timeout <ms>] [--expect-final]
openclaw codex continue <thread-id> [--json] [--url <url>] [--token <token>] [--timeout <ms>] [--expect-final]
openclaw codex archive <thread-id> --confirm-no-other-runner [--json] [--url <url>] [--token <token>] [--timeout <ms>] [--expect-final]

The catalog never includes archived threads and has no archived or include-archived option. Rows appear in the normal Control UI sessions sidebar and open in the normal Chat pane. Transcript history requires a recent Codex App Server with thread/turns/list and is fetched 20 full-item turns at a time through opaque cursors; OpenClaw does not fall back to an unbounded thread/read, and rejects a serialized transcript page above 20 MiB before transport. --limit defaults to 50 sessions per host, --cursor requires --host, and the sessions Gateway timeout defaults to 75,000 ms so cold paired-node catalogs can complete. Continue and archive retain the shared 30,000 ms default. All operator surfaces require operator.write. Paired-node rows can be listed and read; continue and archive operate only on the Gateway-local host, and archive requires the no-other-runner confirmation. Catalog registration does not require supervision.enabled; that setting gates agent-facing supervision tools.

A supervised OpenClaw Chat cannot be deleted while its model-selection lock protects the native binding. Before native archive, OpenClaw checks the exact target and every non-archived spawned descendant reported by Codex; any active OpenClaw binding blocks the operation. Descendant pagination errors, cycles, and safety-limit exhaustion also fail closed. Codex still does not expose a conditional archive operation or cross-process runner lease, so the confirmation covers unknown native clients and the race between the status read and archive request.

Disabling or uninstalling the plugin leaves supervised Chats locked and unavailable rather than rerouting them. Reinstall or re-enable the same plugin and restart the Gateway to resume those Chats.

These shell commands differ from the in-chat /codex runtime commands. In particular, /codex sessions --host <node> lists Codex CLI session files on one node, /codex threads uses the current conversation's App Server connection, and /codex resume or /codex bind changes that conversation's binding. There is no /codex archive runtime command.

Native Codex plugin catalogs are discoverable with /codex plugins available, including repository marketplaces declared in .agents/plugins/marketplace.json in the bound workspace. An owner or operator.admin can install and authorize an exact plugin with /codex plugins install <plugin>@<marketplace>. The owner-scoped codex_plugins agent tool only reads marketplace metadata; installation and policy changes stay on authenticated /codex management commands. Explicitly installing a plugin trusts its skills, apps, MCP servers, and hooks.

For a supervised branch, Codex App Server selects the snapshot fork's model and provider from its current native configuration. OpenClaw starts the canonical harness thread with exactly that returned pair. Codex persists the canonical thread's native selection, and later resumes preserve it because OpenClaw omits model and provider overrides. OpenClaw cannot substitute its outer runtime, model, or fallback. The returned initial pair can differ from the source's last recorded model.

The visible-history mirror keeps at most 200 user or assistant messages, 512 KiB total, and 64 KiB per message. Image inputs become [Image attachment]; image data and local paths are not copied.

See the Codex harness and Codex supervision guides.