Commit Graph

19830 Commits

Author SHA1 Message Date
wanyongstar 16ab7e8b4a fix(slack): route directory cursor pagination through the shared guard (#120445)
listSlackDirectoryPeersLive/listSlackDirectoryGroupsLive hand-rolled
do/while loops with no repeated-cursor detection and no page bound,
while every other users.list/conversations.list consumer goes through
collectSlackCursorPages. A Slack API or proxy edge that keeps
returning the same non-empty next_cursor made directory queries
paginate forever, growing the member/channel arrays without bound.
Both loops now go through collectSlackCursorPages, which throws on a
repeated cursor and caps total pages.
2026-08-08 13:20:37 -07:00
wanyongstar 9c3e4ce431 fix(browser): bound batch action nesting depth in act request normalization (#120274)
* fix(browser): bound batch action nesting depth in act request normalization

normalizeActRequest recursed over batch nesting with no depth bound, so a
~1MB POST /act body with tens of thousands of nested batch levels parsed
fine and then crashed normalization with RangeError: Maximum call stack
size exceeded before the ACT_MAX_BATCH_ACTIONS count check could run,
surfacing an internal stack overflow as the 400 validation message.

Thread the existing ACT_MAX_BATCH_DEPTH limit through normalizeBatchAction/
normalizeActRequest and reject deeper nesting up front with a clear
'batch nesting exceeds maximum depth of 5' error, matching the bound the
Playwright executor already enforces at dispatch time.

* fix(browser): match batch depth executor boundary

Accept the six wrapper levels supported by the Playwright executor and reject the seventh during request normalization.

Co-authored-by: 万拥 0668000723 <wan.yong@xydigit.com>

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-08-08 12:52:59 -07:00
wanyongstar fe908cf309 fix(matrix): ignore out-of-range hex escapes in env account tokens (#120428)
* fix(matrix): ignore out-of-range hex escapes in env account tokens

decodeMatrixEnvAccountToken guarded String.fromCodePoint with
Number.isFinite, which does not bound the Unicode range: a MATRIX_*
env var whose _X<hex>_ escape exceeds 0x10FFFF (e.g.
MATRIX_A_X110000_B_HOMESERVER) threw RangeError out of
listMatrixEnvAccountIds, crashing discovery of every env-backed
Matrix account during startup and doctor checks. Escapes above the
Unicode max are now rejected like any other malformed token.

* chore(matrix): tighten decoder invariant comment

Co-authored-by: 万拥 0668000723 <wan.yong@xydigit.com>

* chore(matrix): clarify decoder invariant

Co-authored-by: 万拥 0668000723 <wan.yong@xydigit.com>

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-08-08 12:24:41 -07:00
Peter Steinberger 0487b6812c fix(qa): honor Crabbox SSH fallback ports (#120539) 2026-08-08 12:01:30 -07:00
Peter Steinberger 47f78a32eb fix(ai): preserve long Responses sessions after server compaction (#120457)
* fix(ai): preserve Responses server compaction state

Persist opaque Responses compaction items as fenced provider replay state so long stateless sessions can resume authoritative compressed history without exposing it in display or diagnostics. Carry state through worker transcripts and prune replay prefixes without splitting tool pairs.

Release note: Preserve long OpenAI Responses sessions across server-side compaction and worker restarts.

Related: #95788

* test(ai): align long-context fixtures with CI contracts

Make tool-result fixtures type-complete, use the canonical model selector helper, remove unused test-helper exports, and route the paid long-context live probe through the dedicated Gateway profile shard.

* test(ai): type mocked Responses terminal events

Give the mock SSE event collection an explicit open event shape so terminal response events coexist with output-item events under the root test typecheck.

* fix(ai): suppress rejected compaction replay

Persist a route-fenced suppression tombstone when encrypted-content recovery rejects a compaction item, so later turns do not retry the same opaque state. Preserve the tombstone through transcript redaction and cover successful fallback followed by the next turn.

* fix(ai): keep compaction suppression transport-private

Keep the suppression contract local to its sole Responses transport owner and make the regression fixture satisfy root type and lint checks without widening the Plugin SDK surface.

* refactor(ai): remove compaction suppression re-export

* fix(ai): scope compaction suppression to replay route

Keep foreign-route rejection tombstones from hiding the newest compatible Responses compaction while preserving same-route suppression.

* fix(ai): harden Responses replay recovery

Stage encrypted replay recovery so compaction is only suppressed after an attributable rejection. Preserve terminal ordering and keep provider replay within worker frame budgets without truncating opaque state.

* refactor(ai): centralize Responses output indexes

Keep normalized output identity tracking in the stream-slot owner, move response failure state to its diagnostic owner, and remove the obsolete replay clone export so exact-head static gates remain shrink-only.

* fix(ai): retain idless terminal tool identity

Use the canonical empty identity only when a provider supplies neither call nor item id, preventing terminal recovery from duplicating a done-only tool call while preserving stronger identities when available.

* fix(sessions): hide provider replay from public events

* fix(ai): stage encrypted replay recovery

* fix(ai): keep replay attempt kind internal

* fix(ai): route Azure through replay recovery

Use the shared encrypted-content retry owner for Azure Responses so compaction suppression and prompt-observer variants stay coherent across transports.

* fix(ai): harden replay persistence boundaries

Fence Azure replay by the resolved request endpoint, drop invalid replay during transcript sanitization, and surface worker-launch replay omissions through the existing redacted diagnostic path.
2026-08-08 11:55:26 -07:00
Peter Steinberger 01ae18c076 fix(clawrouter): honor advertised reasoning efforts (#120631) 2026-08-08 11:01:34 -07:00
Peter Steinberger eecbfcc960 fix(infra): unify env-truthiness, missing-path, realpath, and abort-sleep semantics (#120359)
* fix(infra): unify environment truthiness

* fix(infra): unify missing path classification

* refactor(infra): unify realpath fallbacks

* fix(infra): unify abortable sleep errors

* docs(infra): clarify path fallback semantics

* fix(infra): route realpaths through policy wrapper

* fix(infra): ratchet plugin SDK wildcard budget

* fix(agents): preserve zero-delay abort precedence

* fix(infra): preserve fallback and media recovery contracts

* fix(plugins): share quarantine path resolution
2026-08-08 10:58:57 -07:00
Peter Steinberger 33f082b86a test(codex): align completed tool metadata expectation (#120637) 2026-08-08 10:04:26 -07:00
Glucksberg 6318e30077 fix(codex): prevent session-changed errors after /new (#113429)
* test(codex): prove replies after retained session reset

Co-authored-by: Markus <markuscontasul@gmail.com>
Co-authored-by: Josh Lehman <josh@martian.engineering>

* test(codex): verify reset idempotency

* test(codex): prove reset keeps binding writable

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Josh Lehman <josh@martian.engineering>
2026-08-08 08:55:25 -07:00
Vincent Koc b5180b6816 fix(codex): support app-server 0.147.0 (#120594)
* fix(codex): support app-server 0.147.0

* docs(codex): clarify marketplace version provenance
2026-08-08 23:07:05 +08:00
Peter Steinberger 88975f85ea fix: required background completion silently disappears (#120453)
* fix(agents): surface silent required completions

* refactor(agents): simplify completion fallback handling

* fix(agents): require visible completion delivery

* fix(agents): preserve resolved session patches

* fix(agents): bound resolved session patch output

* fix(agents): reject silent automatic completions

* fix(agents): settle committed completion side effects

* fix(agents): retain ultra thinking profile

* fix(agents): require destination-safe completion evidence

* fix(agents): block invisible completion side effects

* fix(agents): preserve outbound no-replay evidence

* fix(agents): preserve resolved session identifiers
2026-08-08 07:52:12 -07:00
licheer-zte a32e81c8e8 fix(model-fallback): treat empty non-GPT completions as failed candidates (#120132) (#120148)
* fix(model-fallback): treat empty non-GPT completions as failed candidates (#120132)

Empty and whitespace-only completions from non-GPT models were counted as
candidate_succeeded, silently dropping the turn on visible channels. Apply
the empty/reasoning-only classification to every model; deliberate silent
replies and committed outbound deliveries remain successful.

* fix(model-fallback): classify mixed reasoning-plus-blank completions as failed (#120148)

A completion like [{ isReasoning: true, text: "thinking" }, { text: " " }]
carries no user-visible reply: reasoning text is invisible to the shared
visibility test (includeReasoningPayloads: false), so counting it as visible
made the run look successful and silently ended visible-channel turns.

Filter reasoning payloads out of the empty/whitespace predicate so mixed
reasoning-plus-blank results classify as empty_result (fallback-worthy),
while mixed reasoning-plus-visible-text results stay successful.

Regression tests: mixed reasoning+blank -> empty_result; mixed
reasoning+visible -> success.

* fix(model-fallback): require deliverable assistant results

Use one owner-boundary deliverability predicate for fallback classification, preserve intentional terminal outcomes, and add a mock-channel Gateway scenario for mixed reasoning-plus-blank recovery.\n\nCo-authored-by: 李琪0668001400 <li.qi16@xydigit.com>

* chore: preserve contributor credit

Co-authored-by: 李琪0668001400 <li.qi16@xydigit.com>

* test(qa): cover default model fallback scenario

Make the mixed reasoning-plus-blank fixture recover through both the catalog default alternate and the explicit proof model.

Co-authored-by: 李琪0668001400 <li.qi16@xydigit.com>

---------

Co-authored-by: licheer-zte <licheer-zte@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-08-08 07:48:17 -07:00
Peter Steinberger 725fe96bd7 fix(telegram): avoid credential reads in reply policy
Resolve reply mode through Telegram’s config-only account projection so threading policy never reads token files or SecretRefs.
2026-08-08 06:53:37 -07:00
Peter Steinberger f4f5310355 fix(telegram): restore account-scoped reply mode
Resolve reply mode through the selected Telegram account so account overrides and top-level inheritance reach outbound reply context.

Co-authored-by: Tak Hoffman <781889+Takhoffman@users.noreply.github.com>
2026-08-08 06:53:37 -07:00
wonfong 9fdc266f64 fix(subagents): wake the parent when a follow-up finishes a yielded child (#120187)
* fix(subagents): wake the parent when a follow-up finishes a yielded child

A sub-agent that calls sessions_yield on its own behalf parks its run and
correctly withholds the parent's announce. But a later follow-up to that same
child session registered a sibling registry row instead of continuing the paused
one, so the requester defaulted to the child's own main session and the original
parent — itself idle behind sessions_yield — was never woken. The paused row also
stayed an unsettled descendant, deferring the parent's settle batch forever with
nothing recorded explaining the silence.

Follow-up dispatch now adopts the paused row through the existing post-steer
replacement seam, inheriting the requester identity and carrying the settle-wake
credential forward with its frozen batch membership remapped to the new run id.
A follow-up that names its own requester keeps registering separately, since an
explicit requester is a delivery opt-in that adoption would silently drop.

Also stops frozen-result refill from targeting paused rows: a yield clears the
result on purpose, so refilling from the session would attribute a later turn's
text to the paused run.

Closes #120157

* fix(subagents): select the paused owner past a requester-bound sibling

Adoption looked up the newest run for the child session and adopted it only
when that row was itself paused. A requester-bound follow-up deliberately stays
a sibling, but it registers at a higher generation and becomes that newest row,
so any later default follow-up saw an unpaused newest row, declined adoption,
and registered yet another sibling. The original requester stayed parked behind
a paused row that can never announce -- the same silent stall this fix exists to
remove, reached through a valid mixed-delivery sequence.

The latest-run query now takes an optional predicate applied before the
generation comparison, so a caller that owns a specific row class selects the
newest row of that class. Adoption asks for the newest `sessions_yield` row
directly instead of inferring it from generation order.

Docs now state that continuation applies to default delivery, since a follow-up
carrying its own requester runs as a sibling by design.

* test(qa): prove post-yield follow-up delivery through the gateway boundary

The unit and gateway-method tests for paused-run adoption assert on registry
rows, which proves the bookkeeping but not that an operator ever sees the
result. This adds the boundary proof: a real gateway child, the QA mock channel,
and the mock provider driving a subagent that pauses itself and finishes only on
a later follow-up.

A fixture plugin owns both legs. Its `before_dispatch` hook spawns the child with
`completionDelivery: "current-requester"`, so the announce has the operator turn
as its audience. An HTTP route then dispatches the follow-up to that same paused
session using default delivery -- the path adoption is meant to catch. A
requester-bound follow-up would opt into its own audience and run as a sibling
instead, so the two legs must differ here.

The mock provider gains a child that yields on its own behalf. Both of its turns
match on the current prompt rather than the shared transcript, so the yielded
kickoff cannot make the follow-up turn yield a second time.

The scenario asserts both sides of the invariant: no outbound traffic while the
child is paused, and exactly one announce carrying the follow-up marker once it
ends.

Reverting the adoption call site fails this test in the way that matters: the
child still produces its marker and the run still ends with stopReason=stop, but
nothing reaches the requester and the wait times out. The result is computed and
then silently dropped -- which is the failure this repair exists to remove.

* fix(ci): match QA Lab fixture plugin entries as a group in knip

The all-exports pass listed one fixture entry by name, so every new QA Lab
fixture plugin lands as an unused file and turns check-dependencies red until
someone remembers this file. Nothing imports these entries by design: the
Gateway E2E loads them through plugin config paths.

* docs(subagents): scope yield continuation to plugin runtime follow-ups

Adoption is gated on plugin_subagent task tracking, which only
createGatewaySubagentRuntime().run sets, so api.runtime.subagent.run is the
sole route into it. Writing that as one example implied other follow-up paths
to a paused session continue the run too; they are not tracked as sub-agent
runs and announce nobody.

* fix(subagents): reject undurable paused-run adoption

Fail plugin follow-up admission closed when the paused-run ownership swap cannot be persisted, while retaining the existing restart-recovery return-false contract. Trim duplicate tests and keep boundary coverage for requester routing, wake-batch remapping, repeated yield, and persistence rollback.

Co-authored-by: zhou.huanfeng <woundfongv3@163.com>

* docs(subagents): clarify yielded-run steering

Co-authored-by: zhou.huanfeng <woundfongv3@163.com>

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-08-08 06:46:19 -07:00
Peter Steinberger 83900e4683 feat(browser): add relay authentication v2 (#120526)
* feat(browser): add relay authentication v2

* fix(browser): cap relay test WebSocket payload

* fix(browser): keep relay E2E inside extension boundary

* fix(browser): isolate relay admission and cleanup auth

* fix(browser): finish relay auth migration hardening

* fix(browser): keep preauth transport bounded through teardown
2026-08-08 05:48:24 -07:00
Vincent Koc 4eec51a276 merge: align Fish Audio extension directory (#120084) 2026-08-08 20:46:55 +08:00
Vincent Koc 6e100ca56a test(codex): remove duplicate session helper 2026-08-08 18:47:18 +08:00
Peter Steinberger 10c7f005be refactor(github-copilot): narrow runtime facade (#120544)
* refactor(github-copilot): narrow runtime facade

* style(github-copilot): format runtime exports
2026-08-08 02:56:12 -07:00
Vincent Koc cdd621eaee test(telegram): use model selection rejection seam 2026-08-08 02:21:34 -07:00
Vincent Koc daca53c985 test(codex): align file-backed session fixtures 2026-08-08 02:21:34 -07:00
Vincent Koc e5f5030570 fix(plugins): align Fish Audio extension directory 2026-08-08 02:21:33 -07:00
Peter Steinberger c5d00cb47d fix(ui): keep task transcripts in the task sidebar (#120463)
* fix(ui): keep task transcripts in the task sidebar

* refactor(ui): split background task rail rendering

* fix(ui): reset stale task rail views

* fix(sessions): honor context usage provenance

Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>

* fix(qa): split suite runtime agent process below line cap

---------

Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>
2026-08-08 01:45:44 -07:00
Peter Steinberger d9080cfff8 fix(agents): transfer requester wake ownership across yield and deflake survivor registry setup (#120538) 2026-08-08 01:42:33 -07:00
Vincent Koc bf0aadbc40 fix(qa): settle Unix commands after PGID cleanup (#120010)
Punchcard-Session: cobalt-valley-meadow-mg
2026-08-08 15:16:44 +08:00
Peter Steinberger c7b7fe4c32 fix(codex): restore continuity test coverage (#120452)
* test(codex): align run history fixture identity

* test(ci): repair gateway fixture checks
2026-08-07 23:39:26 -07:00
Peter Steinberger fa0fcef9f6 fix(sessions): require provenance for fresh context-size facts (#120497) 2026-08-07 23:06:30 -07:00
Peter Steinberger 947bca5608 fix(browser): harden extension relay authorization (#120390)
* fix(browser): harden extension relay authorization

* fix(browser): validate persisted relay pairing
2026-08-07 21:56:15 -07:00
zengLingbiao 0eeb9db216 fix(voice-call): redact credential material from provider error bodies (#117304) 2026-08-08 11:51:23 +08:00
Alix-007 1575187419 fix(browser): time out Google Docs page exports (#111633) 2026-08-08 11:47:58 +08:00
Alix-007 50ec5570ae fix(synology-chat): prevent duplicate messages after ambiguous failures (#119539)
* fix(synology-chat): avoid ambiguous send replays

* fix(synology-chat): classify aggregate send failures

* test(synology-chat): trim retry loopback proof

---------

Co-authored-by: Dallin Romney <dallinromney@gmail.com>
2026-08-08 11:46:04 +08:00
Peter Steinberger 80ae3506aa test(clickclack): cover native progress default (#120410) 2026-08-07 19:02:35 -07:00
Peter Steinberger 0463d1bd83 test(qa): derive UX producer aggregate status (#120418) 2026-08-07 18:55:19 -07:00
Peter Steinberger 48639663b0 chore(release): prepare 2026.8.1 (#120375) 2026-08-07 18:44:12 -07:00
Peter Steinberger 5a79d19ba1 fix(codex): preserve warm sessions and approvals across conversations (#120405) 2026-08-07 18:38:09 -07:00
Peter Steinberger 96a75be170 fix(discord): default preview streaming to off (#120376)
Make Discord progress drafts and activity receipts explicit opt-ins while preserving configured streaming modes and legacy explicit progress migrations. Related: #87704.
2026-08-07 17:18:12 -07:00
Peter Steinberger fa9626c4e1 refactor(steering): make active-run admission exact and atomic (#120285)
* refactor(steering): make admission exact and atomic

* fix(steering): bind targets by run identity

* fix(steering): inject before reply dispatch

* fix(steering): preserve backend injection modes

* fix(steering): target session suggestions

* fix(steering): preserve sibling delivery contracts

* style(steering): satisfy exact-head lint

* fix(steering): preserve targetless protocol compatibility

* fix(steering): preserve exact targets through replies
2026-08-07 16:58:04 -07:00
joshavant 5cc2fa6249 fix(agents): defer recovered followups until drain 2026-08-07 18:42:57 -05:00
joshavant 7bcdb5823e test: prove repeated request recovery through gateway 2026-08-07 18:42:57 -05:00
joshavant ceeb02d7e1 Revert "refactor(agents): bind tool policy to exact execution identity"
This reverts commit d863dd5f5e.
2026-08-07 18:40:17 -05:00
joshavant 2751bc8991 Revert "fix(agents): restrict harness tool authority"
This reverts commit aacbcaacc8.
2026-08-07 18:40:17 -05:00
joshavant bc197e1b66 Revert "fix(plugins): inject harness tool authority"
This reverts commit 228478bffe.
2026-08-07 18:40:17 -05:00
joshavant 941f4a3ec6 Revert "fix(agents): satisfy tool authority lint"
This reverts commit 63c6cd9a02.
2026-08-07 18:40:17 -05:00
Peter Steinberger bc0c3fcd63 refactor(providers): shared generated-media download guard and binary-response adoption (#120351)
* fix(agents): cancel rejected provider binary bodies

* refactor(providers): share generated media downloads

* refactor(providers): adopt payload patch stream wrapper

* fix(providers): preserve generated media helper contracts
2026-08-07 15:45:53 -07:00
Peter Steinberger 45b1dbf962 test(qa): add managed-worktrees CLI lifecycle scenario coverage (#120335)
* test(qa): add managed-worktrees CLI lifecycle scenario coverage

Managed worktrees had zero QA scenario-pack coverage despite being a
headline feature. Mint agent-runtime.managed-worktrees-lifecycle in the
taxonomy, add a runtime scenario, and prove the real child CLI through
create with .worktreeinclude provisioning and the .openclaw setup hook,
dirty removal pinning a snapshot ref, restore rebuilding tracked,
untracked, and provisioned files with their modes, and gc preserving
manual worktrees.

* fix(qa): align model-switch catalog assertion with expectedAlternate flow

qa/scenarios/models/model-switch-follow-up.yaml switched to
expectedAlternate.model in 5a795f4dda but the catalog test still greps
for the retired alternate?.model literal; the test is outside the PR
change-classification lanes, so the break only surfaces on direct runs.

* test(qa): narrow managed-worktrees taxonomy description to proven manual-owner gc

ClawSweeper P2 on #120335: the scenario proves manual-owner gc retention
only; session and Workboard cleanup lifecycles are not exercised, so the
coverage description must not claim them.
2026-08-07 14:18:12 -07:00
joshavant 87156bab23 fix(typing): keep long active turns visible 2026-08-07 16:11:52 -05:00
Peter Steinberger 32361e749a fix(voice-call): isolate runtime generations (#120289)
Prevent retained plugin closures from recreating, adopting, or stopping successor runtimes across shutdown and restart.

Add deterministic lifecycle regressions for pending startup, exact-owner stop, retained tools, and generation restart.
2026-08-07 13:56:26 -07:00
Peter Steinberger 10e60fa0ce refactor(plugins): shared legacy-state doctor migration and simple secret contracts (#120346)
* refactor(plugins): share legacy JSON doctor migration

* refactor(discord): share account token inspection cascade

* refactor(plugins): share simple channel secret contracts

* refactor(discord): keep token inspector private
2026-08-07 13:55:31 -07:00
Josh Avant c691f2e41c fix(progress): preserve callback acceptance results (#120171)
* fix(progress): preserve callback acceptance results

* fix(progress): require transport acknowledgements

* fix(progress): preserve direct acceptance outcomes
2026-08-07 14:40:33 -05:00
Peter Steinberger 6543e6f7c9 fix(discord): thread archive/delete closes sessions in each agent's store (#120259)
* fix(discord): thread archive/delete closes sessions in each agent's store

closeDiscordThreadSessions resolved the sessions store with the Discord
account id as agentId ('default' on the default path), which points at a
nonexistent agent's store — archiving or deleting a thread silently closed
nothing. The store now resolves per routed agent via listAgentIds and every
agent's matching sessions are deleted.

* fix(discord): type thread session cleanup across agent stores

* fix(discord): scope thread-session scan per agent and keep it read-only

* test(discord): cover thread deletion across agent stores
2026-08-07 12:08:36 -07:00