mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
28ef63a10cdb7323f0aee461ffaeddd9b302d596
50 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
28ef63a10c | fix: revalidate frontend assets across builds | ||
|
|
480a1426b3 |
Fail-closed history-commit handoff (#1005)
* fix(session): fail-closed history-commit handoff (#981) The deleted-workstream discovery is now a terminal, ws_id-keyed latch: keyed conversation commits refuse admission once the durable parent is gone (convergence finalizers and force-abandon are exempt), history handoff refuses to mint a proof token so /history fails closed with a 503 instead of silently wiping the pane, and the SSE stream carries a workstream_gone resync reason. Discarded commits leave a forensic log of commit keys and roles, never content. Conversation rows gain a commit_key (migration 071): keyed saves are idempotent under retry, validated against the full commit identity, and refused when they would cross a workstream deletion. The prune orphan category now requires a NULL alias plus a two-hour updated grace, with cutoffs computed at discovery time and carried into both dialects' rechecks. The mid-turn interjection queue is owner-partitioned with no per-site mode flags: pops take the acting principal's and unowned rows, other participants' rows are structurally retained, and enforcement lives at queue admission plus the shared before_spawn gates. The retraction ledger is bounded by open pop windows: pops open a window atomically with the queue delete, restores close their ids atomically with the ledger consume, every other exit closes through one helper, and misses for unheld ids record nothing. The workstream-gone latch refuses unattended wakes at all three gates (watcher spawn, claim, delivery pre-pop), and the retry dispatcher regained its pre-envelope cancel/error convergence net. Persistence-state reporting derives through the session bound to each UI instead of a registry lookup by id that failed open to healthy during tombstone retention. The dashboard roster no longer re-inserts ghost entries from trailing activity events, the history tool-outcome scan tolerates interleaved non-turn rows, and the shared handoff-deadline handle owns its own retirement. Single-sourced across call sites: keyed-commit row values, attachment save wrappers, tail-truncation and conflict-resolution bodies for both storage dialects; worker-slot lifecycle field sets; the direct-commit admission frame; queued-row layout accessors; the string-aware comment stripper shared by every JS harness suite. Refs #981 #964 * fix(session): sweep handoff fixes to their sibling surfaces The interactive replay loop treated a system row as a tool-batch boundary, so every tool result after an interleaved row vanished from that pane while the coordinator rendered the same history correctly. Only a conversational turn ends the batch window now, matching the shared outcome index. Accepted user turns clear the composer's attachment chips on the same viewer policy that settles optimistic bubbles rather than on having matched a local bubble, so a workstream created with an upload no longer keeps a chip for an attachment the create dispatch already consumed. The coordinator's raced-Stop arm emits the stream-end hook it inherits alongside the idle state, leaving no unfinalized bubble or unflushed tool output. Ending a session surfaces a failure toast when the request never lands or answers with a non-JSON body. The per-second persistence reconcile now probes each session without blocking: a workstream whose generation and handoff locks are held is skipped until the next pass instead of contending the locks every commit needs. The one-shot repair that gates workstream creation at capacity keeps a definite probe — it has no next pass, and the sessions likeliest to be contended are the ones whose unresolved journals emptied its candidate list. Single-sourced: the attachment lane builds its conversation row through the shared commit-identity builder; the ordinary worker exit releases its slot through the lifecycle owner; both operator surfaces snapshot their counters through one non-consuming helper; the replay preamble loses its per-kind wrappers and its config hook; the browser harness suites share one brace walker; and each in-flight history attempt is one record carrying both its abort controller and its deadline. Refs #981 #964 |
||
|
|
0d52b63b50 |
feat(coordinator): nudge a coordinator that goes idle holding unfinished work
Two nudge classes can fire from one IDLE event, tasks first, each asserting only its own domain. idle_tasks (advice) fires when open (pending/in_progress) tasks exist. The body is a counts opener, the open task ids with statuses, and typed branches that each end in a runnable tasks(...) or wait_for_workstream(...) call populated with real server-minted ids. Everything it says about children is governed by one observed fact: live children present adds the caveat sentence and the blocked-on-a-child branch; affirmatively none says nothing about children at all. Any needs_user row parks the class entirely — at the fire gate and the drain predicate — because with no task graph an open task may be gated on a parked one's unanswered question; the operator's answer is the re-arm. Gated on memory.nudges and on the persona actually exposing the tasks tool; carries the per-class cooldown as well as the per-bracket cap. The tasks tool itself gains the needs_user status and a note field — the typed escalation the body's branches point at. idle_children (liveness) fires when children are in a live state — the wake that lets an idle coordinator collect a finished child's results. The body is a roster of workstream id prefixes and states, never names: a child's name is model-authored text and does not enter a system turn. Cap-only and cooldown-free by design, not gated on memory.nudges, and it survives an operator Stop. Fail-closed, event-wide: if any storage read fails while the observer handles an IDLE event, neither nudge is queued and neither cap is charged. Both paths run as side-effect-free plans; the commit tail is storage-free, so no read can fail past the veto point; both drain predicates drop on a failed read. A path's own fault (a generic raise) still costs only that path's fire, so one class's bug cannot strand the other. Task text is stored verbatim and projected per audience at render: the model-facing projection deletes angle brackets, the operator projection keeps them, and both strip newlines and bidi/zero-width runs. Idle cards render what the model was told, formatted for the operator, never augmented with content the model did not receive. |
||
|
|
4007fab855 |
fix(#900): close two vacuity holes the round-2 scenarios left open
Round-3 review, unprimed. Two of its majors were the new scenarios asserting things they did not prove — the false-detector class this campaign keeps returning to. E8 never checked that the held /history was still OUTSTANDING when the redial completed. The disconnect/send/wait_turn/redial sequence is unbounded (wait_turn alone allows 45s), so on a slow box the payload resolves while evtSource is still null, the PRESENCE term declines it, and the run stamps dupes1-healed1 without ever evaluating the generation term. It now fails loudly with the counter values instead. E7 gained the same positive proof its siblings already carried: sse_opens == 0 only means "nothing connected in 8s", which is not the same as "the held load settled and its .finally chose not to reconnect". E5's stated control was simply wrong, in three places. A hide nulls evtSource, and connectSSE early-returns while hidden, so the scenario cannot produce the non-null-but-not-OPEN source that readyState === OPEN exists for — it exercises the presence term only. The earlier control removed both terms at once, which is what disguised it. The readyState half is covered by reasoning plus coord parity, and its correctness twin IS covered through the render-time gate by E6/E8; that scope is now written down rather than overclaimed. Coord's G5 has the same shape. The retry floor becomes a shared export beside its jitter: four sites must move together (both clients' arms, both non-occurrence windows) and it was the only one of them with no single source of truth. Interactive's use of the expression had no pin at all — reverting it to a bare 2000 would have broken cross-client parity with the suite green. Coord's re-anchor still raised ValueError rather than failing on a named assertion, and its first replacement used a fixed window that truncated mid-expression. |
||
|
|
7daf3b782d |
fix(#900): stream generation closes the reconnect-inside-the-await render; jitter both retries
Round-2 review. The render-time cursor-safety gate was point-in-time: a transport that dropped AND finished re-establishing inside the /history await reads back OPEN and is indistinguishable from one that never moved. It is not — the redial re-presented the frozen cursor, the server answered replay_ok, and the quiesce buffered that slice, so the render commits rows the flush then repaints on top. Object identity cannot see it either, since a native reconnect reuses the same EventSource; only a counter can. _connectEpoch is bumped in onopen and nowhere else. Native auto-reconnect calls neither connectSSE nor disconnectSSE, so those two are blind to the exact case this exists for; connectSSE would also false-bump on its document.hidden early return, which establishes no stream; and a closed source can never fire a late open. Captured at dispatch, required unchanged before a seedless render commits. This is original-strata residual, not a regression this branch introduced: before #900 the render was ungated entirely. The branch closed the fire-time half and the still-down cases; these are the drop-and-recover ones that were always open. Two rulings written in at the gate, since neither is closed: a fresh/truncated reconnect inside the await declines a render that would have been safe (one wasted /history, self-healing via the flushed synthetic state_change), and a refetch dispatched between onopen and the replay slice arriving still renders past the frozen cursor — replay_ok emits no end-of-replay marker, so no client-side signal exists (#903). Coord's half of the same gate is #904; its exposure is a race rather than this determinism, so it is not ported blind. Also corrected: the claim that the idle-edge backstop's stream is live by construction. It isn't — handleEvent also runs from the quiesce flush, so a queued idle edge reaches the backstop with the transport down. The render-time gate is what covers it. The clear_ui retry gains additive jitter in BOTH clients from one shared constant: a declined render now leaves the latch set, so a successful fetch can arm the retry, and the decline trigger is herd-shaped. Kept small deliberately — the spread works against #884's single-flight, which coalesces a lockstep herd. test_coordinator_page.py anchored the fire guard on a literal `}, 2000);` and on exact indentation; both would have ERRORED rather than failed once the delay became an expression. |
||
|
|
d8d026394f |
fix(#894): cold flights key on None; typed generation access; abort-Set producer pins
Review round 10 (1 minor bug; 2 major + 2 small quality — the majors both pins-that-cannot-fail). - The flight key's cold fallback was the literal 0, which collides with a live session's generation 0: an eviction/close landing inside a held flight's window let a post-truncation request rejoin a generation-0 pre-truncation flight. Cold/detached workstreams now key on None (rewinds need a live session, so two cold flights are always mutually safe; a rehydrated session restarting at 0 can never share the manager slot with its evicted predecessor — documented at-site). The read is TYPED (live_session.session._history_generation) so mypy carries the shape a getattr chain hid — and the typed access immediately surfaced an unfaithful SimpleNamespace mock in the reasoning-rehydration tests (no .session attr), now made faithful. - Abort-Set producer pins: histCtrls.add exactly once and BEFORE the await, delete exactly once and in the finally — without them the destroy() consumer sweep was satisfiable by an always-empty Set. - _make_session gains ws_id; the generation producer pin uses it. - _coord_stick_latch: G2/G5's inline single-failure prologues RULED deliberate at-site (their baselines/phase timings interleave into the prologue; a per-divergence flag would obscure the choreography). - Stray trailing whitespace stripped. 250 pins green; G2/G5/G7 re-run READY. |
||
|
|
60f6dc07a2 |
fix(#894): drop the unreachable epoch guard; abort-Set; bump-after-delete; producer pins
Review round 9 (4 minor bug, 4 quality, 1 perf nit; security zero). - The r8 clearUiEpoch guard was UNREACHABLE (r9 bug find): clear_ui always dispatches immediately after bumping, so a stale-epoch dispatch is also a stale-seq dispatch and the currency gate discards it before it can paint or clear — the client half of the joined- flight fix was already carried by seq, and the server generation key is the sole load-bearing layer. Machinery removed (decl, bump, capture, conditional clear, section-9 pins); the latch-clear comment now states the two-layer accounting. - destroy()'s abort handle becomes a Set: a newest-wins single slot, nulled by the newer dispatch's finally, left an OLDER overlapping fetch unabortable — the destroyed closure pinned for the bound's remainder. Pinned. - _history_generation now bumps AFTER delete_messages_after: flights rebuild from storage, so old-generation-reads-post-delete is the harmless spuriously-fresh direction while new-generation-reads- pre-delete would be wrongly joinable; the count/floor error paths correctly leave it unbumped. Two-arm producer pin in test_rewind_retry (persisted-rows bump on rewind AND retry; in-memory-only error path must NOT bump) — the flight test's mock can no longer mask a deleted bump. - The harness load_calls increment takes a lock (to_thread workers genuinely overlap under delay_load; a lost update false-fails G7). - G7's viewer B is now a background authenticated GET (a raw request enters load_messages identically; the second browser bought no proof); stale two-tuple key comments and the coalescing matrix line updated; the _send_in_page enumeration dropped for prose. 250 pins green; G1/G6/G7 re-run READY. |
||
|
|
30b6ff7f7b |
fix(#894): rewind-freshness epoch closes the #884 joined-flight window; destroy aborts the bounded fetch
Review round 8 (2 major + 2 minor bug, 2 major + 3 small quality; security/perf zero at five consecutive rounds). - clearUiEpoch (r8 major): the #884 /history single-flight can hand a joiner a payload whose load_messages ran BEFORE the rewind committed (the flight key is (ws_id, limit); joining is invisible to the client, and the client seq stamp cannot see server-side staleness) — reachable single-user (rewind clicked during a truncated-resync fetch joins that flight) and multi-viewer (any concurrent pane's /history). The joined payload rendered as 'success' and CLEARED the latch: the original over-rewind window, resurrected through the server seam. Fix: the epoch bumps at clear_ui arrival, every dispatch captures it pre-await, and only a dispatch that post-dates the latest clear_ui may CLEAR the latch — a pre-rewind payload may still paint (stale-but-real posture, gate holds), the surviving latch arms the retry, and the retry's fresh dispatch starts a new flight with post-rewind truth. Producer/consumer/placement pinned. - destroy() aborts the in-flight bounded fetch (activeHistCtrl): the r7 15s bound alone pinned a destroyed pane's closure until it fired — the same dead-not-inert ruling destroy applies to staleRetryTimer. Pinned. - stop(hard=True) no longer sets force_exit: it skipped the ASGI lifespan teardown and leaked the #885 daemon threads + sse_executor. The 2s graceful-shutdown timeout already force-closes open SSE, and the lifespan runs on both paths (docstring corrected; G6 re-verified — the orphan still manifests). - G6 pacing sized above the scenario's worst-case deadline sum (~200s vs ~95s) so the in-process bash cannot resolve the orphan mid-scenario and degrade the detector to a false READY. - Quality: the r6 reachability comments rewritten to the r7 truth (orphan REAL via hard crash; graceful-close-only synthesis); the bound's WIRING pinned (signal reaches getJSON; getJSON forwards init); _strip_comments deduped (4 inline copies); seq comment re-paired with its asserts; retry_fire window tail-anchored. Full harness (16/16 scenarios) + full suite (9724) green on the prior commit; 136 pins green here. |
||
|
|
412161aa4a |
fix(#894): live-set retirement policy — transport death is not retirement; bound the refetch await; G6 hard-kill detector
Review round 7 (2 major + 1 minor bug, 4 minor quality; security/perf zero). Both majors traced the r6 stratum: - The closeStreamTransport drain of liveToolCalls rested on a false re-announcement premise (verified: replay_ok yields only events past the cursor; the coord fresh/truncated replay yields connected/status/ pending-cards/verdicts, never tool_pending/tool_info). An emptied set fails OPEN — a mid-batch redial plus a slow seedless refetch wiped the live batch. Retirement policy re-derived at the decl: an id leaves on its RESULT, at the SETTLE edge, or with pane death; transport death is NOT a retirement event; a stale id fails CLOSED (skip, latch survives, settle heals). Site-anchored pins: the one drain inside the idle/error block, the delete inside tool_result, the adds inside tool_pending/tool_info, and closeStreamTransport's comment-stripped code may not touch the set. - G6's kill was not a kill: RecoveryServer.stop() gracefully closed workstreams, and session.cancel()'s bash path persisted 'Cancelled by user' BEFORE the reboot — the r6 'recovery synthesizes' ruling was observing the cancel path. stop(hard=True) (skip the close sweep + uvicorn force_exit: a crash does not drain SSE) leaves the orphan genuinely unresulted — REACHABILITY FLIPS: the poisoned-pane state is real, the live-set hardening is reachably load-bearing, and G6 is now its behavioral detector: hard kill -> reload paints the orphan (asserted PRESENT) -> the seedless rewind renders THROUGH the residue (rewind-for-retry truth: the user message stays), negative-controlled against the DOM-probe encoding (stamps orphan1, hist2). Discovered and tracked separately: a hard-crashed reborn node answers stale-high cursors with a silent fresh stream (no replay_truncated — the honest truncation signal rides gracefully-persisted state). - refetchHistory's await is now bounded (AbortController + 15s, the coordSend shape): an accepted-never-answered /history pinned refetchesInFlight and permanently disabled both heals. Pinned. Quality: seq-producer position pinned earlier; the stale section-7 comment corrected; retry/backstop guard windows comment-stripped (vacuous-by-comment-mention foreclosed); the shared G3/G4 double-fail prologue extracted into _coord_stick_latch. G3/G4/G6 re-run READY. |
||
|
|
cc776bfb2d |
fix(#894): event-driven live-tool-call set replaces the DOM liveness probe; G6 synthesis tripwire
Review round 6 (1 bug find + 5 quality; security/perf zero). The bug
finder out-traced r6-perf's dismissal: refetchHistory's own replay path
paints orphan batches (committed tool_calls, no persisted result) with
the same .conv-batch--running class the live path uses, and nothing
ever strips a dead orphan's class — so the r5 DOM-probed gate term
would let one orphan paint poison every seedless heal for the life of
the page (rewind/edit permanently dead; the seeded escape renders
through but REPAINTS the residue).
Reachability ruling (verified empirically): post-kill /history shows
the server synthesizes results for interrupted tool calls at recovery
('Cancelled by user. Outcome UNKNOWN'), so no persisted orphan exists
today and the poisoned state is unreachable — the client-side trace
was right, the server-side producer absent. Hardened regardless:
- liveToolCalls: an event-driven Set — fed ONLY by live tool_pending/
tool_info announces, retired by tool_result, drained at settle edges
and closeStreamTransport, and NEVER touched by any render (pinned:
refetchHistory's comment-stripped body may reference it exactly
once — the gate read). Liveness is read from the channel that
creates the hazard, never from DOM a render can forge.
- G6 coord-orphan-rewind: pins the SERVER invariant the client's
safety rests on — after a mid-bash node kill + reboot the batch must
render RESULTED (no --running residue) and the seedless rewind flow
must work end to end. Honestly scoped in its docstring: with
synthesis present a DOM-probe gate also passes, so the client
discipline is carried by the static pin set.
Quality batch: the seq stamp's producer position pinned (captured
before the await — the twin of the counter-bracket pin); two stale
G5 synthetic-idle comments corrected to the replay_ok-precise shape;
contract-test docstring item 7 restated to the enforced
universal-vs-seedless split; char-count pin windows replaced with
function-boundary slices (both test files); section 6 reuses _fn_slice.
Full coord family C + G1-G6 READY; 136 pins green.
|
||
|
|
fac2393967 |
fix(#894): re-derive the render gate on DOM-live signals; drop the busy conflation
Review round 5 step-back (fix-era critical): the r4 gate's busy term conflated 'a turn is executing' with 'this DOM holds live turn state'. _editAndResend flips busy BEFORE its POST and /rewind emits only clear_ui (no state_change), so the busy term skipped the truncation render the rewind exists to produce and appended the resent bubble onto the PRE-rewind transcript; /retry's regenerated turn likewise raced its own clear_ui refetch. The seam was re-derived once against the caller x state matrix; the gate reads DOM-live signals only, split by scope: - UNIVERSAL: dispatch seq (refetchSeq — overlapping fetches resolve last-DISPATCH-wins; an older snapshot landing late can neither double-render nor clear the latch over newer truth) and the content refs (skipping always beats stranding a ref; seeded callers null theirs before fetching, so it never blocks them). - SEEDLESS-ONLY (keyed on the seedCursor arg): the .conv-batch--running DOM marker for the tool phase (NOT activeBatch — that is the pending-APPROVAL tracker, set only for opts.pending batches; ruled at-site), coordSend's busySource === 'optimistic' flavor (the one busy that marks un-committed DOM), and stream-OPENness (CONNECTING keeps the handle with a frozen cursor and a pending replay; handle-existence was not liveness — also applied to the retry's fire guard). Seedless-only because the SEEDED resync renders over these deliberately: after a node dies mid-batch the --running class is dead residue no result will ever strip, and the resync's render IS the recovery — a universal term wedged the coord-restart scenario outright (family-run find; the r5 finders missed the seeded-path interaction). Backstop comment corrected (r5): fresh/truncated SSE replays DO carry a synthetic state_change (replay_ok does not) — a latched pane pays one refetch per reconnect, bounded by reconnect jitter/backoff and #884's server single-flight; heal-caused triggers remain structurally impossible. Caller fire-time ref guards demoted to the efficiency layer at-site. Contract test re-pins the gate: term presence in comment-stripped CODE, universal-vs-seedless placement, wipe between failure guard and latch-clear, no plain busy. The edit-resend commit gap under a second actor's clear_ui is accepted at-site (re-appears at settle heal). G4's honesty note names the tool-phase branch it behaviorally detects; G5 wording replay_ok-precise. Full coord family (C + G1-G5) green; 136 pins green; 5 static mutants + the G4 behavioral control caught. |
||
|
|
54631d3111 |
fix(#894): render-time gate at the refetch chokepoint; liveness-gate the retry; G4/G5 scenarios
Review round 4 (1 major + 3 minor bug, 1 major + 3 minor quality; bug-4≡q-2). Two correctness findings landed in one seam — the refetch-vs-live-state chokepoint — so the seam was redesigned once against its matrix (caller x stream-state-at-render x refs-at-render) instead of patched per-finding: - RENDER-TIME gate inside refetchHistory, post-await, pre-wipe: the await is a real window (queued sends drain at exactly the idle edges the backstop rides; another operator on a shared coordinator can send any time; hide/suspend can land mid-fetch), and only the chokepoint can see across it. Skip the wipe when a live turn exists (content refs — a wipe strands the bubble and loses the rest of the turn invisibly) or when a seedless render lost its idle/live-stream precondition (busy covers the tool phase the ref check can't see; a dead stream means rendering past the frozen cursor and double-rendering on the show-edge replay). Both requirements key on the seedCursor ARG — seeded callers own their reconnect flows and legitimately rebuild mid-turn. Skips leave the latch set; heals converge at the next organic settle. - The retry's fire guard gains evtSource (close-on-hide keeps the timer armed by design; a hidden firing must not fetch). The backstop needs no term — it runs inside SSE dispatch. - Pins: producer ORDER (inc < await < finally < dec), ref-guard pairs on both heal arms, the else-if exclusivity structure, the render-gate order and terms, the evtSource guard tail. Seven mutants, all caught. - Harness: __esOpens gate in G1-G3 (a pre-connect rewind drops its clear_ui into a channel nobody joined and false-fails the scenario); G4 coord-heal-midturn (a turn started under a held backstop fetch survives its resolution; hist==2 is the discriminating bit — noted honestly in the docstring); G5 coord-hidden-retry (hidden0 non-occurrence + organic-settle heal after show, per the accepted liveness-lag ruling — a quiet reconnect delivers no state_change edge). Negative controls: gate-stripped stamps hist1; guard-less stamps hidden1. Docstring gains the G-family catalog. |
||
|
|
dd1db5c67e |
test(#894): pin the in-flight counter's producer bracketing
Review round 3 (bug/security/perf zero; 1 quality minor): the contract test pinned both CONSUMERS of refetchesInFlight (the backstop and retry-fire yield guards) but not the PRODUCER ++/-- pair — dropping the bracketing would leave the counter at 0 and both consumer pins vacuously green. Count-pinned both sites; mutation-verified (the test fails with the increment stripped). |
||
|
|
ac441471f9 |
fix(#894): teardown-gate the retry ARM; pin both teardown sentinels
Review round 2 (1 minor bug + 1 minor quality, security/perf zero): - The retry's arm site was gated on historyStale alone, so a clear_ui refetch in flight at destroy() that then FAILS re-arms the timer AFTER destroy's clearTimeout — a no-op fire (the visHandler fire guard holds) but the orphan pins the dead closure for its 2s delay, contradicting destroy's dead-not-inert invariant. The arm gate is now historyStale && visHandler, matching the fire guard; the seam matrix (destroy/closeSession/live x arm-and-fire windows) closes with that one term. The edit-resend in the same .then stays deliberately ungated on teardown: the rewind committed server-side and the workstream outlives the pane UI, so the committed edit still delivers (comment at site). - Pins: the arm gate (mutation-verified — the contract test fails against a gate-stripped mutant), the fire guard's visHandler term (sole coordCloseSession protection), and the re-arm clearTimeout. G1/G2/G3 re-run READY; 136 static-pin tests green. |
||
|
|
85214f433f |
test(#894): pin the yield guards; narrow the clear_ui clearTimeout pin
Review round 1 (0 correctness/security/perf; 1 minor + 1 nit) + the suite's collateral: - The latch-contract test now pins !refetchesInFlight on BOTH heal paths (backstop arm + retry fire guard) — the yield guard is load-bearing (same-snapshot double-render stomp without it) and was previously deletable with every test green. Mutation-verified: the backstop pin fails against a guard-stripped coordinator.js. - test_app_js.py's clear_ui pin narrowed from all-clearTimeout to clearTimeout(truncatedResyncTimer): the invariant it protects is that clear_ui carries no path-local cancel of the TRUNCATED repair intent; #894's staleRetryTimer re-arm cancel is the staleness latch's own machinery, deliberately armed there. - _send_in_page's caller enumeration gains G3. |
||
|
|
f7ca4d295d |
fix(coordinator): historyStale latch closes the clear_ui over-rewind window (#894)
From clear_ui arrival until the next SUCCESSFUL refetchHistory render the visible transcript is the stale pre-rewind DOM with busy false, so a second rewind/edit click counted it and POSTed an over-large turn count against the already-restructured server conversation (the #890 sibling, pre-existing since #888 accepted the stale-interactive window). Port of interactive.js's converged #890 latch design, adapted to coord's structure (no load token, no replay quiesce, no ref-resetting render): - historyStale latch: set at clear_ui arrival, cleared ONLY by the success-path render below the if-(!hist) failure guard — a flag would reopen on the failed exit, which is exactly the over-rewind window. - Gates: _rewindToMessage / _editAndResend / _startEdit now require busy || historyStale; _rewindToTurns and _retryLast stay busy-only (explicit-arg / no-DOM-count — rulings at-site). - Heal A: one bounded turn-free retry armed in clear_ui's .then; fire guards read the latch, refetchesInFlight (net-new await-window counter, coord's quiesce-free yield discriminator — a COUNT because overlapping fetches are reachable), busy, the streaming refs (load-bearing: coord's refetch does not reset refs), and visHandler (teardown sentinel). - Heal B: idle-edge backstop as the else-if behind the truncated-resync consumer — TRANSPORT-FREE by ruling (plain seedless refetchHistory; a reconnecting heal draws the synthetic state_change:idle back into its own trigger = zero-backoff storm against a recovering node). Carries ref guards the interactive template omits: this arm also serves error edges where no stream_end nulled the refs. - Teardown: destroy() cancels the retry timer (terminal-only); closeStreamTransport deliberately does not (redials keep heal intent). Static pins: the latch contract (set/clear/gate sites, transport-free backstop, bounded arm, teardown split) + the widened guard-before-wipe window; the contract pin fails against the pre-latch code. |
||
|
|
14c246a569 |
fix(coord): port the truncated-recovery design from interactive (#882)
replay_truncated is now a dead-stream signal, mirroring the converged interactive.js machinery: - loadHistoryThenReconnect: tear the transport down first, drop the live cursor, refetch /history with cursor adoption, reconnect in .finally. The old in-place refetch discarded the /history cursor while /history trims the trailing in-flight turn whenever it returns one — a mid-run truncation wiped the executing turn with no redelivery and later tool results orphaned into top-level bubbles. Both consumption sites (immediate branch and idle-edge deferred consumer) route through it. Dropping the cursor before the fetch is load-bearing, not just parity: a post-restart heal on an idle ws gets no /history cursor, and re-presenting the frozen pre-restart cursor against the reseeded empty ring draws replay_truncated forever — an envelope→resync loop that parks the pane in degraded cooldown cycles (caught by the new browser-level scenario, invisible to source-pattern tests). - truncatedFromCursor: the truncation-time cursor, recorded keep-oldest at the envelope and cleared only by a successful full render; the connect chokepoint presents it over the live cursor so every manual reconnect re-draws the envelope and the repair survives any teardown interleaving (hide/show, degraded cooldown, CLOSED retry, failed fetch). - churn ladder: truncated resyncs feed the same rolling window as overflow closes via the extracted recordChurnAndMaybeTrip(); a trip skips the resync (the degraded wake re-arms via the chokepoint). - herd jitter: resyncs start behind a 0..TRUNCATED_RESYNC_JITTER_MS spread; one pending resync at a time; the fire path nulls its handle before loading; closeStreamTransport owns cancellation. - sidebar refresh: while a truncation gap is on record the gap machinery owns recovery outright — the envelope refreshes once per NEW gap, one heal-time refresh covers the retry window, and onopen's no-cursor / long-gap arm stands down — so a failed-resync retry loop cannot stampede /children + /tasks un-jittered once per reconnect through either path. - a failed /history refetch no longer blanks the pane (wipe + tracking resets sit below the !hist guard); a successful full render supersedes ALL pending repair intent in one place (gap record, deferred latch, pending resync timer) so a heal can never strand a phantom resync. Behavioral coverage: scripts/recovery_e2e.py gains --scenario coord-restart — the REAL coordinator pane (chrome, cookie auth, EventSource, connect chokepoint, resync, churn limiter) mounted against the interactive recovery node (/coord-static + /coord-recovery), driven through hide → node restart → show over CDP, asserting the envelope is drawn, the hidden-window turns heal, the stream re-opens, and the pane converges. Revised the two tests that pinned the in-place shape, added the coordinator mirror of interactive's fresh-connect/churn-limit pins (keep-oldest record, chokepoint consult, clear-on-render, shared churn step, trip-skip, jitter scheduler, cancellation site, cursor drop, per-gap sidebar dedup). |
||
|
|
1224b02d03 |
fix(compaction): review round 8 — seam obligations become primitives
Eight rounds of findings against the defer-and-drain seam shared one
generator: N sites each hand-copying M obligations (spawn discipline,
the order-barrier pair, backpressure, best-effort emission, the client
settle matrix), with every review finding an empty (site x obligation)
cell. This round makes each obligation a single primitive:
- The order barrier is Workstream.send_barrier_active() — one
definition of the two-term pair (pending entries OR drain alive),
consulted by the /send route, the coordinator adapter, and the
queued-nudge wake gate, which previously carried only the list term
and let a synthetic wake jump an acknowledged send during the
claimed-entry window. _PendingSend moved to workstream.py beside the
invariant that justifies the drain-alive term; the pending fields got
precise types and worker_kind became a Literal, so a typo'd
"command" comparison is now a type error instead of a silently
never-firing defer guard.
- _defer_send probes the barrier before constructing anything, bounds
acceptance at 10 pending (the interjection queue's own backpressure
contract — unbounded acceptance pinned message + attachment bytes
per entry for a whole command window and then ran one unattended
turn each), and spawns the drain with rollback: a Thread.start
failure pops the just-accepted entry and answers the retryable
queue_full instead of 500ing after registration (a phantom the
client could neither see nor retract, dispatched later as duplicate
turns). start() deliberately stays inside the lock, unlike
session_worker's outside-lock discipline: this slot is
is_alive()-gated, false for a constructed-but-unstarted thread, so
an outside-lock start would open a double-drain window.
- A /command whose worker never spawned answers 503
{"status": "error"} (spec + docs + a pane error arm) instead of the
generic 200 ok that told SDK callers their /clear ran.
- The compaction lifecycle emitter is raise-proof at its single
dispatch tail: a raising duck-typed hook degrades to a lost render,
never a lost end event — previously a raising on_error or a raising
failed-end emit left every pane a frozen progress bar, and a raising
SUCCESS end after the committed swap fabricated a failed end.
- The client settle matrix lives once: composer_queue's
settleSendResponse owns every /send response arm for both panes
(the near-verbatim twins were already drifting), parsePriority is
shared, and the busy stamp is centralized in setBusy(b, source) with
"server" as the fail-safe default. Deferred sends release the
composer (no worker exists for them; retracting the chip no longer
strands the pane in Stop mode), queue_full on an idle-looking pane
removes the optimistic bubble and restores busy (the refusal can now
fire with no worker and no drain to ever emit a state event), and
the pre-bind settle buffer is TTL-based — a burst of deferred
dispatches parked this tab's own raced settle first, where the old
size cap evicted exactly it.
- The command backstop / console proxy timeout inequality is enforced
by a test importing both named constants (both proxy_client
constructions, startup and the mTLS re-create); the compaction card
wears blue (magenta is reserved for the MCP surface); the redundant
TerminalUI.on_compaction override is gone (the inherited protocol
default is the policy site).
|
||
|
|
7f0e0406b3 |
test(approvals): concurrency matrix + suite migration to the cycle model
New regression matrix for the release blockers: cross-approval independence, lost-wakeup at gate entry, FIFO selector-less resolution, resolve-all sweep, double-resolution no-op, cards/legacy view tracking, and the generation-exactness set — stale delivery rejection, Smart-Approvals origin check, purge keep_origin, the purge-to-register window eviction, late cross-generation "superseded" stamping, concurrent smart+human gates, and the pre-delivered-verdict fast path. Plus sub-agent judge wiring (agent_gate off the main slot, close() firing all generations) and endpoint tests for cycle pinning and the Approve+Always race guard. Gate threads run under one shared mock-patch harness — mock.patch start/stop of the same target from concurrent threads corrupts the patcher's restore stack — with a sweep-until-dead teardown so the conftest leak guard can't trip. Existing suites migrate off the singleton fields to cycle assertions and the pending_approval_details wire shape. |
||
|
|
6424f73da4 |
feat(console): extend cross-user send gate to the coordinator surface
Coordinator workstreams will hit the same cross-user issue once MCP is enabled there, so wire the same protection now. - CoordinatorAdapter.send takes acting_user_id: binds it on a fresh turn (so MCP creds + the acting-user signal are correct) and passes it to queue_message on the interject path (CrossUserInterjectionError block). The create-time initial dispatch passes the creator's id. - ConsoleCoordinatorUI.on_state_change includes acting_user_id (mirrors WebUI); the mid-turn-connect replay already carries it via the shared make_events_handler. The coordinator's user-facing /send route already reuses make_send_handler, so it inherited the interjector guard + 409. - coordinator.js gains the same gate as the interactive pane: tracks the acting user from state_change, blocks send when busy AND acting != viewer, and handles the 409 cleanly. Drive-by hygiene (requested): the _verdictSig join used a raw U+001F byte embedded in the source; replaced with String.fromCharCode(0x1f) — identical runtime, ASCII-clean source (no invisible control char in the file). |
||
|
|
44c11efb53 |
feat(ui): split-view follow-ups — per-pane ✕, child-opens-beside, close-on-ws_closed
Four refinements from first live use: - Per-pane ✕ chip, top-right of every visible pane. Split mode: hide that cell (closeCell — the tab stays, the sibling absorbs the space). Single-pane: close the pane outright (withheld from the unclosable Dashboard). The click decides at click time; the label tracks the mode. Manager-injected into the pane section — content untouched. - Coordinator child links open BESIDE the coordinator (openPaneBeside: split right of the focused cell, seeded with the child pane) instead of replacing it — the parent stays on screen. Degrades to the plain focused-cell swap when the split is denied (cap / narrow viewport). splitFocused() gained an optional explicit-fill parameter for this. - Tier-1 ws_closed now CLOSES the open interactive pane (tab gone, a split cell collapses) — the coordinator-closes-its-child flow, matching the standalone's pane-auto-close. The dead-banner lane stays for streams that die without a ws_closed (node crash/network), where the session may still be revivable. - Paint bug: the focused-cell ring was an inset box-shadow on the section, which paints in the element's own background layer — UNDER opaque children touching the edges, so the status bar / composer strip occluded it. The ring now rides a click-transparent ::after overlay above pane content; the 2px top bar sits above the ring line. The livepass shell surface's demo panes grew a .ws-status-bar footer so the occlusion bug class stays visible to future passes. |
||
|
|
d9f5093a17 |
test(console): make dedupe-pin slice bounds reformat-tolerant
Review feedback: the next-case end markers were exact-indentation string finds that raised a bare ValueError when unmatched. Use whitespace-tolerant regexes with actionable assertion messages, and bound the history-replay window structurally (next role branch, with a generous fallback) instead of a fixed 600 chars. |
||
|
|
c988c9ed1f |
test(console): pin system-turn dedupe wiring on both read paths
The live-SSE/history system-turn dedupe (renderedSystemEventIds / _renderedSystemEventIds) was already in place on both panes and merged to main ( |
||
|
|
30cb9e6097 |
fix(ui): L-shell step 7 — wire coordinator approval keyboard shortcuts (designer P2)
The console twin of the interactive.js approval-key fix: the coordinator's tool-batch card shows kbd hints (Enter approve / D deny / Shift+A approve-all) but they did nothing — the keys were never wired (the standalone routed approval keys through the app.js global keydown + getFocusedPane, retired in the fork collapse). Add a pane-owned keydown on `root` that resolves the current pending batch: - _currentPendingBatch() finds the last .conv-batch with a still-pending [data-needs-approval="1"] row whose actions aren't already disabled — the in-flight double-fire guard (a second key during the resolve is a no-op). - Enter -> approve, D/Esc -> deny, Shift+A -> approve-all, routed to the existing _resolveBatchAction path. - A focus guard skips when an input/textarea/contenteditable is focused, so the keys never hijack composer typing (the coordinator has no feedback field, unlike interactive, so no feedback special-case). Verified end-to-end against the real coord pane (keydown -> _currentPendingBatch -> _resolveBatchAction -> approveWorkstream -> postJSON -> authFetch, stubbed at the HTTP boundary): Enter/D/Shift+A fire the right verb, the double-fire + focus guards hold, errs:[] + a coordinator JS guard. Live keypress confirm rides the merge gate. |
||
|
|
ff7d378235 |
feat(ui): L-shell step 5e.2c — interactive emitter onto shared .conv-* builders
Re-vocabularize interactive.js's approval card onto the shared conversation.js builders, converging it with the coordinator onto ONE neutral .conv-* card: buildToolDiv -> buildConvRow + buildConvCmd, renderVerdictBadge -> buildConvVerdict, _buildOutputWarningEl -> buildConvWarning; showInlineToolBlock / announceToolBlock build the .conv-batch shell + head + rows + buildConvActions (with the inline feedback + recommended glow); resolveApproval / the auto path -> buildConvStatus; updateVerdictBadge replaces the badge via buildConvVerdict; the history-replay branch synthesizes the live `item` shape so replay renders the SAME .conv-row. ~124 ts-approval-* / verdict-* references gone (a final no-stale-vocab assert guarded it); dead toggleVerdictDetail removed. The card chrome converges; interactive's richer post-execution result subsystem (.tool-output collapse/stream + .media-* embeds) is KEPT as-is — those are live-execution affordances the read-only coordinator history doesn't need. Block state classes move to the BEM modifiers (.conv-batch--approved/--denied/ --error/--auto); per-pane keybindings (y/n/a) preserved via the builder kbd hint. The now-dead old card CSS (.coord-tool-*/.ts-approval-*/.verdict-*) is inert (matches nothing) and is removed in 5e.2f's CSS dedup. Verified: node --check both emitters; the no-stale asserts; 60 JS guards green (test_coordinator_page pinned vocab updated coord-tool-batch-> conv-batch). The builders are behavior-tested (5e.2b). Designer + /review + live-backend run once at the merge gate. |
||
|
|
36c60e31e7 |
fix(ui): L-shell step 5e.1b — unify risk-level handling onto conversation.js
Both panes carried their own risk-level logic that disagreed on the fallback: interactive's normalizeRiskLevel sent an unknown level to "medium" (and, lacking the crit/med aliases, rendered a "crit" verdict as medium), while the coordinator's _riskRank sent unknown to "high". Lift one canonical normalize + rank into conversation.js and route both panes through it. - conversation.js: normalizeRiskLevel (aliases crit->critical / med->medium; unknown -> "medium"), riskRank, maxSeverityItem (keeps the no-verdict -> -1 edge so an unassessed item never wins the max-severity pick). - interactive.js: import normalizeRiskLevel, drop the local copy (its 3 callers unchanged); a "crit" verdict now renders critical instead of medium. - coordinator.js: import maxSeverityItem, drop RISK_SEVERITY / _riskRank / _maxSeverityItem; an unknown-level item now ranks medium, not high. - unknown -> medium is the deliberate fallback (per decision), not "high": medium is the neutral default both panes' displays already used. - tests: conversation guards for the fallback + aliases + the no-verdict edge; the two pane guards now check the shared module. The per-site risk->CSS-class display mappings (coordinator's inline chips and renderApprovalBlock's deliberate unknown->crit over-alert pill) are left for the 5e.2 vocabulary reconcile. |
||
|
|
9aad278983 |
refactor(ui): L-shell step 5e.1a — shared conversation.js for the dup'd helpers
Stand up shared_static/conversation.js as the deduplicated conversational-pane substrate both panes import (interactive via ./, the coordinator via /shared/ — both ES modules since 5e.0). First tenants are the byte-identical duplicates the in-file comments flagged for the step-5e lift: stripAnsi, the watch-result card builder, and the system-nudge marker. No visible change — the builders return the same DOM; each caller still appends + scrolls. - conversation.js: stripAnsi (null-safe variant), buildWatchResultCard, buildSystemNudgeMarker. - interactive.js / coordinator.js: import the three, drop their local copies, delegate appendWatchResult + the nudge marker through the shared builders. - stripAnsi unified on the coordinator's null-safe form (interactive's threw on a non-string arg); identical output for string inputs. - tests: new test_conversation_js.py pins the module; the two retry-walk guards now check the watch-result marker in conversation.js (it moved there). - refreshes interactive.js's header comment, stale since 5e.0 made the coordinator an ES module too. |
||
|
|
6614b4a549 |
refactor(ui): L-shell step 5e.0 — migrate coordinator.js to an ES module
Lift coordinator.js off the window.createCoordinatorPane bridge onto a real ESM export, so the upcoming shared conversational module (5e) is import-consumed on both sides rather than through a classic window global. The console shell and the standalone page's bootstrap both import the factory now; zero behaviour change. - coordinator.js: export the factory, drop the window bridge — no classic consumer remains (unlike interactive.js, whose ui/static app.js still uses its global). - shell.js: import the coordinator factory by URL (mirrors the interactive import) and call it directly. - console index.html: stop script-tagging coordinator.js; shell.js's import loads it (a classic tag chokes on the top-level export). - coordinator/index.html: the standalone bootstrap becomes a module that imports the factory — a classic eager IIFE ran before the deferred module loaded it. - tests: pin the new ESM seam (export, shell import, module bootstrap). |
||
|
|
da8ee0b018 |
feat(ui): L-shell step 5c — coordinator child-links open interactive panes
Rewire the coordinator's child ws links (deferred from step 4) to open the child
as a node-proxied interactive pane inside the console L-shell, instead of a
full-page new tab to /node/{id}/?ws_id=.
- A delegated click handler on the pane root catches .ws-link (children tree,
renderChildRow) and .coord-ws-link (linkified tool output, renderToolOutput)
clicks; both link types now carry data-ws-id + data-node-id. When a
PaneManager is present it opens openPane('interactive', child_ws_id,
{nodeId: child_node_id}) — the child's OWN node, so its stream proxies to that
node even though the coordinator lives in the console.
- Progressive enhancement: the link's href (/node/{node}/?ws_id=) stays the
standalone fallback — the standalone coordinator page has no PaneManager, so
the new-tab nav stands. No innerHTML introduced (the tool-output linkifier
still returns a string; only data-* attrs were added).
Verified: a harness running the console stack (coordinator.js + shell.js +
interactive.js) — open a coordinator pane, click a child link -> a node-proxied
interactive pane opens (/node/{child_node}/.../events), zero errors; 106
JS-guard tests (test_coordinator_page step-5c guard), node --check, prettier.
|
||
|
|
6e0901ee1b |
fix(ui): L-shell step 4 — designer-review fixes (pane-aware close, overflow, glyph, a11y)
Four findings from the step-4 designer pass on the coordinator pane.
- P1 (bug): the `end` button ran `window.location.href = "/"`, which inside the
L-shell reloaded the WHOLE console — every other pane destroyed, all their
Tier-2 streams dropped. Thread an `onClose` through the factory; the console
pane passes `() => pm.close(pane.id)` so `end` closes that tab (and runs the
controller teardown via onClose→destroy); the standalone page passes none and
keeps the console redirect.
- P2: the pane root carries both `.pane-body` (overflow:auto) and
`.coord-chrome-root` (flex column), so the generic pane scroller redundantly
wrapped the sticky appbar. `.pane-body.coord-chrome-root { overflow: hidden }`
(scoped to this pane type) — the coord chrome owns its own scroll regions.
- P3: the coordinator tab glyph `●` collided with the rail's running state-dot
vocabulary (a static dot reading as "live"); swap to `◆` (a shape marker that
pairs with dashboard's `◇`), pending the real state-glyph in step 7.
- P3: the destructive `end` button had only a title; add aria-label
"End coordinator session".
Verified via the harness (clicking `end` closes the pane without reloading;
glyph `◆`; aria-label present; zero errors); guards pin the pane-aware close.
test_shell_js + test_coordinator_page (97 green), ruff.
|
||
|
|
05dbcf3818 |
feat(ui): L-shell step 4b — coordinator sessions as console panes
The console can now host coordinator sessions as ws_id-keyed panes alongside
dashboard/admin — step 4 complete (the de-globalization landed in 4a).
- coordinator.js: `buildCoordChrome(root, opts)` builds the coordinator chrome
programmatically (createElement, no innerHTML); the factory builds it on
instantiate, so the SAME factory serves the standalone page and a console pane.
`opts.standalone` adds the page-level bits a pane doesn't want (the Console
back-link, the theme toggle, the shared #toast).
- index.html (standalone): goes thin — a bootstrap calling
createCoordinatorPane(document.body, ws_id, {standalone:true}); the ~500-line
inline <style> is migrated to coord-chrome.css (its lone page-level body rule
scoped to .coord-chrome-root) so the console can load the same chrome CSS.
- shell.js: registerType('coordinator') keyed by ws_id — onMount builds the
controller into the pane body, onActivate opens its Tier-2 SSE once, onClose
destroys it. Plus a window.TS_LOGIN fan-out registry so every pane re-arms its
own stream on re-auth (app.js's single onLoginSuccess becomes one subscriber).
- rail.js: coordinator clicks → openPane('coordinator', ws_id) instead of
full-page nav (interactive sessions stay interim full-page until step 5).
- console/index.html: loads the coordinator controller + chrome CSS + the shared
composer/renderer deps it needs.
Child links → openPane('interactive', ws_id) are deferred to step 5 (the
interactive pane doesn't exist yet); coordinator transport stays console-local
inline (parameterized only when the shared ConversationalPane base is lifted).
Verified end-to-end with a headless harness running the real shell.js + rail.js +
coordinator.js: opening a coordinator pane registers the type, the rail row opens
it, buildCoordChrome populates the pane, the Tier-2 SSE connects, destroy() tears
down — zero uncaught errors; renders cleanly (appbar + chat + children/tasks
sidebar + status bar). test_shell_js + test_coordinator_page (97 green), ruff.
|
||
|
|
0231253c68 |
refactor(ui): L-shell step 4a — de-globalize coordinator.js into a pane factory
coordinator.js was a page-global IIFE keyed off <html data-ws-id>. Make it multi-instantiable so the console shell can host coordinator sessions as panes (one per ws_id) alongside dashboard/admin — the first conversational pane-content. - IIFE -> `createCoordinatorPane(root, wsId)`: the ~40 module-state vars stay closure-local (now automatically per-instance), every #coord-* lookup is root-scoped (27 getElementById -> root.querySelector), ws_id is a constructor arg. - New lifecycle: `connect` (= init), `destroy` (closes the EventSource + clears the 6 timers + the prune interval + the IntersectionObserver — the IIFE had no teardown, so a backgrounded pane would leak an SSE and fire into detached DOM), `onLogin` (re-arm after a 401), `closeSession`. - Drop the page-global collision points: `window.coordSend`/`coordCloseSession` -> local fns (the close button binds per-instance; its inline onclick is removed); `window.onLoginSuccess` -> the returned `onLogin` (the console shell will fan login out to every pane; standalone keeps the single hook). - Standalone coordinator page = one pane filling the body: a thin bootstrap calls `createCoordinatorPane(document.body, ws_id).connect()`. Console-local transport (the coordinator endpoints) stays inline — coordinators always live in the console; transport is parameterized only when the shared ConversationalPane base is lifted (after step 5). The chrome builder, CSS migration, and console pane registration are 4b. Verified: node --check; a headless smoke (the real factory instantiates against a provided root, runs connect()'s snapshot/history/children/tasks/SSE on stubs, then destroy()s — zero uncaught errors); test_coordinator_page.py (13, incl. a new factory-shape guard); ruff. |
||
|
|
09e41d1502 |
fix(coord): extend the system_turn /history+replay dedup to the coordinator pane
The interactive pane (app.js) got the event-id dedup that skips an operator-context system turn already painted from /history when an SSE replay redelivers it; the coordinator pane (coordinator.js) shares the identical /history + live system_turn + last_event_id replay seam but was left without the guard. The backend row/event-id alignment already fixes the actual double for both panes — this restores the defense-in-depth symmetry. - Module-scoped renderedSystemEventIds (persists across reconnects like lastEventId), reset in refetchHistory. - onmessage tags each event with its SSE id; the live system_turn handler skips an already-rendered id; the history loop records the ids it paints. - Parallel static-shape regression test in test_coordinator_page.py. |
||
|
|
107967f548 |
fix: pre-push review — operator-context retry walk, get_content offload, doc cleanups
- A1: add a shared `operator-context` marker to every operator-context row in
both UIs; the retry-skip walk keys on it, so a trailing watch-result /
guard-finding / idle-children card no longer makes retry regenerate the
wrong turn. Pinned by source-grep tests + a headless-DOM self-test.
- A2: wrap get_content's two sync DB gates (get_attachment + the unbounded
ws-scoped attachment_referenced_in_ws LIKE scan) in asyncio.to_thread so a
long scan can't stall the event loop — matching the module's convention.
- A3: add the HistoryEvent attachments docstring bullet; drop the dead
`interjection` class; document the interactive-only system-context label;
correct the SDK attachments-meta docs to {kind, filename, mime_type}
(size_bytes is not carried through the history projection).
|
||
|
|
c6b2288302 |
feat(session): consolidate operator-context into first-class system turns
Replace the two operator-context hacks (the <tool_output>/<system-reminder> content envelope and the transient _reminders side-channel) with one persistent {role: system, _source} trajectory turn. Adds supports_mid_conversation_system (claude-opus-4-8): native models take the turn inline; all others fold it into the preceding turn as a nonce-delimited <system-reminder> block declared in the system prompt as the sole trusted marker. Producers (advisories, metacog nudges, user interjections, idle/watch) emit system turns; the envelope/_reminders machinery, escaping round-trip, replay parser, and reminder SSE events are removed. Eager 060 migration un-wraps legacy envelopes. Net -1662 lines.
Known follow-ups from review (unfixed here): (1) the 060 un-wrap heuristic can irreversibly mis-rewrite bare tool rows that resemble the envelope, so do not run the migration until it is tightened; (2) user_interjection turns lost the user-framing/priority preamble (a regression, and a native-path authority-framing concern); (3) native-path wake nudge can emit empty user content.
|
||
|
|
44f19f1401 |
feat(judge,console,ui): paint pending tool calls before the intent verdict
The intent-validation judge runs before the approval gate resolves, and Smart Approvals (judge.smart_approvals) parks approve_tools on the async LLM verdict for up to judge.timeout — so the tool-call card never reached the UI until the judge had ruled. An operator could not see a committed call, let alone Stop it, during that window. approve_tools now emits a tool_pending event carrying the serialized batch at the top of the gate, before the tool-policy lookup, the verdict wait, and the human prompt. It is a UI paint only — no persistence, audit, or verdict bookkeeping — so it cannot perturb the gate's accounting. The authoritative tool_info / approve_request / tool_result events that follow upgrade the same construct in place, keyed by call_id, and the Last-Event-ID replay slice reconstructs it on reconnect. A ToolPendingEvent joins the SDK registry. Coordinator: appendToolBatch was already idempotent on call_ids; the new handler reuses the --running placeholder it already upgrades, with an "Evaluating" kicker that swaps to "Running" on the auto-approve upgrade. Interactive: showInlineToolBlock was create-only, so a second card would duplicate. Added announceToolBlock + _takeAnnouncedBlock to reuse the announced shell (matched on its call_id set) instead. The announced rail is dashed amber and must out-specify the .msg.ts-approval--inline cyan-hold (specificity 0,2,0) — at 0,1,0 it rendered cyan, indistinguishable from a normal card — so the announced card is the one visually distinct surface in the stream. Screen-reader parity: the early paint announces politely through dedicated off-screen aria-live regions on both surfaces (the messages log is aria-live=off mid-stream, so the appended shell alone is inaudible), and the announced shell carries aria-busy until the upgrade clears it. Polite, not assertive — the human gate keeps its assertive announcement. Tests cover the gate ordering (tool_pending precedes tool_info and the Smart Approvals gate) plus string-guards on both UIs' wiring, the announced-rail specificity, and the screen-reader regions. |
||
|
|
a5570f027b |
fix(sse): consume the resume cursor in the coordinator dashboard
make_history_handler is shared by interactive and coord, so coord /history already trims the executing in-flight orphan turn and returns a cursor. But coordinator.js never read it -- it connected fresh, so the trimmed turn was neither in /history nor delta-replayed and vanished from the dashboard (a regression vs the prior #610 in-flight render). Mirror the ui/static/app.js fix in coordinator.js: refetchHistory takes a seedCursor flag (default false) and seeds lastEventId from hist.cursor only on the initial-connect path; connectSSE gates ?last_event_id= on != null so a cursor of 0 isn't dropped. The clear_ui / replay_truncated re-render callers leave seedCursor false (they run on a live stream and must not rewind the live cursor). Adds a coordinator.js static guard. |
||
|
|
56f230afa2 |
refactor(coordinator): consume the projected /history shape
Migrate the coordinator dashboard's init() history rebuild onto the server-projected wire shape (the prior commit's project_history_messages), retiring its inline raw-storage-shape handling. Field-sourcing only -- the batch render + orphan->--running state machine is unchanged: - callOutcomes reads the server-derived `denied` / `is_error` flags instead of re-sniffing tool content prefixes; - tool_calls are read flat (`tc.name` / `tc.arguments`) now that the projection flattens the nested `function` wrapper; - user content is a plain string and attachments come from the projected `attachments` list, replacing the multipart walk + `_attachments_meta` side-channel read. Fix two latent reload bugs along the way: coord read `m.reminders` / `m.source` but the raw shape carried `_reminders` / `_source`, so metacog reminder bubbles and the system-nudge marker never rendered on a coord history reload. The projection surfaces both top-level, so coord's existing (unchanged) render paths now fire. Repoint the test_coordinator_page deny-classifier guard at the live `m.denied` / `m.is_error` reads instead of comment prose. Refs #549. |
||
|
|
eca4bb79e4 |
fix(replay): seam 1 splice + storage symmetry for queued user messages
Reverses the seam-2-only design from the prior commits on this branch.
Queued user messages arriving DURING a tool batch (Seam 1) splice into
the last tool result's envelope as ``UserInterjection`` advisories via
``wrap_tool_result``. Messages arriving BETWEEN turns (Seam 2) drain
as a single trailing user row via ``_flush_queued_messages`` with
``user_feedback`` (operator text alongside an approval, e.g. "y, use
full path") folded in as a prefix. Cancel/exception drains (Seam 3)
keep the existing ``_flush_queued_messages()`` call unchanged.
Why all three seams:
* Strict-template providers (Mistral, Llama via vLLM with stock chat
templates) reject role-alternation violations. A literal ``user``
row mid-tool-batch breaks ``assistant(tool_calls) → tool → ... →
assistant``; back-to-back ``user → user`` rows on the wire also fail.
* The seam-2-only design produced back-to-back ``user`` whenever
``user_feedback`` and queued items both fired — bug-1 from the round-1
review. Folding ``user_feedback`` as a prefix to the queue-drain
collapses the two into one row.
* During-batch arrivals couldn't ride seam 2 — the splice was the only
way to deliver same-turn without violating role alternation.
Storage symmetry:
Tool DB rows now store the wrapped ``output`` (envelope + advisories)
unconditionally — ``self.messages[i]['content']`` and
``conversations.content`` match exactly. List-typed output (image /
structured MCP results) uses ``wrap_tool_result(raw_joined_text,
advisories)`` at save time so the persisted string is anchored on
``<tool_output>\n`` for the replay parser. ``TOOL_RESULT_STORAGE_CAP``
is removed entirely; tools are responsible for bounding their own
output, storage faithfully represents in-memory. Removing the cap
also simplifies the parser — no truncated-envelope edge case.
Replay extraction:
``decorate_history_messages`` (REST ``/history``) and ``_build_history``
(SSE replay, resume, rewind, retry, post-load, rename re-replay) both
call the public ``extract_advisories_from_tool_envelope`` helper to
pull the envelope back into structured ``advisories`` for JS replay.
Both string content and list-typed content (image+queued-message
combo) covered. JS renders extracted advisories as normal user
bubbles after the tool block via the shared ``replayAdvisoriesAfterTool``
helper in ``shared_static/utils.js``.
Wrapper-tag escape and provider splice:
``escape_wrapper_tags`` now encodes pre-existing ``&`` first using an
``&`` sentinel so tool output containing literal entity strings
(documentation viewers, code analyzers, web scrapers returning entity-
encoded markup) round-trips correctly. Both encode and decode helpers
short-circuit on absence of ``<`` / ``&``.
``_apply_reminders_for_provider`` detects already-wrapped content
(string body and list text-part) by ``startswith("<tool_output>\n")``
and skips re-escape so existing envelopes survive intact when a tool
message also carries ``_reminders`` (the queued-message + tool-error
co-occurrence case is now common).
``decorate_history_messages`` runs in ``asyncio.to_thread`` to keep
MB-scale string work off the event loop.
Other cleanup:
* ``_collect_advisories`` delegates the queue drain to a named helper
``_drain_queued_messages_to_advisories`` so the swap-and-clear pattern
lives next to ``_flush_queued_messages``'s identical pattern and the
side-effect is documented at the call site.
* Preamble strings + body marker for ``UserInterjection`` round-trip
detection moved to module-level constants in ``tool_advisory.py``;
imported by ``history_decoration.py`` so a producer-side rephrase
can't silently desync the parser.
* ``_send_with_mocks`` ctxmgr extracted in ``test_session.py`` — the
six new send-driven tests share an 8-deep ``patch.object`` block.
* ``replayAdvisoriesAfterTool`` shared helper in
``shared_static/utils.js``; ``app.js`` and ``coordinator.js`` both
invoke it.
* Dead truncation-pill CSS removed (``.tool-output-truncated`` and
``.coord-tool-truncated``); the JS that added these elements went
away with ``TOOL_RESULT_STORAGE_CAP``.
* Tautological tests (``TestBuildHistoryAdvisoryPropagation``)
replaced with production-realistic round-trip tests built from
``wrap_tool_result(...)`` envelopes — REST and SSE-replay surfaces
pinned to the same wire shape; full DB round-trip pinned end-to-end.
Negative-tested:
* Reverting the prefix-merge in ``_flush_queued_messages`` produces
back-to-back ``user`` rows, breaking
``test_user_feedback_and_queued_coexistence_single_row_with_prefix``.
* Reverting the ``extract_advisories_from_tool_envelope`` call in
``_build_history``'s tool branch leaves the envelope verbatim in
wire content, breaking the round-trip tests.
* Reverting the wrapper-detection in ``_apply_reminders_for_provider``
entity-encodes the existing envelope's literal tags, breaking both
the string-content and list-content envelope-preservation tests.
* Reverting the ``wrap_tool_result(raw_text, advisories)`` projection
at the DB save site produces a string starting with the original
raw text, breaking
``test_tool_db_row_round_trips_list_output_with_advisories``.
Tests: 5918 passed, 3 deselected. Lint + format + mypy clean on
touched files.
|
||
|
|
a032e71ff3 |
fix(replay): apply review findings q-2 through q-7
Round-1 ``/review`` apply-pass. Drops stale ``UserInterjection`` references from comments and docstrings that no longer describe the post-PR drain shape, asserts the two-stream invariant in the new queued-message persistence test, and pins the ``content.trim()`` + ``renderAssistantToolBatch`` invariants on coord-side so a future refactor can't silently regress the Qwen3 phantom-card fix or the chronological-order render fix. Deferred: * **bug-1** (back-to-back ``user`` row when ``user_feedback`` from the approval-prompt UI callback coexists with a queued-message drain). Reachable on strict OpenAI-compatible local templates (Anthropic and Anthropic-via-merge-consecutive collapse fine; vLLM-hosted Mistral / Llama enforcing role alternation can reject). The pre-PR splice guarded against this case by riding queued items inside the tool result envelope; that guard is what motivated the original UserInterjection design, so the fix lane needs a deliberate decision rather than a quick patch. Sleeping on it. * **q-1** (delete dead ``UserInterjection`` class + tests). Held for the bug-1 decision — if the chosen fix is to resume the splice for the ``user_feedback``+queue coexistence case, the advisory shape stays load-bearing. Class now carries a docstring note marking it retained-pending-decision so a passing reader doesn't grep for producers and assume it's actually dead. Apply-pass content: * ``q-2``: drop "queued user interjections" from the persistent- advisory parenthetical in ``send``'s tool-result loop comment; rewrite to point at ``_flush_queued_messages`` for the queue path. * ``q-3``: ``__init__`` channel-routing comment loses "and ``UserInterjection``" — only ``GuardAdvisory`` remains. * ``q-4``: ``_queue_tool_advisory`` docstring + the tool-error nudge comment lose the user-interjection mentions; the docstring also now describes the side-channel + ``_apply_reminders_for_provider`` splice path (the actual mechanism). * ``q-5``: ``AttachmentsNotQueueableError`` docstring rewritten to describe the post-PR ``_flush_queued_messages`` flow — the single-combined-turn ``\n\n``-join shape can't carry image / file blocks, and per-item separate user turns would expand the strict- template role-ordering surface that the post-batch drain already balances. * ``q-6``: the new ``test_queued_message_persists_as_user_row_after_tool_batch`` in ``test_session.py`` now asserts ``stream_idx == 2`` so a future regression where the post-batch flush runs but the send-loop short- circuits before the next iteration surfaces in CI rather than manual repro. * ``q-7``: ``test_coordinator_page.py`` gets two new string-grep pins mirroring the existing ``test_app_js.py`` shape — ``content.trim()`` on coord's assistant-replay branch and ``renderAssistantToolBatch`` for the hoisted helper that orders content card before tool batch. ## Test plan - [x] ``ruff check`` clean - [x] ``mypy turnstone/`` clean (189 source files) - [x] Affected test surface (``test_session.py`` + ``test_tool_advisory.py`` + ``test_app_js.py`` + ``test_coordinator_page.py``) — 240 passed |
||
|
|
ae4fddfc5a |
fix(console): home composer attachments + coord chat user-message pills (#462)
* fix(console): home composer attachments + coord chat user-message pills Two parity gaps in the console's coordinator surface: - The embedded creator on the home page accepted only text — the paperclip / paste / drop pipeline that the in-coord composer and the interactive new-ws modal both expose was missing, so a user couldn't attach files at create time. Stage Files in memory (no ws_id yet) and ship them multipart on Start; the coord create endpoint already accepts multipart via create_supports_attachments=True. - User messages with attachments rendered as plain text on both live send and history replay — no chip cluster like the interactive pane. Added appendUserMessageWithAttachments and a structured userAttachments list built from _attachments_meta (preferred) or the multipart parts themselves, then rendered the same .msg-user-attach pill strip the interactive pane uses. Polish from a designer pass: - Pill background was --panel-2, equal to the .msg bubble background in both themes (border contrast ≈1.4:1, below WCAG 1.4.11). Switched to --panel so the pill sits on a different surface than the bubble. - Capped chip filename width inside the home composer (max-width 200px + ellipsis) so a long filename doesn't push the strip past the textarea. - aria-live="assertive" → "polite" on #home-coord-error; client-side validation isn't an interrupt-level event. - Reserved min-height on .home-composer-error and dropped the display: none/block toggling so validation messages no longer reflow the active-coordinators list below. * fix(console): address PR #462 review feedback - Block home-composer submit when files are staged but the task field is empty. Server's _coord_create_post_install short-circuits on an empty initial_message, so the multipart upload would create pending attachment rows that never reserve onto a turn — orphaned until the GC sweep. Fail in the browser instead. - Drop the redundant `part &&` guard in coordinator.js's history-replay multipart loop; the earlier `if (!part || ...) continue` already filtered. - Rewrite the home-mount .composer-chip-name CSS comment. shared/chat.css defines .composer-chip{,-size,-remove} but no .composer-chip-name rule — the span inherits the parent chip font with no width cap. - Add smoke-guard string assertions in test_coordinator_page.py for appendUserMessageWithAttachments and msg-user-attach so a future rename can't silently regress the attachment affordance. |
||
|
|
38a0d9c3b6 |
feat(coord): Stage 3 SessionManager Children primitive lift + cluster bus push paths
Lift the Children primitive out of CoordinatorAdapter into universal SessionManager core primitives, replace the fragile poll + state-event piggyback paths with first-class cluster bus event types for inline approval delivery, and clean up the resulting frontend reducer. Architecture - New `turnstone/core/children_registry.py` — universal parent → children + reverse-lookup primitive with atomic `add_child` (returns parent UI for race-free dispatch). Lifted from `CoordinatorAdapter`. - New `turnstone/core/child_source.py` — `ChildSource` Protocol with `SameNodeChildSource` (in-process via SessionManager state observer) and `ClusterChildSource` (cross-node via ClusterCollector listener). - `SessionManager._on_state_change` upgraded to multi-subscriber (`subscribe_to_state` / `unsubscribe_from_state`) under a dedicated lock; CLI consumer migrated. - `CoordinatorAdapter` shrunk: 731 → ~640 LOC. Children data lives in the registry; fan-out lives in ClusterChildSource. Backward-compat property facades dropped; tests updated to use the registry surface. Cluster bus event vocabulary - New event types `intent_verdict`, `approval_resolved`, `approve_request` flow through both `ClusterCollector._apply_delta` (translation from node SSE) and `emit_console_ws_*` (synthesis on console pseudo-node). - `CoordinatorAdapter._dispatch_child_event` re-emits as `child_ws_intent_verdict` / `child_ws_approval_resolved` / `child_ws_approve_request` on the parent coord's SSE stream. - New `_broadcast_intent_verdict` / `_broadcast_approval_resolved` / `_broadcast_approve_request` no-op hooks on `SessionUIBase`. WebUI pushes to the global queue; ConsoleCoordinatorUI pushes to the collector. `approve_tools` calls `_broadcast_approve_request` right after setting `_pending_approval` so the items reach the coord tree immediately, eliminating the bulk-fetch race. Cleanups - `pending_approval_detail` piggyback on `ws_state` / `cluster_state` removed end-to-end. Bulk fetch + explicit verdict / approve-request push are the canonical carriers. - Browser `_judgePollTick` 90-second poll loop deleted; push path is authoritative. - `urgent` flag on `scheduleLiveFetch` deleted (only caller was 409 retry; replaced with `invalidateLiveBadge` + standard schedule). - Console `_fetch_live_block` derives `pending_approval` from a disjunction (`activity_state="approval"` OR `state="attention"` OR detail present) so the bulk fetch can't return false during the state-transition race window. - Coord-side merge guard in `flushLiveFetches` no longer clobbered: `handleChildState` only stamps `sseUpdatedAt` when authoritatively clearing detail. - `child_locality` capability flag removed (was inert dead code). Reliability - Selective drop on listener queue overflow: critical event types (verdicts, approvals, ws_closed, child_ws_*) evict one oldest item to make room rather than dropping themselves on a full queue. Best-effort events (state ticks, content tokens, status, activity) drop as before. Applied to `SessionUIBase._enqueue`, `ClusterCollector._fanout`, and the `WebUI._global_queue` puts in the new broadcast hooks. - `_state_subscribers` snapshot under a dedicated lock so concurrent subscribe / unsubscribe during dispatch can't shift the iterator. UX / a11y - Loading placeholder in renderChildRow keeps row height stable while the bulk fetch is in-flight (sr-friendly aria-label). - Focus preservation across `_renderChildrenNow` (capture + restore by row + marker) and across targeted `_updateChildRow` swaps. - Layout-shift transition on the approval block max-height; respects `prefers-reduced-motion`. - Sidebar pending count: `(N children · M pending)`. - Risk pill `aria-label` spells out level + confidence for SR users. - Per-coord SSE listener queue depth surfaced in the status bar (`queue N/500`) with color escalation (warn at >50%, danger at >80%). Tests - 305+ test changes across 8 files. New unit tests for `ChildrenRegistry`, `ChildSource` (both impls + multi-subscriber observer), the new collector emit + apply_delta cases, the dispatch cases for new event types, the broadcast hook overrides on both WebUI and ConsoleCoordinatorUI, and the focus / placeholder / pending-count frontend assertions in `test_coordinator_page.py`. 5024 passed, ruff + mypy clean. |
||
|
|
88facd260e |
feat(coord): pass pending_approval_detail on child_ws_state SSE events
Inline child approve/deny in the coord tree UI was rendering downstream
of the bulk-live cache (``GET /v1/api/cluster/ws/live``), not the SSE
stream. ``child_ws_state`` events were tiny notifications that fired
an urgent live-bulk fetch on every activity_state transition into/out
of "approval", just to pick up the rich ``pending_approval_detail``
payload. With multiple coord tabs and multi-child workstreams, that
urgent-fetch pattern compounded the SSE-executor pressure Shape A
is unwinding.
Thread the field through every layer so the SSE event itself carries
the rich payload — browser mutates ``liveBadgeCache`` directly,
no urgent fetch:
1. Node ``WebUI._broadcast_state`` emits ``pending_approval_detail``
on ``ws_state`` events. Gated on ``_pending_approval is not None``
so the per-broadcast verdict-cache deepcopy only runs when there
is actually an approval pending. ``_build_node_snapshot`` also
projects the field so the console's reconnect-via-snapshot
resync path delivers it (without this the new collector
forwarding would never see the field on a snapshot row).
2. Console ``ClusterCollector._apply_delta`` (live ``ws_state``
forwarding) and ``_reconcile_node`` (snapshot resync diff) both
forward the field on the emitted ``cluster_state`` event, AND
``_apply_delta`` persists it on the cached ``ws`` dict so the
``get_node_detail`` / ``get_snapshot`` endpoints between
reconciliations don't render stale approve/deny buttons.
3. ``CoordinatorAdapter._dispatch_child_event`` re-emits the field
on the ``child_ws_state`` event sent to coord listener queues.
4. Frontend ``handleChildState`` reads ``ev.pending_approval_detail``
and writes it directly into ``liveBadgeCache``, tagging the
entry with ``sseUpdatedAt``. ``flushLiveFetches`` honors that
tag for ``SSE_AUTHORITATIVE_MS`` (3s) — the upstream
``/dashboard`` cache has its own ~2s TTL, so a bulk-poll
landing right after a transition can otherwise clobber the
fresh SSE-set state with pre-transition data.
The pre-fix ``enteredApproval`` / ``leftApproval`` urgent-fetch
branch is removed. The 409 stale-call_id retry path keeps its own
urgent fetch — that's a different scenario.
Tests cover the forwarding contract at every layer, the broadcast
gate (event includes the field when an approval is pending,
omits it otherwise, and clears after resolution), and the
``flushLiveFetches`` merge-guard structural shape so a refactor
that keeps the symbols but inverts the comparison or drops the
``prev.live`` check can't pass silently.
|
||
|
|
353ff4d18b |
feat(coord): inline tool-batch construct replaces approval dock (#447)
* feat(coord): inline tool-batch construct replaces approval dock
The pinned bottom approval-dock didn't scale: a 10-call spawn_workstream
fan-out filled the whole pane with a wall of repeated verdict chips,
and the call → approval → result lifecycle was split across three
disconnected surfaces (.msg.tool bubble + dock + .msg.tool result).
Replaces it with one chat-stream construct per dispatch turn that
pairs each tool call with its result and embeds the approval gate:
- .coord-tool-batch--solo single-call serial turn
- .coord-tool-batch--parallel ≥2 calls; rows share a left rail
+ per-row tick so they read as
siblings of one assistant decision
Lifecycle: rows render with optional "judge evaluating…" placeholder,
upgrade in place when intent_verdict arrives, and on tool_result the
output lands paired under the originating row. When the batch needs
approval, one Approve/Deny/Always action row renders inside the
construct (envelope-level — server semantics resolve siblings
together). After approval_resolved the action row morphs into a
✓ approved / ✗ denied status pill that stays as a receipt.
Critical bug closed: when a page reload races a pending approval,
pre-scan tool_call_ids in history; turns whose call_ids have no
matching tool result are rendered pending (not resolved-approved).
The SSE approve_request replay then upgrades the existing batch
in place — drops --approved/--denied, adds --pending, swaps the
status pill for actions, and assigns activeBatch. Without this
the operator was locked out of any approval pending at reload.
Defence-in-depth follow-ups from the same review:
- approval_resolved falls back to a DOM lookup if activeBatch
is null (cross-tab resolution where this tab never set it).
- _appendVerdictLineTo dedupes via a row.dataset.verdictSig so
SSE reconnect storms + repeat intent_verdict events don't
tear down + rebuild an unchanged verdict line.
- judgeVerdicts Map soft-capped at 500 entries (FIFO eviction)
via _cacheJudgeVerdict.
- toolRows entries hold {batch, row} only — the originating
item payload is no longer pinned for the page lifetime.
- _scheduleScroll coalesces messagesEl.scrollTop writes through
requestAnimationFrame so history replay doesn't reflow once
per appended message.
- Rationale <details> now inserts immediately after the verdict
line (was tail-appending, breaking ordering once a result
landed below).
- .coord-tool-batch--error wired: _appendResultToRow lifts a
row's error onto the enclosing batch; _renderBatchRow does
the same for policy-blocked rows at construction.
- _buildStatusPill extracted; both _morphBatchResolved and the
appendToolBatch resolved-replay branch route through it.
Removed: ~248 lines of dead .approval-dock CSS, the dock <aside>
element from index.html, and the dead helpers showApproval's
prior body, hideApproval, claimApprovalFocus,
claimApprovalFocusForVerdict, applyJudgeVerdictToRow,
applyJudgePendingToRow, ensureDctxAfterRow, removeRationale,
setApprovalButtonsDisabled, the appendToolCall single-row wrapper,
and window.coordApprove. Five stale comment blocks referencing
the dock as if live also swept.
Children-tree's renderApprovalBlock is independent and untouched
(different surface, different .approval-block / .approval-pill
vocabulary).
* fix(coord): close four Copilot review gaps on PR 447
Copilot review on
|
||
|
|
1f271789b3 |
fix(approve): global judge poll + Copilot round-2 feedback
Bug: LLM judge verdicts stayed stuck on heuristic-only render.
Root cause: per-row poller called scheduleLiveFetch which
short-circuits on non-visible rows — invalidate cleared the
cache, no fetch fired, the row kept rendering its last-cached
heuristic indefinitely. The 12s attempt cap also gave up before
slow LLM judges (>15s with reasoning effort) could land.
Replaced with a single global poller _maybeStartJudgePoll /
_judgePollTick:
- Walks the full childrenState (not just visible rows)
- Bypasses scheduleLiveFetch's visibility + TTL gates by
adding to pendingLiveIds directly + flushing
- One bulk request covers every pending row per tick
- Self-terminates when every verdict lands or 90s elapses
(operator can hit Refresh to retry on a failed judge)
- 90s cap is wall-clock, not attempt count, so an LLM that
takes 60s no longer prematurely gives up
Copilot round-2 feedback:
- _proxy_sse with use_service_auth=True silently fell back to
empty headers when proxy_token_mgr was None, producing a
retry-storm 401/403 loop. Fail fast with a 503 + clear log
so the misconfig surfaces immediately.
- Mobile <700px CSS comment claimed buttons "stretch to full
row width" but the rule keeps flex-direction: row with
flex: 1 on each, giving 50/50 side-by-side. Updated the
comment to match the deliberate side-by-side layout
(stacking would push the action row below preview/disclosure
on tall envelopes; 50/50 keeps both verbs reachable).
|
||
|
|
68e1332c59 |
fix(approve): route child approvals through proxy + poll for late judge verdict
Two bugs reported from local repro on PR #424: 1. Approve/Deny buttons return HTTP 404 on every click. The new approveWorkstream helper hit /v1/api/workstreams/{ws_id}/approve regardless of target — that path is only mounted for coord workstreams (which live on the console process). Child workstreams live on cluster nodes and need to round-trip through the routing proxy at /v1/api/route/workstreams/{ws_id}/approve, which resolves the ws_id to its owning node and forwards the body verbatim. approveWorkstream now picks the path based on whether targetWsId matches the coord's own wsId. 2. LLM judge verdict never populates — rows freeze on the heuristic-tier pill ("⚙ heuristic") even after the judge would have completed. The judge runs async on the child node via a daemon thread and updates _llm_verdicts there, but no signal propagates back to the coord — cluster_state events don't fire on verdict-only changes, and the live-bulk TTL is 5s with no periodic poll. Added _maybePollForJudgeVerdict: when renderChildRow encounters a pending_approval_detail with judge_pending=true and items missing judge_verdict, schedule a recursive 2s urgent live-bulk re-fetch. Self-terminates when the verdict lands, the row closes, the approval clears, or attempts hit the cap (≈12s for a failed/timed-out judge so we don't poll forever). Single timer per ws_id; re-renders are no-ops while a timer is in flight. Smoke-test assertions added for both fixes so a regression on either path surfaces at test-time. |
||
|
|
7e33fc68bb |
fix(approve): apply /review feedback on inline child approvals
Critical:
- coordinator.js RISK_SEVERITY accepted 'crit' only; production
emits 'critical' (per turnstone/core/judge.py:1556 + heuristic
seeds). A risk_level=='critical' verdict ranked as 0 and
rendered with .risk.low (green) styling, never triggering
the crit-risk auto-expand. Now accepts both aliases. Unknown
risk_level falls back to rank 2 ('high') so future schema
drift fails *safe* (over-alert) instead of silently
downgrading. Pill ternary handles both 'crit' and 'critical'
alias to the existing .risk.crit class.
Major:
- Urgent live-badge flush now coalesces N urgent calls in the
same JS tick into one bulk request via queueMicrotask, instead
of firing N single-id fetches. The motivating 10-children-
pending-bash scenario in the design doc now lands on one bulk
/v1/api/cluster/ws/live request.
- Test coverage gap: added test_session_ui_base.py cases for
POLICY-BLOCKED (item.error + needs_approval=False) and
judge-unavailable (no verdict + no judge_pending) matrix rows.
Added literal-string assertions to the smoke list in
test_coordinator_page.py so a refactor dropping either branch
surfaces at test-time.
Minor batch (4 coord.js + 1 CSS + 1 fake-divergence):
- 409 stale-call_id path re-enables both buttons before return
(urgent fetch is best-effort; could also fail).
- judgePending pill no longer conflicts with a present heuristic
verdict — guard changed from !judge to !verdict.
- Empty <div class="approval-reasoning"> no longer appended when
reasoning is absent but evidence is present (evidence still
renders inside the disclosure).
- Dead .ch-row .approval-pill.rec-* CSS rules removed (JS never
combines those classes). Recommendation chip in the disclosure
footer now has its own scoped rules so the chip is actually
styled.
- _FakeUI.serialize_pending_approval_detail call_id selection
aligned to the real impl's "first non-empty" semantics.
- liveBadgeCache reconnect cleanup now preserves permanent
(403/404) entries — denied users no longer pay one wasted
bulk fetch per denied id per reconnect.
All 4465 non-live tests pass. Ruff + mypy clean. node --check OK.
|
||
|
|
a369d5f0d0 |
feat(approve): clear live cache on SSE reconnect — chunk 4 reconnect parity
Closes the stale-button window where a sub-5s SSE gap would leave liveBadgeCache holding pending_approval_detail for a child whose approval was actually resolved during the gap. Without this clear, zombie approve/deny buttons render until either the next child_ws_state event or the natural TTL expiry (whichever comes first). The clear sits beside the existing activeWaits.clear() in the reconnect handler — same posture (drop client-only state that the server's SSE replay doesn't cover) and same blast radius. The 409 race guard in submitChildApproval would catch a stale-call_id POST even without this, but rendering wrong UI until the operator clicks is the worse failure mode. loadChildren's finally block already fires scheduleLiveFetch for every visible row after the replace-mode refresh, so the cache repopulates with authoritative pending_approval_detail in one bulk request within the next debounce window. Plan: docs/design/inline-child-approvals.md (chunk 4 of 4 — last required chunk; 5/6 are stretch). |
||
|
|
54f04496c3 |
feat(approve): inline approve/deny buttons + judge verdict pill on coord tree
Chunk 3 of the inline-child-approvals plan + the SSE pipeline plumbing
needed for sub-second urgent fetches.
JS (coordinator.js):
- approveWorkstream(targetWsId, body) — generic POST helper, callable
for both the coord-self dock and the new per-child inline buttons.
- renderApprovalBlock(child, detail) — risk-level pill (.risk.* per
the design system primitives), tool-name summary with "+ N more"
for envelope-level approvals, intent_summary, ↳ judge reasoning
teaser, ▸ more disclosure carrying the recommendation chip,
evidence list, and items 2..N stacked sub-blocks. Plus matrix
coverage: judge_pending / judge unavailable / tool-policy
blocked / multi-item.
- submitChildApproval — handles the 409 stale call_id race by
invalidating the live cache + urgent-refetching, optimistically
clears pending_approval_detail on success.
- scheduleLiveFetch({ urgent: true }) — bypasses the 5s TTL +
cancels the debounce so attention transitions surface inline UI
immediately instead of after the next polling window.
- handleChildState fires urgent on activity_state="approval"
enter/leave; handleChildClosed eagerly invalidates the live
cache so closed rows can't render stale buttons.
CSS (index.html):
- New .approval-block / pill / preview / actions / disclosure
styles. Inline .act buttons duplicate the dock's colour treatment
(the dock-scoped rules don't reach the children-tree). Mobile
<700px touch targets ≥44px.
Pipeline (collector.py + coordinator_adapter.py):
- All three cluster_state event emitters and the child_ws_state
re-emit now carry activity_state. The previous omission left
the urgent-fetch trigger as dead code — discovered in review.
Tests:
- Static smoke test in test_coordinator_page.py asserting the new
helper names exist + the pending_approval_detail key is read.
Plan: docs/design/inline-child-approvals.md (chunk 3 of 4).
|
||
|
|
e42add1b77 |
feat(coordinator): coordinator workstream kind — phase 1 (#368)
* feat(coordinator): coordinator workstream kind — phase 1
Adds a new ``kind="coordinator"`` workstream that runs inside the
``turnstone-console`` process (first ChatSession hosted on the console)
with a dedicated tool set for spawning and driving child workstreams.
Supersedes the external ``turnstone-coordinator`` MCP side-car for new
installs; the extension is marked deprecated in
``examples/mcp-cluster-ops/README.md`` but still works on 1.4-and-earlier
clusters.
Phase 1 ships: the workstream class, 6 lifecycle tools, console hosting,
9 HTTP endpoints, per-user audit attribution, and a one-pane web UI at
``/coordinator/{ws_id}``. Node/skill discovery tools, task-list tool,
tree-view UI, and routing-proxy audit middleware follow in a later PR.
## Schema
Migration 039 adds ``kind`` / ``parent_ws_id`` columns + indexes to
``workstreams``. Both SQLite and PostgreSQL backends take the new
kwargs on ``register_workstream``; empty-string ``parent_ws_id``
normalises to ``NULL`` at the storage edge. PostgreSQL uses
``INSERT ... ON CONFLICT DO NOTHING`` to match SQLite's ``OR IGNORE``
and close a pre-existing SELECT-then-INSERT TOCTOU window.
``list_workstreams`` gains optional ``parent_ws_id`` / ``kind`` filters;
new ``get_workstream(ws_id)`` returns the full row (the existing
``get_workstream_metadata`` stays untouched for back-compat).
## Core session + kind routing
- ``ChatSession.__init__`` accepts ``kind`` / ``parent_ws_id`` /
``coord_client``. On ``kind="coordinator"`` it swaps
``_tools = COORDINATOR_TOOLS`` and zeros sub-agent tool lists.
- ``Workstream`` dataclass extended with ``user_id`` / ``kind`` /
``parent_ws_id``. Both ``WorkstreamManager`` and the new
``CoordinatorManager`` use the same type — no parallel hierarchy.
- ``_SessionFactory`` Protocol + server / cli factory closures thread
the new kwargs. ``POST /v1/api/workstreams/new`` rejects
``kind != "interactive"`` with 400; ``POST
/v1/api/workstreams/{ws_id}/open`` refuses coordinator rows so a
server node can't accidentally rehydrate one.
## Coordinator tool set
Six tools (``spawn``, ``inspect``, ``send``, ``close``, ``delete``,
``list_workstreams``) with a ``coordinator: true`` metadata flag,
scoped to coordinator-kind sessions only. ``inspect`` and ``list`` are
auto-approved reads; the four mutators need approval. ``list`` returns
``{"children": [...], "truncated": bool}`` so the model can detect
post-filter under-fill and paginate.
## CoordinatorClient (in-process, sync)
Mutating ops HTTP-POST to the console's own ``/v1/api/route/*`` on the
local bind URL so every existing middleware (auth, rate-limit) runs.
Read ops hit ``storage.list_workstreams`` / ``get_workstream`` /
``load_messages`` directly — the routing proxy doesn't expose
list/inspect paths. URL paths are a validated constant table (avoids
an httpx ``base_url``-merge trap). A new
``/v1/api/route/workstreams/delete`` proxy handler joins the existing
route-proxy endpoints.
## Per-session coordinator JWT
``CoordinatorTokenManager`` mints short-lived JWTs with ``sub=<real
user>`` (attribution preserved), ``src="coordinator"``,
``aud="turnstone-console"``, ``coord_ws_id=<ws>`` custom claim.
``_proxy_auth_headers`` preserves ``src`` + ``coord_ws_id`` across the
upstream re-mint so server-side middleware sees coordinator-origin,
not ``console-proxy``. ``AuthResult.extra_claims`` carries
non-reserved claims through validate→remint; ``create_jwt``'s
reserved-claim set (now including ``nbf`` / ``jti``) is symmetric with
``validate_jwt``.
## Console hosts the ChatSession
- New ConfigStore settings: ``coordinator.model_alias`` (required),
``reasoning_effort``, ``max_active`` (default 5),
``session_jwt_ttl_seconds``.
- Console lifespan builds a ``ModelRegistry`` +
``CoordinatorManager``. Missing / unresolvable alias returns **503**
with remediation text — never 500.
- ``CoordinatorManager``: placeholder-slot reservation under lock,
rollback on factory failure, per-ws_id rehydration lock to serialise
concurrent lazy-opens, ``max_active`` enforced via ``close_idle``
eviction semantics.
- ``ConsoleCoordinatorUI`` is a thin ``SessionUI`` implementation — no
global broadcast, no per-node metrics, shared
``_APPROVAL_WAIT_TIMEOUT`` constant across approval + plan paths.
- No eager startup rehydration: persisted coordinator rows load lazily
on first ``GET /v1/api/coordinator/{ws_id}``.
## Console coordinator API
Nine endpoints under ``/v1/api/coordinator/*`` gated by ``approve``
scope + new **``admin.coordinator``** permission (added to
``_VALID_PERMISSIONS``; not in any builtin role — operators opt in
explicitly). Ownership failures return **404, not 403** and use
strict equality so empty-owner rows don't leak across tenants.
Correlation-id masking on every factory-raising path
(``coordinator_create`` + ``coordinator_detail`` lazy rehydrate) — no
stack traces to the client.
## Audit attribution
Three console-side events (``coordinator.create`` / ``.close`` /
``.cancel``) with the real creator's ``user_id`` plus
``detail={coord_ws_id, src="coordinator"}``. No schema migration
required. Per-tool-call audit across the routing proxy is deferred
(needs either a ``source`` column on ``audit_events`` or
``record_audit`` calls wired into the route-proxy handlers).
## Web UI (``/coordinator/{ws_id}``)
One-pane chat served by the console. Reuses ``shared_static``
(``base.css``, ``auth.js``, ``theme.js``, ``toast.js``, ``utils.js``,
``kb.js``) and the server UI's ``renderer.js`` pipeline (KaTeX, Mermaid,
highlight.js already bundled).
- SSE to ``/v1/api/coordinator/{ws_id}/events`` with exponential-
backoff reconnect; status line carries a leading glyph
(● / ○ / ⚠) so state isn't conveyed by colour alone.
- Renders content, reasoning (dimmed italic
``.role-reasoning``), tool_result, approve_request, intent_verdict,
output_warning.
- Child ws_id references auto-wrap to
``/node/{node_id}/?ws_id={child}`` links — both ids regex-validated
before interpolation, everything else HTML-escaped.
- Non-modal approval bar (``role="region"``) with a batch header
("Approve N tool calls"), initial focus on the approve button,
buttons disabled during the in-flight POST, red-bordered deny.
``aria-live`` flips to ``off`` during streaming.
- "New coordinator" button on the dashboard header — permission-gated
on the UI side, matching the backend 403.
- Mobile composer capped under ``@media (max-width: 700px)``.
## Tests
~120 new tests across 8 files: workstream-kind storage + dataclass
semantics, CoordinatorClient URL map + token minting + storage reads +
truncation signalling, tool prepare/exec dispatch and approval gating,
CoordinatorManager create / rollback / eviction / lazy rehydration +
concurrency, HTTP endpoint auth + 404-on-ownership + 503-on-misconfig,
proxy-auth ``src`` preservation, full lifecycle end-to-end, coordinator
page HTML-injection guard. ``test_tools_schema.py`` widened to 25
tools (19 existing + 6 coordinator).
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 files), ``pytest`` 4054 passed (5 pre-existing failures unrelated
to this change — confirmed against ``main``).
* polish(coordinator): address PR review + CI + tool-namespace isolation
CI:
- `ruff format`: two files reformatted, matches the in-repo pre-commit config.
- `wheel-completeness`: add `turnstone/console/static/coordinator/*.html` +
`*.js` to the hatch wheel-include list. Without this the coordinator UI
was missing from published wheels.
- `test (3.11/3.12/3.13)` + `test-postgres`: three `TestExecReadImage`
tests were masking a real bug — my 6 new tool JSONs pushed tool count
19→25, crossing the default `tool_search.auto` threshold (20), which
made `ChatSession.__init__` construct a `ToolSearchManager` and cache
`_cached_capabilities` during init. Tests that later patched
`session._provider.get_capabilities` saw the cached value instead.
Root-cause fix: the tool-search threshold code path now reads
capabilities through `_resolve_capabilities(...)` directly — no cache
populate — so the patch takes.
Tool-namespace isolation (bigger fix than CI symptoms suggested):
- `TOOLS` was the union of all loaded tool JSONs including the 6 new
coordinator tools. Interactive sessions were getting coordinator
tools in their function-calling surface (which is nonsense — they
require a console-hosted `coord_client`), and coordinator sessions
counted against the interactive tool-search threshold. Fix:
- New `INTERACTIVE_TOOLS` / `INTERACTIVE_TOOL_NAMES` in
`turnstone/core/tools.py` exclude anything with `coordinator: true`
metadata. `TOOLS` stays as the union for schema introspection +
eval catalog.
- `ChatSession.__init__` selects tool set by kind: coordinator gets
fixed `COORDINATOR_TOOLS` (no MCP merge, no listeners registered);
interactive gets `INTERACTIVE_TOOLS` (+ MCP if configured).
Coordinators are meta-orchestrators that spawn child workstreams;
MCP tools / resources / prompts live on the children, not on the
coordinator's own surface.
- `_on_mcp_tools_changed` no-ops for coordinator sessions
(defence-in-depth in case listeners were registered).
- `always_on_names` on `ToolSearchManager` is now the set of builtin
tools actually present in the session (kind-aware) rather than the
full `BUILTIN_TOOL_NAMES` frozenset.
- `turnstone/eval.py` uses `INTERACTIVE_TOOLS` (coordinator tools
aren't in scope for the eval harness which tests interactive agent
behaviour).
- Regression tests in `tests/test_workstream_kind.py`:
- `INTERACTIVE_TOOLS ∩ COORDINATOR_TOOLS == ∅` and their union is
`TOOLS`.
- Interactive `ChatSession._tools` does not include any
coordinator tool name.
- Coordinator `ChatSession._tools` contains `spawn_workstream` but
not `bash` / `edit_file` / `memory`; sub-agent lists are empty.
- Coordinator `ChatSession` with an MCP client attached does NOT
merge MCP tools and does NOT register any MCP listeners.
PR review findings:
- **#10 / #11** (Copilot): coordinator UI claimed to reuse the server
renderer pipeline but loaded none of its JS. Mirrored
`turnstone/ui/static/renderer.js` into
`turnstone/console/static/coordinator/renderer.js` (flagged in-file
as a cleanup candidate to promote into `shared_static/`), added
`katex.min.js` / `highlight.min.js` / `renderer.js` script tags to
`coordinator/index.html`. `coordinator.js` now buffers raw markdown
via `textContent` during streaming, then swaps to `renderMarkdown` +
`postRenderMarkdown` on `stream_end`.
- **#7** (Copilot): N+1 query pattern in
`CoordinatorClient.list_children()` — per-row `storage.get_workstream`
just to read `skill_id`. Pushed `skill_id` + `skill_version` into
the `list_workstreams` SELECT projection on both backends; the
client reads them from `row._mapping` directly. New
`test_list_children_skill_filter_avoids_n_plus_one` pins the
behaviour (asserts `storage.get_workstream` call count is 0).
- **#8 / #9** (Copilot): `spawn_workstream` tool JSON said "if empty,
the workstream is created idle" but the prepare method rejected
empty and the field was marked required. Resolved by allowing
empty end-to-end: removed from `required`, prepare builds a
"spawn idle workstream" header + empty preview when empty,
updated `test_spawn_prepare_allows_empty_initial_message`.
- **#1–#5** (github-code-quality): five asserts with side-effecting
method calls in `test_coordinator_manager.py` (`mgr.close`,
`mgr.open`, `mgr.create` in a dead `_c = ...`). Extracted each
call to a local variable so `python -O` can't strip the side
effect.
Verification:
- `ruff check turnstone tests` clean.
- `mypy turnstone` clean (156 source files).
- `pytest -m "not live"` — 4063 passed, 3 deselected (live-backend
tests), 0 failed. The 3 image tests that were failing on this
branch now pass; wheel + lint both green locally.
* polish(coordinator): address Copilot re-review findings
Two findings from the re-review of #368 after the first polish commit.
**user_id wired into `mgr.create()` at the server handlers.** Phase 1
added ``user_id`` to the ``Workstream`` dataclass and
``WorkstreamManager.create()`` signature, but the two call sites in
``turnstone/server.py`` forgot to pass the authenticated caller
through. Result: interactive workstreams created via
``POST /v1/api/workstreams/new`` (including coordinator-spawned
children, which route through this handler) were landing with blank
``user_id``, defeating ownership-based access control on subsequent
sends / approvals / closes (``_require_ws_access`` treats blank
owners as legacy/allowed). Two changes:
- ``server.py:create_workstream`` forwards ``user_id=uid`` — the same
``uid`` already resolved from the auth result (with trusted-service
forwarding preserved).
- ``server.py:open_workstream`` prefers the persisted owner on the
workstream row over the rehydrating caller so reloading someone
else's workstream doesn't silently re-parent it. Falls back to
the authenticated caller when the stored row has no owner
recorded (pre-phase-1 rows).
Regression test in ``tests/test_workstream.py`` pins
``WorkstreamManager.create(user_id=X)`` → ``ws.user_id == X`` so the
manager seam can't regress silently on a future refactor.
**Malformed-JSON recovery allowlist expanded for coordinator args.**
``_prepare_tool()`` has a two-stage salvage path for models that
emit malformed JSON: a regex-extract (fallback 1) and a bare-string
→ primary_key wrap (fallback 2). The fallback-1 key list didn't
include coordinator argument names, so a slightly malformed
``spawn_workstream`` / ``send_to_workstream`` / etc. call would
hard-fail instead of salvaging into a minimal-args dict for retry.
Added ``ws_id`` / ``message`` / ``initial_message`` / ``parent_ws_id``
to the allowlist (kept alphabetised) so the coordinator tools get
the same model-self-correction behaviour as the interactive tools.
Fallback 2 already covers the ``ws_id``-primary-key tools via
``PRIMARY_KEY_MAP``; the regex path matters when the model emits
``{"ws_id": "abc", "message": "..."}`` with a trailing syntax error.
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` → 4065 passed, 3
deselected (live-backend), 0 failed.
* fix(coordinator): address ultrareview findings on coordinator workstream kind
Security
- Cross-tenant leak: CoordinatorClient.inspect/list_children now constrain
to the coordinator's own ws_id + direct children; an LLM coerced via
prompt injection can no longer exfiltrate other tenants' workstreams.
- Empty-owner short-circuit bypass: strict equality at coordinator.py
ownership gate and at the storage-fallback branch in coordinator_history;
orphan/system-owned coordinator rows can no longer be rehydrated by
arbitrary holders of admin.coordinator (DoS + history disclosure vector).
- Closed coordinators no longer silently resurrect on subsequent GET —
the Close button is now actually durable across URL revisits and tab
refreshes; rows with state in {closed, deleted} refuse rehydration.
Correctness
- ChatSession.close() now releases the CoordinatorClient httpx.Client
pool; previously every closed/evicted coordinator dropped a connection
pool on the floor until non-deterministic GC.
- open_workstream rehydration now forwards parent_ws_id + kind, so
coordinator-spawned children survive node restart / idle eviction
with their parent link intact instead of becoming silent orphans.
- list_children truncated flag now signals whenever the SQL fetch hit
the page cap (previously permanently False in the no-filter case,
causing confident-but-incomplete summaries from the coordinator).
- ConsoleCoordinatorUI.approve_tools: per-tool auto-approve now checks
auto_approve_tools independently of the blanket auto_approve flag,
so 'Always approve this tool' actually works on the next invocation.
Concurrency
- _spawn_worker no longer falls through to start a second concurrent
worker thread on the same ChatSession when queue.Full fires; instead
send() returns False and the endpoint surfaces HTTP 429.
- _open_locks entries are now refcounted under self._lock and only
popped when the last waiter releases — eliminates the race where a
rehydration-failure path lets two threads serialize on different lock
instances for the same ws_id and trip the "already tracked" guard.
Tests: +6 regression cases covering closed-coordinator refusal,
empty-owner non-admin refusal, queue.Full no-duplicate-worker,
inspect/list_children cross-tenant rejection, and truncated semantics.
|