12 Commits

Author SHA1 Message Date
Patrick Buckley 480a1426b3 Fail-closed history-commit handoff (#1005)
* fix(session): fail-closed history-commit handoff (#981)

The deleted-workstream discovery is now a terminal, ws_id-keyed latch:
keyed conversation commits refuse admission once the durable parent is
gone (convergence finalizers and force-abandon are exempt), history
handoff refuses to mint a proof token so /history fails closed with a
503 instead of silently wiping the pane, and the SSE stream carries a
workstream_gone resync reason. Discarded commits leave a forensic log
of commit keys and roles, never content.

Conversation rows gain a commit_key (migration 071): keyed saves are
idempotent under retry, validated against the full commit identity, and
refused when they would cross a workstream deletion. The prune orphan
category now requires a NULL alias plus a two-hour updated grace, with
cutoffs computed at discovery time and carried into both dialects'
rechecks.

The mid-turn interjection queue is owner-partitioned with no per-site
mode flags: pops take the acting principal's and unowned rows, other
participants' rows are structurally retained, and enforcement lives at
queue admission plus the shared before_spawn gates. The retraction
ledger is bounded by open pop windows: pops open a window atomically
with the queue delete, restores close their ids atomically with the
ledger consume, every other exit closes through one helper, and misses
for unheld ids record nothing. The workstream-gone latch refuses
unattended wakes at all three gates (watcher spawn, claim, delivery
pre-pop), and the retry dispatcher regained its pre-envelope
cancel/error convergence net.

Persistence-state reporting derives through the session bound to each
UI instead of a registry lookup by id that failed open to healthy
during tombstone retention. The dashboard roster no longer re-inserts
ghost entries from trailing activity events, the history tool-outcome
scan tolerates interleaved non-turn rows, and the shared
handoff-deadline handle owns its own retirement.

Single-sourced across call sites: keyed-commit row values, attachment
save wrappers, tail-truncation and conflict-resolution bodies for both
storage dialects; worker-slot lifecycle field sets; the direct-commit
admission frame; queued-row layout accessors; the string-aware comment
stripper shared by every JS harness suite.

Refs #981 #964

* fix(session): sweep handoff fixes to their sibling surfaces

The interactive replay loop treated a system row as a tool-batch
boundary, so every tool result after an interleaved row vanished from
that pane while the coordinator rendered the same history correctly.
Only a conversational turn ends the batch window now, matching the
shared outcome index.

Accepted user turns clear the composer's attachment chips on the same
viewer policy that settles optimistic bubbles rather than on having
matched a local bubble, so a workstream created with an upload no
longer keeps a chip for an attachment the create dispatch already
consumed. The coordinator's raced-Stop arm emits the stream-end hook it
inherits alongside the idle state, leaving no unfinalized bubble or
unflushed tool output. Ending a session surfaces a failure toast when
the request never lands or answers with a non-JSON body.

The per-second persistence reconcile now probes each session without
blocking: a workstream whose generation and handoff locks are held is
skipped until the next pass instead of contending the locks every
commit needs. The one-shot repair that gates workstream creation at
capacity keeps a definite probe — it has no next pass, and the sessions
likeliest to be contended are the ones whose unresolved journals
emptied its candidate list.

Single-sourced: the attachment lane builds its conversation row through
the shared commit-identity builder; the ordinary worker exit releases
its slot through the lifecycle owner; both operator surfaces snapshot
their counters through one non-consuming helper; the replay preamble
loses its per-kind wrappers and its config hook; the browser harness
suites share one brace walker; and each in-flight history attempt is
one record carrying both its abort controller and its deadline.

Refs #981 #964
2026-08-11 04:18:36 -07:00
Patrick Buckley 7ed5d90a98 test(e2e): E8 observes the double render; honest non-vacuity for E6 (#900 r2)
Until now nothing in this harness could see the artefact the campaign
prevents. Every scenario counts .msg.user rows, and user rows never
travel on the SSE stream — a /send emits none, only /history replay
paints them — so a duplicated assistant bubble was invisible to all of
them. E8 counts a sentinel's occurrences in the transcript text instead,
which is structure-agnostic across duplicate bubbles and tool blocks.

The window it drives is the one readyState cannot see: the retry fires
with the transport OPEN, its /history is held, and inside that await the
transport drops and re-establishes. readyState reads OPEN afterwards
exactly as before. The redial is a real disconnect+connect rather than a
visibility change on purpose — a hide leaves evtSource null, which the
presence term already decides, so a hide-based control would pass for the
wrong reason. Control: stripping the generation term stamps dupes2.

E8's expectation needed correcting once: unlike E6, the stream is live at
flush time here, so the declined render's queued settle fires the
transport-free backstop and the pane converges in one settle. That is
correct behaviour, so the heal is asserted as a convergence leg — a "no
duplicates" verdict must not be earnable by rendering nothing.

E6 gains the non-vacuity it was missing: history_requests counts on
ARRIVAL, before the hold and before any status is chosen, so "the gate
declined a good payload" and "there was no good payload" stamped
identical observables. history_ok — incremented only when the production
route answers 200 — closes that, including the production-side-failure
hole an injected-fail budget cannot see.

Both hidden-window detectors widened for the additive jitter: sized on
the 2000 floor alone they would have closed before a top-of-range firing
and reported hidden0 for the wrong reason. delay_history(0) comments
corrected — it cannot release an in-flight hold.
2026-07-24 18:15:03 -07:00
Patrick Buckley 60f6dc07a2 fix(#894): drop the unreachable epoch guard; abort-Set; bump-after-delete; producer pins
Review round 9 (4 minor bug, 4 quality, 1 perf nit; security zero).

- The r8 clearUiEpoch guard was UNREACHABLE (r9 bug find): clear_ui
  always dispatches immediately after bumping, so a stale-epoch
  dispatch is also a stale-seq dispatch and the currency gate discards
  it before it can paint or clear — the client half of the joined-
  flight fix was already carried by seq, and the server generation key
  is the sole load-bearing layer.  Machinery removed (decl, bump,
  capture, conditional clear, section-9 pins); the latch-clear comment
  now states the two-layer accounting.
- destroy()'s abort handle becomes a Set: a newest-wins single slot,
  nulled by the newer dispatch's finally, left an OLDER overlapping
  fetch unabortable — the destroyed closure pinned for the bound's
  remainder.  Pinned.
- _history_generation now bumps AFTER delete_messages_after: flights
  rebuild from storage, so old-generation-reads-post-delete is the
  harmless spuriously-fresh direction while new-generation-reads-
  pre-delete would be wrongly joinable; the count/floor error paths
  correctly leave it unbumped.  Two-arm producer pin in
  test_rewind_retry (persisted-rows bump on rewind AND retry;
  in-memory-only error path must NOT bump) — the flight test's mock
  can no longer mask a deleted bump.
- The harness load_calls increment takes a lock (to_thread workers
  genuinely overlap under delay_load; a lost update false-fails G7).
- G7's viewer B is now a background authenticated GET (a raw request
  enters load_messages identically; the second browser bought no
  proof); stale two-tuple key comments and the coalescing matrix line
  updated; the _send_in_page enumeration dropped for prose.

250 pins green; G1/G6/G7 re-run READY.
2026-07-24 15:04:31 -07:00
Patrick Buckley b85f792925 test(e2e): G7 joined-flight detector at the flight layer; fix the generation read path it caught (#894 r8)
G7: two browsers on one ws; delay_load parks B's pre-rewind /history
flight open INSIDE load_messages — the flight layer.  (A first cut
held via delay_history, which sleeps in the FAULT layer before the
route: flights never overlapped there and the 'negative control'
passed vacuously — a false detector, caught and rebuilt.  The knob
also sleeps AFTER the load so a parked flight holds the rows it
actually read: its transaction point.)  A rewinds mid-hold; the miss
proof is load_calls growing TWO (a joined request never enters
load_messages — the e2e twin of the unit test's proof) plus A
rendering the post-rewind single row.

The rebuilt detector immediately caught a real bug in the server fix:
mgr.get returns the Workstream WRAPPER, and the route's direct getattr
for _history_generation silently defaulted to 0 forever — joining
stayed enabled while the unit test's mock (attr on the wrong object)
masked the shape.  The route now reads ws.session, and the mock pins
the nested shape so a wrong-object read can never pass again.

Negative control (flight key reverted to (ws_id, limit)): stamps
FAILED-loads1-rows3 — A joins the pre-rewind flight and paints three
stale rows as fresh truth.  Fixed: READY-posts1-loads2-rows1.
2026-07-24 15:04:31 -07:00
Patrick Buckley 30b6ff7f7b fix(#894): rewind-freshness epoch closes the #884 joined-flight window; destroy aborts the bounded fetch
Review round 8 (2 major + 2 minor bug, 2 major + 3 small quality;
security/perf zero at five consecutive rounds).

- clearUiEpoch (r8 major): the #884 /history single-flight can hand a
  joiner a payload whose load_messages ran BEFORE the rewind committed
  (the flight key is (ws_id, limit); joining is invisible to the
  client, and the client seq stamp cannot see server-side staleness) —
  reachable single-user (rewind clicked during a truncated-resync
  fetch joins that flight) and multi-viewer (any concurrent pane's
  /history).  The joined payload rendered as 'success' and CLEARED the
  latch: the original over-rewind window, resurrected through the
  server seam.  Fix: the epoch bumps at clear_ui arrival, every
  dispatch captures it pre-await, and only a dispatch that post-dates
  the latest clear_ui may CLEAR the latch — a pre-rewind payload may
  still paint (stale-but-real posture, gate holds), the surviving
  latch arms the retry, and the retry's fresh dispatch starts a new
  flight with post-rewind truth.  Producer/consumer/placement pinned.
- destroy() aborts the in-flight bounded fetch (activeHistCtrl): the
  r7 15s bound alone pinned a destroyed pane's closure until it fired
  — the same dead-not-inert ruling destroy applies to staleRetryTimer.
  Pinned.
- stop(hard=True) no longer sets force_exit: it skipped the ASGI
  lifespan teardown and leaked the #885 daemon threads + sse_executor.
  The 2s graceful-shutdown timeout already force-closes open SSE, and
  the lifespan runs on both paths (docstring corrected; G6 re-verified
  — the orphan still manifests).
- G6 pacing sized above the scenario's worst-case deadline sum (~200s
  vs ~95s) so the in-process bash cannot resolve the orphan
  mid-scenario and degrade the detector to a false READY.
- Quality: the r6 reachability comments rewritten to the r7 truth
  (orphan REAL via hard crash; graceful-close-only synthesis); the
  bound's WIRING pinned (signal reaches getJSON; getJSON forwards
  init); _strip_comments deduped (4 inline copies); seq comment
  re-paired with its asserts; retry_fire window tail-anchored.

Full harness (16/16 scenarios) + full suite (9724) green on the prior
commit; 136 pins green here.
2026-07-24 15:04:31 -07:00
Patrick Buckley 412161aa4a fix(#894): live-set retirement policy — transport death is not retirement; bound the refetch await; G6 hard-kill detector
Review round 7 (2 major + 1 minor bug, 4 minor quality; security/perf
zero).  Both majors traced the r6 stratum:

- The closeStreamTransport drain of liveToolCalls rested on a false
  re-announcement premise (verified: replay_ok yields only events past
  the cursor; the coord fresh/truncated replay yields connected/status/
  pending-cards/verdicts, never tool_pending/tool_info).  An emptied
  set fails OPEN — a mid-batch redial plus a slow seedless refetch
  wiped the live batch.  Retirement policy re-derived at the decl: an
  id leaves on its RESULT, at the SETTLE edge, or with pane death;
  transport death is NOT a retirement event; a stale id fails CLOSED
  (skip, latch survives, settle heals).  Site-anchored pins: the one
  drain inside the idle/error block, the delete inside tool_result,
  the adds inside tool_pending/tool_info, and closeStreamTransport's
  comment-stripped code may not touch the set.
- G6's kill was not a kill: RecoveryServer.stop() gracefully closed
  workstreams, and session.cancel()'s bash path persisted 'Cancelled by
  user' BEFORE the reboot — the r6 'recovery synthesizes' ruling was
  observing the cancel path.  stop(hard=True) (skip the close sweep +
  uvicorn force_exit: a crash does not drain SSE) leaves the orphan
  genuinely unresulted — REACHABILITY FLIPS: the poisoned-pane state is
  real, the live-set hardening is reachably load-bearing, and G6 is now
  its behavioral detector: hard kill -> reload paints the orphan
  (asserted PRESENT) -> the seedless rewind renders THROUGH the residue
  (rewind-for-retry truth: the user message stays), negative-controlled
  against the DOM-probe encoding (stamps orphan1, hist2).  Discovered
  and tracked separately: a hard-crashed reborn node answers stale-high
  cursors with a silent fresh stream (no replay_truncated — the honest
  truncation signal rides gracefully-persisted state).
- refetchHistory's await is now bounded (AbortController + 15s, the
  coordSend shape): an accepted-never-answered /history pinned
  refetchesInFlight and permanently disabled both heals.  Pinned.

Quality: seq-producer position pinned earlier; the stale section-7
comment corrected; retry/backstop guard windows comment-stripped
(vacuous-by-comment-mention foreclosed); the shared G3/G4 double-fail
prologue extracted into _coord_stick_latch.  G3/G4/G6 re-run READY.
2026-07-24 15:04:31 -07:00
Patrick Buckley 5386d598ef test(e2e): native-transport roster scenario (F2) + strict absence assertions (#881)
Scenario F splits into F1 (manual ?last_event_id= transport) and F2, the
native-header sibling — the only behavioral coverage of two pure-browser
semantics no Tier-1 harness can express: the auto-reconnect header echo,
and id-less frames inheriting the connection's persisted lastEventId
(the mechanism behind app.js's node_snapshot-branch clear).  F2's phase
C is the round-3 fix's discriminator: after the native heal, a forced
manual reconnect must go CURSORLESS with no second truncated round
(pre-fix: cursor1-trunc2), guarded by an idFrames precondition against
the aggregate tick.

Two harness seams earned by F2's first failures, both documented at
site: a failed EventSource reconnect attempt is TERMINAL per WHATWG, so
the restart must never expose a refused window — a SO_REUSEPORT
placeholder binds before the old node stops and hands its backlog to the
successor's uvicorn (make_listen_socket + RecoveryServer sock
injection); and an SSE stream still open at stop() parked uvicorn's
graceful drain indefinitely — timeout_graceful_shutdown=2 bounds it with
the #885 lifespan teardown intact.

Absence assertions tightened (round-4 review): ghost-gone now requires
absence from BOTH the model and the rail via _roster_absent_ws — the
negated AND-membership helper De Morganed into either-surface and could
false-pass a rail-render regression.
2026-07-22 23:44:04 -07:00
Patrick Buckley ea706bf4c0 test(e2e): roster-restart scenario proves the global boot-epoch heal (#881)
Scenario F drives the REAL node dashboard (/ + app.js) through a node
restart on the global stream — no custom page; transport instrumentation
is injected via CDP addScriptToEvaluateOnNewDocument, scoped to
/events/global URLs so per-ws streams can't pollute the counters.
Phase A is the negative control: live roster, live cursor, zero
replay_truncated.  Phase B: hide, force the CLOSED state (a closed
EventSource never auto-retries, making the show edge's manual reconnect
the only reconnect), restart the node re-opening only one of two
workstreams, show.  Asserted: cursor presented via ?last_event_id= and
replay_truncated observed at the transport, the not-reopened
workstream's ghost evicted from the roster model and rail (the dashboard
table's membership refreshes on interaction by design — documented at
_roster_has_ws), and the reborn node's global_events_requests counter
proves the reconnect hit the real endpoint.  The native header
transport differs only in carriage and is pinned by the Tier-1
boot-epoch tests.
2026-07-22 23:44:04 -07:00
Patrick Buckley a935ae3106 fix(server): give the lifespan daemon threads a real shutdown (#885)
_global_fanout_thread, _aggregate_emitter_thread, and
_idle_cleanup_thread were daemon threads with no stop signal — shutdown
abandoned them mid-loop.  The sleep-loop pair now waits on a shared
Event (wait doubles as the tick sleep, so a set wakes them immediately);
the fanout exits on an identity-checked queue sentinel, FIFO-draining
everything enqueued before it (sessions close earlier in the shutdown
tail, so their final events still fan out).  Joins are bounded and
off-loop; daemon=True stays as the backstop for a join timeout, not the
mechanism.  The recovery harness drops its thread-neutering workaround
(module docstring piece 4) — the global lane now runs REAL in harness
boots, which the #881 roster-restart scenario requires.
2026-07-22 23:44:04 -07:00
Patrick Buckley af918c321c docs(interactive): correct #890 gate refs + rule the heal's fire-and-forget
Addresses the Copilot review of #895 (docs/comments only, no behavior change):

- recovery_e2e.py / _sse_recovery_server.py: the mutating affordance gate
  is `busy || _historyStale`, not the superseded `busy || _replayQueue`
  quiesce gate the r3 latch replaced — corrected both docstrings (E2 now
  matches E3).
- interactive.js cross-ws supersession: the branch drops the pending edit
  and releases busy but does NOT clear `_historyStale` (its sole clear
  site is replayHistory) — reworded so it no longer implies the latch is
  released.
- interactive.js idle-edge backstop + bounded retry: documented that the
  fire-and-forget `_refetchHistory` (no `.catch`) is deliberate — no
  composer state to un-strand there, unlike the primary clear_ui caller,
  so a render throw stays loud (peer of the load path's `.finally`).
2026-07-22 17:32:25 -07:00
Patrick Buckley 74eedff1a8 test(e2e): fault-injection knobs + five recovery scenarios for the /history failure paths
RecoveryServer grows an in-process fault layer (pure-ASGI wrapper; the
production app is untouched): fail_history(count) serves minimal 500s
for the next N GET /history requests, delay_history(ms) holds responses
to widen or hold open a refetch window, and per-route request counters
(history_requests, rewind_requests) let scenarios assert backend state
rather than scripted absence.

Five scenarios on that layer, all stamping RECOVERY-READY/FAILED
titles like their siblings:

- fail-refetch: hide mid-turn -> restart -> failed first resync ->
  the stale transcript survives (no wipe, no empty-state) while the
  truncation record stays armed -> the connect-chokepoint retry heals
  (history_requests proves the re-fetch). The #890 acceptance
  contract, browser-observed end to end.
- stale-ref-reload: mid-segment transport death -> turn completes
  during the outage -> failed unarmed same-ws reload -> the next
  turn renders in a FRESH bubble and the stale bubble's text is
  unchanged (regression test for the resumability-gated ref reset).
- rewind-window: a second rewind clicked during a held clear_ui
  refetch window never reaches the server (rewind_requests == 1) and
  the transcript reflects one rewind (regression test for the
  busy-or-latch affordance gate, in-window arm).
- rewind-failed-window: the failed-fetch AFTERMATH sibling — the
  refetch 500s, the staleness latch keeps the gate closed over the
  stale rows (rewind_requests stuck at 1, proven latch-not-quiesce
  via a settle-poll), the bounded turn-free retry heals (3 -> 1 user
  rows), and only then does the gate reopen (rewind_requests == 2).

Negative-control validated: with the interactive.js fixes reverted,
stale-ref-reload stamps fresh0-unchanged0 (the concatenation bug),
rewind-window stamps posts2-rows0 (the in-window over-rewind), and
rewind-failed-window stamps closed2-rows0 (the failed-exit
over-rewind) — every detector observes its bug, then stamps READY
again with the fixes restored.
2026-07-22 17:32:25 -07:00
Patrick Buckley 43561c9b08 test(sse): end-to-end recovery harness (server-contract + browser livepass)
Tier 1 (tests/test_sse_recovery_e2e.py, opt-in e2e_recovery marker): six
scenarios against a real interactive server with a scripted provider and
ephemeral DBs — storm batching without loss, slow-consumer overflow with
lossless ring replay, mid-run truncation with cursor-adoption rebuild,
restart truncated-honesty (exact lost_count; no-loss variant replay_ok),
failed-resync retry via the truncation record, and sub-agent storm
attribution. BrowserlikeSSEClient (tests/_sse_recovery_helpers.py)
implements the browser cursor contract; RecoveryServer
(tests/_sse_recovery_server.py) boots the real app per test.

Tier 2 (scripts/recovery_e2e.py): the livepass idiom against a REAL node
— boots the real InteractivePane over real EventSource/authFetch, with a
dependency-free CDP runner driving the storm and hide-restart-show
scenarios; document.title stamps verdicts so a broken state cannot pass
silently.

Events are produced by the real session engine through the provider
boundary — no synthetic frames; teardown leaves no leaked threads; the
default suite keeps these deselected.
2026-07-20 22:38:32 -07:00