mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
main
12 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
480a1426b3 |
Fail-closed history-commit handoff (#1005)
* fix(session): fail-closed history-commit handoff (#981) The deleted-workstream discovery is now a terminal, ws_id-keyed latch: keyed conversation commits refuse admission once the durable parent is gone (convergence finalizers and force-abandon are exempt), history handoff refuses to mint a proof token so /history fails closed with a 503 instead of silently wiping the pane, and the SSE stream carries a workstream_gone resync reason. Discarded commits leave a forensic log of commit keys and roles, never content. Conversation rows gain a commit_key (migration 071): keyed saves are idempotent under retry, validated against the full commit identity, and refused when they would cross a workstream deletion. The prune orphan category now requires a NULL alias plus a two-hour updated grace, with cutoffs computed at discovery time and carried into both dialects' rechecks. The mid-turn interjection queue is owner-partitioned with no per-site mode flags: pops take the acting principal's and unowned rows, other participants' rows are structurally retained, and enforcement lives at queue admission plus the shared before_spawn gates. The retraction ledger is bounded by open pop windows: pops open a window atomically with the queue delete, restores close their ids atomically with the ledger consume, every other exit closes through one helper, and misses for unheld ids record nothing. The workstream-gone latch refuses unattended wakes at all three gates (watcher spawn, claim, delivery pre-pop), and the retry dispatcher regained its pre-envelope cancel/error convergence net. Persistence-state reporting derives through the session bound to each UI instead of a registry lookup by id that failed open to healthy during tombstone retention. The dashboard roster no longer re-inserts ghost entries from trailing activity events, the history tool-outcome scan tolerates interleaved non-turn rows, and the shared handoff-deadline handle owns its own retirement. Single-sourced across call sites: keyed-commit row values, attachment save wrappers, tail-truncation and conflict-resolution bodies for both storage dialects; worker-slot lifecycle field sets; the direct-commit admission frame; queued-row layout accessors; the string-aware comment stripper shared by every JS harness suite. Refs #981 #964 * fix(session): sweep handoff fixes to their sibling surfaces The interactive replay loop treated a system row as a tool-batch boundary, so every tool result after an interleaved row vanished from that pane while the coordinator rendered the same history correctly. Only a conversational turn ends the batch window now, matching the shared outcome index. Accepted user turns clear the composer's attachment chips on the same viewer policy that settles optimistic bubbles rather than on having matched a local bubble, so a workstream created with an upload no longer keeps a chip for an attachment the create dispatch already consumed. The coordinator's raced-Stop arm emits the stream-end hook it inherits alongside the idle state, leaving no unfinalized bubble or unflushed tool output. Ending a session surfaces a failure toast when the request never lands or answers with a non-JSON body. The per-second persistence reconcile now probes each session without blocking: a workstream whose generation and handoff locks are held is skipped until the next pass instead of contending the locks every commit needs. The one-shot repair that gates workstream creation at capacity keeps a definite probe — it has no next pass, and the sessions likeliest to be contended are the ones whose unresolved journals emptied its candidate list. Single-sourced: the attachment lane builds its conversation row through the shared commit-identity builder; the ordinary worker exit releases its slot through the lifecycle owner; both operator surfaces snapshot their counters through one non-consuming helper; the replay preamble loses its per-kind wrappers and its config hook; the browser harness suites share one brace walker; and each in-flight history attempt is one record carrying both its abort controller and its deadline. Refs #981 #964 |
||
|
|
7ed5d90a98 |
test(e2e): E8 observes the double render; honest non-vacuity for E6 (#900 r2)
Until now nothing in this harness could see the artefact the campaign prevents. Every scenario counts .msg.user rows, and user rows never travel on the SSE stream — a /send emits none, only /history replay paints them — so a duplicated assistant bubble was invisible to all of them. E8 counts a sentinel's occurrences in the transcript text instead, which is structure-agnostic across duplicate bubbles and tool blocks. The window it drives is the one readyState cannot see: the retry fires with the transport OPEN, its /history is held, and inside that await the transport drops and re-establishes. readyState reads OPEN afterwards exactly as before. The redial is a real disconnect+connect rather than a visibility change on purpose — a hide leaves evtSource null, which the presence term already decides, so a hide-based control would pass for the wrong reason. Control: stripping the generation term stamps dupes2. E8's expectation needed correcting once: unlike E6, the stream is live at flush time here, so the declined render's queued settle fires the transport-free backstop and the pane converges in one settle. That is correct behaviour, so the heal is asserted as a convergence leg — a "no duplicates" verdict must not be earnable by rendering nothing. E6 gains the non-vacuity it was missing: history_requests counts on ARRIVAL, before the hold and before any status is chosen, so "the gate declined a good payload" and "there was no good payload" stamped identical observables. history_ok — incremented only when the production route answers 200 — closes that, including the production-side-failure hole an injected-fail budget cannot see. Both hidden-window detectors widened for the additive jitter: sized on the 2000 floor alone they would have closed before a top-of-range firing and reported hidden0 for the wrong reason. delay_history(0) comments corrected — it cannot release an in-flight hold. |
||
|
|
60f6dc07a2 |
fix(#894): drop the unreachable epoch guard; abort-Set; bump-after-delete; producer pins
Review round 9 (4 minor bug, 4 quality, 1 perf nit; security zero). - The r8 clearUiEpoch guard was UNREACHABLE (r9 bug find): clear_ui always dispatches immediately after bumping, so a stale-epoch dispatch is also a stale-seq dispatch and the currency gate discards it before it can paint or clear — the client half of the joined- flight fix was already carried by seq, and the server generation key is the sole load-bearing layer. Machinery removed (decl, bump, capture, conditional clear, section-9 pins); the latch-clear comment now states the two-layer accounting. - destroy()'s abort handle becomes a Set: a newest-wins single slot, nulled by the newer dispatch's finally, left an OLDER overlapping fetch unabortable — the destroyed closure pinned for the bound's remainder. Pinned. - _history_generation now bumps AFTER delete_messages_after: flights rebuild from storage, so old-generation-reads-post-delete is the harmless spuriously-fresh direction while new-generation-reads- pre-delete would be wrongly joinable; the count/floor error paths correctly leave it unbumped. Two-arm producer pin in test_rewind_retry (persisted-rows bump on rewind AND retry; in-memory-only error path must NOT bump) — the flight test's mock can no longer mask a deleted bump. - The harness load_calls increment takes a lock (to_thread workers genuinely overlap under delay_load; a lost update false-fails G7). - G7's viewer B is now a background authenticated GET (a raw request enters load_messages identically; the second browser bought no proof); stale two-tuple key comments and the coalescing matrix line updated; the _send_in_page enumeration dropped for prose. 250 pins green; G1/G6/G7 re-run READY. |
||
|
|
b85f792925 |
test(e2e): G7 joined-flight detector at the flight layer; fix the generation read path it caught (#894 r8)
G7: two browsers on one ws; delay_load parks B's pre-rewind /history flight open INSIDE load_messages — the flight layer. (A first cut held via delay_history, which sleeps in the FAULT layer before the route: flights never overlapped there and the 'negative control' passed vacuously — a false detector, caught and rebuilt. The knob also sleeps AFTER the load so a parked flight holds the rows it actually read: its transaction point.) A rewinds mid-hold; the miss proof is load_calls growing TWO (a joined request never enters load_messages — the e2e twin of the unit test's proof) plus A rendering the post-rewind single row. The rebuilt detector immediately caught a real bug in the server fix: mgr.get returns the Workstream WRAPPER, and the route's direct getattr for _history_generation silently defaulted to 0 forever — joining stayed enabled while the unit test's mock (attr on the wrong object) masked the shape. The route now reads ws.session, and the mock pins the nested shape so a wrong-object read can never pass again. Negative control (flight key reverted to (ws_id, limit)): stamps FAILED-loads1-rows3 — A joins the pre-rewind flight and paints three stale rows as fresh truth. Fixed: READY-posts1-loads2-rows1. |
||
|
|
30b6ff7f7b |
fix(#894): rewind-freshness epoch closes the #884 joined-flight window; destroy aborts the bounded fetch
Review round 8 (2 major + 2 minor bug, 2 major + 3 small quality; security/perf zero at five consecutive rounds). - clearUiEpoch (r8 major): the #884 /history single-flight can hand a joiner a payload whose load_messages ran BEFORE the rewind committed (the flight key is (ws_id, limit); joining is invisible to the client, and the client seq stamp cannot see server-side staleness) — reachable single-user (rewind clicked during a truncated-resync fetch joins that flight) and multi-viewer (any concurrent pane's /history). The joined payload rendered as 'success' and CLEARED the latch: the original over-rewind window, resurrected through the server seam. Fix: the epoch bumps at clear_ui arrival, every dispatch captures it pre-await, and only a dispatch that post-dates the latest clear_ui may CLEAR the latch — a pre-rewind payload may still paint (stale-but-real posture, gate holds), the surviving latch arms the retry, and the retry's fresh dispatch starts a new flight with post-rewind truth. Producer/consumer/placement pinned. - destroy() aborts the in-flight bounded fetch (activeHistCtrl): the r7 15s bound alone pinned a destroyed pane's closure until it fired — the same dead-not-inert ruling destroy applies to staleRetryTimer. Pinned. - stop(hard=True) no longer sets force_exit: it skipped the ASGI lifespan teardown and leaked the #885 daemon threads + sse_executor. The 2s graceful-shutdown timeout already force-closes open SSE, and the lifespan runs on both paths (docstring corrected; G6 re-verified — the orphan still manifests). - G6 pacing sized above the scenario's worst-case deadline sum (~200s vs ~95s) so the in-process bash cannot resolve the orphan mid-scenario and degrade the detector to a false READY. - Quality: the r6 reachability comments rewritten to the r7 truth (orphan REAL via hard crash; graceful-close-only synthesis); the bound's WIRING pinned (signal reaches getJSON; getJSON forwards init); _strip_comments deduped (4 inline copies); seq comment re-paired with its asserts; retry_fire window tail-anchored. Full harness (16/16 scenarios) + full suite (9724) green on the prior commit; 136 pins green here. |
||
|
|
412161aa4a |
fix(#894): live-set retirement policy — transport death is not retirement; bound the refetch await; G6 hard-kill detector
Review round 7 (2 major + 1 minor bug, 4 minor quality; security/perf zero). Both majors traced the r6 stratum: - The closeStreamTransport drain of liveToolCalls rested on a false re-announcement premise (verified: replay_ok yields only events past the cursor; the coord fresh/truncated replay yields connected/status/ pending-cards/verdicts, never tool_pending/tool_info). An emptied set fails OPEN — a mid-batch redial plus a slow seedless refetch wiped the live batch. Retirement policy re-derived at the decl: an id leaves on its RESULT, at the SETTLE edge, or with pane death; transport death is NOT a retirement event; a stale id fails CLOSED (skip, latch survives, settle heals). Site-anchored pins: the one drain inside the idle/error block, the delete inside tool_result, the adds inside tool_pending/tool_info, and closeStreamTransport's comment-stripped code may not touch the set. - G6's kill was not a kill: RecoveryServer.stop() gracefully closed workstreams, and session.cancel()'s bash path persisted 'Cancelled by user' BEFORE the reboot — the r6 'recovery synthesizes' ruling was observing the cancel path. stop(hard=True) (skip the close sweep + uvicorn force_exit: a crash does not drain SSE) leaves the orphan genuinely unresulted — REACHABILITY FLIPS: the poisoned-pane state is real, the live-set hardening is reachably load-bearing, and G6 is now its behavioral detector: hard kill -> reload paints the orphan (asserted PRESENT) -> the seedless rewind renders THROUGH the residue (rewind-for-retry truth: the user message stays), negative-controlled against the DOM-probe encoding (stamps orphan1, hist2). Discovered and tracked separately: a hard-crashed reborn node answers stale-high cursors with a silent fresh stream (no replay_truncated — the honest truncation signal rides gracefully-persisted state). - refetchHistory's await is now bounded (AbortController + 15s, the coordSend shape): an accepted-never-answered /history pinned refetchesInFlight and permanently disabled both heals. Pinned. Quality: seq-producer position pinned earlier; the stale section-7 comment corrected; retry/backstop guard windows comment-stripped (vacuous-by-comment-mention foreclosed); the shared G3/G4 double-fail prologue extracted into _coord_stick_latch. G3/G4/G6 re-run READY. |
||
|
|
5386d598ef |
test(e2e): native-transport roster scenario (F2) + strict absence assertions (#881)
Scenario F splits into F1 (manual ?last_event_id= transport) and F2, the native-header sibling — the only behavioral coverage of two pure-browser semantics no Tier-1 harness can express: the auto-reconnect header echo, and id-less frames inheriting the connection's persisted lastEventId (the mechanism behind app.js's node_snapshot-branch clear). F2's phase C is the round-3 fix's discriminator: after the native heal, a forced manual reconnect must go CURSORLESS with no second truncated round (pre-fix: cursor1-trunc2), guarded by an idFrames precondition against the aggregate tick. Two harness seams earned by F2's first failures, both documented at site: a failed EventSource reconnect attempt is TERMINAL per WHATWG, so the restart must never expose a refused window — a SO_REUSEPORT placeholder binds before the old node stops and hands its backlog to the successor's uvicorn (make_listen_socket + RecoveryServer sock injection); and an SSE stream still open at stop() parked uvicorn's graceful drain indefinitely — timeout_graceful_shutdown=2 bounds it with the #885 lifespan teardown intact. Absence assertions tightened (round-4 review): ghost-gone now requires absence from BOTH the model and the rail via _roster_absent_ws — the negated AND-membership helper De Morganed into either-surface and could false-pass a rail-render regression. |
||
|
|
ea706bf4c0 |
test(e2e): roster-restart scenario proves the global boot-epoch heal (#881)
Scenario F drives the REAL node dashboard (/ + app.js) through a node restart on the global stream — no custom page; transport instrumentation is injected via CDP addScriptToEvaluateOnNewDocument, scoped to /events/global URLs so per-ws streams can't pollute the counters. Phase A is the negative control: live roster, live cursor, zero replay_truncated. Phase B: hide, force the CLOSED state (a closed EventSource never auto-retries, making the show edge's manual reconnect the only reconnect), restart the node re-opening only one of two workstreams, show. Asserted: cursor presented via ?last_event_id= and replay_truncated observed at the transport, the not-reopened workstream's ghost evicted from the roster model and rail (the dashboard table's membership refreshes on interaction by design — documented at _roster_has_ws), and the reborn node's global_events_requests counter proves the reconnect hit the real endpoint. The native header transport differs only in carriage and is pinned by the Tier-1 boot-epoch tests. |
||
|
|
a935ae3106 |
fix(server): give the lifespan daemon threads a real shutdown (#885)
_global_fanout_thread, _aggregate_emitter_thread, and _idle_cleanup_thread were daemon threads with no stop signal — shutdown abandoned them mid-loop. The sleep-loop pair now waits on a shared Event (wait doubles as the tick sleep, so a set wakes them immediately); the fanout exits on an identity-checked queue sentinel, FIFO-draining everything enqueued before it (sessions close earlier in the shutdown tail, so their final events still fan out). Joins are bounded and off-loop; daemon=True stays as the backstop for a join timeout, not the mechanism. The recovery harness drops its thread-neutering workaround (module docstring piece 4) — the global lane now runs REAL in harness boots, which the #881 roster-restart scenario requires. |
||
|
|
af918c321c |
docs(interactive): correct #890 gate refs + rule the heal's fire-and-forget
Addresses the Copilot review of #895 (docs/comments only, no behavior change): - recovery_e2e.py / _sse_recovery_server.py: the mutating affordance gate is `busy || _historyStale`, not the superseded `busy || _replayQueue` quiesce gate the r3 latch replaced — corrected both docstrings (E2 now matches E3). - interactive.js cross-ws supersession: the branch drops the pending edit and releases busy but does NOT clear `_historyStale` (its sole clear site is replayHistory) — reworded so it no longer implies the latch is released. - interactive.js idle-edge backstop + bounded retry: documented that the fire-and-forget `_refetchHistory` (no `.catch`) is deliberate — no composer state to un-strand there, unlike the primary clear_ui caller, so a render throw stays loud (peer of the load path's `.finally`). |
||
|
|
74eedff1a8 |
test(e2e): fault-injection knobs + five recovery scenarios for the /history failure paths
RecoveryServer grows an in-process fault layer (pure-ASGI wrapper; the production app is untouched): fail_history(count) serves minimal 500s for the next N GET /history requests, delay_history(ms) holds responses to widen or hold open a refetch window, and per-route request counters (history_requests, rewind_requests) let scenarios assert backend state rather than scripted absence. Five scenarios on that layer, all stamping RECOVERY-READY/FAILED titles like their siblings: - fail-refetch: hide mid-turn -> restart -> failed first resync -> the stale transcript survives (no wipe, no empty-state) while the truncation record stays armed -> the connect-chokepoint retry heals (history_requests proves the re-fetch). The #890 acceptance contract, browser-observed end to end. - stale-ref-reload: mid-segment transport death -> turn completes during the outage -> failed unarmed same-ws reload -> the next turn renders in a FRESH bubble and the stale bubble's text is unchanged (regression test for the resumability-gated ref reset). - rewind-window: a second rewind clicked during a held clear_ui refetch window never reaches the server (rewind_requests == 1) and the transcript reflects one rewind (regression test for the busy-or-latch affordance gate, in-window arm). - rewind-failed-window: the failed-fetch AFTERMATH sibling — the refetch 500s, the staleness latch keeps the gate closed over the stale rows (rewind_requests stuck at 1, proven latch-not-quiesce via a settle-poll), the bounded turn-free retry heals (3 -> 1 user rows), and only then does the gate reopen (rewind_requests == 2). Negative-control validated: with the interactive.js fixes reverted, stale-ref-reload stamps fresh0-unchanged0 (the concatenation bug), rewind-window stamps posts2-rows0 (the in-window over-rewind), and rewind-failed-window stamps closed2-rows0 (the failed-exit over-rewind) — every detector observes its bug, then stamps READY again with the fixes restored. |
||
|
|
43561c9b08 |
test(sse): end-to-end recovery harness (server-contract + browser livepass)
Tier 1 (tests/test_sse_recovery_e2e.py, opt-in e2e_recovery marker): six scenarios against a real interactive server with a scripted provider and ephemeral DBs — storm batching without loss, slow-consumer overflow with lossless ring replay, mid-run truncation with cursor-adoption rebuild, restart truncated-honesty (exact lost_count; no-loss variant replay_ok), failed-resync retry via the truncation record, and sub-agent storm attribution. BrowserlikeSSEClient (tests/_sse_recovery_helpers.py) implements the browser cursor contract; RecoveryServer (tests/_sse_recovery_server.py) boots the real app per test. Tier 2 (scripts/recovery_e2e.py): the livepass idiom against a REAL node — boots the real InteractivePane over real EventSource/authFetch, with a dependency-free CDP runner driving the storm and hide-restart-show scenarios; document.title stamps verdicts so a broken state cannot pass silently. Events are produced by the real session engine through the provider boundary — no synthetic frames; teardown leaves no leaked threads; the default suite keeps these deselected. |