mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
5de54147e17bcb57a18eef6cb12c5a522c23e967
1028 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5de54147e1 |
fix(832): a prep fault walks the fallbacks it can no longer speak for
Making prepare_wire lane-variant invalidated the premise behind the walk-abort on WirePreparationError: with the fold posture following each lane's capabilities, a preparation fault on one lane no longer implies every lane fails, so aborting the walk skipped healthy fallbacks and the dedicated fatal message was wrong on both of its claims. Preparation faults now keep their no-health rule on every lane but continue the walk — the primary's fault enters it and a fallback's fault yields to the next alias — and the fatal message drops the no-fallback claim. Riding cleanup: the self-surfacing exception pair gets one spelling for the re-issue mask (_SELF_SURFACING_ERRORS; the walk arms stay per-class because auth aborts where prep continues); the tag-scan gate gains a capabilities-shaped form (caps_scan_inline_reasoning) that the lane form delegates to and the title peel now uses, retiring the third spelling; the three streaming provider fakes build on one provider_shell; a comment in session_ui_base names the module function that replaced the deleted session delegate; close_run spells its carry cut as removesuffix; and the prepare_wire docstring paragraph is re-flowed. The walk-continues and per-lane no-health pins are mutation-probed. |
||
|
|
c906776efd |
fix(832): the serving lane's capabilities reach the wire fold
The per-attempt prepare_wire closure folded mid-conversation system turns with the PRIMARY binding's capabilities on every lane, so a fallback whose chat template rejects non-leading system roles failed on the self-inflicted wire shape and burned its own health record — the wrong-dialect class the walk's binding snapshot guards against elsewhere. model_turn now passes the serving lane to prepare_wire, and the session's closure folds with that lane's capabilities; callers without a lane in hand (the token-table re-fold) keep the primary default. Pre-fold prepared once with primary caps for every lane, so this is a named improvement, not a parity break. The arm-duties hook rode the same unguarded two-statement supersession window the _CancelRef docstring accepts only for the stream register: a force-cancel claiming a new generation between the superseded read and the hook let an orphan's late registration null the successor's usage slots and record spurious creation health. on_stream_armed now generation-gates itself, shrinking the accepted window's harm back to the register-only class. Test hygiene: the two overflow-compact tests are one parametrized body; arm_session mints a fresh ArmedHandle per create (provider.handles, _armed_handle = latest) matching the one-handle-per-create rule of real adapters. The duplicate sanitize pass stands as designed (accepted for wire parity); its perf note rides #979. All three product fixes are mutation-probed. |
||
|
|
90e55f92ca |
docs(832): shorten the branch's comments to their constraints
Comment-only sweep over the diff's prose: origin archaeology, next-line narration, and review-thread talk go; each surviving comment states the constraint the code cannot show, re-wrapped to the file's width. The ruled-behavior restatements in the parity transforms and the contract docstrings (eager append, cancel-predicate pairing, carry ownership, the plant call's carve-outs) keep every named invariant. |
||
|
|
06ec1a8629 |
fix(832): the boundary carry belongs to the run owner
The mandated cross-lane interleave angle found the two residual holes in the reasoning-boundary close: the close was gated on not-in_think, so an open inline think block at the boundary never closed and the later state flip relabeled held chain-of-thought as displayed ANSWER text; and the carry parked in the splitter's own pending was re-read under whatever state later flushes hit, relabeling a content-state tail as reasoning. close_run() now closes unconditionally (as the drain does) and RETURNS the partial-tag tail; the consumer owns the carry in a state-immune slot mirroring the drain's separate variable — re-fed when content resumes so a split tag still reassembles, flushed as content at tool, finish, and cancel boundaries, and included in the partial-content rule. The trailing citations footer is now HELD and folded once at stream end over the full answer — structurally the drain's post-loop fold — instead of folding at arrival, which diverged from the commit whenever a lax gateway emitted content after finish. Two non-mirror fixes: the fallback-failure UI line carries the exception class only (its text can embed a credential-bearing base_url; detail goes to the server log, same rule as the re-issue log arm), and a never-armed Stop (creation window, no prior death, zero tokens) writes NO assistant row again — restoring pre-fold semantics; a marker-only row would replay to the model as context on every later turn. Armed zero-token Stops still record their marker. Hygiene riding along: the parity runner zeroes the ladder backoff (the exhaust scenario was sleeping 3.2s of real backoff per suite run, with the retry-notice transform strings updated in step); test_session's porting docstring points at the helper's real module; test_cancel and test_session wrap the shared session factory instead of re-implementing its defaults; arm_session's armed handle is an ArmedHandle with real closed state instead of a MagicMock that satisfies any assertion; and send() derives the tool-call list once for both the persisted mirror and the executed set. All fixes are mutation-probed: re-gating the close, discarding the carry, dropping the promote gate, unredacting the fallback line, and restoring the arrival-time fold each fail their pins. |
||
|
|
6212783e23 |
fix(832): close the content run at a reasoning_delta boundary — display must mirror the drain
Live-caught on a deployed review exercise: the consumer's reasoning_delta arm flipped the splitter's in_think with a buffered content tail still pending, so a flush while in-think (stream finish, tool boundary) relabeled that tail as reasoning. The drain closes each content run at the same boundary, so the committed turn kept the tail as content — display and commit diverged. Worst case: a short answer followed by trailing reasoning displayed as NOTHING while the commit carried the answer plus its citations footer (the display-side blankness gate saw empty content and dropped the footer too). Pre-fold, display and commit came from one continuous splitter and both lost the tail; the fold's drain corrected the commit, leaving the display behind. ThinkTagSplitter.close_run() now closes the run exactly as the drain does — decided text emits at the current state, only a possible partial-tag tail carries into the next run — and the consumer calls it before entering the reasoning phase. This also heals the cancelled- partial rule in the same window, and covers the content-reasoning-tool sequence interleaved-thinking lanes emit. Riding contract fix: partial_tag_tail required only startswith, so a complete <reasoning>/<think> self-matched as a "partial" tail and the drain carried a finished open tag across the run boundary, relabeling the next run. A partial tag is now a PROPER prefix, per the function's own documented contract. Pins: TestDisplayCommitMirror (displayed content must equal committed content across six reasoning-interleave scenarios — the combination the replay-parity grid never scripted), TestPartialTagTail contract rows, TestCloseRun unit pins, and three new interleave rows in the splitter CASES table. Both fixes are mutation-probed: disabling close_run or restoring the self-match fails the pins. |
||
|
|
aa4371ea99 |
fix(832): retire the dead attempt's armed state in the re-create window
Between a mid-stream death and the next begin_attempt there is no live attempt, but the consumer kept the dead attempt's armed _CancelRef: a Stop in that window re-emitted the discarded splitter carry as fresh content behind a duplicate stream_end, and a walk-preamble failure was classified as another armed death, replacing the operator-actionable stream-death error. end_attempt() now pronounces the attempt dead at partial-capture; the consumer gains a single per-attempt initializer (_reset_attempt), a lane-free constructor (one resolve_lane walk per turn), and a saw-chunk classifier fallback so a never-arming adapter's mid-stream death still classifies mid-stream instead of silently double-rendering the same lane. Wire-preparation failures are typed at the seam: model_turn wraps prepare_wire raises in WirePreparationError, both walk arms forward it verbatim (no health record, no fallback walk — a session-data fault would otherwise paint every backend degraded), the fatal formatter gets a dedicated branch, and the re-issue ladder's last-death mask exempts it alongside BackendAuthUnavailableError so an auth outage mid-turn is not misdiagnosed as a network flap. Riding fixes: the tag-scan gate gets its single spelling (lane_scans_inline_reasoning) shared by drain and display; the citations fold's separator+gate become a shared pair in _protocol; _build_main_lane stops passing config_store (dead derivation — the session's own knobs replace both values it feeds); the debug wire dump is ruled per-invocation (the overflow-recovery re-print is the dump that diagnoses the recovery) and pinned; dead delegates _ensure_tool_call_ids and _finalize_provider_blocks deleted; the parity runner adapts to the pre-fold seam signature by inspection and refuses to record a harness-shape TypeError as a baseline; the streaming provider fakes move to tests/_session_helpers (their tree-wide home) and test_cancel's duplicate helper is deleted; committed parity pins restate their rulings in full; architecture.md's circuit-breaker section is replaced by the real passive health-tracker story and the send-flow diagram stops attributing tool-call assembly to the display consumer; stale pre-fold names and ragged comment paragraphs cleaned. New pins are mutation-probed: disabling end_attempt, the saw-chunk fallback, the auth exemption, or the WirePreparationError arm each fails its pin. |
||
|
|
5a20916eec |
docs+test(832): retire pre-fold names from prose; pin the eager-append contract
The docs sweep re-points every stale reference to the deleted seam (architecture.md's flow diagram and ladder inventory, the lowering and anthropic docstrings, the protocol's shared-rule docstrings that described the pre-fold dual-assembler world). The Protocol's cancel_ref contract is strengthened from 'before the first chunk' to 'inside the call body, before the iterator is returned' — the instant the fold's creation-vs-midstream classifier and health recording key on — and three real-SDK-over-mock-transport tripwires pin it per adapter, so a future lazily-issued generator adapter fails loudly instead of silently reclassifying every pre-first-chunk death. |
||
|
|
e103af94e7 |
test(832): port the seam-coupled suites to the folded architecture
Seventeen files, ~1,300 tests, re-pointed or redesigned per the triage ledger's recipes: wholesale turn-scripting moves to ModelTurnResult fakes; streaming-behavior suites drive the REAL wrapper+consumer+drain path through armed provider fakes (tests/_parity_832.arm_session — the eager cancel_ref append every real adapter performs, exception elements for creation-phase failures, sequential per-turn scripts, and the title lane quieted: a provider-level fake otherwise loses its one-shot script to best-effort title generation, which is why the old tests patched at the session level); kwarg-capture suites assert through model_turn's create_streaming call with system-prepend-aware index math; delegate wrappers retired by the fold re-aim at their model_turn module twins. Old-architecture pins are replaced by their new-world equivalents rather than deleted: no shared cancel ref exists (pinned), the handle slot and per-attempt refs carry the cancel surface, the retry gate reads the serving lane's provider, and a superseded generation's death exits send silently as cancelled — a named delta: no arbitrary exception class escapes an orphaned thread anymore. Full suite: 10651 passed, 10 skipped. The wire-payload goldens pass untouched — the fold's lowering composition is byte-equivalent on every provider's request path, as designed. |
||
|
|
58b24de7e6 | test(832): re-aim the extra-params gate pin at the module function (delegate wrapper retired) | ||
|
|
0df9f6e2d4 |
test(832): port test_cancel to the folded seam; add hook + orphan + pre-dispatch pins
Provider-level armed fakes drive the REAL wrapper/consumer/drain path (the seam these tests exist to pin), with title generation quieted — the best-effort title lane consumed one-shot scripts once fakes moved to the provider level. The shared-ref architecture pins become their new-world equivalents (no shared _cancel_ref attribute; _cancel_stream lifecycle via the eager append), and three new pin classes land: on_first_append fires once and never for a superseded arrival; a force-cancelled generation's mid-stream death is never re-issued and touches no UI finalize; a pre-set Stop issues no request and mints no credential on a dynamically authenticated alias. |
||
|
|
ecf14dc001 | test(832): make_result helper for the triage's patched-result recipe | ||
|
|
2e18d159a3 |
feat(session): fold the main streaming loop onto model_turn (#832)
The send path's plant call is now one model_turn invocation per attempt, reached through a lane-swap fallback walk that mirrors the old creation ladder 1:1: an inner per-lane retry (_model_turn_with_retry) inside the two-pass healthy/degraded walk (_model_turn_with_fallback), with health success recorded at the request-accepted instant via the per-attempt _CancelRef's new on_first_append hook and failure once per lane ladder. The hook is also the creation-vs-midstream classifier: an armed attempt's death re-raises to the re-issue ladder on every lane — a fallback stream that died after tokens reached the UI is never swallowed into try-the-next-alias — and carries the per-turn usage-slot resets at the old timing so a reconnecting tab's status bar never blanks mid-walk. Chunk-to-UI translation lives in _StreamTurnConsumer (model_turn's on_chunk body): display-side only, the canonical turn always assembled by drain_stream at the one seam; the inline-tag scan reads the SAME lane capability the drain gate reads (server_parses_reasoning), replacing the creation-time handoff register — which is deleted — so display and commit cannot disagree about a backend's posture, fallback walk included. Cancellation converges: every model-call site now builds fresh generation-scoped refs, closing the force-cancel hole where the old gen-0 shared ref read aborted=False for an orphaned generation and would have let a retry re-issue on its behalf; the pre-dispatch abort read inside model_turn also means a Stop set before the turn no longer mints a credential on a dynamically authenticated alias. send() consumes the result natively: the committed Turn carries minted ids, the finalized native lane, and an accurate producer — fixing the latent mislabel where fallback-served turns were persisted under the primary provider's name, and the fork asymmetry where in-memory turns decoded with producer="". Ruled behavior changes (design D12): the trailing citations footer now folds into committed content (it previously lived only in an ephemeral info bubble and vanished on reload); a stream that exhausts without a finish reason is a retryable mid-stream death instead of a silent partial commit; length-truncated turns keep dropping partial tool calls, now as an explicit post-drain policy. The replay parity harness pins all thirteen scenarios against pre-fold baselines, transformed only where a ruling applies — and caught two real bugs during the fold (the splitter's end-of-stream carry never flushing to the UI, and the footer splicing into the answer's held tail). ChatSession imports no provider module: create_streaming has exactly one caller module, and the protocol types, merge_usage, and create_provider reach the session through model_turn's re-export seam. |
||
|
|
1f3b89610a |
test(832): replay-parity harness + pre-fold baselines
Thirteen scenario scripts drawn from the chunk-field-to-UI grid, each driven through the streaming seam against a scripted provider fake that arms cancel_ref eagerly (the classifier the fold introduces distinguishes creation-vs-midstream failures by that arming, so the fake must mirror the real adapters' eager append). The captured records — ordered UI events, committed-message projection, mid-stream usage, raised class — are the OLD-WORLD baselines: this commit's session.py is byte-identical to main, which is what makes them the record. The assert path applies only the behavior deltas the design table rules, each transform citing its row; a difference outside a ruled transform is a fold regression. |
||
|
|
70165807c7 |
fix(reasoning): close the unmarked chain-of-thought leak, gate the tag scan by backend (#940) (#978)
Some serving setups emit model reasoning inline with no think tags and no
reasoning_content at all — nothing any parser can segregate (measured live
on the dev vLLM: 20/20 sampled completions, streamed and not, proxied and
direct). The drain seam correctly passes unmarked prose through, so it
became the artifact on every bounded-artifact lane: workstream titles
("Thinking Process:"), compaction summaries that were ~90% chain-of-
thought, and the web-fetch tool results #940 reports — which then ride
every following turn as context.
Three coordinated changes:
* Utility lanes ask for no reasoning. _utility_completion (title,
compaction, web-fetch extraction) pins the alias's declared thinking
toggle off and withholds every reasoning-effort channel — the relayed
session knob, the lane rung, the definition default, and the graded
template key — via lane_without_thinking / lane_thinking_suppressed,
the same suppression omni transcription already used (now shared as
thinking_off_template_kwargs). Measured end-to-end: the extraction
that returned 3.7k chars of reasoning returns a 258-char answer.
* server_parses_reasoning capability. A backend that segregates
reasoning into its own channel declares it, and the inline tag scan
turns off on every lane: the drain seam, the interactive splitter
(which now reads the ACTIVE stream's capabilities via the creation-
time handoff register, never the primary alias's), and the title
lane's cosmetic peel — so prose that merely quotes a tag can no
longer be misrouted, and the utility suppression stands down where
reasoning costs the artifact nothing. The built-in commercial
capability tables declare it wholesale (known models and table-miss
defaults); local compat lanes keep the passthrough default the scan
exists for. Bool-typed capability overrides coerce string spellings
instead of truthiness-flipping on hand-edited JSON.
* Title selection follows the prompt's contract, not line position:
the last line within the word cap that ends in a word character —
rejecting explanation sentences, sign-offs, parentheticals, and
reasoning headings in any script (terminal punctuation carries
unspaced scripts where whitespace word counts are meaningless) —
else the last non-empty line. 20/20 captured live responses title
correctly (9/20 before, unchanged since well before the seam
unification: the old and new pipelines scored identically on every
sample, so the regression source was the backend's output shape,
not #965).
Also folded in from the review round: a think tag split across a
reasoning-delta boundary reassembles in the drain (partial-tag tail
carry; tool boundaries still flush), Turn.text joins text blocks with a
newline so multi-block answers stop fusing words in notification bodies
and every flattened read, the notify hook reads final_assistant_text
directly instead of through a one-line shim, web-fetch extraction uses
the shared _non_blank_or fallback, and the judge/output-guard suites use
real ModelCapabilities instead of truthy mock attributes.
Closes #940.
|
||
|
|
14df09a107 |
chore(deps): update vendored js (#974)
* chore(deps): update vendored js * chore: download vendored JS files --------- Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
7076bcf6ef |
fix(session): never dispatch a model call on an aborted cancel_ref (#972) (#976)
* fix(session): never dispatch a model call on an aborted cancel_ref (#972) model_turn consulted cancel_ref.aborted before re-issuing a request after a mid-drain transport death, but never before dispatching one. A caller whose call had already been abandoned — the user hit Stop, or a deadline fired — still lowered its turns, resolved its credentials, and put the request on the wire; the provider registered the stream handle, the ref closed it, and the client discarded a reply the endpoint had already begun producing. The rule was half-present at the seam: don't resurrect an aborted call was enforced, don't start one was not. The predicate is now read before each dispatch through one helper, using the same duck-typed getattr the drain-retry gate uses, so a None ref (perception, title generation, sub-agents, optimizer, eval) and a plain-list ref both stay legal. Two reads, because they buy different things: the entry read skips the lowering and the credential resolve for a call already abandoned when it arrives, while the read immediately before create_streaming is the one that keeps bytes off the wire — a blocking resolve is exactly the window the entry read is too early to see. Cancellation stays cooperative and the docstrings say so: a mint already under way completes, and an abort arriving after the last read still reaches the in-flight call through the ref's own close paths (append for a handle that has not arrived, abort for one that has). The raise is DeadlineCancelledError, the deadline module's abandonment vocabulary. GenerationCancelled would be invisible to the except-Exception arms surrounding these calls, and it lives in session, which imports this module; compaction performs the translation itself, its handler re-checking the session before it reads the error, which is what keeps a Stop mid-summary off the red-error path. That translation holds only while _CancelRef.aborted and _check_cancelled stay the same predicate over the same generation, now recorded on the property that owns it. The raised message deliberately avoids context-window vocabulary: _is_ctx_overflow classifies unrecognized error classes by text, and an overflow reading would send the compaction lane subdividing and re-issuing the very calls this suppresses. The pre-existing abort test keeps its subject, the re-issue gate: its ref now aborts after dispatch, and it asserts that no retry was announced rather than counting calls, which is what separates that gate from the post-backoff one. Two siblings pin the new reads — the resolver is never called for a ref aborted on arrival, and an abort landing inside the resolver still reaches no wire — and a third pins the message against the overflow classifier. * docs(session): disambiguate the abort helper's resolve wording "The credential resolve between them is NOT re-checked" reads as though no abort check follows the resolve, when the second read sits immediately after it — the sentence meant only that nothing interrupts the resolve itself. Left as-is it invites a refactor to delete that second read, which is the one that keeps bytes off the wire when the abort lands mid-mint. States both facts separately now: the mint completes regardless, and the second read is what turns such an abort into a skipped request. |
||
|
|
0150523bb9 |
test(session): pin the both-vocabulary title peel
The title lane's cosmetic peel walks the close-tag vocabularies in sequence, which review read as a double peel that could discard title text between a `</reasoning>` and a `</think>`. It cannot: the remainder of the first cut begins after the last `</think>`, so a `</reasoning>` still found in it is necessarily the later tag — the sequence is equivalent to one cut after whichever close occurs last (verified exhaustively over tag/text arrangements and 200k randomized fragment strings). The equivalence was unpinned, so both orderings join the variants table and the docstring records why the sequence is a single logical cut. |
||
|
|
bc3fa60011 |
fix(providers): segregate inline reasoning at the drain seam
Passthrough servers (parserless vLLM/llama.cpp, LM Studio, bare gateways) emit reasoning as literal <think>/<reasoning> blocks inside content, and only three of nine drained lanes stripped them: web_fetch tool results persisted raw think blocks into every following turn (#940), judge verdicts parsed through tag noise, and a draft verdict inside a think block could shadow the real one at the output guard. One rule at the seam now. drain_stream accumulates content in RUNS bounded by interleaving signals (provider-parsed reasoning deltas, tool-call deltas) with the interactive consumer's within-chunk ordering — reasoning, then content, then the tool-call close — and splits each run through split_inline_reasoning, the one-shot form of the interactive lane's ThinkTagSplitter: a pure raw split, exactly equivalent to the streaming form on every catalog case. One trim policy exists and the drain owns it: blank edge lines are trimmed once over the joined runs when a tag was consumed, so tag residue dies at the edges while genuine inter-run paragraph separators survive. Extracted text is appended to result.reasoning after any server-parsed reasoning with a blank-line boundary and rides the native lane as the reasoning_text synth block. Orphan CLOSE tags deliberately pass through byte-identical: a close whose open never arrived is indistinguishable from prose QUOTING the tag, and drained lanes routinely quote third-party text — reclassifying would let a malicious page containing the literal tag destroy the extraction that cites it. The title lane keeps a local rfind peel as display-string formatting. The citations footer folds only onto non-blank content — sourcing for an answer that does not exist is dropped rather than handed to emptiness checks as a footer-only "answer". Every private strip is deleted: the title lane's strip, the summarizer strip, _strip_reasoning itself, and the optimizer's five regexes (_strip_markdown_fence is now the one fence rule, applied to normalized model output only, never to or-fallback values). Think-only and whitespace-only responses drain to blank content, and every lane's no-answer fallback gates on blankness: web_fetch returns an honest extraction-error card, the intent judge takes the empty-retry ladder, the task-agent synthesis reports "(no output)", and the optimizer keeps the current observer system and prompt verbatim on no-answer passes. Final-say reads (optimizer analyst, eval final_content, the notify hook) use trajectory.final_assistant_text — the last assistant turn only, never an earlier narration presented as the conclusion — while last_assistant_text is the salvage walk (task_agent partial-work recovery), skipping tool-call-only, all-reasoning, and whitespace-only turns. Perception memoizes every completed description immediately, including an empty one — one perceive per key, ever — under a commit-lock guard so an empty result never overwrites a concurrently memoized real description; an all-reasoning perception model pins the placeholder until restart, and the remediation is server-side (a reasoning parser or the template thinking toggle on the perception alias). A true double-reasoning shape (inline-extracted text alongside a native reasoning block) logs chars-only at the drain, where it is distinguishable from the routine reasoning_delta mirror. The dialect's semantics are pinned as one table (tests/_reasoning_dialect.py) driven through shared fixtures (think_tag_stream, seam_provider): one-shot conformance, the exact one-shot/streaming equivalence property, the drain seam rules including quoted-tag safety, run-boundary and separator-preservation pins, per-lane pins for all nine lanes, and the empty-content assistant wire shape. Closes #965. Closes #940. |
||
|
|
1d7db73305 |
fix(models): review feedback — separator vocabulary, constraints stub, import style
The scopes sanitize now shares the registry guard's separator vocabulary: tab/newline/CR read as spaces, and every other C0 byte — including the U+001C–U+001F block str.split() would silently promote to separators — strips like the control it is, so a control byte inside a token can never split it into two valid-looking scopes (pinned alongside the registry's refusal). The livepass auth-constraints stub serves the new app_identity_auth_modes field so the pass exercises the served-data path for the model list's auth badge, and the session-module import in the mint tests drops to the string-path monkeypatch spelling (single-style imports). |
||
|
|
8605c9783d |
feat(models): rfc8693_obo auth mode, per-alias exchange scopes, identity-keyed mint cache
Adds the dedicated `rfc8693_obo` model auth mode (#955): model definitions gain an `obo_scopes` column (migration 069), the mint threads the scopes to the token-exchange leg (RFC 8693), and every dynamic mode pins its grant leg — a mode is a dialect commitment, not a hint the deployment profile resolves. Exchange-capable IdPs refuse an audience whose scope was not requested; this closes the structurally unmintable model-OBO path on token-exchange deployments. The model mint-cache is identity-keyed on the owning definition's alias (`__model_obo__:<alias>` per user, `__model_app__:<alias>` under the shared app principal), matching the MCP discipline where rows key on the unique server name. The bearer's shape lives in the row's audience/scopes columns and the freshness gate compares it on every read, so a re-aimed alias refuses its old row and overwrites the same key in place. Admin lifecycle (rename, re-aim, scope change, delete) purges a definition's own rows through one shared helper — sound because one definition owns each key; a sibling's rows are untouchable by construction. Cooldown and backoff additionally key on the dispatch shape, so an operator's config repair is an instant clean slate. Cause records, cooldowns, locks and memoization are per-alias end to end, and the session heartbeat reads refusal causes under the same keys. Console: default-deny write gating for dynamic rows (value-diff over the full column ladder, admin.mcp escalation, a never-blockable pure-disable carve-out), a two-tier validator (audience allow-list on every write; deployment-posture checks when the pair is chosen), one shared scopes parser whose omit-unchanged arm keeps over-cap DB-direct residue rows disarmable without ungating real changes, and served constraints (dynamic/scopes/app-identity mode lists, mode-to-profile pairing) so the shelf tracks the registry by data. The admin shelf gains the mode option, a scopes input with residue affordances, pairing-aware option greying, and a derived auth badge. Registry load refuses control characters in alias, audience, and scopes — including the C0 separator block that str.split() would silently collapse — and the C0/DEL class has one exported spelling shared by every surface. Profile-mismatch visibility warns at reload and boot with the mode-correct cause, gated on OIDC being enabled. Breaking: a stored `entra_obo` alias on a deployment whose `[oidc] obo_grant_profile` is `rfc8693` (or the inverse pairing) no longer mints via the profile-driven overload — the mint refuses before any IdP traffic with cause `grant_profile_mismatch`, and the `model.auth_fail_closed` policy governs static fallback. Such rows never minted usefully on scope-gating IdPs; the shelf now surfaces the pairing and the per-turn heartbeat names the refusal cause. Live-verified end to end: scoped token exchange mints, the warm cache serves with zero IdP calls, and the mode/profile mismatch refuses with zero IdP traffic (scripts/obo-e2e/keycloak_e2e.sh); the refresh-redemption profile's E1-E7 hold via scripts/obo-e2e/entra_e2e.py. Closes #955. |
||
|
|
7776cc0c2f |
fix(streaming): probe on_stream_discarded for pre-existing UIs and format the hoisted fake
PR feedback round: - on_stream_discarded now follows on_compaction's compat pattern for a hook added after UIs exist in the wild: the protocol member carries a REAL no-op default (an explicit subclass inherits a correct implementation — a UI without server-side turn buffers has nothing to truncate), and both call sites route through a getattr probe, so a duck-typed UI predating the hook degrades to no-truncate instead of raising an AttributeError from the very arm that is handling a stream death — which would replace the wire failure with the attribute error in the retry gate. Pinned with a hook-less-UI retry test. - tests/_session_helpers.py gains the formatting pass the RecordingUI hoist bypassed (the CI lint failure). |
||
|
|
1f9f462b66 |
fix(streaming): gate the dead-segment discard on the backoff surviving the Stop window
Fifth review round — four small correctness edges, none in the retry semantics: - The server-buffer discard now runs only AFTER the backoff survives a Stop: a cancel during the window persists the promoted partial to history, and the idle-state payload (drained from the turn buffer) must carry the same text — discarding first rendered the cancelled turn empty on the dashboard while the transcript had it. Pinned with a real-buffer test; the spinner and fresh segment watermark follow the truncate so a later discard cannot resurrect the dead segment. - stream.retry's dead_content_chars reports THIS death's flushed text only — the Stop-preservation carry retains the previous attempt's partial by design, and logging its length re-attributed the same discarded spend to consecutive retry lines. - The changelog entry for the post-finish-blip rename no longer claims the usage_captured field was dropped; it is emitted and pinned. - The retry suite's module docstring states the shipped finalize contract (stream_end + backoff-gated stream_discarded, never turn_committed) instead of the superseded pair. - RecordingUI is hoisted into tests/_session_helpers next to NullUI — this branch already paid the per-file-fake tax once when a protocol method grew — and a stale deferral sentence is dropped from the fatal-formatter comment. |
||
|
|
961a2017dc |
fix(streaming): delete the retry window's shared slots and gate the send epilogue
Fourth review round. The recurring defect family — cross-frame session slots racing an orphanable window — is removed structurally instead of gated again: - The wire-fold slot is deleted. The fold the stream was actually created from rides the returned message dict on the underscore lane (like _provider_content) and is popped at the single calibration site before commit, so a superseding generation can never alias it and there is nothing left to clear. Plain-dict test fakes fall through the pop to the frame-local fold. - The stream-provider slot is demoted to a creation-time handoff register: _try_stream stamps it, _stream_response copies it into a frame-local immediately after each create returns, and only that local feeds the retry gate. The fatal formatter returns to the consistent PRIMARY identity triple — pairing a fallback's provider name with the primary's base_url and alias sent operators to debug the wrong backend; stamping the full producing identity is #964. - send()'s epilogue is generation-gated: a superseded thread's escaped death no longer records a fatal error over the healthy successor turn (error banner, buffer-wiping error-state drain, wrong last_error for the coord), and a Ctrl-C on an orphan no longer mutates history. - The terminal arm discards as well as finalizes. Keeping the buffers bought nothing — the fatal path's error-state drain wipes them on every server lane — and the skipped discard let a mid-consumption overflow recovered by compact-and-retry concatenate the dead attempt's text with the recovered answer in the idle payload. Pinned with real-buffer tests for the overflow-recovery and orphan-epilogue paths. - stream.post_finish_blip regains usage_captured, tracked by transport_guarded from the chunks it forwards, restoring missing-spend attribution on both lanes. - TerminalUI.on_thinking_start is idempotent at the callee (a live spinner is stopped before being replaced), removing the caller-side stop-first dance and the leak the next unaware call site would have reintroduced. - The think-tag vocabulary in _strip_reasoning and the title lane is derived from ThinkTagSplitter, closing the drift channel that would leak raw reasoning into compaction summaries and titles. - on_stream_discarded's docstring states the true pending-batch semantics (defensive drop; the shipped sequence flushes via the preceding stream_end), and the live-suite recording fake gains the protocol method. |
||
|
|
476cce2e58 |
fix(streaming): scope stream bookkeeping to the send and discard dead segments server-side
Third review round on the retry window: two mediums fixed, one observability gap closed. - New UI-protocol method on_stream_discarded(): on_turn_committed clears only the inflight buffers — it cannot clear _ws_turn_content, the multi-segment buffer the IDLE payload drains, because earlier segments of a tool-looping turn must survive commits — so a dead attempt's text concatenated with the retried text in the dashboard's idle payload. SessionUIBase now truncates the turn buffer to a segment watermark (snapshotted in on_thinking_start, which precedes every stream segment), drops the never-displayed pending batch, and resets the inflight snapshot; the retry arm emits it in place of on_turn_committed. Server-side only — no SSE event, no client change; no-op on the CLI and eval UIs. Pinned with a real-SessionUIBase-buffer test: the recording fakes structurally cannot see this buffer. - _active_stream_provider and _active_wire_msgs are send-scoped: cleared in send()'s finally, after the except arms' fatal formatting (the one legitimate fatal-path reader of the provider field). A later fatal on a utility lane falls back to self._provider instead of wearing a stale interactive-turn binding, and the full-context-sized wire fold no longer outlives its calibration use. - stream.retry carries dead_usage and dead_content_chars: the abandoned generation's billed tokens are otherwise invisible (the wire reports usage only at stream end — Anthropic's early prompt tokens arrive, the OpenAI chat lane's usage chunk trails the finish), so the log line records what the wire delivered plus the discarded completion's char count for spend reconciliation. |
||
|
|
47524654b3 |
fix(streaming): close the retry window's generation, identity, and masking holes
xhigh review round on the mid-stream retry ladder: 14 verified correctness findings, all fixed, plus the verified-but-capped cleanups mined from the review run. Generation safety — the shared-slot class is removed structurally, not gated per site: a dead attempt's partial now rides the raised exception (thread-private by construction) into a wrapper-local variable, and the _midstream_dead_partial session slot is deleted, so an orphaned superseded generation cannot poison a live generation's preservation. The promotion helper is generation-gated, writes the marker row even for a pre-token death (empty content takes the marker-as-message branch), and backfills a recorded-but-empty partial with the previous attempt's text, so a Stop anywhere in the retry window — backoff, re-create, or TTFT wait — preserves the latest text the user actually saw. _record_cancelled_partial is generation-gated too: a superseded thread touches neither the UI nor the shared slot. Identity — the retry gate and the fatal formatter now consult the provider that actually owns the live stream (recorded at creation, covering the fallback walk by construction), so a fallback stream's provider-specific transient is retryable by ITS OWN contract and failures are labeled with the binding that produced them. The mid-retry rebind check compares the full (client, model, provider) binding — reload() keeps the pooled client on model-only swaps — and a re-prepare also re-exports the wire fold that send()'s token-table calibration counts. Masking — a context overflow raised by the mid-retry re-create surfaces as itself so the compact-and-retry arm can recover the turn, and the overflow arm is split: recovery-machinery failures still surface the original overflow (its wording anticipates them), while post-compaction consumption failures surface as themselves instead of a false overflow diagnosis. Cancellation and terminal paths — a Stop that races the trailing-metadata window is re-checked after the chunk loop, so the turn aborts with the marker instead of committing and running its tool calls; the terminal arm finalizes client-side only, deliberately keeping the in-progress snapshot (the unpersisted partial's only copy) for refresh-replay; KeyboardInterrupt gets the same client-side finalize; the retry arm stops the spinner before restarting it (the CLI's on_thinking_start replaces the spinner without stopping it — a thread leak); and the backoff delay is computed from the pre-increment index, matching the sibling ladders' convention. Mined cleanups: the retry suite wraps the shared session factory instead of duplicating its defaults; the usage projection uses dataclasses.asdict; the partial-content rule lives in one closure serving both preservation paths; the two fatal-log tests are parametrized into one; the test import uses the public providers package. |
||
|
|
df81035302 |
fix(streaming): finalize dead attempts on terminal paths and harden the retry window
External-review round on the #937 branch; four confirmed findings fixed, each on a failure path the retry loop itself introduced or made reachable: - The terminal arm (retry exhaustion, non-retryable death) now finalizes the dead attempt with the same stream_end + turn_committed pair the retry path emits, so the last attempt's partial is flushed in every consumer — the CLI was the exposed case (its markdown fence state resets only in on_stream_end; the server workers emit their own after a fatal, the CLI's direct send() does not). The finalize is gated behind the generation check: an orphaned superseded thread must not emit UI events over the new generation's stream. - A Stop landing in the backoff/re-create window now preserves the dead attempt's partial: the attempt stashes its flushed content (plus the content-state carry tail) on a non-cancel death, and the wrapper promotes the stash to the cancelled-partial slot before re-raising, so send()'s cancel handler persists it with the cancellation marker — the same disposition a cancel during the attempt gets. - The fatal-path debug trace logs frames only (format_tb): exc_info rendered the raw exception message, which can carry credentials verbatim — the exact leak the sanitize floor above it exists to hold. The recreate-failure warning drops exc_info for the same reason and logs the exception class name instead. - A mid-retry rebind that replaced the client re-prepares the wire messages against the new binding before re-issuing: the system-turn fold is capability-sensitive, and a registry reload that switched model family would otherwise re-send the old family's wire shape. The cross-thread close boundary pin now accepts ReadError or RemoteProtocolError: which one surfaces is platform/timing-dependent, and both are TransportError members of the stream-death set, which is the property the pin exists for. |
||
|
|
3b9de67e8c |
refactor(session): extract think-tag splitting into ThinkTagSplitter
The interactive chunk consumer's _flush_text/_drain_pending closure pair carried the partial-tag carry buffer and in-think state inline. The tag-scanning half moves to turnstone/core/streaming_text.py as a standalone ThinkTagSplitter (carry buffer, in_think state, earliest- index tag selection, MAX_TAG_LEN safe-flush); dispatch and accumulation stay in the session behind the emit callback, and out-of-band transitions (reasoning_delta path, tool-call starts, cancellation) read/write splitter.in_think and flush_pending() where they previously touched the closure locals. Pure move: table-driven pins covering partial-tag buffering across chunk boundaries, the safe-flush margin, open/close tag precedence, in_think transitions, and reasoning-vs-content dispatch were written against the closure implementation and pass unchanged against the extracted class — byte-identical emitted text, identical UI callback ordering. The session-level _THINK_*/_MAX_TAG_LEN class constants fold into the class. |
||
|
|
a1dfe0bd4f |
refactor(streaming): dedupe transport conversion, usage merge, cancel finalize
Three behavior-preserving consolidations behind the #937 fix, each deleting a hand-rolled twin of a now-shared rule: - drain_stream consumes transport_guarded(chunks) and drops its inline `except httpx.TransportError` arm — one conversion rule for mid-body wire deaths across the drained and interactive lanes. The post-finish tolerance now logs under the wrapper's `stream.post_finish_blip` name (formerly `drain_stream.post_finish_blip`) and no longer carries `usage_captured`; changelog notes the rename for external log filters. The possible usage=None result on a post-finish blip is documented on drain_stream itself. - _stream_attempt's hand-rolled per-chunk usage max-merge becomes a local UsageInfo accumulator folded through merge_usage (drain's rule), re-projected into the _last_usage dict on EVERY usage chunk — that dict has mid-stream readers (_estimated_prompt_tokens, the status line), so the per-chunk write timing is load-bearing and unchanged. - The twin cancelled-partial sequences in _stream_attempt's two cancel arms (cooperative GenerationCancelled, stream-close-converted) merge into one local _record_cancelled_partial helper carrying both arms' tool_calls/_provider_content omission rationale in one place. |
||
|
|
5fb27e8f81 |
fix(session): survive mid-stream transport deaths in interactive turns (#937)
A wire death during body streaming (ReadError on a TLS record failure, peer resets) surfaces after the request has already returned its stream handle, so neither the SDK's request retries nor the creation-time retry ladder ever saw it: the interactive turn died with a bare exception string, the partial output was discarded, and no log trace was left. Utility lanes already survived this through drain_stream's normalization; the interactive loop now gets the same treatment. - transport_guarded() in providers/_protocol.py: drain_stream's transport-death conversion made reusable for consumers that keep streaming semantics. Pre-finish deaths raise the retryable IncompleteStreamError (drain's exact message shape); post-finish blips end the stream cleanly, forfeiting only trailing metadata. - The single-pass chunk consumer renames to _stream_attempt; _stream_response is now the resilient wrapper owning ALL stream acquisition plus a bounded mid-stream re-issue ladder (_MID_STREAM_RETRIES, the shared _stop_retrying predicate with a per-loop cap, cancel-aware exponential backoff). Send()'s overflow compact-and-retry arm now wraps the whole turn and passes re-prepared msgs explicitly. - A dead attempt is finalized across every UI consumer before the retry (stream_end then turn_committed then notice then spinner), so retried text never appends onto the dead attempt's in any surface (browser transcript, CLI markdown fences, Slack/Discord streamed messages, SSE replay ring). - Before re-creating, the session re-resolves its registry binding: a concurrent ModelRegistry.reload() closes cached clients, and the retry must not stream into the closed one. A failing re-create logs stream.retry.recreate_failed and re-raises the ORIGINAL stream-death error rather than masking it. - _format_backend_error gains a stream-death branch naming the provider, endpoint, and model, with a short identity-bearing first sentence. _BACKEND_STREAM_EXC_NAMES joins _BACKEND_KNOWN_EXC_NAMES, which also removes those names from _is_ctx_overflow's text-detection eligibility (deliberate: their texts are fixed transport strings that never carry overflow phrases). - _record_fatal_error now logs session.fatal.recorded (INFO for KeyboardInterrupt, ERROR otherwise) so fatal turns leave a journal trace. - _assistant_pending_tokens resets at stream entry so a post-finish blip that loses the trailing usage chunk cannot append the previous turn's completion count as this turn's estimate. Offline SDK boundary pins (openai/anthropic mid-body death identity and no re-request, cross-thread client close surfacing httpx.ReadError) guard the assumptions the retry gate rests on. |
||
|
|
1e34e19d48 |
refactor: single-style module imports and narrowed JSON body typing
Consolidates the repeated function-local model_registry imports onto one from-style module import per test file (the module object stays available for monkeypatching), converts the e2e script's mcp_oauth import to match, and reads the request body as Any before the isinstance narrow so the declared dict type is earned rather than asserted. Addresses the automated review feedback on the pull request; the two code-scanning flags are dismissed as false positives separately (the missing-key refusal log names config knobs and carries no secret value; the URL assertion is a test expectation, not a sanitizer). |
||
|
|
33ace975d2 |
feat(models): default-deny governance and admin UI for per-alias backend auth
Follow-up to the per-alias Entra OBO/app-identity backend auth: the console write path now applies default-deny field classification, the admin shelf gains full backend-auth support, and the session/registry rebind machinery is hardened for config changes landing under live sessions. Console write gate: - Default-deny classification: any non-neutral change to a row that is or becomes dynamic requires admin.mcp plus validation; the provably auth-neutral columns are enumerated (MODEL_AUTH_NEUTRAL_FIELDS) and a live-schema classification test forces every future column to be classified. The derivation is a pure function (_derive_auth_gate) with unit-pinned exclusivity invariants. - Two-tier validation mirroring the MCP oauth_obo validator: the row tier (audience allow-list) runs on every gated write; the posture tier (OIDC configured, token store present) runs on pair changes and on enable-arming. - Pure-disable carve-out: disabling a dynamic row is de-escalation and is never blocked — admin.models suffices and validation is skipped, including for rows with corrupt or skewed stored values. - Capabilities are compared canonically (key order, integral floats), the audience compare normalizes both sides, and staging an audience on a static row is refused on both write twins. - Calibrate writes the capabilities column under an enforced confinement invariant with a compare-and-swap persist. Admin shelf: - Backend-auth section with a per-open constraints fetch (GET /model-definitions/auth-constraints: audience allow-list, grant profile, dynamic modes), datalist audience suggestions, server-defined modes preserved on round-trip, and permission-aware visibility built on cache-skew-safe helpers shared through auth.js. - Refused live-registry swaps surface as an amber registry_warning on the write, delete, reload, and calibrate responses; audit rows carry auth_gated / auth_disarmed markers visible in the audit view. Registry and sessions: - The encryption-key requirement for dynamic auth is enforced inside ModelRegistry.reload() itself — nodes refuse with 503 and the console records coord_registry_error — and reload bumps the generation before the map swap so a racing reader can never pair a stale generation with new maps. - resolve()/resolve_binding() return the generation from inside the registry lock; sessions rebind per send on generation change with atomic client/provider/config commits, fallback-first handling of removed or unconstructable aliases, and judge/limiter resets only when the binding actually changed. - Mint refusals record per-user causes surfaced in the per-turn heartbeat logs; misconfiguration warnings are deduplicated with bounded state. Verification: 10417 tests (99 added on this branch), a 71-scenario browser harness over the real admin shelf, and a live rfc8693 token-exchange e2e run (MCP legs verified end to end; the model-leg scope gap is tracked as #955 under a narrow known-gap signature). Closes #950. |
||
|
|
9adde920d4 |
feat(models): per-alias backend auth via Entra OBO and app identity (#898)
Adds a per-alias `auth_mode` on model definitions so a model backend can authenticate to an Entra-fronted gateway with a per-request minted token instead of one shared static API key, letting the gateway attribute calls to the actual user or to the app as a machine identity. - `static` (default, unchanged) sends the stored `api_key`. - `entra_obo` mints a per-user On-Behalf-Of token for `obo_audience` from the caller's captured refresh credential. - `entra_app` mints an app-identity token via the client-credentials grant, and covers userless turns that OBO cannot. Reuses the existing OBO grant legs, refresh-token rotation CAS, cluster advisory lock and the `mcp_user_tokens` mint-cache, keyed under synthetic `__model_obo__:<audience>` / `__model_app__:<audience>` rows. The token binds at the call site through `client.with_options(api_key=...)` so each SDK emits it on its own auth path rather than through header injection. Migration 068 adds `auth_mode` and `obo_audience`. Both are additive and existing rows default to `static`, so behaviour is unchanged unless an alias opts in. Operator controls: `model.auth_audience_allowlist` is an exact-match allow-list that gates which audiences may be configured and denies all by default, and changing a mode or audience requires `admin.mcp`. `model.auth_fail_closed` decides whether a failed mint may fall back to an explicitly configured static key. A delegated call with no user, or a dynamic alias with no real static key, always refuses. Two changes here apply regardless of whether any alias opts in: - Storage and app state are now wired into the console MCP client manager. This fixes per-user `oauth_user` / `oauth_obo` dispatch for coordinator-hosted sessions, which previously raised `RuntimeError` on first call because `set_app_state` was only ever called on the node. - Unattended watch restores and `--resume` resolve the persisted workstream owner instead of constructing the session under an empty principal. A workstream with no owner is now a permanent refusal rather than an anonymous, auto-approved run. |
||
|
|
bb9684f505 |
fix(ui): block-copy dismissal listens on documentElement, not document
document-level mouseleave delivery on window exit is flaky in some engines, stranding the floating button until the next in-page pointer event; the <html> element receives the leave event reliably. |
||
|
|
729a02a833 |
feat(ui): copy-to-clipboard for messages and rendered blocks
Three idle-only affordances on every chat surface: a persistent copy button in each assistant bubble's actions bar, a pointer-only floating button over the hovered markdown block (fence, mermaid diagram, table), and Enter on a focused block for keyboard users, with the outcome flashed on the block itself. Copy resolves to SOURCE, not rendered text. The renderer stashes each table's raw markdown in data-md-source at render time — span sentinels restored in reverse mask order, footnote-definition bodies restored to raw before their recursive render — and whole-message copy reads the streaming pipeline's per-frame stash. The clipboard transport falls back to the legacy execCommand path for plain-HTTP LAN nodes, cloning and restoring the user's selection and focus. Outcomes surface button-local only: flash + title + one live-region announcement through the shared makeAnnouncer factory (also adopted by the interactive voice/tool announcers, whose lazily created regions swallowed their first announcement). Busy refusals answer with their own message. Coordinator retry and admin token-copy keep zero-module-dependency degrade paths. |
||
|
|
e526df95d0 |
fix(tool-search): discovery-failure records are per-user, rerank counts honest
Follow-up to #938; closes #941. The unavailable-server advisory fired for users whose own pool was warm: _pool_discovery_error was keyed by server name while pool connections are per-(user, server), so one account's failed prime rendered its exception text into every user's search results. - mcp_client: re-key _pool_discovery_error to (user_id, server_name). Written by the failing user's prime (single sanitize-and-cap pipeline shared with _set_error), cleared by that user's successful connect, retired with the grant on explicit disconnect / dead-grant convergence, and swept name-wide on registration lifecycle (removal, reconcile auth-type flips) via a snapshot-safe helper. Departed users' records are reaped by the eviction tick's orphan sweep — the single tick-side reaper; a live user's record survives its stub's eviction because the advisory has no mid-session re-record path. The eviction loop also starts on record write, so records written before any pool entry exists cannot outlive their users. Status reads scope to the requesting user, with an any-user view under the admin aggregate flag. - tool_search: _status_reason treats discovery_error as an outage only when the requesting user's own status is not connected — with per-user records this is belt-and-braces, since a successful connect clears the user's record. - session: the tool-search status snapshot scopes to the EFFECTIVE user (the acting participant on shared workstreams), matching the get_tools call that builds the search corpus, so an owner's pool state never renders into a non-owner's results. - bm25: with a reranker attached, matches ranked past the recall pool trail in BM25 order (reorder mode), so tool_search's "top N of M" count no longer floors at the pool size; the exception fallback is mode-aware (filter mode keeps its pool bound, byte-for-byte). |
||
|
|
c4b2dd7135 |
feat(tool-search): surface MCP discovery failures & honest result counts (#938)
Tool discovery for a pool-backed (oauth_user/oauth_obo) MCP server that is
down or 5xx-ing was invisible: the server contributed zero tools to the
catalog, so tool_search returned "No matching tools found" —
indistinguishable from a genuine no-match — and matches past max_results
were silently dropped with no signal.
- tool_search: search() ranks the whole deferred corpus and records the
pre-slice match count so format_search_results can report honest
truncation ("top N of M"). An optional status_provider lets results name
servers that are actually failing (open circuit breaker, recorded error,
recorded discovery failure) instead of masquerading as "no such tool".
Un-primed servers are deliberately not flagged, and a provider that
raises never breaks search.
- mcp_client: the previously swallowed pool prime/connect discovery
failure is recorded per server (single-line, bounded), cleared on the
next successful pool connect, on removal, and on reconcile-observed pool
removal or auth-type flips; exposed via get_server_status
as "discovery_error".
- session: wires get_all_server_status(user_id) into both
ToolSearchManager constructions as a lazily-called status provider.
|
||
|
|
57a9041941 |
fix(coordinator): the interjection handoff cannot lose the message, and the fact block is bounded
Review fold-in before push, twelve findings, two of them majors. The handoff popped the interjection queue destructively and handed the text to a send with a non-delivering refusal (the budget latch) and a preamble that can raise before the user turn is appended — a failure destroyed the user's words with a log line, after the charged wake nudges were already cleared. Now: the budget latch is checked before the pop (the message stays queued for a send with a human in front of it, and the wake drain still runs so the worker's exit converges); the pop returns the raw items and any non-cancel escape restores them verbatim — ids and priorities intact — before the failure surfaces; a cancel deliberately does not restore, because the Stop supersedes the queued words. Content-free items (a bare priority marker) are skipped at the shared renderer, so a lone '!!!' no longer buys a content-free turn at the cost of both nudges. The per-child fact block takes the roster formatter's bounds: fact lines cap at the display cap with a counts-only overflow line, and the wait slot keeps its larger handle cap — the body is a persistent system turn replayed on every request, and the block previously grew without bound as finished-but-unclosed children accumulated. The two fact sentences and the overflow line are named template constants, and every test assertion anchors on them; the children projection takes the same drop-never-mangle alteration check as the open-row fields. Eval world seeding: node metadata is JSON-encoded exactly as production writers store it (a raw string never matched a filtered list_nodes lookup), the stub client pins its heartbeat window open so a static world cannot go hollow mid-run, and the world-shape refusals' field branches gain their own tests. Comment accuracy and paragraph wrapping fixed at the sites the review named. |
||
|
|
2ff7c61051 |
feat(coordinator): idle nudges deliver only on the idle wake, and a queued interjection owns the seam
The idle nudges enqueued on the any channel, which every drain seam
serves — a deferred wake left them deliverable at the start of a real
user send or mid-turn at a tool batch, describing an idle moment that
no longer existed. They move to a new wake channel: wake-eligible,
invisible to USER_DRAIN, TOOL_DRAIN, and the quiet ride-along. A user
cancel drops pending wake entries rather than demoting them — the
quiet demote's whole value is later seam delivery, exactly what this
class may never have. Dropping a charged entry is the accepted
fail-closed cost; liveness surviving Stop means the next idle event
fires fresh, not that a queued entry re-wakes the workstream.
A queued user interjection owns the idle seam: at wake delivery, a
non-empty interjection queue drops the wake-channel entries and the
interjection runs as a genuine user send in their place — no wake
tag, so the caps reset as for any real send and the next genuine idle
re-derives both nudges over fresh reads. The check lives in the wake
worker (which owns the slot and can dispatch a full send), not the
watcher's state-transition thread, where skipping would strand the
message. Measured before building: send('') with queued messages
appends an empty user turn and delivers the interjection one
assistant turn late, so the handoff pops first and sends the popped
text — one rendering shared with the flush seams.
External events are not idle nudges: any-channel entries still arm
the wake alone, and with an interjection waiting they ride the
genuine turn's drain seam — both deliver, only the idle nudges drop.
A failed wake send drops wake entries alongside user ones; externals
requeue quiet as before.
|
||
|
|
074b9e02e2 |
feat(eval): cells seed the tool-visible world through production writers
Every surface the model can observe must agree about the world's age and contents. The C1 confirm at n=25 measured models sweeping memory, skills, and list_nodes, finding voids that contradicted a transcript full of referents, and spawning read-only investigators to resolve the contradiction — the forbidden rate was measuring the fixture's hollow tool-world, not dispatch discipline. A cell's world block seeds structured memory rows through the same upsert the memory tool's save action commits (names normalize exactly as model-saved rows do), and node rows through the service registry plus node metadata — the two reads list_nodes intersects, so a seeded node is live inside the heartbeat window by construction. A seed failure raises; a malformed world block is refused at config time before the canary, with its own trip cell in the reachability guard. The approval-stop cell gains the first world: two process-fact memory rows (no coaching — the reservation lives in the transcript only) and one live node. |
||
|
|
76c5519b44 |
feat(coordinator): done branch leads the tasks body; eval worlds survive honest inspection
Three fixes, one per causal mechanism the round-12 baseline exposed. The done branch moves ahead of the escalate branch. The escalate-first order rested on a harm argument — guessing on an operator decision outranks redone bookkeeping, so the escape hatch should be salient — and the baseline measured its cost: 7 of 10 finished-unmarked runs reached for the body's first populated call and escalated visibly finished work, one mode, no tail. The next round measures the reversal both ways: if the legit-stop cells' forbidden rate rises, the harm argument was right and the order flips back (the pin says so in place). The approval-stop cell's transcript anchors its world — named repo, named migration, named artifacts. Its forbidden runs were not sign-off defiance: the model swept empty discovery surfaces, found a void, and spawned explore-the-project children, so the cell was measuring hollow-world exploration rather than dispatch discipline. The co-delivery cell's running child gains an observations-only progress note beside its assignment. A bare-assignment static child cannot survive sustained honest interaction — wait times out, inspect shows nothing, and after patience cycles the model correctly diagnoses a hung child and cancels/respawns, which the forbid list scored as redo. The note makes the child look alive without looking finished. |
||
|
|
48d6b2f84b |
feat(coordinator): the nudge bodies state observed facts, never hedges
The idle-children header drops its opening idleness claim: a queued entry delivers at whichever seam arrives next, and the drain predicate re-verifies that children are active — never that the coordinator is still idle — so the body now opens with the one fact the delivery just verified. The tasks body replaces its hedged children sentence with one observed-fact line per child. The old sentence hedged states the producer's read had just returned and invented activity for an idle coordinator; the producer now threads (ws_id, state) pairs through, and the formatter renders a running child as running (check before redoing what it owns) and a stopped one as stopped, with the tool-behaviour fact that wait_for_workstream returns immediately for it. The line asserts nothing about results: no read observes whether a child produced anything, and the immediate wait is the whole protection — checking is cheap and finds whatever is there. Fact lines are formatter-built beside the counts opener, so no tail override can reach them; the formatter's old indeterminate-read hedge branch is deleted (a failed read renders no body at all), and the open-row status takes the same alteration check as the id. Both bodies hand the model full workstream ids: the resolver refuses truncated ids by design, so the roster's 8-char prefixes were not handles — a model copying a bullet issued a call the resolver rejects. Display prefixing stays on the operator card, derived from the full id in the metadata. Eval alignment: fixture child ids become production-shaped 32-hex (a prefix looked like a different id entirely and the old shape only resolved through the legacy branch); a body-override sweep refuses cells without a live child at config time, keyed on the formatter's own childless condition, so candidate text can never be measured over a world production cannot produce. |
||
|
|
33c82962a2 |
fix(channels): suppress mention resolution and escape untrusted fields
User- and model-authored text (task titles, approval headers, command previews, judge output, error fragments, notification bodies) reaches both channel integrations verbatim, and nothing at the channel boundary neutralised it. The Discord client now carries a client-level allowed-mentions-none default, which every message create inherits — plain sends, edits, and embeds — so broadcast and mention syntax in untrusted text cannot resolve, without mutating the text itself. The Slack adapter escapes each untrusted field into mrkdwn entities at its interpolation site — never the assembled message, so deliberately bot-authored markup like the session-opener mention survives. The policy-deny feedback returned to the server stays verbatim; only the rendered notice escapes. Storage and the shared formatter stay channel-neutral and verbatim: projection happens per audience at the render boundary. |
||
|
|
b8dd5041c4 |
fix(session): force-cancel runs the abandon machinery before emitting idle
The force branch cleared worker ownership and emitted idle from the route thread, while the abandon latch and the queue demote ran only in the stuck worker's own exception handler — a thread force-cancel abandons precisely because it is not making progress. Subscribers on the IDLE fan-out therefore saw an operator-forced idle with the latch unset: the idle observer's operator-Stop gate did not suppress advice, and wake-eligible entries survived un-demoted, so a nudge wake could resume a workstream seconds after the operator forced it to stop. The route now runs the session's abandon machinery first; the abandoned thread re-running it at its eventual death is idempotent. |
||
|
|
556ec793a0 |
feat(eval): behavioral eval for the coordinator idle nudges
turnstone-eval --nudges runs seeded coordinator states against stimulus arms and scores state-first: cells seed a real task envelope through the production tasks_add path into a per-run temp DB, the model's tasks calls really execute, and ground truth is the final envelope plus a forbidden-action list — robust to action-path variation, and never tool_choice-forced. Arms render through the production formatters so the wire carries exactly what production sends; a body-override lane exists for tuning A/B only and skips the ablation arm, whose reading only means anything against the body that ships. Children are seeded with transcripts — an assignment message, plus a completion-with-findings for idle children — so honest inspection finds a world rather than an empty room, and a collect-vs-redo cell measures the real question. Cell authoring is validated up front (including refusing a parked task seeded beside an open one, a state production never sends a body for); mutating calls are scored by what landed, not by what was attempted. Instrument health is measured, not assumed: a canary probes tool-call parsing before and after the sweep, a mid-sweep tripwire aborts a sweep whose parser dies rather than printing a red grid, and empty-log runs are labelled harness: so they can never score against the model. The coordinator idle-observer test file lands in this commit rather than the feature commit: its parity guard imports the eval scenarios to pin the production formatters and the eval wire to one rendering. |
||
|
|
a53bf30428 |
feat(coordinator): carry task and child handles across a compaction
The tasks tool returns an id beside its title, and spawn returns a child's id; that pairing lives only in the transcript, and compaction replaces the transcript. A coordinator that loses it cannot update its own tasks or collect a finished child's results — and it is exactly the coordinator most likely to be sitting idle holding unfinished work. The idle nudge now carries storage-derived ids for the same reason, and this is the other half: the nudge supplies the authoritative set, this preserves what each one means. The harness writes the block itself rather than asking the summariser to preserve ids. Both are available to it — the reads are same-process storage on the thread already running the compaction — so asking the model to transcribe what the controller is holding would be a shortfall in the lowering, and would make every id fallible to no purpose. Neither compactor prompt changes at all, and a test pins that they stay identical across kinds, so a future prose section has to revisit this trade rather than stack on top of it. Interactive sessions have no task envelope and no children, so they take no reads and render nothing — a gate on the kind, not a section that renders empty. Their compaction is unchanged by construction. The block joins the existing carries, which is where the real hazard was: it lands in the same post-compaction prompt as the wind-down spill and the continuation ask, so it is counted as a third carry and rendered against that shared budget, with the count and the render reading one answer. Truncation drops whole rows and names what it dropped; half an id is a call that cannot resolve wearing the costume of one that can. Tasks are served first, with room reserved so a long list cannot starve the children. Titles keep their angle brackets here — unlike the nudge bodies, which delete them because they interpolate into a system turn where a tag-shaped run steers. This is the assistant channel, the titles are the coordinator's own, and they already reach this same model verbatim through its own list results; deleting brackets would only invert the constraints it is working to. The control class is still stripped, so a newline in a title cannot forge a sibling row. A failed read costs the block, never the history: trading a whole history swap for a side read would be the worse failure by far. |
||
|
|
0d52b63b50 |
feat(coordinator): nudge a coordinator that goes idle holding unfinished work
Two nudge classes can fire from one IDLE event, tasks first, each asserting only its own domain. idle_tasks (advice) fires when open (pending/in_progress) tasks exist. The body is a counts opener, the open task ids with statuses, and typed branches that each end in a runnable tasks(...) or wait_for_workstream(...) call populated with real server-minted ids. Everything it says about children is governed by one observed fact: live children present adds the caveat sentence and the blocked-on-a-child branch; affirmatively none says nothing about children at all. Any needs_user row parks the class entirely — at the fire gate and the drain predicate — because with no task graph an open task may be gated on a parked one's unanswered question; the operator's answer is the re-arm. Gated on memory.nudges and on the persona actually exposing the tasks tool; carries the per-class cooldown as well as the per-bracket cap. The tasks tool itself gains the needs_user status and a note field — the typed escalation the body's branches point at. idle_children (liveness) fires when children are in a live state — the wake that lets an idle coordinator collect a finished child's results. The body is a roster of workstream id prefixes and states, never names: a child's name is model-authored text and does not enter a system turn. Cap-only and cooldown-free by design, not gated on memory.nudges, and it survives an operator Stop. Fail-closed, event-wide: if any storage read fails while the observer handles an IDLE event, neither nudge is queued and neither cap is charged. Both paths run as side-effect-free plans; the commit tail is storage-free, so no read can fail past the veto point; both drain predicates drop on a failed read. A path's own fault (a generic raise) still costs only that path's fire, so one class's bug cannot strand the other. Task text is stored verbatim and projected per audience at render: the model-facing projection deletes angle brackets, the operator projection keeps them, and both strip newlines and bidi/zero-width runs. Idle cards render what the model was told, formatted for the operator, never augmented with content the model did not receive. |
||
|
|
96834496c4 |
feat(anthropic): onboard claude-opus-5
The capabilities row is a copy of claude-opus-4-8 — 1M context, 128K output, adaptive thinking, the full low..max effort ladder, mid-conversation system messages. Two of the model's documented breaking changes are unreachable from this lane and stay that way only while thinking_mode is "adaptive": thinking is on by default when the param is omitted, and disabling it is a 400 at effort xhigh or max. We never omit it and never emit "disabled", so both are recorded at the row rather than defended against. The third needed work. A safety classifier can decline with stop_reason="refusal" on a successful HTTP 200, with content either empty or partial, and an unmapped value fell through _normalize_finish_reason as a literal string. The drain gate only raises on an ABSENT finish reason, so a declined turn landed as a complete result with nothing to notice it: the interactive lane matched neither of its two warn arms and stayed silent, and a sub-agent handed the declined partial up to its parent as though it were finished synthesis. Normalizing onto content_filter routes the decline into the arms both lanes already have for that state — the operator gets the warning, and the sub-agent stops rather than passing the fragment on. The raw stop reason is logged where it is still in hand: normalization is lossy and a classifier decline is otherwise indistinguishable from an ordinary content filter. The gate is on the RAW value rather than (normalized != raw), which is true for end_turn and tool_use as well and would fire on every turn in every lane. The anthropic floor moves to 0.117 to track the release current at onboarding. The model needs no new SDK surface — ids are opaque strings and "refusal" has been in the StopReason literal since ~0.95 — so this is hygiene; raise it again when adopting fast mode, server-side fallbacks, advisor, or mid-conversation tool changes, which do need newer typed params. |
||
|
|
04c29c8ef5 |
test(authz): stop the retract test racing the drain's claim window
test_persistently_crashing_entry_is_never_dropped_and_retract_frees_drain fired its DELETE off the attempt counter, which increments as the FIRST statement of the attempt. The counter therefore crossed 2 while the drain still held the entry CLAIMED — popped off _pending_sends, dispatch in flight. Retract only scans that list, and correctly answers not_found for a claimed entry, so the request was racing the crash path's re-insert and the assertion saw not_found instead of removed. Both sides of that race are the same order of magnitude, which is why it read as machine-specific rather than simply broken: the re-insert lands after an intervening log.exception, roughly 5ms under pytest's capture handlers against 0.05ms with none installed, and wait_until polls at 5ms. A fast idle machine loses the race; a slower or busier one wins it. Wait for the state the test is actually about — the entry back on the list, mid-crash-loop and retractable — rather than for the counter. That is also what the docstring already claims is under test. The contract still bites: dropping the entry on the crash path fails the new wait, and removing both retracted-purge sites leaves the drain spinning and fails teardown. |
||
|
|
6f194d338f |
fix(console): restore back-to-console from a proxied node view
The console proxies a node's web UI at /node/{id}/ and injects a shim into
the page it fetches upstream. The only way back was a node-picker menu the
shim built into #ui-header — an element the L-shell renovation (
|
||
|
|
4007fab855 |
fix(#900): close two vacuity holes the round-2 scenarios left open
Round-3 review, unprimed. Two of its majors were the new scenarios asserting things they did not prove — the false-detector class this campaign keeps returning to. E8 never checked that the held /history was still OUTSTANDING when the redial completed. The disconnect/send/wait_turn/redial sequence is unbounded (wait_turn alone allows 45s), so on a slow box the payload resolves while evtSource is still null, the PRESENCE term declines it, and the run stamps dupes1-healed1 without ever evaluating the generation term. It now fails loudly with the counter values instead. E7 gained the same positive proof its siblings already carried: sse_opens == 0 only means "nothing connected in 8s", which is not the same as "the held load settled and its .finally chose not to reconnect". E5's stated control was simply wrong, in three places. A hide nulls evtSource, and connectSSE early-returns while hidden, so the scenario cannot produce the non-null-but-not-OPEN source that readyState === OPEN exists for — it exercises the presence term only. The earlier control removed both terms at once, which is what disguised it. The readyState half is covered by reasoning plus coord parity, and its correctness twin IS covered through the render-time gate by E6/E8; that scope is now written down rather than overclaimed. Coord's G5 has the same shape. The retry floor becomes a shared export beside its jitter: four sites must move together (both clients' arms, both non-occurrence windows) and it was the only one of them with no single source of truth. Interactive's use of the expression had no pin at all — reverting it to a bare 2000 would have broken cross-client parity with the suite green. Coord's re-anchor still raised ValueError rather than failing on a named assertion, and its first replacement used a fixed window that truncated mid-expression. |