mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
cc84f9d176a9a67cccbbce8fd862421fdb9a773a
1041 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cc84f9d176 | fix(memory): harden project scope authorization and consistency | ||
|
|
d2a6c2852e |
Stabilize context-overflow compaction test in Python 3.11 CI (#1006)
* Stabilize overflow compaction test expectations Co-authored-by: eous <13773563+eous@users.noreply.github.com> * Fix ObservedRLock compatibility with Python 3.14 Condition Co-authored-by: eous <13773563+eous@users.noreply.github.com> --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: eous <13773563+eous@users.noreply.github.com> |
||
|
|
480a1426b3 |
Fail-closed history-commit handoff (#1005)
* fix(session): fail-closed history-commit handoff (#981) The deleted-workstream discovery is now a terminal, ws_id-keyed latch: keyed conversation commits refuse admission once the durable parent is gone (convergence finalizers and force-abandon are exempt), history handoff refuses to mint a proof token so /history fails closed with a 503 instead of silently wiping the pane, and the SSE stream carries a workstream_gone resync reason. Discarded commits leave a forensic log of commit keys and roles, never content. Conversation rows gain a commit_key (migration 071): keyed saves are idempotent under retry, validated against the full commit identity, and refused when they would cross a workstream deletion. The prune orphan category now requires a NULL alias plus a two-hour updated grace, with cutoffs computed at discovery time and carried into both dialects' rechecks. The mid-turn interjection queue is owner-partitioned with no per-site mode flags: pops take the acting principal's and unowned rows, other participants' rows are structurally retained, and enforcement lives at queue admission plus the shared before_spawn gates. The retraction ledger is bounded by open pop windows: pops open a window atomically with the queue delete, restores close their ids atomically with the ledger consume, every other exit closes through one helper, and misses for unheld ids record nothing. The workstream-gone latch refuses unattended wakes at all three gates (watcher spawn, claim, delivery pre-pop), and the retry dispatcher regained its pre-envelope cancel/error convergence net. Persistence-state reporting derives through the session bound to each UI instead of a registry lookup by id that failed open to healthy during tombstone retention. The dashboard roster no longer re-inserts ghost entries from trailing activity events, the history tool-outcome scan tolerates interleaved non-turn rows, and the shared handoff-deadline handle owns its own retirement. Single-sourced across call sites: keyed-commit row values, attachment save wrappers, tail-truncation and conflict-resolution bodies for both storage dialects; worker-slot lifecycle field sets; the direct-commit admission frame; queued-row layout accessors; the string-aware comment stripper shared by every JS harness suite. Refs #981 #964 * fix(session): sweep handoff fixes to their sibling surfaces The interactive replay loop treated a system row as a tool-batch boundary, so every tool result after an interleaved row vanished from that pane while the coordinator rendered the same history correctly. Only a conversational turn ends the batch window now, matching the shared outcome index. Accepted user turns clear the composer's attachment chips on the same viewer policy that settles optimistic bubbles rather than on having matched a local bubble, so a workstream created with an upload no longer keeps a chip for an attachment the create dispatch already consumed. The coordinator's raced-Stop arm emits the stream-end hook it inherits alongside the idle state, leaving no unfinalized bubble or unflushed tool output. Ending a session surfaces a failure toast when the request never lands or answers with a non-JSON body. The per-second persistence reconcile now probes each session without blocking: a workstream whose generation and handoff locks are held is skipped until the next pass instead of contending the locks every commit needs. The one-shot repair that gates workstream creation at capacity keeps a definite probe — it has no next pass, and the sessions likeliest to be contended are the ones whose unresolved journals emptied its candidate list. Single-sourced: the attachment lane builds its conversation row through the shared commit-identity builder; the ordinary worker exit releases its slot through the lifecycle owner; both operator surfaces snapshot their counters through one non-consuming helper; the replay preamble loses its per-kind wrappers and its config hook; the browser harness suites share one brace walker; and each in-flight history attempt is one record carrying both its abort controller and its deadline. Refs #981 #964 |
||
|
|
f4fd7e1f67 |
fix(security): classify outbound addresses by what they reach (GHSA-wm4f-79pw-pfr9) (#1003)
* fix(security): classify outbound addresses by what they reach (GHSA-wm4f-79pw-pfr9)
Five guards screened outbound URLs and each hand-rolled its own address
normalization and policy tests, so each had a different hole. An IPv6
transition address carries an IPv4 destination in its low bits and
`ipaddress` classifies the wrapper, not the destination: 64:ff9b::a9fe:a9fe
reports is_global because 64:ff9b::/96 is global unicast, while a NAT64
gateway routes it to the cloud metadata endpoint. CGNAT (100.64.0.0/10) is
neither is_private nor is_global, so a denylist built on is_private missed
it with no gateway involved at all.
Add turnstone/core/ip_classify.py as the single classifier. One function
returns exactly one policy lane — PUBLIC, PRIVATE (operator-approvable) or
NEVER — and every guard branches on the lane rather than re-deriving it.
Two overlapping booleans would make a verdict depend on which one a caller
tested first; several addresses are simultaneously globally routable and
metadata-reaching.
- Decode transition addresses per RFC 6052 §2.2 (NAT64 well-known and
local-use prefixes, 6to4, Teredo, IPv4-mapped, IPv4-compatible) and judge
them by the IPv4 they reach. The local-use prefix does not say which
layout its gateway uses, so every length it can carry is decoded and the
worst result classified.
- Share hostname resolution too. The five copies had already drifted on
which failures they caught, and getaddrinfo raises UnicodeError — not an
OSError — from the IDNA encoder.
- Resolution failure is a refusal, not a pass: the fetch resolves again, so
an authority answering the guard with SERVFAIL and the fetch with an
internal address would otherwise switch the guard off for that hop.
- Screen every redirect hop in every mode. allow_private_origin widens which
lanes are acceptable rather than turning screening off, and the permission
is revoked after any hop that is not wholly private.
- Cleartext http is allowed only for a hostname that RESOLVES to loopback.
*.localhost is ordinary DNS, and trusting the name put an OIDC token
exchange on the wire in the clear.
- Screen doctor and console-probe URLs through the classifier. Both used a
host.startswith("169.254.") string test that never resolved, so any DNS
name pointing at the metadata service passed and its body was returned to
the model.
- Add known vendor metadata prefixes the stdlib does not flag, and place
deprecated IPv6 site-local outside the public lane.
The operator's private-network opt-in still admits the whole home lab,
including IPv6 loopback, CGNAT and split-horizon hosts. Metadata,
link-local, multicast, unspecified and reserved addresses stay refused
regardless of the opt-in, including as a redirect target from an approved
private origin — the settings help and docs now say so.
Reported by @tonghuaroot.
* fix(security): close Azure/Oracle metadata gap and restore dual-stack origins
Review follow-ups on the address-classification rework.
Azure's host-agent endpoint (168.63.129.16) and Oracle Cloud's metadata
endpoint (192.0.0.192) sit in ordinary unicast space, so the stdlib reported
them as globally routable and both classified PUBLIC — reachable with no
opt-in at all, a worse position than the RFC 1918 host beside them, and
directly contradicting the "metadata stays refused even with the opt-in"
guarantee the settings help and docs now advertise. Both join the shared
vendor list.
Revoking the private-hop permission on the ORIGIN hop broke the case
`_screen_tool_url` deliberately admits: a dual-stack or split-horizon
home-lab host answering with both a LAN and a public record was approved,
then refused on its own `302 /login` — one hop was all it ever got. Track
the approved HOST instead, so redirects that stay on it remain covered while
a redirect to any other private host is still refused once the chain is no
longer wholly private.
Also:
- Try several registry candidates for the collector-scope probe instead of
abandoning it when the first is unresolvable, which also stopped a healthy
registry from logging as malformed.
- Bound the probe's name resolution with an explicit timeout matching the
2s the httpx connect deadline used to provide; it runs before the console
lifespan yields and getaddrinfo has no timeout of its own.
- Route doctor and the console probe through `web.screen_url` rather than
keeping a third and fourth copy of parse/resolve/classify/fold, which had
already diverged on default port and empty-hostname wording. An empty
hostname no longer reports as a cloud-metadata refusal.
- Give `screen_url` a scheme-aware default port.
- Stop doubling the word "hostname" in the OAuth resolution refusal.
- Correct the `_screen_tool_url` docstring: it described `private_origin` as
requiring every record to be private, which the mixed-record decision
reversed, and `private_block` as a property of a refusal when it reports
the lane on the success path too.
- Make the preview tests' screening stub opt-in rather than autouse — as a
module-wide fixture it also stubbed the tests whose subject IS the screen,
so one of them would have passed even if screening refused everything.
Verified the module now passes with all name resolution blocked.
* fix(security): refuse mixed-record private origins instead of exempting them
The previous commit let an approved private origin redirect to itself by
exempting its hostname from the chain-wide revocation. That exemption was
wrong three ways: it was captured once and never cleared, so a public hop
could steer the fetcher back into the approved host at a path of its
choosing — reopening the private -> public -> private bypass; it was
re-entrant across same-host redirects with fresh DNS each time, so a
self-redirecting host could walk arbitrary internal addresses; and it
matched on bare hostname, so it spanned every port on the approved box.
All three were reproduced against the parent commit, which refuses them.
Delete the exemption rather than repair it. The case it existed for — a
dual-stack host answering with both a LAN and a public record — is now
refused where it is actually decidable, in `_screen_tool_url`, with the
remedy in the message: point the tool at the LAN address directly. A
granted chain therefore always starts wholly private, so the fetch guard
needs no notion of an approved host and stays one unconditional rule.
That the accommodation could not be expressed safely in the guard is the
signal: the connection may land on either record, so approving such a host
never described where the fetch would go.
Also from the same review:
- Walk the whole service registry for a collector-scope probe candidate
instead of the first three, and split the outcome into three log lines,
so entries that are merely unreachable stop raising the malformed-registry
alarm and skipping the boot check cluster-wide.
- Stop the candidate walk on a resolver timeout. `asyncio.timeout` bounds
the await, not the work, so continuing left one parked thread per timed-out
candidate on the shared executor.
- Move the metadata-hostname denylist into `ip_classify` and enforce it in
`screen_url`, so doctor and the console probe inherit it instead of each
keeping a copy.
- Drop the scheme-aware default port: a numeric service does not change
which addresses resolution returns, and classification reads only those.
`parsed.port` is still touched so an out-of-range value refuses.
- Correct the vendor-metadata comment, which generalized a claim true of
Azure's and Oracle's addresses to Alibaba's CGNAT one.
- Rename a test class that was still named for the rule it no longer tests.
|
||
|
|
9211c4fb29 | chore: download vendored JS files | ||
|
|
766223e774 | feat(judge): parallelize batch evaluations (#991) | ||
|
|
98e96ab5f3 |
Add per-alias model concurrency admission (#990)
* feat(models): add per-alias concurrency admission Add registry-backed FIFO admission limits with queue-aware deadlines and full-stream leases. Expose max_concurrency through storage, admin configuration, OpenAPI, documentation, and diagrams, with role and live backend count coverage. * fix(api): omit null concurrency schema default Keep max_concurrency optional for presence-keyed updates without advertising a null default for its non-null integer OpenAPI shape. |
||
|
|
7a06f5e8bc |
refactor(session): make ModelLane the provider boundary (#979) (#989)
* refactor(session): make ModelLane the provider boundary (#979) ## Summary This closes the model-lane ownership gap left by #832: `ChatSession` no longer stores raw provider/client handles. `ResolvedModelBinding` now carries the provider, client, model, capabilities, registry generation, and backend-auth configuration as one coherent snapshot. - Atomically rebind existing sessions after model-registry changes while pinning each in-flight send, fallback, judge, output guard, task agent, title, compaction, perception, and voice operation to its initiating principal and binding. - Fence UI publication, canonical trajectory folds, durable writes, streams, retries, child scopes, and judge work by generation. Stop can hand off to a successor without accepting late state; cancelled tools retain typed effect receipts, and concurrent approval batches resolve by exact cycle or call. - Make create, fork, open, close, and delete race-safe with hidden `creating` reservations, incarnation-aware state tails, and an ACL-rechecked transaction that clones checkpoint-bounded history, configuration, project/persona state, and attachment references. - Extend REST/OpenAPI and Python/TypeScript SDK contracts for create/fork inputs, routed-create metadata, live-workstream probes, targeted approvals, and structured cancellation results. - Update architecture, storage, authentication, judge, channel, console, API, and SDK documentation, including regenerated architecture diagrams and OpenAPI artifacts. ## Validation - SQLite suite: 11,188 passed, 9 skipped, 10 deselected - PostgreSQL suite: 11,195 passed, 2 skipped, 10 deselected - Live backend: 3 passed - SSE recovery: 6 passed; browser recovery harness passed all scenarios - Ruff: clean; 595 files correctly formatted - mypy: 243 source files clean - TypeScript: typecheck/build and 35 tests passed - OpenAPI artifacts fresh; all 14 changed diagrams reproduce byte-for-byte - `git diff --check` and Git LFS integrity clean Closes #979. * fix(deps): update nanoid for GHSA-2v37-7h3g-55p8 Refresh the transitive lock entry admitted by PostCSS so the TypeScript security gate no longer resolves the vulnerable custom-generator implementation. Validation: - npm ci - npm audit --audit-level=moderate: 0 vulnerabilities - TypeScript typecheck and build - TypeScript tests: 35 passed * fix(test): assert canonical model registry URLs Replace prefix checks with exact canonical base URL assertions so the tests do not model incomplete URL validation. Validation: tests/test_model_registry.py (185 passed); Ruff check/format; mypy. |
||
|
|
0cdb679892 |
Follow-ups on the #832 fold: supersession predicate, wire-prep error hygiene, reasoning-parser tile (#986)
* refactor(session): ask the shared supersession predicate at the older sites ``_check_cancelled`` and ``_compaction_event`` predate ``_generation_superseded`` and each carried its own inline copy of the formula, so the drift the helper exists to prevent had two live places to start from. Both are behaviour-identical today. What the pin protects is the generation-0 convention: a bare ``!=`` reads a direct seam caller as an orphan, which would raise a cancel on a live turn and stamp a live compaction superseded — suppressing the end notice, so an operator watching a real compaction fail would be told nothing at all. * fix(session): render a wire-prep fault's cause class, never its message Every other branch of the fatal formatter tails the backend's own diagnostic text, which is what the operator needs. This branch is different in kind: ``prepare_wire`` is our lowering over the session's stored history, so its exception message can quote that history — and the formatted string is both shown to the operator and persisted to ``last_error``, which a coordinating agent reads. ``redact_credentials`` is a best-effort regex by its own docstring, so it is no floor for arbitrary conversation text. The cause's class still identifies the fault, the guidance is unchanged, and the debug traceback logged in the same function localizes the raise site. * feat(console): surface the server-side reasoning parser capability The inline think-tag scan is a fallback for inference servers with no reasoning parser, and for misconfigured ones. An operator running vLLM or llama.cpp with a parser configured had no way to say so from the model shelf — ``server_parses_reasoning`` was reachable only by hand editing the raw capabilities JSON, and it defaults to off, so the scan stays on and both channels run at once. The tile test is a general invariant rather than a single-key pin: every tile key must render a checkbox, carry a default, and — where the key is a ``ModelCapabilities`` field — agree with the dataclass. The matrix is a hand-maintained mirror, so it drifts silently otherwise. * fix(model_turn): a wire-prep wrapper carries the cause's class, not its text Withholding the message in the fatal formatter was not enough. The wrapper was built as ``WirePreparationError(str(prep_err))``, so ``str(exc)`` IS the cause's message — and the interactive retry arm renders exactly that into the dashboard SSE, one line after the formatter emitted the redacted version. ``sanitize_error_text`` is no floor there: it returns arbitrary stored-history text unchanged. Fixing the exception rather than the one consumer closes every caller that stringifies it, now and later. The message still rides ``__cause__`` for tracebacks and debug logs. * fix(console): coerce lifted capability values the way the backend does The tile lift used bare ``!!``, but the capabilities dict is hand-edited JSON: a stored string "false" is truthy to JS while ``apply_capability_overrides`` reads it as False. Opening such a row rendered the tile CHECKED and saving persisted boolean true — inverting the capability without the operator touching it. For ``server_parses_reasoning`` that silently disables the inline tag scan, the exact typo model_turn's comment already warns about, and this key had just been lifted into the matrix. ``_capBool`` mirrors the backend's spelling table; a value the backend would not coerce stays in the raw JSON rather than being rewritten, which is the policy the modal already applies to thinking_mode. Cases are generated from the Python table and executed under node, so a spelling added on one side fails here. Also tightens two pins the tile test left open: the checkbox must render inside the container the JS actually queries, and a tile key that is not a capability field is exempted by NAME rather than by a blanket hasattr, which was swallowing the consistent-rename case. * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
f15e53dd36 |
test: drop a no-op conditional and splat the pre-fold seam call
Static analysis on the pull request caught two leftovers from the mechanical ports. An `if True:` wrapper survived the conversion of a patch block into the armed-provider fake, adding a nesting level that manages nothing — the same shape as the `nullcontext` leftover removed earlier, and the file now has neither. The parity runner's pre-fold branch calls the seam with two arguments, which is correct only on a tree whose signature still takes the wire list; against the signature this tree has it reads as an arity error to a checker and to a reader. Splatting a named tuple states that the two-argument form belongs to the other world. |
||
|
|
df8a374c3d |
test(session): cover the orphan guards in the streaming arms
Branch coverage showed the supersession guard in the Exception arm never executed and the one in the Ctrl-C arm only ever took its live side. The reason is structural rather than neglect: the ladder converts supersession before these arms can see it, since _model_turn_with_retry re-checks the generation ahead of classifying a death, so on every deterministic path an orphan's failure arrives as GenerationCancelled. The guards exist for the sub-statement race where a force-cancel lands after that check — the same accepted window the cancel ref documents — which no scripted stream can reach. These drive the seam directly to simulate it: the attempt arms, a newer generation claims the session, then the failure surfaces. They pin what the guards protect — an orphaned thread emits nothing, because the successor generation is already streaming into the same UI — plus the live counterpart, where a Ctrl-C still finalizes the display. Deleting either guard, or inverting the Ctrl-C one, fails them. |
||
|
|
50e080c18b |
fix(session): one supersession predicate, asked the same way everywhere
Scoping the arm-duty gate left four sibling gates in the same streaming turn still comparing generations with a bare !=, so one function could reach opposite verdicts for one generation shape: a Stop finalized the display and stashed the partial where a Ctrl-C on the identical shape did neither. _generation_superseded() is now the single predicate and every site asks it — the cancel ref, the streaming consumer, the dead-partial promotion, the Ctrl-C arm, and the orphan arm. Each caller still performs its own read. That is the point rather than an accident: the consumer's read is a genuine second look after the ref's, and a consumer that delegated to the ref would inherit its stale answer and run the arm duties for an orphan — nulling the successor's usage slots and recording health for an abandoned lane. Tests: TestSupersessionVerdictAgreement pins that the arms agree, in both directions. Its orphan case pins the stronger invariant it turned out to hold — a superseded generation never reaches an arm at all, because the ref reads superseded and model_turn refuses to dispatch. The last two hand-rolled dataclasses in the suite are replaced by the real ToolCallDelta, and the prepare_wire docstring paragraph is re-flowed. |
||
|
|
4dd92d150b |
fix(session): scope the arm-duty gate the way the rest of the file scopes generations
The consumer's arm hook and cancel-partial recorder compared generations with a bare !=, while the ref that fires them treats generation 0 as UNSCOPED — so for a direct seam caller the ref armed and fired the hook and the hook refused to act. On a session whose generation had ever been claimed, that left the previous turn's usage in place as this turn's estimate and dropped the serving lane's health success. Both now ask the consumer's own _superseded(), which mirrors the ref's predicate, so the two halves of one decision cannot disagree. The which-errors-speak-for-the-backend policy gets one spelling (_speaks_for_backend over _NON_BACKEND_ERRORS) instead of a matching isinstance in each walk arm, and the length arm stops calling finalize_provider_blocks over an empty list only to discard the result. Tests: the fourteen hand-rolled FakeChunk dataclasses in the cancel suite are replaced by the real StreamChunk its sibling suites already use, so the fakes cannot drift from the shape production emits. |
||
|
|
5de54147e1 |
fix(832): a prep fault walks the fallbacks it can no longer speak for
Making prepare_wire lane-variant invalidated the premise behind the walk-abort on WirePreparationError: with the fold posture following each lane's capabilities, a preparation fault on one lane no longer implies every lane fails, so aborting the walk skipped healthy fallbacks and the dedicated fatal message was wrong on both of its claims. Preparation faults now keep their no-health rule on every lane but continue the walk — the primary's fault enters it and a fallback's fault yields to the next alias — and the fatal message drops the no-fallback claim. Riding cleanup: the self-surfacing exception pair gets one spelling for the re-issue mask (_SELF_SURFACING_ERRORS; the walk arms stay per-class because auth aborts where prep continues); the tag-scan gate gains a capabilities-shaped form (caps_scan_inline_reasoning) that the lane form delegates to and the title peel now uses, retiring the third spelling; the three streaming provider fakes build on one provider_shell; a comment in session_ui_base names the module function that replaced the deleted session delegate; close_run spells its carry cut as removesuffix; and the prepare_wire docstring paragraph is re-flowed. The walk-continues and per-lane no-health pins are mutation-probed. |
||
|
|
c906776efd |
fix(832): the serving lane's capabilities reach the wire fold
The per-attempt prepare_wire closure folded mid-conversation system turns with the PRIMARY binding's capabilities on every lane, so a fallback whose chat template rejects non-leading system roles failed on the self-inflicted wire shape and burned its own health record — the wrong-dialect class the walk's binding snapshot guards against elsewhere. model_turn now passes the serving lane to prepare_wire, and the session's closure folds with that lane's capabilities; callers without a lane in hand (the token-table re-fold) keep the primary default. Pre-fold prepared once with primary caps for every lane, so this is a named improvement, not a parity break. The arm-duties hook rode the same unguarded two-statement supersession window the _CancelRef docstring accepts only for the stream register: a force-cancel claiming a new generation between the superseded read and the hook let an orphan's late registration null the successor's usage slots and record spurious creation health. on_stream_armed now generation-gates itself, shrinking the accepted window's harm back to the register-only class. Test hygiene: the two overflow-compact tests are one parametrized body; arm_session mints a fresh ArmedHandle per create (provider.handles, _armed_handle = latest) matching the one-handle-per-create rule of real adapters. The duplicate sanitize pass stands as designed (accepted for wire parity); its perf note rides #979. All three product fixes are mutation-probed. |
||
|
|
90e55f92ca |
docs(832): shorten the branch's comments to their constraints
Comment-only sweep over the diff's prose: origin archaeology, next-line narration, and review-thread talk go; each surviving comment states the constraint the code cannot show, re-wrapped to the file's width. The ruled-behavior restatements in the parity transforms and the contract docstrings (eager append, cancel-predicate pairing, carry ownership, the plant call's carve-outs) keep every named invariant. |
||
|
|
06ec1a8629 |
fix(832): the boundary carry belongs to the run owner
The mandated cross-lane interleave angle found the two residual holes in the reasoning-boundary close: the close was gated on not-in_think, so an open inline think block at the boundary never closed and the later state flip relabeled held chain-of-thought as displayed ANSWER text; and the carry parked in the splitter's own pending was re-read under whatever state later flushes hit, relabeling a content-state tail as reasoning. close_run() now closes unconditionally (as the drain does) and RETURNS the partial-tag tail; the consumer owns the carry in a state-immune slot mirroring the drain's separate variable — re-fed when content resumes so a split tag still reassembles, flushed as content at tool, finish, and cancel boundaries, and included in the partial-content rule. The trailing citations footer is now HELD and folded once at stream end over the full answer — structurally the drain's post-loop fold — instead of folding at arrival, which diverged from the commit whenever a lax gateway emitted content after finish. Two non-mirror fixes: the fallback-failure UI line carries the exception class only (its text can embed a credential-bearing base_url; detail goes to the server log, same rule as the re-issue log arm), and a never-armed Stop (creation window, no prior death, zero tokens) writes NO assistant row again — restoring pre-fold semantics; a marker-only row would replay to the model as context on every later turn. Armed zero-token Stops still record their marker. Hygiene riding along: the parity runner zeroes the ladder backoff (the exhaust scenario was sleeping 3.2s of real backoff per suite run, with the retry-notice transform strings updated in step); test_session's porting docstring points at the helper's real module; test_cancel and test_session wrap the shared session factory instead of re-implementing its defaults; arm_session's armed handle is an ArmedHandle with real closed state instead of a MagicMock that satisfies any assertion; and send() derives the tool-call list once for both the persisted mirror and the executed set. All fixes are mutation-probed: re-gating the close, discarding the carry, dropping the promote gate, unredacting the fallback line, and restoring the arrival-time fold each fail their pins. |
||
|
|
6212783e23 |
fix(832): close the content run at a reasoning_delta boundary — display must mirror the drain
Live-caught on a deployed review exercise: the consumer's reasoning_delta arm flipped the splitter's in_think with a buffered content tail still pending, so a flush while in-think (stream finish, tool boundary) relabeled that tail as reasoning. The drain closes each content run at the same boundary, so the committed turn kept the tail as content — display and commit diverged. Worst case: a short answer followed by trailing reasoning displayed as NOTHING while the commit carried the answer plus its citations footer (the display-side blankness gate saw empty content and dropped the footer too). Pre-fold, display and commit came from one continuous splitter and both lost the tail; the fold's drain corrected the commit, leaving the display behind. ThinkTagSplitter.close_run() now closes the run exactly as the drain does — decided text emits at the current state, only a possible partial-tag tail carries into the next run — and the consumer calls it before entering the reasoning phase. This also heals the cancelled- partial rule in the same window, and covers the content-reasoning-tool sequence interleaved-thinking lanes emit. Riding contract fix: partial_tag_tail required only startswith, so a complete <reasoning>/<think> self-matched as a "partial" tail and the drain carried a finished open tag across the run boundary, relabeling the next run. A partial tag is now a PROPER prefix, per the function's own documented contract. Pins: TestDisplayCommitMirror (displayed content must equal committed content across six reasoning-interleave scenarios — the combination the replay-parity grid never scripted), TestPartialTagTail contract rows, TestCloseRun unit pins, and three new interleave rows in the splitter CASES table. Both fixes are mutation-probed: disabling close_run or restoring the self-match fails the pins. |
||
|
|
aa4371ea99 |
fix(832): retire the dead attempt's armed state in the re-create window
Between a mid-stream death and the next begin_attempt there is no live attempt, but the consumer kept the dead attempt's armed _CancelRef: a Stop in that window re-emitted the discarded splitter carry as fresh content behind a duplicate stream_end, and a walk-preamble failure was classified as another armed death, replacing the operator-actionable stream-death error. end_attempt() now pronounces the attempt dead at partial-capture; the consumer gains a single per-attempt initializer (_reset_attempt), a lane-free constructor (one resolve_lane walk per turn), and a saw-chunk classifier fallback so a never-arming adapter's mid-stream death still classifies mid-stream instead of silently double-rendering the same lane. Wire-preparation failures are typed at the seam: model_turn wraps prepare_wire raises in WirePreparationError, both walk arms forward it verbatim (no health record, no fallback walk — a session-data fault would otherwise paint every backend degraded), the fatal formatter gets a dedicated branch, and the re-issue ladder's last-death mask exempts it alongside BackendAuthUnavailableError so an auth outage mid-turn is not misdiagnosed as a network flap. Riding fixes: the tag-scan gate gets its single spelling (lane_scans_inline_reasoning) shared by drain and display; the citations fold's separator+gate become a shared pair in _protocol; _build_main_lane stops passing config_store (dead derivation — the session's own knobs replace both values it feeds); the debug wire dump is ruled per-invocation (the overflow-recovery re-print is the dump that diagnoses the recovery) and pinned; dead delegates _ensure_tool_call_ids and _finalize_provider_blocks deleted; the parity runner adapts to the pre-fold seam signature by inspection and refuses to record a harness-shape TypeError as a baseline; the streaming provider fakes move to tests/_session_helpers (their tree-wide home) and test_cancel's duplicate helper is deleted; committed parity pins restate their rulings in full; architecture.md's circuit-breaker section is replaced by the real passive health-tracker story and the send-flow diagram stops attributing tool-call assembly to the display consumer; stale pre-fold names and ragged comment paragraphs cleaned. New pins are mutation-probed: disabling end_attempt, the saw-chunk fallback, the auth exemption, or the WirePreparationError arm each fails its pin. |
||
|
|
5a20916eec |
docs+test(832): retire pre-fold names from prose; pin the eager-append contract
The docs sweep re-points every stale reference to the deleted seam (architecture.md's flow diagram and ladder inventory, the lowering and anthropic docstrings, the protocol's shared-rule docstrings that described the pre-fold dual-assembler world). The Protocol's cancel_ref contract is strengthened from 'before the first chunk' to 'inside the call body, before the iterator is returned' — the instant the fold's creation-vs-midstream classifier and health recording key on — and three real-SDK-over-mock-transport tripwires pin it per adapter, so a future lazily-issued generator adapter fails loudly instead of silently reclassifying every pre-first-chunk death. |
||
|
|
e103af94e7 |
test(832): port the seam-coupled suites to the folded architecture
Seventeen files, ~1,300 tests, re-pointed or redesigned per the triage ledger's recipes: wholesale turn-scripting moves to ModelTurnResult fakes; streaming-behavior suites drive the REAL wrapper+consumer+drain path through armed provider fakes (tests/_parity_832.arm_session — the eager cancel_ref append every real adapter performs, exception elements for creation-phase failures, sequential per-turn scripts, and the title lane quieted: a provider-level fake otherwise loses its one-shot script to best-effort title generation, which is why the old tests patched at the session level); kwarg-capture suites assert through model_turn's create_streaming call with system-prepend-aware index math; delegate wrappers retired by the fold re-aim at their model_turn module twins. Old-architecture pins are replaced by their new-world equivalents rather than deleted: no shared cancel ref exists (pinned), the handle slot and per-attempt refs carry the cancel surface, the retry gate reads the serving lane's provider, and a superseded generation's death exits send silently as cancelled — a named delta: no arbitrary exception class escapes an orphaned thread anymore. Full suite: 10651 passed, 10 skipped. The wire-payload goldens pass untouched — the fold's lowering composition is byte-equivalent on every provider's request path, as designed. |
||
|
|
58b24de7e6 | test(832): re-aim the extra-params gate pin at the module function (delegate wrapper retired) | ||
|
|
0df9f6e2d4 |
test(832): port test_cancel to the folded seam; add hook + orphan + pre-dispatch pins
Provider-level armed fakes drive the REAL wrapper/consumer/drain path (the seam these tests exist to pin), with title generation quieted — the best-effort title lane consumed one-shot scripts once fakes moved to the provider level. The shared-ref architecture pins become their new-world equivalents (no shared _cancel_ref attribute; _cancel_stream lifecycle via the eager append), and three new pin classes land: on_first_append fires once and never for a superseded arrival; a force-cancelled generation's mid-stream death is never re-issued and touches no UI finalize; a pre-set Stop issues no request and mints no credential on a dynamically authenticated alias. |
||
|
|
ecf14dc001 | test(832): make_result helper for the triage's patched-result recipe | ||
|
|
2e18d159a3 |
feat(session): fold the main streaming loop onto model_turn (#832)
The send path's plant call is now one model_turn invocation per attempt, reached through a lane-swap fallback walk that mirrors the old creation ladder 1:1: an inner per-lane retry (_model_turn_with_retry) inside the two-pass healthy/degraded walk (_model_turn_with_fallback), with health success recorded at the request-accepted instant via the per-attempt _CancelRef's new on_first_append hook and failure once per lane ladder. The hook is also the creation-vs-midstream classifier: an armed attempt's death re-raises to the re-issue ladder on every lane — a fallback stream that died after tokens reached the UI is never swallowed into try-the-next-alias — and carries the per-turn usage-slot resets at the old timing so a reconnecting tab's status bar never blanks mid-walk. Chunk-to-UI translation lives in _StreamTurnConsumer (model_turn's on_chunk body): display-side only, the canonical turn always assembled by drain_stream at the one seam; the inline-tag scan reads the SAME lane capability the drain gate reads (server_parses_reasoning), replacing the creation-time handoff register — which is deleted — so display and commit cannot disagree about a backend's posture, fallback walk included. Cancellation converges: every model-call site now builds fresh generation-scoped refs, closing the force-cancel hole where the old gen-0 shared ref read aborted=False for an orphaned generation and would have let a retry re-issue on its behalf; the pre-dispatch abort read inside model_turn also means a Stop set before the turn no longer mints a credential on a dynamically authenticated alias. send() consumes the result natively: the committed Turn carries minted ids, the finalized native lane, and an accurate producer — fixing the latent mislabel where fallback-served turns were persisted under the primary provider's name, and the fork asymmetry where in-memory turns decoded with producer="". Ruled behavior changes (design D12): the trailing citations footer now folds into committed content (it previously lived only in an ephemeral info bubble and vanished on reload); a stream that exhausts without a finish reason is a retryable mid-stream death instead of a silent partial commit; length-truncated turns keep dropping partial tool calls, now as an explicit post-drain policy. The replay parity harness pins all thirteen scenarios against pre-fold baselines, transformed only where a ruling applies — and caught two real bugs during the fold (the splitter's end-of-stream carry never flushing to the UI, and the footer splicing into the answer's held tail). ChatSession imports no provider module: create_streaming has exactly one caller module, and the protocol types, merge_usage, and create_provider reach the session through model_turn's re-export seam. |
||
|
|
1f3b89610a |
test(832): replay-parity harness + pre-fold baselines
Thirteen scenario scripts drawn from the chunk-field-to-UI grid, each driven through the streaming seam against a scripted provider fake that arms cancel_ref eagerly (the classifier the fold introduces distinguishes creation-vs-midstream failures by that arming, so the fake must mirror the real adapters' eager append). The captured records — ordered UI events, committed-message projection, mid-stream usage, raised class — are the OLD-WORLD baselines: this commit's session.py is byte-identical to main, which is what makes them the record. The assert path applies only the behavior deltas the design table rules, each transform citing its row; a difference outside a ruled transform is a fold regression. |
||
|
|
70165807c7 |
fix(reasoning): close the unmarked chain-of-thought leak, gate the tag scan by backend (#940) (#978)
Some serving setups emit model reasoning inline with no think tags and no
reasoning_content at all — nothing any parser can segregate (measured live
on the dev vLLM: 20/20 sampled completions, streamed and not, proxied and
direct). The drain seam correctly passes unmarked prose through, so it
became the artifact on every bounded-artifact lane: workstream titles
("Thinking Process:"), compaction summaries that were ~90% chain-of-
thought, and the web-fetch tool results #940 reports — which then ride
every following turn as context.
Three coordinated changes:
* Utility lanes ask for no reasoning. _utility_completion (title,
compaction, web-fetch extraction) pins the alias's declared thinking
toggle off and withholds every reasoning-effort channel — the relayed
session knob, the lane rung, the definition default, and the graded
template key — via lane_without_thinking / lane_thinking_suppressed,
the same suppression omni transcription already used (now shared as
thinking_off_template_kwargs). Measured end-to-end: the extraction
that returned 3.7k chars of reasoning returns a 258-char answer.
* server_parses_reasoning capability. A backend that segregates
reasoning into its own channel declares it, and the inline tag scan
turns off on every lane: the drain seam, the interactive splitter
(which now reads the ACTIVE stream's capabilities via the creation-
time handoff register, never the primary alias's), and the title
lane's cosmetic peel — so prose that merely quotes a tag can no
longer be misrouted, and the utility suppression stands down where
reasoning costs the artifact nothing. The built-in commercial
capability tables declare it wholesale (known models and table-miss
defaults); local compat lanes keep the passthrough default the scan
exists for. Bool-typed capability overrides coerce string spellings
instead of truthiness-flipping on hand-edited JSON.
* Title selection follows the prompt's contract, not line position:
the last line within the word cap that ends in a word character —
rejecting explanation sentences, sign-offs, parentheticals, and
reasoning headings in any script (terminal punctuation carries
unspaced scripts where whitespace word counts are meaningless) —
else the last non-empty line. 20/20 captured live responses title
correctly (9/20 before, unchanged since well before the seam
unification: the old and new pipelines scored identically on every
sample, so the regression source was the backend's output shape,
not #965).
Also folded in from the review round: a think tag split across a
reasoning-delta boundary reassembles in the drain (partial-tag tail
carry; tool boundaries still flush), Turn.text joins text blocks with a
newline so multi-block answers stop fusing words in notification bodies
and every flattened read, the notify hook reads final_assistant_text
directly instead of through a one-line shim, web-fetch extraction uses
the shared _non_blank_or fallback, and the judge/output-guard suites use
real ModelCapabilities instead of truthy mock attributes.
Closes #940.
|
||
|
|
14df09a107 |
chore(deps): update vendored js (#974)
* chore(deps): update vendored js * chore: download vendored JS files --------- Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
7076bcf6ef |
fix(session): never dispatch a model call on an aborted cancel_ref (#972) (#976)
* fix(session): never dispatch a model call on an aborted cancel_ref (#972) model_turn consulted cancel_ref.aborted before re-issuing a request after a mid-drain transport death, but never before dispatching one. A caller whose call had already been abandoned — the user hit Stop, or a deadline fired — still lowered its turns, resolved its credentials, and put the request on the wire; the provider registered the stream handle, the ref closed it, and the client discarded a reply the endpoint had already begun producing. The rule was half-present at the seam: don't resurrect an aborted call was enforced, don't start one was not. The predicate is now read before each dispatch through one helper, using the same duck-typed getattr the drain-retry gate uses, so a None ref (perception, title generation, sub-agents, optimizer, eval) and a plain-list ref both stay legal. Two reads, because they buy different things: the entry read skips the lowering and the credential resolve for a call already abandoned when it arrives, while the read immediately before create_streaming is the one that keeps bytes off the wire — a blocking resolve is exactly the window the entry read is too early to see. Cancellation stays cooperative and the docstrings say so: a mint already under way completes, and an abort arriving after the last read still reaches the in-flight call through the ref's own close paths (append for a handle that has not arrived, abort for one that has). The raise is DeadlineCancelledError, the deadline module's abandonment vocabulary. GenerationCancelled would be invisible to the except-Exception arms surrounding these calls, and it lives in session, which imports this module; compaction performs the translation itself, its handler re-checking the session before it reads the error, which is what keeps a Stop mid-summary off the red-error path. That translation holds only while _CancelRef.aborted and _check_cancelled stay the same predicate over the same generation, now recorded on the property that owns it. The raised message deliberately avoids context-window vocabulary: _is_ctx_overflow classifies unrecognized error classes by text, and an overflow reading would send the compaction lane subdividing and re-issuing the very calls this suppresses. The pre-existing abort test keeps its subject, the re-issue gate: its ref now aborts after dispatch, and it asserts that no retry was announced rather than counting calls, which is what separates that gate from the post-backoff one. Two siblings pin the new reads — the resolver is never called for a ref aborted on arrival, and an abort landing inside the resolver still reaches no wire — and a third pins the message against the overflow classifier. * docs(session): disambiguate the abort helper's resolve wording "The credential resolve between them is NOT re-checked" reads as though no abort check follows the resolve, when the second read sits immediately after it — the sentence meant only that nothing interrupts the resolve itself. Left as-is it invites a refactor to delete that second read, which is the one that keeps bytes off the wire when the abort lands mid-mint. States both facts separately now: the mint completes regardless, and the second read is what turns such an abort into a skipped request. |
||
|
|
0150523bb9 |
test(session): pin the both-vocabulary title peel
The title lane's cosmetic peel walks the close-tag vocabularies in sequence, which review read as a double peel that could discard title text between a `</reasoning>` and a `</think>`. It cannot: the remainder of the first cut begins after the last `</think>`, so a `</reasoning>` still found in it is necessarily the later tag — the sequence is equivalent to one cut after whichever close occurs last (verified exhaustively over tag/text arrangements and 200k randomized fragment strings). The equivalence was unpinned, so both orderings join the variants table and the docstring records why the sequence is a single logical cut. |
||
|
|
bc3fa60011 |
fix(providers): segregate inline reasoning at the drain seam
Passthrough servers (parserless vLLM/llama.cpp, LM Studio, bare gateways) emit reasoning as literal <think>/<reasoning> blocks inside content, and only three of nine drained lanes stripped them: web_fetch tool results persisted raw think blocks into every following turn (#940), judge verdicts parsed through tag noise, and a draft verdict inside a think block could shadow the real one at the output guard. One rule at the seam now. drain_stream accumulates content in RUNS bounded by interleaving signals (provider-parsed reasoning deltas, tool-call deltas) with the interactive consumer's within-chunk ordering — reasoning, then content, then the tool-call close — and splits each run through split_inline_reasoning, the one-shot form of the interactive lane's ThinkTagSplitter: a pure raw split, exactly equivalent to the streaming form on every catalog case. One trim policy exists and the drain owns it: blank edge lines are trimmed once over the joined runs when a tag was consumed, so tag residue dies at the edges while genuine inter-run paragraph separators survive. Extracted text is appended to result.reasoning after any server-parsed reasoning with a blank-line boundary and rides the native lane as the reasoning_text synth block. Orphan CLOSE tags deliberately pass through byte-identical: a close whose open never arrived is indistinguishable from prose QUOTING the tag, and drained lanes routinely quote third-party text — reclassifying would let a malicious page containing the literal tag destroy the extraction that cites it. The title lane keeps a local rfind peel as display-string formatting. The citations footer folds only onto non-blank content — sourcing for an answer that does not exist is dropped rather than handed to emptiness checks as a footer-only "answer". Every private strip is deleted: the title lane's strip, the summarizer strip, _strip_reasoning itself, and the optimizer's five regexes (_strip_markdown_fence is now the one fence rule, applied to normalized model output only, never to or-fallback values). Think-only and whitespace-only responses drain to blank content, and every lane's no-answer fallback gates on blankness: web_fetch returns an honest extraction-error card, the intent judge takes the empty-retry ladder, the task-agent synthesis reports "(no output)", and the optimizer keeps the current observer system and prompt verbatim on no-answer passes. Final-say reads (optimizer analyst, eval final_content, the notify hook) use trajectory.final_assistant_text — the last assistant turn only, never an earlier narration presented as the conclusion — while last_assistant_text is the salvage walk (task_agent partial-work recovery), skipping tool-call-only, all-reasoning, and whitespace-only turns. Perception memoizes every completed description immediately, including an empty one — one perceive per key, ever — under a commit-lock guard so an empty result never overwrites a concurrently memoized real description; an all-reasoning perception model pins the placeholder until restart, and the remediation is server-side (a reasoning parser or the template thinking toggle on the perception alias). A true double-reasoning shape (inline-extracted text alongside a native reasoning block) logs chars-only at the drain, where it is distinguishable from the routine reasoning_delta mirror. The dialect's semantics are pinned as one table (tests/_reasoning_dialect.py) driven through shared fixtures (think_tag_stream, seam_provider): one-shot conformance, the exact one-shot/streaming equivalence property, the drain seam rules including quoted-tag safety, run-boundary and separator-preservation pins, per-lane pins for all nine lanes, and the empty-content assistant wire shape. Closes #965. Closes #940. |
||
|
|
1d7db73305 |
fix(models): review feedback — separator vocabulary, constraints stub, import style
The scopes sanitize now shares the registry guard's separator vocabulary: tab/newline/CR read as spaces, and every other C0 byte — including the U+001C–U+001F block str.split() would silently promote to separators — strips like the control it is, so a control byte inside a token can never split it into two valid-looking scopes (pinned alongside the registry's refusal). The livepass auth-constraints stub serves the new app_identity_auth_modes field so the pass exercises the served-data path for the model list's auth badge, and the session-module import in the mint tests drops to the string-path monkeypatch spelling (single-style imports). |
||
|
|
8605c9783d |
feat(models): rfc8693_obo auth mode, per-alias exchange scopes, identity-keyed mint cache
Adds the dedicated `rfc8693_obo` model auth mode (#955): model definitions gain an `obo_scopes` column (migration 069), the mint threads the scopes to the token-exchange leg (RFC 8693), and every dynamic mode pins its grant leg — a mode is a dialect commitment, not a hint the deployment profile resolves. Exchange-capable IdPs refuse an audience whose scope was not requested; this closes the structurally unmintable model-OBO path on token-exchange deployments. The model mint-cache is identity-keyed on the owning definition's alias (`__model_obo__:<alias>` per user, `__model_app__:<alias>` under the shared app principal), matching the MCP discipline where rows key on the unique server name. The bearer's shape lives in the row's audience/scopes columns and the freshness gate compares it on every read, so a re-aimed alias refuses its old row and overwrites the same key in place. Admin lifecycle (rename, re-aim, scope change, delete) purges a definition's own rows through one shared helper — sound because one definition owns each key; a sibling's rows are untouchable by construction. Cooldown and backoff additionally key on the dispatch shape, so an operator's config repair is an instant clean slate. Cause records, cooldowns, locks and memoization are per-alias end to end, and the session heartbeat reads refusal causes under the same keys. Console: default-deny write gating for dynamic rows (value-diff over the full column ladder, admin.mcp escalation, a never-blockable pure-disable carve-out), a two-tier validator (audience allow-list on every write; deployment-posture checks when the pair is chosen), one shared scopes parser whose omit-unchanged arm keeps over-cap DB-direct residue rows disarmable without ungating real changes, and served constraints (dynamic/scopes/app-identity mode lists, mode-to-profile pairing) so the shelf tracks the registry by data. The admin shelf gains the mode option, a scopes input with residue affordances, pairing-aware option greying, and a derived auth badge. Registry load refuses control characters in alias, audience, and scopes — including the C0 separator block that str.split() would silently collapse — and the C0/DEL class has one exported spelling shared by every surface. Profile-mismatch visibility warns at reload and boot with the mode-correct cause, gated on OIDC being enabled. Breaking: a stored `entra_obo` alias on a deployment whose `[oidc] obo_grant_profile` is `rfc8693` (or the inverse pairing) no longer mints via the profile-driven overload — the mint refuses before any IdP traffic with cause `grant_profile_mismatch`, and the `model.auth_fail_closed` policy governs static fallback. Such rows never minted usefully on scope-gating IdPs; the shelf now surfaces the pairing and the per-turn heartbeat names the refusal cause. Live-verified end to end: scoped token exchange mints, the warm cache serves with zero IdP calls, and the mode/profile mismatch refuses with zero IdP traffic (scripts/obo-e2e/keycloak_e2e.sh); the refresh-redemption profile's E1-E7 hold via scripts/obo-e2e/entra_e2e.py. Closes #955. |
||
|
|
7776cc0c2f |
fix(streaming): probe on_stream_discarded for pre-existing UIs and format the hoisted fake
PR feedback round: - on_stream_discarded now follows on_compaction's compat pattern for a hook added after UIs exist in the wild: the protocol member carries a REAL no-op default (an explicit subclass inherits a correct implementation — a UI without server-side turn buffers has nothing to truncate), and both call sites route through a getattr probe, so a duck-typed UI predating the hook degrades to no-truncate instead of raising an AttributeError from the very arm that is handling a stream death — which would replace the wire failure with the attribute error in the retry gate. Pinned with a hook-less-UI retry test. - tests/_session_helpers.py gains the formatting pass the RecordingUI hoist bypassed (the CI lint failure). |
||
|
|
1f9f462b66 |
fix(streaming): gate the dead-segment discard on the backoff surviving the Stop window
Fifth review round — four small correctness edges, none in the retry semantics: - The server-buffer discard now runs only AFTER the backoff survives a Stop: a cancel during the window persists the promoted partial to history, and the idle-state payload (drained from the turn buffer) must carry the same text — discarding first rendered the cancelled turn empty on the dashboard while the transcript had it. Pinned with a real-buffer test; the spinner and fresh segment watermark follow the truncate so a later discard cannot resurrect the dead segment. - stream.retry's dead_content_chars reports THIS death's flushed text only — the Stop-preservation carry retains the previous attempt's partial by design, and logging its length re-attributed the same discarded spend to consecutive retry lines. - The changelog entry for the post-finish-blip rename no longer claims the usage_captured field was dropped; it is emitted and pinned. - The retry suite's module docstring states the shipped finalize contract (stream_end + backoff-gated stream_discarded, never turn_committed) instead of the superseded pair. - RecordingUI is hoisted into tests/_session_helpers next to NullUI — this branch already paid the per-file-fake tax once when a protocol method grew — and a stale deferral sentence is dropped from the fatal-formatter comment. |
||
|
|
961a2017dc |
fix(streaming): delete the retry window's shared slots and gate the send epilogue
Fourth review round. The recurring defect family — cross-frame session slots racing an orphanable window — is removed structurally instead of gated again: - The wire-fold slot is deleted. The fold the stream was actually created from rides the returned message dict on the underscore lane (like _provider_content) and is popped at the single calibration site before commit, so a superseding generation can never alias it and there is nothing left to clear. Plain-dict test fakes fall through the pop to the frame-local fold. - The stream-provider slot is demoted to a creation-time handoff register: _try_stream stamps it, _stream_response copies it into a frame-local immediately after each create returns, and only that local feeds the retry gate. The fatal formatter returns to the consistent PRIMARY identity triple — pairing a fallback's provider name with the primary's base_url and alias sent operators to debug the wrong backend; stamping the full producing identity is #964. - send()'s epilogue is generation-gated: a superseded thread's escaped death no longer records a fatal error over the healthy successor turn (error banner, buffer-wiping error-state drain, wrong last_error for the coord), and a Ctrl-C on an orphan no longer mutates history. - The terminal arm discards as well as finalizes. Keeping the buffers bought nothing — the fatal path's error-state drain wipes them on every server lane — and the skipped discard let a mid-consumption overflow recovered by compact-and-retry concatenate the dead attempt's text with the recovered answer in the idle payload. Pinned with real-buffer tests for the overflow-recovery and orphan-epilogue paths. - stream.post_finish_blip regains usage_captured, tracked by transport_guarded from the chunks it forwards, restoring missing-spend attribution on both lanes. - TerminalUI.on_thinking_start is idempotent at the callee (a live spinner is stopped before being replaced), removing the caller-side stop-first dance and the leak the next unaware call site would have reintroduced. - The think-tag vocabulary in _strip_reasoning and the title lane is derived from ThinkTagSplitter, closing the drift channel that would leak raw reasoning into compaction summaries and titles. - on_stream_discarded's docstring states the true pending-batch semantics (defensive drop; the shipped sequence flushes via the preceding stream_end), and the live-suite recording fake gains the protocol method. |
||
|
|
476cce2e58 |
fix(streaming): scope stream bookkeeping to the send and discard dead segments server-side
Third review round on the retry window: two mediums fixed, one observability gap closed. - New UI-protocol method on_stream_discarded(): on_turn_committed clears only the inflight buffers — it cannot clear _ws_turn_content, the multi-segment buffer the IDLE payload drains, because earlier segments of a tool-looping turn must survive commits — so a dead attempt's text concatenated with the retried text in the dashboard's idle payload. SessionUIBase now truncates the turn buffer to a segment watermark (snapshotted in on_thinking_start, which precedes every stream segment), drops the never-displayed pending batch, and resets the inflight snapshot; the retry arm emits it in place of on_turn_committed. Server-side only — no SSE event, no client change; no-op on the CLI and eval UIs. Pinned with a real-SessionUIBase-buffer test: the recording fakes structurally cannot see this buffer. - _active_stream_provider and _active_wire_msgs are send-scoped: cleared in send()'s finally, after the except arms' fatal formatting (the one legitimate fatal-path reader of the provider field). A later fatal on a utility lane falls back to self._provider instead of wearing a stale interactive-turn binding, and the full-context-sized wire fold no longer outlives its calibration use. - stream.retry carries dead_usage and dead_content_chars: the abandoned generation's billed tokens are otherwise invisible (the wire reports usage only at stream end — Anthropic's early prompt tokens arrive, the OpenAI chat lane's usage chunk trails the finish), so the log line records what the wire delivered plus the discarded completion's char count for spend reconciliation. |
||
|
|
47524654b3 |
fix(streaming): close the retry window's generation, identity, and masking holes
xhigh review round on the mid-stream retry ladder: 14 verified correctness findings, all fixed, plus the verified-but-capped cleanups mined from the review run. Generation safety — the shared-slot class is removed structurally, not gated per site: a dead attempt's partial now rides the raised exception (thread-private by construction) into a wrapper-local variable, and the _midstream_dead_partial session slot is deleted, so an orphaned superseded generation cannot poison a live generation's preservation. The promotion helper is generation-gated, writes the marker row even for a pre-token death (empty content takes the marker-as-message branch), and backfills a recorded-but-empty partial with the previous attempt's text, so a Stop anywhere in the retry window — backoff, re-create, or TTFT wait — preserves the latest text the user actually saw. _record_cancelled_partial is generation-gated too: a superseded thread touches neither the UI nor the shared slot. Identity — the retry gate and the fatal formatter now consult the provider that actually owns the live stream (recorded at creation, covering the fallback walk by construction), so a fallback stream's provider-specific transient is retryable by ITS OWN contract and failures are labeled with the binding that produced them. The mid-retry rebind check compares the full (client, model, provider) binding — reload() keeps the pooled client on model-only swaps — and a re-prepare also re-exports the wire fold that send()'s token-table calibration counts. Masking — a context overflow raised by the mid-retry re-create surfaces as itself so the compact-and-retry arm can recover the turn, and the overflow arm is split: recovery-machinery failures still surface the original overflow (its wording anticipates them), while post-compaction consumption failures surface as themselves instead of a false overflow diagnosis. Cancellation and terminal paths — a Stop that races the trailing-metadata window is re-checked after the chunk loop, so the turn aborts with the marker instead of committing and running its tool calls; the terminal arm finalizes client-side only, deliberately keeping the in-progress snapshot (the unpersisted partial's only copy) for refresh-replay; KeyboardInterrupt gets the same client-side finalize; the retry arm stops the spinner before restarting it (the CLI's on_thinking_start replaces the spinner without stopping it — a thread leak); and the backoff delay is computed from the pre-increment index, matching the sibling ladders' convention. Mined cleanups: the retry suite wraps the shared session factory instead of duplicating its defaults; the usage projection uses dataclasses.asdict; the partial-content rule lives in one closure serving both preservation paths; the two fatal-log tests are parametrized into one; the test import uses the public providers package. |
||
|
|
df81035302 |
fix(streaming): finalize dead attempts on terminal paths and harden the retry window
External-review round on the #937 branch; four confirmed findings fixed, each on a failure path the retry loop itself introduced or made reachable: - The terminal arm (retry exhaustion, non-retryable death) now finalizes the dead attempt with the same stream_end + turn_committed pair the retry path emits, so the last attempt's partial is flushed in every consumer — the CLI was the exposed case (its markdown fence state resets only in on_stream_end; the server workers emit their own after a fatal, the CLI's direct send() does not). The finalize is gated behind the generation check: an orphaned superseded thread must not emit UI events over the new generation's stream. - A Stop landing in the backoff/re-create window now preserves the dead attempt's partial: the attempt stashes its flushed content (plus the content-state carry tail) on a non-cancel death, and the wrapper promotes the stash to the cancelled-partial slot before re-raising, so send()'s cancel handler persists it with the cancellation marker — the same disposition a cancel during the attempt gets. - The fatal-path debug trace logs frames only (format_tb): exc_info rendered the raw exception message, which can carry credentials verbatim — the exact leak the sanitize floor above it exists to hold. The recreate-failure warning drops exc_info for the same reason and logs the exception class name instead. - A mid-retry rebind that replaced the client re-prepares the wire messages against the new binding before re-issuing: the system-turn fold is capability-sensitive, and a registry reload that switched model family would otherwise re-send the old family's wire shape. The cross-thread close boundary pin now accepts ReadError or RemoteProtocolError: which one surfaces is platform/timing-dependent, and both are TransportError members of the stream-death set, which is the property the pin exists for. |
||
|
|
3b9de67e8c |
refactor(session): extract think-tag splitting into ThinkTagSplitter
The interactive chunk consumer's _flush_text/_drain_pending closure pair carried the partial-tag carry buffer and in-think state inline. The tag-scanning half moves to turnstone/core/streaming_text.py as a standalone ThinkTagSplitter (carry buffer, in_think state, earliest- index tag selection, MAX_TAG_LEN safe-flush); dispatch and accumulation stay in the session behind the emit callback, and out-of-band transitions (reasoning_delta path, tool-call starts, cancellation) read/write splitter.in_think and flush_pending() where they previously touched the closure locals. Pure move: table-driven pins covering partial-tag buffering across chunk boundaries, the safe-flush margin, open/close tag precedence, in_think transitions, and reasoning-vs-content dispatch were written against the closure implementation and pass unchanged against the extracted class — byte-identical emitted text, identical UI callback ordering. The session-level _THINK_*/_MAX_TAG_LEN class constants fold into the class. |
||
|
|
a1dfe0bd4f |
refactor(streaming): dedupe transport conversion, usage merge, cancel finalize
Three behavior-preserving consolidations behind the #937 fix, each deleting a hand-rolled twin of a now-shared rule: - drain_stream consumes transport_guarded(chunks) and drops its inline `except httpx.TransportError` arm — one conversion rule for mid-body wire deaths across the drained and interactive lanes. The post-finish tolerance now logs under the wrapper's `stream.post_finish_blip` name (formerly `drain_stream.post_finish_blip`) and no longer carries `usage_captured`; changelog notes the rename for external log filters. The possible usage=None result on a post-finish blip is documented on drain_stream itself. - _stream_attempt's hand-rolled per-chunk usage max-merge becomes a local UsageInfo accumulator folded through merge_usage (drain's rule), re-projected into the _last_usage dict on EVERY usage chunk — that dict has mid-stream readers (_estimated_prompt_tokens, the status line), so the per-chunk write timing is load-bearing and unchanged. - The twin cancelled-partial sequences in _stream_attempt's two cancel arms (cooperative GenerationCancelled, stream-close-converted) merge into one local _record_cancelled_partial helper carrying both arms' tool_calls/_provider_content omission rationale in one place. |
||
|
|
5fb27e8f81 |
fix(session): survive mid-stream transport deaths in interactive turns (#937)
A wire death during body streaming (ReadError on a TLS record failure, peer resets) surfaces after the request has already returned its stream handle, so neither the SDK's request retries nor the creation-time retry ladder ever saw it: the interactive turn died with a bare exception string, the partial output was discarded, and no log trace was left. Utility lanes already survived this through drain_stream's normalization; the interactive loop now gets the same treatment. - transport_guarded() in providers/_protocol.py: drain_stream's transport-death conversion made reusable for consumers that keep streaming semantics. Pre-finish deaths raise the retryable IncompleteStreamError (drain's exact message shape); post-finish blips end the stream cleanly, forfeiting only trailing metadata. - The single-pass chunk consumer renames to _stream_attempt; _stream_response is now the resilient wrapper owning ALL stream acquisition plus a bounded mid-stream re-issue ladder (_MID_STREAM_RETRIES, the shared _stop_retrying predicate with a per-loop cap, cancel-aware exponential backoff). Send()'s overflow compact-and-retry arm now wraps the whole turn and passes re-prepared msgs explicitly. - A dead attempt is finalized across every UI consumer before the retry (stream_end then turn_committed then notice then spinner), so retried text never appends onto the dead attempt's in any surface (browser transcript, CLI markdown fences, Slack/Discord streamed messages, SSE replay ring). - Before re-creating, the session re-resolves its registry binding: a concurrent ModelRegistry.reload() closes cached clients, and the retry must not stream into the closed one. A failing re-create logs stream.retry.recreate_failed and re-raises the ORIGINAL stream-death error rather than masking it. - _format_backend_error gains a stream-death branch naming the provider, endpoint, and model, with a short identity-bearing first sentence. _BACKEND_STREAM_EXC_NAMES joins _BACKEND_KNOWN_EXC_NAMES, which also removes those names from _is_ctx_overflow's text-detection eligibility (deliberate: their texts are fixed transport strings that never carry overflow phrases). - _record_fatal_error now logs session.fatal.recorded (INFO for KeyboardInterrupt, ERROR otherwise) so fatal turns leave a journal trace. - _assistant_pending_tokens resets at stream entry so a post-finish blip that loses the trailing usage chunk cannot append the previous turn's completion count as this turn's estimate. Offline SDK boundary pins (openai/anthropic mid-body death identity and no re-request, cross-thread client close surfacing httpx.ReadError) guard the assumptions the retry gate rests on. |
||
|
|
1e34e19d48 |
refactor: single-style module imports and narrowed JSON body typing
Consolidates the repeated function-local model_registry imports onto one from-style module import per test file (the module object stays available for monkeypatching), converts the e2e script's mcp_oauth import to match, and reads the request body as Any before the isinstance narrow so the declared dict type is earned rather than asserted. Addresses the automated review feedback on the pull request; the two code-scanning flags are dismissed as false positives separately (the missing-key refusal log names config knobs and carries no secret value; the URL assertion is a test expectation, not a sanitizer). |
||
|
|
33ace975d2 |
feat(models): default-deny governance and admin UI for per-alias backend auth
Follow-up to the per-alias Entra OBO/app-identity backend auth: the console write path now applies default-deny field classification, the admin shelf gains full backend-auth support, and the session/registry rebind machinery is hardened for config changes landing under live sessions. Console write gate: - Default-deny classification: any non-neutral change to a row that is or becomes dynamic requires admin.mcp plus validation; the provably auth-neutral columns are enumerated (MODEL_AUTH_NEUTRAL_FIELDS) and a live-schema classification test forces every future column to be classified. The derivation is a pure function (_derive_auth_gate) with unit-pinned exclusivity invariants. - Two-tier validation mirroring the MCP oauth_obo validator: the row tier (audience allow-list) runs on every gated write; the posture tier (OIDC configured, token store present) runs on pair changes and on enable-arming. - Pure-disable carve-out: disabling a dynamic row is de-escalation and is never blocked — admin.models suffices and validation is skipped, including for rows with corrupt or skewed stored values. - Capabilities are compared canonically (key order, integral floats), the audience compare normalizes both sides, and staging an audience on a static row is refused on both write twins. - Calibrate writes the capabilities column under an enforced confinement invariant with a compare-and-swap persist. Admin shelf: - Backend-auth section with a per-open constraints fetch (GET /model-definitions/auth-constraints: audience allow-list, grant profile, dynamic modes), datalist audience suggestions, server-defined modes preserved on round-trip, and permission-aware visibility built on cache-skew-safe helpers shared through auth.js. - Refused live-registry swaps surface as an amber registry_warning on the write, delete, reload, and calibrate responses; audit rows carry auth_gated / auth_disarmed markers visible in the audit view. Registry and sessions: - The encryption-key requirement for dynamic auth is enforced inside ModelRegistry.reload() itself — nodes refuse with 503 and the console records coord_registry_error — and reload bumps the generation before the map swap so a racing reader can never pair a stale generation with new maps. - resolve()/resolve_binding() return the generation from inside the registry lock; sessions rebind per send on generation change with atomic client/provider/config commits, fallback-first handling of removed or unconstructable aliases, and judge/limiter resets only when the binding actually changed. - Mint refusals record per-user causes surfaced in the per-turn heartbeat logs; misconfiguration warnings are deduplicated with bounded state. Verification: 10417 tests (99 added on this branch), a 71-scenario browser harness over the real admin shelf, and a live rfc8693 token-exchange e2e run (MCP legs verified end to end; the model-leg scope gap is tracked as #955 under a narrow known-gap signature). Closes #950. |
||
|
|
9adde920d4 |
feat(models): per-alias backend auth via Entra OBO and app identity (#898)
Adds a per-alias `auth_mode` on model definitions so a model backend can authenticate to an Entra-fronted gateway with a per-request minted token instead of one shared static API key, letting the gateway attribute calls to the actual user or to the app as a machine identity. - `static` (default, unchanged) sends the stored `api_key`. - `entra_obo` mints a per-user On-Behalf-Of token for `obo_audience` from the caller's captured refresh credential. - `entra_app` mints an app-identity token via the client-credentials grant, and covers userless turns that OBO cannot. Reuses the existing OBO grant legs, refresh-token rotation CAS, cluster advisory lock and the `mcp_user_tokens` mint-cache, keyed under synthetic `__model_obo__:<audience>` / `__model_app__:<audience>` rows. The token binds at the call site through `client.with_options(api_key=...)` so each SDK emits it on its own auth path rather than through header injection. Migration 068 adds `auth_mode` and `obo_audience`. Both are additive and existing rows default to `static`, so behaviour is unchanged unless an alias opts in. Operator controls: `model.auth_audience_allowlist` is an exact-match allow-list that gates which audiences may be configured and denies all by default, and changing a mode or audience requires `admin.mcp`. `model.auth_fail_closed` decides whether a failed mint may fall back to an explicitly configured static key. A delegated call with no user, or a dynamic alias with no real static key, always refuses. Two changes here apply regardless of whether any alias opts in: - Storage and app state are now wired into the console MCP client manager. This fixes per-user `oauth_user` / `oauth_obo` dispatch for coordinator-hosted sessions, which previously raised `RuntimeError` on first call because `set_app_state` was only ever called on the node. - Unattended watch restores and `--resume` resolve the persisted workstream owner instead of constructing the session under an empty principal. A workstream with no owner is now a permanent refusal rather than an anonymous, auto-approved run. |
||
|
|
bb9684f505 |
fix(ui): block-copy dismissal listens on documentElement, not document
document-level mouseleave delivery on window exit is flaky in some engines, stranding the floating button until the next in-page pointer event; the <html> element receives the leave event reliably. |
||
|
|
729a02a833 |
feat(ui): copy-to-clipboard for messages and rendered blocks
Three idle-only affordances on every chat surface: a persistent copy button in each assistant bubble's actions bar, a pointer-only floating button over the hovered markdown block (fence, mermaid diagram, table), and Enter on a focused block for keyboard users, with the outcome flashed on the block itself. Copy resolves to SOURCE, not rendered text. The renderer stashes each table's raw markdown in data-md-source at render time — span sentinels restored in reverse mask order, footnote-definition bodies restored to raw before their recursive render — and whole-message copy reads the streaming pipeline's per-frame stash. The clipboard transport falls back to the legacy execCommand path for plain-HTTP LAN nodes, cloning and restoring the user's selection and focus. Outcomes surface button-local only: flash + title + one live-region announcement through the shared makeAnnouncer factory (also adopted by the interactive voice/tool announcers, whose lazily created regions swallowed their first announcement). Busy refusals answer with their own message. Coordinator retry and admin token-copy keep zero-module-dependency degrade paths. |
||
|
|
e526df95d0 |
fix(tool-search): discovery-failure records are per-user, rerank counts honest
Follow-up to #938; closes #941. The unavailable-server advisory fired for users whose own pool was warm: _pool_discovery_error was keyed by server name while pool connections are per-(user, server), so one account's failed prime rendered its exception text into every user's search results. - mcp_client: re-key _pool_discovery_error to (user_id, server_name). Written by the failing user's prime (single sanitize-and-cap pipeline shared with _set_error), cleared by that user's successful connect, retired with the grant on explicit disconnect / dead-grant convergence, and swept name-wide on registration lifecycle (removal, reconcile auth-type flips) via a snapshot-safe helper. Departed users' records are reaped by the eviction tick's orphan sweep — the single tick-side reaper; a live user's record survives its stub's eviction because the advisory has no mid-session re-record path. The eviction loop also starts on record write, so records written before any pool entry exists cannot outlive their users. Status reads scope to the requesting user, with an any-user view under the admin aggregate flag. - tool_search: _status_reason treats discovery_error as an outage only when the requesting user's own status is not connected — with per-user records this is belt-and-braces, since a successful connect clears the user's record. - session: the tool-search status snapshot scopes to the EFFECTIVE user (the acting participant on shared workstreams), matching the get_tools call that builds the search corpus, so an owner's pool state never renders into a non-owner's results. - bm25: with a reranker attached, matches ranked past the recall pool trail in BM25 order (reorder mode), so tool_search's "top N of M" count no longer floors at the pool size; the exception fallback is mode-aware (filter mode keeps its pool bound, byte-for-byte). |
||
|
|
c4b2dd7135 |
feat(tool-search): surface MCP discovery failures & honest result counts (#938)
Tool discovery for a pool-backed (oauth_user/oauth_obo) MCP server that is
down or 5xx-ing was invisible: the server contributed zero tools to the
catalog, so tool_search returned "No matching tools found" —
indistinguishable from a genuine no-match — and matches past max_results
were silently dropped with no signal.
- tool_search: search() ranks the whole deferred corpus and records the
pre-slice match count so format_search_results can report honest
truncation ("top N of M"). An optional status_provider lets results name
servers that are actually failing (open circuit breaker, recorded error,
recorded discovery failure) instead of masquerading as "no such tool".
Un-primed servers are deliberately not flagged, and a provider that
raises never breaks search.
- mcp_client: the previously swallowed pool prime/connect discovery
failure is recorded per server (single-line, bounded), cleared on the
next successful pool connect, on removal, and on reconcile-observed pool
removal or auth-type flips; exposed via get_server_status
as "discovery_error".
- session: wires get_all_server_status(user_id) into both
ToolSearchManager constructions as a lazily-called status provider.
|
||
|
|
57a9041941 |
fix(coordinator): the interjection handoff cannot lose the message, and the fact block is bounded
Review fold-in before push, twelve findings, two of them majors. The handoff popped the interjection queue destructively and handed the text to a send with a non-delivering refusal (the budget latch) and a preamble that can raise before the user turn is appended — a failure destroyed the user's words with a log line, after the charged wake nudges were already cleared. Now: the budget latch is checked before the pop (the message stays queued for a send with a human in front of it, and the wake drain still runs so the worker's exit converges); the pop returns the raw items and any non-cancel escape restores them verbatim — ids and priorities intact — before the failure surfaces; a cancel deliberately does not restore, because the Stop supersedes the queued words. Content-free items (a bare priority marker) are skipped at the shared renderer, so a lone '!!!' no longer buys a content-free turn at the cost of both nudges. The per-child fact block takes the roster formatter's bounds: fact lines cap at the display cap with a counts-only overflow line, and the wait slot keeps its larger handle cap — the body is a persistent system turn replayed on every request, and the block previously grew without bound as finished-but-unclosed children accumulated. The two fact sentences and the overflow line are named template constants, and every test assertion anchors on them; the children projection takes the same drop-never-mangle alteration check as the open-row fields. Eval world seeding: node metadata is JSON-encoded exactly as production writers store it (a raw string never matched a filtered list_nodes lookup), the stub client pins its heartbeat window open so a static world cannot go hollow mid-run, and the world-shape refusals' field branches gain their own tests. Comment accuracy and paragraph wrapping fixed at the sites the review named. |