mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-27 22:34:51 -06:00
cf2811fcebec9919db0701df83155eb4ea75bc97
68 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7a06f5e8bc |
refactor(session): make ModelLane the provider boundary (#979) (#989)
* refactor(session): make ModelLane the provider boundary (#979) ## Summary This closes the model-lane ownership gap left by #832: `ChatSession` no longer stores raw provider/client handles. `ResolvedModelBinding` now carries the provider, client, model, capabilities, registry generation, and backend-auth configuration as one coherent snapshot. - Atomically rebind existing sessions after model-registry changes while pinning each in-flight send, fallback, judge, output guard, task agent, title, compaction, perception, and voice operation to its initiating principal and binding. - Fence UI publication, canonical trajectory folds, durable writes, streams, retries, child scopes, and judge work by generation. Stop can hand off to a successor without accepting late state; cancelled tools retain typed effect receipts, and concurrent approval batches resolve by exact cycle or call. - Make create, fork, open, close, and delete race-safe with hidden `creating` reservations, incarnation-aware state tails, and an ACL-rechecked transaction that clones checkpoint-bounded history, configuration, project/persona state, and attachment references. - Extend REST/OpenAPI and Python/TypeScript SDK contracts for create/fork inputs, routed-create metadata, live-workstream probes, targeted approvals, and structured cancellation results. - Update architecture, storage, authentication, judge, channel, console, API, and SDK documentation, including regenerated architecture diagrams and OpenAPI artifacts. ## Validation - SQLite suite: 11,188 passed, 9 skipped, 10 deselected - PostgreSQL suite: 11,195 passed, 2 skipped, 10 deselected - Live backend: 3 passed - SSE recovery: 6 passed; browser recovery harness passed all scenarios - Ruff: clean; 595 files correctly formatted - mypy: 243 source files clean - TypeScript: typecheck/build and 35 tests passed - OpenAPI artifacts fresh; all 14 changed diagrams reproduce byte-for-byte - `git diff --check` and Git LFS integrity clean Closes #979. * fix(deps): update nanoid for GHSA-2v37-7h3g-55p8 Refresh the transitive lock entry admitted by PostCSS so the TypeScript security gate no longer resolves the vulnerable custom-generator implementation. Validation: - npm ci - npm audit --audit-level=moderate: 0 vulnerabilities - TypeScript typecheck and build - TypeScript tests: 35 passed * fix(test): assert canonical model registry URLs Replace prefix checks with exact canonical base URL assertions so the tests do not model incomplete URL validation. Validation: tests/test_model_registry.py (185 passed); Ruff check/format; mypy. |
||
|
|
70165807c7 |
fix(reasoning): close the unmarked chain-of-thought leak, gate the tag scan by backend (#940) (#978)
Some serving setups emit model reasoning inline with no think tags and no
reasoning_content at all — nothing any parser can segregate (measured live
on the dev vLLM: 20/20 sampled completions, streamed and not, proxied and
direct). The drain seam correctly passes unmarked prose through, so it
became the artifact on every bounded-artifact lane: workstream titles
("Thinking Process:"), compaction summaries that were ~90% chain-of-
thought, and the web-fetch tool results #940 reports — which then ride
every following turn as context.
Three coordinated changes:
* Utility lanes ask for no reasoning. _utility_completion (title,
compaction, web-fetch extraction) pins the alias's declared thinking
toggle off and withholds every reasoning-effort channel — the relayed
session knob, the lane rung, the definition default, and the graded
template key — via lane_without_thinking / lane_thinking_suppressed,
the same suppression omni transcription already used (now shared as
thinking_off_template_kwargs). Measured end-to-end: the extraction
that returned 3.7k chars of reasoning returns a 258-char answer.
* server_parses_reasoning capability. A backend that segregates
reasoning into its own channel declares it, and the inline tag scan
turns off on every lane: the drain seam, the interactive splitter
(which now reads the ACTIVE stream's capabilities via the creation-
time handoff register, never the primary alias's), and the title
lane's cosmetic peel — so prose that merely quotes a tag can no
longer be misrouted, and the utility suppression stands down where
reasoning costs the artifact nothing. The built-in commercial
capability tables declare it wholesale (known models and table-miss
defaults); local compat lanes keep the passthrough default the scan
exists for. Bool-typed capability overrides coerce string spellings
instead of truthiness-flipping on hand-edited JSON.
* Title selection follows the prompt's contract, not line position:
the last line within the word cap that ends in a word character —
rejecting explanation sentences, sign-offs, parentheticals, and
reasoning headings in any script (terminal punctuation carries
unspaced scripts where whitespace word counts are meaningless) —
else the last non-empty line. 20/20 captured live responses title
correctly (9/20 before, unchanged since well before the seam
unification: the old and new pipelines scored identically on every
sample, so the regression source was the backend's output shape,
not #965).
Also folded in from the review round: a think tag split across a
reasoning-delta boundary reassembles in the drain (partial-tag tail
carry; tool boundaries still flush), Turn.text joins text blocks with a
newline so multi-block answers stop fusing words in notification bodies
and every flattened read, the notify hook reads final_assistant_text
directly instead of through a one-line shim, web-fetch extraction uses
the shared _non_blank_or fallback, and the judge/output-guard suites use
real ModelCapabilities instead of truthy mock attributes.
Closes #940.
|
||
|
|
bc3fa60011 |
fix(providers): segregate inline reasoning at the drain seam
Passthrough servers (parserless vLLM/llama.cpp, LM Studio, bare gateways) emit reasoning as literal <think>/<reasoning> blocks inside content, and only three of nine drained lanes stripped them: web_fetch tool results persisted raw think blocks into every following turn (#940), judge verdicts parsed through tag noise, and a draft verdict inside a think block could shadow the real one at the output guard. One rule at the seam now. drain_stream accumulates content in RUNS bounded by interleaving signals (provider-parsed reasoning deltas, tool-call deltas) with the interactive consumer's within-chunk ordering — reasoning, then content, then the tool-call close — and splits each run through split_inline_reasoning, the one-shot form of the interactive lane's ThinkTagSplitter: a pure raw split, exactly equivalent to the streaming form on every catalog case. One trim policy exists and the drain owns it: blank edge lines are trimmed once over the joined runs when a tag was consumed, so tag residue dies at the edges while genuine inter-run paragraph separators survive. Extracted text is appended to result.reasoning after any server-parsed reasoning with a blank-line boundary and rides the native lane as the reasoning_text synth block. Orphan CLOSE tags deliberately pass through byte-identical: a close whose open never arrived is indistinguishable from prose QUOTING the tag, and drained lanes routinely quote third-party text — reclassifying would let a malicious page containing the literal tag destroy the extraction that cites it. The title lane keeps a local rfind peel as display-string formatting. The citations footer folds only onto non-blank content — sourcing for an answer that does not exist is dropped rather than handed to emptiness checks as a footer-only "answer". Every private strip is deleted: the title lane's strip, the summarizer strip, _strip_reasoning itself, and the optimizer's five regexes (_strip_markdown_fence is now the one fence rule, applied to normalized model output only, never to or-fallback values). Think-only and whitespace-only responses drain to blank content, and every lane's no-answer fallback gates on blankness: web_fetch returns an honest extraction-error card, the intent judge takes the empty-retry ladder, the task-agent synthesis reports "(no output)", and the optimizer keeps the current observer system and prompt verbatim on no-answer passes. Final-say reads (optimizer analyst, eval final_content, the notify hook) use trajectory.final_assistant_text — the last assistant turn only, never an earlier narration presented as the conclusion — while last_assistant_text is the salvage walk (task_agent partial-work recovery), skipping tool-call-only, all-reasoning, and whitespace-only turns. Perception memoizes every completed description immediately, including an empty one — one perceive per key, ever — under a commit-lock guard so an empty result never overwrites a concurrently memoized real description; an all-reasoning perception model pins the placeholder until restart, and the remediation is server-side (a reasoning parser or the template thinking toggle on the perception alias). A true double-reasoning shape (inline-extracted text alongside a native reasoning block) logs chars-only at the drain, where it is distinguishable from the routine reasoning_delta mirror. The dialect's semantics are pinned as one table (tests/_reasoning_dialect.py) driven through shared fixtures (think_tag_stream, seam_provider): one-shot conformance, the exact one-shot/streaming equivalence property, the drain seam rules including quoted-tag safety, run-boundary and separator-preservation pins, per-lane pins for all nine lanes, and the empty-content assistant wire shape. Closes #965. Closes #940. |
||
|
|
96834496c4 |
feat(anthropic): onboard claude-opus-5
The capabilities row is a copy of claude-opus-4-8 — 1M context, 128K output, adaptive thinking, the full low..max effort ladder, mid-conversation system messages. Two of the model's documented breaking changes are unreachable from this lane and stay that way only while thinking_mode is "adaptive": thinking is on by default when the param is omitted, and disabling it is a 400 at effort xhigh or max. We never omit it and never emit "disabled", so both are recorded at the row rather than defended against. The third needed work. A safety classifier can decline with stop_reason="refusal" on a successful HTTP 200, with content either empty or partial, and an unmapped value fell through _normalize_finish_reason as a literal string. The drain gate only raises on an ABSENT finish reason, so a declined turn landed as a complete result with nothing to notice it: the interactive lane matched neither of its two warn arms and stayed silent, and a sub-agent handed the declined partial up to its parent as though it were finished synthesis. Normalizing onto content_filter routes the decline into the arms both lanes already have for that state — the operator gets the warning, and the sub-agent stops rather than passing the fragment on. The raw stop reason is logged where it is still in hand: normalization is lossy and a classifier decline is otherwise indistinguishable from an ordinary content filter. The gate is on the RAW value rather than (normalized != raw), which is true for end_turn and tool_use as well and would fire on every turn in every lane. The anthropic floor moves to 0.117 to track the release current at onboarding. The model needs no new SDK surface — ids are opaque strings and "refusal" has been in the StopReason literal since ~0.95 — so this is hygiene; raise it again when adopting fast mode, server-side fallbacks, advisor, or mid-conversation tool changes, which do need newer typed params. |
||
|
|
7d4d76e097 |
fix(providers): PR review — orphan deltas arm the finish shim, test style
- Orphan argument deltas count as delivered output for the finish_reason_optional shim, exactly as they count as a streamed signal for the terminal harvest: a lax Responses server that never announces items AND never sends a terminal event still delivered its tool call — with the tolerance declared that is a completion, not an IncompleteStreamError. (Review caught the shim/harvest inconsistency the round-9 fix introduced.) - Test style: single import style for the model_turn module, assert on a local instead of a call expression, drop a pass-through lambda. |
||
|
|
747177a76c |
fix(providers): review round 9 — orphan/harvest collision, shared shim gate, retired-id rationale
Correctness: - Responses: orphan argument deltas (streamed without any output_item.added) now count as a streamed tool-call signal, so the terminal harvest stands down instead of re-emitting the same call onto the same slot — the reproduced collision concatenated the arguments JSON into an unparseable double copy. Cleanup / documentation: - finish_shim_due in _protocol is THE gate for the lax-server finish shim — one predicate (and one definition of 'delivered output') for all three adapter families, so the same capability flag cannot acquire per-family completion semantics. - The Responses error/response.failed branches share one failure tail (only code/message extraction differs) — the same server failure can never become retryable through one event type and fatal through the other, pre- or post-terminal. - _format_refusal pins the refusal rendering the streamed event and the terminal harvest both use. - The capability-table floor comment and CHANGELOG Removed entry now state the real rationale: OpenAI has RETIRED the pruned ids from the API — the rows described unreachable contracts, not unpopular ones. - CHANGELOG names the stream-entitlement break class (verified-org streaming, pre-stream_options gateway api-versions) with its serving-side remediation; deliberately no non-streaming fallback. - docs/architecture.md retry section describes the collapsed transport: the two stacked retry ladders, IncompleteStreamError / ResponsesStreamFailedError retryability, finish_reason_optional remediation; stale non-streaming mentions updated (+ puml). - Anthropic whole-block emission carries its residual hybrid-gateway bet as an explicit comment. Held on standing rulings: post-finish usage forfeiture (keep result + warn, rounds 4/8), session merge_usage twin and StreamAbortRef twin (#832), stream_options wire delta (round 2, caveat now names Azure). |
||
|
|
49d33d8594 |
fix(providers): review round 8 — under-streaming gateway parity, abort-race close, usage-blip visibility
Correctness: - Responses: output that exists ONLY in the terminal payload (buffering gateways that never fire output_text.delta / output_item.added) now reaches CompletionResult.content and tool_calls — the retired non-streaming _parse_response read this same payload, so the drain must too instead of returning a clean-looking empty success (blank compaction summary, silently-skipped tool call). Gated on nothing of that kind having streamed; refusal parts render as the streaming branch does. - Anthropic: content pre-populated inside content_block_start (whole- block lax-gateway emission — the real API sends start blocks empty) is emitted for text, thinking, and tool_use input, type-guarded like _reasoning_text so duck-typed blocks can't leak non-strings. - Responses: a response.completed payload that OMITS status maps to "stop" via the event type, matching the payload-less branch — the empty-string status read as 'length' and fired truncation policies on complete output. - model_turn: the drain-retry loop re-checks cancel_ref.aborted after the backoff sleep — an abort landing mid-sleep now kills the abandoned worker with the original failure instead of issuing one more full request behind the deadline's back. - drain_stream: the post-finish transport-blip tolerance logs a warning naming whether usage was captured — the kept result may report usage=None (chat-lane usage trails the finish reason) and that spend was vanishing from usage accounting with no signal. Cleanup: - ChatSession's inline tool-call fold adopts accumulate_tool_call_delta (drop-in — same ToolCallDelta semantics), so THE merge rule now has one implementation across the chat loop, drain_stream, and the Google capture; the helper's mirror-mandate docstring is retired. - The task-agent _api_call contract comment reconciles the two retry layers (sub-harness owns request-level policy; model_turn owns drain-time re-issue) instead of claiming model_turn is policy-free. - Anthropic's three terminal-emission sites share one _attach_terminal_blocks helper — replay fidelity can't depend on which terminal path a stream took. - Responses create_streaming resolves capabilities once. Held on standing rulings: o-series capability-row removal (4th report; deliberate break, release-noted), StreamAbortRef/_CancelRef unification (#832; docstring mirror-mandate). |
||
|
|
8fa0e7a29e |
fix(providers): review round 7 — id-disciplined slots, all-lane finish tolerance, retry backoff
Correctness:
- ToolCallSlotter: a slot whose id is KNOWN never splits on an id-less
delta — on an id-disciplined server new calls arrive with ids, so an
id-less fragment (the call's FIRST name announcement included) is
always a continuation. Round-6 regression: {id} → {name} → {args}
emission split into an unnamed id-bearing call plus a nameless twin.
Also: a name arriving for a slot with no name yet never splits
(args-first emission), and a bare same-name delta after complete
arguments merges as a redundant footer instead of minting a phantom
zero-argument call that would re-run a side-effecting tool.
- finish_reason_optional is honored on every drained lane, not just
Chat Completions: Anthropic shims a missing message_delta
stop_reason + message_stop pair, Responses a missing terminal event
(both with collected blocks riding the shimmed finish) — the
documented capabilities-JSON remediation now works on the
anthropic-compatible/responses-compat gateways it was written for,
matching the retired non-streaming paths' tolerance.
- Responses: an in-band error/response.failed frame arriving AFTER the
terminal event is teardown noise — log and end the stream instead of
raising away a generation already in hand (the in-band twin of
drain_stream's post-finish transport-blip tolerance).
- model_turn drain retries pace like the SDK request retry they
replace: 0.5s base, doubling, ±50% jitter — instant re-issues
re-hit the still-active rate limit/overload and synchronize into
fleet-scale retry bursts.
- Responses slot bookkeeping survives lax servers: slots minted by a
counter (len(dict) collided calls after a duplicate/empty item-id
overwrite), orphan argument deltas route to the most recently
announced call instead of hardwired slot 0.
Cleanup:
- on_tool_call_delta now receives the normalized ToolCallDelta plus the
raw SDK delta — Google's capture accumulates the exact bytes the
mirror sees (the byte-identical extraction no longer exists twice).
- _ArgsScanner feeds only fully id-less slots (its verdict is never
consulted for id'd slots — dominant-case hot path).
- Anthropic retryable set hoisted to a class constant (per-access
frozenset allocation, same pattern already fixed on Responses).
- GoogleProvider class docstring names the hook-based capture instead
of the deleted _extract_tool_calls override.
|
||
|
|
16647db1b0 |
fix(providers): review round 6 — strict-by-default finish gate, slotter v3, drain retry
Correctness: - The chat-lane finish shim is now armed only by an operator-declared finish_reason_optional capability (model-definition capabilities JSON). Default lanes treat a clean finish-less end as died-mid-generation (retryable) — SSE cannot distinguish lax-server completion from a worker dying behind a clean-closing proxy, and the default must catch truncation rather than bless it. When armed, reasoning-only output counts as a completed generation (parity with the retired non-streaming path's finish_reason-or-stop default). - ToolCallSlotter v3: id-less call-boundary decisions now consult argument JSON completeness (incremental scanner) and name identity instead of a boolean has-args gate. Fixes both residual id-less ambiguities: two zero-argument whole-delta parallel calls no longer fuse (silently dropping an action), and redundant per-fragment name headers no longer split one call into malformed half-JSON calls. - model_turn re-issues transient mid-stream deaths (provider's retryable_error_names, raised while draining) up to twice — the new home of the SDK request-level retry the non-streaming transport gave every single-shot lane (judge, title, perception, compaction). Request-time failures keep the SDK's own policy; an aborted cancel_ref suppresses re-issue (StreamAbortRef gains .aborted). Cleanup: - One slotter drives both the normalized mirror and Google's raw fidelity capture via an on_tool_call_delta hook — raw/mirror slot parity is structural now, not a maintained invariant. - accumulate_tool_call_delta in _protocol.py is THE tool-call merge rule; drain_stream and the Google capture use it (session's copy is #832's tracked adoption). - Responses terminal rebuild only runs when the terminal payload can disagree with the .done-collected items (truncation or count mismatch); on rebuild, annotations are replaced, not re-extended. - Responses retryable set precomputed at class creation. Held on standing rulings: o-series capability-row removal (deliberate, release-noted with remediation), StreamAbortRef/_CancelRef unification (#832; docstrings mandate mirroring until then). |
||
|
|
91c46051d9 |
feat(providers): drop o-series and pre-5.4 GPT-5 capability rows
The OpenAI commercial capability table floor is now gpt-5.4: o1, o1-mini, o3, o3-mini, o3-pro, o4-mini, gpt-5, gpt-5-mini, gpt-5-nano, gpt-5-pro, gpt-5.1, gpt-5.1-codex-max, gpt-5.2, gpt-5.2-pro, and gpt-5.3 are effectively unused in the field. The gpt-5-search-api row (different product surface) and the audio/STT/TTS rows stay. A legacy id now resolves to OPENAI_DEFAULT (temperature sent, no declared effort vocabulary, 200K window) — which those models may reject; the remediation is the model definition's capabilities JSON or a current model, release-noted under Unreleased → Removed. This also retires the transport-collapse review's thrice-reported "stream-rejecting o1-era models are stranded" finding by removing its subject: no row in the table describes a non-streaming model anymore. Tests migrate to 5.4-era equivalents that pin the same behaviors: always-reasoning temperature suppression and off-list effort snap (gpt-5.4-pro for gpt-5-pro/o3), explicit-none forwarding (gpt-5.4 for gpt-5.1), empty-effort-vocabulary knob drop (gpt-5-search-api for o1-mini), and the longest-prefix shadow hazard (gpt-5.4-pro vs gpt-5.4 for codex-max vs gpt-5.1). |
||
|
|
3a28dc2f16 |
fix(providers): review round 5 — same-id fragment merge, post-finish blip tolerance, chat finish shim
Correctness: - ToolCallSlotter's reannounce split is gated to ID-LESS deltas: id equality proves the same call, so compat servers that repeat the id+name header on every argument fragment merge back into one call with valid JSON (round 4's ungated heuristic split them into duplicate half-JSON calls — execution-confirmed by the review). The residual id-less repeat-name-per-fragment shape is documented as inherently ambiguous; ids are the only disambiguator. - drain_stream keeps a completed result when the transport blips AFTER the finish reason (trailing usage chunk / citation footer window): the generation is in hand, so forfeit the trailing metadata instead of discarding a fully-delivered verdict or re-paying a compaction. - The chat iterator shims finish_reason="stop" when a stream ends CLEANLY after delivering content or tool calls — the deleted non-streaming `or "stop"` default for lax finish-reason-less servers, now safe to restore because abrupt deaths surface as httpx.TransportError (round 4) rather than clean exhaustion. This supersedes the round-3 keep-the-gate ruling: the httpx catch changed the calculus, and the Anthropic/Responses lanes already got their marker-based shims. Empty/reasoning-only streams still fail the complete-or-error gate. Two streaming tests gained the shim chunk. Dispositions held: o1-era stream-rejecting models (third re-report) stay a release-note remediation per the earlier ruling. Cleanup: the two Responses terminal branches collapse into one path (status derived from the event type when the payload is missing — also fixes the end-of-stream debug log reporting finish_reason=None for completed lax streams); the annotations walk is one shared helper (the two copies had already diverged on None-content guarding); _raise_responses_failure is annotated NoReturn; scripts/livepass.py drops the phantom supports_streaming key; test_model_registry's capture helpers ride scripted_chat_client; _openai_stream_chunk points at its fake_chat_stream shape-twin for future consolidation. |
||
|
|
cf7cfe8932 |
fix(providers): review round 4 — wire-error retryability, tap/mirror slot parity, terminal completeness
Correctness: - drain_stream chains raw httpx.TransportError from stream iteration into retryable IncompleteStreamError (original type+message preserved via __cause__): streaming moved the body read out of the SDK's APIConnectionError-wrapped request, so mid-body connection drops and read timeouts — retried transparently on 1.7 — were escaping every single-shot retry loop as instantly-fatal raw httpx names. - The index remap is extracted as ToolCallSlotter and GoogleProvider's raw tap slots THROUGH IT over the same delta sequence as the base iterator: round 3's mirror-side de-fusion had left the tap keying by wire index, so a degenerate stream produced 2 mirror calls vs 1 fused raw dict — _prepare_messages' length gate then silently dropped the thought_signature lane (400 on signature-strict Gemini models). - The slotter also splits ID-LESS degenerate parallel calls: a delta announcing a name for a slot that already accumulated arguments is a second whole call, not a fragment (fragmented single calls pinned unaffected). - A payload-less Responses terminal event keeps the provider_blocks already collected from output_item.done events (they came from the stream, not the missing payload); only usage is genuinely lost. - The truncation-rebuild path walks the terminal output's message annotations, so truncated web-search turns keep their Sources footer (the in-flight item never received output_item.done). Cleanup: one _raise_responses_failure ladder serves both in-band failure shapes (error events + response.failed); IncompleteStreamError joins the public providers export (docstrings tell callers to catch it); the de-fusion tests ride the file's existing _openai_stream_chunk helpers instead of a third hand-rolled SSE fake; the dead if-response guard in the terminal branch is gone. Deferred with note: classifying IncompleteStreamError once at the retry-predicate consultation site instead of per-provider strings is #832 territory (the predicate lives in ChatSession); the six-lane parametrized test guards the listing until then. |
||
|
|
56b7674dfa |
fix(providers): review round 3 — in-band error events, terminal-marker tolerance, adapter-owned de-fusion
Correctness: - Responses _iter_stream handles the SDK's in-band `error` SSE event (ResponseErrorEvent is YIELDED, not raised, and no response.failed need follow): the real API code/message now surfaces — code-gated for retryability like response.failed — instead of the stream exhausting finish-less and hiding the cause behind a retried IncompleteStreamError. - Anthropic message_stop supplies a missing stop_reason: it is a genuine terminal marker, so a compat /v1/messages shim whose message_delta omits stop_reason completes (blocks intact) rather than failing a generation that arrived — tolerance the retired non-streaming default provided, restored without weakening the died-mid-response gate. - A Responses terminal event without its response payload still emits the finish reason its type implies (lax compat servers), losing only usage/blocks rather than the whole result. Dispositions held (documented, not re-coded): the complete-or-error gate stays for finish-less Chat Completions streams — indistinguishable in-band from a died generation, and silent partial-storage is the worse failure; CHANGELOG now names the shape and each provider's accepted terminal markers. supports_streaming deletion and the stream_options wire delta were ruled earlier and keep their release-note remediations. Cleanup: index-degenerate de-fusion MOVED from drain_stream into the chat adapter's iterator (mirroring the Anthropic iterator's index assignment) so the interactive loop is fixed too and the drain returns to a plain mirror of the main-loop accumulator; a parametrized test locks "IncompleteStreamError is retryable" across all six provider lanes instead of trusting per-adapter memory; scripted_anthropic_client joins scripted_chat_client (shared _ScriptedClient class, no function attrs) and the two remaining hand-rolled anthropic closures convert. |
||
|
|
3ffa8b9057 |
fix(providers): review round 2 — complete-or-error drain, code-gated retries, truncation-safe blocks
Correctness (3 confirmed + 2 plausible, all fixed): - drain_stream now raises typed, retryable IncompleteStreamError when a stream exhausts without any finish reason — every adapter emits one on a healthy stream, so its absence means the generation died mid-response behind a cleanly-closing proxy. This restores the retired transport's complete-or-error contract (a half-generated compaction summary was previously returned as finish=stop and stored, silently replacing real history) and DELETES round 1's suffix-info fold: with no finish-less success path there is nothing to classify, so a trailing status ping can never be stored as content either. - Index-degenerate parallel tool calls get distinct slots: a delta whose id differs from its slot's opens a new call (id-less fragments still follow their index's current call), so historical compat servers that emit every parallel call at index 0 no longer fuse distinct calls into concatenated garbage arguments. Result order stays index-sorted (stable) like the retired array parse. - response.failed retryability is code-gated: only transient codes (server_error, rate_limit_exceeded) raise the retryable typed error; deterministic rejections (invalid prompt, image fetch, policy) raise plain RuntimeError and stop retry loops on attempt zero instead of running the full backoff ladder against a doomed request. - Terminal Responses events rebuild provider_blocks from response.output when present: the item being generated at max_output_tokens truncation never receives output_item.done, and storing a reasoning item without its required following item made the next turn's replay a 400. - merge_usage's base case uses dataclasses.replace so a future UsageInfo field can't be silently zeroed on drained lanes. Cleanup: run_abortable_with_deadline bundles the three-point abort wiring (ref + cancel_ref + on_abandon) so it cannot be half-wired — both judges converted; scripted_chat_client hoists the 14 chat-lane fake_create closures (call scripts + .calls recording replace per-test counter cells); fake_chat_stream gains reasoning=, collapsing the reasoning-capture suite's hand-rolled chunk shape; FakeAnthropicBlock hoists the duplicated _Block test class; the class and judge PlantUML diagrams drop the retired create_completion flow. Also converts test_model_registry's agent-model fakes, which returned legacy response objects that iterated as EMPTY streams — they only passed through the old drain's silent finish=stop default, exactly the hazard the new gate exists to catch. |
||
|
|
08580f25f9 |
fix(providers): review round 1 — streaming parity gaps the collapse exposed
Correctness (4 confirmed + 1 plausible fixed, 2 accepted+documented): - Anthropic _iter_anthropic_stream handles citations_delta: text-block citations now ride the raw block into provider_blocks, as replay requires (the retired non-streaming lane preserved them via model_dump; the streaming lane dropped them — a pre-existing main-loop gap the collapse would have extended to single-shot lanes). - Anthropic text blocks separate with "\n" at each subsequent block start, restoring the retired lane's "\n".join rendering on drained lanes AND un-fusing streamed web-search responses in the chat loop. - response.failed raises typed ResponsesStreamFailedError, listed in the provider's retryable_error_names — retry loops treat an in-band failure like the wire errors it stands in for instead of hard-stopping on a bare RuntimeError (judges keep their heuristic fallback after retries). - drain_stream folds a finish-less stream's terminal citations footer (suffix rule: pre-finish info invalidated by any later payload), so lax compat servers that never send finish_reason keep their Sources. - usage max-merge extracted as merge_usage() in _protocol.py — the one definition drain uses now and the session's inline consumer adopts on #832. Accepted + release-noted instead of coded around: strict pre-2024 compat servers that 400 on stream_options (such a server already cannot serve the chat loop; CHANGELOG caveat extended), and repeated-index parallel tool-call merging on legacy compat servers (identical to the main loop's accumulator semantics; a shared guard belongs in the #832 unification). Cleanup: run_with_deadline grows on_abandon (best-effort, cannot mask the deadline error) and both judges drop the copy-pasted abort choreography; StreamAbortRef documents the _CancelRef adoption plan; test_model_turn's fake replays through the shared as_stream adapter; docs/architecture.md drops the retired Protocol row. Tests: refusal handler pinned (was advertised, untested); typed-failed retryability; citations capture; text-block separator (plus the mixed text+search expectation updated for the separator chunk); finish-less citation fold; on_abandon firing matrix; StreamAbortRef arrival race. |
||
|
|
1e7ad7bcb6 |
feat(providers): one transport — drain create_streaming, retire create_completion (#831)
Every single-shot lane (model_turn: judges, titles, compaction, web-fetch extraction, perception, eval, optimizer) now samples through the provider's streaming entry and accumulates via a shared drain_stream(), deleting create_completion from the Protocol and all three adapters (xai/google inherit). Request shaping can no longer drift between the two consumption styles, and callers keep the exact CompletionResult contract. The drain mirrors the main loop's proven chunk semantics: per-field max-merge for usage (Anthropic splits prompt/completion across message_start/message_delta), tool-call assembly by delta index, provider_blocks from the terminal emission, trailing citation info folded back into content (byte-matching the old format_citations append), mid-stream status pings dropped. Also in this change: - model_turn grows cancel_ref; both judges wire their run_with_deadline abandon paths to a new StreamAbortRef (deadline.py) that closes the SDK stream — a timed-out judge call now aborts its HTTP read instead of pinning a daemon thread until the next upstream chunk. The append hook covers the arrival race, mirroring ChatSession._CancelRef. - Responses streaming gains the response.incomplete terminal handler (truncated runs were mislabeled finish=stop and lost final usage AND collected provider_blocks) and a refusal handler ([Refused: …] content, matching the retired non-streaming rendering). Both also fix the main chat loop, which shared the gaps. - supports_streaming capability flag deleted (zero readers) along with its admin capability tile; o1-era models that reject streaming need a model alias pointing at a current model (release-noted). - Helpers that existed only for the deleted transport go with it: Responses._parse_response, chat/google._extract_tool_calls. Known behavioral deltas (release-noted): OpenAI-compatible servers that ignore stream_options.include_usage stop producing usage rows on these lanes; multiple Anthropic text blocks concatenate without the old "\n" joint (matching the main loop); model_turn lanes no longer risk client read-timeouts on long generations — the reason the Anthropic adapter already drained a stream internally. Tests: new test_drain_stream.py pins the accumulator rules; shared fakes (as_stream, fake_chat_stream, fake_anthropic_stream) migrate 11 suites to the streaming transport, with the task-agent and adapter suites now exercising the real _iter_stream + drain path end to end. |
||
|
|
660aff6f1e |
fix(task-agent): guard the Google swap against partial lanes; fix a stale comment
The fidelity swap now requires the raw lane to be a faithful counterpart of the mirror — same length, every id present — before replacing tool_calls; a partially-corrupted lane (filtered non-dict elements) would otherwise swap a shorter list over the mirror and orphan a mirrored call whose tool result remains in history. The _run_agent call-site comment now matches the builder's reasoning_text-only blank-id rule. |
||
|
|
dc52bc2b96 |
fix(task-agent): simplify the blank-id rule to reasoning_text-only and heal historical rows
The blank-id gate's strip-then-filter semantics left two residual hazards (surviving Responses reasoning items whose pairing contract needs their original sibling items; an asymmetric Messages-shaped lane surviving when no client block was actually stripped). The rule is now total and simpler: on a blank-id turn only the loose-text reasoning_text synth block survives — it carries no id and is shape-invalid on the Messages translator by design, and real-world blank-id servers are Chat-Completions locals whose reasoning IS that loose text. This also removes the builder's per-call provider import. The Google fidelity swap now skips raw rows carrying a blank id (historical captures that predate the gate would otherwise resurrect the blank id on every replay — the sanitized mirror stays), guards against non-dict lane elements, and legalizes via the new shared lowering.legalize_tool_call_entry — the ONE per-entry legalizer the sanitize pass also uses, so the two seats cannot drift on semantics or the wire.tool_args_legalized breadcrumb. |
||
|
|
98cefc3660 |
fix(task-agent): move the blank-id gate into the shared native-lane builder
The blank-provider-id gate lived only at the _run_agent call site while
the main-loop stream accumulator has the identical back-fill-then-carry
seam — and it over-dropped, discarding the reasoning lane for exactly
the servers that emit blank ids. The gate now lives in
_finalize_provider_blocks as a had_blank_ids parameter both harnesses
thread: client tool blocks (which keep the blank id the mirror back-fill
never reached) are stripped, and when any were present the remaining
Messages-shaped blocks go with them (a surviving native lane REPLACES
the rebuilt content on the Anthropic translator, so a lane missing its
tool_use would orphan every mirrored call) — while shape-invalid
reasoning residuals (reasoning_text, Responses reasoning items) are
kept. This also closes the pre-existing main-loop case: a Gemini
openai-compat turn with a blank tool id no longer persists a raw
fidelity dict whose blank id the swap would resurrect on every replay.
The Google fidelity-swap legalization now reuses the canonical
lowering.legalized_arguments (made public) instead of a hand-rolled
narrower copy: dict-shaped arguments are serialized rather than
collapsed to {}, the standard wire.tool_args_legalized breadcrumb is
logged, and a degenerate non-dict function entry passes through
untouched instead of raising.
|
||
|
|
646bceed52 |
fix(task-agent): review fixes for the native-lane carry
- Skip the native lane on a turn whose provider left a tool-call id blank: the uuid back-fill reaches only the tool_calls mirror, so a carried native tool_use block would replay the blank id and desync from the restored tool_result (Anthropic orphans the result; Google re-fills a fresh uuid). The rebuild path keeps every representation on the back-filled id — the pre-native behaviour, for exactly the degenerate case. - Extract _reasoning_text as the ONE Chat-Completions reasoning extractor shared by the streaming and non-streaming paths: first non-empty STRING of reasoning/reasoning_content wins, so a server putting a structured object in reasoning can neither shadow valid text in reasoning_content nor leak a non-str into the session's reasoning accumulator. - Legalize arguments when GoogleProvider's fidelity swap replaces the sanitized tool_calls mirror with the raw provider dicts — the swap could resurrect a malformed arguments string the upstream sanitize pass had fixed (pre-existing on the main loop; ids and thought_signature untouched). - Drop the redundant emptiness guard on the agent seam's reasoning_parts (the shared finalize helper already guards) and document the wire_id_map lifetime invariant for future resumable/background agents. |
||
|
|
d660819142 |
feat(task-agent): carry the provider-native reasoning lane in the sub-harness
A task agent's replayed turns now carry the native reasoning lane the model produced (Anthropic thinking blocks + signatures, OpenAI Responses reasoning items, Gemini thought_signature blocks, vLLM/llama.cpp parsed reasoning text) instead of being rebuilt from content + tool_calls with the reasoning dropped — restoring reasoning continuity across the agent's own multi-turn tool loop on every provider lane. The prerequisite is the id half: replace legalize_tool_call_ids with restore_provider_tool_ids, a lowering pass that maps the session-minted sub-tool ids back to the provider's own ids on the transient wire copy (from the per-run mint map, never by string-splitting). The native tool_use block is replayed verbatim — its id and signature untouched — and the top-level mirror and tool_result agree with it on every request. The minted id stays the sole internal key (registry, DOM, recall, cancel ledger), #820 unchanged. Chat-Completions lane: non-streaming create_completion now surfaces reasoning/reasoning_content as CompletionResult.reasoning (the twin of the streaming reasoning_delta extraction), and the agent seam runs the Phase 5 vLLM reasoning-field replay against the agent's own provider and alias. The native lane is finalized by a shared helper (_finalize_provider_blocks) so the main loop and the sub-harness cannot drift; replay honors the per-model replay_reasoning_to_model flag on every lane, and llama.cpp stays capture-only, matching the main loop. |
||
|
|
b450b9ad20 |
fix(providers): align GPT-5.6 with the GA API surface
- every 5.6 tier accepts effort "max" and reasoning.mode
"standard"/"pro" (GA docs: pro is a request mode on any GPT-5.6
model) -- drop the Sol-only gating
- GPT-5.6 deprecates prompt_cache_retention; send
prompt_cache_options={"ttl": "30m"} (its only supported lifetime)
and keep the 24h retention policy for pre-5.6 models
- never inject commercial cache params into local lanes: dropped from
the Chat Completions lane (which serves only openai-compatible and
google) and gated off the compat-pinned Responses lane -- a gpt-5*
served-model name is not an OpenAI account
- account cache writes: usage *_tokens_details.cache_write_tokens
flows into cache_creation_tokens (5.6 bills writes at 1.25x the
uncached input rate)
- drop non-string verbosity/reasoning_mode overrides with a warning
instead of raising on unhashable capability-JSON values
- keep ModelCapabilities' public positional prefix stable by appending
the verbosity/pro fields at the tail; pin it with a constructor test
- openai floor 2.44 -> 2.45, the first release with the typed
prompt_cache_options kwarg
|
||
|
|
47f908c9e0 |
feat(providers): add OpenAI GPT-5.6 (Sol/Terra/Luna) support
Onboard the GPT-5.6 family (GA 2026-07-09) to the OpenAI Responses lane. - Capability rows for gpt-5.6 (= Sol alias/catch-all), gpt-5.6-terra, and gpt-5.6-luna: 1.05M context, 128K output, tool_search/vision/pdf/reasoning replay, default effort medium, temperature only at effort=none. - "max" reasoning effort, Sol-only; Terra/Luna cap at xhigh (the knob's "max" snaps to the xhigh ceiling). First commercial OpenAI use of "max" — the ordinal knob already ranked it, so no effort-ladder change was needed. - Verbosity and pro mode as operator-declared capability fields (supports_verbosity/verbosity, supports_pro_mode/reasoning_mode), merged from the model-definition capabilities JSON and emitted on the Responses wire as text.verbosity and reasoning.mode. Both are gated by a supports flag plus an enum guard that drops unknown values with a warning. Pro mode is Sol-only. There is no gpt-5.6-pro model — "pro" is the reasoning.mode param, not a separate model id. - Raise the openai floor to >=2.44 for the 5.6 Responses params. Unit and wire-golden tests cover the rows, max->xhigh snapping, the two levers, and the enum guards. Validated live against the OpenAI API: gpt-5.6 accepts the model id, effort "max", text.verbosity, and reasoning.mode="pro". |
||
|
|
530958e06b |
fix(providers): the session effort level always reaches the local-lane wire
Local lanes dropped the knob's graded value unless the operator declared
reasoning_effort_values (and, on the template channel, an effort key) —
picking Max sent a bare thinking toggle and the effort select
degenerated into seven positions that all meant 'on'. The user's
setting now always rides:
- openai-compatible: the flat reasoning_effort param carries the knob
verbatim (effort_passthrough on the lane default); declared values
still snap ordinally, and a declared effort_param still claims the
template channel and suppresses the flat param.
- anthropic-compatible: the graded value rides chat_template_kwargs
alongside the toggle whenever reasoning control is engaged — under
the operator's effort_param, else the conventional fallback key
(reasoning_effort); templates that don't reference the kwarg ignore
it. thinking_mode=none still injects nothing.
- Commercial lanes untouched: empty declared values still mean 'no
effort control' (o1-mini) and the ordinal snap is unchanged.
Golden writer now pins ensure_ascii=False: the baselines' literal em
dashes came from a hand edit (
|
||
|
|
e136237b63 |
fix(providers): openai-compatible never consults the commercial table
Local-lane model ids are operator-chosen strings (vLLM --served-model-name), so a prefix collision with a cloud model id inherited that model's sampling and effort contract: a box named o3-distill silently lost temperature support, and one named gpt-5.5-my-finetune was sent gpt-5.5's snapped reasoning_effort values it never declared. Both surfaces of the lane now return plain defaults (OPENAI_COMPAT_DEFAULT in _openai_common): the chat class directly, and the responses pin via a compat-mode OpenAIResponsesProvider mirroring AnthropicProvider(compat=True). Everything beyond the defaults is declared by the operator on the model definition, matching the anthropic-compatible lane and lookup_model_capabilities' documented 'no static table for local models' contract. The commercial openai lane (Responses-only) is untouched. Pre-split tests that reached commercial rows through the chat-class OpenAIProvider alias now source them from lookup_openai_capabilities; their subject (registry rows + shared gating helpers) is unchanged. |
||
|
|
3607517814 |
fix(providers): registry effort truth — o-series/gpt-5.5/codex-max/sonnet-5; forward declared none
Capability-registry corrections verified against the official OpenAI reasoning guide, the Azure reasoning-models matrix (2026-06 revision), and the Anthropic models-overview/effort/migration pages (2026-07): OpenAI (vocabulary confirmed none/minimal/low/medium/high/xhigh — no "max" level exists; knob max rides the xhigh ceiling via the ordinal snap): - o1/o3/o3-mini/o3-pro/o4-mini declare low/medium/high (every o-series model except o1-mini) — without declared values the session knob was silently dropped for these models. o1-mini stays effort-free. - gpt-5.5 default corrected none -> medium (5.5 reasons by default, unlike 5.1-5.4). - gpt-5.1-codex-max gets an explicit row: it prefix-matched the gpt-5.1 row (no xhigh), capping the knob's xhigh at high on the one model xhigh was introduced for. Anthropic (effort-page matrix): - claude-sonnet-5 row added — it previously fell through to _ANTHROPIC_DEFAULT (manual budgets, 200k ctx, no effort), all wrong: adaptive-by-default thinking (manual budgets are a 400), sampling params rejected, 1M ctx / 128k out, effort low..max incl. xhigh. - claude-sonnet-4-6 gains its documented "max" effort level (knob xhigh now rides max, not high) and the stale 64k max_output becomes the documented 128k. - fable-5 / opus-4-8 / opus-4-7 / opus-4-6 / opus-4-5 rows verified correct as declared. Knob semantics completed: resolve_reasoning_effort now forwards the knob's "none" position verbatim when the model DECLARES an explicit none level (gpt-5.1+, grok-4.3) — omitting the param there leaves a reasoning-on server default (gpt-5.5: medium) in charge of a knob that promises off. Models without a declared none still omit, and none is never a snap target. Parity harness swaps its synthetic openai shape for the real gpt-5.5 registry row. |
||
|
|
a0e04a8588 |
fix(providers): effort snapping is ordinal — round up, cap at the ceiling
The knob domain grew xhigh/max after the snapping fallbacks were written, which silently inverted their semantics: off-list meant "unrecognized string" then, but now usually means "above the model's ceiling", where falling back to the default tier is directionally wrong (grok-4.3 at knob max got low; values low/medium/high at knob xhigh got medium; Anthropic manual mode gave xhigh/max a 4096 budget while high got 16384). One rule everywhere now, via snap_reasoning_effort in _protocol: exact match wins; otherwise the smallest declared level ranking at or above the knob; above the ceiling, the ceiling. "none" is never a snap target, and default_reasoning_effort only catches values the ordinal snap cannot rank. - resolve_reasoning_effort (flat chat / responses / validated effort_param lanes) snaps ordinally: xhigh over (low, medium, high) now sends high; xhigh over DeepSeek-style (high, max) sends max — matching DeepSeek's official xhigh-to-max aliasing, so a declared values list now reproduces that contract instead of defeating it. - _map_reasoning_to_effort (native output_config) rounds up too: knob xhigh on Opus 4.6 (low, medium, high, max) rides max instead of silently dropping output_config. - EFFORT_BUDGET_MAP is monotone across the whole knob domain: minimal/low 1024 (API floor), medium 4096, high 16384, xhigh 32768, max 65536. Unknown strings still fall to the 4096 default. Google defaults are unaffected (ceiling and default coincide at high); wire goldens unchanged. Parity harness caught the budget clamp interacting with its own max_tokens during development — capture budget raised above the largest manual budget. |
||
|
|
d7941c88be |
fix(providers): thread the session effort knob to Gemini
_GOOGLE_DEFAULT declared no reasoning_effort_values, so resolve_reasoning_effort returned None and the session effort knob was silently dropped for every Gemini model — the same bug class this branch fixed on the local lanes. Gemini's OpenAI-compat surface documents a flat reasoning_effort (2.5: thinking_budget mapping; 3.x: thinking_level), so declaring values lights up the inherited chat-completions path. Values are the safe cross-model set (minimal/low/medium/high): "none" is excluded because 2.5 Pro and the 3.x family reject disabling thinking — and the resolver never forwards the knob's none anyway (the param is omitted, server default applies). Off-list xhigh/max snap to the declared default high. Encoded from the official compatibility docs per the static-caps pattern; not live-verified. |
||
|
|
2cf23b6fe2 |
fix(providers): address high-effort review of the reasoning-knob branch
Verified findings applied: - adaptive thinking_mode never knob-disables: the shared mapping now sends the toggle unconditionally true for adaptive (the native adaptive branch ignores the knob's none), while manual keeps the knob-driven contract. Restores the invariant the deleted chat-lane code upheld. - a set effort_param suppresses the flat top-level reasoning_effort on the chat lane: the template channel replaces it — double-sending could 400 on schema-strict servers and disagree with operator pins. - admin edit-save no longer drops a stored thinking_param when the thinking-mode dropdown is empty: the raw-JSON strip now only fires when a mode value actually round-trips through the dropdown. - effort_param persistence gated on the local-server lanes so a value lingering across a provider switch never lands on commercial rows. - three stale _compat_extra_params references renamed to merge_reasoning_template_kwargs. Documented dispositions (no code change): the knob-none-disables flip on upgrade is intentional and now carries an upgrade note; gateways fronting real Claude belong on provider=anthropic with a custom base_url (the compat lane is vLLM-schema-only); nonstandard thinking_mode strings staying inert is the intended allowlist contract. The real anthropic provider is unaffected throughout — official Claude models keep native thinking/output_config. |
||
|
|
68b22adfa3 |
feat(providers): share the effort-knob→chat_template_kwargs mapping with the openai-compatible lane
Hoist the compat-lane injection into _protocol.merge_reasoning_template_kwargs (next to ModelCapabilities — one implementation for both local-server lanes) and retire OpenAIChatCompletionsProvider._apply_thinking_mode in its favor: _finalize_extra_body now receives the session effort knob, so thinking_mode manual/adaptive maps knob "none" to an explicit thinking_param false (previously the toggle was unconditionally true) and caps.effort_param carries the graded effort key on chat completions too. Operator server_compat pins still win; the Responses surface is untouched (native reasoning handles effort itself). The admin Models form grows an "Effort param" field that round-trips like thinking_param: lifted out of the raw capabilities JSON on edit-load, re-added on save, cleared by emptying the field. Verified live against qwen3.6-27b /v1/chat/completions: knob medium streams reasoning_content, knob none suppresses it. |
||
|
|
dbf389783e |
refactor(personas): file-backed built-in prompts, explicit source column
Built-in persona base prompts move from inline DB text / base.md into prompts/personas/<slug>.md — code-owned, PR-reviewable, drift-proof. base.md / base_coordinator.md become personas/engineer.md / orchestrator.md. Prompt source is now explicit in storage instead of inferred in app logic: a new base_prompt_file column plus CHECK (base_prompt IS NOT NULL OR base_prompt_file IS NOT NULL) — two nullable columns, never both empty. Resolution is a coalesce (base_prompt else load(base_prompt_file)), frozen into the workstream stamp at creation. base_prompt_file marks a persona as built-in (code-only, un-archivable); an operator override on a built-in is allowed and wins over the file. "Inherit the kind default" is a workstream-creation act (is_default), not a persona-row state. Migration 063: - seeds reference their file (base_prompt NULL); no runtime file reads — the backfill's frozen prompt text is inlined as a point-in-time snapshot so migration history stays self-contained and reproducible. - every existing workstream is stamped by kind (creative -> writer, else the kind default), set-based (INSERT..SELECT via temp tables) with the persona column added after the bulk writes to shorten its lock window. Storage guards (both backends): operators must supply base_prompt; built-ins can't be archived or have base_prompt_file set via the API; clearing an operator persona's only source is rejected. Follow-ups reviewed alongside (#756): soft-set visibility docstring scoped to per-process; _apply_persona_snapshot / _current_persona_snapshot own the stamp round-trip; spawn approval-header args (skill/name/target_node) flattened+capped like persona; server-side tool injection generalized to replace-only (client-def gated, incl. the xAI include forwarding). Seed copy revised (researcher soft; de-costumed prose; engineer de-biased). New test_schema_parity asserts create_all matches the alembic head. Closes #683 groundwork; ruff + strict mypy clean, full suite green. |
||
|
|
5d1d34cd82 |
fix(personas): close review findings across the envelope, resume, and RBAC lanes
Provider search gating (replace-only): native web search now stands in for
a client web_search def that survived the persona visibility filter — on
both OpenAI surfaces and both injection lanes (web_search_options, the
server_side_tools loop, and _convert_tools' capability lane). A scribe or
any envelope hiding web_search stays search-free on search-capable models;
coordinators and tool-less utility calls stop receiving search too.
Resume stamp discipline: resume() loads config and parses the target's
stamp BEFORE touching session identity/history, so a corrupt stamp raises
with the session intact instead of half-adopting and then 'repairing' the
target's stamp on the next config save. The MCP lever now follows the
stamp on mid-session adoption: an MCP-off stamp drops the live surface in
place (listeners deregistered, toolsets reset); adopting an MCP-on stamp
into a session whose persona gated the client off is refused loudly (the
surface cannot be rebuilt post-construction). The REPL /resume handler
reports these errors instead of crashing the CLI.
Fail-closed default lane: a FAILED default-persona lookup at create is a
503 (routes) / clear exit (CLI) instead of silently degrading to the
unstamped stock envelope; a clean 'no default configured' still creates
legacy. resolve_persona_for_kind reports storage-unavailable distinctly
from unknown-persona.
Soft-set governance: tool_search expansion under a persona visibility set
recomposes the system prompt so tool-gated policy segments land with the
tool they gate. MCP resource/prompt catalogs gate on read_resource /
use_prompt visibility. Spawn judge/audit projections carry persona (the
human approval header already did). Active-list rows carry persona like
their project_id twin.
RBAC catalogs: persona.{create,read,write} join _VALID_PERMISSIONS and
the roles-editor sections, making the documented grant-outward path real.
Storage hardening: default-persona invariants move to a shared _utils
helper (validate + demote) with a pg advisory xact lock serializing
promotions and a post-promote single-default assertion; create maps the
unique-name race to the same ValueError as the pre-check; reads validate
JSON shape loudly (naming the persona); serialize enforces size caps;
field validation runs before invariant checks so malformed input is a 400,
never a TypeError-500. org_id guards explicit null and caps at 64.
Also: base_override='' means 'no override' at the compose boundary;
persona tag flattened/capped before the spawn approval header; /creative
redirect resolves the writer persona before advertising it; memory-nudge
gating unified through _nudges_enabled.
Provider/row-shape tests updated to the new contracts (the old ones
pinned the injection hole and the pre-persona row shape).
|
||
|
|
114ada791b |
feat(providers): add Claude Fable 5 to the Anthropic provider
- claude-fable-5 capability entry: 1M context / 128K output, adaptive
thinking (summarized display), effort low..max incl. xhigh, no
sampling params, web + tool search, vision, reasoning replay, native
mid-conversation system messages
- document the Fable 5 wire quirk at the capability table: an explicit
thinking={"type": "disabled"} is a 400 on this model; the adaptive
branch never emits "disabled", so adaptive-or-omitted is preserved
- widen the native mid-conversation-system comments from opus-4-8-only
to opus-4-8 + fable-5 (protocol, provider, tool_advisory, prompts,
session)
- raise the anthropic SDK floor 0.39 -> 0.108: 0.39 predates every
named kwarg the provider sends (output_config 0.77, top-level
cache_control 0.83, mid-conversation system blocks 0.105); 0.108
adds claude-fable-5
- tests: capability assertions for claude-fable-5 + dated-variant
prefix match
|
||
|
|
0c3d1d6cd1 |
test: satisfy ruff-format and lift side-effects out of asserts
- ruff format on two test modules that had drifted (a stray blank line and multi-line calls that now fit on one line) — restores a clean `ruff format --check`. - test_attachment_buffer: pull `buf.discard(...)` out of the `assert` expressions into locals so the eviction still runs under `python -O` (CodeQL: assert statement has a side effect). |
||
|
|
16294397c2 |
perf: dict-native wire-prep — drop the per-send Turn<->dict round-trip
The canonical-Turn migration left lowering's fold/drop/repair passes Turn-typed even though they convert to dicts internally and feed dicts to the translators, so _prepare_wire_messages round-tripped the whole history Turn->dict->Turn ~7-8x per send (even on the no-op early-return paths). Make fold_system_turns / drop_empty_user_turns / repair_wire_messages dict-native (list[dict]->list[dict]); _prepare_wire_messages now threads the dict projection _full_messages already produced straight through, with no Turn round-trip. self.messages stays the canonical Turn trajectory. export.py is simplified (it converted to dicts immediately after repair anyway). Equivalence- preserving — test_wire_payload_golden stays byte-identical. |
||
|
|
3bf32d0649 |
refactor(core): lowering operates on canonical Turns
repair_wire_messages / fold_system_turns / drop_empty_user_turns take and return list[Turn] — the neutral lowering layer (A representation + B validity) now speaks the canonical type. Their intricate content-merge / orphan-detect internals run over the dict projection (reading Turn content blocks would only duplicate turn_to_dict's content logic), so each bridges dicts_from_turns ↔ turns_from_dicts at its boundary; byte-identical. ChatSession._prepare_wire_messages lifts the wire dicts into Turns, runs the lowering passes, and lowers the result back to the dict projection the provider translators (the C layer) consume — the dict bridge now lives in the wire layer, not in _full_messages. Export runs the same repair, reordered before the non-canonical reasoning-content attach (a key the Turn model does not carry). The provider translators keep their dict input by design: they are the format layer that emits provider bytes, the vLLM reasoning-attach is a non-canonical wire concern that sits between lowering and the provider on dicts, and feeding the converters the lowered projection is equivalent to — and simpler than — threading Turn content through them. Wire harness byte-identical; full non-live suite green (7130). |
||
|
|
dbf8d88a73 |
refactor(providers): unify orphan tool-call repair into one send-time pass
Synthesizing a cancellation result for an assistant tool_call with no matching tool result was triplicated across the translators: Anthropic's verbatim-replay (pc_tool_ids) and rebuild branches, and sanitize_messages for the OpenAI-compatible lanes (Chat, Responses, Google). The Anthropic pc_tool_ids branch was also the sole repairer of a native tool_use orphan. Lift it to one neutral policy — lowering.repair_wire_messages — run once in ChatSession._prepare_wire_messages before the translator. It reads tool_calls only, which is sound because the native/tool_calls mirror is enforced at save (normalize_native_for_save): a verbatim-replay orphan is caught via its mirrored top-level call. The translators become pure format translation and carry no orphan synthesis. The neutral cancellation turn carries is_error=True; Anthropic renders it on the tool_result block, the OpenAI-compatible tool message has no such field so sanitize_messages drops it (the C-layer translation of the flag). sanitize_messages keeps one orphan synth of its own: a back-filled empty-id tool_call (local servers that omit ids) is id-less when the upstream repair runs and so invisible to it, so that lane owns its cancellation — preserving the pre-refactor behavior for local servers. reconstruct's load-time strip and the runtime-cancel persist-synth are unchanged. Proven byte-identical against the per-provider wire-payload golden harness (including a new native_orphan fixture); the harness applies the same send-side repair the session does. |
||
|
|
99ba82e8ec |
fix(session): operator-turn wire correctness — framing, empty turns, leading system
Phase-2 follow-ups to the mid-conversation-system consolidation: - user_interjection framing (known #2): a queued message that drains mid-turn is re-framed via render_user_interjection ("The user sent … User message: …") so the user's words keep USER authority, not operator authority — the regression mattered most on the native path, where the turn enters as a real role=system message. Empty/whitespace interjections (e.g. a bare "!!!") are dropped (bug-2). - empty-content user turns dropped at the wire boundary after the fold (known #3): the wake pipeline's synthetic empty send("") leaves an empty user turn on the native path (the nudge stays inline); an empty user message is invalid on every provider. The drop runs after the fold so the fold-path wake turn, which the nudge fills, survives. - leading-system guard (_anthropic): a turn that converts to nothing no longer lets a system message become messages[0] (the API requires messages[0]=user). Newly reachable now that the empty-turn drop can expose it on a fresh-session native wake. - refresh stale .msg.watch-result comments (the card was removed) to describe the current operator-bubble rendering. |
||
|
|
c6b2288302 |
feat(session): consolidate operator-context into first-class system turns
Replace the two operator-context hacks (the <tool_output>/<system-reminder> content envelope and the transient _reminders side-channel) with one persistent {role: system, _source} trajectory turn. Adds supports_mid_conversation_system (claude-opus-4-8): native models take the turn inline; all others fold it into the preceding turn as a nonce-delimited <system-reminder> block declared in the system prompt as the sole trusted marker. Producers (advisories, metacog nudges, user interjections, idle/watch) emit system turns; the envelope/_reminders machinery, escaping round-trip, replay parser, and reminder SSE events are removed. Eager 060 migration un-wraps legacy envelopes. Net -1662 lines.
Known follow-ups from review (unfixed here): (1) the 060 un-wrap heuristic can irreversibly mis-rewrite bare tool rows that resemble the envelope, so do not run the migration until it is tightened; (2) user_interjection turns lost the user-framing/priority preamble (a regression, and a native-path authority-framing concern); (3) native-path wake nudge can emit empty user content.
|
||
|
|
1728a4c0af |
feat(web-search): replace Tavily/DuckDuckGo backends with self-hosted SearxNG
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single self-hosted SearxNG service bundled into the docker-compose stacks. Core: - New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to (backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard unchanged. _resolve_search_client follows storage -> toml -> env -> default precedence (explicit "" disables, via ConfigStore.stored_keys()). - Drop the Tavily-era topic=finance (no SearxNG category); topic is now general/news. Settings/config: - Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key. - Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines, with get_searxng_url/get_searxng_engines. Compose + bundled config: - Internal-only searxng service (no published API port, :ro config, /healthz healthcheck, persistent searxng-cache volume) in both stacks; bundle turnstone/deploy/searxng/settings.yml (JSON output on, limiter off). - Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in). - bootstrap extractor + wheel packaging updated. Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the lxml/h2/brotli transitives). Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG; docs/docker.md carries the AGPL-3.0 §13 operator note. BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg"; tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance. Closes #545 |
||
|
|
c58df26a30 | style(tests): ruff-format provider/registry empty-string tests | ||
|
|
8737081373 | add tests for anthropic api | ||
|
|
85db6895e3 |
make sure empty strings don't get passed to the openapi sdk which don't
allow env var fallback |
||
|
|
9d50e90f67 |
feat(providers): add Claude Opus 4.8 model support
Register `claude-opus-4-8` in the Anthropic capability table. Opus 4.8 shares Opus 4.7's request/response surface exactly — adaptive-thinking only (`budget_tokens` rejected), sampling params removed, the low/medium/high/xhigh/max effort levels, `thinking.display` defaulting to omitted, 1M context, and 128K output — so the entry is a verbatim copy of the 4.7 row. `_lookup_capabilities` longest-prefix matching then resolves date-suffixed ids (e.g. `claude-opus-4-8-20260601`) without colliding with the 4.7 key. No provider code paths change: the existing 4.7 handling already covers all of 4.8's behavior. Models are selected via config.toml / the admin ConfigStore UI, so there is no catalog or dropdown to update. - _anthropic.py: new claude-opus-4-8 capability entry + effort comment - tests/test_providers.py: opus 4.8 bare + dated capability tests - turnstone.example.toml: bump the showcased model example to 4.8 |
||
|
|
32fd8f29c7 |
feat(providers): api_surface toggle + mistral medium reasoning fix (#469)
* feat(providers): api_surface toggle + mistral medium reasoning fix
Mistral medium open-weights served by vLLM expects reasoning_effort via
the Responses API (`reasoning.effort`), not as a `chat_template_kwargs`
entry on Chat Completions. The session was unconditionally injecting
`{"reasoning_effort": ...}` into `chat_template_kwargs` for every
openai-compatible request, which corrupted the prompt rendering for any
backend whose chat template didn't consume that key (Mistral medium,
Mistral cloud, Groq, OpenRouter).
Changes:
- Add `api_surface` ("chat" | "responses") to `ModelConfig.server_compat`
and thread it through `create_provider` / `model_registry.get_provider`.
`openai-compatible` defaults to Chat Completions; operators can flip
individual aliases to Responses for endpoints that support it.
- New `vllm-mistral-medium` profile that pre-fills api_surface=responses
on Detect for known Mistral medium model ids.
- Drop the unconditional `reasoning_effort` injection into
`chat_template_kwargs`. Operators running gpt-oss-style local
templates that consume `reasoning_effort` from the chat template now
opt in via `server_compat.extra_body.chat_template_kwargs`.
- New "API Surface" select in the Models admin tab; allowlist-validated
server-side at create/update time; pre-filled by Detect via the
profile suggestion.
- Evict the cached provider singleton in `ModelRegistry.reload()` when
api_surface changes (previously only cfg.provider triggered eviction).
- Fix `_run_agent` fallback path to inherit the session's primary alias
for capability and server_compat resolution; previously the fallback
passed `alias=None`, which silently dropped per-model caps on the
agent path.
Tests: 5117 passed (-m "not live"); ruff + mypy clean.
* fix(providers): don't auto-suggest Responses for Mistral medium
vLLM's Responses API surface for Mistral medium open-weights doesn't
wire up the Mistral tool-call parser as of vLLM 0.x — tool calls leak
into the response as ``[TOOL_CALLS]<name>{...}`` text instead of
structured tool_calls. Chat Completions on the same engine handles
tools cleanly via ``--tool-call-parser mistral``, and reasoning can be
turned on via the vLLM CLI ``--reasoning-parser`` flag.
Drop the auto-suggest mapping so Detect falls back to the generic
``vllm`` profile. Keep the ``vllm-mistral-medium`` profile definition
in place so an operator who specifically wants per-request effort and
accepts the tool-calling limitation can still pick "Responses API"
manually in the admin UI.
* fix(providers): address Copilot review on PR #469
- providers/__init__.py: drop the redundant *_responses_provider /
*_chat_provider names; have create_provider use _openai_provider and
_openai_compat_provider directly so they're not flagged as unused
globals.
- console/server.py: tighten _validate_api_surface to a strict equality
match against the canonical {"chat", "responses"} set. The previous
strip().lower() membership check accepted ' Responses '/'CHAT' but
stored the raw string verbatim, which then failed to round-trip
through the admin <select>.
- console/static/admin.js: gate the entire server_compat block (server
type, api_surface, extra_body) on provider == "openai-compatible" at
save time so toggling provider away can't leave a stale hidden surface
selection in the persisted capabilities JSON.
- tests/test_session.py: splat the bad kwarg via **dict so CodeQL no
longer flags the call as a wrong-name keyword (the point of the test
is the runtime contract, not the static type).
- tests/test_admin_model_registry_refresh.py: add endpoint-level tests
for the api_surface validation on both create and update — covers the
bogus-value rejection, non-canonical-string rejection, and the happy
path persisting through to the refreshed registry.
|
||
|
|
3aa9f53fd8 |
fix(session): metacog reminders ride a side-channel, not user content
User-channel metacognitive nudges (correction, denial, resume, start,
completion) used to be spliced into ``user_msg["content"]`` permanently,
which leaked the ``<system-reminder>`` envelope into every consumer of
``self.messages`` — UI replay (mitigated by a regex strip in /history),
compaction, title generation, and any future channel adapter that
echoes conversation context. The /history strip was a band-aid;
compaction and title-gen still saw the raw spliced text.
Switch to a side-channel: ``_attach_pending_user_reminders`` writes the
rendered reminder list to ``user_msg["_reminders"]`` (sibling key,
leading-underscore convention shared with ``_attachments_meta`` /
``_provider_content``). At the provider boundary, a new
``_apply_reminders_for_provider`` builds a transient shallow-copy with
the reminder spliced into ``content``; the original message dict
stays clean. ``sanitize_messages`` drops the sibling key on the wire.
Once-per-session-not-per-turn semantics for the wire: after stream
success the loop calls ``_mark_reminders_delivered``, which flips a
``_reminders_delivered`` flag on every user message that carried
reminders into that call. ``_apply_reminders_for_provider`` skips
already-delivered messages so the model sees each reminder exactly
once (the turn it advised). ``_build_history`` ignores the delivered
flag entirely, so reconnecting tabs render the same nudge bubble the
originating tab saw via the live ``user_reminder`` SSE event.
UI surface:
- ``SessionUIBase.on_user_reminder`` enqueues a
``{type: "user_reminder", reminders: [...]}`` SSE event with the
same shape ``_build_history`` surfaces.
- ``app.js`` renders a ``.msg.user-reminder`` bubble (yellow accent,
pill-styled) anchored above the user message it advises, both
live and on history replay.
- ``replayHistory`` renders ``addUserMessage`` before
``addUserReminder`` so the anchor lookup finds the just-rendered
turn (not a prior one).
- Multi-tab caveat documented inline: non-originating tabs receive
no ``user_message`` SSE event today, so a reminder may anchor to
a stale prior bubble until ``/history`` reload corrects it.
Pre-existing bug surfaced by the audit: cancel handlers
(``GenerationCancelled`` / ``KeyboardInterrupt`` / generic
``Exception``) in ``ChatSession.send`` cleared
``_pending_tool_advisories`` but not the user-channel buffer. Both
now drain through a shared ``_drain_pending_advisories`` helper.
Removed the ``/history`` regex strip — the side-channel approach
makes it redundant. Hoisted ``escape_wrapper_tags`` +
``render_system_reminder`` imports to module top (called 2-3× per
turn).
Tests:
- ``TestApplyRemindersForProvider`` — pass-through-by-reference,
string + list content splice, escape on user-typed wrapper tags,
multi-reminder ordering, source-untouched invariant, delivered
flag skip path, fallback for unexpected content shape.
- ``TestMarkRemindersDelivered`` — flag idempotency, no-reminders
no-flag, only marks user messages with reminders.
- ``TestUpdateTokenTableMsgsParam`` — calibration uses pre-built
msgs when provided, falls back when not.
- ``TestUserAdvisoryCancelClear`` — all three cancel branches drain
the user buffer.
- ``TestReminderSidechannelIsolation`` — compaction's
``_format_messages_for_summary`` and the title-gen extraction
loop cannot see reminders by construction.
- ``TestSessionUIBaseUserReminderHook`` — ``on_user_reminder``
enqueues the right SSE shape.
- ``TestBuildHistoryReminderPropagation`` — ``entry["reminders"]``
propagation, absent / empty / multi / coexist-with-attachments
cases, malformed input filtering, all-malformed elision.
- ``test_sanitize_messages_strips_underscore_sibling_keys`` covers
``_reminders`` and ``_reminders_delivered``.
|
||
|
|
fa53b414ed |
feat(providers): add gpt-5.5 and gpt-5.5-pro capability entries (#396)
* feat(providers): add gpt-5.5 and gpt-5.5-pro capability entries
OpenAI announced gpt-5.5 on 2026-04-23 (ChatGPT/Codex first, API
"very soon"). Mirror the gpt-5.4 / 5.4-pro capability shape: 1M
context, native tool search, vision, xhigh effort; pro is
always-reasoning with no temperature and medium/high/xhigh only.
No provider-logic changes needed — OpenAI announced no API-surface
changes vs 5.4. Cache retention already covers 5.5 via the existing
startswith("gpt-5") prefix rule.
* test(providers): cover gpt-5.4-pro and gpt-5.5-pro in cache retention test
Pro variants share the same gpt-5 prefix and should keep 24h
retention; explicit coverage guards against regressions if the
prefix rule narrows in the future.
|
||
|
|
30c89f46c6 |
feat: add Claude Opus 4.7 support (#357)
- Add claude-opus-4-7 capability entry (1M ctx, 128K output, adaptive thinking, supports_temperature=False, thinking_display=summarized) - Suppress temperature param for Opus 4.7 (API returns 400) - Add thinking display opt-in via new ModelCapabilities.thinking_display field - Opus 4.7 omits thinking by default, always send summarized - Add xhigh effort level to mapping and Opus 4.7 effort_levels - Add xhigh/max options to skill template dropdowns in admin console - Align reasoning effort label capitalization across all console dropdowns - Update example config to reference claude-opus-4-7 - 10 new tests with regression guards for Opus 4.6 backward compat Verified against live API: streaming and completion calls succeed. |
||
|
|
eb59cdefda |
feat: pass resolved capabilities through to providers, add server com… (#352)
* feat: pass resolved capabilities through to providers, add server compat layer The LLMProvider protocol previously forced providers to re-derive capabilities from static lookup tables, ignoring config overrides set via the admin UI or config.toml (e.g. thinking_mode, token_param). This adds an optional capabilities parameter to create_streaming and create_completion so the session can pass its config-merged ModelCapabilities through to providers. On top of this, adds a server compatibility layer for local model servers (vLLM, llama.cpp). Profiles suggest thinking mode and server workarounds (skip_special_tokens for vLLM, reasoning_format for llama.cpp) during model detection, with structured admin UI fields for server type, thinking mode, and extra body params. Verified against real vLLM (Gemma 4 31B) and llama.cpp (Gemma 4 E4B) servers. * fix: defensive copy in _finalize_extra_body, expose thinking_param in UI Shallow-copy extra_params and its chat_template_kwargs in the provider before _apply_thinking_mode mutates them, so callers that reuse the same dict across models are safe. Replace the hidden thinking_param input with a visible text field that appears when thinking mode is enabled. Shows the default "enable_thinking" and hints that Granite/DeepSeek use "thinking". * fix: address Copilot review feedback on admin UI and server compat - Preserve unrepresentable thinking_mode values (e.g. "adaptive") in raw capabilities JSON instead of silently dropping on edit round-trip - Validate capabilities and extra body JSON are plain objects, not arrays or primitives - Deep-merge chat_template_kwargs from extra_body instead of silently dropping, so operators can extend/override template kwargs * fix: hide server compat section for non-local providers The Server Compatibility fields (server type, thinking mode, extra body) only apply to openai-compatible (local model servers). Hide the entire section when the provider is openai, anthropic, or google. * fix: normalize capsObj to plain object on edit load Defend against DB rows where capabilities is a JSON literal null, an array, or a primitive — previous code would crash on the capsObj.server_compat / capsObj.thinking_mode reads. Same defensive check also applied to the server_compat nested value. * refactor: extract _isPlainObject helper for JSON type checks Consolidates the null/array/typeof check that was inlined at three different call sites into a single helper. Keeps the intent obvious at each use site and avoids the awkward multi-condition ternary. |
||
|
|
50e6e64c3d |
fix: universal tool_call/tool_result orphan detection for OpenAI-comp… (#346)
* fix: universal tool_call/tool_result orphan detection for OpenAI-compat providers The Anthropic provider had orphan detection for mismatched tool_call ↔ tool_result pairs, but OpenAI-compatible providers (Chat Completions, Google, Responses API) had none. When an Anthropic model runs behind an OpenAI-compat API (e.g. Azure) or cancellation creates orphans, the API rejects the malformed request. - Rewrite sanitize_messages() with orphan detection: synthesize error tool results for unmatched tool_calls, drop tool results with no matching tool_call, fill empty tool_call IDs with positional remap - Call sanitize_messages() from Responses API _convert_messages() * fix: address review feedback on orphan detection - Track answered IDs per-turn (local_answered) instead of scanning all of out, preventing false matches from reused IDs across turns - Drop empty-ID tool results that have no remap entry instead of passing them through with invalid empty tool_call_id - Increment empty_result_idx for every empty result, not just remapped - Remove dead result_ids peek-ahead code - Add test for repeated tool_call IDs across turns |