mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
480a1426b31384f434904a9b6c05b9aacfec13fb
111 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
480a1426b3 |
Fail-closed history-commit handoff (#1005)
* fix(session): fail-closed history-commit handoff (#981) The deleted-workstream discovery is now a terminal, ws_id-keyed latch: keyed conversation commits refuse admission once the durable parent is gone (convergence finalizers and force-abandon are exempt), history handoff refuses to mint a proof token so /history fails closed with a 503 instead of silently wiping the pane, and the SSE stream carries a workstream_gone resync reason. Discarded commits leave a forensic log of commit keys and roles, never content. Conversation rows gain a commit_key (migration 071): keyed saves are idempotent under retry, validated against the full commit identity, and refused when they would cross a workstream deletion. The prune orphan category now requires a NULL alias plus a two-hour updated grace, with cutoffs computed at discovery time and carried into both dialects' rechecks. The mid-turn interjection queue is owner-partitioned with no per-site mode flags: pops take the acting principal's and unowned rows, other participants' rows are structurally retained, and enforcement lives at queue admission plus the shared before_spawn gates. The retraction ledger is bounded by open pop windows: pops open a window atomically with the queue delete, restores close their ids atomically with the ledger consume, every other exit closes through one helper, and misses for unheld ids record nothing. The workstream-gone latch refuses unattended wakes at all three gates (watcher spawn, claim, delivery pre-pop), and the retry dispatcher regained its pre-envelope cancel/error convergence net. Persistence-state reporting derives through the session bound to each UI instead of a registry lookup by id that failed open to healthy during tombstone retention. The dashboard roster no longer re-inserts ghost entries from trailing activity events, the history tool-outcome scan tolerates interleaved non-turn rows, and the shared handoff-deadline handle owns its own retirement. Single-sourced across call sites: keyed-commit row values, attachment save wrappers, tail-truncation and conflict-resolution bodies for both storage dialects; worker-slot lifecycle field sets; the direct-commit admission frame; queued-row layout accessors; the string-aware comment stripper shared by every JS harness suite. Refs #981 #964 * fix(session): sweep handoff fixes to their sibling surfaces The interactive replay loop treated a system row as a tool-batch boundary, so every tool result after an interleaved row vanished from that pane while the coordinator rendered the same history correctly. Only a conversational turn ends the batch window now, matching the shared outcome index. Accepted user turns clear the composer's attachment chips on the same viewer policy that settles optimistic bubbles rather than on having matched a local bubble, so a workstream created with an upload no longer keeps a chip for an attachment the create dispatch already consumed. The coordinator's raced-Stop arm emits the stream-end hook it inherits alongside the idle state, leaving no unfinalized bubble or unflushed tool output. Ending a session surfaces a failure toast when the request never lands or answers with a non-JSON body. The per-second persistence reconcile now probes each session without blocking: a workstream whose generation and handoff locks are held is skipped until the next pass instead of contending the locks every commit needs. The one-shot repair that gates workstream creation at capacity keeps a definite probe — it has no next pass, and the sessions likeliest to be contended are the ones whose unresolved journals emptied its candidate list. Single-sourced: the attachment lane builds its conversation row through the shared commit-identity builder; the ordinary worker exit releases its slot through the lifecycle owner; both operator surfaces snapshot their counters through one non-consuming helper; the replay preamble loses its per-kind wrappers and its config hook; the browser harness suites share one brace walker; and each in-flight history attempt is one record carrying both its abort controller and its deadline. Refs #981 #964 |
||
|
|
9211c4fb29 | chore: download vendored JS files | ||
|
|
766223e774 | feat(judge): parallelize batch evaluations (#991) | ||
|
|
98e96ab5f3 |
Add per-alias model concurrency admission (#990)
* feat(models): add per-alias concurrency admission Add registry-backed FIFO admission limits with queue-aware deadlines and full-stream leases. Expose max_concurrency through storage, admin configuration, OpenAPI, documentation, and diagrams, with role and live backend count coverage. * fix(api): omit null concurrency schema default Keep max_concurrency optional for presence-keyed updates without advertising a null default for its non-null integer OpenAPI shape. |
||
|
|
7a06f5e8bc |
refactor(session): make ModelLane the provider boundary (#979) (#989)
* refactor(session): make ModelLane the provider boundary (#979) ## Summary This closes the model-lane ownership gap left by #832: `ChatSession` no longer stores raw provider/client handles. `ResolvedModelBinding` now carries the provider, client, model, capabilities, registry generation, and backend-auth configuration as one coherent snapshot. - Atomically rebind existing sessions after model-registry changes while pinning each in-flight send, fallback, judge, output guard, task agent, title, compaction, perception, and voice operation to its initiating principal and binding. - Fence UI publication, canonical trajectory folds, durable writes, streams, retries, child scopes, and judge work by generation. Stop can hand off to a successor without accepting late state; cancelled tools retain typed effect receipts, and concurrent approval batches resolve by exact cycle or call. - Make create, fork, open, close, and delete race-safe with hidden `creating` reservations, incarnation-aware state tails, and an ACL-rechecked transaction that clones checkpoint-bounded history, configuration, project/persona state, and attachment references. - Extend REST/OpenAPI and Python/TypeScript SDK contracts for create/fork inputs, routed-create metadata, live-workstream probes, targeted approvals, and structured cancellation results. - Update architecture, storage, authentication, judge, channel, console, API, and SDK documentation, including regenerated architecture diagrams and OpenAPI artifacts. ## Validation - SQLite suite: 11,188 passed, 9 skipped, 10 deselected - PostgreSQL suite: 11,195 passed, 2 skipped, 10 deselected - Live backend: 3 passed - SSE recovery: 6 passed; browser recovery harness passed all scenarios - Ruff: clean; 595 files correctly formatted - mypy: 243 source files clean - TypeScript: typecheck/build and 35 tests passed - OpenAPI artifacts fresh; all 14 changed diagrams reproduce byte-for-byte - `git diff --check` and Git LFS integrity clean Closes #979. * fix(deps): update nanoid for GHSA-2v37-7h3g-55p8 Refresh the transitive lock entry admitted by PostCSS so the TypeScript security gate no longer resolves the vulnerable custom-generator implementation. Validation: - npm ci - npm audit --audit-level=moderate: 0 vulnerabilities - TypeScript typecheck and build - TypeScript tests: 35 passed * fix(test): assert canonical model registry URLs Replace prefix checks with exact canonical base URL assertions so the tests do not model incomplete URL validation. Validation: tests/test_model_registry.py (185 passed); Ruff check/format; mypy. |
||
|
|
aa4371ea99 |
fix(832): retire the dead attempt's armed state in the re-create window
Between a mid-stream death and the next begin_attempt there is no live attempt, but the consumer kept the dead attempt's armed _CancelRef: a Stop in that window re-emitted the discarded splitter carry as fresh content behind a duplicate stream_end, and a walk-preamble failure was classified as another armed death, replacing the operator-actionable stream-death error. end_attempt() now pronounces the attempt dead at partial-capture; the consumer gains a single per-attempt initializer (_reset_attempt), a lane-free constructor (one resolve_lane walk per turn), and a saw-chunk classifier fallback so a never-arming adapter's mid-stream death still classifies mid-stream instead of silently double-rendering the same lane. Wire-preparation failures are typed at the seam: model_turn wraps prepare_wire raises in WirePreparationError, both walk arms forward it verbatim (no health record, no fallback walk — a session-data fault would otherwise paint every backend degraded), the fatal formatter gets a dedicated branch, and the re-issue ladder's last-death mask exempts it alongside BackendAuthUnavailableError so an auth outage mid-turn is not misdiagnosed as a network flap. Riding fixes: the tag-scan gate gets its single spelling (lane_scans_inline_reasoning) shared by drain and display; the citations fold's separator+gate become a shared pair in _protocol; _build_main_lane stops passing config_store (dead derivation — the session's own knobs replace both values it feeds); the debug wire dump is ruled per-invocation (the overflow-recovery re-print is the dump that diagnoses the recovery) and pinned; dead delegates _ensure_tool_call_ids and _finalize_provider_blocks deleted; the parity runner adapts to the pre-fold seam signature by inspection and refuses to record a harness-shape TypeError as a baseline; the streaming provider fakes move to tests/_session_helpers (their tree-wide home) and test_cancel's duplicate helper is deleted; committed parity pins restate their rulings in full; architecture.md's circuit-breaker section is replaced by the real passive health-tracker story and the send-flow diagram stops attributing tool-call assembly to the display consumer; stale pre-fold names and ragged comment paragraphs cleaned. New pins are mutation-probed: disabling end_attempt, the saw-chunk fallback, the auth exemption, or the WirePreparationError arm each fails its pin. |
||
|
|
5a20916eec |
docs+test(832): retire pre-fold names from prose; pin the eager-append contract
The docs sweep re-points every stale reference to the deleted seam (architecture.md's flow diagram and ladder inventory, the lowering and anthropic docstrings, the protocol's shared-rule docstrings that described the pre-fold dual-assembler world). The Protocol's cancel_ref contract is strengthened from 'before the first chunk' to 'inside the call body, before the iterator is returned' — the instant the fold's creation-vs-midstream classifier and health recording key on — and three real-SDK-over-mock-transport tripwires pin it per adapter, so a future lazily-issued generator adapter fails loudly instead of silently reclassifying every pre-first-chunk death. |
||
|
|
b38e9be17c |
docs(session): scope the zero-band geometry claims to the default compact threshold
The drain comment and the architecture docs stated the zero-budget band
relative to the auto-compact threshold as if 0.8 were universal
("well below the auto-compact threshold"); with an operator-set
auto_compact_pct under the ~70% zero point the claim reads inverted.
State the geometry against the DEFAULT threshold and make explicit what
was always true of the mechanism: the trigger's predicate is the
exhausted budget itself, never a threshold, so with low thresholds the
owed path compacts first and the trigger is its bail/insufficient
backstop.
|
||
|
|
6c4c848a08 |
fix(session): survive tool-result truncation at zero context budget (#883)
At an exhausted context budget the drain loop replaced every tool result with a placeholder that read as a successful-but-trimmed call. For structural results — spawn_workstream's ws_id, the tasks scratchpad — the model lost the handle orchestration depends on and silently stalled, while the UI (told the real summary before the drain) kept showing success. Worse, the budget zeroes near 70% fullness when max_tokens ≥ context_window/4, well below the 80% auto-compact threshold, so a stalled coordinator could sit in that band indefinitely with no compaction ever firing. Three guarantees at the truncation seam, one renewal trigger at the drain: - structural-tool and error results get a guaranteed 2048-char admission floor (head+tail beyond it) — never the zero-budget drop - any result at or under the floor passes verbatim (denial notices, spawn acks: never destroy what is smaller than the guarantee) - bulky non-structural results get an explicit drop notice stating the call RAN but its output could not be admitted — never a trim impersonation the model cannot distinguish from success - a zero truncation budget triggers one mid-turn compaction (no threshold_pct — none was evaluated, same rule as the ctx-overflow retry), closing the 70-80% band where the budget zeroed but compaction was never owed Background-bash spawn acks ride the small-result pass; a name-keyed floor cannot distinguish them from foreground bash — see #891. |
||
|
|
e4604c278e | chore: download vendored JS files | ||
|
|
747177a76c |
fix(providers): review round 9 — orphan/harvest collision, shared shim gate, retired-id rationale
Correctness: - Responses: orphan argument deltas (streamed without any output_item.added) now count as a streamed tool-call signal, so the terminal harvest stands down instead of re-emitting the same call onto the same slot — the reproduced collision concatenated the arguments JSON into an unparseable double copy. Cleanup / documentation: - finish_shim_due in _protocol is THE gate for the lax-server finish shim — one predicate (and one definition of 'delivered output') for all three adapter families, so the same capability flag cannot acquire per-family completion semantics. - The Responses error/response.failed branches share one failure tail (only code/message extraction differs) — the same server failure can never become retryable through one event type and fatal through the other, pre- or post-terminal. - _format_refusal pins the refusal rendering the streamed event and the terminal harvest both use. - The capability-table floor comment and CHANGELOG Removed entry now state the real rationale: OpenAI has RETIRED the pruned ids from the API — the rows described unreachable contracts, not unpopular ones. - CHANGELOG names the stream-entitlement break class (verified-org streaming, pre-stream_options gateway api-versions) with its serving-side remediation; deliberately no non-streaming fallback. - docs/architecture.md retry section describes the collapsed transport: the two stacked retry ladders, IncompleteStreamError / ResponsesStreamFailedError retryability, finish_reason_optional remediation; stale non-streaming mentions updated (+ puml). - Anthropic whole-block emission carries its residual hybrid-gateway bet as an explicit comment. Held on standing rulings: post-finish usage forfeiture (keep result + warn, rounds 4/8), session merge_usage twin and StreamAbortRef twin (#832), stream_options wire delta (round 2, caveat now names Azure). |
||
|
|
08580f25f9 |
fix(providers): review round 1 — streaming parity gaps the collapse exposed
Correctness (4 confirmed + 1 plausible fixed, 2 accepted+documented): - Anthropic _iter_anthropic_stream handles citations_delta: text-block citations now ride the raw block into provider_blocks, as replay requires (the retired non-streaming lane preserved them via model_dump; the streaming lane dropped them — a pre-existing main-loop gap the collapse would have extended to single-shot lanes). - Anthropic text blocks separate with "\n" at each subsequent block start, restoring the retired lane's "\n".join rendering on drained lanes AND un-fusing streamed web-search responses in the chat loop. - response.failed raises typed ResponsesStreamFailedError, listed in the provider's retryable_error_names — retry loops treat an in-band failure like the wire errors it stands in for instead of hard-stopping on a bare RuntimeError (judges keep their heuristic fallback after retries). - drain_stream folds a finish-less stream's terminal citations footer (suffix rule: pre-finish info invalidated by any later payload), so lax compat servers that never send finish_reason keep their Sources. - usage max-merge extracted as merge_usage() in _protocol.py — the one definition drain uses now and the session's inline consumer adopts on #832. Accepted + release-noted instead of coded around: strict pre-2024 compat servers that 400 on stream_options (such a server already cannot serve the chat loop; CHANGELOG caveat extended), and repeated-index parallel tool-call merging on legacy compat servers (identical to the main loop's accumulator semantics; a shared guard belongs in the #832 unification). Cleanup: run_with_deadline grows on_abandon (best-effort, cannot mask the deadline error) and both judges drop the copy-pasted abort choreography; StreamAbortRef documents the _CancelRef adoption plan; test_model_turn's fake replays through the shared as_stream adapter; docs/architecture.md drops the retired Protocol row. Tests: refusal handler pinned (was advertised, untested); typed-failed retryability; citations capture; text-block separator (plus the mixed text+search expectation updated for the separator chunk); finish-less citation fold; on_abandon firing matrix; StreamAbortRef arrival race. |
||
|
|
b450b9ad20 |
fix(providers): align GPT-5.6 with the GA API surface
- every 5.6 tier accepts effort "max" and reasoning.mode
"standard"/"pro" (GA docs: pro is a request mode on any GPT-5.6
model) -- drop the Sol-only gating
- GPT-5.6 deprecates prompt_cache_retention; send
prompt_cache_options={"ttl": "30m"} (its only supported lifetime)
and keep the 24h retention policy for pre-5.6 models
- never inject commercial cache params into local lanes: dropped from
the Chat Completions lane (which serves only openai-compatible and
google) and gated off the compat-pinned Responses lane -- a gpt-5*
served-model name is not an OpenAI account
- account cache writes: usage *_tokens_details.cache_write_tokens
flows into cache_creation_tokens (5.6 bills writes at 1.25x the
uncached input rate)
- drop non-string verbosity/reasoning_mode overrides with a warning
instead of raising on unhashable capability-JSON values
- keep ModelCapabilities' public positional prefix stable by appending
the verbosity/pro fields at the tail; pin it with a constructor test
- openai floor 2.44 -> 2.45, the first release with the typed
prompt_cache_options kwarg
|
||
|
|
1035fe05eb |
fix(console): effort annotations say in plain words what the request carries
Aliased knob positions were labeled after the lowest sibling sharing
their wire token — a toggle-only model rendered 'Max (= minimal)',
implying a minimal-effort downgrade the wire doesn't contain, and with
declared values 'High (= minimal)' while the wire carries high. Each
position now states its delivered level: exact matches stay plain
('Max'), snapped positions say 'Low — sends high', the adaptive none
position warns 'thinking stays on', budget detail stays in the
tooltip. Effort-param placeholder corrected to the real graded keys
(reasoning_effort / reasoning).
|
||
|
|
530958e06b |
fix(providers): the session effort level always reaches the local-lane wire
Local lanes dropped the knob's graded value unless the operator declared
reasoning_effort_values (and, on the template channel, an effort key) —
picking Max sent a bare thinking toggle and the effort select
degenerated into seven positions that all meant 'on'. The user's
setting now always rides:
- openai-compatible: the flat reasoning_effort param carries the knob
verbatim (effort_passthrough on the lane default); declared values
still snap ordinally, and a declared effort_param still claims the
template channel and suppresses the flat param.
- anthropic-compatible: the graded value rides chat_template_kwargs
alongside the toggle whenever reasoning control is engaged — under
the operator's effort_param, else the conventional fallback key
(reasoning_effort); templates that don't reference the kwarg ignore
it. thinking_mode=none still injects nothing.
- Commercial lanes untouched: empty declared values still mean 'no
effort control' (o1-mini) and the ordinal snap is unchanged.
Golden writer now pins ensure_ascii=False: the baselines' literal em
dashes came from a hand edit (
|
||
|
|
e136237b63 |
fix(providers): openai-compatible never consults the commercial table
Local-lane model ids are operator-chosen strings (vLLM --served-model-name), so a prefix collision with a cloud model id inherited that model's sampling and effort contract: a box named o3-distill silently lost temperature support, and one named gpt-5.5-my-finetune was sent gpt-5.5's snapped reasoning_effort values it never declared. Both surfaces of the lane now return plain defaults (OPENAI_COMPAT_DEFAULT in _openai_common): the chat class directly, and the responses pin via a compat-mode OpenAIResponsesProvider mirroring AnthropicProvider(compat=True). Everything beyond the defaults is declared by the operator on the model definition, matching the anthropic-compatible lane and lookup_model_capabilities' documented 'no static table for local models' contract. The commercial openai lane (Responses-only) is untouched. Pre-split tests that reached commercial rows through the chat-class OpenAIProvider alias now source them from lookup_openai_capabilities; their subject (registry rows + shared gating helpers) is unchanged. |
||
|
|
3607517814 |
fix(providers): registry effort truth — o-series/gpt-5.5/codex-max/sonnet-5; forward declared none
Capability-registry corrections verified against the official OpenAI reasoning guide, the Azure reasoning-models matrix (2026-06 revision), and the Anthropic models-overview/effort/migration pages (2026-07): OpenAI (vocabulary confirmed none/minimal/low/medium/high/xhigh — no "max" level exists; knob max rides the xhigh ceiling via the ordinal snap): - o1/o3/o3-mini/o3-pro/o4-mini declare low/medium/high (every o-series model except o1-mini) — without declared values the session knob was silently dropped for these models. o1-mini stays effort-free. - gpt-5.5 default corrected none -> medium (5.5 reasons by default, unlike 5.1-5.4). - gpt-5.1-codex-max gets an explicit row: it prefix-matched the gpt-5.1 row (no xhigh), capping the knob's xhigh at high on the one model xhigh was introduced for. Anthropic (effort-page matrix): - claude-sonnet-5 row added — it previously fell through to _ANTHROPIC_DEFAULT (manual budgets, 200k ctx, no effort), all wrong: adaptive-by-default thinking (manual budgets are a 400), sampling params rejected, 1M ctx / 128k out, effort low..max incl. xhigh. - claude-sonnet-4-6 gains its documented "max" effort level (knob xhigh now rides max, not high) and the stale 64k max_output becomes the documented 128k. - fable-5 / opus-4-8 / opus-4-7 / opus-4-6 / opus-4-5 rows verified correct as declared. Knob semantics completed: resolve_reasoning_effort now forwards the knob's "none" position verbatim when the model DECLARES an explicit none level (gpt-5.1+, grok-4.3) — omitting the param there leaves a reasoning-on server default (gpt-5.5: medium) in charge of a knob that promises off. Models without a declared none still omit, and none is never a snap target. Parity harness swaps its synthetic openai shape for the real gpt-5.5 registry row. |
||
|
|
a0e04a8588 |
fix(providers): effort snapping is ordinal — round up, cap at the ceiling
The knob domain grew xhigh/max after the snapping fallbacks were written, which silently inverted their semantics: off-list meant "unrecognized string" then, but now usually means "above the model's ceiling", where falling back to the default tier is directionally wrong (grok-4.3 at knob max got low; values low/medium/high at knob xhigh got medium; Anthropic manual mode gave xhigh/max a 4096 budget while high got 16384). One rule everywhere now, via snap_reasoning_effort in _protocol: exact match wins; otherwise the smallest declared level ranking at or above the knob; above the ceiling, the ceiling. "none" is never a snap target, and default_reasoning_effort only catches values the ordinal snap cannot rank. - resolve_reasoning_effort (flat chat / responses / validated effort_param lanes) snaps ordinally: xhigh over (low, medium, high) now sends high; xhigh over DeepSeek-style (high, max) sends max — matching DeepSeek's official xhigh-to-max aliasing, so a declared values list now reproduces that contract instead of defeating it. - _map_reasoning_to_effort (native output_config) rounds up too: knob xhigh on Opus 4.6 (low, medium, high, max) rides max instead of silently dropping output_config. - EFFORT_BUDGET_MAP is monotone across the whole knob domain: minimal/low 1024 (API floor), medium 4096, high 16384, xhigh 32768, max 65536. Unknown strings still fall to the 4096 default. Google defaults are unaffected (ceiling and default coincide at high); wire goldens unchanged. Parity harness caught the budget clamp interacting with its own max_tokens during development — capture budget raised above the largest manual budget. |
||
|
|
59a527f2f2 |
feat(console): surface each model's effective effort ladder
Seven knob positions render as seven behaviors in the UI, but the real
ladder depends on the lane and the model: qwen3.6 has two (off/on),
DeepSeek-V4 three, Claude 4.6 five. Operators had no way to see which
positions alias — the confusion class behind silently-equal effort
levels.
providers/effort_ladder.py projects the knob domain through the same
mapping functions the providers use at request time (resolve_reasoning_
effort, reasoning_template_kwargs, the manual budget map — hoisted to a
shared constant so the projection can't drift), yielding
{value, effective} rows where equal tokens promise identical requests.
/v1/api/models rows now carry the ladder (guarded per row), and
POST /v1/api/admin/models/effort-ladder computes it for the admin
modal's unsaved edits.
The admin per-model effort select and the skill launch-config effort
select annotate aliased positions ("Max (= high)", "None (model
default)") with a sends-tooltip; annotations refresh as thinking-mode /
effort-param / capabilities fields change. The ladder describes what
Turnstone sends — server-side templates may alias further (DeepSeek-V4
folds low/medium into its default high tier).
|
||
|
|
c64dc16319 |
feat(console): Always-on thinking-mode option in the model form
Post effort-knob rework, "Enabled" (manual) means knob-controlled — effort none turns thinking off. Operators who want the pre-#771 always-on behavior (knob never disables) previously had to hand-write thinking_mode "adaptive" into the raw capabilities JSON. The dropdown now offers all three representable modes — None / Effort-knob controlled / Always on — and the edit-load lift captures adaptive instead of relegating it to raw JSON. |
||
|
|
d564cee43d |
docs(providers): ground the effort_param values caution in official template contracts
Cross-checked online: Qwen3.6's template documents enable_thinking + preserve_thinking only — no effort parameter exists (vLLM's flat reasoning_effort convenience boolean-maps to the same toggle). DeepSeek-V4 officially accepts reasoning_effort high/max with Think High as the default thinking tier and low/medium→high, xhigh→max aliasing — so freeform effort_param passthrough matches the contract exactly, and a declared values list omitting xhigh/max would make Think Max unreachable. Live probes on both boxes agree with the official contracts once the default-tier framing is applied. |
||
|
|
2cf23b6fe2 |
fix(providers): address high-effort review of the reasoning-knob branch
Verified findings applied: - adaptive thinking_mode never knob-disables: the shared mapping now sends the toggle unconditionally true for adaptive (the native adaptive branch ignores the knob's none), while manual keeps the knob-driven contract. Restores the invariant the deleted chat-lane code upheld. - a set effort_param suppresses the flat top-level reasoning_effort on the chat lane: the template channel replaces it — double-sending could 400 on schema-strict servers and disagree with operator pins. - admin edit-save no longer drops a stored thinking_param when the thinking-mode dropdown is empty: the raw-JSON strip now only fires when a mode value actually round-trips through the dropdown. - effort_param persistence gated on the local-server lanes so a value lingering across a provider switch never lands on commercial rows. - three stale _compat_extra_params references renamed to merge_reasoning_template_kwargs. Documented dispositions (no code change): the knob-none-disables flip on upgrade is intentional and now carries an upgrade note; gateways fronting real Claude belong on provider=anthropic with a custom base_url (the compat lane is vLLM-schema-only); nonstandard thinking_mode strings staying inert is the intended allowlist contract. The real anthropic provider is unaffected throughout — official Claude models keep native thinking/output_config. |
||
|
|
68b22adfa3 |
feat(providers): share the effort-knob→chat_template_kwargs mapping with the openai-compatible lane
Hoist the compat-lane injection into _protocol.merge_reasoning_template_kwargs (next to ModelCapabilities — one implementation for both local-server lanes) and retire OpenAIChatCompletionsProvider._apply_thinking_mode in its favor: _finalize_extra_body now receives the session effort knob, so thinking_mode manual/adaptive maps knob "none" to an explicit thinking_param false (previously the toggle was unconditionally true) and caps.effort_param carries the graded effort key on chat completions too. Operator server_compat pins still win; the Responses surface is untouched (native reasoning handles effort itself). The admin Models form grows an "Effort param" field that round-trips like thinking_param: lifted out of the raw capabilities JSON on edit-load, re-added on save, cleared by emptying the field. Verified live against qwen3.6-27b /v1/chat/completions: knob medium streams reasoning_content, knob none suppresses it. |
||
|
|
9289693730 |
fix(providers): drive reasoning via chat_template_kwargs on the anthropic-compatible lane
The compat lane sent no reasoning control at all: vLLM's /v1/messages has no thinking request field, thinking_mode stayed "none", and the session effort knob was silently dropped. The reasoning levers live in the chat template, so fold them into extra_body chat_template_kwargs (_compat_extra_params): thinking_mode manual/adaptive maps the knob onto caps.thinking_param (effort "none" = off, mirroring the native manual-mode contract), and caps.effort_param (new ModelCapabilities field) carries a graded effort value for gpt-oss-style templates, validated against reasoning_effort_values when declared. Operator server_compat entries win on key collision; native thinking params, temperature forcing, and output_config never fire on compat. resolve_reasoning_effort moves from _openai_common to _protocol next to ModelCapabilities — importing it into _anthropic would otherwise cross provider families. The admin Models form now shows and round-trips the thinking-mode dropdown for this lane; the #661 hide was premised on thinking_mode being inert here, which this change inverts. Verified live against qwen3.6-27b on vLLM /v1/messages: knob medium streams a thinking block, knob none suppresses it, an operator pin beats the knob. |
||
|
|
7053439e84 |
refactor(eval): split measurement core from prompt optimizer
turnstone-eval was misnamed: it was a prompt optimizer, not a measurement harness. Split the 3252-line turnstone/eval.py into a strictly one-way dependency (optimizer -> eval-core; core never imports the optimizer): - turnstone/eval/core.py measurement substrate — everything up to and including _run_iteration: provider detection, NullUI, HeadlessSession, the test runner, score_run, aggregation, and neutral reporting. - turnstone/eval/cli.py new measure-only `turnstone-eval` — the old --no-optimize path promoted to the whole job (one _run_iteration call, then print the summary table). - turnstone/optimizer.py the UCB self-modify loop and its multi-agent pipeline (analyst/optimizer/observer/diversifier/tool optimizer), now `turnstone-optimizer`; imports from eval.core only. - turnstone/eval/__init__.py re-exports the core public API for back-compat (score_run, _match_action, _run_iteration, HeadlessSession). _apply_tool_overrides lives in core (HeadlessSession needs it) rather than alongside the other tree helpers, so the dependency stays one-way. Breaking change: `turnstone-eval` now measures; use `turnstone-optimizer` to optimize. Both code paths are behaviour-preserving — the moved function bodies are byte-identical. |
||
|
|
0d6d7ebae1 |
docs(personas): concept doc, CHANGELOG 1.7 entry with /creative breaking note
docs/personas.md covers the four levers, the resolve-once/stamp-forever snapshot semantics, the seed matrix, per-surface selection, authoring rules, and RBAC; architecture.md's config-persistence paragraph swaps the removed creative_mode for the persona stamp. |
||
|
|
c0be383f99 |
refactor(doctor): replace turnstone-bootstrap with turnstone-doctor (#718)
* refactor(doctor): replace turnstone-bootstrap with turnstone-doctor turnstone-bootstrap was an LLM setup wizard for Day-0; run.sh now owns install. Repurpose its LLM/conversation plumbing into turnstone-doctor — a diagnose-only tool for a running cluster. - Preflight detects the install kind (docker-compose/systemd/pip/source) from config.toml + TURNSTONE_* env, with secret redaction. - Self-configuring brain resolves the cluster's own model from config/env/storage read-only (no migrations, no create_all), falling back to interactive selection; the attempt itself is the LLM-backend health check. - Deterministic version check: installed version, cluster drift via the console's authoritative /health, and latest upstream stable/experimental (offline-safe). - Read-only diagnostic tools (read_file, compose/systemd/journal, http_health, check_llm_backend, node_health, finish) behind one secret-scrubbing chokepoint; no generic shell, so read-only is structural. - node_health reaches a node the right way for the detected install kind (exec-into-container for compose, direct HTTP otherwise), overridable per node for mixed clusters. - mTLS-aware: forwards [database] SSL params and reports node-mesh mTLS instead of mislabelling healthy nodes "unreachable". init_storage gains a backward-compatible create_tables override for read-only opens. Entry point turnstone-bootstrap -> turnstone-doctor; README/QUICKSTART/ architecture/docker docs, the bundled compose header, run.sh, and the CI smoke updated. CHANGELOG deferred. * fix(doctor): address Copilot + CodeQL review findings on #718 Validated all seven review findings (none false positives) and fixed: - check_llm_backend now applies the same scheme / metadata-host guard as http_health (extracted to _assert_safe_http_url), so a model-supplied base_url can't be steered at the cloud metadata endpoint or a file:// URL. - node_health no longer double-appends the default port when the operator passes host:port (regression: 10.0.0.5:8081 -> http://10.0.0.5:8081:8080). - node_health install_type enum uses "git-source" to match the label the rest of the module and the prompt/report show the model (a schema-strict provider would otherwise reject the value the model is told to use). - _read_api_creds takes base_url + api_key as a unit from the first config source that defines either field, then env-fills, instead of splicing the two across different config files into a pair that exists in no real config. - _mask_secrets masks assignment-shaped content inside comment lines, so a commented-out real secret can't leak through read_file / the report; prose comments (no KEY=value shape) still pass through untouched. - drop the mixed import styles CodeQL flagged in doctor.py and test_doctor.py. Adds 5 tests; ruff + mypy clean; full doctor suite passes (129). |
||
|
|
108714a48d |
fix(auth): isolate server/console session cookies by name
The server (:8080) and console (:8090) both set a cookie named `turnstone_auth`. Cookies ignore port (RFC 6265), so on a shared host (localhost dev, the Electron build, single-box installs) logging into one surface overwrote the other's cookie and 401'd the first session. Give each surface its own cookie name -- `turnstone_auth_server` / `turnstone_auth_console` -- threaded as a required `cookie_name` argument through the cookie builders, `check_request`, `AuthMiddleware`, and the six shared auth handlers (login/logout/setup/whoami/refresh/oidc_callback). Each app passes its own constant; the parameter is required (no default) so a forgotten caller fails loudly instead of silently reverting to the legacy name. Names key on role, not node: the cluster shares one JWT identity and the console->node proxy re-mints a bearer token (dropping Set-Cookie), so per-instance names would break identity portability and aren't used. Hard cutover: the legacy `turnstone_auth` cookie is no longer read and self-expires within its 24h TTL (one forced re-login). JWT audience was already enforced, so the shared cookie was a session clobber, not an auth bypass. |
||
|
|
7ef04e576a |
fix(providers): require base_url for anthropic-compatible
Copilot review on #661: empty base_url let the SDK fall back to https://api.anthropic.com, sending compat-shaped requests to the commercial API. The lane is local-only by definition, and the /v1-strip edge case already established fail-loudly-over-silent-prod-retarget; apply the same principle to the empty case. create_client raises an actionable ValueError; the admin Detect path surfaces it as a clean error string via probe_model_endpoint's existing handler. |
||
|
|
12bd848c68 |
feat(providers): anthropic-compatible lane for local /v1/messages servers
Add provider id "anthropic-compatible": the existing AnthropicProvider pointed at Anthropic-compatible local servers (vLLM /v1/messages), mirroring the openai/openai-compatible split. Registry-only — configured via the admin Models tab or [models.*] toml, not exposed on the bare --provider flag, so the CLI/server prod-URL defaults are unreachable for the lane and real-Anthropic behavior is untouched. Lane behavior (live-verified against vLLM 0.22.1rc1 + DeepSeek-V4-Flash): - Capability defaults replace the Claude static table: token_param max_tokens, thinking_mode none, web_search/tool_search/vision off, reasoning replay on. vLLM rejects Anthropic server-side tool types (tools require input_schema) and ignores the thinking request param, so neither is sent; thinking blocks still stream back and round-trip through the native lane verbatim. - Reasoning toggles via server_compat extra_body chat_template_kwargs (first-class vLLM request field; request-level keys beat server defaults). _build_thinking_and_kwargs forwards non-internal extra_params as SDK extra_body; thinking_budget_tokens stays internal. - No temperature force: thinking_mode none skips the Claude-only temperature=1.0 requirement. Admin UI: provider option + URL placeholder (base_url without /v1 — the SDK appends /v1/messages); the server-compat section shows only the extra-body field for the lane. thinking_mode round-trips through the form dropdown for every provider except anthropic-compatible, where it stays in the raw capabilities JSON — the edit-load lift and save restore use the same predicate so stored overrides are never silently dropped. Docs: architecture.md gains the lane subsection incl. verified quirks (thinking param dropped by vLLM; stop_sequences cut inside thinking and report end_turn; usage has no cache fields; images need a multimodal model; mid-conversation system turns are per-model opt-in). Negative-tested: removing the _INTERNAL_EXTRA_PARAMS exclusion fails test_internal_keys_not_leaked; the live test drives a streamed turn with the chat_template_kwargs toggle and asserts no reasoning deltas. |
||
|
|
b3c1b9c9e0 |
build: promote anthropic, postgres, console, tls to core dependencies
The Anthropic SDK provider was the lone first-class provider gated behind an optional extra, while OpenAI ships in core and Google rides the OpenAI-compatible path. Fold anthropic, psycopg (postgres), croniter (console), and lacme (tls) into the base dependency set so a default `pip install turnstone` yields a complete single- or multi-node deployment; only the Discord/Slack channel gateways stay optional. - pyproject: four extras → base deps; `all` is now discord+slack; drop the redundant croniter from the `test` extra; regenerate uv.lock. - ci: the postgres test job installs `.[test]` (psycopg is base now). - providers: `_ensure_anthropic` becomes a thin SDK accessor for `create_client`; drop the now-redundant eager import-guard calls from the streaming/completion hot path (anthropic is always present). - bootstrap: import anthropic directly. - tests/docs: drop the anthropic importorskips and stale extra-install hints. |
||
|
|
110d44b07e |
refactor(tools): remove man, math, and plan_agent built-in tools
`man` and `math` duplicated capabilities already reachable through `bash`; `plan_agent` is better expressed as a `task_agent` running a planning skill, and carried a large amount of special-case machinery (plan-review gate, refinement loop, per-kind model routing). Removing all three shrinks the tool surface and cuts per-call token cost. Also removed, as dead-once-the-tools-are-gone: - the `math` sandbox executor (`turnstone.core.sandbox`) and its `[sandbox]` extra; the eval analyst now runs bash-only - the read-only `AGENT_TOOLS` sub-agent tool set and the `agent` tool-metadata key (`task_agent`/`TASK_AGENT_TOOLS` retained) - the plan-review protocol end to end: the `on_plan_review` UI hook, `resolve_plan`, `POST /v1/api/plan` + `POST /v1/api/route/plan`, the `plan_review`/`plan_resolved` SSE events, and their Python SDK / TypeScript SDK / OpenAPI / frontend / Discord+Slack bindings - the `model.plan_alias` / `model.plan_effort` settings and the registry `plan_model` / `plan_effort` routing fields TOOLS 31->28, TASK_AGENT_TOOLS 13->11; COORDINATOR_TOOLS unchanged. BREAKING CHANGE: removes the `man`, `math`, `plan_agent` tools, the plan-review SSE/HTTP/SDK surface, and the plan_* model-routing settings from the experimental 1.6 line. |
||
|
|
1728a4c0af |
feat(web-search): replace Tavily/DuckDuckGo backends with self-hosted SearxNG
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single self-hosted SearxNG service bundled into the docker-compose stacks. Core: - New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to (backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard unchanged. _resolve_search_client follows storage -> toml -> env -> default precedence (explicit "" disables, via ConfigStore.stored_keys()). - Drop the Tavily-era topic=finance (no SearxNG category); topic is now general/news. Settings/config: - Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key. - Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines, with get_searxng_url/get_searxng_engines. Compose + bundled config: - Internal-only searxng service (no published API port, :ro config, /healthz healthcheck, persistent searxng-cache volume) in both stacks; bundle turnstone/deploy/searxng/settings.yml (JSON output on, limiter off). - Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in). - bootstrap extractor + wheel packaging updated. Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the lxml/h2/brotli transitives). Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG; docs/docker.md carries the AGPL-3.0 §13 operator note. BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg"; tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance. Closes #545 |
||
|
|
ee8dc7c1c3 |
refactor(history): project the /history wire shape server-side
Collapse the three hand-synced "raw storage -> render shape" projections into one server-side projection. The projection previously lived in a test-only `_build_history` (SSE-era reference impl), a client-side JS normaliser (`history_normalize.js`, the transitional bridge), and coord's inline `init()` handling -- drifting silently with no parity test. Add `project_history_messages` to `history_decoration.py` and run it as the final step of the `make_history_handler` pipeline (load_messages -> decorate -> extract_reasoning -> project), so `GET /history` emits the canonical render shape directly: flat tool_calls (with verdict / output_assessment), top-level source / reminders / attachments, collapsed multipart content, derived denied / is_error / pending, reasoning, and advisories. Interactive `replayHistory` now consumes the payload verbatim. Close two gaps the JS bridge deferred: - list-content <tool_output> advisory extraction (decorate handles only string content; the projection extracts list-carrier advisories, then joins remaining text parts to the string the renderers require); - orphan->pending marks ONLY the last orphan tool-call turn, so a mid-conversation cancelled tool still renders instead of vanishing. Delete `history_normalize.js` (+ its <script> tag and node test) and the test-only `_build_history` (+ orphaned imports); retarget its direct tests onto the projection helpers. Update the WorkstreamHistoryResponse description and the Web UI Resilience architecture note to the projected shape. Coord's `init()` still reads the raw side-channels; migrating it to the projected shape is the next commit, browser-verified separately. Refs #549. |
||
|
|
7404ae46db | chore: download vendored JS files | ||
|
|
6998b442a9 | chore: download vendored JS files | ||
|
|
f28a3533a2 | chore: download vendored JS files | ||
|
|
6f8574eef3 |
fix(reasoning): synthesize reasoning_text alongside non-reasoning provider_blocks
GoogleProvider attaches raw tool_call dicts as ``provider_blocks`` on the finish chunk for ``thought_signature`` round-trip (``_google.py:_iter_stream``). When the same turn streamed Gemini's ``reasoning_content`` as ``reasoning_delta`` chunks, the prior synthesizer bailed out the moment ``provider_blocks`` was non-empty — so the captured reasoning was visible live but lost on page reload. Replace the early-return-if-non-empty check with a reasoning-bearing type test (``thinking`` / ``redacted_thinking`` / ``reasoning`` / ``reasoning_text``). When none of those types appear, append the synthetic ``reasoning_text`` block to the existing list rather than replacing it — preserving Google's tool-call fidelity blocks. Also addresses two doc-accuracy review findings: - ``LLMProvider.extract_reasoning_text`` docstring no longer claims OpenAI Chat / Responses are unwired (Phase 3+4 shipped extractors). - Add the method to the Protocol methods table in ``docs/architecture.md`` (was missing alongside the class diagram). |
||
|
|
20e1e7b110 |
fix(reasoning): apply Copilot review feedback + docs sync
PR #498 round-robin review surfaced 5 findings. 4 applied; 1 rejected with rationale. Applied * **Copilot finding 5** (history_decoration.py:341): dispatcher inspected only ``provider_content[0]['type']``. OpenAI Responses captures EVERY ``output_item.done`` event into ``provider_blocks`` (not just reasoning) — in practice the order is ``[reasoning, message, ...]`` but the API doesn't guarantee that; a hypothetical ``[message, reasoning]`` ordering would silently drop the reasoning under an index-only check. Now walks the list for the first block whose type is in ``_BLOCK_TYPE_PROVIDER_FACTORY``, then dispatches the WHOLE list to that provider's extractor. Each provider's extractor already filters internally by its own block type, so passing the full list is correct. Regression test added (``test_dispatcher_scans_past_unrecognized_first_blocks``). * **Copilot finding 3** (migration 052 docstring): the previous review-fix wave used sed to rename ``persist_reasoning`` → ``surface_persisted_reasoning`` everywhere, which mangled a historical reference in the migration docstring ("The earlier name ``surface_persisted_reasoning`` was renamed..."). Restored to point at the actual pre-rename name (``persist_reasoning``). * **Copilot finding 4** (sdk/typescript/src/events.ts:26): ``HistoryEvent`` JSDoc still referenced ``persist_reasoning`` — the sed rename only walked ``turnstone/`` and ``tests/``, missing the TypeScript SDK. Updated to ``surface_persisted_reasoning``. Also widened the comment to cover all three reasoning-bearing block types (Anthropic ``thinking``, OpenAI Responses ``reasoning``, synthetic ``reasoning_text``) instead of mentioning only Anthropic. * **github-code-quality finding** (session.py:1120): ``_resolve_server_type`` had a bare ``except Exception: pass``. Replaced with a ``log.debug(..., exc_info=True)`` + explanatory comment. Behaviour unchanged (still returns ``""`` on any lookup failure); failures are now observable under DEBUG triage. Rejected (with rationale) * **github-code-quality finding** (_protocol.py:265): ``extract_reasoning_text``'s body is ``...`` per ``LLMProvider`` Protocol convention. Every method in the file uses ``...`` (PEP 544 idiomatic Protocol style). Changing only this one to ``raise NotImplementedError`` would be inconsistent with the rest of the file. CodeQL's "statement has no effect" warning is technically correct for ``...`` as a standalone expression but ignores the documented Python Protocol convention. No fix. Docs sync * docs/api-reference.md: ``history`` SSE event message-shape table gains the optional ``reasoning`` field. * docs/architecture.md: ``ModelCapabilities`` row in the type table gains ``supports_reasoning_replay``; ``StreamChunk`` and ``CompletionResult`` rows gain the existing ``provider_blocks`` field (was missing pre-PR). New "Per-model reasoning persistence" subsection under the Models config section, documenting the two flags + capability gate + three reasoning paths + cross-provider shape filter. * docs/settings.md: new "Reasoning persistence (per-model)" subsection with the two-flag table and capability-gate note. * docs/diagrams/03-core-engine-classes.puml: ``LLMProvider`` interface adds ``extract_reasoning_text`` + the new ``replay_reasoning_to_model`` kwarg; ``ModelCapabilities`` class adds ``supports_reasoning_replay``. PNG regenerated. Lint + test gate * ruff check + ruff format clean. * mypy clean (191 source files). * pytest -m 'not live' — 6116 passed (3 deselected), +1 net new test (``test_dispatcher_scans_past_unrecognized_first_blocks``). |
||
|
|
57cb09c871 |
docs(sse): document state_change + in_progress_snapshot events
Updates the docs that describe the per-workstream SSE event stream and the SessionUI lifecycle to match the refresh-resume changes: - api-reference.md: documented the `state_change` event (previously undocumented despite already being a live event) and the new `in_progress_snapshot` event; rewrote the multi-consumer fan-out paragraph to mention the kind-specific replay tail (state_change + optional in_progress_snapshot) so the "no catch-up needed" claim is no longer misleading. - architecture.md: bumped the SessionUI Protocol stub to 16 methods (added `on_turn_start` / `on_turn_committed`) and pointed at the in_progress_snapshot section in the API reference. - sdk.md: added rows for `state_change`, `in_progress_snapshot`, and `approval_resolved` (preexisting gap) to the per-workstream event table. - coordinator-api-tour.md: added an `in_progress_snapshot` row to the event table and rewrote the reconnection-contract paragraph to cover mid-stream content/reasoning restoration. - diagrams/04-conversation-turn.puml: added `on_turn_start()` before the thinking-start emit and `on_turn_committed()` immediately after `messages.append(assistant_msg)`, with notes explaining the inflight- buffer reset semantics. PNG regenerated. |
||
|
|
eb2a119da9 |
refactor(mcp): remove periodic refresh, add manual refresh/reconnect controls
Deletes the _periodic_refresh task and its supporting state
(_refresh_task, _refresh_failures, _refresh_backoff_until,
_REFRESH_BACKOFF_BASE/MAX, _DEFAULT_REFRESH_INTERVAL, refresh_interval
kwarg) from MCPClientManager. Push notifications and operator-driven
manual refresh now cover all catalog-update needs; the long-running
4-hour timer was dead complexity that obscured the per-user pool
work to come.
Catalog freshness on auto-reconnect is preserved by scheduling an
unblocking _refresh_server task on the mcp-loop after _connect_one
succeeds; the calling thread returns immediately so half-open
recovery latency does not double. Adds MCPClientManager.reconnect_sync
(clears the circuit, closes any existing session, calls _connect_one,
clears stale catalog on failure).
Wires a new pair of operator endpoints —
POST /v1/api/admin/mcp-servers/{name}/refresh and
/v1/api/admin/mcp-servers/{name}/reconnect — that fan out to all
nodes through the existing _internal route family, with per-row
"Refresh" and "Reconnect" buttons in the MCP Servers admin tab.
The new node-internal paths /api/_internal/mcp-{refresh,reconnect}/
are gated to the approve scope to prevent direct unprivileged
reconnects bypassing the console's admin.mcp gate. Internal
endpoints return generic error messages and a filtered status
payload (no command/url) to keep transport details admin-gated.
Drops the [mcp] refresh_interval setting, the
--mcp-refresh-interval CLI flag, and the matching config-mapping
entry; updates docs/architecture.md, docs/tools.md,
docs/settings.md, and the three PlantUML diagrams that referenced
the periodic loop.
Tradeoffs (intentional):
- Idle nodes will not auto-rejoin a recovered MCP server until
traffic arrives or an operator clicks Reconnect. The previous
background reconnection loop is gone by design — push
notifications + operator controls replace it.
- Console fan-out blocks on the slowest node (existing pattern);
not changed here.
This is Phase 1 of the OAuth-MCP series — feature subtraction
ahead of per-user state.
|
||
|
|
d6e615d324 |
fix: apply /review feedback on legacy URL cleanup
Reviewer caught real misses on the consumer-swap claim:
- TypeScript SDK still defined and re-exported `CloseWorkstreamRequest`
(types.ts + index.ts) — drop both. Now matches the Python-side
removal.
- Four `tests/test_auth.py` cases (`test_write_full_token_ok`,
`test_approve_full_token_ok`, `test_bearer_takes_precedence_over_cookie`,
`test_cookie_full_on_write_ok`) were tautological after the legacy
URL removal: they posted to `/api/send` / `/api/approve` and asserted
`allowed is True`, but those paths now classify as `read` so a read
token would also pass — they no longer tested the write/approve
scope enforcement. Swap to path-keyed URLs to restore the original
intent.
- `is_public_path("/api/send")` test renamed + retargeted to a
path-keyed URL.
Doc-table drift the previous commit missed:
- `docs/security.md` path-to-scope mapping rewritten for the
path-keyed verb family (write set, DELETE-on-/send dequeue,
per-ws_id approve).
- `docs/architecture.md` scope-model row text swap from `/api/send`
/ `/api/approve` to the path-keyed equivalents.
- `docs/diagrams/01-system-context.puml` channel→server edge label
swap.
- `docs/diagrams/15-auth-architecture.puml` scope class swap.
Cosmetic comment-only stragglers:
- `tests/test_session_worker.py` module docstring URL update.
- `tests/test_ratelimit.py` ~11 `/api/send` fixture-key strings
retargeted to `/api/workstreams/abc/send` so the URL fixtures
reflect the post-1.5 surface (rate limiter is path-agnostic; the
swap is purely cosmetic).
4557 tests still passing under -m "not live"; ruff + mypy clean.
|
||
|
|
ad0e7ce6eb |
docs: mark 1.5.0 legacy URL surface removal
CHANGELOG [Unreleased] / Removed (BREAKING — 1.5.0) block calling out the legacy URL family removal with the swap table. Doc passes on api-reference.md (per-endpoint sections rewritten with path parameters and slimmer body shapes), architecture.md (handler-list diagram and console-proxy URL example), console.md (URL-rewriting JS shim docstring + SSE proxy example), and the two PlantUML diagrams (11-console-data-flow, 16-channel-architecture). Also picks up two test-side stragglers from step 5 that referenced the legacy adapters in a docstring + a stale /v1/api/events SSE test: turn into path-keyed equivalents. OpenAPI JSON dump regenerated to reflect the catalog edits from step 3. After this commit: - 4557 tests passing under -m "not live" - ruff + mypy clean on turnstone/ tests/ sdk/ - grep for "/v1/api/send", "/v1/api/approve", "/v1/api/cancel", "/v1/api/workstreams/close" returns zero hits across turnstone/ sdk/ docs/ tests/ (excluding CHANGELOG.md, which intentionally documents the old shape). - grep for make_legacy_body_keyed_adapter, make_legacy_query_keyed_adapter, _make_method_dispatch, close_legacy returns zero hits. |
||
|
|
cab57f244d |
refactor(channels): backfill review of Slack/Discord adapters (#382)
* refactor(channels): backfill review of Slack/Discord adapters Retrospective multi-stage review of the Slack (PR #355) and Discord channel adapters — they shipped before the review pipeline existed, so this pass goes back and fixes everything the pipeline would have caught plus a follow-up round of ultrareview findings. ## Security (8 fixes) - Adapter-side owner checks on all interactive flows: Discord ApprovalView / PlanReviewView encode the owner Discord user ID in the embed footer (`{ws_id}|{corr_id}|{owner_id}`) and reject non-owner clicks; Slack plan-approve / request-changes / feedback-modal gain owner tracking in `_pending_plan_review_ts` and a shared `_ensure_plan_review_owner` gate. These closed the two critical authz gaps where the gateway's service-scoped JWT bypassed server-side ownership checks. - Discord thread-message gate: only the registered invoker can drive the workstream (prevents a linked user posting in another user's public thread from injecting into their assistant). Invoker recorded explicitly so `/ask` follow-ups survive the `channel.create_thread` bot-as-owner quirk. - Slack /link flow + per-user identity gate: unlinked Slack users see an ephemeral `/turnstone link <token>` prompt on every message instead of silently creating workstreams under the shared gateway identity. Rate-limited (5/hour) to block online token enumeration. - Gateway `/v1/api/notify` requires `write` scope on the validated JWT; low-scope tokens get 403 + audit. - Thumbnail URL validator DNS-resolves the hostname before fetch and rejects any resolved IP that's loopback / link-local / multicast / reserved, plus an explicit deny-list for IPv6 cloud metadata (`fd00:ec2::/32` — AWS Nitro IMDS + ECS task metadata) that would otherwise slip past the `is_private` allowance. - Per-user rate limit (10 msgs / 60s) + 8 KiB inbound size cap on Slack DMs / channels / notification-reply threads so one user can't exhaust the shared LLM budget. - Discord /link rate limit (5/hour) for token-enumeration defense. ## Bug fixes (9 correctness issues) - Slack DM routing: each top-level DM no longer spawns a fresh workstream (was using per-message `ts` as the route key). - Multi-chunk Slack responses thread correctly under the first chunk's ts instead of fragmenting as independent top-level messages. - Finalize the outgoing StreamingMessage before swapping channel / thread_ts mid-stream, so buffered tokens still land on the old thread. - Redundant `chat_update` on approve/deny eliminated by popping `_pending_approval[ws_id]` after local resolution. - Notification reply tracking on Discord only registers for DMs (guild-channel targets were storing channel IDs where user IDs were expected, so legitimate replies were always rejected). - `get_channel_default_alias` rolls `_channel_default_ts` back on `list_models()` failure so the next caller retries instead of serving an empty alias for the full TTL. - Slack `subscribe_ws` purges dead SSE tasks before the membership short-circuit (previously an unhandled exception left the ws_id in `_subscribed_ws` forever, silently no-opping subsequent subscribes). - ChannelRouter `_create_locks` is now an LRU-bounded OrderedDict that evicts only unheld locks (original dict grew unbounded; naive LRU could evict a held lock and let a second caller race through the critical section, creating duplicate workstreams). - Slack `_parse_ts` pads the fractional field to 6 digits so `"1.2"` and `"1.000002"` stop colliding as `(1, 2)` in the latest-session tiebreaker. ## Performance (6 fixes) - StreamingMessage keeps a rolling truncated display string capped at `max_length` so per-flush cost is O(max_length) instead of O(total_streamed_chars) — long streaming responses no longer do quadratic work every edit interval. - `StreamingMessage.finalize()` caches the joined content so the Discord stream-end DM-forward path doesn't re-join a multi-MB buffer twice. - `PendingApproval` stores the Block Kit payload posted to Slack; `IntentVerdictEvent` appends the verdict in-place and `chat_update`s, skipping an extra `conversations_history` round-trip. - ChannelRouter `lookup_ws_id()` TTL-caches the channel → ws_id resolution (30s TTL, 4096-entry LRU); hot inbound paths skip storage on every message. - Service-discovery startup retry uses exponential backoff (1s → 8s cap) with a 30s deadline instead of 30 × 1s fixed sleep. - `_archive_session` now calls `router.close_workstream` so the `_node_urls` cache entry is dropped (was leaking one entry per archived session). ## Quality / refactors (19 improvements) - `cli.main()` extracted from a 365-line function into focused helpers; imports carefully kept lazy where test patches target source-module paths. - `_run_gateway` finally block now awaits `adapter.stop()` on every adapter so SSE tasks, httpx clients, and the Slack socket handler close cleanly on shutdown. - Shared SSE reconnect loop extracted to `turnstone/channels/_sse.py` (`run_sse_stream` with `on_event` + `on_stale` callbacks); both adapters' `_sse_listener` methods just wire up callbacks. The "404 stops reconnect" invariant is enforced inside the helper so a broken `on_stale` can't livelock. - `_on_ws_event` god-dispatchers split into per-event `_handle_*` methods with a thin isinstance dispatcher at the top. - Slack `_on_approve` / `_on_deny` collapsed into a single `_resolve_approval(*, approved: bool)`. - `ApproveRequestEvent` policy evaluation hoisted into `ChannelRouter.evaluate_tool_policies` returning a `PolicyVerdict`; adapters switch on the verdict kind. - `ChannelAdapter` protocol trimmed to the four methods adapters actually implement; unused `ChannelEvent` dataclass removed. - Shared constants lifted to `turnstone/channels/_config.py`. - `_cleanup_stale_route` and `unsubscribe_ws` share a `_clear_ws_state` helper. - `StreamingMessage` private attrs promoted to `message` / `message_ts` / `accumulated_text` properties so callers don't reach past the `_`-prefix. - Various cleanups: dead var, noqa'd lambdas, renamed `_policy_handled` → `policy_handled`, inlined single-use helpers, added module docstrings, documented `SlackRoute.parse` edge cases. - `chunk_message` plain-text fast path (no backticks → skip fence bookkeeping). ## Test coverage Added 45 tests (178 → 223): - `tests/test_channel_sse.py` (new) — SSE reconnect / backoff / 404-stale-route / on-stale-exception / invalid-JSON-skip / on-event-exception-doesn't-kill-stream / per-connection token refresh / ConnectError retry. - ApprovalView + PlanReviewView owner-check regression tests (owner allowed, non-owner rejected, legacy 2-pipe footer fails closed, modal path rejected for non-owner, `/ask` bot-as-thread-owner follow-up allowed). - Slack `_recover_routes` latest-ts-wins, `_archive_session` drops route + closes workstream. - SSRF tests: DNS rebinding rejected, IPv4 link-local metadata rejected, IPv6 ULA metadata (fd00:ec2::254 / fd00:ec2::23) rejected. - Slack link prefix match (natural-language prompts don't hijack), link rate-limit ceiling. - SlackRoute round-trip across all three shapes + lax-parse behaviour. Lint (ruff) + mypy clean; 210 channel-focused tests pass. * chore(channels): address PR #382 review-bot feedback Three line-level findings from github-code-quality on the backfill review PR. Copilot had no line-level comments. - _sse.py:132 — the `except httpx.HTTPStatusError: pass` branch was flagged as an empty except. The original status was already logged at WARNING inside the try block (we re-raise ourselves after logging), so the handler has real intent. Added a debug log of the exception text + a comment explaining the control flow, so the empty-except lint stops firing and the next reader sees why we fall through to backoff. - discord/bot.py:430, cli.py:354, slack/bot.py:1127 — `await task` inside `contextlib.suppress` was flagged as "statement has no effect". It's a false positive (await is an effect) and the alternative try/except/pass triggers ruff SIM105. Kept the contextlib.suppress pattern and added an explanatory comment above each call so the intent (await CancelledError propagation before state cleanup) is obvious; will reply on the PR thread noting the false positive. No behavior change. Lint + mypy clean; 210 channel tests pass. |
||
|
|
37ed6bbf5b |
feat(core): WorkstreamKind enum + list_workstreams user_id filter (#374)
Foundation PR for the multi-stage-review follow-up. Introduces a single source of truth for workstream kind values and pushes tenant scoping into the storage protocol so list callers can't forget to filter client-side. - WorkstreamKind(StrEnum) replaces bare "interactive" / "coordinator" literals across 17 production modules. Strict mypy narrows every internal call site; raw strings still work at wide boundaries (HTTP body, DB row) via WorkstreamKind(raw) parse at the edge. - StorageBackend.list_workstreams(..., user_id=None) adds a SQL-level WHERE user_id = :user_id gate on both sqlite and postgres impls. Memory wrapper forwards the new filters. - register_workstream now validates kind at the storage edge so SDK / restore / internal callers can't silently corrupt the NOT NULL column with empty / mis-cased / unknown values. - WebUI.__init__ normalizes empty-string parent_ws_id to None, matching the storage-edge and WorkstreamManager invariants. - POST /v1/api/workstreams/new parses body["kind"] through the enum and returns 400 on unknown kinds instead of silent coercion. Absorbs bug-1, bug-2, bug-4/q-6, q-1, q-8, and partial q-2 (wrapper signature forwards the new filters; full deletion of the unused wrapper stays in the cleanup PR). |
||
|
|
a917bf2690 |
docs: apply Copilot review feedback on PR #367
All eight suggestions verified against source before applying: - docs/settings.md — ConfigStore key names are `model.plan_alias` / `model.task_alias` (not `plan_model` / `task_model`); updated in both the overview list and the plan/task overrides table. - docs/security.md — `src` claim values now reflect what actually gets minted: `password`, `database` (from API-token exchange), `oidc`, plus service origins `console`, `cli`, `channel`. - docs/sdk.md — `upload_attachment(ws_id, filename, data, *, mime_type=...)` matches the real SDK signature; `bytes`-returning helper is `get_attachment_content` (not `download_attachment`); code example reordered so it doesn't collide on `filename=` kwarg. - docs/architecture.md — "prior `plan` tool call" → "prior `plan_agent` tool call" so wording stays consistent with the renamed tool. - docs/tools.md — `plan_agent` `primary_key` is `goal`, not `prompt`, in both the primary-key table and the summary table (matches the JSON schema in turnstone/tools/plan_agent.json). |
||
|
|
471d1a3311 |
docs: audit documentation for 1.4 / 1.5 state
Systematic pass over every doc under docs/, the root-level README /
QUICKSTART / CONTRIBUTING, and the PlantUML diagrams. Memory and docs
had drifted against the code since 1.2 — this catches them up to the
1.4.0 release and the 1.5.0a1 experimental line.
User-facing fixes
- README: fix broken docs/mcp.md link (→ mcp-registry.md); channel
gateway entry reflects shipped Discord + Slack adapters instead of
"Slack/Teams planned"; diagrams table mentions both.
- QUICKSTART: docs/*.md relative links were wrong from the repo root;
wizard version bumped from 0.5.4.
- CONTRIBUTING: add dev extra plus the ruff / mypy / pytest commands
we actually expect before push.
Reference docs
- architecture.md: 19 tool schemas (was 15), 18 admin tabs (was 14),
turnstone-bootstrap added to entry-points table, OpenAI provider
file split (chat/responses/common) documented, 38 SDK event
dataclasses (was 27 and referenced deleted mq/protocol.py), Slack
adapter + multi-adapter gateway, plan_agent/task_agent naming,
governance admin-panel rewrite.
- api-reference.md: full attachment endpoints (POST/GET/content/
DELETE on /v1/api/workstreams/{ws_id}/attachments) plus the
multipart mode on POST /v1/api/workstreams/new.
- channels.md: Slack Setup section (Socket Mode app creation, OAuth
scopes, tokens), Slack CLI/env reference in config table, combined-
adapter architecture diagram.
- console.md: 18-tab listing (was 13) with Channels/Models/Nodes/TLS
descriptions and ConfigStore live-edit note.
- docker.md: Slack env vars block; image entry-point list now
includes turnstone / turnstone-bootstrap.
- sdk.md: attachments methods on the server client, attachments
example (upload-then-send and at-creation), event count fixed.
- releasing.md: four-track table (stable/1.0, 1.3, 1.4 + main 1.5);
promotion workflow uses 1.5 / 1.6 numbering.
- settings.md: plan_model / task_model / plan_effort / task_effort
overrides section.
- governance.md: skill naming (/skill, `skill` field — not /template),
Prompts/Judge tabs called out.
- security.md: two-token-types wording; src claim values match the
AuthResult source strings actually emitted.
- mcp-registry.md: SDK package name is @turnstone/sdk.
- tools.md: plan / task renamed to plan_agent / task_agent in the
section headings and summary table; primary-key table matched.
- design/consistent-hash-ring.md: dead direct-http-transport.md
pointer redirected to architecture.md.
Diagrams
- 02-package-structure: drop phantom chat.py entry point, add admin
and bootstrap, add slack/bot.py, rename channels/gateway.py →
channels/cli.py.
- 16-channel-architecture: Slack is no longer "(future)", add a
SlackBot class and the slack-bolt Socket Mode edges; wire the new
bot into ChannelService. PNGs regenerated from both puml sources.
|
||
|
|
934cb075d6 |
feat: per-model sampling parameters (temperature, max_tokens, reasoni… (#350)
* feat: per-model sampling parameters (temperature, max_tokens, reasoning_effort) Model sampling parameters were global-only settings applied uniformly to all models. Different models have fundamentally different requirements (o-series needs no temperature, Anthropic needs temp=1.0 with thinking, local models may need different max_tokens). This adds per-model overrides with global fallback so each model definition can specify its own defaults. Migration 036 adds nullable temperature, max_tokens, reasoning_effort columns to model_definitions. NULL inherits the global default from ConfigStore. The session factory and /model switch command both resolve per-model override → global fallback consistently. The admin UI model create/edit modal now has dedicated form fields for these parameters with client-side validation, a visual section divider, and per-model override hints in the model table rows. Removes vestigial model.name and model.context_window global settings (now handled per-model by the model registry) with startup warnings for existing config.toml users. * fix: defensive parsing for config.toml per-model sampling params Wrap temperature/max_tokens conversions in try/except with range validation. Invalid values log a warning and fall back to None (inherit global default) instead of aborting registry load. |
||
|
|
a3140da3a5 |
docs: update documentation for PRs #312-#316 (#324)
- README: add Google Gemini to multi-provider feature list and requirements - architecture.md: add GoogleProvider, update supported provider values, file listing, config example - judge.md: document cancel_on_approval, fresh-client lifecycle, fallback delivery, Google compatibility - settings.md: add judge.cancel_on_approval, new interface.* section (close_tab_action, theme), update total count - api-reference.md: document 6 new workstream/settings endpoints, add judge_model to workstreams/new - console.md: add judge model to modal fields, add keyboard shortcuts - console_schemas.py: add judge_model field to ConsoleCreateWsRequest - server_spec.py: add 6 new EndpointSpec entries - diagrams: add GoogleProvider to package structure and class diagram |
||
|
|
da5eae5352 |
chore(deps): update dependency katex to v0.16.45 (#309)
* chore(deps): update dependency katex to v0.16.45 * chore: download vendored JS files --------- Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |