Fifth review round — four small correctness edges, none in the retry semantics: - The server-buffer discard now runs only AFTER the backoff survives a Stop: a cancel during the window persists the promoted partial to history, and the idle-state payload (drained from the turn buffer) must carry the same text — discarding first rendered the cancelled turn empty on the dashboard while the transcript had it. Pinned with a real-buffer test; the spinner and fresh segment watermark follow the truncate so a later discard cannot resurrect the dead segment. - stream.retry's dead_content_chars reports THIS death's flushed text only — the Stop-preservation carry retains the previous attempt's partial by design, and logging its length re-attributed the same discarded spend to consecutive retry lines. - The changelog entry for the post-finish-blip rename no longer claims the usage_captured field was dropped; it is emitted and pinned. - The retry suite's module docstring states the shipped finalize contract (stream_end + backoff-gated stream_discarded, never turn_committed) instead of the superseded pair. - RecordingUI is hoisted into tests/_session_helpers next to NullUI — this branch already paid the per-file-fake tax once when a protocol method grew — and a stale deferral sentence is dropped from the fatal-formatter comment.
157 KiB
Changelog
All notable changes to turnstone are documented here.
The format is based on Keep a Changelog,
and this project adheres to PEP 440 for
version numbers (X.Y.Z, with X.Y.ZaN / bN / rcN for pre-releases).
Two active release tracks are maintained — the current stable and the experimental line:
stable/1.7— patch-only (v1.7.x)main— experimental (next major)
Earlier stable lines (stable/1.6, stable/1.5) are frozen.
[Unreleased]
Added
-
Per-model Entra gateway authentication. Model definitions can bind either a caller-delegated OBO token (
entra_obo) or a shared app-identity token (entra_app) through the provider SDK credential surface. Mints reuse the encrypted cluster token cache, refresh-rotation CAS, and advisory locking; add a host-local memo, failure cooldown, long-lived mint HTTP client, audience allow-list/permission boundary, identity-unlink purge, and optionalmodel.auth_fail_closedrefusal policy. Delegated identity now propagates through judge, output-guard, and principal-scoped perception lanes, and unattended watch restoration reacquires the persisted workstream owner. Ownerless OBO calls and dynamic aliases without a real static fallback always fail closed; grant modes are never silently switched. Static authentication remains the default. -
Compaction is visible now: lifecycle events, a progress bar, and a persistent transcript card. Context compaction (manual
/compactand auto) emits a first-classcompactionSSE event (start/progress/end— see the API reference) instead of loose info lines. The web UI renders an in-transcript card with a real progress bar (determinatepart k of Nduring chunked summarization, indeterminate for single-call compactions) that settles into a result card — token delta plus the summary behind a fold — in both the interactive pane and the coordinator viewer. The result survives reloads: the persisted compaction marker now projects through/historyas arole="system",source="compaction"entry (resume/export/search unchanged), stamped with the end event's id so repaint and SSE replay can't double-render. The marker'smetaadditionally recordsbefore_tokens/after_tokens/trigger. Python and TypeScript SDKs gain a typedCompactionEvent. -
One provider transport: every model call now streams (#831). The per-adapter non-streaming entry (
create_completion) is retired; single-shot lanes — judges, titles, compaction, web-fetch extraction, perception, eval, optimizer — sample through the same streaming entry the chat loop uses and accumulate via one shared drain, so request shaping can no longer drift between the two consumption styles. Two operator-visible consequences: long single-shot generations (a thinking model composing a title, a slow local judge) no longer sit in a single blocking read that can hit client read-timeouts — the same reason the Anthropic adapter already streamed internally — and judge timeouts now abort the underlying HTTP read instead of abandoning a worker thread on a dead call. Because every call now streams, an alias pointed at a model or org that cannot stream (OpenAI's verified-org streaming entitlement, a gateway api-version predatingstream_options— e.g. older Azure OpenAI deployments) fails at request time where 1.7's non-streaming single-shot call succeeded; remediation is on the serving side (verify the org, bump the api-version/gateway) — there is deliberately no per-model non-streaming fallback left to configure. These lanes are also complete-or-error now: a stream that ends without any finish signal is treated as a generation that died mid-response and retried, instead of storing the partial text as a clean result (previously a half-generated compaction summary could silently replace real history). Caveats: these lanes now carry the samestream_options: {include_usage: true}the chat loop always sent — OpenAI-compatible servers old enough to ignore it stop producing usage rows on these lanes, and servers strict enough to reject unknown fields (pre-2024 llama.cpp/proxy builds) will 400 — such a server already couldn't serve turnstone's chat loop, but a judge/utility alias pointed at one worked on 1.7 and needs to move to a current server. Transient mid-stream deaths (connection drop, proxy hiccup) are re-issued in place up to twice with exponential backoff — the retry the SDK's request loop used to provide these lanes invisibly. Each lane accepts its own terminal marker (Anthropicmessage_stop, Responses terminal events); a lax server/gateway that never sends any terminal signal needs{"finish_reason_optional": true}in the model definition's capabilities JSON, which restores 1.7's tolerance (clean end-of-stream after output = completion) for that model on every lane — without it such streams fail as died-mid-generation, because SSE gives no way to tell the two apart and the default favors catching truncation. The unreadsupports_streamingcapability flag (and its admin tile) is gone; the o-series models it described are dropped from the capability table entirely (see Removed). -
One turn interface for every model call:
core/model_turn.py(#827). Judges (intent + output guard), perception, title generation, compaction, web-fetch extraction, the eval harness, the optimizer's meta lanes, and task agents all advance a trajectory through the same plant-call primitive the agent seam pioneered — Turn IR in, one shared lowering (argument sanitize → minted-id restore → vLLM reasoning attach), one shared re-ingest (blank-id repair → native-lane finalize). The judges' hand-built OpenAI-dict path is gone, and with it the Gemini judge's tool-blindness: evidence tools now work on Google models because the native lane round-tripsthought_signature(with pairwise repair for blank-id compat responses). Provider adapters still take lowered wire dicts — the transport collapse and main-loop migration are tracked as #831 / #832. -
task_agent keeps its model's reasoning across its own tool loop — on every provider lane. A task agent's replayed turns now carry the provider-native reasoning lane the model produced — Anthropic thinking blocks with their signatures (commercial or an anthropic-compatible server), OpenAI Responses reasoning items, Gemini
thought_signaturefidelity blocks, and the reasoning text a vLLM--reasoning-parser/ llama.cppreasoning_formatsurfaces on the Chat Completions lane — instead of each turn being rebuilt from text + tool calls with the reasoning dropped. On a thinking model this restores reasoning continuity across the agent's own multi-turn tool use. On the wire the agent's session-minted sub-tool ids are mapped back to the provider's own ids (restore_provider_tool_ids), so the native block — replayed verbatim, its signature never touched — thetool_callsmirror, and each tool result always agree; internally the minted ids still key the live card, recall, and the cancel ledger unchanged. Replay honors the same per-modelreplay_reasoning_to_modelflag the main loop uses on every lane: the vLLM Chat-Completions field replay keeps its server-type gate, and llama.cpp stays capture-only, matching main-loop behavior. The native lane is finalized by the same shared builder as the main loop's, so the two harnesses cannot drift. -
Background shells:
bashgainsrun_in_background, plusbash_output/kill_shell. Settingrun_in_background=truestarts the command as a detached shell and returns immediately with abash_Nhandle — "start a dev server, use it in a later call" is back as an explicit opt-in (the shape follows the convention the major coding agents converged on).bash_outputreturns only output produced since the previous read (optionally filtered by a regex) plus status and exit code;kill_shellterminates the shell's whole process group. Output is buffered per shell with a drop-oldest cap, so a chatty server can't grow memory unbounded. When a background shell exits, a system notice lands at the next seam (waking an idle workstream if needed). Shells survive a generation cancel, die with the workstream, and never outlive a task_agent that started them; anything a background shell itself backgrounds is still reaped when that shell exits — the no-leak guarantee below is unchanged.
Changed
-
Log event rename:
drain_stream.post_finish_blipis nowstream.post_finish_blip; itsusage_capturedfield is retained. The single-shot drain normalizes mid-body transport deaths through the sametransport_guardedwrapper the interactive loop uses, so its post-finish-blip tolerance logs under the wrapper's event name. Update any external log filters pinned to the old name; the drained result's possibleusage=Noneon a post-finish blip is unchanged and documented ondrain_stream. -
Breaking (1.8): compaction feedback moved from
infoevents to the typedcompactionSSE event. Pre-1.8 SSE/SDK clients that ignore unknown event types no longer see compaction lines (they are deliberately not dual-emitted — dual emission would double-render on every current client). Consume thecompactionlifecycle event (see the API reference and theCompactionEventSDK type); embedders drivingChatSessionthrough a duck-typedSessionUIare unaffected (the classicon_infolines are restored for them — see Fixed). -
Sampling knobs (temperature, reasoning effort) now ride one assignment scheme: per-model alias value → operator-stored global setting → the model definition's declared default (effort only) → field omitted. Turnstone previously manufactured values onto every unconfigured request — a hidden
temperature: 0.5and areasoning_effort: "medium"baked in at three layers — overriding serving-side defaults like a vLLM model'sgeneration_config. Unconfigured installs now send neither field and the inference engine's own defaults rule;model.temperatureis blank by default ("inherit each model's own default") andmodel.reasoning_effortdefaults to the empty "inherit" choice. The per-model → global resolution lives in one shared resolver used by the session factories, the/modelswitch, and everymodel_turnlane, so the same alias samples identically on every surface. CLI--temperature/--reasoning-effortlikewise default to inherit.Upgrade notes:
- The empty (
"") reasoning-effort choice changed meaning from "explicitly disable thinking" to "inherit the model/serving default". On local manual-thinking models (e.g. Qwen templates withenable_thinking), a stored""previously sentenable_thinking: false; it now sends nothing, so the template's own default (often thinking ON) applies. Usenoneto actually disable reasoning. - Workstreams saved by earlier versions carry the old defaults
(
temperature=0.5,reasoning_effort=medium) in their persisted config and keep that exact behavior on resume; they pick up the new inherit semantics the next time you change the model or a sampling knob in that workstream. New workstreams inherit from the start.
- The empty (
Removed
- O-series and pre-5.4 GPT-5 rows dropped from the OpenAI capability
table.
o1,o1-mini,o3,o3-mini,o3-pro,o4-mini,gpt-5,gpt-5-mini,gpt-5-nano,gpt-5-pro,gpt-5.1,gpt-5.1-codex-max,gpt-5.2,gpt-5.2-pro, andgpt-5.3no longer have built-in capability rows — OpenAI has retired these model ids from the API, so the rows described contracts no request can reach anymore. The table floor is nowgpt-5.4; the search-api and audio/STT/TTS rows are unchanged. An alias still pinning a retired id fails at OpenAI itself; any other unlisted commercial id resolves to the generic commercial defaults (temperature sent, no declared reasoning-effort vocabulary, 200K window) — declare the contract on the model definition's capabilities JSON if you run one, or move to a current model.
Fixed
-
A transport failure mid-generation no longer kills the interactive turn (#937). A wire death during body streaming (TLS record failure, connection reset —
httpx.ReadErrorand kin) surfaces after the request has already returned its stream handle, so neither the SDK's request retries nor the creation-time retry ladder ever saw it: the turn died with a bareReadError: …, the partial output was discarded, and nothing was logged. The interactive loop now normalizes mid-body transport deaths exactly like the single-shot lanes and re-issues the turn (bounded, cancel-aware, exponential backoff), finalizing the dead attempt across every UI surface first so retried text never double-renders (web transcript, CLI markdown fences, Slack/Discord streamed messages). Before re-creating the stream the session re-resolves its registry binding, so a concurrent model-registry reload that closed the old client cannot turn the retry into a misleading closed-client error. On exhaustion the surfaced error names the provider, endpoint, and model with a stream-death message instead of a bare exception string, and every fatal turn now leaves asession.fatal.recordedlog line (INFO for a user Ctrl-C, ERROR otherwise). -
A failed worker-thread spawn no longer wedges the workstream — at either spawn site — and never masquerades as success. If
Thread.start()itself raised (thread exhaustion, out-of-memory), the dispatcher had already claimed the worker slot but the flag's only clearer lived in the never-started thread — the workstream looked idle forever while every subsequent message queued behind a worker that didn't exist, until an operator force-cancel. The claim is now rolled back under the lock and the error propagates, so the workstream is dispatchable again as soon as resources recover. Affected every dispatch path (sends, wakes, retries, deferred-send drain, init). The same failure at the deferred-send drain's own spawn rolls back the just-accepted entry and answers the retryablequeue_full(previously a 500 landed after the entry was registered — an invisible, unretractable phantom that later dispatched as duplicate turns), and a/commandwhose worker never spawned now answers 503{"status": "error"}instead of the generic 200 ok that told SDK callers their/clearor/resumehad applied. -
Manual
/compactfrom the web UI: no phantom user turn, no frozen server, cancellable. A slash command typed into the web composer no longer renders as a user chat bubble (it echoes as a distinct command chip — commands aren't conversation turns and were never persisted as such)./compactitself now dispatches onto the workstream's worker slot instead of running inline on the server's event loop — previously a long compaction froze every SSE stream on the node for its whole duration, which is also why its own progress only ever arrived as one burst after the fact. The manual path carriessend()'s full generation discipline (compact_now()): a force-abandoned compaction goes stale instead of swapping history under a successor turn — and retires at its next checkpoint instead of running out its remaining summary calls, with its late lifecycle events fenced off (compaction_idon every event,supersededon end events — both in the SDKs) so they can't animate, tear down, re-title, or falsely narrate a successor's card or activity pill; a cancel aimed at it is consumed on exit (previously it bricked every/compactretry until the next message); a Stop click on an idle session can't pre-abort the next compaction; a Stop that lands in the completion tail — after the last cancel check, or during a retry backoff (which now aborts immediately instead of sleeping it out) — is honored rather than silently eaten; and Stop now aborts the in-flight summary HTTP call itself (the compaction lane registers its stream in the same abort seam the main loop uses), so cancelling a compaction is immediate instead of waiting out a model call. -
Sends during a command window are deferred, ordered, bounded, and honestly rendered — never silently truncated or lost. Messages sent while any slash command holds the worker slot are deferred: answered
{"status": "queued", "msg_id"}immediately and dispatched as ordinary full-fidelity sends (attachments and sender identity included) when the command finishes — never routed through the mid-turn interjection queue, whose semantics are turn-shaped: previously a send during a manual/compactwas silently truncated to 2,000 characters, a second participant in a shared workstream was locked out with a misleading "another participant's turn" 409 for the whole compaction, and a message queued across a/resume//newcould be answered into the post-swap workstream. Because the response is immediate, timeout-bounded callers — the coordinator'ssend_message, the console proxy, SDKs, anything behind a stock reverse proxy — can no longer lose a message to a multi-minute command window; the deferred send is retractable until dispatch via the sameDELETE .../sendused for queued interjections (node-local, in-memory — the API reference documents the at-most-once durability contract). Deferred responses carry"deferred": true; the pending list is the order authority (a fresh send — or a coordinator dispatch, or a queued-nudge wake — lines up behind acknowledged entries instead of overtaking them, with the two-term barrier defined once on the workstream so the wake gate also honors a claimed entry whose dispatch is mid-flight, and the gate re-arms at the drain's exit even when everything pending was retracted); acceptance is bounded (10 pending per workstream — the interjection queue's own backpressure contract; the 11th answers the retryablequeue_fullinstead of pinning attachment bytes without limit and then running one unattended turn per entry); a dispatch crash re-queues the entry instead of eating an acknowledged message, and a drain thread that fails to start rolls the acceptance back and answersqueue_fullrather than parking a phantom the client can neither see nor retract; each dispatch emits a pane-tiermessage_dispatchedevent (folded: truefor interjection fold-ins) so queued-bubble UI keeps its retract affordance exactly until the message truly leaves — including when the send was accepted by a pane that believed the workstream idle, which now renders a real queued chip instead of a sent-looking bubble, releases the composer (a deferred send has no running worker to wait on), and cleans up fully when the send is refused or the chip retracted instead of stranding the pane in Stop mode. Dismissing a queued bubble — interjection or deferred — is a server-confirmedDELETE, and retracting a deferred send that carried attachments tells the user they were discarded instead of silently expiring them. -
Slash commands hold the worker slot with a loud contract. A
/compactraced against an in-flight turn is refused with an explicit busy response. Every other slash command runs through the same worker slot too — mutual exclusion against sends, a running compaction, and each other, with a busy answer replacing the old silent interleave — while the endpoint still awaits quick commands' completion off-loop (without parking an executor thread per request); the post-command pane refreshes (clear_uiafter/clear//new//resume, the workstream-name sync) ride the worker itself, so a command that outlives the endpoint's 25s response backstop still refreshes every pane on completion (the backstop sits under the console proxy's 30s client timeout so the degradedrunninganswer can actually traverse a proxied pane, which now surfaces it instead of silence; the/commandresponse contract —ok/running, with busy refusals answering a loud HTTP 409 rather than a silent 200 — is now documented in the API reference and the OpenAPI spec). -
Compaction status stays truthful across every UI surface. Manual compaction success also refreshes the status line/context pill immediately (parity with auto-compaction), compaction failures keep feeding the typed
errorevent and the node error counter (while a CLI Ctrl-C reports as cancelled, not a failure), one Stop prints one notice (a cancelled auto-compaction no longer stacks "Compaction cancelled." on top of send's own "[Generation cancelled]"), the workstream activity pill shows "Compacting context…" for the whole summarize phase, restores cleanly afterwards, and can no longer be stranded by a force-stopped compaction (a new turn's generation claim breaks a stale latch). Every retry backoff on the session (stream retries, task agents, notify delivery, compaction) now aborts immediately on Stop via one shared cancel-aware helper instead of sleeping out its exponential delay. -
Compaction failures report exactly once, to the right owner. A compaction failure reports exactly once (auto-compaction errors defer to the turn's fatal handler instead of doubling the red row and the error metric), failed-end notice suppression is computed once by the emitter (a
noticebool on the end event — in the SDKs — replaces hand-synced client policy), and a manual/compactfailure no longer crashes the CLI REPL./compacton a workstream showing theerrorbadge restores the badge on exit instead of stampingidleover it (the compaction neither retried nor resolved the failed turn). A force-cancelled initial send that completes late still delivers its scheduled-run completion notification (the only completion signal unattended workstreams have); the other post-command pane refreshes and error notices remain owner-guarded, so a force-cancelled wedged command that unwedges late can't wipe panes or inject stray notices into a successor turn. -
Pre-1.8 embedder UIs keep their compaction lines. Embedders driving
ChatSessionwith a pre-1.8 duck-typedSessionUI(noon_compactionhook) get the classicon_infocompaction lines back — threshold notice,part k/N, retry waits, token delta + summary box — instead of silent history swaps. (See the breaking event-contract note under Changed for SSE/SDK clients.) -
Static MCP servers: a pushed catalog change no longer wedges the shared session (#839). The static-path
*/list_changedhandler awaited its catalog refresh inline in the SDK's receive loop, but the refresh's own request can only be answered by that (now parked) loop — the refresh never completed, and every user's in-flight calls on the shared per-node session stalled behind it, unbounded, until the health loop's ping timeout tore the transport down (which was also the only way the changed catalog ever landed). Push refreshes now run as spawned tasks — debounced, coalesced per (server, kind), bounded by the connect timeout, and serialized on the per-server connect lock — and the manual and post-reconnect refreshes publish under that same lock, so a slower publisher can no longer land a staler catalog over a fresher one. Every teardown path now also clears the notification debounce stamp, so a reconnected server's first push refreshes immediately. Push-refresh debouncing is now per (server, kind) on BOTH the static and per-user pool paths — a tools push no longer swallows a prompts push arriving in the same 5-second window. A change genuinely lost to the debounce window (a same-kind push landing after the prior refresh finished, which the server will never re-announce) is recovered by an automatic health-tick retry rather than staying invisible until an unrelated push or a reconnect. The resource-refresh fan-out on both paths no longer orphans its sibling list call when one of the pair fails fast — the real error surfaces immediately (not masked as a 30-second timeout) and the surviving sibling is cancelled and reaped, under a bounded grace, inside the scope. A push refresh that fails while the connection stays up is likewise retried on the next health-loop tick until one completes — previously a single transient blip left the shared catalog stale for every user on the node until an operator intervened. An operator/mcp refreshno longer parks behind a busy per-server connect lock (a slow reconnect attempt could eat the whole 30-second refresh budget and fail the pass for every healthy server behind it) — the busy server is skipped on both the connected and disconnected branches, reported distinctly as "skipped" rather than as a false "no changes", the skip arms the automatic retry, and a force-reconnect drops the session up front so queued push refreshes can't starve it. Static-path resource and prompt catalogs are now size-capped like the pool path's (and like static tools) at discovery and on every refresh, so a misbehaving server's push can't balloon the node's merged catalogs. Deleting or reconfiguring a server can no longer leave it half-removed: the config removal and all cleanup are serialized under the connect lock (a cancelled removal completes its cleanup rather than stranding a live session and published catalog with the config already gone), andreconcile_syncretries a removal that timed out instead of marking it done — previously a DB-driven delete of a busy server could be a silent, permanent no-op until process restart. A refresh outcome now threads consistently to every operator surface off one source of truth (the per-serverlast_refresh_outcome): a busy-skip and a genuine failure are each reported distinctly from a real "no changes" —/mcp refreshprints "skipped" or "failed" rather than a false "no changes", and the node-internal refresh endpoint returns202 skippedinstead of a misleading200 okfor a refresh that never ran. A single-kind push refresh no longer paints the whole server healthy: because the error/outcome state is server-scoped, a successful tools push while the prompts catalog is still broken (or vice versa) no longer clears the failure — only a full refresh pass declares "ok". -
OpenAI Responses streaming: truncated and refused responses no longer vanish. A response that hit
max_output_tokensterminates the stream withresponse.incomplete, which the stream consumer did not handle — the turn was mislabeledfinish_reason: stopand its final usage and collected output items were dropped. Refusal parts had no streaming handler at all, so a refusal rendered as empty content instead of the[Refused: …]text the non-streaming path produced. Both now match: truncation maps tolengthwith usage/items intact, refusals render in content. Applies to the chat loop and every drained single-shot lane (#831). -
task_agent: sub-tool ids no longer alias across a local model's reused ids. A local model that reissues per-response sequential tool-call ids (
call_0every turn) made two of a task agent's steps share one id — the live card collapsed both onto one DOM row while/historyrecall kept them apart, so the two views disagreed. Sub-tool ids are now minted{parent}::r{run}s{step}::{id}, unique within the session (across an agent's turns and across concurrent or sequential runs), and that one id keys the nesting registry, the live rows, recall, and the cancel ledger. On the wire the agent's self-built history carries the provider's own ids, restored from the mint map (see the reasoning-lane entry under Added), and malformed tool-call arguments are legalized the same way the main loop's wire prep does. -
bash tool: never hang on a backgrounded child. A command that left a long-lived process running (
server &, a daemon) could wedge the whole workstream forever — the tool read stdout/stderr to EOF, which never arrived because the child inherited the pipe, and the timeout watchdog bailed once the foregroundbashhad exited. The tool now waits on the tracked process (bounded by the tool timeout) and terminates its whole process group on return, so the call always completes. Undecodable output is preserved (errors="replace") instead of being dropped as a spurious error.- Behavior change: a process the command backgrounds no longer survives
the call — nothing persists across bash invocations. (First-class
"run this in the background" support landed separately — see
run_in_backgroundunder Added.)
- Behavior change: a process the command backgrounds no longer survives
the call — nothing persists across bash invocations. (First-class
"run this in the background" support landed separately — see
[1.7.3]
A small feature and maintenance patch for the 1.7 line. No schema migrations and no new configuration knobs.
Added
- OpenAI GPT-5.6 (Sol/Terra/Luna) support — the Responses provider
understands the GPT-5.6 family: the
reasoning.modecontrol, the newmaxeffort tier, andtext.verbosity, with golden wire payloads pinning the request shapes. Theopenaidependency floor moves to>=2.44.
Changed
- Engineer base prompt hardened with process discipline — the default
base prompt for non-coordinator sessions now works in phases scaled to the
size of the change, defaults to red-green for testable work, scopes to the
smallest sufficient diff, stops to report after repeated failed attempts
instead of thrashing, reports only observed results, and delegates
exploration to
task_agent. Persona prompts freeze into the workstream stamp at creation, so this reaches new workstreams only.
Fixed
- Unknown reasoning-mode warnings name the allowed modes — a model definition with an unrecognized reasoning mode now logs the valid options instead of leaving the operator to guess.
Documentation
- HYPOTHESIS.md / PRIMER.md — the control normal form is tightened and the factored Q_E reading is carried into the glossary; the plain-language PRIMER stays in sync.
[1.7.2]
A feature-bearing patch for the 1.7 line. Rather than hold this work for the
larger 1.8 churn, the fixes and the smaller features that had already
stabilised on main are rolled into the stable line now: a rich preview
pane, persona/project settings on scheduled tasks, and a batch of streaming,
rendering, and nudge-delivery hardening.
⚠️ Before upgrading: 1.7.2 adds Alembic migration
066, applied automatically on first start. It adds twoText NOT NULL DEFAULT ''columns (persona,project_id) to thescheduled_taskstable; existing rows migrate to the empty default, which is byte-identical to pre-066 dispatch behaviour. The change is additive and reversible, but — as always — back up your storage before upgrading (pg_dumpfor PostgreSQL; copy the database file for SQLite).
Added
- Rich preview pane +
open_previewtool — a workstream can now open a rendered preview (HTML, Markdown, and other kinds) in a pane beside the conversation via the newopen_previewtool. Guarded fetches stream under a byte budget whose ceiling tracks the widest per-kind cap, preview blob ids are salted, and a preflight probe handles legacy charsets and a remote-assets opt-in. Seedocs/tools.md. allow_private_networkopt-in forweb_fetch/open_preview— private-address fetch and preview targets stay blocked by default; an operator can opt a workstream in through the settings registry when a private endpoint is genuinely intended. (Distinct from the 1.7.1[oidc]flag of the same name, which governs identity-provider discovery.)- Persona + project settings on scheduled tasks (migration
066) — a scheduled task can now pin the persona and project of the workstream it dispatches, matching the levers a manually-created workstream already carries. Both default to empty (kind-default persona / no project), so existing schedules dispatch exactly as before.
Fixed
- Streaming fast-path overflow recovery — fast-stream tokens are now
batched and overflowed SSE listeners recover instead of stalling (and
connectSSEno longer opens into a hidden background tab). The same overflow-recovery companions were carried to the coordinator pane, so a coordinator watching many children recovers dropped listeners the same way the live-session view does. - Renderer containment — markdown sentinel-forgery and recursive-frame content loss are contained, and an indented fence close no longer drags its indent into the enclosed code content.
- Idle nudge / wake delivery — nudge and wake delivery is hardened across
session eviction, cancellation, and identity rebinds; the wake gate now
requires a real nudge queue, refused wakes are logged, and
initial_message_statusis typed as a closed enum on the wire. web_fetchextraction inherits model settings — the completion that extracts content from a fetched page now inherits the workstream's model settings instead of falling back to defaults.- UI panes — ephemeral panes close on split-dismiss instead of orphaning a tab, and an unsplit skips the redundant refresh after an ephemeral pane closes.
- Shared code-highlight CSS — renderer-output CSS is shared so the console and coordinator panes highlight code identically.
Security
Content-Dispositionfilenames made wire-safe — download filenames derived from user-controlled text are sanitised (latin-1- and control-char-safe, quoting-safe) before they reach theContent-Dispositionresponse header, including the fallback path.
Documentation
- HYPOTHESIS.md: daemons + the outer loop, plus a plain-language PRIMER —
the harness north-star document gains its daemon / outer-loop treatment and
a new top-level
PRIMER.md.
[1.7.1]
A maintenance and hardening patch for the 1.7 line. No schema migrations;
the credential-redaction work below is additive and needs no configuration
change. The one new operator-facing knob is the opt-in [oidc] allow_private_network flag (default off).
Security
- Credential redaction hardened across the tool-call surface — the
redactor that scrubs secrets from tool arguments and log previews was
reworked on both the backend and the browser to close several leak paths
and to fix false-positive and performance issues. Malformed tool-call
arguments are now legalised before they reach the wire; the tool-args log
preview scrubs credentials and control characters; and the coordinator's
tool-call cards gain a matching client-side redaction pass so the JS and
backend redactors stay at parity. Pattern coverage now includes
secret_access_key/aws_secret_access_keymulti-segment keys, baretoken=/key=forms (guarded by a negative lookbehind to avoid false positives), and SQLAlchemy+driver-qualified connection-string schemes matched case-insensitively. - OIDC SSRF guard:
[oidc] allow_private_networkopt-in — self-hosted identity providers on private networks can now be reached by settingallow_private_network = trueunder[oidc](default off; the MCP OAuth path stays strict). Rejections of discovered endpoints carry the opt-in hint so the misconfiguration is self-explanatory. Seedocs/oidc.md.
Added
- Persona discoverability + forgiving name resolution — personas are now discoverable by agents, and persona-name resolution tolerates case/whitespace variation; a not-found resolution reports the offending input verbatim instead of a bare error.
Fixed
- MCP transport lifecycles routed through per-entry owner tasks
(#787/#788) — static and pooled MCP transport lifecycles are now driven
by per-server / per-entry owner tasks, with a hardened disarm-sweep loop
guard and targeted exception handling in place of a broad
BaseExceptionarm, so a dying transport can no longer spin the CPU or strand delivery. - Client-construction failures surface as misconfiguration, not raw 500s — a model whose client cannot be constructed now reports a factory misconfiguration, and the raw exception text is kept out of the resulting 503 response.
- Postgres history search survives oversized rows — a conversation row exceeding Postgres' full-text limits no longer aborts history search.
- Agent-tool render is idempotent — tool rendering no longer deep-copies a tool definition until a description actually changes, so no-persona sessions share the tool constant (correctness plus a hot-path allocation win).
- Private-project workstream visibility scoped to members — workstreams in a private project are visible to project members only, not to every admin; coordinator tenancy checks now use request-scoped storage.
- Pane hotkeys work off macOS and match across surfaces — the pane keyboard shortcuts no longer collide with browser accelerators on non-macOS platforms and behave consistently across surfaces.
[1.7.0]
The headline of the 1.7 line is Personas — operator-authored control over how each workstream composes its system message and capability envelope. The rest of the release hardens the pieces a persona leans on: concurrent approvals, cross-provider reasoning-effort control, cooperative compaction, multi-user session safety, and MCP resilience for unattended work.
⚠️ Before upgrading: 1.7.0 adds Alembic migrations
062–065, applied automatically on first start (projects, personas, and two smaller schema tidy-ups). Migration063creates thepersonastable with its six seed personas and converts existingcreative_modeworkstreams to thewriterpersona in place. The changes are additive to your conversation data, but — as always — back up your storage before upgrading (pg_dumpfor PostgreSQL; copy the database file for SQLite).
Breaking changes at a glance (details in the sections below): the
/creative REPL toggle removed (replaced by the writer persona), the
turnstone-bootstrap entry point renamed to turnstone-doctor, and the
approval-status API/SDK field pending_approval_details changed from a
single object to a list (one entry per concurrent approval cycle).
Added
- Personas (#683) — a named, reusable bundle attached to a workstream
at creation, controlling system-message composition and the capability
envelope via exactly four levers: base-prompt override, tool visibility
set, MCP on/off, and memory on/off. The persona is resolved once and
snapshotted into
workstream_config; editing or archiving a persona never changes an existing workstream. Six seed personas ship with migration063(engineerandorchestratorare the per-kind defaults with no overrides, so zero-touch behavior is unchanged;scribe,researcher,writer, andexecutiveare curated envelopes). Selectable on every creation surface (web pickers, the create API/SDKs, coordinatorspawn_workstream/spawn_batch, andturnstone --persona <name>); authored in the console's new Governance → Personas tab (persona.{create,read,write}perms, archive-only lifecycle). Seedocs/personas.md. - Projects — governed resource containers (#724) — group workstreams
and their resources under a project (migration
062), with project-scoped memory, a per-project resources view, a project column on the saved list, and server-enforced private-project workstream visibility. - Task-agent sub-harness (#732) — a spawned task agent now runs on its own Turn-IR sub-harness with parent-tagged step events: its sub-tool steps nest inside an expandable card in the parent trajectory, its sub-trajectory is recallable, and each agent gets read isolation from its siblings.
- MCP static-server autonomous reconnect (#768) — statically configured MCP servers are now kept live by a health loop (capped-jittered backoff, ping-based liveness) instead of silently staying dead after the first transport drop.
- Attachments — capability-gated client-side fallback — when the active model can't natively handle an attachment, the client degrades gracefully (PDF → extracted text, audio → transcript) instead of failing the turn.
- Eval measurement / optimizer split (#763, #765) —
turnstone-evalis now a measure-only substrate with the prompt optimizer factored out, plus a new skill-adherence measurement mode. - Deployment examples — a vLLM + LiteLLM unified-memory inference
example showing a 3-model co-resident stack with an HF loader (#686,
#688), and an Altair +
vl-convert-pythonvisualization stack (#685). - Concurrent approvals and a long-session frontend overhaul (#754,
#755, #773, #775) — the live-session frontend was reworked for long
runs (the pipeline is wedge-proofed and its hot paths de-O(N)'d), and on
top of it a workstream can now hold more than one tool call awaiting
approval at a time. Each parallel batch gets its own approval cycle,
with one card per pending call in the interactive and coordinator UIs,
cycle-keyed tracking in Slack and Discord, and cycle-routed resolution
across the server/console/SDK APIs; sub-agent tool gates run the
intent-judge pipeline as their own generation. The send button no longer
sticks disabled after a batch resolves — orphaned approval cycles are
pruned and the app is the sole owner of the button state.
(BREAKING: the
pending_approval_detailsfield is now a list, oldest first.) - Reasoning-effort control on every provider lane (#771, #774) — the
session effort knob now reaches local backends too: it drives
chat_template_kwargson the anthropic-compatible and openai-compatible lanes and threads through to Gemini and xAI, alongside the commercial providers that handle effort natively. The console surfaces each model's effective effort ladder in plain words and adds an always-on thinking-mode option to the model form. Effort snapping is ordinal — it rounds up and caps at the model's ceiling rather than silently dropping.
Changed
- Skills are capability-context, not identity (#762) — a task agent's identity now comes from its persona; an applied skill's body is demoted to capability context and moved out of the identity system message. Skill-body substitution is unified across every invocation context so the same skill renders identically whether loaded interactively, by the model, or inside a sub-agent.
turnstone-doctorreplacesturnstone-bootstrap(#718) (BREAKING) — the setup/diagnostics entry point is renamed; update any scripts or service units that invoketurnstone-bootstrap.- Honest cancellation dispositions — cancelled or timed-out
side-effecting tools now report an
UNKNOWNdisposition rather than a flat failure, tool dispositions are typed (not just prose), and a coordinator cancel propagates down the sub-tree. - Multi-user shared-workstream context (#750) — in a shared workstream, send is gated to the acting participant while a turn is in flight (both the interactive and coordinator surfaces), cross-user mid-turn interjections are blocked, and shared-workstream state plus fork sender attribution are now durable.
- Cooperative compaction (#730) — the context budget is anchored to
the provider's true capacity, the summary call is chunked so it can't
overflow, and the active plan and the outstanding ask are carried across
compaction verbatim. The
recalltool is scoped to the compacted-away past. - Intent judge sees the full tool arguments (#760) — the judge's
argument projection is no longer narrowed, so it stops issuing confident
false denials on a partial view. The output-guard judge sources its real
context window, and
context_window = 0inconfig.tomlnow means auto-detect.
Fixed
- Compaction resume hardening (#731) — checkpoint markers are persisted so resume rehydration is bounded, context-overflow on resume is recovered across providers, and a recognized rate-limit is no longer misclassified as context overflow.
- MCP unattended-work resilience (#706, #742, #767) — dead-transport
handling is completed, consented OAuth (OBO) tokens are refreshed
proactively so autonomous runs don't strand on an expired grant, the
Entra ID on-behalf-of impersonation flow blockers are closed (migration
065adds the OIDCoid), and OAuth refresh failures are classified so a transient blip never revokes consent nor a dead grant strands the user. - Memory writes (#735) — save/update is a single atomic upsert, and writing a memory no longer recomposes the system prefix mid-session.
Removed
/creativeremoved (BREAKING) — subsumed by the Personas feature above: the REPL toggle (and its tab completion) is gone, and thewriterseed persona replaces it — start a session withturnstone --persona writeror pick Writer in the web pickers. Unlike the old fork, the writer persona composes the full system message, so session context and mandatory prompt policies now apply to prose-only sessions too. Thecreative_modekey inworkstream_configis no longer read or written. Migration063converts existing creative-mode workstreams to thewriterpersona automatically, so they resume as writing sessions rather than as legacy defaults.
Security
- High-risk skill activation is gated (#762) — a model-initiated load
of a
high- orcritical-risk skill is gated and fails closed when the backing storage is unavailable, so an untrusted turn can't silently pull in a dangerous capability. - Dependency security floors —
cryptographyandstarletteare pinned to security-fixed minimums. - CI publish hardening — the vendored-JS dispatch path refuses fork
PRs, and
workflow_runpublishing is gated to same-repo tag pushes, so a fork can't trigger a release build.
[1.6.0]
The first stable release of the 1.6 line — and the first under Apache 2.0.
⚠️ Before upgrading from 1.5.x: 1.6.0 changes the internal conversation storage schema (Alembic migration
060, applied automatically on first start). The migration converts existing workstreams and attachments in place — back up your storage before upgrading (pg_dumpfor PostgreSQL; copy the database file for SQLite). Background: discussion #631.
Breaking changes at a glance (details in the sections below):
web_search backend overhaul (Tavily/DuckDuckGo removed, topic →
category), the man / math / plan_agent built-in tools and the
plan-review protocol removed, and the body-keyed /v1/api/command
endpoint replaced by path-keyed workstream verbs.
License
- Relicensed to Apache 2.0 — from BUSL-1.1, effective with this
release (#546, contributor assent record in #548). Versions 1.5.x and
earlier remain under BUSL-1.1 as shipped, and the
stable/1.5branch keeps its original LICENSE. NewNOTICEandCONTRIBUTORS.mdfiles;THIRD-PARTY-NOTICESrefreshed to match the bundled library versions.
Added
- Mid-conversation system messages — advisories, watch results,
skill hints, and operator interjections are now first-class
role=systemturns in the trajectory instead of ad-hoc reminder envelopes. Models with native mid-conversation system support receive them verbatim; for everything else they fold into a nonce-fenced wrapper. The one-shot_remindersside-channel is gone. - Self-hosted SearxNG web search — the
web_searchbackend for local/vLLM models is now a bundled SearxNG service (in both compose stacks; internal network only). Configure viatools.searxng_url/tools.searxng_engines. Commercial providers keep their native server-side search; the model can target a corpus by passingcategory(general,news,it,science). Operators exposing the bundled SearxNG publicly: see the AGPL-3.0 §13 note in docs/docker.md. - Endpoint-backed reranking — a reranker is now a per-model
definition (Cohere/Jina-compatible wire: vLLM, TEI, llama.cpp, or a
commercial endpoint), disabled by default. When configured it scores
web_searchresults and the BM25 retrieval surfaces (deferred tools, skills, memory) behind atools.rerank_bm25toggle with a relevance floor; a calibration CLI (and calibrate-on-detect) tunes the floor per model. - Proactive memory relevance — injected memories are selected by BM25 + reranker against the recent user messages instead of recency alone, and first composition defers to the first user turn so fresh sessions select against a real query.
- Smart Approvals — opt-in (default off): high-confidence
approveverdicts from the intent judge auto-approve the tool call instead of waiting for a human, with a confidence threshold and verdict bookkeeping designed so a denied or reset judge never auto-fires. - Early-painted tool calls — committed tool calls render immediately
as pending cards (both UIs upgrade the card in place by
call_id) instead of waiting for the judge verdict, so big parallel batches no longer sit invisible during judging. - Voice I/O v1 — speech-to-text and text-to-speech as model roles speaking the OpenAI audio wire protocol (#618); the interactive composer grows a mic button.
- Rewind / retry / edit-first-message — full UX in both the interactive UI and the coordinator pane, backed by shared path-keyed verb handlers (#549).
- Workstream export — download a conversation as OpenAI-format messages JSON.
- Skills platform round —
SKILL.mdingestion learnswhen_to_use/model/effort/paths; prompt substitution supports$ARGUMENTS,$N,$<name>, and${CLAUDE_*}(#572); per-skilldisable-model-invocationanduser-invocableflags (#571);skill+list_skillsunify into one dual-kind tool; newmodel.skills.writepermission. - Coordinator hardening for small models — workstream references in
coordinator tool calls are validated with did-you-mean recovery, and
wait_for_workstreamfails fast with uniformnot_foundentries instead of hanging on a hallucinatedws_id. - Provider support — Claude Fable 5 and Claude Opus 4.8; xAI/Grok via the OpenAI Responses lane; vLLM reasoning-field replay completes the reasoning-persistence work (#537).
- Cluster-by-default deployment — the compose stack fronts
everything with Caddy and supports bare-metal node join; a one-line
curl | bashinstaller bootstraps a node; nodes with no configured models boot into a degraded state instead of crash-looping; channel gateways stand by when no adapter token is set. - MCP OAuth tokens encrypted at rest.
turnstone-adminreadsconfig.toml— same[database]section and precedence as the server (`CLI / config.toml > TURNSTONE_DB_* envdefaults
), includingpool_sizeand thessl*knobs it previously dropped; new--config PATH` flag.
Changed
- Conversation storage and the provider wire are rebuilt around a
canonical trajectory (migration
060— see the upgrade note). Internally a conversation is now a provider-neutralTurnsequence lowered to each provider's wire format at send time; provider-specific tool-call metadata rides an opaque producer-tagged lane (replayed verbatim to the producing provider, rebuilt for others); attachments become content-addressed, reference-counted rows resolved at the provider boundary; orphan tool-call repair happens once, at send time. Wire-visible behavior is unchanged for OpenAI-compatible providers; histories are preserved across the migration. - The console and web UI share one L-shell — a left glyph rail, a
tab bar, and a pane host now frame interactive chats, coordinator
sessions, dashboards, and the admin panel as tabs in a single window;
the standalone web UI adopts the same shell and the old split-pane
layout is retired. Coordinator and interactive conversations render
through shared
.conv-*card builders, the rail collapses to a glyph strip (remembered per browser), mobile gets an off-canvas drawer, and the frontend is now ES modules end to end. - Admin panel modals → the Service Hatch shelf — all ~35 admin modals are replaced by pane-scoped shelves plus a small dialog tier for confirmations. Schedules gain a cron builder with a next-3-runs preview endpoint, model capabilities render as an LED tile matrix, and the legacy modal machinery is deleted.
- SSE delivery is resumable end to end — per-workstream ring buffer
with
Last-Event-IDreplay (cap raised 2,000 → 50,000), fresh-connect and reconnect unified on one event-id cursor (in-flight tool batches included), persistedlast_errorreplays on connect, the console proxy forwardsLast-Event-ID, and panes close their connections onbeforeunloadto stop multi-pane refresh from exhausting the browser's per-host connection cap (#539). - Workstream verbs are path-keyed (BREAKING) —
rewind/retry/edit-first-messagelive at/v1/api/workstreams/{ws_id}/<verb>alongside the other session verbs; the body-keyed/v1/api/commandendpoint is removed (#549). /historyis projected server-side — both UIs consume the same REST-first wire shape instead of re-deriving it client-side.- Saved workstreams & coordinators: card grid → sortable table with model/skill/context columns, pagination, and a unified selector across both dashboards.
tools.web_search_backendaccepted values (BREAKING) — now""(auto),"searxng", or"mcp:server:tool". The old"tavily"and"ddg"values are gone; a config still set to either disables web search and logs a warning. Auto-detect resolves to SearxNG whensearxng_urlis set.web_searchtool:topic→category(BREAKING) — renamed LLM-facing parameter; values map to SearxNG categories. The Tavily-erafinancetopic is gone.- Core install includes what most deployments use —
anthropic,postgres,console, andtlsare core dependencies rather than extras. - NODES table → bottom-bar node picker in the console.
Fixed
- Cluster mTLS actually survives operations — certificate identity keys on the advertised host rather than the container ID, renewals are scoped per node, reloaded certs hot-swap into the live SSL context, and healthchecks/boot retries are mTLS-aware.
- Intent-verdict lifecycle — history replay ships risk-none verdict
rows (live/replay parity), late verdicts persist as
supersededfor the audit trail instead of vanishing, bulk verdict insert tolerates per-row conflicts, and cancel-on-approval honors its run-to-completion contract. - Usage accounting — dashboard totals were under-counting; auxiliary LLM spend (judge, rerank, memory) is now recorded.
- Concurrent first-boot migrations no longer deadlock on the advisory lock.
- Output renderer — single-
$inline math no longer false-positives in prose;strip_htmlpreserves block structure and drops a ReDoS risk. - Model registry orders versions numerically (no more
1.10 < 1.9selection).
Removed
- Tavily and DuckDuckGo
web_searchbackends (BREAKING) — replaced by the bundled SearxNG service. Removed:tools.tavily_api_key,$TAVILY_API_KEY,[api].tavily_key, and theddginstall extra. PointTURNSTONE_SEARXNG_URLat an existing instance or use the bundled one; no database migration required. man,math, andplan_agentbuilt-in tools (BREAKING) —man/mathduplicatedbash; planning is better expressed as atask_agentrunning a planning skill. Also removed: themathsandbox executor, the read-onlyAGENT_TOOLSsub-agent set, the plan-review protocol (/v1/api/plan,plan_review/plan_resolvedSSE events,on_plan_reviewhooks), and themodel.plan_*settings. Interactive built-in tool count: 19 → 16.stable/1.4track retired — the maintenance policy is now the current stable plus one prior (stable/1.6+stable/1.5as of this release). 1.4's final release wasv1.4.0; its tags and released artifacts remain available, under BUSL-1.1 as shipped.
Security
- Zero direct-HTML frontend — every
innerHTMLsink across the console and web UI is replaced with DOM construction orsetSafeHtml, inline handlers became delegated bindings, and CI lints pin the invariant (plusvar-free and const-reassign checks) across all swept bundles. - Output guard grows an LLM stage — merged with the heuristics as escalate-only (an LLM verdict can raise but never lower a heuristic positive), with annotated findings, a capability gate, and hardening against domain-camouflaged injection (#560, #573).
- One trust-fence primitive — operator and judge envelopes share a nonce-fenced wrapper (64-bit nonces, host-escaping); the output guard flags nonce forgery, and skill hints no longer echo model-controlled filter values into trusted text.
- RBAC — built-in role overrides get an editor, and several under-enforced permission gates are tightened (#585).
- Permissive
config.tomlwarns — a single startup warning when the resolved config file is group- or world-readable; operators usually want0600. - Dependency floors —
starlette>=1.0.1(PYSEC-2026-161 host-header path injection) andaiohttp>=3.14.0(security release).
[1.5.17]
Backports a clutch of coordinator-tool clarity fixes plus a watch-delivery
correctness fix from main to the stable/1.5 track, plus a previously-
latent intent-verdicts persistence bug exposed by the new heuristic-verdict
INSERT paths. No schema changes.
Fixed
intent_verdictsPK collisions on every llm_fallback delivery — async LLM-tier "llm_fallback" verdicts (turnstone/core/judge.py—_deliver_fallbacksand the in-loop fallback path) deliberately reuse the heuristic verdict'sverdict_idso the row gets "upgraded in place" fromtier="heuristic"→tier="llm_fallback"when the LLM judge times out, is cancelled, or returns no content. The consumer_persist_intent_verdictwas doing a plain INSERT, hitting theintent_verdicts_pkeyconstraint on every fallback delivery; Postgres logged the duplicate-key error, the application try/except swallowed it atlog.debug, and the row never actually got upgraded — the LLM judge's annotation ("(LLM judge did not return a verdict)") was lost. The collision rate exploded on this release because the new heuristic-INSERT paths in the auto-approve early-return branches ofapprove_tools(introduced below) leave no gap for the fallback to land cleanly into. Fix: newupsert_intent_verdictstorage method usingON CONFLICT (verdict_id) DO UPDATEthat updates onlytier,reasoning,judge_model— the three fields that genuinely change between heuristic and llm_fallback. Every other column (identity, carried-verbatim, anduser_decision) is excluded;user_decisionin particular would otherwise be clobbered back to"pending"when a fallback arrives after the operator has already resolved the approval. The bulk-INSERT path stays as plain INSERT — fresh UUIDs injudge.evaluatemake in-turn dups impossible; the inverse race (fallback wins before bulk lands) is reachable but unchanged in observable behavior by this fix, documented at the bulk site for a future hardening pass.- Coordinator LLM re-spawn loops on large fan-outs — the spawn-tool
return JSON used
ws_idas its key, which primed the model's recency bias to feed the spawn result straight back into anotherspawn_workstream(ws_id=...)call instead of progressing towait_for_workstream(ws_ids=[...]). On 10+ child fan-outs this cascaded into self-inflicted re-spawn loops. The LLM-facing tool result now emitschild_ws_id(the storage column / HTTP API contract is unchanged); the field name is already an existing project term so the rename aligns rather than introduces new vocabulary. Also handles the silent upstream-omits-ws_id success-shape edge that previously emitted{"child_ws_id": null}to the LLM — now surfaces a tool error so the model retries rather than chasing a null id. inspect_workstreamblowing the coordinator context budget — a coord doing a fan-out wave against tool-heavy children could land100 KB of raw output per inspect call, and the previous safety net (
_truncate_output's head+tail strategy) silently dropped middle messages — exactly the wrong shape for understanding a child's trajectory (the FIRST sets the brief, the LAST shows the conclusion, the middle is the connective tissue). Output now goes through a three-tier degradation ladder mirroring the search tool's_format_search_results:_tier="full"(every message verbatim) →_tier="compact"(per-message head/tail-snipped content + snippedtool_calls.arguments, falling through a(20,30)/(10,20)/(5,10)message-list trim ladder) →_tier="skeleton"(counts, role distribution, last-assistant preview). Budget 32 KiB matches the search tool's; the chosen tier is annotated on the response so the model can recall with a tightermessage_limitif signal was lost.- Auto-approved verdicts indistinguishable from pending review —
intent_verdictrows for auto-approved tool calls landed withuser_decision="", which read identically to "still waiting for the operator" in the audit trail and led to a real misdiagnosis incident. The column now carries an explicit vocabulary at insert:pending/approved/denied/timeout/policy/blanket/skill/always/auto_approve_tools. The auto-approve early-return branches inapprove_toolsnow persist heuristic verdicts stamped with their reason (previously dropped on the floor), and late LLM-tier verdicts that arrive for an already-auto-approved call_id are stamped via a TTL-pruned lookup map — so the audit row carries the auto-approve reason even when the LLM judge daemon completes after the synchronous approval cycle finished.resolve_approvalgains atimeoutkwarg writing"timeout"(the previous shape collapsed passive timeouts and active denials into the same column). list_skillsemptyallowed_toolsmisread as "no tool access" — the response previously emitted"allowed_tools": []for every skill that hadn't declared an auto-approve allowlist, which a coordinator model read as "this skill can't use any tools" (real misdiagnosis: a code-review child appeared to have been spawned with zero tool access). The field is now omitted entirely when empty — absence carries the unambiguous meaning "no tool is pre-approved for this skill", presence (non-empty list) keeps the standard Claude Code skill-spec shape. The tool description rewrite makes the auto-approve-allowlist semantics explicit so a future reader doesn't re-derive the gating misread.- Watch terminal-fires silently dropped on backpressure — delivery now routes terminal events through the same path as normal fires instead of being filtered out when the consumer was saturated.
Documentation
- Storage
LIKE_ESCAPEcontract — clarify that callers passing.like(escape=...)must use the same escape character that the storage helper assumes; previous wording let a reader pass a different escape and silently produce no matches.
[1.5.15]
Fixed
- Admin console blank-page on MCP server rows with consented users — a
Phase 9 (1.5.14) regression in
admin.jsused double-quote string delimiters on the bulk-revoke button HTML literal, but the literal embeds a"mid-attribute. JS closed the string early, turnedbulk-revoke (into bare tokens, and the resultingSyntaxErrorwiped out every global inadmin.js—showAdminand all other admin entry points became undefined, so the console UI was non-functional whenever the rendered MCP server list contained at least one row withconsented_users_count > 0. Switch the literal to single-quote delimiters to match the surrounding block.
[1.5.14]
Backports OAuth-MCP Phase 9 from main to the stable/1.5 track.
Added
-
OAuth-MCP Phase 9 — admin status, deferred-consent persistence, operator docs — completes the per-(user, server) OAuth-MCP build-out. The sync pool dispatchers now upsert into a new
mcp_pending_consenttable onmcp_consent_required/mcp_insufficient_scope, so a non-interactive run (scheduled / channel) that hits an unconsented server surfaces the deferred prompt to the user on their next dashboard load via the gear-icon badge — rows are cleared automatically by the OAuth callback handler on consent completion, or via new DELETE endpoints for manual dismiss. The MCP Servers admin row gains aconsented_users_countpill and a two-step-confirm bulk-revoke button forauth_type=oauth_userservers (upstream RFC 7009 revoke is intentionally not attempted in bulk to avoid N synchronous round-trips against the provider). Operator-facing docs land atdocs/mcp-oauth.mdanddocs/operations/mcp-oauth-headless.md.Introduces forward-only migrations
054_mcp_pending_consentand055_mcp_user_tokens_server_index.
[1.5.13]
This release introduces one forward-only schema migration:
053_services_notify_trigger — installs the services_notify PostgreSQL
trigger that backs the new LISTEN/NOTIFY dispatcher (no-op on SQLite, where
the dispatcher uses in-process fan-out).
Added
- Reactive node discovery via PG LISTEN/NOTIFY — the console gains a
NotifyDispatcherthat holds a dedicated session-mode PostgreSQLLISTENconnection (bypasses pgbouncer transaction pooling) and fans wake-ups out to per-channel handlers on a separate dispatch thread. The cluster collector subscribes to a newserviceschannel and reacts to node register / deregister within ~500 ms instead of waiting up to 60 s for the next discovery loop; the 60 s loop is retained as the backstop for crash-shaped loss (NOTIFY only fires on real writes). The storage layer also gains a uniformnotify/listenAPI with an SQLite synthetic-sweep fallback so consumer code is identical across backends.TURNSTONE_DB_LISTEN_URL(or[database] listen_urlinconfig.toml) points the dispatcher at a direct-to-Postgres URL; defaults to the main DB URL when unset. - Event-driven
wait_for_workstream— coord's block-wait tool no longer polls storage every 500 ms. A new in-processChildEventBusnotifies waiters whenever a child state change is dispatched to the UI, and the wait loop blocks onthreading.Event.waitwith a 2 s heartbeat cap (matching the existingwait_progressSSE cadence). A 600 s wait that previously hit storage ~2400 times now wakes only on real state transitions, with ~4× lower SSE traffic in the quiescent case. - Memory tool audit trail — the memory tool now emits
memory.save,memory.update, andmemory.deleteaudit events (the admin-console DELETE route previously emitted onlymemory.delete, so tool-initiated mutations had no audit footprint). All emissions are best-effort and never break the tool call itself. task_agentper-call personas viaskill=—task_agentnow accepts an optionalskill=<name>argument that loads the named skill's content as the sub-agent's persona in place of the hardcoded identity statement. The fixed operating-guidance block (one-shot, tool-use over narration, no follow-up questions) is still layered on top of every persona. High- and critical-risk skills surface their risk tier in the approval header and emit atask_agent.high_risk_skillwarning, matching the existing session-load gate.
Fixed
- Per-role plan / task model overrides could be bypassed by the LLM — the
back-compat
defaultalias auto-synthesised byload_model_registryremained visible to the model even when an operator had configuredmodel.task_alias/model.plan_alias, sotask_agent(model="default")routed to whichever backend the synthesised alias was attached to at boot instead of the configured per-role default. The synthesised alias is now only added when neither the DB nor[models.*]populates the registry, filtered out of the LLM-visible alias list, and explicitly rejected at the validator chokepoint as defense-in-depth. - Mermaid streaming parse errors + progressive
hljs— live-streamed mermaid blocks with bare(,[,{inside unquoted edge or rectangle node labels were re-entering the shape parser and producingParse error, got 'PS'messages. The renderer now autoquotes the two affected label forms (|content|andID[content]) before the SVG cache lookup; shapes whose syntax already nests delimiters (cylinders, subroutines, trapezoids, etc.) are intentionally left alone. The companionhljschange highlights code blocks progressively as they stream rather than only after completion. - Re-auth from inside the proxy-prefixed UI — on a proxied node page
(
/node/{id}/...), an expiring JWT triggered an in-page login modal whose POST went to/v1/api/auth/loginand was rewritten to/node/{id}/v1/api/auth/login. Two latent bugs both blocked re-auth: the console'sAuthMiddlewaredidn't recognise the/node/{id}/prefix over a public path, andproxy_apiwould have forwarded the login request to the upstream node (which mintsJWT_AUD_SERVERtokens the console then rejects). Both fixed: proxied public paths stay public, andproxy_apinow dispatches every entry in_PROXY_AUTH_LOCAL_HANDLERS(login, logout, setup, refresh, status, whoami, oidc/authorize, oidc/callback) to the console's own auth handlers. The dispatch table is a single(method, path) → handlermapping so the test parametrize list can't drift from the implementation. - Appbar visibility + gear-icon dropdown on the dashboard — the dashboard
overlay was covering the entire appbar, hiding the proxy-injected node
picker. The overlay now starts at
top: 48pxand the dashboard's role downgrades fromdialog+aria-modaltoregionso the appbar above it remains reachable. The gear icon converts from a direct settings-panel click into a dropdown with "MCP connections" and "Logout" (the latter with.destructivestyling). The settings-menu keydown handler is now attached synchronously soEscapecan't fall through the brief window between the menu opening and its listeners being installed. - PostgreSQL test backend on the notify dispatcher suite — migration 053's
services_notifytrigger lives only in the alembic chain, but the test fixture creates tables viametadata.create_all. The trigger function + trigger are now declared in_schema.pyand attached viasa.event.listen(services, "after_create", ...)DDL events gated on the PostgreSQL dialect, with the same SQL constants imported by migration 053 so there's a single source of truth.
[1.5.12]
Added
- Enriched backend error messages — provider name and attempted URL are now included in session error responses, so operators can triage connectivity failures without enabling debug logging.
Fixed
/rewindalways emits ahistorySSE event — pre-fix, if the session had no messages remaining after a rewind the history event was skipped, leaving connected UIs with stale content and blocking edit-and-resend flows.
[1.5.11]
This release introduces one forward-only schema migration:
052_model_reasoning_persistence — surface_persisted_reasoning and
replay_reasoning_to_model flag columns on model_definitions.
Added
- SSE refresh-resume — clients that reload mid-stream (browser refresh, tab
restore) now receive an
in_progress_snapshotevent carrying the buffered partial response, so the UI can resume rendering the in-flight turn without losing content. The snapshot is keyed by a monotonic_ws_inflight_seqcounter so a reconnecting client can skip events it already saw. - Reasoning persistence (Phases 1–4) — model reasoning text can now be
persisted to conversation history and optionally replayed to the model on
subsequent turns. Phase 1 persists reasoning text on the history payload.
Phase 2 wires a build-time shape filter and a per-model
replay_reasoning_to_modelflag. Phases 3+4 add full OpenAI Responses API (include=["reasoning.encrypted_content"]) and Chat Completions support; anANTHROPIC_VALID_BLOCK_TYPESshape filter guards the Anthropic path. Two new per-model capability flags (surface_persisted_reasoning,replay_reasoning_to_model) both defaultFalseon unknown and local-server models. - Console home composer: placeholders + toggle — the console landing-page composer now shows context-aware placeholder text and a toggle component for advanced options; an admin polish pass tightened spacing and focus behaviour across the form.
Changed
judge.modelnow requires a named alias — raw provider model IDs onjudge.modelin config are no longer accepted; the judge must reference an alias registered in the model registry. The session-provider raw-model fallback is removed. Existing configs using an unregistered model ID need a corresponding alias entry.
Fixed
replay_reasoning_to_modelAND-gated with model capability — setting the flag for a model that does not declare reasoning-replay support now silently no-ops instead of forwarding reasoning blocks and triggering a provider error.- Coordinator alias resolution unified across placeholder + factory — a placeholder coordinator and the real coordinator factory could previously resolve to different model aliases, producing a visible mismatch in the model display. Both paths now share the same resolution logic.
- Console
cs=Nonefallback in/v1/api/modelsplaceholder — an under-initialised coordinator state no longer 500s when the models endpoint is hit before the coordinator subsystem is fully bootstrapped. - SSE
_ws_inflight_seqalways advances — sequence numbers were previously skipped when an emit was past the buffer cap, leaving gaps in the monotonic counter that brokestate_change/in_progress_snapshotordering on reconnect. - Reasoning persistence shape + replay fixes — per-block
ANTHROPIC_VALID_BLOCK_TYPESfilter applied;reasoning_textis now synthesised alongside non-reasoningprovider_blocksso both appear together in the history payload.
[1.5.10]
This release introduces one forward-only schema migration:
051_skill_notify_on_complete_array_default — backfills
prompt_templates.notify_on_complete from '{}' to '[]'.
Added
- Skills unlock action — operators can unlock an installed skill to allow local customisation. Once unlocked, the skill's resource content, system prompt additions, and notify configuration are editable through the admin UI. Skills shipped as part of a bundle remain locked (read-only) until explicitly unlocked; the unlock is logged to the audit trail. A lock icon in the top-right of the Skills detail pane doubles as the unlock trigger.
Fixed
skills.shinstall endpoint — the install script was targeting an endpoint removed in an earlier refactor; switched to/api/download.- Skills
notify_on_completedefault — the field defaulted to{}(object) instead of[](array), causing notify configurations to be rejected at schema validation. - Skills admin UI modal errors —
.is-visibleclass used consistently instead of inlinestyle.display; stale error text is cleared on submit; designer-review lock-icon UX applied.
[1.5.9]
Fixed
repair=Falseon all display-readload_messagescall sites — passingrepair=Trueon display paths was silently mutating the stored message list, causing divergence between what the UI showed and what the model received on the next turn.
[1.5.8]
This release introduces two forward-only schema migrations:
049_mcp_oauth_schema — OAuth token + consent tables for MCP servers;
050_conversations_source_and_reminders — _source and _reminders columns
on conversations.
Added
-
MCP OAuth 2.1 + PKCE — MCP servers that require OAuth can now be configured with a client ID and secret through the admin UI. The full token lifecycle (acquire → refresh → rotate) is managed automatically; tokens are stored encrypted at rest using a key derived from the JWT secret. The consent flow runs in-browser via a provider redirect. Rolled out in phases:
- Minimum admin form and OAuth schema (
21663d15). - Token-at-rest AES-GCM encryption layer (
a4c335d7). - Per-(user, server) OAuth 2.1 + PKCE flow (
b0f7029f). - Per-(user, server)
ClientSessionpool with OAuth dispatch (1a1043c4). - SDK 401/403 introspection via httpx response hook (
bde09134). - Phase 7 — per-user tool catalog scoping: each user sees only the tools
their OAuth token is permitted to call (
cfc8a6c8). - Phase 7b — per-user resource + prompt pool dispatch (
b368bdee). - Phase 8 — per-user MCP consent UX: users see a consent dialog on first
use of an OAuth-gated server and can revoke consent from their profile;
admins see per-server consent counts in the MCP Servers tab (
61051339).
- Minimum admin form and OAuth schema (
-
Metacognition NudgeQueue — all advisory channels (repeat-tool nudges, watch reminders, wake triggers) are unified into a pull-model
NudgeQueuethat delivers at most one nudge per turn, preventing multi-channel pile-ups that inflate context. Observable changes:- Watch results carry metadata (watch ID,
valid_until, trigger type) through to the system message so the model can reason about recency. - Coordinator idle-children observer: a coordinator with no in-flight children for longer than the configured idle threshold receives a nudge.
- Wake trigger (
IdleNudgeWatcher): sessions waiting on an external event can be unblocked viaChatSession.deliver_wake_nudge_from_queue. - Watch switchover: watch results are now enqueued on the
NudgeQueuerather than the previous_watch_pendinglist, giving them the same delivery guarantees and priority handling as other advisories.
- Watch results carry metadata (watch ID,
-
Structured watch-result card — the UI renders watch results as a styled card with a system-nudge marker, distinct from the assistant message body. On history replay, system-nudge turns are visually distinguished from normal assistant turns.
-
Side-channel persistence —
_sourceand_remindersside-channel fields are persisted to theconversationsstorage table and restored on session resume, so metacognitive context survives process restarts. AREMINDER_TEXT_STORAGE_CAPbyte clamp prevents unbounded growth.
Fixed
- Replay consistency — queued user messages captured mid-loop are now
persisted and replayed in the correct order on a subsequent
eventssubscription. Coordinator history replay fixed: blank assistant cards and out-of-order tool results on the coordinator tree no longer occur when the coordinator has mixed queued + delivered messages. - Session reminder preservation on fork + resume —
_sourceand_remindersare carried through workstream fork and restored from storage on resume. - NUL-byte sanitization in storage — PostgreSQL rejects
\x00in text columns;_sourceand_remindersnow strip NUL bytes on write. - Console coordinator subsystem bootstrap — the coordinator subsystem is now committed atomically on first model add; startup teardown is offloaded to avoid blocking the event loop.
- MCP
asyncio.timeoutoverasyncio.wait_for— Python 3.11'swait_forwraps the coroutine in a fresh task, breaking anyio'saclosescope exit. Replaced withasync with asyncio.timeout(N)for safe cleanup. - MCP pool-reuse 401 recovery — a reused
ClientSessionreturning 401 now replaces the pool entry with a fresh session; the carrier token is owned by the pool entry to prevent a race between the 401 handler and a concurrent request. - OIDC hardening — multiple security and correctness fixes:
SSRF + plaintext credential exfil via discovery document (sec-1, sec-3);
TURNSTONE_OIDC_REDIRECT_BASEnow required, Host-header fallback removed (sec-2); atomic user + identity provisioning prevents orphan rows (bug-1); callback robustness — typed exceptions, shape checks, log sanitization, JS race (bug-4–6, sec-4); role-mapping concurrency serialized (bug-2, perf-1); stranded-user self-heal on role-mapping failure (cumulative bug-1).
[1.5.7]
Added
- Inline node picker — a compact node-switcher dropdown in the console header replaces the "← Back to console" banner, so operators can switch between nodes without a full navigation.
Fixed
- Queued user messages injected mid-loop — messages queued while a generation was in progress were not being delivered at the correct seam and could be dropped or reordered when the worker consumed the queue.
- Search tool output bounded — pathological inputs (very long lines with no whitespace) could produce search results exceeding the context budget. Output is now clamped before reaching the message.
[1.5.6]
Added
api_surfacetoggle — model definitions gain anapi_surfacefield ("chat"|"responses") that selects which OpenAI-compatible API surface the provider client uses. Enables Mistral Medium reasoning via the Responses surface; Chat Completions remains the default for all other models.- Healthy model aliases per node —
GET /v1/api/cluster/nodesnow includes ahealthy_aliaseslist per node, so the coordinator and operators can see which model aliases are currently reachable without a separate per-model health probe. - Plan/task agent settings in Models → Roles — the Models admin tab's
Roles sub-tab gains
plan_agentandtask_agentrows so operators can configure per-kind reasoning effort and alias overrides from the UI rather than editingconfig.toml. Live-refresh dropdowns update in place when model definitions change.
Fixed
- Memory candidate selection — recall now uses OR-of-terms BM25 with query-aware candidate-set selection, dramatically improving recall for queries whose terms span multiple stored entries.
- Workstream model + config preserved on rehydrate — reopening a closed workstream no longer overwrites the model alias and per-workstream config with session defaults.
- Console home composer: attachments + user-message pills — multipart attachments in the home composer were not forwarded correctly; user-message pills in the coordinator chat pane were missing.
[1.5.5]
Fixed
- Saved-workstream tool result rendering — tool results in closed workstreams were not rendering on history replay. Audit-trail decoration for tool calls is now applied on the replay path.
[1.5.4]
Added
- Stage 3 SessionManager Children primitive lift — child workstreams are
first-class citizens in the cluster event bus.
child_ws_stateevents are pushed through the cluster SSE stream so the console tree view updates in real time without polling.list_childrenandget_childprimitives onSessionManagerprovide a consistent cross-node view of the coordinator's spawn tree. - Multi-select delete for Saved Coordinators — the Saved Coordinators grid in the console admin panel now supports checkbox multi-select with a bulk-delete action.
[1.5.3]
This release introduces one forward-only schema migration:
048_workstream_reaper_index — partial composite index on workstreams for
the orphan-reaper query.
Fixed
- Coordinator orphan reaping scoped by heartbeat — the session manager's
close_idlepass now scopes the DB-orphan reaper byservices.last_heartbeatso workstreams belonging to a live node are not incorrectly reaped.bulk_close_stale_orphansandtouch_workstreamstorage primitives added; a partial composite index keeps the reaper scan cheap. - Coordinator pool idle cleanup — a periodic task on the console now closes coordinator pool entries whose session has gone idle past the configurable threshold, preventing pool exhaustion on long-running consoles.
[1.5.2]
Added
- Metacognition themed reminder bubble — repeat-tool and user-reminder
nudges are rendered as a distinct styled bubble rather than being injected
inline into the assistant message, making it easier to distinguish model
output from metacognitive annotations. The CLI REPL gains matching
on_user_reminder/on_tool_remindercallbacks.
Fixed
- Metacog streak detector — the N≥3 sequential-same-call streak detector now fires correctly on the third repetition; a write-success-clear that reset the counter after a successful tool call (preventing streaks across mixed-outcome sequences) was removed.
- Metacog reminders isolated to side-channel — reminder text no longer appears in the user content turn; it flows through a dedicated side-channel the session injects into the system context, preventing the model from attributing it to the user.
[1.5.1]
Added
pending_approval_detailon childws_stateSSE events — coordinators now receive the child's pending approval detail inchild_ws_stateevents, enabling the coordinator to surface approval prompts without a separate poll.
Fixed
- Coordinator registry auto-refresh — the console coordinator registry now refreshes when model definitions change, so a newly added alias is visible to coordinators without restarting.
- Coordinator fan-out default — coordinators now fan out to independent child workstreams by default instead of serialising them, matching the documented contract for parallel-work patterns.
wait_for_workstreammessage cap raised to 10 KiB — large plan summaries and tool results from child workstreams were silently truncated at the previous 4 KiB cap.- Coordinator SSE isolated on dedicated thread pool — coordinator SSE
polling now runs on a dedicated 200-thread executor, matching interactive's
sse_executor, so coordinator long-poll blocking no longer contends with storage and routing workers on the default pool.
[1.5.0]
User-visible additions: a unified workstream HTTP surface (interactive and coordinator under one URL family), inline child approvals, coordinator composer parity, progressive rendering, OIDC authentication, MCP OAuth foundations, and a redesigned UI built on the Design System v1 token layer.
This release removes the pre-1.5 body-keyed and query-keyed URL family. See Removed (BREAKING) below before upgrading from a 1.x stable line.
This release introduces the following forward-only schema migrations that the server applies automatically on first startup. All are additive; no data loss.
039_workstream_kind—kind+parent_ws_idcolumns onworkstreams.040_coord_cluster_admin_perms— grantsadmin.coordinator+admin.cluster.inspectto the builtin-admin role.041_workstream_index_tuning— refined indexes for the workstream query mix introduced by 039.042_coord_trust_send_perm— addscoordinator.trust.sendpermission to builtin-admin.043_skill_description_required— backfills emptydescriptionrows inprompt_templates.044_skill_kind— addskindclassifier column toprompt_templates(interactive/coordinator/any).045_skill_risk_level_rename— renamesprompt_templates.scan_status→risk_level.046_drop_hash_ring_tables— drops the hash-ring bucket tables superseded by rendezvous routing in 1.4.047_drop_coord_spawn_quota_settings— removes the spawn-quota settings rows removed from the coordinator in 1.5.0a4.
Added
- Inline child approvals — pending tool approvals on coordinator child
workstreams surface directly in the coordinator tree view. A risk pill shows
the judge verdict (or "pending" while the judge evaluates); Approve/Deny
buttons appear inline so operators do not need to navigate to the child's
workstream.
pending_approval_detailis exposed onGET /v1/api/dashboardand passed through the cluster live-bulk SSE payload so all connected clients render approval prompts simultaneously. LLM judge verdicts are cached client-side and replayed on SSE reconnect. - Coordinator composer parity — the coordinator composer now supports Stop, Send-to-queue, and Attach (file upload), matching the interactive workstream composer feature set.
- Per-call model and judge override on coordinator composer — operators can override the model alias and judge model for a single coordinator send from the composer, without changing the node-wide or role-wide defaults. Bad aliases return a corrective error listing available choices.
- Coordinator status bar + richer history replay — each coordinator workstream gains a per-coordinator status bar showing active children, token spend, and generation state. History replay in the coordinator panel is extended to include tool results and thinking blocks.
- Coordinator child error surfacing + memory tool — child workstream
errors are surfaced as distinct error rows in the coordinator tree view
rather than disappearing silently. The coordinator gains access to a
memorytool (same interface as interactive) for retrieving stored facts. - Coordinator inline tool-batch construct — the coordinator tool approval UI replaces the separate approval dock with an inline batch construct that groups all pending tool calls for a given turn into a single review card.
- Node capability auto-detection — nodes report kernel-level capabilities
(available memory, CPU count, accelerator presence) via
/v1/api/node/capabilitiesat startup, enabling the console to filter model aliases offered to coordinators routing to that node. - Skills: paste
SKILL.mdto auto-fill the Create Skill modal — pasting aSKILL.mdfile's content into the modal auto-populates the name, description, and configuration fields. - Progressive mermaid rendering — Mermaid diagrams begin rendering as soon as a complete diagram block is detected in the stream rather than waiting for the full response; the diagram re-renders in place as the model extends it.
- LaTeX and MathML delimiter support —
\(…\)inline and\[…\]block math delimiters are now recognised alongside the existing$$fences.
Removed (BREAKING — 1.5.0)
-
Legacy body-keyed and query-keyed URL family for the workstream interaction verbs. Pre-1.5 interactive shipped both a path-keyed and a body-keyed surface for the same five verbs; this release drops the body-keyed and query-keyed mounts (and the
make_legacy_body_keyed_adapter/make_legacy_query_keyed_adaptershims that backed them). External SDK consumers on stable 1.0/1.3/1.4 must move to the path-keyed shape:Removed (1.0/1.3/1.4) Use instead GET /v1/api/events?ws_id=XGET /v1/api/workstreams/{ws_id}/eventsPOST /v1/api/send(bodyws_id)POST /v1/api/workstreams/{ws_id}/sendDELETE /v1/api/send(bodyws_id)DELETE /v1/api/workstreams/{ws_id}/sendPOST /v1/api/approve(bodyws_id)POST /v1/api/workstreams/{ws_id}/approvePOST /v1/api/cancel(bodyws_id)POST /v1/api/workstreams/{ws_id}/cancelPOST /v1/api/workstreams/close(body)POST /v1/api/workstreams/{ws_id}/closeCalls to the old URLs return 404 on 1.5.0+. Bodies on the new URLs no longer carry
ws_id(the path provides it); theSendRequest/ApproveRequest/CancelRequestPydantic schemas drop the field, andCloseWorkstreamRequestslims to a single optionalreasonfield (the body is still required to be valid JSON — send{}when omitting all fields)./v1/api/planand/v1/api/commandare unaffected and remain body-keyed in this release. The bundled web UI, channel adapters, Python SDK, TypeScript SDK, and console routing-proxy SDK ship the new URLs automatically; pinning to ≥ 1.5.0 is enough.The console routing proxy's
/v1/api/route/...family is updated alongside:/v1/api/route/workstreams/{ws_id}/<verb>replaces the pre-1.5/v1/api/route/{send,approve,cancel,workstreams/close}mounts.DELETEis now passed through (client.request(method, ...)instead ofclient.post(...)) so the new dequeue route works through the proxy. Audit attribution forDELETEon/sendis logged asroute.workstream.dequeuerather thanroute.workstream.send.Auth scope wiring (
WRITE_PATHS/APPROVE_PATHSliterals plus the path-keyed verb match inrequired_scope) updated to grantwritefor path-keyedsend/cancel/close,approvefor path-keyedapprove, andwriteforDELETEon path-keyed/send. The/node/*proxy branch mirrors all four.
Changed
-
Dashboard row shape:
id→ws_id. TheGET /v1/api/dashboardrow dict now keys the workstream identifier asws_id(matching the rest of the v1 workstream surface — active list, saved list, history, detail). The Stage 2 list-verb lift converged/v1/api/workstreamsand/v1/api/workstreams/savedonws_idbut left dashboard alone to keep that PR's diff focused; this lands the same rename on the remaining endpoint so the v1 row shape is consistent across the family. PydanticDashboardWorkstreamand the TypeScript SDKDashboardWorkstreaminterface both rename the field accordingly. The bundled web UI is the only consumer that readsdashboard.workstreams[].idand is updated atomically; no external SDK on a stable line reads the field, so the swap is bounded by normal static-asset reload. Console_fetch_live_block(cluster-inspect's projection over a remote node's dashboard payload) is updated to match. -
Coordinator gains rich
ws_statepayload + live activity broadcast ([§ Post-P3 reckoning item #2 follow-up]). Pre-lift coord's cluster broadcast was state-only — the dashboard's coord rows showed the state column flipping but thetokens,context_ratio,activity, and per-turncontentfields were all hardcoded to zero / empty. The lift turnson_status/on_content_token/on_thinking_start/on_thinking_stop/on_stream_end/on_tool_resultinto shared bodies on :class:SessionUIBaseso coord populates the same per-ws metric fields interactive does (the fields were already declared on the base; only the writes were WebUI-specific).coord_adapter.emit_statenow reads the UI's snapshot under_ws_lockvia the new :meth:SessionUIBase.snapshot_and_consume_state_payloadhelper and passes the rich kwargs through tocollector.emit_console_ws_state; the cluster dashboard's coord rows now render with the same tokens / activity / content / context_ratio fields interactive rows do.Three observable behaviour changes (all CHANGELOG-callout-worthy):
- Coord persists
usage_eventstorage rows. Pre-lift only WebUI did. The liftedon_statusbody unifies usage tracking so governance dashboards / token-spend queries see coordinator consumption alongside interactive. Operators queryingusage_eventbyws_idwill see coord rows for the first time. - Coord broadcasts live activity transitions. New
ClusterCollector.update_console_ws_activity(ws_id, *, activity, activity_state)method (namedupdate_*rather thanemit_*to flag the no-fan-out asymmetry vs. the rest of theemit_console_ws_*family — it updates the in-memory pseudo-node row but intentionally does NOT fan out a separate SSE event). The cluster dashboard's per-ws polling reads the in-memory pseudo-node row, so activity ticks land on the next snapshot fetch (matches WebUI's behaviour where activity events are observational; not fanned out through the cluster SSE stream). - Cluster
cluster_stateevents for coord rows now carry non-zerotokens/contentfields. Frontend rendering that conditionally hid these on coord rows can drop the branch.
Architecture changes:
_MAX_TURN_CONTENT_CHARSmoved fromturnstone.servertoturnstone.core.session_ui_baseso coord enforces the same per-turn content cap interactive does.- WebUI keeps
on_status/on_tool_result/on_erroroverrides that layer Prometheus_metrics.record_*calls (node-only) on top of the shared body viasuper()— the Prometheus surface stays node-scoped (the console isn't a node and has no /metrics endpoint). ConsoleCoordinatorUIadds a_broadcast_activityoverride that fans out via the cluster collector instead of the global SSE queue (which is node-only on interactive).coord_endpoint_configwires a new_coord_spawn_metricshook (mirrors interactive's) so the per-spawn_ws_messagesincrement +_ws_turn_tool_callsreset happen on coord too.
Test additions: 23 new tests in
tests/test_coord_rich_ws_state_payload.pypin the per-ws metric writes (status, content accumulation, activity tracking, tool-result counters, stream-end activity clear), the snapshot helper's IDLE/ERROR drain semantics + single-lock-acquisition guarantee, the adapter's rich-payload pass-through + defensive None-UI handling, the activity broadcast (collector wire + failure swallow + no-op-when- collector-unset + dedup against last-emitted state), the spawn_metrics hook, and a concurrent-writes-during-snapshot stress case (cycles through running / idle / error so the drain branches actually run against a concurrent writer). Plus WebUI-override regression tests confirming_metrics.record_*still fires on top of the lifted bodies. Existingtests/test_webui_content.pyupdated to import_MAX_TURN_CONTENT_CHARSfrom its new home inturnstone.core.session_ui_base;tests/test_coordinator_adapter.pyupdated to expect the rich-payload kwargs (tokens=0defaults) onemit_console_ws_state.Two deferred follow-ups (out-of-scope for this lift, flagged for tracking):
- Synchronous
record_usage_eventINSERT on coord worker thread. The liftedon_statusbody persists usage rows on every provider response — same shape WebUI uses, but coord workers can fire multi-step plan/task agent loops where each response blocks the worker for a write transaction. Parity with WebUI is the explicit goal here; if coord throughput becomes a concern, batch usage_event writes onto a background flusher thread (one batch INSERT per N events / per K ms) on both kinds. - Coord assistant turn content now flows on the cluster SSE
stream (
/v1/api/cluster/events). Pre-lift the broadcast wascontent=""; post-lift it carries the joined assistant output. The cluster SSE stream has no per-user filter today — extends an existing cross-tenant exposure (interactivecluster_stateevents already carry content) to a previously-empty channel (coord rows). Proper fix needs the SSE endpoint gated onadmin.cluster.inspect(matching/v1/api/cluster/ws/{ws_id}/detail) or per-listener user_id filtering. Tracked as a separate security-tightening project; not gating this lift since it inherits an existing exposure rather than introducing a new mechanism.
- Coord persists
-
history/detailverb bodies lifted across both kinds ([Stage 2 Verb Lift —history/detail]). The coordGET /v1/api/workstreams/{ws_id}/historyandGET /v1/api/workstreams/{ws_id}handlers now share two factory bodies viamake_history_handler(cfg)andmake_detail_handler(cfg). The lift adds both endpoints to the interactive surface as a feature gain (pre-lift only coord exposed them; interactive consumers had to subscribe to/eventsSSE just to read history rows or display fields). No newSessionEndpointConfigfields — the factories reusepermission_gate,manager_lookup,not_found_label,audit_action_prefix, and (for history's storage-fallback kind check)list_kind— all already wired by both production lifespans.Three observable behaviour changes (all documented per kind):
- Interactive gains
GET /v1/api/workstreams/{ws_id}. Pre-lift interactive had no detail endpoint — SDK consumers had to read display fields from the SSE replay on/eventsor scrape the active list. The lifted body lazy-rehydrates a closed/evicted workstream viamgr.open()so the response shape is stable across loaded / persisted-only states. Same{ws_id, name, state, user_id, kind}shape coord exposed pre-lift, now available on both surfaces. - Interactive gains
GET /v1/api/workstreams/{ws_id}/history. Same?limit=query param contract as coord (default 100, max 500, malformed values fall back to 100, out-of-range clamps to [1, 500]). Persisted-but-not-loaded interactives serve history without rehydrating — the lifted body falls back to a storage-row- kind check (via
cfg.list_kind) whenmgr.getreturnsNone, mirroring coord's pre-lift_resolve_coordinator_or_404ladder.
- kind check (via
- Storage / manager-lock work moved off the event loop on coord.
The lifted
historybody always runsstorage.get_workstream(storage-fallback path) andstorage.load_messagesthroughasyncio.to_thread; pre-lift coord ran them inline on the event loop. Long-tail message reads on a saturated console no longer stall every other async handler for the duration of the SQL.
Pydantic schemas:
CoordinatorDetailResponseandCoordinatorHistoryResponseremoved; both folded intoWorkstreamDetailResponse/WorkstreamHistoryResponseon the sharedserver_schemas.py(mirrors the list lift's pattern forWorkstreamInfo). Both server and console OpenAPI specs reference the unified schemas;server_spec.pygainsEndpointSpecentries for the new interactive endpoints. TS SDK gainsWorkstreamDetailResponse/WorkstreamHistoryResponseinterfaces insdk/typescript/src/types.ts;openapi-{server,console}.jsonregenerated.GET /v1/api/workstreams/{ws_id}/historyis the only verb whose lifted body keeps a kind-aware storage fallback (viacfg.list_kind);detaildefers cross-kind isolation tomgr.open()itself. - Interactive gains
-
list/savedverb bodies lifted across both kinds ([Stage 2 Verb Lift —list/saved]). The interactiveGET /v1/api/workstreams+GET /v1/api/workstreams/savedand coordGET /v1/api/workstreams+GET /v1/api/workstreams/savedhandlers now share two factory bodies viamake_list_handler(cfg)andmake_saved_handler(cfg). Four newSessionEndpointConfigfields capture the per-kind divergence:list_resolve_titles: ListResolveTitles | None— interactive wires :func:turnstone.core.memory.get_workstream_display_names(new bulk helper added on the storage layer +memory.py) so the active-list endpoint resolves every user-set alias in ONESELECT ... WHERE ws_id IN (...)instead of the pre-lift per-row N+1. Coord wiresNone(no alias surface today).list_kind: WorkstreamKind | None— required storage-side kind classifier passed tolist_workstreams_with_history. Interactive wiresWorkstreamKind.INTERACTIVE; coord wiresWorkstreamKind.COORDINATOR. Distinct fromaudit_action_prefix(audit-action namespacing) so adding a third kind doesn't have to overload the audit prefix as a classifier; missing value surfaces as 500 with a clear log line rather than silently filtering for the wrong kind.saved_state_filter: str | None— coord wires"closed"so only explicitly-closed coordinators surface in the saved-card grid. Interactive wiresNone(the storage layer already excludesstate='deleted'tombstones).saved_loaded_lookup: SavedLoadedLookup | None— coord-only defence-in-depth filter that excludes ws_ids currently in the in-memory pool (a row can bestate='closed'for a few seconds while the close-emit sequence races the in-memory pop). Interactive wiresNone.
Five observable behaviour changes (all documented per kind):
- Active-list top-level key converges on
"workstreams". Pre-lift coord returned{"coordinators": [...]}; the lifted body returns{"workstreams": [...]}for response-shape parity with interactive. Coord is a 1.5.0aN-only surface — never shipped stable — so SDK / frontend consumers swap once and there's no compat shim or fallback (the convergence MUST land before v1.5.0 stable perproject_unification_before_stable.md). - Saved-list top-level key converges on
"workstreams". Same shape change as the active list, applied toGET /v1/api/workstreams/savedon coord. Coord-only surface; no compat shim. - Active-list row key renames
"id"→"ws_id"on interactive. Pre-lift interactive used the bareidfield while every other shared verb on this surface (cancel, open, events, create, saved-list) usesws_id. Convergence eliminates the internal inconsistency. Frontend consumers readingws.idfrom the active-list response swap tows.ws_id. Interactive HAS shipped stable across 1.0 / 1.3 / 1.4, but the active-list endpoint is consumed by the bundled JS only — there's no external SDK on those stable lines reading the field. Browser-cache staleness is bounded by normal static-asset reload on next page load. - Active-list row gains always-include fields.
user_idwas coord-only;kind+parent_ws_idwere interactive-only. Both kinds now populate all three.parent_ws_iddefaults toNonefor coord (coordinators have no parent). - Storage / manager-lock work moved off the event loop on
interactive. The lifted
savedbody always usesasyncio.to_threadforlist_workstreams_with_history; pre-lift interactive ran it inline (correlated COUNT subquery can stall every other async handler on a cluster with thousands of saved rows). Coord already usedto_thread(perf-2 from the saved-coordinators review); convergence lifts interactive up. The active-list body also movesmgr.list_all+ per-row title resolution off the event loop on both kinds.
Pydantic schemas:
WorkstreamInfo.idrenamed →ws_id,WorkstreamInfo.user_idfield added.CoordinatorInfoandCoordinatorListResponseremoved (folded into the unifiedWorkstreamInfo/ListWorkstreamsResponse);console_specactive-list endpoint now points atListWorkstreamsResponse. OpenAPI spec snapshots regenerated.GET /v1/api/dashboardis not in the lift's scope and still returns rows keyed onid. A separate cleanup PR will converge the dashboard row shape with the rest of the v1 surface. -
SessionManager.creategains a deferred-emit option; liftedcreateHTTP handler eliminates the phantom create→close pair on coord rollback.SessionManager.createnow acceptsdefer_emit_created: bool = False(default preserves the legacy "advertise immediately" contract for direct callers); two new methods complete the deferred-create bracket:SessionManager.commit_create(ws)fires the deferredemit_createdevent after the caller's post-create work confirms the workstream should be advertised.SessionManager.discard(ws_id)releases the in-memory slot- cleans up the UI WITHOUT firing
emit_closed— the workstream's existence was never advertised, so there's nothing to advertise on rollback. Storage-row deletion stays a separate concern (caller invokesdelete_workstreamfor a complete rollback), mirroringmgr.create's split between slot reservation andregister_workstream. Logs awarning(session_mgr.discard.after_emit_created) when invoked on a workstream that's already been advertised (non-deferred create or post-commit_create); the slot is still released so capacity isn't stranded, but the warning surfaces the caller-bug case whereclosewould have been the right call.
- cleans up the UI WITHOUT firing
The lifted
make_create_handlernow uses this bracket: passdefer_emit_created=True, validate uploaded attachments, thenmgr.commit_create(ws)on success /mgr.discard(ws.id)on failure. Pre-fix, coord'smgr.createfiredemit_createdsynchronously — a rollback then calledmgr.closewhich firedemit_closed, surfacing a quick create→close pair on the cluster events stream that the collector's diff-reconcile had to handle. Post-fix, a rejected upload produces zero events. Interactive'semit_createdis a documented no-op stub so the deferral is observably a no-op there; thews_createdbroadcast on the global SSE queue continues to fire from the kind's post_install callback after attachment validation passes (unchanged).Direct callers of
mgr.create(test fixtures, the CLI REPL, channel adapters) keep the defaultdefer_emit_created=Falseand see no behaviour change. -
Coordinator HTTP surface unified under
/v1/api/workstreams/([Stage 2 Priority 0]). The experimental/v1/api/coordinator/*URL tree from 1.5.0aN is removed; coord verbs now mount at the same shape as interactive workstreams via a shared route registrar (turnstone.core.session_routes). Path mapping:Was (1.5.0aN) Now POST /v1/api/coordinator/newPOST /v1/api/workstreams/newGET /v1/api/coordinatorGET /v1/api/workstreamsGET /v1/api/coordinator/savedGET /v1/api/workstreams/savedGET /v1/api/coordinator/{ws_id}GET /v1/api/workstreams/{ws_id}POST /v1/api/coordinator/{ws_id}/{verb}POST /v1/api/workstreams/{ws_id}/{verb}Permission scopes, request / response bodies, and SSE event shapes are unchanged. Callers on the experimental 1.5.0aN coord SDK must swap their URL prefix; the legacy paths are gone with no compat shim. Stable releases (1.0 / 1.3 / 1.4) never exposed
/v1/api/coordinator/, so this change is a no-op for anyone upgrading from a stable line.Two handler bodies (
approve,close) lifted into the shared registrar with kind branching behindSessionEndpointConfig— both kinds share one implementation per verb. Two related behavior changes on the interactive close path:mgr.close()race-loss returns 404 (was 500 on coord; "popped between .get() and .close()" is a not-found semantic, not a server error).- Audit-write failures (
record_auditraising on the storage write) are now caught and logged atwarninglevel; the close still returns 200. Previously the interactive path let the exception propagate as HTTP 500. Coord previously already swallowed; convergence is intentional — operators monitor thews.close.audit_failedlog line in both kinds the same way.
Other shared verbs (
send,cancel,open,events,create,list,saved,history,detail) keep their per-kind handlers — body convergence for those requires SessionManager- side refactors (e.g. Priority 1's worker-dispatch unification forsend) or coordinated frontend changes (response-shape unification forlist/saved) that fall outside Priority 0 scope. -
TypeScript SDK bumped to 0.4.0 to flag the URL change for any 1.5.0aN-era consumer of the experimental coord client. The
openapi-{server,console}.jsonreference specs ship with the unified path tree. -
Worker dispatch unified across interactive + coordinator ([Stage 2 Priority 1]). The atomic check-and-(spawn-or-queue) decision for
ChatSession.sendnow lives inturnstone.core.session_worker.sendand is shared by both paths. Interactive/v1/api/send, the coordinator adapter, the watch-result dispatch, the rewind/retry path, and the initial-message-on-create path all gate onWorkstream._worker_running(set/cleared atomically underws._lock) instead ofThread.is_alive()— closes a race where two senders could spawn parallel workers on the same ChatSession.The
/sendHTTP body itself stays per-kind in this PR. Verb-shape convergence (one shared factory body with capability flags for attachments / queue priorities / metric increments) is tracked as P1.5 and MUST land before 1.5.0 stable — letting the fork ship into the stable line bakes the duplication in for the lifetime of the 1.5 track. -
/sendbody lift + coordinator attachments + queue surface parity ([Stage 2 Priority 1.5]). The/sendHTTP handler is now ONE factory body (make_send_handler(cfg)) wired with capability flags on both kinds; the four attachment endpoints (upload/list/get_content/delete) are also unified viamake_attachment_handlers(cfg). Coord workstreams light up:POST/GET /v1/api/workstreams/{ws_id}/attachments,GET .../attachments/{aid}/content,DELETE .../attachments/{aid}— same shape, same caps, same reservation flow as interactive.POST /v1/api/workstreams/{ws_id}/sendacceptsattachment_ids(or auto-consumes pending) and returnsattached_ids/dropped_attachment_idsfor surfacing partial reservations. Live-worker reuse path also returnspriority/msg_id(parity with the interactivestatus: queuedshape).
Backend parity is end-to-end: storage layer was already kind-agnostic; the route registrar's
AttachmentHandlersslot has been there since Stage 2 P0; the multi-node attachment routing-proxy on the console (route_attachment_proxy) was already shipping. P1.5 is the wiring + verb-shape lift that lets these primitives surface on the coord side.Coord dashboard rendering surfaces an attachment-count badge on past messages with attachments; full chip rendering with click-to-view is deferred (the coord dashboard is diagnostic-leaning and chip parity isn't on the critical path for the unification thesis). Python SDK adds
coordinator_send/coordinator_upload_attachment/coordinator_list_attachments/coordinator_get_attachment_content/coordinator_delete_attachmentonAsyncTurnstoneConsole+TurnstoneConsole. TS SDK regenerated; bumped to 0.5.0.Three lifted helpers (
sniff_image_mime,classify_text_attachment,upload_lock) moved fromturnstone/server.pytoturnstone/core/attachments.pyso both processes use the canonical implementation. The interactive surface keeps the same behaviour; the helpers are simply imported from their new home.coordinator_sendno longer returns429on a full worker queue — the unified body returns200 {"status": "queue_full"}for parity with interactive. Existing callers checking for429should switch to the status-code shape.Coord
GenerationCancellednow emitsstate=idle+stream_end(parity with interactive); pre-P1.5 a cancel-killed coord worker would have terminated silently with no state event. Cluster fanout / alerting keyed onstate=errorfor cancelled coord workers should switch to monitoringstream_end/state=idletogether. -
SessionKindAdapterProtocol split into construction + emission ([Stage 2 Priority 3]). The adapter Protocol now covers only what every kind must implement (kind/build_ui/build_session/cleanup_ui); the four lifecycle emit methods (emit_created/emit_state/emit_rehydrated/emit_closed) move to a separateSessionEventEmitterProtocol wired through a new optionalevent_emitter: SessionEventEmitter | Nonekwarg onSessionManager. Both production adapters (interactive onserver.py, coordinator onconsole/server.py) implement both Protocols and are passed as bothadapterandevent_emitterat lifespan-construction time, so production behavior is unchanged. The interactive adapter's threeemit_created/emit_state/emit_rehydratedmethods remain documented no-op stubs (those events fire from out-of-band paths — the create handler enqueuesws_createdafter attachment validation,WebUI._broadcast_stateemitsws_state);emit_closedstays load-bearing as the sole transport path forws_closedonto the global SSE queue. -
cancelverb body lifted across both kinds ([Stage 2 Verb Lift —cancel]). The interactive/v1/api/canceland coord/v1/api/workstreams/{ws_id}/cancelhandlers now share one body viamake_cancel_handler(cfg, *, audit_emit=None); per-kind divergence captured by a newcancel_forensics: CancelForensics | Nonefield onSessionEndpointConfig(interactive wires_capture_cancel_forensics; coord wiresNone). Three observable behaviour changes for coord callers:- Coord cancel now accepts a
forceflag. Same shape as interactive: posting{"force": true}abandons the worker thread and emitsstream_endso a stuck coord generation can be recovered without waiting for the daemon thread to exit. Pre-lift coord ignoredforce. - Coord cancel response always includes
"dropped". Pre-lift coord returned bare{"status": "ok"}; the lifted body returns{"status": "ok", "dropped": {}}(always-include parity with interactive). SDK consumers don't need to branch on kind to readdropped. - Coord cancel returns 400 when the workstream's session is
None. Pre-lift coord calledcoord_mgr.cancelwhich silently no-op'd on a placeholder/build-failed workstream; the lifted body 400s with{"error": "No session"}for parity with interactive's pre-existing branch.
Two observable changes for interactive (asymmetric — coord pre-lift already had this behaviour):
resolve_plannow runs on every cancel (previously gated onwas_running).resolve_planhas an internal_pending_plan_review is Noneguard, so the call is no-op when no plan review is pending. Lift gives interactive coord's pre-lift recovery path: a stuck plan-pending state from a crashed worker can be cleared viacancelinstead of requiring a workstream close + rehydrate.resolve_approvalruns on every cancel only whenui._pending_approval is not None(the lifted body gates the call).resolve_approvalis not idempotent — it always broadcastsapproval_resolvedand overwrites_approval_result— so the gate prevents a stale resolution event from leaking on idle cancels while preserving the recovery path when an approval really is pending.
Coord
coordinator.cancelaudit detail now includesforceso operator-driven recovery is distinguishable from a routine cancel in the audit log.Three /review fixes folded into the same commit:
- No more stale
approval_resolvedSSE event on idle cancel. The lifted body'sresolve_approvalcall is now gated onui._pending_approval is not None. Pre-fix, the unconditional call would broadcast a phantomapproval_resolvedto every SSE listener even when no prompt was pending — listener UIs that key on the event would dismiss prompts they didn't have. - Force-cancel now clears
_worker_runningalongsideworker_thread. Previously the force path left the half- state(_worker_running=True, worker_thread=None), which routed any follow-upsendthrough the queue-enqueue path onto the abandoned worker (where the cancel flag short-circuits the queue-drain seam, leaving the message orphaned until the next spawn). Restores the(worker_thread, _worker_running)invariantsession_worker.senddocuments. coordinator_stop_cascadenow treats child cancel400 + "No session"asskipped(was previouslyfailed). Lifted coord cancel returns 400 on placeholder / build-failed children — matching the pre-lift outcome where those children were silently no-op'd, so the cascade response'sfailedbucket no longer fires spurious operator alerts.
- Coord cancel now accepts a
-
openverb body lifted across both kinds ([Stage 2 Verb Lift —open]). The interactivePOST /v1/api/workstreams/{ws_id}/openand coordPOST /v1/api/workstreams/{ws_id}/openhandlers now share one body viamake_open_handler(cfg, *, audit_emit=None). Per-kind divergence captured by two newSessionEndpointConfigfields:open_resolve_alias: AliasResolver | None— interactive wires :func:turnstone.core.memory.resolve_workstreamso callers can pass user-friendly aliases ("my-debug-ws") in the path param. Coord wiresNone(hex ids only).open_post_load: OpenPostLoad | None— interactive wires the UI-replay (clear_ui+ history) + handler-sidews_createdenqueue onto the global SSE queue. Coord wiresNoneand relies on the cluster collector fan-out fromCoordinatorAdapter.emit_rehydrated.
Load-bearing fix (§ Post-P3 reckoning item #3): interactive
open_workstreampreviously calledmgr.create(ws_id=resolved_id)+ws.session.resume(...)to rehydrate, bypassingmgr.open()entirely. After the lift both kinds route throughmgr.open()— which makesInteractiveAdapter.emit_rehydratedreachable on interactive (it had been dead-by-routing) and gives the manager a single rehydrate code path to maintain.emit_rehydratedstays a documented no-op stub on the interactive adapter (the handler-sidews_createdenqueue fromopen_post_loadis the load-bearing emission).Two observable behaviour changes for interactive callers:
- Cross-kind open returns 404 (was 400). Pre-lift had a
pre-mgr storage probe that returned
400with"Workstream is not an interactive kind"for coord rows; the lift consolidates onmgr.open()'s singleNone- return contract for missing / wrong-kind / tombstoned rows. Security boundary unchanged. - Already-loaded response uses
ws.namedirectly (wasget_workstream_display_name(resolved_id) or resolved_id). A workstream renamed viaset_workstream_aliasafter being loaded into memory will surface the storage-row name in the open response'snamefield instead of the latest alias. The dashboard listing endpoint still resolves aliases on its own pass, so the user-visible workstream name in the tab strip isn't affected.
Coord behaviour unchanged.
Two /review fixes folded into the same commit:
- Resume failures now return 5xx instead of broken-200.
SessionManager.open()previously caught andlog.debug- swallowed exceptions fromChatSession.resume(which assignsself.messagesbefore the config-restore block, so a partial-failure resume — corruptedworkstream_configrow, model-registry mismatch on a saved alias, malformedtemperature/max_tokens— would leave the session with history but with default config). Pre-lift, the interactive open handler calledws.session.resume(...)directly and let exceptions propagate as 500. The lift accidentally inherited the swallow because it routed throughmgr.open(). Restored pre-lift behaviour:mgr.open()now re-raises resume exceptions after rolling back the slot (cleanup_ui+_remove_locked), so the lifted handler returns 500 with a correlation id and the storage row stays available for a retry instead of silently 200'ing with broken state. except Exceptionin the lifted body documents intent. The bare exception catch aroundmgr.open(ws_id)is intentional — the kind's session factory has no documented exception spec, and resume can propagate fromChatSession.resume. A one-line rationale comment in the handler body keeps a future contributor from narrowing it incorrectly.
-
eventsverb body lifted across both kinds ([Stage 2 Verb Lift —events]). The interactiveGET /v1/api/events?ws_id=...and coordGET /v1/api/workstreams/{ws_id}/eventsSSE handlers now share one body viamake_events_handler(cfg). Per-kind divergence captured by a newevents_replay: EventsReplay | Nonecfg field — a Protocol- typed callback yielding the kind-specific initial replay payload that the lifted body iterates and sends asdata:lines before starting the live event loop. Interactive's_interactive_events_replayyields the pre-lift sequence (connected+status+history+pending_approval- cached intent verdicts +
pending_plan_review); coord's_coord_events_replayyields justpending_approval+pending_plan_review(matches pre-lift coord behaviour).
The legacy interactive query-keyed URL is preserved via a new
make_legacy_query_keyed_adapterhelper (sister tomake_legacy_body_keyed_adapterfrom earlier lifts) — it readsws_idfrom the query string and splices intorequest.path_paramsbefore delegating to the lifted body.GET /v1/api/events?ws_id=...continues to work for any 1.x SDK consumer.Two convergence wins:
- Coord gains SSE connect/disconnect metrics. Pre-lift
coord didn't record per-stream metrics; the lifted body
always calls
metrics.record_sse_connect()/...disconnect(), giving the cluster dashboard the same per-stream observability interactive's had since 1.0. - Both kinds now check
request.is_disconnected()AND thews_closedevent to terminate. Pre-lift interactive relied solely onws_closed(which never fires if the client just goes away without closing the workstream); pre-lift coord relied solely onis_disconnected. The lifted body uses both — whichever fires first wins.
One observable shape change for coord callers: the lifted body returns 409
"session has no UI"whenws.uiis missing (placeholder / build-failed UI), matching pre-lift coord. Pre-lift interactive returned 404 in this case; the lift converges on 409 across kinds because the workstream EXISTS (404 would imply it doesn't).Item #2 from § Post-P3 reckoning split out of this lift during scoping (rich
ws_statepayload parity for coord — lifting coord'sConsoleCoordinatorUIto broadcasttokens + context_ratio + activity + contentlikeWebUI._broadcast_statedoes). The body lift touchessession_routes.py+server.py+console/server.py; the rich-payload work touchescoordinator_ui.py+collector.py+session_ui_base.py(different files, different reviewer concern). Tracked as standalone follow-upfeat/coord-rich-ws-state-payload.Two /review fixes folded into the same commit:
- Restored interactive's dedicated SSE thread pool. The
initial draft of
make_events_handlerusedasyncio.to_thread(default executor, capped atmin(32, cpu_count + 4)) for the per-connectionclient_queue.getblocking wait. Pre-lift interactive used a dedicated 200-threadsse_executor(created in the lifespan withthread_name_prefix="sse") precisely to avoid this — under high concurrent SSE counts the default pool starves and SSE polling contends with every otherasyncio.to_threadcaller in the process (storage, router, audit). Restored isolation via a newsse_executor_lookup: SseExecutorLookup | Nonecfg field; interactive returnsrequest.app.state.sse_executor, coord wiresNoneand falls through to the default executor. - Restored 5s queue.get poll (was shortened to 1s in the
initial draft). The 5x wakeup-rate bump compounded the thread-
pool starvation; the
request.is_disconnected()probe between polls already covers cancel-detection latency the timeout would otherwise gate. - Replay phase streams events directly from the generator
instead of pre-building into a list. The initial draft
materialised the entire kind-specific replay payload
(
connected+status+history+ pending prompts) into a list before constructing theEventSourceResponse, delaying time-to-first-byte until the heaviest replay event (_build_historyfor long-running interactive workstreams) finished serialising AND letting the per-UI listener queue accumulate over its 500-slot cap on a chatty mid-generation workstream. The lifted body now iteratescfg.events_replayinside the async generator so each event ships as soon as the callback yields it; the existing observational-failure swallow semantics are preserved by wrapping the iteration in the same try/except.
- cached intent verdicts +
-
createverb body lifted across both kinds ([Stage 2 Verb Lift —create]). The interactivePOST /v1/api/workstreams/newand coordPOST /v1/api/workstreams/newhandlers now share one body viamake_create_handler(cfg, *, audit_emit=None). Per-kind divergence captured by five newSessionEndpointConfigfields:create_supports_attachments: bool— multipart body parsing- attachment validation+save+rollback. Both kinds wire
True.
- attachment validation+save+rollback. Both kinds wire
create_supports_user_id_override: bool— trusted-source bodyuser_idoverride (interactiveTruefor console- proxied creates; coordFalse).create_validate_request: CreateRequestValidator | None— per-kind pre-create gates (interactive: ws_id format, kind, parent ownership, attachments+resume_ws combo; coord: 401-on- empty-uid).create_build_kwargs: CreateKwargsBuilder | None— per-kind kwargs dict formgr.create.create_post_install: CreatePostInstall | None— per-kind tail end (interactive: WebUI auto_approve + watch_runner +ws_createdglobal broadcast + atomic resume + skill session config + notify_targets + routing override + initial-message worker thread; coord:coord_adapter.sendfor the optional initial_message).
The pure helper
_validate_and_save_uploaded_fileslifted fromturnstone.servertoturnstone.core.attachmentsasvalidate_and_save_uploaded_filesso both processes can call the same kind-agnostic implementation.§ Post-P3 reckoning item #1 done — coord gains create-time attachments. Pre-lift
coordinator_createaccepted JSON only and ignored uploads; the lifted body parsesmultipart/form-dataon coord and saves attachments through the kind-agnostic storage layer.CoordinatorAdapter.sendgained optionalattachments+send_idkwargs so when a create request carries bothinitial_messageand uploads, the attachments are reserved onto the dispatched first turn — the worker'sChatSession.send(..., send_id=...)consumes them on dequeue exactly the way interactive's create-with-attachments worker thread does. Thesend_idreservation token soft-locks the rows, and the adapter's failure path unreserves so a worker crash returns them to pending. The pure helper_reserve_and_resolve_attachmentslifted fromserver.pytoturnstone.core.attachmentsasreserve_and_resolve_attachmentsso both kinds call one kind-agnostic implementation.Note on broadcast timing: coord's
mgr.createfiresemit_created(cluster collector fan-out) BEFORE the lifted body runs attachment validation. If validation fails on coord and the rollback (mgr.close→emit_closed) fires, the cluster events stream sees a phantom create→close pair. Cluster consumers handle this gracefully (same shape as any quick-create-close); decouplingemit_createdfrommgr.createwould be a bigger refactor that doesn't belong in the verb lift. Interactive's broadcast (gq.put_nowait("ws_created")) is held until after attachment validation by the post-install callback, so interactive never sees the phantom pair.Five observable behaviour changes on the create response:
- Both kinds converge on 200 OK. Pre-lift interactive
returned 200 (default JSONResponse status); pre-lift coord
returned 201. Picked 200 over 201 for response-shape parity
with every other shared verb at the cost of REST-strict
correctness — a one-time release note rather than ongoing
client churn (the rest of the v1 SDK already uses
response.okperfeedback_test_frontend_locally.md). SDK consumers that branched onstatus == 201for coord must switch toresponse.ok. - Always-include response shape. Pre-lift interactive
returned
{ws_id, name, resumed, message_count, attachment_ids}(5 fields); pre-lift coord returned{ws_id, name}(2). The lifted body always returns the full shape, withresumed=False/message_count=0/attachment_ids=[]on kinds whose post-install doesn't populate them. Coord callers will see the parity fields appear with default values. - Both kinds converge on the manager-at-capacity 429
semantic. Pre-lift interactive translated
mgr.create'sRuntimeErrorto 400; coord already translated to 429. The documented contract onSessionManager.createis "raises RuntimeError when the manager is at capacity" — 429 (rate- limit / try-later) is the correct shape. - Both kinds converge on the factory-misconfig 503 semantic.
Pre-lift interactive let
ValueErrorpropagate as 500 with a stack trace; coord already translated to 503 with the factory's remediation text. Operators get the actionable message instead of the trace. - Both kinds get a correlation_id'd 500 on unexpected
mgr.createfailure. Pre-lift interactive let unexpected exceptions propagate as 500 with a stack trace (potential information leak via frame names / file paths); coord already returned a correlation_id'd 500 with the message redacted. The lifted body adopts coord's safer pattern on both kinds.
Two coord-specific parity gains:
- Coord rejects disabled skills. Pre-lift
coordinator_createsilently allowed disabled skills to flow through tomgr.create— the row would create with a skill the operator had marked inert, surprising both the operator and the next user. The lifted body returns 400 "Skill not found or disabled" matching interactive's behaviour. - Coord audit-emit failures no longer 500. Pre-lift
coordinator_createalready swallowed; pre-lift interactive let the failure propagate as 500. The lifted body wrapsaudit_emitin try/except +warninglog, returning the successful 200 to the caller. Mirrors the close / cancel / open / events lift contracts.
No legacy adapter is needed for create — both kinds already mounted
POST {prefix}/newpre-lift; the lifted handler slots in at the same path on each kind.Three /review fixes folded into the same commit:
- Pre-lift's 400 on malformed
notify_targetspreserved. The initial draft surfacednotify_targetsvalidation errors from inside the interactivepost_installcallback, which the factory had no return-the-400 channel for — the only signal was toraise, which the factory's generic exception handler turned into a redacted 500. Worse, by the timepost_installran the workstream was fully built (audit row written,ws_createdbroadcast emitted), so a malformed-input request surfaced as "create failed" with the workstream actually live. Fixed by moving thenotify_targetsvalidation into :func:_interactive_create_validate_request(the pre-create gate), which returns the 400 beforemgr.createruns and keeps storage clean. New regression test:test_create_lift_400s_on_malformed_notify_targets. - Skill-lookup storage failure now correlation_id'd. The
initial draft swallowed
get_skill_by_nameexceptions intoskill_data = Noneand returned a 400 "Skill not found or disabled" — masking storage outages as user-input misses and making operator triage of skill-related reports impossible. The lifted body now lets the storage exception propagate to the same correlation_id'd 500 path thatmgr.createfailures use; the skill-lookup + version count +mgr.createall live inside onetry / exceptso storage outages anywhere in the create-prelude get the redacted-message-with-correlation-id treatment instead of a stack-traced 500 leak. - Whitespace-only
skillfield treated as empty. The initial draft tookbody.get("skill") or ""literally — a payload with"skill": " "would have hitget_skill_by_name(" ")and 400'd as "Skill not found". Pre-lift coord stripped via(body.get("skill") or "").strip() or None; the lifted body now strips for both kinds (interactive never received whitespace-only skills from the web UI but the convergence is the safer default). - Canonical skill name persisted to
mgr.create. The initial draft's_interactive_create_build_kwargs/_coord_create_build_kwargspassed the rawbody["skill"]through, so a whitespace-padded request would have persisted" my-skill "even though the lookup was done on the stripped name. The build_kwargs callbacks now threadskill_data["name"](the resolved row's canonical name) so the persistedWorkstream.skillmatches the row that was actually applied — keeps later session-sideskilllookups working regardless of how dirty the inbound payload was.
-
Coordinator scratchpad tool renamed:
task_list→tasks. The tool name on the LLM-facing schema, the audit event name (task_list.update→tasks.update), the SSEtool_resultevent name (the coord-tree UI keysev.name === "tasks"for /tasks-refetch debounce), and the log tag (task_list.corrupt_envelope→tasks.corrupt_envelope) all switch together. Operators with audit dashboards / SIEM filters / log greps that pinned the old prefix should update; the rename is observable on the wire, not just internal. Internal Python surface follows:CoordinatorClient.task_list_*→tasks_*,ChatSession._prepare_task_list/_exec_task_list→_prepare_tasks/_exec_tasks,_TASK_LIST_MAX→_TASKS_MAX. The previous name compounded the bare wordtask(which collides with chat-template channels on local models — the same reasontask_agentcarries the suffix); the plural form sidesteps the collision and is more accurate, since the tool acts on the whole list rather than a single task.
Security
-
Coord attachment endpoints are now kind-strict ([Stage 2 P1.5]). The coord
attachment_owner_resolverresolves through the in-memorycoord_mgronly — it does NOT fall back to storage. Without this, anadmin.coordinator-scoped caller could pass an interactive workstream ws_id to the new coord attachment endpoints; the genericget_workstream_ownerstorage call (kind-agnostic) would resolve cleanly and grant cross-kind read / write access to interactive attachments. The kind-strict resolver returns 404 for any ws_id not currently held by the coord manager, closing the cross-kind path. Persisted-but-not-loaded coordinators must beopened before their attachment endpoints respond. Caught by /review pre-merge; no exploit observed. -
Workstream state writes are now buffered through
StateWriter.SessionManager.set_stateno longer holdsws._lockacross a synchronous PostgresUPDATEfor non-terminal transitions; instead aStateWriter(constructed at app startup, started / shutdown by the lifespan) coalesces transient transitions per ws_id and flushes every ~1s. Observable behavior change: transient state (thinking/running/idle/attention) shows up in storage up to ~1s late; SSE consumers see it immediately via the adapter'semit_state. TerminalERRORtransitions andclose()write synchronously and remain durable on return. The bug-3 invariant — a closed row can't be resurrected by a buffered transient — is preserved byclose()callingstate_writer.discard(ws_id)(drops pending + waits for any in-flight flush) before its syncstate='closed'write.
[1.4.0]
User-visible additions: a full attachment system (images + text documents, including pre-creation uploads), a unified dashboard composer, a Slack channel adapter, per-call plan/task model selection with an admin UI, and provider capability passthrough.
This release introduces two forward-only schema migrations
(037_workstream_attachments, 038_workstream_attachments_reserved_at)
that the server applies automatically on first startup against an
existing 1.3.x database. Both are additive; no data loss. See
Database migrations below for details.
Added
- Workstream attachments — images (png/jpeg/gif/webp, 4 MiB cap) and
text documents (any
text/*MIME, allowlisted application MIMEs, or known text extensions; 512 KiB cap; UTF-8 enforced). Magic-byte image sniffing on upload; per-(ws, user) pending cap of 10. Three-state lifecycle (pending → reserved → consumed) with reservation tokens threaded through/v1/api/sendso queued multimodal turns can't lose files to overlapping sends. Provider-side translation: Anthropic emits native document blocks; OpenAI Chat Completions inlines them as escaped<document>text blocks; Responses API emitsinput_textwith the same wrapper. (#356) - Attachments at workstream-creation time —
POST /v1/api/workstreams/newacceptsmultipart/form-data(onemetaJSON field plus 0..Nfileparts). Files are validated and reserved onto the first turn before the dispatch worker fires; failure rolls back the fresh workstream so no orphan rows leak. Web UI (new-workstream modal + dashboard composer), Python SDK, and TypeScript SDK all gained attachment support. Cluster routing (/v1/api/route/workstreams/{ws_id}/attachments) extended to forward multipart bodies + preserve upstream headers (CSP, Content-Disposition). SDKs auto-generatews_idclient-side so cluster-routed callers can bind the body to the owning node before it lands. (#362) - Slack channel adapter (Socket Mode) — mirrors the Discord adapter:
per-user channel sessions via configurable slash command, DM routing
without slash command, SSE event consumption, tool approval buttons
with per-user owner enforcement, plan-review approve / request-changes
modal, notification reply routing back into the workstream, and
session recovery after restart via persisted recoverable route keys
(the bot re-subscribes to existing Slack-routed workstreams when it
comes back). Install with
pip install 'turnstone[slack]'. (#355) - Console admin UX support for Slack — channel-link modal offers
Slack alongside Discord; skill notify-on-complete forms expose a
per-row channel-type dropdown (and no longer hardcode
discord); per-platform.scope-discord/.scope-slackbadge classes with theme-aware tokens (--discord/--slack) so light theme passes WCAG AA. (#365) - Per-call plan/task model selection —
plan_modelandtask_modelare now distinct from the conversation model and from each other, with configurable reasoning effort per agent. Three layers:- Backend split (
#54dd557) —ModelRegistrygainsplan_model,task_model,plan_effort,task_effort; per-kind overrides win over the legacyagent_model, which still works as the single-knob fallback.resolve_agent_alias(kind)andresolve_agent_effort(kind)centralise resolution. Loader validates effort against{none, minimal, low, medium, high, xhigh, max}with warn+drop on typos. - Runtime configurability (
#360) —ConfigStoreadmin tab in the console UI lets operators switch alias and reasoning effort per agent without restarting.INHERIT_EMPTY_LABEL_KEYSshows(inherit)for empty effort selections — distinct from the literalnonechoice which actually disables reasoning. Routing overrides apply on/v1/api/_internal/config-reload(admin saves), andmodel-reloadshort-circuits when nothing changed so no in-flight clients churn. - Per-call override (
#361) — the calling LLM can passmodel="<alias>"toplan_agentortask_agentto override the operator-configured per-kind model for that one invocation. Tool descriptions list the live registered aliases (refreshed when the operator hits "sync to nodes"), so the LLM always sees current options. Bad aliases return a corrective error dict listing the available choices. No whitelist — cost control is intentionally ceded to the model. Plan-retry path reuses the alias so coaching reflects real model behaviour. (#360, #361)
- Backend split (
- Provider capability passthrough — resolved per-model capabilities
(vision, reasoning, native web search, thinking_mode, token_param,
etc.) flow through to provider clients via a new
capabilitiesparameter oncreate_streaming/create_completion, so feature gating no longer relies on string matching and admin-UI / config.toml overrides actually reach the provider. Defensive shallow-copy in_finalize_extra_bodyso callers reusing the same dict across models are safe; deep-merge ofchat_template_kwargsso operators can extend instead of silently overwriting. (#352) - Server compatibility layer for local model servers — vLLM and
llama.cpp profiles suggest the right thinking mode and per-server
workarounds (
skip_special_tokensfor vLLM,reasoning_formatfor llama.cpp) during model detection. Admin UI gains structured fields for server type, thinking mode, and extra body params, hidden for non-local providers (openai/anthropic/google). Newthinking_paramtext field surfaces the alias name (defaultenable_thinking; Granite/DeepSeek usethinking). Verified end-to-end against real vLLM (Gemma 4 31B) and llama.cpp (Gemma 4 E4B) servers. (#352) - Claude Opus 4.7 support —
claude-opus-4-7capability entry (1M ctx, 128K output, adaptive thinking,supports_temperature=False,thinking_display=summarized). NewModelCapabilities.thinking_displayfield — Opus 4.7 omits thinking by default but always sends summarized blocks back through the provider boundary. Addsxhigheffort level to the global mapping and to Opus 4.7'seffort_levels; admin-console skill-template dropdowns gainedxhighandmaxoptions. Reasoning effort label capitalization aligned across all console dropdowns. (#357 — also in 1.3.1) - Dashboard composer refactor — unified single-flow create from the
per-node dashboard. Multi-line textarea + collapsible Options panel
(model / judge / skill) + paperclip + drag-drop / paste-image + chip
strip. Submit-button label dynamically toggles between
Create(empty) andSend(text or attachments staged); Enter and click both go through the samedashboardSubmit(). Replaces the inconsistent prior split where Enter created+sent raw and the button opened a separate modal. Options panel state persists inlocalStorage; active non-default selections render as an inline summary chip beside the Options button; drag-over shows an explicit "Drop to attach" overlay. The tab-bar+new-workstream modal also gained a paperclip- chip strip + first-message field so the same flow is reachable from both entry points. (#362, #366)
- Workstream attachments — orphan reservation sweep — periodic
background sweep clears
reserved_for_msg_idon rows whosereserved_atexceeds a 1-hour threshold, self-healing reservations leaked by process crashes between reserve and consume. Backed by a partial index on(reserved_at) WHERE reserved_at IS NOT NULLso the scan stays cheap as the consumed-history grows. Threshold tracks reservation age, not upload age, so a long-pending fresh send can't be racially unreserved. (#363) SendResponseextended —attached_ids,dropped_attachment_ids,priority,msg_idfields exposed in Pydantic + TypeScript SDKs so attachment-aware clients can detect partial reservations and dequeue queued messages. (#365)
Changed
plan_modelandtask_modelnow split from the conversation model and from each other — operators who rely on a single model for all three should set bothplan_modelandtask_modelexplicitly in their config; otherwise both default to the conversation model so behaviour is unchanged. (#54dd557)- Channel notify-on-complete
channel_typeis no longer hardcoded in the admin UI — operators creating notify targets through the skill admin form previously gotchannel_type: "discord"regardless of what they wanted. Existing skill JSON values are unaffected; only newly created targets through the form differ. (#365) - Slack adapter approval previews — capped at 600 chars per item
with a 2700-char total budget so multi-tool approval batches never
exceed Slack's 3000-char
section.textlimit. Truncated batches show a…and N more (preview truncated)suffix. (#365) - PostgreSQL deployment image swapped from
bitnami/pgbouncertoedoburu/pgbouncerto track upstream releases and reduce image size. Environment variables remapped to the edoburu naming, ports updated to match documented expectations, and the Kubernetes Helm Chart link in the deployment docs now points at the same container. Review your helm values if you depend onbitnami-specific environment variable conventions. (#353)
Fixed
plan_resolvedSSE broadcast — when one client resolved a plan approval, other clients viewing the same workstream now have the approval card dismissed in sync. (#87a9af1)- Slack notification reply routing — one notification reply
previously pinned every later assistant response for that workstream
to the notification thread until the bot restarted. Reply-route
override now clears on
StreamEndEvent. (#365) - Slack plan-review mrkdwn fence — plan content containing triple
backticks (very common — plans often quote code) no longer breaks the
surrounding fence and lets later content render as live markup. The
shared
_sanitize_slack_previewhelper splices a zero-width space inside any`` sequence while keeping single backticks readable. (#365) - Slack-routed workstreams now load the chat-specific system prompt
via
client_type="chat", matching Discord. (#365) /v1/api/workstreams/newno longer emits a phantomws_created/ws_closedSSE pair when attachment validation rejects a multipart create. Validation runs before the broadcast so failed creates are silent on dashboards. (#362)- Multipart Content-Type boundary preservation in console routing
proxy —
boundary=parameter is case-sensitive and was being lowercased before forwarding to the upstream node, breaking parsing for clients that used mixed-case boundaries (most browsers). (#362) - Local-theme contrast for new badge colors —
.scope-discordand.scope-slackfirst shipped with raw hex that failed WCAG AA on light theme (1.8:1 / 2.4:1). Theme-aware--discord/--slacktokens with proper light variants now pass. (#365) - Cross-user attachment fetch hardening —
get_attachment_contentnow scopes the row byuser_idin addition tows_id, so an unowned workstream can't be a vector for cross-user blob fetches via attachment-id guessing. (#356) - Attachment-list DoS guard —
/v1/api/sendrejectsattachment_idslists longer than the per-(ws, user) pending cap with a 400, preventing hostile clients from blowing up the storageIN (...)clause. (#356) - Bounded LRU for upload locks — the per-(ws, user) attachment upload-lock map now evicts the oldest unlocked entries past a soft cap, so the in-process map can't grow unbounded on long-running nodes. (#356)
- 3.12 CI deadlock on attachment uploads — the upload-lock was
initially an
asyncio.Lock, but Starlette'sTestClientruns each request on a fresh anyio task / event loop, so the cached lock's_waitersbound to the first loop and a later request would block on a Future from a closed loop (silent deadlock). Switched tothreading.Lock— loop-agnostic, and the critical section is one COUNT + one INSERT. Same root cause is reproducible against any Starlette TestClient harness on Python ≥ 3.10; 3.12 surfaces it more often. Production users on a single event loop weren't affected, but the test environment was. (#356)
Security
- Slack approval per-user authentication — only the session owner can click Approve/Deny on a Slack tool-approval card. Without this, any channel member with view access could approve dangerous tool calls initiated by someone else. (#355)
- Attachment ownership masking — cross-user/cross-workstream attachment ID lookups return 404 (not 403) so non-owners can't enumerate workstream existence by response code. (#356)
- Bumped Debian base image; remaining unfixable
jqCVEs are documented and exception-listed. (#aaea4d3)
Database migrations
037_workstream_attachments— newworkstream_attachmentstable with the lifecycle columns described above. Indexes for ws_id, pending lookups, message linkage, and reservation scoping.038_workstream_attachments_reserved_at— addsreserved_atcolumn for the orphan-sweep staleness signal, plus a partial index onreserved_at IS NOT NULLso the periodic scan is cheap.
Both migrations are additive and idempotent, and the server applies
them automatically on first startup against an existing 1.3.x database.
No manual alembic upgrade step is required — though running it
manually beforehand (e.g. as part of a phased deploy) remains safe.
SDK
Python + TypeScript clients gained:
AttachmentUploadtypeupload_attachment(ws_id, filename, data, mime_type=None)list_attachments(ws_id)get_attachment_content(ws_id, attachment_id) → bytes / Blobdelete_attachment(ws_id, attachment_id)send(message, ws_id, attachment_ids=...)(extended)create_workstream(..., attachments=[...])— multipart variant with client-sidews_idgeneration for cluster-routed callers- Console SDK:
route_create_workstream(attachments=...),route_upload_attachment,route_list_attachments,route_get_attachment_content,route_delete_attachment - Refusal of
attachments + target_nodecombination at the SDK boundary (the multipart routing layer doesn't honortarget_node, so silently picking the wrong node is now an explicit error) PlanResolvedEventSSE event with type guard, dispatched when one client (e.g. mobile) resolves a plan so other connected clients can dismiss their plan-approval modal in sync. Available in both the Python and TypeScript SDKs. (#87a9af1)
Operational
- CI vendor-asset auto-download covers
hls.js— thevendor-js.ymlworkflow previously only iterated katex/hljs/mermaid, so Renovate bumps forhls.jsfailed the wheel-completeness check and required manual file downloads. Detection loop now includeshls, so future Renovate bumps are merge-ready without intervention. (#354)
Contributors
Thanks to the people who made this release happen — especially the external contributors who picked up substantial pieces of work:
- @daoxley — designed and shipped the Slack channel adapter (Socket Mode bot, per-user sessions, approvals, plan-review, notification routing). Major new feature surface in #355.
- @pizzaandcheese — replaced the deprecated bitnami pgbouncer image with the edoburu image, remapped environment variables, ports, and helm chart references. Operationally important for anyone running our reference Postgres deployment (#353).
- Renovate kept dependencies and the JS vendor tree current via several automated bumps.
If you're interested in contributing, channel-attachment ingest from Discord + Slack is the headline 1.4.1 feature and a solid place to start — see the open issues on GitHub or open one to scope a piece.
[1.3.1]
Added
- Backport: Claude Opus 4.7 support (provider capabilities, tokenizer, adaptive thinking). (#357)