* fix(server): trusted-team workstream visibility on listing endpoints
The per-user filter on /v1/api/workstreams, /v1/api/dashboard, and
/v1/api/workstreams/saved (PR #375's _visible_workstreams helper) was
written for a multi-tenant SaaS threat model that doesn't match how
turnstone gets deployed. In a self-hosted, trusted-team install the
filter created friction without preventing the relevant threats — and
hid the auto-created name="default" startup workstream from every
web user, leaving fresh installs staring at a blank dashboard.
Listing endpoints now return the cluster-wide set to any authenticated
caller. Per-workstream MUTATIONS (/send, /close, /open, /title,
/delete, /refresh-title) keep their independent ownership checks — the
cross-tenant guards from PR #375 stay in force on those handlers (see
TestCrossTenant{Delete,Approve,Close,Title,Open}). Listing only
exposes metadata (name, state, kind, message_count); message history
still requires the per-workstream gate on /history.
Resuming a saved workstream still goes through /open's owner check, so
the metadata-leak surface ends at "you can see workstream X exists" —
not at any actionable cross-user capability.
The console collector's service-scope is now load-bearing only for the
SSE event stream gate (/v1/api/events/global); kept anyway as belt-
and-braces.
If turnstone is ever deployed as a true multi-tenant SaaS, the right
boundary is a real ``tenant_id`` column with row-level filtering at
the storage layer, not the empty-user_id heuristic this used to apply.
Tests updated to assert the new contract: listing returns all owners;
mutation gates unchanged.
* fix(server): repair test mocks + tighten docstrings on listing endpoints
- tests/test_auth.py: TestServerAuth + TestServerLogin mocks now set
kind / parent_ws_id / user_id explicitly so /v1/api/workstreams JSON-
serializes them. Bare MagicMock attributes return another MagicMock
that fails json.dumps and surfaces as 500.
- turnstone/server.py: list_saved_workstreams docstring corrected to
describe what the endpoint actually returns (summary metadata, not
history) and to spell out that ownerless persisted rows are claimable
by any authenticated caller via /open — consistent with the trusted-
team model the listing endpoints assume. Same callout added next
to the open_workstream ownership-gate block. Comments throughout
rewritten to be timeless (no "previously" / PR-number references).
- tests/test_server_authz.py: TestSaved... docstring matches the actual
/open behavior for orphan rows (claimable by any authenticated
caller, not a separate admin path).
* fix(chat): collapse phantom whitespace + tighten paragraph rhythm in markdown body
The assistant chat body was rendering 30-50px gaps between every
section. Two compounding causes:
1. ``.ts-msg-body`` had ``white-space: pre-wrap`` on the markdown
container. The custom regex-based markdown converter
(renderer.js) leaves ``\n`` text nodes between block siblings —
pre-wrap rendered every one of those as visible vertical space,
stacking ~14-16px between every heading/paragraph/katex-display.
2. No ``.ts-msg-body p`` margin override, so paragraphs fell back to
browser-default 1em top + 1em bottom (~28px stacked between any
two paragraphs). Headings already had a tight ``8px 0 4px`` rule;
paragraphs were the outlier.
Switched the body to ``white-space: normal`` and added a
``.ts-msg-body p { margin: 6px 0 }`` rule that matches the heading /
list / blockquote rhythm. Mirrored the paragraph rule on the design-
v1 ``.msg-body`` selector so both legacy and v1 surfaces stay in sync.
``<pre>`` blocks have ``white-space: pre`` built in so fenced code
still preserves formatting. Mid-stream partial fences (before the
closing ``\`\`\`` arrives) render as collapsed text for one frame and
then snap back when the next render tick wraps them in ``<pre>`` —
acceptable trade vs. the persistent gap regression.
User-typed messages render through ``.msg-user-text`` (a separate
DOM path), so this only affects assistant markdown output.
* fix(chat): preserve inline <code> whitespace under white-space: normal body
The body's ``white-space: normal`` (which collapses phantom inter-block
``\n`` text nodes from the markdown converter) inherits to inline
``<code>`` and silently collapses multiple spaces inside backtick
spans. ``<pre>`` blocks rely on the user-agent ``pre { white-space:
pre }`` rule and are unaffected; only bare inline code needs an
explicit override.
Adds ``white-space: pre-wrap`` to ``.ts-msg-body code`` (chat.css) and
the design-v1 ``.msg-body code`` selector so backtick-wrapped code
spans render verbatim while still wrapping on long lines.
Addresses Copilot review feedback on PR #401.
* feat(coord): saved coordinators surface + shared session-card primitives
The console home view now lists explicitly-closed coordinators in a
"Saved Coordinators" card grid below the active list. Click a card →
POST /v1/api/coordinator/{ws_id}/open then navigate; capacity issues
surface as a toast instead of a broken detail page. Card click is
de-duped by an `is-busy` class so rapid double-clicks don't fire
parallel resurrects.
GET /v1/api/coordinator/saved is the new backend endpoint (mirrors the
interactive list_saved_workstreams shape). Filters at the SQL layer
to state='closed' via a new optional `state` parameter on
list_workstreams_with_history (added to the protocol + both backends);
also drops any rows currently loaded into coord_mgr as defence in
depth. The blocking storage call + the lock-acquiring list_all are
offloaded via asyncio.to_thread to match coordinator_create's pattern.
CoordinatorManager._open_impl now allows resurrect of state='closed'
rows (deleted is still a tombstone). The DB state-flip on resurrect
that the first cut had is gone — it raced concurrent close()s and the
next set_state() call syncs the DB naturally; the saved list filters
already keep a still-loaded coordinator from appearing as a saved
card even when its on-disk state lags.
Frontend dedup that paid for the saved surface ships in the same diff:
- shared_static/cards.css: lifted from ui/static/style.css so both
surfaces share the basic card primitive (delete-mode rules stay
interactive-only until coordinator gets the same UX)
- shared_static/cards.js: new renderSessionCard(sess, opts) helper
used by both renderSavedWorkstreams (interactive) and
renderSavedCoordinators (console)
- shared_static/utils.js: formatRelativeTime moved here from
ui/static/app.js
Coordinator landing visual fixes folded in:
- .home-section-title now uses var(--accent) so the COORDINATORS
heading reads as a peer of the NODES heading
- .home-panel dropped its bg/border/padding so the composer is no
longer double-framed (matching the dashboard-composer feel)
- "Active coordinators" → "Saved Coordinators" rename + "Coordinators"
on the active list
ws_closed SSE handler now gates on the closed ws's kind so interactive
closes don't spam /v1/api/coordinator/saved on busy clusters.
loadSavedCoordinators in-flight de-dup coalesces close-event bursts to
one fetch instead of N.
Tests cover: caller-scoping, admin sees-all, blank-uid fail-closed,
loaded-coordinator filtering, state filter (idle rows excluded), plus
the manager-level open-resurrect / open-refuses-deleted contracts.
Closes the bug-{1,2,3}, perf-{1,2,3,4}, sec-{1,2}, q-{1,2,3,4,5,6,7}
findings from the prior multi-stage review.
* fix(design): restore amber accent on the v1 design system
The Claude Design handoff swapped the accent hue to teal (h=182).
Walking back to amber (h=75) — turnstone's original brand colour.
Lightness + chroma bumped slightly (0.62→0.7, 0.10→0.13) so the
restored gold matches the visual weight of the legacy #e5a042 token.
Hue map header comment updated to record what happened so the next
person doesn't repeat the swap. Only surfaces with data-design="v1"
on <html> pick this up — currently just turnstone-server's webui.
* chore: gitignore design_ideas/ and .claude/ dev directories
design_ideas/ holds personal Claude Design handoff scratch + reference
HTML; .claude/ holds per-user Claude Code state (worktrees, settings,
plugin caches). Neither belongs in version control.
* fix(coord): address PR #399 review nits
- tests/test_coordinator_endpoints.py: split `assert mgr.close(ws.id)`
in `_seed_closed_coord_with_history` so the close call always runs
even under `python -O` (asserts stripped). Same fix in
test_coordinator_manager.py's `test_open_refuses_deleted_coordinator`
for the open() and open_admin() calls.
- shared_static/cards.css: `.card-wsid` now reads `var(--font-mono, "IBM
Plex Mono", monospace)` so design-v1 surfaces pick up the JetBrains
Mono token while console (still pre-v1) keeps the literal fallback.
* feat(auth): inline refresh response + sessionStorage rehydrate hardening
The proactive refresh path now consumes the /refresh response body
inline (permissions + exp), eliminating the chained /whoami round-trip
and the brief stale-sessionStorage window after refresh succeeds but
before whoami completes.
Adds AbortController + _loggedOut guards to the whoami fetch so a
logout fired mid-flight cannot re-populate sessionStorage after it
clears. A non-OK whoami on tab restore now explicitly clears
sessionStorage instead of silently leaving stale cosmetic permissions
(server-side identity gone → UI gating reflects it on next render).
Surfaces window.permissionsReady (one-shot promise) so permission-
gated UI can await the initial whoami's completion instead of guessing
a setTimeout duration.
Tests cover the new refresh response shape, the existing leeway path,
the storage-failure fallback, and the no-perms 403 path.
Closes the bug-3 / perf-4 / sec-1 / q-6 findings from the multi-stage
review of the prior uncommitted change set.
* fix(auth): guard whoami superseding race in _scheduleRefreshFromWhoami
_scheduleRefreshFromWhoami is invoked from several entry points
(initial page load, _onSuccess, BroadcastChannel "login"/"refresh",
_tryRefresh fallback). Two firing in quick succession could let an
older slow whoami land after a newer one and clobber its effects —
clearing permissions right after a successful login, or rescheduling
the refresh timer off stale exp.
Now aborts any prior _whoamiAbort before starting a new request and
guards the .then's _storePermissions / _scheduleRefreshAt with a
`_whoamiAbort === ctrl` check so a late arrival from a superseded
call is fully neutralised.
Addresses Copilot review feedback on PR #398.
* feat(providers): add gpt-5.5 and gpt-5.5-pro capability entries
OpenAI announced gpt-5.5 on 2026-04-23 (ChatGPT/Codex first, API
"very soon"). Mirror the gpt-5.4 / 5.4-pro capability shape: 1M
context, native tool search, vision, xhigh effort; pro is
always-reasoning with no temperature and medium/high/xhigh only.
No provider-logic changes needed — OpenAI announced no API-surface
changes vs 5.4. Cache retention already covers 5.5 via the existing
startswith("gpt-5") prefix rule.
* test(providers): cover gpt-5.4-pro and gpt-5.5-pro in cache retention test
Pro variants share the same gpt-5 prefix and should keep 24h
retention; explicit coverage guards against regressions if the
prefix rule narrows in the future.
Remove the weekly Trivy scan job and the .trivyignore exclusion file.
The scanner has been flagging base-image CVEs that require no action
on our part (upstream-only fixes) and has provided no actionable
signal, while breaking CI on an ongoing basis.
* feat(auth): cookie refresh endpoint, JWT leeway, coord-token observability
Three robustness wins around the auth/JWT layer.
1. POST /v1/api/auth/refresh — handle_auth_refresh in core/auth.py,
wired in both console/server.py and server.py. Sliding-window
re-mint of the auth cookie. Re-resolves the user's permissions
from storage so a role change propagates within one refresh cycle
instead of persisting until the original cookie's natural expiry.
Returns the same JSON shape as /api/auth/login plus a fresh
Set-Cookie header. Refuses to extend a session for a deleted /
role-stripped user (403).
Resolves the user-visible "401 after browser tab open >24h"
symptom: previously the only refresh path was a full re-login,
now a single POST extends the session.
2. validate_jwt now passes leeway=30 to PyJWT. Absorbs minor
clock skew between hosts (multi-replica console deployments) and
between mint-time and validate-time within the same process.
Standard tolerance for short-lived tokens.
3. CoordinatorTokenManager._mint logs at debug. Mirrors the pattern
in ServiceTokenManager._mint (auth.py). Premature-401 diagnostics
would have been an order of magnitude faster with this in place
the first time around.
Frontend (shared_static/auth.js):
- _scheduleRefreshFromWhoami() reads the JWT exp surfaced via /whoami
and sets a setTimeout at 90% of remaining cookie life to call
/refresh. Floor 30s, ceiling 24h. Fires on initial page load
(silent if not authenticated) and after every successful login.
- _tryRefresh() de-dupes concurrent callers via a shared in-flight
promise — many parallel authFetch's hitting 401 at once still only
fire one /refresh.
- authFetch on-401 now attempts a single reactive refresh-then-retry
before falling through to the login overlay. Covers cases where
the proactive timer didn't fire (tab restored from disk-cache after
expiry, system clock jump, page first-load with stale cookie).
- BroadcastChannel "refresh" message keeps sibling tabs in sync so
they don't redundantly hit /refresh themselves.
- logout() cancels the proactive timer.
Tests:
- validate_jwt accepts 10s-expired tokens (within 30s leeway).
- validate_jwt rejects 60s-expired tokens (past leeway).
- /whoami includes exp claim with sane bounds.
- /refresh returns ok + Set-Cookie + the refreshed cookie keeps
working on subsequent authenticated requests.
- /refresh without a cookie returns 401.
Not addressed: the coordinator.session_jwt_ttl_seconds ceiling
(currently 1h) — that's a separate, preventative concern for very-
quiet long-running coordinators, orthogonal to the user-visible 401
this PR fixes. Can bump in a follow-up if it actually surfaces.
* fix(auth): address Copilot PR #395 feedback
Two real bugs caught by Copilot, both fixed.
1. Storage failure was indistinguishable from "user deleted" in
handle_auth_refresh. _load_user_permissions() swallows exceptions
and returns set(), so a transient DB hiccup looked like
"user has no permissions" and returned 403 — logging the user out.
Now calls storage.get_user_permissions() directly with try/except.
- Exception → log + fall through to in-token claims (refresh succeeds
with stale-but-valid permissions; better than fail-closed mid-
session for a hiccup).
- Empty set returned (no exception) → 403 (legitimate signal: user
deleted or role-stripped).
Tests:
- test_refresh_storage_failure_falls_back: storage raises → 200 +
in-token permissions.
- test_refresh_user_with_no_perms_403: storage returns empty → 403.
2. Logout race: a /refresh in flight when the user clicks Logout could
land AFTER /logout's clear-cookie response and re-set the cookie
from /refresh's Set-Cookie header, silently undoing the logout.
Fix in shared_static/auth.js:
- Add a _loggedOut latch + _refreshAbort AbortController.
- logout() sets _loggedOut = true synchronously and aborts any
in-flight /refresh BEFORE the /logout fetch fires.
- _tryRefresh() bails on its post-fetch effects (don't store perms,
don't reschedule, don't broadcast) when _loggedOut is set. The
stale Set-Cookie from /refresh is harmless because /logout's
response overwrites it on the way back.
- _onSuccess() (re-login) clears the latch so subsequent refreshes
work again.
Race window is small but real on slow networks / contested CPU.
* feat(design-system): DS phase 2 — opt server chat UI into v1 primitives
ui/static/index.html:
- data-design="v1" on <html> opts this view into design system tokens
and primitives scoped under the attribute selector.
- Link DS stylesheets after the legacy cascade: tokens + typography +
appbar (chrome) + panel / buttons / pills / message / field
(primitives). Legacy /shared/base.css, /shared/ui-base.css,
/shared/chat.css, and /static/style.css stay linked to handle
anything not yet migrated (rich markdown, tabs, dashboard, split
panes, approvals, modals).
- Header <div id="header"> picks up .appbar + .appbar-title +
.appbar-status + .appbar-spacer + .appbar-actions alongside the
legacy .ts-header classes. Theme-toggle gets .btn for DS pill shape
while keeping .header-btn for palette continuity.
ui/static/app.js:
- Chat message elements emit both legacy and DS class names so the
DS primitive picks up the message surface while legacy .ts-msg--*
rules keep view-specific markdown styling (tables, callouts, katex,
mermaid, hljs). Pairs:
ts-msg ts-msg--user → + msg user
ts-msg ts-msg--assistant → + msg assistant
ts-msg ts-msg--reasoning → + msg reasoning
ts-msg ts-msg--info → + msg info
ts-msg ts-msg--error → + msg error
ts-msg-body → + msg-body
- Approval blocks keep legacy-only styling — their shape is distinct
from the DS .msg primitive (the DS approval-dock pattern is a
fixed bottom dock, not inline-in-chat).
No backend or wire-format changes. SSE events, POST bodies, endpoint
URLs, ARIA attributes, and keyboard shortcuts all unchanged.
* feat(ui/static): DS-skin tool-call + approval + verdict internals
The outer .ts-msg.ts-approval--inline picked up DS .msg styling via
PR #2's dual-class approach, but the inner structure kept rendering
with legacy yellow/green/red colours and legacy chip shapes. Result:
a DS-accent-bordered card containing a mustard tool-name, a clunky
uppercase-yellow verdict chip, and a mismatched auto-approved pill.
Add [data-design="v1"]-scoped overrides that reskin the inner
vocabulary onto DS tokens:
.ts-approval-tool panel-over-panel-2 card with hair border
.tool-name accent (teal) for tool-kind identity
.tool-cmd / .tool-diff ink-2 text; diff-del/add/warn → err/ok/warn
.verdict-badge.verdict-* chip aesthetic matching DS k-badge —
low=ok-tinted, medium=warn-tinted,
high/critical=err-tinted, with a
3px left-border semantic stripe
.verdict-detail panel-2 callout with structured rows
.verdict-judge-spinner ts-pulse animation (reuses primitive)
.ts-approval-badge--* pill shape hugging max-content, matches
DS approve-button-family colour palette
(ok-text, err-text-mix)
.tool-output panel bg, hair border, accent stream
left-border, fade-gradient on collapse
.ts-verdict-glow--* soft ring on the corresponding action
button (approve=ok, deny=err, review=warn)
No JS changes; DOM shape unchanged. CSS-only reskin so approval flow,
tool streaming, and verdict expand/collapse behaviour all stay intact.
* fix(ui/static): consistent tool-card width + flat badge aesthetic
Two fixes to the DS-skinned tool-call rendering:
1. Tool-call cards were sizing to their content (short output →
narrow card, long output → full-width), producing a jagged column.
Force .ts-msg.ts-approval--inline to width: 100%; align-self:
stretch; box-sizing: border-box; so the chat column reads evenly.
2. The "approved" / "auto-approved" pill was styled as a button
(pilled shape, 1px bg-tinted border, 4x10 padding) which read
as clickable. Switched to a flat badge aesthetic matching the
.risk primitive: 3px-squared, 2x6 padding, 10px mono uppercase
on a --ok-soft / --err-soft tinted surface, no border. Reads
as a status tag, not a call-to-action.
* fix(ui/static): address Copilot PR #392 feedback
Copilot findings, all applied:
- Drop the legacy Outfit + IBM Plex Mono Google Fonts link. DS
typography.css @imports Inter + JetBrains Mono; loading both stacks
on opted-in pages wastes downloads and triggers FOIT/FOUT differences.
- Drop the .ts-header-title class on the <h1>. Its legacy rule forces
font-family: var(--font-display) (Outfit) which overrides the DS
appbar typography. .appbar-title alone is sufficient under v1.
- Override .ts-msg font-family under [data-design="v1"] when .msg is
also present (and not the .tool variant). Legacy .ts-msg forces
mono; DS user/assistant/reasoning/info/error messages should use
the UI font. .msg.tool keeps mono via the primitive's own rule.
- Replace the inline name.style.color = "var(--red)" in buildToolDiv
with a .tool-name--error class. Inline styles win over CSS rules
and broke the DS token mapping (legacy --red is not the DS --err).
- Correct the header comment in style.css for the approval-block
overrides. Prior comment claimed the outer wrapper picks up DS .msg
styling; it doesn't — the dual-class approach wasn't extended to
approval blocks. Updated comment to match actual DOM.
* feat(design-system): DS phase 3 — coordinator chat migration
Opt the per-session coordinator view into data-design="v1" and migrate
its rendering to the DS primitives + patterns shipped in phase 1. This
is the larger of the two parallel chat migrations (the other being the
server UI under turnstone/ui/static/).
Scope — this PR touches two files only:
turnstone/console/static/coordinator/index.html
- data-design="v1" on <html>; DS stylesheets linked after the legacy
base so primitives win on specificity and legacy styles keep
covering anything not-yet-migrated.
- Header rewired from .ts-header to .appbar with .appbar-back,
.appbar-title + .dim subtitle, .appbar-spacer, .appbar-status for
SSE state, and .appbar-actions wrapping the cancel / end / theme
buttons (now .btn pills).
- Approval bar replaced with the .approval-dock pattern. Signature
change: amber Approve becomes an ok-family (green) filled button
with 1.5px border + --r-md squared shape. .dcall rows frame each
pending call like a mini inspectable code line. Action cluster
sits in a .drow with Deny (.act.danger) / Always (.act.always) /
Approve (.act.primary) and the preview's kbd affordances
(D / ⇧A / ⏎). role="region" + aria-live="assertive" preserved;
the dock stays non-modal (no focus trap), focus moves to the
primary Approve button on open via the existing handler.
- Sidebar shell adopts .sidebar + .side-section + .side-label +
.ghost refresh buttons. Coordinator-only .sidebar overrides unset
the DS sticky-left-column defaults (which assume an admin-shell
grid) so the aside continues to flex into the right column of
#coord-body. Tree-row + task-row styling stays view-local,
rehomed to DS tokens (--hair-2 hover, --accent focus, --ok/--warn/
--err + -soft task-status tints).
- Inline <style> trimmed of rules now covered by DS primitives;
only the coordinator-specific flex wiring, tree-row visuals, and
<700px responsive accordion remain.
turnstone/console/static/coordinator/coordinator.js
- appendMsg() emits .msg + role variant (.msg.user / .msg.assistant /
.msg.reasoning / .msg.tool / .msg.error / .msg.info) and .msg-body.
_TS_ROLE_VARIANTS renamed _MSG_VARIANTS.
- Streaming helpers query .msg-body; SSE dedup-by-call-id query
updated to .msg[data-call-id=...].
- showApproval() renders the .approval-dock DOM shape: .dhead count
in a .dcount, one .dcall per pending call with .risk index pill +
.dfn function name + .dargs preview. approvalBar.hidden toggles
visibility (the DS pattern is position: fixed and always-rendered;
[hidden] is the show/hide hook).
- setSseStatus() keeps .appbar-status as the base; semantic colour
tracks OK / ERR via inline --ok / --err. Leading glyph (●/○/⚠)
preserves the WCAG 1.4.1 non-colour-only cue.
- Wait indicator uses .appbar-status instead of the legacy
.ts-header-status BEM; styling from the inline page rules colours
it --think.
Contracts preserved:
- SSE wire format and event names unchanged (approve_request,
child_ws_created, wait_progress, batch_started, state_change,
stream_end, ...).
- POST /approve body shape unchanged: {approved, always, call_id}.
No per-item feedback field is added (that's phase 9 PR C).
- Keyboard behaviour unchanged: Enter continues to approve via the
primary-button focus shift in showApproval(); the D and ⇧A kbd
labels are rendered per the pattern spec but the global key
handlers (if any) remain untouched.
- ARIA attributes (role, aria-label, aria-live) preserved on the
approval dock, messages log, and sidebar.
- All shared-static JS imports and order unchanged; composer module
continues to own its own DOM inside #coord-composer-mount.
No backend changes. Legacy CSS (/shared/base.css, /shared/ui-base.css,
/shared/chat.css, /static/style.css) stays linked as the compatibility
layer — DS selectors [data-design="v1"] beat legacy where applied.
* fix(coordinator): inline approval dock above composer, not viewport-pinned
The .approval-dock DS pattern defaults to position: fixed; bottom: 22px
— designed for the fleet dashboard where the dock overlays content. In
the coordinator chat that rule pinned the dock to the viewport bottom,
covering the composer input area.
Move the dock DOM back inside #coord-main between #coord-messages and
the composer mount so it flex-stacks naturally above the input. Add a
view-local override that neutralises the fixed positioning (position:
static, z-index/box-shadow auto) while preserving the visual pattern
(warm top stripe, head/call/actions rows, dashed Always button).
Drop the 160px bottom-padding hack on #coord-messages since the dock
is now in-flow and naturally pushes the message log up.
Also likely resolves the Firefox initial-render issue — position:fixed
+ [hidden] toggle had cross-browser quirks where the dock wouldn't
appear on first SSE approval event until a separate DOM mutation
forced a reflow. In-flow layout makes it boring and predictable.
* fix(coordinator): integrate judge verdicts into approval dock, not chat
The judge's intent_verdict is evaluation context for the pending
approval, not a chat message. Previously each verdict appended a
"[judge] deny (risk=low)" tool message into the transcript even when
the corresponding approval was visible in the dock — two separate
surfaces showing related decision context, neither one complete.
Now:
- Each .dcall row gets data-call-id from the approve_request item
- intent_verdict looks up the matching row and renders a .dctx sibling
below it with "judge: <recommendation> (risk: <level>)" + optional
"confidence: <score>" chips. Reasoning attaches as title tooltip.
- Verdicts cache in a Map<call_id, verdict> so late-arriving
approve_request events can still apply verdicts that came early
- Fallback to the old chat-message surface only when the approval isn't
visible (call_id missing, or resolved before we could render) so the
verdict isn't silently dropped
* feat(coordinator): judge verdict polish — colour-coded chips, spinner, reasoning
Three refinements to the approval-dock judge integration:
1. Colour-code verdict chips by recommendation — approve=green (--ok),
review=amber (--warn), deny=red (--err). Reviewers can triage at a
glance without reading the chip text; complements the text label
for WCAG 1.4.1 (non-colour-only signaling).
2. Spinner while evaluating — when showApproval builds a .dcall row
without a cached verdict, render a "judge evaluating…" chip with
a spinner. Replaced in-place when intent_verdict arrives. Reuses
the ts-spin keyframe from primitives/feed.css.
3. Justification inline — judge.reasoning is delivered in every
intent_verdict event but was hidden behind a title tooltip. Now
renders as a wrapped prose block (.drationale) below the .dctx
chips, styled like the .msg-body .evi callout (left-rule + mono
+ --ink-3). Full text, no truncation — justification is the whole
point.
View-local styling; the approval-dock pattern itself is unchanged.
If these patterns turn out to be broadly useful, they can promote to
shared_static/design/patterns/approval-dock.css in a later PR.
* fix(coordinator): defer approve-button focus until judge verdict arrives
The Approve button was getting focus the instant the approval dock
opened, which lit up the green focus ring and made the filled-green
button look pre-confirmed. A reviewer could mistake that for "already
approved" before the judge has even returned a verdict.
Now focus is deferred until the intent_verdict for the first-pending
call arrives, then moves to:
- Deny button when judge recommends "deny" (safety default)
- Approve button for "approve" / "review" / anything else
Fallback timer (3s) claims focus anyway if no verdict arrives — covers
disabled judge and slow judge cases so keyboard users still land on a
button within a beat.
Focus claim is idempotent so batch approvals don't bounce focus across
buttons as trickling verdicts arrive. hideApproval clears the timer
and the claimed flag so re-open cycles start fresh.
* fix(coordinator): drop approve-focus fallback timer
Previous commit added a 3s fallback that focused Approve if no verdict
arrived. Ambiguous — a focus ring that lands "eventually" looks the
same as one that lands because the judge recommended approve.
Now focus only ever moves when a real intent_verdict arrives. If the
judge is disabled or the verdict never comes, focus stays put and
keyboard users tab from the composer to reach the buttons. An absent
focus ring is a clearer signal than an ambiguous one.
* fix(coordinator): address Copilot PR #393 feedback
Copilot findings, applied:
- Restore <h2> for Children / Tasks sidebar section labels (were
changed to <span>). .side-label class still applies; screen readers
recover heading-level structure + rotor navigation.
- Mount the wait-indicator into #coord-header (the appbar container)
instead of #coord-status. #coord-status is reset via
statusEl.textContent = ... on every state_change event, which was
clobbering the wait indicator between ticks. As a sibling inside
the appbar, it survives state updates.
- Route `info` SSE events to appendText("info", ...) so they render
with .msg.info (think-indigo) styling. Prior routing to "tool"
gave info events accent-tinted tool-call styling, miscategorising
them visually.
- Define @keyframes ts-spin locally in the coordinator's <style>.
Canonical definition lives in primitives/feed.css but this page
doesn't link feed.css (no .feed-item usage), so the "judge
evaluating…" spinner wasn't animating.
- Clear judgeVerdicts Map in hideApproval. Map was growing unbounded
across resolve cycles — fine for short sessions, leaks memory on
long-lived coordinators with many approvals.
Not applied: Copilot's suggestion to restore focus-on-open or add a
fallback timer. User explicitly requested no fallback — the design
decision is that the focus ring should only ever appear when the
judge has returned a verdict, so an absent ring reliably means "no
recommendation yet." An auto-focus fallback would produce an
ambiguous ring that could be misread as "judge approved."
* feat(design-system): DS phase 1 — chat primitives for view migrations
Three new primitives enabling the chat-surface migrations (server UI +
coordinator):
primitives/message.css .msg + variants (user / assistant /
reasoning / tool / error / info / system),
.msg-meta author/timestamp slot, .msg-body
markdown target, .msg-actions hover-revealed
row, data-streaming="true" blinking caret.
Replaces .ts-msg* family in chat.css.
primitives/field.css .field wrapper with label/help/error, element
selectors for text/email/password/url/number/
search/tel/date/time/datetime/month/week +
textarea + select. .field.inline for checkbox/
radio rows, .field.invalid for error state.
Native-control focus-visible handled for
checkbox+radio so box-shadow ring remains
visible on unframed controls.
chrome/appbar.css chat-app header: back link + title + status +
action cluster. Distinct from the admin-style
.topbar (brand mark + nav + env metadata).
min-width:0 on .appbar-title so .dim subtitle
ellipsis fires under narrow viewports.
Preview.html extended with three demo sections exercising every variant
(plus a data-streaming example with live caret).
Fixes carried in from code review:
- @media (hover: none) and (pointer: coarse) to match chat.css
convention (hover-none alone is too broad, catches styluses)
- .field-help uses --ink-3 (not --ink-4 which fails AA on --panel)
- .msg-meta slot added so downstream PRs don't invent a custom class
- Tool message pre/code on --panel-2 (parent is --panel; same-bg
would make inline code disappear)
- Checkbox/radio :focus-visible override (native controls lack a
border for the default box-shadow ring to wrap)
- Message.css comment corrected: "accent-tinted" not "cyan"
All rules scoped under [data-design="v1"]. Nothing existing modified.
* fix(design-system): address Copilot PR #391 feedback
- .msg-actions: add pointer-events: none when hidden, auto when visible.
opacity:0 alone still intercepts clicks in the top-right corner —
broke text selection on short one-line messages. Toggle applied in
both hover/focus-within and the touch-media-query visible states.
- .field.inline comment: rewrite to match behaviour. Old comment said
".field stays flex-column" but the rule sets flex-direction: row.
- preview.html appbar demo: swap <a tabindex="0"> back-link to
<button type="button">. tabindex-only anchors without href have
inconsistent focus + screen-reader semantics; button is the correct
native element for "navigate back via JS."
Live-preview-driven tuning pass following PR #389:
Palette
- Accent hue 70 (amber) → 182 (teal). Amber collided with warn on
same-surface k-badges; teal gives the brand accent its own hue.
- ok / warn / err / think unified at L=0.50 light / L=0.68 dark and
C=0.13-0.17 for palette coherence. err holds higher chroma so red
doesn't wash; warn stays in the gold 80 lane (never 90+ / "puke").
- Soft variants unified at L=0.94 / L=0.29, C=0.05-0.07.
New tokens
--ok-live brighter green for liveness signals (running dot)
--ok-text theme-aware text colour for filled green surfaces,
dark forest in light / bright mint in dark, ~7.5:1
against the approve-button bg in both themes
--err-fill darker red specifically for filled destructive
surfaces (.risk.crit) — bright --err as a fill
reads as alarm-loud
--warn-tint, directly-defined gold tints for k-tools k-badge —
-tint-border skips the color-mix-through-dark-cool-panel mud
that would otherwise render warm low-L mixes brown
Approve / Always / Deny
- Approve filled green (color-mix --ok 28% into panel); text uses
--ok-text for theme-correct contrast. Matches the pre-refactor
turnstone/shared_static/chat.css convention where approve = green.
Deviates from the Claude Design spec which had warn-tinted approve.
- Always outlined dashed green (same --ok hue family); four non-colour
cues for WCAG 1.4.1: fill state, border style, label, position.
- Deny unchanged (err-outlined).
k-badge glyphs
Replaced generic shapes with semantic symbols:
tools ⚙ approval ⚠\FE0E policy § role ◉
oidc ⌘ token ◆ judge ⚖\FE0E query ?
step ⇧ session ◈ skill ★ workstream ⇉
fanout ⇶ default ·
⚠ and ⚖ carry \FE0E to force text-presentation (avoid emoji
promotion to coloured yellow triangle / blue scales on iOS Safari).
token uses ◆ instead of ⬢ for universal font coverage.
k-approval split from k-tools
k-tools stays gold (--warn family) — "tool call" kind.
k-approval moves to green (--ok family) — matches the Approve button
visually, completing the "⚠ approval → Approve" same-family story.
Running pill
Text uses --ok (passes AA on pale --ok-soft); dot uses --ok-live +
pulse. Liveness signal lives in the dot, not the text.
All changes stay under [data-design="v1"] — existing views untouched.
* feat(design-system): DS-A — tokens + typography scaffold
Adds turnstone/shared_static/design/{tokens.css,typography.css} as the
first phase of a multi-PR design refactor seeded by Claude Design.
- tokens.css: full palette + shape + rhythm, light default with
[data-theme="dark"] override. oklch() raw colours, color-mix kept out
of DS-A entirely (reserved for primitives in DS-B).
- typography.css: Inter + JetBrains Mono via Google Fonts; six-step
scale (10/11/12/13/14/20-24). Utility classes .t-kicker/.t-meta/
.t-btn/.t-row/.t-body/.t-stat/.t-h1.
Signature accent stays warm amber (oklch hue 70) rather than Claude
Design's teal — preserves turnstone's "Instrument Panel" identity.
All other tokens match the spec verbatim.
Additive: both files gate under [data-design="v1"] so existing views
(base.css, per-view stylesheets) are untouched. DS-B will opt views in
one at a time.
* feat(design-system): DS-B — chrome + primitives + preview page
Adds the reusable primitive kit that DS-C and DS-Cluster will build on:
primitives/
panel.css .panel, .panel-head (.tools pinned right), .ghost
buttons.css .btn (pill 999px), .primary, .deny, .approve (amber)
pills.css .pill (running/thinking/attn/idle/err), .k-badge
(glyph-prefixed per WCAG 1.4.1), .chip, .risk
stats.css .stat + .stat-row, .mini-bar, .spark
feed.css .feed-item (grid ts/body/acts + .evi callout)
chrome/
topbar.css 48px sticky, conic-gradient brand mark
sidebar.css 240px sticky, .shell layout, semantic swatches
preview.html renders every primitive in both themes with an
in-page theme toggle (tracks prefers-color-scheme)
Additive: every selector scopes under [data-design="v1"] so existing
views (base.css + per-view stylesheets) stay untouched.
Spec deviations from the Claude Design prototype:
- `color-mix(in srgb, …)` throughout; prototype had two `in oklab`
usages — srgb per the spec's hard rule
- `.btn.approve` is warn-tinted amber, not green
(approvals signal "needs attention"; amber resolves on approval)
- k-badge tint uses `color-mix` instead of oklch relative-colour syntax
for broader browser support
- `@keyframes pulse/spin` renamed to `ts-pulse/ts-spin` to avoid
clashing with keyframes in base.css on pages that load both
- `prefers-reduced-motion` disables pulse + spin animations
- Text-on-accent-soft + text-on-warn-tinted darkened via color-mix
with ink to pass WCAG AA at 12px (fixes the classic same-hue trap)
- `.risk.crit` uses `#fff` text (dark-mode --panel on bright err fails)
- `.feed-item .acts button:not(.btn)` — compact action styling now
skips .btn-classed buttons so they keep their pill shape
- Focus-visible rings on .btn, .ghost, .stat, .topnav, .side-item
* feat(design-system): DS-C — patterns (approval-dock, fleet-grid, live-feed)
Completes the design library with three patterns that compose primitives
into the signature product surfaces described in the Claude Design handoff.
patterns/
approval-dock.css bottom-pinned approval strip. 1.5px-border,
--r-md squared action cluster: amber Approve
(primary), dashed Always, red Deny. kbd hints
and focus-visible rings on all three acts.
Call row (.dcall) framed as an inline code-
panel to emphasize "this is the exact call."
fleet-grid.css 14-col grid of .node squares. State modifiers
(.s-ok/.s-thinking/.s-attn/.s-err/.s-idle/
.s-unreach) + --pct load fill. Hover uses
outline, not box-shadow (neighbour bleed is
the intended density cue). .fleet-legend
swatch row below.
live-feed.css thin scroll-container wrapper over the
.feed-item primitive with a sticky top fade.
preview.html imports the three patterns, extends the
fleet demo to use the real .fleet class +
legend, adds a live-feed panel, renders the
approval dock fixed at the bottom with
aria-live="polite".
Spec notes:
- Approve button is amber (warn-tinted), never green
- Dock action buttons are 1.5px-bordered 6px-radius squares — NOT
pills — signaling "primary-action surface"
- All three dock actions clear WCAG AA in both themes via the same
color-mix-with-ink darkening pattern used in .btn.approve
- kbd hint color matches primitives/buttons.css (--ink-3, not --ink-4)
View-level rewrites (coordinator.html + coordinator.js opt-in,
admin/cluster dashboard rebuild) are follow-up PRs — they need a
running server to test SSE streams + the approval POST contract.
* fix(design-system): scope DS-A tokens to [data-design="v1"]
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* fix(design-system): scope DS-A font vars to [data-design="v1"]
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* fix(design-system): align dark-mode selector with theme.js convention
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* fix(packaging): add shared_static/design/** to wheel includes
Agent-Logs-Url: https://github.com/turnstonelabs/turnstone/sessions/74d0939a-c55f-46b0-92f4-14d0cbfb7084
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* docs(coordinator): phase 8 PR C — API tour, skills guide, bulk-endpoints contract
Four deliverables that close out the phase 8 doc debt carried since
phase 1:
- docs/coordinator-api-tour.md — 9-step lifecycle walkthrough
(create → subscribe → send → inspect children / detail → wait for
fan-out → govern (trust / restrict / stop_cascade / close_all_children)
→ approve / cancel → close), one request + response per step, every
SSE event type the UI has to handle, and every operation id cross-
referenced against the live /openapi.json. Integrators driving a
coord session from a custom UI or SDK can work end-to-end from this
doc without reverse-engineering the console page.
- docs/coordinator-skills.md — writing a SkillKind=COORDINATOR skill.
Tool-surface diff (13 orchestration tools, no bash / edit / web /
sub-agent), persona diff (orchestrator vs maker, composing on
base_coordinator.md), SkillKind enum + migration 044, task_list
integration, ws_id handling, wait vs inspect cost profile, three
orchestration patterns (delegate-and-summarise, fan-out-and-
synthesise, plan-then-delegate), testing surface.
- docs/bulk-endpoints.md — codifies the two shape idioms that shipped
across phases 6–8: {results, denied, truncated} for bulk-read /
bulk-create-with-payload (cluster/ws/live, spawn_batch); {<bucket>,
failed, skipped} for cascade-mutation (stop_cascade,
close_all_children). Picks-by-semantics guidance so the next bulk
endpoint author doesn't coin a third shape.
- docs/diagrams/27-coordinator-wait-for-workstream.puml + rendered
PNG — sequence diagram covering spawn → wait (blocking, with
bounded progress emission) → inspect → close. Embedded in the
API tour doc's §6 so the "why is my coord session blocking?"
question has a visible answer.
No code changes. All operation ids in the API tour verified against
a live build of the console spec; all markdown internal links
resolve; PlantUML renders clean on the system plantuml jar.
* docs(coordinator): address PR #388 copilot review
- api-tour.md child-event payload keys: events stamp `ws_id` as the
coord's own id and carry the child's id separately as
`child_ws_id`. Doc previously listed `ws_id` as the child
identifier on all four child_ws_* events, which would send SDK /
UI implementers parsing the wrong field.
- api-tour.md SSE table: add the `status` event emitted by
ConsoleCoordinatorUI.on_status (token usage + context_window +
effort snapshot; fires on every streaming tick). Previously
omitted from the "every event type a UI has to handle" list.
- api-tour.md /children response key: server returns `{items,
truncated}`, not `{children, truncated}`. Also drop the
`state=closed` query-param claim — the endpoint has no state
filter; clients filter locally on the returned `state` field.
- skills.md task_list shape: the persisted row uses `id` (not
`task_id` — the input schema uses `task_id`, the row uses `id`),
has `child_ws_id` / `created` / `updated` (no `notes` field),
and supports a 5th `reorder` action alongside add/update/remove/
list. Adds the parallel-dispatch caveat from the tool
description.
- skills.md tenant-guard behaviour: foreign / hallucinated ws_ids
don't return an empty result — they return explicit
error/not-found/denied shapes that differ by op (mutating ops
return `{error, status: 404}`; inspect returns `{error}`; wait
reports state=denied). Important distinction — a skill that
expects empty on mismatch will mishandle every single case.
Docs-only; no code / schema / SDK changes. All internal links
still resolve.
* feat(coordinator): phase 8 PR B — spawn budget + rate limit + /quota endpoint
Adds two complementary controls so a runaway coordinator can't saturate a
cluster's max_active without anyone noticing:
- **Spawn budget** (hard quota) — cap on concurrently active children.
Default 20 per coord. spawn_workstream returns a tool error guiding
the model to close idle children; spawn_batch routes overflow rows to
`denied[]` with partial-success semantics.
- **Spawn rate limit** (soft pacing) — classic token bucket, defaults
5 tokens/minute with burst 10. A rate-limited spawn surfaces a tool
error carrying `retry after Ns` so the model paces itself. Zero
refill rate is honoured as "disable refill" (bucket still honors the
initial burst).
Shipped infra:
- `turnstone/core/spawn_quota.py` — thread-safe `SpawnBudget` +
`TokenBucket`. 15 unit tests.
- `turnstone/core/session.py` — coord-only state built from settings at
__init__. Shared `_eval_spawn_quota(active)` helper drives both the
single-spawn path (wraps the denial reason in `_coord_tool_error`) and
the batch path (annotates `spec["_error"]`). `_count_active_children`
routes through `coord_client.list_children(include_closed=False)` and
fails *open* on lookup error (budget is operator-safety, not security).
- `POST/GET /v1/api/coordinator/{ws_id}/quota` — partial-update admin
endpoint mirroring the /trust + /restrict shape. Accepts either the
nested `spawn_rate` object or flat aliases — supplying both for the
same field returns 400 so the admin UI can't half-migrate silently.
Overrides are in-memory only (die on session reopen). Audits via
`coordinator.quota.updated` with before/after snapshots.
- Settings: `coordinator.spawn_budget`, `coordinator.spawn_rate.tokens_per_minute`,
`coordinator.spawn_rate.burst` with ranges 1..500 / 0..600 / 1..500.
The range bounds are the single source of truth — the endpoint
validators and Pydantic schema both import from `settings_registry.SETTINGS`
so bumping a cap in one place lights up everywhere.
- OpenAPI: `CoordinatorQuotaRequest` / `CoordinatorQuotaResponse` /
`CoordinatorSpawnRateState` schemas + endpoint specs. TS SDK regenerated.
Tests: +15 unit (SpawnBudget + TokenBucket), +17 endpoint (GET + POST
happy paths, range edges, mixed-body rejection, non-object spawn_rate,
service-token refusal), +11 session-side (budget blocks single spawn,
budget batch partial-success, rate batch partial-success, empty-body
reject, mutator live-update, non-coord session has no quota state).
Deferred (not this PR): per-skill scoping via migration 047 +
`prompt_templates.spawn_budget` column. Count-only storage helper
(opportunistic — list_children at budget ≤ 500 is fine behind a
human-gated approval flow).
* fix(coordinator): address PR #387 copilot review
- Budget undercount: _count_active_children used list_children's
LIMIT-then-Python-filter path, so a fan-out with many recently-closed
children could push live rows past the SQL LIMIT and silently
undercount, leaking spawn slots past the budget. Replace with a new
CoordinatorClient.count_active_children that uses
storage.count_workstreams_by_state (SQL aggregate, no pagination,
sums non-terminal states). Tenant-guarded; fails open on storage
error (budget is operator-safety, not a security gate). New client
tests cover the non-terminal count, the closed/deleted exclusion,
the foreign-parent guard, and the fail-open path.
- Service-token bypass on /quota: both GET and POST used the default
allow_service_bypass=True, so a service token whose user_id matched
the coord owner could read or *raise* spawn capacity without the
explicit admin.coordinator grant. Flip both to
allow_service_bypass=False for consistency with /restrict,
/stop_cascade, and /close_all_children.
- OpenAPI contract leak: CoordinatorSpawnRateState was used for both
the request and response shapes, which let generated SDKs imply
clients could POST tokens_available (a read-only bucket reading the
handler ignores). Split into CoordinatorSpawnRateInput (request:
tokens_per_minute + burst only) and CoordinatorSpawnRateState
(response: adds tokens_available). No runtime behaviour change;
SDKs regenerate with two distinct types.
Drops the _ACTIVE_COUNT_SLACK / _ACTIVE_COUNT_MIN_LIMIT constants in
session.py — no longer needed since the new helper takes no limit
argument. Updates the 5 session-side quota tests to stub
count_active_children instead of list_children.
* feat(coordinator): phase 8 PR A — spawn_batch + close_all_children batch tools
Adds two model-facing batch tools so a coordinator can fan out without burning one approval per child:
- `spawn_batch` — create up to 10 child workstreams in a single approval. Serialised
spawns so sibling ordering (by created_at) stays deterministic. Returns
`{results: {idx: {ws_id, name, node_id, status}}, denied: [{idx, reason}]}`.
Per-item validation / spawn failures surface in `denied[]`; the batch hard-errors
on >10 rather than silent truncation.
- `close_all_children` — soft-close every direct child in one approval. Server-side
Sem(16) fan-out via `coord_client.close_workstream`; `reason` propagates to every
closed child's audit + workstream_config. Response mirrors `stop_cascade`'s cascade
idiom: `{closed, failed, skipped}` where `skipped` is upstream-404 / already-gone.
Shipped infra:
- New console endpoint `POST /v1/api/coordinator/{ws_id}/close_all_children`
(gated `admin.coordinator`, `allow_service_bypass=False`, 512-char reason cap,
`coordinator.closed_all_children` audit).
- Shared `_fanout_on_children` helper — both `stop_cascade` and `close_all_children`
now delegate to it (one place to own the snapshot → semaphore-gather → bucket-split
skeleton).
- `CoordinatorClient.close_all_children(reason)` plus a `_post_url` seam that
`_post` now reuses (no more duplicated transport-error handling).
- `_emit_batch_event` — best-effort SSE emitter modelled on `_emit_wait_event`.
Emits `batch_started` / `batch_ended` pairs keyed by call_id. Throttled
`batch_progress` deferred to a follow-up.
- OpenAPI request + response schemas, endpoint spec entry, TS SDK regenerated.
- Persona doc (`tools_coordinator.md`) covers the two new patterns.
Bulk-endpoint shape policy (codified in PR C later): split by semantic category —
`{results, denied, truncated}` for bulk-read / bulk-create-with-payload (cluster/ws/live,
spawn_batch), `{<bucket>, failed, skipped}` for cascade-mutation (stop_cascade,
close_all_children). No retrofit needed on stop_cascade.
Tests: new `test_coordinator_close_all_children.py` (8 endpoint tests), expanded
`test_coordinator_tools.py` (session-side prepare/exec, coord_client=None guards,
batch SSE events), expanded `test_coordinator_client.py` (route map, client method,
transport errors), tool-count assertions updated.
Deferred (not this PR): per-item selective-deny approval UI, throttled batch_progress
SSE, coordinator-skills doc + bulk-endpoints doc (PR C), spawn budget / rate limit (PR B).
* fix(coordinator): address PR #386 copilot review
- coordinator_client.close_all_children: pass the unformatted path template
as log_path so telemetry aggregates don't fragment per session (ws_id
still lives in the real URL).
- session.py: drop dead spawned_ids accumulator in _exec_spawn_batch —
leftover from an eager-register path that got removed earlier.
- close_all_children tool JSON: document the 512-char server-side cap on
reason and that reason is echoed back in the response payload. Added
maxLength:512 on the schema property so the LLM sees the constraint.
- CoordinatorCloseAllChildrenRequest: add Field(max_length=512) so the
OpenAPI schema reflects the runtime 400-on-overflow constraint.
* refactor(ui): shared composer widget (pane / coordinator / coord-create)
The interactive workstream pane (turnstone-server), the coordinator
session view (turnstone-console), and the console home's "start a new
orchestration task" form had drifted into three unrelated composer
implementations with different DOM, different class names, and
different behaviour sets. All three now build on a single
`shared_static/composer.js` widget parameterised by feature flags.
The widget owns the textarea, send button, optional stop button,
optional attach button + file input + chip container, optional
drag-drop / paste-image wiring, optional queue-while-busy send-label
rotation, optional touch-aware Enter-to-send, and an optional
collapsible Options panel with input / select fields, live summary
chip, and localStorage-persisted open/closed state. A stacked layout
puts the textarea above the action row for creation-form consumers;
the inline layout keeps the chat-style single row for send composers.
Consumer wiring:
- Pane: attachments + stopBtn + queueWhileBusy + drag-drop; keeps
its own attachment-upload pipeline and routes file events through
the composer's onAttach callback. Pane-specific CSS (.pane-stop
/ .pane-send.queue-mode) retired; `.ts-composer-stop` / `.ts-
composer-send--queue` in shared/chat.css take their place.
- Coordinator send: just textarea + send with touchEnterSends=true
to preserve the pre-refactor tap-to-send behaviour on tablets.
The header-mounted coord-cancel-btn stays (different semantics
than a per-generation stop).
- Coord-create: stacked layout, rows=3, Start-labelled send, Options
dropdown holding Name + Skill; Ctrl/Cmd+Enter handler scoped to
the composer mount. `_createCoordinator` lost its DOM-ref shape
in favour of raw values + a setBusy callback; a single
`_refreshHomeCoordSubmitEnabled` reconciler owns the submit
button's disabled flag so the 503 probe and in-flight submit
can't race each other.
Shared chat.css grew the .ts-composer-stop, .ts-composer-options-*,
and .ts-composer--stacked blocks; the coord-create consumer dropped
its custom .home-composer-task/-row/-name/-skill/-submit selectors
and the "Start a new orchestration task" panel title so the
placeholder text carries its own context, matching the webui
dashboard's clean look.
The visible behaviour on each surface is intentionally the same as
before; the change is structural — the three composers can no longer
drift apart silently.
* fix(composer): review fixes — Enter guard, single disable owner, widget-owned stop reset
Review of the squashed whole caught issues that the piecewise reviews
missed because they only become visible with all three consumers
together:
- **Enter bypassed sendBtn.disabled.** Composer's Enter keydown
handler called _fireSend() without checking sendBtn.disabled. In
the coord-create flow submitHomeCoord doesn't clear the textarea
before the POST completes (it redirects on success), so two rapid
Enter presses both fired _createCoordinator and could create two
coordinators. Enter now mirrors the click path.
- **Two writers to sendBtn.disabled.** Composer.setBusy and
_refreshHomeCoordSubmitEnabled both wrote the flag. They agreed
in sequence today but it was the exact drift hazard the reconciler
was meant to eliminate. Added an externalDisable option; when
true Composer's setBusy rotates labels / placeholder / stop button
but leaves sendBtn.disabled to the caller's reconciler. The
coord-create composer opts in.
- **setSendLabel dead weight.** Called on every busy transition
with static "Start" / "Starting…". Composer.setBusy now rotates
labels universally (not just in queueWhileBusy mode); the busy
label goes at construction via the existing busyLabel option and
setSendLabel is removed.
- **Pane reached through composer to reset stopBtn.** Pane.setBusy
was writing stopBtn.textContent / aria-label / dataset after
delegating — internals leaking through. Composer.setBusy now
resets the stop button's standard label + clears forceCancel on
every transition (matching the comment that used to live in Pane);
Pane drops the reach-through.
- **destroy() left detached DOM reachable.** Back-refs (inputEl,
sendBtn, etc.) are nulled out so post-destroy access fails loudly
instead of silently mutating detached nodes.
- **_maybeAutoResize dead indirection.** The enabled-check folded
into autoResize itself.
* fix(composer): busyLabel context-sensitive default + busyPlaceholder universal
Round-two review caught three related loose ends:
- Default `busyLabel="Queue"` was fine when label rotation was queue-
mode-only, but became misleading after the earlier review fix made
rotation universal: non-queue consumers calling setBusy(true)
without explicit busyLabel would flash "Queue" on the disabled
button. Default is now context-sensitive — "Queue" when
queueWhileBusy=true, sendLabel otherwise (no rotation). The
coord-send composer no longer needs to touch the label at all.
- `busyPlaceholder` JSDoc implied universal swap on busy but the
implementation gated it on queueWhileBusy. Decoupled — the
placeholder swaps whenever busy, with callers that don't set
busyPlaceholder seeing no visible change because it defaults to
the idle placeholder.
- `options.toggleLabel` and `options.onChange` were supported by the
implementation but undocumented. JSDoc for the options shape
enumerates every supported key + its default.
* refactor(routing): replace hash-ring rebalancer with rendezvous (HRW) hashing
Routing was a stored bucket table maintained by a central rebalancer
daemon, which shared its liveness primitive (services.last_heartbeat)
with the collector — when a heartbeat-fresh node went into a zombie
HTTP-handler-broken state, neither the collector nor the rebalancer
could self-correct, and the router kept directing traffic at it.
Rendezvous hashing makes the route a pure function of (ws_id,
live_services) so the heartbeat is the single source of truth and any
liveness-eviction propagates to the next route call without a separate
state-publication step.
The rebalancer's central state has no analogue: the new router computes
the per-key node winner on every call, the collector pushes membership
updates into the router cache from its discovery thread, and per-route
overrides survive on workstream_overrides. Eager workstream migration
goes away; in-flight workstreams lazily rehydrate from storage on the
new owner — already the dead-node behaviour.
* fix(tools): describe rendezvous re-routing on spawn/inspect node_id
The first pass overclaimed `node_id` "stays canonical for this
workstream's lifetime" — under rendezvous routing the active owner
re-derives per-call from live membership, so a node join/drop after
spawn can shift it. Tool descriptions now say `node_id` is the
spawn-time binding; subsequent ops re-route via rendezvous over the
current live-node set; the new owner lazily rehydrates from shared
storage; coordinators should re-read with inspect_workstream rather
than caching the value.
* feat(coordinator): phase 7 — governance + skill metadata + cross-cutting invariants
Combines three stacked sub-PRs into a single coordinator phase-7
shipment against the phase-7 plan doc. The sub-PR structure (0 / A /
B) preserved on individual branches for reviewer drill-down; this
branch is the one reviewers should merge.
## Sub-PR 0 — service-auth boundary invariants
Shared helpers and contracts that lock the console ↔ node service-auth
boundary so later authz surfaces use them by construction.
- ``_effective_user_filter(request)`` in both ``turnstone.console.server``
and ``turnstone.server`` with a shared ``DENY_EMPTY_SUB`` sentinel
on ``turnstone.core.auth``. Three-way return — admin/service
bypass, scoped caller uid, or fail-closed sentinel on blank sub.
Four callsite migrations (``_coordinator_rows``,
``coordinator_children``, ``coordinator_metrics``,
``cluster_ws_live_bulk``).
- ``StorageBackend`` class docstring codifies the tenancy contract
(every list/count/aggregate method must accept ``user_id: str |
None = None`` and push ``WHERE user_id = :user_id`` into SQL) and
the ``_mapping`` row-access contract. New
``turnstone.testing.row_contract`` ships ``assert_row_like()``.
- ``_verify_collector_service_scope`` probes an upstream node at boot
with ``expected_node_id=_scope-probe_``; a 409 proves the scope
gate was passed, a 403/401 sets ``collector_scope_error`` and
causes ``cluster_snapshot`` / ``cluster_events_sse`` to return 503
with a remediation hint. Probe URL allowlist rejects non-http(s)
schemes and 169.254.0.0/16 hosts.
- 4xx log-level floor on ``_NodeDashboardCache.get``,
``_fetch_live_block``, and ``_proxy_sse`` — dotted-hierarchy
prefixes with bounded body previews. ``_bounded_body_preview`` and
``_bounded_stream_preview`` strip control chars.
## Sub-PR A — coordinator governance core
Mid-session governance surface for coordinator workstreams.
- **Trusted-session mode.** New ``coordinator.trust.send``
permission (migration 042). ``ChatSession.set_trust_send`` /
``revoke_tools`` methods with a ``_governance_lock``. ``POST
/v1/api/coordinator/{ws_id}/trust {send: bool}`` double-gated on
``admin.coordinator`` AND ``coordinator.trust.send`` with
``allow_service_bypass=False`` so service tokens can't escalate.
``_prepare_send_to_workstream`` auto-approves sends whose target is
in the coordinator's own subtree; foreign ws_ids still require
approval. ``_is_own_subtree`` checks both ``parent_ws_id`` AND
``user_id`` to defend against cross-tenant row corruption.
- **Audit-layer credential redaction.** ``record_audit`` walks
``detail`` (dicts, lists, tuples, sets, frozensets; keys too)
and routes every string through ``redact_credentials`` + a C0
control-char scrub. New kw-only ``raw_detail=True`` opt-out.
``_has_any_string`` fast-path. Audit action registry extended
with the four new governance sub-prefixes.
- **Mid-session revocation + cascading stop.** ``POST
/v1/api/coordinator/{ws_id}/restrict {revoke: [...]}`` caps 256
entries / 128 chars; ``_prepare_tool`` short-circuits with a
tool-error. ``POST /v1/api/coordinator/{ws_id}/stop_cascade``
cancels the coord's in-flight generation then dispatches
``cancel_workstream`` for every direct child in parallel via
``asyncio.gather`` bounded by ``Semaphore(16)``. Per-child
outcomes split into ``cancelled`` / ``failed`` / ``skipped``
(404 = already-gone rather than dispatch-broken). Both endpoints
apply ``allow_service_bypass=False`` on the admin gate.
- **Shared plumbing.** ``_resolve_coord_session`` helper collapses
the handler prelude three endpoints shared. ``_emit_coord_audit``
wraps ``record_audit`` in a dedicated ``ThreadPoolExecutor``
(``app.state.audit_executor``) so audit bursts don't starve cancel
dispatches. ``_require_json_object`` guards body parsing so non-
object JSON returns 400 instead of 500.
## Sub-PR B — skill metadata governance
- **Description validator (migration 043).** ``prompt_templates``
rows now require a non-empty ``description``. Existing empty rows
get backfilled with a ``"Skill: <name>"`` placeholder on upgrade.
The installer (``admin_skill_discover``) and MCP prompt sync both
synthesise a placeholder when the upstream description is blank
so non-admin write paths satisfy the invariant.
- **Skill kind classifier (migration 044).** New
``prompt_templates.kind`` column (``interactive`` / ``coordinator``
/ ``any``; defaults to ``any``). New
``turnstone.core.skill_kind.SkillKind`` StrEnum is the single
source of truth; Pydantic schemas type ``kind`` as ``SkillKind``
(OpenAPI advertises the enum) and the handler validator catches
the ValueError. ``list_skills_filtered`` gains a
``kinds: list[str] | None = None`` SQL filter.
``CoordinatorClient.list_skills`` defaults to
``kinds=["coordinator", "any"]`` so interactive-only skills are
hidden from the orchestrator.
- **``scan_status`` → ``risk_level`` rename (migration 045).**
Lossless column rename to align with ``IntentVerdict.risk_level``
terminology. Swept storage (both backends + schema + protocol),
handlers, API schemas, tool JSON, generated OpenAPI specs,
TypeScript SDK types, frontend (``governance.js``), tests, and
English prose in ``docs/judge.md`` + ``docs/tools.md``. The
user-facing on-load warning now reads ``has risk level:
{risk_tier}``. Tool JSON's ``risk_level`` enum corrected to the
scanner's actual taxonomy (``safe / low / medium / high /
critical``; was the never-shipped ``clean / flagged / unscanned /
pending``). Historical migration 021 left untouched.
## Migrations
042 (``coordinator.trust.send`` perm — PR A)
043 (description backfill — PR B)
044 (``kind`` column add — PR B)
045 (``scan_status`` → ``risk_level`` rename — PR B)
All four use position-anchored permission strings / host-side
parse-filter-rejoin on downgrade where SQL ``REPLACE`` could
corrupt prefix-overlapping values.
## Verification
- ``ruff check turnstone tests`` clean.
- ``mypy turnstone`` clean on 165 source files.
- ``pytest -m "not live"``: 4431 passed (+85 over the phase-6
baseline). Includes +32 tests in ``tests/test_service_auth_boundary.py``
and +38 in ``tests/test_coordinator_governance.py``; shared fixtures
extracted to ``tests/_coord_test_helpers.py``.
- Generated OpenAPI JSON (``sdk/typescript/openapi-{console,server}.json``)
regenerated via ``sdk/typescript/scripts/generate-types.py``; zero
``scan_status`` occurrences remaining outside the historical
migration 021 and the rename migration 045.
## Security reviews
Both reviews flagged by the phase-7 plan (items 1 + 5, plus 0a's
refuse-to-serve gate) ran through the multi-stage ``/review``
pipeline twice per sub-PR; all confirmed findings landed in-branch.
* fixup(phase-7): CI lint + PR #383 review fixups
Addresses the lint CI failure (ruff format) plus 12 findings from the
two automated PR reviewers.
Copilot:
- ``_sqlite.list_installed_skill_urls`` / ``_postgresql.list_installed_skill_urls``
used positional row indexing (``r[0]``/``r[1]``/``r[2]``) while this
same PR's ``StorageBackend`` class docstring forbids it. Switched
both to ``r._mapping["..."]`` access.
- ``list_skills.json`` previously advertised ``risk_level=""`` as a
filter for unscanned skills, but the implementation treats empty
strings as "no filter". Clarified the tool description to say
omit the filter entirely to include unscanned rows, and added an
explicit ``enum`` on the parameter restricting it to the scanner
tiers. ``_prepare_list_skills`` keeps the ``strip() or None``
normalisation — unscanned filtering now has an unambiguous contract.
- ``test_storage_skills_filtered.test_risk_level_filter`` used the
legacy ``clean`` / ``flagged`` values from the pre-rename column.
Rewritten with the scanner's actual taxonomy (``safe`` / ``high``).
github-code-quality (CodeQL):
- ``test_deny_sentinel_is_singleton`` previously asserted
``cs.DENY_EMPTY_SUB is cs.DENY_EMPTY_SUB`` — an identical-expression
comparison. Rewritten as two separate ``from ... import ... as`` aliases
(``FIRST_READ`` / ``SECOND_READ``) so the identity check is between
distinct bindings.
- ``test_restrict_empty_revoke_is_noop_but_audits`` unpacked ``state``
without using it. Renamed to ``_state``.
- Mixed import styles in ``test_service_auth_boundary.py`` — the
file previously used both ``import turnstone.console.server as cs``
and ``from turnstone.console.server import ...`` for the same
module (same story for ``turnstone.core.auth`` and
``turnstone.server``). Consolidated to the ``from X import Y`` style
used elsewhere in the file; the ``_fetch_live_block`` test now
patches via pytest's ``monkeypatch`` fixture instead of a manual
rebind through a module alias.
CI:
- ``ruff format`` reformatted one line in
``tests/test_coordinator_endpoints.py``.
Verification: ruff check + mypy clean (166 files); 4459 non-live
pytest pass.
* fix(tests): swap asyncio marker for anyio in service-auth boundary tests
PR #383 CI caught that the 13 ``@pytest.mark.asyncio`` decorators I
added in ``test_service_auth_boundary.py`` are an off-convention
choice — the rest of the repo uses ``@pytest.mark.anyio`` (148 sites
vs my 13). The CI environment pulls in ``anyio`` but not
``pytest-asyncio``, so every async test in this one file was failing
with "async def functions are not natively supported". It passed
locally by accident — my dev venv happens to have pytest-asyncio
installed ambiently.
Swapped all 13 marker sites to ``@pytest.mark.anyio``. No functional
change; the tests run under the same default asyncio backend anyio
provides.
Verification: ruff + mypy clean (166 files); 4459 non-live pytest
pass.
* refactor(channels): backfill review of Slack/Discord adapters
Retrospective multi-stage review of the Slack (PR #355) and Discord
channel adapters — they shipped before the review pipeline existed,
so this pass goes back and fixes everything the pipeline would have
caught plus a follow-up round of ultrareview findings.
## Security (8 fixes)
- Adapter-side owner checks on all interactive flows: Discord
ApprovalView / PlanReviewView encode the owner Discord user ID in
the embed footer (`{ws_id}|{corr_id}|{owner_id}`) and reject
non-owner clicks; Slack plan-approve / request-changes /
feedback-modal gain owner tracking in `_pending_plan_review_ts`
and a shared `_ensure_plan_review_owner` gate. These closed the
two critical authz gaps where the gateway's service-scoped JWT
bypassed server-side ownership checks.
- Discord thread-message gate: only the registered invoker can
drive the workstream (prevents a linked user posting in another
user's public thread from injecting into their assistant).
Invoker recorded explicitly so `/ask` follow-ups survive the
`channel.create_thread` bot-as-owner quirk.
- Slack /link flow + per-user identity gate: unlinked Slack users
see an ephemeral `/turnstone link <token>` prompt on every
message instead of silently creating workstreams under the
shared gateway identity. Rate-limited (5/hour) to block online
token enumeration.
- Gateway `/v1/api/notify` requires `write` scope on the validated
JWT; low-scope tokens get 403 + audit.
- Thumbnail URL validator DNS-resolves the hostname before fetch
and rejects any resolved IP that's loopback / link-local /
multicast / reserved, plus an explicit deny-list for IPv6 cloud
metadata (`fd00:ec2::/32` — AWS Nitro IMDS + ECS task metadata)
that would otherwise slip past the `is_private` allowance.
- Per-user rate limit (10 msgs / 60s) + 8 KiB inbound size cap on
Slack DMs / channels / notification-reply threads so one user
can't exhaust the shared LLM budget.
- Discord /link rate limit (5/hour) for token-enumeration defense.
## Bug fixes (9 correctness issues)
- Slack DM routing: each top-level DM no longer spawns a fresh
workstream (was using per-message `ts` as the route key).
- Multi-chunk Slack responses thread correctly under the first
chunk's ts instead of fragmenting as independent top-level
messages.
- Finalize the outgoing StreamingMessage before swapping channel /
thread_ts mid-stream, so buffered tokens still land on the old
thread.
- Redundant `chat_update` on approve/deny eliminated by popping
`_pending_approval[ws_id]` after local resolution.
- Notification reply tracking on Discord only registers for DMs
(guild-channel targets were storing channel IDs where user IDs
were expected, so legitimate replies were always rejected).
- `get_channel_default_alias` rolls `_channel_default_ts` back on
`list_models()` failure so the next caller retries instead of
serving an empty alias for the full TTL.
- Slack `subscribe_ws` purges dead SSE tasks before the
membership short-circuit (previously an unhandled exception left
the ws_id in `_subscribed_ws` forever, silently no-opping
subsequent subscribes).
- ChannelRouter `_create_locks` is now an LRU-bounded OrderedDict
that evicts only unheld locks (original dict grew unbounded;
naive LRU could evict a held lock and let a second caller race
through the critical section, creating duplicate workstreams).
- Slack `_parse_ts` pads the fractional field to 6 digits so
`"1.2"` and `"1.000002"` stop colliding as `(1, 2)` in the
latest-session tiebreaker.
## Performance (6 fixes)
- StreamingMessage keeps a rolling truncated display string capped
at `max_length` so per-flush cost is O(max_length) instead of
O(total_streamed_chars) — long streaming responses no longer do
quadratic work every edit interval.
- `StreamingMessage.finalize()` caches the joined content so the
Discord stream-end DM-forward path doesn't re-join a multi-MB
buffer twice.
- `PendingApproval` stores the Block Kit payload posted to Slack;
`IntentVerdictEvent` appends the verdict in-place and
`chat_update`s, skipping an extra `conversations_history`
round-trip.
- ChannelRouter `lookup_ws_id()` TTL-caches the channel →
ws_id resolution (30s TTL, 4096-entry LRU); hot inbound paths
skip storage on every message.
- Service-discovery startup retry uses exponential backoff
(1s → 8s cap) with a 30s deadline instead of 30 × 1s fixed
sleep.
- `_archive_session` now calls `router.close_workstream` so the
`_node_urls` cache entry is dropped (was leaking one entry per
archived session).
## Quality / refactors (19 improvements)
- `cli.main()` extracted from a 365-line function into focused
helpers; imports carefully kept lazy where test patches target
source-module paths.
- `_run_gateway` finally block now awaits `adapter.stop()` on
every adapter so SSE tasks, httpx clients, and the Slack socket
handler close cleanly on shutdown.
- Shared SSE reconnect loop extracted to `turnstone/channels/_sse.py`
(`run_sse_stream` with `on_event` + `on_stale` callbacks); both
adapters' `_sse_listener` methods just wire up callbacks. The
"404 stops reconnect" invariant is enforced inside the helper
so a broken `on_stale` can't livelock.
- `_on_ws_event` god-dispatchers split into per-event `_handle_*`
methods with a thin isinstance dispatcher at the top.
- Slack `_on_approve` / `_on_deny` collapsed into a single
`_resolve_approval(*, approved: bool)`.
- `ApproveRequestEvent` policy evaluation hoisted into
`ChannelRouter.evaluate_tool_policies` returning a
`PolicyVerdict`; adapters switch on the verdict kind.
- `ChannelAdapter` protocol trimmed to the four methods adapters
actually implement; unused `ChannelEvent` dataclass removed.
- Shared constants lifted to `turnstone/channels/_config.py`.
- `_cleanup_stale_route` and `unsubscribe_ws` share a
`_clear_ws_state` helper.
- `StreamingMessage` private attrs promoted to `message` /
`message_ts` / `accumulated_text` properties so callers don't
reach past the `_`-prefix.
- Various cleanups: dead var, noqa'd lambdas, renamed
`_policy_handled` → `policy_handled`, inlined single-use
helpers, added module docstrings, documented
`SlackRoute.parse` edge cases.
- `chunk_message` plain-text fast path (no backticks → skip
fence bookkeeping).
## Test coverage
Added 45 tests (178 → 223):
- `tests/test_channel_sse.py` (new) — SSE reconnect / backoff /
404-stale-route / on-stale-exception / invalid-JSON-skip /
on-event-exception-doesn't-kill-stream / per-connection token
refresh / ConnectError retry.
- ApprovalView + PlanReviewView owner-check regression tests
(owner allowed, non-owner rejected, legacy 2-pipe footer fails
closed, modal path rejected for non-owner, `/ask`
bot-as-thread-owner follow-up allowed).
- Slack `_recover_routes` latest-ts-wins, `_archive_session`
drops route + closes workstream.
- SSRF tests: DNS rebinding rejected, IPv4 link-local metadata
rejected, IPv6 ULA metadata (fd00:ec2::254 / fd00:ec2::23)
rejected.
- Slack link prefix match (natural-language prompts don't
hijack), link rate-limit ceiling.
- SlackRoute round-trip across all three shapes + lax-parse
behaviour.
Lint (ruff) + mypy clean; 210 channel-focused tests pass.
* chore(channels): address PR #382 review-bot feedback
Three line-level findings from github-code-quality on the backfill
review PR. Copilot had no line-level comments.
- _sse.py:132 — the `except httpx.HTTPStatusError: pass` branch was
flagged as an empty except. The original status was already logged
at WARNING inside the try block (we re-raise ourselves after
logging), so the handler has real intent. Added a debug log of the
exception text + a comment explaining the control flow, so the
empty-except lint stops firing and the next reader sees why we
fall through to backoff.
- discord/bot.py:430, cli.py:354, slack/bot.py:1127 — `await task`
inside `contextlib.suppress` was flagged as "statement has no
effect". It's a false positive (await is an effect) and the
alternative try/except/pass triggers ruff SIM105. Kept the
contextlib.suppress pattern and added an explanatory comment above
each call so the intent (await CancelledError propagation before
state cleanup) is obvious; will reply on the PR thread noting the
false positive.
No behavior change. Lint + mypy clean; 210 channel tests pass.
* feat(coordinator): phase 6 — polish, observability, active-coords via SSE, frontend cleanup
Squashed from two working commits:
1. phase-6 backend polish + active-coords SSE
2. phase-6 frontend cleanup (legacy chat-view classes + designer nits)
Both tier-A/B observability items and tier-C frontend consolidation
ship together — the shared-vocabulary migration touches surfaces the
backend polish already had its hands in, so one combined commit keeps
the diff reviewable as a coherent phase.
Observability
-------------
- **Coordinator-side wait dashboard** — `_exec_wait_for_workstream`
emits `wait_started` / `wait_progress` / `wait_ended` SSE events
via a new `progress_callback` hook on
`CoordinatorClient.wait_for_workstream`; coordinator.js renders a
"⧗ waiting · N ws · Ts" header indicator keyed by call_id so
overlapping waits coexist. Progress throttled to emit only on
snapshot-diff or 5s heartbeat; full results dict attached only on
transitions so a 600s wait doesn't flood SSE listener queues.
Indicator only attaches when a proper header host exists (no
floating document.body fallback) and is cleared on SSE reconnect
so a dropped `wait_ended` can't pin the badge.
- **`cancel_workstream` forensics** — `server.cancel_generation`
captures `ui._pending_approval` tool names +
`session._queued_messages` count / preview before invoking
`session.cancel`, returning the snapshot as `dropped`; routing
proxy passes it through to the tool result. Preview runs through
`redact_credentials` before the 120-char truncate so pasted
secrets / connection strings don't land verbatim in the
coordinator's conversation history.
- **Per-coordinator metrics** — `GET /v1/api/coordinator/{ws_id}/metrics`
returns `spawns_total` / `spawns_last_hour` / `child_state_counts`
/ `judge_fallback_rate` (substring match on verdict.tier) plus
zero placeholders for wait_* pending dedicated instrumentation.
Derived from new `storage.count_workstreams_by_state` +
`count_workstreams_since` aggregate helpers — no 10k-row
hydrated-select to compute a histogram. Ownership 404-mask
matches `coordinator_detail`.
- **Coordinator skill in inspect** — `CoordinatorManager.create`
resolves `skill` → `template_id` / `applied_version` via
`get_skill_by_name` + new `storage.count_skill_versions`
(replacing the SELECT-all-for-COUNT anti-pattern) and persists
them on the workstreams row. `/new` handler dispatches via
`asyncio.to_thread` so blocking storage calls don't stall the
event loop.
- **`wait_for_workstream(since=…)`** — optional prior-snapshot
hint; when supplied, the wait loop diffs each polled ws_id that
IS in `since_map` and exits on any change, independent of mode.
ws_ids absent from `since_map` fall through to the normal mode
condition — a disjoint since dict no longer silently exits the
wait on tick one.
- **`task_list.child_ws_id` referential cleanup** —
`CoordinatorClient.cleanup_dead_task_child_refs(ws_id)` holds the
same per-ws `_task_lock` as `task_list_*` so a close racing a
task_list write can't lose the mutation. `CoordinatorManager.close`
delegates. Final save-failure logs at `warning` instead of
`debug`.
Home view live-updates
----------------------
- **Active-coordinators via SSE** instead of a 5s poll —
`ClusterCollector.ensure_console_pseudo_node` +
`emit_console_ws_created / _closed / _state / _rename` plumbing;
`CoordinatorManager.create / open / close / eviction` +
`ConsoleCoordinatorUI.on_state_change / on_rename` all fan out
through the collector. The pseudo-node is exempt from the
discovery-loop eviction; rehydrate-path eviction now also emits
`console_ws_closed` for the evicted row so other tabs drop it
live. `app.js` reads coordinators from
`clusterState.nodes["console"]`; poller + back-compat shims
deleted (9 call sites). Overview / nodes list skip the
pseudo-node so it doesn't inflate cluster totals. Tenant-filtering
preserved by excluding the pseudo-node from
`collector.get_workstreams` so `/v1/api/cluster/workstreams` still
uses the existing tenant-filtered `_coordinator_rows` path.
`CoordinatorManager.NODE_ID` bound from
`ClusterCollector.CONSOLE_PSEUDO_NODE_ID` so the two literals
can't drift.
Frontend perf
-------------
- **Bulk cluster-ws live endpoint** — `GET /v1/api/cluster/ws/live?ids=`
returns `{results, denied, truncated}` (cap 50); coordinator.js
batches visible-row live-badge fetches into one bulk request per
~250ms window (replaces per-row /detail polling). Ownership
check routes through the empty-string-safe pattern (non-admin
with empty `caller_uid` doesn't match empty-owner rows).
Legacy chat-view class cleanup
------------------------------
- Drop the `.msg` / `.msg-user` / `.msg-assistant` / `.msg-tool` /
`.msg-error` / `.msg-info` / `.approval-block` / `.approval-tool`
/ `.approval-btn` / `.approval-badge` / `.approval-prompt` /
`.approval-feedback-input` / `.approval-actions` / `.pane-input` /
`.pane-input-area` / `.pane-input-row` / `.pane-attach` /
`.pane-attach-chip` / `.coord-msg` / `.coord-body` / `role-*` /
`btn-approve` / `btn-deny` / `btn-always` / `verdict-glow-*`
legacy dual-class names left over from the phase-4 migration.
Every JS className concatenation + querySelector + CSS selector
now uses the `ts-*` vocabulary from `shared_static/chat.css` (and
`ui/static/style.css` where the interactive-page extensions
live). Feature-specific class names that don't map to `ts-*`
stay — `msg-queued` / `msg-editing` / `msg-actions` / `msg-edit-*`
/ `msg-user-attach*` / `msg-user-text` / `queued-badge` /
`queued-dismiss` / `tool-name` / `tool-cmd` / `tool-diff` /
`tool-header` / `tool-preview`.
Designer nits
-------------
- **`.ui-btn--icon:focus-visible`** — new rule matching `.ui-btn`'s
`outline: 2px solid var(--accent); outline-offset: 1px` so the
compact icon variant gets the accent ring instead of the
browser-default outline.
- **Dropped speculative 701-880px composer wrap rule** — the flex
math at ≥701px fits comfortably in every desktop viewport, so the
mid-zone break rule was forcing a 2-line layout where the browser
wouldn't have wrapped naturally. The existing `<700px` full-stack
covers the original wrap observation.
- **`.verdict-badge` border-top** + **`.ch-row.highlight`
prefers-reduced-motion** — confirmed already in main; no
additional code change needed for phase 6.
Follow-up designer-review findings
----------------------------------
- Dropped `border-top` from `.ts-approval-badge` + `.ts-approval-body`
(chat.css's max-content width / flex-gap made them read as
truncated / floating lines).
- `var(--muted)` → `var(--fg-dim)` on denied tool names (undefined
token was silently failing).
- Dropped 3 dead `.ts-approval-badge.badge-*` rules + duplicated
`.ts-approval-btn:focus-visible` + dead `.reasoning` CSS rule +
`contains("reasoning")` JS guard.
- Dropped `tool-row` / `approval-header` / `btn-row` / `label` dead
legacy classes in the coordinator.
- Added `:focus-visible` to `.ch-row a.ws-link` + `.task-row` so
keyboard users get the accent ring on sidebar rows.
Cleanups
--------
- `_WAIT_REAL_TERMINAL_STATES` / `_WAIT_TERMINAL_STATES` /
`_WAIT_MAX_*` / `_WAIT_POLL_INTERVAL` hoisted to module level on
`coordinator_client` so `session.py` no longer reads a class
internal; ClassVar aliases kept for back-compat.
- `ConsoleCoordinatorUI` state/rename observers typed as
`Callable[[str], None] | None` instead of `Any`.
- SSE error renderer verified end-to-end (coordinator.js already
handles `case "error"` → `appendText`; no code change).
Tests
-----
- 26 new test cases: 11 for `_diff_since` + `cleanup_dead_task_child_refs`,
15 for `cluster_ws_live_bulk` + `coordinator_metrics`. Full suite
4345 passing (4319 base + 26 phase-6).
Gate: ruff + mypy + pytest -m "not live" (4345 passed) all clean.
* fix(coordinator): address PR #381 review feedback
Copilot comments:
- Cross-tenant aggregate leak in coordinator_metrics — the new
count_workstreams_by_state / count_workstreams_since aggregates
took parent_ws_id but not user_id, so a non-admin caller could
observe drifted / forged child rows that share parent_ws_id with
their coord but whose user_id drifted to another tenant. The
404-mask on coord ownership (_resolve_coordinator_or_404) is the
primary defense; this is defense-in-depth inside the aggregate
queries. Pass filter_user_id (None for admin, caller_uid for
non-admin) — matches coordinator_children's tenant-push-into-SQL
pattern.
- wait_for_workstream(since=…) docstring + tool schema were stale —
claimed "A missing entry counts as changed on first observation"
but the implementation ignores ws_ids absent from since_map to
prevent a disjoint since dict from silently exiting on tick one.
Rewrote both doc sites to match the actual semantics: only ws_ids
present in `since` participate in the diff-exit check; others fall
back to the normal mode-based completion condition.
- WAIT_TERMINAL_STATES comment drift — the comment claimed it was
"used by the resolved-count summary" but the summary counts only
WAIT_REAL_TERMINAL_STATES (denied is a rejection, not a
resolution). Rewrote the comment to describe the real usage:
mode='any' pure-denied short-circuit + mode='all' settle check.
github-code-quality (CodeQL):
- coordinator.js — dropped the dead typeof _renderWaitIndicator
guard + the typeof activeWaits guard around the reconnect clear.
Both symbols are defined in the same IIFE; the onopen handler
fires strictly AFTER IIFE execution finishes, so the guards
always evaluated to true. Removing the dead branching also
removes a CodeQL nit.
- Protocol-method `...` statements — the bot flagged the three new
methods (count_workstreams_by_state / count_workstreams_since /
count_skill_versions) with "statement has no effect". Left as
`...` to match the file's universal convention (216 `...` bodies
/ 0 `pass` bodies pre-change); swapping just the new methods to
`pass` would introduce inconsistency with every other Protocol
method. Resolved as non-actionable.
Test: new test_metrics_tenant_filter_excludes_forged_cross_tenant_child
covering the aggregate-query tenant filter with both a legitimate
alice child and a forged bob child sharing parent_ws_id. Non-admin
alice sees 1; admin sees 2.
Gate: ruff + mypy + pytest -m "not live" (4346 passed) all clean.
* fix(server,console): kind filter on saved-workstreams + closed coords on landing
Two independent bugs folded into one hotfix:
1. Coordinators leaking into the interactive UI's "saved workstreams"
sidebar. ``list_workstreams_with_history`` (SQLite + postgres) was
kind-agnostic — every coordinator row with conversation history came
back alongside interactive rows, and ``list_saved_workstreams``
serialized them uniformly with no kind field so the interactive UI
rendered coordinators as regular interactive entries.
Fix: add optional ``kind: WorkstreamKind | str | None = None`` kwarg
on ``list_workstreams_with_history`` (storage protocol + both
backends + the ``turnstone.core.memory`` helper). Pass
``kind=WorkstreamKind.INTERACTIVE`` from the /v1/api/workstreams/saved
handler so the interactive surface only sees interactive rows.
Default ``None`` preserves legacy all-kinds behaviour for any
other caller that wants both.
2. Closed coordinators vanish from the console landing page.
``_coordinator_rows`` in console/server.py built dashboard rows
exclusively from the in-memory ``CoordinatorManager`` registry,
which pops rows on ``close()``. The persisted storage row stays
(state='closed') but never reached the landing-page poller at
/v1/api/cluster/workstreams?node=console.
Fix: two-lane merge in ``_coordinator_rows``. The in-memory lane
(manager) stays authoritative for live session state (model /
model_alias / current state / tokens). A new persisted lane queries
``storage.list_workstreams(kind=COORDINATOR, user_id=uid, limit=200)``
and appends rows NOT already in the in-memory set — surfacing
closed / error / deleted coordinators so the operator can still
see them on the landing page. Ownership semantics unchanged —
non-admin callers only see their own tenant, admin-bypass via
admin.users/admin.roles honored on both lanes, empty-string
defense-in-depth matches _check_row_owner_or_404.
Tests:
- tests/test_storage_sqlite.py — two new tests: kind filter excludes
coordinators from the history list; string form of kind accepted
(matches the memory.py forwarding shape).
- tests/test_coordinator_endpoints.py — four new tests:
- closed coordinators from storage surface alongside active ones.
- in-memory row wins on ws_id dedup (live state authoritative).
- persisted rows respect tenant filter (non-admin, admin bypass).
- orphan rows (empty user_id) never leak to empty-sub callers.
Gate: ruff + mypy + pytest -m "not live" (4315 passed) all clean.
* fix(server,console): address Copilot review on PR #380
Three review comments folded in:
1. Tenancy leak in /v1/api/workstreams/saved — the handler called
list_workstreams_with_history without a user_id filter, so any
authenticated user could see every other user's saved workstream
aliases / titles / names. Fix:
- Add ``user_id: str | None = None`` kwarg to
list_workstreams_with_history on the protocol + both backends
(SQLite + postgres). Pushes the filter into SQL.
- memory.py helper forwards the kwarg.
- /v1/api/workstreams/saved reads ``_auth_scopes(request)``: a
service-scoped caller gets cluster-wide visibility (None), a
non-service caller with a blank ``sub`` returns an empty list,
otherwise the SQL filter is scoped to the caller's uid. Matches
the _visible_workstreams pattern used on /workstreams and
/dashboard.
2. Loose type annotation on the memory.py helper — ``kind: Any``
tightened to ``WorkstreamKind | str | None`` so mypy catches
invalid callers. WorkstreamKind was already imported in the
module.
3. Brittle positional indexing in _coordinator_rows persisted-rows
lane — ``row[10]`` for user_id encoded a column offset that would
silently corrupt the projection on any future SELECT reorder.
Drop the test-double fallback entirely; the storage-protocol
contract already requires SQLAlchemy Row with _mapping, and every
real caller (SQLite + postgres) provides it.
Tests:
- test_server_authz.py TestSavedWorkstreamsTenantScoping — four new
regression tests covering: non-service caller sees only own rows,
service scope sees cluster-wide, blank-sub non-service returns
empty, and coordinator rows excluded even for service callers.
Gate: ruff + mypy + pytest -m "not live" (4319 passed) all clean.
* fix(console): service scope on collector token + surface upstream 4xx
CRITICAL: the console's ClusterCollector ServiceTokenManager was
configured with only frozenset({"read"}) scope, but every upstream
node's /v1/api/events/global hard-gates on "service" scope (added in
PR #375 for cross-tenant authz hardening). Every console→upstream
SSE connect 403'd, the collector never populated node state, and the
failure was silent — node health, idle workstreams, and interactive-
kind workstream rows all disappeared from the console dashboard with
no user-visible error. The only surface was a log.debug line in the
collector's _node_sse_task that operators had to opt into via DEBUG
logging or browser DevTools.
Fix:
- Add "service" to the collector_token_mgr scopes
(turnstone/console/server.py). Matches the proxy_token_mgr (which
already has it) and the existing cli / admin / channel-gateway
service tokens. Restores /v1/api/events/global SSE subscription
and /v1/api/dashboard visibility (which silently tenant-filters
non-service callers to zero rows).
- Upgrade the 4xx path in _node_sse_task to log.warning with the
status code + 200-char body preview, so configuration-level
failures (scope misconfig, JWT secret mismatch, expired token)
show up in operator logs instead of being masked by the generic
except-block debug line. Keep transient network errors
(CancelledError, ConnectError) at debug so the log doesn't flood
during brief node restarts.
- Add reachable_reason field to NodeSnapshot + surface via
get_nodes / get_node_detail / get_snapshot (and the browser's
buildNodeInfoFromSnapshot). Operators now see the failure cause
on the cluster node list without tailing the log. Cleared on
successful reconnect in _apply_snapshot.
- Test coverage: test_server_authz.py TestGlobalEventsServiceGate
gains a positive-path test asserting that a token with exactly
the collector's scope set ({"read", "service"}) is accepted by
/v1/api/events/global. Locks in the scope contract so any future
rename breaks the test before it breaks the dashboard.
Gate: ruff + mypy + pytest -m "not live" (4309 passed) all clean.
* fix(console): address Copilot review on PR #379
Two review comments folded in:
- collector.py — bounded body read for 4xx SSE error previews. The
prior ``await source.response.aread()`` buffered the entire
upstream error body into memory just to log a 200-char preview; a
malicious / oversized upstream response (HTML error page, proxy-
generated body) could have forced the collector to download an
arbitrary amount of bytes. Iterate ``aiter_bytes()`` and stop once
the preview cap (256 bytes, ~200 chars after UTF-8 decode) is
satisfied.
- test_server_authz.py — tighten the service-scope positive test.
The prior ``assert resp.status_code != 403`` could pass on
unrelated 500s AND left an SSE stream open indefinitely. Send
``?expected_node_id=definitely-wrong-node-id`` so the handler
passes the scope gate, hits the post-auth node-identity check, and
returns 409. Now ``assert resp.status_code == 409`` proves the
scope contract precisely and terminates the request immediately.
Gate: ruff + mypy + pytest -m "not live" (4309 passed) all clean.
* feat(coordinator): phase 5 — harness-test polish + wait_for_workstream + judge fix
Closes the bug list surfaced by the 2026-04-17 coordinator harness test
plus the post-phase-4 wait_for_workstream ask, and folds in three
adjacent cleanups that landed in the same window. Tightens defense-in-
depth on the model-invoked mutating ops, fixes the LLM judge silent
no-op, kills the inspect-poll token burn, and rounds out a handful of
observability / docstring / spec gaps.
The session-factory pre-resolve at console/session_factory.py and
server.py was rewriting `judge.model` from an alias (e.g. `judge-mini`)
to the resolved underlying id (e.g. `gpt-5-mini`). IntentJudge then
checked `model_registry.has_alias(config.model)`, found nothing, and
fell back to the SESSION's provider/client with that bare model id —
silent `llm_fallback / "did not return a verdict"` whenever the
coordinator and judge alias resolved to different providers.
Pass the alias through unchanged; IntentJudge's existing alias-
resolution path picks up the matching client + provider. Validate
the alias exists so an obvious typo still surfaces, but don't replace
the model field.
Regression: `test_alias_uses_registry_provider_not_session_provider`
constructs an alias whose provider differs from the session's and
asserts the judge picks up the alias's provider/client/model;
`test_coordinator_tool_call_returns_llm_verdict_not_fallback` asserts
the verdict tier is `llm` (not `llm_fallback`) on the happy path.
New `cancel_workstream` tool (approval required, primary_key=ws_id) —
cancels in-flight generation, unblocks any pending approval / plan,
moves the child to idle, leaves the row in storage so a fresh
send_to_workstream lands cleanly. Re-uses the existing
`/v1/api/route/cancel` route + `route.cancel` audit namespace; no
new server endpoint.
`CoordinatorClient.cancel/close_workstream/delete/send` now enforce
a tenant guard inline (`_is_own_subtree`) — only the coordinator
itself or one of its own children is targetable. Foreign ids return
the same 404-shape inspect/wait_for_workstream use, so the model
can't distinguish foreign from missing (no existence oracle).
Defense-in-depth — the upstream node enforcement is the perimeter,
this is the second line.
`list_workstreams` advertised `state="deleted"` and an
`include_closed=true` that surfaced deleted rows. Hard-deletes
cascade the workstream + conversation rows out of storage, so
deleted is unreachable in normal operation. Doc-only fix; the
synthetic-test path that registers `state="deleted"` rows still
works (terminal-state filter still excludes them via
`_terminal_states = {"closed", "deleted"}` in list_children).
Documented that the 120s service-registry heartbeat window means a
node returned by list_nodes can drop out before a follow-up
`spawn_workstream(target_node=…)` lands — the spawn fails with "No
available node for routing" rather than falling back. Two-line
clarification on each tool. No code change (a code fallback is a
bigger discussion deferred to 1.6).
`close_workstream` accepts `reason`; the upstream server handler now
persists it to `workstream_config.close_reason` (capped at 512 BYTES,
sliced on UTF-8 not code points so a CJK / emoji-heavy payload can't
4× the documented budget). `CoordinatorClient.inspect()` reads it
and surfaces as `close_reason` in the result dict — only for
terminal-state children (closed/error/deleted) so the live-child hot
path doesn't pay a per-inspect DB round-trip.
Tests: server-side persistence covers success / no-reason /
length-cap / non-string / storage-failure / multi-byte-utf8 paths;
client-side surface covers terminal vs. live workstreams.
For idle children whose node-dashboard live counter is 0 (the live
block only surfaces in-flight token counters), fall back to
`SUM(prompt_tokens + completion_tokens)` from `usage_events` so the
inspect output reflects cumulative spend.
New `storage.sum_workstream_tokens(ws_id) -> int` on the protocol +
both backends. The fallback is folded INTO `_fetch_cluster_live` so
the merged live block (with persisted total applied) is what gets
cached — back-to-back inspects of an idle child amortize through
the existing 2s LRU cache instead of each firing a fresh aggregation.
`CoordinatorClient.list_skills()` now projects `allowed_tools` per
skill — capped at 20 with a `+N more` sentinel so a skill that
whitelists a wide MCP surface doesn't bloat the per-row payload.
Reads the existing `prompt_templates.allowed_tools` column; no
storage change. Coordinators no longer have to guess what tools a
skill brings.
`route_create` now sets `routing_strategy: "hash_ring" | "target_node"
| "resume"` on the spawn response so the coordinator's spawn
response (and the `spawn_workstream` tool output) carries why a
given node was chosen. 3 lines + 3 covering tests in
test_console_routing_proxy.py.
New coordinator tool `wait_for_workstream(ws_ids, timeout=60,
mode='any'|'all')` that absorbs the wait into a single tool call —
the model sees one call + one result regardless of how long the
children take. Kills the busy-poll inspect loop that burned 20+
turns on a 3-child fan-out.
Storage-poll loop with batched primitives —
`get_workstreams_batch` + `sum_workstream_tokens_batch` issue exactly
two storage calls per tick regardless of N. At the cap (32 ws_ids /
600s / 0.5s tick) that's ~2400 round-trips for a full wait, down
from ~38k under the naive per-id shape.
Validation single-source-of-truth: the client owns mode whitelist,
ws_ids dedup + cap, timeout coerce + clamp. The session preparer
is a thin pass-through that builds the header + dispatches; bad
input surfaces at exec time as a tool error via `result.get("error")`.
Tenant-isolation collapse: missing-row and cross-tenant cases both
return `state="denied"` so wait can't be used as an existence oracle
(matches the 404-mask contract `inspect` uses).
Prompt-side: tools_coordinator.md adds a `wait_for_workstream`
pattern + an explicit "PREFER wait_for_workstream OVER a loop of
inspect_workstream" line in the workflow-shape section.
Replaces the quote-bracketed substring LIKE/ILIKE pattern with proper
JSON-array containment. The previous shape effectively did
`LOWER(tags) LIKE '%"<lower-tag>"%'`, which broke for tag values
containing `"` (the JSON encoder escapes it to `\"` and the literal-
substring search misses), `\` (encoded as `\\`), or non-ASCII
characters that the encoder rendered as `\uXXXX`. Also exposed a
small spoofing surface — `tags=["foo\","bar"]` would have matched a
query for `bar`. Real-world tag values are alphanumeric+dash today
so it hadn't fired in production, but the fix is small.
- SQLite: `EXISTS (SELECT 1 FROM json_each(prompt_templates.tags)
WHERE lower(value) = lower(:tag))` (JSON1 extension; SQLite 3.38+).
- PostgreSQL: `EXISTS (SELECT 1 FROM jsonb_array_elements_text(
prompt_templates.tags::jsonb) AS jat(elem) WHERE lower(jat.elem) =
lower(:tag))`.
Three new tests prove the substring pattern was broken for
quoted / backslash / unicode tag values; the existing case-fold +
wildcard tests continue to pin the contract.
Phase 1 added the coordinator workstream API; phase 2 added only
`/open` to the OpenAPI catalog and missed every other coordinator
endpoint plus phase 3's `/children`, `/tasks`, and the
`/cluster/ws/{ws_id}/detail` aggregator. SDK consumers + operators
browsing `/docs` couldn't discover the surface. Doc-only addition:
12 endpoints + 9 new Pydantic models, all under the `Coordinator`
OpenAPI tag so /docs groups them together.
Sidebar re-fetches `GET /tasks` on every `task_list` `tool_result`
SSE event. A model that runs `add → list` (or any back-to-back
mutation pair) double-fetches the same envelope. Coalesced into
one fetch per 150ms window via a new `loadTasksDebounced` wrapper;
direct UI actions (refresh button, page load) keep calling
`loadTasks` directly so user clicks aren't delayed.
- `ruff check turnstone tests` — clean
- `mypy turnstone` — clean (157 source files)
- `pytest -m "not live"` — 4284 passed, 3 deselected (was 4226 on
main; +58 new tests across coordinator client, tools, judge,
storage, console routing proxy, server close-handler,
storage_skills_filtered, OpenAPI catalog, server close-reason
persistence)
- New tools added: 2 (cancel_workstream, wait_for_workstream) —
TOOLS count 28 → 30; coordinator subset 9 → 11; auto_approve adds
wait_for_workstream; primary_key adds cancel_workstream
- New OpenAPI endpoints: 12 (every phase-1/2/3 coordinator route +
the cluster-inspect aggregator)
- New storage protocol methods: 3 (sum_workstream_tokens,
sum_workstream_tokens_batch, get_workstreams_batch)
All phase 1 / 2 / 3 / 4 invariants preserved: COORDINATOR_TOOLS /
INTERACTIVE_TOOLS disjoint; coordinator sessions have no MCP surface;
list-style tools return {items, truncated}; route-proxy emits
route.<action> audit on 2xx; 404-mask on ownership failures; tenant
filters pushed into SQL; per-coordinator JWT carries scope context.
* fix(coordinator): address Copilot review on PR #378
Three valid Copilot findings on the wait_for_workstream surface:
1. ``wait_for_workstream.json`` description claimed the tool returns a
top-level mapping ``ws_id -> {state, tokens, updated}`` plus
elapsed/complete/mode at the same level, but the actual shape is
``{results: {ws_id: {...}}, elapsed, complete, mode}``. Description
now matches the implementation. Also adds ``deleted`` to the
advertised terminal-state list (it's in ``_WAIT_REAL_TERMINAL_STATES``;
the doc and runtime now agree).
2. ``CoordinatorClient.wait_for_workstream`` docstring listed
``idle / error / closed`` as the real terminal set but the constant
includes ``deleted``. Same fix — list ``deleted`` with a parenthetical
noting it's unreachable in normal operation (hard-delete cascades the
row).
3. Storage protocol docstring math: ``sum_workstream_tokens_batch``
claimed "from ~38k to ~1200" round-trips per wait at the cap, but
``wait_for_workstream`` issues TWO storage calls per tick
(``get_workstreams_batch`` + this one), so 1200 ticks × 2 = ~2400.
Updated to "~2400" with the math spelled out.
Also a clean rebase onto today's main (PR #377 — the rebalancer node_id
snapshot doc — landed since phase 5's last push). Single conflict in
``inspect_workstream.json`` resolved by keeping both notes (rebalancer
node_id binding semantics + the new ``close_reason`` surface from phase
5); ``spawn_workstream.json`` auto-merged.
The github-code-quality bot also flagged three items on
``_protocol.py`` asking to replace ``...`` with ``pass`` in Protocol
method bodies. Refuted: ``...`` is the canonical PEP 544 idiom for
Protocol method bodies and the rest of the file uses it consistently.
The bot's lint rule misfires for ``Protocol`` classes.
Verification:
- ``ruff check turnstone tests`` clean
- ``mypy turnstone`` clean (158 source files)
- ``pytest -m "not live"`` — 4308 passed, 3 deselected (no test count
change; pure doc/comment edits)
Phase 3 fixed spawn_workstream's response to return the storage-
authoritative node_id at spawn time, but neither tool description
mentioned that the cluster rebalancer can migrate the workstream to
a different node afterwards. A coordinator that cached the
spawn-time node_id for a long-running callback would silently dispatch
to a node that no longer owns the workstream.
- spawn_workstream: ``node_id`` is a POINT-IN-TIME snapshot at spawn;
re-read with inspect_workstream when you need the current binding.
- inspect_workstream: ``node_id`` is the CURRENT (storage-authoritative)
binding; reflects any rebalancer migration that happened since spawn.
Pure description edit — no schema or runtime change.
Third and final PR of the retrospective-review series. Addresses the
remaining bug / perf / doc findings from the original multi-stage review
plus the three inline comments left on #374 and #375.
From the original review:
- bug-3: delete_workstream now nulls out parent_ws_id on every child
row before dropping the target — previously, deleting a coordinator
left orphaned parent_ws_id pointers and list_workstreams(parent_ws_id=
<deleted>) kept returning ghost-parented rows. Fix lives at the
storage edge so both SQLite and PostgreSQL benefit without a schema
migration.
- perf-1 / perf-2 / perf-3: new migration 041 drops the low-cardinality
idx_workstreams_kind outright, rebuilds idx_workstreams_parent as a
partial index (WHERE parent_ws_id IS NOT NULL) to halve its btree,
and uses CREATE INDEX CONCURRENTLY on postgres so the rebuild
doesn't take ACCESS EXCLUSIVE on populated tables. Dialect-guarded;
sqlite path is a straight partial CREATE INDEX.
- perf-5: _rebuild_children_from_storage bumps its limit sentinel to
10_000 and logs a warning when the cap is hit instead of silently
truncating the tail on every console cold-start.
- q-2: turnstone.core.memory.list_workstreams wrapper deleted (zero
live callers; PR #374 kept it forward-compatible with the new
kwargs as a stepping stone).
- q-5: migration 039's docstring now warns operators that downgrade
drops parent_ws_id irreversibly and notes the 041 dependency.
- q-7: GET /v1/api/workstreams row shape now includes kind +
parent_ws_id to match /v1/api/dashboard; the Pydantic
WorkstreamInfo schema follows so SDK consumers see the same fields.
Inline review comments:
- #374 (copilot): console/server.py::coordinator_children now pushes
user_id into the SQL filter for non-admin callers, so forged /
migration-era rows with matching parent_ws_id but a different
owner can't leak through. Admins bypass the filter — they're
expected to see the full subtree.
- #375 (copilot, delete handler): storage.get_workstream(ws_id) for
the audit snapshot moved inside the try: block so a transient DB
error surfaces through the endpoint's redacted 500 handler instead
of an unhandled exception.
- #375 (copilot, _require_ws_access): added optional mgr= kwarg —
when the workstream is live in the in-memory manager, trust its
cached user_id instead of round-tripping storage. In-memory-only
handlers (approve / plan / cancel / command / close / events_sse /
refresh-title / set-title) pass mgr= so they stay functional
during transient DB outages and skip one query on the hot path.
Storage-backed handlers (/delete, /open) omit mgr= and keep the
storage path for persisted-but-not-loaded rows.
Tests:
- tests/test_workstream_kind.py adds regression tests for the cascade
null-out on delete and the new user_id SQL filter.
- tests/test_workstream_endpoints.py updated so the title-handler
tests exercise the in-memory fast path (MagicMock manager returning
None falls through to storage; explicit ws.user_id set where the
mock ws is used).
Lint (ruff), typecheck (strict mypy), pytest -m 'not live' all green
(4209 passing).
Second of three PRs addressing the retrospective review of the
turnstone-server interactive-kind feature. The first (PR #374) put
the structural pieces in place — WorkstreamKind enum + user_id
kwarg on the storage protocol. This PR uses them to close the
handler-level ownership gaps that shipped under the prior design.
- sec-1: approve / plan_feedback / cancel_generation / command now
call _require_ws_access before touching the target UI. Previously
any authenticated user could resolve pending tool-approvals on
another tenant's workstream — RCE-adjacent because the attacker
could approve destructive operations the victim would have denied.
- sec-2: /v1/api/workstreams/{ws_id}/delete now gates on ownership
AND writes a workstream.deleted audit event. Previously any
authenticated user could destroy any other tenant's workstream,
conversations, and attachments in one call with no tamper-evident
trail.
- sec-3: /v1/api/events (per-ws SSE) gates before _register_listener
so non-owners can't subscribe to another tenant's message / tool /
approval stream.
- sec-4 / sec-5: /v1/api/workstreams and /v1/api/dashboard filter
to the caller's tenant view via a new _visible_workstreams helper;
service-scoped tokens (cluster / routing proxy) keep the full view.
- sec-6: /v1/api/events/global requires service scope. The global
snapshot carries cross-tenant workstream inventory and was never
intended for end-user browsers.
- sec-7: /v1/api/workstreams/{ws_id}/open verifies the caller is
the stored owner (or holds service scope) before rehydrating.
Returns 404 on mismatch — existence isn't enumerable by response
code.
- sec-8 / sec-9: /workstreams/close, /refresh-title, /title all gate
on ownership. Cross-tenant close aborts the victim's running
generation; cross-tenant rename is a phishing / denial-of-use
vector in list / dashboard responses.
- sec-11: workstream.created / .deleted / .closed / .opened now
land in the audit_events table with kind + parent_ws_id detail,
so forensic review can reconstruct lifecycle even after the row
is gone.
- q-4: new tests/test_server_authz.py covers every gate above via
TestClient, plus the PR #1 HTTP-boundary kind-validation branches
that had no regression coverage (coordinator / unknown-kind / 400,
cross-tenant parent_ws_id / 403, non-interactive open / 400).
- q-3: test_workstream_kind.py now uses the conftest storage fixture
so it runs against both SQLite and PostgreSQL under
--storage-backend=postgresql, closing the sqlite↔postgres drift
risk the prior review flagged. Added storage-edge ValueError and
user_id SQL filter tests alongside.
Tests, lint (ruff), typecheck (strict mypy) all green. Stacked on
PR #374 — merges after that lands.
Foundation PR for the multi-stage-review follow-up. Introduces a
single source of truth for workstream kind values and pushes tenant
scoping into the storage protocol so list callers can't forget to
filter client-side.
- WorkstreamKind(StrEnum) replaces bare "interactive" / "coordinator"
literals across 17 production modules. Strict mypy narrows every
internal call site; raw strings still work at wide boundaries
(HTTP body, DB row) via WorkstreamKind(raw) parse at the edge.
- StorageBackend.list_workstreams(..., user_id=None) adds a SQL-level
WHERE user_id = :user_id gate on both sqlite and postgres impls.
Memory wrapper forwards the new filters.
- register_workstream now validates kind at the storage edge so SDK /
restore / internal callers can't silently corrupt the NOT NULL
column with empty / mis-cased / unknown values.
- WebUI.__init__ normalizes empty-string parent_ws_id to None, matching
the storage-edge and WorkstreamManager invariants.
- POST /v1/api/workstreams/new parses body["kind"] through the enum
and returns 400 on unknown kinds instead of silent coercion.
Absorbs bug-1, bug-2, bug-4/q-6, q-1, q-8, and partial q-2 (wrapper
signature forwards the new filters; full deletion of the unused
wrapper stays in the cleanup PR).
* feat(ui): phase 4 — chat-UX unification + coordinator-first console landing
Phase 4 unifies the three turnstone UIs (server-node chat, console
dashboard, coordinator page) around a shared design-system layer,
promotes coordinator sessions to first-class citizens on the console
landing, and folds the chat-view itself onto a shared vocabulary so
the two chat pages no longer reinvent messages / approvals / composer /
header / sidebar chrome from scratch.
## Shared static consolidation
- turnstone/shared_static/renderer.js — consolidates the two copies
(ui/static/ + console/static/coordinator/) into one. Adds
streamingRender / streamingRenderFinalize helpers with
requestAnimationFrame coalescing + per-element buffer cache so both
chat views re-render the streamed markdown smoothly without the
prior "plain-text → final pop" on the coordinator page and without
thrashing renderMarkdown + DOM replacement faster than the paint
cycle. renderMarkdown stays the trust boundary for innerHTML
assignment (escapeHtml internal); postRenderMarkdown (syntax
highlighting, mermaid, KaTeX) is deferred to finalize.
- turnstone/shared_static/ui-base.css — flat form-control + button +
state-glyph + pill + panel vocabulary on top of base.css. Sizes in
px to match the 11/12/13px scale used elsewhere. Namespace rubric
documented inline (.ui-* shared controls; .dash-* dashboard legacy;
.ts-* chat vocabulary; page-local stays unprefixed).
- turnstone/shared_static/chat.css (new) — chat-view component
vocabulary: .ts-msg (user / assistant / reasoning / tool / error /
info), .ts-msg-actions floating toolbar, .ts-approval (inline +
batch layout hooks sharing a visual language), .ts-verdict-badge,
.ts-composer shell, .ts-header shell, .ts-sidebar shell. Mobile +
reduced-motion covered.
## Interactive server UI migration
- turnstone/ui/static/app.js — dual-class adoption of .ts-msg + .ts-
approval + .ts-composer + .ts-msg-actions alongside existing class
names so feature-specific rules (.msg-user-text, .msg-queued,
.msg-editing, .msg-action-btn toolbar, attachment chips, verdict
details, media embeds, plan inline) keep working while the shared
chat.css baseline takes over padding / border / typography.
- Left-aligned user messages: .msg-user loses align-self: flex-end
and .msg-assistant loses align-self: flex-start. Both roles now
render as single-column blocks distinguished by left-border colour
(amber for user, neutral for assistant, dashed for reasoning,
mono + code-bg for tool, red for error) per the locked design.
- style.css trimmed: .msg / .msg-info / .msg-error baseline rules
dropped (chat.css provides); all other feature rules intact.
- index.html links /shared/chat.css and tags the header with
.ts-header + .ts-header-title.
## Coordinator page migration
- coordinator.js appendMsg drops the visible .role-label <div> per
the hybrid no-labels design, preserving the role text on
data-ts-role + aria-label so screen readers and SSE dedup-by-call-
id continue to see meaningful labels. Adds .ts-msg + .ts-msg--*
variants + .ts-msg-body onto the existing .coord-msg / .coord-body
elements.
- coordinator/index.html adopts .ts-header, .ts-header-title,
.ts-header-spacer, .ts-header-status on the header; .ts-approval +
.ts-approval--batch on the pinned bar with .ts-approval-btn
variants on the buttons; .ts-composer + .ts-composer-input +
.ts-composer-send on the composer; .ts-sidebar + .ts-sidebar-
section + .ts-sidebar-section-heading on the children + tasks
sidebar. Inline <style> pared from ~300 to ~130 lines — only
genuinely coordinator-specific layout (flex wiring, sidebar list
rows, mobile accordion breakpoint) remains.
- Nits fixed along the way: .ch-row .glyph-thinking recoloured cyan
to match the shared .ui-glyph vocabulary; .task-row .status-done
lost its 0.7 opacity (colour already signals done; opacity
reduced contrast for no gain).
## Console landing + admin panel redesign (phase 4 scope-expansion)
Replaces the node-list-first console landing with a coordinator-first
layout on a new #view-home pane:
- #coord-composer-panel — persistent "Start a new coordinator task"
composer (textarea + optional name + skill dropdown + submit).
Permission-gated on admin.coordinator (same rule the +coordinator
header button uses). Pre-probes GET /v1/api/coordinator on init
and after login (bug-1 fix) so a 503 (no coordinator.model_alias
resolvable) surfaces as a remediation banner linking to Admin →
Models instead of failing on submit. Probe gates on r.ok instead
of r.status !== 503 (bug-2 fix) so auth/permission errors don't
incorrectly flip the banner to ready. Composer and modal share a
_createCoordinator helper (q-1 fix) — POST + redirect + error-
handling tail is not forked.
- #active-coordinators — SSE-driven list of kind=="coordinator"
workstreams, rendered through the shared _renderWsRow helper so
state glyphs + child-count badges match the existing tree view.
- #cluster-summary-compact — one-line aggregate. Clicking expands
into the legacy #view-overview via showOverview() so deep-link
callers of ?view=overview / ?view=node / ?view=filtered keep
working unchanged.
View switching consolidated into a _setLandingView helper so every
show* / drillDown* function toggles the four landing panes through
one call path. Default currentView flipped from "overview" to
"home"; popstate + init history.replaceState land on {view: "home"}.
patchClusterState preserves kind / parent_ws_id / user_id on
ws_created events (phase 3 invariant) so the active-coordinators
list picks up new coordinators immediately without a snapshot
refetch.
Header H1 is now a home link so operators have a single-click path
back to the coordinator landing from any drill-down / admin view.
## Design polish (review pipeline fixes)
- .home-panel-title dropped from 13px/accent to 11px/fg-dim so it
sits in the same heading tier as .ui-section-heading / .dash-
header-title / .home-section-title instead of outweighing them
(dsn-3).
- .home-composer-banner recoloured from amber-on-amber-glow to
bg-surface + 1px yellow border + fg-bright text + accent link
with thicker underline (dsn-1).
- .ui-pill--done dropped the 0.75 opacity — colour signals done,
opacity reduced contrast for no gain (dsn-6).
- .ui-heading fleshed out with --sm/--md/--lg tiers so the utility
actually conveys size (dsn-12).
- ui-base.css size scale moved from rem to px matching 11/12/13px
(dsn-2).
* fixup: address Copilot feedback on PR #373
- admin.js: drop the stale `#view-overview` display:none mutation in
showAdmin. #view-overview is now nested inside #view-home and
toggled via the `hidden` attribute; inline display:none here would
stick after returning to home and suppress the cluster-details
expand.
- app.js: reword the _renderHomeView token-bucket fingerprint comment
to match the actual `Math.floor(tokens / 100)` bucketing — the
prior comment said "thousands / sub-thousand drift".
* feat(coordinator): tree-view UI, cluster-wide live inspect, dashboard grouping — phase 3
Closes out the 1.5 coordinator UX surface: a right-sidebar tree view at
/coordinator/{ws_id} showing spawned children + task list, a new
cluster-wide live inspect endpoint that powers the tree's live badges,
and 2-level dashboard tree grouping that nests spawned children under
their coordinator parent.
## Cluster-wide live `inspect_workstream`
New `GET /v1/api/cluster/ws/{ws_id}/detail` on the console, gated by a
new `admin.cluster.inspect` permission (unassigned to any builtin role;
operators opt in). Aggregates `storage.get_workstream` with a
short-timeout (2s) HTTP fetch against the owning node's
`/v1/api/dashboard`. Coordinator-hosted workstreams get their `live`
block from the in-process `CoordinatorManager` instead of a proxy hop.
Response shape `{persisted, live, messages}` — `live: null` on node
unreachability / 5xx / missing-entry with status 200 so the UI can
degrade gracefully without an error state. Correlation-id masks
unexpected exceptions. 404-masks cross-tenant reads (non-admin
callers see only their own workstreams).
`CoordinatorClient.inspect()` best-effort merges the `live` block onto
its storage snapshot so the model-facing `inspect_workstream` tool
gains a `live` key without any schema change. Model-facing tool
schema stays identical.
## Tree-view UI
New right sidebar at `/coordinator/{ws_id}` with a 2-level children
tree + the phase-2 task list.
Backend:
- New `GET /v1/api/coordinator/{ws_id}/children` returns
`{items, truncated}` — identical row shape to the `list_children`
tool — filtered via `storage.list_workstreams(parent_ws_id=..., kind=None)`.
- New `GET /v1/api/coordinator/{ws_id}/tasks` returns the
`{version, tasks}` envelope via the shared module-level
`load_task_envelope` decoder (extracted from `CoordinatorClient`
so both the tool path and the UI read share corruption semantics).
Corrupt envelopes return an empty list for UI resilience — the
`task_list` tool remains the authoritative write + error path.
- `CoordinatorManager` subscribes to the `ClusterCollector`'s
listener channel from the console lifespan and dispatches filtered
`child_ws_created / child_ws_state / child_ws_closed / child_ws_rename`
events onto each coordinator's SSE stream. Filter authoritative
on the server via a per-coordinator child-ws_id registry populated
lazily on `open()` from storage and incrementally on `ws_created`
events; cleared on `close()` / eviction. One SSE connection per
client, no client-side filtering.
Frontend:
- DOM-method-only child-row rendering (no innerHTML of user content).
- State glyph vocabulary (● running / ◐ thinking / ⚠ attention /
✗ error / ○ idle) plus text labels — WCAG 1.4.1 carries info in
both glyph and label.
- Live badges (tokens + pending-approval pip) fetched via
`/cluster/ws/{ws_id}/detail` with a 5s TTL cache and 250ms debounce
per child. One request per state change, not per second.
- SSE child events update in place; renderChildren() re-sorts.
- Mobile (<700px) sidebar collapses to an accordion above the chat
with a toggle button flipping aria-expanded; a `.highlight` flash
marks task→child scroll targets; `prefers-reduced-motion` respected.
- Deep-link child rows to `/node/{node_id}/?ws_id=<child>` via
`<a target="_blank" rel="noopener">` with encodeURIComponent on
regex-validated ids.
## Dashboard tree grouping
Cluster dashboard rows now group by `parent_ws_id`. Coordinator rows
(`kind == "coordinator"` or children present) get an expand/collapse
caret (button with `aria-expanded`); collapsed shows "(N children)".
Expanded renders children indented as sibling rows with a left-border
gutter. Orphaned children (parent missing or closed) render at top
level with a muted "orphan" badge. Expansion state persisted in
`localStorage` keyed per coordinator ws_id so operator preference
survives reloads. Coordinator rows deep-link to `/coordinator/{id}`;
node-backed workstreams keep their existing proxy deep-link.
Per-node `ws_created / ws_state / ws_activity` SSE event payloads
gained `parent_ws_id` + `kind` so the collector can propagate them
through its fan-out to browser clients without a second lookup;
`_build_node_snapshot` and `/v1/api/dashboard` rows include the
same. Coordinators (which don't live on cluster nodes) merge into
`/cluster/workstreams` via a new `_coordinator_rows` helper that
threads them through the collector's `get_workstreams(extra_rows=...)`
parameter — extras share the filter / sort / paginate pipeline with
node-backed rows.
## Tests
- `tests/test_coordinator_endpoints.py` — 19 new cases covering
children (empty / populated / ownership 404 / admin bypass /
invalid ws_id / truncation), tasks (empty / round-trip / corrupt /
ownership), and cluster-inspect (auth gates / 400 / 404 / ownership /
coordinator self-path / unloaded-live-null / message-limit clamp).
- `tests/test_coordinator_manager.py` — 8 new cases covering registry
bootstrap on create + open, dispatch for each event type,
unrelated-parent filtering, shutdown idempotency.
- `tests/test_console.py` — existing `cluster_workstreams` assert
updated for the new `extra_rows` kwarg.
## Verification
- `ruff check turnstone tests` clean.
- `mypy turnstone` clean.
- `pytest -m "not live"` — 4184 passed, 3 deselected.
* fix(coordinator): race in dispatch + ui_factory kwarg filtering — PR #370 review
Addresses feedback from the GitHub Copilot + code-quality bot review
passes on PR #370.
## Race in _dispatch_child_event ws_created branch
Copilot flagged a TOCTOU where the lock-free read of
``self._active_coords`` (line 912) could see the parent coordinator,
then ``close()`` / eviction pops ``_children[parent]`` + drops the
coord from ``_active_coords`` before we acquire ``_children_lock``,
and then ``setdefault(parent, set())`` resurrects the entry —
leaking the registry key forever and fanning events to a closed UI.
Fix: re-check ``parent in self._active_coords`` inside
``_children_lock``. The reference swap is still atomic; holding
``_children_lock`` and re-reading the snapshot catches the race
without serializing back through ``self._lock``.
Regression test: create → close → dispatch a ws_created → assert
neither ``_children`` nor ``_active_coords`` regained the entry.
## ui_factory kwarg filtering via inspect.signature
code-quality bot flagged that the previous ``try ui_factory(…, kind=,
parent_ws_id=) except TypeError`` dance fired on every call with
legacy test factories (``lambda wid: WebUI(ws_id=wid)``) — wasteful
and masks real signature mismatches.
Fix: inspect the factory's signature and only pass kwargs it
actually accepts (explicit param name OR ``**kwargs`` absorber).
Keep a conservative ``except TypeError`` fallback for C-callables
and odd signatures ``inspect`` can't introspect.
Copilot also flagged a comment mismatch (the old comment said
"KeyError on **kwargs" — it's ``TypeError``, which is what the code
caught). The rewritten comment is correct.
## Nit: side-effect in assert
code-quality bot flagged ``assert mgr.close(ws.id)`` in
test_coordinator_manager.py. Split into two statements.
## Verification
- ``ruff check`` clean.
- ``mypy turnstone`` clean.
- ``pytest -m "not live"`` — 4223 passed, 3 deselected, 0 failed.
* feat(coordinator): audit middleware on routing proxy — phase 2
Adds per-tool-call audit attribution to the multi-node routing proxy
handlers so coordinator → server hops land observable rows in
``audit_events``. Phase 1 preserved the ``src="coordinator"`` claim
through ``_proxy_auth_headers``'s upstream re-mint; this commit
closes the recording side. Was the last real security gap from
phase 1 — an enterprise deployment with ``admin.coordinator``
granted got only the three console-side
``coordinator.{create,close,cancel}`` rows; per-tool-call
attribution was missing.
## Action-naming scheme
route.workstream.create POST /v1/api/route/workstreams/new
route.workstream.send POST /v1/api/route/send
route.workstream.close POST /v1/api/route/workstreams/close
route.workstream.delete POST /v1/api/route/workstreams/delete
route.approve POST /v1/api/route/approve
route.cancel POST /v1/api/route/cancel
route.command POST /v1/api/route/command
route.plan POST /v1/api/route/plan
Action-name conventions documented in ``turnstone/core/audit.py``
module docstring alongside the existing namespaces — the docstring
is now ``<resource>.<verb>`` shaped (non-exhaustive) rather than
trying to enumerate every prefix.
## Recording rules
- ``record_audit()`` fires only on a 2xx upstream response. 4xx/5xx
are observable via ``_record_route``'s metrics path; doubling the
audit-events table size for failure rows would dilute signal
without giving operators much extra value.
- ``detail`` JSON carries ``{src, node_id, coord_ws_id?}`` — ``src``
lands verbatim from ``auth.token_source`` so non-coordinator
origins (``"jwt"``, ``"console-proxy"``) also get attribution;
``coord_ws_id`` only appears when the inbound JWT carried it.
- Wrapped in ``try/except`` + ``log.debug("route.audit_failed", ...)``
defence-in-depth. ``record_audit`` itself is fire-and-forget;
the outer try guards against a programmer error in the call site.
## Routing-proxy specifics
- ``route_create``: emits at the post-multipart/JSON convergence
``if resp.status_code == 200`` block. Both branches set
``audit_ws_id`` correctly — multipart from the query-string ws_id,
JSON from ``body["ws_id"]`` (post-503-retry) or ``body["resume_ws"]``.
- ``route_proxy``: emits the URL-method-mapped action. ``ref`` is
reassigned to ``new_ref`` after a successful 404→cache-refresh
retry so audit attribution uses the retried node, not the failed
first node.
- ``route_workstream_delete``: emits on 2xx using the ws_id from
the request body.
- ``route_attachment_proxy``: out of scope (upstream attachment
endpoints emit their own ``workstream.attachment.*`` rows;
auditing here would double-count).
## Tests
16 new tests in ``tests/test_route_proxy_audit.py`` covering:
- Coordinator-origin emission with full detail payload.
- 502 / 400 / 503-retry-final-node-id paths.
- Parametrised method→action mapping for the 6 ``route_proxy`` URLs.
- Plain-JWT origin (no ``coord_ws_id`` in detail).
- Delete handler 2xx + 502.
- Audit-storage exception swallowed (proxied response unchanged).
- ``auth_storage`` absent → no-op (existing route-handler tests
unaffected).
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` 4087 passed,
3 deselected (live-backend), 0 failed.
* feat(coordinator): discovery tools and /open parity — list_nodes, list_skills, POST /coordinator/{ws_id}/open
Adds the read-side surface coordinators need to make informed
orchestration decisions plus an explicit rehydration endpoint
matching the server's ``POST /v1/api/workstreams/{ws_id}/open``.
## list_nodes (auto-approved)
``list_nodes(filters={key: value, ...})`` reads ``node_metadata`` via
``storage.filter_nodes_by_metadata`` + ``get_all_node_metadata`` —
one query each, no N+1. Each row carries its full metadata dict so
the coordinator has both auto keys (``arch`` / ``cpu_count`` /
``fqdn`` / ``hostname`` / ``os`` / ``os_release`` / ``python``;
always present) and operator-supplied user keys (``capability`` /
``region`` / ``tenant`` / ``role``) without a second round-trip.
Tool description enumerates the auto keys explicitly so the model
knows what's always available vs deployment-specific.
Storage stores metadata values as JSON-encoded strings (the write
path in ``server.py`` / ``admin.py`` / ``console/server.py`` all go
through ``json.dumps``). The client re-encodes filter values
before the stored-text comparison and decodes stored values before
returning them to the model — so ``{"capability": "gpu"}`` is the
natural form the model uses, not ``{"capability": "\"gpu\""}``.
Ints round-trip as ints.
Returns ``{nodes, truncated}``; ``truncated=True`` when the page
was full.
## list_skills (auto-approved)
``list_skills(category?, tag?, scan_status?, enabled_only?, limit?)``
surfaces the skill registry so coordinators can discover worker
profiles. New storage protocol method ``list_skills_filtered(...)``
on both SQLite and PostgreSQL backends pushes filters into SQL.
``tag`` filter matches against the JSON-array ``tags`` column with
quote-bracketed substring (``%"tag"%``) — quote-safe against
``foo`` vs ``foobar`` collisions on both backends.
Returns ``{skills, truncated}`` with ``name`` / ``category`` /
``tags`` (decoded to list) / ``version`` / ``description`` /
``model`` / ``enabled`` / ``scan_status`` / ``activation`` — the
discovery projection, not the full row.
## POST /v1/api/coordinator/{ws_id}/open
Explicit rehydration endpoint. Lazy ``GET`` rehydration works for
the UI; this gives SDK callers and operators a way to warm a
coordinator without browsing to it. Same ownership / 404-on-
mismatch / correlation-id-masked error semantics as
``coordinator_detail``. Returns ``{ws_id, name, already_loaded?}``.
Registered in ``turnstone/api/console_spec.py`` with a dedicated
``CoordinatorOpenResponse`` Pydantic model so the OpenAPI schema
matches the wire shape.
## Tests
- ``tests/test_storage_skills_filtered.py`` — 8 cases validated on
BOTH SQLite and PostgreSQL backends (``pytest --storage-backend
postgresql``). Covers no-filter ordering, category exact-match,
tag quote-safety (``"foo"`` matches ``["foo","bar"]`` but not
``["foobar"]``), scan_status, enabled_only, limit, AND semantics,
empty result.
- ``tests/test_coordinator_client.py`` — 11 new cases covering
node/skill shape decoding, JSON-encoded filter round-trip (the
``"gpu"`` vs ``'"gpu"'`` case), int filter encoding, truncation,
no-match empty, no N+1 (``get_prompt_template`` /
``get_node_metadata`` call counts asserted zero).
- ``tests/test_coordinator_tools.py`` — 11 new cases for
``_prepare``/``_exec`` dispatch, filter type-drop, limit clamping
(``limit=0`` falls back to 100, negatives clamp to 1),
truncation-signal summary.
- ``tests/test_coordinator_endpoints.py`` — 8 new cases for
``/open``: ``already_loaded`` on in-memory hit, 404 on ownership
mismatch, lazy rehydrate on miss, admin bypass, unknown ws_id,
503 on ``coord_mgr`` unavailable, 500 with correlation-id mask on
factory failure, 503 passthrough on ``ValueError``.
- ``tests/test_workstream_kind.py`` / ``test_tools_schema.py``
updated to include ``list_nodes`` and ``list_skills`` in the
disjoint-namespace regression guard and the tool-count check.
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` 4122 passed, 3
deselected (live-backend), 0 failed. Postgres backend storage
tests green (``pytest --storage-backend postgresql
tests/test_storage_skills_filtered.py`` 8 passed).
* feat(coordinator): task_list tool — persistent planning state
Adds a coordinator-only ``task_list`` tool persisted on the
coordinator's own ``workstream_config`` row. Gives coordinators a
scratch surface for work decomposition that survives restarts so the
UI can render planned-vs-done state once the tree view lands.
## Tool surface
``task_list(action, ...)`` with five actions:
- ``list`` auto-approved read. Returns ``{tasks, truncated}``;
truncated=True when the list exceeded the 200-row
page cap.
- ``add`` needs approval. ``title`` required; optional
``status`` and ``child_ws_id``. Title clamped at
200 chars. Capacity cap at 500 tasks — hitting the
cap is an explicit signal to prune done/blocked rows.
- ``update`` needs approval. Mutate by ``task_id``; fields
``title`` / ``status`` / ``child_ws_id`` optional.
- ``remove`` needs approval. Drop by ``task_id``.
- ``reorder`` needs approval. Pass ``task_ids``; validated as an
exact permutation of the current set (rejects
partial, extra, or substituted ids — prevents silent
task loss).
Status enum: ``pending`` / ``in_progress`` / ``done`` / ``blocked``.
``child_ws_id`` links a task to the child workstream spawned for it
(no enforcement; the coordinator owns the relationship).
## Persistence
Stored as a single JSON-envelope value on ``workstream_config`` —
``{"version": 1, "tasks": [...]}``. No new table; the kanban v2
work will supersede this row via a format migration keyed on
``version``. ``_save_task_list`` writes only the ``tasks`` key so
concurrent writers to other ``workstream_config`` keys (e.g. the
admin Settings UI updating ``reasoning_effort``) aren't clobbered
by a read-modify-write on the full row.
## Corrupt-envelope safety
A hand-edited or legacy config row that doesn't parse as the
expected shape logs a warning and returns an empty envelope from
``task_list_get``. Mutators refuse to overwrite corrupt data —
they detect the sentinel and return a clear error so the operator
can inspect or clear the row rather than losing work silently.
## Concurrency
Per-(ws) ``threading.Lock`` cached on the client. The worker
thread is single-threaded for tool execs so this is mostly
defence-in-depth against future maintenance-script call sites.
Cache never grows beyond one entry per coordinator session because
the scope guard short-circuits foreign ``ws_id`` before the lock
is acquired.
## Malformed-JSON recovery
``_prepare_tool`` fallback-1 regex-extract allowlist extended with
``action`` / ``status`` / ``task_id`` / ``title`` (alphabetized) so
slightly-malformed ``task_list`` calls get the same
self-correction behaviour as the other coordinator tools.
## Tests
- ``tests/test_coordinator_client.py`` — 15 new cases covering:
fresh-envelope shape, add/get roundtrip, empty-title + invalid-
status rejection, 200-char title clamp, update by id + missing
id, remove semantics, reorder permutation validation (partial +
extra + wrong id + valid), cross-ws scope violation, corrupt-
JSON read recovery, corrupt-envelope write refusal (all four
mutators), 500-task capacity cap, workstream_config key
preservation across ``_save_task_list``.
- ``tests/test_coordinator_tools.py`` — 12 new cases covering the
dispatch layer: list auto-approved, each mutating action needs
approval, unknown-action / missing-required-arg errors, list
returns tasks, page-cap at 200 with truncated signal, add
dispatches to client, reorder surfaces permutation error,
remove-not-found.
- ``tests/test_tools_schema.py`` / ``tests/test_workstream_kind.py``
extend the tool-count + disjoint-namespace + primary-key
regression guards with ``task_list``.
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` 4148 passed,
3 deselected (live-backend), 0 failed.
* feat(coordinator): coordinator workstream kind — phase 1
Adds a new ``kind="coordinator"`` workstream that runs inside the
``turnstone-console`` process (first ChatSession hosted on the console)
with a dedicated tool set for spawning and driving child workstreams.
Supersedes the external ``turnstone-coordinator`` MCP side-car for new
installs; the extension is marked deprecated in
``examples/mcp-cluster-ops/README.md`` but still works on 1.4-and-earlier
clusters.
Phase 1 ships: the workstream class, 6 lifecycle tools, console hosting,
9 HTTP endpoints, per-user audit attribution, and a one-pane web UI at
``/coordinator/{ws_id}``. Node/skill discovery tools, task-list tool,
tree-view UI, and routing-proxy audit middleware follow in a later PR.
## Schema
Migration 039 adds ``kind`` / ``parent_ws_id`` columns + indexes to
``workstreams``. Both SQLite and PostgreSQL backends take the new
kwargs on ``register_workstream``; empty-string ``parent_ws_id``
normalises to ``NULL`` at the storage edge. PostgreSQL uses
``INSERT ... ON CONFLICT DO NOTHING`` to match SQLite's ``OR IGNORE``
and close a pre-existing SELECT-then-INSERT TOCTOU window.
``list_workstreams`` gains optional ``parent_ws_id`` / ``kind`` filters;
new ``get_workstream(ws_id)`` returns the full row (the existing
``get_workstream_metadata`` stays untouched for back-compat).
## Core session + kind routing
- ``ChatSession.__init__`` accepts ``kind`` / ``parent_ws_id`` /
``coord_client``. On ``kind="coordinator"`` it swaps
``_tools = COORDINATOR_TOOLS`` and zeros sub-agent tool lists.
- ``Workstream`` dataclass extended with ``user_id`` / ``kind`` /
``parent_ws_id``. Both ``WorkstreamManager`` and the new
``CoordinatorManager`` use the same type — no parallel hierarchy.
- ``_SessionFactory`` Protocol + server / cli factory closures thread
the new kwargs. ``POST /v1/api/workstreams/new`` rejects
``kind != "interactive"`` with 400; ``POST
/v1/api/workstreams/{ws_id}/open`` refuses coordinator rows so a
server node can't accidentally rehydrate one.
## Coordinator tool set
Six tools (``spawn``, ``inspect``, ``send``, ``close``, ``delete``,
``list_workstreams``) with a ``coordinator: true`` metadata flag,
scoped to coordinator-kind sessions only. ``inspect`` and ``list`` are
auto-approved reads; the four mutators need approval. ``list`` returns
``{"children": [...], "truncated": bool}`` so the model can detect
post-filter under-fill and paginate.
## CoordinatorClient (in-process, sync)
Mutating ops HTTP-POST to the console's own ``/v1/api/route/*`` on the
local bind URL so every existing middleware (auth, rate-limit) runs.
Read ops hit ``storage.list_workstreams`` / ``get_workstream`` /
``load_messages`` directly — the routing proxy doesn't expose
list/inspect paths. URL paths are a validated constant table (avoids
an httpx ``base_url``-merge trap). A new
``/v1/api/route/workstreams/delete`` proxy handler joins the existing
route-proxy endpoints.
## Per-session coordinator JWT
``CoordinatorTokenManager`` mints short-lived JWTs with ``sub=<real
user>`` (attribution preserved), ``src="coordinator"``,
``aud="turnstone-console"``, ``coord_ws_id=<ws>`` custom claim.
``_proxy_auth_headers`` preserves ``src`` + ``coord_ws_id`` across the
upstream re-mint so server-side middleware sees coordinator-origin,
not ``console-proxy``. ``AuthResult.extra_claims`` carries
non-reserved claims through validate→remint; ``create_jwt``'s
reserved-claim set (now including ``nbf`` / ``jti``) is symmetric with
``validate_jwt``.
## Console hosts the ChatSession
- New ConfigStore settings: ``coordinator.model_alias`` (required),
``reasoning_effort``, ``max_active`` (default 5),
``session_jwt_ttl_seconds``.
- Console lifespan builds a ``ModelRegistry`` +
``CoordinatorManager``. Missing / unresolvable alias returns **503**
with remediation text — never 500.
- ``CoordinatorManager``: placeholder-slot reservation under lock,
rollback on factory failure, per-ws_id rehydration lock to serialise
concurrent lazy-opens, ``max_active`` enforced via ``close_idle``
eviction semantics.
- ``ConsoleCoordinatorUI`` is a thin ``SessionUI`` implementation — no
global broadcast, no per-node metrics, shared
``_APPROVAL_WAIT_TIMEOUT`` constant across approval + plan paths.
- No eager startup rehydration: persisted coordinator rows load lazily
on first ``GET /v1/api/coordinator/{ws_id}``.
## Console coordinator API
Nine endpoints under ``/v1/api/coordinator/*`` gated by ``approve``
scope + new **``admin.coordinator``** permission (added to
``_VALID_PERMISSIONS``; not in any builtin role — operators opt in
explicitly). Ownership failures return **404, not 403** and use
strict equality so empty-owner rows don't leak across tenants.
Correlation-id masking on every factory-raising path
(``coordinator_create`` + ``coordinator_detail`` lazy rehydrate) — no
stack traces to the client.
## Audit attribution
Three console-side events (``coordinator.create`` / ``.close`` /
``.cancel``) with the real creator's ``user_id`` plus
``detail={coord_ws_id, src="coordinator"}``. No schema migration
required. Per-tool-call audit across the routing proxy is deferred
(needs either a ``source`` column on ``audit_events`` or
``record_audit`` calls wired into the route-proxy handlers).
## Web UI (``/coordinator/{ws_id}``)
One-pane chat served by the console. Reuses ``shared_static``
(``base.css``, ``auth.js``, ``theme.js``, ``toast.js``, ``utils.js``,
``kb.js``) and the server UI's ``renderer.js`` pipeline (KaTeX, Mermaid,
highlight.js already bundled).
- SSE to ``/v1/api/coordinator/{ws_id}/events`` with exponential-
backoff reconnect; status line carries a leading glyph
(● / ○ / ⚠) so state isn't conveyed by colour alone.
- Renders content, reasoning (dimmed italic
``.role-reasoning``), tool_result, approve_request, intent_verdict,
output_warning.
- Child ws_id references auto-wrap to
``/node/{node_id}/?ws_id={child}`` links — both ids regex-validated
before interpolation, everything else HTML-escaped.
- Non-modal approval bar (``role="region"``) with a batch header
("Approve N tool calls"), initial focus on the approve button,
buttons disabled during the in-flight POST, red-bordered deny.
``aria-live`` flips to ``off`` during streaming.
- "New coordinator" button on the dashboard header — permission-gated
on the UI side, matching the backend 403.
- Mobile composer capped under ``@media (max-width: 700px)``.
## Tests
~120 new tests across 8 files: workstream-kind storage + dataclass
semantics, CoordinatorClient URL map + token minting + storage reads +
truncation signalling, tool prepare/exec dispatch and approval gating,
CoordinatorManager create / rollback / eviction / lazy rehydration +
concurrency, HTTP endpoint auth + 404-on-ownership + 503-on-misconfig,
proxy-auth ``src`` preservation, full lifecycle end-to-end, coordinator
page HTML-injection guard. ``test_tools_schema.py`` widened to 25
tools (19 existing + 6 coordinator).
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 files), ``pytest`` 4054 passed (5 pre-existing failures unrelated
to this change — confirmed against ``main``).
* polish(coordinator): address PR review + CI + tool-namespace isolation
CI:
- `ruff format`: two files reformatted, matches the in-repo pre-commit config.
- `wheel-completeness`: add `turnstone/console/static/coordinator/*.html` +
`*.js` to the hatch wheel-include list. Without this the coordinator UI
was missing from published wheels.
- `test (3.11/3.12/3.13)` + `test-postgres`: three `TestExecReadImage`
tests were masking a real bug — my 6 new tool JSONs pushed tool count
19→25, crossing the default `tool_search.auto` threshold (20), which
made `ChatSession.__init__` construct a `ToolSearchManager` and cache
`_cached_capabilities` during init. Tests that later patched
`session._provider.get_capabilities` saw the cached value instead.
Root-cause fix: the tool-search threshold code path now reads
capabilities through `_resolve_capabilities(...)` directly — no cache
populate — so the patch takes.
Tool-namespace isolation (bigger fix than CI symptoms suggested):
- `TOOLS` was the union of all loaded tool JSONs including the 6 new
coordinator tools. Interactive sessions were getting coordinator
tools in their function-calling surface (which is nonsense — they
require a console-hosted `coord_client`), and coordinator sessions
counted against the interactive tool-search threshold. Fix:
- New `INTERACTIVE_TOOLS` / `INTERACTIVE_TOOL_NAMES` in
`turnstone/core/tools.py` exclude anything with `coordinator: true`
metadata. `TOOLS` stays as the union for schema introspection +
eval catalog.
- `ChatSession.__init__` selects tool set by kind: coordinator gets
fixed `COORDINATOR_TOOLS` (no MCP merge, no listeners registered);
interactive gets `INTERACTIVE_TOOLS` (+ MCP if configured).
Coordinators are meta-orchestrators that spawn child workstreams;
MCP tools / resources / prompts live on the children, not on the
coordinator's own surface.
- `_on_mcp_tools_changed` no-ops for coordinator sessions
(defence-in-depth in case listeners were registered).
- `always_on_names` on `ToolSearchManager` is now the set of builtin
tools actually present in the session (kind-aware) rather than the
full `BUILTIN_TOOL_NAMES` frozenset.
- `turnstone/eval.py` uses `INTERACTIVE_TOOLS` (coordinator tools
aren't in scope for the eval harness which tests interactive agent
behaviour).
- Regression tests in `tests/test_workstream_kind.py`:
- `INTERACTIVE_TOOLS ∩ COORDINATOR_TOOLS == ∅` and their union is
`TOOLS`.
- Interactive `ChatSession._tools` does not include any
coordinator tool name.
- Coordinator `ChatSession._tools` contains `spawn_workstream` but
not `bash` / `edit_file` / `memory`; sub-agent lists are empty.
- Coordinator `ChatSession` with an MCP client attached does NOT
merge MCP tools and does NOT register any MCP listeners.
PR review findings:
- **#10 / #11** (Copilot): coordinator UI claimed to reuse the server
renderer pipeline but loaded none of its JS. Mirrored
`turnstone/ui/static/renderer.js` into
`turnstone/console/static/coordinator/renderer.js` (flagged in-file
as a cleanup candidate to promote into `shared_static/`), added
`katex.min.js` / `highlight.min.js` / `renderer.js` script tags to
`coordinator/index.html`. `coordinator.js` now buffers raw markdown
via `textContent` during streaming, then swaps to `renderMarkdown` +
`postRenderMarkdown` on `stream_end`.
- **#7** (Copilot): N+1 query pattern in
`CoordinatorClient.list_children()` — per-row `storage.get_workstream`
just to read `skill_id`. Pushed `skill_id` + `skill_version` into
the `list_workstreams` SELECT projection on both backends; the
client reads them from `row._mapping` directly. New
`test_list_children_skill_filter_avoids_n_plus_one` pins the
behaviour (asserts `storage.get_workstream` call count is 0).
- **#8 / #9** (Copilot): `spawn_workstream` tool JSON said "if empty,
the workstream is created idle" but the prepare method rejected
empty and the field was marked required. Resolved by allowing
empty end-to-end: removed from `required`, prepare builds a
"spawn idle workstream" header + empty preview when empty,
updated `test_spawn_prepare_allows_empty_initial_message`.
- **#1–#5** (github-code-quality): five asserts with side-effecting
method calls in `test_coordinator_manager.py` (`mgr.close`,
`mgr.open`, `mgr.create` in a dead `_c = ...`). Extracted each
call to a local variable so `python -O` can't strip the side
effect.
Verification:
- `ruff check turnstone tests` clean.
- `mypy turnstone` clean (156 source files).
- `pytest -m "not live"` — 4063 passed, 3 deselected (live-backend
tests), 0 failed. The 3 image tests that were failing on this
branch now pass; wheel + lint both green locally.
* polish(coordinator): address Copilot re-review findings
Two findings from the re-review of #368 after the first polish commit.
**user_id wired into `mgr.create()` at the server handlers.** Phase 1
added ``user_id`` to the ``Workstream`` dataclass and
``WorkstreamManager.create()`` signature, but the two call sites in
``turnstone/server.py`` forgot to pass the authenticated caller
through. Result: interactive workstreams created via
``POST /v1/api/workstreams/new`` (including coordinator-spawned
children, which route through this handler) were landing with blank
``user_id``, defeating ownership-based access control on subsequent
sends / approvals / closes (``_require_ws_access`` treats blank
owners as legacy/allowed). Two changes:
- ``server.py:create_workstream`` forwards ``user_id=uid`` — the same
``uid`` already resolved from the auth result (with trusted-service
forwarding preserved).
- ``server.py:open_workstream`` prefers the persisted owner on the
workstream row over the rehydrating caller so reloading someone
else's workstream doesn't silently re-parent it. Falls back to
the authenticated caller when the stored row has no owner
recorded (pre-phase-1 rows).
Regression test in ``tests/test_workstream.py`` pins
``WorkstreamManager.create(user_id=X)`` → ``ws.user_id == X`` so the
manager seam can't regress silently on a future refactor.
**Malformed-JSON recovery allowlist expanded for coordinator args.**
``_prepare_tool()`` has a two-stage salvage path for models that
emit malformed JSON: a regex-extract (fallback 1) and a bare-string
→ primary_key wrap (fallback 2). The fallback-1 key list didn't
include coordinator argument names, so a slightly malformed
``spawn_workstream`` / ``send_to_workstream`` / etc. call would
hard-fail instead of salvaging into a minimal-args dict for retry.
Added ``ws_id`` / ``message`` / ``initial_message`` / ``parent_ws_id``
to the allowlist (kept alphabetised) so the coordinator tools get
the same model-self-correction behaviour as the interactive tools.
Fallback 2 already covers the ``ws_id``-primary-key tools via
``PRIMARY_KEY_MAP``; the regex path matters when the model emits
``{"ws_id": "abc", "message": "..."}`` with a trailing syntax error.
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` → 4065 passed, 3
deselected (live-backend), 0 failed.
* fix(coordinator): address ultrareview findings on coordinator workstream kind
Security
- Cross-tenant leak: CoordinatorClient.inspect/list_children now constrain
to the coordinator's own ws_id + direct children; an LLM coerced via
prompt injection can no longer exfiltrate other tenants' workstreams.
- Empty-owner short-circuit bypass: strict equality at coordinator.py
ownership gate and at the storage-fallback branch in coordinator_history;
orphan/system-owned coordinator rows can no longer be rehydrated by
arbitrary holders of admin.coordinator (DoS + history disclosure vector).
- Closed coordinators no longer silently resurrect on subsequent GET —
the Close button is now actually durable across URL revisits and tab
refreshes; rows with state in {closed, deleted} refuse rehydration.
Correctness
- ChatSession.close() now releases the CoordinatorClient httpx.Client
pool; previously every closed/evicted coordinator dropped a connection
pool on the floor until non-deterministic GC.
- open_workstream rehydration now forwards parent_ws_id + kind, so
coordinator-spawned children survive node restart / idle eviction
with their parent link intact instead of becoming silent orphans.
- list_children truncated flag now signals whenever the SQL fetch hit
the page cap (previously permanently False in the no-filter case,
causing confident-but-incomplete summaries from the coordinator).
- ConsoleCoordinatorUI.approve_tools: per-tool auto-approve now checks
auto_approve_tools independently of the blanket auto_approve flag,
so 'Always approve this tool' actually works on the next invocation.
Concurrency
- _spawn_worker no longer falls through to start a second concurrent
worker thread on the same ChatSession when queue.Full fires; instead
send() returns False and the endpoint surfaces HTTP 429.
- _open_locks entries are now refcounted under self._lock and only
popped when the last waiter releases — eliminates the race where a
rehydration-failure path lets two threads serialize on different lock
instances for the same ws_id and trip the "already tracked" guard.
Tests: +6 regression cases covering closed-coordinator refusal,
empty-owner non-admin refusal, queue.Full no-duplicate-worker,
inspect/list_children cross-tenant rejection, and truncated semantics.
All eight suggestions verified against source before applying:
- docs/settings.md — ConfigStore key names are `model.plan_alias` /
`model.task_alias` (not `plan_model` / `task_model`); updated in
both the overview list and the plan/task overrides table.
- docs/security.md — `src` claim values now reflect what actually
gets minted: `password`, `database` (from API-token exchange),
`oidc`, plus service origins `console`, `cli`, `channel`.
- docs/sdk.md — `upload_attachment(ws_id, filename, data, *,
mime_type=...)` matches the real SDK signature; `bytes`-returning
helper is `get_attachment_content` (not `download_attachment`);
code example reordered so it doesn't collide on `filename=` kwarg.
- docs/architecture.md — "prior `plan` tool call" → "prior
`plan_agent` tool call" so wording stays consistent with the
renamed tool.
- docs/tools.md — `plan_agent` `primary_key` is `goal`, not
`prompt`, in both the primary-key table and the summary table
(matches the JSON schema in turnstone/tools/plan_agent.json).
Systematic pass over every doc under docs/, the root-level README /
QUICKSTART / CONTRIBUTING, and the PlantUML diagrams. Memory and docs
had drifted against the code since 1.2 — this catches them up to the
1.4.0 release and the 1.5.0a1 experimental line.
User-facing fixes
- README: fix broken docs/mcp.md link (→ mcp-registry.md); channel
gateway entry reflects shipped Discord + Slack adapters instead of
"Slack/Teams planned"; diagrams table mentions both.
- QUICKSTART: docs/*.md relative links were wrong from the repo root;
wizard version bumped from 0.5.4.
- CONTRIBUTING: add dev extra plus the ruff / mypy / pytest commands
we actually expect before push.
Reference docs
- architecture.md: 19 tool schemas (was 15), 18 admin tabs (was 14),
turnstone-bootstrap added to entry-points table, OpenAI provider
file split (chat/responses/common) documented, 38 SDK event
dataclasses (was 27 and referenced deleted mq/protocol.py), Slack
adapter + multi-adapter gateway, plan_agent/task_agent naming,
governance admin-panel rewrite.
- api-reference.md: full attachment endpoints (POST/GET/content/
DELETE on /v1/api/workstreams/{ws_id}/attachments) plus the
multipart mode on POST /v1/api/workstreams/new.
- channels.md: Slack Setup section (Socket Mode app creation, OAuth
scopes, tokens), Slack CLI/env reference in config table, combined-
adapter architecture diagram.
- console.md: 18-tab listing (was 13) with Channels/Models/Nodes/TLS
descriptions and ConfigStore live-edit note.
- docker.md: Slack env vars block; image entry-point list now
includes turnstone / turnstone-bootstrap.
- sdk.md: attachments methods on the server client, attachments
example (upload-then-send and at-creation), event count fixed.
- releasing.md: four-track table (stable/1.0, 1.3, 1.4 + main 1.5);
promotion workflow uses 1.5 / 1.6 numbering.
- settings.md: plan_model / task_model / plan_effort / task_effort
overrides section.
- governance.md: skill naming (/skill, `skill` field — not /template),
Prompts/Judge tabs called out.
- security.md: two-token-types wording; src claim values match the
AuthResult source strings actually emitted.
- mcp-registry.md: SDK package name is @turnstone/sdk.
- tools.md: plan / task renamed to plan_agent / task_agent in the
section headings and summary table; primary-key table matched.
- design/consistent-hash-ring.md: dead direct-http-transport.md
pointer redirected to architecture.md.
Diagrams
- 02-package-structure: drop phantom chat.py entry point, add admin
and bootstrap, add slack/bot.py, rename channels/gateway.py →
channels/cli.py.
- 16-channel-architecture: Slack is no longer "(future)", add a
SlackBot class and the slack-bolt Socket Mode edges; wire the new
bot into ChannelService. PNGs regenerated from both puml sources.
Audit pass against the actual commit messages between v1.3.0 and
v1.4.0 turned up several substantive items the initial CHANGELOG
under-described or omitted entirely. Fix-forward expansion plus a
Contributors section recognizing external contributors.
Added detail / coverage:
- New "Server compatibility layer for local model servers" entry —
the vLLM / llama.cpp profiles + admin UI fields shipped in #352
alongside the capabilities passthrough; previously buried under one
bullet.
- Per-call plan/task model selection split into three sub-bullets
(backend split, runtime configurability without restart via
ConfigStore admin tab, per-call override) — three PRs that build on
each other deserve to be discoverable independently.
- Opus 4.7 entry expanded with 1M ctx / 128K output, the new
thinking_display capability field, xhigh effort level, and admin
dropdown updates.
- Dashboard composer note: tab-bar `+` modal also gained the paperclip
+ chip strip + first-message field.
- Slack adapter: explicit "session recovery via persisted recoverable
route keys" — ops-relevant promise for restart behaviour.
- pgbouncer swap: helm chart link + ports updates noted.
- Provider capabilities entry: defensive shallow-copy + chat_template
deep-merge follow-ups.
New Fixed entries:
- Cross-user attachment-fetch hardening (get_attachment_content
scopes by user_id).
- Attachment-list DoS guard on /v1/api/send.
- Bounded LRU for upload locks.
- 3.12 CI deadlock root-cause writeup (asyncio.Lock vs Starlette
TestClient loop teardown).
New SDK entry:
- PlanResolvedEvent type + guard, dispatched cross-client when one
client resolves a plan so others dismiss in sync.
New Operational subsection:
- vendor-js workflow now auto-downloads hls.js for future Renovate
bumps so they're merge-ready without manual file fetches.
Contributors:
- Recognise @daoxley (Slack adapter, #355) and @pizzaandcheese
(pgbouncer swap, #353) — the two external contributors with
meaningful net-new work in this release — plus the Renovate bot.
- Pointer to channel-attachment ingest as the headline 1.4.1 feature
for would-be contributors.
Repo previously had no CHANGELOG. Establishes the file with full
1.4.0 coverage (attachments end-to-end, dashboard composer refactor,
Slack adapter, per-call plan/task model, provider capability
passthrough, Opus 4.7) plus a one-line 1.3.1 entry for the Opus 4.7
backport. Format follows Keep a Changelog 1.1.0; release-track
guidance up top covers the three stable branches + main.
Operator-relevant call-out at the top of [1.4.0]: migrations 037 +
038 must be applied before starting 1.4.0 against an existing 1.3.x
database. Both are additive and idempotent.
* feat(ui): dashboard composer polish from PR #362 designer review
Three deferred items from the prior designer pass on the unified
dashboard composer. Pure UX affordances; no server change.
- **Persist Options open/closed in localStorage.** Power users who
routinely set non-default model/skill don't have to click "Options"
on every page load. Key: `turnstone.dashboard.options_open`.
Defaults closed for first-time users. Falls back gracefully when
localStorage is unavailable (private mode, quota).
- **Active-options summary chip.** Renders the non-default model /
judge / skill values inline next to the Options button (mono, dim,
separated by middots). Hidden via `[hidden]` when everything is at
default — no chrome cost in the common case. Updates on any select
change via a single delegated handler on the panel. Hidden on
narrow viewports (the action row stacks vertically there and the
chip would push the layout further).
- **"Drop to attach" overlay during drag.** CSS pseudo-element on
`.dashboard-composer-drop` overlays a centered "Drop to attach"
label so dragging a file makes the action explicit instead of just
showing the dashed-border highlight. pointer-events: none keeps
the underlying composer controls reachable; visual only.
* fix(ui): address Copilot review on dashboard composer polish
- _restoreDashboardOptionsState() forced the panel closed every time
showDashboard() ran when localStorage was unavailable (private mode,
storage quota), contradicting the comment that promised a per-session
fallback. Add a module-scoped _dashOptionsOpenSession variable
updated by _setDashboardOptionsOpen / _toggleDashboardOptions, and
only override the visible state from localStorage when the read
genuinely succeeded. The session value now preserves the user's
choice across hide/show cycles in environments where localStorage
throws.
- Fold the duplicated `.dashboard-composer { position: relative; }`
block into the existing rule above. The position context is needed
for the .dashboard-composer-drop::before overlay; the comment now
says so.
PR #355 added the Slack adapter on the server but missed the console
admin surfaces that talk to channel_type. Three concrete gaps + a
designer-review polish pass.
Functional bug + UI parity:
- _collectNotifyTargets() in admin.js hardcoded `channel_type: "discord"`
— even on a Slack-only deployment the skill notify-on-complete form
always wrote Discord targets, sending notifications to the wrong
adapter (or nowhere). Add a per-row channel-type <select> driven by
a small _NOTIFY_CHANNEL_TYPES table that's the one place to register
a new platform; collector and populator both read from the dropdown.
ID-input placeholder updates dynamically when the platform changes.
- The "Link Channel Account" modal only offered Discord — users
couldn't link a Slack account through the UI at all. Add a Slack
<option> and reuse the same dynamic-placeholder helper. Drop the
static Discord-shaped HTML placeholder so the JS-driven hint doesn't
flash a Discord example before the dropdown initializes.
- Skill create/edit modals only showed Discord in the notify-on-complete
placeholder example. Show both adapters.
- Per-platform .scope-discord / .scope-slack badge classes so the
linked-accounts list distinguishes platforms visually instead of all
rendering as the generic .scope-channel magenta. Falls back to
.scope-channel for any future channel_type the stylesheet doesn't
yet know about.
Designer review polish:
- Theme-aware --discord / --slack / --discord-glow / --slack-glow
tokens in base.css. The first pass shipped raw hex (#818cf8 /
#f472b6) that fails WCAG AA on light theme (1.8:1 and 2.4:1); the
light variants (#4f46e5 indigo, #be185d rose) pass. Badge classes
now reference tokens, matching every other .scope-* rule.
- Notify-row mobile layout: three controls in a row left ~80px for
the ID input at 360px viewport, truncating snowflakes. Tighten
platform select to 76px (labels are short), add flex-wrap, and at
≤700px drop the ID input to its own row so it gets full width.
- Per-platform classes apply alone (not co-classed with scope-channel)
so winning the cascade doesn't depend on stylesheet source order.
- Replace "Discord snowflake" jargon with "Discord ID"; give Slack
ids concrete examples (C01234567 / U01234567) instead of an
ambiguous "C0…".
Combines the substantive bot.py fixes flagged in both review trails on
PR #355. Discord parity items grouped here too since they're the same
surface (slack/bot.py).
From Copilot:
- _notify_reply_routes was read on StreamEndEvent but never popped on
the success path. Result: one notification reply pinned every later
response for that ws_id to the notification thread until the bot
restarted. Pop after read; combine the surrounding ifs (SIM102).
- PlanReviewEvent embedded raw event.content inside a triple-backtick
mrkdwn fence without escaping. A plan with ``` (very common — plans
often quote code) would break the fence and let later content render
as live markup, including unintended Slack mentions/links. Rewrite
_sanitize_slack_preview to splice a zero-width space inside any ```
sequence (Slack stops recognizing it as a delimiter) instead of
escaping every single backtick — keeps single-backtick code snippets
readable while still protecting the fence. Apply to plan-review.
- _send_approval_request joined unbounded tool_lines into one mrkdwn
section, but Slack section.text caps at 3000 chars. Multi-tool
batches with large previews silently failed chat_postMessage,
leaving the user unable to approve/deny. Cap each preview to 600
chars under a 2700-char total budget; append "+N more" when truncated.
From eous (parity with Discord):
- Pass `client_type="chat"` from both `get_or_create_workstream` call
sites (slash-command session + DM). Without it Slack-routed
workstreams loaded the web-default prompt; the chat-specific
system prompt now applies as it does for Discord.
- Add `exc_info=True` to the eleven `log.debug(...)` exception handlers
so underlying tracebacks are available when debug logging is on
instead of being silently dropped. Level stays debug — these are
benign-by-default sites (chat_update on a deleted message, etc.) so
only the visibility changes. Typed-exception handlers
(RemoteProtocolError, etc.) keep their bare debug log.
- Module docstring on slack/__init__.py so pydoc / import errors have
human-readable context.
Tests: rewrite the sanitizer test to match the new (more permissive)
single-backtick behaviour; add coverage for the triple-backtick
neutralization + short-input passthrough; patch httpx.AsyncClient at
all five TurnstoneSlackBot construction sites so each test doesn't
leak an unclosed real client.
- cli.py: ChannelAdapter import is annotation-only; move into
TYPE_CHECKING block and switch the two cast() calls to string-form
so the runtime import isn't required (TC001).
- slack/{config,routes}.py: ruff format fixes (whitespace + drop
redundant string-form annotation now that __future__ annotations
is in effect).
- pyproject.toml: drop the unused `tests.*` mypy override — `mypy
turnstone` (the only invocation in CI + local) never matches it,
so it was pure noise in the "unused section(s)" report. Other
optional-dep overrides stay; they're real safety nets when running
mypy without the [all] extras (e.g. on the test job).
- uv.lock: regenerate to match the slack-bolt + transitive deps the
pyproject changes resolve to (lock-check was failing on stale hash).
* fix(ui): rehydrate chip strip after queued-message dequeue not_found
The dequeue handler only refreshed the per-pane chip strip when the
DELETE returned status="removed". On status="not_found" (the queued
message already dispatched), chips stayed stale: any reservations that
raced the dispatch could leave the UI showing a different pending set
than the server actually had.
Re-fetch on both paths so the chip strip always reflects the
authoritative server state. The queued-message bubble itself stays
visible on not_found, same as before — the promote loop strips the
queued styling on idle.
* feat: sweep orphan attachment reservations periodically
Process crashes between reserve_attachments and consume/unreserve can
leave attachment rows soft-locked forever (reserved_for_msg_id NOT NULL
with no consumer ever coming back). The worker-thread exception path
in /v1/api/send already handles in-process failures, but a hard kill
or oom mid-send escapes that.
Add sweep_orphan_reservations(older_than_seconds) to the storage
protocol — clears reserved_for_msg_id on rows where message_id IS NULL
and created < now() - threshold. Implemented for SQLite + PostgreSQL
using the same string-comparison form (created is ISO-8601 text in
both backends, lexicographic order matches chronological).
Wire into the server lifespan: run once at startup (catches anything
left over from the previous process), then every 30 minutes as
defense-in-depth. Threshold is 4 hours so we don't race a long-running
dispatch and unreserve rows the worker is still about to consume.
Tests cover sweep semantics: clears old reserved rows, leaves fresh
ones alone, skips already-consumed rows, no-ops on zero/negative
threshold.
* fix: track reserved_at for orphan-reservation sweep
Copilot review on PR #363 flagged a real correctness bug: the sweep
used the attachment row's `created` timestamp (upload time) as the
staleness signal. An attachment uploaded hours ago but reserved fresh
could be unreserved mid-send, after which mark_attachments_consumed
silently drops the row because reserved_for_msg_id no longer matches
the send_id.
Add a dedicated `reserved_at` column (migration 038) set on
reserve_attachments and cleared on mark_attachments_consumed /
unreserve_attachments. The sweep now scopes by `reserved_at < cutoff`,
so reservation age is what's measured, not upload age. Backed by a
partial index `(reserved_at) WHERE reserved_at IS NOT NULL` so the
periodic scan stays cheap as the consumed-history grows.
Threshold dropped from 4h to 1h since it now means "longest realistic
single send" rather than "longest plausible time between upload and
send" — a tighter, more defensible bound.
Tests cover the regression (uploaded long ago + reserved fresh must
not be swept), plus reserved_at clearing on both consume and unreserve.
* feat: workstream attachments at creation time + SDK + UI parity
Closes the two big deferred items from PR #356: attaching files as part
of the initial workstream-creation request, and full SDK coverage of the
attachment surface.
Server: POST /v1/api/workstreams/new now accepts multipart/form-data
(meta JSON + 0..N file parts). Files are validated and saved as pending
under the new ws; when initial_message is also set the create handler
reserves them onto that turn before the dispatch worker fires, mirroring
the /v1/api/send pattern. Validation failure rolls back the workstream
via delete_workstream so we don't leak orphan rows or emit a phantom
ws_created/ws_closed pair on SSE. JSON path is unchanged.
Console routing: route_create accepts multipart with ?ws_id=<hex> as a
query parameter (the console hashes the id before the body lands).
Added /v1/api/route/workstreams/{ws_id}/attachments POST/GET/DELETE +
.../{attachment_id}/content GET proxies that forward raw bytes and
preserve upstream headers (Content-Disposition, X-Content-Type-Options,
CSP sandbox).
Python + TypeScript SDKs: AttachmentUpload type, upload_attachment,
list_attachments, get_attachment_content, delete_attachment, and
send(attachment_ids=...). create_workstream(attachments=...) sends
multipart and pre-generates a ws_id client-side so cluster routing
works. SDKs reject attachments+target_node combinations since the
multipart route doesn't honor target_node.
Web UI: dashboard composer refactored to a single unified create flow.
Replaced the inconsistent split (Enter created+sent raw, "New Chat"
opened a modal) with one rich composer carrying a textarea, paperclip
+ chip strip, drag-drop, paste-image, and a collapsible Options panel
for model/judge_model/skill. Submit button dynamically labels Create
vs Send. New-workstream modal also gained the same paperclip + chip
strip + first-message field for the tab-bar + entry point.
Tests: 30 new tests across server multipart create, console route
multipart + attachment proxies, Python + TS SDK attachment surfaces,
plus regressions for the three review-flagged bugs (Content-Type
boundary preservation, attachments+target_node rejection, no phantom
ws_created on validation failure).
* fix: address Copilot review feedback on PR #362
- web_helpers: docstring now matches behaviour — read_multipart_create_or_400
does enforce the optional max_per_file_bytes cap as defense-in-depth.
- app.js: drop the duplicated _formatAttachSize definition (one already
exists earlier for pane chips); add a shared _isAttachmentAllowed helper
that mirrors the server's classifier (png/jpeg/gif/webp images, text/*
MIMEs, allowlisted application/* MIMEs, known text extensions) and call
it from both _newWsAddFiles and _addDashboardFiles so unsupported files
fail fast client-side instead of after a server roundtrip.
- app.js: dashboardSubmit catch now suppresses the redundant error toast
on authFetch's "auth" Error and falls back to a generic message when
err.message is undefined, instead of rendering "Connection error: undefined".
- SendResponse (Pydantic + TS): document and expose attached_ids,
dropped_attachment_ids, priority, and msg_id so attachment-aware SDK
callers can detect partial reservations and dequeue queued messages.
- test_server_attachments_on_create: drop the dual `import turnstone.server`
+ `from turnstone.server import` style — use monkeypatch.setattr by
dotted path for module-level mutation and `from … import …` for the
helpers, keeping a single import style.
* feat: per-call model selection on plan_agent / task_agent
The calling LLM can now pass `model="<alias>"` to plan_agent or
task_agent to override the operator-configured per-kind model for
that one invocation. Useful when subtask difficulty varies within a
session: the model can downgrade to a cheap alias for trivial work
and reach for a stronger one when the problem is hard.
Tool descriptions list the live registered aliases (refreshed when
the operator hits "sync to nodes" / internal_model_reload), so the
calling LLM always sees the current options. Bad aliases return a
corrective error dict with the available choices so the LLM retries
cleanly rather than failing silently.
No whitelist — any alias the registry knows is acceptable; cost
control is intentionally ceded to the model. No per-call effort
override (out of scope; effort stays operator-configured).
Resolution precedence in _run_agent: explicit per-call agent_alias
override > registry per-kind (plan_model/task_model) > legacy
agent_model > session model. The plan retry path (when
_validate_plan fails) reuses the same alias so coaching reflects
real model behaviour rather than a different model masking the
signal.
Implementation:
- plan_agent.json / task_agent.json: optional `model` parameter.
- ChatSession._validate_agent_model_override extracts and validates
the arg; mirrors the existing empty-prompt error pattern.
- _prepare_plan / _prepare_task stash the override in
item["model_override"]; _exec_* pass it through.
- _run_agent gains agent_alias kwarg with defence-in-depth
ValueError on unknown alias.
- _render_agent_tool_descriptions deep-copies plan/task entries
before mutating description so the module-level TOOLS constant
stays untouched across sessions; rebuilds the BM25 tool-search
index when active so its text matches what the LLM sees.
- server._broadcast_agent_tool_schema_refresh walks active
workstreams on internal_model_reload so descriptions update
without restart.
* fix: clarify no-registry placeholder + avoid double BM25 rebuild
Addresses Copilot feedback on PR #361.
1. plan_agent.json / task_agent.json placeholder said the parameter
falls back to the "operator-configured plan/task model". That
text is what no-registry sessions see (registry-bearing sessions
get the templated description with the live alias list); for
those single-model sessions, omitting the param falls back to
the current session model, not an operator-configured one.
Reword so the no-registry user gets accurate guidance.
2. _on_mcp_tools_changed already calls _rebuild_tool_search after
merging MCP tools. _render_agent_tool_descriptions also
rebuilt the BM25 index when active, so the MCP refresh path
was rebuilding twice per refresh. Move the BM25 rebuild out
of the private render helper into the public
refresh_agent_tool_schemas wrapper — _on_mcp_tools_changed
keeps calling the render helper directly (no double rebuild),
and registry-reload callers go through the wrapper which
still keeps the index in sync.
* feat: ConfigStore + admin UI for plan/task agent model and effort
Per-kind sub-agent routing was added in #359 but only via config.toml.
Operators can now switch the plan_agent / task_agent model and reasoning
effort at runtime from the admin Model tab without restarting.
Adds four ConfigStore-backed settings:
model.plan_alias — alias for plan_agent
model.task_alias — alias for task_agent
model.plan_effort — reasoning effort for plan_agent
model.task_effort — reasoning effort for task_agent
Server startup and internal_model_reload both apply these as overrides
on top of the registry's config.toml-loaded values; the new logic
computes "effective" values for all five model-routing fields and only
calls registry.reload() when at least one differs.
Admin UI: extracts ALIAS_SETTING_KEYS to a const used by both the
dynamic-alias-choice injection and the empty-option label rendering.
Adds INHERIT_EMPTY_LABEL_KEYS so plan_effort / task_effort show
"(inherit)" for empty — distinct from the literal "none" choice (which
actually disables reasoning, very different from leaving unset).
Also fixes Copilot review feedback from #359:
- _validate_effort treats empty / whitespace as unset rather than
warning on benign explicit-empty configs (with .strip().lower()
normalisation; "HIGH" and " low " now parse correctly)
- turnstone.example.toml's reasoning_effort comment lists the full
set of accepted values (none, minimal, low, medium, high, xhigh, max)
* fix: apply routing overrides on config-reload + skip no-op model-reload
Addresses Copilot feedback on PR #360.
1. Admin settings updates fan out via /_internal/config-reload, which
only reloaded the ConfigStore — plan/task routing changes weren't
visible until a model-reload or restart, defeating the runtime
configurability this PR is meant to add.
2. /_internal/model-reload always called registry.reload(), churning
cached clients even when nothing changed. Risky when fanned out
across nodes (could close in-flight clients).
Extracts two helpers in server.py:
- _effective_routing(cs, ...) pure function: overlay CS values on base
- _apply_routing_overrides(reg, cs) reload only when something differs
Used by the startup path, config_reload (new), and model_reload (now
short-circuits with a noop response when models + routing are unchanged).
plan_agent and task_agent previously shared a single agent_model knob and
plan_agent hardcoded reasoning_effort="high" in three call sites. They
have different cost/latency profiles — plan is rare and benefits from a
stronger model, task is frequent and benefits from a cheaper one — so
sharing the knob undertunes both.
ModelRegistry gains plan_model, task_model, plan_effort, task_effort.
Per-kind overrides win over the legacy agent_model, which still works
as the single-knob fallback for both. resolve_agent_alias(kind) and
resolve_agent_effort(kind) centralise the resolution; PLAN_DEFAULT_EFFORT
captures the back-compat "high" default in one place rather than at
every call site.
session._run_agent delegates resolution by label ("plan" vs "task").
The three hardcoded reasoning_effort="high" arguments are removed —
behaviour is identical when no plan_effort is configured.
Loader validates effort against {none,minimal,low,medium,high,xhigh,max}
and warns + drops typos rather than passing them to the provider.
ConfigStore parity and admin UI for the new knobs are deferred to a
follow-up — config.toml-only is enough for the backend split.
Previously, resolving a plan on one client (e.g. phone) cleared the
server's pending state and unblocked the worker, but emitted no event
to other connected clients. Their plan-approval modal stayed stuck.
resolve_plan() now enqueues a plan_resolved frame (mirroring the
approval_resolved pattern in resolve_approval) before clearing
_pending_plan_review, so a reconnecting client cannot receive both
the replayed plan_review and the live plan_resolved. Skips the frame
on the cancel-with-no-plan path.
Client adds a plan_resolved handler that dismisses the modal without
re-firing /v1/api/plan, restores keyboard context (skipped on touch
to avoid soft-keyboard pop on mobile), labels the inline plan summary
"(synced)" so remote dismissal is unambiguous, announces via the
existing aria-live #toast for screen-reader parity, and falls back
to an info message if plan_resolved races ahead of plan_review.
Adds PlanResolvedEvent to the Python and TypeScript SDKs with
deserialization and type-guard tests.
- Add claude-opus-4-7 capability entry (1M ctx, 128K output, adaptive
thinking, supports_temperature=False, thinking_display=summarized)
- Suppress temperature param for Opus 4.7 (API returns 400)
- Add thinking display opt-in via new ModelCapabilities.thinking_display
field - Opus 4.7 omits thinking by default, always send summarized
- Add xhigh effort level to mapping and Opus 4.7 effort_levels
- Add xhigh/max options to skill template dropdowns in admin console
- Align reasoning effort label capitalization across all console dropdowns
- Update example config to reference claude-opus-4-7
- 10 new tests with regression guards for Opus 4.6 backward compat
Verified against live API: streaming and completion calls succeed.
Trivy flags two HIGH CVEs in jq/libjq1 1.7.1-6+deb13u1 with no fixed
version yet from Debian:
- CVE-2026-39979: out-of-bounds read in jv_parse_sized() on non-NUL-
terminated buffers
- CVE-2026-40164: DoS via crafted JSON causing hash collisions
jq is invoked only on trusted CLI/admin paths against
process-controlled JSON input in turnstone — never on untrusted
network bytes — so the NUL-terminated invariant holds and the DoS
vector is not reachable.
Will revisit when Debian publishes a patched libjq1.
* feat: workstream attachments (images + text documents)
Adds end-to-end support for attaching images (png/jpeg/gif/webp) and
plain-text documents (markdown, source, JSON, etc.) to a workstream's
next user turn via the web UI.
Storage: new workstream_attachments table (migration 037) with a
three-state lifecycle — pending → reserved → consumed — scoped by
(ws_id, user_id) and linked to conversations.id on consume. Rewind/
truncation cascades attachment rows; delete_workstream does too.
Session: ChatSession.send(attachments, send_id) builds multipart user
content (text + image_url + document parts) and persists text-only to
conversations with attachments joined on load via message_id. Queue
path carries ordered attachment_ids plus a reservation token so
queued multimodal turns can't lose files to overlapping sends.
Providers: internal document content parts translate at the API
boundary — Anthropic emits native document blocks (text/plain
coerced, original MIME folded into title); OpenAI Chat Completions
and the Google OpenAI-compat endpoint inline them as escaped
<document> text blocks (XML-attr escape + </document> neutralization);
Responses API emits input_text with the same wrapper.
Server: POST/GET/DELETE /v1/api/workstreams/{ws_id}/attachments with
multipart upload (magic-byte image sniffing, UTF-8 enforcement for
text, per-kind size caps, Content-Length pre-check, per-(ws,user)
pending cap + TOCTOU lock). /v1/api/send reserves before dispatch
using a full-UUID token, threads it into session.send / queue_message,
releases on worker-thread failure, and reports attached/dropped ids
so the UI can reflect partial reservations. GET /content sets
X-Content-Type-Options, CSP sandbox, inline Content-Disposition, and
forces text/plain for text kinds. Ownership failures mask as 404.
UI: paperclip button, hidden file input with accept allowlist, chip
strip above textarea, drag/drop + paste-image handlers. Chips
rehydrate on ws switch and on queued-message dequeue; send clears
only attached ids and shows a toast when some dropped. Historical
user messages render filename pills via a _attachments_meta sibling
populated on both live-send and reconstruct paths.
530 tests covering CRUD, reservation lifecycle, races (TOCTOU cap,
reserve-then-dispatch overlap), provider translation, XSS headers,
cascade delete, history round-trip, and service-scoped actor flow.
* fix(attachments): address PR review feedback
- get_attachment_content now scopes the row by user_id too, so an
unowned workstream can't be a vector for cross-user blob fetches
via attachment_id guessing (Copilot, server.py:2676)
- send_message rejects attachment_ids lists longer than the pending
cap with 400 — prevents hostile clients from blowing up the
storage IN (...) clause (Copilot, server.py:1515)
- _attachment_upload_locks switched to a bounded LRU OrderedDict;
evicts the oldest unlocked entries past the soft cap so the map
can't grow unboundedly on long-running nodes (Copilot, server.py:2417)
- Pane.dragleave handler uses relatedTarget instead of target so the
drop-zone styling clears correctly when the cursor moves through
child elements; dragend listener added as a fallback for cancelled
drags (Copilot, app.js:297)
- uploadAttachment always cleans up the placeholder chip on failure,
including auth errors — no more stuck "uploading..." chips after
re-auth (Copilot, app.js:427)
- New _swapPlaceholderChip / _removeAttachmentChip helpers preserve
user-selection order through the placeholder→real-id swap; the
pendingAttachments Map is rebuilt in place rather than naïvely
delete+set, which would have moved the entry to iteration end
(Copilot, app.js:420)
- Drop unused `var self = this;` in removeAttachment (github-code-quality)
- Two regression tests: cross-user fetch on an unowned workstream,
and oversized attachment_ids list rejection
* fix(attachments): switch upload-lock to threading.Lock to avoid 3.12 CI hang
The per-(ws, user) upload lock was a module-cached asyncio.Lock.
Starlette's TestClient runs each request on a fresh anyio task /
event loop, so the cached lock's internal _waiters bind to the first
loop that acquired it. When a later request runs in a different
loop, await lock.acquire() blocks on a Future from a closed loop —
silent deadlock.
This surfaced as test (3.12) hanging indefinitely in CI on one push
while the same suite passed on 3.11/3.13 and on the next push. Same
root cause is reproducible against any Starlette TestClient harness
on 3.10+; 3.12 just happens to surface it more often given changes
in how anyio + asyncio.Future interact across loop teardown.
Switched to threading.Lock — loop-agnostic, and the critical section
is one COUNT + one INSERT, short enough that briefly blocking the
event loop is fine. Updated the LRU-eviction probe accordingly
(threading.Lock has no public .locked(), so use a non-blocking
acquire+release as the "is it free?" probe).
TOCTOU pending-cap test still passes; full attachment suite passes
on both 3.12 and 3.13.
* replace bitnami pgbouncer wit edoburu
replaced bitnami pgbouncer with edoburu pgbouncer container and updated environment variables to fit
* updated ports & Kubernetes
Updated ports to fit existing documentation. Also updated the Kubernetes Helm Chart link to use the same container.
* chore(deps): update dependency hls.js to v1.6.16
* chore: download vendored hls.js files + add hls to workflow detection loop
The wheel-completeness check failed on the Renovate bump because
vendor-js.yml only iterated katex/hljs/mermaid — so hls.js PRs
never got their files auto-downloaded. Adding hls to the loop so
future Renovate bumps are merge-ready without manual intervention.
Also running the update now to fix this specific PR.
---------
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: Patrick Buckley <buckleypm@gmail.com>
* feat: pass resolved capabilities through to providers, add server compat layer
The LLMProvider protocol previously forced providers to re-derive
capabilities from static lookup tables, ignoring config overrides set
via the admin UI or config.toml (e.g. thinking_mode, token_param).
This adds an optional capabilities parameter to create_streaming and
create_completion so the session can pass its config-merged
ModelCapabilities through to providers.
On top of this, adds a server compatibility layer for local model
servers (vLLM, llama.cpp). Profiles suggest thinking mode and server
workarounds (skip_special_tokens for vLLM, reasoning_format for
llama.cpp) during model detection, with structured admin UI fields
for server type, thinking mode, and extra body params.
Verified against real vLLM (Gemma 4 31B) and llama.cpp (Gemma 4 E4B)
servers.
* fix: defensive copy in _finalize_extra_body, expose thinking_param in UI
Shallow-copy extra_params and its chat_template_kwargs in the provider
before _apply_thinking_mode mutates them, so callers that reuse the
same dict across models are safe.
Replace the hidden thinking_param input with a visible text field
that appears when thinking mode is enabled. Shows the default
"enable_thinking" and hints that Granite/DeepSeek use "thinking".
* fix: address Copilot review feedback on admin UI and server compat
- Preserve unrepresentable thinking_mode values (e.g. "adaptive") in
raw capabilities JSON instead of silently dropping on edit round-trip
- Validate capabilities and extra body JSON are plain objects, not
arrays or primitives
- Deep-merge chat_template_kwargs from extra_body instead of silently
dropping, so operators can extend/override template kwargs
* fix: hide server compat section for non-local providers
The Server Compatibility fields (server type, thinking mode, extra
body) only apply to openai-compatible (local model servers). Hide
the entire section when the provider is openai, anthropic, or google.
* fix: normalize capsObj to plain object on edit load
Defend against DB rows where capabilities is a JSON literal null,
an array, or a primitive — previous code would crash on the
capsObj.server_compat / capsObj.thinking_mode reads. Same defensive
check also applied to the server_compat nested value.
* refactor: extract _isPlainObject helper for JSON type checks
Consolidates the null/array/typeof check that was inlined at three
different call sites into a single helper. Keeps the intent obvious
at each use site and avoids the awkward multi-condition ternary.
* feat: per-model sampling parameters (temperature, max_tokens, reasoning_effort)
Model sampling parameters were global-only settings applied uniformly to
all models. Different models have fundamentally different requirements
(o-series needs no temperature, Anthropic needs temp=1.0 with thinking,
local models may need different max_tokens). This adds per-model overrides
with global fallback so each model definition can specify its own defaults.
Migration 036 adds nullable temperature, max_tokens, reasoning_effort
columns to model_definitions. NULL inherits the global default from
ConfigStore. The session factory and /model switch command both resolve
per-model override → global fallback consistently.
The admin UI model create/edit modal now has dedicated form fields for
these parameters with client-side validation, a visual section divider,
and per-model override hints in the model table rows.
Removes vestigial model.name and model.context_window global settings
(now handled per-model by the model registry) with startup warnings for
existing config.toml users.
* fix: defensive parsing for config.toml per-model sampling params
Wrap temperature/max_tokens conversions in try/except with range
validation. Invalid values log a warning and fall back to None
(inherit global default) instead of aborting registry load.
* fix: use gethostname() instead of getfqdn() for advertise URLs
socket.getfqdn() does a reverse DNS lookup that often returns a
truncated hostname (e.g. "flat" instead of "flat-blck-io"). Use
gethostname() for advertise URLs in both server and console. For TLS
SANs, include both names so certs cover all variations.
* docs: clarify advertise URL comment re Docker/k8s
* fix: standardize database env vars on TURNSTONE_DB_* naming
compose.yaml used DB_BACKEND/DATABASE_URL in .env which got mapped to
TURNSTONE_DB_BACKEND/TURNSTONE_DB_URL inside containers. Running bare-
metal required the TURNSTONE_ prefix, but docs didn't explain this.
Eliminate the indirection — use TURNSTONE_DB_BACKEND and TURNSTONE_DB_URL
everywhere (compose, .env, bare-metal, docs, bootstrap wizard).
* fix: update .env.example to use TURNSTONE_DB_* naming
* fix: universal tool_call/tool_result orphan detection for OpenAI-compat providers
The Anthropic provider had orphan detection for mismatched tool_call ↔
tool_result pairs, but OpenAI-compatible providers (Chat Completions,
Google, Responses API) had none. When an Anthropic model runs behind
an OpenAI-compat API (e.g. Azure) or cancellation creates orphans,
the API rejects the malformed request.
- Rewrite sanitize_messages() with orphan detection: synthesize error
tool results for unmatched tool_calls, drop tool results with no
matching tool_call, fill empty tool_call IDs with positional remap
- Call sanitize_messages() from Responses API _convert_messages()
* fix: address review feedback on orphan detection
- Track answered IDs per-turn (local_answered) instead of scanning
all of out, preventing false matches from reused IDs across turns
- Drop empty-ID tool results that have no remap entry instead of
passing them through with invalid empty tool_call_id
- Increment empty_result_idx for every empty result, not just remapped
- Remove dead result_ids peek-ahead code
- Add test for repeated tool_call IDs across turns
* fix: accurate token usage tracking for compaction across all providers
Anthropic's input_tokens excluded cached tokens, causing massive
under-reporting (e.g. 327 vs 9000 actual) when prompt caching was
active. This prevented auto-compaction from triggering.
- Normalize Anthropic prompt_tokens to total input (input_tokens +
cache_creation + cache_read), matching OpenAI semantics
- Reset _last_usage per API call so tool-chain iterations get fresh
usage instead of max()-merging with stale values
- Add mid-turn compaction check during tool chains to prevent context
overflow before end-of-turn
- Anchor _remaining_token_budget() on provider-reported prompt_tokens
with local estimates only for the delta since last API call
- Improve _msg_char_count() to include structural overhead (role,
tool_call_id, tool call IDs) and handle image tokens in calibration
- Emit status after every API call, not just end of turn
* fix: defensive null coercion and index clamping from review feedback
- Add `or 0` to all getattr calls for input_tokens/output_tokens in
Anthropic provider (streaming + non-streaming) to handle SDK nulls
- Use getattr for non-streaming input_tokens/output_tokens instead of
direct attribute access for consistency
- Clamp _calibrated_msg_count with min() in _remaining_token_budget()
to prevent stale state from over-slicing after compaction
The ws-check-hint animation clobbered the fadein's forwards fill,
making the checkbox invisible for 0.6s on card-body click — appearing
as a deselect-then-reselect. Remove the hint, the unused role=checkbox
on the card, and restore the original symmetric toggle behavior.
* fix(ui): improve delete workstream UX and accessibility
Card body click no longer deselects (prevents confusing red border loss);
checkbox pulse hint guides users to deselect affordance. Adds keyboard
navigation, aria-labels, hover feedback, animations, and neutral Close
button styling after deletion.
* fix(ui): remove duplicate a11y checkbox from delete-mode cards
Hide the visual checkbox from the a11y tree and tab order so the card
(role=checkbox) is the sole keyboard/screen-reader target. Addresses
Copilot review feedback about nested interactive elements.
* perf: reduce initial rebalance from ~1.5s to ~50ms on PostgreSQL
Increase seed_ring_buckets chunk sizes (PG 500→16k, SQLite 500→8k) to
cut network round-trips from 131 to 5. Add ConsoleRouter.populate_from_assignments()
to build the routing cache directly from computed assignments, eliminating the
65 536-row DB read-back. Router becomes ready in <1ms; DB persistence follows.
* fix: address review — populate after seed write, sync router version
Move router cache population after seed_ring_buckets() so the router
is never "ready" with an unpersisted ring. Pass the new rebalancer
version to populate_from_assignments() so check_version() on the
collector thread does not trigger a redundant 65 536-row refresh.
- Remove role="status" and aria-label during _promoteQueuedMessages
so screen readers don't announce stale "queued" context
- Mark element with pendingDismiss when user dismisses before msg_id
arrives; send deferred DELETE when the send response provides the ID
If the model responds without tool calls, the main loop exits
immediately — no tool-result seam exists for advisory injection.
Queued messages were silently orphaned in the OrderedDict. Now
flushed as regular user messages before emitting idle state.
Bug 1: Extract _promoteQueuedMessages() — removes badge, dismiss
button, queued classes, and data-msgId. Called from setBusy(false)
on state_change: idle.
Bug 2: _dequeueMessage no longer removes the DOM element when server
returns not_found (message already injected). Only removes on
"removed" (actually dequeued). Network errors also preserve the
element. The promote loop handles cleanup on idle instead.
* feat: tool result advisory system with user message queuing
General-purpose advisory injection for tool results — when advisories
are present, tool output is wrapped in <tool_output> tags with
<system-reminder> blocks appended. Two initial producers:
- Output guard advisories: model sees why content was flagged/redacted
- User message interjections: users can queue messages mid-execution
via the web UI, injected at the next tool-call seam
Queued messages use !!! prefix for important priority. Advisory
injection is gated by ModelCapabilities.supports_tool_advisories
(default true for commercial models, false for local/vLLM).
On cancel/error, queued messages are flushed as regular user messages
so nothing is silently lost. Raw tool output (pre-wrap) is persisted
to the DB to keep history clean of ephemeral advisory XML.
* fix: frontend UX for queued messages — rollback, discoverability, a11y
- Send button changes to "Queue" (outline style) during busy state,
visually distinct from filled red Stop button
- Placeholder updates to hint at !!! priority convention
- addQueuedMessage returns element ref for optimistic UI rollback
- Remove queued element on queue_full, busy, or connection error
- Add role="status" and aria-label to queued message elements
- Promote queued messages to normal appearance when generation ends
* feat: queued message removal via dismiss button
Switch backing store from queue.Queue to OrderedDict + Lock for O(1)
removal by ID. Each queued message gets a UUID, returned to the
frontend and stored as data-msg-id on the DOM element.
Dismiss button (x) on queued messages calls DELETE /v1/api/send with
the msg_id. If the message was already injected (race), server returns
not_found and the UI removes the element anyway.
No new endpoint — DELETE method added to the existing /v1/api/send
route. dequeue_message() on ChatSession is O(1) under the lock.
* fix: address PR review — escaping, types, list output, message cap
- Escape </tool_output> and <system-reminder> in tool output to prevent
wrapper tag injection from untrusted tool results
- Change _collect_advisories return type from list[Any] to list[ToolAdvisory]
- Drain queued messages on list/structured output (append as text part)
so they aren't silently stuck until a str result appears
- Cap queued message length at 2000 chars to prevent context bloat
- Remove unused var in _dequeueMessage
* feat: replace workstream action buttons with per-tab dropdown menu
Move refresh-title, edit-title, fork, close, and delete actions from
the header toolbar into a dropdown menu on each workstream tab,
triggered by a ▾ chevron that replaces the × close button.
Dropdown follows the existing pane context menu pattern: keyboard
navigation, mutual exclusion, click-outside/Escape dismiss, toggle
on re-click, aria-expanded + aria-haspopup, and focus restoration.
Delete is visually distinct (red text + wash + red focus ring, 6px
separator). Mobile hides "Refresh title" and sizes the chevron to
36px touch targets.
Removes updateWsActionButtons(), _applyTitleButtonState(), and
_wsTitleState tracking (dead code after button removal).
* fix: remove Ctrl+Shift+R shortcut that overrides browser hard refresh
Refresh title is a low-frequency action accessible from the tab
dropdown; no replacement keybind needed.
* fix: address tab dropdown review findings
- Pass wsId through dropdown actions so they target the correct
workstream even when opened on a non-active tab
- Fix setTimeout race where closeTabDropdown before timeout fires
could leave stale listeners
- Guard Close and Delete on last workstream (dropdown, keyboard
shortcuts, and defense-in-depth in confirmDeleteWorkstream)
- Use aria-disabled instead of disabled so screen reader users can
discover unavailable items via arrow keys
- Enlarge chevron hit target, add hover affordance with subtle
background highlight
- Add 0.1s dropdown open animation (respects prefers-reduced-motion)
- server.py: catch (Exception, GenerationCancelled) instead of
BaseException so KeyboardInterrupt/SystemExit propagate normally
- judge.py: log client close failures instead of bare pass
Gemini's OpenAI-compat endpoint requires thought_signature to survive
the tool-call round-trip. Previously dropped because the Chat Completions
provider cherry-picks only standard fields (id, type, function).
Fix: GoogleProvider now captures raw tool-call dicts (including
thought_signature) via provider_blocks — the same fidelity lane the
Anthropic provider uses for signature round-tripping. On the next turn,
_prepare_messages reconstructs tool_calls from the stored raw data and
strips _provider_content so it never reaches the wire.
Changes:
- _openai_chat.py: add _prepare_messages and _extract_tool_calls hooks
- _google.py: override hooks + tap-pattern _iter_stream for streaming
- model_registry.py: auto-detect .googleapis.com → google provider
- session.py: read cancel_on_approval from ConfigStore
- console/server.py: add PUT/DELETE to proxy route methods
- server.py: fix fork naming (don't inherit source display name)
The channel gateway registers with its Docker-internal hostname
(e.g. http://channel:8091) which is unreachable from a host-side
server. Publish port 8091 and set TURNSTONE_CHANNEL_ADVERTISE_URL
to localhost so the server can reach it for schedule notifications.
GenerationCancelled extends BaseException, not Exception, so it bypassed
the except handler in _run_initial. The finally block ran but
_extract_last_assistant_content returned "" (response never appended to
messages), and _fire_notify_targets bailed on the empty content guard.
Fixes:
- Catch BaseException (not just Exception) in _run_initial so
GenerationCancelled is handled and the UI state is cleaned up
- Remove the empty-content suppression in _fire_notify_targets —
scheduled tasks should always deliver, even with a fallback message
when no output was captured
- Move action buttons (refresh/edit/fork/delete) from header to tab bar,
grouped in #ws-action-group with separators. Contextually adjacent to
the workstream tabs they operate on.
- Toggle group visibility via CSS class (.hidden) instead of per-button
inline style.display — makes media query overrides reliable.
- Call updateWsActionButtons() from renderTabBar() so buttons appear on
initial load and ws_created, not just on tab switch.
- Fix theme loss between nodes: loadInterfaceSettings no longer overwrites
localStorage with server defaults — preserves user's theme choice when
switching nodes via console proxy.
- Add flex-shrink:0 on +/split buttons to prevent squeeze with many tabs.
The model-reload handler read model.default_alias from ConfigStore's
in-memory cache, which could be stale if the earlier best-effort
config-reload notification failed or hadn't arrived yet. Force a
cs.reload() from DB before reading the alias. Also publish config
changes from the console before dispatching model-reload, and
downgrade the misleading "No 'default' model alias" log to debug.
Ctrl+Shift+R Refresh title (regenerate via LLM)
Ctrl+Shift+E Edit title
Ctrl+Shift+F Fork workstream
Ctrl+Shift+X Delete workstream (X not D — avoids Chrome DevTools conflict)
Shortcuts are blocked when any modal is open (edit-title, delete-ws,
batch-delete, new-ws). Help dialog (?) updated with the new bindings.
* feat: add per-node metadata with auto-collection, admin API, and console UI
Adds a normalized node_metadata table for structured per-node key/value
metadata with source tracking (auto/user/config). Auto-populated fields
(hostname, OS, arch, interfaces, cpu_count) are collected at server startup
via stdlib; user-defined fields are managed through the admin API, CLI, or
config.toml [metadata] section.
Storage: migration 035, 7 new protocol methods (get, get_all, set,
set_bulk, delete, delete_by_source, filter), both SQLite and PostgreSQL
backends. Filtering uses single-query GROUP BY/HAVING for efficiency.
Console API: GET/PUT/DELETE endpoints under /admin/nodes/{node_id}/metadata
with auto-source protection. cluster_nodes gains meta.* query param
filtering; cluster_node_detail attaches metadata to responses.
Frontend: new Nodes admin tab with collapsible per-node sections, inline
add form, delete with confirmation. Read-only metadata panel in node
detail drill-down. Proper design token usage, accessibility (ARIA,
keyboard nav, screen reader labels), and mobile responsiveness.
CLI: turnstone-admin list-node-metadata, set-node-metadata, and
delete-node-metadata subcommands.
64 tests (25 storage, 19 node_info, 20 existing unaffected).
* fix: resolve CI typecheck and test failures
- Fix mypy error: use %-style format string instead of structlog kwargs
for standard Logger.warning() in console server
- Fix test_get_nodes assertion to include new node_ids=None parameter
- Add debug logging to _collect_interfaces empty except block
* fix: address Copilot review feedback on node metadata
- Clear stale auto/config metadata before upserting on startup
- Wrap metadata filter in try/except with graceful fallback
- Add metadata field to NodeDetailResponse schema
- Use _VALID_NODE_ID regex for consistent node_id validation
- Defensive JSON decode in admin_get_node_metadata
- Switch to read_json_or_400 and require_storage_or_503 helpers
- Add SetNodeMetadataValueRequest for single-key PUT endpoint
- Add bulk GET /admin/node-metadata endpoint (replaces N+1 fetches)
- Update frontend to use single bulk metadata fetch
* feat: add admin.nodes permission scope for node metadata
- Add admin.nodes to builtin-admin role via migration 035
- Switch all node metadata handlers from admin.settings to admin.nodes
- Register admin.nodes in the admin panel permission set
- Node detail metadata panel fetches from cluster endpoint (no admin
permission needed) instead of admin endpoint
* fix: address second round of Copilot feedback
- Replace inline onclick handlers with data-* attributes and event
delegation to prevent JS string context XSS
- Move NodeMetadataEntry before NodeDetailResponse and use it as the
typed metadata field (was list[dict[str, Any]])
- Clean up config metadata on shutdown (was only cleaning auto)
Add save_messages_bulk() to StorageBackend protocol and both backends.
Fork path now inserts all messages in a single transaction instead of
N individual save_message() calls — for a 200-message workstream this
goes from 200 connection/insert/commit cycles to 1.
FTS5 indexing is intentionally skipped for bulk fork data (historical
messages indexed on rebuild). Ordering preserved via auto-increment id
with a shared timestamp across all rows in the batch.
Also adds 22 endpoint tests covering the 6 new workstream management
endpoints (delete, open, title, refresh-title, list/update interface
settings) and 4 storage-level tests for the bulk insert path.
- Replace inline-style console banner with CSS classes + light/dark theme
- Node ID in banner is now a clickable link back to the node UI
- Add judge_model parameter to create_workstream flow
- Add Google to model provider list with default URL
- Provider-specific placeholder hints in model editor
- Detect results populate model name suggestions datalist
- Theme changes in admin settings apply immediately
- Persist theme selection to server via settings API
- Use workstream title field (with name fallback) in collector SSE events
- Add judge model dropdown to new-workstream modal
Add workstream forking (resume with fork=True keeps new ws_id), custom
naming via aliases, title refresh via LLM, and workstream deletion.
New server endpoints: delete, refresh-title, set-title, open-workstream,
list/update interface settings. Verdict caching with SSE replay on
reconnect, display name fallback (alias→title→name) across all
endpoints, judge_model override per workstream, and settings_changed
broadcast on config reload.
New settings: judge.cancel_on_approval, interface.close_tab_action,
interface.theme. Storage backends updated with name in
list_workstreams_with_history and new get_workstream_metadata method.
Add workstream action buttons in header (refresh title, edit title, fork,
delete) with supporting modals and keyboard shortcuts.
Workstream tabs: always-visible close button, ws_id badge, configurable
close-tab-action (last_used/nearest/dashboard) via interface settings.
Dashboard: batch delete mode with multi-select, saved workstream cards
with ws_id badge, open endpoint for resuming sessions.
Judge display: late-arriving verdict toast when DOM element is gone,
worst-case verdict glow across all tool calls in approval block.
Theme: server-persisted via admin settings API, real-time sync across
clients via SSE settings_changed events.
New workstream modal: judge model dropdown for per-workstream judge
model selection.
- Create fresh HTTP client per evaluation run to avoid stale connections
- Store client factory args instead of client instance for on-demand creation
- Add cancel_on_approval config: when True, abort remaining items on user
approval; when False (default), run all evaluations to completion
- Always deliver LLM verdicts via callback (or fallback when LLM returns None)
- Add _deliver_fallbacks helper for cancelled/incomplete evaluations
- Skip read-only tools for Google provider (requires thought_signature)
- Flatten conversation history to plaintext transcript in _prepare_context
to avoid multi-turn role sequence errors with strict providers like Google
- Use per-turn timeout instead of shared budget so slow turns don't starve
later ones
- Add empty-response retry logic (up to 3 retries without consuming turns)
- Enhanced structured logging throughout judge pipeline
- Update tests to match new signatures and behavioral changes
Add GoogleProvider that extends OpenAIChatCompletionsProvider for
Gemini models via the OpenAI-compatible /v1beta/openai/ endpoint.
- New _google.py with 2M context window defaults and vision support
- Lazy-initialized singleton in create_provider() (thread-safe)
- Route 'google' through OpenAI SDK in create_client()
- Return empty list from list_known_models() (Google models change frequently)
* feat: reconcile judge admin rule UX with edit, disable, and reset actions
Replace the misleading "Customize" button on built-in rules with a
logically consistent 4-state action model: pure built-in (Disable/Edit),
overridden built-in (Disable/Edit/Reset), disabled built-in
(Enable/Edit/Reset), and custom rule (Enable-Disable/Edit/Delete).
Add edit modals for both heuristic rules and output guard patterns,
reusing the existing create modal form structure. Introduce amber
"Reset" button styling to visually distinguish reversible resets from
permanent deletes. Fix source badge redundancy (disabled built-ins now
show grey "built-in" in SOURCE, red "disabled" in STATUS only). Add
aria-labels and role="listitem" for screen reader support.
* fix: preserve built-in pattern_flags and priority on override
Derive pattern_flags from compiled regex for built-in output guard
patterns in the list API so IGNORECASE and other flags survive the
disable/edit/override round-trip. Carry priority through edit modals
via hidden fields so built-in evaluation order is preserved.
* feat: auto-invalidate JWT and static assets on version upgrade
Add a `ver` claim (major.minor) to user-facing JWTs so tokens from
previous versions are rejected after upgrade, triggering re-login.
Service tokens are excluded for rolling-deployment safety. Tokens
without a `ver` claim (pre-upgrade) are accepted for backward compat.
Inject `?v={__version__}` query strings into static asset URLs at
startup so browsers fetch fresh JS/CSS after any release. Vendored
libraries (KaTeX, Highlight.js, etc.) are skipped since they already
carry version numbers in directory paths. HTML responses now include
`Cache-Control: no-cache` to ensure browsers always revalidate.
Frontend detects upgrade-specific 401s and shows a contextual subtitle
("The server was updated — please sign in again"), then performs a full
page reload after re-auth to load the new versioned assets.
* refactor: address PR review — public API name, single decode, idempotent regex
Rename _version_slot() → jwt_version_slot() to make the cross-module
import explicit rather than relying on a private name.
Move version gating from validate_jwt() into check_request() via a new
AuthResult.token_version field. This eliminates the double JWT decode
that occurred on version-mismatch detection — the token is now decoded
once and the version compared afterward.
Guard version_html() regex against double-apply by excluding URLs that
already contain a query string ([^"?]+ instead of [^"]+).
* feat: structured version_mismatch code, ETag, cross-tab auth sync
Add structured "code": "version_mismatch" field to the 401 response
so the frontend detects upgrade-triggered re-auth without string
matching on the error message.
Add ETag headers to HTML index responses (server, console, and proxied
node UI). Combined with Cache-Control: no-cache, browsers send
conditional GETs and receive 304 between upgrades, saving bandwidth.
Add BroadcastChannel-based cross-tab auth sync so logging in on one
tab dismisses the login modal on all other tabs (and vice-versa for
logout).
Add a reminder to the vendored JS update script about the
version_html() regex lookahead.
* fix: remove unused import in test_web_helpers
* feat: Discord /ask model alias, channel default setting, admin UX
Add optional 'model' parameter to Discord /ask command with
autocomplete from available aliases. Model precedence:
explicit > channels.default_model_alias > CLI --model > server default.
- Add channels.default_model_alias to settings registry
- Extend /v1/api/models response with default_alias and
channel_default_alias fields (both server and console)
- Add list_models() to async + sync SDK clients and ChannelRouter
- TTL-cached channel default in ChannelRouter (5min, fail-open)
- @mention path also respects channel default
- Admin Settings tab: model alias settings render as dropdowns
populated from enabled model definitions
- Admin Settings tab: is_secret settings render as write-only
password inputs with save button (replaces static label)
- Update OpenAPI schemas for new response fields
- Validate alias defaults against enabled models on both endpoints
* fix: address PR #306 review feedback
- Move TTL timestamp update before await in get_channel_default_alias
to prevent concurrent duplicate fetches
- Add 30s TTL cache for list_models() to avoid per-keystroke HTTP
traffic during Discord autocomplete
- Type SDK list_models() with ListAvailableModelsResponse instead
of raw dict (both server and console, async + sync)
- Fix IntentJudge.__init__() control flow: model override block was
dangling inside try/except instead of being a separate branch
- Remove provider/base_url/api_key kwargs from server.py and cli.py
JudgeConfig construction (fields removed in prior commit)
- Remove stale TOML mapping entries from config.py
- Remove --judge-provider CLI argument
- Fix Judge settings font sizes to match Settings tab (12px keys,
11px descriptions, tighter spacing, --fg instead of --accent)
Judge model config now uses model aliases exclusively via ModelRegistry.
The separate provider, base_url, and api_key fields on JudgeConfig were
redundant with what's already stored in model definitions. Removes the
fields from JudgeConfig, the explicit-provider resolution path from
IntentJudge.__init__(), and the 3 settings from the registry.
- Replace all raw fetch() + _adminToken with authFetch() helper
- Fix URL paths to use /v1/api/admin/judge/ prefix
- Add r.ok checks on all GET fetches (match existing tab pattern)
- Load model definitions before settings to fix picker race condition
- Escape secret input values with escapeHtml
- Use Mapping type for evaluate_output patterns param (mypy)
- Clean up stale blank lines and comment references
* feat: configurable judge rules with dedicated admin tab
Externalize heuristic intent validation rules and output guard patterns
from hard-coded module constants into the storage abstraction with full
admin UI CRUD. Introduces a dedicated Judge tab in the admin panel that
consolidates all judge configuration (scalar settings, heuristic rules,
output guard patterns) under a single admin.judge permission scope.
- Add heuristic_rules and output_guard_patterns tables (migration 033)
- Add RuleRegistry with thread-safe merge of built-in + DB rules
- Refactor output_guard.py patterns into structured OutputGuardPatternDef
- evaluate_heuristic() and evaluate_output() accept optional rules/patterns
- IntentJudge resolves model aliases via ModelRegistry
- 15 admin API endpoints under /api/admin/judge/ with regex validation
- Judge tab with Settings, Heuristic Rules, and Output Guard sub-panels
- Filter judge.* settings from generic Settings tab
- ConfigStore.storage public property for backend access
* fix: align Judge tab with admin panel design system
- Replace raw <table> with grid-based admin-row/admin-colheaders pattern
- Replace dynamic innerHTML modals with static overlays using focus traps
- Replace confirm() with styled showConfirmModal()
- Replace inline badge styles with scope-badge classes
- Add mobile responsive breakpoints for Judge tab grids
* fix: Judge tab accessibility and polish
- Extract sub-section switcher inline styles to CSS classes
- Add focus-visible outline and reduced-motion support
- Add tab button IDs and fix aria-labelledby on tabpanels
- Add tabindex roving and arrow key navigation for sub-tabs
- Add role=list and aria-live to table containers
- Replace status text with scope-badge classes for scannability
* fix: address CodeQL and Copilot review feedback
- Remove unused validation constants from rule_registry.py (CodeQL)
- Return MappingProxyType from output_patterns for immutability
- Fix ThreadPoolExecutor shutdown(wait=False) to prevent hangs
- Use separate _VALID_OG_RISK_LEVELS (no "critical") for output guard
- Pass pattern_flags to regex validation in update endpoint
- Chain redactions in configurable mode (compose pattern + complex)
- Initialize RuleRegistry on console app.state
- Fix test fixtures to use valid enum values (approve/review/deny)
* fix: use Mapping type for evaluate_output patterns param (mypy)
* feat: multi-model health tracking with runtime default and DB-only startup
Replace active-probe circuit breaker with passive per-backend health
tracking. Backends are marked degraded after consecutive failures and
recover when a request succeeds — requests are never blocked.
- Add model.default_alias ConfigStore setting for runtime default model
- Make load_model_registry CLI args optional for DB-only startup
- Per-(provider, base_url) health trackers via HealthTrackerRegistry
- Two-pass fallback: prefer healthy backends, then try degraded
- Remove BackendHealthMonitor, CircuitState, probe threads, cooldown
- Remove circuit_state from API schema, SDK events, metrics, frontends
* feat: add "Set Default" button to Model Definitions admin panel
Show a "default" badge on the current default model alias and a
"set default" action button on all other models. Clicking it writes
model.default_alias via the settings API. The list endpoint now
includes default_alias in the response so the UI can highlight it.
* fix: address review feedback — metric scoping, effective default, session alias
- Move turnstone_backend_up metric out of BackendHealthTracker into
server callback; only the effective default backend drives the gauge
- _build_health_dict resolves effective default via ConfigStore override
- session_factory computes selected_alias once before registry.resolve
- admin model-definitions endpoint returns effective default (not just
override) so UI shows correct badge when ConfigStore is empty
- Rename circuitTitle → healthTitle in console JS
- Fix ruff SIM117 lint in test
* fix: validate effective default against enabled models, degraded label, log normalization
- admin model-definitions endpoint validates default_alias against
enabled models using same fallback rules as load_model_registry
- UI text "backend down" → "backend degraded" to match advisory semantics
- Health tracker log uses normalized base_url from key, not raw argument
Migration 031 created the prompt_policies table but never registered
admin.prompt_policies in _VALID_PERMISSIONS or granted it to the
builtin-admin role, causing 403 on all prompt-policy admin endpoints.
* fix: capacity-aware tool output truncation and context overflow recovery
Large tool results (e.g. 593K-char search output) could overflow the
context window in a single turn when the conversation was already
partially full. The fixed 50%-of-context truncation limit didn't
account for current usage.
Changes:
- _truncate_output() now accepts remaining token budget and uses
min(tool_truncation, remaining_budget_chars) as the effective limit
- _remaining_token_budget() helper calculates available capacity with
reserves for max_tokens response and 5% safety margin
- Safety truncation at tool-result append: every string tool result is
clamped to remaining budget before entering the message array
- _exec_web_search() now calls _truncate_output() (was missing)
- Context overflow recovery: catches provider errors indicating context
length exceeded (OpenAI + Anthropic patterns), auto-compacts, retries
once. Falls back to original error if compact-and-retry fails.
* fix: address review — zero-budget floor, nested spinner, Anthropic patterns, tests
- Remove 256-char floor from budget truncation — zero budget now returns
a placeholder instead of allowing 256 chars through
- Stop thinking spinner before compact to avoid nested start/stop
- Add Anthropic error patterns (prompt is too long, input tokens)
- Wrap compact-and-retry so failures re-raise the original error
- Add 15 tests covering budget calculation, capacity-aware truncation,
and overflow recovery for both providers
* fix: cap response reservation at 25% of context window
Reserving the full max_tokens in _remaining_token_budget() zeroed the
budget for common configs like max_tokens=32768 on a 32K context,
collapsing all tool output to a placeholder. max_tokens is a ceiling,
not guaranteed consumption — cap the reserve at context_window // 4.
Adds regression test for max_tokens >= context_window.
* fix: skip chat_template_kwargs for commercial OpenAI API
OpenAI rejects chat_template_kwargs as an unknown parameter — it's only
meaningful for local model servers (vLLM, llama.cpp, SGLang).
Split OpenAIProvider into separate singletons for "openai" vs
"openai-compatible" so _provider_extra_params can gate on provider_name
instead of inspecting base_url. Also deduplicates agent inline code into
the same method and fixes pre-existing test pollution where
get_capabilities was mutated on the singleton without cleanup.
* feat: add OpenAI Responses API provider for commercial models
Split the OpenAI provider into three concrete implementations behind the
LLMProvider protocol:
- _openai_chat.py: Chat Completions API for local model servers
(vLLM, llama.cpp, SGLang)
- _openai_responses.py: Responses API for commercial OpenAI
(GPT-5.x, O-series)
- _openai_common.py: shared capability table, temperature/reasoning
gating, cache retention, citations, usage extraction
The Responses API handles reasoning_effort as a {"effort": value} dict,
system messages as an instructions field, and tool format translation at
the provider boundary. ChatSession is unchanged — the provider abstracts
the API difference.
Also fixes diff_file direction when comparing against provided content.
* fix: Responses API input format and local model provider routing
- Assistant input messages use plain string content (not output_text)
- Tool call argument deltas match on item_id, not call_id
- Auto-detect openai-compatible provider for non-api.openai.com URLs
- Fix diff_file direction when comparing against provided content
* fix: resolve env vars before provider auto-detection in config.toml models
Config-file model entries using ${ENV_VAR} placeholders in base_url were
not resolving env vars before _resolve_openai_provider(), causing
commercial OpenAI configs to be misclassified as openai-compatible.
* fix: replace empty except blocks with diagnostic logging
Add log.debug/warning to 7 bare except-pass blocks that silenced
failures in security-relevant or operationally-important paths:
- Channel route lookup, CLI policy evaluation, OIDC JWKS fetch,
prompt policy loading, plan file write, routing override, username
resolution.
Plan write now reports failure to user instead of falsely claiming
"Plan saved."
* fix: replace assert-with-side-effect and narrow BaseException catch
- Convert 4 assert isinstance() to explicit TypeError raises — assertions
are stripped under python -O, removing runtime type checks
- Narrow except BaseException to except Exception in fallback handler —
KeyboardInterrupt/SystemExit should not record as health failures
- Plan write failure now reports error to user instead of "Plan saved"
* fix: wire up toast error type and remove useless conditional
- showToast() now accepts optional type param ("error") with red border
styling — 3 call sites were passing "error" that was silently ignored
- Remove always-true if (q) guard after early-return on empty query
* fix: remove unreachable return None after return self._judge
* fix: parenthesize multi-line string concatenations in dev_parts list
Explicit parens make intentional concatenation unambiguous to static
analysis (CodeQL implicit-string-concatenation-in-list rule).
* fix: remove constant-true filter in test mock — return list directly
* fix: extract side-effecting calls from assert in tests
store.delete() and mgr.close() have side effects that would be
stripped under python -O. Assign to variable first, then assert.
* fix: remove unused local variables in tests
Drop assignments to unused workstream/variable references created
solely for side effects. Use _ for unused tuple unpacking.
* fix: use admin.prompt_policies permission for prompt policy endpoints
All 5 prompt-policy endpoints (list, create, get, update, delete)
were checking admin.policies (the tool-policy permission) instead of
admin.prompt_policies. This caused a mismatch with the admin UI which
gates the tab on admin.prompt_policies — users could see the tab but
get 403, or reach the endpoint but never see the tab.
* fix: use caplog instead of capsys for structlog warning assertion
structlog output goes through the logging system, not stdout/stderr.
* fix: address review — remove dead isinstance, module-level import, unnecessary lambdas
- session.py: remove unreachable isinstance check (has_batch already
validates raw_edits is a list)
- cli.py: move logging import to module level
- test_workstream.py: replace lambda wid: FakeUI(wid) with FakeUI
* fix: harden MCP client against misbehaving servers
Misbehaving/failed/misconfigured MCP servers could peg CPU at 100% due
to anyio cancel-scope busy-loops (SDK #2147), uncancelled orphaned
futures, and missing application-layer resilience.
Five fixes:
1. Cancel orphaned futures on timeout — future.cancel() in all sync
bridge methods prevents coroutine accumulation on the event loop
2. Per-server circuit breaker — 3-failure threshold with exponential
cooldown (30s–5min), per-server jitter, auto-reconnect on half-open
probe, McpError excluded (protocol errors from healthy servers)
3. Safe transport stream pre-close — store stream refs and close them
before stack teardown in all error/shutdown paths, preventing the
anyio zero-buffer CPU busy-loop
4. Notification debounce — 5s per-server rate limit on list_changed
refresh storms from buggy servers
5. Periodic refresh backoff with auto-reconnect — disconnected servers
get reconnection attempts with exponential backoff (60s–1hr) instead
of being silently skipped forever
* docs: add MCP resilience section to architecture docs and diagram
Document the circuit breaker, future cancellation, stream pre-close,
notification debounce, and periodic refresh backoff in the architecture
guide and the MCP architecture PlantUML diagram.
* fix: address review — stack leak on transport error, half-open comment
- Widen _connect_one guard to check _per_server_stacks too, not just
_sessions. Transport errors in sync dispatch methods evict the session
but left the stack behind, leaking anyio tasks on reconnect.
- Clarify half-open design: multiple callers are intentionally allowed
through (reconnects serialize on the event loop, first failure re-trips).
* fix: mobile UX for console sidebar drawer and server chat input
Console admin sidebar: add box-shadow elevation, close button with
focus return, 44px touch targets, focus-into-drawer on open, flip
active indicator to left border, cubic-bezier easing, aria-expanded,
fix resize handler state desync, guard toggle injection for panels
without toolbars.
Server chat input: on touch devices Enter inserts newline (tap Send
button to send), hide Shift+Enter hint from placeholder.
* fix: preserve first group label spacing when close header is injected
Add sibling combinator selector so the first sidebar group keeps its
reduced top padding regardless of whether the close header div is
present as first-child.
* feat: render rich media embeds for MCP tool results
Detect structured media JSON (stream_url, results, sessions) in MCP
tool output and render interactive cards instead of plain text.
Web UI: media cards with thumbnail, title, metadata, and click-to-play
video/audio. HLS via lazy-loaded hls.js with direct-stream preference.
Collapsed raw JSON (API keys redacted) for inspection.
Discord: rich embeds with proxied thumbnail images (fetched by the bot
since Discord CDN cannot reach private media servers). Search results
as numbered lists, session state as "Now Playing" cards. Stream URLs
never exposed in embeds — web_url used for safe clickable links.
CI: vendor hls.js 1.6.15 with renovate tracking and update script.
* fix: address PR #292 review — SSRF guards, streaming fetch, tests
- URL validation: reject non-http(s) schemes and userinfo in thumbnail
URLs. Private IPs intentionally allowed (media servers are on LAN).
- Streaming fetch: use http.stream() with aiter_bytes() and a running
byte count to enforce the 2MB cap without buffering the full response.
Validate content-type is image/* before downloading.
- Resilience: wrap try_build_media_embed in try/except in bot.py so a
media embed failure falls through to the code-block path.
- LICENSE: download hls.js LICENSE from npm on update instead of only
copying from old dir.
- Tests: add 19 new tests — try_parse_media (8 cases), _is_safe_image_url
(7 cases), embed builders (4 cases including stream_url exclusion and
string season/episode safety).
* chore: add LICENSE file for vendored hls.js
* fix: remove ANSI escape codes from tool preview fields
Preview text (tool args, URLs, queries) was wrapped in DIM/RESET ANSI
codes at the source in session.py, which leaked into SSE events and
rendered as raw escape sequences in Discord and the web UI.
Move ANSI styling to the CLI consumer (cli.py) where it belongs. Also
escape markdown in Discord tool name titles to prevent __ from being
interpreted as underline formatting.
* fix: drop [MCP: server] prefix from tool descriptions
The prefix made MCP tools look second-class compared to builtins,
causing models to hesitate using them. The server name is already
encoded in the tool name (mcp__server__tool).
* feat: pretty-print JSON tool output, player error state, broader key redaction
- JSON tool results are detected and pretty-printed with 2-space indent
instead of rendering as a wall of text
- API key redaction extended to cover api_key, apiKey, api-key, and
token query params across all tool output (not just media embeds)
- Video/audio player shows styled error message when stream fails to
load instead of leaving a broken player element
- Both appendToolOutput and replayHistory use shared renderToolOutput()
* fix: designer review — player error retry, contrast, tool-cmd cap
- Player error: role="alert" for screen readers, retry button that
reuses existing play handler, includes media title in error message
- Light theme: darken --red from #dc2626 to #b91c1c (5.7:1 contrast
on --code-bg, was 4.3:1 failing WCAG AA at 12px)
- Pretty-print collapsed raw JSON in media embeds (was missed earlier)
- Cap .tool-cmd at 120px to prevent tools with many args from making
approval blocks disproportionately tall in history replay
- Dedicated .media-player-error class instead of reusing .tool-output
* fix: Discord tool info name matching regression, suppress deprecation warning
The escape_markdown call on tool names was stored for matching against
ToolResultEvent.name, but event.name is raw/unescaped. The escaped name
never matched, so the "Running → Done" transition silently failed and
previews disappeared from the status embed.
Fix: store raw name for matching, use escaped name only for display.
Also suppress discord.py's re.sub count deprecation warning (Python
3.13+ issue, fixed upstream).
* fix: update MCP tool description tests to match prefix removal
* fix: address PR #292 review round 2
- Retry button: handle missing span children in click handler so retry
buttons from player error state don't throw
- Footer count: use len(lines) instead of min(len(results), 10) to
reflect actual rendered count after char budget truncation
- Null display: use "null" instead of "None" in JS tool arg preview
- Broader redaction: also redact JSON "api_key": "..." patterns
- SSRF hardening: block loopback and link-local IPs plus cloud metadata
hostnames in thumbnail fetch (private LAN IPs still allowed)
* fix: bundle production compose.yaml for pipx users (#293)
Users who install via pipx don't have a git clone, so there's no
compose.yaml or Dockerfile. Bootstrap now extracts a bundled production
compose file that uses pre-built ghcr.io images instead of local builds.
- Add turnstone/deploy/compose.yaml (ghcr.io images, no build blocks,
single-node production profile only)
- Add write_compose tool to bootstrap wizard
- Update bootstrap system prompt to check for and write compose.yaml
- Remove stale ddgCluster profile references from system prompt
- Include turnstone/deploy/*.yaml in wheel
* fix: use postgresql+psycopg:// DSN scheme in compose fallbacks
The Docker image ships psycopg3, not psycopg2, so the bare
postgresql:// scheme fails. Also clarify PG usage comment in
production compose.
* fix: improve web_fetch reliability — strip scripts, dynamic truncation, more tokens
- strip_html() now removes <script>, <style>, <template>, <noscript>
element content instead of just their tags
- Truncation budget scales with context window (75% in chars, 50k floor)
and takes from the beginning only instead of head+tail splice
- max_tokens bumped from 2000 to 8192 so thinking models don't starve
the visible extraction answer
- reasoning_effort="low" on summarization call to avoid wasting tokens
- Empty responses and empty extractions now report as tool errors
* refactor: extract _utility_completion to fix reasoning_effort duplication
Callers previously had to pass reasoning_effort both as a direct keyword
(for commercial providers) and via _provider_extra_params (for local
model servers). This duplication was easy to get wrong — web_fetch was
already missing the direct keyword.
_utility_completion threads it through both paths from a single call,
used by title generation, compaction, and web_fetch extraction.
* fix: disable thinking when max_tokens too small, cap extraction at 500k
_reasoning_params now returns empty dict when max_tokens can't fit a
thinking budget (e.g. title gen with max_tokens=200). Previously
produced budget_tokens >= max_tokens which is an API error on
manual-thinking Anthropic models.
Also caps web_fetch content truncation at 500k chars — the dynamic
context-window calc was producing 3M chars on 1M-context models.
* fix: clamp utility max_tokens to model output limit, add strip_html tests
_utility_completion now clamps max_tokens to the model's advertised
max_output_tokens so small/local models don't reject 8192-token
requests.
Adds 8 tests for invisible element stripping (script, style, template,
noscript) including multiline, case-insensitive, and attribute cases.
* fix: mock get_capabilities in title retry tests for _utility_completion
_utility_completion calls _get_capabilities to clamp max_tokens. The
existing title tests mocked _provider as a bare MagicMock, so
caps.max_output_tokens was a truthy MagicMock instead of an int. Set
get_capabilities to return a real ModelCapabilities instance.
Build the image once via the profileless console service and reference
it as turnstone:local from server/channel. Prevents stale images when
users run docker compose build without --profile.
* fix: include prompt .md files in wheel, add wheel-completeness CI (#289)
Prompt markdown files were missing from PyPI wheels since the modular
prompts refactor, causing FileNotFoundError on startup for pip-installed
users. Add the missing include pattern and a new CI job that diffs
source-tree data files against wheel contents so omissions are caught
before merge.
* fix: sanitise ALLOW patterns in wheel-completeness check
Strip blank lines and leading whitespace from the allowlist before
passing to grep -vFxf so empty patterns cannot silently match all lines.
* fix: log clean one-liner when PostgreSQL becomes unavailable
Wrap all 174 connection sites in PostgreSQLBackend through a _conn()
context manager that catches OperationalError, emits a single
database.unavailable log line (with connection URL), and suppresses
repeats until the connection is restored (database.connection_restored).
* fix: add StorageUnavailableError and cover all heartbeat loops
Address review feedback:
- Separate connect-phase from execution-phase in _conn() so that
OperationalError during caller code (e.g. BEGIN IMMEDIATE lock
contention) is not misclassified as a connectivity failure.
- Add StorageUnavailableError exception class so callers can
distinguish transient DB outages without redundant tracebacks.
- Apply the same _conn() wrapper to SQLiteBackend for consistency.
- Catch StorageUnavailableError in all 7 periodic loops: watch
runner, server heartbeat, channel heartbeat, console heartbeat,
collector discovery, rebalancer, and scheduler.
- Guard dedup flag with threading.Lock.
- Add tests for dedup logging and PostgreSQL path.
* fix: chunk IN clauses to stay within DB parameter limits
psycopg caps query parameters at 65 535 and SQLite defaults to 999.
assign_buckets, prune_workstreams, and count_skill_resources_bulk were
passing unbounded lists into single IN(...) clauses, causing
OperationalError during rebalancer runs on full-size hash rings.
Chunk sizes: 10 000 (PostgreSQL), 500 (SQLite).
* fix: deduplicate assign_buckets input, add chunking regression tests
Address review feedback: deduplicate bucket list before chunking to
prevent inflated rowcount from cross-chunk duplicates. Add tests that
exercise the multi-chunk path (1200 buckets > SQLite chunk_size of 500)
and verify dedup preserves accurate counts.
Multiple CI completions for the same commit (tag push + branch push)
caused duplicate publish and docker runs. Concurrency group keyed on
head_sha ensures only one publish runs per commit.
* chore: release infrastructure for dual-track stable/experimental
CI/CD changes for the 1.0 release:
- Gate PyPI publish and Docker publish on CI success via workflow_run
- Add docker-publish.yml: builds and pushes to GHCR with smart tagging
(stable gets :X.Y.Z/:X.Y/:stable/:latest, pre-release gets :experimental)
- Add stable/* and v* tags to CI and docker-scan triggers
- Remove stale [mq] extra and types-redis from CI (Redis MQ deleted)
- Remove stale redis from Renovate package rules
Release tooling:
- scripts/release.sh: bump version, uv lock, commit, tag (with --push)
- docs/releasing.md: documents stable/experimental workflow
Docker:
- Add /workspace mount point (WORKSPACE_MOUNT env var, defaults to empty volume)
- Update .env.example: remove stale Redis/auth-token refs, add workspace/model/discord
README:
- Remove beta warning, add hero image and release tracks table
* fix: derive release tag from git instead of workflow_run.head_branch
Use git tag --points-at HEAD after checkout to resolve the release
tag instead of relying on workflow_run.head_branch, which may not
reliably be the tag name for tag-triggered CI runs. Both publish
and docker-publish workflows now skip cleanly when no v* tag exists
at the checked-out commit.
* fix: replay plan review prompt on SSE reconnection
Plan approval prompts were lost when a user navigated to the server
web UI from the console dashboard (triggering a new SSE connection).
Tool approvals stored pending state in _pending_approval and replayed
it on reconnection, but plan reviews used fire-and-forget _enqueue
with no persistent state.
Mirror the _pending_approval pattern: store _pending_plan_review
before blocking, replay it in events_sse for new SSE clients, and
clear it on resolution. Without this fix, plan reviews silently
timed out after 1 hour and were treated as approval.
* test: add plan review SSE replay regression tests
Covers pending state lifecycle: stored during on_plan_review, cleared
on resolve_plan, available for SSE reconnection replay.
The async variants accepted token_factory for auto-rotating JWTs via
ServiceTokenManager, but the sync wrappers did not expose or forward
the parameter. External SDK users calling the sync clients with
token_factory got a TypeError.
Settings rows and admin table rows used align-items: center, which
caused inputs and source badges to drift away from their labels when
descriptions wrapped to multiple lines. Switch to align-items: start
so controls stay next to their label names regardless of row height.
Add 2px top margin on settings toggles to pixel-align with text input
top padding in start-aligned rows.
* fix(examples): rewrite mcp-cluster-ops to use console SDK for cluster routing
The example MCP server was broken after the direct HTTP transport
refactor — it used TurnstoneServer (single-node) for cluster ops that
require TurnstoneConsole (cluster gateway). Rewrites dispatch flow to:
route via console → SSE stream from node → cleanup via console.
- Switch from TurnstoneServer to TurnstoneConsole for node listing and
workstream routing (TURNSTONE_CONSOLE_URL replaces TURNSTONE_SERVER_URL)
- Add proper workstream lifecycle: create via routing proxy, stream from
node, close in finally block with leak-safe ws_id guard
- Catch dispatch exceptions in run_on_node for structured JSON errors
- Extract _extract_node_ids helper, remove dead n.get("id") fallback
- Normalise _console_kwargs to always include token key
- Rewrite tests against Console+Server mocks (36 → 44 tests)
* fix(examples): paginate node listing and clarify auth in README
Address Copilot review feedback on #278:
- _list_nodes_sync now paginates via offset/limit loop so clusters
with >100 nodes are fully discovered
- README step 2 now mentions token passthrough for authenticated clusters
- New test_paginates_large_clusters verifies multi-page fetch (45 tests)
* fix: remove non-auth support from bootstrap wizard
Auth is now mandatory for all deployments. Remove the
TURNSTONE_AUTH_ENABLED toggle and make JWT_SECRET and AUTH_TOKEN
required in the wizard's system prompt.
* fix: remove auth disable support from runtime and infra
Remove AuthConfig.enabled field — auth is always on. Drop
TURNSTONE_AUTH_ENABLED env var, config toggle, and the
check_request bypass. Update compose.yaml, Helm chart,
Terraform, docs, and tests to match.
* feat: deprecate config tokens, require JWT secret, prefer JWT auth
Phase 1 of config-token removal:
- load_jwt_secret() now exits with error if no secret is configured
(was: silently auto-generated ephemeral secret)
- _authenticate_token() logs deprecation warning on config token use
- CLI /cluster commands use ServiceTokenManager when JWT secret is set
- turnstone-admin tls-list uses ServiceTokenManager when JWT secret is set
- Update bootstrap wizard, docker.md, security.md to mark
TURNSTONE_AUTH_TOKEN as deprecated and JWT_SECRET as required
- Console test fixtures use auth token + headers (auth always enforced)
* feat: add service scope for inter-service JWT auth
Add "service" to VALID_SCOPES and SCOPE_HIERARCHY. Service tokens
bypass require_permission() RBAC checks, replacing the old
empty-user-id bypass that config tokens relied on.
All ServiceTokenManager instances that need admin access now include
"service" in their scopes (console proxy, channel gateway, CLI,
admin CLI). Read-only services (collector, notification) unchanged.
* feat: phase 2 config token deprecation
- SDK doc examples now show API tokens (ts_) instead of config tokens
- Remove _get_config_token() from admin CLI (dead code)
- Block config token exchange in handle_auth_login — only password
and API token login allowed
- Update login tests to use password-based auth instead of config
token exchange
* feat: phase 3 — remove config tokens entirely
Complete removal of config-file token authentication:
- Delete AuthConfig.tokens, check(), _ROLE_TO_SCOPES, hmac dispatch
branch, and config token loading from load_auth_config()
- Remove auth_config parameter from _authenticate_token() and
check_request() — callers updated throughout
- Remove TURNSTONE_AUTH_TOKEN from compose.yaml, Helm charts,
Terraform, turnstone.example.toml
- Remove --auth-token CLI flags from turnstone, turnstone-admin,
and turnstone-console
- Simplify console main() — always use ServiceTokenManager
(no fallback to static tokens)
- Delete config-token-specific tests, rewrite check_request and
integration tests to use JWT auth with proper audience claims
- Remove all config token references from docs (security.md,
docker.md, sdk.md, console.md, architecture.md, bootstrap prompt)
* fix: address code review findings
- Fix 33 broken tests: add JWT auth to test_api_versioning,
test_console_routing_proxy, test_tls_admin, test_tls_manager,
test_server_live (jwt_secret + audience-scoped auth headers)
- Add TestRequirePermissionServiceScope: 4 tests covering the
service scope RBAC bypass path
- Remove stale comments referencing config tokens in auth.py and
console/server.py
- Remove dead proxy_auth_token parameter from console create_app()
and static token fallback in _proxy_auth_headers()
- Remove TURNSTONE_AUTH_TOKEN from env.py scrub list
* fix: address Copilot review — JWT audience, compose require secret
- CLI /cluster: add audience=JWT_AUD_CONSOLE to ServiceTokenManager
(console validates audience, JWTs without it were rejected)
- Admin CLI tls-list: same audience fix
- compose.yaml: TURNSTONE_JWT_SECRET now uses :? to fail fast if unset
- SDK console: fix default port from 8081 to 8090
* test: add auth enforcement tests for TLS admin endpoints
5 new tests: unauthenticated requests return 401 (list, renew,
delete), read-only-scoped requests return 403 (renew, delete).
Closes the TLS auth enforcement test gap noted in PROGRESS.md.
* fix: address remaining Copilot review feedback
- Fix token_source="config" → "test" in TLS test fixtures
- Fix AuthResult.token_source docstring to include service origins
- Require TURNSTONE_JWT_SECRET in cluster compose profile (:?)
- Helm: add auth.jwtSecret + auth.existingSecret values, wire
TURNSTONE_JWT_SECRET into secret.yaml and both deployments
- Terraform: replace auth_token with jwt_secret variable + secret,
remove orphaned auth_token resources and IAM reference
- Remove [[auth.tokens]] from security.md config example
* fix: address full code review — 10 findings
Critical:
- Terraform: replace concat(common_env, auth_env) with common_env
(auth_env local was removed but still referenced)
- Channel gateway: remove hmac static token auth from _check_auth(),
use JWT-only validation. Remove --auth-token CLI arg from channel
- Rebalancer: add token_manager support so migration requests carry
JWT auth (was sending unauthenticated POST to /internal/migrate)
Major:
- Guard _permissions_to_scopes() against "service" privilege
escalation from DB role permissions
- Remove dead AuthConfig class, load_auth_config(), and all
auth_config parameters from create_app() signatures
- Helm: inject JWT secret for both inline and existingSecret paths
Minor:
- Remove dead auth_token param from ClusterCollector
- Remove empty TestLoadAuthConfig class
- Short JWT secret now exits instead of warning
- Compose: add generation command comment above JWT_SECRET
- Clean stale config token references from 6 doc files
- Clean stale AUTH_TOKEN reference from bootstrap wizard prompt
* fix: remove remaining stale config token references from docs
- channels.md: remove --auth-token from options table
- oidc.md: remove "config-file tokens still work" claim
- security.md: remove config token section, fix JWT secret docs
(now required/exits, no ephemeral fallback), remove hmac from
ASCII diagram, remove --auth-token reference
* fix: populate model in _last_usage so usage-by-model records correctly
_last_usage was built purely from UsageInfo token counts, never
including a "model" key. server.py's on_status() fell back to
model="" for every record_usage_event call, so GROUP BY model
collapsed all rows into a single empty-key bucket.
* fix: inject model at emission time, preserve dict[str, int] typing
Address Copilot review: keep _last_usage as dict[str, int] for type
safety, inject "model" from self.model when passing to on_status().
This also fixes stale model after /model switch since the value is
read fresh each time.
* fix: MCP tools not surfacing after Sync to Nodes, update Anthropic tool search
Three fixes:
1. session_factory closure captured mcp_client=None when no --mcp-config
was passed at startup. internal_mcp_reload created a new MCPClientManager
on app.state but the factory never saw it. New workstreams got 0 MCP tools.
Fix: mutable _mcp_ref list shared between factory and reload handler.
2. Anthropic dropped the date suffix from tool_search_tool_bm25_20251119
and now requires name == type. Updated constant and tool definition.
3. Add diagnostic logging around API errors (provider, model, base_url,
message counts, full exception chain) and workstream resume (pre/post
provider state, alias resolution warnings).
Also adds Node.js 24 LTS to Dockerfile via multi-stage copy for npx-based
MCP servers.
* fix: address Copilot review — set_storage on reload, sanitize log output
- Call mcp_mgr.set_storage(storage) when internal_mcp_reload creates a
new MCPClientManager so prompt sync works for post-startup servers
- Strip query params from base_url before logging (may contain API keys
in some vLLM deployments)
- Split API error logging: concise warning (type names only) + separate
debug with exc_info=True for full traceback when needed
* chore: remove DDG MCP sidecar, web_search uses built-in ddgs client
The DuckDuckGo MCP server container is redundant — the built-in
DuckDuckGoClient (via ddgs package, included in all extras) auto-detects
when no Tavily key is configured. Removes the ddg-search service,
ddgCluster profile, and mcp-ddg.json config file.
* fix: materialize skill resources to disk for subprocess access
Skill-bundled scripts stored in skill_resources were loaded into memory
but never written to disk, causing FileNotFoundError when the model
tried to execute them. Write resources to a per-workstream temp directory
on skill load, expose via SKILL_RESOURCES_DIR env var and PATH, clean up
on skill change or session close.
* fix: pre-flight validation warns when skill references missing resources
Scan rendered skill content for path references (scripts/foo.py, etc.)
and compare against bundled skill_resources. Warn via on_info if any
referenced paths are not bundled, so operators see the gap at skill
activation rather than at runtime FileNotFoundError.
* fix: address PR #271 review feedback
- Fix trailing colon in PATH when $PATH is empty (cwd-on-PATH risk)
- Move try/except inside per-resource loop so one bad write doesn't
abort all resources
- Explicit encoding="utf-8" for deterministic writes across locales
* fix: address PR #271 review round 2
- Normalize available paths in _validate_skill_resources() to match
referenced paths (both sides use os.path.normpath now)
- Fix flaky traversal test: assert inside base dir, not escaped path
Multiple browser tabs open to the same server could see workstream
names, states, and content mixed up between workstreams when creating,
closing, and switching tabs rapidly.
Root causes and fixes:
- Global SSE ws_created events were never handled — other tabs never
learned about new workstreams, causing blank names and stale tab bars
- SSE reconnection assigned all stale panes to the first workstream
instead of deduplicating; now uses two-pass assignment with tracking
- switchTab left the old EventSource open while reassigning pane.wsId,
creating a window for events to leak; now disconnects SSE first
- Per-workstream events carried no ws_id — server now stamps ws_id on
all events via _enqueue (shallow copy); client handleEvent drops
events with mismatched ws_id as defense-in-depth
- Plan dialog used pane.wsId at resolve time (could drift after tab
switch); now captures ws_id when the dialog opens
- Global ws_closed could reassign panes before per-ws SSE finished
draining; now disconnects per-ws SSE immediately on close
The 5 prompt policy admin endpoints required "admin.prompt_policies"
but the builtin-admin role only grants "admin.policies". Changed to
match the existing permission used by tool policy endpoints.
* fix: harden Discord bot against gateway disconnects and SSE failures
- Isolate Discord API failures from SSE stream — _on_ws_event exceptions
no longer kill the SSE connection and cause missed events
- Fix broken exponential backoff on 4xx/5xx (delay was reset on every
attempt); skip aiter_sse() on error responses
- Add read timeout (90s) to SSE httpx client so half-open TCP
connections are detected and recovered
- Re-resolve node URL on each SSE reconnect attempt
- Add on_resumed handler to recover SSE tasks that died during brief
gateway disconnects (on_ready is not called on session resume)
- Sync slash commands only on first on_ready to avoid Discord rate limits
* fix: SSE backoff on 4xx/5xx and retrieve dead task exceptions
- Replace `continue` with raise+catch so 4xx/5xx errors hit the
exponential backoff path instead of tight-looping
- Retrieve task exceptions in _purge_dead_sse_tasks to suppress
"Task exception was never retrieved" warnings and log the cause
* feat: modular system message composition with admin prompt policies (#267)
Replace the monolithic persona+tools block in _init_system_messages()
with a modular composition harness (turnstone/prompts/). System messages
are now assembled from five typed layers: BASE (persona), ENV (client
surface — web/cli/chat), CONTEXT (datetime, timezone, username), TOOLS
(usage patterns), and POLICIES (behavioral rules with tool gating).
Prompt policies are admin-managed via a new Prompts tab in the Governance
group (CRUD with modal forms, tool gating, priority ordering, enable/disable).
DB policies override file-based defaults by name; file-based policies serve
as deployment defaults. Migration 031 adds the prompt_policies table.
ClientType is threaded end-to-end from channel adapters through the SDK,
HTTP API, WorkstreamManager, and session factory to ChatSession. Discord
sessions now receive chat-optimized formatting (no tables, no Mermaid,
concise output) instead of the web UI's rich markdown instructions.
* fix: address CI failures and Copilot review feedback
- Add client_type param to CLI session_factory (mypy protocol match)
- Add prompt policy CRUD to PostgreSQL backend (test-postgres CI)
- Fix ClientType resolution: compare against enum values, not members
- Fix null client_type coercion (body.get returns None, not "")
- Use local time with astimezone() instead of UTC with local tz name
- Sanitize tool_gate in update endpoint (coerce null to empty string)
* feat: replace console HTTP polling with persistent SSE streams
Console collector now subscribes to each server node's /v1/api/events/global
SSE stream for real-time state updates instead of polling /v1/api/dashboard
and /health every 15 seconds.
Server changes:
- Emit ws_created/ws_closed events on global queue from create/close handlers
- Add node_snapshot on SSE connect (workstreams, health, aggregate)
- Add ?expected_node_id= identity verification (409 on mismatch)
- Add health_changed callback to BackendHealthMonitor circuit breaker
- Add periodic aggregate emitter thread (10s)
Console collector changes:
- Single asyncio event loop on one thread multiplexes all SSE connections
(scales to 1000+ nodes vs thread-per-node)
- Discovery loop spawns/cancels async SSE tasks per node
- Snapshot reconciliation on connect, delta application for live events
- Fix ws_state→cluster_state event type mismatch
- Remove polling code (poll_interval, max_poll_workers, --poll-interval CLI)
SDK changes:
- Add NodeSnapshotEvent, HealthChangedEvent, AggregateEvent dataclasses
- Add stream_node_events() method (async + sync)
* fix: address review feedback on node event streams
- Fix stop() to let SSE manager exit naturally instead of force-stopping
the event loop (ensures finally cleanup runs)
- Guard against empty/invalid SSE data from ping frames
- Treat missing node_id as identity mismatch (409) when expected_node_id
is provided
- Fix stale docstring on _update_metrics
* feat: show thinking indicators, tool calls, and results in Discord threads
Discord threads now surface real-time activity during multi-tool chains
instead of appearing idle. ThinkingStart/Stop events display a transient
italic status message. ToolInfoEvent sends a per-tool "running" embed
that ToolResultEvent edits in-place with the result (FIFO matching by
tool name, fallback to new message). Includes backtick-injection escaping
in tool output. Visibility respects auto-approve config so tool calls
always appear somewhere.
* fix: address review feedback on Discord action visibility
Delete thinking messages in unsubscribe/stale-route cleanup (not just
pop state). Sanitize tool-call previews (escape backticks, strip
mentions). Fix format_tool_result docstring re ellipsis line count.
Add regression test for triple-backtick escaping.
* fix: disable approval buttons on server-side resolution (timeout)
Handle ApprovalResolvedEvent in _on_ws_event to disable buttons and
grey out the approval embed when the server resolves the approval
externally (timeout, auto-approve from another client). Extract
disable_message_buttons helper from views.py so it works on a plain
Message (not just an Interaction).
* fix: reply with guidance when user DMs the bot directly
Non-reply DMs were silently ignored. Now sends a message directing the
user to /ask or @mention in a server channel.
* fix: address round 2 review feedback
- Show error items (policy-denied) in ToolInfoEvent unconditionally
- Match ToolResultEvent to ToolInfoEvent by call_id (deterministic),
fall back to name-based FIFO when call_id is absent
- Escape triple backticks before truncating in format_tool_result so
the 500-char limit holds after expansion
* fix: edit thinking message in-place instead of delete-and-recreate
ThinkingStopEvent now preserves the message for the next event to reuse.
ContentEvent seeds StreamingMessage with the thinking message so the
first flush edits it. ToolInfoEvent edits the thinking message into the
first tool embed. Eliminates the visible delete → gap → new message
flicker during thinking → tool call transitions.
* feat: separate tool call and result into distinct Discord messages
ToolInfoEvent sends a "running" embed (light grey, tool name + preview).
ToolResultEvent marks it "Done"/"Error" (color + title update) and sends
the result as a separate message. This gives clear lifecycle tracing in
chat-style threads where verbosity aids readability.
* fix: show running embed for all tools and remove redundant name prefix
ToolInfoEvent now shows a running embed for every tool regardless of
needs_approval — the running indicator and approval dialog serve
different purposes. Removes the needs_approval/auto_approve filter
that caused missing running embeds when tools were approved via
"Always Approve" or server-side auto-approve.
Also drops the redundant **name** prefix from format_tool_result since
the embed title already carries the tool name.
* fix: concise logging for SSE connection failures
Catch httpx.ConnectError/ConnectTimeout separately from the generic
exception handler. Logs url and error string instead of the full
httpx/httpcore stack trace, which is noise for expected transient
connection failures during node restarts.
* fix: address round 3 review feedback
- Pop _pending_approval_msgs on button click so ApprovalResolvedEvent
doesn't double-update the embed title (e.g. "Approved - Approved")
- Remove unused name/is_error params from format_tool_result — embed
title carries the name, embed color carries the error status
- Fix _disable_buttons docstring to mention title update
- Fix missing /v1 prefix on SSE endpoint URL (caused all SSE connections
to get text/plain 404 responses instead of event streams)
- Stop treating StreamEndEvent as session-terminal (it fires per-segment,
not per-workstream) so multi-turn conversations work in Discord
- Bail on 404 instead of retrying forever for gone workstreams, and clean
up stale routes from storage
- Check response status before iterating SSE events to avoid retrying
non-retryable upstream errors
- Default rebalancer.enabled to True so hash ring routing works without
manual ConfigStore setup
- Add one-shot cache refresh fallback on route endpoints to handle the
startup race between rebalancer and first routed request
- Remove stale "message queues" language from tagline
- Remove duplicated content covered by docs (governance details,
judge config, config.toml reference, health/rate-limit details,
monitoring metrics, tool table, multi-model config)
- Replace tool table with summary + link to docs/tools.md
- Add documentation index table linking to all doc pages
- Add architecture summary (single-node vs multi-node routing)
- Add component table for entry points
- Trim diagram table to most useful subset
- Consolidate quickstart section
README is now a concise landing page that directs to docs for
details, not a duplicated reference manual.
When TURNSTONE_ADVERTISE_URL is set (Docker deployments), the
_advertise_host variable was never assigned. The TLS upgrade path
tried to use it to construct the https:// URL, causing an
UnboundLocalError that made TLS init fail silently.
Fix: derive the TLS URL from _advertise_url (replace http → https)
instead of reconstructing from _advertise_host.
Address Copilot PR feedback:
- Exit with clear error if neither console_url nor server_url is
available after discovery (prevents cryptic failures downstream)
- Fix log field names: console → console_url, server → server_url
for consistency with other channel log events
Add token_factory parameter to SDK clients (_BaseClient, server,
console) — a Callable[[], str] invoked before each request to get
the current auth token. Supports ServiceTokenManager for auto-rotating
JWTs that re-mint transparently before expiry.
Channel gateway creates dual token managers:
- console-audience JWT for routing proxy calls (via AsyncTurnstoneConsole)
- server-audience JWT for direct SSE connections to server nodes
Both _request() and _stream_sse() inject the factory header per-call,
so long-lived connections get fresh tokens on reconnect.
Also adds TURNSTONE_CONSOLE_URL to console compose service for
DNS-resolvable service discovery.
- Default --server-url is now empty (not localhost:8080) to avoid
unreachable fallback URLs inside Docker containers
- Auto-discovery retries for up to 30s waiting for console or server
to register in the services table (handles startup ordering)
- Log discovery progress (discovering, discovered_console, discovered_server)
and warn on timeout or failure
- Wrap discovery in try/except so storage init failures don't crash startup
- Default --server-url is now empty (not localhost:8080) to avoid
unreachable fallback URLs inside Docker containers
- Auto-discovery retries for up to 30s waiting for console or server
to register in the services table (handles startup ordering)
- Both console_url and server_url are discovered from DB when not
explicitly set via CLI flags or env vars
- 404 retry: use blocking lock acquire so retry waits for cache refresh
to complete instead of skipping on contention
- 404 retry: surface httpx.HTTPError as 502 instead of suppressing it
and returning the original 404
- channel router: pass auto_approve_tools to create_workstream calls
(was silently dropped for console-routed creates)
- api-reference.md: document all /v1/api/route/* console routing proxy
endpoints and console /metrics
- router.route(): validate ws_id length and hex format before bucket
extraction, raise NoAvailableNodeError instead of ValueError
- router: expose version as public property, collector uses it instead
of accessing _version directly
- memory.py: deduplicate _bucket_of with canonical bucket_of from
hash_ring module
- architecture SVG: reroute direct/SSE lines below console to avoid
crossing over the console box
The collector's _apply_poll only detected workstream additions and
removals (set diff on ws_ids). State changes within existing
workstreams (idle → running, running → attention, etc.) were not
emitted to the SSE stream, so the dashboard only updated on manual
page refresh.
Now _apply_poll compares state and name fields between old and new
poll snapshots and emits ws_state and ws_rename events for any
changes. These flow through _fanout to the browser SSE stream,
giving real-time dashboard updates without page refresh.
Bug: Server nodes registered with container ID hostnames (e.g.,
http://a236323a92f6:8080) which aren't DNS-resolvable by other
containers. The console collector failed to poll nodes, causing
stale health/error status on the dashboard.
Fix: Add TURNSTONE_ADVERTISE_URL env var support. In compose, each
server sets it to the Docker service name (http://server-1:8080 etc).
Falls back to socket.getfqdn() when not set.
Also: remove the 100-node stress cluster (ddgStressCluster profile)
from compose.yaml. It was 720 lines of boilerplate from the old
simulator era. The simulator is being rebuilt separately (task #5).
Compose goes from 1028 to 304 lines.
Bug 1: Server's create_workstream handler ignored initial_message from
the request body. The old bridge sent it as a follow-up SendMessage
via Redis, but with direct HTTP nobody was sending it. Now the server
spawns a worker thread to send the initial message after creation,
matching the bridge's behavior.
Bug 2: Channel gateway compose config used --server-url=http://server:8080
which doesn't exist in cluster/ddgCluster profiles. Removed the hardcoded
URL — the channel gateway auto-discovers the console from the services
table via shared PostgreSQL. Added TURNSTONE_DB_URL and auth token to
the channel environment so DB-based service discovery works.
Server SDK create_workstream: add initial_message, auto_approve_tools,
user_id, ws_id params (all optional, omitted when empty).
Console SDK: add auto_approve, auto_approve_tools, user_id to
create_workstream. Add 8 route_* methods for the routing proxy path
(/api/route/*): route_create_workstream, route_send, route_approve,
route_plan_feedback, route_close, route_cancel, route_command,
route_lookup. Sync mirrors for all.
Prepares for channel gateway and scheduler to use SDK clients instead
of raw httpx calls.
Move the consistent hash ring implementation (FNV-1a, virtual nodes,
bisect lookup) from code to docs/design/consistent-hash-ring.md as a
forward-looking reference for future scalability work.
The current rebalancer uses weight-proportional distribution (simpler,
exact splits, no hash variance). The ring algorithm is documented with
test vectors, stability properties, and a comparison table for when
the ring approach becomes advantageous (large clusters, decentralized
routing, cross-language determinism).
hash_ring.py retains: RING_SIZE, bucket_of(), RingNode, NoAvailableNodeError
(all actively used by router and rebalancer).
Replace the full-rehash algorithm (diff ideal vs current across all
65536 buckets) with a donor/recipient algorithm that only moves
buckets from overloaded nodes to underloaded nodes.
Key improvements:
- Adding node C to {A, B} only moves buckets TO C, never between
A and B. Previously the HashRing rehash could shuffle between
existing nodes.
- Seeding uses weight-proportional distribution instead of HashRing
virtual nodes, producing an exact split that doesn't trigger
immediate correction on the next cycle.
- Dead-node buckets are redistributed to the most underloaded
survivors, not rehashed across the whole ring.
- HashRing class is no longer used by the rebalancer (still
available for other uses like the Go rewrite reference).
The threshold check still gates live-to-live moves. Dead-node
recovery remains unconditional.
set_bucket_stat: single-upsert storage method replacing the N-loop
reconciliation in the rebalancer. Reduces DB round-trips from
|ws_delta| per bucket to exactly 1.
Console metrics: /metrics endpoint on the console exposing 6 routing
and ring metrics in Prometheus text format:
- turnstone_router_requests_total (method, status)
- turnstone_router_request_duration_seconds (method)
- turnstone_ring_membership_size
- turnstone_ring_version
- turnstone_ring_rebalance_total (result)
- turnstone_ring_migrations_total
Instrumented in route_create, route_proxy, route_lookup handlers.
Ring gauges updated on collector discovery loop. Rebalance/migration
counters recorded after each rebalancer pass.
When rebalancer.eager_migrate is enabled, the rebalancer POSTs
/_internal/migrate to source nodes after reassigning buckets,
triggering immediate workstream eviction instead of waiting for
lazy resume on the next request.
Only idle workstreams are eagerly migrated — active ones (running,
thinking, attention) are left alone to avoid disrupting in-flight
work. Failed migrations are logged and skipped (the lazy path
handles them eventually).
Rebalancer: daemon thread in the console process that maintains
bucket-to-node assignments in hash_ring_buckets. Seeds the ring on
first run (empty table → 65536 rows via consistent hash). Periodically
checks for membership changes and rebalances: moves cheapest buckets
first (empty > idle > active), respects imbalance threshold, reconciles
bucket_stats against actual workstream counts before each pass.
Uses DB-based leader election (rebalancer_lock in system_settings) for
multi-console deployments. Increments rebalancer_version after writes
so console routers refresh their caches.
Add 6 settings: ring.vnodes_per_unit, rebalancer.enabled/interval/
threshold/eager_migrate, node.weight.
Add /_internal/migrate endpoint on server for eager workstream eviction.
Add routing proxy endpoints to the console server:
- POST /v1/api/route/workstreams/new — hash-ring-routed create with
503 retry, target_node pinning, and node_url injection
- POST /v1/api/route/{send,approve,cancel,command,close} — generic
proxy to workstream owner via O(1) bucket lookup
- GET /v1/api/route?ws_id=X — node URL lookup for direct SSE
Wire ConsoleRouter into console lifespan (cache refresh on startup)
and collector discovery loop (version-based cache invalidation).
Add --console-url to channel gateway CLI for multi-node routing.
ChannelRouter routes control-plane through console when set, SSE
connections go direct to server nodes via node_url from create response.
HashRing: FNV-1a virtual nodes, immutable, computes ideal bucket-to-node
distribution. Used by the rebalancer (next commit) to seed and maintain
the assignment table.
ConsoleRouter: in-memory flat array of 65536 NodeRef entries loaded from
hash_ring_buckets table. O(1) routing via ws_id prefix. Supports
per-workstream overrides, version-based cache refresh, and targeted
ws_id generation.
Both are pure library code with no server integration yet.
Delete the entire turnstone/mq/ package (broker, bridge, protocol,
client) and turnstone/sim/ package. Remove Redis as a dependency.
Channel gateway and console now communicate with server nodes via
direct HTTP (httpx + httpx-sse) instead of Redis pub/sub and queues.
Single-node deployments work with zero infrastructure beyond the
database.
Key changes:
- Channel adapters use httpx POST for create/send/approve/close
and httpx-sse for per-workstream event streaming
- Console collector discovers nodes via services table instead of
Redis SCAN
- Console scheduler dispatches tasks via HTTP POST with DB-based
leader election
- Server registers in services table with 30s heartbeat
- Server accepts optional ws_id in create request (for Phase 2
console-generated routing)
- SDK events gain IntentVerdictEvent and OutputWarningEvent types
- All docs, examples, bootstrap wizard updated
63 files changed, -5968 net lines (Redis transport fully removed)
* fix: sync actual TLS state to ConfigStore on console startup
The console writes tls.enabled to the DB but never clears it when TLS
init fails or isn't configured. Server nodes read the stale DB value
and attempt TLS negotiation with a non-TLS console, producing noisy
SSL errors on every startup.
Console now syncs the actual TLS state after init: if TLS succeeded,
tls.enabled=true; if it failed or wasn't attempted, tls.enabled=false.
Server TLS failure log reduced from full traceback to one-line warning.
* fix: sync TLS state to ConfigStore on console startup
Console now writes the definitive TLS state to ConfigStore so server
nodes don't attempt TLS against a non-TLS console:
- TLS init succeeded → write true
- TLS not configured (DB false/unset) → write false (definitive)
- TLS configured (DB true) but init failed → don't overwrite
(transient failure shouldn't permanently disable)
Server TLS warning reduced to one line with exception type, full
traceback available at debug level.
* feat: auto-detect model changes when LLM backend swaps models
The BackendHealthMonitor already probes /v1/models every 30s but
discarded the response. Now compares the detected model against the
last known one and triggers a registry reload when it changes.
- Extract _extract_context_window() helper for reuse across
detect_model, probe_model_endpoint, and the health monitor
- BackendHealthMonitor: new provider/initial_model/on_model_changed
params; _check_model_change() fires callback on model swap
- Server: wire _handle_model_change callback that updates cli_model_args
and calls registry.reload(); guarded by _user_specified_model flag
so --model overrides are never auto-replaced
- Session: _refresh_model_from_registry() called at top of send();
two string compares when nothing changed, full re-resolve on swap
- 7 new tests for _extract_context_window and model change detection
* fix: address Copilot review on model re-detection
- server: update cli_model_args only after successful reload (not
before), add finally block for new_reg.shutdown(), guard against
cli_model_args not yet initialized
- session: wrap registry lookup in try/except for concurrent reload
race, reset judge on model change, recompute tool_truncation when
context_window changes in auto mode
Replace "You are an expert software engineer" with a grounded
narrative persona: a resident engineer on a focused team with real
tools, real code, and real consequences. Sets expectations about
boundaries, judgment calls, and working within constraints.
* fix: scope-filter memory list/search to current workstream and user
Unscoped memory(action='list') and memory(action='search') returned all
memories across all workstreams. Now applies the same 3-query pattern
(global + current workstream + current user) used by system prompt
injection.
* fix: validate user scope on memory search/list for unauthenticated sessions
Adds _validate_scope guard to search and list prepare paths, matching
save/get/delete. Prevents explicit scope='user' from returning all
user-scoped memories when session is unauthenticated.
* fix: update _get_visible_memories references to _list_visible_memories
* fix: defense-in-depth guard for empty scope_id on search/list
Copilot review: if scope is 'user' or 'workstream' with empty
scope_id, the storage query returns all memories in that scope
across all users/workstreams. The prepare step already validates
via _validate_scope, but add exec-level guard to reject scoped
queries with empty scope_id as defense-in-depth.
* fix: detect context window from vLLM max_model_len field
vLLM exposes the context window as max_model_len on the model object,
not meta.n_ctx_train (llama.cpp format). Both detect_model() and
probe_model_endpoint() now check max_model_len first, falling back
to meta.n_ctx_train for llama.cpp. Fixes 32768 fallback on vLLM
servers that report 262144+ token context windows.
* test: add vLLM max_model_len detection tests
Copilot review: new vLLM context window path had no test coverage.
Add tests for probe_model_endpoint (max_model_len detected, preferred
over meta.n_ctx_train) and detect_model (vLLM model object with
max_model_len).
Experimental multi-node AI orchestration platform. Deploy tool-using AI agents across a cluster of servers, driven by message queues or interactive interfaces.
Multi-node AI orchestration platform. Deploy tool-using AI agents across a cluster of servers with direct HTTP routing, interactive interfaces, and enterprise governance.
> **Beta — Use at your own risk.** Turnstone is under active development and has not reached a stable release. APIs, configuration formats, and database schemas may change between versions without migration paths. We make no guarantees of determinism, reliability, or backward compatibility. Evaluate thoroughly before deploying to any environment where these properties matter.
<p align="center">
<img src="docs/assets/hero.png" alt="Turnstone console — multi-workstream AI orchestration with mermaid diagrams" width="960"/>
</p>
Named after the [Ruddy Turnstone](https://en.wikipedia.org/wiki/Ruddy_turnstone) (*Arenaria interpres*) — a shorebird that flips stones to discover what's hiding underneath.
| **Experimental** | `pip install turnstone --pre` | `ghcr.io/turnstonelabs/turnstone:experimental` | New features. May have rough edges. |
See [docs/releasing.md](docs/releasing.md) for the full release process.
## What it does
Turnstone gives LLMs tools — shell, files, search, web, planning — and orchestrates multi-turn conversations where the model investigates, acts, and reports. It runs as:
Turnstone gives LLMs tools — shell, files, search, web, planning — and orchestrates multi-turn conversations where the model investigates, acts, and reports.
- **Interactive sessions** — terminal CLI or browser UI with parallel workstreams
- **Queue-driven agents** — trigger workstreams via message queue, stream progress, approve or auto-approve tool use
- **Multi-node clusters** — generic work load-balances across nodes, directed work routes to a specific server
- **Cluster dashboard** — real-time view of all nodes and workstreams, reverse proxy for server UIs
- **Intent validation** — an LLM judge evaluates every tool call before approval, presenting risk assessments and evidence-based recommendations so users can make informed decisions instead of blindly approving raw tool calls
- **Cluster simulator** — test the stack at scale (up to 1000 nodes) without an LLM backend
Works with any OpenAI-compatible API (vLLM, llama.cpp, NVIDIA NIM) or Anthropic's native Messages API. Supports [MCP](https://modelcontextprotocol.io/) for external tool servers with native deferred tool loading on Anthropic and OpenAI APIs (BM25 fallback for local models).
- **Cluster dashboard** — real-time view of all nodes and workstreams with console routing proxy
- **Intent validation** — LLM judge evaluates every tool call with risk assessments and evidence
- **Multi-provider** — OpenAI-compatible APIs (vLLM, llama.cpp, NIM), Anthropic Messages API, and Google Gemini
- **MCP support** — external tool servers with native deferred loading (Anthropic/OpenAI) or BM25 fallback
<p align="center">
<img src="docs/diagrams/architecture-overview.svg" alt="Turnstone system architecture — data flow from clients through gateways, Redis MQ, cluster nodes, to LLM providers" width="960"/>
<img src="docs/diagrams/architecture-overview.svg" alt="Turnstone system architecture" width="960"/>
Then open `http://localhost:8090` for the cluster-wide dashboard. Create workstreams from the console and interact with any node's server UI through the built-in reverse proxy — no direct server port access required.
### Docker
```bash
cp .env.example .env # edit LLM_BASE_URL, OPENAI_API_KEY, etc.
docker compose up # starts redis + server + bridge + console (SQLite)
docker compose --profile production up
```
For production with PostgreSQL:
See [QUICKSTART.md](QUICKSTART.md) for the bootstrap wizard and [docs/docker.md](docs/docker.md) for Docker configuration and profiles.
```bash
# Requires POSTGRES_PASSWORD and DB_BACKEND=postgresql in .env (or exported)
docker compose --profile production up # adds PostgreSQL, uses it as database
Turnstone includes a built-in governance layer for enterprise deployments — manage who can do what, which tools run unattended, and where every token goes.
- **OIDC SSO** — single sign-on via any OpenID Connect provider (Okta, Azure AD, Google, Keycloak); Authorization Code Flow with PKCE, auto-provisioning, claim-based role mapping with demotion propagation; see [docs/oidc.md](docs/oidc.md)
- **Tool policies** — glob-pattern rules (`allow` / `deny` / `ask`) with priority ordering; automate approvals or lock down dangerous tools
- **Skills** — reusable behavioral profiles with system prompts, `{{variable}}` substitution, session config (model, temperature, token budget), install-time security scanning, version history, external discovery (skills.sh / GitHub), and runtime `skill` tool for model-driven skill activation
- **Usage tracking** — per-request token and tool metrics, aggregation by day / model / user, automatic 90-day pruning
- **Audit logging** — append-only event trail for all admin mutations, IP-aware, 365-day retention
All governance features are managed through the console admin panel (13 tabs) and the full REST API. Runtime settings (model, tools, rate limiting, health, judge, memory) are configurable via the admin Settings tab — no config file edits or restarts needed for most changes. See [docs/governance.md](docs/governance.md) for setup and [docs/settings.md](docs/settings.md) for the settings reference.
### Intent Validation (LLM Judge)
Every tool call that requires human approval is evaluated by an intent validation judge that provides a structured risk assessment alongside the approval prompt — so instead of "approve this bash command?", users see a verdict with risk level, confidence, recommendation, and reasoning.
The system uses a two-tier evaluation pipeline:
1.**Heuristic tier** (instant, free) — 36 pattern-based rules classify tool calls by severity. Catches destructive commands (`rm -rf /`, `DROP TABLE`), privilege escalation (`sudo`), credential access, supply chain risks, browser data export, cloud infrastructure mutations, and more. Results appear immediately.
2.**LLM judge tier** (async) — A full LLM evaluation runs in the background with access to `read_file` and `list_directory` for evidence gathering. The judge can inspect files that a write would overwrite, check directory contents before a delete, and cite specific evidence in its reasoning. Results update the UI progressively when ready.
The judge defaults to the same model as the session (self-consistency) but can be configured to use a separate model — useful when running a small local model for tasks but wanting a commercial model for safety evaluation.
```toml
[judge]
enabled=true# on by default
model=""# empty = same as session model
provider=""# empty = same as session provider
timeout=60.0# generous for local models
```
Verdicts are persisted for audit and exposed via Prometheus metrics (`turnstone_judge_verdicts_total`, `turnstone_judge_llm_latency_seconds`).
Skills are also scanned at install time — the scanner evaluates content, supply chain, vulnerability, and declared capability risk across four independent axes. Results populate `scan_status` (tier) and `scan_report` (structured JSON breakdown) on the skill record so administrators can assess risk before enabling a skill.
Tool execution results are evaluated by an output guard before entering the conversation — detecting prompt injection payloads in fetched content, credential leakage in command output, and encoded payloads. Detected credentials are automatically redacted.
See [docs/judge.md](docs/judge.md) for the full guide.
## Multi-node routing
Each Turnstone server runs a bridge process. Bridges share a Redis instance for coordination:
| Redis Key | Purpose |
|-----------|---------|
| `turnstone:inbound` | Shared work queue — generic tasks, any node |
| `turnstone:events:global` | Global event pub/sub |
| `turnstone:events:cluster` | Cluster-wide state changes (for turnstone-console) |
**Routing rules:**
1. Message has `target_node` → routes to that node's queue
2. Message has `ws_id` → looks up owner, routes to owning node
3. Neither → shared queue, next available bridge picks it up
Bridges BLPOP from their per-node queue (priority) then the shared queue. Directed work always takes precedence.
## Tools
15 built-in tools, 2 agent tools, plus external tools via MCP:
Built-in tools for shell, files, search, web, memory, notifications, and autonomous sub-agents — plus external tools via [MCP](https://modelcontextprotocol.io/) with native deferred loading. See [docs/tools.md](docs/tools.md) for the full reference and [docs/mcp-registry.md](docs/mcp-registry.md) for MCP configuration.
| Tool | Description | Auto-approved |
|------|-------------|:---:|
| `bash` | Execute shell commands | |
| `read_file` | Read file contents (text or images with vision models) | yes |
| `write_file` | Write/create files | |
| `edit_file` | Fuzzy-match file editing | |
| `search` | Search files by name/content | yes |
| `math` | Sandboxed Python evaluation | |
| `man` | Read man pages | yes |
| `web_fetch` | Fetch URL content | |
| `web_search` | Web search (provider-native or Tavily) | |
| `watch` | Periodic command polling with conditions | |
| `task` | Spawn autonomous sub-agent | |
| `plan` | Explore codebase, write .plan.md | |
| `mcp__*` | External tools from MCP servers | |
## Architecture
When the total tool count exceeds a configurable threshold (default 20), MCP tools are automatically deferred using native `defer_loading` on Anthropic and OpenAI APIs, or a transparent client-side BM25 search for local models. The LLM discovers deferred tools on demand via a `tool_search` capability — no configuration needed beyond `--tool-search auto` (the default).
**Single-node**: Client → Server (direct HTTP + SSE). No external dependencies beyond the database.
### MCP Tool Servers
**Multi-node**: Client → Console (rendezvous routing proxy) → Server nodes. The console picks the target node for each workstream via rendezvous (HRW) hashing over the live service registry — pure function of `(ws_id, live_nodes)`, no stored bucket state, deterministic across readers. A node join or drop only re-routes the keys that score highest on the affected node.
Turnstone supports the [Model Context Protocol](https://modelcontextprotocol.io/) (MCP) for connecting external tool servers. MCP tools are discovered at startup, converted to OpenAI function-calling format, and merged with built-in tools. Each MCP tool is prefixed with `mcp__{server}__{tool}` to avoid name collisions. Tool lists stay fresh via push notifications (`tools.listChanged`), periodic polling for servers without push, and manual `/mcp refresh`.
| Component | Purpose |
|-----------|---------|
| `turnstone` | Terminal CLI (REPL) |
| `turnstone-server` | Web UI + REST API + SSE events |
Use `/mcp` in the REPL to list connected tools, `/mcp refresh` to re-fetch tool lists from servers. MCP tools require user approval by default (overridden by `--skip-permissions` or UI auto-approve).
### Multi-Model and Multi-Provider Support
Turnstone supports multiple model backends per server instance, including different LLM providers. `ChatSession` delegates all API communication to pluggable `LLMProvider` adapters — the internal message format stays OpenAI-like, and each provider translates at the API boundary. Define named models in `config.toml` and select per-workstream or switch mid-session with `/model <alias>`.
```toml
[models.local]
base_url="http://localhost:8000/v1"
model="qwen3-32b"
# provider defaults to "openai" (works with vLLM, llama.cpp, etc.)
[models.claude]
provider="anthropic"
api_key="sk-ant-..."
model="claude-opus-4-6"
context_window=200000
[models.openai]
base_url="https://api.openai.com/v1"
api_key="sk-..."
model="gpt-5"
context_window=400000
[model]
default="local"# which model to use by default
fallback=["claude","openai"]# try these if the primary is unreachable
agent_model="claude"# optional: separate model for plan/task sub-agents
```
Supported providers: `"openai"` (default -- OpenAI, vLLM, llama.cpp, any OpenAI-compatible API) and `"anthropic"` (Anthropic Messages API, requires `pip install turnstone[anthropic]`).
Use `/model` to show available models, `/model claude` to switch. Workstreams created via the API accept an optional `model` parameter.
## Configuration
All entry points read `~/.config/turnstone/config.toml`. CLI flags override config values.
```toml
[api]
base_url="http://localhost:8000/v1"
api_key=""
# tavily_key = "" # only needed for local/vLLM models without native search
[model]
name=""# empty = auto-detect
temperature=0.5
reasoning_effort="medium"
default="default"# model alias for new workstreams
fallback=[]# ordered list of fallback model aliases
agent_model=""# model alias for plan/task sub-agents
[tools]
timeout=30
skip_permissions=false
search="auto"# "auto" (enable when >threshold tools), "on", "off"
search_threshold=20# min tools before tool search activates
search_max_results=5# max tools returned per search query
[server]
host="0.0.0.0"
port=8080
max_workstreams=50# auto-evicts oldest idle when full
[redis]
host="localhost"
port=6379
password=""
[bridge]
server_url="http://localhost:8080"
node_id=""# empty = hostname_xxxx
[console]
host="0.0.0.0"
port=8090
url="http://localhost:8090"# used by CLI /cluster commands
poll_interval=10
[health]
backend_probe_interval=30
backend_probe_timeout=5
circuit_breaker_threshold=5
circuit_breaker_cooldown=60
[ratelimit]
enabled=true
requests_per_second=10.0
burst=20
[database]
backend="sqlite"# "sqlite" (default) or "postgresql"
path=".turnstone.db"# SQLite file path (relative to working directory)
Parallel independent conversations, each with its own session and state:
| Symbol | State | Meaning |
|--------|-------|---------|
| `·` | idle | Waiting for input |
| `◌` | thinking | Model is generating |
| `▸` | running | Tool execution in progress |
| `◆` | attention | Waiting for approval |
| `✖` | error | Something went wrong |
Idle workstreams are automatically cleaned up after 2 hours (configurable). In multi-node deployments, workstream ownership is tracked in Redis — follow-up messages auto-route to the owning node.
-`turnstone_judge_enabled` — whether the intent validation judge is active (0/1)
Per-workstream metrics are labeled by `ws_id` (bounded by `[server].max_workstreams`).
### Health & Rate Limiting
**Health degradation.** A background `BackendHealthMonitor` probes the LLM backend every `backend_probe_interval` seconds. When the backend is unreachable, `/health` reports `"status": "degraded"` (HTTP 200) and the `turnstone_backend_up` gauge drops to 0.
**Circuit breaker.** After `circuit_breaker_threshold` consecutive probe failures the circuit opens (CLOSED -> OPEN). While open, `ChatSession._create_stream_with_retry` skips the backend entirely and returns an error. After `circuit_breaker_cooldown` seconds the circuit enters HALF_OPEN, allowing a single probe. A successful probe closes the circuit; a failure re-opens it.
**Per-IP rate limiting.** When `[ratelimit].enabled` is true, each client IP is tracked with a token-bucket limiter (`requests_per_second` / `burst`). Rate limiting is applied in `do_GET`/`do_POST` after authentication but before route dispatch. `/health` and `/metrics` are exempt. Requests that exceed the limit receive HTTP 429 with a `Retry-After` header.
**Workstream eviction.** When `WorkstreamManager.create()` would exceed `max_workstreams`, the oldest IDLE workstream is automatically evicted and the `turnstone_workstreams_evicted_total` counter is incremented. Configure via `[server].max_workstreams` (default 50).
- An OpenAI-compatible API endpoint ([vLLM](https://github.com/vllm-project/vllm), [NVIDIA NIM](https://build.nvidia.com/), [llama.cpp](https://github.com/ggml-org/llama.cpp), etc.) or an Anthropic API key
JWTs are the recommended credential for browser sessions. API tokens are suitable for programmatic access and CI/CD. Config tokens are a simple option for single-node deployments.
JWTs are the recommended credential for browser sessions. API tokens are suitable for programmatic access and CI/CD.
### `POST /v1/api/auth/login`
@@ -523,7 +520,7 @@ inactivity.
Each SSE connection to a workstream receives its own delivery queue. Events
produced by the worker thread are fanned out to all registered listener queues,
so multiple consumers (browser, bridge, console proxy, SDK) can connect
so multiple consumers (browser, console proxy, SDK) can connect
simultaneously and each receives every event. On reconnect the client receives
a full history replay, so no catch-up mechanism is needed.
@@ -845,6 +842,15 @@ button automatically.
Creates a new workstream. The server supports up to 10 concurrent workstreams.
The endpoint accepts **either**`application/json` (legacy shape) **or**
`multipart/form-data` when you want to upload attachments at creation
time. Multipart requests carry one `meta` field containing the JSON body
shown below plus zero-or-more `file` parts; each file is validated and
reserved onto the new workstream's first turn before the dispatch worker
runs, so queued multimodal turns cannot lose files to racing sends. If
validation fails the fresh workstream is rolled back so no orphan rows
leak.
**Request body:**
```json
@@ -860,6 +866,7 @@ All fields are optional. The body can be empty or an empty JSON object.
| `auto_approve` | bool | false | Auto-approve all tool calls for this workstream |
| `resume_ws` | string | "" | Workstream ID to resume atomically during creation (empty = fresh)|
| `skill` | string | "" | Skill name. Applies content (system prompt), model, temperature, reasoning effort, max tokens, auto-approve policy, token budget, and other session config from the skill. Returns 400 if not found or disabled. Ignored when `resume_ws` is set (resumed sessions restore their own skill). |
| `judge_model` | string | "" | Optional model alias for the judge (overrides default judge model for this workstream) |
> **Skill behavior:** When `skill` is specified, the skill's content is injected as a system message and its session config fields (model, temperature, auto-approve, token budget, etc.) override system defaults for the new workstream.
*.json 15 tool schemas (OpenAI function-calling format + turnstone metadata)
*.json 19 tool schemas (OpenAI function-calling format + turnstone metadata)
```
Both UIs share a common design system extracted into `turnstone/shared_static/`: design tokens, login overlay, toast notifications, theme toggle, keyboard shortcuts, and utility functions. Each UI imports `base.css` and the shared JS modules at `/shared/`, then adds only page-specific code at `/static/`.
@@ -448,13 +448,15 @@ from each schema and builds:
-`PRIMARY_KEY_MAP` -- `{name: primary_key}` for JSON fallback recovery
`turnstone-console` is a cluster management service that provides cluster-wide visibility and control across all turnstone nodes. It connects to the shared Redis broker, discovers nodes via heartbeat keys, polls each node's HTTP API for workstream data, and subscribes to a cluster event channel for real-time state changes.
`turnstone-console` is a cluster management service that provides cluster-wide visibility and control across all turnstone nodes. It discovers nodes via the`services` database table and subscribes to each node's SSE event stream for real-time workstream, health, and metric updates.
The console also supports **workstream creation** (dispatched via MQ to target nodes) and a **reverse proxy** that serves each node's server UI through the console port — so users only need network access to the console, not to individual server nodes.
The console also supports **workstream creation** (dispatched via HTTP proxy to target nodes) and a **reverse proxy** that serves each node's server UI through the console port — so users only need network access to the console, not to individual server nodes.
## Architecture
> See also: [Console Data Flow diagram](diagrams/png/11-console-data-flow.png)
- **Inbound (monitoring):**Bridges publish state changes to `{prefix}:events:cluster` on Redis pub/sub. The console subscribes for real-time updates and periodically polls each node's `GET /v1/api/dashboard` for full workstream snapshots.
- **Outbound (control):** The console pushes `CreateWorkstreamMessage` to Redis inbound queues targeting specific nodes. Bridges pick up these messages and create workstreams on their local servers.
- **Inbound (monitoring):**The console discovers nodes via the `services` database table (nodes register on startup and send periodic heartbeats). It opens a persistent SSE connection to each node's `GET /v1/api/events/global` endpoint, receiving a full snapshot on connect followed by real-time delta events (state changes, health transitions, aggregate metrics).
- **Outbound (control):** The console proxies workstream creation requests to target nodes via HTTP.
- **Proxy (pass-through):** The console reverse-proxies each node's server UI at `/node/{node_id}/`, forwarding HTTP and SSE traffic so the browser never contacts server nodes directly.
| Node HTTP API | `GET/POST {server_url}/*` | Proxy | Server UI, API requests, SSE streams |
### Redis Key: Cluster Event Channel
Bridges publish to `{prefix}:events:cluster` whenever a workstream state change, creation, closure, or rename occurs. Events include `node_id` so the console can attribute them to the correct node.
| `ws_created` | ws_id, name, node_id | New workstream created |
| `ws_closed` | ws_id | Workstream closed |
| `ws_rename` | ws_id, name | Workstream renamed |
---
## ClusterCollector
The collector (`turnstone/console/collector.py`) maintains an in-memory snapshot of all nodes and workstreams. Three daemon threads handle data acquisition:
The collector (`turnstone/console/collector.py`) maintains an in-memory snapshot of all nodes and workstreams. Two daemon threads handle data acquisition:
1. **Event subscriber** — subscribes to`{prefix}:events:cluster` via `RedisBroker.subscribe_cluster()`. Applies state changes, creates, closes, and renames to the in-memory model immediately.
1. **Node discovery** — queries the`services` database table every 60 seconds. Adds newly discovered nodes, removes expired ones (stale heartbeats), emits `node_joined` / `node_lost` events to SSE listeners, and spawns/cancels SSE tasks for new/lost nodes.
2. **Node discovery** — scans heartbeat keys every 15 seconds via `broker.list_nodes()`. Adds newly discovered nodes, removes expired ones, emits `node_joined` / `node_lost` events to SSE listeners.
3. **Poll loop** — fetches `GET /v1/api/dashboard` and `GET /health` from each known node every 10 seconds. Uses `ThreadPoolExecutor(max_workers=50)` for parallelism. Each poll replaces the node's workstream list with the authoritative server data.
2. **SSE manager** — a single asyncio event loop on one thread multiplexes persistent SSE connections to all discovered nodes via `GET /v1/api/events/global`. Each connection receives a `node_snapshot` on connect (workstreams, health, aggregate) followed by real-time delta events (`ws_state`, `ws_created`, `ws_closed`, `ws_rename`, `health_changed`, `aggregate`). On disconnect, the node is marked unreachable and the connection is retried with exponential backoff (1s–30s). An `?expected_node_id=` query parameter provides identity verification against IP reuse (server returns 409 on mismatch).
A `get_snapshot()` method builds the full cluster state under a single lock acquisition — overview aggregates and per-node workstream lists in one atomic read. This is served both as a REST endpoint and as the initial SSE event on client connect.
@@ -70,7 +53,7 @@ All reads and writes to the node/workstream map are protected by a single `threa
### Scale Considerations
- **50,000 workstreams** (1,000 nodes × 50 per node) at ~500 bytes each = ~25 MB in memory
- **1,000 nodes**polled in parallel — fan-out concurrency is configurable via `cluster.node_fan_out_limit` (default 200), yielding 5 batches at ~100ms each = ~0.5 second poll cycle
- **1,000 nodes**connected via persistent SSE — a single asyncio event loop multiplexes all connections with negligible overhead. Ensure `ulimit -n` >= 4096 for fd headroom
- **Filtering and pagination** run in-memory on the full workstream list — sub-millisecond at this scale
- **SSE fan-out** uses per-client queues (2,000 events) — backed-up clients get events dropped, not blocking
- **Database** — for clusters sharing PostgreSQL, use [PgBouncer](pgbouncer.md) in transaction pooling mode
@@ -183,7 +166,7 @@ Full cluster state in a single response — all nodes with their workstreams plu
### `POST /v1/api/cluster/workstreams/new`
Create a new workstream on a target node. Dispatches a `CreateWorkstreamMessage` through the Redis MQ pipeline — the bridge on the target node picks it up and creates the workstream on the server. Requires `write` scope.
Create a new workstream on a target node. The console proxies the creation request to the target node's HTTP API. Requires `write` scope.
Request:
@@ -197,9 +180,9 @@ Request:
All fields are optional:
- `node_id` — targeting mode:
- **omitted or `"auto"`** — console picks the reachable node with the most available capacity (max_ws - ws_total) and pushes to its directed queue.
- **`"pool"`** — pushes to the shared inbound queue; the next available bridge picks it up (true general-pool dispatch).
- **specific node ID** — pushes to that node's directed queue.
- **omitted or `"auto"`** — console picks the reachable node with the most available capacity (max_ws - ws_total) and proxies the request to it.
- **`"pool"`** — console picks a reachable node with available capacity using round-robin selection.
- **specific node ID** — proxies the request to that node directly.
- `name` — workstream display name. Auto-generated if omitted.
- `model` — model alias from the target node's registry. Uses the node's default model if omitted.
@@ -213,7 +196,7 @@ Response:
}
```
Creation is asynchronous — the response confirms the MQ message was dispatched. A `ws_created` event on the cluster SSE stream confirms the workstream was actually created.
The response confirms the workstream creation request was proxied to the target node. A `ws_created` event on the cluster SSE stream confirms the workstream was actually created.
Triggered by the "+ new" header button. A modal dialog with:
- **Node selector** — dropdown with three targeting modes: "Auto (best available)" picks the node with the most headroom, "General pool (any node)" pushes to the shared queue for any bridge to pick up, or a specific node from the list (showing capacity).
- **Node selector** — dropdown with three targeting modes: "Auto (best available)" picks the node with the most headroom, "General pool (any node)" picks a node with available capacity using round-robin, or a specific node from the list (showing capacity).
- **Profile** — optional dropdown listing enabled skills. Applies the skill's model, auto-approve policy, token budget, and other behavioral settings at creation time.
- **Name** — optional text input. Auto-generated if left empty.
- **Model** — optional text input for a model alias from the target node's registry.
- **Judge Model** — optional text input for the judge model alias (overrides the default judge model for this workstream).
Keyboard shortcuts: Ctrl+Shift+R (refresh title), Ctrl+Shift+E (edit title), Ctrl+Shift+F (fork), Ctrl+Shift+X (delete). Press ? for full shortcut help.
On submit, `POST /v1/api/cluster/workstreams/new` dispatches the creation request. A toast confirms success; the SSE stream delivers the `ws_created` event to update the dashboard.
@@ -410,10 +396,19 @@ The browser maintains a local `clusterState` object that mirrors the cluster sna
Accessed via the "admin" button in the header (visible when authenticated
with `approve` scope). Provides user, API token, channel link, MCP server,
and skill management with 13 tabs (see also
[Governance](governance.md) for
the Roles, Policies, Skills, Usage, and Audit tabs, and
[Settings](settings.md) for the database-backed configuration editor):
and skill management with 18 tabs (Users, API Tokens, Channels, Schedules,
Audit, Memories, Models, Nodes, Settings, TLS). See also
[Governance](governance.md) for the Roles, Policies, Skills, Usage, and
Audit tabs, and [Settings](settings.md) for the database-backed
configuration editor.
The **Channels** tab links users to either a Discord or Slack account
via a per-row channel-type selector. The **Models** tab is a CRUD
editor for `model_definitions`, the **Nodes** tab edits per-node
metadata, and the **TLS** tab manages CA and leaf certificates for the
internal mTLS fabric. The **Settings** tab edits ConfigStore values
live; edits apply without restart.
**Users tab:**
@@ -482,17 +477,17 @@ to create the initial admin user and receive a JWT in one step. See
## Scheduled Tasks
The console includes a background **TaskScheduler** daemon that creates workstreams on a timed basis via the MQ broker. It supports cron-based recurring schedules and one-shot `at` schedules.
The console includes a background **TaskScheduler** daemon that creates workstreams on a timed basis via HTTP proxy to target nodes. It supports cron-based recurring schedules and one-shot `at` schedules.
### Architecture
The scheduler runs as a daemon thread inside the console process. Every `check_interval` seconds (default 15) it:
1. Acquires a distributed lock via Redis `SET NX EX` (prevents duplicate dispatch in multi-console deployments)
1. Acquires a distributed lock via the `system_settings` table (prevents duplicate dispatch in multi-console deployments)
2. Queries the storage backend for tasks whose `next_run <= now` and `enabled = true`
3. Dispatches each due task as one or more `CreateWorkstreamMessage` via MQ
3. Dispatches each due task as one or more workstream creation requests via HTTP proxy
4. Updates `last_run` and computes the next `next_run` (or disables one-shot `at` tasks)
5. Releases the lock via Lua script (safe conditional delete)
5. Releases the lock
Run history is automatically pruned (runs older than 90 days) approximately once per hour.
@@ -508,7 +503,7 @@ Run history is automatically pruned (runs older than 90 days) approximately once
| Mode | Behavior |
|------|----------|
| `auto` | Picks the reachable node with the most available capacity |
| `pool` | Pushes to the shared inbound queue (any bridge picks it up) |
| `pool` | Picks a reachable node with available capacity using round-robin |
| `all` | Fan-out to all reachable nodes (capped at `max_fan_out`, default 20) |
| `<node_id>` | Targets a specific node by ID |
@@ -645,12 +640,6 @@ CLI flags for `turnstone-console`:
Open `http://localhost:8090` for the cluster dashboard. Create workstreams via the "+ new" button. Click any workstream to open the proxied server UI — no direct access to server ports required.
| `approve_request` | One or more tool calls need operator approval | `items: [{call_id, header, preview, func_name, approval_label, needs_approval}]` |
| `rename` | Session's display name changed | `name` |
| `intent_verdict` | Intent judge produced a verdict on a pending tool call | `risk_level`, `recommendation`, `reasons` |
| `output_warning` | Output guard flagged a tool result | `call_id`, `risk_level`, `flags` |
| `child_ws_created` | A direct child of this coord was just created (fan-out from the cluster bus) | `child_ws_id`, `node_id`, `name`, `parent_ws_id` (`ws_id` in the envelope is always the coord's own id) |
| `child_ws_state` | A direct child transitioned state | `child_ws_id`, `state` |
| `child_ws_closed` | A direct child closed | `child_ws_id` |
| `child_ws_rename` | A direct child's name changed | `child_ws_id`, `name` |
component [**turnstone:node:{node_id}**\n\nNode heartbeat + metadata.\nValue: JSON {server_url, started, ...}\nTTL: 60s (refreshed every 30s)\n\nOps: SET with EX, GET, SCAN] as node_hb <<STRING>>
}
package "Event Channels (Redis PUBSUB)" #FFF3E0 {
component [**turnstone:events:global**\n\nGlobal event broadcast.\nAll state changes, ws lifecycle.\n\nOps: PUBLISH, SUBSCRIBE] as evt_global <<PUBSUB>>
component [**turnstone:events:{ws_id}**\n\nPer-workstream events.\nContent, tools, status.\n\nOps: PUBLISH, SUBSCRIBE] as evt_ws <<PUBSUB>>
component [**turnstone:events:cluster**\n\nCluster-wide state changes.\nUsed by Console dashboard.\n\nOps: PUBLISH, SUBSCRIBE] as evt_cluster <<PUBSUB>>
**Default** (no flag) — starts `redis`, `server`, `bridge`,`console`. Requires an OpenAI-compatible LLM API running on the host (default: `http://localhost:8000/v1`).
**Default** (no flag) — starts `server` and`console`. Requires an OpenAI-compatible LLM API running on the host (default: `http://localhost:8000/v1`).
```bash
docker compose up
@@ -46,22 +39,12 @@ docker compose up
docker compose --profile production up
```
**Cluster** — 10-node server/bridge fleet sharing PostgreSQL and Redis. Access all nodes via the console at `:8090`. Requires `POSTGRES_PASSWORD`:
**Cluster** — 10-node server fleet sharing PostgreSQL. Access all nodes via the console at `:8090`. Requires `POSTGRES_PASSWORD`:
```bash
docker compose --profile cluster up
```
**Sim** — adds the simulator. Can run alongside the full stack or standalone with just Redis and the console:
```bash
# Sim + console (no LLM needed)
docker compose --profile sim up redis console sim
# Everything including sim
docker compose --profile sim up
```
## Configuration
All configuration is via environment variables in `.env` (copy from `.env.example`):
@@ -74,13 +57,6 @@ All configuration is via environment variables in `.env` (copy from `.env.exampl
| `OPENAI_API_KEY` | `dummy` | API key (`dummy` for local servers) |
| `TAVILY_API_KEY` | — | Web search API key (only needed for local/vLLM models; Anthropic and OpenAI search models use native search) |
| `TURNSTONE_DB_URL` | — | Database URL (e.g. `postgresql://user:pass@db:5432/turnstone`). For SQLite, defaults to `/data/.turnstone.db` |
| `TURNSTONE_DB_URL` | — | Database URL (e.g. `postgresql+psycopg://user:pass@postgres:5432/turnstone`). For SQLite, defaults to `/data/.turnstone.db` |
| `TURNSTONE_DB_POOL_SIZE` | `2` | PostgreSQL connection pool size per process (default: 2 base + 3 overflow = 5 max) |
| `POSTGRES_USER` | `turnstone` | PostgreSQL container username (used in default `TURNSTONE_DB_URL` for cluster/channel) |
| `POSTGRES_PASSWORD` | — | PostgreSQL container password (required for production and cluster profiles) |
The database stores workstream history, user accounts, and API tokens. When using JWT auth, a database backend is required for user storage.
> **Upgrading from <1.3.0a4:** Earlier versions used `DB_BACKEND` and `DATABASE_URL` in `.env`, which `compose.yaml` mapped to the `TURNSTONE_`-prefixed names internally. These short aliases have been removed. Rename `DB_BACKEND` → `TURNSTONE_DB_BACKEND` and `DATABASE_URL` → `TURNSTONE_DB_URL` in your `.env` file.
> **Large clusters:** Each turnstone process maintains a small connection pool (5 max). At hundreds of nodes this adds up — use [PgBouncer](pgbouncer.md) in transaction pooling mode between turnstone and PostgreSQL.
> **First-time setup:** After deploying with auth enabled, create an initial admin user by running `turnstone-admin create-user` inside the container:
@@ -129,30 +108,26 @@ The database stores workstream history, user accounts, and API tokens. When usin
| `TURNSTONE_SLACK_SLASH_COMMAND` | `/turnstone` | Slash command registered in the Slack app |
The channel service runs in the `production` profile. When`TURNSTONE_DISCORD_TOKEN` is set, the Discord adapter connects to the Discord Gateway and routes messages through Redis MQ to the bridge and server. See [Channel Integrations](channels.md) for full setup instructions including Discord application creation and user account linking.
### Simulator
| Variable | Default | Description |
|----------|---------|-------------|
| `SIM_NODES` | `100` | Number of simulated nodes |
The channel service runs in the `production` profile. When
`TURNSTONE_DISCORD_TOKEN` or the Slack pair is set the gateway starts the
corresponding adapter; both can run in one process. See
[Channel Integrations](channels.md) for platform app setup and user
account linking.
## Scaling
For multi-node testing, use the `cluster` profile which provides 10 dedicated server+bridge pairs with unique node IDs (`node-1` through `node-10`), resource limits, and shared PostgreSQL:
For multi-node testing, use the `cluster` profile which provides 10 server instances with unique node IDs (`node-1` through `node-10`), resource limits, and shared PostgreSQL:
```bash
POSTGRES_PASSWORD=secret docker compose --profile cluster up
```
The default `server` and `bridge` also run alongside the cluster nodes (11 total). All nodes are accessible via the console dashboard at `:8090`.
The default `server` also runs alongside the cluster nodes (11 total). All nodes are accessible via the console dashboard at `:8090`.
For production clusters beyond ~50 nodes, add PgBouncer between turnstone services and PostgreSQL. See [PgBouncer Connection Pooling](pgbouncer.md) for Docker Compose and Helm configuration.
@@ -160,7 +135,6 @@ For production clusters beyond ~50 nodes, add PgBouncer between turnstone servic
All entry points are installed in a single image: `turnstone-server`, `turnstone-bridge`, `turnstone-console`, `turnstone-channel`, `turnstone-admin`, `turnstone-sim`, `turnstone-eval`.
All entry points are installed in a single image: `turnstone`,
| `TURNSTONE_OIDC_PROVIDER_NAME` | No | `SSO` | Display name for the login button (e.g. "Google", "Okta") |
| `TURNSTONE_OIDC_ROLE_CLAIM` | No | — | ID token claim containing role/group values (see [Role Mapping](#role-mapping)) |
| `TURNSTONE_OIDC_ROLE_MAP` | No | — | Mapping from claim values to Turnstone role IDs (see [Role Mapping](#role-mapping)) |
| `TURNSTONE_OIDC_PASSWORD_ENABLED` | No | `true` | Set to `false` to hide the password form and block all username/password logins (including admin). API tokens and config-file tokens still work. |
| `TURNSTONE_OIDC_PASSWORD_ENABLED` | No | `true` | Set to `false` to hide the password form and block all username/password logins (including admin). API tokens continue to work. |
| `TURNSTONE_OIDC_REDIRECT_BASE` | No | — | Externally-reachable origin for the OIDC redirect URI (e.g. `https://app.example.com`). Recommended when running behind a reverse proxy. When unset, derived from the request Host header. |
OIDC is enabled when all three required fields (issuer, client ID, client
@@ -246,10 +246,10 @@ password) before OIDC is enabled. The setup wizard always works
regardless of this setting because it is only available when zero users
exist in the database.
API token login (`POST /v1/api/auth/login` with a `ts_` token) and
config-file tokens (`Authorization: Bearer tok_xxx`) continue to work
regardless of this setting. OIDC-only mode affects password-based
authentication only.
API token login (`POST /v1/api/auth/login` with a `ts_` token)
continues to work regardless of this setting. JWTs and API tokens are
the supported authentication methods. OIDC-only mode affects
This bumps `pyproject.toml` + `turnstone/__init__.py`, regenerates `uv.lock`, commits, tags `v1.5.0a2`, and pushes. CI runs, then publish + Docker workflows fire automatically.
## Releasing a Stable Patch (from stable/X.Y)
```bash
git checkout stable/1.4
git cherry-pick <commit-hash> # bugfix from main
scripts/release.sh 1.4.1 --push
```
## Promoting Experimental to Stable
When `main` is ready for a stable release:
```bash
# 1. Tag the stable release on main
scripts/release.sh 1.5.0 --push
# 2. Create the stable maintenance branch from that tag
git branch stable/1.5 v1.5.0
git push origin stable/1.5
# 3. Start the next experimental cycle on main
scripts/release.sh 1.6.0a1 --push
```
The previous stable branch (`stable/1.4`) continues to receive
security-only patches; older tracks (`stable/1.0`, `stable/1.3`) are
retired when they fall out of support.
## CI/CD Pipeline
All releases are gated on CI success:
1. `git push` with `v*` tag triggers **CI** (lint, typecheck, test, test-postgres, lock-check, security audit)
2. On CI success, **Publish to PyPI** fires via `workflow_run`
3. On CI success, **Publish Docker Image** fires via `workflow_run`
- `client.logout()` clears the stored JWT from the client.
- If a request returns 401, the SDK raises `TurnstoneAPIError` — the caller is responsible for re-authenticating.
### Backward Compatibility
### Token Types
The config-file token (`TURNSTONE_AUTH_TOKEN`) still works as a simple Bearer token for environments that do not use the user/JWT system. When the server receives a non-JWT Bearer token, it falls back to the legacy token check.
The SDK accepts any Bearer token — JWTs (from `ServiceTokenManager` or login) and API tokens (`ts_` prefix) are both supported. Use `token_factory` for auto-rotating JWTs or a static `token` for API tokens.
When a per-model override is `NULL` (empty in the UI), the global default is
used. Switching models via `/model <alias>` re-resolves sampling parameters
from the new model's overrides or global defaults.
**Removed settings:** `model.name` and `model.context_window` have been removed
from ConfigStore. Model names and context windows are now configured per-model
in the Models tab. A startup warning is logged if these keys appear in
`config.toml`.
### Plan / task agent overrides
`plan_agent` and `task_agent` sub-sessions resolve independently from the
conversation model so operators can pick a cheaper/faster model for
autonomous loops:
| Setting | Purpose |
|---------|---------|
| `model.plan_alias` | Alias used for `plan_agent` sub-sessions. Falls back to `[model].plan_model` in config.toml, then `[model].agent_model`, then the session's active model. |
| `model.task_alias` | Alias used for `task_agent` sub-sessions. Same fallback chain as `plan_alias`. |
The simulator (`turnstone-sim`) creates lightweight simulated nodes that talk to a real Redis instance using the standard turnstone protocol. External observers — `TurnstoneClient`, `turnstone-console`, real bridges — see identical behavior. No LLM backend is needed.
Use `--metrics-file report.json` to write the full report as JSON.
## Console Integration
The simulator's nodes appear in `turnstone-console` exactly like real nodes. Run them together to see the dashboard populate with simulated workstreams:
```bash
# Terminal 1: start Redis and console
docker compose up redis console
# Terminal 2: run simulator
docker compose --profile sim up sim
```
Or all at once:
```bash
SIM_NODES=50 SIM_DURATION=120 docker compose --profile sim up redis console sim
```
Open http://localhost:8090 to see simulated nodes, workstream states, token counts, and load bars updating in real time.
## Architecture
> See also: [Simulator Architecture diagram](diagrams/png/10-simulator-architecture.png)
```
turnstone/sim/
├── __init__.py # Public API: SimCluster, SimConfig
├── config.py # SimConfig — all simulation parameters
**Key design:** The `InboundDispatcher` batches ~50 node queues into a single Redis `BLPOP` call, keeping connection count bounded at ~20 regardless of node count. All nodes share a single `ConnectionPool(max_connections=64)`.
An MCP server that exposes tools for executing commands across a [Turnstone](https://github.com/turnstonelabs/turnstone) cluster. Serves as a reference implementation for both MCP server patterns and Turnstone MQ client SDK usage.
An MCP server that exposes tools for executing commands across a Turnstone cluster. Serves as a reference implementation for both MCP server patterns and Turnstone SDK usage.
> [!NOTE]
> **Superseded by the built-in coordinator workstream in Turnstone 1.5.**
>
> This MCP side-car is the pre-1.5 pattern for cluster-wide orchestration.
> Turnstone 1.5 promotes coordinator behaviour to a first-class workstream
> kind hosted inside `turnstone-console` — no external MCP server to
> install or operate, proper per-user audit attribution, and a dedicated
> UI at `/coordinator/{ws_id}`.
>
> The extension continues to work for 1.4-and-earlier clusters. On 1.5+:
> grant the `admin.coordinator` permission, set `coordinator.model_alias`
> in the admin Settings tab, and create sessions via the dashboard's
> "new coordinator" button or `POST /v1/api/coordinator/new`. Full
> removal of this example (including docker / compose references) is
> curl -X POST https://console.example/v1/api/coordinator/new \
> -H "Authorization: Bearer $TOKEN" \
> -H "Content-Type: application/json" \
> -d '{"name":"planner","initial_message":"Spawn a worker to check the build"}'
> ```
>
> The response carries `ws_id`; open
> `https://console.example/coordinator/{ws_id}` to watch the session.
## How it works
This server uses Turnstone's MQ client (`TurnstoneClient`) to dispatch shell commands to specific nodes via Redis. Remote agents execute the command and the raw bash output is captured directly from the `ToolResultEvent` stream — bypassing the costly "agent reads output → re-generates output as completion tokens" round-trip.
This server uses the Turnstone console SDK (`TurnstoneConsole`) for node discovery and routing, and `TurnstoneServer` for per-node SSE streaming. The dispatch flow for each command is:
1. **Route** — `TurnstoneConsole.route_create_workstream(target_node=..., auto_approve=True)` creates a workstream pinned to the target node via the console's rendezvous routing proxy, returning `ws_id` and `node_url`.
2. **Execute** — `TurnstoneServer(node_url, token=...)` connects directly to the node's SSE stream using the same `TURNSTONE_API_TOKEN`. `send_and_wait(prompt, ws_id)` runs the command and the raw bash output is captured from the `ToolResultEvent` — bypassing the costly "agent reads output then re-generates output as completion tokens" round-trip.
3. **Cleanup** — `TurnstoneConsole.route_close(ws_id)` closes the workstream.
Multi-node dispatches run in parallel via `asyncio.gather`, so total wall time is bounded by the slowest node rather than the sum.
@@ -19,8 +60,7 @@ Multi-node dispatches run in parallel via `asyncio.gather`, so total wall time i
## Prerequisites
- A running Turnstone cluster (at least one `turnstone-server`+`turnstone-bridge`)
- Redis accessible from wherever this MCP server runs
- A running Turnstone cluster with at least one `turnstone-server`and a`turnstone-console`
- Python 3.11+
## Installation
@@ -28,10 +68,6 @@ Multi-node dispatches run in parallel via `asyncio.gather`, so total wall time i
```bash
# From the turnstone repo root:
pip install -e ./examples/mcp-cluster-ops
# Or install turnstone with MQ support first, then the example:
The HTTP SDK (`TurnstoneServer`) talks to a single server instance. The MQ client (`TurnstoneClient`) routes through Redis with `target_node` support, which is the entire point of cross-node cluster operations.
## Security Considerations
**This MCP server grants the calling agent shell access to cluster nodes.**
@@ -104,8 +135,8 @@ The HTTP SDK (`TurnstoneServer`) talks to a single server instance. The MQ clien
is returned through the MCP tool result and becomes part of the LLM context.
- The security boundary is at the MCP host layer -- use Turnstone's tool
policy system to restrict which agents can invoke these tools.
- Set `REDIS_PASSWORD` via your environment or a secrets manager -- avoid
hardcoding passwords in config files.
- Set `TURNSTONE_API_TOKEN` via your environment or a secrets manager -- avoid
echo"Run 'git push origin <branch> $TAG' to publish"
fi
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.