5 follow-up comments from Copilot, all valid:
1. **CRITICAL — snap_seq race with split writer** (concurrency, 001).
Round-1's fix lifted snapshot capture into
register_listener_with_replay under nested locks, but the
WRITER side (on_content_token / on_reasoning_token) still
released _ws_lock before calling _enqueue (which bumps
_event_id under _listeners_lock). A reader could
interleave between writer's release and writer's _enqueue:
capture inflight WITH the new text, read STALE _event_id,
return snap_seq < new_event_id. The new event's live emit
then has _seq > snap_seq, slips past the dedup filter, and
double-renders text the snapshot already contained.
Fix: move self._enqueue(...) INSIDE the with self._ws_lock:
block in both token writers. The inflight mutation and the
_event_id advancement are now atomic against any snapshot
reader. Lock order _ws_lock (outer) → _listeners_lock
(inner via _enqueue) matches the snapshot helpers, so no
deadlock. Fan-out's put_nowait calls happen under
_ws_lock for token writers — microsecond cost per listener,
acceptable for the correctness guarantee.
2. **NIT — stale comment ref to buffered[-1]._event_id** (docs, 002).
The comment referenced a local var (buffered) that lives in
register_listener_with_replay, not in the events handler.
Reworded to describe the cutoff in terms of the last replayed
event id and the atomic-against-writers registration.
3. **MODERATE — 401 branch leaves reconnect loop** (bug, 003).
The coord's onerror schedules a 5 s CLOSED-state recovery timer
unconditionally. In the 401-expired-session branch we close
evtSource and showLogin — but the timer still fires 5 s later,
observes !evtSource, and calls scheduleReconnect(),
which opens a new EventSource that 401s again → infinite
reconnect loop while the login overlay is up. Fix: cancel
reconnectTimer in the 401 branch.
4. **MODERATE — race test was vacuous** (test_coverage, 004).
The previous regression test drained the listener queue after
register_listener_with_replay returned, but the helper
doesn't backfill buffered events into the queue, so the loop
was almost always a no-op and the assertion never executed.
Rewrote with a monkey-patched _enqueue that sleeps 50 ms
before bumping _event_id — widens the race window
deterministically. Verified: the test FAILS on pre-fix code
(snap.content has marker but snap.seq=0 < final_event_id=1)
and PASSES on post-fix code (writer holds _ws_lock through
_enqueue, so the reader blocks until writer fully done).
Also pinned the no-backfill contract so a future change adding
listener-queue backfill remembers to keep snap_seq the
high-water mark.
5. **NIT — except Exception too broad in test** (best_practices, 005).
Tightened except Exception: to except queue.Empty: so
unexpected exceptions aren't silently swallowed in the drain
loop.
Tests:
- 86 tests in test_sse_reconnect_replay.py + test_session_ui_base.py
pass (existing 84 + 2 new race regressions).
- Full non-live suite: 6347 passed, 15 skipped, no regressions.
- Ruff + mypy clean on changed .py files; JS parses.
Four issues raised on the merged PR #542, evaluated and fixed:
1. **Truncated-path snap_seq bug (Copilot low-confidence, VALID).**
make_events_handler's truncated branch set snap_seq = 0,
disabling the live-drain _seq <= snap_seq dedup filter. Any
token writer racing between register_listener_with_replay
returning and the live drain's first read would land in BOTH the
listener queue AND the captured snapshot text (the snapshot is
emitted via in_progress_snapshot as the recovery floor), so
the client double-renders. Fix lifts the snapshot capture INTO
register_listener_with_replay under the same nested-lock
acquire as the listener registration + buffer slice + counter
read, so the returned snapshot["seq"] is the exact
high-water mark the snapshot text corresponds to. Handler now
uses snapshot["seq"] as snap_seq on truncated, dropping
any token event with _seq <= snap_seq from the live emit.
2. **Lock-held string join in truncated path (Copilot, VALID).**
"".join(ui_base._ws_inflight_content) ran inside the
with ui_base._ws_lock: block, holding the lock for the
duration of the join and blocking on-token writers. Fix (folded
into #1's refactor): the new register_listener_with_replay
copies the inflight lists under lock and joins outside, matching
the existing pattern in
register_listener_with_in_progress_snapshot.
3. **_strip_js_comments docstring misclaim (Copilot, VALID).**
Docstring claimed the helper preserves "string/regex literals"
but the implementation only tracks string delimiters. Fix:
docstring updated to call out the regex-literal limitation
explicitly + note that current callers don't scan regions
containing regex literals. Extending the tracker is left for
a future caller that needs it.
4. **Coord scheduleReconnect dead-code regression (Copilot, VALID).**
After the PR-D refactor, scheduleReconnect() had no remaining
call sites — which meant reconnectAttempts never incremented,
wasReconnecting was always false, AND there was no
fallback when the browser transitioned the source to CLOSED
(hard 4xx after retries, intermediary tearing the connection
down with prejudice, etc.). The first failure mode silently
broke the post-gap replace-mode refresh of children / tasks /
wait indicator / live-badge cache; the second left the coord
permanently disconnected on non-transient failures. Fix:
- Introduce disconnectedSinceLastOpen flag set in onerror,
cleared in onopen. wasReconnecting reads it (with the
legacy reconnectAttempts > 0 fallback for the
scheduleReconnect-driven case), so the post-gap refresh fires
after every reconnect including the common native-reconnect
path.
- Re-introduce CLOSED-state recovery: onerror schedules a 5 s
delayed check via reconnectTimer; if the source is still
CLOSED at that point, call scheduleReconnect(), which
opens a new EventSource (threading the saved
lastEventId via the URL query param so replay still works
across the manual reconnect). Cancel/replace successive
timers so onerror floods don't pile up multiple checks for
the same source.
Tests:
- New test_truncated_path_snapshot_captures_real_snap_seq pins
the snap_seq fix at the helper boundary.
- New test_truncated_path_filters_already_in_snapshot_tokens
pins the end-to-end dedup invariant — would have caught the
double-render under the old code.
- Existing test_sse_reconnect_replay.py call sites updated for
the new 6-tuple return of register_listener_with_replay.
- All 86 tests in those two files pass; full non-live suite (6347
tests) passes; ruff + mypy clean on changed files; both JS files
parse-check.
Four /review findings collapsed to one code chokepoint + two
documentation fixes:
1. find's `kind` arg now validated against ``SkillKind`` (matching
create / update's existing pattern at session.py:8298 / :8512).
Closes two failure modes that shared the same root:
- typos (`kind="interactivee"`) silently produced
`kinds=["interactivee", "any"]` filtering to literal-`any` rows
only and masquerading as a narrowed catalog — now returns an
explicit "kind must be one of: ..." error;
- the documented enum value `kind="any"` degenerated to
`kinds=["any", "any"]` which narrowed to literal-`any` rows
instead of returning "every kind" — now collapses to ``None``
so the documented semantic holds.
2. docs/coordinator-skills.md "two-surface model" section rewritten
to reflect the post-flatten reality: kind is metadata, not an
enforcement boundary. The line-67 tools-table row updated from
the long-dead `list_skills` to `skills (action=find)` with the
opt-in kind-filter framing.
3. Three stale "interactive-only" comments in session.py
(:5514, :7857, :8210) that directly contradicted the
`_prepare_skills_load` docstring ("Both kinds can load") — drop
the qualifier so future grep-and-encode hazards don't reintroduce
the rejection.
Tests:
- test_find_kind_invalid_errors — typo case (replaces the silent
degenerate to literal-any-only)
- test_find_kind_any_means_no_filter — documented enum value matches
documented semantic (collapses to None at prepare)
- test_find_kind_narrow_passes_through — valid narrowing values
reach exec as expected
Deferred to release notes (no code change, intentional policy shift):
- skills(action='get') / load can now read full content + scan_report +
allowed_tools on cross-kind rows from any session. Operators with
pre-existing kind=coordinator skills authored under the prior
implicit visibility contract should audit those bodies for
sensitive content (allowed_tools allowlists, embedded credentials,
internal hostnames in examples) before upgrade.
Closes#557. SkillKind was authored audience metadata that the
discoverability filter dressed up as a runtime visibility gate. Real
access control is allowed_tools + auto_approve, which apply identically
across kinds. The kind-scoping chokepoints scaled linearly with every
new model-write surface for zero security payoff.
Drop kind consultation from:
- ChatSession._skills_kinds (deleted) and ._lookup_visible_skill
(deleted; callers inlined to storage.get_prompt_template_by_name).
- _exec_skills_find: no longer auto-threads kinds=. The opt-in `kind`
arg is a passable filter (threads [<kind>, "any"]) so the
discoverability win survives without enforcement.
- _exec_skills_get / _exec_skills_load: row lookup is name-only.
Disabled-row gate stays on load (admin quarantine is the actual
boundary). _prepare_task already uses unscoped get_skill_by_name;
session_routes.py already calls storage directly with no kind
check. Both confirmed by the spike, no source change needed.
- tools/skills.json: drop "Coord sessions see / interactive sees"
language; kind arg description re-cast as opt-in discoverability
narrowing.
- storage Protocol docstring + console_schemas.py kind field
description: refresh to reflect passive-metadata role.
Keep:
- SkillKind enum, kind column on prompt_templates, admin Skills tab
editing, kind field in skills.find / skills.get projection. The
field is useful for sorting/grouping at the model layer and as
authored intent.
- storage.list_skills_filtered(kinds=...) parameter — admin-filter
only now; docstring updated to note it's no longer auto-threaded
from the model-tool path.
Design calls:
1. find accepts opt-in `kind` arg: YES. ~5 lines on prepare + exec.
Threads kinds=[<kind>, "any"] only when supplied. Preserves the
model's ability to narrow a browse without enforcing.
2. kind field stays in find/get projection: YES. Already pulled
directly from the row dict in _skills_project_row (session.py
line 8187); the projection survives the flatten unchanged.
Tests:
- Delete TestLookupVisibleSkill (helper gone), the two
TestExecSkillsLoadKindScoping cross-kind reject branches, the two
test_find_kind_scoping_* tests, and test_get_cross_kind_returns_not_found
— the rejections those pinned are gone.
- Add test_find_default_threads_no_kind_filter (kinds=None by default
for both session kinds), test_find_returns_all_kinds_for_session
(interactive sees both interactive- and coord-tagged rows),
test_find_filters_by_kind_when_supplied (opt-in narrowing works),
test_get_returns_row_across_kinds (cross-kind get succeeds),
test_load_works_across_kinds (cross-kind load succeeds in both
directions — the flatten contract), test_load_rejects_missing_skill
(missing-row hint coverage). Keep test_load_rejects_disabled_skill
(admin quarantine still applies), test_load_works_for_coord_on_*
(coord-side load still works on more kinds now).
Storage tests untouched: tests/test_storage_skills_filtered.py
keeps its kinds= coverage (the parameter still works, just no longer
auto-threaded from the model-tool path).
Closes review findings from PR #555 that motivated the rethink:
sec-1 (_prepare_task unscoped) moot, sec-2 (HTTP create unscoped)
moot, sec-3 (no audit on cross-kind probes) moot — there is no
cross-kind concept anymore.
Boundary spike (verified against fresh main at 9a98d07d):
- _skills_kinds defined at session.py:7933, callers exactly two:
_lookup_visible_skill (7978) + _exec_skills_find (8057, 8109).
Verified via grep across turnstone/.
- _lookup_visible_skill defined at session.py:7943, callers exactly
two: _exec_skills_get (8213) + _exec_skills_load (8277). Verified
via grep.
- _skills_project_row reads kind directly from the row dict
(r.get("kind") or "any") at session.py:8187 — no helper call;
projection survives flatten.
- _prepare_task at session.py:6353 calls unscoped get_skill_by_name
— no kind check, no change needed.
- session_routes.py:1985 calls storage.get_prompt_template_by_name
directly with no kind check — already flat.
- storage.list_skills_filtered(kinds=...) parameter is identical
across _sqlite.py:2986, _postgresql.py:2827, and _protocol.py:1364.
- tests/test_skills_tool.py: TestLookupVisibleSkill (5 tests, lines
481-536) and TestExecSkillsLoadKindScoping (5 tests, lines 539-639)
pre-flatten. Total: 61 collected → 57 collected post-flatten.
Net production LOC: -15.
The local-escape posture from the prior commit double-encoded values
that inlineMarkdown had already escaped: leading escapeHtml(text)
turns `&` into `&`, the local escapeHtml(url) then turned that
into `&amp;`, which breaks query-string URLs after browser parse
+ getAttribute + new URL round-trip.
Switch to convention-rename: regex callback params renamed to
safeAlt / safeUrl / safeLabel to signal the upstream-escape
invariant. The attribute-context lint enforces all future
attribute-context concat sites maintain the safe* convention or
call escapeHtml explicitly — defense-in-depth preserved without
the regression. Two added pin tests verify `&` survives with
single (not double) entity encoding through image data-src and
link href.
Also addresses two test issues from the same review:
- Docstring listed `safe[A-Z_]…` but code only checked isupper().
Drop the underscore option (JS uses camelCase anyway).
- `_all_attr_names` only recorded attribute-bearing tags, so a
bare `<script>` injection would have false-negatived the link-
label pin test. Refactored to `_parse_renderer_html` returning
both start tags and (tag, attr) pairs.
inlineMarkdown's image and link renderers now escapeHtml each
interpolated value (url, alt, label, domain) at the call site
instead of relying on the upstream escape pass. Defence-in-depth:
a future refactor calling those renderers from outside
inlineMarkdown would otherwise silently regress.
New CI lint scans renderer.js for `attr="' + ident` patterns; ident
must be escapeHtml(...), safe*, or in the reviewer-approved
allowlist. Four pin tests use html.parser.HTMLParser to verify
attacker URLs and labels don't materialize event-handler attributes
on rendered DOM.
Three independent fixes flagged by Copilot's review on PR #555:
2. ``update`` auto_approve self-escalation warning false-positive
(turnstone/core/session.py:_prepare_skills_update)
- The warning was computed against ``existing_auto_approve or
proposed_auto_approve`` — meaning an update that explicitly
turned auto_approve OFF still triggered the warning because the
existing row had it ON. Now computes against the *final state*
(``updates["auto_approve"]`` if present, else
``existing.get("auto_approve")``) combined with the final
``allowed_tools`` value. False-positives gone; the inverse case
(existing auto_approve=False, update turns it ON without
touching allowed_tools) now correctly fires the warning against
the inherited allowlist.
3. ``temperature`` validator silent-coerce → explicit error
(turnstone/core/skill_field_validation.py:parse_skill_session_config)
- Non-numeric temperature input silently coerced to ``None``,
unlike ``max_tokens`` / ``token_budget`` which return an error.
Numeric-field consistency: temperature now errors on
unparseable input with "temperature must be a number between 0
and 2". Range check unchanged; blank / None still → None.
4. Version-snapshot uses max+1, not count+1
(turnstone/core/session.py:_exec_skills_update)
- ``count_skill_versions + 1`` re-uses version numbers when any
row has been deleted via the existing
``storage.delete_skill_versions`` method, and the schema has no
``(skill_id, version)`` unique constraint to catch the
collision. Switched to ``max(list_skill_versions)`` + 1,
matching the ``storage.unlock_skill`` pattern. A storage-side
atomic allocator is the right architectural fix and is tracked
for a future PR.
Tests cover both the false-positive and inverse-positive auto_approve
cases, the new temperature error path, and the version-numbering edge
case where prior versions have been deleted (max diverges from count).
Two changes that share the same kind-scoping touch point.
Lookup unification (closes the bypass Copilot flagged on _exec_skills_load):
- New ChatSession._lookup_visible_skill(name) — single source of truth for
"find me a skill by name, if it's visible to this session". Combines
storage.get_prompt_template_by_name with the kind filter in one call;
returns None for both the missing-row and out-of-kind cases so callers
don't have to branch on the reason.
- _exec_skills_get refactored from inline two-step to one helper call.
- _exec_skills_load refactored from the unscoped memory.get_skill_by_name
to the new helper — the kind-scoping bypass it had (interactive could
load a kind=coordinator skill by name) goes away by construction
because the unscoped path no longer exists on the model-tool surface.
- memory.get_skill_by_name stays available for admin / sub-agent /
rehydrate paths that need full-catalog visibility — those are
deliberate cross-kind callers, not bypass surfaces. Storage exceptions
now propagate from _lookup_visible_skill by design (distinct from the
legacy swallow-and-return-None) so the operator gets a clear signal on
DB outage rather than a misleading "not found".
Coord-side load support:
- _prepare_skills_load no longer rejects coordinator sessions. Parity
with the admin / HTTP create path that already accepts a `skill` body
field on kind=coordinator workstreams — what the operator can do at
create time, the model can now do on its own session. Visibility is
still kind-scoped via _lookup_visible_skill at exec (a coord can only
load {coordinator, any}-tagged skills; interactive can only load
{interactive, any}), matching what `find` / `get` enforce.
The kind-scoping itself is queued for a separate cleanup PR: the marker
turned out to be a discoverability hint that never gated runtime
capability, and the combinatorial complexity (every new model-tool /
HTTP path needs kind awareness) isn't worth the squeeze at this team
size. Follow-up issue to land.
Test coverage:
- TestLookupVisibleSkill — 5 cases: visible / cross-kind / missing /
storage-unavailable / kind=any-on-both-surfaces.
- TestExecSkillsLoadKindScoping — kind-rejection from both directions
(interactive→coord-only, coord→interactive-only), disabled-skill
caller-side gate, and the two new positive coord-load cases (coord
loads kind=coordinator and kind=any).
- Removed test_load_on_coord_session_errors (the rejection it pinned
is gone).
Plus the /review-suggested doc fixes that came with the unification:
- Comment in _exec_skills_load now correctly attributes the disabled
collapse to the caller's enabled check rather than implying the
helper handles it.
- _lookup_visible_skill docstring documents the deliberate
exception-propagation behavior.
Replaces the legacy `skill` (load + search) and `list_skills` tools with a
single `skills(action=...)` tool serving both interactive and coordinator
sessions. Stacks on the model.skills.write permission introduced in PR 1.
Tool surface
- `find`: filter by category/tag/risk_level/enabled_only/limit with
optional BM25 query ranking; auto-approved on both kinds; kind-scoped at
the storage filter (interactive sees interactive+any, coord sees
coordinator+any).
- `get`: fetch a single skill including content; cross-kind misses
collapse to "not found" so a model can't enumerate the other surface
by name-probing.
- `load`: activate a skill in the current session (interactive-only;
coord sessions get an explicit hint pointing at spawn_workstream).
- `create`/`update`/`enable`/`disable`: require approval AND
model.skills.write; permission re-checked at exec time to catch a
revocation between approval and write.
- No `delete` — hard-delete stays admin-UI exclusive; tool description
documents the soft-delete-via-disable pattern.
Defenses on the write surface
- Approval cards surface projected risk_level (scanner re-run against
the proposed final state) and warn explicitly when allowed_tools +
auto_approve combine (auto-fire-on-load consequence is spelled out,
not just shown as raw field values).
- Toggle preview surfaces existing risk_level + allowed_tools count so
re-enabling a critical-tier skill is never a one-click bypass.
- Update path now re-fetches the row at exec to catch a readonly flip
between approval and write, filters updates back to the runtime-only
set if so, refuses if no fields survive.
- Update path rejects empty content (hollow-out via emptying bypassed
the soft-delete-via-disable invariant), non-list tags, and empty
category — failures are loud rather than silent.
- Permission denials audit `skill.write_denied` with actor_source=model
so probing the permission state leaves a trail. Audit failures log
at error (not warning) — a successful write without a row is the
exact gap the trail exists to surface.
- `_skill_hint` routes both message and system_reminder through
escape_wrapper_tags so caller-controlled values can't close the
<system-reminder> envelope and let the model fabricate directives in
its own future context.
Shared validation
- `parse_skill_session_config` lifted from console/server.py to
turnstone/core/skill_field_validation.py; both the HTTP admin path and
the model-tool path consume it. Single source of truth so field rules
can't drift between layers.
- `SKILL_RUNTIME_CONFIG_FIELDS` lifted similarly (was duplicated as
_SKILL_RUNTIME_CONFIG_FIELDS in server.py and _SKILLS_READONLY_FIELDS
on ChatSession).
- `notify_on_complete` validator now accepts list input from the JSON
schema's `array` type — previously rejected because str() of a list
yields Python repr that json.loads then refuses.
Performance
- Update prepare skips the projected-risk scan when neither content nor
allowed_tools is changing (storage re-scans on write authoritatively).
Metadata-only updates no longer pay the ~25 regex-pass scan cost.
Cleanup
- CoordinatorClient.list_skills deleted (-91 lines); model-tool path
talks to storage directly via list_skills_filtered.
- Roles admin UI gains a Model section exposing model.skills.write.
- tests/test_load_skill.py renamed to tests/test_skills_tool.py and
rewritten for the new tool — 48 tests covering registration, prepare
dispatch, permission gating (including TOCTOU-revoked exec deny),
audit actor_source on create + disable + permission-denied probe,
BM25 ranking, invalid-kind branches, audit-failure swallow, and
<system-reminder> envelope injection resistance.
In-process permission check for model-facing tool exec paths that need
to gate a write capability without HTTP middleware in the loop. Foundation
for the upcoming skills tool refactor: the merged
skills(action=create|update|enable|disable) tool will gate on
model.skills.write before reaching storage.
- Add model.skills.write to _VALID_PERMISSIONS (default-ungranted on every
role including builtin-admin — operators opt themselves in explicitly)
- Add user_has_permission(user_id, permission, *, storage=None) helper
that fails-closed on storage outages and short-circuits on empty user_id
- Document service-scope asymmetry with require_permission (no AuthResult
in the model-tool path → no bypass; explicit guidance if a legitimate
service-scope caller ever needs to reach here)
- Pin the "no implicit cache" contract with a regression test asserting
every helper call hits storage (call_count == 2 after two calls)
- Lock the "builtin-admin default-ungranted" invariant with an alembic
migration test that drives the chain to head and asserts the role's
permission string omits model.skills.write
- Plus the role-create end-to-end test proving the constant flows through
the admin endpoint's validator
Roles admin UI changes deferred to the PR that lands the gated tool — no
operator action needed until the capability exists.
Per-call DB hit + warning-log spam on outage deferred to a follow-up PR;
the helper is dead code in this commit, so cache TTL would be sized
against guesswork — better to wait for a real call-rate signal from the
first caller.
Starlette 1.0.0 reconstructs request URLs without validating the Host
header, allowing path-injection that can bypass authentication on apps
comparing reconstructed URL paths instead of `request.url.path`. Fixed
in 1.0.1.
- pyproject.toml: bump `starlette>=0.45` to `starlette>=1.0.1` so the
CVE floor is explicit at the dependency declaration, not just in the
lockfile. Annotated with the advisory ID so the rationale survives
a future floor relax.
- uv.lock: regenerated via `uv lock --upgrade-package starlette`;
starlette 1.0.0 -> 1.0.1, no transitive bumps.
Locally verified `pip-audit --strict` returns clean after the bump and
the auth + service-boundary test suites (250 tests covering the URL/
host-header reconstruction surface) continue to pass.
2000 was sized for the cloud-provider regime (50–200 events/sec)
and was too small for the two regimes that actually shape PR-D's
recovery floor:
1. **Local inference**: vLLM / llama.cpp hit 500–2000 tok/s per
active stream. Each token is an _enqueue call, so a single
busy workstream burns through 2000 events in ~1 s. Reconnects
after any disconnect longer than a network blip immediately
fall through to the replay_truncated recovery path on a
stream that was supposed to be transparently resumable.
2. **Backgrounded tabs**: Chrome (and Firefox to a lesser extent)
throttle the SSE-drain microtask aggressively when a tab isn't
visible — Chrome's background-tab budget drops to ~1 wake/min
after ~5 min hidden, so a backgrounded pane can legitimately
sit on tens of seconds of un-drained events. PR-G (drop-pings-
let-it-die) deliberately closes those connections on hide and
re-opens on focus return; reconnect-with-replay is the only
recovery path, and if the buffer evicted in the interim, the
snapshot floor is all that's left for past-turn structural
events (tool calls, state changes, approvals).
50000 at the 2000-tok/s local-inference rate buys ~25 s of pure
token streaming before truncation; at cloud rates it's minutes of
coverage. Memory cost is ~200–500 bytes per event (deque node +
dict + payload), so 50000 × 100-ws design ceiling caps at roughly
2.5 GB worst-case — and practically nowhere close because the cap
is per-ws ceiling, not per-ws steady-state. Operators on heavier
workloads can raise via TURNSTONE_SSE_EVENT_BUFFER_MAX.
Considered and rejected: in-buffer coalescing of consecutive
content/reasoning tokens. A naive text-merge breaks the replay-
slice semantic — a coalesced entry has the latest _event_id
but text that includes content the client already received under
an earlier id, so any consumer with last_event_id falling
INSIDE the coalesced span would double-render on replay. A
correctness-preserving coalesce would need a per-consumer high-
water tracker we deliberately don't maintain. Bigger cap +
simple per-event storage avoids the trap; the rationale is
captured inline in _resolve_event_buffer_max.
The browser-side completion of PR-D reconnect-with-replay. Today's
`onerror` handlers on `Pane.connectSSE`, `connectGlobalSSE`, and
the coordinator's `connectSSE` all explicitly call
`evtSource.close()` on the transient-error path — that forces the
source into the terminal CLOSED state, defeating EventSource's
native auto-reconnect (which would otherwise reconnect with the
`Last-Event-ID` header that PR-D commit 1 now honours server-side).
Three handler refactors share the same shape:
- Remove the unconditional `close()` from the transient-error
branch. Native EventSource handles CONNECTING -> CONNECTING ->
OPEN with replay automatically.
- Keep UI updates (status bar dim, Reconnecting… text) — those
are orthogonal visualizations of the disconnected state.
- Keep terminal-branch closes: a 401 expired-session still does an
explicit close + showLogin (the user must re-authenticate); a
workstream-reassignment to a different ws still disconnects +
connects on the new wsId (it's a different stream, not a same-
stream replay).
- Capture `lastEventId` in `onmessage` BEFORE `JSON.parse` so a
malformed event doesn't desync the manual-reconnect fallback
from native auto-reconnect.
- Thread `?last_event_id=N` on the URL when constructing a fresh
`new EventSource(url)` — the constructor can't set custom
headers so the query-param fallback covers the manual-reconnect
path (initial connect with a saved id, scheduleReconnect after
an explicit close, etc.).
For `Pane.connectSSE`, the long focused-pane workstream-refetch
body inside `onerror` is lifted to a dedicated
`_refetchWorkstreamsAndReassign` method so it survives the
refactor as an orthogonal trigger (handles the workstream-evicted-
during-disconnect recovery case, which is independent of the SSE
reconnect mechanics). The reassignment branch's existing
`disconnectSSE + connectSSE(newWsId)` sequence stays — different
workstream genuinely needs a fresh stream. When reassigning, the
saved `_lastEventId` is dropped because replay is per-ws and an
id from ws-A is meaningless against ws-B.
Tests in `tests/test_app_js.py` add 3 static lint guards that
fail loudly if any future refactor reintroduces a naked
`evtSource.close()` in a transient-error path of any of the three
handlers. The guards understand the allowed terminal-branch
exceptions (401, login overlay, reassignment) and ship with an
escape hatch (functions that explicitly reference `last_event_id`
have taken explicit responsibility for the replay header and are
exempt). A small `_strip_js_comments` helper handles the
apostrophe-in-comment hazard that pre-existing
`_slice_balanced_body` doesn't (comments are stripped before
brace-walking; offsets preserved by space substitution).
The console SSE proxy (`_proxy_sse`) is the inbound SSE path for
multi-node deployments — every browser EventSource that targets a
per-node route traverses it. Today's proxy strips client request
headers (only `Accept`, `Cache-Control`, and the re-minted auth
token make it upstream), so the per-ws / global SSE handlers'
`Last-Event-ID` resume (PR-D commit 1) never sees the header in
the multi-node shape — every reconnect would be a fresh connect
and silently drop events from the disconnect window.
Builds the upstream headers dict conditionally: copy `Last-Event-ID`
from the incoming request when present, omit otherwise (no
fabricated value on fresh connects). Starlette's header dict is
case-insensitive so the `request.headers.get("last-event-id")`
lookup catches both the spec-recommended capitalization and any
intermediary normalisation.
The query-param fallback (`?last_event_id=N`) needs no proxy
change — `request.url.query` is already forwarded verbatim at the
top of the function.
Tests in `tests/test_service_auth_boundary.py::TestProxySseLastEventIdForwarding`:
- Positive: browser header → upstream header (value preserved).
- Negative: browser sends nothing → upstream gets nothing (no
fabricated value).
Adds the server-side foundation for SSE reconnect-with-replay (PR-D
in issue #540's sequencing): a per-ws monotonic ring buffer that
holds the last N events for replay against a client's
`Last-Event-ID` header (or `?last_event_id=N` query-param
fallback for manual reconnect paths that can't set custom headers).
Per-ws lane (SessionUIBase + make_events_handler):
- `_event_buffer` deque (cap 2000, env-overridable via
`TURNSTONE_SSE_EVENT_BUFFER_MAX`) holds (event_id, event_dict)
tuples; `maxlen` evicts the oldest automatically.
- Existing `_ws_inflight_seq` renamed to `_event_id` and lifted
to live alongside the listeners — one monotonic counter drives
both the new replay slice AND the existing `_seq`/`snap_seq`
snapshot dedup (byte-identical contract on token events).
- `_enqueue` now stamps every event with `_event_id` (and `_seq`
on `content`/`reasoning` token events) under
`_listeners_lock`, so the buffer append + listener fan-out + new
listener registration are all atomic against each other.
- New `register_listener_with_replay` returns
(queue, replay_events, status, lost_count, earliest_id) where
status ∈ {replay_ok, truncated}. `make_events_handler` reads
`Last-Event-ID` (header or query), branches three ways
(fresh / replay_ok / truncated), and emits the SSE `id:` field
on every event sourced from the buffer. On `replay_ok` the
in-progress snapshot is skipped (the buffered events already
cover it); on `truncated` an explicit envelope precedes the
fresh-style recovery path.
- Every events stream emits a jittered `retry:` in [2500, 4500] ms
on first yield so 6-pane reconnects don't lockstep on
EventSource's default ~3 s interval.
Global lane (server.py / _global_fanout_thread / global_events_sse):
- Parallel buffer + counter on `app.state.global_event_buffer` and
`app.state.global_event_id_holder`; fanout thread stamps each
event with `_event_id` and appends to the buffer under
`global_listeners_lock`. `global_events_sse` branches on
`Last-Event-ID` with the same three shapes.
Tests:
- 16 new tests in `tests/test_sse_reconnect_replay.py` cover the
ring buffer semantics (empty-listeners hold, last_event_id
slicing, truncation, atomic registration), the counter
invariants (monotonic under concurrent writers, no skip on
queue.Full, persists across turn boundaries, cross-thread
consistency), and the handler branching (retry on first yield,
id: on buffered events, snapshot-skip on replay_ok, envelope on
truncated, query-param fallback, malformed header → fresh).
- Existing `tests/test_session_ui_base.py` updated for the
`_ws_inflight_seq` → `_event_id` rename and the new
`_event_id` field on enqueued events.
Backward-compat: all consumers that don't send `Last-Event-ID`
(today's browser, Python SDK, TypeScript SDK, channel adapter) see
behaviour identical to pre-PR — the server change is purely
additive on the request side.
Mirror the \s* posture used by the other unsafe-sink clauses (eval\s*\(,
Function\s*\(, setTimeout\s*\() so a regression like
``el.insertAdjacentHTML ("beforeend", x)`` — or a multi-line form with a
newline before the paren — still trips the lint. The trailing ``HTML``
literal continues to discriminate against insertAdjacentElement and
insertAdjacentText.
Caught by Copilot review on #541.
Extend _UNSAFE_CODE_SINK_RE with an `.insertAdjacent` + `HTML\(`
alternation so insertAdjacentHTML(...) is flagged across all 8 tracked
JS bundles. The `HTML\(` suffix excludes insertAdjacentElement, which
takes a DOM node and is not an XSS sink — the five remaining sites in
ui/static/app.js (lines 170, 216, 328, 330, 1578) stay clear.
Retire the two carve-out paragraphs (file-level comment + function
docstring) that named ui/static/app.js's verdict-badge writers as the
reason the lint hadn't already broadened. Commit 1 of this PR cleaned
both writers, so the carve-out is no longer load-bearing.
After this commit the DOM-cleanup arc (started in #532) is complete:
every unsafe-write sink family — inner/outer-HTML assignment (plain +
concat), insertAdjacentHTML, document.write, string-eval, dynamic-
Function, string-first-arg setTimeout/setInterval — is forbidden
across all 8 LLM-rendering bundles.
Rewrite the verdict-badge HTML builder from string-concat into DOM
construction (createElement + textContent + setAttribute + append).
The helper now returns a DocumentFragment of two top-level siblings
(.verdict-badge and .verdict-detail), which appendChild expands into
the parent — preserving the sibling-traversal invariants relied on by
Pane.updateVerdictBadge, toggleVerdictDetail, and the d-key keyboard
shortcut.
Inline onclick="toggleVerdictDetail(this)" replaced with an
addEventListener click handler; the non-arrow callback keeps the
`this`→button binding the old inline form had.
Both call sites (replayHistory + the live approval flow) swap from
el.insertAdjacentHTML("beforeend", X) to el.appendChild(X).
This is the last unsafe-write site in the DOM-cleanup arc started in
#532; commit 2 broadens the test_app_js.py lint regex to forbid the
insertAdjacent-HTML sink across all 8 tracked JS bundles.
The pre-push /review pass surfaced a second const-reassign that mirrors
the original `redacted` bug but in prefix-increment form:
const _paneCounter = 0; // turnstone/ui/static/app.js:10
class Pane {
constructor(wsId) {
this.id = "p" + ++_paneCounter; // line 14 — TypeError at runtime
…
}
}
`new Pane(...)` throws `TypeError: Assignment to constant variable.`
on every pane construction. The first iteration of the const-reassign
guard in tests/test_app_js.py missed it because the regex matched
postfix `X++` / `X--` but not prefix `++X` / `--X`.
Two changes:
1. Change `const _paneCounter = 0` to `let _paneCounter = 0` at
turnstone/ui/static/app.js:10. Same fix shape as the `redacted`
bug — original walker tightened to const because its reassignment
regex also only matched postfix forms.
2. Extend the reassignment regex in test_swept_bundle_has_no_const_reassign
to detect prefix `++X` / `--X` so a third repeat of this class
can't ship. Verified by injection: temporarily reverting (1)
makes the new guard fire with a clear source-text diagnostic.
Quality polish on the same test (q-1/q-2 from the pre-push pass):
- Failure message now prints the offending decl + reassignment line
text alongside line numbers, so CI failures are self-contained
(was: opaque tuples requiring two file-jumps to interpret).
- Comment on `_SWEPT_BUNDLES` documents the maintenance contract
(add only after sweeping; coordinator.js intentionally excluded).
After the var → const/let sweep, four guards keep the post-sweep state
honest in CI:
1. node --check per bundle (parse-level smoke; catches a future edit
that drops a brace or mis-balances a string before it reaches the
browser).
2. Static var-free assertion per bundle pins the keyword-swap result —
any future `var X = …` in these 7 files fails CI loudly.
3. Scope-aware static const-reassign guard per bundle. For each
`const X = …`, scans only the enclosing block (innermost { … } via
brace tracking with regex/string/comment awareness) for X
reassignments, so a same-named `let X` in an unrelated function
doesn't false-positive against a `const X` in this one. Catches
the bug class that shipped through the original sweep:
_redactApiKeys's `const redacted; redacted = …` threw TypeError
at call-time, invisible to node --check.
4. Runtime smoke for _redactApiKeys via `node -e` — calls the
function with both query-string (`api_key=…`) and JSON
(`"api_key": "…"`) shapes. This is the bit that would have
caught the actual shipped TypeError; (3) is the equivalent
static check that catches the class without needing a runtime
invocation.
Bundle list:
- turnstone/ui/static/app.js
- turnstone/console/static/admin.js
- turnstone/console/static/governance.js
- turnstone/console/static/app.js
- turnstone/shared_static/auth.js
- turnstone/shared_static/kb.js
- turnstone/shared_static/utils.js
Verified by injection: temporarily reverting `let redacted` to
`const redacted` makes both guard (3) and guard (4) fail loudly.
Follow-up to the initial var-sweep commits. The walker used a flat,
file-wide reassignment check to decide let vs const, which was
conservative when the same name appeared in multiple unrelated
scopes — e.g. `let i` as a loop counter in one function and an
unrelated `let i` reassigned in another would both stay `let`.
This second pass uses brace-tracking block-scope analysis (regex
literal aware) so tightening considers only reassignments within
the same block:
- console/static/app.js: +5 const -5 let
- console/static/governance.js: +15 const -15 let
- console/static/admin.js: +26 const -26 let
Mirrors q-2 from the /review pipeline. ui/static/app.js was
tightened in the same way already in its sweep commit.
Mechanical; no behavioural change.
754 line-start var + 50 for-loop counters converted: 663 const, 96 let
(line-start), plus 50 for-init let counters.
Hand-fix sites (surfaced by spike § 2):
- 4 try-block hoists where the var was referenced from outside the
try (var hoists out, let does not):
- tryParseMedia()'s `obj`
- _tryPrettyJson()'s `obj`
- tryParseMcpError()'s `obj`
- inline-plan render's `action` (used in the catch handler)
- showNewWsModal() cleanup (was: 2 same-scope redeclarations):
- submitBtn — first lookup at the top of the modal kept; the
redundant re-fetch + duplicate textContent at the bottom
dropped; submitBtn.disabled = false now sits as a bare
property write
- defaultOpt → renamed second occurrence to tplDefaultOpt
(genuinely distinct DOM element — modelSelect vs tplSelect),
both can be const
Post-review fix to the walker output:
- _redactApiKeys(): the walker tightened `let redacted` to `const`
but missed the `redacted = redacted.replace(...)` reassignment
on the JSON-style pass. Root cause was the walker's
find_decl_extent not recognising JS regex literals — the
unescaped " inside the character class [^&\s"] opened an
in_str state that never closed on the same line, spilling
the declaration span past `);` and pulling the reassignment
line into the skip set. The /review pipeline's bug finder and
security finder both caught it (rendering would have thrown
TypeError on every tool-output render). Now `let redacted`.
Scope-aware const-tightening pass on top of the walker (mirrors q-2
from /review): 45 additional `let` → `const` flips where the walker
was conservative because the name happened to be reassigned in an
unrelated function elsewhere in the file. Examples: `let pane` in
the 4 plan-dialog helpers; `let el` in the small Pane class methods.
Each tightening is verified safe by a brace-tracking block-scope
analysis (regex-literal aware).
Mechanical; no behavioural change.
744 line-start var + 87 for-loop counters converted: 606 const, 138 let.
Includes 2 multi-decl sites (counter accumulators at 3389 and 4021,
both `let` because the names are reassigned via += in the loop body),
and the spike-identified `indicator` redeclaration in `_toggleOidcPanel`
(now two `const indicator` declarations in disjoint block scopes —
inner if-block at 455 and function body at 482, so block-scoping makes
them independent).
Mechanical; no behavioural change.
470 line-start var + 64 for-loop counters converted: 337 const, 133 let.
Includes 2 multi-decl sites (`let url, method;` at 3951 and 4504 — both
uninitialised pairs that stay `let`) and two sibling `for (var k …)`
loops at lines 112/119 in the same function (now `for (let k …)` —
block-scoped to each loop init, no collision).
Mechanical; no behavioural change.
67 line-start var + 4 for-loop counters converted: 55 const, 12 let.
The 12 let cases are all genuine reassignments:
- Top-level state (`_loginBusy`, `_authMode`, `_refreshTimer`, etc.)
- `let delay` in `_scheduleRefreshAt` (clamped to min/max)
- `let data` inside `_tryRefresh` (assigned from inner try-catch)
- For-loop counters `let attempt`, `let i`
Mechanical; no behavioural change.
12 line-start var + 1 for-loop counter converted: 15 const, 1 let.
File previously had 4 const from the DOM-cleanup helpers; sweep finishes
the conversion.
Walker is scope-aware: when checking if name X is reassigned anywhere
in the file, lines that themselves declare X (`let X = ...`, function
parameters `(X)`, etc.) are skipped — `X = ...` in another scope is a
new binding, not a reassignment of the original. This lets variables
like `min`/`hr` (declared inside two different formatter functions)
both become `const` correctly.
8 line-start var declarations converted: 6 const, 2 let.
Walker rules:
- var X = init → const X = init when X is never reassigned in the file
- var X = init → let X = init when X is reassigned (e.g. _kbPreviousFocus
assigned in showKbHelp, html accumulated via +=)
- Reassignment check uses negative lookbehind to skip property writes
(obj.X = ...).
Mechanical; no behavioural change.
* feat(reasoning): Phase 5 — vLLM Chat Completions reasoning-field replay
Multi-turn CoT replay for vLLM-served reasoning models (Qwen3, DeepSeek-R1)
via the non-standard `reasoning` field on assistant messages. Closes the
PR #498 gap claiming Chat Completions has no replay surface — vLLM's
ChatMessage.reasoning input field is that surface (verified in
vllm/entrypoints/openai/chat_completion/protocol.py:54-64).
Session-level attach (no provider class changes). Three-gate composite:
provider isinstance OpenAIChatCompletionsProvider AND
server_compat.server_type == "vllm" AND operator-set
ModelConfig.replay_reasoning_to_model. Deliberately drops the
supports_reasoning_replay capability gate that protects Paths 1+2 —
vLLM's failure mode is silent (template-drop), not loud (server 400),
so the static gate would add operator friction without preventing the
silent failure. Server-type pin bounds blast radius — canonical OpenAI,
llama.cpp, sglang never see the non-standard field.
Also fixes a pre-existing _resolve_server_type bug: it read
cfg.capabilities.get("server_compat") but the model_registry loader pops
server_compat OUT of capabilities into the dedicated cfg.server_compat
dataclass field (model_registry.py:401, 485). Pre-fix the function
returned "" for every production ModelConfig, silently degrading PR #498
Path 3's synth-block source tag and would have made Phase 5 dead-on-
arrival. Test stubs across 3 files updated to mirror production shape
(empty capabilities + populated top-level server_compat) so the same
stub-drift can't hide future regressions.
The agent _run_agent path is deliberately excluded from Phase 5 hoists:
agent assistant messages don't carry _provider_content (rebuilt per
invocation from CompletionResult.content + tool_calls), so the helper
would no-op every turn. Comment at session.py inside _api_call documents
the exclusion.
OpenAI SDK version pin raised to >=2.37 to match the version verified
by the cross-boundary regression test
(test_reasoning_field_present_in_wire_body_when_attached) — drives a
real OpenAI client through httpx MockTransport and asserts the
non-standard field reaches the captured POST body, catching any future
SDK version that adds runtime field filtering.
Tests: 10 helper unit + 17 session integration (incl. SDK boundary
round-trip + per-gate negative tests + call-site wiring tests) + 2
audit-log discipline tests extending the PR #498 logging contract.
* docs(reasoning): apply PR #537 review on Phase 5 docstrings
Two nits from PR #537 review:
1. `_resolve_server_type` docstring claimed Phase 5 (`_maybe_attach_vllm_chat_reasoning`) called it; in fact Phase 5 reads `cfg.server_compat["server_type"]` directly off the single cfg it fetches for the operator-flag check, to avoid a second `registry.get_config` round-trip. Rewrite the paragraph: name `_maybe_synth_reasoning_block` as the sole caller (informational metadata for UI rehydration), then a separate paragraph noting Phase 5 reads the same field path directly and that both readers MUST stay aligned on changes.
2. `_maybe_attach_vllm_chat_reasoning` docstring referenced `project_reasoning_replay_capability_gate.md` which lives in personal memory store, not the repo. Replace the dead-link reference with an inline summary of the asymmetry rationale (Paths 1+2 keep the dual-gate because loud server-side failures; Path C drops the static gate because vLLM's failure mode is template-drop silent).
Three quality findings from the multi-stage /review pass on the
preceding 4-commit class-refactor stack. Bundled into one commit
because each is sub-20-line documentation/test-hygiene with no
behavioral surface.
1. **Drop stale verdict-badge line numbers in tests/test_app_js.py.**
Two comments cited ``ui/static/app.js:1287`` and ``app.js:1538``
as the ``insertAdjacentHTML`` + ``renderVerdictBadge`` consumer
sites. The class refactor moved them to 1440 and 1655 (and any
future nearby edit will move them again). Drop the numbers; cite
the helper name (``renderVerdictBadge`` / "the verdict-badge
writers") instead.
2. **Drop ``.prototype`` from 3 coord comment cross-refs.**
``coordinator.js:339, 439, 566`` referenced
``Pane.prototype.addUserMessage`` / ``addUserReminder`` /
``addToolReminder`` / ``replayHistory`` — but ``app.js`` has zero
``Pane.prototype.X`` after the refactor (it's all ``Pane.X``
class methods now). A reader following the breadcrumb hits a
grep dead-end.
3. **Introduce indent-agnostic _pane_method_offset() helper.**
The four test slices switched from ``"Pane.prototype.X = function"``
to ``"\n X("`` in commits 3 + 4 — that's brittle against the
deferred PR-B/C/D/E/F modernization (IIFE / module wrap shifts
indent to 4 spaces, breaks all four slices silently with a bare
``ValueError``). The new helper uses ``re.MULTILINE`` + ``\s{2,}``
to match the method header at any leading-whitespace depth and
``assert``s on miss so a renamed method fails loudly at the
pinning slice instead of further downstream.
Replaces 8 ``body.index("\n X(")`` call pairs across the 4 anchored
tests (replayHistory ×3, appendToolOutput ×1).
Tests: 27/27 ``tests/test_app_js.py`` green. No other suites touched.
Fourth and final commit of the ES6-class refactor (~/pane-class-refactor.md).
Migrates the remaining 5 prototype methods into the class block,
dissolving the last 4 `var self = this` workarounds, and updating the
appendToolOutput-anchored test in lockstep.
Methods migrated INTO the class body:
showInlineToolBlock(items, autoApproved, judgePending)
resolveApproval(approved, always, feedback, skipPost)
appendToolOutput(callId, name, output, isError)
sendMessage()
cancelGeneration()
Two of these (`showInlineToolBlock`, `resolveApproval`) had multi-line
header decls; the conversion script joins their arg lines back into
a single-line class-method header.
Test anchor update (`tests/test_app_js.py`):
body.index("Pane.prototype.appendToolOutput = function")
→ body.index("\n appendToolOutput(")
body.index("Pane.prototype.", start + 10)
→ body.index("\n sendMessage(", start)
The new upper-bound anchors on the next class method's header (which,
by the source-file order preserved through the refactor, is
`sendMessage`). The slice's inner assertions
(tryParseMcpError-before-renderToolOutput offset comparison) are
untouched — only the outer anchor pattern changes.
Final state:
* `class Pane { ... }`: 1 declaration with 39 members
(constructor + 38 methods)
* `Pane.prototype.X = function`: 0 occurrences (was 38)
* `var self = this`: 0 occurrences (was 16)
* Arrow callbacks (`=>`): 55 (was 0)
* 3 module-level helpers (_buildWatchResultBubble,
_buildDefaultReminderBubble, _buildOutputWarningEl) cluster
immediately after the class block.
The framing goal — "coord speaks a more modern JavaScript than
interactive" — collapses on this axis: Pane is now ES6 class shape
with arrow-function callbacks and `this`-lexical inner scopes, on par
with coord's ES6+ idioms. The remaining var → const/let sweep and
template-literal pass are deferred to follow-up PRs B-F per §8 of the
refactor brief.
Tests: 27/27 `tests/test_app_js.py` + 258 broader (renderer + console
suites) green. All 4 historically anchored test slices now use
class-method anchors and pass cleanly.
Third of the four planned commits in ~/pane-class-refactor.md.
Migrates the two test-anchored history-rebuild methods into the
class block, updates the three pytest assertions that sliced them
by `Pane.prototype.X = function` literal, and relocates the last
nested module-level helper to live alongside the other two.
Methods migrated INTO the class body:
replayHistory(messages) — 304-line method, the largest single
method in the file. Dissolves
2 of the remaining `var self = this`
sites (the method-scope one + the
inner-callback one inside the
`tool` role branch's
replayAdvisoriesAfterTool callback).
_attachRetryToLastAssistant() — small leaf method that the
replayHistory tests use as the
lower-bound sentinel for their
slice.
Helper relocated to just after the class block:
_buildOutputWarningEl(assessment) — was nested between
replayHistory's `};` and
`_attachRetryToLastAssistant`'s
header. Joins the two helpers
that already moved in commit 1
(_buildWatchResultBubble,
_buildDefaultReminderBubble) —
all three module-level helpers
now cluster immediately after
the class.
Test anchor updates (`tests/test_app_js.py`):
body.index("Pane.prototype.replayHistory = function")
→ body.index("\n replayHistory(")
body.index("Pane.prototype._attachRetryToLastAssistant", start)
→ body.index("\n _attachRetryToLastAssistant(", start)
Three assertions touched: `test_replay_history_renders_content_before_tool_block`,
`test_replay_history_renders_persisted_verdict_badge`,
`test_replay_renders_user_interjection_advisory_after_tool_block`.
The slice's inner assertions (`msg.content`-vs-`msg.tool_calls` offset
ordering, `renderVerdictBadge` regex, `replayAdvisoriesAfterTool` +
`addUserMessage` substring matches) are untouched — only the outer
anchor pattern changes.
After this commit 5 prototype declarations remain: showInlineToolBlock,
resolveApproval, appendToolOutput, sendMessage, cancelGeneration —
all five migrate in commit 4 alongside the appendToolOutput test
anchor update.
Tests: 27/27 `tests/test_app_js.py` green. Class block now spans
lines 12 → 1634; all 3 module-level helpers cluster at 1642 / 1676 /
1698 just after.
Second of the four planned commits in ~/pane-class-refactor.md.
Migrates the callback-heavy methods that don't anchor any test
slice, dissolving 12 of the 16 `var self = this` workarounds into
arrow-function lexical-this along the way.
Migrated INTO the class body:
_createDOM, connectSSE, handleEvent,
_addUserMsgActions, _addRetryAction, _retryLast,
_rewindToMessage, _startEdit, _editAndResend
Each migration applies the same mechanical transformation:
* `Pane.prototype.X = function (args) {` header → class-method
`X(args) {` form, body re-indented +2 spaces.
* Every `var self = this;` declaration removed.
* Every inner `function (...) {` callback rewritten as
`(...) => {` — arrow functions inherit `this` lexically, so the
outer-self capture pattern dissolves without behavioural change.
* Every `\bself\b` identifier rewritten as `this`.
* Closer `};` → `}` (no semicolon on class methods).
`sendMessage` and `cancelGeneration` are deliberately deferred to
commit 4 even though they're shape-eligible for this commit: they sit
AFTER `appendToolOutput` in the file, and
`test_phase8_appendtooloutput_dispatches_mcp_error_before_renderer`
slices `appendToolOutput` by looking for the next `Pane.prototype.`
declaration as an upper bound. Migrating sendMessage + cancelGeneration
now would leave `appendToolOutput` as the LAST prototype declaration
in the file, breaking that slice. Co-migrating all three in commit 4
keeps every intermediate commit green.
Spike-verified (§2.10.5): all 16 `var self = this` sites in app.js
are Type 1 (outer-`this` capture only — no event-target-`this`, no
delayed semantic capture). Mechanical conversion is safe for every
site touched here.
Diff: 805 insertions / 816 deletions (net −11 lines) — the arrow
form is more compact than `function (args) {`, partly offsetting the
class-body indent overhead.
Tests: 27/27 `tests/test_app_js.py` + 238 broader (renderer + console
suites) green. Anchored tests (replayHistory ×3, appendToolOutput ×1)
remain on their existing `Pane.prototype.X = function` literals —
their methods migrate in commits 3 + 4.
Opens the ES6 modernisation of ui/static/app.js (per the spike in
~/pane-class-refactor.md §2.10). This first commit lays the class
scaffolding and migrates the 22 callback-free leaf methods — the
remaining 16 callback-heavy / test-anchored methods land in commits
2-4.
Migrated INTO the class body:
constructor(wsId) (was `function Pane(wsId)` at line 12)
reset, updateWsName, disconnectSSE, setBusy,
showEmptyState, removeEmptyState,
addThinkingIndicator, removeThinkingIndicator,
addSystemNudgeMarker, addUserReminder, addToolReminder,
addUserMessage, getFeedback, appendToolOutputChunk,
showOutputWarning, updateVerdictBadge, updateVerdictGlow,
addInfoMessage, addErrorMessage, updateStatus,
isNearBottom, scrollToBottom
Two module-level helpers (`_buildWatchResultBubble`,
`_buildDefaultReminderBubble`) were previously nested between leaf
methods. Class bodies can't hold free function declarations, so they
relocate to immediately after the class closing `}`. Function
declarations are module-hoisted so the relocation is semantically
free.
The other 16 prototype methods continue as `Pane.prototype.X =
function (...)` below the class — they each add to `Pane.prototype`
exactly as before, so the prototype shape is unchanged.
Spike-verified guardrails:
* Hoisting: only `new Pane()` site is `createPane` at line 2152
(renumbered), well after the new class block ends at line 458.
* Strict mode: `app.js` is already clean of `with`,
`arguments.caller`, `arguments.callee` — class bodies' implicit
strict mode is a no-op.
* Method enumerability: the 14 `for (var pid in panes)` loops
iterate the module-level `panes` ID-map, not method names on an
instance — class methods being non-enumerable on the prototype
doesn't affect them.
No `var self = this` sites are touched in this commit — the 22 leaf
methods all have zero inner callbacks. Commits 2-4 will dissolve the
16 var-self-this sites as their parent methods migrate.
Tests: 27/27 in `tests/test_app_js.py` green (test anchors on
`replayHistory` and `appendToolOutput` are untouched — commits 3 + 4
will co-migrate them with the assertion updates).
Adds ``console/static/app.js`` to ``_UNSAFE_CODE_SINK_LINT_TARGETS``.
The parametrized scan now covers 8 static JS bundles — all admin-side
bundles are clean.
Docstring updates:
- Fully qualify the bundle paths in both posture lists
(``ui/static/app.js``, ``shared_static/utils.js``, etc.) so the two
``app.js`` files are unambiguous now that both are in the targets
list.
- ``console/static/app.js`` joins the **strict DOM-construction**
list (alongside the interactive surface + shared helpers + coord
chat entry) — the cluster-dashboard renderer is now full
createElement / textContent construction, no HTML strings ever
interpolated.
- ``console/static/admin.js`` + ``console/static/governance.js``
remain in the **sink-free string-concat** list — they retain the
escapeHtml + concat builder shape with the unsafe sink off the call
site.
The ``insertAdjacent`` carve-out note (verdict-badge writers in
``ui/static/app.js``) is now qualified to avoid ambiguity.
Initial AST-light swap routed 12 ``innerHTML =`` sites through
``setSafeHtml`` (8 sites) or ``replaceChildren()`` (4 empty-string
clears). The /review pipeline flagged the node-table render path as
a hot loop where DOMParser-per-row costs scale with cluster size (the
project's 100-node design ceiling × per-RAF render frame on SSE
churn), and the static colHeaders literal was being re-parsed every
render.
This commit lifts the entire console-dashboard renderer to true
``createElement`` + ``textContent`` + ``append`` construction —
matching the strict-posture lane used by ``ui/static/app.js`` rather
than the sink-free string-concat posture admin/governance use.
Net result: **zero** ``innerHTML`` and **zero** ``setSafeHtml`` calls
remain in ``console/static/app.js``.
Three new module-local helpers carry the heavy structural fragments:
- ``buildColHeaders()`` returns a DocumentFragment with the 7-span
column-header layout used at the top of the table and inside each
multi-node group body. Built once via createElement so the static
literal isn't re-parsed every render — and the duplicated literal
between the top-table and per-group sites collapses into one helper.
- ``_buildNodeNumCell(value, highlighted, cellClass)`` — the
``<span class="X num [has-value]">N</span>`` shape used 8× across
buildNodeRow and the group header.
- ``_buildHealthCell(cellClass, healthPct, healthFillClass)`` — the
health-bar trailing cell used both per-row and per-group.
The 4 ``setSafeHtml`` sites that remained after the initial sweep
(error-state placeholders + state-pill builder) also flip to
``createElement`` / ``makeEmptyState`` for consistency with the rest
of the file's new posture. ``makeEmptyState`` is the helper added
during the interactive cleanup (#532).
Two Copilot findings, both pre-existing on main since 2026-04-05 but
preserved by this PR's mechanical refactor. Addressing them here
since they're appropriately in scope (the renderJudgeSettings
function is the focus of the refactor) and Copilot ranked them high.
- **Dead ``shortKey === "model"`` branch removed.** The loop at the
top of ``renderJudgeSettings`` does ``if (s.key === "judge.model")
continue;`` because ``judge.model`` is rendered by the cross-cutting
model-alias picker in ``admin.js`` (line 4895, ``aliasKey: "judge.
model"`` registry entry). No other setting key has the form
``judge.X`` where ``X === "model"``, so the conditional branch was
provably dead — Copilot's confusion ("operators won't see a model
picker") is the same confusion a future reader would hit. Drop
the branch + the corresponding ``SELECT``-dispatch in the binding
loop (no SELECT inputs remain in this renderer).
- **Float-input value/min/max now escapeHtml'd.** Previously
``currentVal`` and ``s.min_value`` / ``s.max_value`` were
interpolated raw into ``value="..."``, ``min="..."``, ``max="..."``
attributes. Numbers stringify safely, but
``admin_list_judge_settings`` can fall back to returning a raw
stored string when deserialization fails (Copilot's flag) — a
non-numeric fallback containing a ``"`` would break out of the
attribute boundary. Wrap with ``escapeHtml(String(...))`` for
defense-in-depth + null-safety.
Adds ``console/static/governance.js`` to ``_DOM_WRITE_LINT_TARGETS``.
The parametrized scan now covers 7 static JS bundles. Docstring
updated: governance.js joins admin.js in the "sink-free string-concat"
posture; ``console/static/app.js`` (cluster dashboard / node table) is
the last admin-side bundle still pending — same posture once cleaned.
Same posture sweep as the admin DOM cleanup (#533) applied to
turnstone/console/static/governance.js:
1. **46 innerHTML sites → setSafeHtml**. Mechanical swap via the
same AST-light Python walker the admin PR used. HTML strings are
still built with escapeHtml + concat — same defence as before, just
no innerHTML sink at the call site. `node --check` clean; prettier
formatted.
2. **6 inline event handlers in renderJudgeSettings refactored to
delegated bindings**. The previous code embedded the setting key
as a JS-string inside an HTML attribute
(`onclick="saveJudgeSettingFromInput('KEY')"`), the same footgun
addressed in admin.js: escapeHtml turns `'` into `'`, but the
HTML parser decodes that before the JS parser runs, so a key with
an apostrophe would escape the JS string. Keys today come from a
static judge-settings registry without apostrophes, so no live
vuln — but the pattern is brittle.
Inputs now carry a single `data-judge-key`; the binding loop
dispatches on `this.type === "checkbox"` / `this.tagName ===
"SELECT"` to wire change-listeners for the auto-save inputs, and
leaves text/number/password inputs alone (they commit via the
adjacent Save button). Save and Reset buttons carry their own
`data-judge-save-key` / `data-judge-reset-key` and bind via click
delegation — mirrors the admin.js settings-tab shape.
3. **1 inline handler in audit pagination converted** for
consistency: `<button onclick="loadMoreAudit()">` → addEventListener
after setSafeHtml.
4. **4 stale safety-narrator comments stripped** ("// values escaped
via escapeHtml above", "// NOTE: innerHTML usage below is safe",
etc.). The safety now lives in setSafeHtml; the per-call-site
narration is tombstone-shaped and removed per the project
no-tombstone-comments convention.
5. **`saveJudgeSettingFromInput` hardened**. The
`document.querySelector('[data-judge-key="' + key + '"]')` lookup
now wraps the key in `cssEscape` (from shared/utils.js) so a
future key containing `"` or `\` doesn't break the selector.
console/static/app.js (12 sites) remains pending — same posture once
cleaned, separate PR.
Two small follow-ups from the Copilot PR review:
- ``admin.js:3112-3115`` — the docstring on ``_onSettingChange`` said
``inp`` is "passed in by the delegated handler", but the wiring at
the call site is a per-element ``addEventListener`` rather than a
single delegated handler on the container. Reword to
"per-input event-listener callback" so a future reader doesn't
read "delegated handler" and refactor under that mistaken premise.
- ``test_app_js.py`` — ``_DOM_WRITE_LINT_TARGETS`` constant name
was missed in the earlier sweep that renamed ``_UNSAFE_DOM_WRITE_RE``
→ ``_UNSAFE_CODE_SINK_RE`` and the test to
``test_no_unsafe_code_sinks_in_static_assets``. Rename to
``_UNSAFE_CODE_SINK_LINT_TARGETS`` so the three names align.
Extend ``_UNSAFE_DOM_WRITE_RE`` to also flag the JS string-to-code
constructors that share the same XSS / RCE-on-injection threat model
as innerHTML:
- ``eval(...)`` — string-eval
- ``new Function(...)`` — dynamic-Function constructor
- ``setTimeout(string, ...)`` / ``setInterval(string, ...)`` — the
string-first-arg form (function-first-arg remains unflagged)
Verified that none of these sinks exist in the six currently-scanned
files (interactive app.js + shared utils/auth/kb + coord chat entry +
console admin.js). Pre-existing parametrized test
``test_no_unsafe_dom_writes_in_static_assets`` extends naturally to
the broader pattern; all 25 cases pass.
``insertAdjacent`` + HTML continues to be excluded — two existing
verdict-badge sites in ui/static/app.js consume
``renderVerdictBadge``'s HTML-string output, so broadening that
specific sink first needs the upstream helper cleaned.
The settings panel renderer in admin.js previously emitted inline
``onclick``/``onkeydown``/``oninput``/``onchange`` attributes that
embedded the setting key as a JS-string inside an HTML-attribute
context:
'<button ... onclick="_saveSettingValue(\'' + escapedKey + '\')">'
``escapeHtml`` escapes apostrophes to ``'``, but the HTML parser
decodes that *before* the JS parser runs — so a key containing an
apostrophe would break out of the JS string. Today the keys come
from a static settings registry without apostrophes, but the pattern
is brittle: a future maintainer adding operator-controlled values to
the attribute would discover the footgun the hard way.
Every inline handler in admin.js is now replaced with a delegated
``addEventListener`` set up after ``setSafeHtml(container, html)``.
Handlers read their context off ``data-*`` attributes (which
``setAttribute`` correctly escapes), so the HTML-attribute /
JS-string double-context is eliminated.
Touched renderers:
- ``_renderSettings`` — section headers, help buttons, per-key inputs
(input/change), per-key save + reset buttons
- ``_renderNodeMetadata`` — section headers (delete/add buttons
already used delegation)
- MCP install source selector — radio-change handler for
``_updateInstallFields``
The named handlers (``_saveSettingValue``, ``_toggleSettingsSection``,
etc.) are unchanged in signature and behaviour; only their wiring
moved from inline-attribute to ``addEventListener``.
Two changes to ``tests/test_app_js.py``'s DOM-write lint:
1. Add ``console/static/admin.js`` to the scan target list. Now
covers all 6 static JS bundles that render LLM output, tool
results, operator-supplied data, or user input. ``governance.js``
and ``console/static/app.js`` remain pending follow-ups (will land
as separate cleanup PRs).
2. Parametrize the lint test over the target list. Each file is now
its own pytest case (e.g.
``test_no_unsafe_dom_writes_in_static_assets[turnstone/console/static/admin.js]``),
so a failure attributes precisely to the offending file instead of
masking offenders behind the first-file's assertion.
Rename the test from ``..._in_interactive_assets`` to
``..._in_static_assets`` — the coverage now spans more than the
interactive surface, and the surface-neutral name leaves room for
governance.js / console-app.js without another rename.
3. Restructure the docstring to surface the two distinct postures
(strict DOM-construction surfaces vs. sink-free string-concat
admin.js) up front, instead of burying the admin caveat after the
main contract claim.
48 sites in turnstone/console/static/admin.js previously assigned
HTML strings directly to .innerHTML. Every site is now routed
through the shared setSafeHtml helper (added in the interactive
cleanup PR), which parses the trusted HTML via DOMParser and installs
the result via replaceChildren — no innerHTML sink at the call site.
Two distinct postures across the admin pages:
- 47 sites: ``setSafeHtml(el, html_built_with_escapeHtml)`` — admin
builders construct HTML strings via string concatenation, running
every interpolated value through escapeHtml first. Defence still
depends on escapeHtml at the builder; the lint catches the sink but
cannot catch a missing escape. Full DOM-construction rewrites
(createElement + textContent) would be structurally safer but are
out of scope — 136 escapeHtml call sites + several thousand lines
of builder code is a separate effort.
- 1 site: ``srcEl.replaceChildren()`` for the MCP-install package
panel's empty-state branch — equivalent to the old
``srcEl.innerHTML = ""`` clear, slightly more idiomatic.
No user-visible behaviour change. DOMParser parses the same HTML the
prior innerHTML assignment did; the new DOM is identical, and the
container.querySelectorAll("[data-X]") event-binding pattern still
finds the newly-installed nodes the same way it did before.
console/static/governance.js (46 sites) and console/static/app.js
(12 sites) remain pending follow-ups. The verdict-badge writers'
two insertAdjacentHTML sites in ui/static/app.js still need the
upstream helper cleaned first — separate effort.
The pyjwt 2.12.1 advisory (\"weak encryption\") is disputed by the
supplier — the key length is the calling application's
responsibility, not the library's. Turnstone generates its JWT
signing keys via the standard ``secrets`` module at
operator-controlled strength (see ``turnstone/core/auth.py``), so the
advisory does not apply to this codebase.
No fix version is available — pyjwt 2.12.1 is the current PyPI
latest as of 2026-05-21. Adding ``--ignore-vuln PYSEC-2025-183``
with the rationale documented in-line so a future reviewer can
re-evaluate when an upstream fix or a non-disputed re-issue lands.
The advisory was published between main's last CI pass (2026-05-19)
and the interactive-cleanup PR's CI run (2026-05-21); main's
security job will fail next push without this fix.
Two Copilot-suggested improvements to the regression scan:
- Allow optional ``+`` before ``=`` in the regex so a future
regression that switches sinks from ``el.innerHTML = X`` to
``el.innerHTML += X`` is still caught. The trailing ``(?!=)``
negative-lookahead still excludes ``===`` / ``==`` reads.
- Switch the scan from line-by-line ``splitlines()`` iteration to a
whole-body ``finditer`` so ``\\s*`` can span newlines. Multi-line
sinks like ``el.innerHTML\\n = X`` (an artifact of formatter
line-wrapping at the assignment) are now caught. Match positions
map back to line numbers for the failure message.
Verified locally with representative test cases including
``el.innerHTML += X``, ``el.outerHTML += X``, multi-line variants,
and the ``===`` / ``==`` reads that must remain unflagged.
Adds two regression tests in tests/test_app_js.py:
- test_no_unsafe_dom_writes_in_interactive_assets: whole-file scan
for inner/outerHTML assignment + doc-write sinks across all five
interactive surfaces (app.js, shared utils/auth/kb, coord chat).
Includes line + content in the failure message so a regression
fails loudly with location.
- test_shared_utils_defines_set_markdown_helper: pins setMarkdown's
signature and the DOMParser path so a refactor that drops the
parser (e.g. swap to Range.createContextualFragment) forces an
explicit reviewer decision.
The lint regex is tightened with a negative-lookahead so equality
comparisons (``===`` / ``==``) don't false-positive, and broadened
to cover ``outerHTML`` and the legacy doc-write sink in addition to
``innerHTML``. ``insertAdjacent`` + HTML is *not* covered yet — two
existing verdict-badge sites (app.js:1287, 1538) consume the HTML
output of renderVerdictBadge and would need that helper cleaned
first.
Three adjacent sites that all assign pre-trusted HTML strings (built
from escapeHtml + static template literals, no caller-supplied raw
HTML) get the same DOMParser + replaceChildren treatment via the
shared setSafeHtml helper:
- coordinator.js:327 (appendMsg's body) — callers pass either
esc(text) or renderToolOutput(...) output, both pre-escaped.
- auth.js:302 (login overlay) — _buildLoginHTML() returns a static
template with no caller-supplied interpolation.
- kb.js:30 (keyboard-help overlay) — html is built from a static
keys-and-bindings table.
Eliminates the only remaining innerHTML site in coord and the two
shared-overlay sites. Console admin / governance JS bundles
(106 sites in console/static/{app,admin,governance}.js) remain
outside this PR — different threat model (admin-only behind auth
gate), separate effort.
Routes every direct-HTML assignment in turnstone/ui/static/app.js
through the helpers added in the previous commit, or through native
DOM construction (createElement + textContent + append /
replaceChildren). Net result: zero ``.innerHTML =`` sites in app.js.
Breakdown of the 26 sites:
- 2 renderer-output sites (replayHistory message body, plan-inline
body) now use setMarkdown — DOMParser keeps the audit at zero
innerHTML sites, stricter than the prior centralise-not-eliminate
plan.
- 8 empty-string clears (pane reset, layout rebuild, dashboard
refresh, etc.) become replaceChildren().
- 5 keyboard-shortcut button labels (y/n/a/Esc) collapse onto
makeKeyLabel(hint, label).
- 5 dashboard placeholders (Loading / Failed / No active workstreams)
use makeEmptyState(text).
- 6 escapeHtml-interpolated HTML strings (command preview, judge
evidence, dashboard state cells, footer node, diff lines) become
createElement + textContent + append; escapeHtml drops out because
textContent escapes intrinsically.
No user-visible behaviour change — DOMParser parses the same HTML
the prior innerHTML assignment did, and DOM construction with
textContent produces equivalent rendered output. Mermaid / hljs
post-render scope is unchanged (now scoped to the body element
rather than the wrapper for the two setMarkdown sites; both
contain the same code blocks).
Adds four helpers that move the unsafe HTML-string sinks off the call
site:
- setSafeHtml(el, html): parses a trusted HTML string via DOMParser
and installs the result via replaceChildren — no innerHTML.
- setMarkdown(el, content): renderMarkdown -> setSafeHtml ->
postRenderMarkdown (hljs + mermaid).
- makeEmptyState(text): builds a <div class="dashboard-empty"> card.
- makeKeyLabel(hint, label): keyboard-hint + label fragment for
approve/deny/always/amend/reject buttons.
The helpers are unused at this point; subsequent commits route the
26 app.js sites, coord:327, auth.js, and kb.js through them.
* feat(admin): align turnstone-admin DB config with server (config.toml + env)
turnstone-admin previously read TURNSTONE_DB_* env vars only, forcing
operators with credentials in config.toml to re-export them just to
run admin commands. Wire add_config_arg + apply_config(["database"])
into main() so admin honors the same precedence as turnstone-server:
CLI / config.toml [database] > TURNSTONE_DB_* env > hardcoded defaults.
Also exposes pool_size + sslmode/sslrootcert/sslcert/sslkey to admin,
which previously dropped any such config silently.
Hardening: load_config() now warns once when config.toml is group- or
world-readable, since DB password and TLS key paths live in [database].
Tests cover precedence (default / config / env / partial fallback /
empty-string-in-config-beats-env), the real init_storage boundary on
a tmp sqlite path, the sys.argv -> main() pre-parser path, and the
new permission check (mode 0644 warns, 0600 quiet).
* test(admin): unify config import style in test_admin_db_config
Use module alias (config_mod.apply_config) instead of mixing
'import turnstone.core.config as config_mod' with 'from
turnstone.core.config import apply_config'. Addresses
github-code-quality bot feedback on PR #531.
Async LLM-tier "llm_fallback" verdicts (judge.py:1073, judge.py:1131
via _deliver_fallbacks) deliberately reuse the heuristic verdict's
``verdict_id`` so the row gets "upgraded in place" from heuristic →
llm_fallback when the LLM judge times out, is cancelled, or returns
no content. The consumer ``_persist_intent_verdict`` was doing a
plain INSERT via ``create_intent_verdict``, hitting the
``intent_verdicts_pkey`` constraint on every llm_fallback delivery.
Postgres logged the duplicate-key error; the application try/except
swallowed it at log.debug — so the row never actually got upgraded
and the LLM judge's annotation ("(LLM judge did not return a
verdict)") was lost.
The collision rate exploded on stable/1.5 smoke tests because
PR #527 (just merged) added two new heuristic-INSERT paths in the
auto-approve early-return branches of ``approve_tools`` — previously
those branches dropped heuristic verdicts on the floor, leaving no
row for the fallback to collide with.
Fix:
- New ``upsert_intent_verdict`` method on the storage protocol +
sqlite + postgres impls, using dialect-specific
``insert(...).on_conflict_do_update(index_elements=["verdict_id"],
set_={...})``. Set_ clause updates ONLY the three fields that
genuinely change between heuristic and llm_fallback: ``tier``,
``reasoning``, ``judge_model``.
- Every other column is excluded from set_: identity columns
(verdict_id, ws_id, call_id, func_name, func_args), carried-
verbatim columns (intent_summary, risk_level, confidence,
recommendation, evidence, latency_ms), and ``user_decision``.
- ``user_decision`` exclusion is load-bearing: ``IntentVerdict
.to_dict()`` doesn't project it, so a fallback verdict reaching
``_persist_intent_verdict`` carries the kwarg's ``"pending"``
default. If the operator already resolved the approval between
heuristic INSERT and fallback delivery, the row's user_decision
has been stamped to ``"approved"``/``"denied"``/``"timeout"`` (or
an auto-approve reason at heuristic-INSERT time per PR #527).
Including ``user_decision`` in set_ would silently clobber that
back to ``"pending"``.
- ``_persist_intent_verdict`` switched from ``create_*`` to
``upsert_*``. Bulk path ``create_intent_verdicts_bulk`` stays as
plain INSERT — every heuristic ``verdict_id`` is freshly minted
in ``judge.evaluate`` so in-turn dups can't happen. The inverse
race (daemon-judge verdict lands BEFORE the bulk write) IS
reachable today but its observable behavior is unchanged by the
per-row UPSERT switch; documented at the bulk site for a future
hardening pass.
Test coverage:
- TestIntentVerdictUpsert × 4 — fresh-id insert, conflict-upgrade,
user_decision preservation across heuristic→approved→fallback,
identity + carried-field preservation.
- Existing tests in test_session_ui_base.py updated to mock the
new upsert method instead of create_intent_verdict.
Two dead module-level constants flagged by github-code-quality on
PR #529: ``_INSPECT_MSG_SNIP_THRESHOLD`` and
``_INSPECT_TOOL_ARG_SNIP_THRESHOLD`` lost their callers when the
content-snip logic moved into the ``_snip_head_tail`` helper. The
helper now reads ``head + tail + _INSPECT_ELISION_MARGIN`` so the
"reserve bytes for the elision marker" rationale that the dead
constants documented stays named instead of becoming a bare ``64``.
A coord doing a fan-out wave of inspect_workstream calls against
tool-heavy children could blow the context budget on raw output
alone (one child with a 100 KB bash result × N children). The
previous safety net was ``_truncate_output``'s head+tail strategy,
which silently drops *middle* messages — exactly the wrong shape
for a coordinator trying to understand a child's trajectory (the
LAST message tells the model what the child concluded; the FIRST
sets the brief; the middle is the connective tissue).
Three-tier degradation modeled on the search tool's pattern at
``session.py:_format_search_results``:
Tier 1 (full): every message verbatim — used when size fits.
Tier 2 (compact): per-message head/tail-snipped content (600/300
chars) plus snipped ``tool_calls.arguments``
(300/100 chars). When content snipping alone
doesn't fit, fall through a message-list trim
ladder ((20,30) → (10,20) → (5,10)) that keeps
head + tail messages and elides the middle as
``{"_omitted": N}``.
Tier 3 (skeleton): no messages — counts + role distribution +
verdicts-by-risk + last assistant preview.
Budget 32 KB (matches ``_SEARCH_OUTPUT_BUDGET``). First emission
whose JSON serialization fits the budget wins. ``_tier`` lands on
every non-error emission so the coordinator LLM and audit readers
can see which compression rung was selected; ``_tier_note`` carries
actionable advice (re-call with a smaller ``message_limit`` etc.).
Error-shape results bypass tiering — they're already small.
Bug fixes caught during review:
- ``_compact_message`` now preserves the assistant-side ``tool_calls``
list with snipped ``function.arguments``; the pre-fix shape left
audit readers with tool-result orphans against invisible calls.
- The intermediate Tier-2 list-trim ladder fixes a size-monotonicity
bug where Tier-2 with un-snippable content (per-message body
under the 964-char threshold) plus the added ``_tier_note`` came
out STRICTLY larger than Tier-1, falling through to skeleton
when a head+tail trim would have preserved dozens of messages.
- ``_inspect_skeleton`` reads ``result["skill_id"]`` (production
storage row key) with a ``skill`` fallback; pre-fix it read
``skill`` only and emitted ``null`` for every real workstream.
Three Copilot threads from PR #526:
1. ``_exec_spawn_workstream`` success path emitted
``{"child_ws_id": null}`` when the upstream response unexpectedly
omitted ``ws_id`` (200-shape with no error field, no id field).
Adds the missing guard — mirrors ``_exec_spawn_batch`` which
already surfaces ``"spawn returned no ws_id"`` as a denied row.
The LLM now sees a tool error and can retry instead of chasing
a null id through follow-up tools.
2. ``docs/coordinator-skills.md`` UI render note said "keep the
ws_id as the click-through key" in a paragraph that had just
introduced ``child_ws_id`` — readable as "the ws_id value" but
confusable as a field-name claim. Clarifies that the value
class is the same regardless of which key carried it.
3. ``docs/bulk-endpoints.md`` ``spawn_batch`` example shows
``child_ws_id`` (coord-tool output shape). The doc title and
the "model tool" column label already disambiguate it from HTTP
API responses, but a reader landing at the example section
directly could miss the framing. Adds one explicit sentence.
Coordinator LLMs on large fan-outs recency-bias on seeing `ws_id`
in a `spawn_workstream` / `spawn_batch` return -- calling
`spawn_workstream(ws_id=...)` again instead of progressing to
`wait_for_workstream(ws_ids=[...])`. On 10+ child fan-outs this
cascades into self-inflicted re-spawn loops.
Rename to `child_ws_id` (already an existing project term -- see
`tasks` tool, `child_event_bus.py`) defuses the recency bias.
Scope is the LLM-facing JSON only -- the server HTTP API at the
spawn endpoint still returns `ws_id`, and the internal reads of
that HTTP response are unchanged.
Also updates the two tool descriptions, the operator-facing skill
doc, and the bulk-endpoints example so docs don't undo the rename.
A skill with `allowed_tools=[]` in the coordinator's `list_skills`
response read as "no tools are usable by this skill" to a model that
didn't know the semantics — but the actual meaning is "no tools are
pre-approved for auto-approval (auto-approve exemption list)". Real
misdiagnosis incident: a code-review child appeared to have been
spawned with zero tool access when in fact the skill simply hadn't
declared an auto-approve allowlist.
Two-part fix:
- `coordinator_client.list_skills` omits the `allowed_tools` key from
the per-skill dict when empty. Absence now carries the unambiguous
meaning "no tool is pre-approved for this skill"; presence (with a
non-empty list) keeps the standard Claude Code skill-spec shape.
- `turnstone/tools/list_skills.json` description rewrites the field
doc so the LLM sees: "tool names exempt from the operator approval
gate ... the field is OMITTED when empty: a skill without
`allowed_tools` still has access to every tool in its session's
toolset; absence of the field means no tool is pre-approved for
this skill, not that the skill has no tools."
Field name stays `allowed_tools` — matches the upstream Claude Code
skill frontmatter (`allowed-tools` hyphenated, stored as
`allowed_tools` internally per `skill_parser.py:241-242`). Parser,
storage column, admin UI, and SDK unchanged.
Auto-approved tool calls left intent_verdict rows with `user_decision=""`,
indistinguishable from rows still pending manual review. Real misdiagnosis
incident: a coord with `recommendation="review"` and `user_decision=""` was
read as "stuck waiting for approval" when in fact the tools had been
auto-approved and the child was running normally.
New vocabulary at the storage API boundary (column server_default stays
`""` so pre-fix legacy rows are still distinguishable as such):
- `pending` — at insert, before any resolution
- `approved` / `denied` — manual user resolution
- `timeout` — approval-event timeout (split from `denied` so the
audit column alone tells them apart; the feedback
string used to carry this distinction)
- `policy` / `blanket` / `skill` / `always` / `auto_approve_tools` —
auto-approve reasons (mirror `AutoApproveReason`)
Heuristic verdicts on the two auto-approve early-return branches are now
persisted with `user_decision=<reason>` (previously dropped on the floor).
Late LLM verdicts for already-auto-approved call_ids look up the reason via
a TTL-pruned `_auto_approve_reasons` map (lazy 60s prune at write time, so
no fixed cap can silently regress the fix on the N+1th auto-approve; LLM-
disabled sessions don't leak entries because prune fires whenever auto-
approves happen).
Bug fixes caught during review:
- `on_intent_verdict` early-returns when the verdict already carries an
auto_reason — without this, a manual `resolve_approval` on a mixed batch
would overwrite the auto-stamped row with `approved`/`denied`.
- `_record_auto_approves` runs BEFORE `_persist_auto_approved_heuristic_*`
so the lookup map is populated before any concurrent LLM verdict can
fire and miss it.
- `resolve_approval(timeout=True, approved=True)` now raises ValueError
to make the split-brain shape unrepresentable.
- Approval-timeout feedback string derives from `_APPROVAL_WAIT_TIMEOUT`
rather than the hardcoded "1 hour".
The previous comment described "\\" as "non-default", which is
backwards — "\\" is the SQL standard escape character. The
actually-non-default part is SQLAlchemy's ``.like()`` itself: it
defaults to no escape character, so ``escape_like``'s output is only
interpreted correctly when callers pass ``escape=LIKE_ESCAPE``
explicitly. Reword to put the caller-side requirement first.
WatchRunner._poll_watch committed active=False to the row BEFORE
calling _dispatch_result for a terminal fire, and the dispatch closure
registered by ChatSession.set_watch_runner enqueued each reminder with
a valid_until=is_watch_active predicate that re-read the row at drain
time. Since the runner already flipped active to 0, the predicate
returned False for every dispatched fire and NudgeQueue.drain silently
dropped the entry — the model never saw a watch result. Then a
subsequent action=cancel call hit list_watches_for_ws (filters
active==1), the now-inactive row was invisible, and the cancel
returned 'Watch "X" not found.' regardless of whether the watch had
actually run.
Reorder _poll_watch to dispatch before the row write, drop the
valid_until predicate from the watch closure (its only effect was the
bug above), and add a _terminal_dispatched guard on the runner so a
transient storage failure between dispatch and row-write doesn't
re-fire the reminder on the next tick. Add WatchRunner.forget_terminal_dispatched
and call it from the cancel path so an out-of-band deactivate (next_poll='')
doesn't leak the watch_id from the runner's pending-retry set indefinitely.
Cancel-by-name now routes through a new find_watch_by_name storage
method that ignores the active filter and prefers active rows over
newer-inactive same-name siblings. The session.py cancel branch
distinguishes 'already completed (auto-cancelled)' from 'not found'
so the model can tell apart 'this watch ran and finished' from
'no such watch.' Consolidate the two byte-identical _escape_like
/ _escape_ilike helpers in the storage backends into a single
turnstone.core.storage._utils.escape_like and apply it to the new
find_watch_by_name LIKE pattern so a model-supplied watch name
containing % or _ can't redirect a cancel to a sibling watch.
NudgeQueue.drain previously dropped predicate-failed entries without
logging anything, which is what hid this bug for so long. Drain now
emits nudge_queue.predicate_dropped: info for reason=predicate_false
(the normal lifecycle case — idle_children when every active child
finished between enqueue and drain), warning with exc_info for
reason=predicate_raised (a misbehaving predicate).
Tests: new test_poll_watch_terminal_fire_survives_drain (parametrized
stop_on_fired + max_polls_reached) drives the real WatchRunner._poll_watch
against a real tmp_db row and confirmed to fail against pristine main.
test_poll_watch_retry_deactivate_after_update_watch_failure exercises
the _terminal_dispatched retry-deactivate branch end to end.
test_cancel_clears_pending_terminal_dispatched_entry covers the cancel-
path leak case. test_find_by_name_prefers_active_over_newer_inactive
catches the ordering regression. test_find_by_name_treats_percent_as_literal
+ test_find_by_name_treats_underscore_as_literal pin the LIKE escape.
The shared_static exclude in scripts/update-vendored-js.sh was meant to
skip self-references inside vendored libraries, but it also hid
shared_static/renderer.js — which loads the vendored libs and pinned
mermaid-11.14.0 across every renovate bump since #426. Tests under
tests/test_web_helpers.py were similarly invisible because the include
list omitted *.py.
Replace the broad shared_static exclude with the specific old-versioned
vendor directory (about to be rm -rf'd next anyway), and add *.py to the
include list. Bump renderer.js to mermaid-11.15.0 to repair the live
404, and refresh the test fixtures to current vendor versions so they
stop drifting.
Line 3446 used double-quote string delimiters with an embedded ">
that terminated the string mid-attribute, leaving "bulk-revoke (" as
bare tokens. The rest of the surrounding block uses single-quote
delimiters; switch the broken line to match so the embedded > and "
sit safely inside the string.
The parse error wiped out every global in admin.js, so showAdmin and
the rest of the admin entry points were undefined — the console was
non-functional whenever an MCP server row had consented_users_count > 0.
Backports OAuth-MCP Phase 9 (#516) to the stable/1.5 track. Introduces
forward-only migrations 054_mcp_pending_consent and
055_mcp_user_tokens_server_index.
* feat(mcp): admin status, deferred-consent persistence, operator docs (Phase 9)
Completes the OAuth-MCP build-out (Phases 0-8 shipped) by closing the
operator + deferred-consent gaps:
1. **Per-(user, server) deferred-consent persistence** — when a
non-interactive run (scheduled / channel) hits ``mcp_consent_required``
or ``mcp_insufficient_scope``, the sync pool dispatchers now upsert a
row into a new ``mcp_pending_consent`` table. The dashboard hydrates
the gear-icon badge from this table on load, so users who weren't
online to see the in-flight SSE prompt still surface the deferred
work on next login. Cleared automatically by the OAuth callback
handler on consent completion; manual user dismiss via new DELETE
endpoints. Composite PK ``(user_id, server_name)`` collapses repeat
occurrences for the same server — no NULLs-not-distinct trap.
2. **Admin status pill + bulk-revoke** — the MCP Servers admin row now
shows ``consented_users_count`` for ``auth_type=oauth_user`` rows
when ≥1, with a two-step-confirm ``bulk-revoke`` button that drops
every user's token for the server via the existing
``delete_mcp_oauth_rows_by_server_name`` primitive. Upstream RFC
7009 revoke is intentionally NOT attempted in bulk (avoids N
upstream HTTP calls per admin click); audit detail records
``upstream_revoke_outcome=bulk_admin_no_upstream``. A "last
refresh" pill (age + outcome) renders on each row, sourced from a
new ``_last_refresh`` dict populated by ``_refresh_server`` on every
call (both manual ``refresh_sync`` and the ``_cb_auto_reconnect``
follow-up).
3. **ClientType.SCHEDULED** added to the prompts module + scheduler
passes it through to ``create_workstream``. ``ChatSession`` now
computes ``_is_interactive_for_consent`` at construction (WEB / CLI
are interactive; CHAT / SCHEDULED are not) and plumbs the flag
through ``call_tool_sync`` / ``read_resource_sync`` /
``get_prompt_sync`` to the three sync dispatchers. The wrap at the
``_is_structured_error`` gate routes consent codes to the new
``_record_pending_consent_best_effort`` helper for non-interactive
callers only; interactive sessions stay on the in-flight SSE path
Phase 8 ships unchanged.
4. **Operator docs** — ``docs/mcp-oauth.md`` (operator guide, parallel
to ``docs/oidc.md``: ``auth_type`` choice, OAuth client setup,
encryption-key rotation, troubleshooting matrix) and
``docs/operations/mcp-oauth-headless.md`` (one-paragraph runbook
per ``feedback_runbook_trust_llm.md``: pre-consent recipe for
scheduled / channel-driven runs).
Schema
- Migration 054_mcp_pending_consent.py — composite PK
``(user_id, server_name)``, ``occurrence_count`` + ``first_seen_at`` /
``last_seen_at`` for recency metadata, ``idx_mcp_pending_consent_user``
for the badge-load query. No FKs (matches the rest of the
oauth_user schema).
- Migration 055_mcp_user_tokens_server_index.py — adds
``idx_mcp_user_tokens_server`` on ``(server_name, expires_at)`` so
the admin pill's ``count_mcp_consented_users_*`` queries don't
full-scan against the leading-``user_id`` composite PK.
- Cross-backend: works on SQLite + PostgreSQL via dialect-specific
``on_conflict_do_update`` (PG ``postgresql.insert`` / SQLite
``sqlalchemy.dialects.sqlite.insert``). No ``NULLS NOT DISTINCT``
needed — the simplified PK eliminates the cross-version trap.
Endpoints
- ``GET /v1/api/mcp/oauth/pending`` — list deferred-consent records for
the authenticated user. Install-level gate via cached
``any_oauth_user_mcp_servers`` short-circuits to ``{pending: 0}`` on
installs with no oauth_user MCP servers — local-auth deployments
exercise zero new storage queries on this path. The gate result is
cached on ``app.state`` with a 60s TTL to spare repeat dashboard
loads.
- ``DELETE /v1/api/mcp/oauth/pending/{server_name}`` — single dismiss.
Returns 204 in both existed-and-deleted and never-existed cases
(no cross-tenant existence leak); audits
``mcp_server.oauth.pending_consent_dismissed`` with
``mode=single`` + ``cleared=0|1`` so a session-hijack attacker
scrubbing breadcrumbs leaves an audit trail.
- ``DELETE /v1/api/mcp/oauth/pending`` — bulk dismiss; audits
``mode=bulk`` + ``cleared=N``.
- ``POST /v1/api/admin/mcp-servers/{name}/bulk-revoke`` — admin
bulk-revoke for the named server's per-user tokens. Requires
``admin.mcp`` permission + 400s when the row isn't ``oauth_user``.
All four registered on both ``turnstone-server`` and
``turnstone-console`` (mirrors the Phase 8 ``/connections`` endpoint
shape).
Performance
- Admin list handler now uses a single ``GROUP BY`` bulk-count query
(``count_mcp_consented_users_grouped_by_server``) wrapped in
``asyncio.to_thread`` rather than N per-row sync DB round-trips
inside the async handler. Skipped entirely when no row is
oauth_user.
Frontend
- ``ui/static/app.js``: ``loadPendingConsents()`` hydrates the
existing ``_pendingConsentServers`` set on dashboard init + after
the user opens the settings modal. Endpoint failures stay silent
— the badge will be re-driven by the next in-flight tool error.
- ``console/static/admin.js``: ``consented_users_count`` pill +
``bulk-revoke`` button on each MCP row (only when ≥1 consented),
two-step confirm matching the existing delete pattern. ``last-
refresh`` age + outcome pill in the per-row status cell, sourced
from the freshest per-node entry in ``status[*].last_refresh_at`` /
``last_refresh_outcome``. CSS for the pills in ``style.css``.
Tests
- ``test_mcp_pending_consent_storage`` — 13 tests covering upsert
idempotency, list ordering, per-user isolation, single/bulk delete,
count-by-server + grouped variant, install-level gate.
- ``test_mcp_pending_consent_dispatch`` — 9 tests, including the
boundary-cross gate per ``feedback_tests_through_boundaries.md``:
drives the real ``call_tool_sync`` → ``_dispatch_pool_sync`` →
``_is_structured_error`` → ``_record_pending_consent_best_effort``
with a mocked classified-lookup so the structural plumb-through is
verified end-to-end. Includes a storage-failure test that pins
the docstring's "envelope unchanged on storage failure" promise.
- ``test_mcp_pending_consent_endpoints`` — 11 tests: install gate,
list-for-self, no-cross-user-leak, single/bulk delete, idempotent
not-found, audit emission on single + bulk + cross-tenant dismiss.
- ``test_chat_session_interactivity_flag`` — 7 tests pinning the
``ClientType`` → ``_is_interactive_for_consent`` mapping against
the module-level ``INTERACTIVE_CONSENT_CLIENT_TYPES`` frozenset.
- ``test_mcp_admin_bulk_revoke`` — 7 tests covering admin.mcp
permission gate, 404 on missing, 400 on non-oauth_user, 200 with
``rows_deleted`` + ``consented_users_before``, audit row with
``upstream_revoke_outcome=bulk_admin_no_upstream``, cross-server
isolation.
- ``test_mcp_oauth_handlers`` — 2 new callback tests pin the post-
callback ``delete_mcp_pending_consent`` invocation: success-clears
+ storage-failure-still-redirects.
- 636 tests pass on the impacted surface (47 new + Phase 0-8 OAuth-MCP
+ session + prompts + storage admin). ruff + mypy clean.
Hard invariants honored
- Static path byte-identical for ``auth_type ∈ {none, static}`` — the
flag flows only through the pool dispatchers, which only fire when
the row resolves to ``oauth_user``.
- ``asyncio.timeout`` (not ``asyncio.wait_for``) preserved on every
AS / SDK / pool-loop await — no new awaits added to the hot path.
- Install-level gate on the badge endpoint: cached
``any_oauth_user_mcp_servers`` returns False on a row-less
deployment → endpoint short-circuits without touching the pending-
consent table; 60s TTL bounds the staleness window after admin
flips ``auth_type``.
- Operator-actionable codes (key-unknown, url-insecure, *_forbidden)
explicitly filtered out of persistence — they're outside the
user-facing consent badge scope.
- Best-effort write: the structured-error envelope returned to the
agent is identical whether the persistence write succeeds or fails
(storage exception is logged with type name only — no chained
context that could carry an ``httpx.Request`` bearer header).
- No ``exc_info=True`` on any new path that can chain a bearer-bearing
``httpx.Request``.
- Defensive parsing: ``_parse_pending_consent_envelope`` mirrors
``_is_structured_error``'s ``isinstance(decoded, dict)`` guard plus
filters scope tokens through ``is_valid_scope_token`` capped at
``MAX_INSUFFICIENT_SCOPE_REPORTED`` — defense-in-depth even though
production callers already validate upstream.
- Audit events on every dismiss endpoint so a session-control attacker
scrubbing dashboard breadcrumbs still leaves a trail.
Cross-backend
- Tested on SQLite via the conftest backend fixture.
- PostgreSQL path uses ``postgresql.insert(...).on_conflict_do_update``
parallel to the existing ``mcp_user_tokens`` upsert in Phase 3.
Deferred (not Phase 9 blockers)
- Multi-node pool eviction on bulk-revoke: only local-node sessions
would be evicted if we built it, and there's no bulk-by-server
primitive on MCPClientManager today; remote nodes will surface as
a 401 on next dispatch which refreshes through the (now empty)
token row.
- RFC 8693 / Azure OBO ``auth_type=oauth_token_exchange`` — captured
in the design doc as a future architectural direction (~600 LOC +
IdP-side admin work); requires OIDC token capture and per-MCP-server
resource-trust configuration that v1 does not ship.
* docs(mcp): address Copilot review feedback on Phase 9
- Fix misleading admin.js comment that claimed the refresh pill rendered
"<short-relative> <outcome>" — the pill actually renders only the short
age, with outcome reflected via CSS class and tooltip.
- Replace broken feedback_secrets_not_in_env.md repo-root link in
mcp-oauth.md with the inlined rationale (env-borne secrets reachable
via shell tools / os.environ; TOML secrets are not).
Adds notes for the 14 patches cherry-picked to stable/1.5 since 1.5.12:
reactive PG LISTEN/NOTIFY node discovery + event-driven wait_for_workstream,
memory tool audit trail, task_agent skill personas, plus fixes for the
LLM-visible default alias bypass, mermaid streaming parse errors,
proxy-prefixed re-auth, dashboard appbar visibility, and the PG test
backend on the notify dispatcher suite.
Introduces forward-only migration 053_services_notify_trigger.
- Put ``skill`` back in the access-denial list in the tool
description with a clarification — TASK_AGENT_TOOLS does not
include the skill tool, so sub-agents cannot switch personas
mid-task. Removing the disclaimer entirely created an ambiguity
the LLM could misread.
- Minimize the skill_data carried on the approval item dict to
``name`` / ``content`` / ``risk_level`` only. ``get_skill_by_name``
returns the full ~30-column prompt_templates row including
``scan_report``, ``installed_by``, ``source_url`` — none of those
flow through ``_exec_task`` / ``_evaluate_intent``, and they
shouldn't ride along any future audit serializer that reads the
approval item shape.
- Regression test for ``skill=""``, whitespace-only, and ``\t\n``
values — pins the documented "empty value is acceptable" contract
at the ``(args.get("skill") or "").strip()`` chokepoint.
The task_agent tool now accepts an optional ``skill=<name>`` argument
that loads the named skill's content as the sub-agent's persona,
substituting the hardcoded "# Task Agent" identity statement. The
operating-guidance numbered list (one-shot, tool-use over narration,
no follow-up questions) is layered on top of every persona and always
applies — those are sub-agent semantics that a persona should ride on
top of, not replace.
Validation lives in ``_prepare_task`` so the approval surface tells
the operator what they're consenting to: the validated skill dict
(including content) rides on the item dict from prepare to exec to
defeat TOCTOU between consent and execution. An unknown skill
returns a clean error item with a hint pointing at
``skill(action='search')``; a disabled skill returns a distinct error
so the LLM's recovery path can tell "not found" from "quarantined",
mirroring the enabled gate that ``_exec_skill(action='load')`` and
skill-search already apply.
High and critical skills now surface their risk tier on the approval
header (``, risk: critical``) and emit a
``task_agent.high_risk_skill`` warning — same signal ``_load_skills``
emits for session-level skills, so the operator sees the same flag
whether the skill is loaded session-wide or per-call. ``_exec_task``
emits a ``task_agent.skill_invoked`` info log on the skill branch for
forensic traceability — the approval row captures the choice at
consent time, this log captures it at exec time so post-incident
search doesn't have to cross-walk approval and exec tables.
The ``_evaluate_intent`` func_args projection now includes the skill
name — without it, heuristic ``arg_pattern`` rules targeting a risky
persona name on ``task_agent`` silently no-op and the audit row loses
the choice. Mirrors the long-standing ``spawn_workstream``
projection.
Caught by Copilot on PR #514. openSettingsMenu sets _settingsMenu
synchronously, but the menu's keydown handler was registered inside
setTimeout(0). The previous-commit guard in the global keydown
handler returns early when _settingsMenu is set (so dashboard isn't
hidden by Escape over the menu), which created a window where
Escape had no handler at all — the global skipped, the menu's own
listener wasn't ready yet, and the menu got stuck open until the
next interaction.
Attach keydown synchronously; keep mousedown + initial focus in
setTimeout (mousedown to avoid the opening click triggering its own
outside-click close, focus because the menu DOM needs a tick to
settle layout).
Two related changes that surfaced when the user pointed out the proxy's
node-picker pill was unreachable from the proxied dashboard view: the
dashboard overlay was covering the entire appbar.
- Dashboard overlay now starts at top: 48px so the appbar (with the
proxy-injected node picker) stays visible and interactive while the
dashboard is open. showDashboard no longer marks ui-header inert
(tab-bar and split-root still are). The dashboard's role downgrades
from dialog+aria-modal to region — the appbar being reachable above
it would otherwise contradict aria-modal's "ignore everything else"
semantics.
- Gear icon converts from a direct openSettingsPanel() click into a
dropdown menu with two items: "MCP connections" (existing modal) and
"Logout". Reuses the .ws-tab-dropdown shell for visual consistency
with the workstream tab chevron menu and the proxy node-picker.
Logout uses .destructive styling to reduce misclick risk.
Bug fixes caught by the merged code-review pipeline:
- Global Escape handler skips when _settingsMenu is open, otherwise it
fires hideDashboard() before the menu's own handler — wiping the
composer text + staged attachments out from under the user.
- Menu-item click refocuses the trigger before close, so
openSettingsPanel captures the gear (not <body>) as the eventual
return-focus target.
- ArrowUp keyboard cycling uses idx <= 0 ? len - 1 : idx - 1 instead
of (idx - 1 + len) % len so the no-focus case wraps to the last
item rather than the second-to-last. Same fix backported to
showTabDropdown which had the identical modulo bug.
- Position clamps reordered: right-edge override now runs before the
left-edge floor so a menu wider than the viewport still clamps to
mx >= 4 instead of going negative.
- openSettingsMenu caches _settingsMenuTrigger so closeSettingsMenu
can reset ARIA without re-querying the gear by id.
- aria-controls lifecycle wired both ways (set on open, removed on
close).
The LLM was passing ``task_agent(model="default")`` (and the same for
plan_agent) and routing to whichever backend the auto-created
``default`` alias was attached to at boot — flatspark in the verified
case (ws_id 7dde674) — silently bypassing the operator-configured
``model.task_alias`` / ``model.plan_alias`` (gh200).
Root fix:
- ``load_model_registry`` only synthesises the back-compat ``default``
alias when neither DB nor ``[models.*]`` populate the registry. The
shim was only ever meant for single-CLI-model setups; with a multi-
model DB it became a phantom routing target aliasing ``LLM_BASE_URL``.
- ``_render_agent_tool_descriptions`` filters ``default`` out of the
LLM-visible alias list. The English reading of "default" trips the
model into picking it explicitly even when the description tells it
to omit ``model=`` for the per-role default.
Defense-in-depth at the validator chokepoint
(``_validate_agent_model_override``): explicit rejection of
``alias == "default"`` (post-strip) with corrective guidance;
``default`` filtered out of the unknown-alias retry list so an LLM
probing with a bogus alias can't enumerate it back; the no-alternatives
wording is distinguished from the no-registry-configured wording. The
render path also always rewrites tool descriptions instead of returning
early on filter-empty, so a reload that drops the registry to only
``default`` clears stale alias names left over from a prior render.
Previously only the admin-console DELETE route emitted memory.delete
audit rows, so a long-running session whose memory was deleted via
the admin UI had no log trail showing what happened — masking
out-of-band deletes as apparent tool bugs.
The save branch now stamps memory.save (new row) or memory.update
(upsert); the delete branch does a lookup-then-delete-by-id pair so
the audit can record the resolved memory_id and type. All emissions
are best-effort: failures log at debug and swallow so an audit hiccup
never breaks the tool call itself. Reads (get/search/list) remain
un-audited.
Copilot review on #511 flagged that the dispatch chain and the tests
both claimed to be in lockstep with one another, but only the comment
text said so — the parametrize list and the if/elif chain were two
independent hand-maintained copies, and the comments still referenced
the (long-reverted) ``_PROXY_AUTH_LOCAL_HANDLERS`` symbol.
Make the lockstep guarantee real by collapsing both copies onto one
``_PROXY_AUTH_LOCAL_HANDLERS: dict[tuple[str, str], str]`` mapping
``(method, path)`` to handler-name strings. ``proxy_api`` resolves
the name through ``globals()`` at call time so ``patch(...)`` in
tests still observes the override — a dict of function refs would
have captured the originals at module load (which is why the first
attempt at this dispatch broke the tests and got reverted). Test
cases now derive directly from ``_PROXY_AUTH_LOCAL_HANDLERS.items()``,
so adding or removing an entry in the dispatch table flows through
to the parametrize list automatically and the two can't drift.
When the user is on a proxied node page (``/node/{id}/...``) and the
JWT expires, the in-page login modal POSTs to ``/v1/api/auth/login``
which the proxy shim rewrites to ``/node/{id}/v1/api/auth/login``.
Two latent bugs both had to be fixed for the user to be able to
re-authenticate from inside the proxied UI:
1. ``is_public_path`` didn't recognise the ``/node/{id}/`` prefix
over a public path, so the console's ``AuthMiddleware`` 401'd the
login POST before any handler ran. Extended via the existing
``_extract_proxied_path`` helper so a proxied public path stays
public.
2. Even if the path had been public, ``proxy_api`` would have
forwarded the request to the upstream node. The upstream mints
``JWT_AUD_SERVER`` tokens; the console's ``AuthMiddleware``
(expecting ``JWT_AUD_CONSOLE``) would reject those on the next
proxied call, and ``_proxy_post`` drops ``Set-Cookie`` when
forwarding anyway. ``proxy_api`` now dispatches every entry in
``_PROXY_AUTH_LOCAL_PATHS`` (login, logout, setup, refresh,
status, whoami, oidc/authorize, oidc/callback) to the console's
own auth handlers, and short-circuits non-canonical methods on
those paths with 405 instead of letting them slip through with
the service-token fallback.
Tests parametrize across all eight local-dispatch entries so a future
refactor that drops a branch (or routes it through ``_proxy_post``)
fails loudly, plus a no-auth-header reproduction for the original
lockout and a 405 regression guard for the method-mismatch surface.
* fix(renderer): mermaid streaming parser errors + progressive hljs
Live streaming was rendering mermaid diagrams with `Parse error,
got 'PS'` messages — bare `(`, `[`, `{` inside unquoted edge / node
labels re-entered Mermaid's shape parser. Two unrelated streaming-
specific issues in the renderer pile-up here; this commit addresses
both plus a follow-on UX improvement for code highlighting.
## Mermaid label autoquoter
`_normalizeMermaidSource` wraps two label forms that Mermaid rejects
when they contain bare shape-delimiter chars:
1. Edge labels: `|content|` → `|"content"|`
2. Rectangle node labels: `ID[content]` → `ID["content"]`
Shapes whose syntax already nests delimiters — cylinders `[(...)`,
subroutines `[[...]]`, trapezoids `[/.../]` `[\...\]`, circles
`((...))`, hexagons `{{...}}`, diamonds `{...}` — are intentionally
left alone (their inner delimiters are part of the shape syntax;
quoting would corrupt them). Labels already wrapped in `"..."` are
also left alone. The rewrite is idempotent and runs before the
mermaid SVG cache lookup so identical malformed input hits the
cache on re-render rather than re-quoting per tick.
## Markdown fence-pair regex
The old fence regex `/(```+)([^\s`]*)\n([\s\S]*?)\1/g` would, mid-
stream, pair an unclosed ```mermaid open with the OPENING backticks
of a later ```python fence as the "close", handing mermaid a
truncated source. New regex:
/(```+)([^\s`]*)\n((?:(?!\1)[\s\S])*?)\1[ \t]*(?=\n|$)/g
Two constraints close the gap:
- `(?!\1)` inside the content quantifier blocks the lazy matcher
from extending across another N-backtick run. Smaller inner
counts (e.g. 3-backtick inner inside a 4-backtick outer) still
pass since `\1` is the open's actual count.
- `[ \t]*(?=\n|$)` after `\1` forces the close to a line
boundary, so ```python (open with a language tag) can't
masquerade as a previous fence's close.
Together: an unclosed fence stays as plain markdown until its true
close arrives, so neither mermaid nor hljs ever sees a mid-stream
truncated source.
## Progressive hljs
Extracted `postRenderHljs` from `postRenderMarkdown` with a source-
keyed `_hljsCache` (FIFO, cap 64, keyed on `language:source`) and
wired it into `_streamingRenderApply`. Closed code fences are now
syntax-highlighted as they stream in, matching the progressive
mermaid pattern from #426. Per-tick cost stays cheap because the
cache returns the pre-tokenized HTML synchronously on hit; only
unique (language, source) pairs pay `hljs.highlightElement`.
## Internal cleanup from the review pipeline
- `_cacheFifoEntry(cache, key, value, max)` replaces the duplicated
`_cacheHljsEntry` and `_cacheMermaidEntry`. Single tested
implementation across four caches (hljs, mermaid svg, mermaid
error, mermaid normalize memo). The "don't evict on overwrite"
invariant is pinned per-cache in tests.
- `_mermaidNormalizeCache` memoizes raw textContent → normalized
output so the per-rAF-tick autoquoter split + regex doesn't
repeat for unchanged diagrams. Eviction shares
`_MERMAID_CACHE_MAX` with the SVG cache it feeds.
## Tests
The fake DOM in tests/test_renderer_js.py grew a few capabilities
to drive these paths:
- `classList` is now array-like (length + indexed access) so the
hljs language-extraction loop works.
- `textContent` setter mirrors the real-DOM side effect of
entity-escaping into innerHTML, so `escapeHtml()` round-trips
(otherwise every `renderMarkdown` returns empty `<p>` tags).
- `querySelectorAll` handles both `pre code.language-mermaid`
and `pre code[class*='language-']`.
Added: 6 fence-pairing regression cases, 9 hljs-progressive cases
(cache hit / distinct sources / language separation / NO_HIGHLIGHT
langs / terminal class / eviction / overwrite / postRenderMarkdown
wraps hljs / _streamingRenderApply invokes hljs), 11 autoquoter
cases including both diagram sources from the live screenshot
encoded verbatim as parametrized regressions, and 3 normalize-memo
cases (populates on first call, consulted before normalize via
sentinel pre-seed, distinct sources cache separately).
Total: 104 renderer tests pass (was 67).
* fix(renderer): apply Copilot review feedback on #510
Two doc / harness adjustments from the PR review — no behavior
change in production code.
- The `_mermaidNormalizeCache` comment claimed eviction "stays in
lockstep with the SVG cache". That was misleading: the two
caches key on different things (raw textContent vs normalized
source) and evict independently. Updated the comment to describe
what they actually share (the cap, for memory footprint) and
what they don't (positional coupling), and to note that the memo
deliberately survives `_initMermaid` since normalize output is
theme-independent.
- The fake DOM in tests/test_renderer_js.py had `innerHTML` setter
clear `children` but leave `_textContent` intact, so subsequent
`textContent` reads could return stale data after an innerHTML
mutation (real DOM invalidates textContent on innerHTML write).
No current test triggered this, but it would mask future bugs
that depend on innerHTML/textContent consistency. Setter now
clears `_textContent`; the children-derived fallback in the
getter returns `''` after the wholesale replace.
All 104 renderer tests still pass; ruff + mypy clean.
CI's postgres-backend run failed 11 of the new notify tests from #505.
Three independent issues:
1. Migration 053's ``services_notify`` trigger lives only in the
alembic chain, but the test fixture in conftest.py calls
``init_storage(..., run_migrations=False)`` for speed. That path
skips migrations and relies on ``metadata.create_all`` for the
table tree. Previous alembic-only DDL (migrations 041 / 048
``CREATE INDEX CONCURRENTLY`` on workstreams) is performance-only,
so tests never depended on it. 053's trigger is the first
behaviorally-required alembic-only DDL in the project — without it
``register_service`` doesn't fire NOTIFY and the trigger-filter
tests time out.
Fix: declare the trigger function + trigger in ``_schema.py`` and
attach them via ``sa.event.listen(services, "after_create", ...)``
DDL events, gated on ``dialect == "postgresql"``. The same SQL
constants are imported by migration 053 so there's a single source
of truth. Test fixture stays unchanged — ``create_all`` now
installs the trigger on fresh PG test DBs. Migration covers the
upgrade-on-existing-DB path; the two are mutually exclusive given
``create_tables = not run_migrations`` in ``init_storage``.
2. NotifyDispatcher tests fired ``storage.notify(...)`` immediately
after ``d.start()`` and hit a race: the listener thread is
concurrently calling ``psycopg.connect(listen_url)`` + ``LISTEN
<channel>`` over the network, so the notify can land before any
session is listening on the channel and PG drops it (pg_notify
only routes to sessions LISTEN'ing at COMMIT time).
Fix: dispatcher gains a ``_listener_ready: threading.Event`` set
inside ``_listener_loop`` after each successful ``storage.listen``
open and cleared on disconnect, plus a public
``wait_until_ready(timeout)`` method. Tests use a new
``_start_ready(d)`` helper that calls ``start()`` + asserts ready.
Production callers don't need this (real reactive traffic arrives
well after startup), but it's the right primitive for any future
"start dispatcher, immediately send" call site too.
3. ``TestSqliteNotify`` is misnamed — its tests run against whichever
backend the ``storage`` fixture provides (PG by default in CI).
Two of its assertions were SQLite-specific:
``assert got.pid == 0`` only holds for the synthetic in-process
path (PG carries real backend PIDs), and
``test_synthetic_sweep_emits_after_interval`` is fundamentally
SQLite-only (no sweep on the PG path).
Fix: drop the pid assertion (channel + payload are the
backend-agnostic invariants), add an ``_is_sqlite`` fixture mirror
of ``_is_postgres``, and gate the sweep test on it. The sweep
test also moves from monkey-patching ``stream._sweep_interval`` to
passing the ``sweep_interval`` kwarg that ``SQLiteBackend.listen``
now accepts (from the earlier Copilot review fix).
Validated locally against a fresh ``turnstone_test`` PG DB: 263
storage + console + notify tests pass on PG, 257 on SQLite, mypy +
ruff clean.
Retire two polling patterns in coord that have clean event sources.
PR 2 of 3 in the coord-completion stack; sits on top of PR #505
(reactive node discovery via PG LISTEN/NOTIFY).
`wait_for_workstream` (coord's block-wait tool) polled storage every
0.5 s in a worker thread regardless of whether anything had changed —
a 600 s wait incurred ~2400 round-trips. Now subscribes to a new
in-process `ChildEventBus` (`turnstone/core/child_event_bus.py`) and
blocks on `threading.Event.wait(min(remaining, WAIT_HEARTBEAT_INTERVAL))`:
- `CoordinatorAdapter` owns the bus; `_dispatch_child_event` calls
`bus.notify(child_ws_id)` after each `_enqueue_on_ui` for the
state-class branch (cluster_state, ws_closed, ws_rename,
intent_verdict, approval_resolved, approve_request).
- Wait loop clears the Event BEFORE the storage snapshot to close
the subscribe/check race; a notify between clear and the next
`wait()` leaves the Event set so the loop re-reads without
losing the wake-up.
- 2 s heartbeat cap preserves the existing `wait_progress` SSE
cadence for the sidebar UI while cutting SSE traffic ~4x vs the
pre-bus 500 ms cadence in the quiescent case.
- Worst-case completion latency is 2 s (vs pre-bus 0.5 s) because
`set_state` buffers non-ERROR writes through `StateWriter`
(async-flushed) while `emit_state` fans out immediately — a
bus-driven wake can beat the flusher and read pre-transition
state, then re-block until heartbeat. Deliberate trade-off; the
SSE-traffic reduction outweighs the regression on the most
common terminal transition.
- Defense-in-depth: ownership-filter `cleaned` to own-subtree
before `register_waiter` so a foreign ws_id passed by an
untrusted coord LLM (prompt injection) can't observe wake-up
timing as a side channel. Predicate (`_row_in_own_subtree`)
requires both `parent_ws_id == coord_ws_id` AND `user_id ==
coord_user_id` parity — same gate strength as the existing
`_is_own_subtree` mutating-op guard, so a corrupted /
cross-tenant `parent_ws_id` alone can't satisfy it. Shared with
`_snapshot_all` so the snapshot's `denied` shape stays in
lockstep with the bus filter (Copilot review on #506).
Coord idle-cleanup thread polled the storage scan every
`check_every` seconds (~30 s on default 2 h timeout) even when no
coord was anywhere near idle. Now subscribes to
`SessionManager._state_subscribers` with a `tick_now` event and
blocks on `tick_now.wait(check_every)` — any state change wakes
the sweeper without waiting a full interval, AND the timeout still
fires the periodic sweep for the DB-orphan-only case. A
`min_sweep_interval=5 s` floor bounds DB-call traffic at ~0.2/s
under sustained activity so the loop can't tight-spin `close_idle`
at the rate of its own DB latency (6x improvement over the
pre-refactor fixed 30 s cadence under any activity, and prompt
state-change-driven wakes when below the floor).
`CoordinatorClient` constructor takes `child_event_bus` as a
required kwarg — there's no external SDK shape to preserve and
keeping it optional would silently mask a wiring bug in any future
caller. Tests construct their own `ChildEventBus()` per fixture.
Tests: 16 unit tests for `ChildEventBus` (register / unregister
symmetry, multi-waiter fan-out, multi-child waiter, subscribe/check
race, concurrent register / notify smoke); 7 new adapter tests
(bus notify fires for all 6 state-class events, drops for unknown
child / wrong ws_id); 7 new coord-client wait tests (subscribe-
after-terminal, notify wakes, unrelated notify doesn't wake,
heartbeat fires without notify, unregister on exit, multi-waiter
independence, cross-tenant denial via the user_id-parity filter);
8 idle-cleanup tests (initial sweep, heartbeat cadence, exception
swallowing, stop_event clean exit, state-change wake, subscriber
cleanup, mid-sweep wake, `min_sweep_interval` floor). All pass;
ruff + mypy clean. Full non-live suite: 6227 passed (+2 vs prior
baseline).
* feat(console): reactive node discovery via PG LISTEN/NOTIFY dispatcher
Add a console-side `NotifyDispatcher` that holds a dedicated PostgreSQL
`LISTEN` connection and fans wake-ups out to per-channel handlers on a
separate dispatch thread. Cluster collector subscribes to a new
`services` channel and runs node discovery reactively — new-node /
graceful-deregister visibility drops from up-to-60 s to ~500 ms on
Postgres, with the 60 s discovery loop retained as the backstop for
crash-shaped node loss (NOTIFY only fires on real writes).
Storage layer gains a uniform `notify` / `listen` API:
- PostgreSQL: real `pg_notify` / `LISTEN` on a dedicated session-mode
connection that bypasses pgbouncer (mandatory: pgbouncer is required
in transaction-pool mode per docs, which is incompatible with LISTEN).
- SQLite: in-process fan-out + synthetic-sweep fallback so consumer
code is identical across backends.
`TURNSTONE_DB_LISTEN_URL` (or `[database] listen_url` in config.toml)
points the dispatcher's connection direct-to-Postgres. Defaults to the
main DB URL when unset.
Migration 053 installs the `services_notify` trigger; it filters
heartbeat-only UPDATEs in-trigger so the 30 s × N-nodes heartbeat tick
stays quiet, while INSERT, DELETE, and url/metadata-changing UPDATE
still fire.
Dispatcher detail:
- Two threads: listener (drains stream → bounded queue) and dispatch
(invokes handlers under exception suppression). Same-channel notifies
coalesce per dispatch batch so an N-node deploy burst is one
`_discover_nodes` per channel.
- Reconnect uses exponential backoff (1 s → 30 s cap). After any
successful reopen — whether the prior failure was a stream-poll error
or a connect / initial-LISTEN error — one synthetic Notify with
payload="reconcile" is enqueued per channel so handlers re-read on
the same code path they use for real events.
Future consumers (ConfigStore live reload, scheduler immediate
dispatch, audit live-tail) plug in by adding their channel to the
dispatcher's construction list.
Tests: 22 dispatcher tests (incl. reconnect + coalescing under stub
storage), 7 SQLite notify-stream tests, 4 PG-gated trigger-filter
tests, 4 collector wire-in tests. All pass; ruff + mypy clean.
* fix(notify): address Copilot review on #505
- _sqlite.py: SQLiteBackend.listen() now de-dupes channel names via
dict.fromkeys before constructing the stream — duplicates would
otherwise register the queue twice and double-deliver each notify.
- _sqlite.py: SQLiteBackend.listen() gains a keyword-only sweep_interval
parameter (defaults to _SQLITE_NOTIFY_SWEEP_INTERVAL) — matches what
the comment at the constant already promised, and lets future
consumers without their own polling timer pick a tighter cadence
without reaching into private stream attributes.
- _sqlite.py: documented the `except queue.Empty: pass` end-of-drain
termination so it's not mistaken for swallowing an unexpected error.
- _postgresql.py: docstring referenced :func:`_pg_listen_url` which
was renamed to _resolve_pg_listen_url during PR development.
- notify_dispatcher.py: module docstring referenced a non-existent
_bootstrap_console_subsystem; wire-in is at console/server.py::main.
Refuted (no change, false positives from github-code-quality bot):
- 4× "Statement has no effect" on Protocol-method `...` ellipsis bodies
(idiomatic Python Protocol declaration, not dead code).
- 2× "Mixed import style" in tests — `import ... as nd_mod` is
intentional to allow attribute assignment for monkey-patching the
module's `_RECONNECT_BACKOFF_INITIAL` constant inside try/finally.
Converts [Unreleased] to [1.5.0] and adds individual sections for
1.5.1 – 1.5.12. Covers: MCP OAuth 2.1 + PKCE (Phases 1–8), OIDC
hardening, metacog NudgeQueue + wake trigger, SSE refresh-resume,
reasoning persistence (Phases 1–4), structured watch-result cards,
skills unlock, inline child approvals, Stage 3 Children primitive
lift, coordinator composer parity, node capability auto-detection,
progressive mermaid rendering, and the full schema migration list
for each release.
Updates the track list to stable/1.4, stable/1.5, and main.
Editing the first message in a workstream sends /rewind N where N is the
total user turns, leaving session.messages empty. The handler guarded the
history event with `if history:`, so only clear_ui was emitted. The
frontend dispatches the queued edit-and-resend from the history event
handler (app.js _pendingEditSend), so an empty history orphaned the
pending text and left the composer stuck in busy.
replayHistory already handles the empty case via showEmptyState(), so
emitting the event unconditionally is safe and unblocks the dispatch.
A bare ``httpx.ReadTimeout`` previously surfaced as ``ReadTimeout: timed
out`` — no provider, no base URL, no model — leaving the user with no
signal to tell whether a model server hung, the URL was wrong, or the
model isn't loaded on the backend.
``ChatSession._format_backend_error`` now rewrites known boundary
exceptions (httpx ``ReadTimeout`` / ``ConnectError`` / etc. and OpenAI /
Anthropic SDK ``APITimeoutError`` / ``APIConnectionError`` /
``NotFoundError`` / ``AuthenticationError`` / ``RateLimitError``) into
operator-actionable text that names the provider, base URL (query
string stripped before ``sanitize_error_text`` redacts credentials),
and model. Matching is by class name so the helper carries no SDK
imports. Unrecognised exceptions fall through to the legacy
``f"{type(exc).__name__}: {exc}"`` shape, preserving existing grep
targets.
The Anthropic call sites in session.py passed the operator-side
`replay_reasoning_to_model` flag through without checking the
model's static `supports_reasoning_replay` capability. The OpenAI
Responses path AND-gated both flags in `_build_kwargs` so a model
without a reasoning lane (gpt-4o, etc.) silently skipped replay even
when the operator flag was set. The Anthropic path had no such gate.
For all current Claude entries this was a no-op asymmetry - every
`_ANTHROPIC_CAPABILITIES` row sets `supports_reasoning_replay=True`,
so `True AND op == op`. But:
- The capability flag was dead code on the Anthropic path
- A future Claude entry (or any Anthropic-shaped surface) shipping
with the cap left at its False default would have replay fire
anyway, against the cap declaration
- The asymmetry made `supports_reasoning_replay` an unreliable
signal - readers couldn't tell if it gated anything per-provider
Move the AND-gate into `_resolve_replay_reasoning_to_model` via a
new optional `caps=` kwarg. When caps is provided, the resolver
returns `operator_on AND caps.supports_reasoning_replay`; when
omitted (back-compat for any caller not yet updated), it returns
the operator flag unchanged.
Thread caps through the three call sites: `_utility_completion`
(non-streaming), `_try_stream` (streaming, hoisted resolution out
of the retry loop since caps are attempt-invariant), and the
agent `_api_call` closure in `_run_agent`.
With the AND-gate now living at the session resolver, the redundant
in-provider gate in `OpenAIResponsesProvider._build_kwargs` is
removed. The provider now trusts the resolved bool it receives,
matching the AnthropicProvider shape and giving the cap a single
source of truth across providers. The two provider-level tests
that pinned the in-provider gate
(`test_include_omitted_when_capability_false`,
`test_include_omitted_by_default`) drop out; the session-level
boundary test
`TestSessionToOpenAIResponsesBoundaryIntegration::test_capability_false_omits_include_even_when_flag_true`
already covers the same end-to-end invariant.
Tests added:
- 4 resolver-level tests pinning the AND-gate semantics +
back-compat when caps is omitted
- 1 wire-boundary integration test mirroring the OpenAI Responses
`test_capability_false_omits_include_even_when_flag_true` -
drives session._try_stream through the real AnthropicProvider
with operator flag True + capability False and asserts the
thinking block does NOT reach the SDK boundary
Existing `TestUtilityCompletionPassesFlag` test had its caps mock
upgraded from `SimpleNamespace` to a real `ModelCapabilities`
instance to satisfy the new attribute read and stay robust to
future capability fields.
Copilot review feedback on #500. The original
``list_available_models`` had an implicit cs=None branch where the
placeholder still advertised ``registry.default`` (filtered against
enabled rows) when ``app.state.config_store`` was None but
``coord_registry`` was bound — useful in the rare degraded state
where lifespan wired the registry but the ConfigStore failed to
initialise. The PR #500 refactor accidentally dropped that branch:
the helper requires a config_store, so the cs=None case fell out as
"blank coordinator default".
Add an explicit ``elif coord_registry is not None`` branch that
mirrors the helper's tier 3 with the placeholder's enabled-rows
filter applied. New test exercises this path by passing
``config_store=False`` to the test fixture.
Previously /v1/api/models (home composer placeholder) and
console/session_factory.py walked separate two-/three-tier chains for
the coordinator alias. session_factory was missing the
``model.default_alias`` tier, so admins who set the system default in
the Models tab would see it advertised but new coordinator sessions
would silently keep launching on ``registry.default``.
This commit:
- Extracts the chain into ``turnstone/console/coordinator_alias.py``.
``resolve_coordinator_alias`` returns the effective alias under a
shared three-tier policy: explicit pin → ``model.default_alias`` →
``registry.default``. Tier 2 is validated against
``registry.has_alias`` and falls through to tier 3 with a logged
warning if unknown. Tier 1 is intentionally passed through
unvalidated so an explicit operator pin surfaces as 503 at
``registry.resolve`` rather than being silently swapped out.
- Wires both call sites through the helper. The placeholder supplies
an ``alias_filter`` that restricts every tier to enabled DB rows so
the home composer never advertises a model the workstream picker
can't actually offer; the session factory uses no filter (matches
prior 503-on-typo behaviour for explicit pins).
- Adds direct integration tests for the session factory's chain
(``tests/test_console_session_factory.py``) and updates the
placeholder tests' fixture to provide a stub coord_registry, since
the helper now requires one.
Light-review followup on 389400c8.
The "mirrors session_factory.py:109-110" claim was inaccurate —
session_factory's chain is two tiers (coordinator.model_alias →
registry.default) and skips model.default_alias entirely. The
placeholder handler extends that chain with model.default_alias as
tier 2 so admins who set the default in the Models tab see it
advertised in the home composer. Comment now lists the three tiers
explicitly and flags the session_factory-vs-placeholder drift case
(where model.default_alias ≠ registry.default) as a separate issue
to track.
Also lifts the ``from types import SimpleNamespace`` import in the
test fixture to module level — minor readability cleanup.
Two Copilot-review followups on /v1/api/models default resolution.
- console/server.py: coordinator_default_alias now mirrors the full
fallback chain in console/session_factory.py:109-110 — explicit
coordinator.model_alias → model.default_alias → registry.default.
The registry tier was missing, so the home composer placeholder went
blank whenever an operator never set model.default_alias in the admin
UI even though new coordinator sessions still launch on
registry.default (loaded from config.toml [model].default by
load_model_registry). Two new tests cover the registry-default
branch and the disabled-alias guard.
- console/static/app.js: _resolveModelLabel returns "" (not the bare
alias) when the alias isn't found in the dropdown's model list, so
callers can rely on the documented "fall back to neutral placeholder"
contract. Matches the existing doc comment.
Bundles the click-around polish on the console admin UX.
Home composer + schedule modals
- /v1/api/models now exposes coordinator_default_alias + judge_default_alias,
resolved through the same chain console/session_factory.py uses. Both the
home composer's MODEL / JUDGE MODEL placeholders and the schedule create /
edit modal model placeholders rewrite to "Default — alias (model)" once
the API responds. The `models_changed` SSE refresh keeps placeholders
current as operators edit per-role assignments.
- Composer.setOptionPlaceholder added so callers can update just the first
option's text without disturbing the rest of the choice list.
Admin → Models → Roles
- Channel adapter row added (channels.default_model_alias) — the migration
to the Roles sub-tab missed it. Key added to
_MODEL_AFFECTING_SETTING_KEYS so edits fire the SSE refresh, and to the
settings-tab roleKeys skip-list so it only renders in one place.
- Plan/Task agent rows now display "(inherit)" instead of the misleading
"(default — <alias>)" — those roles cascade through plan_model →
agent_model → session model, not a single concrete default.
- coordinator.reasoning_effort accepts "" (inherit), matching
model.plan_effort / model.task_effort.
- Blank options in each role's MODEL select now match the "alias (model)"
shape used by the other rows.
Toggle-switch component
- New .toggle-switch component (visually-hidden native checkbox + styled
track + label). 40×22 hit target meets WCAG 2.5.5 (AAA), inset ring on
the off state for ≥1.5:1 contrast against the modal surface.
- .toggle-stack groups toggles in a column with .toggle-group-divider for
conceptual grouping (used in the Add Model modal between "Active" and the
paired Reasoning toggles).
- .toggle--flush modifier zeroes the default top margin for toggles that
sit flush against a heading or a dynamically-rendered row.
Sweep — every admin-modal boolean checkbox is now a toggle:
schedule (cs/es-autoapprove, es-enabled), policy (ep/epp-enabled),
tool-mode (ctm/etm-default), skill (csk/esk-auto-approve, csk/esk-enabled),
MCP (mcp-auto-approve, mcp-enabled), Add Model (Active, surface-persisted-
reasoning, replay-reasoning), judge bool settings (cancel_on_approval et
al.), and the user-roles-modal role assignment list. The two
ogp-cred / eogp-cred inline credential checkboxes stay as compact inline
boxes since they sit beside text inputs in tight horizontal rows.
Add Model modal — the "Enabled" toggle promoted to "Active" and moved to
the very top of the form. Tooltip explains it gates dropdown visibility
without removing the definition.
MCP authorization — the three radio buttons replaced with a vertical
.segmented-control option list. Selected row paints --accent-dim plus a
filled .segmented-indicator; focus ring uses --accent so it stays visible
on the currently-selected option.
Role permissions modal — the 19 permission checkboxes are now
.toggle-switch.perm-toggle (monospace lowercase identifiers preserved).
The permissions are split into Scopes / Admin / Workstreams & Tools
sections under caps-styled section headers so the row-flow grid no longer
slices `admin.*` mid-column.
Judge bool toggles use a static "Enabled" caption rather than flipping
text on `.checked`; flipping lagged 50–300 ms behind the slider position
because the caption was sourced from the post-save reload.
CSS cleanup — dead `.admin-checkbox` / `.perm-checkbox` rules removed.
Specificity audit (scripts/css_specificity_audit.py) returns no conflicts
on any new component class.
Tests — 525 pass on the affected slices; new tests/test_console_available_
models.py pins each branch of the resolution chain in /v1/api/models so the
home composer placeholder stays correct as precedence rules evolve.
`IntentJudge.__init__` previously had a 3-way resolution chain: registered
alias → raw model id pinned onto the session provider → session model. The
middle branch was a footgun documented in `console/session_factory.py:130-137`
— pinning the literal `judge.model` string onto the coordinator's session
provider silently broke every verdict whenever that provider didn't recognise
the model id (e.g. coordinator on Anthropic, `judge.model = "gpt-5-mini"` →
uniform `llm_fallback`).
Tightens to alias-only, matching `coordinator.model_alias` /
`model.plan_alias` / `model.task_alias`. An unknown `config.model` now logs
a warning and inherits the session model — same path as empty. Help text on
`judge.model` updated to clarify the contract.
Adds two regression tests in `TestModelAliasResolution` covering the
session-model inheritance for unknown values and the empty-model self-
consistency case.
GoogleProvider attaches raw tool_call dicts as ``provider_blocks`` on
the finish chunk for ``thought_signature`` round-trip
(``_google.py:_iter_stream``). When the same turn streamed Gemini's
``reasoning_content`` as ``reasoning_delta`` chunks, the prior
synthesizer bailed out the moment ``provider_blocks`` was non-empty
— so the captured reasoning was visible live but lost on page reload.
Replace the early-return-if-non-empty check with a reasoning-bearing
type test (``thinking`` / ``redacted_thinking`` / ``reasoning`` /
``reasoning_text``). When none of those types appear, append the
synthetic ``reasoning_text`` block to the existing list rather than
replacing it — preserving Google's tool-call fidelity blocks.
Also addresses two doc-accuracy review findings:
- ``LLMProvider.extract_reasoning_text`` docstring no longer claims
OpenAI Chat / Responses are unwired (Phase 3+4 shipped extractors).
- Add the method to the Protocol methods table in
``docs/architecture.md`` (was missing alongside the class diagram).
The earlier all-or-nothing shape check on ``_provider_content`` discarded
every valid Anthropic block in a message the moment a single foreign
block (OpenAI ``reasoning``, Gemini thought parts, the synthetic
``reasoning_text`` from path-3 capture) appeared. In the cross-model
resumption edge case that meant ``server_tool_use`` /
``web_search_tool_result`` blocks lost their ``encrypted_content``
silently, breaking web-search round-trip continuity on subsequent turns.
Replaced with a per-block walk: foreign blocks are dropped individually,
valid blocks ride the verbatim path, and an identity-preserving fast
path reuses the source list reference when nothing was filtered or
stripped (pinned by the ``is`` assertions in test_providers.py).
Also addresses validation-pass review findings:
- Document the single-tier vs three-tier ``surface_persisted_reasoning``
resolution divergence between server.py:_build_history and
session_routes.make_history_handler.
- Document why OpenAIResponsesProvider._convert_messages defaults
``replay_reasoning_to_model=False`` while Anthropic's defaults True.
- Document the ``source`` metadata field on synthetic ``reasoning_text``
blocks as reserved-for-future-use, not dead code.
- Add edge tests for non-dict / missing-type-key blocks in
_provider_content (defensive branches in the per-block walk).
CI test job installs `[test]` extras, which omits `anthropic`. The two
TestSessionToWireBoundaryIntegration cases drive the real
AnthropicProvider.create_streaming, which calls _ensure_anthropic() and
raises ImportError. Match the repo convention (test_channel_discord,
test_channel_slack, test_tls_*) by gating the helper with
pytest.importorskip("anthropic").
PR #498 round-robin review surfaced 5 findings. 4 applied; 1 rejected
with rationale.
Applied
* **Copilot finding 5** (history_decoration.py:341): dispatcher
inspected only ``provider_content[0]['type']``. OpenAI Responses
captures EVERY ``output_item.done`` event into ``provider_blocks``
(not just reasoning) — in practice the order is
``[reasoning, message, ...]`` but the API doesn't guarantee that;
a hypothetical ``[message, reasoning]`` ordering would silently
drop the reasoning under an index-only check. Now walks the list
for the first block whose type is in ``_BLOCK_TYPE_PROVIDER_FACTORY``,
then dispatches the WHOLE list to that provider's extractor. Each
provider's extractor already filters internally by its own block
type, so passing the full list is correct. Regression test added
(``test_dispatcher_scans_past_unrecognized_first_blocks``).
* **Copilot finding 3** (migration 052 docstring): the previous
review-fix wave used sed to rename ``persist_reasoning`` →
``surface_persisted_reasoning`` everywhere, which mangled a
historical reference in the migration docstring ("The earlier name
``surface_persisted_reasoning`` was renamed..."). Restored to
point at the actual pre-rename name (``persist_reasoning``).
* **Copilot finding 4** (sdk/typescript/src/events.ts:26):
``HistoryEvent`` JSDoc still referenced ``persist_reasoning`` —
the sed rename only walked ``turnstone/`` and ``tests/``, missing
the TypeScript SDK. Updated to ``surface_persisted_reasoning``.
Also widened the comment to cover all three reasoning-bearing
block types (Anthropic ``thinking``, OpenAI Responses ``reasoning``,
synthetic ``reasoning_text``) instead of mentioning only Anthropic.
* **github-code-quality finding** (session.py:1120): ``_resolve_server_type``
had a bare ``except Exception: pass``. Replaced with a
``log.debug(..., exc_info=True)`` + explanatory comment. Behaviour
unchanged (still returns ``""`` on any lookup failure); failures
are now observable under DEBUG triage.
Rejected (with rationale)
* **github-code-quality finding** (_protocol.py:265):
``extract_reasoning_text``'s body is ``...`` per ``LLMProvider``
Protocol convention. Every method in the file uses ``...`` (PEP
544 idiomatic Protocol style). Changing only this one to
``raise NotImplementedError`` would be inconsistent with the rest
of the file. CodeQL's "statement has no effect" warning is
technically correct for ``...`` as a standalone expression but
ignores the documented Python Protocol convention. No fix.
Docs sync
* docs/api-reference.md: ``history`` SSE event message-shape table
gains the optional ``reasoning`` field.
* docs/architecture.md: ``ModelCapabilities`` row in the type table
gains ``supports_reasoning_replay``; ``StreamChunk`` and
``CompletionResult`` rows gain the existing ``provider_blocks``
field (was missing pre-PR). New "Per-model reasoning persistence"
subsection under the Models config section, documenting the two
flags + capability gate + three reasoning paths + cross-provider
shape filter.
* docs/settings.md: new "Reasoning persistence (per-model)"
subsection with the two-flag table and capability-gate note.
* docs/diagrams/03-core-engine-classes.puml: ``LLMProvider`` interface
adds ``extract_reasoning_text`` + the new ``replay_reasoning_to_model``
kwarg; ``ModelCapabilities`` class adds ``supports_reasoning_replay``.
PNG regenerated.
Lint + test gate
* ruff check + ruff format clean.
* mypy clean (191 source files).
* pytest -m 'not live' — 6116 passed (3 deselected), +1 net new test
(``test_dispatcher_scans_past_unrecognized_first_blocks``).
Multi-stage /review on the full Phase 1+2+3+4 stack surfaced 9 findings
(0 critical, 3 major, 5 minor, 1 nit, 1 uncertain). All applied.
Major
* perf-1 (session_routes.py:2402): make_history_handler ran sync
storage.load_workstream_config inside async def history on the cold-
workstream path, blocking the event loop on every dashboard /history
request for non-resident workstreams. Every other storage call in
the same handler correctly used asyncio.to_thread. Wrap the sync
call in asyncio.to_thread (preserving the existing try/except so a
DB failure still degrades to the conservative-default branch instead
of bubbling out).
* q-2 (test_reasoning_audit_log_discipline.py): the security-sensitive
test (reasoning text never lands at INFO+ severity) only covered the
4 Phase 1 surfaces. Phase 2 added the strip predicate in
AnthropicProvider._convert_messages and Phase 3 added 3 more code
paths that touch reasoning text — none guarded. Added 4 parallel
tests using the existing capture-and-walk infrastructure:
OpenAIResponsesProvider.extract_reasoning_text,
OpenAIChatCompletionsProvider.extract_reasoning_text,
ChatSession._stream_response (drives the synth-block stamp via a
fake reasoning-emitting stream), AnthropicProvider._convert_messages
with replay_reasoning_to_model=False (drives the Phase 2 strip
predicate).
* q-1 (model_registry.py:42): the persist_reasoning flag name implied
storage-control but actually gates UI rehydration only — operators
flipping it could reasonably expect "stop persisting reasoning" but
storage of reasoning bytes happens in provider_data regardless.
Renamed everywhere to surface_persisted_reasoning: ModelConfig
field, migration 052 column (renaming in-place since 052 is not yet
on main), schema, MODEL_DEFINITION_MUTABLE allowlist, _postgresql.py
+ _sqlite.py CRUD impls, _protocol.py create_model_definition
signature, 3 console_schemas Pydantic models, console/server.py
admin POST + PUT, model_registry row mapper, history_decoration.py
helper parameter, server.py _build_history local var,
session_routes.py make_history_handler local var, sdk/events.py
HistoryEvent docstring, admin.js form id + override pill label,
index.html form input id + UI label + tooltip, coordinator.js (none
needed), and every test that referenced the old field name. The
admin tooltip now reads "Storage of reasoning bytes is unaffected
by this flag — they ride in provider_data regardless" so the
decoupling stays explicit at the operator surface.
Minor
* bug-1 (history_decoration.py:336): dispatcher discriminated on
provider_content[0]["type"] only. Anthropic's redacted_thinking
blocks (sealed by the safety system) can appear before, after, or
interleaved with regular thinking blocks per the API docs. When a
redacted block lands first, the dispatcher returned "" and the UI
silently lost the surrounding thinking text. Registered
"redacted_thinking" as a second key in _BLOCK_TYPE_PROVIDER_FACTORY
pointing at the same AnthropicProvider factory — the existing
extractor's type=="thinking" filter already correctly skips redacted
blocks while walking the full list. Regression test added.
* q-3 (_protocol.py:155): replay_reasoning_to_model defaults split
across 9 sites — operator-side defaults to False (matches DB
server_default), provider-API defaults to True (back-compat with
direct callers). Original "pick False everywhere" fix would have
silently flipped behaviour for any direct provider caller. Instead
documented the intentional bifurcation in the Protocol's
create_streaming docstring.
* q-4+q-5 (_protocol.py:107 + 3 providers): MAX_REASONING_DISPLAY_BYTES
was enforced via Python str slicing which counts code points, not
UTF-8 bytes — 4-byte CJK/emoji glyphs would blow past the byte
ceiling. Renamed to MAX_REASONING_DISPLAY_CHARS to match actual
behaviour. Hoisted the 4-line truncation pattern into a shared
_join_reasoning_with_cap helper in _protocol.py; each provider's
extractor becomes a single line at the tail.
* q-6 (tests/_session_helpers.py): _NullUI + _make_session were
duplicated verbatim between test_session_replay_reasoning.py and
test_session_synth_reasoning_block.py. Hoisted to a shared
tests/_session_helpers.py module (importable, leading underscore so
pytest doesn't try to collect it). test_model_registry.py's
_make_session has a different signature (registry/model_alias args
+ _FakeUI) and is not a candidate for sharing.
Nit
* q-7 (history_decoration.py:286): _make_provider_factory used a
dict-as-cell workaround for closure read-only scope. Replaced with
the more idiomatic nonlocal pattern.
Lint + test gate
* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6115 passed (3 deselected). Net +5 tests
(4 audit-log discipline + 1 redacted_thinking dispatcher).
Refinements vs the dedupe output (caught during sanity rendering
the report)
* perf-1 fix preserved the try/except wrapper. The original "wrap in
to_thread" one-liner would have let an OperationalError bubble out
instead of degrading to the fallback branch.
* q-3 fix explicitly documented the bifurcation rather than
collapsing both sides to False. "Pick False everywhere" would
silently flip back-compat behaviour for direct provider callers.
* q-1 fix included the admin.js:5292 fallback site
(m.persist_reasoning !== false) that the original threaded-change
list missed.
* q-6 fix verified the third _make_session in test_model_registry.py
is structurally different (different signature + different UI
helper) and intentionally NOT a dedupe target.
Wire reasoning capture and (where the API supports it) replay for the
two remaining provider paths. Phase 3 was originally scoped as
"OpenAI Responses + Gemini" but a spike against the OpenAI SDK source
revealed that Gemini routes through the OpenAI-compatible endpoint
(``/v1beta/openai/``), which is structurally identical to vLLM /
llama.cpp / any other Chat-Completions-shaped local model. Phase 3
and Phase 4 collapse into one feature with two distinct sub-paths:
* **Path 2 (OpenAI Responses)** — full capture+replay. ``include=
["reasoning.encrypted_content"]`` on the request makes the API
surface ``encrypted_content`` on reasoning items in
``provider_blocks``; ``_convert_messages`` round-trips them as
``ResponseReasoningItemParam`` input items on subsequent turns.
Verified against the OpenAI Python SDK 2.33.0 source
(``response_reasoning_item.py:31-62``,
``response_reasoning_item_param.py:33-37``,
``response_create_params.py:70-74``). Even with ``store=False``,
``encrypted_content`` round-trips correctly per the SDK's own
documentation.
* **Path 3 (Chat Completions / vLLM / llama.cpp / Gemini-compat)** —
persist-only. Canonical OpenAI Chat Completions has no reasoning
field on the wire, but several local-model servers tack on
``delta.reasoning_content`` as Pydantic extras. ``ChatSession.
_maybe_synth_reasoning_block`` stamps a synthetic ``{type:
"reasoning_text", text, source?}`` block onto ``_provider_content``
at end-of-stream when no native ``provider_blocks`` were emitted but
``reasoning_parts`` accumulated text. The ``source`` field carries
``server_compat.server_type`` (vllm, llama.cpp, sglang, …) for
diagnostic value — informational only, doesn't gate behaviour.
Reasoning text NEVER replays back to the model on this path; it
rides ``_provider_content`` only for ``/history`` UI rehydration
and gets stripped from the wire by the existing
``sanitize_messages`` underscore-prefix strip on every request.
What this change does
* ``ModelCapabilities.supports_reasoning_replay: bool = False`` added
to the dataclass. Set True on every OpenAI reasoning model
(gpt-5* + o-series via the Responses API) and every Anthropic
Claude entry (default + 6 model-specific). Path-2 wire-build does
``replay_active = bool(replay_reasoning_to_model and caps.supports_
reasoning_replay)`` so an operator who flips the flag on a
non-reasoning model (gpt-4o via Responses) silently no-ops rather
than emit a malformed ``include=`` request.
* ``OpenAIResponsesProvider`` gains:
- ``_build_kwargs`` accepts ``replay_reasoning_to_model: bool``
(threaded from ``create_streaming``/``create_completion``);
adds ``include=["reasoning.encrypted_content"]`` when active.
- ``_convert_messages`` accepts the same flag, captures
``_provider_content`` reasoning items pre-sanitization, and
emits them as input items immediately before the assistant
message they belong to. Position is tracked by ASSISTANT
ORDINAL (not raw index) — ``sanitize_messages`` drops orphan
tool results and inserts synthesized error tool messages, but
NEVER drops or duplicates assistant messages, so the n-th
assistant in the original list is invariably the n-th in the
sanitized list. Index-based lookup would have silently
misrouted reasoning attachments after any tool-message repair.
- ``extract_reasoning_text`` walks ``type=="reasoning"`` items and
returns ``summary[*].text`` + ``content[*].text`` concatenation.
- ``_reasoning_item_for_input`` projects a stored item into
``ResponseReasoningItemParam`` shape (drops server-only
``status``). Returns ``None`` when ``id`` is missing or non-
string per the SDK ``Required[str]`` schema; caller skips
appending, preventing malformed input items from reaching the API.
* ``OpenAIChatCompletionsProvider`` gains:
- ``extract_reasoning_text`` walks synthetic
``type=="reasoning_text"`` blocks and returns the concatenated
text directly (no underlying provider semantics — the synth
block IS the surface).
* ``ChatSession`` gains:
- ``_resolve_server_type(alias)`` reads ``server_compat.server_type``
from the active model's capabilities dict.
- ``_maybe_synth_reasoning_block(provider_blocks, reasoning_parts)``
creates the synthetic ``reasoning_text`` block when no native
blocks were emitted but reasoning was captured. Wired at the
end of ``_stream_response`` immediately before the
``_provider_content`` stamp.
* ``history_decoration.py`` dispatcher collapses three near-identical
lazy-init singleton getters (one per recognised block type) into a
single ``_BLOCK_TYPE_PROVIDER_FACTORY`` dict + helper. Adding a
fourth provider becomes a one-line dict entry.
* Constants hoist: ``MAX_REASONING_DISPLAY_BYTES = 64 * 1024`` moved
from three sibling provider modules into ``_protocol.py`` so a
tuning change propagates uniformly to every provider's display path.
Cross-provider safety
The synthetic ``reasoning_text`` block type is intentionally NOT in
``ANTHROPIC_VALID_BLOCK_TYPES`` (Phase 2 constant). Cross-model
resumption (operator switches from a local model to Anthropic mid-
workstream) falls through Phase 2's shape filter cleanly to the
text+tool_calls rebuild path rather than reaching Anthropic with a
malformed block. Pinned by ``test_synthetic_block_falls_through_
anthropic_shape_filter``.
Same protection applies in reverse: OpenAI Responses
``type=="reasoning"`` items reaching Anthropic mid-workstream fail
the shape filter and rebuild from text+tool_calls.
Tests (49 net new tests)
* ``tests/test_provider_openai_responses_reasoning.py`` (21 tests):
- Extractor unit tests: empty/none/no-reasoning/single/mixed/
truncation/malformed/non-list (8).
- ``_reasoning_item_for_input`` projection (4 tests including the
new None-on-missing-id guard).
- ``_build_kwargs`` include= gating: flag+capability/flag-false/
capability-false/default-omits (4).
- ``_convert_messages`` reasoning round-trip: emit-before-assistant/
drop-on-replay-false/foreign-shape-skipped/default-replay-false (5).
* ``tests/test_session_synth_reasoning_block.py`` (23 tests):
- ``_maybe_synth_reasoning_block`` direct unit tests (6).
- Cross-provider safety regression — synthetic block falls through
Anthropic shape filter (2).
- ``OpenAIChatCompletionsProvider.extract_reasoning_text`` for the
new synthetic block type (6).
- ``_resolve_server_type`` direct unit tests (5).
- ``_stream_response`` integration tests driving fake reasoning-
emitting streams through the actual session method (3 tests
— added in response to a code-review finding that pinned the
wire-up at session.py needs an integration test).
* ``tests/test_session_replay_reasoning.py`` extended with 4
``TestSessionToOpenAIResponsesBoundaryIntegration`` tests driving
``session._try_stream`` -> real ``OpenAIResponsesProvider`` ->
captured ``client.responses.create`` SDK boundary call. Negative-
tested: temporarily reverting the ``include=`` step in
``_build_kwargs`` makes ``test_replay_true_adds_include_to_
responses_request`` fail; restoring makes it pass.
* ``tests/test_history_decoration.py`` extended with the new
``reasoning_text`` dispatcher branch test, and the Phase 1 stub
test for the OpenAI Responses dispatcher branch was tightened
(it now asserts real text extraction instead of the empty-string
stub).
* ``tests/test_provider_anthropic_reasoning.py`` had its Phase 1
``OpenAIResponses returns "" for reasoning blocks`` stub test
retitled and updated to assert the real Phase 3 behaviour.
Code-review pass
Multi-stage ``/review`` pipeline (4 finders + verify + dedupe) ran
on this diff. 6 findings (1 major, 3 minor, 2 nit), 0 critical, 0
security, 0 performance. All applied:
* MAJOR (bug-1+bug-4+q-1): ``_convert_messages`` enumerate-index
lookup was unsound under ``sanitize_messages`` length changes.
Fixed by switching to assistant-ordinal-keyed lookup.
* MINOR (q-2+q-3): ``_MAX_REASONING_DISPLAY_BYTES`` duplicated
across three provider modules + declared after first use.
Fixed by hoisting to ``_protocol.py``.
* MINOR (q-4): three near-identical singleton getters in dispatcher.
Fixed by collapsing to ``_BLOCK_TYPE_PROVIDER_FACTORY`` dict.
* MINOR (q-5): ``_maybe_synth_reasoning_block`` wire-up not pinned
by integration test. Fixed by adding three
``TestStreamResponseSynthBlockIntegration`` tests.
* NIT (bug-2): ``_reasoning_item_for_input`` fell back to ``id=""``;
fixed to return ``None`` on missing/non-string id.
* NIT (q-6): four naming variants for the same concept; renamed
``_convert_messages`` kwarg to match the operator-flag name.
* REFUTED (bug-3): SDK distinguishes summary vs content as separate
fields; no double-counting concern.
Briefing departures
The briefing's Phase 3 plan grouped Gemini with OpenAI Responses on
the assumption that Gemini reasoning had its own native shape (like
Anthropic's ``thinking``). The spike confirmed Gemini-via-OpenAI-
compat is path-3 (Chat Completions shape, no native reasoning
items). Phase 3+4 merger handles Gemini for free via the synthetic
``reasoning_text`` block — same mechanism used for vLLM and
llama.cpp. Whether Gemini's specific endpoint actually emits
``reasoning_content`` deltas is server-dependent and not yet
empirically verified; capture is best-effort (server-emission-driven,
no flag gate).
The briefing's Phase 4 plan stamped reasoning as Anthropic-shaped
``thinking`` blocks ``{type: "thinking", thinking: <text>}``. This
PR uses a distinct ``{type: "reasoning_text", text, source?}`` shape
to avoid a cross-model resumption hazard the briefing missed: an
unsigned synthetic Anthropic-shape block reaching Anthropic's wire
would 400 the API. The distinct shape falls through Phase 2's shape
filter cleanly without needing signature validation in the filter.
Lint + test gate
* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6110 passed (3 deselected). Phase 3+4
added 49 net new tests.
Make ``replay_reasoning_to_model=False`` actually suppress prior-turn
thinking blocks on the Anthropic wire (Phase 1 stored the operator
flag but the wire path always re-sent ``_provider_content``
verbatim). As a side benefit, close a pre-existing latent bug where
foreign-shaped ``_provider_content`` (e.g. an OpenAI Responses
``type="reasoning"`` block reaching Anthropic on a mid-workstream
model switch, post-Phase-3) would have 400'd the API.
Why now: Phase 1 shipped the operator knob and UI rehydration but
the wire payload still always carried thinking blocks for
Anthropic-with-thinking turns. Operators flipping replay=False saw
no behaviour change on the actual API call -- the flag only affected
``/history`` rendering. Phase 2 closes that gap.
What this change does
* ``ANTHROPIC_VALID_BLOCK_TYPES`` (frozenset of 8 block types
Anthropic's input boundary accepts) and
``ANTHROPIC_REASONING_BLOCK_TYPES`` (the strip subset) added at
the top of ``_anthropic.py``. The strip set is intentionally
narrow: ``{"thinking", "redacted_thinking"}`` -- ``tool_use`` /
``server_tool_use`` / ``web_search_tool_result`` (which carry
web-search ``encrypted_content``) MUST survive for round-trip
continuity, and a regression test pins this.
* ``_convert_messages`` signature gains
``replay_reasoning_to_model: bool = True`` (back-compat default
-- production call sites pass the resolved value explicitly).
The verbatim ``_provider_content`` replay path is now wrapped by
a shape-validity check using ``ANTHROPIC_VALID_BLOCK_TYPES``;
foreign-shaped payloads fall through to the existing text+
tool_calls rebuild path rather than reaching the API. When
shape is valid AND replay=False, a list comprehension drops
thinking blocks from ``wire_blocks`` while preserving
tool_use / web_search blocks. When all blocks are stripped
(message had only thinking, no text or tool_calls), the message
also falls through to the rebuild path -- which silently skips
if both content and tool_calls are empty (correct: stripped
reasoning has nothing to replay).
* Orphan-tool detection still walks the ORIGINAL ``provider_content``
(not ``wire_blocks``) so the strip cannot accidentally lose the
source-of-truth tool_use IDs. The implementation comment pins
this invariant.
* Protocol surface grows the kwarg on both ``create_streaming`` and
``create_completion``. ``OpenAIChatCompletionsProvider``,
``OpenAIResponsesProvider``, and ``GoogleProvider`` (via
inheritance) accept the kwarg and ignore it -- they have no
first-class reasoning shape on the wire today. Phase 3 will use
it on the OpenAI Responses adapter to gate
``include=["reasoning.encrypted_content"]``.
* ``ChatSession._resolve_replay_reasoning_to_model(alias)`` reads
``ModelConfig.replay_reasoning_to_model`` from the registry,
defaulting to ``False`` on lookup failure (the conservative
miss-fallback: replaying reasoning text against an unknown
operator preference is worse than missing the strip). Threaded
into the three production call sites:
``ChatSession._try_stream`` (streaming), ``_utility_completion``
(title gen / compaction / extraction), and the agent provider
call site (plan / task agents).
Token calibration deferred to Phase 4
The briefing's optional Phase 2 step (extending ``_msg_text_chars``
to count ``_provider_content`` bytes that survive the strip)
required either invasive flag-threading through every call site
of the static method or a lossy approximation that picked the wrong
direction for the default case. Per the briefing's ``pick a
phase'' guidance, this is bumped to Phase 4. The pre-existing
silent under-count on Anthropic-thinking turns persists when
replay=True. Strip-when-False naturally fixes the under-count by
keeping the bytes off the wire entirely; the residual case is the
opt-in replay path.
Tests (28 new, all driving through real boundary objects)
* ``tests/test_provider_anthropic_replay.py`` (19 tests):
- Strip vs preserve under both flag values (3 tests including
redacted_thinking).
- Default-kwarg back-compat preserves verbatim replay (1 test).
- Web-search tool_use + server_tool_use + web_search_tool_result
survive strip with encrypted_content intact (2 tests, edge 14).
- Orphan-tool synthesis after strip -- pins the
``provider_content`` source-of-truth read at lines 397-433
(1 test).
- Foreign-shape fallthrough: OpenAI ``type="reasoning"`` block
rebuilds via text+tool_calls (1 test).
- Mixed-shape fallthrough: even one foreign block forces
rebuild (1 test).
- Empty / None / non-list ``_provider_content`` fallthrough
(3 tests).
- Legacy Anthropic-thinking row pre-Phase-2 stays in verbatim
path -- no regression on existing conversations (2 tests).
- All-blocks-stripped fallthrough behaviour: rebuild from text
if available, silently skip if not (2 tests).
- Constants pinning: strip set is narrow, valid set includes
web search, strip is subset of valid (3 tests).
* ``tests/test_session_replay_reasoning.py`` (12 tests):
- Resolver: 6 tests covering miss / default / set / explicit /
fallback alias / exception.
- Streaming call site: 3 tests pinning the kwarg propagates
through ``_try_stream`` to a stub provider.
- Non-streaming call site: 1 test pinning
``_utility_completion`` propagates the flag.
- End-to-end boundary integration: 2 tests driving
``_try_stream`` -> real ``AnthropicProvider`` -> captured
Anthropic SDK ``client.messages.stream`` boundary, asserting
on the ACTUAL wire payload shape. Negative-tested:
temporarily reverting the kwarg-thread at
``_anthropic.py:create_streaming`` makes the wire test fail
with ``Strip predicate did not fire at wire boundary``;
restoring makes it pass.
The boundary integration tests were added in response to a code
review finding that the bare-stub call-site tests would not catch
a regression where the provider stops reading the kwarg or
``_convert_messages`` silently drops the strip. The integration
tests close that gap by inspecting what reaches the (mocked) SDK,
not just what the provider was called with.
Lint + test gate
* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6061 passed (3 deselected). Phase 2
added 28 net new tests.
Surface stored Anthropic thinking blocks on /history responses so
refreshing the page rehydrates the reasoning bubble. Wire payloads
unchanged. Per-model operator knobs added to model_definitions for
both UI rehydration and (Phase 2) wire-build replay.
Why now: reasoning is already round-tripped via _provider_content for
Anthropic-with-thinking turns, but never surfaces on the history wire,
so a tab reload showed only the final answer with no rationale.
Operators also have no per-model lever to opt out of UI display or to
opt in to replay-to-model on subsequent calls.
What this change does
* Migration 052 adds two boolean columns to model_definitions:
persist_reasoning (default 1) controls UI rehydration; replay_
reasoning_to_model (default 0) reserved for Phase 2's wire-build
shape filter. Mirrors the enabled column pattern (NOT NULL +
integer server_default).
* LLMProvider Protocol gains extract_reasoning_text(provider_blocks)
with concrete impls on AnthropicProvider (walks type=='thinking'
blocks, joins with newline, caps at 64 KiB) and no-op stubs on
OpenAIChatCompletionsProvider + OpenAIResponsesProvider. Google
inherits the no-op via OpenAIChat. Phase 3 will wire the OpenAI
Responses extractor once include=['reasoning.encrypted_content']
is requested.
* turnstone.core.history_decoration gains a structural dispatcher
extract_reasoning_text_from_provider_content keyed off the first
block's type field (Anthropic 'thinking' / OpenAI Responses
'reasoning' / Gemini 'thought' are non-overlapping by API design).
Both history surfaces use it: _build_history calls the dispatcher
directly (the SSE-replay path builds entry dicts from scratch),
and the lifted make_history_handler runs the list-helper variant
in the existing to_thread block.
* make_history_handler resolves persist_reasoning via three tiers:
live session -> workstream_config.model_alias (the same key
SessionManager uses to rehydrate the original model after process
restart) -> conservative True default. Operator flag-flip takes
effect uniformly on both warm and cold workstreams.
* Frontend: app.js replayHistory and coordinator.js role==='assistant'
branch each call the existing reasoning-bubble construction (for
app.js, the document.createElement pattern from the live SSE
handler; for coord, the appendMsg('reasoning') helper) when
msg.reasoning is non-empty. Reasoning bubbles render before the
content bubble, matching live SSE order.
* Admin UI: two checkboxes ('Persist reasoning', 'Replay reasoning
to model') in the model edit modal, plus override-pill display in
the model row when set to non-default values.
What is intentionally out of scope
* Phase 2 -- ANTHROPIC_VALID_BLOCK_TYPES shape filter at
_anthropic.py:312-316, _convert_messages replay_reasoning_to_model
parameter, thinking-strip branch, _msg_text_chars token-calibration
extension. The replay flag is stored but not consumed on the wire.
* Phase 3 -- OpenAI Responses include=['reasoning.encrypted_content'],
Gemini include_thoughts spike, ModelCapabilities.supports_
reasoning_replay.
* Phase 4 -- Local-model / chat-template reasoning persistence
(session.py:3486 reasoning_parts accumulator).
Tests
* AnthropicProvider.extract_reasoning_text -- 13 unit tests covering
None / empty / mixed / multi-block / cap / malformed / non-list
inputs plus other-provider no-op verification (real provider
instances, no mocks).
* extract_reasoning_for_history -- 10 dispatcher tests including
block-type discriminator routing (thinking vs reasoning vs
unknown), strip-when-flag-false, empty / non-dict guards, and
cross-role isolation.
* _build_history -- 6 boundary tests through the real Anthropic
extractor with stub sessions, including the registry-lookup
failure default-True branch.
* make_history_handler -- 5 round-trip tests through real storage:
the storage layer's reconstruct_messages decodes provider_data
into _provider_content, and the helper extracts through the real
AnthropicProvider. Includes the live-session flag honoring path,
the cold-workstream workstream_config lookup path, and the
no-alias default-True fallback path.
* Audit-log discipline -- 4 structural mock-and-assert tests that
capture every Logger.info / warning / error call across the
pipeline (extractor, dispatcher, list-helper, _build_history)
and assert no captured payload contains a marker reasoning string.
* model_definitions storage -- 6 round-trip tests: default flags,
explicit create with both flags, individual update of each flag,
and list-includes-flags assertion.
* model_registry -- 4 tests: dataclass defaults, dataclass with
explicit flags, DB-row-mapping with both flags, and pre-052
legacy-row default-fallback.
Edge cases pinned by the test suite
* Pre-052 DB rows missing the new columns degrade to dataclass
defaults (test_db_reasoning_flags_default_when_absent).
* Live session in memory has its flag honored (test_history_handler_
with_persist_flag_false_via_live_session).
* Cold workstream resolves the flag via workstream_config +
app.state.registry (test_history_handler_cold_workstream_resolves_
via_workstream_config) -- this closes the gap where a process
restart would have silently un-honored an operator flag-flip.
* Cold workstream without persisted model_alias falls through to
default True (test_history_handler_cold_workstream_no_alias_
defaults_true).
* Foreign / unknown / missing block types degrade silently to no
reasoning field rather than misroute or crash.
Lint + test gate
* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6030 passed (3 deselected).
Doc-debt cleanup flagged by /review on 9dc29db7. The cap+seq fix
flipped the seq-advance rule but left two doc sites describing the
old "incremented only on actual append" shape — exactly the buggy
invariant the previous commit removed. Future readers trusting the
stale docs would be one wrong assumption away from re-introducing
the silent-drop bug.
Updates the field-init comment block and the docstring on
register_listener_with_in_progress_snapshot (which sits at the
snap_seq capture site, so its contract is consumer-facing).
Also drops the now-dead `seq: int = 0` initializer in
on_reasoning_token and on_content_token — under the new shape, the
unconditional `seq = self._ws_inflight_seq` inside the lock makes
the initializer unreachable. Was load-bearing under the old
else-branch; harmless now but signals "some path leaves seq at 0"
to a reader.
Copilot caught a real bug in the cap+seq interaction: the previous
shape only advanced ``_ws_inflight_seq`` when the buffer actually
appended, on the theory that "every _seq corresponds to a buffered
fragment" was a useful invariant. It wasn't — once the buffer hit
its cap, seq stalled at the high-water-pre-cap, so a subscriber that
registered AFTER the cap was hit would capture
``snap_seq == stalled_seq``, and every subsequent live token (also
tagged with the stalled seq) would be filter-dropped by the events
handler's ``seq <= snap_seq`` dedup. Silent loss of the entire
post-cap stream for refresh-past-cap tabs.
Fix: advance seq on every emit, regardless of buffer cap. The cap
is a buffer-size limit, not a stop-streaming signal. Past-cap tokens
are absent from the snapshot's text payload (the buffer was
truncated at cap) but the live stream past them is now correctly
delivered — refresh-after-cap renders snapshot-up-to-cap then live
tokens past it, with a visual gap equal to the past-cap chunk and
no silent drop of subsequent tokens.
Test ``test_inflight_seq_increments_only_on_actual_append`` enforced
the buggy invariant and is renamed/flipped to
``test_inflight_seq_advances_on_every_emit_even_at_cap``. Added
``test_subscriber_after_cap_hit_receives_subsequent_tokens`` (and
the reasoning equivalent) as direct regressions for the
silent-token-loss scenario.
Updates the docs that describe the per-workstream SSE event stream and
the SessionUI lifecycle to match the refresh-resume changes:
- api-reference.md: documented the `state_change` event (previously
undocumented despite already being a live event) and the new
`in_progress_snapshot` event; rewrote the multi-consumer fan-out
paragraph to mention the kind-specific replay tail (state_change +
optional in_progress_snapshot) so the "no catch-up needed" claim
is no longer misleading.
- architecture.md: bumped the SessionUI Protocol stub to 16 methods
(added `on_turn_start` / `on_turn_committed`) and pointed at the
in_progress_snapshot section in the API reference.
- sdk.md: added rows for `state_change`, `in_progress_snapshot`, and
`approval_resolved` (preexisting gap) to the per-workstream event
table.
- coordinator-api-tour.md: added an `in_progress_snapshot` row to the
event table and rewrote the reconnection-contract paragraph to
cover mid-stream content/reasoning restoration.
- diagrams/04-conversation-turn.puml: added `on_turn_start()` before
the thinking-start emit and `on_turn_committed()` immediately after
`messages.append(assistant_msg)`, with notes explaining the inflight-
buffer reset semantics. PNG regenerated.
Refreshing a coordinator or interactive workstream pane while the LLM
is mid-stream now restores the partial assistant text + reasoning
immediately and flips the composer back to stop-mode, instead of
showing nothing until the response completes.
Per-turn inflight buffers (`_ws_inflight_content`, `_ws_inflight_reasoning`,
`_ws_inflight_seq`) on `SessionUIBase` are kept separate from the
existing multi-turn `_ws_turn_content` buffer that drives the
dashboard's IDLE-piggyback payload. New `on_turn_start` (top of
send-loop, defensive) and `on_turn_committed` (right after
`messages.append(assistant_msg)`, primary) lifecycle hooks reset
inflight at turn boundaries. The seq counter is monotonic across
turns so a long-lived subscriber's `snap_seq` cutoff stays valid for
the lifetime of the connection — resetting per-turn would silently
drop turn N+1's first M tokens (M = whatever was streamed pre-snapshot
in turn N).
`snapshot_and_consume_state_payload` also drains inflight at idle/error
so cancel and exception paths don't leak stale text. New
`register_listener_with_in_progress_snapshot` atomically registers a
listener and snapshots the inflight buffers; `make_events_handler`
emits a `state_change` event (so the JS busy machine flips to
stop-mode) followed by a one-shot `in_progress_snapshot` after the
kind-specific replay, then strips the internal `_seq` field from
yielded live events while filtering against `snap_seq`. A per-listener
shallow `dict` copy in the live drain prevents the multi-tab race
where one listener's `del event["_seq"]` would corrupt another
listener's filter view.
`_synthesize_cancelled_results` now emits synthetic `on_tool_result`
events for each cancelled tool so live coord tabs can drop the
newly-additive `coord-tool-batch--running` indicator cleanly. The
indicator now coexists with `--auto`/`--approved` (applied on
`tool_info` and `approval_resolved` approved; removed when every row
in the batch has a result), making live tool execution visually
parallel to the replay-time orphan rendering.
Frontend handlers in `app.js` (interactive) and `coordinator.js` (coord)
absorb EventSource auto-reconnect re-replays via a length-based
prefix check on the in-progress buffer. New `InProgressSnapshotEvent`
+ `StateChangeEvent` dataclasses in the Python and TypeScript SDKs
with type guards.
`_MAX_TURN_CONTENT_CHARS` lifted 256 KiB → 512 KiB (single constant
for both buffers — headroom for current commercial models).
Regression tests cover race-free composition under concurrent writers,
seq-filter dedup invariants, the cross-turn seq monotonic invariant,
idle/error inflight drain, synthesized `on_tool_result` on cancel
(including UI-hook failure isolation), and the multi-listener
shared-dict invariant.
Two findings, both confirmed against the source:
1. Migration 051's downgrade rewrote every '[]' row back to '{}',
which would (a) destroy operator-written empty arrays and
(b) reintroduce the known-invalid sentinel that every consumer
rejects. Pre-migration '{}' rows and operator-authored '[]' rows
are indistinguishable after upgrade — there is no clean inverse
for the data state. Made downgrade an explicit no-op with the
rationale documented inline; '[]' is the correct shape under any
consumer's interpretation, so leaving the data untouched on
downgrade is strictly safer than reversing it. Updated the
module docstring to call this out.
2. admin_update_skill's notify_on_complete validator short-circuited
on empty string: `if nc and nc != "[]":` skipped the JSON-parse
branch when nc=="" and persisted the empty string straight to
storage, leaving a non-JSON value behind. Folded the empty case
into the existing "{}" coercion so any blank/whitespace/legacy
value normalises to "[]" before the array-validation gate.
Tests: three new regressions in TestSkillAPI — empty-string
normalises, "{}" sentinel coerces, non-array JSON 400s. The third
locks in the array-only validator that the previous "valid JSON"
gate would have accepted.
Every consumer of prompt_templates.notify_on_complete treats it as a
JSON-array string (the admin form's array editor, the JSON.isArray
validator in submitEditTemplate, _validate_notify_targets in
server.py, the documented "list of channel/contact identifiers"
shape). But the column's server_default — set in migration 011 and
inherited through 021's lift into prompt_templates — has been "{}"
(an empty JSON object) since day one.
Newly-installed remote skills inherit the schema default, so every
unlock-then-edit flow trips the array validator on the inherited
"{}" and the request never leaves the browser. The user-visible
symptom was "click Save, nothing happens"; the latent symptom was
silent shape divergence between every install and every operator-
authored skill.
Migration 051: rewrites every legacy "{}" row to "[]". Operator-
edited values (anything that's neither "{}" nor NULL) are left
intact. Downgrade restores "{}" only on rows still holding the
post-migration "[]" so any later operator edits stick.
Server-side defaults flipped to "[]" in the same PR so new rows
land correct without depending on the column's server_default:
- _schema.py prompt_templates.notify_on_complete server_default
- StorageBackend protocol create_prompt_template kwarg
- sqlite + postgres create_prompt_template kwargs
- console_schemas.py SkillCreateRequest / SkillUpdateRequest /
SkillInfo Pydantic defaults
- core/session.py ChatSession._notify_on_complete initial value
- server.py initial-message worker fallback when skill_data omits
the field
admin_update_skill validator now also rejects non-array JSON (was
"valid JSON" only — would have accepted "{}" or "{\"a\": 1}").
_skill_to_response coerces legacy "{}" rows to "[]" on read so the
admin UI sees a consistent shape even before migration 051 runs.
The frontend's `tmpl.notify_on_complete || "[]"` fallback already
handled empty-string but not "{}" — the read-side coercion makes
it moot.
Designer-review follow-up to the .is-visible sweep. With role=alert
+ aria-live=assertive, AT engines re-announce when the element's
text content changes — but without aria-atomic some engines only
read the diff between old and new content. With aria-atomic=true
the entire updated message is read each time, which matters when a
validation error is replaced by a server error on retry (or
vice-versa).
Added aria-atomic=true to all 24 modal error elements (every
role=alert with aria-live=assertive). Same accessibility uplift
across the board — no per-modal exceptions.
Also dropped the stale `style="display: none"` attribute from the
three MCP error elements (mcp-create-error, mcp-import-error,
mcp-install-error). The CSS rule
.admin-modal [role="alert"] { display: none; }
already hides them by default — the inline attribute was redundant
and would have overridden the .is-visible toggle if the class-based
contract is ever changed.
Sweep of the latent bug PR #494 fixed for the skill modals: the
project's CSS contract for modal errors is
.admin-modal [role="alert"] { display: none; }
.admin-modal [role="alert"].is-visible { display: block; }
…but ~30 sites across governance.js and admin.js were toggling
`style.display = ""` instead of the .is-visible class. The "show"
side broke silently — clearing the inline style fell back to the
CSS `display: none` so the error never rendered, and any
validation failure looked like an unresponsive button.
Mechanical conversion of every show/hide site for these modal
error elements:
governance.js
create-role-error, edit-role-error
create-policy-error, edit-policy-error
github-import-error
cpp-error, epp-error (custom + eval prompt policies)
create-hr-error, edit-hr-error (heuristic rules)
create-ogp-error, edit-ogp-error (output-guard patterns)
admin.js
mcp-create-error, mcp-import-error, mcp-install-error
Plus the global `_showModalError` helper in admin.js — its
`style.display = "block"` happened to work today (inline display
beats the CSS rule), but normalising it to .is-visible keeps every
modal on a single canonical path. The five modals that route their
show side through that helper (create-user, create-token,
create-channel, create-schedule, edit-schedule) had their hide
sides converted in lockstep.
Added a comment on `_showModalError` documenting the contract so
the next contributor doesn't reintroduce the bug.
Out of scope: model-create-error (already canonical), home-coord-error
(not in .admin-modal), edit/create-template-error (fixed in #494).
No CSS or HTML changes; behaviour-equivalent for hide sides; show
sides go from broken-silent-no-render to correct-render-with-AT-
announcement.
Once edit-template-error is actually visible (the visibility fix in
this same PR), a stale error now persists across resubmit cycles:
the user sees a red message, fixes the input, clicks Save, the
validator passes, the PUT goes out — and the previous error stays
on-screen the whole time, only clearing when the modal closes on
success.
Fix at the start of submitEditTemplate / submitCreateTemplate:
clear .is-visible AND empty textContent. Cheaper than tracking
every validator branch and every .catch path; a fresh submit is a
clean slate.
Smoke-testing the unlock flow surfaced a latent bug: clicking Save
on the edit-skill modal silently no-op'd whenever the
notify-on-complete field had non-JSON content. The error div was
DOM-correct (text content set, role=alert, aria-live=assertive),
but invisible — because the project's modal-error CSS contract is:
.admin-modal [role="alert"] { display: none; }
.admin-modal [role="alert"].is-visible { display: block; }
…and the JS in submitEditTemplate / submitCreateTemplate was
clearing the inline `display: none` via `el.style.display = ""`.
That falls back to the CSS rule, which still says `display: none`,
so the error never rendered. The user saw no error and the click
felt unresponsive (compounded by the early-return before the
disabled-state reset, which also made Save look broken).
Fixed both skill-modal flows (create + edit) by toggling the
canonical `.is-visible` class instead. Six sites in governance.js:
the two early-return show paths, the two .catch show paths, and
the two modal-open hide-resets.
Scope note: this same bug pattern exists in ~20 other modal error
sites across governance.js and admin.js (create-role, edit-role,
create-policy, edit-policy, github-import, cpp, epp, create-hr,
edit-hr, create-ogp, edit-ogp, mcp-create, mcp-import, mcp-install,
plus admin.js sites that don't go through _showModalError). All
pre-existing, broken silently for who knows how long. Out of scope
for this PR — recommend a follow-up sweep that also normalises
_showModalError's `style.display = "block"` to the same convention.
Designer review of the cb5fa1b lock-icon iteration flagged five
items; four are addressed here, one was a deliberate trade-off
documented below.
- Glyph hardening (#2): the lock character is now 🔒︎ — U+1F512 with
the U+FE0E text variation selector — paired with the existing
font-variant-emoji: text rule. font-variant-emoji shipped late
and isn't universal yet (Chrome 131+, Safari 16.4+, Firefox 132+);
the explicit text VS is belt-and-braces so older Chromium / most
Linux don't fall back to a coloured emoji that would clash with
the monochrome instrument-panel aesthetic.
- Accent-line de-conflict (#3): top:14px → 18px so the lock button
sits below the modal's ::before accent-line decoration's visual
band rather than competing with it horizontally. h2's
padding-right reservation (44px) still gives the title clearance.
- Mobile touch target (#4): @media (max-width: 700px) bumps the
button to 44×44 (WCAG 2.5.5 / Apple HIG / Material minimum) and
shifts it to top:8px right:8px, with h2 padding-right widened to
56px to match.
- Keyboard discoverability (#6): on readonly open, focus lands on
the lock button instead of Cancel. Keyboard users hit the unlock
affordance immediately instead of having to Tab past every
disabled spec input to reach it. Cancel is one Shift-Tab away.
Deferred:
- (#1) Reviewer flagged top-right placement as risking confusion
with the universal × close-button convention. Keeping the
icon-only design per product direction; the bordered chip styling
+ accent-coloured hover make it visually distinct from the
thin-stroke unbordered × pattern, and the confirm dialog catches
any misclick safely.
- (#5) Optional empty-corner indicator after unlock — the
"Customized from upstream" badge text already carries the signal;
not adding new chrome.
Three issues from manual smoke-testing the unlock flow:
1. Confirm dialog rendered behind the edit-skill modal. Both
overlays sat at z-index 600, and confirm-overlay is earlier in
the DOM than edit-template-overlay — so DOM order put the parent
modal on top of its own confirm. Bumped confirm-overlay to 650
(still below toasts at 700) since confirm dialogs are launched
FROM other overlays and need to sit above them.
2. Save button stayed disabled (or non-functional) after unlock.
submitEditTemplate disables etm-submit on click and re-enables in
.finally, but a stale disabled=true survives the mutate-in-place
re-render that runs after unlock. Always reset
submitBtn.disabled = false in showEditTemplateModal so the
re-render path can never inherit a stuck disabled state.
3. UX redesign — moved the unlock affordance from a "Customize…"
button at the bottom of the footer to a 🔒 icon button at the
top-right of the modal. The lock glyph is the universal "this is
locked, click to unlock" affordance and reads more clearly than
a footer button next to Cancel/Save. font-variant-emoji: text
keeps it monochrome on browsers that support it (instrument-panel
aesthetic) with graceful fallback to coloured emoji elsewhere.
admin-modal-skill h2 reserves padding-right so a long title can
never collide with the absolute-positioned button.
Cleanup: removed the now-unused .modal-secondary and
.modal-buttons-spacer rules; the bottom etm-unlock button + flex
spacer are gone from the modal footer.
Copilot caught that prompt_templates.readonly is an Integer column
(_schema.py: sa.Column("readonly", sa.Integer, nullable=False,
server_default="0")) and create_prompt_template stores it as 1/0,
but unlock_skill in the postgres backend was passing a Python bool
(readonly=False). The sqlite impl already uses 0; this aligns the
two backends and matches the 0/1 idiom used for the sibling flag
columns (is_default, auto_approve, enabled).
The other Copilot findings on this PR (loadGovSkills race, NBSP
double-space, list_skill_versions O(history_size), ignored
set_skill_readonly return value + None re-read) were all closed by
the prior review-feedback commit (eea795d): the snapshot+flip is
now an atomic unlock_skill() that uses SELECT MAX(version)+1
internally, the handler guards both the unlock_skill return and the
post-flip get_prompt_template re-read, the JS chains
showEditTemplateModal off loadGovSkills's promise, and the badge
NBSP matches the sibling pattern.
Code review caught a race + a missing None guard; designer review
caught a window.confirm regression and a button-hierarchy issue.
Backend:
- Race fix (bug-2): replace set_skill_readonly+create_skill_version
with a single atomic unlock_skill(template_id, snapshot, changed_by)
-> int|None on the storage protocol (sqlite + postgres). Snapshot
insert + readonly flip happen in one transaction; the next version
number is computed via SELECT MAX(version)+1 inside the txn rather
than len(list)+1 outside, closing the (skill_id, version)
collision window where two concurrent admin actions could both pick
the same version.
- None guard (bug-3): check the post-flip get_prompt_template re-read;
return 404 instead of letting _skill_to_response(None) raise.
- Audit body: also record snapshot_version, and harden None-vs-empty
with `or ""` on the existing.get(...) calls.
Frontend:
- D-1: replace window.confirm with the existing showConfirmModal
(admin.js:2350) — themed dialog, focus-trap, can render the source
URL with consistent typography. The native dialog could collapse
the multi-paragraph copy depending on browser.
- D-2: mutate-in-place on success rather than hide → reload → reopen.
loadGovSkills now returns its fetch promise so unlockSkill can
chain showEditTemplateModal after the cache refresh — no flicker,
no focus bounce, and it kills bug-1 (the reopen was reading stale
_govSkills before loadGovSkills resolved). showEditTemplateModal
is idempotent when already open: it skips the trigger-element
capture and the focus-trap reinstall.
- D-3: button hierarchy. Drop flex:1 from .modal-secondary so the
Save button keeps a stable width whether or not Customize is
rendered; insert a flex-spacer between Customize and Save so the
destructive-ish detach groups left next to Cancel and the primary
action floats right.
- D-4: NBSP normalized to match the existing escape pattern
on the sibling badge line (was an actual NBSP byte).
- D-5: success toast now reads "Skill unlocked — fields are now
editable" so the operator gets a positive affirmation that the
edit affordance is live.
- D-10: aria-describedby="etm-origin-badge" on disabled spec inputs
so screen-reader users get the same "this came from upstream"
context that sighted users see in the cyan badge.
Tests: + test_unlock_skill_versions_after_existing_history seeds an
out-of-order version (3) and asserts unlock picks 4, defending
against the len()-based version computation regressing.
skills.sh / GitHub installs land with readonly=True so admins can only
tune runtime config (model, temperature, etc.); the SKILL.md spec is
locked. In practice, upstream skills aren't always tuned for turnstone,
so locking the spec adds friction without a real safety win — every
edit is audited and version-snapshotted regardless.
This adds an explicit unlock so the boundary stays visible (multi-user
audit trail benefits from a discrete event, vs. silently dropping the
gate). Behaviour:
- POST /v1/api/admin/skills/{id}/unlock — flips readonly=False on a
readonly row. Snapshots the pre-unlock state into skill_versions so
the upstream-pristine version is recoverable from the History tab.
Records skill.unlock audit with {name, source_url, origin}. 400 on
already-unlocked, 404 on missing.
- origin stays "source" after unlock so the UI keeps a "Customized
from upstream" provenance badge — the readonly flag is the gate, the
origin field is the lineage.
- Storage: dedicated set_skill_readonly writer on the protocol +
sqlite + postgres backends. readonly is intentionally absent from
SKILL_MUTABLE so the generic update path can't piggyback on a
provenance flip — the dedicated writer pattern matches what's
already used for set_mcp_oauth_client_secret_ct.
- Frontend: "Customize…" button in the edit modal (visible only when
readonly), with a confirm dialog explaining the upstream-detach.
Once unlocked the existing edit-skill flow handles spec edits with
no other changes. Origin badge updates to show "Customized from"
the upstream URL when a source-origin row is unlocked.
Tests cover: unlock flips readonly + persists, pre-unlock snapshot
written to skill_versions, 400 on already-unlocked, 404 on missing,
post-unlock PUT can edit name/content/description (the readonly gate
no longer fires).
Three issues caught by Copilot on the initial PR:
1. SKILL.md size cap was measured in code points, not UTF-8 bytes.
`len(str)` is a *lower* bound on encoded byte length — multi-byte
chars (emoji, CJK) inflate up to 4×, so a 100k-emoji SKILL.md
(400KB encoded) would slip past the 256KB cap. Switch to
`len(contents.encode("utf-8"))` and surface lone-surrogate failures
as SkillSourceError instead of dropping them silently. New
regression test feeds emoji content.
2. _skills_sh_source_url did not normalize the skill_id, so a sloppy
id from `/api/search` (whitespace, surrounding slashes) would pass
`_split_skills_sh_id`'s charset check (which strips first) and
produce a malformed persisted source_url that broke the
discover-UI dedup contract. Strip the id inside the helper, and
reconstruct the canonical id from validated parts in
download_skill's listing so downstream callers never see the raw
input.
3. The catch-all `except Exception:` around create_prompt_template
relabeled every storage failure (DB connection, disk full,
permission errors) as "conflict", masking operational issues.
Translate IntegrityError → StorageConflictError at the storage
shim (matching the pattern already used for OIDC user
provisioning) in both sqlite and postgres backends, then catch
StorageConflictError specifically in the install handler. Real
conflicts → "conflict" + warning; other exceptions → new
"internal error" reason + log.exception.
Tests: +3 (oversized multibyte SKILL.md, source_url normalization,
storage-layer conflict translation). 226 passing.
The skills.sh install path was failing with 404s because their public
API surface changed: /api/skills/{id} is gone, replaced by
/api/skill/[owner]/[repo]/[skill] (auth-walled) and
/api/download/[owner]/[repo]/[skill] (unauthenticated, returns the
SKILL.md + bundled resources inline as JSON). The error was not
surfacing in logs because admin_skill_install had a silent
`except Exception:` around create_prompt_template that relabeled every
storage failure as "conflict" with no log entry.
- Replace SkillsShClient.resolve_github_url with download_skill that
hits /api/download/{owner}/{repo}/{skill} and returns a SkillPackage
directly. No GitHub round-trip; no rate-limit surface.
- Add _split_skills_sh_id with strict per-segment charset validation
([A-Za-z0-9._-]+) so URL-hostile content can't produce a malformed
request or divergent persisted source_url.
- Use len(contents) instead of len(contents.encode("utf-8",
errors="ignore")) for the SKILL.md size cap — errors='ignore' was
silently dropping invalid units, making the cap bypassable.
- Extract _accept_resource(rel_path, byte_size) gate predicate; share
it between download_skill and the GitHub _find_resource_files helper.
- Have search() derive a deterministic source_url from the skill id
when /api/search omits one (which it currently always does), so the
discover-UI "already installed" check matches what download_skill
persists.
- Add structured logging across admin_skill_install and
admin_skill_discover: a shared _log_install_failure helper for the
four except branches (was four near-duplicate log calls with one
drift), plus per-resource failure tallying — partial-resource
installs now surface failed_resources in the response and audit
record instead of silently committing the skill row with missing
assets.
Tests: 7 new — empty/non-list files, oversized SKILL.md, resource
cap, non-text extension filtering, plus _split_skills_sh_id charset
rejection (whitespace, query chars). Verified end-to-end against
live skills.sh with tavily-search.
Four Copilot findings on c6041c6 — all confirmed valid, all bounded
to authenticated-user prompt-injection scenarios but worth closing
before merge.
Wrapper-detect bypass (string + list branches of
``_apply_reminders_for_provider``):
The round-2 fix used ``content.startswith("<tool_output>\\n")`` to
detect already-wrapped content and skip ``escape_wrapper_tags``. A
tool whose RAW output starts with that prefix (e.g. ``echo
'<tool_output>'``) would match and have its escape skipped, letting
literal ``<tool_output>`` / ``<system-reminder>`` tags reach the model
and impersonate a system envelope. Replace the prefix check with
``extract_advisories_from_tool_envelope(content) is not None`` —
parsing requires the open AND matching close tags AND a structurally
valid envelope, raising the bypass bar significantly.
Mirror fix in the list-content branch so a tool emitting an unmatched
envelope as a text part can't bypass the per-text-part escape.
``_build_history`` legitimate-envelope drop:
The list-content drop path previously removed any text part starting
with ``<tool_output>\\n``. A tool that legitimately outputs a
well-formed envelope (documentation viewer, code analyzer demoing the
wrapper, an echo tool) would have that part silently disappear on
replay. Tighten the drop heuristic to require BOTH ``cleaned_text ==
""`` AND at least one extracted advisory — the structural signature of
the injected ``wrap_tool_result("", advisories)`` carrier we produce
in ``session.py`` for list-typed tool output. A legitimate envelope
has non-empty inner body or no advisory blocks and survives the
projection.
Empty advisory body:
``queue_message`` accepts any non-None text including ``""`` and
whitespace-only strings. ``_classify_advisory`` would return a
``user_interjection`` advisory with empty / whitespace body, which
``replayAdvisoriesAfterTool`` then renders as a featureless empty user
bubble. Filter empty / whitespace-only bodies at classification time
so the wire-shape contract is uniform: no empty advisories ever ride
the wire.
Tests:
* ``test_apply_reminders_escapes_tool_output_starting_with_envelope_prefix``
pins the structural-parser bypass close: a string starting with the
envelope prefix but lacking a close tag still gets escaped.
* ``test_apply_reminders_escapes_list_text_part_with_unmatched_envelope_prefix``
mirrors for the list-content branch.
* ``test_build_history_keeps_legitimate_envelope_text_part_with_body``
pins that legitimate envelope output stays in the projected list.
* ``test_decorate_suppresses_empty_advisory_body`` and
``test_decorate_suppresses_whitespace_only_advisory_body`` pin the
empty-body filter in ``_classify_advisory``.
Tests: 5923 passed, 3 deselected. Lint + format + mypy clean.
Reverses the seam-2-only design from the prior commits on this branch.
Queued user messages arriving DURING a tool batch (Seam 1) splice into
the last tool result's envelope as ``UserInterjection`` advisories via
``wrap_tool_result``. Messages arriving BETWEEN turns (Seam 2) drain
as a single trailing user row via ``_flush_queued_messages`` with
``user_feedback`` (operator text alongside an approval, e.g. "y, use
full path") folded in as a prefix. Cancel/exception drains (Seam 3)
keep the existing ``_flush_queued_messages()`` call unchanged.
Why all three seams:
* Strict-template providers (Mistral, Llama via vLLM with stock chat
templates) reject role-alternation violations. A literal ``user``
row mid-tool-batch breaks ``assistant(tool_calls) → tool → ... →
assistant``; back-to-back ``user → user`` rows on the wire also fail.
* The seam-2-only design produced back-to-back ``user`` whenever
``user_feedback`` and queued items both fired — bug-1 from the round-1
review. Folding ``user_feedback`` as a prefix to the queue-drain
collapses the two into one row.
* During-batch arrivals couldn't ride seam 2 — the splice was the only
way to deliver same-turn without violating role alternation.
Storage symmetry:
Tool DB rows now store the wrapped ``output`` (envelope + advisories)
unconditionally — ``self.messages[i]['content']`` and
``conversations.content`` match exactly. List-typed output (image /
structured MCP results) uses ``wrap_tool_result(raw_joined_text,
advisories)`` at save time so the persisted string is anchored on
``<tool_output>\n`` for the replay parser. ``TOOL_RESULT_STORAGE_CAP``
is removed entirely; tools are responsible for bounding their own
output, storage faithfully represents in-memory. Removing the cap
also simplifies the parser — no truncated-envelope edge case.
Replay extraction:
``decorate_history_messages`` (REST ``/history``) and ``_build_history``
(SSE replay, resume, rewind, retry, post-load, rename re-replay) both
call the public ``extract_advisories_from_tool_envelope`` helper to
pull the envelope back into structured ``advisories`` for JS replay.
Both string content and list-typed content (image+queued-message
combo) covered. JS renders extracted advisories as normal user
bubbles after the tool block via the shared ``replayAdvisoriesAfterTool``
helper in ``shared_static/utils.js``.
Wrapper-tag escape and provider splice:
``escape_wrapper_tags`` now encodes pre-existing ``&`` first using an
``&`` sentinel so tool output containing literal entity strings
(documentation viewers, code analyzers, web scrapers returning entity-
encoded markup) round-trips correctly. Both encode and decode helpers
short-circuit on absence of ``<`` / ``&``.
``_apply_reminders_for_provider`` detects already-wrapped content
(string body and list text-part) by ``startswith("<tool_output>\n")``
and skips re-escape so existing envelopes survive intact when a tool
message also carries ``_reminders`` (the queued-message + tool-error
co-occurrence case is now common).
``decorate_history_messages`` runs in ``asyncio.to_thread`` to keep
MB-scale string work off the event loop.
Other cleanup:
* ``_collect_advisories`` delegates the queue drain to a named helper
``_drain_queued_messages_to_advisories`` so the swap-and-clear pattern
lives next to ``_flush_queued_messages``'s identical pattern and the
side-effect is documented at the call site.
* Preamble strings + body marker for ``UserInterjection`` round-trip
detection moved to module-level constants in ``tool_advisory.py``;
imported by ``history_decoration.py`` so a producer-side rephrase
can't silently desync the parser.
* ``_send_with_mocks`` ctxmgr extracted in ``test_session.py`` — the
six new send-driven tests share an 8-deep ``patch.object`` block.
* ``replayAdvisoriesAfterTool`` shared helper in
``shared_static/utils.js``; ``app.js`` and ``coordinator.js`` both
invoke it.
* Dead truncation-pill CSS removed (``.tool-output-truncated`` and
``.coord-tool-truncated``); the JS that added these elements went
away with ``TOOL_RESULT_STORAGE_CAP``.
* Tautological tests (``TestBuildHistoryAdvisoryPropagation``)
replaced with production-realistic round-trip tests built from
``wrap_tool_result(...)`` envelopes — REST and SSE-replay surfaces
pinned to the same wire shape; full DB round-trip pinned end-to-end.
Negative-tested:
* Reverting the prefix-merge in ``_flush_queued_messages`` produces
back-to-back ``user`` rows, breaking
``test_user_feedback_and_queued_coexistence_single_row_with_prefix``.
* Reverting the ``extract_advisories_from_tool_envelope`` call in
``_build_history``'s tool branch leaves the envelope verbatim in
wire content, breaking the round-trip tests.
* Reverting the wrapper-detection in ``_apply_reminders_for_provider``
entity-encodes the existing envelope's literal tags, breaking both
the string-content and list-content envelope-preservation tests.
* Reverting the ``wrap_tool_result(raw_text, advisories)`` projection
at the DB save site produces a string starting with the original
raw text, breaking
``test_tool_db_row_round_trips_list_output_with_advisories``.
Tests: 5918 passed, 3 deselected. Lint + format + mypy clean on
touched files.
Round-1 ``/review`` apply-pass. Drops stale ``UserInterjection``
references from comments and docstrings that no longer describe the
post-PR drain shape, asserts the two-stream invariant in the new
queued-message persistence test, and pins the ``content.trim()`` +
``renderAssistantToolBatch`` invariants on coord-side so a future
refactor can't silently regress the Qwen3 phantom-card fix or the
chronological-order render fix.
Deferred:
* **bug-1** (back-to-back ``user`` row when ``user_feedback`` from the
approval-prompt UI callback coexists with a queued-message drain).
Reachable on strict OpenAI-compatible local templates (Anthropic and
Anthropic-via-merge-consecutive collapse fine; vLLM-hosted Mistral /
Llama enforcing role alternation can reject). The pre-PR splice
guarded against this case by riding queued items inside the tool
result envelope; that guard is what motivated the original
UserInterjection design, so the fix lane needs a deliberate decision
rather than a quick patch. Sleeping on it.
* **q-1** (delete dead ``UserInterjection`` class + tests). Held for
the bug-1 decision — if the chosen fix is to resume the splice for
the ``user_feedback``+queue coexistence case, the advisory shape
stays load-bearing. Class now carries a docstring note marking it
retained-pending-decision so a passing reader doesn't grep for
producers and assume it's actually dead.
Apply-pass content:
* ``q-2``: drop "queued user interjections" from the persistent-
advisory parenthetical in ``send``'s tool-result loop comment;
rewrite to point at ``_flush_queued_messages`` for the queue path.
* ``q-3``: ``__init__`` channel-routing comment loses "and
``UserInterjection``" — only ``GuardAdvisory`` remains.
* ``q-4``: ``_queue_tool_advisory`` docstring + the tool-error nudge
comment lose the user-interjection mentions; the docstring also now
describes the side-channel + ``_apply_reminders_for_provider``
splice path (the actual mechanism).
* ``q-5``: ``AttachmentsNotQueueableError`` docstring rewritten to
describe the post-PR ``_flush_queued_messages`` flow — the
single-combined-turn ``\n\n``-join shape can't carry image / file
blocks, and per-item separate user turns would expand the strict-
template role-ordering surface that the post-batch drain already
balances.
* ``q-6``: the new ``test_queued_message_persists_as_user_row_after_tool_batch``
in ``test_session.py`` now asserts ``stream_idx == 2`` so a future
regression where the post-batch flush runs but the send-loop short-
circuits before the next iteration surfaces in CI rather than
manual repro.
* ``q-7``: ``test_coordinator_page.py`` gets two new string-grep pins
mirroring the existing ``test_app_js.py`` shape — ``content.trim()``
on coord's assistant-replay branch and ``renderAssistantToolBatch``
for the hoisted helper that orders content card before tool batch.
## Test plan
- [x] ``ruff check`` clean
- [x] ``mypy turnstone/`` clean (189 source files)
- [x] Affected test surface (``test_session.py`` +
``test_tool_advisory.py`` + ``test_app_js.py`` +
``test_coordinator_page.py``) — 240 passed
Three independent rehydrate / replay regressions reported on long
multi-turn conversations after the pull-model wake stack landed.
**1. coord history replay rendered tool_calls above the assistant
narration that announced them.**
In ``coordinator.js``'s loadHistory loop, the ``role === "assistant"``
``tool_calls`` branch sat above the role switch — every assistant turn
with both narration AND tool dispatch produced ``[tool batch][content
card]`` in the DOM, even though chronological order is content first.
On a parallel fan-out (e.g. four ``close_workstream`` calls in one
turn) operators saw the assistant text "Let me close them out and
summarize" with NO tool batch between it and the next assistant
message — the four-row batch had been rendered above the announcing
text and was scrolled out of view.
Hoisted the ``tool_calls`` synthesis into a local
``renderAssistantToolBatch(m)``, called from inside the assistant
branch AFTER the content card. Live SSE order (text → dispatch →
results) now matches replay order.
**2. Whitespace-only assistant content rendered as a blank card on
replay.**
Models with vLLM's ``--reasoning-parser`` (Qwen3 in production)
strip ``<think>…</think>`` and emit only the trailing ``"\n\n"`` as
``content`` before a tool call. ``content_parts = ["\n\n"]`` saves
``content = "\n\n"`` to the conversations row. Live the user only
sees ``.msg.reasoning`` (the thinking content) — the empty
``.msg.assistant`` card lives next to it but reads as a thin
divider. On rehydrate the reasoning bubble is gone (not persisted)
and the empty assistant card is the only thing left, surfacing as
"blank cards where the assistant message was."
Both UIs now check ``content && content.trim()`` before rendering
the body — whitespace-only content skips the card entirely instead
of showing a phantom row. Live render unchanged.
**3. Queued user messages disappeared on reconnect.**
PR #474 routed queued user messages into the tool-result envelope
via ``UserInterjection`` advisories — same-turn delivery, but no
persisted user row. On page reload / cross-tab replay the
optimistic ``.msg-queued`` bubble vanished: there was no DB row to
rehydrate it.
Dropped the ``UserInterjection`` splice in ``_collect_advisories``;
the queue drains through ``_flush_queued_messages`` AFTER the tool
batch completes instead. Sequence becomes
``assistant(tool_calls) → tool … tool → user(drained)``, which is
valid for Mistral and Anthropic strict role validators (the only
forbidden shape was user injected mid-batch BEFORE the tool result,
which this still avoids). Persists a real user row → bubble survives
reconnect, and stays in the session's wire-side context window on
the next turn.
## Test plan
- [x] ``ruff check`` clean
- [x] ``mypy turnstone/`` clean (189 source files)
- [x] ``pytest -m "not live"`` — 5798 passed, 3 deselected
- [x] Updated ``test_collect_advisories_does_not_drain_queued_messages``
(was pinning the old UserInterjection shape)
- [x] Added ``test_queued_message_persists_as_user_row_after_tool_batch``
(drives ``send`` end-to-end with a queued message arriving during
the tool batch; asserts the user row lands in self.messages AND
hits ``save_message``)
- [x] Updated ``test_replay_history_renders_content_before_tool_block``
to tolerate the new ``msg.content && msg.content.trim()`` guard
- [ ] Live browser pass on coord (close_workstream parallel fan-out
rehydrates with the 4-row batch BETWEEN the announcing assistant
text and the summary) and interactive (Qwen3 ``"\n\n"`` rows no
longer paint blank cards on reload; queued bubble survives a tab
refresh)
PR #489 review feedback (Copilot + github-code-quality):
- closeSettingsPanel now closes nested revoke modal first on close-button
path (Escape was already handled by the parent keydown trap deferring
to the inner trap; missing-modal-on-close-button was an orphan-modal
hazard).
- _refreshConsentBadge now updates the settings button's aria-label +
title dynamically with the pending-consent count for screen readers
(badge stays aria-hidden — the count is in the label).
- _MAX_INSUFFICIENT_SCOPE_REPORTED promoted to public
MAX_INSUFFICIENT_SCOPE_REPORTED in mcp_http_parsers; drops cross-module
private import in mcp_oauth's /start handler.
- Stale test comment in test_session_mcp_dispatch_error.py corrected:
_exec_read_resource does not log with exc_info=True (bearer-leak
invariant).
- Rejected the protocol-method ellipsis warning: rest of _protocol.py
uses ... consistently per Protocol convention.
Lint:
- ruff format applied to test_mcp_pool_auth_integration.py and
test_mcp_pool_auth_resource_integration.py (combined `with` grammar —
pure formatting).
Flake fix — test_integration_pool_reuse_401_refresh_and_retry_succeeds
on Python 3.11 / resource-constrained CI:
Same cross-task scope hazard f6a3b66 fixed at the close side, surfacing
at the connect side. asyncio.wait_for at mcp_client.py:1206 wraps
streamablehttp_client.__aenter__ in a fresh asyncio.Task. That fresh
task enters anyio cancel scopes, completes, and dies. The eventual
stack.aclose() during eviction or auth_401 retry runs from a different
task and tries to exit scopes whose entering task is dead — anyio
raises RuntimeError, the wedged anyio state blocks the retry's stack
teardown + reconnect, and the call exceeds the 15s budget on slow
workers.
Fix: replace asyncio.wait_for with `async with asyncio.timeout(...)` so
the streamablehttp_client.__aenter__ runs in the dispatch task itself,
no fresh-task scope ownership. Aligns with invariant 18 (asyncio.timeout
not asyncio.wait_for for any SDK / AS / pool-loop await crossing anyio
scopes).
Static path (_connect_one) at lines 905 and 1000 deliberately retains
asyncio.wait_for — auth_type ∈ {none, static} is byte-identical
(invariant 1) and the narrow connect-once / no-eviction-then-reuse
pattern doesn't trigger the cross-task hazard. Anchor comments pin
both directions: a future migration there would break invariant 1; a
future revert at 1206 would re-introduce the flake.
The cited test is the symptom (non-deterministically times out under
load), not a structural gate (no deterministic asyncio.timeout
assertion exists). The comment block at line 1206 records this so a
maintainer who reverts and finds green on a fast machine doesn't
conclude the fix is unneeded.
Verified on Python 3.11.14 (/tmp/venv311) and 3.13.7 (.venv): ruff
format clean, ruff check clean, mypy clean. 368 unit tests + 30 pool
integration tests pass on both interpreters; the previously-flaky test
passed 20× in isolation on 3.11.
Multi-stage /review (4 finders × verify × dedupe): bug/security/perf
returned zero findings; quality returned 3 confirmed minor/nit items
all of which are applied here (q-1 anchor comments at 905+1000, q-2
symptom-vs-gate clarification at 1206, q-3 module-docstring sentence
in mcp_http_parsers).
Wires the structured-error envelopes produced by Phase 7b's pool
dispatcher (mcp_consent_required / mcp_insufficient_scope /
mcp_*_forbidden / mcp_token_undecryptable_key_unknown /
mcp_oauth_url_insecure) through to the user-facing dashboard, and
adds a per-user settings panel for managing MCP server consents.
Changes
- ``_dispatch_pool_sync`` and ``_dispatch_pool_resource_sync`` wrap
structured-error string returns as ``RuntimeError(json_str)`` via
``_is_structured_error()`` so the session-layer ``except Exception``
branch fires uniformly across tool / resource / prompt dispatchers
(the prompt path's ``isinstance(result, str)`` shortcut works only
because prompts return ``list[dict]`` on success). Without this,
the consent UX silently does not render for tool / resource calls.
- ``_structured_error`` extended with an optional ``consent_url``
field; ``_build_consent_url`` produces ``/v1/api/mcp/oauth/start``
query strings (path-relative; the dashboard appends ``return_url``
at click time). Wired to all 12 ``mcp_consent_required`` and the
``mcp_insufficient_scope`` emit sites.
- New endpoints ``GET /v1/api/mcp/oauth/connections`` and
``DELETE /v1/api/mcp/oauth/connections/{server_name}`` registered
on both ``turnstone-server`` and ``turnstone-console``. The DELETE
handler runs local delete + audit + 204 first, then schedules the
RFC 7009 upstream revoke as a fire-and-forget ``asyncio.create_task``
with strong-ref tracking via ``_revoke_upstream_tasks`` (mirrors
the ``_pg_refresh_drain_tasks`` pattern). Soft cap of 256 concurrent
in-flight revokes prevents pile-up under coordinated mass-revoke;
the audit detail records ``upstream_revoke_outcome`` as
``scheduled | no_refresh_token | no_http_client | shed_by_cap``.
- ``ASMetadata`` extended with ``revocation_endpoint`` parsed from
RFC 8414 metadata. ``revoke_token_at_as`` helper posts the form
body under ``asyncio.timeout`` (not ``asyncio.wait_for``) and
never raises; ``_attempt_upstream_revoke`` is wrapped in an outer
``try/except Exception`` so unhandled exceptions don't surface as
``Task exception was never retrieved``.
- ``/v1/api/mcp/oauth/start`` accepts an optional ``scopes=`` query
param; tokens are validated against RFC 6749 §3.3 grammar via
``is_valid_scope_token`` (promoted to ``mcp_http_parsers``),
capped at ``_MAX_INSUFFICIENT_SCOPE_REPORTED`` (32), and unioned
with the configured server scopes for the step-up consent flow.
- Storage primitive ``list_mcp_user_token_metadata_by_user`` projects
the metadata columns at the SQL boundary so ciphertext blobs never
cross the wire on the settings-list path. New
``MCPUserTokenMetadataRow`` TypedDict in ``_protocol.py``;
``MCPTokenStore.list_user_token_metadata`` re-types to the existing
``MCPUserTokenMetadata`` shape.
- Dashboard renderer (``app.js``): ``tryParseMcpError`` detects the
envelope shape on ``tool_result`` SSE events with ``is_error=True``
and ``buildMcpErrorEmbed`` renders an action card mirroring the
existing ``buildMediaEmbed`` pattern. Three categories: actionable
(consent_required / insufficient_scope) with a ``Connect`` button
that opens ``/v1/api/mcp/oauth/start`` in a popup with a scheme
guard, forbidden (mcp_*_forbidden) with a static notice, operator
(key-mismatch / url-insecure) with an operator-action notice.
- New gear button in the appbar opens an MCP-connections settings
modal driven by ``loadMcpConnections`` / ``confirmRevokeMcp``
(two-step revoke confirmation matching the existing delete-ws
pattern). Pending-consent badge tracks unresolved consent prompts
in this tab; cleared after the connections list returns. Console
proxy collision-checked: the IIFE only prepends a node-id pill to
``header.firstChild``, so the right-anchored gear button is safe.
Bearer-leak invariant
- No ``exc_info=True`` on any new path that can carry a chained
``httpx.Request`` (revoke handler, dispatch sites, exec sites).
The two pre-existing ``exc_info=True`` calls in
``_exec_read_resource`` / ``_exec_use_prompt`` were replaced with
structured-field logs as a Phase 8 sibling fix.
Tests
- 440 pytest passes on both Python 3.13 (.venv) and 3.11
(/tmp/venv311); ruff + mypy clean.
- 5 new test files: ``test_mcp_consent_url_sibling_audit`` (structural
gate that every ``code="mcp_consent_required"`` / ``mcp_insufficient_scope``
site carries ``consent_url=``), ``test_mcp_oauth_connections``,
``test_mcp_oauth_revoke``, ``test_mcp_token_store_metadata``,
``test_session_mcp_dispatch_error``.
- End-to-end regression coverage for the bug-1 sibling pattern:
``test_call_tool_sync_raises_on_structured_error_envelope``,
``test_read_resource_sync_raises_on_structured_error_envelope``,
``test_get_prompt_sync_raises_on_structured_error_envelope``, plus
``test_call_tool_sync_does_not_wrap_non_structured_string`` as the
defensive gate (only ``mcp_*`` envelopes are wrapped).
Hard invariants honored
- Static path byte-identical for ``auth_type ∈ {none, static}``: the
wrap fires only when the dispatcher returns a structured-mcp-error
string, which only happens on the oauth_user pool path.
- ``asyncio.timeout`` (not ``asyncio.wait_for``) on every new
AS / SDK / pool-loop await per Python 3.11 anyio cancel-scope
hazard.
- Scope cap ``_MAX_INSUFFICIENT_SCOPE_REPORTED = 32`` enforced at
every output / merge site.
- Cross-user isolation on the revoke endpoint: a non-owner DELETE
returns 404 with the same body shape as a never-existed row;
``http_client_mock.post.assert_not_called()`` pins this in 3 tests.
Deferred (not Phase 8 blockers)
- perf-2 (``asyncio.gather`` parallelisation in revoke handler) —
superseded by perf-1's fire-and-forget pattern.
- q-4 (prompt-path ``isinstance(str)`` vs sibling ``_is_structured_error``
asymmetry) — already documented in the function docstring.
- q-9 (``_pendingConsentServers`` → ``_serversNeedingConsent``
rename) — pure naming taste.
Apply sanitize_text() to the new _source and _reminders columns in
both save_message and save_messages_bulk on SQLite + PostgreSQL,
mirroring the existing pattern used for content and provider_data.
Producers (sanitize_payload on the watch dispatch path,
format_nudge constants on the standard nudge path) already strip
NUL bytes today so nothing in production reaches this clamp — but
the storage layer is opaque to those invariants, and PostgreSQL
TEXT columns reject NUL outright. Without this clamp, a future
producer that forgets sanitize_payload (or hand-builds the column
string) hard-fails the chat-loop persist path on PostgreSQL.
Cost is negligible — sanitize_text early-exits on the common
no-NUL case via 'if value and "\x00" in value'.
Surfaced by Copilot's PR #486 review.
Closes round-2 review finding q-7 (nit).
The kwarg was added to close round-1 perf-2 cosmetically — the
storage backend's signature already accepted ``limit``, but the
single in-tree caller (``ChatSession.resume``) doesn't pass it and
other tail-load consumers go direct to ``storage.load_messages``.
Adding signature surface to mark a perf finding closed without an
actual consumer is API-surface bloat.
When a tail-load consumer is written (e.g. a heuristic in
``session.resume`` to skip ancient wake rows), the kwarg can come
back — at that point with a real caller driving the contract.
Closes round-2 review findings q-6 (nit) and perf-1 (nit).
* **q-6:** ``_WATCH_REMINDER_OPTIONAL_KEYS`` carried a leading
underscore (Python's module-private convention) but was imported
from two other modules — clearly a public contract between
``build_watch_reminder`` and its consumers
(``ChatSession._dispatch`` + ``server._build_history``). Drop the
underscore so the import sites match the constant's documented
cross-module role.
* **perf-1:** The dispatch closure imported the constant inside its
body, paying ``IMPORT_NAME`` + ``IMPORT_FROM`` bytecode on every
watch fire. ``server.py`` already imports at module scope; hoist
the same way in ``session.py``. Microsecond savings per dispatch,
but the in-closure form was just an oversight from the apply-pass.
Closes round-2 review findings q-1 (minor), q-3 (nit), q-4 (nit), q-5
(nit).
* **q-1:** Drop the ``post-migration 050`` clause from the fork-block
comment — the apply-pass relocated rather than removed the
tombstone-style temporal reference round-1 q-2 was supposed to fix.
The bulk-row dict shape and ``_encode_reminders`` are
self-explanatory; the WHY is pinned by
``test_fork_preserves_source_and_reminders``.
* **q-3:** Replace ``DOES persist now`` framing on the wake-row save
comment with a present-tense invariant. The ``now`` implies the
reader knows the prior state, same family as the temporal
tombstones.
* **q-4:** Trim the 12-line WHAT-narration block above the
resume-time ``_reminders_delivered = True`` loop to two lines
stating the WHY only. The new regression test pins the contract.
* **q-5:** Reframe ``test_fork_preserves_source_and_reminders``
docstring as a forward-looking invariant; drop the
``Dropping them was the original bug`` and ``post-migration 050``
fix-narration.
Project convention: invariant statements, present tense; don't
reference the current task / fix / migration number.
Closes round-2 review findings bug-1 (minor) and q-2 (minor).
* **bug-1:** ``_encode_reminders`` clamped each entry's ``text`` field
with Python ``str`` slicing, which counts codepoints. Multi-byte
UTF-8 input (CJK, emoji) could land 4 bytes per character past the
cap, defeating the row-width / FTS5-index protection by up to 4x.
Switch to UTF-8 byte clamping with ``errors="ignore"`` on the
decode boundary so a slice mid-codepoint drops the partial
character cleanly.
* **q-2:** Both the constant block-comment and the ``_encode_reminders``
docstring referenced ``docs/design/watch-card-ux-briefing.md`` —
local-only per project convention (``feedback_no_design_doc_commits``)
so the canonical repo reads as a dead reference. The cap value
stands by itself; the row-width / FTS5 WHY is enough.
Closes round-1 review findings q-2 (minor), q-5 (minor), q-6 (nit), q-7
(nit), sec-1 (nit), perf-4 (nit).
* **q-5:** Export ``_WATCH_REMINDER_OPTIONAL_KEYS`` from
``turnstone/core/watch.py`` and import in the dispatch closure
(session.py) and the replay filter (server.py:_build_history). The
three-place duplication of the literal tuple
``("watch_name", "command", "poll_count", "max_polls", "is_final")``
is gone; future field adds touch one constant.
* **sec-1:** Run ``sanitize_payload`` over string-typed metadata fields
(``watch_name`` / ``command``) before they enter the queue. Today's
consumers all use ``textContent``, but the asymmetry — sanitised
``text`` alongside unsanitised metadata — would survive forever in
DB rows and resurface if a future consumer used a non-textContent
sink (aria-label, copy-to-clipboard, markdown render).
* **q-7:** Drop the per-iteration ``isinstance(reminder, dict)`` from
the dispatch closure's metadata comprehension. By the time the
block runs, ``text = reminder.get("text", "") if isinstance(...)``
+ the ``if not sanitized: return`` guard above already established
``reminder`` is a non-empty dict.
* **q-2:** Strip tombstone-style references — "post-#482", "post-#484",
"Step 7 of the watch-card UX plan", "Post-Step-7 dispatch surface",
and the brittle line-anchor "session.py:2685-2686" — across
``session.py``, ``test_session.py``, ``test_watch.py``,
``test_watch_dispatch.py``, ``test_watch_integration.py``. Comment
intent preserved; historical anchors gone.
* **q-6:** Drop the ``del source`` line in ``cli.py``'s
``on_user_reminder``; the parallel ``on_tool_reminder`` ignores
``tool_call_id`` without ``del`` and the comment alone is enough.
* **perf-4:** Document the SQLite ``render_as_batch=True`` recreate
cost in migration 050's docstring — first deployment after upgrade
copies the conversations table twice (one per ``add_column``).
PostgreSQL is unaffected.
5734 non-live tests pass; ruff + mypy clean.
Closes round-1 review findings q-3 + q-4 (minor, merged) and bug-3 + bug-4
(nit, merged).
* **q-3 + q-4:** The new ``.msg.user-reminder .msg-body { white-space:
pre-wrap }`` rule was a no-op on the interactive UI because that
frontend's ``_buildDefaultReminderBubble`` appended label + text spans
directly to the outer ``.msg.user-reminder`` element with no
``.msg-body`` wrapper. Coord rendered the same shape with a wrapper.
The two implementations diverging on DOM structure also meant a
shared-helper extraction was harder than necessary. Reconciled by
wrapping interactive's spans in ``.msg-body`` to match coord; the CSS
rule now applies to both UIs and the shared-extraction follow-up to
``shared_static/cards.js`` is mechanical (deferred per the review
report — out of scope for this commit).
* **bug-3 + bug-4:** The reminder anchor lookup ``.msg.user`` also
matched ``.msg.user.system-nudge`` markers because the marker carries
both classes. A non-wake reminder fired between a wake marker and
the next real user message would anchor below the wake marker rather
than the previous real user message. Edge case (``/history`` reload
corrects), but the fix is mechanical: change the selector to
``.msg.user:not(.system-nudge)`` in both files.
Closes round-1 review finding perf-2 (minor).
Storage backends accept ``*, limit: int | None = None`` (see
:meth:`StorageBackend.load_messages` at storage/_protocol.py:146) but
the in-memory wrapper at memory.py:82-85 dropped the kwarg, so
callers that wanted to tail-load (e.g. ``session.resume`` against a
long-running coord with hundreds of wake rows + persisted reminder
JSON) were forced to pull every row through the wrapper anyway.
Wraparound is mechanical: signature widens, default leaves existing
callers unaffected.
Closes round-1 review finding q-1 (major).
The comment block above ``self._attach_pending_user_reminders(user_msg)``
asserted that reminders "stay in-memory only and don't persist across
reloads" — directly contradicted by the comment block immediately below
(at the save_message call site) that explains the new persistence
semantics, plus the actual code that now writes ``_source`` and
``_reminders`` to the conversations row. Future readers hitting both
blocks would lose trust in the surrounding comments.
The lower block already documents the persistence contract, so the
upper block is just deleted rather than rewritten.
Closes round-1 review findings bug-2 (major), perf-1 (minor), perf-6 (nit).
* **bug-2:** ``ChatSession.resume(..., fork=True)``'s bulk-row builder
silently dropped the ``_source`` and ``_reminders`` side-channel
data the source workstream had persisted via ``_append_user_turn``.
Both backends' ``save_messages_bulk`` already accept these keys
(the columns exist post-migration 050) — the bulk builder just
didn't supply them. The fork's resumed transcript would then look
like the assistant turn answered out of nowhere: every wake marker
and every reminder bubble that survived to disk on the source got
dropped on the fork. New regression test
``test_fork_preserves_source_and_reminders`` pins the contract.
* **perf-6:** Extracts ``_encode_reminders(reminders) -> str | None``
near ``_apply_reminders_for_provider`` so the user-turn save path,
the tool-turn save path, and the new fork bulk builder share one
encoder. Eliminates the drift risk between three near-identical
``json.dumps(..., separators=(",", ":")) if X else None`` patterns.
* **perf-1:** The new helper clamps each entry's ``text`` field at
``REMINDER_TEXT_STORAGE_CAP = 8192`` characters before encoding so
a single rogue producer (a watch streaming unbounded shell output,
a corruption-class steering payload) can't blow the conversations
row width or the FTS5 index. The in-memory side-channel keeps the
full body — only the persisted JSON is clamped. Mirrors
``TOOL_RESULT_STORAGE_CAP`` on tool result rows.
5734 non-live tests pass; ruff + mypy clean.
Persisted ``_reminders`` survive ``load_messages`` but the in-memory
``_reminders_delivered`` flag does not (it's session-scoped — set by
``_mark_reminders_delivered`` after each successful provider stream,
never persisted alongside the JSON column). Without a re-splice
guard at resume time, ``_apply_reminders_for_provider`` would walk
every loaded message, see ``_reminders`` set + the flag falsy, and
splice every historical ``<system-reminder>`` envelope onto the wire
on the very next user turn — leaking each reminder a second time, the
turn after it had already advised.
Mirror the post-stream hook in ``resume()``: every loaded message
that carries reminders has already been delivered (it survived to
disk), so flag it accordingly so ``_apply_reminders_for_provider``
short-circuits on the pass-through path.
Test pins the contract end-to-end — stage a workstream with a
persisted reminder, resume into a fresh session, append a live user
turn, run the wire transform, and assert the historical reminder
body does NOT land in the rendered output.
User-visible slice of the watch-card UX workstream — combines the
replay-path widening, both frontend renderers, the CSS, and the
cross-cutting Python tests.
server._build_history widens the reminder filter from {type, text} to
project on a known set of optional fields (watch_name, command,
poll_count, max_polls, is_final) and surfaces _source as
entry["source"] when set. The known-key filter narrows the blast
radius if a future producer accidentally stuffs sensitive fields
into the dict.
SessionUIBase.on_user_reminder takes a new source: str | None kwarg
that rides on the SSE event when set. _attach_pending_user_reminders
forwards user_msg["_source"] so non-originating tabs see the wake's
"system_nudge" tag and render the thin marker. Protocol + cli + eval
implementations widen accordingly.
Frontend (coordinator.js + app.js — touched in lockstep per project
memory's "logic that lands in BOTH UIs must touch both files"):
* Branch on r.type === "watch_triggered" for a structured
.msg.watch-result card with header / $ command / <pre> body /
poll N/M [· final] footer.
* New addSystemNudgeMarker (interactive) + appendSystemNudgeMarker
(coord) renders a thin .msg.user.system-nudge anchor for
wake-driven reminders, both live (source === "system_nudge" on the
SSE event) and replay (msg.source === "system_nudge").
* Default .msg.user-reminder rendering preserved for every other
metacog nudge type.
CSS (shared_static/chat.css):
* New .msg.watch-result rules — full-width treatment, cyan accent,
monospace body with word-break: break-word for mobile.
* New .msg.user.system-nudge rule — thin yellow marker.
* Bonus newline-collapse fix: .msg.user-reminder .msg-body now sets
white-space: pre-wrap so multi-line shell output / bulleted lists
stay readable inside the advisory bubble.
Plan reference: docs/design/watch-card-ux.md §4 Steps 9-12 + bonus
CSS §11 (Commit 4).
WatchRunner._dispatch_result now takes a structured reminder dict
produced by build_watch_reminder() — text matches format_watch_message
verbatim (so compaction / channel adapters / wire splice keep their
behaviour), and watch_name / command / poll_count / max_polls /
is_final ride alongside as queue-entry metadata.
The dispatch closure registered in ChatSession.set_watch_runner pulls
the optional fields out of the dict and passes them to enqueue via
the new metadata kwarg. Drain seams already merge metadata into the
rendered reminder dict (Commit 2), so the SSE event for a watch fire
now carries the structured fields without further plumbing.
* turnstone/core/watch.py — new build_watch_reminder() helper, _poll_watch
switches from format_watch_message + dispatch(str) to build_watch_reminder
+ dispatch(dict). set_dispatch_fn / get_dispatch_fn / restore_fn
signatures widen from Callable[[str, str], None] to
Callable[[dict[str, Any], str], None].
* turnstone/core/session.py — dispatch closure builds the metadata dict
via {k: reminder[k] for k in ("watch_name", "command", ...) if k in reminder}
and passes it to nudge_queue.enqueue.
* tests/test_watch.py — new TestBuildWatchReminder class pinning the
builder shape; existing dispatch_fn_registry / restore_fn tests
updated to dict shape.
* tests/test_watch_dispatch.py — every dispatch(...) call updated to
pass a structured reminder dict via _reminder() helper; new
TestMetadataPropagation class pins the metadata-on-enqueue contract.
* tests/test_watch_integration.py — _dispatch_result calls updated to
dict shape.
Plan reference: docs/design/watch-card-ux.md §4 Step 7 + Step 8 watch-test
subset (Commit 3).
Producers (today only watch_triggered) can now attach a metadata dict
to a queued nudge so the rendered reminder dict on the user/tool side
carries fields beyond {type, text}. Wire shape stays additive: the
SSE event picks up the optional fields when present, and producers
without metadata leave it None.
* _Entry grows from 4 fields to 5 — metadata: dict[str, Any] | None.
* enqueue accepts metadata=... as a kwarg.
* drain returns list[tuple[str, str, dict | None]] (was 2-tuples).
* pending stays narrow at (type, text) for legacy callers; new
pending_with_metadata projects the third slot for tests that need
to assert producer-specific fields.
* Three drain consumers in session.py — _collect_advisories,
_attach_pending_user_reminders, deliver_wake_nudge_from_queue —
unpack the new 3-tuple shape and merge metadata into each
reminder dict.
* on_user_reminder / on_tool_reminder protocol signatures widen
from list[dict[str, str]] to list[dict[str, Any]] across
ChatSession.UI, SessionUIBase, CLI, eval harness.
Plan reference: docs/design/watch-card-ux.md §4 Step 6 + Step 8 _Entry
subset (Commit 2).
Adds two TEXT-NULL columns to the conversations table so multi-tab /
multi-device replay sees the same metacognitive bubble shape the
originating tab saw live. Until now, reminders lived only on the
in-memory ChatSession.messages dict, and the wake-driven empty user
turn was not persisted at all (skip at session.py:2685-2686) — a
second tab connecting via /history saw the assistant turn with no
preceding wake context, and missed every other tab's reminder
bubbles besides.
Single Alembic revision 050 (head was 049) adds:
* conversations._source — today only "system_nudge" for wake rows
* conversations._reminders — JSON-encoded reminder list
Both backends (sqlite + postgresql) thread the columns through
save_message / save_messages_bulk / load_messages. reconstruct_messages
unpacks the row tuple as 9 elements (was 7), JSON-decoding _reminders
on the user AND tool branches with the same contextlib.suppress guard
the existing provider_data / tool_calls decode uses. Tool-row
reminders ride the same column so tool_error / repeat replay shape
matches user-channel parity.
session.py:2685-2686 wake-row persist skip is dropped; _append_user_turn
JSON-encodes user_msg["_reminders"] and passes both source + reminders
to save_message. The tool-message save site at session.py:3014-3020
mirrors with metacog_reminders.
Plan reference: docs/design/watch-card-ux.md §4 Steps 1-5 (Commit 1).
Address Copilot review feedback on PR #487:
1. **Atomic commit invariant**: ``_bootstrap_coord_subsystem`` previously
stamped ``coord_mgr`` ~50 lines before the final ``coord_registry``
commit, and started threads + subscriptions in between. A concurrent
dashboard request running through ``_require_coord_mgr`` during the
runtime-bootstrap window could observe ``coord_mgr`` set with
``coord_registry`` still ``None`` and surface the misleading
"Restart the console after adding a model definition" 503.
Refactored to two phases: (a) build everything as locals, (b) start
side-effects (StateWriter / observer / nudge watcher / child fan-out
/ cleanup thread), then atomic commit at the end with ``coord_mgr``
stamped LAST. The build-phase ``try/except`` rolls back any started
side-effects from local handles before re-raising — no daemon thread
or subscription leaks across retries, and ``app.state`` is never
stamped on a partial failure.
2. **Class-attr cleanup symmetry**: ``_teardown_partial_coord_subsystem``
now also clears ``ConsoleCoordinatorUI._coord_mgr`` /
``_collector`` / ``_console_metrics`` to match the lifespan shutdown
path (server.py ~line 4629). A failed bootstrap (or test teardown
reuse) no longer leaks process-global pointers at a half-built
subsystem.
3. **Lifespan startup offload**: the lifespan startup error path used
to call ``_teardown_partial_coord_subsystem`` synchronously, which
in turn calls ``StateWriter.shutdown(timeout=2.0)`` — a thread-join
+ sync DB writes that could block the event loop for up to 2s
while the console is still coming up. Wrapped the whole
load-and-bootstrap in ``asyncio.to_thread`` via the new
``_load_and_bootstrap_coord_subsystem`` synchronous helper, so all
blocking work (including any rollback) runs on a worker thread.
Mirrors the pattern the regular lifespan shutdown (line ~4620) and
the runtime CRUD-triggered path already use.
Tests:
- ``test_bootstrap_atomic_commit_no_partial_visibility``: a polling
thread in tight loop watches ``coord_mgr`` / ``coord_registry``
during a real bootstrap and asserts no observation has ``coord_mgr``
set with ``coord_registry`` still ``None``.
- ``test_real_bootstrap_rolls_back_partial_state_on_side_effect_failure``:
monkeypatches ``install_idle_nudge_watcher`` to raise mid-build,
asserts ``app.state`` shows the clean fresh-install state and the
builder-failure error string surfaces ``RuntimeError`` (not the
stale "no models" boot-time message).
A freshly-installed console with no model rows in the DB at boot
caught the ``ValueError`` from ``load_model_registry()`` in the
lifespan and skipped the entire coord subsystem build, leaving
``coord_mgr`` ``None``. ``_refresh_coord_registry`` then bailed
out at ``existing is None`` rather than building the subsystem on
first model add — operators had to restart the console after
configuring their first model in the admin panel for the
"Coordinator subsystem not initialized" banner to clear.
Extract the lifespan's coord build into a reusable
``_bootstrap_coord_subsystem`` and add ``_maybe_bootstrap_coord_subsystem``
that runs as an ``asyncio.to_thread`` follow-on after every admin
model-CRUD endpoint (create/update/delete/reload). The helper:
- fast-paths to a no-op when ``coord_mgr`` is already set;
- guards concurrent first-install attempts with
``_COORD_BOOTSTRAP_LOCK`` + double-checked re-test inside the lock;
- pre-computes config-derived integers BEFORE any thread starts so
``int(config_store.get(...))`` failures don't strand a started
``StateWriter`` daemon;
- stamps ``coord_state_writer`` to ``app.state`` immediately after
``.start()`` so the new ``_teardown_partial_coord_subsystem`` can
shut it down on a partial failure (no thread leaks across retries);
- atomically commits ``coord_registry`` + clears
``coord_registry_error`` as the final step so callers can rely on
the invariant ``coord_registry`` is set iff ``coord_mgr`` is set;
- replaces the stale boot-time "no model definitions" message with
a builder-failure-specific diagnosis (carrying ``type(exc).__name__``)
on construction failure so the dashboard's 503 banner reflects the
actual cause.
Both the lifespan path and the runtime-bootstrap path now route
through the same helper and the same teardown on failure.
Tests: 12 new tests covering the helper-level wiring (idempotent
fast-path, missing-prereq parametrised over ``config_store`` /
``collector`` / ``console_metrics``, no-rows error recording, builder
failure error replacement, partial-state teardown), the endpoint
integration, the deterministic concurrent-call lock test (uses an
instrumented lock wrapper that signals when a second acquirer arrives,
so the test fails fast on slow CI rather than depending on a
wall-clock sleep), and a real-builder end-to-end case constructing a
working ``SessionManager`` against a real ``ConfigStore`` + real
``ClusterCollector``.
Two of five Copilot comments on PR #485 were valid; this commit applies
both. The other three (one duplicate of comment 1, plus the INFO-logging
and `_pending`-naming nits) get rationale on-thread and resolution.
1. emit_oauth_failure_audit action now derived from `code` (#485 bug-1)
The Phase 7b refactor generalized `emit_insufficient_scope_audit` →
`emit_oauth_failure_audit`, routing both `mcp_insufficient_scope` AND
generic-403 (`mcp_*_forbidden`) through the same helper. The audit
`action` field stayed hardcoded as
`"mcp_server.oauth.insufficient_scope_emitted"`, mislabeling generic
forbidden events under the insufficient_scope bucket — downstream
alerting / analytics filtering on `action` would silently fold both
categories together.
The action is now selected from `code`:
* `mcp_insufficient_scope` →
`mcp_server.oauth.insufficient_scope_emitted` (preserves existing
alerting consumers)
* `mcp_tool_call_forbidden` / `mcp_resource_read_forbidden` /
`mcp_prompt_get_forbidden` →
`mcp_server.oauth.forbidden_emitted` (new, distinct label)
Detail row continues to carry both `code` and `kind` so operators get
sub-bucket distinction within either action.
2. Resource-listener docstrings cite RFC §3.2 (#485 doc-1)
Per the codebase convention established in Phase 7b round-1 q-1
(`_rebuild_user_prompt_map` corrected §3.2 → §3.3 because prompts are
§3.3 in the MCP spec), resource-related docstrings should cite §3.2.
The three resource-listener docstrings were citing §3.3, and the
"Mirrors `_notify_listeners` for tools (RFC §3.3)" parenthetical in
both `_notify_resource_listeners` and `_notify_prompt_listeners` read
as "tools are at §3.3" — confusing twice over. All four sites now
carry the correct catalog-kind citation explicitly:
* resource-listener docstrings → "RFC §3.2 (resources)"
* prompt-listener docstrings → "RFC §3.3 (prompts)"
Tests / lint:
* 119 passed on 3.13 + 3.11 (targeted MCP OAuth pool tests)
* ruff + mypy clean on both files
Extends the Phase 7 per-(user, server) ClientSession pool to cover
RFC §3.2 (resources/read) and §3.3 (prompts/get) on the same shape
already proven for tools/call. Pool discovery is capability-gated so
servers without resources/ or prompts/ stay free of extra round-trips.
API additions / widenings (MCPClientManager):
- ``read_resource_sync(uri, *, user_id=None, timeout=120)`` —
per-user-first dispatch; falls through to the byte-identical static
path when ``user_id`` is None or the URI doesn't resolve to an
``oauth_user`` pool entry.
- ``get_prompt_sync(prefixed_name, arguments=None, *, user_id=None,
timeout=30)`` — same dispatch shape; structured-error responses
surface via ``RuntimeError`` so the agent-loop's ``except Exception``
block renders the JSON without polluting the prompt-protocol return
shape.
- ``get_resources(user_id=None)`` / ``get_prompts(user_id=None)`` —
per-user merged catalogs (admin/global call still passes None).
- ``add_{resource,prompt}_listener`` /
``remove_{resource,prompt}_listener`` — ``user_id`` keyword scopes
the listener so a pool-only catalog change for one user does not
wake another user's session.
- ``resource_count_for_user(user_id=None)`` /
``prompt_count_for_user(user_id=None)`` — method-form variants used
by ChatSession's ``read_resource`` / ``use_prompt`` tool gating; the
legacy ``resource_count`` / ``prompt_count`` properties remain
static-only for admin paths.
- ``_dispatch_pool_resource`` / ``_dispatch_pool_prompt`` async coros
— mirror ``_dispatch_pool`` for the new SDK calls; share the
carrier-race-and-cancel core via ``_dispatch_pool_with_entry_call``.
- ``_handle_auth_403`` extended with ``kind=Literal["tool",
"resource", "prompt"]`` so the per-operation ``mcp_*_forbidden``
code surfaces (kind="tool" remains the default for back-compat).
- Pool notification handler now refreshes resources / prompts on
``ResourceListChangedNotification`` / ``PromptListChangedNotification``
via ``_refresh_pool_server_resources`` / ``_refresh_pool_server_prompts``.
ChatSession (``turnstone/core/session.py``) call-site updates:
- 12 sites threaded the session-bound ``user_id`` through
``add_*_listener`` / ``remove_*_listener``, ``get_resources`` /
``get_prompts``, gating, ``read_resource_sync`` /
``get_prompt_sync``, and ``is_mcp_prompt`` so the per-user merged
catalog drives both the visible-tool set and dispatch.
- ``/mcp`` slash command now lists this user's pool resources and
prompts alongside tools (Phase 7 already scoped tools).
Scope decisions:
- Per-user-first URI ordering (decision 0.1): the dispatcher attempts
the user's pool catalog first, falling back to the static catalog
only when no pool entry resolves the URI / prefixed name. Pool-only
users never see the static catalog leak into their resolution.
- Method-form ``*_count_for_user`` (vs property) keeps the legacy
``resource_count`` / ``prompt_count`` properties intact for admin
endpoints whose contract is "static catalog size only".
- Shared ``_dispatch_pool_with_entry_call`` helper accepts an
``sdk_call: Callable[[ClientSession], Awaitable[Any]]`` closure,
keeping the entry-locked carrier-race / classification / retry
plumbing single-source instead of a 3x copy across tool / resource
/ prompt paths.
R6 (anyio uniformity): every pool-side list / read / get path uses
``async with asyncio.timeout(...)`` — ``asyncio.wait_for`` is
forbidden in those paths because it wraps the inner awaitable in a
fresh task and surfaces ``CancelledError`` from inside
``streamablehttp_client``'s anyio TaskGroup on Python 3.11
(per ``feedback_asyncio_timeout_vs_wait_for.md``).
Tests:
- ``test_mcp_pool_auth_resource_integration.py`` — 9 real-transport
resource tests (FastMCP upstream + ``BehaviorMiddleware``):
401-refresh-retry success, persistent 401 -> consent_required,
403+insufficient_scope, 403 generic -> mcp_resource_read_forbidden,
breaker-isolation under repeated auth failures, missing-token,
decrypt-failure, http:// URL guard, unknown-URI ValueError.
- ``test_mcp_pool_auth_prompt_integration.py`` — 9 mirror tests for
the prompt path; structured-error responses verified via
``RuntimeError`` payload shape.
- ``test_mcp_user_catalog.py`` — extended unit coverage for per-user
resource / prompt rebuild + collision policy + symmetric eviction.
- ``test_sessions.py::TestMCPToolGating`` — pool-only-user canary
asserts ``read_resource`` / ``use_prompt`` stay visible when the
static catalog is empty but the user has pool entries.
Round-1 review fixes (4-finder review applied, no push yet):
- bug-1: ``_exec_use_prompt`` was hardcoding ``"MCP prompt error: failed
to invoke prompt"`` — discarding the structured-error JSON that
``_dispatch_pool_prompt_sync`` raises via ``RuntimeError``. Now uses
``f"MCP prompt error: {e}"`` mirroring ``_exec_mcp_tool``; pool-prompt
consent_required / insufficient_scope / forbidden errors now reach
the LLM as intended.
- bug-2 + bug-3: resource template discovery was uncapped —
``_cap_server_resources`` covered ``res_result.resources`` but the
separate ``tmpl_result.resourceTemplates`` loop appended every
template a server returned. Added ``_MAX_RESOURCE_TEMPLATES_PER_SERVER``
(1000) + ``_cap_server_resource_templates`` helper, applied at both
the initial discovery site (``_connect_one_pool``) and the refresh
site (``_refresh_pool_server_resources``). Mirrors the existing
``_MAX_TOOLS_PER_SERVER`` / ``_MAX_PROMPTS_PER_SERVER`` defensive
ceilings.
- sec-1 + sec-2: ``emit_insufficient_scope_audit`` generalized to
``emit_oauth_failure_audit(kind, code, ...)``, called from both the
insufficient_scope branch AND the previously-silent generic 403
branch. Audit detail now records ``{"kind": kind, "code": code,
"scopes_required": [...]}`` so operators can distinguish tool-call
vs resource-read vs prompt-get 403s in audit logs and so cross-
tenant probing on the generic 403 path leaves a trail. The Phase 7
inherited gap (``mcp_tool_call_forbidden`` had the same silence) is
closed in the same refactor.
- perf-1: pool resource discovery now uses ``asyncio.gather(
list_resources, list_resource_templates)`` inside the existing
``async with asyncio.timeout(...)`` budget — disjoint catalogs, no
ordering dependency. Typical-case 2-RTT cold-connect resource block
collapses to 1-RTT. Same change applied at ``_refresh_pool_server_resources``.
- q-1: ``_rebuild_user_prompt_map`` docstring corrected RFC §3.2 →
§3.3 (resources are §3.2; prompts are §3.3).
- q-2: ``_refresh_pool_server_prompts`` docstring now carries the
R6 / mcp-loop note that the resource sibling already had — both
refresh paths now declare the asyncio.timeout invariant explicitly.
- q-5: added the ``_user_resource_map`` / DB-mismatch guard to
``read_resource_sync`` for parity with ``get_prompt_sync``. A stale
per-user map entry with no matching oauth_user row now raises a
specific ValueError instead of silently falling through to a
generic ``Unknown MCP resource``.
- q-6: ``_dispatch_pool_with_entry`` (now a single-caller wrapper
after the ``_dispatch_pool_with_entry_call`` extraction) gains a
one-line docstring explaining why the wrapper is preserved
(tool-decode localization + stack-trace identity for debugging).
- q-7: added 1 resource + 1 prompt end-to-end integration test that
drive REAL discovery + dispatch in the same connect (no
``_seed_pool_*_map`` shortcuts), mirroring the tool path's
``test_integration_pool_reuse_401_refresh_and_retry_succeeds``.
The seeded-map tests stay (faster, focused on dispatch); the new
e2e tests cover the connect-discover-dispatch composition that
caught Phase 6's carrier-on-entry bug.
Pre-push round-1 review fixes (3-finder review on the final state —
the lesson from Phase 7 round-3's q-1 regression: round-2 catches
what the round-1 apply pass missed):
- q-1 (MAJOR): the bug-1 sibling that round-1 missed —
``_exec_read_resource`` was hardcoding ``"MCP resource error: failed
to read resource"`` while ``_exec_use_prompt`` (post-bug-1) preserved
the structured-error JSON via ``f"... error: {e}"``. The round-1
apply pass patched the prompt side but not the resource side. q-5's
per-user-map / DB-mismatch ValueError was being swallowed at the
agent loop boundary, defeating the operator-diagnostic intent. Now
``_exec_read_resource`` mirrors ``_exec_mcp_tool`` and ``_exec_use_prompt``.
- q-6 (nit): defensive-cap comment block at module-level cited
"(RFC §3.2)" while covering both resource and prompt list paths;
prompts are §3.3. Now reads "(RFC §3.2 for resources, §3.3 for
prompts)" matching the convention the q-1 apply established.
- q-5 (rejected with better justification): the reviewer flagged
``_dispatch_pool_with_entry`` as a single-caller wrapper that should
be inlined. After examination — the autouse fixture
``tests/test_mcp_pool_auth_introspection.py::_install_capture_intercept``
monkeypatches this method to stash ``entry.auth_capture`` for the
fake call_tool stubs in dispatcher-asserting tests. Inlining would
redirect the patch to ``_dispatch_pool_with_entry_call`` (different
kwargs shape) and require re-validating every test that depends on
the interception. The wrapper IS load-bearing; q-6 docstring updated
to cite the test-fixture rationale instead of the thin "stack-trace
identity" claim.
Deferred to follow-up (documented rationale):
- perf-2: single-pass partition for system-message resource list
(concrete vs templates). Sub-microsecond at expected scale;
opportunistic-only.
- q-2 (pre-push): ~200 lines of fixture infrastructure
(``BehaviorMiddleware``, ``_build_server``, ``_seed_oauth_server``,
``running_loop_mgr``, etc.) duplicated across three pool-integration
test files. Real maintenance cost, but a 200-line conftest extraction
is a focused refactor that earns its own commit / PR. Tracking as
follow-up rather than balloon Phase 7b's diff further.
- q-3 / q-4 (refactor): extract shared dispatcher / scheduler
helpers to compress three near-identical 90-line bodies (round-1
q-3 was the same root cause; the pre-push q-3/q-4 reviewer
reaffirmed it concretely). Three named methods preserve readability
for the codebase's hottest correctness path; follow-up if
duplication grows further or if a per-path divergence ships.
- q-4 (round-1, distinct from pre-push q-4): split pool concerns
into ``mcp_pool.py``. Out-of-scope per finder; future refactor as
the file approaches the navigation/merge-conflict threshold.
3.13: 5590 passed (5541 baseline -> +49 net; pre-review +47, q-7
e2e tests added +2). Existing audit-detail tests updated in-place
to expect the new ``kind`` and ``code`` fields.
3.11: 5590 passed (parity gate per ``feedback_pytest_env_parity.md``).
Closes PR #484 review findings (Copilot): the soft-cap pattern in
``ChatSession.set_watch_runner``'s dispatch closure was a non-atomic
two-call pair (``count_by_type`` then ``drop_oldest_by_type``) with
two separate lock acquisitions. A concurrent drain on the worker
thread (``USER_DRAIN`` / ``TOOL_DRAIN`` consuming ``"watch_triggered"``
entries via the ``"any"`` channel) could slip between the two calls,
making the drop a no-op. The dispatch closure also discarded
``drop_oldest_by_type``'s return value and unconditionally logged
``dropped_oldest=True``, so a no-op drop got reported as a successful
drop.
* New ``NudgeQueue.cap_at_or_drop_oldest(nudge_type, max_depth,
channel=None) -> bool`` does the count+drop in a single critical
section. Returns the actual outcome.
* Dispatch closure (``session.py:1410-1416``) now calls the helper and
uses its return value to gate the WARNING log line, so the log is
accurate when a drop did NOT happen.
* ``drop_oldest_by_type``'s docstring no longer overstates the
per-call lock as covering a count+drop pair — it points readers
to ``cap_at_or_drop_oldest`` for that contract.
7 new tests in ``TestCapAtOrDropOldest`` cover: below-cap no-op,
at-cap drop-oldest, above-cap drop-only-one (per-call), channel
filter, other-type isolation, ``max_depth <= 0`` defensive no-op,
no-match.
5708 non-live tests pass; ruff + mypy clean.
The github-code-quality bot finding ("Statement has no effect" on
``_protocol.py:939``'s ``...`` body) is a false positive — every
Protocol method in ``_protocol.py`` uses ``...`` as its body, which
is the canonical Python Protocol pattern. Replacing with ``pass``
would diverge from the file's existing style. No code change.
Closes round-2 review findings q-3, q-4, q-5, q-7.
* **q-4:** ``_NAME_CONTROL_CHARS`` and ``_PAYLOAD_CONTROL_CHARS`` shared
7 lines of Unicode-steering character classes (zero-width / bidi /
separators / BOM / tag chars above BMP). Factored into a single
``_CONTROL_CHARS_TAIL`` constant; each regex now differs only in its
leading ASCII range. Future bidi or zero-width additions edit one
place.
Side effect: this corrects a latent bug where ``_NAME_CONTROL_CHARS``
had two literal ASCII spaces in place of U+2028 / U+2029 (line and
paragraph separators) — visible as ``r" "`` in source but rendered
as the actual codepoints in ``_PAYLOAD_CONTROL_CHARS``. After the
factoring both regexes correctly include U+2028 / U+2029, closing
the gap that would have let a workstream name with embedded line
separators forge a sibling bullet (the same vector ``\n`` was
blocked for in the original bug-1 fix).
Switched to ``\u`` escapes for readability (and to keep future Edit
tool runs against this block reliable).
* **q-3:** Tombstone clause "standing in for the deleted
``_watch_pending`` maxsize bound" survived in
``ChatSession.set_watch_runner``'s docstring after the apply-pass
trim cleaned the inline soft-cap comment. Dropped.
* **q-5:** ``test_newline_in_name_does_not_forge_extra_bullet`` carried
five WHAT-narration comments restating what the immediately-following
asserts already say. Dropped — the docstring carries the security
invariant; the assertions speak for themselves.
* **q-7:** ``patch_session_storage`` had a 14-line docstring including
fallback-guidance and self-justification ("accumulated 7 near-duplicate
sites"). Trimmed to a 3-line contract.
Closes round-2 review findings q-1, q-2, q-6.
* **q-1:** ``test_valid_until_drops_when_watch_missing`` collapsed to the
same code path as ``test_valid_until_drops_when_watch_inactive`` after
the apply-pass switched the predicate from ``get_watch[active]`` to
``is_watch_active`` (both stubbed via ``patch_session_storage(active=False)``).
The "missing" case has no distinguishable branch at the dispatch
layer, so dropping it removes a tautological duplicate. The
missing-row mapping moves to the storage layer (q-2 below) where it
IS distinguishable.
* **q-2:** ``is_watch_active`` was a new public storage primitive with
zero direct backend coverage — only via-session-via-stub coverage.
New ``TestIsWatchActive`` in ``tests/test_watch_storage.py`` covers
active row → True, inactive row → False, missing row → False.
Pinned at the storage boundary so future backend changes fail loudly
there instead of in the dispatch tests.
* **q-6:** Concurrency test had ``n_threads = 2`` alongside two literal
Thread objects and a tautological ``assert len(threads) == n_threads``.
Threads are now built from a labels tuple, so ``len(threads)`` drives
the slack bound; the redundant assertion is gone.
Closes review findings bug-4 and q-6.
bug-4 — the watch dispatch concurrency test bounded depth at
``_WATCH_QUEUE_SOFT_CAP + 2 * per_thread`` (= 250) which is
tautologically true: two threads × 100 fires can append at most 200
entries above the cap, so the bound asserted nothing more than what
``depth <= 2 * per_thread`` already says. Tighten to
``_WATCH_QUEUE_SOFT_CAP + N_THREADS`` (= 52): the count-then-drop window
admits at most one slip per concurrent thread.
q-6 — 7 near-duplicate ``monkeypatch.setattr(session_mod, "get_storage",
lambda: _StubStorage())`` sites across ``test_watch_dispatch.py`` +
``test_watch_integration.py`` (4 different stub shapes, mostly trivial
variations on the active flag). Lift a ``patch_session_storage``
helper into the existing ``tests/_helpers.py`` with kwargs for the
common cases (``active``, ``raise_on_is_active``), returns the call list
so call-shape assertions still work. Tests collapse from ~10-line
inline-class blocks to one-line helper calls.
Closes review findings q-2 and q-5.
q-2 — ``bound_watch_id = watch_id`` rebind was unnecessary. ``_dispatch``
is constructed fresh per fire (not in a loop), so ``_still_active``
closes over the function parameter directly without any
loop-variable-capture risk. Drop the rebind.
q-5 — the inline soft-cap comment restated rationale already covered by
the ``_WATCH_QUEUE_SOFT_CAP`` block-comment at module scope and dragged
in a tombstone reference to the deleted ``_watch_pending`` path. Trim
to one line stating only the WHY (drop-oldest because latest output is
most useful). Leave the ``set_watch_runner`` docstring's operational
detail at lines 1356-1378 alone — trimming further risks losing the
``valid_until`` predicate semantics.
Closes review finding q-4.
The closure built inside ``server.py``'s ``_watch_restore_fn`` is the
new contract surface introduced by the switchover — it constructs a
fresh ChatSession, calls ``session.resume(ws_id)`` to adopt the
original ws_id, re-registers the dispatch closure via
``set_watch_runner``, and returns ``WatchRunner.get_dispatch_fn`` for
the runner to invoke directly. No automated coverage exists today;
a future refactor (e.g. swapping ``manager.create + session.resume``
for ``manager.open``) could silently break the watch-restore pipeline.
Adds ``test_watch_dispatch_through_restore_fn_lands_on_rehydrated_session``
to ``tests/test_watch_integration.py`` — drives the full restore path:
persists a kickoff message for the original ws_id, fires
``_dispatch_result`` against a runner with no registered dispatch fn,
asserts the restore_fn ran exactly once, the rehydrated session is a
distinct object that adopted the original ws_id, and the watch payload
landed on the rehydrated session's NudgeQueue (not on the original).
Closes review finding perf-1.
The watch dispatch closure's ``valid_until`` predicate fires once per
watch entry at every drain seam — on the chat-loop hot path. It only
needs the ``active`` flag, but ``storage.get_watch`` runs a full-row
``SELECT *`` and marshals the result into a dict. At the typical drain
depth (cap-50 + a busy chat loop) that's ~50 throwaway dict allocations
per drain pass for one boolean.
Adds ``StorageProtocol.is_watch_active(watch_id) -> bool`` plus
SQLite + Postgres implementations doing a single-column
``SELECT active FROM watches WHERE watch_id = ?`` (returns False on
missing row). ``_still_active`` in ``ChatSession.set_watch_runner``
now calls that instead of indexing into the full row.
Test stubs that mocked ``get_watch`` for the predicate are converted
to mock ``is_watch_active`` directly. Bulk variant deferred — single-row
fix is sufficient at typical drain depths.
Closes review findings perf-2, q-3, bug-3.
The watch dispatch closure's soft-cap pre-check materialised the whole
queue snapshot via ``pending(channel="any")`` only to throw away the
text and count the type — wasteful at typical drain depths (cap-50 +
mixed producers means a 50-tuple allocation per fire just to read a
length). The other half of the cap pair (``drop_oldest_by_type``)
walked the *whole* queue regardless of channel, so a future producer
that enqueued ``"watch_triggered"`` on a different channel could be
dropped by the watch cap, and vice versa — silently surprising once
that producer existed.
Adds ``NudgeQueue.count_by_type(nudge_type, channel=None) -> int`` that
walks ``_items`` once under the queue lock without materialising
tuples; extends ``drop_oldest_by_type`` to take an optional ``channel``
filter so both halves can agree on the entry set being capped. The
watch dispatch closure now passes ``channel="any"`` to both —
consistent with where the closure enqueues — so a future channel split
can't bleed across producers.
Adds ``TestCountByType`` mirroring the existing ``TestDropOldestByType``
shape, plus a ``test_drop_oldest_by_type_channel_filter`` case pinning
the new optional argument's behaviour.
Closes review finding q-1.
The live-marker scaffold in ``tests/test_watch_live.py`` couldn't actually
run as written: the ``live_client`` / ``live_model_id`` fixtures it
referenced live in ``tests/test_server_live.py`` at ``scope="module"``,
not on a shared ``conftest.py``, so the file would have ImportError'd
at collection if anyone ever tried ``pytest -m live`` against it.
Lifting the fixtures into a shared conftest is a larger refactor
than R9 justifies — the deterministic envelope-arrival contract is
already pinned end-to-end by ``test_watch_fires_then_user_send_drains_envelope``
and ``test_three_back_to_back_watch_fires_drain_into_one_turn`` in
``test_watch_integration.py`` (real ChatSession + real WatchRunner +
real chat-loop drain). The model-quality-of-response leg is genuinely
manual; the plan doc's R9 entry is updated locally to reflect that
deferral.
Closes review finding bug-1.
The shared ``sanitize_payload`` regex preserved TAB/LF/CR so multi-line
watch shell output kept its layout — necessary for the watch path, but a
correctness gap for the idle_children formatter, which renders the
user-controlled ``name`` field as a single bullet item. A child name
with an embedded ``\n`` would split the bullet across two rendered rows
and let a hostile name forge a fake sibling entry in the listing.
Splits the regex in two: ``_NAME_CONTROL_CHARS`` strips TAB/LF/CR
(used by the new ``sanitize_name`` helper for single-line name fields),
``_PAYLOAD_CONTROL_CHARS`` keeps the existing permissive shape (used by
``sanitize_payload`` for multi-line watch payloads).
``format_idle_children_nudge`` now calls ``sanitize_name``.
Adds ``test_newline_in_name_does_not_forge_extra_bullet`` — feeds a
hostile name with embedded ``\n`` + bullet-shaped continuation, asserts
the rendered listing still has exactly N bullet rows for N children
(no forged sibling), and the hostile newline got flattened to an inline
space. Adds a ``TestSanitizeName`` class mirroring the existing
``TestSanitizePayload`` shape for the new strict variant.
The deleted comment claimed the closure may be registered "under the
rehydrated workstream's id, which may differ from the original ws_id we
restored against" — but ``ChatSession.resume(ws_id, fork=False)`` adopts
the parameter as the session's id at session.py:1682, so they match
exactly post-resume. The lookup works because the ids are equal, not
because they may differ.
The accessor name ``get_dispatch_fn`` is self-explanatory; no replacement
comment is needed (per the project's "default to no comments" rule).
Adds two boundary-crossing integration tests and one live-marker
scaffold for the watch switchover landed in the previous commits:
tests/test_watch_integration.py — drives a real ChatSession + real
WatchRunner end-to-end (LLM stubbed) through the unified pull-model
chat-loop drain seam. Pins:
- test_watch_fires_then_user_send_drains_envelope: a synchronous
WatchRunner.dispatch fire enqueues "watch_triggered" on "any";
session.send drains the entry into the user message's _reminders
side-channel — confirms the envelope splice path.
- test_three_back_to_back_watch_fires_drain_into_one_turn: pins the
intentional behavioural delta from the plan section 3.4 / risk
register R3 — N back-to-back fires now produce ONE assistant turn
with N _reminders entries, not N successive turns.
tests/test_watch_live.py (new file, single test, marked @pytest.mark.live):
risk register R9 verification recipe — confirm a real LLM handles a
<system-reminder>-framed watch payload sensibly. Collects under the
regular -m "not live" run; the user runs it on demand against an
Anthropic-backed config.
Implements watch-switchover plan section 5.2 (integration) and step 11
(live scaffold).
Replaces the deleted tests/test_watch_dispatch.py with a focused
14-test suite exercising the closure that ChatSession.set_watch_runner
now constructs (per the previous commit's switchover). Each test
pins one assertion:
- enqueue shape: ("watch_triggered", text, "any") on the per-session
NudgeQueue; not on user / tool channels
- producer-side sanitisation strips control / bidi / zero-width chars
and angle-bracket tag breakers; preserves TAB/LF/CR so multi-line
shell output keeps its layout (R8); empty-after-strip → no enqueue
- soft-cap drop-oldest at _WATCH_QUEUE_SOFT_CAP with a queue_full
WARNING log; non-watch entries on the same queue are not collateral
damage
- valid_until predicate drops on inactive / missing / storage-raises;
delivers when active (counter-test)
- concurrent enqueues across two threads stay bounded under the
3-acquisition count-then-drop window
Implements watch-switchover plan section 5.1 / step 9. No production
changes — pure test rewrite.
Replaces the bespoke _make_watch_dispatch / _watch_pending /
_dispatch_pending_watch / _MAX_WATCH_CHAIN machinery with a single
NudgeQueue.enqueue("watch_triggered", ...) call inside
ChatSession.set_watch_runner. Watch results now drain at the same
<system-reminder> envelope seams as every other metacog nudge
(USER_DRAIN, TOOL_DRAIN, IdleNudgeWatcher IDLE wake) — no separate
worker-spawn, no recursive watch chain, no per-session queue.Queue.
The dispatch closure built inside set_watch_runner carries:
- producer-side sanitize_payload over the whole formatted message
before enqueue, so steering-vector / control-char shell output
can't tamper with the envelope at interpolation time
- a soft cap of 50 entries on per-session "watch_triggered" depth
via the new NudgeQueue.drop_oldest_by_type, replacing the prior
_watch_pending maxsize=20 + _MAX_WATCH_CHAIN=5 bounds; drop policy
is drop-oldest (latest output most useful), logged at WARNING
- a valid_until predicate that re-checks
storage.get_watch(watch_id)["active"] at drain time so a cancelled
watch's last splat doesn't ride out a future wake
Behavioural delta documented in the plan section 3.4: N back-to-back
watch fires now drain into ONE assistant turn responding to all N
(via the envelope splice) instead of N separate send turns. This is
intentional — fewer model invocations for noisy watches, and uniform
with the rest of the metacog pull-model surface introduced by #482.
Implements watch-switchover plan steps 5-8. Server-side simplifications
let the previously-load-bearing _make_watch_dispatch (47 lines), its
session_worker.send import, and the chat-loop _dispatch_pending_watch
seam at the no-tools IDLE branch all disappear. The obsolete
tests/test_watch_dispatch.py and the wake-tag test in test_session.py
(both pinning contracts that no longer exist) are removed; the
NudgeQueue-based replacement plus an integration test land in the
following commit.
Widens the per-workstream dispatch fn signature from ``(message,)``
to ``(message, watch_id)``. The runner now passes the originating
``watch_id`` through ``_dispatch_result`` so dispatch closures can
capture per-watch metadata at fire time — the upcoming switchover
needs this for the ``valid_until`` predicate that re-checks
``storage.get_watch(watch_id)["active"]`` before a stale entry rides
out a wake.
Also adds ``WatchRunner.get_dispatch_fn(ws_id)`` as the public
accessor used by the server-side restore path to retrieve the
closure that ``set_watch_runner`` constructed during workstream
rehydrate (avoiding private-attr access into ``_dispatch_fns``).
Implements watch-switchover plan step 4 plus risk register R4.
The pre-existing single-arg callers (``_make_watch_dispatch`` and
``set_watch_runner``'s ``dispatch_fn=`` fallback) get replaced
in the next commit; their mypy types are ``Any`` today so the
type mismatch isn't caught at this step.
Renames _sanitize_child_name to sanitize_payload and widens it to be
the shared producer-side sanitiser for both idle_children and the
incoming watch_triggered nudges. The regex now skips TAB / LF / CR
so multi-line shell output rendered into a watch payload keeps its
line structure when sanitised as a whole formatted message — the
pre-switchover code path collapsed multi-line output to one line.
Adds the watch_triggered entry to _NUDGE_MAP alongside idle_children
so ``_NUDGE_MAP``-as-registry consumers (should_nudge gating, future
audit / UI tagging) recognise the type. Body is empty — payload
comes from the producer (the watch dispatch closure), same shape as
idle_children.
Implements watch-switchover plan section 3.2 plus risk register R8
(TAB/LF/CR exclusion) and step 3 (_NUDGE_MAP registration).
Adds an atomic drop-oldest-by-type operation to NudgeQueue used by
producers that need a per-type soft cap on their own queue depth.
The watch dispatcher (next commit in this stack) is the first user:
when "watch_triggered" saturates, the dispatch closure drops its
oldest entry under the queue lock so the count snapshot and drop
can't interleave with a concurrent enqueue from the same producer.
Implements watch-switchover plan section 3.1 — the producer-side soft
cap takes the place of the deleted _watch_pending maxsize=20 bound.
Other producers (idle_children, advisories) have natural rate limiters
already, so the helper is opt-in per producer rather than a global cap
in enqueue itself.
Three Copilot findings on PR #483 (commit dad98c0); one rejected as a
false positive.
- mcp_client.py:1189 — pool notification handler's exception path
used ``log.warning(..., exc_info=True)`` which serializes the
chained ``httpx.Request.headers`` carrying ``Authorization: Bearer
<token>`` into Sentry / faulthandler frame captures. Same threat
model as the round-1 sec-1 dispatch-path fix, applied to a site
the original review missed. Now logs structured fields only
(server, user, exc type) without ``exc_info``.
- mcp_client.py:1202 — ``_connect_one_pool``'s handshake step used
``asyncio.wait_for(session.initialize(), ...)``, the same Python
3.11 + anyio cross-task-cancel-scope anti-pattern that the
Phase 7 round-3 q-1 fix removed from the discovery step (and that
f6a3b66 originally addressed for ``_safe_close_stack``). Pre-
existing Phase 5 code, but the same latent bug class — a 401
during initialize() under 3.11 would surface ``RuntimeError:
Attempted to exit cancel scope in a different task`` as the
SDK's TaskGroup unwinds. Switched to ``async with asyncio.timeout(...)``
matching the discovery step's pattern.
- mcp_client.py:1522 — renamed loop tuple-unpack variable
``_server_name`` → ``server_name`` in ``_rebuild_user_tool_map``.
The leading underscore conventionally signals "intentionally
unused", but the variable is read at the assignment a few lines
below. Two other ``_server_name`` unpacks in this file (1410,
3111) genuinely don't use the value and keep the underscore.
Rejected as false positive:
- test_mcp_user_catalog.py:58 (github-code-quality bot, "Statement
has no effect"): ``await task`` inside ``contextlib.suppress(
BaseException)`` is the standard pattern for cleanly draining a
cancelled task. The bot's static analysis treats ``await`` of a
result that's discarded as a no-op statement, but ``await`` here
triggers cancellation propagation and waits for the task to
finish — load-bearing in the fixture's teardown. No change.
Verified on Python 3.11 (``/tmp/venv311``) and 3.13 (``.venv``):
ruff + mypy clean, full test suite green.
Light up production reachability of pool dispatch (RFC §3, invariant 8)
by widening the public catalog API to optionally take a ``user_id``:
- ``MCPClientManager.get_tools(user_id=None)`` returns the merged
static + per-user pool view when ``user_id`` is supplied; the default
preserves the legacy global-only contract.
- ``is_mcp_tool(name, *, user_id=None)`` extends the lookup to the
per-user ``_user_tool_map``. Pool tools become reachable from
``ChatSession._prepare_tool`` only when the session-bound user_id
flows through — flipping invariant 8 from "must hold" to "satisfied".
- Listener identity becomes ``(user_id, callback)``. Static-path
changes fire ALL listeners (admin + every user); pool-entry
changes fire only matching-user + admin (``None``) listeners.
RFC §3.3.
- Pool sessions discover their tool list on first connect
(``_connect_one_pool`` → ``await session.list_tools()``); the
notification closure binds to ``(user_id, server_name)`` so
push-driven ``list_changed`` updates target the correct user's
catalog. R6 verified empirically: ``list_tools()`` 401 propagates
through anyio TaskGroup unwinding, no hang — plain ``await`` is
fine, no carrier-race shape needed for discovery.
- ``_evict_session`` drops ``entry.tools`` and rebuilds the user's
index so an evicted-then-reconnected session doesn't carry
stale catalog state.
- ``web_search.resolve_web_search_client`` refuses
``auth_type=oauth_user`` backends (per-node web search can't
carry per-user tokens).
Resources / prompts pool dispatch deferred to Phase 7b — invariant 8
is satisfied by the tool path alone, and the resource/prompt path
needs sibling ``_dispatch_pool_resource_sync`` /
``_dispatch_pool_prompt_sync`` helpers each with their own
carrier-race plumbing (~400 LOC). Phase 7b will follow the patterns
established here.
CLI sessions default ``user_id=""`` and so cannot use oauth_user
MCP servers — documented limitation; users must use the web UI.
Round-1 review fixes (4-finder review applied, no push yet):
- bug-1: get_tools(user_id) was iterating _user_pool_entries from sync
threads while the mcp-loop concurrently mutated it (RuntimeError:
dictionary changed size during iteration). Now reads from a sibling
_user_tools dict updated atomically by _rebuild_user_tool_map.
- bug-2: _close_pool_entry_if_idle (LRU/TTL eviction) skipped the
catalog cleanup that _evict_session does — stale tools persisted
in _user_tool_map and ChatSession's tool list never rebuilt. Now
mirrors _evict_session.
- perf-1: _last_pool_notification_refresh debounce dict was never
pruned in either eviction path. Now popped alongside the entry.
- perf-3: web_search resolver was issuing a sync SQL query per LLM
turn to gate oauth_user backends. Now reads from the cached
in-memory config.
- sec-1: bearer token could leak into exc_info-rendered tracebacks
via Sentry/faulthandler. log.debug now uses structured fields,
not exc_info.
- sec-2: tools-per-server response now capped at 1000 (defensive,
mirrors _MAX_ERROR_LEN / _MAX_INSUFFICIENT_SCOPE_REPORTED).
- Test cleanup: dropped two listener fan-out tests duplicating
test_mcp_client.py coverage; renamed test_pool_session_notification_handler
to match its actual scope (_refresh_pool_server_tools); removed
stale comments referencing /tmp/r6-spike*.py scratchpads and a
misleading "copy-on-write" comment.
Round-2 pre-push review fixes (focused single-pass review applied):
- round2-1: bug-2's catalog-cleanup block in _close_pool_entry_if_idle
had no integration test (exactly the failure mode flagged in
feedback_tests_through_boundaries.md). Added
test_close_pool_entry_if_idle_clears_catalog_and_fires_listener
driving the LRU/TTL eviction path through real streamablehttp_client +
MockTransport. Negative-test verified: reverting the
_rebuild_user_tool_map / _notify_user_tool_listeners calls makes
the new test fail.
- round2-3: documented the _oauth_user_server_names cache invariant
in add_server_sync / remove_server_sync docstrings. Cache is
reconcile_sync's sole owner — direct callers leave it stale, but
_db_servers_to_config strips oauth_user rows so production paths
are unaffected. Static→oauth_user transitions correctly leave the
name in the cache because remove_server_sync drops the static
connection, not the cache identity.
- round2-6: strengthened test_rebuild_user_tool_map_populates and
test_rebuild_user_tool_map_drops_empty_user to assert on the
_user_tools sibling cache (bug-1 fix). Without this, a future
revert dropping the sibling write would still pass the unit
tests because get_tools coverage lives in separate tests.
Round-3 full-stack review fixes (multi-stage review on the final
state caught what the layered apply passes missed):
- q-1 REGRESSION: pool tool-discovery used asyncio.wait_for around
session.list_tools(), the exact pattern the f6a3b66 fix (and
feedback_asyncio_timeout_vs_wait_for.md) put in place to avoid.
Python 3.11's asyncio.wait_for wraps the inner coroutine in a
fresh task → cross-task scope-exit when the SDK's anyio TaskGroup
unwinds on a 401. Switched to `async with asyncio.timeout(...):`
pattern used by _safe_close_stack.
- sec-2: TOCTOU in _connect_one_pool — entry.tools was published
(via _rebuild_user_tool_map + listener fan-out) BEFORE entry.session
was assigned. A sync-thread reader could observe a tool whose
backing entry has session=None. Defence-in-depth — dispatch
re-fetches its own token and lazy-reconnects on session=None — but
reordering catches the race at the source. entry.session now
publishes BEFORE catalog visibility.
- bug-1: _close_pool_entry_if_idle's _user_pool_locks.pop ran
unconditionally after the try/finally, but the early-return
branches (entry None on re-check, in_flight > 0 under lock) skip
it via Python's return-through-finally semantics. The lock was
never popped on those paths. Now gated behind an `evicted` flag
set only on the success path; in_flight > 0 leaves the lock for
the active dispatcher to reuse, entry-None races leave the lock
for re-allocation by _ensure_pool_entry. Comment now describes
the actual semantics, not the original promise.
- bug-2: softened the _rebuild_user_tool_map docstring's atomicity
claim. The two-dict write is technically non-atomic across Python
statements; in practice the window is sub-microsecond on the
mcp-loop with no awaits between writes, and the listener fan-out
fires AFTER both writes complete. Docstring now says "back-to-back
on the mcp-loop" instead of "atomically alongside".
- q-3: dropped `hasattr(mcp_client, "server_auth_type")` defensive
check in web_search.py. The method ships in this commit; the
hasattr created a silent fallthrough that would let a future
rename silently re-enable oauth_user backends.
- q-4: surfaced the CLI / empty-user_id limitation in a docstring
comment at ChatSession.__init__'s self._user_id assignment. The
note previously lived only inside is_mcp_tool's docstring — a
future maintainer wiring CLI features against MCP pool servers
wouldn't think to read is_mcp_tool to find the constraint.
- q-2 + q-5: deleted a tautological duplicate test in
test_mcp_user_catalog.py whose docstring claimed to test
ChatSession.close but never instantiated a ChatSession (the
manager-level identity semantics are already covered by
test_listener_identity_includes_user_id in the same file and by
test_session_close_removes_listener_with_same_user_id in
test_mcp_client.py which DOES drive a ChatSession). Reworded a
misleading "fixture provides only 5s" comment to point at the
actual `_run_on_loop(..., timeout=5)` site.
- q-6: the `self._user_id or None` collapse repeated at 8 sites
across session.py. Cached once at __init__ as
``self._mcp_user_id`` (since ``_user_id`` is set once and never
mutated); 8 call sites now read the cached value. The empty-
string-is-CLI-sentinel invariant is documented at the assignment
site, not re-asserted at each consumer.
Deferred to follow-up:
- sec-1: a hostile MCP server bound to user-A could craft a
tool.name containing `__` to synthesize a prefixed-name collision
in user-A's own catalog. Bounded impact: cross-tenant dispatch is
prevented by the per-tenant token gate in _dispatch_pool, and
user-B's get_tools(user_id="B") never includes user-A's pool
entries. The fix needs policy decisions (reject vs. sanitize)
and touches _mcp_to_openai which is shared between static and
pool paths; better discussed in its own follow-up where the
policy applies uniformly to static-path servers too. The threat
model already requires user-A to have consented to a malicious
server, who has many more dangerous vectors than tool-name
shenanigans.
Test count delta: +31 tests (5435 → 5466, ``-m "not live"``; one
test deleted in round-3 apply per q-2):
- ``tests/test_mcp_client.py`` +20 (per-user catalog state, listener
identity, session thread-through)
- ``tests/test_mcp_user_catalog.py`` +9 NEW (integration tests
driving real ``streamablehttp_client`` + ``httpx.MockTransport`` per
invariant 14: discovery on connect, user isolation, eviction +
reconnect, LRU/TTL eviction (round2-1), R6 401-propagation
regression, static byte-identical canonical regression; review
passes dropped duplicate listener fan-out tests from earlier
drafts whose coverage lived in test_mcp_client.py)
- ``tests/test_web_search.py`` +2 (oauth_user backend rejection +
static backend acceptance regression; updated to use the new
``server_auth_type`` in-memory accessor)
Three confirmed findings from the PR #482 bot review pass.
* **Copilot (idle_nudge_watcher.py)**: ``IdleNudgeWatcher`` was gating
wake dispatch on ``len(_nudge_queue) == 0`` (any channel), but
``deliver_wake_nudge_from_queue`` only drains ``USER_DRAIN``. A
``"tool"``-channel entry queued by ``_queue_tool_advisory`` would
pass the gate, spawn a wake daemon, and immediately no-op at the
drain guard — repeating on every IDLE event for as long as the
tool entry sat unconsumed. No correctness bug (the no-op return
prevents bad state) but a wasted thread spawn per IDLE. Fixed by
gating on ``has_pending(USER_DRAIN)``; tool-only queues no longer
trigger the wake path.
* **Copilot (coordinator_idle_observer.py)**: docstring referenced
the old module path ``turnstone.core.metacognition.IdleNudgeWatcher``;
the class moved to ``turnstone.core.idle_nudge_watcher`` in q-3 of
the apply-pass.
* **Copilot (nudge_queue.py)**: ``has_pending`` docstring cited
``ChatSession.deliver_wake_nudge_from_queue`` as its caller, but
that method calls ``drain(USER_DRAIN)`` directly — no production
caller used ``has_pending`` until this commit. Updated to point
at the now-actual caller (``IdleNudgeWatcher``).
* **github-code-quality (test_nudge_queue.py)**: false positive on
``test_channel_is_required`` — the no-channel ``q.enqueue("a", "1")``
call is wrapped in ``pytest.raises(TypeError)`` to verify the
validation contract. No code change.
5571 non-live tests pass; ruff + mypy clean.
Round-2 review caught 11 confirmed findings on the 3-commit metacog stack;
this commit applies them.
* **bug-1 (major)**: Wake source tag was leaking onto real user messages
flushed during a wake send. ``_append_user_turn`` and ``send`` now
take an explicit ``from_wake: bool`` parameter — only the wake's
synthesized first turn passes True, so ``_flush_queued_messages``'s
real user input no longer inherits the audit tag. Regression test
pins the contract.
* **perf-1 (major)**: ``CoordinatorIdleObserver._maybe_enqueue`` was
issuing list_workstreams + visible_memory_count storage queries
before the cheap cooldown gate could short-circuit. New
``_cooldown_allows`` read-only peek runs first; storage queries only
fire when cooldown actually allows the nudge.
* **q-1 (major)**: Added the missing coord-side integration test that
exercises ``CoordinatorIdleObserver`` + ``IdleNudgeWatcher`` together
in the production install order against a real ``SessionManager``,
protecting the subscription-order contract from silent regression.
* **perf-2/3 (minor)**: Cap check moved above ``_last_assistant_used_wait``;
``_fire_counts`` restructured as ``dict[str, dict[str, int]]`` keyed by
ws_id so the leave-IDLE existence check is O(1).
* **perf-4 (minor)**: ``NudgeQueue.drain`` fast-paths the all-match
case (the common one for chat-loop drain seams) by swapping
``self._items`` directly instead of allocating a fresh ``kept``
deque + per-entry append.
* **perf-5 (minor)**: Wake's synthesized empty user turn no longer
writes a content-empty row to the conversations table — the
``_source`` audit tag isn't column-backed and the side-channel
reminder is stripped before persist, so the row would carry nothing.
* **q-3 (minor)**: Split ``IdleNudgeWatcher`` + ``install_*`` /
``shutdown_*`` helpers out of ``metacognition.py`` into the new
``turnstone/core/idle_nudge_watcher.py``; metacog stays a
static-template module.
* **sec-1 (nit)**: Widened ``_sanitize_child_name``'s control-char
regex to cover Unicode bidi-overrides, zero-width chars,
line/paragraph separators, BOM, and tag chars.
* **q-4/q-5 (nits)**: Docstring referenced the wrong peek primitive
(``has_pending`` → ``len()``); ``_last_assistant_used_wait``'s
``session`` parameter now typed ``ChatSession``.
5571 non-live tests pass; ruff + mypy clean.
Adds the first concrete consumer of the wake trigger: when a coordinator
goes IDLE while interactive children are still running, a
``CoordinatorIdleObserver`` enqueues an ``idle_children`` nudge that the
``IdleNudgeWatcher`` then dispatches as a synthetic empty-user-turn
``send``. The model receives a system-reminder body listing the active
children (capped at 6 inline + 32 in the suggested ``wait_for_workstream``
call) and a nudge to block on them rather than reply prematurely.
Observer gates (in order): coord-only filter, skip if last assistant
turn used ``wait_for_workstream``, per-(ws, nudge_type) hard cap (3)
that resets only on non-wake leave-IDLE, active-children query,
``should_nudge`` cooldown. Console lifespan registers the observer
BEFORE the watcher so subscriber-fire order has the observer
enqueueing first on the same IDLE event.
Adds an opt-in ``valid_until`` predicate on ``NudgeQueue.enqueue``
(R9 from the design risk register) — drain re-checks the predicate
outside the queue lock; falsy / raising drops the entry without
delivering it. ``deliver_wake_nudge_from_queue`` now drains inline
before synthesizing the empty user turn so a stale predicate-drop
doesn't leave the wake send with empty content; ``_attach_pending_user_reminders``
consumes the pre-drained reminders via ``_wake_drained_reminders``.
The observer's ``valid_until`` uses ``count_workstreams_by_state``
(boolean check, no row fetch) instead of full ``list_workstreams``,
keeping the chat-loop user-attach path off the heavy query.
User-controlled child workstream names are sanitized
(``_sanitize_child_name``) before interpolation so a name like
``</thinking>...`` can't steer the model's reasoning channels through
the rendered body — the wire-boundary ``escape_wrapper_tags`` only
covers ``<system-reminder>`` / ``<tool_output>`` envelopes.
Adds the third metacog channel: an out-of-band wake that converts a
workstream's IDLE transition into a synthetic empty-user-turn ``send``
when the session has any-channel nudges queued. The ``IdleNudgeWatcher``
subscribes to ``SessionManager.subscribe_to_state``; on IDLE it dispatches
via ``session_worker.send`` with a no-op ``enqueue`` callback so a
busy-worker race silently drops without spawning a competing worker.
Wake-source-tag plumbing on ``ChatSession`` short-circuits metacog
detection on the synthetic empty input, suppresses queue producers
during the wake's own tool dispatch, and stamps ``_source = "system_nudge"``
on the synthetic user-message for audit / replay distinction. The tag
is saved / restored across ``_dispatch_pending_watch`` so watch chains
recursing off the wake are processed as normal user turns rather than
inheriting the wake's guards.
Generic ``install_idle_nudge_watcher`` / ``shutdown_idle_nudge_watchers``
helpers wire the watcher into both the interactive and coord lifespans
via a single ``app.state`` registry so both surfaces share the same
teardown contract.
Foundation for PR 3 (CoordinatorIdleObserver + idle_children formatter)
and PR 4 (watch dispatcher switchover).
Replaces the dual `_pending_user_advisories` / `_pending_tool_advisories`
list pair with a single channel-tagged `NudgeQueue` per session.
Producers tag entries with a channel ("user", "tool", or "any");
consumers drain by channel filter at their existing seams. Foundation
for the wake trigger (PR 2) and coordinator idle-children nudge (PR 3).
Existing nudges (start, correction, completion, denial, resume,
tool_error, repeat) keep their wire shape and drain timing — zero
behavior change. Cancel paths now `clear()` the unified queue.
Python 3.11's ``asyncio.wait_for`` wraps its inner coroutine in a fresh
``asyncio.Task`` via ``ensure_future``. When the inner is
``stack.aclose()`` on an ``AsyncExitStack`` containing
``streamablehttp_client(...)`` (anyio cancel scopes entered in the
calling task), the fresh task's attempt to exit those scopes raises
``RuntimeError('Attempted to exit cancel scope in a different task
than it was entered in')``. Python 3.12+ rewrote ``wait_for`` to use
``asyncio.timeout`` internally — runs in the current task — so 3.13
ran the same code path successfully.
Symptom on 3.11: integration tests where ``session.initialize()``
returns 4xx (e.g., 403 insufficient_scope tests) hit
``_connect_one_pool``'s ``except Exception:`` handler →
``_safe_teardown_on_connect_failure`` → ``_safe_close_stack`` → cross-
task RuntimeError. The ``concurrent.futures._base.CancelledError``
that surfaces in ``future.result(timeout=...)`` is the cascade
fallout from the asyncio loop's exception handler reacting to the
unretrieved-task-exception.
Fix: use ``asyncio.timeout`` instead of ``asyncio.wait_for`` for the
5s aclose bound. Equivalent semantics, current-task execution, works
on 3.11+. The 5s guard against ``aclose()`` hanging on a broken stack
is preserved.
Verified on Python 3.11.14 (full suite 5427 passed) and 3.13.7 (full
suite 5427 passed); all 9 integration tests pass on both.
Pre-existing bug — surfaced only after the marker fix in 5c9850c
let CI's test (3.11) actually run the 4xx tests.
Two pre-existing defects in the Phase 6 pool dispatch path that only
manifest when a pooled session is reused for a second dispatch:
1. The per-dispatch _AuthCapture allocated in _dispatch_pool was wired
into the httpx response hook only at first connect (via
_connect_one_pool). On a reused session no fresh connect runs, so
the hook continues writing to the original-connect's carrier while
the new dispatch inspects an empty carrier — auth_401/403 silently
misclassified to "other", refresh-and-retry never fires.
2. Even with the carrier on the entry (so the hook writes to a stable
reachable object), session.call_tool itself hangs forever on
upstream 4xx for reused sessions. Trace: SDK's spawned
handle_request_async raises HTTPStatusError, the outer
streamablehttp_client TaskGroup cancels post_writer, post_writer's
finally aclose's read_stream_writer, BaseSession's _receive_loop
exits and enters its CONNECTION_CLOSED-fanout finally. anyio's
send_nowait skips waiting receivers with pending_cancellation; the
dispatch task (created by run_coroutine_threadsafe for the reuse
case) is NOT in any cancel-scope chain, so the send "delivers" but
the receiver's Event is set on stale state — receive() never
wakes. Test 21 doesn't hit this because its 401 happens during
initialize, in the same task that opens streamablehttp_client, so
the cancel scope DOES propagate.
Fix:
- Move _AuthCapture ownership to PoolEntryState (and asyncio.Event
alongside, allocated lazily on the mcp-loop). The hook closes over
entry.auth_capture at first connect and stays valid across
dispatches; reset under open_lock before each call_tool.
- Race session.call_tool against the carrier's fired_event in
_dispatch_pool_with_entry. If the event wins (hook captured 4xx
before SDK propagated), cancel call_tool and raise an internal
_CarrierAuthSignal — _classify_failure resolves to auth_401/403
via the carrier's status, the dispatcher evicts the broken
session, and the cross-task retry handshake reconnects on a fresh
bearer.
Adds tests/test_mcp_pool_auth_integration.py::test_integration_pool_reuse_401_refresh_and_retry_succeeds
which drives the reuse path through real upstream + real SDK and is
the structural gate against this class regressing. Negative-tested
twice: revert PoolEntryState.auth_capture → test fails (carrier
empty); revert the race → test times out (SDK hang).
Also drops the @pytest.mark.asyncio decorator (replaced with
@pytest.mark.anyio) on four tests in test_mcp_pool_auth_introspection.py.
The project depends on anyio's pytest plugin (anyio is in deps);
pytest-asyncio is NOT a project dep and CI's test (3.13) failed on
those four. Local pytest happened to pick it up via system Python.
Found via Copilot review on PR #481.
Phase 6 of OAuth-MCP. Recovers upstream 401/403 from MCP servers via a
capturing httpx_client_factory: an async response hook records 4xx
status + WWW-Authenticate header into a per-dispatch carrier before
the SDK's post_writer swallows the underlying httpx.HTTPStatusError.
Splits _classify_failure into auth_401 (refresh-and-retry once) vs
auth_403 (parse insufficient_scope, emit mcp_insufficient_scope with
parsed scope set). The 401 retry runs on a fresh asyncio.Task via
run_coroutine_threadsafe in _dispatch_pool_sync, escaping the anyio
cancel-scope state of the prior dispatch's TaskGroup.
WWW-Authenticate parsing extracted to a new mcp_http_parsers module
with an RFC 7235 challenge tokenizer (replaces hand-rolled substring
scanners). Two-layer defense against multi-Bearer-challenge injection:
the hook uses get_list("www-authenticate")[0] to drop attacker's
second challenge, the parser truncates at challenge boundary as
belt-and-braces. Scope set capped at 32 entries before hitting the
audit row or the LLM-visible structured-error JSON.
Auth failures (401/403) never trip the per-server circuit breaker
(server-only breaker invariant). Static path remains byte-identical.
_PgRefreshLock untouched. Pool dispatch still reachable from the
agent loop only via Phase 7 catalog scoping; Phase 6 behaviour is
testable via direct call_tool_sync.
5557 tests pass. 33 tokenizer unit tests in tests/test_mcp_http_parsers
cover the RFC 7235 grammar + the scope/error wrappers + the 4 KB input
cap. 7 integration tests in tests/test_mcp_pool_auth_integration drive
real upstream 401/403 through streamablehttp_client + a FastMCP
subprocess fixture — the structural exit gate that makes
HTTPStatusError-injection-only unit tests insufficient.
Models often emit page references in the standard man-page form
(``printf(3)``, ``open(2)``, ``perlfunc(3pm)``) rather than splitting
them into ``page`` + ``section`` args. The page-name sanitizer was
rejecting the parens as invalid input, killing the call. Parse the
section out of the page string before sanitization (explicit
``section`` arg still wins) and widen the section validator to accept
multi-letter suffixes like ``3pm`` / ``3perl`` that already appear on
real systems.
Phase 5 PR #479 review fix-up. Three review rounds (bot + two internal
multi-stage /review) caught:
- _PgRefreshLock now allocates a per-instance ThreadPoolExecutor instead of
a module-global single-worker one. The global shape preserved psycopg2
thread-affinity but serialized every advisory-lock acquire on the node
behind one thread, even for unrelated (user, server) keys.
- get_user_access_token_classified flips to `async with lock, pg_lock:` so
concurrent same-key callers serialize on the in-process asyncio.Lock
before allocating the pg_lock's per-instance executor + spin loop. N
concurrent same-key callers collapse to one executor allocation.
- _drain_orphan_pg_lock no longer re-awaits the cancelled asyncio Future
from `__aenter__`. It receives the underlying concurrent.futures.Future
and re-wraps it via asyncio.wrap_future, getting an independent asyncio
Future tied to the worker outcome. This way cancellation of the awaiter
doesn't poison the drain's wait, and the drain genuinely waits for the
worker to settle before deciding whether to call cm.__exit__.
- Module-level _pg_refresh_drain_tasks set holds strong refs to in-flight
drains (asyncio's task set is weak — fire-and-forget tasks could be GC'd
mid-cleanup; RUF006 hazard).
- Drain narrows except clauses to Exception so a drain-task cancellation
records as cancelled instead of being silently logged as 'completed
normally with no acquire'.
Test integrity (was a major finding in round 2 — old generator-based cm
let the test pass via GC finalization timing rather than drain logic):
- New _ObservableLockCm class-based context manager whose __exit__ is a real
observable method (records call args + thread). Distinguishable from
GeneratorExit thrown by GC of a generator-based cm.
- Strong external ref to the cm via created_cms list — keeps cm alive past
the test's awaits, so a no-op drain genuinely fails the assertion rather
than papering over via GC timing.
- Deterministic drain wait via _pg_refresh_drain_tasks gather — no
fixed-duration sleeps.
- _run_cancel_scenario helper drops the duplicated setup between the two
cancellation tests.
Negative-test verified: replacing _drain_orphan_pg_lock body with `return`
makes test_pg_refresh_lock_cancellation_releases_on_same_thread fail with
'drain did NOT call cm.__exit__ — orphan Postgres lock + open transaction'.
Other fixes: protocol docstring corrected to describe pg_try_advisory_xact_lock
spin + retry (was claiming pg_advisory_xact_lock blocking acquire);
get_user_access_token_classified docstring rewritten for new lock order;
narrow `except BaseException` -> `except Exception` in
test_mcp_user_pool.py concurrent-dispatch helper.
882 tests pass (MCP + auth + storage). ruff + mypy clean.
Phase 5 of OAuth-MCP — adds a per-(user, MCP-server) ClientSession
pool to MCPClientManager alongside the existing static-server path,
gated entirely on the per-server `auth_type='oauth_user'` config.
Pool architecture:
- `_user_pool_entries: dict[(user_id, server_name), PoolEntryState]`
with lazy connect on first dispatch, per-key asyncio.Lock allocated
on the mcp-loop, idle eviction coroutine (default 600s TTL, LRU cap
200), and an `in_flight` counter as the eviction interlock so live
calls can never be torn down mid-flight.
- `_dispatch_pool` runs the token-state machine: missing token →
`mcp_consent_required`; key-rotation decrypt failure →
`mcp_token_undecryptable_key_unknown` with NO consent prompt and NO
auto-delete; expired token → silent refresh under per-(user, server)
advisory lock; refresh failure → revoke + consent.
- `_classify_failure` separates transport (trips breaker) from auth
401/403 (does NOT trip breaker — server-only invariant) from
protocol (no breaker change).
- `entry.open_lock` held only across connect-or-reuse and released
before the `await session.call_tool` so concurrent calls from one
user against one server overlap (validated by Spike 1 scenario 2).
Auth-class failures are fail-soft in Phase 5: any 401/403 surfaced by
the SDK propagates to the agent as a tool error and the next dispatch
reconnects on a fresh refresh. Real introspection of upstream 401/403
is a Phase 6 concern — the MCP SDK's `streamable_http` post_writer
swallows `httpx.HTTPStatusError` upstream, so detecting status from
the response chain requires `McpError(CONNECTION_CLOSED)` payload
parsing or a custom httpx middleware around `streamablehttp_client`.
The mid-flight 401 refresh-retry path and the `mcp_insufficient_scope`
structured error for 403 step-up land together in Phase 6, gated by
an integration test that drives a real upstream 401/403 (the unit-
test injection of `HTTPStatusError` is what masked the production gap
on the first apply-findings pass — the integration test is the
structural gate so the gap can't reopen). RFC §1.5 steps 4-5 and the
phase table in §Implementation phases reflect this scope split.
Multi-node refresh contention:
- New `StorageBackend.acquire_advisory_lock_sync` Protocol method.
SQLite returns nullcontext (single-node, in-process asyncio.Lock
is sufficient). Postgres uses `pg_try_advisory_xact_lock` with
retry on a fresh per-attempt connection, so waiters don't pin pool
connections during the AS roundtrip. Inner try/except + nested
finally ensures conn is always returned to the pool, even when
begin / execute / yield / commit raises mid-body.
- Lock ordering: pg_advisory outer, asyncio.Lock inner. Re-read after
lock collapses cluster-wide contention to one HTTP roundtrip per
(user, server) per refresh window.
- `_PgRefreshLock` enter/exit pinned to a single-worker
ThreadPoolExecutor so SQLAlchemy connection state stays
thread-affine across cancellations.
Token storage refactor:
- `get_user_access_token_classified` returns a tagged TokenLookupResult
(Token / MissingToken / DecryptFailure / RefreshFailed) so the
dispatcher maps each state to the right user-facing error.
- `get_user_access_token` is now a thin wrapper around the classified
variant; the previous duplicated state machine is gone.
Security:
- Pool dispatch + admin endpoints reject `http://` URLs for
`auth_type='oauth_user'` servers (only exact loopback hostnames are
exempt — `*.localhost` is intentionally NOT honored because RFC 6761
localhost-zone resolution is configuration-dependent and could route
bearers to non-loopback IPs via custom resolvers / hosts file /
Docker overlays). Validated at three layers:
`_dispatch_pool` (structured `mcp_oauth_url_insecure` error),
`_connect_one_pool` (defensive ValueError), and
`admin_create_mcp_server` / `admin_update_mcp_server` (400 before
storage write).
- Admin URL change on an oauth_user row purges per-user OAuth tokens
bound to the old URL: bearers are bound (via OAuth resource /
audience) to the URL active at consent time, so silently rebinding
them to a new URL is a token-binding violation. Re-consent forces
fresh issuance for the new resource.
- Encryption-key fingerprints stay in audit logs only; no longer
surfaced in agent-facing error payloads.
User_id thread-through:
- `MCPClientManager.call_tool_sync(..., user_id=None)` (additive;
default None preserves the static path byte-identically).
- `ChatSession._exec_mcp_tool` passes `self._user_id or None`.
- `set_app_state(app_state)` setter wires OAuth state at lifespan
startup, called from both turnstone-server and turnstone-console.
Performance:
- LRU cap eviction iterates `_user_pool_entries` (not
`_user_pool_last_used`) so pre-dispatch entries are eligible.
- Eviction batch closes via `asyncio.gather` instead of serial await.
- `_resolve_pool_target` returns the resolved server row to
`_dispatch_pool` to eliminate the second DB lookup.
- Production reachability of pool dispatch is gated on Phase 7
(catalog scoping) wiring pool tools into `_tool_map`; until then
pool dispatch is reachable only via direct `call_tool_sync` with a
prefixed name (the path the new pool tests exercise).
Hardening parity preserved:
- Static path (auth_type ∈ {none, static}) byte-identical; PR #296
hardening (SDK #2147 mitigations, anyio cancel-scope, stale-session-
and-stack guard, server-only circuit breaker) intact.
- `test_reconnect_preserves_static_state_identity` unchanged + green.
- `MCPTokenStore.get_user_token` does not auto-delete on
MCPTokenDecryptError (key-rotation safety).
- Notification debounce stays manager-level.
- Connect-failure cleanup factored into
`_safe_teardown_on_connect_failure` shared by both connect paths.
Tests: 5475 → 5493 (+18). New file `tests/test_mcp_user_pool.py`
plus additions to test_mcp_oauth_refresh.py, test_mcp_admin_api.py,
and test_mcp_client.py covering: pool data structures, lazy connect,
eviction TTL + LRU + lock interlock, dispatch state machine (token
states), failure classification, http-rejection at dispatch and
admin layers, URL-change-purges-tokens (sec), concurrent dispatch on
one (user, server), pg_advisory lock parity, and user_id threading.
Phase exit criterion (synthetic load test 50 users × 3 servers × LRU
30 × 1000 calls × 200 evictions) deferred to a post-Phase-5 fitness
spike that runs against a staging deployment with real FDs and real
network behaviour, not a CI mock — same shape as Spike 1's
pre-Phase-0 SDK validation.
Out-of-scope for Phase 5 (Phase 6+): SDK-level 401 refresh-retry +
403 `mcp_insufficient_scope` (Phase 6), per-user catalog scoping
(Phase 7), consent UX SSE event + dashboard renderer (Phase 8),
admin UI status indicators (Phase 9).
Spike artifact validating MCP SDK behavior before Phase 5 builds the
per-(user, MCP-server) ClientSession pool. Three scenarios, all pass:
1. N=20 concurrent ClientSession instances against the same URL — no
FD blow-up, no shared transport state, each session's tools/list
returns independently.
2. Two concurrent tools/call on a shared ClientSession with
interleaving payloads — request_id demux works under contention.
3. Per-session Authorization header isolation across 5 sessions —
httpx connection pooling does not cross headers between sessions,
so per-session bearer tokens reach the server unmixed.
Outcome gates the Phase 5 architecture (lazy dict[(user_id,
server_name), ClientSession] + per-key asyncio.Lock + LRU eviction).
Had any scenario failed, the fallback was per-call header injection
(Alternative F in the OAuth-MCP RFC).
Spike-only — not collected by pytest. Run manually:
uv run python tests/spike_sdk_concurrency.py
Addresses ten findings on the Phase 4 OAuth-MCP commit: four from the
PR #478 review surface, plus six surfaced by a follow-up multi-stage
review of the first round of fixes. Two of the latter were genuine
security regressions in the very code that claimed to close those
holes.
Security
--------
- _validate_return_url now pins return_url same-origin against the
configured oidc_config.redirect_base instead of request.url. Behind
a permissive front proxy that did not normalise Host, an attacker
could spoof Host and provide a matching absolute return_url to mint
an open redirect off /api/mcp/oauth/start. Same fix pattern as
PR #476 OIDC.
- Reject return_url values containing literal backslashes or starting
with `//` up front. urlparse leaves backslashes inside `path`, so a
value like `/\evil.example/foo` slipped through the path-only branch
and became the protocol-relative `//evil.example/foo` after WHATWG-
conformant browsers normalised the backslash — re-introducing the
open redirect the same-origin pin was meant to close.
- internal_mcp_status (read-scoped) projects through a new
_strip_server_status_for_read helper that drops the verbose `error`
text and replaces it with a coarse `has_error` boolean. The error
string is built as `f"{type(exc).__name__}: {exc}"` and so carries
stdio binary paths (FileNotFoundError) or internal MCP URLs
(httpx.ConnectError) — equivalent to leaking command/url, which
this same patch deliberately strips. Approve-scoped refresh and
reconnect callers continue to receive the full `error` text via
the existing _strip_server_status helper.
- internal_mcp_status now returns the projected (sanitised) entries
for every server in mcp_mgr.get_all_server_status() instead of
emitting the un-sanitised dict that included `command` (stdio argv)
and `url` (remote MCP endpoint). Sibling refresh/reconnect endpoints
already used _public_server_status to strip these.
- internal_mcp_status docstring documents the trust boundary — server
enumeration to read scope is intentional so dashboards can render
per-server indicators; verbose error detail and command/url remain
approve-scoped.
Correctness / UX
----------------
- _validate_return_url comparison normalises (scheme, host, port)
before equality. Lowercases hostname and collapses the scheme's
default port, so `https://App.Example.COM/x` and
`https://app.example.com:443/x` are recognised as same-origin
with `redirect_base = https://app.example.com` instead of being
silently downgraded to the `/` fallback.
- mcp_crypto startup-gate error message now names both
`mcp_token_encryption_keys` (rotation list) and
`mcp_token_encryption_key` (single) so an operator using rotation
isn't misled into thinking only the singular form is valid.
Cleanup
-------
- Delete the unused _KNOWN_TRUSTED_ENDPOINT_HOSTS legacy re-export
shim in oidc.py (zero callers — a no-op that survived the Phase 4
oauth_ssrf extraction). Sphinx :data: docstring reference at
validate_discovered_endpoint updated to point at
turnstone.core.oauth_ssrf.KNOWN_TRUSTED_OAUTH_ENDPOINT_HOSTS
directly. The Google multi-origin allowlist is unaffected — it
lives at the canonical name and is read from oauth_ssrf.py:164.
- test_mcp_oauth_handlers TestValidateReturnUrl imports
_validate_return_url at module level instead of repeating the
import inside each test method.
- test_server_lifespan_mcp_crypto replaces a fragile
`messages.count("mcp_token_encryption_key") >= 2` substring trick
with `re.search(r"mcp_token_encryption_key(?!s)", messages)` —
asserts the singular form directly via negative lookahead.
Tests
-----
5448 pass (+13 vs the prior tip):
- TestValidateReturnUrl gains backslash-bypass, protocol-relative,
default-port, uppercase-host, and explicit-port-mismatch cases
alongside the original same-origin / cross-origin / scheme-
mismatch / path-only cases.
- TestInternalMcpStatusEndpoint asserts the `error` text never
reaches the read-scope wire (binary-path FileNotFoundError no
longer appears anywhere in the rendered response) and that the
coarse `has_error` boolean lights up correctly on the failed
server.
- TestInternalMcpStatusEndpoint also pins the no-mcp-client path to
`{"servers": {}}`.
- _routes_with_internal extended to include the
/api/_internal/mcp-status route so the new tests can exercise it
through TestClient.
- Existing test_startup_aborts_with_oauth_user_row_and_no_key
strengthened to require both singular and plural key names appear
in the error log.
Lands the OAuth flow that uses the token-at-rest store from the prior
commit: discovery (RFC 9728 PRM + RFC 8414 AS metadata with operator-
override precedence), PKCE S256 (mandatory — refuse AS without it),
RFC 8707 resource indicator on every authorize and token request,
RFC 7591 minimal one-shot dynamic client registration, authorization-
code exchange, refresh-token grant with re-read-after-acquire single-
flight lock, and the /v1/api/mcp/oauth/{start,callback} endpoints
mounted on both server and console.
Refactored:
- validate_url_no_ssrf, validate_discovered_endpoint, is_localhost,
effective_port, sanitize_log_text moved out of oidc.py into a shared
oauth_ssrf module; oidc.py re-exports for compatibility. The shared
helpers also expose async wrappers (validate_url_no_ssrf_async,
validate_discovered_endpoint_async) so OAuth-MCP discovery — invoked
from async handlers — does not block the event loop on the
synchronous socket.getaddrinfo call.
- MCPTokenStore.get_oauth_client_secret reader path added (the prior
commit was write-only)
- Storage protocol gains create/pop/cleanup_*_mcp_oauth_pending_state
and get_mcp_oauth_client_secret_ct (mirror OIDC pending-state
pattern: SQLite BEGIN IMMEDIATE select-then-delete, Postgres atomic
DELETE...RETURNING)
Refresh-grant correctness:
- When the AS omits refresh_token (RFC 6749 §6 — MAY rotate), the
existing refresh value is preserved at the OAuth-flow layer rather
than cleared, so production ASes (Google, Auth0 default, Okta) don't
force re-consent every hour
- expires_in accepts int, float, str-with-decimal — earlier int-coerce
through str() failed on float and silently dropped expiry tracking
- The refresh-grant `resource=` parameter (RFC 8707) is the canonical
MCP server URL, not the audience. Audience and resource are distinct
concepts; using audience as resource would mismatch the AS RS
allowlist.
Audience handling:
- _validate_token_audience accepts str or tuple; the callback resolves
accepted_audiences = {server_url, oauth_audience} and validates
against the set, so Auth0-style ASes that honor `audience=` (not
RFC 8707 `resource=`) issue tokens that pass audience-bound
validation
- build_authorize_url emits both `resource=` (RFC 8707) and
`audience=` (Auth0-style) per server config; comment documents which
AS implementations need which form
Security hardening:
- redirect_uri pinned to oidc_config.redirect_base instead of the
request Host header — closes the same Host-header injection PR #476
fixed for OIDC. Both /start and /callback return 503 with operator-
actionable hint when redirect_base is unset
- DCR registration runs under per-server asyncio.Lock with re-fetch
inside the lock, so concurrent /start callers don't both register
and overwrite each other's client_id (the second user's code is no
longer rejected on callback)
- /callback error branch pops the pending state row before redirecting
so a leaked state can't be replayed against a separately-obtained
code in the 60s cleanup window
- WWW-Authenticate Bearer parser handles RFC 7235 quoted-string
escapes (\" and \\) instead of the naive [^"]+ regex
- AS-controlled response bodies and error_description query params go
through sanitize_log_text before reaching exception messages or
audit details. AS error responses are parsed for the standard
RFC 6749 fields (error, error_description, error_uri), each
capped at 80 chars and run through redact_credentials to defend
against ASes that echo the request body back into their error
payload.
- oauth_as_issuer_cached is re-validated against the SSRF guard on
read; on rejection the column is cleared and PRM rediscovery runs
- DCR / token-endpoint / refresh-endpoint response bodies cap at 64
KiB (PRM/AS metadata cap stays at 256 KiB) so a hostile or
malfunctioning AS can't exhaust client memory.
- oauth_client_secret operator input capped at 1024 chars at the
admin-form boundary; longer plaintext rejected with 400.
- /start and /callback responses stamp `X-Frame-Options: DENY` so the
redirected pages can't be framed by attacker sites.
- delete_user cascades to mcp_user_tokens and mcp_oauth_pending so
user deletion no longer leaves dangling per-user OAuth state.
- Renaming or deleting an oauth_user MCP server purges per-user
tokens and pending OAuth state for the previous server name
(delete_mcp_oauth_rows_by_server_name). The OAuth tables key on the
mutable server_name; without this purge, a future server with the
same name (and an attacker-controlled URL) would silently rebind
prior user tokens. A future schema migration will replace the
server_name key with a server_id FK + ON DELETE CASCADE.
- get_user_access_token catches MCPTokenDecryptError (raised when no
installed key can decrypt the row, e.g. after key rotation) and
falls through to None so dispatch surfaces a re-consent rather than
crashing.
- oauth_user MCP server rows are skipped in the static auto-connect
path. Auto-connecting them at startup with empty headers fails the
AS check and trips the circuit breaker; per-user tokens come online
lazily once the user has consented.
Audit (mcp_server.oauth.* prefix):
- consent_started, consent_completed, consent_failed, token_refreshed,
token_revoked, dcr_registered. _audit_event is async and wraps
record_audit in asyncio.to_thread so the audit write doesn't block
the event loop. resource_id on the audit row is the immutable
server_id (PK UUID) so admin-driven server renames don't break
event correlation; server_name is exposed in detail for cross-
reference. dcr_registered detail.has_secret reflects whether the
DCR-issued secret was actually persisted (the prior code reported
has_secret=true even on persistence failure).
- _admin_mcp_action audits the immutable server_id, not the mutable
server_name (which is what the column is — the table's PK was
always server_id).
- All OAuth-flow log keys use the mcp_server.oauth.* prefix to match
the audit-action taxonomy.
Lifespan close-order in turnstone.server and turnstone.console.server
is reversed (LIFO) — mcp_oauth → mcp_crypto → oidc — to match init
order.
Deferred until the upcoming per-user pool integration:
- Multi-node refresh-lock contention via pg_advisory_lock
- DCR re-register on token-endpoint 401 (the dispatch path surfaces
those 401s)
- TTL-LRU caching of decrypted plaintext access tokens
- DNS-rebinding hardening (httpx Transport pin) — documented as
limitation in oauth_ssrf module docstring
Tests: 7 new test files / ~85 new tests covering discovery precedence
+ PRM quoted-string parsing, PKCE round-trip, SSRF helper extraction,
authorize/callback handlers including 503-on-no-redirect-base + DCR
concurrency + JWT audience polymorphism + callback-error-pops-pending,
refresh single-flight lock, refresh resource-vs-audience regression,
decrypt-error fallthrough, _db_servers_to_config skipping oauth_user,
pending-state CRUD round-trip.
Phase 3 of docs/design/oauth-mcp.md. Adds the Fernet/MultiFernet wrapper,
[security] config loader with rotation support, MCPTokenStore CRUD facade,
typed MCPTokenDecryptError that maps to the RFC's mcp_token_undecryptable_
key_unknown class, and a startup gate that fails loud when auth_type=
'oauth_user' rows exist without a configured encryption key.
Crypto module (turnstone/core/mcp_crypto.py):
- MCPTokenCipher wraps cryptography.fernet.Fernet + MultiFernet for
rotation; encrypt with first key, decrypt by trying each in order
- load_mcp_token_cipher_config reads [security] mcp_token_encryption_keys
(plural list) or mcp_token_encryption_key (singular), validates each
key is base64-decodable to exactly 32 bytes
- MCPTokenCipherConfig is repr=False with custom __repr__ that redacts
raw key bytes (defense in depth against accidental log/traceback leak)
- _key_fingerprint produces an 8-hex-char SHA-256 prefix for audit
attribution without exposing the key
- MCPTokenStore handles encrypt-on-write / decrypt-on-read for
mcp_user_tokens and mcp_servers.oauth_client_secret_ct
- get_user_token MUST NOT auto-delete the row on MCPTokenDecryptError
(test_get_user_token_with_wrong_key_raises_decrypt_error verifies
the row stays intact across a key-mismatch read)
- initialize_mcp_crypto_state / close_mcp_crypto_state lifespan helpers
shared between server and console
Storage protocol (5 new ciphertext-only methods):
- set_mcp_oauth_client_secret_ct (dedicated writer; deliberately NOT
added to MCP_SERVER_MUTABLE so generic update_mcp_server cannot write
the secret column)
- create_mcp_user_token, get_mcp_user_token,
update_mcp_user_token_after_refresh, delete_mcp_user_token
Server + console lifespans (turnstone/server.py + console/server.py):
- after OIDC init, count auth_type='oauth_user' rows; if any exist and
no encryption key is configured, log an actionable error and
raise SystemExit(1)
- without oauth_user rows, missing key is fine (lazy validation; admin
flip without restart returns 503 from the admin handler)
- app.state.mcp_token_cipher / .mcp_token_store populated when key
configured; None otherwise
Admin handlers:
- _require_token_store_for_oauth_secret pre-mutation gate validates
token_store availability and oauth_client_secret type BEFORE
storage.create_mcp_server / update_mcp_server runs, so a 503 from a
missing key never leaves an orphan row or partial-update state
- _apply_oauth_client_secret encapsulates the encrypt + audit write
used after the storage mutation; rolled out across both create and
update handlers
- 503 message references both mcp_token_encryption_key (singular) and
mcp_token_encryption_keys (plural for rotation)
- non-string oauth_client_secret payloads (false / 0 / lists / dicts)
are rejected with 400 instead of being str()-coerced
- when auth_type transitions away from oauth_user, the encrypted
secret column is cleared in the same admin call (with audit), so
flipping back doesn't silently resurrect a stale credential
Audit events (mcp_server.oauth.* per audit.py taxonomy; RFC's
mcp.oauth.* renamed for consistency):
- mcp_server.oauth.client_secret_set fired from admin handlers with
cleared:bool and key_fingerprint
- mcp_server.oauth.token_decrypt_failure fired from MCPTokenStore
.get_user_token when no installed key can decrypt; carries
key_fingerprints_attempted
Tests: 35 new tests across test_mcp_crypto, test_mcp_token_store,
test_server_lifespan_mcp_crypto, plus 6 admin-API tests covering the
no-orphan-row, no-partial-update, secret-clear-on-transition, and
non-string-secret-rejection invariants. Suite at 5337 (Phase 3 added
~50 tests including the rebase-imported skill suite).
cryptography>=42 promoted from transitive (lacme[tls]) to direct dep
since the encryption layer is now core, not optional.
Phase 4 (OAuth flow) wires the actual callers; Phase 3 adds only the
crypto layer and is exercised entirely by tests.
Adds the data model and admin UI surface required by the OAuth-MCP flow.
Phase 2 of the per-user delegation initiative.
Schema:
- migration 049 creates mcp_user_tokens (PK user_id, server_name) and
mcp_oauth_pending (PK state, indexed by created_at)
- eight new columns on mcp_servers: auth_type ('none' / 'static' /
'oauth_user', NOT NULL DEFAULT 'static') plus six oauth_* config
fields and oauth_as_issuer_cached
- post-upgrade UPDATE normalises auth_type to 'none' for streamable-http
rows whose headers are NULL/empty/'{}'; stdio rows are left at the
'static' default (auth_type is HTTP-auth-only)
- _schema.py kept in lockstep with the migration so metadata.create_all
and alembic upgrade produce identical shapes
- mcp_user_tokens / mcp_oauth_pending TypedDicts in _protocol.py for
Phase 3/4 use (no CRUD methods yet)
Storage / API:
- create_mcp_server gains the eight kwargs across protocol + sqlite +
postgresql
- MCP_SERVER_MUTABLE picks up auth_type and the six text oauth_* fields;
oauth_client_secret_ct is intentionally NOT in the whitelist — Phase 3
will own ciphertext writes via a dedicated method
- McpServerInfo + Create/Update Pydantic schemas extended; oauth_client_secret
accepted as plaintext input but discarded (Phase 3 wires encryption)
Admin handlers:
- _parse_auth_type validates against {'none', 'static', 'oauth_user'} and
rejects empty / unknown values; shared between create and update
- when auth_type changes away from 'oauth_user', the oauth_* config
columns are explicitly nulled in the same UPDATE so the row stays
consistent
- _clean_oauth_text caps text fields at 512 chars (URLs at 2048) to bound
admin write surface
- _mask_mcp_secrets now masks oauth_client_secret_ct to '***' regardless
of reveal=true (write-only field)
- audit detail dict redacts oauth_client_secret if present
Frontend:
- new "Multitenant Authorization" fieldset on the MCP-server modal with
three radio buttons (None / Shared / Per-user OAuth 2.1)
- conditional OAuth subform: AS URL, registration mode (preregistered /
dcr; cimd is future), client ID, client secret, scopes, audience
- secret input is autocomplete=off and never round-trips on edit
- audience auto-populates from the MCP server URL on blur
- headers textarea hidden and submitted as {} when auth_type is 'none' or
'oauth_user' so flipping the radio cleans up server-side state
Tests: storage round-trip for the new columns, oauth_pending table smoke,
migration 049 upgrade/downgrade with stdio-vs-http normalisation, four
admin-API tests for auth_type validation and oauth_*-clear-on-flip-away.
Suite passes 5284 (matched pre-Phase-2 baseline 5267 + 17 new).
Stacks on Phase 0; no behavioural change for existing rows.
Phase 0 of the OAuth-MCP RFC: prepare MCPClientManager for the per-(user,
server) session pool that lands in Phase 5, without changing static-path
behavior.
Two changes:
1. Hardening helpers _pre_close_streams and _tcp_probe rename their first
parameter from `name` to `key`. Type stays `str` for now; widening to
`str | tuple[str, str]` happens in Phase 5 when callers actually pass
tuples. _safe_close_stack takes the stack directly and is unchanged.
2. The eleven parallel name-keyed dicts (_sessions, _per_server_stacks,
_per_server_tools, _per_server_resources, _per_server_prompts,
_supports_list_changed, _supports_resources, _supports_resource_list_changed,
_supports_prompts, _supports_prompt_list_changed, _server_streams) are
consolidated into _static_servers: dict[str, StaticServerState]. Server-
level state (circuit breaker, notification debounce, last-error,
db-managed, merged catalog maps, listener lists) stays on the manager,
unchanged.
PoolEntryState is defined for Phase 5 use but no code instantiates it. The
typed map declarations (dict[str, StaticServerState] vs dict[tuple[str, str],
PoolEntryState]) make accidental cross-keying lookups easier to catch.
PR #296 hardening preserved exactly:
- pre-close-streams atomic take-and-clear before stack teardown
- stale-session-and-stack guard at _connect_one top: both state.session and
state.stack checked, cleared independently, entry preserved (not popped)
- transport-error session-eviction in dispatch sets state.session=None only,
leaving stack/streams for the next connect-time guard sweep
- _safe_close_stack CancelledError suppression unchanged
- TCP probe before streamablehttp_client unchanged
- future.cancel() after TimeoutError in all sync bridges unchanged
- notification debounce stays manager-level (not migrated into the dataclass)
Refresh helpers (_refresh_server_tools/_resources/_prompts) snapshot
state.session into a local immediately after the None guard so concurrent
transport-error eviction during await cannot null the session reference
mid-call.
Tests: shared _seed_static_state helper in tests/conftest.py replaces eleven
direct dict mutations; new test_reconnect_preserves_static_state_identity
guards the entry-preservation invariant. Pass count rises 5266 → 5267.
Deletes the _periodic_refresh task and its supporting state
(_refresh_task, _refresh_failures, _refresh_backoff_until,
_REFRESH_BACKOFF_BASE/MAX, _DEFAULT_REFRESH_INTERVAL, refresh_interval
kwarg) from MCPClientManager. Push notifications and operator-driven
manual refresh now cover all catalog-update needs; the long-running
4-hour timer was dead complexity that obscured the per-user pool
work to come.
Catalog freshness on auto-reconnect is preserved by scheduling an
unblocking _refresh_server task on the mcp-loop after _connect_one
succeeds; the calling thread returns immediately so half-open
recovery latency does not double. Adds MCPClientManager.reconnect_sync
(clears the circuit, closes any existing session, calls _connect_one,
clears stale catalog on failure).
Wires a new pair of operator endpoints —
POST /v1/api/admin/mcp-servers/{name}/refresh and
/v1/api/admin/mcp-servers/{name}/reconnect — that fan out to all
nodes through the existing _internal route family, with per-row
"Refresh" and "Reconnect" buttons in the MCP Servers admin tab.
The new node-internal paths /api/_internal/mcp-{refresh,reconnect}/
are gated to the approve scope to prevent direct unprivileged
reconnects bypassing the console's admin.mcp gate. Internal
endpoints return generic error messages and a filtered status
payload (no command/url) to keep transport details admin-gated.
Drops the [mcp] refresh_interval setting, the
--mcp-refresh-interval CLI flag, and the matching config-mapping
entry; updates docs/architecture.md, docs/tools.md,
docs/settings.md, and the three PlantUML diagrams that referenced
the periodic loop.
Tradeoffs (intentional):
- Idle nodes will not auto-rejoin a recovered MCP server until
traffic arrives or an operator clicks Reconnect. The previous
background reconnection loop is gone by design — push
notifications + operator controls replace it.
- Console fan-out blocks on the slowest node (existing pattern);
not changed here.
This is Phase 1 of the OAuth-MCP series — feature subtraction
ahead of per-user state.
* feat(skills): paste SKILL.md to auto-fill the Create Skill modal
When a user pastes an Anthropic-style SKILL.md (YAML frontmatter +
markdown body) into the Create Skill content textarea, the frontend
sniffs the leading ``---``, posts the raw text to a new backend parse
endpoint, and populates name / description / tags / author / version /
license / compatibility / allowed_tools from the parsed fields. The
textarea is left with the body only (frontmatter stripped), and a toast
reports how many fields were set vs. kept (already-typed values are
preserved).
Backend
- ``POST /v1/api/admin/skills/parse`` (admin.skills permission) wraps
the existing ``turnstone.core.skill_parser.parse_skill_md`` so admin
imports and external installs share one parser. ``ParseSkillRequest``
/ ``ParseSkillResponse`` schemas added; OpenAPI spec + sync/async
console SDK methods updated.
- Hardening: 32 KiB cap on ``raw`` (Pydantic ``max_length`` + handler
enforcement); ``Content-Length`` pre-check returns 413 before any body
buffering; parse offloaded via ``asyncio.to_thread`` so deeply-nested
YAML cannot stall the event loop.
Frontend (turnstone/console/static)
- New paste handler with optimistic paint (raw text shown immediately,
textarea disabled + ``aria-busy`` flipped, hint switches to
"Parsing...") so the round-trip is visible on slow networks.
- ``AbortController`` + generation guard (``_ctmPasteController``) so a
fresh paste or modal close cancels a stale fetch — the previous
handler's callbacks see the controller has been replaced and bail
before touching the DOM.
- Non-destructive overwrite: ``_setSkillFormField`` returns "filled" /
"skipped" / "absent" and refuses to clobber non-empty values. Toast
reports counts.
- Bumps ``#toast`` z-index above modal overlays (was 200 vs. modal 600
— toasts fired while a modal was open were invisible). Console-wide
fix exposed by this being the first feature to fire toasts mid-modal.
HTML / CSS
- New ``.skill-paste-hint`` line above the textarea announcing the
affordance, sized to match surrounding ``.label-hint`` text.
- ``aria-describedby`` ties the hint to the textarea; ``aria-live=
"polite"`` announces the busy-state transition to screen readers.
- "Skill Content" heading hint reworded "system message — ..." →
"available: ..." and the variables row label "Variables" → "Used"
to disambiguate available vs. in-use template variables.
Tests
- 11 new cases in ``tests/test_skill_parse_api.py``: happy paths
(full / minimal / nested-metadata / unquoted-colon recovery),
malformed YAML 400, missing/blank/missing-name 400, RBAC 403, raw
body 32 KiB cap (Content-Length pre-check), chunked-encoding bypass
forces the application-layer cap. Test pins ``raw_frontmatter``
omission so a future ``dataclasses.asdict`` refactor can't silently
leak the full YAML dict back to clients.
Validation
- 5146 / 5146 ``pytest -k "not live"`` pass.
- ``ruff`` + ``mypy`` clean on changed sources.
- ``node -c`` clean on governance.js.
- Two-stage code review (full pipeline + bug+quality re-review of the
fix patches) applied; all confirmed findings addressed.
* fix(skills): Copilot PR #477 review fixes (cumulative bug-1, bug-2, q-1)
bug-1 (server.py): Content-Length pre-check was clamped to 32 KiB —
the same number as the per-string char cap on ``raw``. A legitimate
``raw`` of exactly 32 KiB produces a JSON body well above 32 KiB once
the ``{"raw":"..."}`` wrapper and any escaping is added, so valid
near-max requests were 413'd. New constant
``_PARSE_SKILL_MAX_BODY_BYTES = _PARSE_SKILL_MAX_CHARS * 4`` admits the
wrapper + multibyte expansion while still refusing obviously oversized
payloads early; the per-string ``len(raw)`` check stays authoritative.
bug-2 (governance.js): hideCreateTemplateModal aborted the inflight
paste controller and nulled the global, but the handler's ``.catch``
and ``.finally`` guard each DOM mutation behind ``_isCurrent()`` —
both bail when the controller has been nulled, leaving the textarea
``disabled`` + ``aria-busy`` and the hint stuck on "Parsing…".
Reopening the modal landed on a poisoned state. The second-pass
review's q-2 cleanup that dropped the show-side defensive reset
missed this scenario — the verifier's reachability argument confused
"controller is null" with "UI state is reset"; the two are
independent. Hide now resets the paste-induced visible state
alongside the abort.
q-1 (console_spec.py): error_codes for the parse endpoint listed only
400; handler also returns 413 for oversized bodies. Added 413; kept
403 implicit per the convention sibling admin endpoints follow.
Test fixup: bumped the Content-Length test payload to 200 KB so it
clearly exceeds the new 128 KB pre-check threshold; otherwise it was
falling through to the per-string check and duplicating
test_oversized_raw_chunked_returns_413's coverage.
PR #476 review feedback (Copilot, oidc.py:584,616):
1. initialize_oidc_state's docstring claimed "on any failure
enabled is False" but the JWKS-prefetch failure branch
intentionally keeps enabled=True so the callback's lazy-fetch
retry can recover from a transient IdP issue at startup.
Docstring rewritten to spell out the three post-conditions:
disable, JWKS-failure-keeps-enabled, success.
2. The long-lived httpx.AsyncClient was created up front, then
three disable branches (discovery exception, discovery-returned-
disabled, missing redirect_base) returned without closing it,
leaving sockets held until shutdown.
Restructured: discovery now uses a transient AsyncClient inside
a context manager (closed at exit). The long-lived client is
only created after the disable checks pass. The JWKS-failure
branch still legitimately keeps the client open because the
lazy-retry path needs it.
The pre-existing single-client-passthrough test was replaced
with three more specific tests: long-lived client only goes to
fetch_jwks (not discover_oidc); discovery-exception path leaves
http_client=None; missing-redirect_base path leaves
http_client=None.
q-4: tests/test_oidc.py's _make_config and tests/test_oidc_handlers.py's
_make_oidc_config built the same OIDCConfig with sensible defaults but
had drifted — only the handlers helper set redirect_base. After b3
made redirect_base operationally required, every test_oidc.py test
that exercised redirect_base had to override it explicitly. A future
test could omit redirect_base and silently exercise the wrong
production path.
Moves make_oidc_test_config to tests/conftest.py with the more
complete handler-version defaults (including redirect_base). Both
test files import it under their existing local alias
(_make_config / _make_oidc_config) so the 60+ call sites in
test_oidc.py and the handler tests don't have to change.
q-5: section banner '# Exception' (singular) at oidc.py:79 became
inconsistent after b5 (callback robustness) added OIDCKeyNotFoundError.
Renamed to '# Exceptions'.
The OIDC perf batch added storage.count_users() and migrated the two
OIDC handlers (handle_oidc_authorize, handle_oidc_callback) but missed
handle_auth_status — which still ran storage.list_users() then
len(users) > 0 for the same has-any-users gate.
count_users() is one COUNT(*) round-trip vs list_users() rehydrating
every row dict. Wrapped in asyncio.to_thread to match the OIDC handler
pattern; the async handler no longer blocks the event loop on storage
I/O for what's effectively an existence probe.
bug-2 (Postgres) — replace_oidc_roles read existing rows under default
READ COMMITTED with no row lock. Two concurrent OIDC callbacks for the
same user_id (racing token refreshes with differing claim sets) could
both observe the same baseline and produce a final role state matching
neither caller's intent. Adds .with_for_update() to the SELECT so the
existing rows for this user are locked for the duration of the
transaction.
The lock is per-user_id, not table-wide; unrelated user writes are
unaffected. Empty result sets acquire no locks, so a brand-new user
with no rows yet still allows two callers to proceed and merge via
ON CONFLICT DO NOTHING — that's a permissive race that self-heals on
the next reconciliation cycle, documented in code.
perf-1 (SQLite) — replace_oidc_roles took the SQLite global write
lock unconditionally via BEGIN IMMEDIATE before reading. Steady-state
re-logins (claims unchanged, no INSERT/DELETE needed) paid the lock
cost for nothing and serialised against unrelated writers.
Replaces with a double-check pattern: phase 1 reads under the default
deferred transaction (no write lock), computes the diff, and returns
(set(), set()) on no-op. Phase 2, only when mutation is needed,
commits the read txn, escalates to BEGIN IMMEDIATE, RE-READS, and
re-computes the diff under the lock before writing. The returned
(added, removed) reflects what was actually written, so caller logging
in apply_role_mapping stays truthful even when concurrent writers
shifted state between the two reads.
The OR IGNORE on insert is now defense-in-depth (the lock makes it
unnecessary) but kept as a safety net.
The 8-commit OIDC stack added TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS
(operator allow-list for cross-host IdP discovery endpoints) and
promoted TURNSTONE_OIDC_REDIRECT_BASE to required, but the docs drifted
in two places:
q-1 — Troubleshooting > "OIDC not configured" still listed three
required env vars. An operator hitting the missing-redirect-base
startup error landed on a debugging entry that didn't mention the
variable they were missing. Fixed; added a separate troubleshooting
entry naming the exact log message produced by initialize_oidc_state
when redirect_base is unset.
q-2 — TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS was undocumented entirely.
Added a row to the env-var table and a new "Cross-host endpoints"
section explaining when the knob is needed (Google is the canonical
multi-origin IdP, but it's auto-handled; the env var is for any other
IdP whose discovery doc legitimately references hosts beyond the
issuer's origin). Added a troubleshooting entry pointing at the new
section.
If apply_role_mapping raised after create_oidc_user committed (transient
storage failure, race with role deletion, etc.), provision_oidc_user's
inline safety-net was skipped — and on retry the existing-identity
branch never reached the safety-net code, leaving the user permanently
stranded with zero roles.
Extracts _ensure_default_role(storage, user_id, desired_role_ids=None)
helper. Calls it on BOTH the new-user and existing-identity paths so a
user stranded by a transient failure recovers on next login.
desired_role_ids is a hint that lets the helper skip list_user_roles
when claim-driven mapping populated at least one role; the new-user
path was already paying that query, the existing-identity path now
pays it only when claim mapping returned an empty desired set.
Documents the admin-strip behavior in the helper docstring: stripping
all roles from an OIDC user no longer locks them out, since the next
login will re-grant builtin-viewer (assigned_by='oidc-default'). The
documented way to deny an OIDC user is to unlink their OIDC identity
via the admin endpoint, not to strip roles. The pre-fix behavior
(stripped user actually locked out) was the bug.
The 'oidc-default' vs 'oidc' assigned_by distinction is preserved:
apply_role_mapping's revocation lane only touches 'oidc' rows, so the
safety-net role survives every subsequent login regardless of claims.
Six new tests cover both paths, the hint short-circuit, the
list_user_roles fallback, the missing-builtin-viewer no-op, and the
self-heal regression case for already-stranded users.
q-5: _derive_username's UUID-retry tier (oidc.py:923-933) was untested.
After perf-6 collapsed tier-2 to a single find_existing_usernames call,
the only remaining tail was the 3-attempt UUID-retry loop and the final
raise. New TestDeriveUsername class covers:
- falls into UUID retry when all 10 suffix candidates are taken
- UUID retry succeeds on the second attempt after one collision
- UUID retry exhausted -> raises OIDCError
q-8: filled the unit-level coverage holes the multi-stage review flagged:
- test_validate_id_token_retry_after_kid_rotation — direct unit test of
the OIDCKeyNotFoundError path with real RS256 keys + JWKS rotation
(previously only exercised end-to-end through the handler).
- test_callback_uses_pending_audience_not_handler_audience — pins down
the bug-3 fix by decoding the issued JWT cookie and asserting aud
matches the audience stored at /authorize time, not the handler param.
- test_apply_role_mapping_int_claim / _dict_claim — exercises the
else: values = [str(claim_value)] branch for non-string non-list
claim shapes.
- TestFetchJWKS — non-200 status, non-dict body, dict-missing-keys,
keys-not-list, transport network error.
- TestExchangeCode network/4xx/5xx error tests (the non-dict-body case
already shipped in batch 5).
Also a small production hardening that fell out of writing the
TestFetchJWKS::test_fetch_jwks_non_dict_body_raises test: fetch_jwks now
guards isinstance(result, dict) before result.get("keys"), matching the
shape-check pattern that discover_oidc and exchange_code already use.
A list/null body now surfaces as OIDCError("...not a JSON object") rather
than AttributeError leaking up to the lifespan.
Eleven small maintenance fixes; no behavior change beyond bug-3.
bug-3: pending.get('audience', audience) couldn't fall back because
pop_oidc_pending_state always returns a dict with the audience key
set verbatim from a non-null TEXT column. Replaced with
pending.get('audience') or audience to cover the empty-string case
defensively. Comment explains the security rationale.
q-1: extract _env_or_cfg_str / _env_or_cfg_bool helpers in oidc.py;
load_oidc_config's six near-identical env-or-config blocks collapse
to one-liners. role_map / trusted_endpoint_hosts / redirect_base
retain bespoke parsing.
q-3: discover_oidc narrows except (httpx.HTTPError, ValueError, KeyError)
with exc_info=True.
q-4: OIDC_STATE_TTL_SECONDS = 300 constant in oidc.py; auth.py imports
and passes it explicitly. Storage signatures keep the literal default
(storage layer doesn't know OIDC TTL semantics).
q-6: hoist runtime imports (OIDCError, OIDCKeyNotFoundError, exchange_code,
fetch_jwks, provision_oidc_user, validate_id_token, build_authorize_url,
generate_pkce_verifier) to module scope in auth.py. The genuine cycle
is only oidc._derive_username -> auth.is_valid_username, kept
function-scoped. test_oidc_handlers.py mock targets repointed to
turnstone.core.auth.X to match the new binding.
q-7: comment + docs explain the 'oidc' vs 'oidc-default' assigned_by
marker distinction.
q-9: OIDCIdentity / OIDCPendingState TypedDicts in storage protocol.
Implementations construct via TypedDict syntax so mypy structurally
verifies all required fields.
q-10: fetch_jwks narrows except (httpx.HTTPError, ValueError); docstring
matches.
q-11: rename generate_pkce_pair -> generate_pkce_verifier; return only
the verifier (build_authorize_url already recomputes the challenge).
q-12: extract _buildOidcRow helper in admin.js so future field additions
go in one place.
q-13: OIDCConfig docstring lists startup-config vs discovery-derived
field groups.
Eight independent perf wins on the OIDC hot path:
perf-1: list_users() full-scan setup-gate replaced with new count_users()
on both authorize and callback. Saves a full users-table fetch per login.
perf-2: handle_oidc_callback's sync DB chain wrapped in asyncio.to_thread
for cleanup, pop_oidc_pending_state, count_users, and provision_oidc_user.
handle_oidc_authorize gets the same treatment for count_users and
create_oidc_pending_state. Event loop no longer blocks for the full
callback duration on Postgres deployments.
perf-3: apply_role_mapping N+1 collapsed via new replace_oidc_roles
storage method. One transaction handles the diff + insert + delete
instead of 2N+1 commits per login. Returns (added, removed) so the
caller can still emit per-role audit logs.
The diff respects the documented invariant "manually-assigned roles
are never touched" — desired_role_ids is filtered against rows where
assigned_by != 'oidc' before computing added/removed. This prevents a
PK conflict (Postgres lockout) or silent OR-IGNORE no-op (SQLite lying
return) when admin-ui or oidc-default already holds the same role_id.
perf-4: provision_oidc_user no longer re-queries list_user_roles after
apply_role_mapping. The new-user builtin-viewer fallback is gated on
desired_role_ids being empty, which is information apply_role_mapping
already returned.
perf-5: JWKS refetch dedup via asyncio.Lock on app.state. Both lazy-fetch
(cold-start recovery) and rotation paths share the same lock with a
double-check pattern: re-resolve kid against the current cache before
issuing a new GET. N concurrent callbacks during rotation now produce
at most 1 fetch.
perf-6: _derive_username's 9-suffix loop collapsed via new
find_existing_usernames(candidates) -> set query. Worst case drops
from 13 sequential queries to 1 + up-to-3 UUID-retry queries.
perf-7: cleanup_expired_oidc_states gated to once-per-60s per process
via app.state.oidc_last_cleanup_monotonic. The pop already deletes
the consumed row; the bulk cleanup is only relevant for abandoned
authorize flows, so frequency was overkill.
perf-8: Long-lived httpx.AsyncClient stashed on app.state.oidc_http_client
by initialize_oidc_state. discover_oidc/fetch_jwks/exchange_code accept
an optional client= kwarg; when set, skip the per-call AsyncClient
context-manager. New close_oidc_state lifespan teardown closes it.
Tests pass client=None to keep the transient-client legacy path.
New storage methods (sqlite + postgresql):
- count_users() -> int
- find_existing_usernames(candidates) -> set[str]
- replace_oidc_roles(user_id, desired) -> (added, removed)
Four small hardening fixes on the OIDC callback hot path:
bug-4: JWKS rotation retry was matching the substring 'not found in JWKS'
inside an OIDCError message. A future rephrasing would silently break
key rotation. Adds OIDCKeyNotFoundError(OIDCError); validate_id_token
raises the subclass at the kid-not-found site; handle_oidc_callback
catches it explicitly. Other 'not found' errors in validate_id_token
remain as plain OIDCError.
bug-5: tokens['id_token'] raised KeyError if the IdP returned 200 without
id_token. exchange_code now rejects non-dict response bodies; the
callback validates id_token shape (must be non-empty str) before
passing to validate_id_token. Both raise OIDCError, surfaced as the
standard 'Authentication failed' redirect.
bug-6: shared_static/auth.js — the OIDC error display raced showLogin's
/v1/api/auth/status fetch via a 300ms setTimeout. showLogin now takes
an optional oidcError parameter and paints it after _switchMode clears
the error, in both the success and catch branches of the fetch.
sec-4: oidc.py exchange_code's non-200 OIDCError interpolated up to 500
bytes of attacker-controlled IdP body, which then went to log.warning
via 'OIDC callback failed: %s'. CRLF in resp.text could forge log
lines. New _sanitize_log_text helper escapes control chars via
unicode_escape and caps at the rendered length.
provision_oidc_user previously called create_user (INSERT OR IGNORE
on SQLite — silent no-op on UNIQUE conflict), then create_oidc_identity
(also INSERT OR IGNORE), then apply_role_mapping which writes user_role
rows for the supposedly-new user_id. On a username TOCTOU race or
concurrent (issuer, sub) double-create, both inserts no-opped but
user_role rows were already written — leaving orphan rows pointing
at a user_id that doesn't exist.
PostgreSQL's create_user raised IntegrityError instead of silently
no-opping so it produced a misleading 'Authentication failed' error
without orphans, but the user-facing UX was equally poor.
Adds StorageConflictError to the storage protocol and create_oidc_user
that does both inserts in one transaction. Username collision and
(issuer, subject) collision both raise StorageConflictError, mapped
to OIDCError by provision_oidc_user. Crucially the new code does not
silently bind a colliding-username new identity to the existing user
— that would be an account-takeover vector. It raises.
SQLite uses BEGIN IMMEDIATE inside the try block so lock-contention
errors surface as StorageConflictError instead of leaking the raw
sqlalchemy OperationalError.
PostgreSQL relies on SQLAlchemy 2.x begin-on-demand semantics; the
explicit conn.commit()/rollback() in the catch block is the only
materialization path. Discrimination on PG uses
exc.orig.diag.constraint_name with message-substring fallback.
_build_oidc_redirect_uri previously fell back to the request Host
header when redirect_base was unset. With a permissive reverse proxy
or direct backend access, a spoofed Host minted an authorize URL
pointing to attacker-controlled host — combined with a permissive
IdP redirect_uri allowlist this enables auth-code interception.
There is no production scenario where a Host-derived redirect_uri is
correct, so this fails closed:
- initialize_oidc_state checks redirect_base after discovery succeeds
and disables OIDC (with an explicit error log naming the env var)
if it's empty. Runs before fetch_jwks so a misconfigured deploy
doesn't make a wasted JWKS call.
- _build_oidc_redirect_uri simplifies to f"{redirect_base}/v1/api/auth/oidc/callback".
request parameter dropped; both call sites (handle_oidc_authorize,
handle_oidc_callback) updated.
- docs/oidc.md promotes TURNSTONE_OIDC_REDIRECT_BASE from "Recommended"
to "Required" with the security rationale.
The OIDC discovery + JWKS prefetch block was duplicated byte-for-byte
between turnstone/server.py and turnstone/console/server.py. The bare
except branch in that block also left app.state.oidc_config unchanged
on unexpected exceptions — leaving the runtime with enabled=True and
empty endpoints, producing malformed authorize URLs.
Extracts initialize_oidc_state(app_state) into turnstone/core/oidc.py
which guarantees a coherent post-condition on every code path:
- discovery exception -> oidc_config replaced with enabled=False, jwks_data=None
- discovery returns enabled=False -> jwks_data=None
- JWKS prefetch fails -> jwks_data=None but enabled=True preserved (the
callback's lazy-fetch retry path remains the recovery)
- success -> oidc_config + jwks_data both populated
Also hardens discover_oidc against non-dict discovery responses
(list/null/string/int) — previously these raised AttributeError out
of doc.get and propagated past the lifespan's bare except.
server.py and console/server.py lifespan blocks collapse to a single
await initialize_oidc_state(app.state) call.
OIDC discovery-document endpoints (token_endpoint, jwks_uri,
userinfo_endpoint) were stored verbatim in OIDCConfig and later passed
to httpx without revalidation. Only the issuer URL was checked. A
hostile or compromised IdP could return token_endpoint pointing to an
internal IP (169.254.169.254, 10.0.0.0/8, etc.) and Turnstone would
POST the client_secret there.
Extracts the existing scheme/userinfo/SSRF check into
_validate_url_no_ssrf, adds validate_discovered_endpoint that runs the
same checks plus an issuer-binding check, and wires it into
discover_oidc for authorization_endpoint, token_endpoint, jwks_uri,
and userinfo_endpoint (when present).
Issuer binding accepts:
- Same (scheme, hostname, effective port) as the issuer.
- A hostname in _KNOWN_TRUSTED_ENDPOINT_HOSTS for the issuer (Google's
multi-origin discovery is in the allow-map by default).
- A hostname in OIDCConfig.trusted_endpoint_hosts, settable via
TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS env var or config.toml, for
IdPs not in the static map.
Effective port comparison treats https://host and https://host:443 as
the same origin (urllib.parse.urlparse leaves the explicit form's port
as 443 and the implicit form's as None).
24 new tests cover the validator, the Google known-hosts path, the
operator allow-list, default-port equivalence, foreign-host
rejection, private-IP rejection, embedded credentials, and DNS
rotation between issuer check and endpoint use.
* feat(console): inline node picker replaces back-to-console banner
Drops the 32px banner the console proxy used to inject above proxied
server-UI pages and replaces it with an inline node-id pill in the
existing #ui-header. Click the pill to open a dropdown that lists
healthy nodes (health dot, ws count, reachable/degraded/unreachable
text) plus a top-row link back to the console.
Reuses the .ws-tab-dropdown shell from ui/static/style.css for
animation, shadow, theme override, and item layout, so the picker
visually matches the workstream-tab chevron menu it sits next to.
Keyboard nav (ArrowDown/Up/Home/End/Tab/Escape) mirrors the chevron
menu's handler with cross-reference comments at both sites.
Lazy-fetches /v1/api/cluster/nodes against the console origin
(bypassing the prefix shim) on first open.
Reclaims 32px of vertical space, consolidates three separate
"you're on node X via console" indicators into one, and turns the
wayfinding chrome into a real cluster-nav primitive.
* fix(console): address Copilot review on node picker
- Request /v1/api/cluster/nodes?limit=1000 (collector's hard cap)
instead of relying on the default 100 — clusters with more than
100 nodes were silently dropping rows from the picker.
- Hand off focus to the first menu item after the async fetch
resolves: openMenu()'s deferred focus hook ran while only the
skeleton was in the DOM, so first-open keyboard users were
stranded on the trigger until they pressed an arrow key.
- Tab now closes the menu without preventDefault, so focus moves
to the next focusable element on the first press (ARIA APG menu
pattern). Escape still preventDefault + returns to the pill.
- Cap pill max-width at 240px and ellipsize the id span; node ids
are accepted up to 256 chars upstream and could otherwise push
the title and right-side controls off the appbar. Pill carries
a title attribute so the full id is still legible on hover.
* fix(session): properly inject queued user messages mid-loop
Two queued-user-message bugs in ``ChatSession.send()``.
**Mid-tool-call: ``Unexpected role 'tool' after role 'user'`` on Mistral.**
The ``supports_tool_advisories`` capability flag (default False for
unknown openai-compatible models) routed cap-off providers down a
short-circuit branch in ``_collect_advisories`` that called
``_flush_queued_messages`` directly. That appended a ``user`` turn
between ``assistant(tool_calls)`` and ``tool``, which mistral-common's
``_validate_message_order`` rejects with a 400.
Drop the flag. All providers now run the unified path: queued user
messages become ``UserInterjection`` advisories that ride inside the
tool result envelope via ``wrap_tool_result``, splicing
``<system-reminder>`` text into the tool message's content. Role
sequence stays ``assistant → tool``. Live-confirmed on Mistral
medium and Qwen3 — both correctly distinguish system-reminder from
tool stdout in their reasoning.
**Mid-stream: queued message orphaned until next user send.**
After a no-tool assistant turn, ``_flush_queued_messages`` would
append the queued user message to history and the loop would
``break``, leaving the message at the tail of history with no
model response. Visible as "two sends to get one reply".
``_flush_queued_messages`` now returns ``bool``. The no-tool branch
``continue``s on drain instead of ``break``ing, so the model gets a
turn over the extended history.
Tests:
- ``test_collect_advisories_drains_text_queued_messages_to_persistent``
pins the unified-path drain (text-only queue → ``UserInterjection``,
no separate user turn appended to ``self.messages``).
- ``test_send_continues_when_messages_queued_during_streaming`` pins
the loop-continue behavior (fails with 1 stream call pre-fix,
passes with 2 post-fix).
* fix(session,ui): reject queued attachments + paperclip busy state
Copilot pointed out that the attachment-bearing branch in
``_collect_advisories`` had the same role-ordering bug as the
text-only path that 802658f fixed: an attachment-bearing queued
item would still call ``_append_user_turn`` mid-tool-call,
injecting ``user`` between ``assistant(tool_calls)`` and ``tool``.
Pragmatic fix: don't allow attachments to be queued at all.
**Backend.** ``ChatSession.queue_message`` raises a new
``AttachmentsNotQueueableError`` when called with non-empty
``attachment_ids``. The interactive ``/send`` route catches it,
releases reservations via the existing ``_release_reservation_on_fail``
hook, and surfaces ``status: "attachments_busy"`` to the caller
with the IDs in ``dropped_attachment_ids``. The coord adapter
mirrors the cleanup (releases the soft-locked reservation taken
for ``_send_id``) so the create-with-attachments path can't leak.
Now that the queue can never carry attachments, the per-item
``att_ids`` slot is gone:
- Queue tuple slimmed ``(cleaned, priority, att_ids)`` →
``(cleaned, priority)``.
- ``_flush_queued_messages`` collapses to a single combined-text
user turn (no attachment branch).
- ``_collect_advisories`` queue-drain pushes ``UserInterjection``
advisories only (no ``attachment_items`` list).
- ``dequeue_message`` no longer unreserves (queue can't reserve).
- ``_resolve_attachment_ids`` had no remaining production callers
and is deleted along with the tests that exercised it in
isolation.
**Frontend.** ``Composer.setBusy`` disables the paperclip whenever
busy (regardless of ``queueWhileBusy``) — text still queues,
attachments don't. ``chat.css`` gains a ``.composer-attach:disabled``
rule (mirrors the existing ``.composer-send:disabled`` treatment)
so the affordance actually looks unclickable instead of falling
through to the UA default. ``title`` and ``aria-label`` are kept in
sync for AT users (WCAG 4.1.2).
Both interactive and coordinator UIs handle the new
``attachments_busy`` response with a chat-surface error bubble:
> Attachments can't be sent while the assistant is working.
> Send a text-only message now, or wait and resend with attachments.
Chips stay in the composer so the user can retry once idle.
**Tests.** Replaced the now-impossible ``TestQueuedWithAttachments``
class with a rejection-coverage class. Rewrote the
``_queue_with_attachment`` route-test fixture to reserve directly
via ``reserve_attachments`` (the queue path no longer reaches the
reserved state). Added a route-level test for the new
``attachments_busy`` contract.
* Bound search tool output against pathological inputs
Replaces the per-line truncation with a fully bounded pipeline so the
search tool can no longer overflow the LLM context — or OOM the parent —
on minified bundles, multi-GB JSONL records, or huge result sets.
Backend:
- Prefer ripgrep when on PATH; grep is the fallback. Detection is
cached via functools.cache.
- ripgrep flags do most of the bounding natively: --max-columns 1024
+ --max-columns-preview, --max-filesize 10M, --max-count 100,
--no-config, --no-messages, plus negative globs for the same
noisy directories grep has been excluding.
- ripgrep added to the Dockerfile.
Streaming subprocess (_search_capture):
- subprocess.Popen with a streaming, byte-capped stdout read (4 MB).
Defends against single-line files (training data, minified bundles)
that would have OOM'd the previous subprocess.run capture.
- threading.Timer watchdog enforces tool_timeout even when the
pipe read is blocked in the kernel — proc.wait(timeout=…) alone
was insufficient because the read sat ahead of it.
- Stderr drained in a daemon thread to avoid pipe-deadlock when the
child writes to stderr while we're still reading stdout. Cap on
captured stderr keeps a hostile child from growing the buffer.
Tier-based formatter (_format_search_results):
- Tier 1: full path:line:content output, stream-emitted with a
running-cost short-circuit so we never materialize past the budget.
- Tier 2: K samples per file with overflow notes; K is computed
analytically from budget / file_count / avg-line-length so we hit
the right ladder rung in a single pass.
- Tier 3: per-file counts only, also budget-bounded with a tail line
reporting the omitted files. Sorted by descending count.
- Total output budget (32 KB) is well under tool_truncation, so the
head+tail _truncate_output strategy never silently drops middle
files in a search result.
Argument injection fix:
- The ripgrep arg list was missing the `--` separator that the grep
branch already had. With auto_approve on the search tool, that was
exploitable: path='--pre=COMMAND' would have made ripgrep run the
script as a per-file preprocessor and surface its stdout. Added
`--` and a regression test.
State-machine cleanup in _exec_search:
- rc < 0 (signal-killed by something other than us) now surfaces a
dedicated 'killed by signal N' message instead of being parsed as
success.
- capped + zero parsed records (e.g. one multi-MB line with no \n)
now returns a dedicated byte-cap message instead of the malformed-
output message that previously masked the real cause.
- _report_tool_result descriptions now match the returned payload
(no more 'no matches' tag on a 'malformed' payload).
Defence-in-depth on env scrub:
- RIPGREP_CONFIG_PATH, GIT_CONFIG, GIT_CONFIG_GLOBAL, GIT_CONFIG_SYSTEM
added to _EXPLICIT_SCRUB. We pass --no-config on the rg CLI today,
but if a future caller forgets the flag, an attacker who can set
one of these env vars could plant a config containing --pre=… and
recreate the same RCE shape.
Tests:
- TestSearchLineTruncation rewritten to mock _search_capture instead
of subprocess.run (the previous tests passed ChatSession kwargs
that no longer satisfy the constructor).
- TestSearchBackendSelection covers rg/grep detection and arg
construction, including the --pre flag-injection regression.
- TestSearchOutputBudget exercises Tier 1/2/3 directly.
- TestSearchCaptureStreaming spawns real Python subprocess writers
to exercise the byte-cap trim, mega-line-no-newline edge case, the
watchdog timeout when the child writes nothing, and the stderr
drain under load.
- test_env_scrub picks up the new tool-config keys.
* Address Copilot review on #473
- Budget the Tier 2/3 header up front so the formatter's emission stays
strictly within _SEARCH_OUTPUT_BUDGET. Previously the fit checks only
counted body bytes, letting the final string overflow by ~120 chars
(header + separator) and triggering _truncate_output's head+tail
dropout — exactly the shape this code was trying to avoid.
- Restore the (5, 3, 1) ladder in Tier 2: the analytical K from perf-2
is kept as a starting estimate, but if that K's actual emission
doesn't fit (the estimate ignores the header and overweights shared-
path compression) we step down through the ladder before falling
through to Tier 3. The previous one-shot K could collapse to counts-
only when 3/file or 1/file would have fit.
- Only normalise rc to 0 in the capped-output path when rc < 0 (our
SIGKILL). There's a narrow race where the child can exit naturally
between our read and our kill; preserving a non-negative rc means
rg's rc=2 ('matches found but some files had errors') no longer
silently turns into a clean success when the byte cap also fires.
- Clarify _MAX_SEARCH_LINE_LENGTH doc: the cap applies to the content
portion (after path:lineno:), not the whole emitted line.
- Add explanatory comments on the two intentional `except Exception:
pass` blocks in _search_capture (stderr drain, pipe close in the
cleanup finally) so static analysis and future readers can see the
silence is deliberate.
- Tighten the budget tests: now assert strict `<= _SEARCH_OUTPUT_BUDGET`
instead of the +512-char slack that was masking the header overflow.
- New regression tests:
- Tier 2 ladder step-down (K=5 over budget, K=3 fits, no Tier 3 fall-through)
- capped + rc=2 surfaces stderr instead of being normalised to success
- capped + rc<0 (our SIGKILL) flows through as a partial-result success
* chore(search): post-review cleanup
Follow-up to the Copilot-review fixes in 39d2aa2 — these are all small
quality items (no behaviour change, no new tests).
- q-1: collapse the Tier 2 candidates filter to a single expression.
Drops the redundant inner ``max(estimated_k, 1)`` and the unreachable
``if not candidates`` branch (the ladder ends in 1 and ``estimated_k``
is already floored at 1, so the comprehension always yields ≥ ``[1]``).
``or [...]`` is kept as defence against future ladder changes.
- q-2: update _format_search_results docstring to match the new ladder
semantics (analytical seed → step down through (5, 3, 1) from the
highest rung ≤ the estimate). The previous wording suggested every
Tier 2 attempt started at 5.
- q-3: combine the two ``from turnstone.core.session import ...``
statements in test_tier2_steps_down_ladder_before_falling_to_tier3
into a single top-of-function import (matches the surrounding tests).
- q-4: shorten the explanatory comments on the two best-effort cleanup
paths in _search_capture to one line each. Both sites now read with
the same shape ("# best-effort: pipe may be torn down by ...").
- q-5: trim the _MAX_SEARCH_LINE_LENGTH comment from 7 lines back to 3.
Keeps the load-bearing semantic (cap is on the content portion only)
and the pathological-line defence; drops the paths-aren't-bounded
parenthetical, which was background reading rather than WHY.
* feat(providers): api_surface toggle + mistral medium reasoning fix
Mistral medium open-weights served by vLLM expects reasoning_effort via
the Responses API (`reasoning.effort`), not as a `chat_template_kwargs`
entry on Chat Completions. The session was unconditionally injecting
`{"reasoning_effort": ...}` into `chat_template_kwargs` for every
openai-compatible request, which corrupted the prompt rendering for any
backend whose chat template didn't consume that key (Mistral medium,
Mistral cloud, Groq, OpenRouter).
Changes:
- Add `api_surface` ("chat" | "responses") to `ModelConfig.server_compat`
and thread it through `create_provider` / `model_registry.get_provider`.
`openai-compatible` defaults to Chat Completions; operators can flip
individual aliases to Responses for endpoints that support it.
- New `vllm-mistral-medium` profile that pre-fills api_surface=responses
on Detect for known Mistral medium model ids.
- Drop the unconditional `reasoning_effort` injection into
`chat_template_kwargs`. Operators running gpt-oss-style local
templates that consume `reasoning_effort` from the chat template now
opt in via `server_compat.extra_body.chat_template_kwargs`.
- New "API Surface" select in the Models admin tab; allowlist-validated
server-side at create/update time; pre-filled by Detect via the
profile suggestion.
- Evict the cached provider singleton in `ModelRegistry.reload()` when
api_surface changes (previously only cfg.provider triggered eviction).
- Fix `_run_agent` fallback path to inherit the session's primary alias
for capability and server_compat resolution; previously the fallback
passed `alias=None`, which silently dropped per-model caps on the
agent path.
Tests: 5117 passed (-m "not live"); ruff + mypy clean.
* fix(providers): don't auto-suggest Responses for Mistral medium
vLLM's Responses API surface for Mistral medium open-weights doesn't
wire up the Mistral tool-call parser as of vLLM 0.x — tool calls leak
into the response as ``[TOOL_CALLS]<name>{...}`` text instead of
structured tool_calls. Chat Completions on the same engine handles
tools cleanly via ``--tool-call-parser mistral``, and reasoning can be
turned on via the vLLM CLI ``--reasoning-parser`` flag.
Drop the auto-suggest mapping so Detect falls back to the generic
``vllm`` profile. Keep the ``vllm-mistral-medium`` profile definition
in place so an operator who specifically wants per-request effort and
accepts the tool-calling limitation can still pick "Responses API"
manually in the admin UI.
* fix(providers): address Copilot review on PR #469
- providers/__init__.py: drop the redundant *_responses_provider /
*_chat_provider names; have create_provider use _openai_provider and
_openai_compat_provider directly so they're not flagged as unused
globals.
- console/server.py: tighten _validate_api_surface to a strict equality
match against the canonical {"chat", "responses"} set. The previous
strip().lower() membership check accepted ' Responses '/'CHAT' but
stored the raw string verbatim, which then failed to round-trip
through the admin <select>.
- console/static/admin.js: gate the entire server_compat block (server
type, api_surface, extra_body) on provider == "openai-compatible" at
save time so toggling provider away can't leave a stale hidden surface
selection in the persisted capabilities JSON.
- tests/test_session.py: splat the bad kwarg via **dict so CodeQL no
longer flags the call as a wrong-name keyword (the point of the test
is the runtime contract, not the static type).
- tests/test_admin_model_registry_refresh.py: add endpoint-level tests
for the api_surface validation on both create and update — covers the
bogus-value rejection, non-canonical-string rejection, and the happy
path persisting through to the refreshed registry.
* fix(memory): query-aware candidate selection + OR-of-terms search
The system-message memory composition path used a recency-ordered
candidate set (`_list_visible_memories(limit=fetch_limit)`). On
deployments with more than `fetch_limit` (default 50) visible
memories, BM25 only ever ranked the 50 most-recently-touched memories
— a relevant memory written months ago was silently invisible
regardless of how well it matched the recent context. Multi-word
search at the SQL layer used AND-of-terms, killing recall on any
multi-word query without an exact field overlap.
## Functional changes
- `_init_system_messages` (`turnstone/core/session.py`): extract
recent context first, then `_search_visible_memories(context)` to
pull query-aware candidates. Search hits below `fetch_limit` union
with the recency list (deduped by memory_id) so the BM25 candidate
pool is always a SUPERSET of the prior recency-only pool — even on
noisy queries where the cap fills with stopwords, the recency-50
the original bug surfaced still reaches BM25. Empty context skips
search entirely. Candidate-selection logic extracted into
`_select_memory_candidates`.
- `search_structured_memories` (PostgreSQL + SQLite): per-term
clauses join with OR instead of AND. A row matches if ANY term
matches ANY of name/description/content. Downstream BM25 narrows
back down by relevance.
## Perf hardening
- Collapse the 1-3 fanned scope queries into a single SQL. New
backend methods `list_visible_structured_memories` /
`search_visible_structured_memories` union the visibility scopes
into one WHERE OR-group, so a composition rebuild now hits the DB
at most twice (search + recency) instead of up to six times.
- Cap and normalize search terms. Composition can hand a multi-KB
pasted message to ILIKE-based search; without a cap, every distinct
token would emit one unindexable predicate per scope-fanned query.
`normalize_search_terms` (`storage/_utils.py`) de-dupes
case-insensitively, drops <2-char tokens, and hard-caps at 16.
- Per-turn search cache. `_init_system_messages` fires from many
call sites within one turn (state transitions, MCP refresh, tool
results) and the recent-context query is identical across them.
Session-instance cache keyed by (query, mem_type, limit) absorbs
the duplicates; invalidated in `_append_user_turn` and after
memory save/delete tool actions.
- Stable secondary sort by `memory_id`. `updated` is second-precision
and `touch_structured_memories` can land a batch on identical
timestamps; without a tie-breaker SQL returns rows in
implementation-defined order, BM25 input shuffles, and the
LLM-side prompt cache misses across calls. All four backend ORDER
BYs now break ties on `memory_id ASC`.
## Quality cleanups
- Coalesce `memory.search.term_count` + `memory.search.zero_results`
into a single `memory.search` log carrying both `term_count` and
`result_count`.
- New `memory.composition` log: source / candidates / injected.
- Promote a shared `make_chat_session` factory to `tests/_helpers.py`.
- Rename SQL builder local `extra` -> `scope_filters` for clarity.
- Add docstrings on `search_structured_memories` so the AND->OR flip
survives future readers.
## Tests
Adds 20 tests across `tests/test_structured_memory.py`,
`tests/test_structured_memory_storage.py`, and
`tests/test_memory_relevance.py`: recency-ceiling regression,
empty-query fallback, sparse-match union, recency-preserved-when-
search-returns-noise (locks in the pool-superset invariant),
OR-of-terms on both backends, scope filtering preserved,
search-facade multi-word behavior, term-cap normalization, the new
visible-scope helpers (list + search + empty-scopes guard),
coord-scope composition isolation, end-to-end
`memory(action='search')` tool execution, per-turn cache hit +
invalidation, and stable ordering under tied `updated` timestamps.
Memory test sweep: 102/102. Broader regression
(session, storage, coordinator, load_skill): 411/411.
* fix(memory): address Copilot review on PR #468
Three follow-ups from Copilot's inline review:
1. SUPERSET invariant violation (Copilot, session.py:5510).
`(search_hits + extra)[:fetch_limit]` capped the union back down to
fetch_limit, evicting the recency tail when search added distinct
hits. Recency tail is exactly where ancient-but-recently-touched
memories live — the recall this PR is supposed to improve — so
tail eviction recreated the bug for the narrow case where a query
term fell off the 16-cap and the matching memory sat in
recency[40-49]. Drop the cap; both halves are already SQL-capped
at fetch_limit, so the union is at most 2 × fetch_limit (~100 with
defaults). BM25 over 100 candidates in pure Python is sub-ms;
irrelevant recency fillers get score=0 and don't pollute ranking.
Updates the docstring to actually be honest about the invariant.
Adds `test_recency_tail_preserved_when_search_adds_distinct_hits`
that locks the behavior in: 5 search hits + 10 recency = 15-item
pool, every recency item present, source="union".
2. Unbounded `query.split()` in normalize_search_terms (Copilot,
_utils.py:74). `str.split()` allocates the full token list before
the cap-after-16 break, so a 100KB pasted query did MB of throwaway
work even though only 16 tokens entered SQL. Switch to
`re.finditer(r'\S+', query)` — streaming iterator, stops scanning
at the first 16 normalized terms regardless of input size.
3. Misleading + unbounded log term_count (Copilot, session.py:8571).
`len(item["query"].split())` had two problems: same unbounded
split as #2, and the value reported the raw input token count
rather than the normalized term count that actually hit the SQL
WHERE clause — misleading metric for an operator trying to
understand storage-side behavior. Switch to
`len(normalize_search_terms(item["query"]))` — accurate count, and
bounded for free via #2.
Refuted: github-code-quality flagged `...` bodies in the new Protocol
methods as "statement has no effect." False positive — `...` is the
canonical Protocol body convention, used 213 other times in the same
file.
Memory test sweep: 103/103. Broader regression: 411/411.
CI failure on main: test_publish_records_metric_outcome saw an empty
calls list — its monkeypatch was patching a different metrics
instance from the one `_publish_models_metadata` reads.
Two changes:
- test_close_reason_persistence.py: replace the bare
`srv_mod._metrics = MetricsCollector()` assignment in `_make_app`
with an autouse `monkeypatch.setattr(srv_mod, "_metrics", ...)`
fixture so the test's metrics swap auto-restores. Other test
files (test_auth.py, test_server_attachments_endpoints.py) carry
the same anti-pattern; left for a follow-up since they're not on
the critical path here.
- test_server_node_models_metadata.py: switch the publish-helper
metric test to a string-form `monkeypatch.setattr("turnstone.
server._metrics", FakeMetrics())` so it replaces whatever binding
the live module currently holds, regardless of what other tests
did to it. Robust against future leaks of the same shape.
* feat(coord): expose healthy model aliases per node on list_nodes
Surfaces a `model_aliases` field on each `list_nodes` row so a
coordinator can discover which model aliases each cluster node will
accept on `spawn_workstream(model=...)` without an HTTP fan-out.
Each server projects its registry into a `models` entry on
`node_metadata` (`{alias, provider, healthy}` per alias) at lifespan
startup, on every 30s heartbeat tick, and after `internal_model_reload`.
The publish helper short-circuits on a payload-equality cache so a
stable cluster doesn't pay UPSERT churn — exposed via the new
`turnstone_node_models_publish_total{outcome="written|skipped"}`
Prometheus counter so operators can graph cache hit-rate.
Coord client filters the per-alias rows to healthy aliases only and
drops the provider-side model identifier (`cfg.model`) — coords kept
reaching for it when they should pass the local alias.
* fix(coord): address Copilot+CodeQL feedback on list_nodes models work
- internal_model_reload: reuse a single get_storage() local across the
registry load and the metadata publish (Copilot:3047)
- _collect_node_models_metadata: iterate sorted aliases so two
structurally identical registries built in different insertion orders
serialize to the same JSON — directly improves the publish-cache hit
rate exposed via turnstone_node_models_publish_total (Copilot:3105)
- tests: drop mixed turnstone.server import style flagged by CodeQL —
hoist _metrics into the from-import block, and use sys.modules in
the shutdown-race regression test instead of `import as srv`
Address Copilot feedback on PR #465:
1. The has_alias fallback in both session_factories silently rewrote
any unknown caller-supplied alias to the default, including on the
fresh-create path where the create handler maps the factory's
ValueError to a 503 with operator-friendly text. A typo in
body.model would now silently start a workstream on the default
instead of telling the caller their requested model could not be
resolved. Move the fallback out of the factories: each factory
raises again on unknown aliases, and SessionManager filters stale
aliases out of the rehydrate path via a new ``model_validator``
constructor kwarg (production wiring passes ``registry.has_alias``
on both interactive and coordinator).
2. ChatSession.resume()'s elif branch flipped self.model to the
persisted model name even when the alias was unresolvable, leaving
the session paired with the constructor's default provider/client
but a removed model name — a broken state whose next API call
fails. Drop the model copy: keep the constructor's coherent
default (provider + model + capabilities) and just log the
unreachable saved values so the missing alias is auditable.
Tests:
- Move stale-alias coverage from the factory level into
SessionManager (tests/test_session_manager.py): validator drops
stale aliases before reaching build_session; live aliases pass
through unchanged.
- tests/test_sessions.py renamed test_resume_restores_model →
test_resume_keeps_defaults_when_alias_unresolvable to match the new
contract.
SessionManager.open() was calling build_session(ws) without a model
arg on the rehydrate path. The session_factory then resolved the
*current* default alias, ChatSession.__init__'s _save_config() (INSERT
OR REPLACE per-key) clobbered the persisted workstream_config with
those defaults, and the subsequent resume() "restored" what was now
the default — silently resetting model_alias, model, temperature,
reasoning_effort, max_tokens, skill, creative_mode, instructions,
token_budget, and notify_on_complete on every reopen and every
service restart, for both interactive and coordinator workstreams.
Three layers:
1. SessionManager.open() now reads workstream_config via
self._storage.load_workstream_config(ws_id) and threads the saved
model_alias into build_session(ws, model=saved_alias).
2. ChatSession.__init__ now skips its initial _save_config() when a
workstream_config row already exists for self._ws_id — protects
every other persisted knob without having to plumb each one
through the adapter signature, and catches any future construction
path that forgets to thread model through build_session.
3. Both session_factories (server.py interactive, console
session_factory.py coordinator) now treat an unknown caller-
supplied alias the same as an unset alias: fall back to the
runtime default rather than raising. Without this, a workstream
pinned to an alias an operator has since removed from the registry
would 500 on every reopen — defeating the "best effort restore,
default if the original is gone" contract this fix is meant to
deliver. Mirrors _effective_default_alias's existing has_alias
guard against a stale ConfigStore default.
Three changes from PR review:
- Permission gating: hide the Roles sub-tab button when the user
lacks ``admin.settings``. The sub-tab loads/saves through
``/v1/api/admin/settings``, so an admin with ``admin.models`` but
no ``admin.settings`` would otherwise see a perpetual 403 loader.
When Roles is the active sub-tab and the permission check fails,
snap the panel back to Definitions so the user lands somewhere
usable.
- Drop the redundant ``/v1/api/admin/model-definitions`` fetch from
``loadAdminModelRoles``. Both entry points (initial Models-tab
open + ``models_changed`` SSE refresh) flow through
``loadAdminModels`` first, which already populates ``_modelDefs``
+ ``_modelDefaultAlias``; ``_saveModelRole`` doesn't touch model
definitions, so the cached snapshot stays accurate when the save
chains back here. Halves the per-render request count and
removes a wasted round-trip on every cluster-wide model edit.
- Add ``test_models_changed_event.py`` covering the SSE fanout the
prior commit introduced: each model-definition CRUD endpoint
emits exactly one ``models_changed``, settings PUT/DELETE only
emit for keys in ``_MODEL_AFFECTING_SETTING_KEYS`` (parametrised
over all eight), and unrelated settings (e.g.
``session.retention_days``) don't trigger spurious refreshes.
The expected key set is pinned in the test so a stray addition
to the allowlist doesn't silently bypass coverage.
Same shape as the coordinator/judge rows already there: alias dropdown
+ reasoning_effort dropdown sourced from the existing
``model.plan_alias`` / ``model.plan_effort`` and
``model.task_alias`` / ``model.task_effort`` settings. Adds the four
keys to the SSE ``models_changed`` allowlist so changes from the
Settings API also trigger a live dropdown refresh, and filters them
out of the Settings tab so they only render in one place.
Lifts judge and coordinator model assignments out of their respective
admin tabs and into a new Models → Roles sub-tab so role overrides live
next to the model definitions they reference. Forward-looking shape for
the upcoming perception.{audio,image,video} model settings — adding a
new role is one entry in the declarative MODEL_ROLES array.
Also drops the misleading "Coordinator subsystem not configured" home
banner. The session factory already falls back to the registry's
default model when coordinator.model_alias is unset, so the banner was
nagging on fresh installs where the system was actually working. The
related _probeCoordSubsystem / _homeCoordReady plumbing went with it.
Wires SSE-driven live refresh: the console now emits a models_changed
event when a model definition is created/updated/deleted/reloaded, or
when a model-affecting setting (model.default_alias, judge.model,
coordinator.model_alias, coordinator.reasoning_effort) changes.
Connected browsers refetch /v1/api/models on receipt so the home
composer's model dropdown and the Roles sub-tab stay accurate without
a manual reload — fixes the case where editing the underlying model
for an existing alias left the dropdown showing the old model id.
Companion cleanups:
- Renamed .judge-section-* CSS classes to .admin-subtab-* and shared
them with the Models sub-tab switcher (same a11y attrs, arrow-key
nav). Old names had no other callers.
- Filtered judge.model out of the Judge Settings sub-tab and
coordinator.model_alias / coordinator.reasoning_effort out of the
Settings tab — they live exclusively under Models → Roles now.
- Reworded the _require_coord_mgr 503 messages to point operators at
the Models tab instead of suggesting they set coordinator.model_alias.
* fix(console): home composer attachments + coord chat user-message pills
Two parity gaps in the console's coordinator surface:
- The embedded creator on the home page accepted only text — the
paperclip / paste / drop pipeline that the in-coord composer and the
interactive new-ws modal both expose was missing, so a user couldn't
attach files at create time. Stage Files in memory (no ws_id yet) and
ship them multipart on Start; the coord create endpoint already accepts
multipart via create_supports_attachments=True.
- User messages with attachments rendered as plain text on both live
send and history replay — no chip cluster like the interactive pane.
Added appendUserMessageWithAttachments and a structured userAttachments
list built from _attachments_meta (preferred) or the multipart parts
themselves, then rendered the same .msg-user-attach pill strip the
interactive pane uses.
Polish from a designer pass:
- Pill background was --panel-2, equal to the .msg bubble background in
both themes (border contrast ≈1.4:1, below WCAG 1.4.11). Switched to
--panel so the pill sits on a different surface than the bubble.
- Capped chip filename width inside the home composer (max-width 200px +
ellipsis) so a long filename doesn't push the strip past the textarea.
- aria-live="assertive" → "polite" on #home-coord-error; client-side
validation isn't an interrupt-level event.
- Reserved min-height on .home-composer-error and dropped the
display: none/block toggling so validation messages no longer reflow
the active-coordinators list below.
* fix(console): address PR #462 review feedback
- Block home-composer submit when files are staged but the task field is
empty. Server's _coord_create_post_install short-circuits on an empty
initial_message, so the multipart upload would create pending
attachment rows that never reserve onto a turn — orphaned until the
GC sweep. Fail in the browser instead.
- Drop the redundant `part &&` guard in coordinator.js's history-replay
multipart loop; the earlier `if (!part || ...) continue` already
filtered.
- Rewrite the home-mount .composer-chip-name CSS comment. shared/chat.css
defines .composer-chip{,-size,-remove} but no .composer-chip-name rule
— the span inherits the parent chip font with no width cap.
- Add smoke-guard string assertions in test_coordinator_page.py for
appendUserMessageWithAttachments and msg-user-attach so a future
rename can't silently regress the attachment affordance.
* fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration
Loading a saved workstream silently dropped tool results and missed
verdict / output-guard / truncation signals on replay. Root cause was
in `Pane.prototype.replayHistory`: an assistant message carrying both
content and tool_calls cleared the `lastToolBlock` anchor before the
following tool-result iteration could attach. The fix reorders content
to render before the tool block (matching live SSE order) and
restructures the tool-result branch to anchor by `data-call-id` so
multi-tool batches render `[hdr A][out A][hdr B][out B]` rather than
bunching outputs at the bottom.
Beyond the bug, replay now reaches near-parity with the live UX:
- Persisted intent verdicts and output_assessments flow through both
the SSE replay (`_build_history`) and the `/history` REST endpoint
used by coord. Single shared helper module owns the wire shape.
- Memory/recall calls persist instead of being filtered at storage
time — full audit trail; UI dims them by default with hover-reveal
so heavy memory usage doesn't crowd the narrative.
- Truncation indicator surfaces as a sibling pill (consistent across
interactive + coord) when a tool result hit the 2000-char cap.
- `replayHistory` wraps DOM work in `aria-busy` so screen readers
don't get a chatty announce-flood on long replays.
- `_build_history`'s storage I/O moves off the event loop via a new
`events_replay_prepare` async hook for the SSE path; other async
callers wrap in `asyncio.to_thread`.
Coord parity:
- `/history` REST endpoint decorates tool_calls with verdict +
output_assessment + truncation flag (was previously raw
`load_messages` output).
- Coord JS stamps `judge_verdict` / `heuristic_verdict` from
history-loaded `tc.verdict` so the existing batch render paints
the persisted pill, seeds the verdict cache to dedupe later live
SSE events, and emits an inline `.coord-tool-row-warning` chip
per call instead of a generic chat line.
- Memory/recall dim rule mirrored on `.coord-tool-row[data-tool-name=...]`.
* fix(replay): address PR #461 review feedback + raise tool-result storage cap
Copilot review feedback:
- Sibling-chain dim rule (memory/recall) now adds :focus-within
alongside :hover for .tool-output / .media-embed / .output-warning
/ .tool-output-truncated — keyboard users tabbing into a faded
subtree now get full opacity.
- ``cfg.open_post_load`` is now invoked via ``await asyncio.to_thread``
so its sync ``_build_history`` call (storage I/O for verdict
indexes + message reconstruction) doesn't block the event loop on
every workstream open. Mirrors the SSE replay path that's already
protected via ``events_replay_prepare``.
- Replaced the hardcoded ``2000`` literal in server.py and session.py
with ``TOOL_RESULT_STORAGE_CAP`` from the shared decoration module
so the UI truncation-pill detection can't silently desync from the
storage write side.
While here:
- Raised ``TOOL_RESULT_STORAGE_CAP`` from 2000 → 10000. A 2000-char
clip routinely cut grep / file-read bodies mid-line, leaving the
audit trail useless for retrospective debugging. FTS5 + row size
grow proportionally; the per-tool upper bound is still bounded
upstream by ``_truncate_output``'s context-budget clamp.
- Updated the user-visible truncation-pill tooltip on both
interactive and coord to reflect the new cap.
- ``test_decorates_tool_calls_and_marks_truncated`` now references
the constant instead of a literal so it stays correct on future
cap changes.
Speculative reliability machinery from the Stage 3 push that turned
out not to address any user-visible bug. The actual fixes (state /
activity disjunction in handleChildState, bulk-fetch race fix in
_fetch_live_block, push approve_request via cluster bus) are what
resolved the wedged-row issues. Manual testing showed the per-tab
SSE listener queue depth never climbed past single digits even when
rows were stuck — overflow was never the cause.
Removed
- ``_CRITICAL_EVENT_TYPES`` + ``_put_with_priority`` helper.
- Per-tab listener queue selective drop (back to plain
``contextlib.suppress(queue.Full)`` everywhere).
- ``ClusterCollector._fanout`` reverts to the same.
- WebUI ``_broadcast_intent_verdict`` / ``_broadcast_approval_resolved``
/ ``_broadcast_approve_request`` revert to plain ``put_nowait``.
- ``_queue_stats`` periodic SSE emit + frontend status-bar indicator
+ the supporting CSS rules.
- Broken ``.approval-block`` ``transition: max-height`` /
``max-height: 80vh`` / ``overflow: hidden`` rules — the transition
never fired (nothing toggled max-height) and ``overflow: hidden``
clipped long verdict reasoning. Layout-shift on auto-expand jumps
again, which is preferable to clipped content (Copilot review).
Tidied
- ``_CollectorProtocol`` / ``_ManagerProtocol`` method bodies switch
from ``...`` ellipsis to docstring-only bodies, silencing four
CodeQL "statement has no effect" warnings without changing the
Protocol contract.
5024 passed, ruff + mypy clean.
Lift the Children primitive out of CoordinatorAdapter into universal
SessionManager core primitives, replace the fragile poll + state-event
piggyback paths with first-class cluster bus event types for inline
approval delivery, and clean up the resulting frontend reducer.
Architecture
- New `turnstone/core/children_registry.py` — universal parent → children
+ reverse-lookup primitive with atomic `add_child` (returns parent UI
for race-free dispatch). Lifted from `CoordinatorAdapter`.
- New `turnstone/core/child_source.py` — `ChildSource` Protocol with
`SameNodeChildSource` (in-process via SessionManager state observer)
and `ClusterChildSource` (cross-node via ClusterCollector listener).
- `SessionManager._on_state_change` upgraded to multi-subscriber
(`subscribe_to_state` / `unsubscribe_from_state`) under a dedicated
lock; CLI consumer migrated.
- `CoordinatorAdapter` shrunk: 731 → ~640 LOC. Children data lives in
the registry; fan-out lives in ClusterChildSource. Backward-compat
property facades dropped; tests updated to use the registry surface.
Cluster bus event vocabulary
- New event types `intent_verdict`, `approval_resolved`,
`approve_request` flow through both `ClusterCollector._apply_delta`
(translation from node SSE) and `emit_console_ws_*` (synthesis on
console pseudo-node).
- `CoordinatorAdapter._dispatch_child_event` re-emits as
`child_ws_intent_verdict` / `child_ws_approval_resolved` /
`child_ws_approve_request` on the parent coord's SSE stream.
- New `_broadcast_intent_verdict` / `_broadcast_approval_resolved` /
`_broadcast_approve_request` no-op hooks on `SessionUIBase`. WebUI
pushes to the global queue; ConsoleCoordinatorUI pushes to the
collector. `approve_tools` calls `_broadcast_approve_request` right
after setting `_pending_approval` so the items reach the coord tree
immediately, eliminating the bulk-fetch race.
Cleanups
- `pending_approval_detail` piggyback on `ws_state` / `cluster_state`
removed end-to-end. Bulk fetch + explicit verdict / approve-request
push are the canonical carriers.
- Browser `_judgePollTick` 90-second poll loop deleted; push path is
authoritative.
- `urgent` flag on `scheduleLiveFetch` deleted (only caller was 409
retry; replaced with `invalidateLiveBadge` + standard schedule).
- Console `_fetch_live_block` derives `pending_approval` from a
disjunction (`activity_state="approval"` OR `state="attention"`
OR detail present) so the bulk fetch can't return false during the
state-transition race window.
- Coord-side merge guard in `flushLiveFetches` no longer clobbered:
`handleChildState` only stamps `sseUpdatedAt` when authoritatively
clearing detail.
- `child_locality` capability flag removed (was inert dead code).
Reliability
- Selective drop on listener queue overflow: critical event types
(verdicts, approvals, ws_closed, child_ws_*) evict one oldest item
to make room rather than dropping themselves on a full queue.
Best-effort events (state ticks, content tokens, status, activity)
drop as before. Applied to `SessionUIBase._enqueue`,
`ClusterCollector._fanout`, and the `WebUI._global_queue` puts in
the new broadcast hooks.
- `_state_subscribers` snapshot under a dedicated lock so concurrent
subscribe / unsubscribe during dispatch can't shift the iterator.
UX / a11y
- Loading placeholder in renderChildRow keeps row height stable while
the bulk fetch is in-flight (sr-friendly aria-label).
- Focus preservation across `_renderChildrenNow` (capture +
restore by row + marker) and across targeted `_updateChildRow` swaps.
- Layout-shift transition on the approval block max-height; respects
`prefers-reduced-motion`.
- Sidebar pending count: `(N children · M pending)`.
- Risk pill `aria-label` spells out level + confidence for SR users.
- Per-coord SSE listener queue depth surfaced in the status bar
(`queue N/500`) with color escalation (warn at >50%, danger at >80%).
Tests
- 305+ test changes across 8 files. New unit tests for
`ChildrenRegistry`, `ChildSource` (both impls + multi-subscriber
observer), the new collector emit + apply_delta cases, the dispatch
cases for new event types, the broadcast hook overrides on both
WebUI and ConsoleCoordinatorUI, and the focus / placeholder /
pending-count frontend assertions in `test_coordinator_page.py`.
5024 passed, ruff + mypy clean.
* feat(console): multi-select delete UX for Saved Coordinators
Mirror the per-server "Saved Workstreams" multi-select delete onto the
console's "Saved Coordinators" section. Coordinator deletes go through
the existing routing proxy at POST /v1/api/route/workstreams/delete
(body-keyed by ws_id, since coordinators live on the node that owns
them) — no backend change required.
Pagination caps the visible page (and therefore the Select-All fan-out)
at 24. Without it, a Select-All on a busy cluster would pin the
console proxy pool with hundreds of parallel deletes through the
fan-out router. While in delete mode the saved-coordinators list is
frozen against SSE re-renders so visible cards don't shuffle out from
under the user's selections (drained on cancel / post-delete close).
Refactor: shared logic now lives in turnstone/shared_static/cards.{css,js}.
* .ws-delete-* CSS moved out of ui/static/style.css into the shared
sheet alongside .dashboard-card; the existing ui/static modal
markup picks up class hooks instead of id-scoped rules.
* createSavedCardsController() owns mode state, checkbox decoration,
toolbar wiring, focus trap, modal lifecycle, and batch fan-out.
Both ui/static (Saved Workstreams) and console/static (Saved
Coordinators) instantiate one controller; ui/static is now ~300
LOC lighter as a result.
* Internalises stale-selection prune across SSE re-renders, the
wsId->item lookup map (was O(selected x N)), and the aria-hidden
wrap on the toggle button's emoji glyph.
Designer review tightened the affordance:
* Modal close restores focus to the toggle button (was landing on
<body>) — WCAG 2.4.3.
* Modal [role="alert"] gets a red-chip treatment when populated,
stays invisible at rest via :not(:empty).
* Pagination consolidated onto the existing .pagination control
(terse "X / Y" label + arrow-glyph buttons) instead of a parallel
.coord-pagination treatment.
* Filled destructive buttons darkened to #dc2626 in dark theme so
the white label clears WCAG AA contrast (was 3.0:1 on --red).
Light theme keeps --red unchanged (5.9:1 already passes).
* Toolbar wraps below 700px viewport — Delete Selected drops to its
own full-width row underneath count + Cancel + Select All for
thumb-target separation.
* .ws-card-check:focus-visible outline + word-break on
.ws-delete-item for narrow-modal long aliases.
* fix(cards): address Copilot review feedback on PR #458
* closeModal focus restore now falls back to the section toggle button
(opts.buttonId) when prevFocus is hidden or detached. The post-delete
Close path runs cancel() before closeModal(), which puts the bar at
display:none — so the captured prevFocus (the bar's "Delete Selected"
button) is no longer focusable and focus would land on <body>,
defeating the WCAG 2.4.3 fix. Esc / Cancel paths still land on the
original focus owner because the bar stays visible in those flows.
* Saved Coordinators onClose drains _savedCoordsRetry before reloading.
Without it, SSE events that arrived during the delete-mode freeze
leave the retry flag true, so loadSavedCoordinators's .finally()
re-fires a second fetch immediately after the first resolves. Mirrors
the same idiom in cancelCoordDeleteMode.
Three issues from the Copilot review on PR #457:
1. SQLite race in bulk_close_stale_orphans (Copilot): the SELECT-then-
UPDATE flow doesn't re-apply the eligibility predicates on the
UPDATE, so a row that gets touch_workstream-bumped (or set_state-
transitioned) between the two statements would still be flipped
to closed. Postgres dodges this via UPDATE...RETURNING (one atomic
statement); SQLite needs the explicit re-application. Fix: rebuild
the WHERE conditions list once, apply on both SELECT and UPDATE,
then SELECT-back by ``state='closed' AND updated=now`` to get the
accurate closed-id list. A row that became fresh between the two
statements skips the UPDATE entirely.
2. SQLite IN-clause bind-parameter limit (Copilot): default 999 cap
could be exceeded on a backlog reap (e.g. after a long outage).
Chunked the candidate id list at 500 — same chunk size
prune_workstreams (line 453) uses for the same reason.
3. Wall-clock-dependent test asserts (Copilot, two locations): the
tests asserted ``updated > '2024-01-01T00:00:00'`` which is fragile
on systems with skewed clocks or pre-2024 dates. Replaced with
``updated != stale_seed`` — captures the same intent (the value
was bumped) without depending on wall-clock date.
Two ``...``-as-no-op flags from github-code-quality were false
positives — ``...`` is the standard Python idiom for Protocol method
bodies and matches every other method in _protocol.py. No code change.
Replaces the ``node_id == self_node_id`` orphan-scoping heuristic from
earlier on this branch with liveness-based scoping using
``services.last_heartbeat``. The heuristic was wrong for the post-#384
world: PR #384 (refactor: replace hash-ring rebalancer with rendezvous
hashing) deleted the rebalancer that used to keep workstreams.node_id
pointing at a live node. Without it, ``workstreams.node_id`` is now
stamped at create time and never updated, so in containerized
deployments with dynamic hostnames a dead pod's rows have ``node_id``
matching no surviving service — they'd accumulate forever under the old
heuristic.
services.last_heartbeat is the same primitive the rendezvous router
uses for routing. Reusing it here keeps reap scoping aligned with
routing: dead pods' rows fall out of the live set after the heartbeat
window and become reapable; alive pods' rows stay protected as long as
they heartbeat.
Mechanics:
- ``bulk_close_stale_orphans`` parameter renamed
``node_id: str | None`` → ``live_node_ids: list[str] | None``. The
WHERE clause becomes ``(node_id IS NULL OR node_id NOT IN
live_node_ids)``. ``None`` skips the filter entirely (single-process
/ tests / operator backfill). ``[]`` treats every row as
unprotected.
- ``SessionManager.close_idle`` pass 2 calls
``storage.list_services(self._service_type)`` to enumerate live
peers, passes their service_ids as ``live_node_ids``. ``_service_type``
is derived from ``self.kind`` (INTERACTIVE→"server",
COORDINATOR→"console") via a module-level mapping — no constructor
param, so production wiring can't miswire the kind/service_type
pairing.
- list_services failure → pass 2 is skipped this tick (conservative;
never reap when liveness state is unknown). Pass 1 still runs.
- ``workstreams.node_id`` with NULL value is always eligible — defends
against ANSI ``NULL NOT IN (...)`` evaluating to NULL (not TRUE) and
silently protecting orphans forever.
- Migration 048 simplified to ``(kind, updated)``; the new query's
``NOT IN (small list)`` predicate against an unbounded-cardinality
column doesn't index well, so leading ``node_id`` would just add
write cost.
Tests cover the live-services protection (own/dead/null cases), the
empty-peers reap-all case, the list_services-failure conservative
fallback, both kind/service_type pairings (interactive→"server",
coordinator→"console"), and the combined live_node_ids +
exclude_ws_ids filter matrix.
bulk_close_stale_orphans runs every min(300s, idle_timeout/4) on
every server and console process. Its WHERE shape is:
WHERE kind = ?
AND state IN ('idle','thinking','attention','running')
AND updated < ?
AND node_id = ? -- multi-node interactive only
At current scale the existing single-column indexes are sufficient —
idx_workstreams_state prunes to non-closed and the planner filters the
rest sequentially. At 100k+ rows that filter becomes a tablescan-
shaped cost.
A partial index covering only BULK_CLOSE_STATE_VALUES rows matches the
reaper's query exactly while staying tiny — closed rows (typically
95%+ of the table) and error rows are excluded, so the index is
roughly 5% the size a full multi-column index would be. Write
amplification only kicks in for transitions touching one of the four
covered states.
Column order (node_id, kind, updated): node_id is the most selective
filter for multi-node interactive (each server prunes to its own
node's rows), kind second so coord-only and interactive-only queries
within a node still get index-only scans, updated last so the range
comparison rides the trailing column.
Postgres uses CREATE INDEX CONCURRENTLY so the build is non-blocking
on a live system; SQLite has no concurrent concept and the table-
level write lock already serializes, so a plain CREATE INDEX is fine.
The console's coord SessionManager had no idle thread — close_idle was
never called for coordinator workstreams. This is the worse half of
the lifecycle leak: the dashboard filters via the in-memory pool, so
DB-only orphan coords were invisible. At empirical diagnosis,
coord closure was 16% (10 closed / 64 total) vs interactive 63%.
Adds _coord_idle_cleanup_thread mirroring turnstone/server.py's
_idle_cleanup_thread but skipping the rate-limiter / global-queue arms
the console doesn't have. Started from the lifespan when coord_mgr is
constructed and server.workstream_idle_timeout > 0 (reuses the
existing setting — same cadence works for both kinds).
Initial sweep runs INSIDE the thread before the first sleep, not
synchronously in the lifespan: cold-start orphans are reaped without
blocking Starlette boot. Important because cold start with many DB
orphans (the precise condition this code targets) is exactly when the
UPDATE is most likely to be slow.
Helper takes an optional stop_event parameter purely for tests —
production callers pass None and the daemon runs for process lifetime.
This avoids the SystemExit-from-stub + module-wide filterwarnings
fragility a previous iteration relied on.
Four tests: initial sweep runs before first sleep, ticks fire each
loop, exceptions don't kill the thread, stop_event exits cleanly.
Real bug: workstream rows accumulate in non-closed states (idle,
thinking, attention, running) when their owning process restarts or
crashes. Empirical diagnosis on a live deployment found ~60 stuck
coord rows in DB invisible to the in-memory-keyed dashboard, plus
100+ interactive rows older than the 2h timeout (one stuck "thinking"
for 2 weeks — impossible across a process restart).
Root cause: close_idle iterates self._workstreams.values() — only the
loaded subset. Anything left behind by a prior process incarnation
sits in DB forever because nothing ever re-loads it.
This commit gives close_idle a second pass.
Pass 1 (existing, unchanged): close loaded IDLE rows whose
ws.last_active (monotonic) is past timeout. IDLE-only so legitimately-
attentive rows (waiting for user response) stay live.
Pass 2 (new): bulk-close DB rows of this manager's kind whose updated
is past the wall-clock cutoff and which aren't currently loaded.
Closes the broader BULK_CLOSE_STATE_VALUES set — any matching row is
by definition not loaded by any process and cannot be in a live
interaction. Scoped by self._node_id so a sibling node can't reap
rows we own (multi-node interactive correctness). No emit_closed —
never-loaded rows have no SSE listeners expecting them.
Lock invariant: pass 1 holds self._lock briefly to snapshot victims
and pop them (existing behavior). Pass 2 holds self._lock briefly to
snapshot the loaded keys, then releases before the DB UPDATE so a slow
reaper query can't block create/get/set_state.
Also fixes a same-process race in open(): the rehydrate path read DB,
released the manager lock, then re-acquired to install — a concurrent
pass 2 between the two acquisitions snapshots loaded keys without the
in-flight ws_id, and could clobber its DB row to closed. open() now
calls touch_workstream(ws_id) on rehydrate so the row's updated is
fresh against any pass-2 cutoff. Pure timestamp write is safe against
concurrent close() (close still wins on the state column).
Three new tests cover the DB orphan pass (basic, exclude-loaded, kind
filter) plus node_id scoping (own/foreign rows, None-skips-filter) and
the open() rehydrate touch.
Two new methods on the StorageBackend Protocol, with implementations on
both Postgres (UPDATE ... RETURNING) and SQLite (SELECT-then-UPDATE in
one transaction). No callers yet — wiring lands in subsequent commits.
bulk_close_stale_orphans(kind, cutoff, exclude_ws_ids, node_id=None)
flips rows in BULK_CLOSE_STATE_VALUES (idle/thinking/attention/running)
to closed when their updated timestamp is lex-older than cutoff. The
node_id filter scopes the reap to a single node's partition — required
for multi-node interactive deployments where each node only has
authority over its own workstreams.node_id rows. Excludes loaded ids
so the in-memory pass owns those.
touch_workstream(ws_id) bumps updated without changing state. Used by
the open() rehydrate path to defend against the orphan reaper clobbering
a freshly-loaded row whose DB updated is older than the cutoff. Pure
timestamp write is safe against concurrent close() because close still
wins on the state column.
BULK_CLOSE_STATE_VALUES is centralized in workstream.py so the two
backend implementations and FakeStorage all agree; if a new transient
state is added to WorkstreamState, deciding whether it joins this set
is part of the change rather than an after-the-fact audit across three
files.
Storage tests (run against both backends via the conftest fixture) cover
the kind/state/cutoff/exclude/node_id matrix plus touch_workstream.
The themed ``tool_reminder`` bubble below the tool block already
shows the metacog text, and the tool block immediately above it
carries the tool name — so a separate gray ``[repeat: list_workstreams()
called with same arguments]`` info line was just duplicate visual
noise (operator-visible in the screenshot below the bubble).
Drop the ``ui.on_info`` call inside ``_apply_post_execute_advisories``
that emitted the diagnostic line. Update the docstring to reflect
that the bubble is the canonical signal. Rename
``test_emit_repeat_ui_line_on_streak_fire`` →
``test_no_legacy_repeat_info_line_on_streak_fire`` and invert the
assertion.
CI typecheck failed because ``WorkstreamTerminalUI(TerminalUI)``
inherits from ``SessionUI`` (the Protocol), and the Protocol's
``on_user_reminder`` / ``on_tool_reminder`` declarations have empty
bodies — mypy treats those as implicitly abstract, so the subclass
became un-instantiable.
Add real implementations on ``TerminalUI`` that render reminders as
``[metacognition · type] text`` lines in yellow. This also restores
the metacog signal on the CLI surface (the legacy
``[metacognition: nudge injected — …]`` info-line went away with
``_emit_nudge_ping``; without this commit the CLI showed no signal
at all for metacog nudges). Tool-channel and user-channel render
identically because terminal output is anchored by stdout flow
rather than by DOM anchor — the line lands directly after the
message it advises.
Address Copilot's review feedback on PR #456 — the docstrings and
inline comments hadn't all caught up with the architectural shift
across the branch:
- ``_apply_reminders_for_provider`` docstring: "every user message"
→ role-agnostic, since tool messages also carry ``_reminders``
(tool_error / repeat).
- ``_mark_reminders_delivered`` docstring: same role-agnostic
update; explicitly note both channels.
- ``_append_user_turn`` callsite comment near
``_attach_pending_user_reminders``: still described splicing
``<system-reminder>`` blocks into user content; updated to
reflect the side-channel attach + transient-copy splice at the
provider boundary.
- ``_build_history`` block comment: was user-message-only; now
mentions tool messages and both ``user_reminder`` /
``tool_reminder`` SSE events.
- ``_build_history`` propagation comment: same role-agnostic note
on the per-entry surface.
- ``app.js`` ``user_reminder`` SSE handler comment: said the
bubble renders "above" the user message, but
``insertAdjacentElement('afterend', el)`` drops it BELOW.
- ``app.js`` ``replayHistory`` comment: said "insertBefore drops
the reminder directly above the just-rendered user bubble";
same fix — bubble lands BELOW.
No behaviour change.
The repeat-detection block in ``_apply_post_execute_advisories`` had
a leftover "clear streak when a write tool succeeded" branch from
when ``RepeatDetector`` tracked cumulative counts. With the
consecutive-streak semantics introduced earlier in the branch the
branch became:
1. Redundant — any different (name, args) signature already resets
the streak via ``RepeatDetector.record``, so an intervening
read/write naturally breaks the streak.
2. Actively wrong — the clear runs ONCE at the top of each
``_apply_post_execute_advisories`` call, before the per-result
loop records sigs. In a single parallel batch
``[bash, bash, bash]`` the clear runs once and then three
``record`` calls accumulate to count=3 in the same call → fires.
But across three sequential turns, each turn calls
``_apply_post_execute_advisories`` fresh, the clear runs at the
top of each call, and only one ``record`` per call follows — so
the count never gets above 1 and the canonical
"small local model stuck on ``bash('echo test')``" pattern
never triggered the nudge.
The asymmetry only existed for successful calls — failures don't
satisfy the ``not _tool_error_flags.get(tc["id"])`` predicate, so
the clear didn't fire and sequential failures already worked. The
fix is to drop the clear entirely; ``RepeatDetector``'s
consecutive-streak semantics handle every case uniformly.
Tests:
- ``test_successful_write_clears_streak`` →
``test_intervening_different_call_resets_streak`` —
rewords the assertion to reflect the actual mechanism (any
different sig resets, write-or-otherwise) since "writes clear"
was the bug, not the contract.
- ``test_failed_write_does_not_clear_streak`` →
``test_sequential_bash_failures_fire_repeat`` — same shape, just
framing fixed.
- New ``test_sequential_bash_same_command_fires_repeat`` —
regression for the bug user hit (three sequential successful
``bash('echo test')`` calls now correctly fire the nudge).
The yellow themed reminder card introduced for user-channel nudges
(correction / denial / resume / start / completion) now also fronts
tool-channel nudges (tool_error / repeat). Pre-fix the tool channel
shipped its reminders inside the tool-result envelope via
``wrap_tool_result``, leaking the ``<system-reminder>`` block into
``self.messages`` content (same problem the user channel had before
the side-channel refactor) and surfacing the legacy gray
``[metacognition: nudge injected — …]`` info line as the only
operator-visible signal — duplicated alongside the new themed bubble
for user-channel nudges.
Tool-channel parity:
- ``_collect_advisories`` now returns
``(persistent_advisories, metacog_reminders)``. Persistent
advisories (``GuardAdvisory`` / ``UserInterjection``) keep
riding ``wrap_tool_result`` because they ARE conversation
history. Metacognitive reminders extract to the second tuple
element; the caller attaches them to the tool message dict's
``_reminders`` side-channel and emits ``on_tool_reminder``.
- ``_apply_reminders_for_provider`` already handles ``_reminders``
on any role, so the tool-channel splice into wire content is
free. ``_build_history`` also already propagates
``entry["reminders"]`` regardless of role, so reload renders the
bubble too.
- ``SessionUI`` Protocol gains ``on_tool_reminder(reminders,
tool_call_id)``; ``SessionUIBase`` enqueues a ``tool_reminder``
SSE event with the ``tool_call_id`` anchor.
- ``_emit_nudge_ping`` had no remaining callers and was removed —
the themed bubble (live SSE + ``/history`` reload) is the
canonical operator signal for both channels now.
UI polish (the four fixes the screenshot caught for the user
channel + their tool-channel mirror):
- Bubble renders BELOW the message it advises (semantically: a
hint to the model right before its turn). ``addUserReminder``
swaps ``insertBefore`` for ``insertAdjacentElement('afterend',
el)``; ``addToolReminder`` anchors below the ``.ts-approval``
block whose tool result triggered the batch's reminder.
- Label uses the full feature name ``metacognition`` (was the
``metacog`` shorthand).
- Card width / alignment inherits from the base ``.msg`` rule —
``align-self: flex-end`` and the explicit ``max-width`` are
gone, so the card matches the user / assistant column instead
of pinning right-aligned narrow.
- The legacy ``[metacognition: nudge injected — …]`` gray info
line is gone for both channels.
Frontend additions:
- ``Pane.prototype.addToolReminder(reminders, toolCallId)``
anchors below the ``.ts-approval`` block (live: by
``data-call-id``; replay: by "last block in messagesEl"
fallback, which is correct because messages render in order).
- SSE switch case ``"tool_reminder"`` calls ``addToolReminder``.
- ``replayHistory``'s tool-message branch now calls
``addToolReminder`` when ``msg.reminders`` is present.
- ``addUserReminder`` advances its anchor on each loop iteration
so multiple reminders stack in queued order rather than
reversed.
Coord console parity:
- ``coordinator.js`` gains ``appendReminderBubble`` /
``appendUserReminderLive`` / ``appendToolReminderLive`` mirroring
the interactive UI. The tool-channel anchor walks
``toolRows[callId].batch`` to attach below the
``.coord-tool-batch`` construct (one bubble per dispatch turn,
matching the "one nudge per batch even with many failing tools"
drain).
- SSE switch handles ``user_reminder`` and ``tool_reminder`` on
the coord conversation surface.
- ``/history`` replay propagates ``msg.reminders`` for user and
tool messages — same wire shape as the interactive pane.
- ``.msg.user-reminder`` styles moved to
``shared_static/chat.css`` so both surfaces inherit the same
yellow themed bubble from the shared base.
Defensive read on ``_apply_reminders_for_provider`` (per Copilot
review on the closed PR): a malformed ``_reminders`` entry (string,
None, etc. — corruption / partial state) used to abort ``send`` via
AttributeError on the ``.get("text", "")`` call. Filter to dicts
before building the block, mirroring the same filter
``_build_history`` already applies on the wire-out side; an
all-malformed list passes through as no-reminders.
Tests:
- ``test_collect_advisories_drains_tool_buffer_on_last_result``
rewritten to assert the ``(persistent, metacog)`` tuple shape
and that ``MetacognitiveAdvisory`` no longer appears in the
persistent list.
- ``test_collect_advisories_holds_*`` and ``_drops_*`` updated for
tuple return.
- ``test_attach_emits_visibility_ping`` /
``test_collect_advisories_emits_visibility_ping`` inverted to
assert the legacy gray line is gone on both channels.
- ``TestSessionUIBaseToolReminderHook`` covers the new SSE event
shape with the ``tool_call_id`` anchor.
- ``test_malformed_reminders_filtered_out`` and
``test_all_malformed_reminders_passes_through`` cover the
Copilot-flagged defensive filter.
User-channel metacognitive nudges (correction, denial, resume, start,
completion) used to be spliced into ``user_msg["content"]`` permanently,
which leaked the ``<system-reminder>`` envelope into every consumer of
``self.messages`` — UI replay (mitigated by a regex strip in /history),
compaction, title generation, and any future channel adapter that
echoes conversation context. The /history strip was a band-aid;
compaction and title-gen still saw the raw spliced text.
Switch to a side-channel: ``_attach_pending_user_reminders`` writes the
rendered reminder list to ``user_msg["_reminders"]`` (sibling key,
leading-underscore convention shared with ``_attachments_meta`` /
``_provider_content``). At the provider boundary, a new
``_apply_reminders_for_provider`` builds a transient shallow-copy with
the reminder spliced into ``content``; the original message dict
stays clean. ``sanitize_messages`` drops the sibling key on the wire.
Once-per-session-not-per-turn semantics for the wire: after stream
success the loop calls ``_mark_reminders_delivered``, which flips a
``_reminders_delivered`` flag on every user message that carried
reminders into that call. ``_apply_reminders_for_provider`` skips
already-delivered messages so the model sees each reminder exactly
once (the turn it advised). ``_build_history`` ignores the delivered
flag entirely, so reconnecting tabs render the same nudge bubble the
originating tab saw via the live ``user_reminder`` SSE event.
UI surface:
- ``SessionUIBase.on_user_reminder`` enqueues a
``{type: "user_reminder", reminders: [...]}`` SSE event with the
same shape ``_build_history`` surfaces.
- ``app.js`` renders a ``.msg.user-reminder`` bubble (yellow accent,
pill-styled) anchored above the user message it advises, both
live and on history replay.
- ``replayHistory`` renders ``addUserMessage`` before
``addUserReminder`` so the anchor lookup finds the just-rendered
turn (not a prior one).
- Multi-tab caveat documented inline: non-originating tabs receive
no ``user_message`` SSE event today, so a reminder may anchor to
a stale prior bubble until ``/history`` reload corrects it.
Pre-existing bug surfaced by the audit: cancel handlers
(``GenerationCancelled`` / ``KeyboardInterrupt`` / generic
``Exception``) in ``ChatSession.send`` cleared
``_pending_tool_advisories`` but not the user-channel buffer. Both
now drain through a shared ``_drain_pending_advisories`` helper.
Removed the ``/history`` regex strip — the side-channel approach
makes it redundant. Hoisted ``escape_wrapper_tags`` +
``render_system_reminder`` imports to module top (called 2-3× per
turn).
Tests:
- ``TestApplyRemindersForProvider`` — pass-through-by-reference,
string + list content splice, escape on user-typed wrapper tags,
multi-reminder ordering, source-untouched invariant, delivered
flag skip path, fallback for unexpected content shape.
- ``TestMarkRemindersDelivered`` — flag idempotency, no-reminders
no-flag, only marks user messages with reminders.
- ``TestUpdateTokenTableMsgsParam`` — calibration uses pre-built
msgs when provided, falls back when not.
- ``TestUserAdvisoryCancelClear`` — all three cancel branches drain
the user buffer.
- ``TestReminderSidechannelIsolation`` — compaction's
``_format_messages_for_summary`` and the title-gen extraction
loop cannot see reminders by construction.
- ``TestSessionUIBaseUserReminderHook`` — ``on_user_reminder``
enqueues the right SSE shape.
- ``TestBuildHistoryReminderPropagation`` — ``entry["reminders"]``
propagation, absent / empty / multi / coexist-with-attachments
cases, malformed input filtering, all-malformed elision.
- ``test_sanitize_messages_strips_underscore_sibling_keys`` covers
``_reminders`` and ``_reminders_delivered``.
Cleanup pass on the metacognitive nudge stack — restores pre-split
errored-counts-toward-repeat behaviour and tightens the is_error
plumbing through the per-batch advisory hook.
The per-batch hook in ``_run_loop`` was duplicating the is_error
signal: ``self._tool_error_flags`` (set by ``_report_tool_result``)
and a string-prefix tuple (``Error`` / ``JSON parse error`` / …).
Two truth sources is what got us here — bash commands that exit
non-zero with normal stdout matched the flag but not the prefix,
the deny path matched the prefix but not the flag, and the result
was that stuck-loop detection silently broke for the most common
failure mode (the model bashing the same broken command).
Single source of truth now:
- ``_execute_tools.run_one`` deny branch routes through
``_report_tool_result(is_error=True)`` so denied calls populate
``_tool_error_flags`` like every other error path.
- The error-prefix tuple is gone; the write-success-clear gate and
the tool-error-nudge gate both read ``_tool_error_flags`` only.
Repeat-detection state moves from a ``set[str]`` (fired on the second
identical call, ignored errors entirely) to a ``RepeatDetector``
helper in ``metacognition.py`` with consecutive-streak semantics:
- Threshold raised from 2 to 3 — two-in-a-row was noisy on
legitimate transient retries; three is the cheapest stuck-loop
signal.
- Recording a different signature resets the count, so [A, A, B, A]
is two short streaks of 2 and not a streak of 4. Bounded by O(1)
state regardless of session length.
- Errored calls now count toward the streak (the split into a
separate metacog module unintentionally introduced a "skip errors"
branch — restored).
While there:
- ``metacognition._COOLDOWN_SECS`` default aligned to 300s (matches
``MemoryConfig.nudge_cooldown`` and the ``memory.nudge_cooldown``
config-store default; was set to 30 by an earlier investigation).
- The per-batch advisory block (~80 lines of mixed orchestration
inside ``_run_loop``) is extracted to
``ChatSession._apply_post_execute_advisories`` so the wired
behaviour is testable without driving ``_run_loop`` end-to-end.
Producer extraction to a dedicated module is deferred to a
follow-up; advisory producers all live on ``ChatSession`` for
now per existing convention.
- Frontend ``appendToolOutput`` (turnstone/ui/static/app.js) now
skips rendering when the parent approval block is denied or
the output starts with ``Denied by user`` / ``Blocked``,
mirroring the history-replay guard at ``_build_history``.
Previously the live SSE path didn't need this guard because
the deny path never emitted a ``tool_result`` event; the
is_error routing change above means it does now, so without
this guard the badge from ``resolveApproval`` and the SSE
output would both render.
Tests: 8 unit tests for ``RepeatDetector`` covering streak,
threshold, clear, and intervening-sig reset; 9 integration tests
for ``_apply_post_execute_advisories`` covering the wired
behaviour (3-identical fires warning + advisory + UI line, errored
calls count toward streak as a regression guard, intervening sig
resets streak, successful write clears, failed write does not,
JSON outputs tracked but not inline-warned, tool_error nudge gates
on memory_count, repeat UI line emitted on streak fire).
The pre-existing comment said pending_approval_detail "rides on
every ws_state event" — that overstated the case. The node-side
emit is gated on ``_pending_approval is not None`` so the field is
absent on the steady-state broadcast and possibly null on a node
mid-rolling-upgrade. The handleChildState fallback already
handles both cases; only the comment was wrong.
Inline child approve/deny in the coord tree UI was rendering downstream
of the bulk-live cache (``GET /v1/api/cluster/ws/live``), not the SSE
stream. ``child_ws_state`` events were tiny notifications that fired
an urgent live-bulk fetch on every activity_state transition into/out
of "approval", just to pick up the rich ``pending_approval_detail``
payload. With multiple coord tabs and multi-child workstreams, that
urgent-fetch pattern compounded the SSE-executor pressure Shape A
is unwinding.
Thread the field through every layer so the SSE event itself carries
the rich payload — browser mutates ``liveBadgeCache`` directly,
no urgent fetch:
1. Node ``WebUI._broadcast_state`` emits ``pending_approval_detail``
on ``ws_state`` events. Gated on ``_pending_approval is not None``
so the per-broadcast verdict-cache deepcopy only runs when there
is actually an approval pending. ``_build_node_snapshot`` also
projects the field so the console's reconnect-via-snapshot
resync path delivers it (without this the new collector
forwarding would never see the field on a snapshot row).
2. Console ``ClusterCollector._apply_delta`` (live ``ws_state``
forwarding) and ``_reconcile_node`` (snapshot resync diff) both
forward the field on the emitted ``cluster_state`` event, AND
``_apply_delta`` persists it on the cached ``ws`` dict so the
``get_node_detail`` / ``get_snapshot`` endpoints between
reconciliations don't render stale approve/deny buttons.
3. ``CoordinatorAdapter._dispatch_child_event`` re-emits the field
on the ``child_ws_state`` event sent to coord listener queues.
4. Frontend ``handleChildState`` reads ``ev.pending_approval_detail``
and writes it directly into ``liveBadgeCache``, tagging the
entry with ``sseUpdatedAt``. ``flushLiveFetches`` honors that
tag for ``SSE_AUTHORITATIVE_MS`` (3s) — the upstream
``/dashboard`` cache has its own ~2s TTL, so a bulk-poll
landing right after a transition can otherwise clobber the
fresh SSE-set state with pre-transition data.
The pre-fix ``enteredApproval`` / ``leftApproval`` urgent-fetch
branch is removed. The 409 stale-call_id retry path keeps its own
urgent fetch — that's a different scenario.
Tests cover the forwarding contract at every layer, the broadcast
gate (event includes the field when an approval is pending,
omits it otherwise, and clears after resolution), and the
``flushLiveFetches`` merge-guard structural shape so a refactor
that keeps the symbols but inverts the comparison or drops the
``prev.live`` check can't pass silently.
``coordinator_children`` was calling ``storage.list_workstreams``
directly on the event loop, ``coordinator_tasks`` did the same with
``load_task_envelope``, and ``_resolve_coordinator_or_404`` (called
from both handlers, plus ``coordinator_history`` and
``_resolve_coord_session``) did the same with
``storage.get_workstream`` on its cold-cache path.
The cold-cache resolver path is hit on every console restart,
coordinator eviction, and console proxy hop — exactly when the
event loop is most contended. Three coord tabs reconnecting after a
brief network blip = three serial event-loop blocks per call site.
Other lifted handlers in this file already use
``asyncio.to_thread``; bring all four call sites onto the same
pattern.
Convert ``_resolve_coordinator_or_404`` to ``async def`` and update
its four call sites to ``await``. Exception flow is unchanged.
Each coord ``events`` SSE listener parks a thread on
``client_queue.get(timeout=5)`` for the connection lifetime. The
console's coord endpoint was wiring no ``sse_executor_lookup`` on
``coord_endpoint_config``, so those parks landed on Python's default
ThreadPoolExecutor (~min(32, cpu_count+4)) and competed with every
other ``asyncio.to_thread`` caller (storage, router, audit). A few
coord tabs against a multi-child workstream would stall new request
handlers waiting for a worker thread.
Mirror the interactive-side precedent (the ``sse_executor`` /
``sse_executor_lookup`` pattern in ``turnstone/server.py``) — build a
dedicated 200-thread ``coord_sse_executor`` in the console lifespan
and wire ``sse_executor_lookup`` onto ``coord_endpoint_config``.
Drain order matters: shut the pool down AFTER ``coord_adapter.shutdown()``
so no new listeners arrive at a dying pool. ``cancel_futures=True``
discards queued-but-not-started futures during teardown.
Update the stale comment on the interactive-side wiring that claimed
"coord wires None and falls back to the default executor" — it now
points at the console's matching wire.
Three follow-ups from Copilot's round-2 review on #453.
ValueError logging surfaced the wrong reason
The catch-all ``except ValueError:`` logged ``reason=no_enabled_rows``
unconditionally, but ``ModelRegistry.__init__`` raises ValueError for
five distinct config issues (empty models, default / fallback / agent /
plan / task alias not present). Operator looking at logs for a
config.toml typo would see the wrong cause. Switch to
``log.warning("...reason=%s", exc)`` so the actual error message
threads through. Behavior unchanged — existing registry still
preserved on every ValueError path.
Misleading shutdown() comment
The ``finally`` comment claimed shutdown() was closing clients the
throwaway registry created during DB load. ``load_model_registry`` only
constructs ModelConfigs and the bare ``ModelRegistry(...)``;
``ModelRegistry.__init__`` leaves ``_clients`` / ``_providers`` empty
and they populate lazily on first resolve. Today shutdown() iterates
empty dicts. Comment now says so explicitly while keeping the call
(and its try/except) for forward-compat against an eager-init future.
Stale "probe" wording in test docstring
``test_helper_preserves_registry_when_db_probe_fails`` →
``test_helper_preserves_registry_when_strict_load_fails``. The
explicit probe was removed in commit 1ba17ed when the helper switched
to ``load_model_registry(..., strict=True)``; the test name and
docstring still talked about a probe. Updated wording reflects that
the loader's strict-mode re-raise is what the helper catches now.
132 tests pass.
Hygiene follow-ups from the multi-stage code review on #453.
perf-1 — sync helper called from async route handlers
``_refresh_coord_registry`` runs two sync DB reads and a registry reload
that takes ``_client_lock``; calling it directly from an async handler
held the event loop for the duration. All four call sites now
``await asyncio.to_thread(_refresh_coord_registry, ...)``, matching the
pattern from commit ``1f7d6ad`` (offloaded ``tenant_check``).
perf-3 — ModelRegistry.reload() tore down all clients unconditionally
The reload always closed every cached client and provider, even when
the changed fields (``model``, ``temperature``, ``context_window``)
didn't touch the connection target. Now selective: clients drop only
when alias removed or ``(base_url, api_key, provider)`` differs;
providers drop only when alias removed or ``provider`` string differs.
Keeps connection pools warm across the common admin-edit case where
only metadata changed. Two new ``test_model_registry`` cases lock the
keep-warm vs drop-on-change behaviour, and the existing
``test_reload_clears_clients`` was updated (it asserted the old
overly-aggressive contract) into
``test_reload_keeps_clients_when_connection_target_unchanged``.
q-5 — helper rename
``_refresh_console_coord_registry`` → ``_refresh_coord_registry``. The
``console_`` prefix was redundant given the function lives in
``turnstone/console/server.py`` and sibling helpers there
(``_notify_nodes_model_reload``, ``_publish_config_change``,
``_collect_model_status``) all omit it.
q-1 — shared test middleware
``tests/test_admin_model_registry_refresh`` now imports the
header-driven ``_AuthMiddleware`` from ``tests/_coord_test_helpers``
and sets default ``X-Test-User`` / ``X-Test-Perms`` headers on the
``TestClient``. The local hardcoded variant duplicated infrastructure
the helper module exists to centralise.
q-3 — multi-alias test registry
``_make_registry`` extracted a ``_make_config`` helper and gained an
``extras={alias: model}`` param so multi-alias scenarios stop
hand-building ``ModelConfig`` literals.
``test_delete_endpoint_refreshes_registry`` now uses the helper.
310 tests pass across the related coordinator + model surfaces.
bug-3 / q-2 from the multi-stage review on #453: the previous test
``test_update_endpoint_with_empty_body_does_not_blow_up`` asserted only
that the registry's model name was unchanged after an empty PUT, which
holds whether or not the refresh ran (DB row matches registry → refresh
is idempotent). A regression that always called
``_refresh_console_coord_registry`` — exactly the gate this test was
meant to lock — would have left the assertion green.
Rename to ``test_update_endpoint_skips_refresh_on_empty_body`` and spy
on the helper via ``monkeypatch.setattr``. Empty-body PUT must register
zero calls; any future change that drops the ``if updates:`` gate now
fails loudly.
Two correctness follow-ups from the multi-stage code review on #453.
bug-2 / perf-2 (DB probe was theatre + double scan)
The previous probe defended nothing the loader didn't already swallow
on the next line: ``load_model_registry``'s row-loop catches Exception
internally, so a transient DB error after the probe still degrades to
a config.toml-only registry that ``existing.reload()`` would apply,
silently dropping every DB-sourced alias. And on the happy path each
CRUD paid for two scans of ``model_definitions``.
Add a ``strict: bool = False`` flag to ``load_model_registry``. When
strict, the row-loop's except re-raises instead of swallowing. The
helper passes ``strict=True`` and drops the probe — single DB scan,
real failure isolation, the loader's silent fallback can no longer
mask a partial-result regression. Default ``strict=False`` so CLI /
lifespan callers keep their boot-with-config-fallback behaviour.
bug-1 (shutdown could escape after a successful reload)
``ModelRegistry.shutdown()`` calls ``client.close()`` unguarded, and the
helper's ``finally`` block ran it outside the try/except. A raising
close() after a successful ``existing.reload()`` would surface as 500
with the registry already mutated and the audit row already recording
success. Wrap ``new_registry.shutdown()`` in its own try/except that
matches the helper's belt-and-suspenders error policy elsewhere.
The helper's docstring also drops the obsolete probe paragraph; the
``if existing is None: return`` branch gets a one-line inline comment
about the boot-from-empty case (the multi-paragraph version restated
behaviour the line itself documents).
129 tests pass (test_admin_model_registry_refresh + test_model_registry).
Two follow-ups from Copilot review of #453:
1. ``load_model_registry`` swallows storage read errors internally
(logs + continues with config.toml-only models). Without a strict
probe in the helper, a transient DB outage on an admin CRUD would
apply a truncated registry that drops every DB-sourced alias —
silently, since the loader returns a non-empty registry built from
``[models.*]`` config.toml entries. Add an explicit
``storage.list_model_definitions(enabled_only=True)`` probe before
the loader call so the failure is visible here and the existing
registry is preserved on outage.
2. The previous docstring claimed ``admin_model_reload`` "has its own
boot-from-empty story." It doesn't — it just calls this helper,
which no-ops when ``coord_registry`` is None. When no model rows
existed at boot, lifespan leaves the entire coord subsystem
uninitialized (no ``coord_mgr``, no ``coord_adapter``, no
``session_factory``), and a console restart remains required after
the operator adds the first row. Tighten the docstring to admit
that limitation rather than overstating the helper's reach.
New test ``test_helper_preserves_registry_when_db_probe_fails``
monkeypatches ``list_model_definitions`` to raise and asserts the
existing registry stays intact.
The console builds ``app.state.coord_registry`` once at lifespan startup
and the coordinator session factory closes over that exact instance.
Until now, the model-definition admin endpoints (create/update/delete)
wrote to the DB but never touched the in-process registry — and the
explicit reload button only fanned out to nodes via HTTP, also leaving
the console's own registry stale.
Symptom: an operator who changed the underlying model name behind a
local-LLM alias (same alias, same endpoint) saw the DB row update
immediately, but coordinator sessions kept calling the prior model
name until the console process was restarted.
Fix: a new helper ``_refresh_console_coord_registry`` rebuilds a fresh
ModelRegistry from DB and applies it to ``app.state.coord_registry``
via the existing thread-safe ``ModelRegistry.reload()`` — in-place
mutation preserves object identity so the factory closure keeps
working, and active coord sessions auto-pick up the swap on their
next ``send()`` via ``ChatSession._refresh_model_from_registry``.
Wired into four endpoints in ``console/server.py``:
- ``admin_create_model_definition`` — after the DB write
- ``admin_update_model_definition`` — after the DB write, gated on
``if updates:`` so a no-op PUT skips the rebuild
- ``admin_delete_model_definition`` — after the DB write
- ``admin_model_reload`` — between ``_publish_config_change`` and
``_notify_nodes_model_reload`` so the console mirrors what the
reload broadcasts to nodes
Failure isolation: a load or reload error leaves the existing registry
intact (logged + swallowed). Coord stays usable while the operator
investigates; the explicit reload remains the user-facing recovery path.
No node fan-out on CRUD — the explicit reload button continues to gate
cluster-wide HTTP propagation, preserving today's UX semantics on shared
clusters.
Tests in ``tests/test_admin_model_registry_refresh.py`` cover:
- helper-level: rebuild from DB, identity preservation, no-op when
registry is None, preservation on load failure / no-enabled-rows /
reload validation error
- endpoint-level: create / update / delete / explicit-reload all
refresh the registry; an empty PUT skips the rebuild
Production fan-outs are frequently hitting the 6 KiB per-child cap by
just 1-2 KiB, forcing the coordinator into a follow-up inspect_workstream
round-trip per truncated child to recover the tail. Bumping the cap to
10 KiB absorbs the common overshoot without changing the truncation
semantics — truncated=True still fires for genuinely oversized messages,
and inspect_workstream remains the unbounded follow-up.
Worst-case context impact: a 32-child fan-out at the cap is now ~320 KiB
(was ~192 KiB), still well within commercial model context windows.
Typical fan-outs of 1-5 children land at 10-50 KiB.
LAST_ERROR_MAX_LEN (1 KiB) is unchanged — it's intentionally smaller
than the wait cap so error truncation happens at write time, and
1 KiB still sits well below 10 KiB.
WAIT_MESSAGE_MAX_BYTES is referenced by name (not literal 6144) in the
truncation test, so no test value needs updating.
The coordinator system message was descriptive about parallelism rather
than prescriptive — "while multiple children run in parallel" framed
fan-out as incidental, and "a tasks entry, a child to own it" primed
singular delegation. The spawn_batch example (benchmark A, benchmark B,
prototype the winner) showed dependent work under a fan-out framing,
teaching the wrong shape.
In practice the coordinator failed to decompose enumerable requests
("top stories on HN, Lobsters, /r/programming, …") without explicit
"please fan this out" instructions, on both GPT-5.5 and Claude Opus.
base_coordinator.md
- Replace singular "a tasks entry, a child to own it" with plural
"enumerate the independent units of work, spawn one child per unit,
run them in parallel by default. Sequential only when one child's
output feeds the next."
- Tighten the delegation paragraph.
tools_coordinator.md
- Drop the persona repetition that duplicated base_coordinator.md.
- Drop the prescriptive "## Workflow shape" section (the cost note is
already in wait_for_workstream's tool description; the edit-X
redirect is already in the persona).
- Drop "in one approval" / "single approval" mentions to avoid
surfacing approval mechanics to the model.
- Replace the misleading spawn_batch example with truly independent
items; drop "(up to 10)" which overstated the cap (it's per-call,
not global, and is documented in the tool schema).
- Add a course-correction example to send_to_workstream — the pattern
coordinators most often replace with cancel-and-respawn.
- Drop the read action from the tasks examples to keep the lifecycle
(add → update → remove) coherent.
Coord system message ~16% shorter (4440 → 3722 chars). Both GPT-5.5
and Claude Opus now naturally decompose the news-board prompt without
explicit fan-out instructions. 29 prompt-composition tests pass.
The server's --help epilog and compose.yaml both reference
--skip-permissions, but the argparser never defined it, so any
container started with SKIP_PERMISSIONS=1 exited with
"unrecognized arguments: --skip-permissions".
Wire the flag through to app.state.skip_permissions, OR-ing it
with the existing tools.skip_permissions config-store setting so
the stored value still works on its own.
Replaces the old mermaid-rendering shot with a coordinator session
mid-attention — parallel tool batches, judge-graded approval,
children + tasks side panels — which more accurately represents
what the platform does today.
* perf(api): offload tenant_check to thread on lifted session handlers
Every make_*_handler factory in turnstone/core/session_routes.py invoked
cfg.tenant_check(request, ws_id, mgr) synchronously inside its async
handler. For the interactive surface tenant_check chains through
_interactive_tenant_check → _require_ws_access → resolve_workstream_owner,
which short-circuits on mgr.get(ws_id) for warm cache but falls through
to a synchronous get_workstream_owner SQL call on a cold cache,
blocking the event loop for the duration of the storage round-trip.
Wrap each of the 8 call sites (approve, close, cancel, events, history,
detail, send, dequeue) in await asyncio.to_thread(...) — mirroring the
existing storage-offload pattern at make_history_handler's other call
sites. Coord wires tenant_check=None and is unaffected. Five handlers
gain a local import asyncio (matching the per-handler lazy-import
convention in this module). Centralizes the offload rationale on
SessionEndpointConfig.tenant_check's field docstring.
Adds two regression tests in TestTenantCheckOnReadEndpoints that wire
the real resolve_workstream_owner as tenant_check and force the
storage fall-through path the existing class only stubbed past with
fake allow/deny callables.
* test(api): spy asyncio.to_thread to pin tenant_check offload
Copilot flagged the cold-cache regression tests for asserting the
response shape but not the offload itself: reverting
await asyncio.to_thread(cfg.tenant_check, ...) to the sync call shape
would still leave the storage fall-through working and the tests
green. Patch asyncio.to_thread inside both tests with an async spy
that records every offloaded callable, then assert cold_check is in
the call list — sanity-checked by reverting the history wrap locally
and watching the assertion bite (offloaded only contained
storage.get_workstream + storage.load_messages, missing cold_check).
* feat(coord): inline tool-batch construct replaces approval dock
The pinned bottom approval-dock didn't scale: a 10-call spawn_workstream
fan-out filled the whole pane with a wall of repeated verdict chips,
and the call → approval → result lifecycle was split across three
disconnected surfaces (.msg.tool bubble + dock + .msg.tool result).
Replaces it with one chat-stream construct per dispatch turn that
pairs each tool call with its result and embeds the approval gate:
- .coord-tool-batch--solo single-call serial turn
- .coord-tool-batch--parallel ≥2 calls; rows share a left rail
+ per-row tick so they read as
siblings of one assistant decision
Lifecycle: rows render with optional "judge evaluating…" placeholder,
upgrade in place when intent_verdict arrives, and on tool_result the
output lands paired under the originating row. When the batch needs
approval, one Approve/Deny/Always action row renders inside the
construct (envelope-level — server semantics resolve siblings
together). After approval_resolved the action row morphs into a
✓ approved / ✗ denied status pill that stays as a receipt.
Critical bug closed: when a page reload races a pending approval,
pre-scan tool_call_ids in history; turns whose call_ids have no
matching tool result are rendered pending (not resolved-approved).
The SSE approve_request replay then upgrades the existing batch
in place — drops --approved/--denied, adds --pending, swaps the
status pill for actions, and assigns activeBatch. Without this
the operator was locked out of any approval pending at reload.
Defence-in-depth follow-ups from the same review:
- approval_resolved falls back to a DOM lookup if activeBatch
is null (cross-tab resolution where this tab never set it).
- _appendVerdictLineTo dedupes via a row.dataset.verdictSig so
SSE reconnect storms + repeat intent_verdict events don't
tear down + rebuild an unchanged verdict line.
- judgeVerdicts Map soft-capped at 500 entries (FIFO eviction)
via _cacheJudgeVerdict.
- toolRows entries hold {batch, row} only — the originating
item payload is no longer pinned for the page lifetime.
- _scheduleScroll coalesces messagesEl.scrollTop writes through
requestAnimationFrame so history replay doesn't reflow once
per appended message.
- Rationale <details> now inserts immediately after the verdict
line (was tail-appending, breaking ordering once a result
landed below).
- .coord-tool-batch--error wired: _appendResultToRow lifts a
row's error onto the enclosing batch; _renderBatchRow does
the same for policy-blocked rows at construction.
- _buildStatusPill extracted; both _morphBatchResolved and the
appendToolBatch resolved-replay branch route through it.
Removed: ~248 lines of dead .approval-dock CSS, the dock <aside>
element from index.html, and the dead helpers showApproval's
prior body, hideApproval, claimApprovalFocus,
claimApprovalFocusForVerdict, applyJudgeVerdictToRow,
applyJudgePendingToRow, ensureDctxAfterRow, removeRationale,
setApprovalButtonsDisabled, the appendToolCall single-row wrapper,
and window.coordApprove. Five stale comment blocks referencing
the dock as if live also swept.
Children-tree's renderApprovalBlock is independent and untouched
(different surface, different .approval-block / .approval-pill
vocabulary).
* fix(coord): close four Copilot review gaps on PR 447
Copilot review on caa07e6 flagged four follow-ups:
1. History replay was rendering EVERY orphan tool_calls turn (one
that lacks a matching tool result message) as `pending: true,
judgePending: true`. That paints Approve/Deny on turns that
could be just running — auto-approved-and-still-in-flight, or
already-approved-and-still-in-flight — and clicking would 409
because the call_id isn't in `pending_items`. Add a new
`--running` state for the orphan case (no actions, neutral
accent stripe). SSE then upgrades in place: `--running` →
`--pending` when `approve_request` replays, or `--running` →
`--auto` when `tool_info` replays. Tool_result events still
route into the rows for the third case (already-approved + in
flight) since `toolRows` is populated. Kicker text reads
"Running · Parallel N" while ambiguous, so the operator can
tell the in-flight-replay state apart from a fresh "Parallel ·
N tools" auto-approved batch.
2. Removing the dock also removed its `aria-live="assertive"`
region — pending tool-batches now append into the polite
`#coord-messages` log (which gets flipped to `aria-live="off"`
during streaming), so a screen reader could miss the
action-required signal. Add an off-screen
`aria-live="assertive"` `#coord-sr-announcer` region and route
"Approval required: <name> + N more" through it whenever a
pending batch is created OR an upgrade-in-place promotes a
running batch to pending. Also mark pending batches with
`role="region"` + a matching `aria-label` so SR landmark
navigation surfaces them; both are dropped on resolve so the
resolved batch stops claiming the landmark.
3. `_resolveBatchAction` was selecting the first row whose
`data-call-id` was set and that wasn't `.error` — but
`approve_request` envelopes carry the FULL items list,
including auto-approved siblings whose `needs_approval=false`
means the server's `pending_items` won't recognise their
call_id (→ 409 on submit, or resolves the wrong gate). Tag
rows that are genuinely in `pending_items` with
`data-needs-approval="1"` at construction (and during
upgrade-in-place when SSE arrives), and select against that
selector specifically. Restores the legacy
`pendingApprovalCallId` contract that filtered on
`needs_approval` before the dock was retired.
4. The `.coord-tool-row-result` comment claimed the styles applied
a click-to-expand "collapsed" affordance like the interactive
UI's `.tool-output.collapsed`, but the implementation only set
`max-height: 240px; overflow: auto` (a scroll pane, not a
collapse with expand control). Update the comment to describe
what the rules actually do and explain the deliberate
divergence from interactive (coord is a diagnostic-leaning
read-once surface; an internal scroll pane reads with lower
friction than a click-to-expand control on the operator's
primary monitoring view).
No Python touched; node --check on coordinator.js clean.
* fix(coord): restore reload-time pending approval gate
Agent-Logs-Url: https://github.com/turnstonelabs/turnstone/sessions/30f630fe-3ded-4abe-991b-b5a95f699127
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* feat(api): expose pending_approval on workstream detail response
PR 447 / 93cb3d9 (Copilot autonomous follow-up) added a JS path that
reads ``wsSnapshot.pending_approval_detail`` off the
``GET /v1/api/workstreams/{ws_id}`` snapshot in coordinator.js
init() so a freshly-loaded chat tab can paint the inline approval
gate immediately at reload, without waiting for the SSE
approve_request replay (which leaves a brief --running flash on the
inflight orphan placeholder).
But the server's ``WorkstreamDetailResponse`` schema only declared
``{ws_id, name, state, user_id, kind}`` and the lifted
``make_detail_handler`` matched: nothing was populating
``pending_approval`` or ``pending_approval_detail`` on the wire.
The frontend block silently no-op'd at runtime; Copilot's
accompanying assertion only grep'd the JS source for the literal
strings, so it stayed green while the actual contract was missing.
Extend the contract to match the Copilot frontend:
- Add ``pending_approval: bool`` + ``pending_approval_detail:
PendingApprovalDetail | None`` to ``WorkstreamDetailResponse``,
same shape as the dashboard / cluster live projection.
- ``make_detail_handler`` reads ``ws.ui._pending_approval`` (only
treats it as live when ``isinstance(_, dict)`` so MagicMock-
based unit tests don't trip the path) and calls
``ui.serialize_pending_approval_detail()`` to fill the detail.
A serializer raise falls back to ``pending_approval=True`` +
``detail=None`` instead of 500ing the whole response — SSE
replay still carries the authoritative payload.
- ``test_returns_workstream_fields`` updated for the two extra
fields (False / None on a MagicMock UI).
- ``test_pending_approval_fields_propagate_from_ui`` is the new
behavioural test: stub a UI with a realistic
``_pending_approval`` dict + serializer return, assert the JSON
surfaces ``pending_approval=True`` + the items list.
- ``test_pending_serializer_failure_falls_back_to_bool_only``
pins the defensive degradation so a future serializer
regression can't 500 every reload.
Tests: 4822 pass (3 deselected live). Ruff + mypy clean.
* fix(coord): three regressions on PR 447 inline tool-batch refactor
Three regressions reported during operator harness shakedown, all
landed by the inline tool-batch refactor in caa07e6:
1. ``stripAnsi`` ReferenceError on every ``tool_result``.
``_appendResultToRow`` called ``stripAnsi(output || "")`` but the
helper only existed in ``ui/static/app.js`` — coord.js never
imported or defined it. The thrown ReferenceError propagated up
through ``appendToolResult``, aborting the SSE handler before
``loadTasksDebounced()`` could fire, AND the result block never
appended to the row, AND history replay's tool-message loop
bailed out at the first orphan-tool-result. Three reported
bugs (tasks pane stops auto-refreshing, tool output missing in
the modal, reload only rebuilds the conversation up to the first
tool result), one root cause.
Fix: hoist a local ``stripAnsi`` mirroring the interactive UI's
regex. Keep it local rather than centralised — coord and
interactive tool-output paths have different rendering
strategies, and the interactive helper isn't on the shared
module surface today.
2. JSON tool output rendered as a single unreadable line. Coord
tool surfaces (``list_nodes``, ``tasks``, ``spawn_workstream``,
...) emit JSON by default, and ``textContent = stripAnsi(raw)``
showed the whole envelope on one line. The parent
``.coord-tool-row-result`` already has ``white-space: pre-wrap``
so a ``JSON.stringify(parsed, null, 2)`` body lays out as
intended without a nested ``<pre>``. Non-JSON / unparseable
output falls through to the raw cleaned string.
3. Header tier badge stuck on ``⚙ heuristic`` after the LLM judge
landed an upgraded verdict. ``_pickBatchTier(items)`` ran once
at batch-creation time; later ``intent_verdict`` SSE events
updated the per-row chip via ``_appendVerdictLineTo`` but never
refreshed the head.
Fix: persist the verdict's tier on ``row.dataset.verdictTier``
(+ ``verdictModel`` when set), add ``_refreshBatchTier(batch)``
that scans the rows and computes the cross-row best tier (LLM
beats heuristic), and call it from ``_appendVerdictLineTo``
whenever a row writes a verdict. ``_pickBatchTier`` gets the
same prefer-LLM scan so the initial render is consistent. The
``intent_verdict`` cache entry tags ``tier: "llm"`` so a late
verdict landing on a previously heuristic-only row escalates
the badge correctly.
No Python touched; node --check on coordinator.js clean.
* fix(coord): close five Copilot review gaps on PR 447
Five distinct findings from the second Copilot pass on the inline
tool-batch refactor (the sixth — stripAnsi ReferenceError — already
shipped in 77dc24e):
1. CSS rail tucks never matched. The ``--first / --last`` row trims
used ``:first-of-type`` / ``:last-of-type``, but the batch
contains other ``<div>`` siblings (.coord-tool-batch-head,
.coord-tool-actions / .coord-tool-status) — the
structural-pseudo-class is type-based (``div``), not class-
based, so the first .coord-tool-row is not the first ``<div>``
in the parent. Selector silently no-op'd, leaving the rail
butting against the inner top/bottom edges of the batch. Fix:
apply explicit ``.coord-tool-row--first`` / ``--last`` markers
in JS at row-build time and key the CSS off them.
2. Upgrade-in-place left stale ``data-needs-approval`` markers on
non-pending sibling rows. The original block only added the
attribute for items where ``needs_approval=true``, never
clearing it for rows whose earlier (replay-time) shell tagged
them. ``_resolveBatchAction`` could then pick a non-pending
row's call_id, yielding a 409 stale call_id on approve / deny.
3. Upgrade-in-place left row-level status pills out of sync with
the SSE-authoritative item shape. When a ``--running`` orphan
gained a ``tool_info`` envelope, the ✓ auto pill never
appeared; when it gained an ``approve_request`` envelope with
policy-blocked siblings, the ✗ blocked pill / ``.error`` class
were missed. Batch-level state classes flipped, but per-row
visual cues lagged.
Fix for 2 + 3: extract ``_refreshRowStatus(row, item)`` from
``_renderBatchRow``. It clears prior ``data-needs-approval`` +
pills and re-applies from the item, preserving runtime
``tool_result`` errors via the new
``.coord-tool-row-result--error`` marker on the result block.
Both ``_renderBatchRow`` (initial render) and the
upgrade-in-place loop now route through it, so the two paths
can't drift.
4. History replay defaulted ``item.needs_approval = true`` on
every synthesized tool call. ``_renderBatchRow`` then tagged
the row with ``data-needs-approval="1"`` regardless of whether
the call genuinely needed approval. Combined with the missing
clear in finding 2, an SSE upgrade with a mixed envelope kept
incorrect markers on auto-approved siblings. Drop the
replay-time default; let SSE supply the authoritative bit when
the upgrade fires (``_refreshRowStatus`` reads it from the
item).
5. Tool result routed into an existing batch row didn't trigger
``_scheduleScroll()``. Result blocks grow ``scrollHeight``;
without the rAF-coalesced scroll the user pinned at the bottom
loses their pin when the row inflates. Add the call after
``_appendResultToRow`` in the early-return path so this branch
matches ``appendMsg``'s pinning behaviour.
Plus comment-only:
6. Detail-handler comment claimed "the JSON omits the section"
when the UI doesn't expose ``serialize_pending_approval_detail``,
but the response always includes both keys (with ``False`` /
``null`` for the bool / detail). Updated to match the actual
shape.
Tests: ``test_workstream_endpoints.TestDetailInteractive`` +
coordinator-detail + page tests pass (14 / 0 failed). Ruff +
mypy clean. ``node --check`` on coordinator.js clean.
* fix(coord): close 17 review findings on PR 447
Second /review pipeline pass surfaced 16 confirmed findings (1 sec
major, 1 bug major, several minor + nit); operator harness shakedown
+ this commit's stale-comment sweep adds one more. All addressed
here.
Security:
sec-1 (major) — make_detail_handler + make_history_handler in
session_routes.py now invoke ``cfg.tenant_check`` after ws_id
validation, matching every other lifted session verb (send /
approve / close / cancel / events / attachments). Pre-fix the
detail response carried 5 low-data fields and history exposed
message rows; PR 447 added pending_approval_detail to detail
(tool previews + LLM judge reasoning) which made cross-tenant
reads via the missing gate a real disclosure on the interactive
surface (coord wires tenant_check=None and is unaffected). Plus
4 new regression tests in TestTenantCheckOnReadEndpoints that
wire a tenant_check function into the test cfg and assert the
gate fires on detail + history.
Bug fixes:
bug-1 (major) — history replay used to render every fully-
resolved tool batch as ``resolved: { approved: true }`` regardless
of the persisted tool result content. A denied tool round-trip
showed the green "✓ approved" pill alongside the persisted
"Denied by user" result text — directly contradictory state. Fix:
pre-scan classifies each tool message via a ``callOutcomes`` Map
by inspecting content prefix ("Denied by user" / "Blocked by
tool policy" / "Error:") and ``m.is_error``. Assistant tool_calls
render ``resolved.approved=false`` when any call's outcome is
"denied"; the existing --running fallback covers orphan turns
(any call lacking an outcome).
bug-2 — _verdictSig joined recommendation/risk_level/confidence/
reasoning only. When a late LLM verdict text-matched the earlier
heuristic verdict, the dedupe early-return fired before the
row's dataset.verdictTier was updated, so _refreshBatchTier
never escalated the header from "⚙ heuristic" to "⚖ llm".
Fix: include verdict.tier and verdict.judge_model in the
signature (with a "\x1f" separator instead of the empty join,
reducing field-boundary collision risk).
bug-3 — history replay's tool-result rendering hardcoded
isError=false. A runtime tool error on reload rendered without
the .error class, --error stripe, or "✗ error:" lead. Fix:
the same callOutcomes pre-scan that drives bug-1's denial path
also classifies "Error:" prefixes; appendToolResult now receives
isError=callOutcomes.get(callId) === "error".
bug-4 — approval_resolved derived ``wasAlways`` exclusively from
this tab's ``batch.dataset.requestedAlways``; cross-tab "Always"
click never propagated to peer tabs' status pill. Fix: server's
resolve_approval now takes a keyword ``always`` arg and includes
it on the SSE event body; client prefers ``ev.always`` and falls
back to the dataset stash for the hot-deploy window where the
SSE event might briefly omit the field.
bug-5 (nit) — appendToolBatch's create-new path overwrote
toolRows entries unconditionally. A partial-mapped envelope
(some call_ids previously seen, some new) silently orphaned the
prior batch's row pointers. Fix: detect the partial overlap,
console.warn, unmap the stale entries before the new batch
claims them.
Performance:
perf-1 — _refreshBatchTier did a querySelectorAll per verdict
insertion; for an N-row batch upgrade this was O(N²) DOM walks.
Coalesce via queueMicrotask + a _tierDirtyBatches Set so a burst
of N verdict updates collapses into ONE tier scan. Synchronous
body extracted to _refreshBatchTierImmediate (called from the
microtask flush).
perf-2 — _appendResultToRow pretty-printed JSON via
JSON.parse + JSON.stringify(parsed, null, 2) on every tool
result with no size cap. A 100KB JSON output stalled the main
thread; 10 parallel tool_result events compounded. Fix: gate
on cleaned.length <= 32 KiB AND a first-char check (0x7B / 0x5B)
so plain text + oversized payloads skip the parse. Parent CSS
is white-space: pre-wrap so raw text still wraps.
Quality:
q-1 — deleted dead row.dataset.funcName write (no readers).
q-2 — extracted _formatTierLabel(llmModel, hasHeuristic) shared
by _pickBatchTier (item-driven) and _refreshBatchTierImmediate
(dataset-driven). Single source of truth for the tier label
literals.
q-3 — extracted _pendingKickerText(items) used by both the
upgrade-in-place and fresh-build paths in appendToolBatch.
q-4 — added string-presence assertions to
test_coordinator_js_exposes_inline_approval_helpers covering
the new tool-batch helpers (appendToolBatch, _morphBatchResolved,
_resolveBatchAction, _refreshBatchTier, _refreshRowStatus), the
--running / --pending state classes, and the callOutcomes
outcome classifier.
q-5 — renamed _announcePolitelyAssertive → _announceAssertive.
Function unconditionally writes into the aria-live="assertive"
region; "politely assertive" was contradictory.
q-6 — rescoped the test docstring to acknowledge it covers two
layers (Chunk 3 children-tree + PR 447 tool-batch).
q-7 — tightened pending_approval_detail: Any → dict[str, Any]
| None in make_detail_handler. Mypy-confirmed.
Plus the third /review pass's q-1 stale-comment sweep:
_resolveBatchAction's comment still claimed the server doesn't
echo ``always`` on approval_resolved — wrong post-bug-4-fix.
Updated to reflect that the dataset stash is now backward-compat
fallback only, not the primary source.
Tests: 4826 pass (+4 new from TestTenantCheckOnReadEndpoints, plus
expanded assertions in TestDetailInteractive). Ruff + mypy clean.
``node --check`` on coordinator.js clean.
Verifier confirmed all 16 findings; pass-3 /review on the
addressing-commit surfaced only 0 critical / 0 major / 2 minor /
2 nit, none blocking. The two pass-3 minor findings are
pre-existing patterns across all lifted verbs (sync tenant_check
inside async handlers) and best addressed in a dedicated follow-up
PR auditing the whole lifted-verb surface.
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
The system-message developer block was rebuilt every turn with two
unstable inputs: minute-precision current_datetime in the middle of
the composed prefix, and _pending_nudge entries appended-then-cleared
at the bottom. Both invalidated prompt-cache reuse on Anthropic /
OpenAI for the entire prefix, every turn.
current_datetime now rounds to the top of the hour. Nudges no longer
ride on the system message at all — they drain through two channels:
- tool_error and repeat ride the existing tool-result <system-reminder>
envelope via a new MetacognitiveAdvisory ToolAdvisory subtype, drained
in _collect_advisories alongside GuardAdvisory and UserInterjection.
- correction, denial, resume, start, completion splice as
<system-reminder> blocks at the trailing edge of the next user
message via a new _splice_pending_user_advisories helper.
User content passes through escape_wrapper_tags before concatenation
so a user typing literal <system-reminder> tags cannot fabricate an
envelope; the same escape now runs on advisory.render() output inside
wrap_tool_result for defense-in-depth across all advisory types.
Cancel handlers (GenerationCancelled / KeyboardInterrupt / bare
Exception) now clear _pending_tool_advisories alongside the existing
_flush_queued_messages so a queued nudge from an aborted batch
cannot leak into the next generation.
Visibility ping ([metacognition: nudge injected — ...]) preserved at
both new attach points via a single _emit_nudge_ping helper.
Also adds a "Session kind" line (interactive | coordinator) to the
composed Session Context so the model can see which manager hosts
its session.
Tests: 4847 passing (+7 new in TestMetacognitiveBuffers and
test_tool_advisory). ruff + mypy clean.
Copilot review on c5fe3e7 flagged three follow-ups:
1. tasks-write batch order was scheduler-dependent. Prior comment
claimed the result was "deterministic against the input set even
if the dispatch order isn't" — true for the SET of tasks, but
``tasks_add`` appends under a per-ws lock, so the FINAL list
ordering (and order-derived timestamps) varied with whichever
thread happened to acquire the lock first. Fix: when a batch
contains any tasks-write, the dispatcher runs the WHOLE batch
serially in input order. Other batches stay parallel.
2. ``tasks_add`` test stubs were the wrong shape. ``CoordinatorClient.
tasks_add()`` returns the task dict directly with top-level
``id`` / ``title`` / ``status`` / ``child_ws_id`` / ``created`` /
``updated`` — the previous stubs wrapped it as ``{"ok": True,
"task": {...}}`` and weakened the tests since
``_exec_tasks``'s summary path reads ``result.get("id")`` and
would have seen ``"?"`` against the wrong shape. Both stubs
updated to match the real contract.
3. New regression test pins the input-order property on tasks-write
batches. ``test_tasks_writes_run_in_input_order`` captures the
``tasks_add`` call sequence and asserts it matches the model's
emit order exactly — pre-fix this would be scheduler-dependent.
Plus ``test_tasks_writes_serial_when_mixed_with_non_tasks_siblings``
pins the same property when the batch interleaves a
``list_nodes`` call with two ``tasks(add)`` calls.
Tests: 4820 pass (+2 net since the prior PR 446 push). Ruff +
mypy clean.
Operator's harness shakedown found list_nodes filters silently
ignored on every call:
list_nodes(os="Linux") → returns ALL 10 nodes
list_nodes(has_gpu=true) → returns ALL 10 nodes
list_nodes(memory_gb=751) → returns ALL 10 nodes (no
node has 751 GiB; should be 0)
Storage filter pipeline is fine (pinned by an existing
test_list_nodes_filter_uses_natural_value_not_quoted). The bug is
upstream in ``_prepare_list_nodes``: it only honoured
``args["filters"]`` (the canonical nested shape). Several models
drop the nesting and emit each filter as a top-level kwarg —
``list_nodes(os="Linux")`` instead of ``list_nodes(filters={"os":
"Linux"})`` — and the strict prepare silently degraded those calls
to "no filter" → full-cluster return.
Fix: any top-level kwarg that ISN'T one of the four reserved control
parameters (``filters``, ``limit``, ``include_network_detail``,
``include_inactive``) is now treated as a flat filter. Nested entries
still win on key collision so the canonical shape stays
deterministic. Tool description unchanged so well-behaved models
keep using ``filters={...}``; the relaxation is purely receiver-side.
Tests: 4818 pass (+5 net). Five new tests pin both shapes plus the
collision-precedence rule and the prepare→exec wiring. Ruff + mypy
clean.
Operator observed the prior rule rejecting a natural decompose-the-
plan turn:
[tasks(add×4), list_nodes, list_skills, list_workstreams]
The 4 tasks(add) calls landed (per-ws lock serialised them) but the
guard blanket-rejected EVERY tasks(...) regardless of what its
siblings actually were. All-write batches converge under the
per-ws lock; all-read batches can't race. The only genuinely-
hazardous shape is the read+write mix where tasks(list)
paralleled with tasks(add=...) inside ``run_one``'s
ThreadPoolExecutor has unspecified ordering and the read can land
on either side of the write.
The rule now scopes precisely:
- All ``tasks`` writes in a batch — permitted.
- All ``tasks`` reads in a batch — permitted.
- ``tasks`` paralleled with non-``tasks`` siblings — permitted
in either direction. Non-tasks tools don't touch the tasks
state, so there's no read-after-write surface.
- ``tasks`` read AND ``tasks`` write in the same batch — REJECTED
(still, because that IS the actual hazard).
Tests: 4813 pass (+4 net). Six new tests pin the relaxation
(all-write OK, all-read OK, write+sibling OK, read+sibling OK,
non-tasks-only batch unaffected) and the one tightened rejection
case (read+write mixed in tasks specifically). Ruff + mypy clean.
* feat(node): auto-detect node capabilities via kernel interfaces
Closes the operator-burden gap the harness shakedown surfaced — the
list_nodes capability/region/role filtering surface that nodes were
launching with empty. Auto-detection runs at server startup and
populates ``node_metadata`` rows with sensible defaults that
operators can still override via the ``[metadata]`` section of
config.toml (operator-config writes win on the per-key upsert).
What's detected, all from kernel interfaces (no userspace binaries
on PATH — works the same way regardless of whether nvidia-smi /
rocm-smi / lspci is installed):
- ``gpu_count`` / ``gpu_vendor`` / ``gpu_vendors`` / ``gpus`` —
walks ``/sys/class/drm/cardN/device/{vendor,device}`` and decodes
PCI vendor IDs to friendly names (NVIDIA / AMD / Intel / Apple).
Heterogeneous-GPU nodes get the first KNOWN vendor in the flat
``gpu_vendor`` key — never ``"unknown"`` when known vendors are
present — so a coord filtering on ``gpu_vendor=nvidia`` matches
nodes whose first card happened to be exotic.
- ``memory_gb`` — reads ``/proc/meminfo``, rounds GiB down so
``filters={"memory_gb": 32}`` doesn't match a 31.5 GiB node.
- ``cpu_model`` — first ``model name`` line from ``/proc/cpuinfo``.
- ``cloud_provider`` / ``cloud_region`` / ``cloud_zone`` /
``cloud_instance_type`` / ``cloud_instance_id`` — DMI sysfs
identifies the cloud provider from BIOS/SMBIOS strings (no
network call) and only THEN does the IMDS probe fire. Baremetal
hosts pay zero startup latency on the cloud path.
Hardening highlights:
- IMDS probes target the link-local IP literal ``169.254.169.254``
for AWS, GCP, AND Azure — no DNS-resolvable hostname for any
vendor, so a host with attacker-controlled DNS can't redirect
the probe even when its DMI claims a cloud provider.
- Response bodies capped at 64 KiB on read; per-field strings
capped at 256 chars and stripped of control characters before
persistence. Stops a hostile IMDS responder from spraying
multi-megabyte / newline-injected payloads into ``node_metadata``
and from there into coord-LLM ``list_nodes`` context.
- ``isinstance(doc, dict)`` guards on every JSON IMDS response so
a non-conformant body (list / scalar / null) returns clean ``{}``
instead of raising.
- ``collect_node_info()`` runs via ``asyncio.to_thread`` from the
server's lifespan handler so the IMDS probe latency never blocks
the event loop.
- GCP fans the three zone/machine-type/id probes concurrently so a
misidentified host's worst case is one timeout window (~1 s)
instead of three (~3 s).
- Operator opt-out via ``TURNSTONE_AUTO_CLOUD_METADATA=0`` skips the
IMDS phase entirely; the DMI-derived ``cloud_provider`` still
populates because that's a kernel interface.
Tests: 4807 pass (+11 net, 73 in test_node_info.py). Ruff + mypy
clean on every modified file. New tests pin the heterogeneous-GPU
flat-key fix, the IMDS hardening (non-dict JSON, control-char
sanitisation, body cap, per-field cap), and the GCP IP-literal
property.
* fix(node): filter synthetic display adapters + per-vendor GPU flags
PR review on 68cb0ab flagged two real issues with the GPU surface:
1. Hyper-V synthetic display adapter (vendor 0x1414, device 0x06)
registers a /sys/class/drm/cardN entry on Linux but is NOT a
compute GPU. A CI runner reproduced this and came back with
gpu_count=1 on a CPU-only VM. Same hazard for AWS Nitro VGA,
QEMU virtio-gpu, and any other hypervisor synthetic display
adapter. Fix: ``_detect_gpus`` filters DRM cards by PCI vendor
against the GPU allow-list (NVIDIA / AMD / Intel / Apple); cards
from other vendors are skipped entirely. Operators with exotic
accelerators that don't match any known vendor can still set
``gpu_count`` + the relevant flags via [metadata] config.
2. ``gpu_vendor`` (singular flat key) was sorted-alphabetical-first
of the unique known vendors. list_nodes() filtering does
exact-equality JSON matching, so a mixed AMD+NVIDIA node ended
up with ``gpu_vendor=amd`` and was invisible to a coord
filtering ``gpu_vendor=nvidia``. Fix: drop the singular key
entirely; emit per-vendor booleans (``gpu_has_nvidia=true``,
``gpu_has_amd=true``) so multi-vendor nodes match EITHER vendor.
Also add ``has_gpu=true`` for "any compute GPU at all" filtering.
Tests: 4809 pass (+2 net). New / updated tests pin both behaviors:
- _detect_gpus: Hyper-V synthetic + arbitrary unknown-vendor card
now filter out; mixed-known-and-unknown keeps only the known card.
- collect_node_info integration: multi-vendor node has both
``gpu_has_amd`` and ``gpu_has_nvidia`` set; ``gpu_vendor`` (singular)
is asserted absent so a future regression that re-introduces it
fails loudly.
Ruff + mypy clean.
* fix(coord): close gaps an operator's harness shakedown surfaced
Operator-driven shakedown of the coordinator tool surface flagged
five issues; this commit addresses all of them plus the review
findings against the initial fix.
1. Cancelled-mid-stream partial assistant content now carries a
"[generation cancelled before completion]" marker. Without it,
``inspect_workstream`` / ``wait_for_workstream`` callers and the
next coord-LLM turn read the truncated text as a complete answer.
``_cancelled_partial_msg`` no longer ships ``_provider_content``
(Anthropic would otherwise read that lane verbatim and bypass the
marker; partial tool_use blocks could also leak through).
2. ``spawn_workstream`` / ``spawn_batch`` no longer surface the
routing-proxy ``status`` field (always HTTP 200 on the success
path). The tool description claimed it was "lifecycle state at
creation"; code that did ``if result["status"] == "idle"``
silently never matched. Lifecycle state lives on the workstream
row — ``inspect_workstream`` is the read. Tool JSON descriptions
plus docs/coordinator-skills.md and docs/bulk-endpoints.md
examples updated to match.
3. ``inspect_workstream`` not-found error string is bare ("workstream
not found"); the structured ``ws_id`` field carries the queried
id. Pre-fix the error STRING echoed the id back at the caller
who just sent it — redundant and out of step with the rest of the
surface. Cross-tenant + missing rows still return the same shape,
preserving the existence-leak guarantee.
4. ``tasks(...)`` is now rejected when called in a parallel tool
batch. The prior shape relied on a docstring warning ("a list
paralleled with writes can reflect pre-write state") that put
cognitive overhead on every model invocation; turning the silent
footgun into an explicit error means the model only thinks about
the rule the moment it actually breaks it. Warning dropped from
the tasks tool description. ``_PARALLEL_INCOMPATIBLE_TOOLS``
constant in session.py is the extension point for any future
tool with the same read-after-write hazard.
Plus the multi-stage code review's findings against the initial
fix (q-1 / q-2 docs drift, q-3 idiom, q-4 keys-assertion, q-5
duplicate guard) — all addressed in the same pass.
Tests: 4752 pass, +6 net since the pre-fix baseline. Ruff + mypy
clean. Three new tests pin the parallel-batch-rejection behaviour
on tasks (rejected when batched, runs alone, sibling tools
unaffected); existing cancel + spawn + inspect tests updated to
match the new shape.
* fix(coord): close two copilot review gaps on PR 444
Copilot review on PR 444 flagged two follow-ups:
1. Empty-content cancel divergence — when ``GenerationCancelled``
races BEFORE the first content token, the prior shape skipped
``save_message`` and only appended an empty-content msg in
memory. In-memory and storage diverged: a rehydrate would see
nothing in storage but the session would carry an empty
assistant turn. Both branches now persist; on the empty-content
shape the marker becomes the entire message
("[generation cancelled before completion]") so storage matches
the in-memory history.
2. Test stub cleanup — three new tests injected ``ui.approve_tools``
via ad-hoc ``lambda + type: ignore[attr-defined]``. Replaced
with a permissive ``approve_tools`` method on ``_StubUI`` so the
stub matches the SessionUI surface the dispatcher actually
reads. Tests that exercise approval pathways can still override
per-instance.
Tests: 4752 pass. Ruff + mypy clean.
* feat(coord): surface child errors, isolate tool exceptions, add memory tool
Closes four coordinator gaps identified during operator triage:
1. Child workstream errors now surface in inspect/wait. Worker-thread
exception text is sanitized (URL userinfo masked, sk-/Bearer/ghp_/
github_pat_/AKIA tokens redacted, capped at 1024 chars) and persisted
to workstream_config.last_error before _emit_state("error") fires, so
coord polling never sees state=error with a missing cause. The row
is cleared on recovery transitions (idle/running) so a once-leaked
exception body doesn't outlive the failure. inspect_workstream and
wait_for_workstream return last_error for state=error rows; the
wait surface prefers it over the assistant-tail walk.
2. Tool exceptions now return as tool_results with sibling-aware
guidance. ChatSession._safe_prepare_tool wraps every per-call
_prepare_tool invocation; a buggy preparer becomes an error item
for that call only — sibling parallel tool_calls keep going,
never orphaning the assistant message's tool_calls block.
run_one's runtime exception path includes the exception class
and a short note that other tool calls in the batch completed
independently so the model can recover.
3. Memory tool exposed to coordinator with a coord-only scope.
memory.json gains coordinator: true + interactive: true + per-kind
kind_variants. Coord sessions see scope enum ["coordinator"] and
an orchestration-flavored description; IC sessions see ["global",
"workstream", "user"] and the existing flavor. Coord-scope rows
are private to the coordinator session (children cannot read or
write them), closing the cross-session prompt-injection lane that
an adversarially-steered child would otherwise have. Coord
visibility is also restricted to coord-scope only — coords no
longer see global / workstream / user memories that belong to the
user's interactive sessions.
4. Per-call exception isolation in tool batches. _safe_prepare_tool
was previously the implicit shield; now it's an explicit method
with documented invariants. KeyboardInterrupt / GenerationCancelled
re-raise so the cooperative cancel path still works.
Other notable changes:
- LAST_ERROR_CONFIG_KEY + persist_last_error / clear_last_error /
load_last_error / sanitize_error_text moved to turnstone.core.memory
(the storage facade hub) — readers in coordinator_client.py import
the constant.
- Memory scope tuples extracted to module constants
_VALID_MEMORY_SCOPES and _IMPLICIT_SCOPE_WALK; seven inline
duplicates collapsed.
- tools.py grows _apply_kind_variant for the per-kind tool surface;
tools without kind_variants pass through unchanged (no spurious
deep-copies).
- Session adds _coordinator_scope_id, _default_memory_scope,
_implicit_scope_walk, and _record_fatal_error chokepoints so the
worker-thread fatal path is one site rather than three.
- Removed duplicate on_error / on_state_change emits from
session_routes.py and coordinator_adapter.py — session.send()'s
_record_fatal_error owns the sequence now.
Tests: 4742 pass (no live), +30 net since the baseline. Ruff + mypy
clean on every modified production file.
* fix(coord): redact secrets in tool error paths via output_guard
Copilot review flagged two paths where ``str(exc)`` flowed back into
the model-facing tool_result without going through the credential-
redaction the new fatal-error path applies:
- ``ChatSession._safe_prepare_tool``: a preparer-side exception
becomes an error item whose ``error`` field embedded the raw
exception text.
- ``ChatSession._execute_tools.run_one``: a runtime tool exception
became an ``Error executing X: <e>`` tool_result, again with
the raw exception text.
Both now route through ``sanitize_error_text`` (sanitised log line +
sanitised tool_result), and ``sanitize_error_text`` itself was
refactored to delegate to ``output_guard.redact_credentials`` instead
of carrying its own parallel regex catalog — the audit log + post-tool
guard already use that pattern set, so the credential definition
stays in one place.
Also extended ``_RE_CONNECTION_STRING`` in ``output_guard`` to cover
``http(s)://user:pass@host`` so a misconfigured ``OPENAI_BASE_URL``
that lands in an httpx ``ConnectError.__str__`` is redacted by every
caller of ``redact_credentials`` (audit details, close-reason
persistence, last_error, the two tool error paths). The
host (useful for triage) survives; only the password is replaced
with the standard ``[REDACTED:password]`` marker.
Tests: full suite (4745 pass), ruff + mypy clean. Two new tests pin
the redaction behaviour in both tool error paths so a future refactor
can't drift back to leaking ``str(exc)`` verbatim.
Bring the coord dashboard toward parity with the interactive pane on
two operator-visible surfaces:
- Status bar pinned above the composer. Same four cells as the
interactive pane (model, token / context-window usage with effort
suffix, tool calls this turn, conversation turn) driven by the
same on_status SSE events. ws-status-bar CSS hoisted from
ui/static/style.css to shared_static/chat.css so both UIs read one
copy. StatusBar.paint helper extracted to
shared_static/status_bar.js; both Pane.prototype.updateStatus and
the new coord updateStatusBar delegate to it so warn/danger
thresholds, prefix glyphs, and effort-suffix rules can't drift.
CTX_WARN_PCT / CTX_DANGER_PCT now named constants on a single line.
- _coord_events_replay now yields the connected + status preamble
via a shared session_replay_preamble helper in
turnstone/core/session_replay.py. _interactive_events_replay
routes through the same helper so a future field add lands once.
Coord still skips conversation history in the SSE replay (the
dashboard fetches it via GET /history); only the status preamble
is shared.
- History replay reconstructs tool calls. Pre-fix, an assistant
turn that only dispatched tools rendered as an empty bubble
followed by raw tool-result text — the call's intent and
parameters were lost on reload. synthesizeHistoricalToolCall
builds an appendToolCall-shaped item from the persisted
function.name + function.arguments (special-casing bash so the
shell line shows in the header). Tool result rows now resolve
their label from the matching tool_call_id instead of always
printing "tool".
- onopen restores the tokens placeholder when no prior status was
seen, so a transient SSE blip on a fresh coord doesn't leave the
dim "Reconnecting…" copy stuck until the next live tick.
Tests: 4 new tests for the shared replay preamble (connected first,
status only when last_usage present, status payload shape, no-session
fallthrough); existing approval/verdict ordering tests refactored
through a shared make_replay_mocks helper in tests/_replay_helpers.py
that both interactive and coord suites import.
tasks(update) is the only mutation that allows title to be omitted,
so _prepare_tasks stores ``item["title"] = None`` for an update that
only changes status/child_ws_id. _evaluate_intent then projected via
``it.get("title", "")[:100]`` — but dict.get returns the stored None
(the default kicks in only when the key is absent), and the slice
crashed with ``TypeError: 'NoneType' object is not subscriptable``.
The exception fired before any tool in the parallel batch executed,
so the assistant's tool-call message was already on the wire while
no tool-result entries followed. Reconstruction/sanitisation later
synthesised "Tool execution was cancelled" for every sibling — the
visible symptom that masked the real None-slice failure.
- Switch tasks/notify/task_agent/plan_agent/spawn_workstream/
spawn_batch/send_to_workstream/close_workstream/close_all_children
projections to ``(it.get(x) or "")[:N]`` so absent and explicit-None
both fall back to the empty string. The other tools weren't
observed crashing, but the bug shape is identical at every site;
hardening the projection layer once costs one extra ``or`` per line
and removes the foot-gun for any future preparer that stores None.
- Regression tests reproduce the original TypeError on
``tasks(update)`` without title both standalone and in a parallel
batch alongside ``tasks(add)``.
* feat(console): per-call model + judge_model on coord composer
Brings the landing-page coordinator composer toward parity with the
interactive new-ws modal — operators can now pick a model and judge
model per session without round-tripping through the Models admin tab.
- Add Model + Judge Model selects to the home composer's options
panel, populated from /v1/api/models. Empty / non-string fields
collapse to None so the factory falls back to ConfigStore defaults
(coordinator.model_alias, judge.model).
- _coord_create_build_kwargs threads the body fields onto mgr.create.
- Console session factory accepts judge_model and overrides the
JudgeConfig via dataclasses.replace, mirroring the server-side
interactive factory's pattern (alias preserved for IntentJudge's
provider/client resolution).
- Sanitise the 503 factory-misconfig response across make_open_handler,
make_create_handler, and make_detail_handler: a new
_safe_factory_misconfig_message helper strips control characters
and caps at 200 chars before echoing exc text. Operators still get
the full alias in the warning log; clients see a bounded printable
string. Defends the user-controlled body["model"] reflection
surface on the create path.
- _build_mgr_with_factory test helper extracted from _build_mgr so
tests that need to capture factory kwargs don't reconstruct the
CoordinatorAdapter + SessionManager scaffolding inline.
- Tests cover: passthrough of model + judge_model, empty / whitespace
/ non-string body fields collapsing to None, and the 503 sanitiser
truncating + scrubbing a hostile alias payload.
* fixup: address PR #440 Copilot review
- _safe_factory_misconfig_message: hard-cap return at
_FACTORY_MISCONFIG_MAX_LEN total (was MAX_LEN+1 because the slice
was MAX_LEN long with the ellipsis appended on top). Reserve one
codepoint for the ellipsis so the cap is honoured. Update the
regression test to assert the tighter bound.
- Composer judge_model placeholder: "Default (agent model)" was
misleading when ConfigStore judge.model is set — the actual fallback
is judge.model when set, IntentJudge's agent-model fallback when
not. Use "Default judge model" instead so the label matches both
configs.
- Drop the duplicate "N nodes · M workstreams" header span — same data is
already on the page.
- Drop the "+ new" workstream header button + modal; the coordinator
composer is now the primary entry point on the landing page.
- Always render the NODES list inline; remove the cluster-summary
compact toggle since the list already self-collapses same-prefix
nodes into groups.
- Replace the meta node-detail page (#view-node) with direct navigation
to /node/{node_id}/. Removes drillDownToNode, loadNodeDetail,
_loadNodeMetadataPanel, the popstate "node" branch, and the
currentNodeId/currentServerUrl state.
- popstate now falls back to showHome() for unknown state shapes so a
back-nav from a tab on an older build doesn't no-op.
- test_index_landing_surfaces guards the removed IDs from
reintroduction.
* feat(coord): composer parity with interactive — stop/queue/attach
Bring the coordinator one-pane UI to feature parity with the
interactive composer: in-composer Stop button replaces Send during a
turn, queue-while-busy with !!! priority + dismiss, paperclip attach
+ drag/drop/paste. The coord backend already supported all three
(lifted send/cancel/attachment handlers, emit_message_queued=True,
supports_attachments=True); this wires the UI through.
Backend:
- Wire make_dequeue_handler(coord_endpoint_config) so DELETE
/v1/api/workstreams/{ws_id}/send works for coord-kind workstreams.
- Add the matching OpenAPI EndpointSpec.
- Five new test_dequeue_* tests (success, not_found, missing msg_id,
unknown ws, scope gate) pin the URL/method/scope contract.
Frontend extraction:
- New shared modules composer_attachments.js (createAttachmentController)
and composer_queue.js (createQueueController) replace ~300 LOC of
pre-existing duplication between the interactive Pane and the coord
IIFE. Both panes now share one source of truth for the chip pipeline,
optimistic queue bubble, and busy-edge promote sweep.
Coordinator pane:
- Composer constructor adds attachments/stopBtn/queueWhileBusy/
busyPlaceholder/dragDrop options.
- setBusy now drives off SSE state_change (running/thinking/attention →
busy; idle/error → idle), with composer.setBusy unconditional and the
edge-only work (timer cleanup + queue.onIdleEdge) gated on the actual
transition.
- Cancel uses the in-composer Stop with a 2s "Force Stop" affordance +
10s safety auto-recover; the legacy header-mounted #coord-cancel-btn
is removed.
- coordCloseSession suspends SSE before close and re-establishes it on
any failure path so the UI never goes dark on a still-alive session.
- Race handling: bind() releases the queued slot server-side when the
bubble was already dismissed or promoted; rehydrate re-checks getWsId
in its .then so a stale-tab response can't clobber the new tab's
chips.
Interactive pane:
- Pane class adopts the same controllers via this.attachments /
this.queue. Pane.prototype.uploadAttachment, _renderAttachmentChip,
_swapPlaceholderChip, _removeAttachmentChip, removeAttachment,
rehydrateAttachments wrapper, addQueuedMessage, _dequeueMessage, and
_promoteQueuedMessages are all gone — the controllers own the state.
- setBusy collapses to the same shape as coord: composer.setBusy +
edge calc + queue.onIdleEdge on idle.
CSS:
- Move .msg-queued / .queued-badge / .queued-dismiss styles from
ui/static/style.css into shared_static/chat.css so both panes share
one rendering.
- Add .coord-drop-target overlay rule so the coord pane shows the
drag-and-drop affordance.
Tests pass: 160 in the impacted suites (coord endpoints + attachments
+ session routes), including 5 new dequeue tests for coord.
* fix(coord): Copilot review + lint follow-ups
Lint:
- ruff: cast(MagicMock, ...) → cast("MagicMock", ...) under
``from __future__ import annotations`` (UP037).
Copilot review (PR #438):
- composer_queue _sendDelete now invokes onAfterDequeue on success
so a bind() race-DELETE (queued bubble dismissed pre-bind or
promote sweep raced ahead) still rehydrates the caller's chip pile;
released attachment reservations no longer linger invisibly until
the next page load.
- Coord's createQueueController gains onAfterDequeue: attachments.
rehydrate(). The previous omission was a v2 review carry-over from
before coord supported attachments — now it does, so the same
contract as interactive applies.
- Both panes' send-response handler now accepts status:queued without
a queuedEl (SSE-not-yet-connected race on initial load): flips busy
so subsequent sends queue correctly. The current message keeps its
optimistic user bubble — accepted UX gap (no in-UI dismiss for
THIS message) since flipping a rendered user bubble into a queued
one mid-stream would be jarring.
- Doc updates: chat.css comment + composer_queue.js module docstring
refer to the renamed onIdleEdge() instead of the removed
promote()/promoteQueuedMessages.
Cluster nodes were OOM-killing under MCP child-process load with the
old 384M/0.5cpu budget chosen for a leaner, pre-MCP turnstone. Bump
each cluster server to 4G/4cpu and postgres to 4G/4cpu. The single-node
server, console, and channel services remain uncapped.
Four themes from a coordinator-feature shakedown:
1. Correctness fixes (return shapes / examples / behavior)
- tools_coordinator.md: drop fake skill names from spawn examples;
fix wrong kwarg ``node_id=`` → ``target_node=``.
- wait_for_workstream.json: document ``message`` + ``truncated``
per-ws fields (always enriched in the client; the JSON shape
lagged the docstring).
- cancel_workstream.json: document the conditional ``dropped``
payload — ``was_running`` always present when ``dropped`` is,
``pending_approval`` and ``queued_messages`` conditional sub-shapes.
- spawn_workstream.json: document full return shape including
``routing_strategy ∈ {rendezvous, target_node, resume}`` and
``status``.
- close_all_children.json: clarify ``skipped`` covers BOTH
hard-deleted children AND already-closed-and-evicted children
(wire shape doesn't distinguish); drop incorrect "echoed back
in response" claim — server returns ``{status, closed, failed,
skipped}``, never echoes ``reason``.
- console/server.py: comment in ``_fanout_on_children`` clarifying
that the 400 "No session" branch fires for cancel-cascade
callers and is unreachable from close_all_children (close
handler 404s instead).
- coordinator_client._utc_now_iso(): switch to bare ISO format
matching the rest of the storage row format used in the codebase.
2. Tightened the 11 longest tool descriptions (~23% cut on the
coord set). Removed ALL-CAPS emphasis, normalised em-dashes,
dropped informal phrasing. No new claims.
3. Removed static approval annotations from descriptions.
Approval is governed at runtime by the unified ``approve_tools``
body and admin-defined ``tool_policies`` (#436); static
"Auto-approved" / "Approval required" / per-action approval
tags become a stale signal. Field names (``pending_approval``)
and operational verb behaviour ("cancel unblocks pending
approvals") stay.
4. Renamed ``task_list`` coord tool → ``tasks``. The previous name
compounded the bare word ``task`` (which collides with chat-template
channels on local models — same reason ``task_agent`` carries
the suffix); the plural form sidesteps the collision and reads
more accurately, since the tool acts on the whole list rather
than a single task. Sweep covers tool JSON, Python methods (5
client methods + 2 session methods + 1 helper + 1 constant),
audit event name (``task_list.update`` → ``tasks.update``), log
tag (``task_list.corrupt_envelope`` → ``tasks.corrupt_envelope``),
frontend SSE event matcher, prompts, docs, and tests. CHANGELOG
entry added.
Plus: dropped the ENV block (Output Environment / Available
rendering / Formatting principles) from coordinator system
prompts. Coordinators orchestrate rather than render rich output
to the user, so the rendering capability matrix is not actionable
for them. Coord prompt drops ~29% (6309 → 4493 chars).
SDK regeneration via ``generate-types.py`` updates both
``openapi-console.json`` (the rename's downstream change) and
``openapi-server.json`` (PR #436 drift — its merge added
``pending_approval_detail`` + ``recent_auto_approvals`` fields to
the Python schemas but didn't regenerate the JSON artifact).
## Behavior changes (operator-visible)
- Audit event name: ``task_list.update`` → ``tasks.update``.
Audit dashboards / SIEM filters / log greps that pinned the old
prefix should update.
- SSE ``tool_result`` events now ship ``name="tasks"`` for the
scratchpad tool. The bundled coord-tree UI is updated atomically;
external consumers reading SSE events by tool name need to update.
- Existing task envelopes in production storage have ``+00:00``
timestamps from the old ``_utc_now_iso``. New writes are bare;
old rows are not backfilled. Within an envelope you may briefly
see mixed formats until each row is re-touched. No code path
string-compares timestamps within an envelope, so this is
cosmetic.
## Validation
- ``ruff check`` + ``ruff format --check`` clean
- ``mypy turnstone/`` clean (175 source files)
- ``pytest -m "not live"`` — 4679 passed, 3 deselected
* refactor(core): unify approve_tools across kinds + judge visibility + perf
Lift WebUI.approve_tools to SessionUIBase so both interactive and
coordinator workstreams run the same body. The shared body now owns
tool-policy gating, per-tool auto-approve, blanket carve-out for
__budget_override__, activity tagging, heuristic-verdict persistence,
and the approve_request/approval_event blocking pattern. Subclass
hooks layer kind-specific surfaces on top.
This closes the drift the LLM-judge audit flagged on coord — the
judge (heuristic + LLM tier) now sees actual tool args for every
coord tool call instead of empty func_args. spawn_batch projects
the full children list so a malicious mid-batch entry is no longer
hidden.
= Unification core =
- SessionUIBase.approve_tools: lifted body covering policy / per-tool
auto-approve / blanket / activity tagging / heuristic-verdict
persistence / approval gate
- _APPROVAL_WAIT_TIMEOUT class constant + _record_judge_metric hook
- WebUI.approve_tools deleted; _record_judge_metric override fires
per-node MetricsCollector.record_judge_verdict
- ConsoleCoordinatorUI.approve_tools deleted; _record_judge_metric
+ on_intent_verdict overrides fire ConsoleMetrics.record_judge_verdict
- ConsoleMetrics.record_judge_verdict + turnstone_judge_verdicts_total
in /metrics text output (cluster PromQL rolls coord+interactive up
uniformly)
- _console_metrics class attribute wired in console lifespan
- Frontend: coord SSE event tools_auto_approved -> tool_info for parity
= Judge args visibility =
- _evaluate_intent populates func_args for all coord tools that hit
approval (spawn_workstream / spawn_batch / send_to_workstream /
close_workstream / close_all_children / cancel_workstream /
delete_workstream / task_list)
- spawn_batch projects every child's skill / initial_message[:200] /
target_node so the judge sees the full fan-out (was first child only)
- fire_judge_verdict_metric helper collapses 4 sites of identical
record_judge_verdict shape across WebUI + ConsoleCoordinatorUI
= Hardening =
- __budget_override__ carve-out reads from pre-filter items list, not
post-filter pending; policy block skips matching the synthetic
name entirely so a wildcard `*: allow` cannot strip the override
before the gate sees it
- _persist_intent_verdict default_tier parameter so heuristic + llm
paths share the storage write helper
= Performance =
- TTL cache on list_tool_policies in turnstone/core/policy.py
(60s, keyed by org_id, lock-free hits)
- Storage-layer invalidation: create/update/delete_tool_policy on
both SQLite and PostgreSQL backends call invalidate_policy_cache
(covers admin-API path + direct test fixtures + any future caller)
- Admin-API handlers also call invalidate_policy_cache as
defense-in-depth
- storage.create_intent_verdicts_bulk on both backends: one
multi-row INSERT + one commit instead of N round-trips. approve_tools
switches to the bulk path so a fan-out turn no longer pays N x commit
before the approval prompt enqueues
- _persist_intent_verdicts_bulk helper on SessionUIBase
= Test coverage =
- tests/test_coord_ui_approve_tools.py (NEW, 17 cases): inheritance
regression, tool-policy deny/allow/mixed on coord, heuristic verdict
persistence (bulk path), activity tagging on auto-approve and pending,
judge_pending dynamic flag (true + false), event-name parity,
per-tool auto-approve, __budget_override__ carve-out under blanket
+ wildcard policy, _record_judge_metric wired/unwired, on_intent_verdict
llm-tier metric
- tests/test_console_metrics.py: 3 cases for the new
record_judge_verdict counter
- tests/test_judge_storage.py: 3 cases for create_intent_verdicts_bulk
- tests/test_coordinator_tools.py: 3 cases pinning the spawn_batch
full-children projection (truncation, mid-batch visibility, empty
defensive)
- tests/conftest.py: autouse _clear_policy_cache fixture so the
process-level cache doesn't leak between tests with distinct storage
instances
= Drift fixes (review feedback) =
- Refresh stale "no-op on coord" comments now that coord overrides
the hook
- WebUI.on_plan_review timeout uses self._APPROVAL_WAIT_TIMEOUT
instead of literal 3600
- Drop redundant bool() wrapper around any() in judge_pending
- Rephrase broken docstring grammar in _coord_spawn_metrics
- Hoist redundant get_storage import out of approve_tools per-item loop
(folded into _persist_intent_verdicts_bulk helper)
= Validation =
- pytest -m "not live": 4679 passed, 3 deselected
- ruff check + ruff format: clean
- mypy: no issues in 175 source files
* fix(approval): apply Copilot feedback on PR #436
- Policy-cache invalidation now drops both the org-scoped slot AND the
default ``""`` slot on ``create_tool_policy`` for both SQLite and
PostgreSQL backends. ``list_tool_policies("")`` returns rows from
every org_id, and the production evaluators (SessionUIBase.approve_tools
/ cli.py) read with the default ``org_id=""``, so an org-scoped insert
that only invalidated its own slot would leave the default cache slot
stale until the TTL window expired.
- Cap ``reason`` to 200 chars in ``_evaluate_intent`` for ``close_workstream``
and ``close_all_children`` — both fields are LLM/user-provided and the
preparer doesn't size-limit them, so an unbounded reason could bloat
the persisted verdict row's func_args. Matches the cap applied to other
free-form coord tool fields (initial_message, message, title).
- Refresh ``_PolicyCache`` docstring: it claimed lock-free reads on
cache hit but ``get()`` always acquires ``self._lock``. Updated to
reflect that the lock is held briefly to copy the policies reference.
Validation: targeted suite 201/201, ruff + mypy clean.
Follow-up to #434. That PR unified the chat-message primitive on .msg
and noted that the parallel .ts-composer prefix on the shared composer
widget was still in place; this drops it so the widget sits in the
shared/* vocabulary the same way .msg does.
Mechanical 1:1 rename (`ts-composer` -> `composer`) across:
shared_static/chat.css — 48 selectors
shared_static/composer.js — 19 className strings
ui/static/style.css — 11 per-node UI overrides
ui/static/app.js — 7 chip queries / className strings
Pre-rename collision check confirmed clean: the only `composer`-substring
matches in the codebase were IDs (#coord-composer-mount, #coord-composer-
panel, #coord-composer-503, #home-coord-composer-mount — IDs are a
different namespace from classes) and the unrelated console
.home-composer-banner / .home-composer-error pair (different prefix).
CSS specificity audit (scripts/css_specificity_audit.py): 26 findings on
origin/main, 26 on this branch — no new cascade flips.
Tests: 189 affected tests pass (test_app_js, test_webui_content,
test_webui_auto_approve_visibility, test_html, test_web_helpers,
test_coordinator_adapter, test_coordinator_client).
Manual visual verification of composer surfaces (textarea, send button,
stop button, attach button + file picker, chip pills + remove buttons,
options panel toggle, paste-image and drag/drop attach paths, stacked
layout used by creation forms) recommended before merge.
* refactor(ui): drop legacy .ts-msg* dual-classing, chat surfaces share .msg primitive
Third and final follow-up after #431 stripped the data-design="v1"
gate. This drops the parallel .ts-msg* family that had been kept as a
transitional bridge during the gated rollout. Per-node UI now renders
pure .msg classes (previously dual-classed as
"ts-msg ts-msg--user msg user"), matching the coordinator chat view
which already used pure .msg*.
chat.css: deleted the ~210-line legacy .ts-msg* rule block (Messages +
floating action toolbar + mobile + reduced-motion sections); renamed
.ts-msg.ts-approval--inline to .msg.ts-approval--inline; restored the
streaming-markdown rationale (white-space: normal intent + partial-fence
behavior + .msg-user-text path) on .msg-body that previously lived on
the deleted .ts-msg-body, with white-space: normal now declared
explicitly so a future "simplification" can't silently break streaming.
ui/static/app.js: dropped the ts-msg* half of every dual-class string
and updated querySelector callsites (.ts-msg--user -> .msg.user,
.ts-msg--assistant -> .msg.assistant).
ui/static/style.css: renamed all .ts-msg--* selectors to .msg.*; removed
two now-dead override rules (.ts-msg.msg:not(.tool) and
.ts-msg-body.msg-body font-family overrides) that existed solely to
unwind the legacy .ts-msg font-mono default that's now gone.
The .msg.ts-approval--inline selector intentionally keeps the .msg
qualifier (rather than bare .ts-approval--inline) so its (0,2,0)
specificity ties with .ts-approval.approved/.denied/.error and the
later-cascade rule wins; without the qualifier those state classes
would suddenly flip the inline-approval card colour based on state.
Composer rename (.ts-composer-* -> .composer-*) deferred to a follow-up
PR; ~80 occurrences across composer.js + chat.css would have obscured
this verification.
Tests: 538 affected tests pass (test_app_js, test_webui_content,
test_webui_auto_approve_visibility, test_web_helpers, test_html,
test_auth, test_console, test_api_versioning, test_coordinator_*).
Visual verification (message cards, hover toolbar, approval/denial/error
cards in light + dark themes) recommended before merge.
* docs(ui): clarify .msg.reasoning emission comment per Copilot review
The previous wording — ".reasoning as a bare role class is no longer
emitted" — implied .reasoning is never emitted, but the new className
is "msg reasoning" so .reasoning IS emitted, just always alongside .msg.
Reword to make the actual invariant (never on its own) explicit.
* fix(css): two cascade-flip bugs found by specificity audit
PR #431 stripped [data-design="v1"] from ~400 rules, dropping each by a
specificity tier; two cascade flips (#header outranking .appbar, #header h1
outranking .appbar-title) were caught visually during that PR's review and
fixed by renaming id="header" → id="ui-header" on the per-node UI page.
This is the audit follow-up; it found two more:
- textarea.skill-content-area (was .skill-content-area) — bumped to (0,1,1)
so the rule ties with `.admin-modal textarea` (0,1,1) and wins on source
order. Without the bump, min-height: 220px was clobbered to 40px by the
modal default and the spec-content textarea rendered short. The three
!important markers (font-family/size/line-height) are now redundant
against the modal's font: inherit shorthand and are dropped.
- h3.skill-spec-heading — removed `font-size: inherit;`. The author wrote
it to "reset UA defaults" but it locked font-size to the parent's
(~14-16px) at (0,1,1), silently overriding `.skill-spec-heading`'s 10px
at (0,1,0). The bare class already beats UA `h3` on specificity (class >
tag), so no font-size reset was needed; the `margin-block: 0` line stays
because the bare class's `margin: 14px 0 6px` shorthand may not reset
the UA's logical margin-block-start/end on every engine.
Adds scripts/css_specificity_audit.py — the audit tool. It parses every
CSS file referenced from the project's three HTML entry points, computes
selector specificity (incl. :not/:is/:has math, attribute selectors, and
!important), and flags every place an unscoped legacy rule could outrank
a bare-class designed primitive. Honours per-page stylesheet manifests,
state-pseudo subset gating (a `:hover` rule overriding a resting-state
base rule is intentional, not a flip), and shorthand→longhand expansion
for font/padding/margin/border/background. Triage of remaining findings
(26 id-tier in default mode, 74 total at --all-tiers) confirmed all are
intentional designer overrides — id-scoped buttons, BEM modifier classes,
contextual ancestor selectors, last-child margin reset, [hidden] toggle.
* fix(css-audit): correct two cascade-resolution bugs flagged by Copilot
1. _parse_declarations dict insertion order didn't update on overwrite, so
a sequence like `font-size: 13px; font: inherit; font-size: 12px;` would
iterate as (font-size=12px, font=inherit) and the shorthand expansion
then clobbered font-size back to `inherit` — wrong. Delete-then-insert
on overwrite so the last occurrence lands at the dict's tail and the
shorthand expansion sees the real source order.
2. The cascade-winner tie-break used `rule.line_no` only, ignoring the
stylesheet load order. A rule at line 1000 of `base.css` looked
"later" than a rule at line 50 of `style.css`, even though the page
loads `base.css` BEFORE `style.css`. Sort by `(file_index, line_no)`
keyed off the element's per-page stylesheet manifest instead.
* fix(ui): preserve approval pill when tool errors
When an approved (or auto-approved) tool subsequently failed during
execution, both `replayHistory` and `appendToolOutput` located the
existing `.ts-approval-badge` and overwrote its className + textContent
with the `--error` variant — losing the record that the user had
approved the call.
Append a separate `--error` pill as a sibling of the existing approval
pill instead. The `.ts-approval` parent is `flex-direction: column` with
a 6px gap, so the two pills stack vertically and read as a small
status timeline ("you approved this, then it errored"). Idempotency
guard via `querySelector(".ts-approval-badge--error")` so duplicate
fires don't stack badges.
CSS classes are unchanged (the `--error` modifier already exists in
both per-node and shared chat stylesheets).
Adds a static-string guard in `tests/test_app_js.py` that pins both
call sites and forbids the mutate-in-place anti-pattern via a regex
that pairs a queried `.ts-approval-badge` handle with an `--error`
className overwrite.
Deferred from #431.
* refactor(ui): extract appendToolErrorBadge helper, broaden test guard
Address Copilot feedback on #432:
- Extract the duplicated 5-line error-pill construction into a single
module-level `appendToolErrorBadge(blockEl)` helper next to the
other approval-related helpers (`buildToolDiv`, `renderVerdictBadge`,
`toggleVerdictDetail`). Reduces drift risk on ARIA / class / text
string between the two call sites.
- Loosen the affirmative test check from a literal substring keyed on
the local variable name to a regex matching any
`querySelector(".ts-approval-badge--error")` lookup, in either
quote style, in either guard idiom (`if (!q) {...}` at a call site
or `if (q) return;` inside the helper). A future refactor that
preserves behaviour shouldn't trip CI on cosmetics.
- Broaden the anti-pattern regex to accept single quotes and to
catch the `classList.add("ts-approval-badge--error")` form on a
queried badge handle, not only `className = "..."`.
* refactor(css): unify design system, eliminate data-design="v1" gating
Strip the [data-design="v1"] attribute that was wrapping every DS rule
since #389 and never came back out. Result: every styled element on
coord/ui pages had two CSS rules (default + v1-gated), reviewers
couldn't tell which one rendered, and the bundle shipped duplicates.
Changes:
* Strip [data-design="v1"] prefix from ~400 gated rules. Remove the
attribute from coordinator/index.html, ui/index.html, preview.html.
* Merge shared_static/design/* into pre-v1 sheets:
tokens + typography → base.css (:root, dark default)
appbar + panel + buttons + pills + field → ui-base.css
message primitives → chat.css
sidebar + approval-dock → console/static/coordinator/coordinator.css
(new file linked from coord only — admin no longer ships ~9KB of
coordinator-only chrome on first load).
* Delete preview.html + 5 preview-only orphan stylesheets (topbar /
stats / feed / fleet-grid / live-feed). Drop the empty
shared_static/design/ directory.
* Drop unused primitives the merge dragged in: .pill / .k-badge /
.chip / .field / .t-* utilities / .side-item / .shell. Drop unused
tokens (--accent-c / --accent-l / --row-h / --density / --gap /
--font-display alias). Find-replace var(--font-display) →
var(--font-ui) across 6 files (103 sites).
* Standardize on the DS font stack: Inter body, JetBrains Mono code.
Admin's body shifts from IBM Plex Mono → Inter via the alias rename.
* Re-tune legacy --bg/--bg-surface/--bg-highlight/--bg-elevated from
steel-blue to neutral charcoal so admin and v1 pages share one
palette. Drop the cyan radial-gradient overlay on body that added a
blue tint to the formerly-blue bg.
* Polish:
- .msg.tool / .ts-msg--tool / .ts-approval / inline-approval all use
--cyan instead of amber, removing the user/tool colour collision.
- .msg-action-btn reverts to icon-button styling (transparent,
28x24) after the merge gave it text-button chrome that dwarfed the
13-14px icon glyphs inside.
- Light-mode composer contrast: flip .ts-composer surface roles
(wrapper recessed, textarea elevated) so the textarea reads
against its container; bump .dashboard-composer textarea
border-bottom + options panel surface so they're visible on white.
- .pane-messages padding 20px -> 16px 12px and gap 14px -> 0 (the
flexbox gap was stacking with .ts-msg margin-bottom for ~18px
inter-card spacing); .ts-msg/.msg margin-bottom 8px -> 4px.
- Restore WCAG 2.5.5 36x36 touch target on .msg-action-btn.
- Rename ui's <div id="header"> to id="ui-header" so the legacy
#header chrome no longer outranks the .appbar primitive on the
per-node page (3 getElementById calls in app.js updated).
- Drop chat.css link from coord (coord renders pure DS classes; the
composer.js consumer of chat.css is on admin + ui only).
Verified: ruff + mypy clean (175 files); 4653 non-live tests pass;
zero data-design / --font-display / shared_static/design hits remain.
Net source change: -1885 lines (924 added, 2809 deleted across 28 files).
* build: include coordinator.css in wheel; drop dead design/ glob
The previous commit added turnstone/console/static/coordinator/coordinator.css
(coord-only chrome moved out of console/static/style.css) but didn't
update the [tool.hatch.build.targets.wheel] include list, so CI's
wheel-completeness check failed.
Also drop the now-stale 'turnstone/shared_static/design/**/*' glob —
that directory was deleted in the same v1-elimination commit.
Verified locally: replicating the CI step's source-vs-wheel diff
returns MISSING: none.
* fix(css): address Copilot review feedback on PR #431
* coordinator/index.html — comment now correctly points to the moved
.sidebar rules at console/static/coordinator/coordinator.css (was
console/static/style.css before the perf-1 split-out).
* ui/static/index.html — restore <h1 class="appbar-title">; the cascade
conflict that motivated the h1→div change is gone now that the
wrapper id was renamed away from #header (the legacy #header h1 rule
no longer matches). Page semantics + accessibility regain the
top-level heading.
* governance.js — drop the inline font-family:var(--font-ui) on the
config-key <code> elements; let them inherit the global mono default
from base.css. The inline style was an artifact of the
--font-display → --font-ui find-replace; the original Outfit was
already odd on a <code> tag.
* ui-base.css — typography-helpers comment said "Body text still
inherits var(--font-mono) at 13px" but base.css now sets var(--font-ui)
at 14px. Reword to match current defaults.
* fix(approve): visibility for child tool calls bypassing operator gate
When a coord LLM spawns a child with `skill="X"`, the skill template's
`allowed_tools` JSON list silently populates the child UI's
`auto_approve_tools` set. Tool calls whose names are in that set
short-circuit the approval gate without prompting the operator —
matching the user-reported bug "tool calls of children occasionally
getting approved instead of waiting for approve/deny".
The auto-approve paths themselves are unchanged (Option C — visibility
only). Surfaces:
- Per-item annotations: each pending tool gets `auto_approved=True` +
`auto_approve_reason` ("skill" / "always" / "policy" / "blanket" /
"auto_approve_tools") at the four gate-bypass paths.
- Per-ws ring buffer (cap 10) of recent bypasses, exposed via
`/dashboard` and the cluster live-bulk projection so the coord-
tree row can render an "auto-approved by ..." pill.
- `tool.auto_approved` audit row per `approve_tools` call —
forensic durability beyond the in-memory ring buffer.
- Per-ws WebUI page: inline "auto: <reason>" badge next to each
tool name, so an operator who clicks through from the coord tree
to the child's page sees the same bypass signal.
Persistence across UI rebuilds:
- The ring buffer is in-memory only; a saved-workstream rehydrate /
coord→node click-through / process restart all build a fresh UI.
`replay_recent_auto_approvals_from_audit` runs at the end of
`SessionUIBase.__init__` and re-seeds the buffer from recent
`tool.auto_approved` audit rows scoped to this ws_id.
- Adds `resource_id` filter to `list_audit_events` (protocol +
SQLite + Postgres) so the replay is a single indexed query.
Source provenance:
- `_auto_approve_tools_source: dict[str, str]` per UI tracks which
writer added each tool name to `auto_approve_tools` ("skill" at
skill-template setup time, "always" on Approve+Always click).
Lets the dashboard pill distinguish a skill-driven bypass from
an explicit operator-Always click — those are very different
signals that previously rendered the same.
Magic-string drift mitigation:
- `AutoApproveReason` constants in `core/session_ui_base.py` lift
the five reason strings into a single source of truth.
- `KNOWN_AUTO_APPROVE_REASONS` JS constant + validator render
unknown reasons as "unknown" with a console.warn instead of
rendering raw (a typo would otherwise silently desync wire ↔
pill).
Recording-leak fixes (q-2 from review):
- Policy `allow` partial-resolve now records the policy-tagged
items at two previously-leaking branches: the early-return-on-
deny path and the still_pending-non-empty fall-through to the
prompt path.
Other review fixes:
- Heuristic verdict surfaces consistently as `heuristic_verdict`
in both `_serialize_approval_items` and the dashboard
serializer (was inconsistent: one emitted `verdict`, the other
`heuristic_verdict`). app.js updated to read either key for
mid-deploy compatibility.
- `_tag_auto_approved` helper on SessionUIBase replaces the
verbatim tag loops previously copy-pasted across WebUI and
ConsoleCoordinatorUI.
* fix(approve): apply Copilot review feedback on PR #430
- coordinator_ui: use ``approval_label or func_name`` for the
``auto_approve_tools`` subset check, matching WebUI. Pre-fix
an "Approve + Always" entry whose approval_label differs from
func_name (skill__name, mcp_resource__uri) wouldn't match on
the coord page and the operator would be re-prompted.
- _parse_audit_timestamp: treat naive ISO strings as UTC. Audit
rows are written via ``datetime.now(UTC).strftime(...)`` with
no timezone marker; ``datetime.fromisoformat`` returns a naive
datetime, and ``.timestamp()`` on a naive datetime interprets
it in the server's local timezone — wrong on any non-UTC
server. Stamp UTC explicitly before converting.
- server.py: drop the dead ``pending = []`` after the blanket
tag — the function returns inside the same block without
reading ``pending`` again.
- _protocol.py: fix docstring reference from
``_replay_recent_auto_approvals`` to
``replay_recent_auto_approvals_from_audit`` (the actual
method name).
* fix(coord): tree UI not updating when LLM deletes workstream
The coord LLM's `delete_workstream` tool wiped the storage row but
fired no SSE event, so a long-lived dashboard tab kept the deleted
child visible (with its last-known idle/closed state) until a full
reload. A coordinator that spawns→completes→deletes children would
leave an ever-growing tree.
Fix: add `SessionManager.delete()` that drops the in-memory slot if
present and emits `ws_closed` with `reason="deleted"` (mirrors
`close()`'s shape). Wire `delete_workstream_endpoint` to call it
after the storage delete succeeds, snapshotting the workstream's
name into the event payload before the row is wiped. The cluster
collector → coord adapter chain re-emits as `child_ws_closed`; the
browser's existing `handleChildClosed` already keys on
`reason === "deleted"` to mark the row, so no JS changes needed.
Event emit is best-effort — a fan-out failure logs a warning but
doesn't roll back the storage delete (the row is already gone).
* fix(coord): apply Copilot review feedback on PR #429
- server.py: clarify that ``name`` is forwarded to mgr.delete only
(not into the audit detail) — comment previously claimed both.
- test_session_manager.py: extract ``mgr.delete(ws_id)`` to a local
before asserting (CodeQL: no side-effecting calls inside ``assert``,
which would be stripped under ``python -O``).
- test_workstream_endpoints.py: docstring said "Yield" but the
fixture ``return``s; switch to "Return".
Copilot caught a doc/code mismatch from the q-2 cleanup: the
docstring still claimed `closed` / `deleted` / `denied` all return a
sentinel, but the `deleted` branch was dropped (hard deletes cascade
rows out of storage so the state is unreachable). Update the
docstring to align with `_wait_message_for`'s actual behaviour —
`deleted` falls into the same null-message shape as a still-running
entry.
Each per-ws snapshot now carries `message` + `truncated` so the
coordinator LLM doesn't need a follow-up `inspect_workstream`
round-trip per child to read what came back. idle/error states
return the last assistant turn (capped at 6 KiB UTF-8 bytes,
truncated from the end); closed/denied return a sentinel; running
children carry null. Storage reads for idle/error parallelize
across an 8-worker thread pool so a 32-child fan-out lands in 4
batches instead of 32 sequential round-trips.
Creating or opening a workstream from the dashboard left the chat
UI blank until the operator refreshed: switchTab early-returned at
``if (!pane) return;`` because getFocusedPane was null on a fresh-
loaded page that had no workstreams. The freshly-created ws was
added to the workstreams dict and the dashboard was hidden, but no
pane was bootstrapped, no SSE connected, and the chat area sat
empty until refresh — at which point initWorkstreams saw the
populated list and bootstrapped the pane via the existing
"if (!Object.keys(panes).length)" branch.
switchTab now mirrors that bootstrap when no focused pane exists:
createPane + splitRoot leaf + setFocusedPane + renderLayout. The
rest of switchTab (disconnectSSE / reset / connectSSE) runs as
before — no-ops on the just-constructed pane up to the connectSSE
call which is exactly what we want.
Subsequent creations on the same node already worked because the
first create populated panes and switchTab found a focused one.
Static smoke test in tests/test_app_js.py guards against the
early-return regressing.
* feat(renderer): progressive mermaid rendering during streaming
Mermaid diagrams used to materialize all-at-once at stream_end via
streamingRenderFinalize, which felt laggy on long responses with
multiple diagrams. Now closed mermaid fences render progressively
as each fence completes during streaming.
The blocker was streamingRender's wholesale `el.innerHTML = html`
on every rAF tick, which destroys any rendered SVG nodes — without
caching, calling postRenderMermaid per tick would re-trigger an
async mermaid.render every time, thrashing the renderer.
Added a source-keyed SVG cache (_mermaidSvgCache, FIFO-bounded at
64 entries):
- Cache hit on identical source: synchronous innerHTML swap, no
loading flash, no async work. Mermaid is deterministic for a
given init, so identical source ⇒ identical SVG, safe to reuse.
- Cache miss: queue async render, populate cache on success.
- Errored sources cached separately (_mermaidErrorCache) so a
syntactically-broken diagram doesn't re-thrash mermaid on every
tick. The user can fix the diagram and the new source string
misses the cache, triggering a fresh render.
_streamingRenderApply now calls postRenderMermaid after the
innerHTML replace. Per-stream cost: each unique mermaid source
pays mermaid.render once, then synchronous cache hits for every
subsequent rAF tick. hljs syntax highlighting stays deferred to
streamingRenderFinalize (it's a separate pass and benefits less
from progressive rendering — code blocks tend to be short and
already legible without color).
Tests: built a richer Node-driven harness with a fake DOM that
tracks attributes / classList / parent chain / replaceWith, plus
a stubbed mermaid.render with a call counter. 5 new tests cover:
cache-hit skips render, distinct sources render independently,
errors cache to avoid thrash, FIFO eviction at cap, and a static
guard that _streamingRenderApply actually calls postRenderMermaid.
* fix(renderer): apply Copilot feedback on PR #426
Six review items, all real:
1. _cacheMermaidEntry evicted on overwrite — overwriting an
existing source unnecessarily dropped the oldest entry.
Now: only evict when inserting a new key.
2. _initMermaid didn't clear caches — a theme change via
reRenderAllMermaid (which calls _initMermaid) would serve
stale SVG keyed by source-only, since rendered output
depends on themeVariables. Now clears both caches on
(re-)init.
3. bindFunctions never re-applied on cache hits — mermaid's
bindFunctions attaches link/click handlers to each rendered
SVG instance. Pre-fix, only the first render got bindings;
subsequent cache hits via raw innerHTML left the SVG inert.
Cache value is now {svg, bindFunctions}; cache hits go
through _applyMermaidSvg which re-applies bindings on each
new container instance.
4. Truthiness checks on cache lookups — empty-string SVG / error
would have masqueraded as a miss. Switched to cache.has()
(and then .get) so intent is explicit.
5. Concurrent mermaid.render — postRenderMermaid now fires on
every streaming rAF tick, so multiple ticks could overlap
while earlier render Promises pend. mermaid.render uses
module-level state internally — concurrent calls clobber it.
Two layers of serialization fix this:
- _mermaidPending: per-source. While a render is in flight
for source X, additional containers asking for X are queued
and the single render result fans out to all pending
containers when it lands.
- _mermaidRenderChain: across-source. Promises chain so
mermaid.render runs at most one at a time globally.
- Detached containers (no longer in the DOM by the time the
render completes) are skipped via isConnected guard —
wholesale innerHTML replace during streaming detaches them
and a later tick is already taking care of the live one.
6. Test brittleness — _streamingRenderApply guard used
body.index("\\n}\\n", start) which would stop at the first
inner-block closing brace inside the function. Switched to a
bounded-window string search (Copilot's suggestion).
Three new tests added: overwrite doesn't evict; _initMermaid
clears caches; cache hit re-applies bindFunctions. Existing tests
updated for the new {svg, bindFunctions} cache shape and the
async serialization (drain via setTimeout hops instead of bare
microtask resolves).
Test harness fix: fake DOM elements now have an isConnected
getter derived from the parent chain, so the new guard
exercises correctly under test.
* fix(renderer): handle LaTeX-style \(...\) and \[...\] math delimiters
The browser renderer at turnstone/shared_static/renderer.js only
recognized TeX-style $...$ / $$...$$ delimiters. Most modern LLMs
(GPT-5 / o-series, Claude with reasoning effort) emit LaTeX-style
\(...\) for inline math and \[...\] for display by default — those
slipped through as raw text in the coord + interactive WebUIs,
making KaTeX appear "broken when nested inside a markdown block"
(actually broken everywhere, the surrounding markdown just made
the failure noticeable).
Added a second pass for each delimiter style alongside the
existing $...$ / $$...$$ patterns. Both styles now feed the same
mathBlocks / inlineMaths placeholder pipeline so all the existing
nested-block handling (lists, blockquotes, tables, bold, headings,
details, post-render KaTeX markup) Just Works.
Edge cases verified by the new test_renderer_js.py harness:
- \(...\) inside inline code stays literal
- \(...\) inside fenced code blocks stays literal
- Solo \[ with no closing \] doesn't trigger spurious math
- Markdown links [text](url) untouched (regex uses \[ \], not [ ])
- Mixed TeX + LaTeX delimiters in one message both render
The harness drives renderer.js through Node via vm.runInThisContext
with stubbed document/katex globals — first JS-side regression
guard for the renderer; previously it had no test coverage at all.
* fix(renderer): apply Copilot feedback on PR #425
Three review items from Copilot:
1. Display-math sentinel could leak through inline-code spans.
The original ordering ran $$...$$ / \[...\] extraction BEFORE
inline code, so a backtick span around math (e.g. `$$x$$` or
`\[x\]`) had its delimiters consumed by the math regex and
replaced with \x00MB…\x00. Inline code then captured the
sentinel; restore order put MB after IC, leaving the null-byte
placeholder visible inside the rendered <code>. Reorder: inline
code first, then display math, then inline math. Code spans
now seal their content before any math regex sees it. The
reverse edge case (math containing backticks, e.g. \verb|`x`|)
is much rarer and KaTeX rejects \verb anyway.
2. Inline LaTeX-style \(...\) regex used [\s\S]+? which allowed
newlines, so an unterminated \( on one line would eat the
next paragraph until it found a closing \). Aligned with the
existing $...$ behavior by switching to [^\n]+? — display
math (\[...\] / $$...$$) stays multi-line by design.
3. tests/test_renderer_js.py was guarded with a node-availability
skip, but CI's test + test-postgres jobs didn't explicitly
install Node, so the suite would have silently no-op'd if the
runner image dropped Node. Added actions/setup-node@v5 to
both jobs.
Four new regression tests cover the leak (both delimiter styles
inside backticks must stay literal) and the cross-paragraph span
(both \(...\) and $...$ must not eat newlines).
Bug: LLM judge verdicts stayed stuck on heuristic-only render.
Root cause: per-row poller called scheduleLiveFetch which
short-circuits on non-visible rows — invalidate cleared the
cache, no fetch fired, the row kept rendering its last-cached
heuristic indefinitely. The 12s attempt cap also gave up before
slow LLM judges (>15s with reasoning effort) could land.
Replaced with a single global poller _maybeStartJudgePoll /
_judgePollTick:
- Walks the full childrenState (not just visible rows)
- Bypasses scheduleLiveFetch's visibility + TTL gates by
adding to pendingLiveIds directly + flushing
- One bulk request covers every pending row per tick
- Self-terminates when every verdict lands or 90s elapses
(operator can hit Refresh to retry on a failed judge)
- 90s cap is wall-clock, not attempt count, so an LLM that
takes 60s no longer prematurely gives up
Copilot round-2 feedback:
- _proxy_sse with use_service_auth=True silently fell back to
empty headers when proxy_token_mgr was None, producing a
retry-storm 401/403 loop. Fail fast with a 503 + clear log
so the misconfig surfaces immediately.
- Mobile <700px CSS comment claimed buttons "stretch to full
row width" but the rule keeps flex-direction: row with
flex: 1 on each, giving 50/50 side-by-side. Updated the
comment to match the deliberate side-by-side layout
(stacking would push the action row below preview/disclosure
on tall envelopes; 50/50 keeps both verbs reachable).
#422's legacy URL adapter removal deleted the body-keyed
/v1/api/route/{verb} endpoints (with ws_id in JSON body) but
turnstone/console/coordinator_client.py still pointed at them.
The coord LLM's close_workstream / close_all_children tools
404'd; send / approve / cancel were equally broken though
exercised less often.
_ROUTE_PATHS now uses {ws_id}-templated path-keyed forms:
send → /v1/api/route/workstreams/{ws_id}/send
approve → /v1/api/route/workstreams/{ws_id}/approve
cancel → /v1/api/route/workstreams/{ws_id}/cancel
close → /v1/api/route/workstreams/{ws_id}/close
_post() interpolates {ws_id} at call time when the template has
the slot; body-keyed paths (delete, close_all_children) still
work via the same code path. Each affected caller (send, approve,
cancel, close_workstream, close_all_children) was updated to pass
ws_id as the kwarg and drop ws_id from the body.
Added test_route_paths_match_actual_console_mounts: walks the
real Starlette app's routes and asserts every _ROUTE_PATHS entry
corresponds to an actually-mounted route. Catches the next URL
unification drift before runtime. Updated the existing literal
assertions + path-checking tests for the new shape.
Pre-existing bug surfaced while testing inline-child-approvals.
The interactive WebUI's app.js opens an EventSource against
/v1/api/events/global on load (cluster-wide tab indicators,
ws_state for the dashboard). When loaded via the console proxy
at /node/{node_id}/, the JS shim rewrites that to
/node/node-X/v1/api/events/global and the proxy forwards using
the user's re-minted JWT.
Upstream global_events_sse requires `service` scope by design
— the stream carries cross-tenant cluster inventory, intended
for the cluster collector, not browsers. End-user JWTs don't
carry service scope, so every proxied call returned 403, the
browser auto-retried with exponential backoff, and the console
log filled with proxy.sse.non_200 warnings.
_proxy_sse gains a use_service_auth flag. proxy_api flips it on
for events/global only, swapping the user JWT for the console's
proxy_token_mgr bearer token. Per-ws events stay on user auth
(tenant filtering on the upstream still requires user identity).
The upstream-side privacy posture is unchanged — the data on
events/global is the same cluster-wide inventory the console's
own /v1/api/cluster/events endpoint already serves to any
read-scoped caller under the trusted-team posture. The console's
AuthMiddleware on /node/{node_id}/v1/api/ remains the gate that
decides who can use the proxy at all.
The console's node-API passthrough at /node/{node_id}/v1/api/{path}
detected SSE only on the bare events / events/global paths. After
#422 removed the legacy /v1/api/events?ws_id= shape and moved
per-workstream SSE under /v1/api/workstreams/{ws_id}/events, the
proxy never got updated to match the new path — per-ws events
fell through to the regular GET branch, the upstream returned a
text/event-stream payload that the regular GET response couldn't
hold open, and Firefox surfaced the failure as "can't establish a
connection to the server".
Extend the SSE detection to also match
``workstreams/{ws_id}/events``. Pre-existing bug surfaced while
testing inline-child-approvals (operator clicks through from the
coord tree to the per-child interactive WebUI) but affects every
caller hitting a node's per-ws events stream via the console
proxy.
Two new tests in TestConsoleProxy: per-ws events route to
_proxy_sse with the correct upstream path; existing
events/global routing still works.
Previously the 409 stale-call_id branch in submitChildApproval
re-enabled both buttons synchronously before kicking off the
urgent live-bulk refresh. That opened a window where rapid clicks
on an already-resolved approval (or a row whose call_id had
rolled) each re-armed the click handler, fired another POST, and
collected another 409. Operators rage-clicking saw a network 409
storm and a stack of warning toasts.
Keep the buttons disabled in the 409 path. The row is about to be
re-rendered wholesale via the urgent refresh — the disabled DOM
gets dropped along with it. If the row's approval truly resolved,
the new render has no buttons. If a new round started, the new
render has fresh enabled buttons. Either way the operator-facing
signal IS the row updating, not the toast.
Drop the toast.warn (noisy on every rapid-click race) in favour
of a single console.warn for diagnostics.
If the urgent refresh fails entirely, the buttons stay disabled
on that row — but the operator can hit the Refresh button on the
children panel to force a full reload. Acceptable degraded state
vs the previous 409 loop.
The coord's _coord_events_replay re-yielded _pending_approval on
connect but not the cached _llm_verdicts entries. A tab refreshing
mid-approval saw the approve_request prompt without the judge chip
because intent_verdict is a one-shot SSE event with no
late-subscriber push — the chip would only ever land if the operator
re-invoked the tool call.
Mirrored the interactive path at turnstone/server.py:875-878:
after re-injecting the pending_approval prompt, walk
ui._llm_verdicts under _ws_lock and yield each cached verdict as
an intent_verdict event. Pre-existing bug surfaced during the
inline-child-approvals work but the coord-self dock UX was always
affected on reconnect — not introduced by this PR.
Two new tests: cached verdicts replay after pending_approval; stale
verdicts from a prior round don't replay when no approval is pending.
Two bugs reported from local repro on PR #424:
1. Approve/Deny buttons return HTTP 404 on every click. The new
approveWorkstream helper hit /v1/api/workstreams/{ws_id}/approve
regardless of target — that path is only mounted for coord
workstreams (which live on the console process). Child
workstreams live on cluster nodes and need to round-trip
through the routing proxy at
/v1/api/route/workstreams/{ws_id}/approve, which resolves the
ws_id to its owning node and forwards the body verbatim.
approveWorkstream now picks the path based on whether targetWsId
matches the coord's own wsId.
2. LLM judge verdict never populates — rows freeze on the
heuristic-tier pill ("⚙ heuristic") even after the judge would
have completed. The judge runs async on the child node via a
daemon thread and updates _llm_verdicts there, but no signal
propagates back to the coord — cluster_state events don't fire
on verdict-only changes, and the live-bulk TTL is 5s with no
periodic poll.
Added _maybePollForJudgeVerdict: when renderChildRow encounters
a pending_approval_detail with judge_pending=true and items
missing judge_verdict, schedule a recursive 2s urgent
live-bulk re-fetch. Self-terminates when the verdict lands,
the row closes, the approval clears, or attempts hit the cap
(≈12s for a failed/timed-out judge so we don't poll forever).
Single timer per ws_id; re-renders are no-ops while a timer is
in flight.
Smoke-test assertions added for both fixes so a regression on
either path surfaces at test-time.
Copilot review on PR #424 flagged three items:
1. Schema drift on /v1/api/dashboard — DashboardWorkstream didn't
declare the new pending_approval_detail field, so generated
OpenAPI / typed clients were out of sync. Added
PendingApprovalItem + PendingApprovalDetail Pydantic models
and referenced PendingApprovalDetail from DashboardWorkstream.
2. deepcopy under _ws_lock in serialize_pending_approval_detail
could extend lock hold under contention with on_intent_verdict
(daemon judge thread) and per-token activity writes that also
take _ws_lock. _llm_verdicts entries are only assigned/cleared,
never mutated in place, so a snapped reference is stable after
the lock drops. Snapshot refs under lock; deepcopy after release.
3. Plan doc removed from the branch — design docs are local-only
working artifacts, same posture as PROGRESS.md.
Critical:
- coordinator.js RISK_SEVERITY accepted 'crit' only; production
emits 'critical' (per turnstone/core/judge.py:1556 + heuristic
seeds). A risk_level=='critical' verdict ranked as 0 and
rendered with .risk.low (green) styling, never triggering
the crit-risk auto-expand. Now accepts both aliases. Unknown
risk_level falls back to rank 2 ('high') so future schema
drift fails *safe* (over-alert) instead of silently
downgrading. Pill ternary handles both 'crit' and 'critical'
alias to the existing .risk.crit class.
Major:
- Urgent live-badge flush now coalesces N urgent calls in the
same JS tick into one bulk request via queueMicrotask, instead
of firing N single-id fetches. The motivating 10-children-
pending-bash scenario in the design doc now lands on one bulk
/v1/api/cluster/ws/live request.
- Test coverage gap: added test_session_ui_base.py cases for
POLICY-BLOCKED (item.error + needs_approval=False) and
judge-unavailable (no verdict + no judge_pending) matrix rows.
Added literal-string assertions to the smoke list in
test_coordinator_page.py so a refactor dropping either branch
surfaces at test-time.
Minor batch (4 coord.js + 1 CSS + 1 fake-divergence):
- 409 stale-call_id path re-enables both buttons before return
(urgent fetch is best-effort; could also fail).
- judgePending pill no longer conflicts with a present heuristic
verdict — guard changed from !judge to !verdict.
- Empty <div class="approval-reasoning"> no longer appended when
reasoning is absent but evidence is present (evidence still
renders inside the disclosure).
- Dead .ch-row .approval-pill.rec-* CSS rules removed (JS never
combines those classes). Recommendation chip in the disclosure
footer now has its own scoped rules so the chip is actually
styled.
- _FakeUI.serialize_pending_approval_detail call_id selection
aligned to the real impl's "first non-empty" semantics.
- liveBadgeCache reconnect cleanup now preserves permanent
(403/404) entries — denied users no longer pay one wasted
bulk fetch per denied id per reconnect.
All 4465 non-live tests pass. Ruff + mypy clean. node --check OK.
Closes the stale-button window where a sub-5s SSE gap would leave
liveBadgeCache holding pending_approval_detail for a child whose
approval was actually resolved during the gap. Without this clear,
zombie approve/deny buttons render until either the next child_ws_state
event or the natural TTL expiry (whichever comes first).
The clear sits beside the existing activeWaits.clear() in the
reconnect handler — same posture (drop client-only state that the
server's SSE replay doesn't cover) and same blast radius. The 409
race guard in submitChildApproval would catch a stale-call_id POST
even without this, but rendering wrong UI until the operator clicks
is the worse failure mode.
loadChildren's finally block already fires scheduleLiveFetch for
every visible row after the replace-mode refresh, so the cache
repopulates with authoritative pending_approval_detail in one bulk
request within the next debounce window.
Plan: docs/design/inline-child-approvals.md (chunk 4 of 4 — last
required chunk; 5/6 are stretch).
Chunk 3 of the inline-child-approvals plan + the SSE pipeline plumbing
needed for sub-second urgent fetches.
JS (coordinator.js):
- approveWorkstream(targetWsId, body) — generic POST helper, callable
for both the coord-self dock and the new per-child inline buttons.
- renderApprovalBlock(child, detail) — risk-level pill (.risk.* per
the design system primitives), tool-name summary with "+ N more"
for envelope-level approvals, intent_summary, ↳ judge reasoning
teaser, ▸ more disclosure carrying the recommendation chip,
evidence list, and items 2..N stacked sub-blocks. Plus matrix
coverage: judge_pending / judge unavailable / tool-policy
blocked / multi-item.
- submitChildApproval — handles the 409 stale call_id race by
invalidating the live cache + urgent-refetching, optimistically
clears pending_approval_detail on success.
- scheduleLiveFetch({ urgent: true }) — bypasses the 5s TTL +
cancels the debounce so attention transitions surface inline UI
immediately instead of after the next polling window.
- handleChildState fires urgent on activity_state="approval"
enter/leave; handleChildClosed eagerly invalidates the live
cache so closed rows can't render stale buttons.
CSS (index.html):
- New .approval-block / pill / preview / actions / disclosure
styles. Inline .act buttons duplicate the dock's colour treatment
(the dock-scoped rules don't reach the children-tree). Mobile
<700px touch targets ≥44px.
Pipeline (collector.py + coordinator_adapter.py):
- All three cluster_state event emitters and the child_ws_state
re-emit now carry activity_state. The previous omission left
the urgent-fetch trigger as dead code — discovered in review.
Tests:
- Static smoke test in test_coordinator_page.py asserting the new
helper names exist + the pending_approval_detail key is read.
Plan: docs/design/inline-child-approvals.md (chunk 3 of 4).
Threads the field added by Chunk 1 through the console's live-bulk
endpoint so coord tree UI can read it without a separate per-child
fetch. Three touchpoints:
- _CLUSTER_WS_LIVE_KEYS gains the new key so _fetch_live_block's
projection forwards it from the upstream /dashboard response on
node-backed child rows.
- _coordinator_live_snapshot synthesizes the same shape from
ConsoleCoordinatorUI._pending_approval for in-process coord
rows (no upstream /dashboard exists on the console pseudo-node).
- One source of truth: SessionUIBase.serialize_pending_approval_detail.
Both branches now emit the same 12-key live block; coord judge isn't
wired today so coord-self judge_verdict is always None — flagged in
the plan as a stretch follow-up.
Plan: docs/design/inline-child-approvals.md (chunk 2 of 4).
Lays the server-side groundwork for inline approve/deny buttons + judge
verdict on the coordinator children-tree UI. Two surgical changes:
1. SessionUIBase.serialize_pending_approval_detail() merges the active
_pending_approval items[] with per-call_id verdicts from
_llm_verdicts. The dashboard handler embeds this on every per-ws
row so cluster live-bulk callers can render inline UI without an
extra per-child round-trip.
2. make_approve_handler now returns 409 when the body sends a call_id
that doesn't match any currently-pending item. Closes the stale
call_id race where an operator clicks approve on a row showing
call A while the child has rolled over to call B. Empty/missing
call_id preserves backwards compatibility with CLI + channel
adapters that don't track it.
Cross-tenant exposure on /dashboard is consistent with the trusted-team
posture already in place for activity / tokens — documented in the new
method's docstring so the choice survives the next reviewer.
Plan: docs/design/inline-child-approvals.md (chunk 1 of 4).
Copilot review on PR #422 flagged that the DELETE-on-send (dequeue)
EndpointSpec declared no request_model, so the generated OpenAPI
showed no requestBody for an operation that *requires* a JSON body
with ``msg_id`` and 400s when it's missing.
- Add ``DequeueRequest`` to ``server_schemas.py`` with the single
required ``msg_id: str`` field.
- Wire ``request_model=DequeueRequest`` and ``response_model=
StatusResponse`` on the DELETE EndpointSpec; trim the now-redundant
inline body example from the description.
- Re-import the schema in ``server_spec.py`` and add the entry to
``_ALL_MODELS`` so the OpenAPI components list carries it.
- Regenerate ``openapi-server.json``.
Sibling thread on the close EndpointSpec was already addressed in
4000ae2 (request_model=CloseWorkstreamRequest).
4558 tests passing under ``-m "not live"``; ruff + mypy clean.
Copilot caught three real issues in PR #422 review, all clustered
around the close request body contract:
1. The interactive close handler runs with
``supports_close_reason=True``, which calls
``read_json_or_400(request)`` — an empty / non-JSON body returns
``400 {"error": "Invalid JSON body"}``. The previous SDK fix
sent NO body via ``json_body=None``, which would 400 against a
real server. The mock-transport test silently masked it because
the mock answered without inspecting the body.
2. The doc said the body was empty (or ``{}``), with no mention
of the optional ``reason`` field, its 512-byte cap, or the
credential-redaction guard.
3. The Pydantic schema for close was deleted outright; OpenAPI
and SDKs lost their typed shape for the optional ``reason``.
Changes:
- ``turnstone/api/server_schemas.py``: reintroduce
``CloseWorkstreamRequest`` with a single optional
``reason: str | None = None`` field. Docstring documents the
must-be-valid-JSON contract and notes that coord ignores the body
(``supports_close_reason=False``).
- ``turnstone/api/server_spec.py``: re-import the schema, point the
close ``EndpointSpec`` at it via ``request_model=``, restore the
``_ALL_MODELS`` entry. OpenAPI JSON regenerated.
- ``turnstone/sdk/server.py``: ``close_workstream`` (sync + async)
gains an optional ``reason: str | None = None`` parameter and
always sends ``json_body={}`` (or ``{"reason": ...}``) so the
body is never empty. Adds a regression test
(``test_close_workstream_sends_valid_json_body``) that inspects the
raw transport content rather than relying on a path-keyed mock —
the kind of check that would have caught this bug pre-merge.
- ``sdk/typescript/src/server.ts``: ``closeWorkstream`` gains an
optional ``opts.reason`` parameter; reintroduce
``CloseWorkstreamRequest`` interface in ``types.ts`` and re-export
from ``index.ts``.
- ``docs/api-reference.md``: close section documents the JSON-body
requirement, the ``reason`` field, the 512-byte cap, the
multibyte-safe behavior, the credential-redaction guard, and the
non-string-coercion path.
- ``CHANGELOG.md``: amend the 1.5.0 BREAKING block to reflect the
schema reintroduction (slim form, ``reason`` optional) instead of
the prior "removed outright" claim.
4558 tests passing under ``-m "not live"`` (was 4557 — +1 from the
regression test). ruff + mypy clean.
Reviewer caught real misses on the consumer-swap claim:
- TypeScript SDK still defined and re-exported `CloseWorkstreamRequest`
(types.ts + index.ts) — drop both. Now matches the Python-side
removal.
- Four `tests/test_auth.py` cases (`test_write_full_token_ok`,
`test_approve_full_token_ok`, `test_bearer_takes_precedence_over_cookie`,
`test_cookie_full_on_write_ok`) were tautological after the legacy
URL removal: they posted to `/api/send` / `/api/approve` and asserted
`allowed is True`, but those paths now classify as `read` so a read
token would also pass — they no longer tested the write/approve
scope enforcement. Swap to path-keyed URLs to restore the original
intent.
- `is_public_path("/api/send")` test renamed + retargeted to a
path-keyed URL.
Doc-table drift the previous commit missed:
- `docs/security.md` path-to-scope mapping rewritten for the
path-keyed verb family (write set, DELETE-on-/send dequeue,
per-ws_id approve).
- `docs/architecture.md` scope-model row text swap from `/api/send`
/ `/api/approve` to the path-keyed equivalents.
- `docs/diagrams/01-system-context.puml` channel→server edge label
swap.
- `docs/diagrams/15-auth-architecture.puml` scope class swap.
Cosmetic comment-only stragglers:
- `tests/test_session_worker.py` module docstring URL update.
- `tests/test_ratelimit.py` ~11 `/api/send` fixture-key strings
retargeted to `/api/workstreams/abc/send` so the URL fixtures
reflect the post-1.5 surface (rate limiter is path-agnostic; the
swap is purely cosmetic).
4557 tests still passing under -m "not live"; ruff + mypy clean.
CHANGELOG [Unreleased] / Removed (BREAKING — 1.5.0) block calling out
the legacy URL family removal with the swap table. Doc passes on
api-reference.md (per-endpoint sections rewritten with path
parameters and slimmer body shapes), architecture.md (handler-list
diagram and console-proxy URL example), console.md (URL-rewriting
JS shim docstring + SSE proxy example), and the two PlantUML
diagrams (11-console-data-flow, 16-channel-architecture).
Also picks up two test-side stragglers from step 5 that referenced
the legacy adapters in a docstring + a stale /v1/api/events SSE
test: turn into path-keyed equivalents. OpenAPI JSON dump regenerated
to reflect the catalog edits from step 3.
After this commit:
- 4557 tests passing under -m "not live"
- ruff + mypy clean on turnstone/ tests/ sdk/
- grep for "/v1/api/send", "/v1/api/approve", "/v1/api/cancel",
"/v1/api/workstreams/close" returns zero hits across turnstone/
sdk/ docs/ tests/ (excluding CHANGELOG.md, which intentionally
documents the old shape).
- grep for make_legacy_body_keyed_adapter, make_legacy_query_keyed_adapter,
_make_method_dispatch, close_legacy returns zero hits.
Mechanical updates across the test suite to swap legacy
/v1/api/{send,approve,cancel,events,workstreams/close} URLs for the
path-keyed equivalents under /v1/api/workstreams/{ws_id}/<verb>, and
to drop ws_id from request bodies (the path provides it now).
Per file:
- test_session_routes.py: deletes test_close_legacy_mounts_when_handler_provided
(the close_legacy slot is gone); test_send_mounts_post_and_delete_when_dequeue_provided
(added in PR commit 1) stays.
- test_openapi.py: expected-paths set swaps to path-keyed shape;
test_send_endpoint_has_request_body now asserts the OpenAPI for
/v1/api/workstreams/{ws_id}/send.
- test_auth.py / test_auth_identity.py: required_scope and
check_request fixtures swap to path-keyed shape; new tests cover
write/approve/read scope assignment for the path-keyed verbs +
the /node/* proxy mirror.
- test_sdk_server.py / test_sdk_console.py: mock-transport URL keys
swap; bodies drop ws_id.
- test_server_attachments_endpoints.py: ~17 send sites migrated to
/v1/api/workstreams/<ws>/send (a small Python script ran the bulk
rewrite — body ws_id stripped, URL rebuilt).
- test_server_authz.py: cross-tenant approve/close/cancel/events
tests retargeted to path-keyed URLs;
test_events_legacy_query_keyed_url_still_resolves_to_404_for_unknown_ws
renamed to test_events_path_keyed_url_resolves_to_404_for_unknown_ws
with the docstring updated to note the legacy adapter is gone.
- test_close_reason_persistence.py: 7 close sites all swap.
- test_console_routing_proxy.py: route-proxy tests swap to
/v1/api/route/workstreams/{ws_id}/<verb>; the upstream-URL
assertion now reads from .request (route_proxy uses
client.request(method, url, ...) for method passthrough); _wire_proxy
helper installs both .post and .request mocks for compatibility.
- test_route_proxy_audit.py: parametrized URLs migrated;
_make_proxy now also exposes a .request side-effect that delegates
to .post for the same compatibility surface.
- test_api_versioning.py: openapi.json path assertion swaps to the
path-keyed shape.
4557 passing under -m "not live"; ruff + mypy clean.
All in-tree consumers of the legacy /v1/api/send | /approve | /cancel |
events?ws_id= | /workstreams/close URLs now hit the path-keyed shape
under /v1/api/workstreams/{ws_id}/<verb>. Bodies drop ws_id (the path
provides it). The SSE event stream URL likewise moves to the path-keyed
form; channel adapters drop the params={"ws_id": ...} kwarg on
aconnect_sse.
Touched:
- turnstone/ui/static/app.js: 7 call sites (send×3, dequeue, approve,
cancel, close + EventSource SSE URL).
- turnstone/sdk/server.py (Python SDK): close_workstream, send,
approve, cancel, stream_events, send_and_wait's internal SSE
consumer.
- sdk/typescript/src/server.ts: closeWorkstream, send, approve,
cancel, streamEvents + sendAndWait's internal SSE consumer.
- turnstone/sdk/console.py: route_send, route_approve, route_close,
route_cancel — proxy URLs swap to /v1/api/route/workstreams/{ws_id}/<verb>.
route_plan_feedback / route_command remain body-keyed (out of scope).
- turnstone/console/server.py:
- Proxy mount table swaps the four legacy /api/route/{send,approve,
cancel,workstreams/close} mounts for path-keyed equivalents under
/api/route/workstreams/{ws_id}/<verb>; /send accepts both POST
and DELETE for dequeue.
- route_proxy reads ws_id from path_params (with body-fallback for
the surviving plan/command body-keyed mounts), uses
client.request(request.method, ...) so DELETE on /send proxies
through correctly, and audits DELETE-on-/send as a separate
"route.workstream.dequeue" action via _ROUTE_PROXY_AUDIT_ACTIONS.
- Internal `method` variable renamed to `verb` to avoid confusion
with HTTP method now that the two diverge.
- turnstone/channels/_sse.py: SSE URL builder swaps to path-keyed.
- turnstone/channels/{discord,slack}/bot.py: docstring URL updates.
- turnstone/server.py, turnstone/core/session_worker.py,
turnstone/sdk/events.py, turnstone/api/server_spec.py: comment /
docstring URL updates only.
Test fixtures still reference legacy URLs and will be swapped in step
5 of this PR.
- WRITE_PATHS / APPROVE_PATHS in turnstone/core/auth.py drop the four
legacy literal entries (/api/send, /api/cancel, /api/workstreams/close,
/api/approve). The path-keyed verb match for write expands from
{delete, open, refresh-title, title, attachments} to also include
{send, cancel, close}; a sibling branch maps POST /workstreams/{ws_id}/approve
to the approve scope, and a DELETE branch maps DELETE
/workstreams/{ws_id}/send (dequeue) to write. The /node/* proxy
block mirrors all four expansions so the console routing proxy
stays in lockstep.
- server_schemas.py drops the body-keyed ws_id field from SendRequest,
ApproveRequest, CancelRequest. CloseWorkstreamRequest deleted in
full (its only field was ws_id, now provided by the path).
- server_spec.py: drops CloseWorkstreamRequest from imports and
_ALL_MODELS, swaps the five legacy EndpointSpec entries to their
path-keyed equivalents (POST/DELETE workstreams/{ws_id}/send, POST
/approve, POST /cancel, POST /close, GET /events). Catalogue retains
/api/plan and /api/command unchanged (out of scope).
Tests still reference the legacy URLs and will fail at this commit;
test fixture updates land in step 5 of this PR. Step 4 swaps the
UI / SDK / console proxy / channels callers next.
Removes the pre-1.5 interactive URL family that mounted body- and
query-keyed shapes on top of the lifted path-keyed handlers via
make_legacy_body_keyed_adapter / make_legacy_query_keyed_adapter.
Path-keyed equivalents under /v1/api/workstreams/{ws_id}/<verb>
already serve every consumer; coord never used the legacy URLs.
Removed:
- make_legacy_body_keyed_adapter / make_legacy_query_keyed_adapter
from turnstone/core/session_routes.py.
- _make_method_dispatch from turnstone/server.py (zero callers
after legacy /api/send POST+DELETE block goes — its only purpose
was to bridge that single dual-method legacy URL).
- 5 legacy Route mounts in turnstone/server.py:
/api/events?ws_id, /api/send POST+DELETE, /api/approve, /api/cancel,
/api/workstreams/close.
- close_legacy field on SharedSessionVerbHandlers and its mount in
register_session_routes — the only surviving body-keyed slot in
the registrar, no longer needed.
Tightened make_dequeue_handler to read ws_id from the path only;
the body-fallback existed solely for the legacy DELETE /api/send
path and is now dead.
Test-suite updates and consumer call-site swaps (UI / SDK /
console proxy / channels) follow in subsequent commits in the same
PR — main stays broken across this commit until step 4 lands.
External SDK consumers on stable 1.0/1.3/1.4 calling these URLs
will receive 404s on upgrade to 1.5.0; CHANGELOG breaking-change
call-out lands with the docs commit.
Pre-flight for the legacy URL adapter removal: the path-keyed
`/v1/api/workstreams/{ws_id}/send` route only mounted POST today;
the dequeue handler was reachable only via the legacy
`DELETE /v1/api/send` body-keyed URL through `_make_method_dispatch`.
Add a new `dequeue: Handler | None = None` slot on
`SharedSessionVerbHandlers` next to `send`, mounted as a second
`Route` on the same path with `methods=["DELETE"]` (two distinct
Routes rather than collapsing methods on one Route — different
handler callables, and collapsing would force the same
method-dispatch wrapper this cleanup is tearing out).
Wire `dequeue=dequeue_handler` in `turnstone/server.py`'s
`SharedSessionVerbHandlers(...)` call so DELETE on the path-keyed
shape works in the same merge as the legacy mount removal.
Adds a regression-locking test covering both the POST+DELETE and
the dequeue-alone cases.
Switch fenced-code language tag from `json` to `http` on the seven
example blocks that mix an HTTP request line with a JSON body
(/trust, /restrict, /stop_cascade, /close_all_children, /approve,
/cancel, /close). Pure JSON response blocks stay tagged `json`.
Pre-existing pattern in the doc that Copilot flagged on the lines
this PR touched; fixed across all instances for consistency. No
content / URL changes — only fence-tag adjustment for correct
syntax highlighting.
The Stage 2 verb-shape lift converged coord and interactive on the
unified /v1/api/workstreams/{ws_id}/<verb> URL tree; the
/v1/api/coordinator/* tree was removed in P0. Two docs still
documented the pre-lift surface:
- coordinator-api-tour.md (the integrator's lifecycle walk-through):
rewrites all 9 step URLs to the post-lift paths, keeps a one-block
callout noting the historical /v1/api/coordinator/* tree and why
it converged, and drops the operation-id column (operation ids
shifted with the URL move and are now best looked up live via
/openapi.json + Swagger UI rather than baked into prose).
- bulk-endpoints.md (the cascade-mutation shape contract): two table
rows for stop_cascade / close_all_children fixed.
No code changes. CHANGELOG entry kept implicit since this is doc-only
and the URL convergence itself was already documented under the P0
verb-lift CHANGELOG block.
The /v1/api/dashboard endpoint was the last workstream-listing surface
keyed on `id` rather than `ws_id`. The Stage 2 list-verb lift converged
the active list (`/v1/api/workstreams`) and saved list
(`/v1/api/workstreams/saved`) on `ws_id` but explicitly left dashboard
alone to keep that PR's diff focused. This lands the same rename on
the remaining endpoint so v1 row shape is consistent across the family.
Scope kept narrow:
- Pydantic `DashboardWorkstream` and TS SDK `DashboardWorkstream`
interface both rename `id: str/string` → `ws_id`.
- The bundled web UI (`turnstone/ui/static/app.js`) is the only consumer
reading `dashboard.workstreams[].id` and is updated atomically.
- Console `_fetch_live_block` (cluster-inspect's projection over a
remote node's dashboard payload at `turnstone/console/server.py`)
flips its `entry.get("id")` lookup to `entry.get("ws_id")`.
- Drive-by: stale `id` example in `docs/api-reference.md` for the
earlier `/v1/api/workstreams` rename also fixed.
`_build_node_snapshot` (the global-events SSE node_snapshot payload
consumed by the cluster collector) deliberately stays on `id` — it's
part of a separate cluster-row family (collector → cluster_workstreams
→ console UI) that is internally consistent on `id` and would need its
own coordinated sweep. CHANGELOG documents the bounded blast radius.
Tests: 4554 passing (-m "not live"). ruff + mypy clean.
* feat(console): coord rich ws_state payload + live activity broadcast (Stage 2 follow-up)
Pre-lift coord's cluster broadcast was state-only — the dashboard's
coord rows showed the state column flipping but ``tokens`` /
``context_ratio`` / ``activity`` / ``content`` were all hardcoded
to zero / empty. The lift makes coord populate the same per-ws
metric fields interactive does and broadcasts them through the
cluster collector with the rich kwargs.
**Architecture changes:**
- Lift ``on_status`` / ``on_content_token`` / ``on_thinking_start`` /
``on_thinking_stop`` / ``on_stream_end`` / ``on_tool_result`` /
``on_reasoning_token`` / ``on_tool_output_chunk`` / ``on_info`` /
``on_error`` from ``WebUI`` to :class:`SessionUIBase` as base
implementations. Coord inherits the bodies; the per-ws metric
fields it had at the base but never populated now flow.
- ``WebUI`` keeps overrides for ``on_status`` / ``on_tool_result`` /
``on_error`` to layer Prometheus ``_metrics.record_*`` calls
on top of ``super()`` (node-only — the console isn't a node).
``WebUI._broadcast_state`` now uses the new
:meth:`SessionUIBase.snapshot_and_consume_state_payload` helper
for the rich-payload snapshot read.
- ``ConsoleCoordinatorUI`` adds a ``_broadcast_activity`` override
that calls the new
:meth:`ClusterCollector.update_console_ws_activity` (in-memory
pseudo-node row update; named ``update_*`` rather than ``emit_*``
to flag the no-fanout asymmetry vs. the rest of the
``emit_console_ws_*`` family).
- ``coord_adapter.emit_state`` reads ``ws.ui``'s snapshot under
``_ws_lock`` and passes the rich kwargs to the extended
:meth:`ClusterCollector.emit_console_ws_state`. Defensive when
``ws.ui is None`` mid-eviction (broadcasts state-only).
- ``coord_endpoint_config`` wires a new ``_coord_spawn_metrics``
hook so per-spawn ``_ws_messages`` / ``_ws_turn_tool_calls``
bookkeeping fires on coord too.
- ``_MAX_TURN_CONTENT_CHARS`` moved from ``turnstone.server`` to
``turnstone.core.session_ui_base`` so coord enforces the same
per-turn content cap.
**Three observable behaviour changes** (CHANGELOG-callout-worthy):
- Coord persists ``usage_event`` storage rows on every status
emission (governance dashboards / token-spend queries gain
coord visibility).
- Coord broadcasts live activity transitions to the cluster
collector (dashboard's coord rows show activity ticks between
state changes the same way interactive does), with last-emitted
dedup so a tool-heavy turn's repeated ``activity=""`` clears
don't hammer the collector lock.
- Cluster ``cluster_state`` events for coord rows now carry
non-zero ``tokens`` / ``content``. Frontend rendering that
conditionally hid these on coord can drop the branch.
**Tests:** 23 new tests in ``tests/test_coord_rich_ws_state_payload.py``
(per-ws metric writes, snapshot helper drain semantics +
single-lock-acquisition, adapter rich-payload pass-through +
None-UI defensive handling, activity broadcast wire + dedup +
failure swallow + no-op-when-collector-unset, spawn_metrics
hook, concurrent-writes-during-snapshot stress with reader
cycling through running/idle/error so drain branches actually
run, on_stream_end activity-clear pin). Plus WebUI override
regression tests confirming ``_metrics.record_*`` still fires
on top of the lifted bodies. Existing
``tests/test_webui_content.py`` updated to import
``_MAX_TURN_CONTENT_CHARS`` from its new home;
``tests/test_coordinator_adapter.py`` updated to expect the
rich-payload kwargs (default zeros) on
``emit_console_ws_state``. Total: ``4491 → 4514``.
``ruff check`` clean, ``mypy`` clean on touched files.
**/review pipeline** (4 finders → verify → dedupe) caught 14
findings → 12 unique (3 collapsed as duplicates of the lockless
``on_content_token`` writer):
- bug-1 Minor: ``on_status`` regressed coord's defensive
``usage.get(...)`` indexing → restored ``.get(..., 0)`` for
``prompt_tokens`` / ``completion_tokens`` on both base + WebUI
override.
- bug-2 Nit: concurrent-snapshot reader only used ``"running"`` →
cycled through ``("running", "idle", "error")`` so drain
branches run; also captures + re-raises thread exceptions
instead of silently passing.
- bug-3 + sec-2 + perf-3 Nit (merged): ``on_content_token``
mutated ``_ws_turn_content`` lockless while the snapshot drained
under lock → wrapped the cap-check + append + size-update in
``_ws_lock``.
- perf-2 Minor: collector lock contention from per-event activity
broadcasts → cached last-emitted ``(activity, activity_state)``
on the UI; subsequent identical ticks return early without
acquiring the collector lock.
- perf-4 Nit: join-under-lock in snapshot helper → swap-then-join
pattern (capture list reference under lock, reassign to empty,
join the captured list outside the lock). Halves the lock
hold and decouples the join walk from concurrent appenders.
- q-1 Minor: ``emit_console_ws_activity`` was misleading (no
``_fanout`` call, unlike the rest of the ``emit_console_ws_*``
family) → renamed to ``update_console_ws_activity`` + docstring
call-out for the asymmetry.
- q-2 + q-3 Minor/Nit: stale docstrings on
``coordinator_ui.py`` (still claimed "no per-node metrics —
Phase D") and ``_interactive_spawn_metrics`` (still claimed
"counters live on WebUI only") → both updated to reflect the
lifted base class + coord's new hook.
- q-4 Nit: broken Sphinx cross-ref
``:meth:\`_snapshot_and_consume_state_payload\``` → dropped
the leading underscore.
- q-5 Nit: missing ``test_coord_on_stream_end_clears_activity``
→ added.
**Two findings explicitly deferred** (out-of-scope follow-ups,
documented in CHANGELOG):
- perf-1: synchronous ``record_usage_event`` INSERT on coord
worker thread per status tick. Parity with WebUI is the lift's
goal; if throughput becomes a concern, batch usage_event writes
on a background flusher (would apply to both kinds).
- sec-1: coord assistant content now flows on the cluster SSE
stream, which has no per-user filter today. Pre-existing
exposure for interactive ``cluster_state`` events; the lift
extends to coord rows. Proper fix needs SSE auth gating
(``admin.cluster.inspect``) or per-listener user_id filtering
— separate security project, doesn't gate this lift.
* fix(console): apply review feedback on PR #420
Three review findings, all confirmed against source:
1. **Copilot — dedup-state-vs-failure race in `_broadcast_activity`**
(correctness bug): pre-fix ``self._last_broadcast_activity = current``
was assigned inside the ``_ws_lock`` block BEFORE the collector call.
If the collector raised mid-broadcast, the exception was swallowed
but the dedup state was already updated, so subsequent identical
activity ticks would be deduped and never retried — leaving the
dashboard's coord row stranded at the pre-failure activity until
the activity actually changed.
Fix: move the dedup-state update OUT of the lock and place it AFTER
a successful collector call. On failure, ``_last_broadcast_activity``
stays unchanged so the next identical tick retries. Two new
regression tests pin both the failure-recovery (``test_coord_ui_
broadcast_activity_failure_does_not_strand_dedup``) and the
happy-path dedup behavior (``test_coord_ui_broadcast_activity_
dedup_skips_identical_after_success``).
2. **Copilot — stale `emit_console_ws_activity` reference in
CHANGELOG**: the method was renamed to ``update_console_ws_activity``
per /review's q-1 finding before the original commit landed, but the
CHANGELOG entry was written ahead of the rename. Updated to match
the actual API + added the no-fanout asymmetry rationale inline so
readers don't have to chase the method name.
3. **code-quality bot ×2 — `except BaseException` in test workers**:
the concurrent-snapshot stress test caught thread-worker exceptions
with ``except BaseException`` (with a noqa to suppress BLE001).
``BaseException`` is overkill for a thread worker — ``SystemExit``
/ ``KeyboardInterrupt`` are main-thread signals and ``Exception``
is the right scope. Narrowed to ``except Exception`` on both
workers; ``writer_exc`` / ``reader_exc`` types narrowed from
``list[BaseException]`` to ``list[Exception]``.
Tests: ``4514 → 4516`` (+2 regression tests for the dedup race fix).
``ruff check`` clean, ``mypy`` clean. No code-path changes outside
the dedup-state placement; the rich-payload broadcast surface is
unchanged.
Server-side history endpoint declared ``error_codes=[404]`` but the
lifted ``make_history_handler`` factory can also return:
- ``400`` on empty ``ws_id`` (defensive — Starlette routing makes
it unreachable in practice, but the factory has the branch).
- ``500`` on the ``cfg.list_kind is None`` misconfig gate added in
the /review fix-up (defense-in-depth fail-loud; both production
cfgs wire ``list_kind`` so the gate doesn't fire today).
- ``503`` via ``cfg.manager_lookup`` when the kind's manager isn't
available (interactive's lookup never returns 503; coord's can).
Updated ``server_spec.py`` to ``[400, 404, 500, 503]`` per Copilot's
suggestion — matches the existing detail entry's shape so the two
endpoints document the same possible-error envelope.
Caught the parallel asymmetry on ``console_spec.py``: history was
``[403, 404, 503]`` but the lifted factory's misconfig + empty-
ws_id branches reach coord too. Updated to
``[400, 403, 404, 500, 503]`` — same factory body, same possible
responses, plus ``403`` from coord's ``admin.coordinator``
permission gate.
Regenerated ``openapi-{server,console}.json``. No code changes;
spec metadata only. Tests + lint + mypy unchanged.
Last verb-shape lift before v1.5.0 stable can tag. Adds two new
factories to ``turnstone/core/session_routes.py``:
- ``make_history_handler(cfg)`` — body lifted from coord's
``coordinator_history`` near-verbatim. ``?limit=`` query param
defaults to 100, clamps to [1, 500], malformed values fall back
to 100. Storage operations (``get_workstream`` on the
storage-fallback path, ``load_messages`` for the row read) now
run via ``asyncio.to_thread`` (was inline pre-lift on coord).
- ``make_detail_handler(cfg)`` — body lifted from coord's
``coordinator_detail``. Lazy-rehydrates a closed/evicted
workstream via ``mgr.open()`` on miss; mirrors
:func:`make_open_handler`'s exception envelope (``ValueError``
→ 503 with the session-factory's remediation text; bare
``Exception`` → correlation_id'd 500 with the per-kind noun
via ``cfg.audit_action_prefix``).
NO new ``SessionEndpointConfig`` fields — the factories reuse
``permission_gate``, ``manager_lookup``, ``not_found_label``,
``audit_action_prefix``, and (for history's storage-fallback
kind check) ``list_kind`` — all already wired by both production
lifespans for the list/saved factories.
Coord side: ``coordinator_history`` and ``coordinator_detail``
standalone handler bodies removed from ``console/server.py``;
``register_session_routes`` now wires
``history=make_history_handler(coord_endpoint_config)`` and
``detail=make_detail_handler(coord_endpoint_config)``.
Interactive side: GAINS both endpoints as a feature gain. Pre-lift
interactive had no ``GET /v1/api/workstreams/{ws_id}`` and no
``GET /v1/api/workstreams/{ws_id}/history`` — SDK consumers had to
subscribe to ``/events`` SSE just to read display fields or
message rows. The same lifted factories are wired with the
interactive endpoint config; cross-kind isolation is preserved on
both sides (history via ``cfg.list_kind`` storage-fallback gate
+ fail-loud-on-misconfig 500; detail via ``mgr.open()``'s internal
kind check).
Pydantic schemas: ``CoordinatorDetailResponse`` /
``CoordinatorHistoryResponse`` removed from ``console_schemas.py``;
``WorkstreamDetailResponse`` / ``WorkstreamHistoryResponse`` added
to ``server_schemas.py`` (mirrors the list lift's pattern for
``WorkstreamInfo``). Both server and console OpenAPI specs
reference the unified schemas; ``server_spec.py`` gains
``EndpointSpec`` entries for the new interactive endpoints. TS
SDK gains both interfaces in ``sdk/typescript/src/types.ts``;
``openapi-{server,console}.json`` regenerated.
Tests: 6 new coord regression/parity tests in
``test_coordinator_endpoints.py`` (limit clamping, cross-kind 404
on storage fallback, storage-only history, detail 503 on
session-factory misconfig, detail 500 with correlation_id on
unexpected rehydrate failure, history swallows
``load_messages`` exception → 200 with empty messages). 10 new
interactive parity tests in ``test_workstream_endpoints.py``
(``TestHistoryInteractive`` + ``TestDetailInteractive``). 1 new
openapi spec test pinning the server-side ``?limit=`` query param.
Total: ``4490 → 4491`` after the new exception-swallow
regression test landed. ``ruff check`` clean, ``mypy`` clean on
touched files.
/review pipeline (4 finders → verify → dedupe) caught 1 Minor
defense-in-depth (bug-1/sec-1, merged: ``make_history_handler``
fail-closed gate when ``cfg.list_kind is None``, mirroring
``make_saved_handler``'s same gate) + 1 Minor test-helper rename
(q-1: ``_interactive_history_cfg`` → ``_interactive_endpoint_cfg``)
+ 4 Nits (q-2 unused fixture parameter, q-3 CHANGELOG TS SDK
mention, q-4 missing exception-swallow regression test, q-5
misleading test comment) — all addressed in the same commit.
Three docstring + CHANGELOG drift items from the post-review
M3 + Mi1 fixes:
- ``make_list_handler`` docstring referenced ``cfg.list_resolve_title``
(singular) but the field renamed to ``list_resolve_titles``
(bulk variant) when the N+1 fix landed. Updated to the plural
name + a one-line note about the bulk SELECT pattern.
- ``make_saved_handler`` docstring still claimed kind was derived
from ``cfg.audit_action_prefix`` string-compare. The Mi1 fix
replaced that with the explicit ``cfg.list_kind`` field +
fail-loud-on-missing semantic; docstring now describes the
current contract.
- CHANGELOG ``[Unreleased]`` entry said "Three new
``SessionEndpointConfig`` fields" and listed the singular
``list_resolve_title`` wired to ``get_workstream_display_name``.
Updated to "Four" + the bulk plural names + the new
``list_kind`` field with its rationale (distinct from
``audit_action_prefix``; fail-loud on misconfig).
The fourth review comment — code-quality bot flagging the ``...``
ellipsis body on the new ``get_workstream_display_names`` Protocol
method as "statement has no effect" — is a false positive.
``...`` is the canonical Protocol method body throughout
``turnstone/core/storage/_protocol.py`` (every other method uses
it). Refuting; the file's pattern wins over the bot's per-method
suggestion.
No code changes; docstring + CHANGELOG only. Tests + lint + mypy
unchanged.
New ``make_list_handler(cfg)`` and ``make_saved_handler(cfg)``
factories in ``turnstone/core/session_routes.py`` replace four
pre-lift bodies (interactive ``list_workstreams`` +
``list_saved_workstreams``; coord ``coordinator_list`` +
``coordinator_saved``). Same factory + capability-flag pattern as
the merged cancel / open / events / create lifts.
Four new ``SessionEndpointConfig`` fields:
- ``list_resolve_titles: ListResolveTitles | None`` — bulk lookup
``(ws_ids) -> {ws_id: title-or-None}``. Interactive wires
``get_workstream_display_names`` (new bulk helper added on the
storage layer + memory.py); the lifted body resolves every active
row in ONE ``SELECT ... WHERE ws_id IN (...)`` instead of the
pre-lift N+1 (one SELECT per row).
- ``list_kind: WorkstreamKind | None`` — explicit kind classifier
for the saved-list storage filter. Replaces the initial draft's
``audit_action_prefix == "coordinator"`` string compare which
would have silently leaked INTERACTIVE rows for any future kind
whose audit prefix didn't match. Required when a kind mounts
list/saved; misconfig surfaces as a 500 with a clear log line.
- ``saved_state_filter: str | None`` — coord wires ``"closed"``;
interactive wires ``None``.
- ``saved_loaded_lookup: SavedLoadedLookup | None`` — coord-only
defence-in-depth filter that excludes ws_ids in the warm pool.
Behaviour changes (all observable in CHANGELOG):
- **Active-list row shape converges on always-include** ``{ws_id,
name, state, kind, parent_ws_id, user_id}``. Interactive renames
``id`` → ``ws_id``; both kinds populate every field (coord adds
kind + parent_ws_id; interactive adds user_id).
- **Top-level response key converges on ``"workstreams"``** on
both endpoints. Coord ``coordinators`` key removed — coord is a
1.5.0aN-only surface (never shipped stable) so the convergence
has no compat shim; SDK / frontend consumers swap once.
- **Storage + manager-lock work moved off the event loop on
interactive**. ``list_workstreams_with_history`` runs through
``asyncio.to_thread`` on both kinds (matches coord's pre-existing
perf-2 pattern from the saved-coordinators review); ``mgr.list_all``
+ per-row work also offloaded.
- **N+1 storage round-trips on /v1/api/workstreams eliminated**.
Pre-lift interactive resolved the alias for every active row in a
separate SELECT (up to 50 round-trips per dashboard refresh on a
saturated node). Lifted body issues one bulk SELECT.
Pydantic schemas: ``WorkstreamInfo.id`` renamed → ``ws_id``,
``WorkstreamInfo.user_id`` field added. ``CoordinatorInfo`` and
``CoordinatorListResponse`` removed (folded into the unified
``WorkstreamInfo`` / ``ListWorkstreamsResponse``). OpenAPI spec
snapshots regenerated. TS SDK types updated (``WorkstreamInfo``
interface gains ws_id + the always-include fields); TS test
mock + assertion updated to match.
``GET /v1/api/dashboard`` is intentionally NOT in this PR's scope
and still returns rows keyed on ``id``. Tracked as a separate
cleanup PR (tombstone-note added at the dashboard handler).
/review pipeline run; the four Major findings + one Minor + six
nits all addressed in the same commit:
- M1: TS SDK ``WorkstreamInfo`` interface stale (id: string) →
renamed + fields added.
- M2: TS SDK test masked the type-mismatch with stale mock → updated.
- M3: N+1 alias resolution on active list → bulk
``get_workstream_display_names`` helper + ``list_resolve_titles``
bulk cfg hook.
- M4: Missing interactive parity regression test for unified row
shape → mirror of coord's added in test_server_authz.py.
- Mi1: ``audit_action_prefix`` string-compare deriving kind →
explicit ``cfg.list_kind: WorkstreamKind`` field.
- Six nits: redundant inner asyncio import, forward-ref quotes on
Awaitable, duplicated frontend comments, dashboard ``id`` field
has no tombstone-note, empty-coord_mgr short-circuit on
``saved_loaded_lookup``.
4512 tests passing; ruff + mypy clean.
* refactor(core): defer emit_created on SessionManager.create + commit_create / discard pair
Eliminates the phantom create→close pair on coord rollback that was
documented as a known limitation in PR #416. The pair surfaced on the
cluster events stream when a multipart workstream-create request
failed attachment validation: coord's ``mgr.create`` fired
``emit_created`` synchronously, then the rollback called
``mgr.close`` which fired ``emit_closed``. Cluster consumers had to
reconcile via the collector's diff path. Post-fix, a rejected upload
produces zero events.
API changes on ``SessionManager``:
- ``create(..., defer_emit_created: bool = False)`` — when True,
skip the trailing ``emit_created`` so the caller can run additional
post-create work (attachment validation in the lifted HTTP handler)
before advertising the workstream. Default preserves the existing
"advertise immediately" contract for direct callers (test fixtures,
CLI REPL, channel adapters).
- ``commit_create(ws)`` — fires the deferred ``emit_created`` event
after the caller's post-create work confirms the workstream should
be advertised. Synchronous; the wrapped work is in-memory and
non-blocking on every kind (interactive: documented no-op stub;
coord: dict updates under a lock + ``queue.put_nowait`` fan-out).
- ``discard(ws_id)`` — releases the in-memory slot + cleans up the UI
WITHOUT firing ``emit_closed``. Distinct from ``close`` which
advertises the transition; ``discard`` is for the rollback case
where the workstream's existence was never advertised. Storage-row
deletion stays a separate concern (caller invokes
``delete_workstream``), mirroring ``mgr.create``'s split between
slot reservation and ``register_workstream``.
Caller-bug detection: ``Workstream._emit_created_fired`` is set
inside ``create`` (non-deferred path) and ``commit_create``;
``discard`` logs ``session_mgr.discard.after_emit_created`` warning
when invoked on an already-advertised workstream. Slot is still
released so capacity isn't stranded.
Lifted ``make_create_handler`` updated to use the deferred bracket:
pass ``defer_emit_created=True``, validate uploaded attachments,
then ``mgr.commit_create(ws)`` on success / ``mgr.discard(ws.id)``
on failure. Ordering invariants (``commit_create`` BEFORE
``audit_emit`` and ``post_install`` so any state events the worker
fires reach the cluster collector for an already-known ws_id) are
documented in the handler docstring.
Tests:
- 5 new ``SessionManager`` unit tests (defer skips emit, commit
fires it, commit no-ops without emitter, discard releases without
emit_closed, discard returns False on unknown id).
- 2 caller-bug regression tests (commit_create after discard pins
the silent re-emit behaviour; discard after non-deferred create
asserts the warning fires + slot still releases).
- 1 coord regression test asserting the cluster collector sees zero
events when attachment validation fails.
``/review`` pipeline run; M1 (test gap on caller-bug paths) +
Mi1 (no runtime guard for already-advertised) + Mi2
(``_make_manager`` event_emitter override) + Mi3 / N2 (duplicated
comments + ordering invariant) + N1 (drop ``to_thread`` on
``commit_create``) all addressed.
4509 tests passing; ruff + mypy clean.
* fix(core): apply Copilot + code-quality review feedback on PR #417
Copilot review:
- ``Workstream._emit_created_fired`` comment claimed the flag was
"set under the manager's _lock-protected emit", but the actual
ordering set it OUTSIDE the lock. Comment updated to describe the
real synchronization (non-deferred ``create`` sets it immediately
before ``emit_created``; ``commit_create`` sets it under the
manager lock alongside the tracked-ws check).
- ``commit_create`` had no guard against duplicate calls,
post-discard calls, or calls on workstreams not tracked by this
manager — any of those would have fired duplicate or phantom
``ws_created`` events. Added a guard symmetric to ``discard``'s
after-emit warning: under ``self._lock``, check ``_emit_created_fired``
+ ``_workstreams.get(ws.id) is ws``, no-op + log a warning
(``session_mgr.commit_create.already_fired`` /
``session_mgr.commit_create.untracked``) on either failure. The
emit itself still runs outside the lock so coord's collector
fan-out doesn't couple to the manager mutex.
- ``test_commit_create_after_discard_is_caller_bug_no_op`` was
internally inconsistent — name + docstring said "must not re-emit"
but the assertion expected the re-emit. Renamed to
``test_commit_create_after_discard_is_no_op`` and updated to
assert the new no-op + warning behaviour.
New test ``test_commit_create_is_idempotent_on_duplicate_call``
pins the second-commit-call code path: exactly one ``ws_created``
event fires, second call short-circuits via the guard with a
``commit_create.already_fired`` warning.
Code-quality bot review (3 findings, identical pattern):
- Three test ``assert`` statements wrapped side-effecting calls
(``assert mgr.discard(ws_id) is True/False``); under ``python -O``
the asserts strip and the side-effect strips with them. Refactored
all three to assign the result to a local first, assert on the
local. No behaviour change.
4510 tests passing; ruff + mypy clean.
Coord initial-message + create-time-attachments coordination:
- ``CoordinatorAdapter.send`` gains optional ``attachments`` + ``send_id``
kwargs so the worker dispatched at create time can carry the uploaded
files onto the first turn. Mirrors interactive's pre-existing
worker-thread pattern. The ``send_id`` reservation token soft-locks
the rows; the adapter's failure path unreserves so a worker crash
returns them to pending.
- ``_coord_create_post_install`` reserves any uploaded ``attachment_ids``
via the lifted ``reserve_and_resolve_attachments`` helper before
dispatching through the adapter — closes the parity gap with
interactive's create-with-attachments+initial_message flow.
- ``_reserve_and_resolve_attachments`` lifted from ``turnstone/server.py``
to ``turnstone/core/attachments.py`` as ``reserve_and_resolve_attachments``
so both processes use one kind-agnostic implementation.
Copilot review fixes on PR #416:
- Skill lookup now calls ``storage.get_prompt_template_by_name`` directly
rather than going through ``turnstone.core.memory.get_skill_by_name``;
that helper swallows storage exceptions into ``None`` which would have
masked outages as the 400 "Skill not found" branch. Calling storage
directly lets exceptions bubble to the lifted body's correlation_id'd
500 path so operators chasing skill-related reports can distinguish
real misses from registry outages.
- ``_interactive_create_build_kwargs`` /
``_coord_create_build_kwargs`` thread ``skill_data["name"]`` (the
canonical row name) into ``mgr.create`` instead of the raw
``body["skill"]`` value. Pre-fix a whitespace-padded request body
``"skill": " my-skill "`` would have persisted the dirty name even
though the lookup ran on the stripped key.
- ``make_create_handler`` docstring corrected: audit-emit failures
return 200 (not 201).
- ``_audit_workstream_created`` docstring corrected: factory keeps the
successful 200 response on audit-emit failure (was 201).
New regression test:
``test_create_with_multipart_attachments_and_initial_message_reserves``
asserts attachments are reserved (not pending) when both
``initial_message`` and uploads land in the same coord create request.
Updated ``_SendSession`` stub in ``test_coordinator_adapter.py`` to
match the new ``send`` / ``queue_message`` signatures.
4501 tests passing; ruff + mypy clean.
New ``make_create_handler(cfg, *, audit_emit=None)`` factory in
``turnstone/core/session_routes.py`` consumes five new ``SessionEndpointConfig``
fields (``create_supports_attachments``, ``create_supports_user_id_override``,
``create_validate_request``, ``create_build_kwargs``, ``create_post_install``)
and replaces both ``create_workstream`` and ``coordinator_create`` bodies.
Same factory + capability-flag pattern as the merged cancel / open / events
lifts. ``_validate_and_save_uploaded_files`` lifted to
``turnstone.core.attachments`` so both processes call one kind-agnostic
implementation.
Coord parity gains (§ Post-P3 reckoning item #1 + carry-forward):
- Create-time attachments: multipart parsing, validate+save+rollback,
``attachment_ids`` on the response. Coord adapter ``send`` doesn't yet
reserve attachments at create time, so the rows save as pending and the
next ``/send`` picks them up via the standard send-with-attachments path.
- Disabled-skill rejection (matches interactive's pre-lift gate).
- Always-include response shape ``{ws_id, name, resumed, message_count,
attachment_ids}`` populated with default ``False``/``0``/``[]`` on the
fields coord doesn't fill.
- 200 status (was 201).
- Audit-emit failures swallow + warning log instead of 500.
Both kinds converge on the manager-at-capacity 429, factory-misconfig 503,
and correlation_id'd 500 for unexpected ``mgr.create`` failure (interactive
lifted up to coord's safer error envelope).
Three /review fixes folded in:
- ``notify_targets`` malformed input gates at the validator (400) instead
of bubbling out of post_install as a 500 — pre-fix the workstream had
already been created + audited + broadcast by the time the validation
raised.
- Skill-lookup storage failures now share the correlation_id'd 500 path
with ``mgr.create`` (was masquerading as 400 "Skill not found").
- Whitespace-only ``skill`` field treated as empty (matches pre-lift coord).
CHANGELOG entry under [Unreleased] documents every observable behaviour
change. OpenAPI spec regenerated. Three new coord regression tests
(create-time-attachments save pending rows, always-include parity fields,
disabled-skill rejection) plus one interactive regression test
(notify_targets 400). 4500 tests passing.
* refactor(core): lift events verb body across both kinds (Stage 2 verb lift)
The interactive ``GET /v1/api/events?ws_id=...`` and coord
``GET /v1/api/workstreams/{ws_id}/events`` SSE handlers now share
one body via ``make_events_handler(cfg)``. Per-kind divergence
captured by two new ``SessionEndpointConfig`` fields:
* ``events_replay: EventsReplay | None`` — Protocol-typed callback
that yields the kind-specific initial replay payload. Interactive
wires ``_interactive_events_replay`` (connected + status + history
+ pending_approval + cached intent verdicts + pending_plan_review);
coord wires ``_coord_events_replay`` (just pending_approval +
pending_plan_review). The lifted body iterates the callback
before starting the live event loop.
* ``sse_executor_lookup: SseExecutorLookup | None`` — per-kind
executor for the live loop's blocking ``client_queue.get``.
Interactive returns the dedicated 200-thread ``sse_executor``
from app state so SSE polling stays isolated from every other
``asyncio.to_thread`` caller in the process; coord returns
``None`` and the lifted body falls through to the default executor.
Also adds ``make_legacy_query_keyed_adapter(handler)`` (sister to
``make_legacy_body_keyed_adapter`` from earlier lifts): reads
``ws_id`` from the query string and splices into ``request.path_params``
before delegating to the lifted body. Preserves the
``GET /v1/api/events?ws_id=...`` legacy URL shape so any 1.x SDK
consumer keeps working.
Old ``events_sse`` (server.py) + ``coordinator_events``
(console/server.py) bodies deleted.
Two convergence wins for coord:
* **SSE connect/disconnect metrics** — pre-lift coord didn't record
per-stream metrics; the lifted body always calls
``metrics.record_sse_connect()`` / ``record_sse_disconnect()``,
giving the cluster dashboard the same per-stream observability
interactive's had since 1.0.
* **Both kinds now check ``request.is_disconnected()`` AND the
``ws_closed`` event** to terminate. Pre-lift interactive relied
solely on ``ws_closed`` (which never fires if the client just
goes away without closing the workstream); pre-lift coord relied
solely on ``is_disconnected``. The lifted body uses both.
One observable shape change for coord callers: the lifted body
returns 409 ``"session has no UI"`` when ``ws.ui`` is missing
(placeholder / build-failed UI), matching pre-lift coord.
Pre-lift interactive returned 404 in this case; the lift converges
on 409 because the workstream EXISTS (404 would imply it doesn't).
Item #2 from § Post-P3 reckoning (rich ``ws_state`` payload parity
for coord) split out during scoping — touches different files
(``coordinator_ui.py`` + ``collector.py`` + ``session_ui_base.py``)
with different reviewer concerns. Tracked as standalone follow-up
``feat/coord-rich-ws-state-payload``.
Two /review fixes folded in:
* **Dedicated SSE thread pool restored.** Initial draft used
``asyncio.to_thread`` (default executor, ~32 workers). Pre-lift
interactive deliberately used a dedicated 200-thread
``sse_executor`` to avoid pool starvation; the
``sse_executor_lookup`` cfg field above restores that isolation.
* **5s poll timeout restored.** Initial draft shortened to 1s,
multiplying thread-wakeup rate 5x while the pool was already
starving. ``is_disconnected()`` between polls covers cancel-
detection latency.
Plus minor cleanups: stale ``coordinator_events`` comment
references in coordinator.js refreshed; ``TestInteractiveEventsLifted``
gets a ``_make_interactive_replay_mocks`` fixture so per-test
intent stays clear; live-loop coverage gap documented in the
test class docstring.
Lint + mypy clean. 4497 tests passing (+8 new events tests).
* fix(core): stream events replay from inside the generator instead of pre-building
PR #415 review caught that ``make_events_handler`` pre-built the
full replay payload (``connected`` + ``status`` + ``history`` +
pending prompts) into a list before constructing the
``EventSourceResponse``. Two real costs:
* **TTFB delay** — the client saw nothing until the heaviest
replay event finished serialising (``_build_history`` on a
long-running interactive workstream can take 10s of ms). With
pre-build, the ``connected`` event was buried at the end of
the materialisation pass instead of streaming first.
* **Listener-queue accumulation** — registering the per-UI
listener BEFORE building the replay let live events queue
during the build window. On a chatty mid-generation
workstream that window can fill the 500-slot listener queue
and drop events before the live loop starts draining.
Fix: iterate ``cfg.events_replay`` inside the async generator
so each event ships as soon as the callback yields it. The
observational-failure swallow semantics are preserved by
wrapping the iteration in the same try/except + log.debug as
before — partial replay is still acceptable; the live loop
continues either way.
Resolves the Copilot review thread on PR #415. Lint + mypy
clean. 4497 tests passing (no test changes — the replay
callbacks themselves are unchanged; only the lifted body's
consumption pattern flipped from eager-build to lazy-stream).
* refactor(core): lift open verb body across both kinds (Stage 2 verb lift)
The interactive ``POST /v1/api/workstreams/{ws_id}/open`` and coord
``POST /v1/api/workstreams/{ws_id}/open`` handlers now share one
body via ``make_open_handler(cfg, *, audit_emit=None)``. Per-kind
divergence captured by two new ``SessionEndpointConfig`` fields:
* ``open_resolve_alias: AliasResolver | None`` — interactive wires
``resolve_workstream`` so callers can pass user-friendly aliases
in the path param. Coord wires ``None``.
* ``open_post_load: OpenPostLoad | None`` — interactive wires
``_interactive_open_post_load`` (display-name sync + UI replay
via ``clear_ui`` + history + handler-side ``ws_created`` enqueue
onto the global SSE queue). Coord wires ``None`` and relies on
the cluster collector fan-out from
``CoordinatorAdapter.emit_rehydrated``.
Plus an optional ``audit_emit`` parameter (interactive wires
``_audit_workstream_opened``; coord wires ``None`` — coord doesn't
audit open today). Old ``open_workstream`` (server.py) +
``coordinator_open`` (console/server.py) bodies deleted.
**Load-bearing fix** (§ Post-P3 reckoning item #3 from the planning
docs): pre-lift interactive's ``open_workstream`` called
``mgr.create(ws_id=resolved_id)`` + ``ws.session.resume(...)`` to
rehydrate, bypassing ``mgr.open()`` entirely. After the lift both
kinds route through ``mgr.open()`` — which makes
``InteractiveAdapter.emit_rehydrated`` reachable on interactive
(it had been dead-by-routing) and gives the manager a single
rehydrate code path to maintain. ``emit_rehydrated`` stays a
documented no-op stub on the interactive adapter; the handler-side
``ws_created`` enqueue from the post-load callback is the
load-bearing emission for the SSE consumers.
Behaviour changes for interactive callers (documented in CHANGELOG):
* **Cross-kind open returns 404** (was 400 with
``"Workstream is not an interactive kind"``). The lift consolidates
on ``mgr.open()``'s single ``None``-return contract for missing /
wrong-kind / tombstoned rows. Security boundary unchanged.
* **Already-loaded response uses ``ws.name`` directly** (was
``get_workstream_display_name(resolved_id) or resolved_id``).
The dashboard listing endpoint still resolves aliases on its own
pass, so the user-visible name in the tab strip isn't affected.
Two /review fixes folded in:
* **Resume failures now return 5xx instead of broken-200.**
``SessionManager.open()`` previously caught and ``log.debug``-
swallowed exceptions from ``ChatSession.resume``. Since
``ChatSession.resume`` assigns ``self.messages`` *before* the
config-restore block, a partial-failure resume (corrupted
``workstream_config`` row, model-registry mismatch on a saved
alias, malformed ``temperature`` / ``max_tokens``) would leave
the session with history but with default config. Pre-lift the
interactive open handler called ``ws.session.resume`` directly
and let exceptions propagate as 500. Restored that behaviour:
``mgr.open()`` now re-raises resume exceptions after rolling
back the slot (``cleanup_ui`` + ``_remove_locked``), so the
lifted handler returns 500 with a correlation id and the storage
row stays available for a retry.
* **Bare ``except Exception`` documents intent.** A one-line
rationale in the handler body explains why the catch is broad
(no documented exception spec on ``adapter.build_session``;
resume can propagate via the new contract above). Keeps a future
contributor from narrowing it incorrectly.
Test scaffolding:
* ``tests/test_workstream_endpoints.py`` — fixture rebuilt to
use ``make_open_handler`` + a minimal cfg with a lazy alias
resolver so per-test ``@patch`` calls take effect. Added 5 new
tests: already-loaded uses ws.name, alias resolution runs first,
``mgr.open`` is called (NOT ``mgr.create``), post-load callback
fires with (request, ws) only on the load-from-storage path
(not the already-loaded shortcut), post-load exception swallowed
→ 200.
* ``tests/test_coordinator_endpoints.py`` — fixture imports
updated to ``make_open_handler``.
* ``tests/test_server_authz.py`` — ``TestOpenKindGate`` now expects
404 (not pre-lift's 400) for cross-kind open attempts. Docstring
explains the consolidation.
Two nit cleanups: dropped the unnecessary ``import secrets as
_secrets`` aliasing in the exception handler; refreshed the stale
``open_workstream`` reference in the ``AliasResolver`` doc-comment.
Lint + mypy clean. 4488 tests passing (was 4475; +13 new open
tests).
* fix(core): use cfg.audit_action_prefix for the per-kind noun in open's 500 error
PR #414 review caught the hardcoded ``"failed to open workstream"``
in ``make_open_handler``'s 500 path: coord callers got misleading
text (pre-lift coord said ``"failed to open coordinator"``).
The fix derives the noun from ``cfg.audit_action_prefix``
("workstream" interactive, "coordinator" coord) — a field both
production lifespans already construct, and which the previous
/review pipeline (q-5) flagged as dead config (set but read by
no factory). Reusing it here both fixes the wording AND gives
the field its first runtime reader.
Pinned by a new test
(``test_open_500_message_uses_kind_noun_from_cfg``) that wires a
coord-shaped cfg, forces ``mgr.open`` to raise, and asserts the
500 body contains ``"failed to open coordinator"`` + the
correlation id, without echoing the exception text.
Lint + mypy clean. 4489 tests passing (+1 new).
* refactor(core): lift cancel verb body across both kinds (Stage 2 verb lift)
The interactive ``/v1/api/cancel`` (body-keyed ws_id) and coord
``/v1/api/workstreams/{ws_id}/cancel`` (path-keyed) handlers now
share one body via ``make_cancel_handler(cfg, *, audit_emit=None)``
in ``turnstone.core.session_routes``. Per-kind divergence captured
by a new ``cancel_forensics: CancelForensics | None`` field on
``SessionEndpointConfig`` (interactive wires
``_capture_cancel_forensics``; coord wires ``None``) plus an
optional ``audit_emit`` (coord wires ``_audit_cancel_coordinator``;
interactive wires ``None`` — pre-lift interactive didn't audit
cancel).
Same factory + capability-flag pattern as P1.5's ``make_send_handler``
+ make_attachment_handlers. Old ``cancel_generation`` body deleted
from ``server.py``; old ``coordinator_cancel`` body deleted from
``console/server.py``.
Behavior changes (documented in CHANGELOG):
* **Coord gains the ``force`` flag.** Pre-lift coord ignored
``force``; the lifted body honours it on both kinds. Stuck-worker
recovery becomes available on coord (parity gain — coord workers
hang the same way interactive's can).
* **Coord cancel response always includes ``"dropped"``.** Pre-lift
returned bare ``{"status": "ok"}``; lifted returns
``{"status": "ok", "dropped": {}}``. Always-include parity with
interactive so SDK consumers don't branch on kind.
* **Coord cancel returns 400 ``"No session"``** on placeholder /
build-failed workstreams (was a silent 200 no-op pre-lift). Parity
with interactive's existing 400 branch.
* **Coord ``coordinator.cancel`` audit detail now includes
``force``** so operator-driven recovery is distinguishable from
routine cancels.
Three /review fixes folded in:
* **bug-1**: lifted body's ``resolve_approval`` is now gated on
``ui._pending_approval is not None``. Pre-fix, the unconditional
call leaked a stale ``approval_resolved`` SSE event on every
idle cancel — listener UIs that key on the event would dismiss
prompts they didn't have. ``resolve_plan`` keeps its existing
internal no-pending guard so the unconditional call is still
safe there.
* **bug-2**: force-cancel now clears ``_worker_running`` alongside
``worker_thread`` inside the same ``with ws._lock`` block. Prior
half-state ``(_worker_running=True, worker_thread=None)`` routed
follow-up sends through the queue-enqueue path onto the abandoned
worker (whose cancel flag short-circuits the queue-drain seam,
leaving messages orphaned until next spawn). Restores the
``(worker_thread, _worker_running)`` invariant
``session_worker.send`` documents.
* **bug-3**: ``coordinator_stop_cascade._fanout_on_children`` now
treats child cancel ``400 + "No session"`` as ``skipped`` (was
``failed``). Lifted coord cancel returns 400 on placeholder
children; matches the pre-lift outcome where those children were
silently no-op'd, so the cascade response's ``failed`` bucket
stops firing spurious operator alerts.
Test scaffolding:
* ``tests/test_coordinator_endpoints.py`` — replace ``coordinator_cancel``
fixture with ``make_cancel_handler(...)`` wiring; add 6 new
tests covering always-include shape, force-flag worker-abandon,
400-on-null-session, cancel_forensics swallowed-exception,
audit_emit swallowed-exception, no-stale-approval-resolved-on-idle.
* ``tests/test_server_authz.py`` — new ``TestInteractiveCancelLifted``
class with HTTP-level coverage of ``/v1/api/cancel`` for the
dropped shape, force-flag + ``_worker_running`` clearing, and
400-on-null-session. Pre-lift ``cancel_generation`` had no
HTTP-level test; this is the first.
One observable change for interactive (pre-existing call site):
``resolve_approval`` / ``resolve_plan`` now run on every cancel
regardless of ``was_running`` (was gated). Lifts coord's
unconditional behaviour onto interactive — a stuck approval-pending
state from a crashed worker can now be cleared via cancel without
requiring close + rehydrate.
Lint + mypy clean. 4484 tests passing (was 4475; +9 new cancel
tests minus the moved one that became part of the new suite).
* docs(core,changelog): correct cancel-lift behaviour description for resolve_approval
Two review comments on PR #413 caught the same drift between the
implementation and its documentation: my bug-1 fix gated
``resolve_approval`` on ``_pending_approval is not None`` (because
it broadcasts ``approval_resolved`` unconditionally), but the
``make_cancel_handler`` docstring and the CHANGELOG entry still
claimed both ``resolve_approval`` and ``resolve_plan`` "run on
every cancel" and "the calls are idempotent and no-op when
nothing is blocked".
Reality:
* ``resolve_plan`` does run on every cancel and its no-op-when-
nothing-pending behaviour is real (the method has an internal
``_pending_plan_review is None`` short-circuit).
* ``resolve_approval`` runs only when ``ui._pending_approval is
not None``. Without the gate, every idle cancel would broadcast
a stale ``approval_resolved`` SSE event and overwrite
``_approval_result``.
Updated:
* ``make_cancel_handler`` docstring (turnstone/core/session_routes.py
in the "Behavior changes vs the pre-lift handlers" section) —
splits the two methods into separate bullets, explains why
``resolve_approval`` is gated and ``resolve_plan`` isn't.
* CHANGELOG.md ``[Stage 2 Verb Lift — cancel]`` entry — same
split + rationale; the asymmetric coord pre-lift parity is
still flagged as the recovery path that drove the lift.
Docs-only change; lint + mypy clean; cancel test suite (59 tests)
unchanged.
* style(core): replace CancelForensics ellipsis stub with docstring
github-code-quality bot flagged the ``...`` body of
``CancelForensics.__call__`` as "Statement has no effect". The
ellipsis is the canonical Protocol method-body idiom (no real
issue), but switching to a one-line docstring satisfies the bot
AND adds a small piece of method-level documentation. The class-
level rationale (why Protocol-typed instead of a plain Callable
alias) moves from a wall of leading ``#`` comments into a proper
class docstring at the same time.
Style-only change; the Protocol semantics are identical.
* refactor(core): split SessionKindAdapter Protocol into construction + emission (Stage 2 P3)
The single ``SessionKindAdapter`` Protocol that ``SessionManager``
takes is split into two:
* ``SessionKindAdapter`` — kind / build_ui / build_session /
cleanup_ui. Required for every kind. The shared lifecycle
manager always delegates here for construction + cleanup.
* ``SessionEventEmitter`` — emit_created / emit_state /
emit_rehydrated / emit_closed. **Optional**, wired through a new
``event_emitter: SessionEventEmitter | None = None`` kwarg on
``SessionManager``. Reserved for future kinds whose lifecycle
transitions don't fan out anywhere; both production kinds wire
one today.
Both production adapters implement both Protocols. The interactive
lifespan (``server.py``) and console lifespan
(``console/server.py``) pass their adapter as both ``adapter`` and
``event_emitter`` — production behaviour is unchanged. Six lifecycle
sites in ``SessionManager`` (create / open eviction / open rehydrate /
close / set_state / close_idle / _reserve_and_install_locked unwind)
now call ``self._event_emitter.emit_*(...)`` guarded by
``if self._event_emitter is not None``.
InteractiveAdapter asymmetry preserved + documented:
* ``emit_closed`` stays load-bearing — it's the **sole** transport
path for ``ws_closed`` onto the process-wide global SSE queue
(Stage 1 consolidated emission from the create handler here so
there's exactly one emission point; ``name`` powers the
frontend's eviction toast).
* ``emit_created`` / ``emit_state`` / ``emit_rehydrated`` are
documented no-op stubs (``del ws[, state]``). Those events fire
from out-of-band paths — the create HTTP handler enqueues
``ws_created`` directly onto ``global_queue`` *after* attachment
validation (so a rejected upload doesn't surface a phantom
create→close pair); ``WebUI._broadcast_state`` emits the full
``ws_state`` payload (tokens + context_ratio + activity) via the
``SessionUI.on_state_change`` callback chain. The stubs exist
solely to satisfy ``SessionEventEmitter`` Protocol so the
adapter can be wired as the manager's ``event_emitter`` for the
``emit_closed`` path. Each stub has a 1-line inline rationale to
match the in-repo convention (``coordinator_adapter.py:210``).
Test scaffolding:
* ``tests/test_session_manager.py`` — ``_make_manager`` and
``_make_with_writer`` wire ``FakeAdapter`` as both ``adapter``
and ``event_emitter`` for production parity; the standalone
``test_create_uses_configured_node_id`` does the same.
``FakeAdapter.emit_rehydrated`` now records as
``_Event("rehydrated", ...)`` rather than conflating with
``"created"``, and ``test_open_resurrects_closed_state`` asserts
against ``events_of("rehydrated")`` so a regression where the
manager fires the wrong call on the open path actually fails.
* ``tests/_coord_test_helpers.py`` and
``tests/test_coordinator_end_to_end.py`` — wire
``CoordinatorAdapter`` as both args.
* Six interactive test fixtures (``test_skills.py``,
``test_prompt_templates_runtime.py`` x2, ``test_model_registry.py``,
``test_server_authz.py``, ``test_server_attachments_on_create.py``)
— wire ``event_emitter=adapter`` so they match the production
wiring, removing the footgun where a future contributor adds a
``gq.get_nowait()`` assertion and silently loses the only
``ws_closed`` transport.
* ``tests/test_interactive_adapter.py`` — drops the three
tautological no-op-emit_* tests (``test_emit_created_is_noop``,
``test_emit_state_is_noop``, ``test_emit_rehydrated_is_noop``);
keeps the four ``emit_closed`` tests (real behaviour).
Lint + mypy clean. 4475 tests passing.
* docs(core): correct SessionKindAdapter + SessionEventEmitter docstrings to match implementation
Two Copilot review threads on PR #412 caught the same real
discrepancy: my P3 docstrings on ``SessionKindAdapter`` and
``SessionEventEmitter`` described an *intent* — "interactive
doesn't implement ``SessionEventEmitter``; the manager skips emit
calls when no emitter is wired" — that doesn't match the actual
wiring. ``InteractiveAdapter`` does implement both Protocols and
``server.py`` does pass it as ``event_emitter``; only the three
no-op stubs (``emit_created`` / ``emit_state`` / ``emit_rehydrated``)
are dead, while ``emit_closed`` is load-bearing.
Updated both docstrings to:
* State that both production adapters implement both Protocols.
* Explain the asymmetry is in *which* emit methods carry real
bodies (coord: 4; interactive: 1, with 3 documented stubs because
the out-of-band paths — create handler ``ws_created`` after
attachment validation, ``WebUI._broadcast_state`` carrying the
richer ``ws_state`` payload — fire those events).
* Clarify the ``if self._event_emitter is not None`` guard exists
for the kwarg-omitted case (tests that don't care about events,
reserved for future kinds whose transitions don't fan out
anywhere).
Docstring-only change. Lint + mypy clean; the 75 tests in
test_session_manager + test_interactive_adapter + test_coordinator_adapter
pass.
Resolves the two Copilot review threads on PR #412 (commits
PRRC_kwDORcMomM67VyPD, PRRC_kwDORcMomM67VyPI).
Five fixes from Copilot's review of Stage 2 P1.5 — all preserve
behaviour, narrow docstring claims, and round out the response shape:
* **session_routes.py:supports_attachments docstring** — claimed
the handler "accepts only ``{"message": ...}``" when ``False``,
but the implementation silently ignores ``attachment_ids``
rather than rejecting. Updated wording to say the
attachment-resolution block short-circuits and any
``attachment_ids`` are silently ignored. Behaviour unchanged
(silent-ignore is the right choice for forward compat — clients
passing ``attachment_ids`` speculatively to a not-yet-lit-up
kind shouldn't get a 400).
* **session_routes.py:queue_full response shape** — restored the
always-include guarantee for ``attached_ids`` /
``dropped_attachment_ids``. The queue_full path now returns
``attached_ids: []`` and ``dropped_attachment_ids: list(requested_ids)``
so SDK consumers don't have to branch on status.
* **server.py:_interactive_spawn_metrics guard** — added
``_ws_turn_tool_calls`` to the ``hasattr`` chain. Previously
the guard checked ``_ws_lock`` + ``_ws_messages`` and then
unconditionally assigned ``_ws_turn_tool_calls`` — would
raise on a SessionUI subclass with the first two but not the
third.
* **console_spec.py:coord_send error_codes** — added 409
(the 'session UI not available' branch in
``make_send_handler`` returns 409, but the spec didn't list
it). OpenAPI spec regenerated; TS SDK types refreshed.
* **session_routes.py:tenant_check docstring** — claimed
interactive uses ``_require_ws_access`` with "404 on owner
mismatch", but the helper now delegates to
``resolve_workstream_owner`` which explicitly does NOT enforce
row-level ownership (trusted-team semantics; 404s only on
missing rows). Updated wording to match.
Six fixes from the local /review pipeline (find-bug + find-security +
find-quality, all confirmed by verify):
* **sec-1 (major)** — coord ``attachment_owner_resolver`` now
resolves through ``coord_mgr.get(ws_id)`` only and does NOT fall
back to storage. Without the kind-strict check, an
``admin.coordinator``-scoped caller could pass an *interactive*
workstream ws_id to the new coord attachment endpoints; the
generic ``get_workstream_owner`` storage call (kind-agnostic)
would resolve and grant cross-kind read / write access to
interactive attachments. New regression test
``test_coord_attachment_endpoints_404_on_interactive_ws_id``
pins the surface.
* **bug-1 (minor)** — UI hook calls in the spawn-path ``_run``
closure are now wrapped per-hook (via ``_emit_ui``) so a failure
in ``ui.on_error`` doesn't suppress the subsequent
``ui.on_stream_end`` / ``ui.on_state_change`` calls. Mirrors the
pre-P1.5 coord_adapter.send per-hook defense.
* **bug-2 (minor)** — ``make_dequeue_handler`` now 404s when
``ws.ui is None`` (preserves the pre-P1.5 ``_get_ws`` contract;
a partially-constructed or close-window workstream shouldn't
answer DELETE).
* **bug-3 (minor)** — ``coordinator.js`` gains a
``case "message_queued":`` handler that surfaces the queueing
as an info row. Coord wires ``emit_message_queued=True`` for
parity with interactive but the dashboard had no router branch
for these events, silently dropping them.
* **bug-4 (minor)** — error-message format on coord regressed
from ``f"{type(exc).__name__}: {exc}"`` to ``f"Error: {e}"``
(lost the exception class name, which coord operators rely on
to triage failures). Restored.
* **q-1 (major)** — duplicate ``_auth_user_id`` and
``_require_ws_access`` helpers in ``server.py`` and
``console/server.py`` now delegate to the lifted
``turnstone.core.web_helpers.auth_user_id`` /
``resolve_workstream_owner``. The lifted versions are the
canonical implementations; the shims keep existing call sites
working without a sweeping rename.
CHANGELOG entry adds a Security section noting the kind-strict
resolver fix and a behaviour callout for the cancel-state semantic.
Five new TestCoordinatorAttachments tests in
``tests/test_coordinator_endpoints.py`` exercising the lifted
attachment surface end-to-end on coord:
* upload → list round-trip
* get_content returns raw bytes with text/plain forced for text
* delete removes pending entries and clears them from the listing
* send with attachment_ids consumes pending under the send_id token
* send response carries attached_ids / dropped_attachment_ids even
on plain-text sends (unified shape parity)
The existing ``_coord_endpoint_config`` fixture grew capability
flags to mirror the production console wiring, and ``_make_client``
now mounts the four coord attachment routes via
``make_attachment_handlers``.
OpenAPI specs regenerated; TS SDK bumped to 0.5.0. CHANGELOG entry
under [Unreleased] documents the verb-shape lift, the coord
attachment surface coming online, the response-shape change for
``coordinator_send``, the unification of the three lifted classifier /
lock helpers under ``turnstone.core.attachments``, and the new SDK
helpers.
Replaces per-kind ``send_message`` / ``coordinator_send`` and the
four interactive attachment handlers with calls to the shared
factories from ``turnstone.core.session_routes``. Net deletion of
~660 LOC from ``server.py`` (the lifted body lives in
``session_routes`` and is mounted twice — once interactive, once
coord).
Interactive (``turnstone/server.py``):
* ``SessionEndpointConfig`` now carries ``supports_attachments=True``,
``attachment_owner_resolver`` (delegates to ``_require_ws_access``
via storage path to preserve test fixtures using MagicMock
managers), ``attachment_helpers`` (the lifted classifiers +
upload-lock), ``spawn_metrics`` (records the per-conversation
WebUI counters that coord doesn't have), and
``emit_message_queued=True``.
* New ``_make_method_dispatch`` adapter lets the legacy body-keyed
``/v1/api/send`` URL serve both POST (send) and DELETE (dequeue)
via the lifted handlers.
* The four attachment handler bodies (``upload_attachment`` etc.)
are deleted; the shared registrar mounts them via
``make_attachment_handlers(cfg)``.
Coord (``turnstone/console/server.py``):
* Same wiring with coord-specific resolvers
(``_coord_attachment_owner`` via the lifted
``resolve_workstream_owner``). ``spawn_metrics=None`` since the
coord dashboard doesn't have per-conversation counters; cluster
metrics fan out via the collector.
* Old ``coordinator_send`` body deleted.
* Console-side coord attachment endpoints come up automatically
through the shared ``AttachmentHandlers`` slot — no per-kind
attachment handler bodies needed at all.
Coord dashboard (``coordinator.js``): user messages with
attachments arriving on history replay now extract just the text
portion + a ``📎 N attachment(s)`` count badge instead of
JSON-stringifying the multipart content. Full chip-rendering with
click-to-view stays deferred.
Python SDK adds coord-side helpers on
``AsyncTurnstoneConsole`` + ``TurnstoneConsole``:
``coordinator_send`` (with ``attachment_ids``),
``coordinator_upload_attachment``,
``coordinator_list_attachments``,
``coordinator_get_attachment_content``,
``coordinator_delete_attachment``. URL prefix is direct
``/v1/api/workstreams/`` since coord workstreams live on the
console — no routing-proxy hop needed.
Behaviour change for coord callers:
* Worker-queue-full responses are now ``200 {"status": "queue_full"}``
for parity with interactive (was ``429 {"error": "..."}``). SDK
consumers checking for 429 should switch to the status field.
* Send response now always carries ``attached_ids`` /
``dropped_attachment_ids`` (empty arrays on plain text sends);
the live-worker reuse path also surfaces ``priority`` /
``msg_id``.
Stage 2 P1.5 — verb-shape unification at the HTTP layer for both
``send`` and the four attachment endpoints. New factories in
``turnstone.core.session_routes``:
* ``make_send_handler(cfg)`` — single body covering the
attachment-resolution dance, dispatcher hand-off, queue/spawn
outcome surfacing, and metrics increment. Capability flags on
``SessionEndpointConfig`` (``supports_attachments``,
``attachment_owner_resolver``, ``attachment_helpers``,
``spawn_metrics``, ``emit_message_queued``) toggle the per-kind
bits without forking the body.
* ``make_dequeue_handler(cfg)`` — DELETE branch (cancel a queued
message by ``msg_id``). Path-keyed; mountable on both new
``/v1/api/workstreams/{ws_id}/send`` and the legacy body-keyed
``/v1/api/send`` URL via ``make_legacy_body_keyed_adapter``.
* ``make_attachment_handlers(cfg)`` — quartet of upload / list /
get_content / delete with shared scope checks and 404 masking.
Per-kind classification + locking comes in via the new
``AttachmentUploadHelpers`` bundle so the cfg stays declarative.
Three pure helpers (``sniff_image_mime``,
``classify_text_attachment``, ``upload_lock``) moved from
``turnstone/server.py`` to ``turnstone/core/attachments.py`` so the
console process can wire them into the lifted attachment endpoints
without depending on the node-side server module. Behaviour is
unchanged.
``turnstone.core.web_helpers`` gains ``auth_user_id`` and
``resolve_workstream_owner`` so both kinds share the owner-resolution
helper underpinning attachment scoping. The interactive ``trusted-team``
404-on-missing semantics are preserved; ``not_found_label`` is
parameterised so coord can return ``coordinator not found``.
Console spec adds ``CoordinatorSendResponse`` (parity with interactive
``SendResponse``) and four new endpoint declarations for the coord
attachment surface.
PR #410 review pass:
* **session_worker**: ``except BaseException`` → ``except Exception``
in ``_runner`` (code-quality bot). Daemon threads don't receive
SystemExit/KeyboardInterrupt, so the wider catch was unjustified
defensive style. Same defense-in-depth for unexpected ``run()``
exceptions; doesn't widen scope to runtime signals.
* **session_worker**: ``threading.Thread()`` construction moved
inside the spawn branch under ``ws._lock`` (Copilot). The
enqueue path no longer allocates and then discards a Thread
object on each call against a busy workstream. Thread()
construction is microsecond-cheap, so the lock-window growth is
negligible vs. the saved allocation churn.
* **lifespans**: ``state_writer.shutdown()`` (and the console
equivalent) now run via ``asyncio.to_thread`` so the daemon-
thread join + sync DB drain don't block the event loop and
delay other teardown tasks (Copilot, ×2).
* **tests**: five remaining ``writer._flush_once()`` calls
switched to the public ``writer.flush()`` API across
test_session_manager.py (4) and test_state_writer.py (1)
(Copilot, ×5). Tests no longer depend on private internals.
Six fixes from the local /review pipeline (find-bug + find-perf +
find-quality, all confirmed by verify):
* **bug-1 (critical)** — ``StateWriter.record(flush_now=True)`` now
drops any pending buffered transient for the same ws_id AND waits
on the flush_lock before its sync UPDATE. Without this, an
earlier buffered 'running' could flush AFTER the sync 'error'
write and clobber the terminal state — same shape as the
close-vs-buffered-transient race ``discard`` was already
guarding. New regression tests cover both the drop and the
in-flight wait.
* **bug-3 (major)** — ``session_worker.send`` now assigns
``ws.worker_thread = t`` AND sets ``ws._worker_running = True``
under the same ``ws._lock`` acquisition. Previously
``worker_thread`` was assigned outside the lock, so a reader
holding ``ws._lock`` could observe ``_worker_running=True``
paired with a stale (already-exited) ``worker_thread`` —
defeating every ``ws.worker_thread is me`` identity check
downstream.
* **bug-2 (major)** — rewind/retry busy gate in
``server.py:command`` now reads ``ws._worker_running`` instead
of ``ws.worker_thread.is_alive()``. The is_alive() gate could
see a stale dead thread under ws._lock while a new worker was
in the middle of starting (post bug-3 fix the window narrows
but the gate-mismatch was independent — ``_worker_running`` is
the canonical gate post-Stage-2-P1).
* **perf-2 (major)** — ``StateWriter.discard`` now waits on
``_flush_lock`` with a 5s timeout (configurable). Without a
bound, a stuck Postgres connection inside an in-flight flush
would block ``close()`` and ``close_idle()`` indefinitely while
they hold ``ws._lock`` — a system-wide hang on every close
path. On timeout we log + proceed; the worst-case degrades to
"buffered transient flushes shortly after sync 'closed'"
(eventual consistency) rather than process hang.
* **q-1** — inline comments on ``run_retry`` and ``_run_initial``
now explain why those two spawn sites don't go through
``session_worker.send``: retry-when-busy is a hard reject (no
fallback queue), and init-on-create can't have a pre-existing
worker by construction (enqueue branch is dead code). Both
still set ``_worker_running`` + ``ws.worker_thread`` together
under ws._lock for parity with the dispatcher.
* **q-3** — ``state_writer.discard`` callsite comments in
``session_manager.py`` no longer reference 'bug-3' (which lived
only in untracked working notes). Now describe the invariant
inline by what it prevents.
* **q-5** — ``StateWriter._flush_once`` promoted to public
``flush()``. Tests now drive flushes via the public API.
Two new bullets under [Unreleased]:
* Worker dispatch unified — ``session_worker.send`` shared by
interactive ``/v1/api/send``, the coord adapter, watches, retry,
and initial-message paths. Gate is ``_worker_running`` (atomic
under ws._lock) instead of ``Thread.is_alive()``. Closes a
parallel-worker race that any concurrent path (watch + /send,
retry + /send, init + /send) could trigger pre-P1.
* Buffered ``StateWriter`` for set_state — non-terminal transitions
now show up in storage up to ~1s late (SSE consumers see them
immediately via the adapter). Terminal ERROR + close still write
sync; bug-3 invariant preserved via state_writer.discard before
the sync 'closed' write.
The `/send` HTTP body convergence stays out of scope — interactive's
attachments / reservations / queue-outcome distinctions diverge from
coord's response shape too far for a clean factory split until
coord grows attachments parity (post-1.5.0).
Five new tests under ``TestSessionManagerWithStateWriter`` exercise
the bug-3 invariant under write-behind:
* set_state buffers via the writer (long flush_interval → no sync
write until drain).
* set_state(ERROR) flushes synchronously.
* close after a buffered transient writes 'closed' as the final
state — the buffered 'running' must NOT be flushed to storage
AFTER close's sync 'closed' write.
* close_idle exhibits the same invariant.
* set_state arriving AFTER close short-circuits on ws._closed and
never reaches the buffer.
``SessionManager.__init__`` accepts an optional ``state_writer``;
when present, ``set_state`` for non-terminal transitions records via
the buffered writer instead of holding ``ws._lock`` across a sync DB
UPDATE. Terminal ERROR transitions still flush sync (error-surfacing
paths need durability before any observer sees the state).
``close()`` and ``close_idle()`` call ``state_writer.discard(ws_id)``
under ws._lock BEFORE their sync 'closed' write — drops any pending
buffered transient and waits on the flush_lock for any in-flight
flush to complete. Without this, a buffered 'running' could land in
storage AFTER the sync 'closed' write and resurrect the closed row
(bug-3 invariant under write-behind).
Lifespan wiring on both servers: build the StateWriter alongside the
SessionManager, ``state_writer.start()`` on enter, ``shutdown()`` on
teardown (drains any pending writes synchronously). Tests can leave
``state_writer=None`` and get the legacy direct-write behaviour.
``turnstone.core.state_writer.StateWriter`` buffers non-terminal
``update_workstream_state`` writes (last state per ws_id wins) and
flushes them on a ~1s cadence (configurable). Terminal ERROR
transitions and close()'s 'closed' write bypass the buffer.
Bounded buffer (``max_buffer=10000`` default) evicts the oldest
ws_id on insertion overflow — protects against unbounded growth
when storage is unreachable. ``discard(ws_id)`` drops any pending
buffered transition AND waits on a flush_lock for any in-flight
write to complete; this is the close-path hook that preserves the
bug-3 invariant (a closed row can't be resurrected by a buffered
transient writing AFTER close's sync 'closed').
13 unit tests cover coalescing, flush_now, bounded buffer, the
discard / in-flight-flush wait, lifecycle (start/shutdown
idempotence), wake-on-record latency, and resilience to storage
errors poisoning subsequent flushes.
Five spawn sites in turnstone/server.py now share the worker dispatch:
* ``send_message`` (``POST /v1/api/send``) — the main path. Now uses
``session_worker.send`` with separate ``_enqueue`` / ``_run``
closures; the queue-vs-spawn outcome is conveyed via a captured
``queue_outcome`` dict so the existing response shapes
(``status: queued`` vs ``status: ok``) survive.
* ``_make_watch_dispatch`` — watch results dispatch.
* ``run_retry`` (post-rewind) and ``_run_initial`` (initial-message
on workstream creation) — set ``_worker_running`` directly under
ws._lock instead of going through session_worker (their structural
shape doesn't fit a queue-vs-spawn decision) but stay consistent
with the shared gate so they can't race with /send into parallel
workers.
* ``cancel_generation``'s ``was_running`` snapshot now reads
``_worker_running`` for parity with the dispatcher.
Pre-dispatch cancel-await also gates on ``_worker_running`` for
consistency. The ``busy_error`` / ``status: busy`` legacy branch
(reached only when worker is alive but ws.session is None) is gone
— the new path checks ws.session up front and returns the same
500 shape.
Test fixtures in test_server_attachments_endpoints.py and
test_watch_dispatch.py updated to set ws._worker_running explicitly
(MagicMock auto-truthifies the field, which would otherwise mis-route
all idle paths into queue mode).
Introduces ``turnstone.core.session_worker.send`` — the atomic
check-and-(spawn-or-queue) decision both interactive and coordinator
HTTP paths use to drive ``ChatSession.send``. Callers pass no-arg
``enqueue`` / ``run`` closures; the shared module owns only the
``ws._worker_running`` lifecycle.
CoordinatorAdapter.send now delegates to the shared module — its
``_spawn_worker`` body is gone. Workstream._worker_running's
docstring updated to note both kinds use it post-Stage-2-P1.
PR #409 line-level review feedback. Three of four findings valid;
the fourth (code-quality bot's "unused TYPE_CHECKING imports")
verified as false-positive — removing the imports breaks mypy on
the string-form annotations in ``ManagerLookup`` / ``TenantCheck``
/ ``CloseAuditEmitter``.
CI lint failure (ruff format on ``tests/_coord_test_helpers.py``)
addressed alongside.
Findings addressed:
- **Copilot #1** (``session_routes.py`` SessionEndpointConfig
docstring): said the config is "stored on
``app.state.session_endpoint_config``" and "handler bodies pull
this config from app.state". Stale after the previous fixup
switched the factories to capture ``cfg`` via closure. Rewrote
the class docstring + the lifted-handler comment block + the
module docstring + the ``create_app`` block comments in both
``server.py`` and ``console/server.py``.
- **Copilot #2** (``server.py:_interactive_manager_lookup``
docstring): referenced ``:data:SessionRouteHandlers`` which was
renamed to ``SharedSessionVerbHandlers`` AND wasn't the right
reference anyway — the callable matches
``SessionEndpointConfig.manager_lookup``. Fixed.
- **Bonus**: dropped the now-dead
``app.state.session_endpoint_config = ...`` assignments in both
servers (nothing reads them since the closure-capture switch).
- **Bonus**: dropped the stale "close (interactive caps + redacts +
persists close_reason)" entry from the deferred-verbs comment in
``session_routes.py`` — close was lifted in the previous commit
and is no longer in the deferred set.
- **CI lint**: ``ruff format`` joined the
``MockStorage.list_services`` signature in
``tests/_coord_test_helpers.py`` to a single line (95 chars,
fits the 100-char limit).
ruff + ruff format + mypy clean. 88 affected tests pass.
Addresses the eight verified findings from the second review pass on
the body-convergence work (one bug-flagged behavior change, one
defensive-style nit, six quality items). One quality item (q-6,
``request.scope[\"path_params\"]`` mutation in the legacy adapter)
is documented but not refactored — restructuring the lifted handler
signatures to take ``ws_id`` as an explicit param is bigger than
this fixup's scope; the adapter docstring already explains the
choice.
Findings addressed:
- **bug-1 + q-5**: hoist module-level ``log = get_logger(__name__)``
in ``session_routes.py``; bump audit-failure log from ``debug``
to ``warning`` (compliance signal). Document the interactive
500-on-audit-failure → 200+log behavior change in CHANGELOG +
in ``make_close_handler``'s docstring.
- **bug-2**: switch ``_audit_close_workstream`` to
``getattr(request.app.state, \"auth_storage\", None)`` for
consistency with the upstream gate. Same fix on coord side.
- **q-1**: pass ``SessionEndpointConfig`` into
``make_approve_handler(cfg)`` and
``make_close_handler(cfg, *, audit_emit, supports_close_reason)``
via closure capture. Removes the implicit ``app.state`` contract
and parallels the two factory signatures. Tests + production
wiring updated.
- **q-2**: promote ``_audit_close_coordinator`` to a module-level
function in ``turnstone/console/server.py``. Both test fixtures
import it instead of duplicating the body. The previous three
near-identical implementations collapse to one.
- **q-3**: lift ``_interactive_tenant_check`` and
``_audit_close_workstream`` from nested ``create_app`` closures
to module-level functions in ``turnstone/server.py``, beside the
other ``_audit_*`` / ``_require_*`` helpers. Add
``_interactive_manager_lookup`` so the config doesn't need a
lambda. ``create_app`` shrinks accordingly.
- **q-4**: merge the bottom ``if TYPE_CHECKING`` block into the
one at the top of ``session_routes.py``.
- **q-7**: replace ``assert mgr is not None`` with
``mgr = cast(\"SessionManager\", mgr_opt)`` in both lifted
handlers — survives ``python -O`` and makes the type-checker-only
intent explicit.
- **q-8**: update ``test_coordinator_endpoints.py`` file docstring
to mention the lifted-handler wiring.
ruff + mypy + 4366 pytest pass. Live console smoke against the
unified URLs returns 503 (no coord_mgr in smoke env) — proves the
factory-captured config is reachable + manager_lookup fires.
CHANGELOG ``[Unreleased]`` entry expanded to flag the audit-failure
swallow as an interactive behavior change alongside the existing
500→404 standardization.
Adds two paragraphs:
- Notes the two verbs (``approve``, ``close``) whose bodies were
successfully lifted into the shared registrar, plus the close-
failure status-code standardization (500 → 404 on coord). Calls
out the verbs whose bodies are intentionally NOT lifted, with
the underlying reason (Priority 1 dependency, response-shape
unification, etc.) so the next-session reader doesn't re-litigate.
- Notes the TS SDK 0.4.0 bump and the regenerated reference specs.
No code change.
Stage 2 Priority 0 Step 0.2 body-convergence — second verb.
``make_close_handler(audit_emit=..., supports_close_reason=...)``
factory in ``turnstone/core/session_routes.py`` produces the lifted
body; both interactive and coord pass their kind-specific audit
emitter at app construction.
The two body-keyed close URL aliases on the interactive side reach
the same lifted body:
- ``POST /v1/api/workstreams/{ws_id}/close`` (new, path-keyed)
via ``register_session_routes(handlers.close=...)``.
- ``POST /v1/api/workstreams/close`` (legacy, body-keyed) via
``make_legacy_body_keyed_adapter(close_handler)``.
Coord exposes only the path-keyed shape.
Behavior gains:
- ``supports_close_reason=True`` (interactive only) keeps the 512-
byte UTF-8 cap + credential redaction + ``workstream_config``
persistence path. Coord stays at ``False``; if coord ever wants
close-reason metadata, flipping the flag is a one-line change.
- ``audit_emit`` is per-kind so each owns its detail dict shape
(``{kind, parent_ws_id, reason}`` vs ``{coord_ws_id, src}``) and
audit action name (``workstream.closed`` vs ``coordinator.close``).
- Standardizes the close-failure status code to 404 across both
kinds. The coord code previously returned 500 on a
``mgr.close()`` race-loss, which was overly pessimistic — the
semantic is "the ws was popped between .get() and .close()", i.e.
not-found.
Coord-side test fixtures (``test_coordinator_endpoints``,
``test_coordinator_end_to_end``) swap the imported
``coordinator_close`` for the lifted handler + a local audit_emit
adapter so the tests exercise the same code path the live console
does.
ruff + mypy + 4366 pytest pass. Live console smoke against
``POST /v1/api/workstreams/abc/close`` returns 503 (no coord_mgr
loaded in the smoke env) — proves the lifted handler is reachable
+ the manager_lookup callable fires correctly.
Two verbs converged so far (``approve`` + ``close``); the remaining
pairs (``send``, ``cancel``, ``open``, ``events``, ``create``,
``list``, ``saved``, ``history``, ``detail``) have substantive
behavior divergence that doesn't factor cleanly into the
SessionEndpointConfig + factory-handler pattern — see the
session_routes module docstring for the per-verb status.
Stage 2 Priority 0 Step 0.2 body-convergence — first verb. Both
interactive ``approve`` and coord ``coordinator_approve`` handler
bodies collapse into ``make_approve_handler()`` in
``turnstone/core/session_routes.py``. Each kind sets a
``SessionEndpointConfig`` on ``app.state`` carrying the kind-
specific policies (auth gate, manager lookup, tenant check, audit
prefix, not-found label) the lifted body consults at request time.
The two interactive URLs converge:
- ``POST /v1/api/workstreams/{ws_id}/approve`` (new, path-keyed)
reaches the lifted body directly via ``register_session_routes``.
- ``POST /v1/api/approve`` (legacy, body-keyed) keeps shipping;
``make_legacy_body_keyed_adapter`` peeks the body for ``ws_id``,
splices it into ``request.path_params``, and forwards to the same
lifted body. Frontend can keep using the legacy URL — no caller
churn.
Coord exposes only the path-keyed shape (its URLs were experimental
in 1.5.0aN; the URL-shape commit already removed the ``coordinator/``
prefix).
Tenant-check is split out from permission-gate so interactive's
``_require_ws_access`` (404 on cross-owner) and coord's
``_require_admin_coordinator`` (cluster-wide scope) coexist without
either kind triggering the wrong gate.
Coord-side test fixture (``test_coordinator_endpoints._make_client``)
swaps the imported ``coordinator_approve`` for the lifted handler
and seeds ``app.state.session_endpoint_config`` so the tests
exercise the same code path the live console does.
Net delta: ~−25 LOC for this verb on top of the SessionEndpointConfig
+ legacy-adapter scaffolding (~80 LOC paid once). Subsequent verb
lifts amortize against that scaffolding.
ruff + mypy + 4366 pytest pass. Live console smoke against the
unified URL returns 503 (no coord_mgr loaded in the smoke env) —
proves the lifted handler is reachable + the manager_lookup callable
fires correctly.
Verbs still kind-specific (deferred — bodies have substantive
behavior divergence, not just naming): ``send`` (Priority 1
worker dispatch), ``cancel`` (interactive forensics + force flag),
``close`` (interactive close-reason cap+redact+persist), ``open``
(interactive resume vs coord rehydrate), ``events`` (different SSE
replay shapes), ``create`` (interactive attachments vs coord
initial_message), ``list`` / ``saved`` (different response keys).
Stage 2 Priority 0 Step 0.5 follow-on. The handwritten Python
OpenAPI spec already moved to ``/v1/api/workstreams/`` in the URL
sweep commit; this just regenerates ``sdk/typescript/openapi-{server,console}.json``
from those specs so generated TS callers see the new paths.
Bumps the TS SDK to 0.4.0 to flag the URL-shape break for any
1.5.0aN-era consumer of the experimental coord client. Python SDK
needs no change — it never exposed the coord HTTP surface.
TS typecheck + 32 vitest tests pass.
Addresses the eight quality findings the per-priority /review pass
flagged on the registrar refactor. All confirmed by the verifier;
none blocking. Net −186 LOC in this fixup.
q-1, q-9: trim ``session_routes.py`` module docstring + console
``create_app`` comments to the timeless explanation. The Step 0.1 →
0.4 narrative was already stale within the PR that introduced it
(every step had landed by the final commit) and would rot further as
the body-convergence follow-on lands.
q-2, q-5: group the four attachment handlers into an
``AttachmentHandlers`` dataclass exposed as
``handlers.attachments: AttachmentHandlers | None``. The type system
now carries the all-or-none invariant; the parallel four-condition
chain + bare ValueError disappear.
q-3: drop the ``mgr`` and ``adapter`` placeholder kwargs from both
``register_session_routes`` and ``register_coord_verbs``. Pre-
threading them so a future commit avoids "callsite churn" violated
the project's "don't pre-build for the next step" norm — the
body-convergence follow-on will edit the callsites anyway. Drops
``SessionManager.adapter`` for the same reason.
q-4: drop ``SessionRouteConfig`` outright. It existed solely to
carry ``supports_legacy_close``; the registrar now mounts the legacy
close route whenever ``handlers.close_legacy is not None``, matching
the all-Optional convention used for every other handler.
q-6: move ``MockStorage`` from ``tests/test_console.py`` into the
shared ``tests/_coord_test_helpers.py`` and re-import in test_console
+ test_session_routes. No more cross-test-module import.
q-7: trim the two exhaustive route-table set-equality assertions
(``test_coord_shape_mounts_expected_verbs``,
``test_register_coord_verbs_mounts_expected_paths``); replaced with
focused ``test_attachment_routes_mount_when_quartet_provided`` and
``test_close_legacy_mounts_when_handler_provided``. The targeted
ordering tests still catch the actual registrar bugs.
q-8: delete the tombstone comment block where the legacy
``/api/coordinator/`` Routes used to be — per the user's
``feedback_no_tombstone_comments`` norm, deletions don't get
narrated inline.
q-10: rename ``SessionRouteHandlers`` → ``SharedSessionVerbHandlers``
and ``CoordVerbHandlers`` → ``CoordOnlyVerbHandlers`` so the
"shared verbs vs coord-only verbs" symmetry is visible at the type
names. Drop the back-compat aliases since nothing uses them.
The plan called the CHANGELOG callout for the
``/v1/api/coordinator/`` → ``/v1/api/workstreams/`` move
non-negotiable since experimental SDK consumers from 1.5.0aN lose
the URL outright. Adds the path-mapping table under [Unreleased]
so the line lands in the 1.5.0 release notes when the version cuts.
Stable upgraders (1.0 / 1.3 / 1.4) never saw the coord URL prefix
so the change is a no-op for them — the entry says so explicitly.
Stage 2 Priority 0 Steps 0.4–0.7 — collapses the four migration
steps into one commit since they have to land together. The legacy
``/v1/api/coordinator/`` URL prefix never shipped in a stable release
(it appeared in 1.5.0aN experimental), so there's no compat carry-
forward — just rip and replace.
What moves:
- Step 0.4: deletes the eighteen ``Route("/api/coordinator/...")``
entries from ``console/server.py``. Coord traffic now flows
exclusively through the unified ``/v1/api/workstreams/`` shape
mounted via ``register_session_routes`` + ``register_coord_verbs``
(Steps 0.2 and 0.3).
- Step 0.5: rewrites the OpenAPI spec (``console_spec.py``) and
schemas (``console_schemas.py``, ``server_schemas.py``) to
document the new paths. ``test_openapi.py`` parity assertions
swap with them.
- Step 0.6: mechanical URL sweep across the frontend
(``console/static/app.js`` — 9 sites; ``coordinator/coordinator.js``
— 16 sites; ``index.html`` — 1 comment).
- Step 0.7: same sweep across the test suite
(``test_coordinator_endpoints.py``, ``test_coordinator_end_to_end.py``,
``test_coordinator_governance.py``, ``test_coordinator_close_all_children.py``,
``test_coordinator_client.py``, ``test_phase6_endpoints.py``).
Also touched:
- Server-side ``CoordinatorClient`` (``coordinator_client.py``) —
the coord agent's HTTP path for ``close_all_children`` updates
to the new shape.
- Handler docstrings in ``console/server.py`` say
``POST /v1/api/workstreams/...`` not ``/coordinator/...`` so a
``grep`` for a verb's URL lands on the right line.
- ``settings_registry.py`` setting descriptions, ``server.py``
cross-process error message, migration 042 docstring — all
updated to the unified shape.
The handler functions stay named ``coordinator_*`` until the
body-convergence follow-on lifts them into ``session_routes`` with
kind branching behind ``SessionRouteConfig`` flags. URL surface is
the only thing that changes here.
The ``test_session_routes`` route-walk now asserts the legacy paths
are GONE — previously it asserted both shapes coexisted. A future
accidental remount of ``/api/coordinator/`` would fail that test.
Stage 2 Priority 0 Step 0.3 — adds ``CoordVerbHandlers`` +
``register_coord_verbs`` to ``turnstone.core.session_routes`` and
wires the seven coord-only verbs (``children`` / ``tasks`` /
``metrics`` / ``trust`` / ``restrict`` / ``stop_cascade`` /
``close_all_children``) through it on the console.
These verbs are legitimately kind-specific — they read or mutate
state (children registry, parent quota, trust / restrict policy,
cascade controls) that doesn't exist on interactive workstreams —
so they live on a Protocol distinct from
``SessionRouteHandlers``. The unified URL prefix
``/api/workstreams/{ws_id}/`` is shared with the session verbs;
the separate registrar call keeps the kind separation explicit at
the wiring site.
Legacy ``/api/coordinator/{ws_id}/{verb}`` paths stay live during
the transition; both URL shapes resolve to the same handler
function. Step 0.4 deletes the legacy shape.
The route-table walk in ``test_session_routes`` now covers all
eighteen verb pairs (eleven session + seven coord-only) so a
future drift between legacy and unified handlers fails CI.
Stage 2 Priority 0 Step 0.2 — extends ``register_session_routes`` to
cover the per-``{ws_id}`` interaction verbs (``send`` / ``approve`` /
``plan`` / ``cancel`` / ``close`` / ``events`` / ``history`` /
``detail``) and wires ``console/server.py`` to mount coord at the
unified ``/v1/api/workstreams/`` shape.
The legacy ``/v1/api/coordinator/`` paths stay live during the
transition; both URL shapes resolve to the same handler functions.
Step 0.4 deletes the legacy shape outright once the frontend
(Step 0.6) and tests (Step 0.7) move off it.
Handler bodies still live in their server modules — body
convergence (kind branching behind ``SessionRouteConfig`` flags) is
the next Step 0.2 follow-on. Splitting "URL surface unified" from
"handler bodies converged" keeps the soak-able diffs small.
``mgr`` and ``adapter`` registrar arguments are now Optional because
the console builds its coord ``SessionManager`` inside the lifespan
(after app construction). They become required again in the
body-convergence follow-on once the lifted handlers read them.
New ``tests/test_session_routes.py`` covers the registrar's mounting
rules (route ordering, attachment-quartet enforcement, legacy-close
gate) and asserts that each unified path on the console points at
the SAME handler object as its legacy counterpart — the transition
is a pure URL alias, not a fork.
Stage 2 Priority 0 Step 0.1 — introduces
``turnstone/core/session_routes.py`` (``SessionRouteConfig`` +
``SessionRouteHandlers`` + ``register_session_routes``) and rewires
``server.py``'s ``/v1/api/workstreams/*`` route table to mount through
it. Pure scaffolding: handler bodies stay where they are, URL surface
is byte-identical, tests pass unchanged.
Sets up Step 0.2 to lift handler bodies into the registrar and have
the console mount the same shape against its coord manager.
Adds ``SessionManager.adapter`` accessor so the registrar can pick
up the kind adapter without callers re-threading it through every
construction layer.
* feat(core): scaffold SessionManager + SessionKindAdapter Protocol
Stage 1 step 1 — pure addition, no production wiring. Defines the
shape later steps will port the shared mechanics onto: slot
accounting, per-ws-id refcounted rehydrate locks, kind-agnostic
lifecycle; kind-specific event transport + session construction on
the adapter.
Pruned from the earlier Protocol draft (see design brief): per-kind
permission_scope (static handler map is simpler), allows_child_spawn /
quota_policy (deleted in #403), on_child_spawned (coordinator tool
owns children registry), allows_active_focus / active_id / switch
(frontend owns the active-tab state).
* feat(core): port shared session-lifecycle mechanics onto SessionManager
Stage 1 step 2. Adds create / open / close / set_state / close_idle /
get / list_all / count on top of the Step 1 scaffolding. Pure
addition — still no production wiring; the new class doesn't replace
any call sites yet.
Concurrency shape is ported from CoordinatorManager (the more-
complete side): single-phase slot reservation under the manager
lock, per-ws refcounted open-lock to serialize concurrent lazy
rehydrate, placeholder workstreams count toward max_active but can't
evict each other. WSM's two-phase eviction outside the lock is not
carried over; it had a window where a burst of creates could silently
exceed max_active.
Deletions (vs. the union of the two old managers):
- "refuse to close last workstream" guard — handled by the
dashboard; only existed to protect the now-deleted default startup
workstream.
- active_id / switch / get_active — frontend owns focus; server-side
duplicate state is gone.
- _active_coords presence cache — defer measurement to Step 4; if it
pays for itself at realistic cluster sizes, the CoordinatorAdapter
can maintain it by observing emit_* calls.
- Children registry + reverse index — coordinator tool owns this,
manager stays kind-agnostic.
Skill resolution (name → template_id + applied_version) is now
shared via SessionManager._resolve_skill, so WSM's pre-resolve-at-
callsite pattern and CM's internal-lookup pattern converge. Callers
pass the skill name; the manager does the lookup once.
26 smoke tests cover create eviction + overflow, concurrent-create
cap, persist/session rollback, open for missing/deleted/wrong-
kind/wrong-user rows, concurrent-open serialization, close unblocks
UI + emits closed, set_state + storage + adapter observer,
close_idle, list_all ordering, count, eviction fires adapter
transport, node_id passthrough.
* feat(core): add InteractiveAdapter for SessionManager
Stage 1 step 3. Adapter that bridges SessionManager to the node's
interactive transport:
- emit_created/state/closed → pushes onto the process-wide SSE
global_queue (same shape current server.py handlers produce inline)
- cleanup_ui → ports WorkstreamManager._cleanup_ui body: unblock
_approval_event / _plan_event / _fg_event, broadcast ws_closed to
per-UI listener queues (with full-queue fallback), cancel + close
the session
- build_ui/build_session → delegate to injected factories
(ui_factory builds WebUI, session_factory is the existing closure
from server.py with judge_model + memory_config captures)
Also extends SessionKindAdapter.build_session with **extra passthrough
so interactive callers can pass judge_model per-call without polluting
the manager API; and adds a reason= kwarg to emit_closed so the
frontend's "evicted" special-case keeps working (frontend doesn't
differentiate "idle" from "closed", so close_idle collapses into
close()).
14 new adapter tests cover wire payload shape, queue.Full tolerance,
cleanup_ui event unblocking + listener broadcast + queue-full
fallback, session cancel+close, graceful handling of stub UIs / None
session, and kwarg passthrough to the session factory.
* feat(console): add CoordinatorAdapter for SessionManager
Stage 1 step 4. Coordinator-side SessionKindAdapter implementation:
- emit_created/state/closed → delegate to the existing
ClusterCollector.emit_console_ws_* methods (same wire shape the old
CoordinatorManager emitted inline)
- cleanup_ui → ports the listener-queue + approval/plan event
unblocks from CoordinatorManager._cleanup, with queue-full
fallback so an unresponsive browser tab can't wedge close
- build_ui/build_session → delegate to injected factories; session
factory doesn't accept client_type so we strip it at the adapter
boundary
Collector emission exceptions are swallowed (same policy as today's
inline fan-out — dashboard lag on one tick is preferable to breaking
the lifecycle path).
Intentionally out of scope: the children registry (_children /
_child_to_coord) stays in the coordinator tool when wired in Step 5;
the _active_coords lock-free presence cache is deferred pending a
measurement at realistic cluster sizes. 10 new tests cover transport
payloads, collector-exception tolerance, cleanup_ui event unblock +
listener broadcast + queue-full eviction, construction passthrough.
* feat(server): wire interactive server.py to SessionManager
Stage 1 step 5a. Production-path swap: WorkstreamManager →
SessionManager(InteractiveAdapter(...)).
- Construction at server startup: build the adapter with the
process-wide global_queue, a WebUI ui_factory closure, and the
existing session_factory. SessionManager gets storage + max_active.
- Default startup workstream wiring removed (the CLI-REPL leftover
flagged in the handoff's "Convergence is also a pruning
opportunity" section). --resume now lazily creates a workstream
scoped to the resumed content; no workstream at all if --resume
isn't given. The dashboard handles the 0-ws state.
- HTTP handler mgr.create() calls switched to the new kw-only
signature (user_id, name, model, skill, ws_id, client_type,
judge_model, parent_ws_id). ui_factory/skill_id/skill_version/kind
no longer threaded through — adapter handles UI construction and
manager resolves skill internally.
- Dropped the mgr.last_evicted block in the /new handler (adapter
emits ws_closed:evicted automatically on capacity eviction).
- mgr.max_workstreams → mgr.max_active.
- Added active_id / switch / switch_by_index / get_active / index_of
/ eviction_count to SessionManager because turnstone/cli.py uses
them extensively; the handoff's "delete unless there's a live
caller" rule flips here — CLI is a live caller.
Test fixtures across 9 files updated to build SessionManager +
InteractiveAdapter rather than WorkstreamManager. test_workstream.py
stays unchanged (it tests WSM directly; it'll be deleted in step 5d
alongside the class itself).
Full pytest: 4528 passed. Ruff + mypy clean. Next: 5b (console-side
wiring, with the children-registry relocation to the coordinator
tool).
* feat(console): wire console server to SessionManager
Stage 1 step 5b. Production-path swap: CoordinatorManager →
SessionManager(CoordinatorAdapter(...)).
- CoordinatorAdapter now owns the coord-specific bits that were bolted
onto the old CoordinatorManager: the children registry (forward +
reverse index), the lock-free active-coords presence cache, the
cluster-event fan-out thread, and the worker-dispatch path
(send / _spawn_worker). The shared SessionManager stays kind-agnostic.
- Added CoordinatorAdapter.attach(mgr) for late-binding the owning
manager (the manager's ctor takes the adapter, so the dependency has
to break here). Used inside _rebuild_children_registry for the tenant-
filtered SQL query, inside send/dispatch for mgr.get(ws_id), and
inside the fan-out seed path for mgr.list_all().
- emit_created now seeds the children registry + active-coords slot AND
calls _rebuild_children_registry (covers both create — empty query —
and open/rehydrate, where the subtree is persisted). emit_closed
drops both entries. Collapses the three old call-sites in
CoordinatorManager's create/open/close into one per-event hook.
- Console server.py builds the manager via:
coord_adapter = CoordinatorAdapter(collector=..., ...)
coord_mgr = SessionManager(coord_adapter, storage=..., max_active=...,
node_id=ClusterCollector.CONSOLE_PSEUDO_NODE_ID)
coord_adapter.attach(coord_mgr)
ConsoleCoordinatorUI._coord_mgr = coord_mgr
app.state.coord_adapter = coord_adapter
- HTTP handler call-site updates:
- coord_mgr.create drops initial_message; the handler now calls
coord_adapter.send(ws.id, initial_message) after create so the
worker spawn stays out of the shared manager.
- coord_mgr.open_admin(ws_id) → coord_mgr.open(ws_id, user_id="",
admin=True). Matches SessionManager.open's unified signature.
- coord_mgr.list_for_user(uid) inlined as a list comp on list_all()
(SessionManager doesn't expose the filter; two callers).
- coord_mgr.children_snapshot / send → coord_adapter.*.
- coord_mgr.cancel stays (now lives on SessionManager from 5a).
- ConsoleCoordinatorUI.on_state_change now flows state transitions
through ConsoleCoordinatorUI._coord_mgr.set_state, mirroring the
WebUI pattern. The old _on_state_observer / _on_rename_observer
closures the manager used to install are dead code now; leaving the
fields in place for 5d cleanup.
- Lifespan shutdown calls coord_adapter.shutdown() (was coord_mgr.
shutdown()) and resets ConsoleCoordinatorUI._coord_mgr on teardown.
Test fixture updates in _coord_test_helpers, test_coordinator_end_to_end,
test_coordinator_endpoints, test_phase6_endpoints: build SessionManager
+ CoordinatorAdapter in _build_mgr, set app.state.coord_adapter, switch
mgr.register_children / mgr.children_snapshot tests to mgr._adapter.*,
and rewrite test_open_admin_uses_open_admin to assert the unified
open(user_id="", admin=True) call shape.
Full pytest: 4486 passed. Ruff + mypy clean. Next: 5d (remove
CoordinatorManager + WorkstreamManager class bodies and their test
files).
* feat(core): delete WorkstreamManager + CoordinatorManager classes
Stage 1 step 5c + 5d. Final step of the unification — the legacy
classes and their test files go away now that every production
caller has been ported.
- Delete turnstone/console/coordinator.py entirely (CoordinatorManager
class + the _enqueue_on_ui helper, which CoordinatorAdapter now hosts
its own copy of).
- Trim turnstone/core/workstream.py to just the Workstream dataclass +
WorkstreamKind + WorkstreamState. ~385 lines of WorkstreamManager
logic gone; the remaining shape is pure data types shared by both
managers.
- Delete tests/test_workstream.py (WSM-specific) and
tests/test_coordinator_manager.py (CM-specific).
- Wire turnstone/cli.py to SessionManager + InteractiveAdapter, same
pattern as turnstone/server.py. The CLI's WorkstreamTerminalUI uses
manager.set_state + manager.active_id — both preserved on
SessionManager (CLI is a live caller that keeps the focus API
honest, per the handoff's "delete unless it pulls its weight" rule).
- Add an optional manager-level ``_on_state_change`` observer hook
restored for the CLI's background-attention notification (the web
path uses the adapter's emit_state; this hook covers callers that
don't consume SSE).
- Drop dead ``_on_state_observer`` / ``_on_rename_observer`` fields
from ConsoleCoordinatorUI — the old CoordinatorManager installed
them; SessionManager/CoordinatorAdapter handle fan-out directly.
Vulture @ 80% confidence: zero unused symbols across the new
SessionManager + adapter files. Ruff + mypy clean (170 files).
Full pytest (excluding tests/live): 4414 passed.
Net across the whole Stage 1 branch: one unified SessionManager +
adapter Protocol replaces two ~500-line parallel managers + a
~600-line CoordinatorManager, and the interactive + coordinator
transports stay cleanly separated at the adapter boundary.
* refactor(auth): drop workstream row-level ownership gates
Turnstone is a trusted-team tool (per #400). user_id stays as
metadata for audit + display; it no longer rejects requests. Scope-
level auth via admin.workstreams / admin.coordinator tokens is the
only gate now.
Solves sec-1 (cross-tenant delete via collision on caller-supplied
ws_id, because the gate was half-implemented) and sec-2 (blank-sub
JWT bypass on empty-owner rows). Net: 359 lines of defensive
empty-string comparisons and admin=True bypass plumbing deleted.
* fix(core): serialize set_state vs close + worker spawn
Three concurrency fixes from the multi-stage review:
- bug-3: set_state now looks up ws under self._lock and gates its
storage write on ws._closed (a new tombstone flag). close() sets
ws._closed=True and does its storage write under ws._lock. A
set_state that acquires ws._lock after close sees the tombstone
and skips its write instead of resurrecting the closed row.
- bug-1: _spawn_worker wraps the check-and-spawn in ws._lock so two
concurrent send() HTTP requests can't both observe "no live worker"
and start duplicate worker threads on the same ChatSession.
- bug-2: replaces Thread.is_alive() as the reuse gate with an
explicit ws._worker_running flag. The flag is set before the worker
thread starts and cleared in its finally block — both under
ws._lock. Using is_alive() left a narrow window where the worker
could exit between the check and a queue_message call, stranding
the user's message with no consumer.
perf-2 (lock-held-across-DB-write) is accepted as-is: per-ws
serialization of state transitions behind a DB round-trip is real
cost but bounded — a given ws's state flips happen sequentially on
its worker thread anyway. Dropping ws._lock around the DB write
would reintroduce the bug-3 race.
Full pytest: 4401 passed. Ruff + mypy clean.
* refactor(core): drop _resolve_skill from SessionManager
Skill resolution (name → template_id + applied_version) moves out of
the shared manager and back to the HTTP handlers that own the
create request. The interactive handler already resolved skill_data
+ applied_skill_version for other purposes (model override, judge
config, post-create session seed) and was passing the name to
SessionManager which then redundantly re-resolved via
get_skill_by_name + count_skill_versions — two wasted DB round-trips
per create on a user-visible latency path.
- SessionManager.create: accepts skill_id + skill_version as
already-resolved kwargs; _resolve_skill helper deleted.
- turnstone/server.py create_workstream: passes the skill_id /
applied_skill_version it already computed.
- turnstone/console/server.py coordinator_create: pre-resolves
inline (parity with interactive) before calling coord_mgr.create.
Fixes perf-1 (redundant skill queries per create), q-4 (divergent
skill-version computation between manager and handler), q-5
(coordinator-specific lookup on the shared manager surface).
Full pytest: 4401 passed. Ruff + mypy clean.
* refactor(adapters): extract shared cleanup_ui + drop dead child-registry methods
Both InteractiveAdapter.cleanup_ui and CoordinatorAdapter.cleanup_ui
(plus their _broadcast_ws_closed_to_listeners helpers) were byte-identical.
Pull them into turnstone/core/adapters/_ui_cleanup.py:cleanup_session_ui
so the two adapters delegate to one implementation.
Also drop CoordinatorAdapter.register_children (only test callers — now
use _seed_children in tests/_coord_test_helpers.py) and _add_child
(zero callers anywhere).
* refactor(adapters): symmetric attach() + fail-loud on unattached manager
Add InteractiveAdapter.attach(manager) + .manager property mirroring
the coord-side pattern. CLI (cli.py) now uses cli_adapter.attach(manager)
instead of the _mgr_ref list-ref late-binding hack; server.py picks up
the same call for consistency.
CoordinatorAdapter.send / _rebuild_children_registry /
_prime_children_from_snapshot no longer silently return when
self._manager is None — raise RuntimeError so a forgotten attach() at
startup fails loud instead of dropping the whole fan-out.
* docs: replace stale WorkstreamManager / CoordinatorManager references
Both classes were deleted in 965e0b6; prose docstrings across the
codebase still named them. Update to SessionManager (or describe the
collapsed-into-one-class architecture where the distinction matters).
Leaves the 'Ported from …' historical markers in session_manager.py /
coordinator_adapter.py / interactive_adapter.py intact — those are
deliberate pointers back to the pre-unification code.
* fix(core): atomic close_if_idle + batch pop under one lock
bug-5: SessionManager.close_idle re-checked ws.state == IDLE outside
the lock, so a pending tool result could flip state IDLE→RUNNING
between the snapshot and close() acquiring self._lock. Add
_close_if_idle_locked that tests state + pops under self._lock.
perf-5: drop the per-victim self._lock acquisition; collect + pop the
whole batch in one acquisition, then run cleanup_ui / storage write /
emit_closed outside the lock.
* perf(coord): split emit_created / emit_rehydrated to skip storage query on fresh creates
CoordinatorAdapter.emit_created was unconditionally calling
_rebuild_children_registry (storage.list_workstreams with
parent_ws_id=... limit=10001) on every create, even for fresh-create
paths that provably have zero children.
Add emit_rehydrated to the SessionKindAdapter Protocol. SessionManager
.create still calls emit_created; .open (lazy rehydrate) now calls
emit_rehydrated. CoordinatorAdapter.emit_created seeds the registry +
fan-out but skips the rebuild; emit_rehydrated seeds + rebuilds + fans
out. InteractiveAdapter.emit_rehydrated delegates to emit_created (no
children-registry on the interactive transport).
* perf(coord): fold _active_coords into _children_lock + mutate payload in place
perf-4: _active_coords used a copy-on-write dict-swap pattern so the
fan-out dispatch could read it lock-free, but _dispatch_child_event
already re-validates the parent under _children_lock anyway — the
lock-free snapshot was premature. Replace with a plain dict read+write
both under _children_lock; install and remove collapse to one-liners.
Value also drops the user_id half — dead after a46dab1 removed
row-level ownership gates — so _active_coords is now just
coord_ws_id → ui.
perf-6: _enqueue_on_ui was doing {**payload, "ws_id": coord_ws_id} on
every dispatch. The dispatch path owns payload and doesn't reuse it —
mutate in place.
* test(coord): add adapter tests for worker dispatch + children registry + fan-out
Fills the coverage gap on CoordinatorAdapter — the review (q-3) flagged the
coord-specific concurrency paths ported from the deleted CoordinatorManager
as untested. Three new test classes:
- TestCoordinatorAdapterWorkerDispatch: _spawn_worker reuse gate, queue.Full
backpressure, concurrent-call bug-1 reproducer (two threads → exactly one
worker via ws._lock + _worker_running), finally-clears-flag.
- TestCoordinatorAdapterChildrenRegistry: registry seed on emit_created vs
emit_rehydrated rebuild, _pop_coord_registry_locked reverse-index cleanup,
_merge_child_ids_locked idempotency, _prime_children_from_snapshot merge.
- TestCoordinatorAdapterDispatchChildEvent: unknown-parent drop, ws_created
fan-out, cluster_state / ws_closed reverse-index routing, perf-6 in-place
ws_id stamp.
* fix: regressions flagged by ultrareview
Verify stage of the cloud review surfaced 6 confirmed regressions
from Stage 1's adapter layer. Fixing together since they share the
same root cause (plumbing moved into adapters without retiring the
old emission paths).
- Interactive adapter emit_created / emit_state / emit_rehydrated
become no-ops. The create_workstream HTTP handler still fires
ws_created (after attachment validation, per the pre-Stage-1
"no phantom events on rejected upload" contract); WebUI
_broadcast_state still fires ws_state with the full payload
(tokens + context_ratio + activity). Firing from the adapter too
was duplicating both events. Also closes the phantom-ws-created
regression (adapter fired before attachment validation ran).
- emit_closed Protocol gains a ``name`` kwarg; the adapter is the
sole emitter for ws_closed on interactive now, and the frontend
eviction toast needs the name. Manager passes ws.name from
close() / create()+open() eviction / close_idle paths.
- _idle_cleanup_thread stops firing its own reason="idle" ws_closed
— close_idle already fires via the adapter with reason="closed",
and the frontend never differentiated the two anyway.
- close_workstream_endpoint fix: "Cannot close last workstream" 400
was a stale error (the guard went away with the default-startup
workstream). Return 404 on close() == False (which now means the
ws was already closed or unknown). Also switches the audit actor
from _require_ws_access's stored owner to _auth_user_id — the
stored owner is metadata post-#400, so attributing actions to it
misrepresents who actually did them.
- CLI /ws close mirrors the same stale-error fix.
- SessionManager.close now calls storage.delete_workstream_override
alongside update_workstream_state, same as the old
WorkstreamManager.close did. Without it overrides leak until
tombstone cleanup. close_idle does the same.
- SessionManager._reserve_and_install_locked records the eviction
on turnstone.core.metrics so the global eviction counter keeps
working. Old WSM did this inline; the unification dropped it.
- ConsoleCoordinatorUI.on_rename now fans out to the cluster
collector via a new class attribute ``_collector`` (set at
console startup alongside ``_coord_mgr``). The old
``_on_rename_observer`` plumbing went away with
CoordinatorManager and the "adapter emit_console_ws_rename runs
from whichever code path renames" comment was aspirational —
nothing actually did it.
Full pytest: 4375 passed (tests/live + test_server_live.py excluded;
both pre-existing live-backend failures unrelated to this branch).
Ruff + mypy clean.
* refactor(ui): extract SessionUIBase for shared UI scaffolding
Direct response to review feedback that the unification wasn't
merging enough of the two workstream kinds. WebUI (node) and
ConsoleCoordinatorUI (console) both:
- Keep a per-UI list of SSE listener queues guarded by a lock
- Block a worker thread on _approval_event / _plan_event
- Fan enqueued events out with the same ws_id-stamping pattern
- Resolve approvals / plans with the same broadcast-then-signal
pattern
All of that now lives once in turnstone/core/session_ui_base.py.
Both UIs subclass SessionUIBase; kind-specific bodies (WebUI's
per-UI metrics + _broadcast_state + intent-verdict bookkeeping,
ConsoleCoordinatorUI's collector fan-out) stay in the subclasses.
WebUI.resolve_approval still overrides the base (it adds intent-
verdict updates) but now calls super() for the shared broadcast +
event-set steps. Same shape as the other approval/plan hooks:
subclasses extend, base provides skeleton.
Net file-level: +156 LOC for the base, -144 LOC across the two
subclasses. The raw number is unexciting — but there's now a
single source of truth for the listener + blocking-gate machinery,
and bugs (like the duplicate ws_created / ws_state events that
prompted this refactor) can't arise from the two implementations
drifting.
Full pytest: 4375 passed. Ruff + mypy clean.
* refactor(ui): move metrics + verdict bookkeeping into SessionUIBase
Second pass at unifying the two UIs. Per-workstream metrics
accumulators (token counts, tool-call counts, context ratio,
activity tracking), intent-judge verdict cache + pending-decision
list, and the verdict-persistence path all move to SessionUIBase.
Before: WebUI tracked all of it; ConsoleCoordinatorUI tracked none
of it (a comment on the old on_intent_verdict literally admitted
the deferral — "skip the persistence + late-decision plumbing that
WebUI does"). Coord sessions never got verdict rows in storage, never
had a user_decision stamped, and the dashboard had no way to show
coord token usage because the data wasn't captured.
Now the base class captures the data and persists the rows for
every kind. Kind-specific broadcast (WebUI's _broadcast_state with
rich per-UI payloads) stays on WebUI; prometheus counters on the
node (_metrics.record_judge_verdict) stay on WebUI's on_intent_verdict
override. Everything else shared.
Behaviour change worth flagging: coord sessions now write
intent_verdicts and output_assessments rows for every judge call
and every output-guard warning. Previously silent; the storage rows
now exist and any future coord-dashboard surface can read them.
Shape of the unification:
- resolve_approval: was overridden on WebUI (intent-verdict decision
propagation); now lives on the base. Both kinds inherit unchanged.
- on_intent_verdict: WebUI overrides only to add _metrics.record_*;
rest of the body is the base.
- on_output_warning: was on both separately; fully base-shared now.
Full pytest: 4375 passed. Ruff + mypy clean.
* fix: regressions flagged by second-pass review
Three confirmed findings with direct fixes + a dedicated test file
for SessionUIBase (was previously uncovered).
bug-1 — Coord approve_tools didn't reset _last_verdict_decision or
clear _llm_verdicts between approval rounds. WebUI did (inline).
Coord inherited SessionUIBase.on_intent_verdict which stamps via
the decision flag, so after the first resolve every subsequent
round's verdicts were stamped with the prior round's user_decision
before the user had decided the new round.
Fix: add SessionUIBase._reset_approval_cycle() clearing both under
_ws_lock; call from the top of both subclass approve_tools methods.
Single-source invariant — can't drift again.
sec-1, sec-2 — delete_workstream_endpoint and open_workstream's
rehydrate path recorded the audit row under the stored ws.user_id
("owner_uid") rather than the authenticated caller. With row-level
ownership gating gone (a46dab1), any team member acting on a peer's
workstream produced an audit row naming the victim as the actor.
Fix: pass _auth_user_id(request) as the audit actor, matching the
pattern close_workstream already follows.
q-2 — SessionUIBase had no direct tests. The new
tests/test_session_ui_base.py covers listener fan-out, approval +
plan blocking gates, intent-verdict cache + FIFO eviction, verdict
persistence paths, output-guard persistence, the reset-between-rounds
invariant (bug-1 regression test), a cross-subclass test that
verifies BOTH WebUI.approve_tools and ConsoleCoordinatorUI.approve_tools
call _reset_approval_cycle (verified it fails without the fix), and
a concurrent enqueue/register smoke.
Full pytest: 4395 passed (+20 new). Ruff + mypy clean.
* fix: PR #408 review findings from copilot + code-quality
Three substantive fixes + mechanical side-effect-in-assert cleanup.
Copilot findings:
- session_ui_base.py: on_intent_verdict had a race with
resolve_approval. Previously acquired _ws_lock twice (read decision
→ release → if unset, acquire again to append). resolve_approval
could interleave between the two acquisitions, swap-and-clear the
pending list and set the decision — our verdict then got appended
to the fresh (empty) list and stamped with the NEXT round's
decision on the following resolve. Fix: decision-check + append
under ONE acquisition; storage UPDATE (if decision already set)
runs outside the lock. New regression test counts lock
acquisitions during on_intent_verdict and fails if the two-phase
pattern returns.
- server.py close_workstream_endpoint: comment said "treat as
already-closed success" but handler returned 404. Comment
rewritten to match the 404 behaviour ("the ws isn't tracked here"
is the only reachable meaning for close() → False now).
- test_session_ui_base.py concurrency smoke: the test ended with
``pytest.assume = lambda ...`` — a leftover that mutates pytest
globals and can surprise other tests. Replaced with explicit
``not is_alive()`` assertions so the "threads completed cleanly"
intent survives -O optimization stripping.
Code-quality (assert side-effects):
Six ``assert mgr.open(...)`` / ``assert mgr.close(...)`` in
test_session_manager.py stripped under ``python -O``. Mechanical
fix: extract to local before asserting.
Ignored the two "Protocol method body is `...`" flags — that's the
standard Protocol idiom; replacing with ``pass`` or
``NotImplementedError`` changes typing semantics.
Full pytest: 4396 passed.
* chore(coord): remove spawn-quota subsystem
The quota gate was operator-level safety per its own comments, not a
security boundary, and never fired in a week of heavy use. Runaway
coordinator spawns are already bounded by max_active slot exhaustion,
which surfaces to the coord LLM as a tool error — same operational
shape, one fewer moving part. Precedes the Stage 1 SessionManager
unification so the coord tool doesn't inherit quota bookkeeping.
Upgraded deployments with the three removed settings persisted will
log three "Skipping invalid setting" warnings on startup and
otherwise degrade cleanly; a follow-up migration to delete the rows
would silence that noise.
* chore(migrations): drop stale coord spawn-quota settings rows (047)
Clears persisted rows for the three ConfigStore keys removed in the
previous commit so upgraded deployments don't log "Skipping invalid
setting" warnings on every startup. Downgrade is a no-op — the rows
were operator-set values, and a rollback to pre-1.5.0 code falls back
to the registry defaults for any key not present.
* fix(coord): render markdown on history reload
The coordinator chat's history-load path piped assistant content
through ``appendText`` → ``appendMsg(role, esc(text))``, which dumps
escaped raw text into the message body without ever calling the
markdown converter or the post-render hooks (highlight.js, mermaid,
KaTeX). Live streaming uses ``streamingRender`` /
``streamingRenderFinalize`` which DO render markdown, so a fresh
stream looked correct but a page-reload / reconnect surfaced every
table, code fence, and math block as literal characters.
Now the history loop dispatches by role: assistant + reasoning go
through ``streamingRenderFinalize`` (mirrors what live streaming does
on stream_end); tool messages keep ``appendToolResult``; user / system
stay on ``appendText`` since they're typed verbatim and don't carry
markdown structure.
* fix(coord): keep reasoning role on plain-text path on history replay
The history loop routed reasoning role through streamingRenderFinalize,
but live streaming renders reasoning tokens via textContent
(appendReasoningToken). History replay would render reasoning as
markdown while a fresh stream rendered it as plain text — inconsistent
look and unnecessary hljs / mermaid / KaTeX work on reasoning content.
Reasoning now uses appendText on replay, matching the live path.
Addresses Copilot review feedback on PR #402.
* fix(server): trusted-team workstream visibility on listing endpoints
The per-user filter on /v1/api/workstreams, /v1/api/dashboard, and
/v1/api/workstreams/saved (PR #375's _visible_workstreams helper) was
written for a multi-tenant SaaS threat model that doesn't match how
turnstone gets deployed. In a self-hosted, trusted-team install the
filter created friction without preventing the relevant threats — and
hid the auto-created name="default" startup workstream from every
web user, leaving fresh installs staring at a blank dashboard.
Listing endpoints now return the cluster-wide set to any authenticated
caller. Per-workstream MUTATIONS (/send, /close, /open, /title,
/delete, /refresh-title) keep their independent ownership checks — the
cross-tenant guards from PR #375 stay in force on those handlers (see
TestCrossTenant{Delete,Approve,Close,Title,Open}). Listing only
exposes metadata (name, state, kind, message_count); message history
still requires the per-workstream gate on /history.
Resuming a saved workstream still goes through /open's owner check, so
the metadata-leak surface ends at "you can see workstream X exists" —
not at any actionable cross-user capability.
The console collector's service-scope is now load-bearing only for the
SSE event stream gate (/v1/api/events/global); kept anyway as belt-
and-braces.
If turnstone is ever deployed as a true multi-tenant SaaS, the right
boundary is a real ``tenant_id`` column with row-level filtering at
the storage layer, not the empty-user_id heuristic this used to apply.
Tests updated to assert the new contract: listing returns all owners;
mutation gates unchanged.
* fix(server): repair test mocks + tighten docstrings on listing endpoints
- tests/test_auth.py: TestServerAuth + TestServerLogin mocks now set
kind / parent_ws_id / user_id explicitly so /v1/api/workstreams JSON-
serializes them. Bare MagicMock attributes return another MagicMock
that fails json.dumps and surfaces as 500.
- turnstone/server.py: list_saved_workstreams docstring corrected to
describe what the endpoint actually returns (summary metadata, not
history) and to spell out that ownerless persisted rows are claimable
by any authenticated caller via /open — consistent with the trusted-
team model the listing endpoints assume. Same callout added next
to the open_workstream ownership-gate block. Comments throughout
rewritten to be timeless (no "previously" / PR-number references).
- tests/test_server_authz.py: TestSaved... docstring matches the actual
/open behavior for orphan rows (claimable by any authenticated
caller, not a separate admin path).
* fix(chat): collapse phantom whitespace + tighten paragraph rhythm in markdown body
The assistant chat body was rendering 30-50px gaps between every
section. Two compounding causes:
1. ``.ts-msg-body`` had ``white-space: pre-wrap`` on the markdown
container. The custom regex-based markdown converter
(renderer.js) leaves ``\n`` text nodes between block siblings —
pre-wrap rendered every one of those as visible vertical space,
stacking ~14-16px between every heading/paragraph/katex-display.
2. No ``.ts-msg-body p`` margin override, so paragraphs fell back to
browser-default 1em top + 1em bottom (~28px stacked between any
two paragraphs). Headings already had a tight ``8px 0 4px`` rule;
paragraphs were the outlier.
Switched the body to ``white-space: normal`` and added a
``.ts-msg-body p { margin: 6px 0 }`` rule that matches the heading /
list / blockquote rhythm. Mirrored the paragraph rule on the design-
v1 ``.msg-body`` selector so both legacy and v1 surfaces stay in sync.
``<pre>`` blocks have ``white-space: pre`` built in so fenced code
still preserves formatting. Mid-stream partial fences (before the
closing ``\`\`\`` arrives) render as collapsed text for one frame and
then snap back when the next render tick wraps them in ``<pre>`` —
acceptable trade vs. the persistent gap regression.
User-typed messages render through ``.msg-user-text`` (a separate
DOM path), so this only affects assistant markdown output.
* fix(chat): preserve inline <code> whitespace under white-space: normal body
The body's ``white-space: normal`` (which collapses phantom inter-block
``\n`` text nodes from the markdown converter) inherits to inline
``<code>`` and silently collapses multiple spaces inside backtick
spans. ``<pre>`` blocks rely on the user-agent ``pre { white-space:
pre }`` rule and are unaffected; only bare inline code needs an
explicit override.
Adds ``white-space: pre-wrap`` to ``.ts-msg-body code`` (chat.css) and
the design-v1 ``.msg-body code`` selector so backtick-wrapped code
spans render verbatim while still wrapping on long lines.
Addresses Copilot review feedback on PR #401.
* feat(coord): saved coordinators surface + shared session-card primitives
The console home view now lists explicitly-closed coordinators in a
"Saved Coordinators" card grid below the active list. Click a card →
POST /v1/api/coordinator/{ws_id}/open then navigate; capacity issues
surface as a toast instead of a broken detail page. Card click is
de-duped by an `is-busy` class so rapid double-clicks don't fire
parallel resurrects.
GET /v1/api/coordinator/saved is the new backend endpoint (mirrors the
interactive list_saved_workstreams shape). Filters at the SQL layer
to state='closed' via a new optional `state` parameter on
list_workstreams_with_history (added to the protocol + both backends);
also drops any rows currently loaded into coord_mgr as defence in
depth. The blocking storage call + the lock-acquiring list_all are
offloaded via asyncio.to_thread to match coordinator_create's pattern.
CoordinatorManager._open_impl now allows resurrect of state='closed'
rows (deleted is still a tombstone). The DB state-flip on resurrect
that the first cut had is gone — it raced concurrent close()s and the
next set_state() call syncs the DB naturally; the saved list filters
already keep a still-loaded coordinator from appearing as a saved
card even when its on-disk state lags.
Frontend dedup that paid for the saved surface ships in the same diff:
- shared_static/cards.css: lifted from ui/static/style.css so both
surfaces share the basic card primitive (delete-mode rules stay
interactive-only until coordinator gets the same UX)
- shared_static/cards.js: new renderSessionCard(sess, opts) helper
used by both renderSavedWorkstreams (interactive) and
renderSavedCoordinators (console)
- shared_static/utils.js: formatRelativeTime moved here from
ui/static/app.js
Coordinator landing visual fixes folded in:
- .home-section-title now uses var(--accent) so the COORDINATORS
heading reads as a peer of the NODES heading
- .home-panel dropped its bg/border/padding so the composer is no
longer double-framed (matching the dashboard-composer feel)
- "Active coordinators" → "Saved Coordinators" rename + "Coordinators"
on the active list
ws_closed SSE handler now gates on the closed ws's kind so interactive
closes don't spam /v1/api/coordinator/saved on busy clusters.
loadSavedCoordinators in-flight de-dup coalesces close-event bursts to
one fetch instead of N.
Tests cover: caller-scoping, admin sees-all, blank-uid fail-closed,
loaded-coordinator filtering, state filter (idle rows excluded), plus
the manager-level open-resurrect / open-refuses-deleted contracts.
Closes the bug-{1,2,3}, perf-{1,2,3,4}, sec-{1,2}, q-{1,2,3,4,5,6,7}
findings from the prior multi-stage review.
* fix(design): restore amber accent on the v1 design system
The Claude Design handoff swapped the accent hue to teal (h=182).
Walking back to amber (h=75) — turnstone's original brand colour.
Lightness + chroma bumped slightly (0.62→0.7, 0.10→0.13) so the
restored gold matches the visual weight of the legacy #e5a042 token.
Hue map header comment updated to record what happened so the next
person doesn't repeat the swap. Only surfaces with data-design="v1"
on <html> pick this up — currently just turnstone-server's webui.
* chore: gitignore design_ideas/ and .claude/ dev directories
design_ideas/ holds personal Claude Design handoff scratch + reference
HTML; .claude/ holds per-user Claude Code state (worktrees, settings,
plugin caches). Neither belongs in version control.
* fix(coord): address PR #399 review nits
- tests/test_coordinator_endpoints.py: split `assert mgr.close(ws.id)`
in `_seed_closed_coord_with_history` so the close call always runs
even under `python -O` (asserts stripped). Same fix in
test_coordinator_manager.py's `test_open_refuses_deleted_coordinator`
for the open() and open_admin() calls.
- shared_static/cards.css: `.card-wsid` now reads `var(--font-mono, "IBM
Plex Mono", monospace)` so design-v1 surfaces pick up the JetBrains
Mono token while console (still pre-v1) keeps the literal fallback.
* feat(auth): inline refresh response + sessionStorage rehydrate hardening
The proactive refresh path now consumes the /refresh response body
inline (permissions + exp), eliminating the chained /whoami round-trip
and the brief stale-sessionStorage window after refresh succeeds but
before whoami completes.
Adds AbortController + _loggedOut guards to the whoami fetch so a
logout fired mid-flight cannot re-populate sessionStorage after it
clears. A non-OK whoami on tab restore now explicitly clears
sessionStorage instead of silently leaving stale cosmetic permissions
(server-side identity gone → UI gating reflects it on next render).
Surfaces window.permissionsReady (one-shot promise) so permission-
gated UI can await the initial whoami's completion instead of guessing
a setTimeout duration.
Tests cover the new refresh response shape, the existing leeway path,
the storage-failure fallback, and the no-perms 403 path.
Closes the bug-3 / perf-4 / sec-1 / q-6 findings from the multi-stage
review of the prior uncommitted change set.
* fix(auth): guard whoami superseding race in _scheduleRefreshFromWhoami
_scheduleRefreshFromWhoami is invoked from several entry points
(initial page load, _onSuccess, BroadcastChannel "login"/"refresh",
_tryRefresh fallback). Two firing in quick succession could let an
older slow whoami land after a newer one and clobber its effects —
clearing permissions right after a successful login, or rescheduling
the refresh timer off stale exp.
Now aborts any prior _whoamiAbort before starting a new request and
guards the .then's _storePermissions / _scheduleRefreshAt with a
`_whoamiAbort === ctrl` check so a late arrival from a superseded
call is fully neutralised.
Addresses Copilot review feedback on PR #398.
* feat(providers): add gpt-5.5 and gpt-5.5-pro capability entries
OpenAI announced gpt-5.5 on 2026-04-23 (ChatGPT/Codex first, API
"very soon"). Mirror the gpt-5.4 / 5.4-pro capability shape: 1M
context, native tool search, vision, xhigh effort; pro is
always-reasoning with no temperature and medium/high/xhigh only.
No provider-logic changes needed — OpenAI announced no API-surface
changes vs 5.4. Cache retention already covers 5.5 via the existing
startswith("gpt-5") prefix rule.
* test(providers): cover gpt-5.4-pro and gpt-5.5-pro in cache retention test
Pro variants share the same gpt-5 prefix and should keep 24h
retention; explicit coverage guards against regressions if the
prefix rule narrows in the future.
Remove the weekly Trivy scan job and the .trivyignore exclusion file.
The scanner has been flagging base-image CVEs that require no action
on our part (upstream-only fixes) and has provided no actionable
signal, while breaking CI on an ongoing basis.
* feat(auth): cookie refresh endpoint, JWT leeway, coord-token observability
Three robustness wins around the auth/JWT layer.
1. POST /v1/api/auth/refresh — handle_auth_refresh in core/auth.py,
wired in both console/server.py and server.py. Sliding-window
re-mint of the auth cookie. Re-resolves the user's permissions
from storage so a role change propagates within one refresh cycle
instead of persisting until the original cookie's natural expiry.
Returns the same JSON shape as /api/auth/login plus a fresh
Set-Cookie header. Refuses to extend a session for a deleted /
role-stripped user (403).
Resolves the user-visible "401 after browser tab open >24h"
symptom: previously the only refresh path was a full re-login,
now a single POST extends the session.
2. validate_jwt now passes leeway=30 to PyJWT. Absorbs minor
clock skew between hosts (multi-replica console deployments) and
between mint-time and validate-time within the same process.
Standard tolerance for short-lived tokens.
3. CoordinatorTokenManager._mint logs at debug. Mirrors the pattern
in ServiceTokenManager._mint (auth.py). Premature-401 diagnostics
would have been an order of magnitude faster with this in place
the first time around.
Frontend (shared_static/auth.js):
- _scheduleRefreshFromWhoami() reads the JWT exp surfaced via /whoami
and sets a setTimeout at 90% of remaining cookie life to call
/refresh. Floor 30s, ceiling 24h. Fires on initial page load
(silent if not authenticated) and after every successful login.
- _tryRefresh() de-dupes concurrent callers via a shared in-flight
promise — many parallel authFetch's hitting 401 at once still only
fire one /refresh.
- authFetch on-401 now attempts a single reactive refresh-then-retry
before falling through to the login overlay. Covers cases where
the proactive timer didn't fire (tab restored from disk-cache after
expiry, system clock jump, page first-load with stale cookie).
- BroadcastChannel "refresh" message keeps sibling tabs in sync so
they don't redundantly hit /refresh themselves.
- logout() cancels the proactive timer.
Tests:
- validate_jwt accepts 10s-expired tokens (within 30s leeway).
- validate_jwt rejects 60s-expired tokens (past leeway).
- /whoami includes exp claim with sane bounds.
- /refresh returns ok + Set-Cookie + the refreshed cookie keeps
working on subsequent authenticated requests.
- /refresh without a cookie returns 401.
Not addressed: the coordinator.session_jwt_ttl_seconds ceiling
(currently 1h) — that's a separate, preventative concern for very-
quiet long-running coordinators, orthogonal to the user-visible 401
this PR fixes. Can bump in a follow-up if it actually surfaces.
* fix(auth): address Copilot PR #395 feedback
Two real bugs caught by Copilot, both fixed.
1. Storage failure was indistinguishable from "user deleted" in
handle_auth_refresh. _load_user_permissions() swallows exceptions
and returns set(), so a transient DB hiccup looked like
"user has no permissions" and returned 403 — logging the user out.
Now calls storage.get_user_permissions() directly with try/except.
- Exception → log + fall through to in-token claims (refresh succeeds
with stale-but-valid permissions; better than fail-closed mid-
session for a hiccup).
- Empty set returned (no exception) → 403 (legitimate signal: user
deleted or role-stripped).
Tests:
- test_refresh_storage_failure_falls_back: storage raises → 200 +
in-token permissions.
- test_refresh_user_with_no_perms_403: storage returns empty → 403.
2. Logout race: a /refresh in flight when the user clicks Logout could
land AFTER /logout's clear-cookie response and re-set the cookie
from /refresh's Set-Cookie header, silently undoing the logout.
Fix in shared_static/auth.js:
- Add a _loggedOut latch + _refreshAbort AbortController.
- logout() sets _loggedOut = true synchronously and aborts any
in-flight /refresh BEFORE the /logout fetch fires.
- _tryRefresh() bails on its post-fetch effects (don't store perms,
don't reschedule, don't broadcast) when _loggedOut is set. The
stale Set-Cookie from /refresh is harmless because /logout's
response overwrites it on the way back.
- _onSuccess() (re-login) clears the latch so subsequent refreshes
work again.
Race window is small but real on slow networks / contested CPU.
* feat(design-system): DS phase 2 — opt server chat UI into v1 primitives
ui/static/index.html:
- data-design="v1" on <html> opts this view into design system tokens
and primitives scoped under the attribute selector.
- Link DS stylesheets after the legacy cascade: tokens + typography +
appbar (chrome) + panel / buttons / pills / message / field
(primitives). Legacy /shared/base.css, /shared/ui-base.css,
/shared/chat.css, and /static/style.css stay linked to handle
anything not yet migrated (rich markdown, tabs, dashboard, split
panes, approvals, modals).
- Header <div id="header"> picks up .appbar + .appbar-title +
.appbar-status + .appbar-spacer + .appbar-actions alongside the
legacy .ts-header classes. Theme-toggle gets .btn for DS pill shape
while keeping .header-btn for palette continuity.
ui/static/app.js:
- Chat message elements emit both legacy and DS class names so the
DS primitive picks up the message surface while legacy .ts-msg--*
rules keep view-specific markdown styling (tables, callouts, katex,
mermaid, hljs). Pairs:
ts-msg ts-msg--user → + msg user
ts-msg ts-msg--assistant → + msg assistant
ts-msg ts-msg--reasoning → + msg reasoning
ts-msg ts-msg--info → + msg info
ts-msg ts-msg--error → + msg error
ts-msg-body → + msg-body
- Approval blocks keep legacy-only styling — their shape is distinct
from the DS .msg primitive (the DS approval-dock pattern is a
fixed bottom dock, not inline-in-chat).
No backend or wire-format changes. SSE events, POST bodies, endpoint
URLs, ARIA attributes, and keyboard shortcuts all unchanged.
* feat(ui/static): DS-skin tool-call + approval + verdict internals
The outer .ts-msg.ts-approval--inline picked up DS .msg styling via
PR #2's dual-class approach, but the inner structure kept rendering
with legacy yellow/green/red colours and legacy chip shapes. Result:
a DS-accent-bordered card containing a mustard tool-name, a clunky
uppercase-yellow verdict chip, and a mismatched auto-approved pill.
Add [data-design="v1"]-scoped overrides that reskin the inner
vocabulary onto DS tokens:
.ts-approval-tool panel-over-panel-2 card with hair border
.tool-name accent (teal) for tool-kind identity
.tool-cmd / .tool-diff ink-2 text; diff-del/add/warn → err/ok/warn
.verdict-badge.verdict-* chip aesthetic matching DS k-badge —
low=ok-tinted, medium=warn-tinted,
high/critical=err-tinted, with a
3px left-border semantic stripe
.verdict-detail panel-2 callout with structured rows
.verdict-judge-spinner ts-pulse animation (reuses primitive)
.ts-approval-badge--* pill shape hugging max-content, matches
DS approve-button-family colour palette
(ok-text, err-text-mix)
.tool-output panel bg, hair border, accent stream
left-border, fade-gradient on collapse
.ts-verdict-glow--* soft ring on the corresponding action
button (approve=ok, deny=err, review=warn)
No JS changes; DOM shape unchanged. CSS-only reskin so approval flow,
tool streaming, and verdict expand/collapse behaviour all stay intact.
* fix(ui/static): consistent tool-card width + flat badge aesthetic
Two fixes to the DS-skinned tool-call rendering:
1. Tool-call cards were sizing to their content (short output →
narrow card, long output → full-width), producing a jagged column.
Force .ts-msg.ts-approval--inline to width: 100%; align-self:
stretch; box-sizing: border-box; so the chat column reads evenly.
2. The "approved" / "auto-approved" pill was styled as a button
(pilled shape, 1px bg-tinted border, 4x10 padding) which read
as clickable. Switched to a flat badge aesthetic matching the
.risk primitive: 3px-squared, 2x6 padding, 10px mono uppercase
on a --ok-soft / --err-soft tinted surface, no border. Reads
as a status tag, not a call-to-action.
* fix(ui/static): address Copilot PR #392 feedback
Copilot findings, all applied:
- Drop the legacy Outfit + IBM Plex Mono Google Fonts link. DS
typography.css @imports Inter + JetBrains Mono; loading both stacks
on opted-in pages wastes downloads and triggers FOIT/FOUT differences.
- Drop the .ts-header-title class on the <h1>. Its legacy rule forces
font-family: var(--font-display) (Outfit) which overrides the DS
appbar typography. .appbar-title alone is sufficient under v1.
- Override .ts-msg font-family under [data-design="v1"] when .msg is
also present (and not the .tool variant). Legacy .ts-msg forces
mono; DS user/assistant/reasoning/info/error messages should use
the UI font. .msg.tool keeps mono via the primitive's own rule.
- Replace the inline name.style.color = "var(--red)" in buildToolDiv
with a .tool-name--error class. Inline styles win over CSS rules
and broke the DS token mapping (legacy --red is not the DS --err).
- Correct the header comment in style.css for the approval-block
overrides. Prior comment claimed the outer wrapper picks up DS .msg
styling; it doesn't — the dual-class approach wasn't extended to
approval blocks. Updated comment to match actual DOM.
* feat(design-system): DS phase 3 — coordinator chat migration
Opt the per-session coordinator view into data-design="v1" and migrate
its rendering to the DS primitives + patterns shipped in phase 1. This
is the larger of the two parallel chat migrations (the other being the
server UI under turnstone/ui/static/).
Scope — this PR touches two files only:
turnstone/console/static/coordinator/index.html
- data-design="v1" on <html>; DS stylesheets linked after the legacy
base so primitives win on specificity and legacy styles keep
covering anything not-yet-migrated.
- Header rewired from .ts-header to .appbar with .appbar-back,
.appbar-title + .dim subtitle, .appbar-spacer, .appbar-status for
SSE state, and .appbar-actions wrapping the cancel / end / theme
buttons (now .btn pills).
- Approval bar replaced with the .approval-dock pattern. Signature
change: amber Approve becomes an ok-family (green) filled button
with 1.5px border + --r-md squared shape. .dcall rows frame each
pending call like a mini inspectable code line. Action cluster
sits in a .drow with Deny (.act.danger) / Always (.act.always) /
Approve (.act.primary) and the preview's kbd affordances
(D / ⇧A / ⏎). role="region" + aria-live="assertive" preserved;
the dock stays non-modal (no focus trap), focus moves to the
primary Approve button on open via the existing handler.
- Sidebar shell adopts .sidebar + .side-section + .side-label +
.ghost refresh buttons. Coordinator-only .sidebar overrides unset
the DS sticky-left-column defaults (which assume an admin-shell
grid) so the aside continues to flex into the right column of
#coord-body. Tree-row + task-row styling stays view-local,
rehomed to DS tokens (--hair-2 hover, --accent focus, --ok/--warn/
--err + -soft task-status tints).
- Inline <style> trimmed of rules now covered by DS primitives;
only the coordinator-specific flex wiring, tree-row visuals, and
<700px responsive accordion remain.
turnstone/console/static/coordinator/coordinator.js
- appendMsg() emits .msg + role variant (.msg.user / .msg.assistant /
.msg.reasoning / .msg.tool / .msg.error / .msg.info) and .msg-body.
_TS_ROLE_VARIANTS renamed _MSG_VARIANTS.
- Streaming helpers query .msg-body; SSE dedup-by-call-id query
updated to .msg[data-call-id=...].
- showApproval() renders the .approval-dock DOM shape: .dhead count
in a .dcount, one .dcall per pending call with .risk index pill +
.dfn function name + .dargs preview. approvalBar.hidden toggles
visibility (the DS pattern is position: fixed and always-rendered;
[hidden] is the show/hide hook).
- setSseStatus() keeps .appbar-status as the base; semantic colour
tracks OK / ERR via inline --ok / --err. Leading glyph (●/○/⚠)
preserves the WCAG 1.4.1 non-colour-only cue.
- Wait indicator uses .appbar-status instead of the legacy
.ts-header-status BEM; styling from the inline page rules colours
it --think.
Contracts preserved:
- SSE wire format and event names unchanged (approve_request,
child_ws_created, wait_progress, batch_started, state_change,
stream_end, ...).
- POST /approve body shape unchanged: {approved, always, call_id}.
No per-item feedback field is added (that's phase 9 PR C).
- Keyboard behaviour unchanged: Enter continues to approve via the
primary-button focus shift in showApproval(); the D and ⇧A kbd
labels are rendered per the pattern spec but the global key
handlers (if any) remain untouched.
- ARIA attributes (role, aria-label, aria-live) preserved on the
approval dock, messages log, and sidebar.
- All shared-static JS imports and order unchanged; composer module
continues to own its own DOM inside #coord-composer-mount.
No backend changes. Legacy CSS (/shared/base.css, /shared/ui-base.css,
/shared/chat.css, /static/style.css) stays linked as the compatibility
layer — DS selectors [data-design="v1"] beat legacy where applied.
* fix(coordinator): inline approval dock above composer, not viewport-pinned
The .approval-dock DS pattern defaults to position: fixed; bottom: 22px
— designed for the fleet dashboard where the dock overlays content. In
the coordinator chat that rule pinned the dock to the viewport bottom,
covering the composer input area.
Move the dock DOM back inside #coord-main between #coord-messages and
the composer mount so it flex-stacks naturally above the input. Add a
view-local override that neutralises the fixed positioning (position:
static, z-index/box-shadow auto) while preserving the visual pattern
(warm top stripe, head/call/actions rows, dashed Always button).
Drop the 160px bottom-padding hack on #coord-messages since the dock
is now in-flow and naturally pushes the message log up.
Also likely resolves the Firefox initial-render issue — position:fixed
+ [hidden] toggle had cross-browser quirks where the dock wouldn't
appear on first SSE approval event until a separate DOM mutation
forced a reflow. In-flow layout makes it boring and predictable.
* fix(coordinator): integrate judge verdicts into approval dock, not chat
The judge's intent_verdict is evaluation context for the pending
approval, not a chat message. Previously each verdict appended a
"[judge] deny (risk=low)" tool message into the transcript even when
the corresponding approval was visible in the dock — two separate
surfaces showing related decision context, neither one complete.
Now:
- Each .dcall row gets data-call-id from the approve_request item
- intent_verdict looks up the matching row and renders a .dctx sibling
below it with "judge: <recommendation> (risk: <level>)" + optional
"confidence: <score>" chips. Reasoning attaches as title tooltip.
- Verdicts cache in a Map<call_id, verdict> so late-arriving
approve_request events can still apply verdicts that came early
- Fallback to the old chat-message surface only when the approval isn't
visible (call_id missing, or resolved before we could render) so the
verdict isn't silently dropped
* feat(coordinator): judge verdict polish — colour-coded chips, spinner, reasoning
Three refinements to the approval-dock judge integration:
1. Colour-code verdict chips by recommendation — approve=green (--ok),
review=amber (--warn), deny=red (--err). Reviewers can triage at a
glance without reading the chip text; complements the text label
for WCAG 1.4.1 (non-colour-only signaling).
2. Spinner while evaluating — when showApproval builds a .dcall row
without a cached verdict, render a "judge evaluating…" chip with
a spinner. Replaced in-place when intent_verdict arrives. Reuses
the ts-spin keyframe from primitives/feed.css.
3. Justification inline — judge.reasoning is delivered in every
intent_verdict event but was hidden behind a title tooltip. Now
renders as a wrapped prose block (.drationale) below the .dctx
chips, styled like the .msg-body .evi callout (left-rule + mono
+ --ink-3). Full text, no truncation — justification is the whole
point.
View-local styling; the approval-dock pattern itself is unchanged.
If these patterns turn out to be broadly useful, they can promote to
shared_static/design/patterns/approval-dock.css in a later PR.
* fix(coordinator): defer approve-button focus until judge verdict arrives
The Approve button was getting focus the instant the approval dock
opened, which lit up the green focus ring and made the filled-green
button look pre-confirmed. A reviewer could mistake that for "already
approved" before the judge has even returned a verdict.
Now focus is deferred until the intent_verdict for the first-pending
call arrives, then moves to:
- Deny button when judge recommends "deny" (safety default)
- Approve button for "approve" / "review" / anything else
Fallback timer (3s) claims focus anyway if no verdict arrives — covers
disabled judge and slow judge cases so keyboard users still land on a
button within a beat.
Focus claim is idempotent so batch approvals don't bounce focus across
buttons as trickling verdicts arrive. hideApproval clears the timer
and the claimed flag so re-open cycles start fresh.
* fix(coordinator): drop approve-focus fallback timer
Previous commit added a 3s fallback that focused Approve if no verdict
arrived. Ambiguous — a focus ring that lands "eventually" looks the
same as one that lands because the judge recommended approve.
Now focus only ever moves when a real intent_verdict arrives. If the
judge is disabled or the verdict never comes, focus stays put and
keyboard users tab from the composer to reach the buttons. An absent
focus ring is a clearer signal than an ambiguous one.
* fix(coordinator): address Copilot PR #393 feedback
Copilot findings, applied:
- Restore <h2> for Children / Tasks sidebar section labels (were
changed to <span>). .side-label class still applies; screen readers
recover heading-level structure + rotor navigation.
- Mount the wait-indicator into #coord-header (the appbar container)
instead of #coord-status. #coord-status is reset via
statusEl.textContent = ... on every state_change event, which was
clobbering the wait indicator between ticks. As a sibling inside
the appbar, it survives state updates.
- Route `info` SSE events to appendText("info", ...) so they render
with .msg.info (think-indigo) styling. Prior routing to "tool"
gave info events accent-tinted tool-call styling, miscategorising
them visually.
- Define @keyframes ts-spin locally in the coordinator's <style>.
Canonical definition lives in primitives/feed.css but this page
doesn't link feed.css (no .feed-item usage), so the "judge
evaluating…" spinner wasn't animating.
- Clear judgeVerdicts Map in hideApproval. Map was growing unbounded
across resolve cycles — fine for short sessions, leaks memory on
long-lived coordinators with many approvals.
Not applied: Copilot's suggestion to restore focus-on-open or add a
fallback timer. User explicitly requested no fallback — the design
decision is that the focus ring should only ever appear when the
judge has returned a verdict, so an absent ring reliably means "no
recommendation yet." An auto-focus fallback would produce an
ambiguous ring that could be misread as "judge approved."
* feat(design-system): DS phase 1 — chat primitives for view migrations
Three new primitives enabling the chat-surface migrations (server UI +
coordinator):
primitives/message.css .msg + variants (user / assistant /
reasoning / tool / error / info / system),
.msg-meta author/timestamp slot, .msg-body
markdown target, .msg-actions hover-revealed
row, data-streaming="true" blinking caret.
Replaces .ts-msg* family in chat.css.
primitives/field.css .field wrapper with label/help/error, element
selectors for text/email/password/url/number/
search/tel/date/time/datetime/month/week +
textarea + select. .field.inline for checkbox/
radio rows, .field.invalid for error state.
Native-control focus-visible handled for
checkbox+radio so box-shadow ring remains
visible on unframed controls.
chrome/appbar.css chat-app header: back link + title + status +
action cluster. Distinct from the admin-style
.topbar (brand mark + nav + env metadata).
min-width:0 on .appbar-title so .dim subtitle
ellipsis fires under narrow viewports.
Preview.html extended with three demo sections exercising every variant
(plus a data-streaming example with live caret).
Fixes carried in from code review:
- @media (hover: none) and (pointer: coarse) to match chat.css
convention (hover-none alone is too broad, catches styluses)
- .field-help uses --ink-3 (not --ink-4 which fails AA on --panel)
- .msg-meta slot added so downstream PRs don't invent a custom class
- Tool message pre/code on --panel-2 (parent is --panel; same-bg
would make inline code disappear)
- Checkbox/radio :focus-visible override (native controls lack a
border for the default box-shadow ring to wrap)
- Message.css comment corrected: "accent-tinted" not "cyan"
All rules scoped under [data-design="v1"]. Nothing existing modified.
* fix(design-system): address Copilot PR #391 feedback
- .msg-actions: add pointer-events: none when hidden, auto when visible.
opacity:0 alone still intercepts clicks in the top-right corner —
broke text selection on short one-line messages. Toggle applied in
both hover/focus-within and the touch-media-query visible states.
- .field.inline comment: rewrite to match behaviour. Old comment said
".field stays flex-column" but the rule sets flex-direction: row.
- preview.html appbar demo: swap <a tabindex="0"> back-link to
<button type="button">. tabindex-only anchors without href have
inconsistent focus + screen-reader semantics; button is the correct
native element for "navigate back via JS."
Live-preview-driven tuning pass following PR #389:
Palette
- Accent hue 70 (amber) → 182 (teal). Amber collided with warn on
same-surface k-badges; teal gives the brand accent its own hue.
- ok / warn / err / think unified at L=0.50 light / L=0.68 dark and
C=0.13-0.17 for palette coherence. err holds higher chroma so red
doesn't wash; warn stays in the gold 80 lane (never 90+ / "puke").
- Soft variants unified at L=0.94 / L=0.29, C=0.05-0.07.
New tokens
--ok-live brighter green for liveness signals (running dot)
--ok-text theme-aware text colour for filled green surfaces,
dark forest in light / bright mint in dark, ~7.5:1
against the approve-button bg in both themes
--err-fill darker red specifically for filled destructive
surfaces (.risk.crit) — bright --err as a fill
reads as alarm-loud
--warn-tint, directly-defined gold tints for k-tools k-badge —
-tint-border skips the color-mix-through-dark-cool-panel mud
that would otherwise render warm low-L mixes brown
Approve / Always / Deny
- Approve filled green (color-mix --ok 28% into panel); text uses
--ok-text for theme-correct contrast. Matches the pre-refactor
turnstone/shared_static/chat.css convention where approve = green.
Deviates from the Claude Design spec which had warn-tinted approve.
- Always outlined dashed green (same --ok hue family); four non-colour
cues for WCAG 1.4.1: fill state, border style, label, position.
- Deny unchanged (err-outlined).
k-badge glyphs
Replaced generic shapes with semantic symbols:
tools ⚙ approval ⚠\FE0E policy § role ◉
oidc ⌘ token ◆ judge ⚖\FE0E query ?
step ⇧ session ◈ skill ★ workstream ⇉
fanout ⇶ default ·
⚠ and ⚖ carry \FE0E to force text-presentation (avoid emoji
promotion to coloured yellow triangle / blue scales on iOS Safari).
token uses ◆ instead of ⬢ for universal font coverage.
k-approval split from k-tools
k-tools stays gold (--warn family) — "tool call" kind.
k-approval moves to green (--ok family) — matches the Approve button
visually, completing the "⚠ approval → Approve" same-family story.
Running pill
Text uses --ok (passes AA on pale --ok-soft); dot uses --ok-live +
pulse. Liveness signal lives in the dot, not the text.
All changes stay under [data-design="v1"] — existing views untouched.
* feat(design-system): DS-A — tokens + typography scaffold
Adds turnstone/shared_static/design/{tokens.css,typography.css} as the
first phase of a multi-PR design refactor seeded by Claude Design.
- tokens.css: full palette + shape + rhythm, light default with
[data-theme="dark"] override. oklch() raw colours, color-mix kept out
of DS-A entirely (reserved for primitives in DS-B).
- typography.css: Inter + JetBrains Mono via Google Fonts; six-step
scale (10/11/12/13/14/20-24). Utility classes .t-kicker/.t-meta/
.t-btn/.t-row/.t-body/.t-stat/.t-h1.
Signature accent stays warm amber (oklch hue 70) rather than Claude
Design's teal — preserves turnstone's "Instrument Panel" identity.
All other tokens match the spec verbatim.
Additive: both files gate under [data-design="v1"] so existing views
(base.css, per-view stylesheets) are untouched. DS-B will opt views in
one at a time.
* feat(design-system): DS-B — chrome + primitives + preview page
Adds the reusable primitive kit that DS-C and DS-Cluster will build on:
primitives/
panel.css .panel, .panel-head (.tools pinned right), .ghost
buttons.css .btn (pill 999px), .primary, .deny, .approve (amber)
pills.css .pill (running/thinking/attn/idle/err), .k-badge
(glyph-prefixed per WCAG 1.4.1), .chip, .risk
stats.css .stat + .stat-row, .mini-bar, .spark
feed.css .feed-item (grid ts/body/acts + .evi callout)
chrome/
topbar.css 48px sticky, conic-gradient brand mark
sidebar.css 240px sticky, .shell layout, semantic swatches
preview.html renders every primitive in both themes with an
in-page theme toggle (tracks prefers-color-scheme)
Additive: every selector scopes under [data-design="v1"] so existing
views (base.css + per-view stylesheets) stay untouched.
Spec deviations from the Claude Design prototype:
- `color-mix(in srgb, …)` throughout; prototype had two `in oklab`
usages — srgb per the spec's hard rule
- `.btn.approve` is warn-tinted amber, not green
(approvals signal "needs attention"; amber resolves on approval)
- k-badge tint uses `color-mix` instead of oklch relative-colour syntax
for broader browser support
- `@keyframes pulse/spin` renamed to `ts-pulse/ts-spin` to avoid
clashing with keyframes in base.css on pages that load both
- `prefers-reduced-motion` disables pulse + spin animations
- Text-on-accent-soft + text-on-warn-tinted darkened via color-mix
with ink to pass WCAG AA at 12px (fixes the classic same-hue trap)
- `.risk.crit` uses `#fff` text (dark-mode --panel on bright err fails)
- `.feed-item .acts button:not(.btn)` — compact action styling now
skips .btn-classed buttons so they keep their pill shape
- Focus-visible rings on .btn, .ghost, .stat, .topnav, .side-item
* feat(design-system): DS-C — patterns (approval-dock, fleet-grid, live-feed)
Completes the design library with three patterns that compose primitives
into the signature product surfaces described in the Claude Design handoff.
patterns/
approval-dock.css bottom-pinned approval strip. 1.5px-border,
--r-md squared action cluster: amber Approve
(primary), dashed Always, red Deny. kbd hints
and focus-visible rings on all three acts.
Call row (.dcall) framed as an inline code-
panel to emphasize "this is the exact call."
fleet-grid.css 14-col grid of .node squares. State modifiers
(.s-ok/.s-thinking/.s-attn/.s-err/.s-idle/
.s-unreach) + --pct load fill. Hover uses
outline, not box-shadow (neighbour bleed is
the intended density cue). .fleet-legend
swatch row below.
live-feed.css thin scroll-container wrapper over the
.feed-item primitive with a sticky top fade.
preview.html imports the three patterns, extends the
fleet demo to use the real .fleet class +
legend, adds a live-feed panel, renders the
approval dock fixed at the bottom with
aria-live="polite".
Spec notes:
- Approve button is amber (warn-tinted), never green
- Dock action buttons are 1.5px-bordered 6px-radius squares — NOT
pills — signaling "primary-action surface"
- All three dock actions clear WCAG AA in both themes via the same
color-mix-with-ink darkening pattern used in .btn.approve
- kbd hint color matches primitives/buttons.css (--ink-3, not --ink-4)
View-level rewrites (coordinator.html + coordinator.js opt-in,
admin/cluster dashboard rebuild) are follow-up PRs — they need a
running server to test SSE streams + the approval POST contract.
* fix(design-system): scope DS-A tokens to [data-design="v1"]
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* fix(design-system): scope DS-A font vars to [data-design="v1"]
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* fix(design-system): align dark-mode selector with theme.js convention
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* fix(packaging): add shared_static/design/** to wheel includes
Agent-Logs-Url: https://github.com/turnstonelabs/turnstone/sessions/74d0939a-c55f-46b0-92f4-14d0cbfb7084
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
* docs(coordinator): phase 8 PR C — API tour, skills guide, bulk-endpoints contract
Four deliverables that close out the phase 8 doc debt carried since
phase 1:
- docs/coordinator-api-tour.md — 9-step lifecycle walkthrough
(create → subscribe → send → inspect children / detail → wait for
fan-out → govern (trust / restrict / stop_cascade / close_all_children)
→ approve / cancel → close), one request + response per step, every
SSE event type the UI has to handle, and every operation id cross-
referenced against the live /openapi.json. Integrators driving a
coord session from a custom UI or SDK can work end-to-end from this
doc without reverse-engineering the console page.
- docs/coordinator-skills.md — writing a SkillKind=COORDINATOR skill.
Tool-surface diff (13 orchestration tools, no bash / edit / web /
sub-agent), persona diff (orchestrator vs maker, composing on
base_coordinator.md), SkillKind enum + migration 044, task_list
integration, ws_id handling, wait vs inspect cost profile, three
orchestration patterns (delegate-and-summarise, fan-out-and-
synthesise, plan-then-delegate), testing surface.
- docs/bulk-endpoints.md — codifies the two shape idioms that shipped
across phases 6–8: {results, denied, truncated} for bulk-read /
bulk-create-with-payload (cluster/ws/live, spawn_batch); {<bucket>,
failed, skipped} for cascade-mutation (stop_cascade,
close_all_children). Picks-by-semantics guidance so the next bulk
endpoint author doesn't coin a third shape.
- docs/diagrams/27-coordinator-wait-for-workstream.puml + rendered
PNG — sequence diagram covering spawn → wait (blocking, with
bounded progress emission) → inspect → close. Embedded in the
API tour doc's §6 so the "why is my coord session blocking?"
question has a visible answer.
No code changes. All operation ids in the API tour verified against
a live build of the console spec; all markdown internal links
resolve; PlantUML renders clean on the system plantuml jar.
* docs(coordinator): address PR #388 copilot review
- api-tour.md child-event payload keys: events stamp `ws_id` as the
coord's own id and carry the child's id separately as
`child_ws_id`. Doc previously listed `ws_id` as the child
identifier on all four child_ws_* events, which would send SDK /
UI implementers parsing the wrong field.
- api-tour.md SSE table: add the `status` event emitted by
ConsoleCoordinatorUI.on_status (token usage + context_window +
effort snapshot; fires on every streaming tick). Previously
omitted from the "every event type a UI has to handle" list.
- api-tour.md /children response key: server returns `{items,
truncated}`, not `{children, truncated}`. Also drop the
`state=closed` query-param claim — the endpoint has no state
filter; clients filter locally on the returned `state` field.
- skills.md task_list shape: the persisted row uses `id` (not
`task_id` — the input schema uses `task_id`, the row uses `id`),
has `child_ws_id` / `created` / `updated` (no `notes` field),
and supports a 5th `reorder` action alongside add/update/remove/
list. Adds the parallel-dispatch caveat from the tool
description.
- skills.md tenant-guard behaviour: foreign / hallucinated ws_ids
don't return an empty result — they return explicit
error/not-found/denied shapes that differ by op (mutating ops
return `{error, status: 404}`; inspect returns `{error}`; wait
reports state=denied). Important distinction — a skill that
expects empty on mismatch will mishandle every single case.
Docs-only; no code / schema / SDK changes. All internal links
still resolve.
* feat(coordinator): phase 8 PR B — spawn budget + rate limit + /quota endpoint
Adds two complementary controls so a runaway coordinator can't saturate a
cluster's max_active without anyone noticing:
- **Spawn budget** (hard quota) — cap on concurrently active children.
Default 20 per coord. spawn_workstream returns a tool error guiding
the model to close idle children; spawn_batch routes overflow rows to
`denied[]` with partial-success semantics.
- **Spawn rate limit** (soft pacing) — classic token bucket, defaults
5 tokens/minute with burst 10. A rate-limited spawn surfaces a tool
error carrying `retry after Ns` so the model paces itself. Zero
refill rate is honoured as "disable refill" (bucket still honors the
initial burst).
Shipped infra:
- `turnstone/core/spawn_quota.py` — thread-safe `SpawnBudget` +
`TokenBucket`. 15 unit tests.
- `turnstone/core/session.py` — coord-only state built from settings at
__init__. Shared `_eval_spawn_quota(active)` helper drives both the
single-spawn path (wraps the denial reason in `_coord_tool_error`) and
the batch path (annotates `spec["_error"]`). `_count_active_children`
routes through `coord_client.list_children(include_closed=False)` and
fails *open* on lookup error (budget is operator-safety, not security).
- `POST/GET /v1/api/coordinator/{ws_id}/quota` — partial-update admin
endpoint mirroring the /trust + /restrict shape. Accepts either the
nested `spawn_rate` object or flat aliases — supplying both for the
same field returns 400 so the admin UI can't half-migrate silently.
Overrides are in-memory only (die on session reopen). Audits via
`coordinator.quota.updated` with before/after snapshots.
- Settings: `coordinator.spawn_budget`, `coordinator.spawn_rate.tokens_per_minute`,
`coordinator.spawn_rate.burst` with ranges 1..500 / 0..600 / 1..500.
The range bounds are the single source of truth — the endpoint
validators and Pydantic schema both import from `settings_registry.SETTINGS`
so bumping a cap in one place lights up everywhere.
- OpenAPI: `CoordinatorQuotaRequest` / `CoordinatorQuotaResponse` /
`CoordinatorSpawnRateState` schemas + endpoint specs. TS SDK regenerated.
Tests: +15 unit (SpawnBudget + TokenBucket), +17 endpoint (GET + POST
happy paths, range edges, mixed-body rejection, non-object spawn_rate,
service-token refusal), +11 session-side (budget blocks single spawn,
budget batch partial-success, rate batch partial-success, empty-body
reject, mutator live-update, non-coord session has no quota state).
Deferred (not this PR): per-skill scoping via migration 047 +
`prompt_templates.spawn_budget` column. Count-only storage helper
(opportunistic — list_children at budget ≤ 500 is fine behind a
human-gated approval flow).
* fix(coordinator): address PR #387 copilot review
- Budget undercount: _count_active_children used list_children's
LIMIT-then-Python-filter path, so a fan-out with many recently-closed
children could push live rows past the SQL LIMIT and silently
undercount, leaking spawn slots past the budget. Replace with a new
CoordinatorClient.count_active_children that uses
storage.count_workstreams_by_state (SQL aggregate, no pagination,
sums non-terminal states). Tenant-guarded; fails open on storage
error (budget is operator-safety, not a security gate). New client
tests cover the non-terminal count, the closed/deleted exclusion,
the foreign-parent guard, and the fail-open path.
- Service-token bypass on /quota: both GET and POST used the default
allow_service_bypass=True, so a service token whose user_id matched
the coord owner could read or *raise* spawn capacity without the
explicit admin.coordinator grant. Flip both to
allow_service_bypass=False for consistency with /restrict,
/stop_cascade, and /close_all_children.
- OpenAPI contract leak: CoordinatorSpawnRateState was used for both
the request and response shapes, which let generated SDKs imply
clients could POST tokens_available (a read-only bucket reading the
handler ignores). Split into CoordinatorSpawnRateInput (request:
tokens_per_minute + burst only) and CoordinatorSpawnRateState
(response: adds tokens_available). No runtime behaviour change;
SDKs regenerate with two distinct types.
Drops the _ACTIVE_COUNT_SLACK / _ACTIVE_COUNT_MIN_LIMIT constants in
session.py — no longer needed since the new helper takes no limit
argument. Updates the 5 session-side quota tests to stub
count_active_children instead of list_children.
* feat(coordinator): phase 8 PR A — spawn_batch + close_all_children batch tools
Adds two model-facing batch tools so a coordinator can fan out without burning one approval per child:
- `spawn_batch` — create up to 10 child workstreams in a single approval. Serialised
spawns so sibling ordering (by created_at) stays deterministic. Returns
`{results: {idx: {ws_id, name, node_id, status}}, denied: [{idx, reason}]}`.
Per-item validation / spawn failures surface in `denied[]`; the batch hard-errors
on >10 rather than silent truncation.
- `close_all_children` — soft-close every direct child in one approval. Server-side
Sem(16) fan-out via `coord_client.close_workstream`; `reason` propagates to every
closed child's audit + workstream_config. Response mirrors `stop_cascade`'s cascade
idiom: `{closed, failed, skipped}` where `skipped` is upstream-404 / already-gone.
Shipped infra:
- New console endpoint `POST /v1/api/coordinator/{ws_id}/close_all_children`
(gated `admin.coordinator`, `allow_service_bypass=False`, 512-char reason cap,
`coordinator.closed_all_children` audit).
- Shared `_fanout_on_children` helper — both `stop_cascade` and `close_all_children`
now delegate to it (one place to own the snapshot → semaphore-gather → bucket-split
skeleton).
- `CoordinatorClient.close_all_children(reason)` plus a `_post_url` seam that
`_post` now reuses (no more duplicated transport-error handling).
- `_emit_batch_event` — best-effort SSE emitter modelled on `_emit_wait_event`.
Emits `batch_started` / `batch_ended` pairs keyed by call_id. Throttled
`batch_progress` deferred to a follow-up.
- OpenAPI request + response schemas, endpoint spec entry, TS SDK regenerated.
- Persona doc (`tools_coordinator.md`) covers the two new patterns.
Bulk-endpoint shape policy (codified in PR C later): split by semantic category —
`{results, denied, truncated}` for bulk-read / bulk-create-with-payload (cluster/ws/live,
spawn_batch), `{<bucket>, failed, skipped}` for cascade-mutation (stop_cascade,
close_all_children). No retrofit needed on stop_cascade.
Tests: new `test_coordinator_close_all_children.py` (8 endpoint tests), expanded
`test_coordinator_tools.py` (session-side prepare/exec, coord_client=None guards,
batch SSE events), expanded `test_coordinator_client.py` (route map, client method,
transport errors), tool-count assertions updated.
Deferred (not this PR): per-item selective-deny approval UI, throttled batch_progress
SSE, coordinator-skills doc + bulk-endpoints doc (PR C), spawn budget / rate limit (PR B).
* fix(coordinator): address PR #386 copilot review
- coordinator_client.close_all_children: pass the unformatted path template
as log_path so telemetry aggregates don't fragment per session (ws_id
still lives in the real URL).
- session.py: drop dead spawned_ids accumulator in _exec_spawn_batch —
leftover from an eager-register path that got removed earlier.
- close_all_children tool JSON: document the 512-char server-side cap on
reason and that reason is echoed back in the response payload. Added
maxLength:512 on the schema property so the LLM sees the constraint.
- CoordinatorCloseAllChildrenRequest: add Field(max_length=512) so the
OpenAPI schema reflects the runtime 400-on-overflow constraint.
* refactor(ui): shared composer widget (pane / coordinator / coord-create)
The interactive workstream pane (turnstone-server), the coordinator
session view (turnstone-console), and the console home's "start a new
orchestration task" form had drifted into three unrelated composer
implementations with different DOM, different class names, and
different behaviour sets. All three now build on a single
`shared_static/composer.js` widget parameterised by feature flags.
The widget owns the textarea, send button, optional stop button,
optional attach button + file input + chip container, optional
drag-drop / paste-image wiring, optional queue-while-busy send-label
rotation, optional touch-aware Enter-to-send, and an optional
collapsible Options panel with input / select fields, live summary
chip, and localStorage-persisted open/closed state. A stacked layout
puts the textarea above the action row for creation-form consumers;
the inline layout keeps the chat-style single row for send composers.
Consumer wiring:
- Pane: attachments + stopBtn + queueWhileBusy + drag-drop; keeps
its own attachment-upload pipeline and routes file events through
the composer's onAttach callback. Pane-specific CSS (.pane-stop
/ .pane-send.queue-mode) retired; `.ts-composer-stop` / `.ts-
composer-send--queue` in shared/chat.css take their place.
- Coordinator send: just textarea + send with touchEnterSends=true
to preserve the pre-refactor tap-to-send behaviour on tablets.
The header-mounted coord-cancel-btn stays (different semantics
than a per-generation stop).
- Coord-create: stacked layout, rows=3, Start-labelled send, Options
dropdown holding Name + Skill; Ctrl/Cmd+Enter handler scoped to
the composer mount. `_createCoordinator` lost its DOM-ref shape
in favour of raw values + a setBusy callback; a single
`_refreshHomeCoordSubmitEnabled` reconciler owns the submit
button's disabled flag so the 503 probe and in-flight submit
can't race each other.
Shared chat.css grew the .ts-composer-stop, .ts-composer-options-*,
and .ts-composer--stacked blocks; the coord-create consumer dropped
its custom .home-composer-task/-row/-name/-skill/-submit selectors
and the "Start a new orchestration task" panel title so the
placeholder text carries its own context, matching the webui
dashboard's clean look.
The visible behaviour on each surface is intentionally the same as
before; the change is structural — the three composers can no longer
drift apart silently.
* fix(composer): review fixes — Enter guard, single disable owner, widget-owned stop reset
Review of the squashed whole caught issues that the piecewise reviews
missed because they only become visible with all three consumers
together:
- **Enter bypassed sendBtn.disabled.** Composer's Enter keydown
handler called _fireSend() without checking sendBtn.disabled. In
the coord-create flow submitHomeCoord doesn't clear the textarea
before the POST completes (it redirects on success), so two rapid
Enter presses both fired _createCoordinator and could create two
coordinators. Enter now mirrors the click path.
- **Two writers to sendBtn.disabled.** Composer.setBusy and
_refreshHomeCoordSubmitEnabled both wrote the flag. They agreed
in sequence today but it was the exact drift hazard the reconciler
was meant to eliminate. Added an externalDisable option; when
true Composer's setBusy rotates labels / placeholder / stop button
but leaves sendBtn.disabled to the caller's reconciler. The
coord-create composer opts in.
- **setSendLabel dead weight.** Called on every busy transition
with static "Start" / "Starting…". Composer.setBusy now rotates
labels universally (not just in queueWhileBusy mode); the busy
label goes at construction via the existing busyLabel option and
setSendLabel is removed.
- **Pane reached through composer to reset stopBtn.** Pane.setBusy
was writing stopBtn.textContent / aria-label / dataset after
delegating — internals leaking through. Composer.setBusy now
resets the stop button's standard label + clears forceCancel on
every transition (matching the comment that used to live in Pane);
Pane drops the reach-through.
- **destroy() left detached DOM reachable.** Back-refs (inputEl,
sendBtn, etc.) are nulled out so post-destroy access fails loudly
instead of silently mutating detached nodes.
- **_maybeAutoResize dead indirection.** The enabled-check folded
into autoResize itself.
* fix(composer): busyLabel context-sensitive default + busyPlaceholder universal
Round-two review caught three related loose ends:
- Default `busyLabel="Queue"` was fine when label rotation was queue-
mode-only, but became misleading after the earlier review fix made
rotation universal: non-queue consumers calling setBusy(true)
without explicit busyLabel would flash "Queue" on the disabled
button. Default is now context-sensitive — "Queue" when
queueWhileBusy=true, sendLabel otherwise (no rotation). The
coord-send composer no longer needs to touch the label at all.
- `busyPlaceholder` JSDoc implied universal swap on busy but the
implementation gated it on queueWhileBusy. Decoupled — the
placeholder swaps whenever busy, with callers that don't set
busyPlaceholder seeing no visible change because it defaults to
the idle placeholder.
- `options.toggleLabel` and `options.onChange` were supported by the
implementation but undocumented. JSDoc for the options shape
enumerates every supported key + its default.
* refactor(routing): replace hash-ring rebalancer with rendezvous (HRW) hashing
Routing was a stored bucket table maintained by a central rebalancer
daemon, which shared its liveness primitive (services.last_heartbeat)
with the collector — when a heartbeat-fresh node went into a zombie
HTTP-handler-broken state, neither the collector nor the rebalancer
could self-correct, and the router kept directing traffic at it.
Rendezvous hashing makes the route a pure function of (ws_id,
live_services) so the heartbeat is the single source of truth and any
liveness-eviction propagates to the next route call without a separate
state-publication step.
The rebalancer's central state has no analogue: the new router computes
the per-key node winner on every call, the collector pushes membership
updates into the router cache from its discovery thread, and per-route
overrides survive on workstream_overrides. Eager workstream migration
goes away; in-flight workstreams lazily rehydrate from storage on the
new owner — already the dead-node behaviour.
* fix(tools): describe rendezvous re-routing on spawn/inspect node_id
The first pass overclaimed `node_id` "stays canonical for this
workstream's lifetime" — under rendezvous routing the active owner
re-derives per-call from live membership, so a node join/drop after
spawn can shift it. Tool descriptions now say `node_id` is the
spawn-time binding; subsequent ops re-route via rendezvous over the
current live-node set; the new owner lazily rehydrates from shared
storage; coordinators should re-read with inspect_workstream rather
than caching the value.
* feat(coordinator): phase 7 — governance + skill metadata + cross-cutting invariants
Combines three stacked sub-PRs into a single coordinator phase-7
shipment against the phase-7 plan doc. The sub-PR structure (0 / A /
B) preserved on individual branches for reviewer drill-down; this
branch is the one reviewers should merge.
## Sub-PR 0 — service-auth boundary invariants
Shared helpers and contracts that lock the console ↔ node service-auth
boundary so later authz surfaces use them by construction.
- ``_effective_user_filter(request)`` in both ``turnstone.console.server``
and ``turnstone.server`` with a shared ``DENY_EMPTY_SUB`` sentinel
on ``turnstone.core.auth``. Three-way return — admin/service
bypass, scoped caller uid, or fail-closed sentinel on blank sub.
Four callsite migrations (``_coordinator_rows``,
``coordinator_children``, ``coordinator_metrics``,
``cluster_ws_live_bulk``).
- ``StorageBackend`` class docstring codifies the tenancy contract
(every list/count/aggregate method must accept ``user_id: str |
None = None`` and push ``WHERE user_id = :user_id`` into SQL) and
the ``_mapping`` row-access contract. New
``turnstone.testing.row_contract`` ships ``assert_row_like()``.
- ``_verify_collector_service_scope`` probes an upstream node at boot
with ``expected_node_id=_scope-probe_``; a 409 proves the scope
gate was passed, a 403/401 sets ``collector_scope_error`` and
causes ``cluster_snapshot`` / ``cluster_events_sse`` to return 503
with a remediation hint. Probe URL allowlist rejects non-http(s)
schemes and 169.254.0.0/16 hosts.
- 4xx log-level floor on ``_NodeDashboardCache.get``,
``_fetch_live_block``, and ``_proxy_sse`` — dotted-hierarchy
prefixes with bounded body previews. ``_bounded_body_preview`` and
``_bounded_stream_preview`` strip control chars.
## Sub-PR A — coordinator governance core
Mid-session governance surface for coordinator workstreams.
- **Trusted-session mode.** New ``coordinator.trust.send``
permission (migration 042). ``ChatSession.set_trust_send`` /
``revoke_tools`` methods with a ``_governance_lock``. ``POST
/v1/api/coordinator/{ws_id}/trust {send: bool}`` double-gated on
``admin.coordinator`` AND ``coordinator.trust.send`` with
``allow_service_bypass=False`` so service tokens can't escalate.
``_prepare_send_to_workstream`` auto-approves sends whose target is
in the coordinator's own subtree; foreign ws_ids still require
approval. ``_is_own_subtree`` checks both ``parent_ws_id`` AND
``user_id`` to defend against cross-tenant row corruption.
- **Audit-layer credential redaction.** ``record_audit`` walks
``detail`` (dicts, lists, tuples, sets, frozensets; keys too)
and routes every string through ``redact_credentials`` + a C0
control-char scrub. New kw-only ``raw_detail=True`` opt-out.
``_has_any_string`` fast-path. Audit action registry extended
with the four new governance sub-prefixes.
- **Mid-session revocation + cascading stop.** ``POST
/v1/api/coordinator/{ws_id}/restrict {revoke: [...]}`` caps 256
entries / 128 chars; ``_prepare_tool`` short-circuits with a
tool-error. ``POST /v1/api/coordinator/{ws_id}/stop_cascade``
cancels the coord's in-flight generation then dispatches
``cancel_workstream`` for every direct child in parallel via
``asyncio.gather`` bounded by ``Semaphore(16)``. Per-child
outcomes split into ``cancelled`` / ``failed`` / ``skipped``
(404 = already-gone rather than dispatch-broken). Both endpoints
apply ``allow_service_bypass=False`` on the admin gate.
- **Shared plumbing.** ``_resolve_coord_session`` helper collapses
the handler prelude three endpoints shared. ``_emit_coord_audit``
wraps ``record_audit`` in a dedicated ``ThreadPoolExecutor``
(``app.state.audit_executor``) so audit bursts don't starve cancel
dispatches. ``_require_json_object`` guards body parsing so non-
object JSON returns 400 instead of 500.
## Sub-PR B — skill metadata governance
- **Description validator (migration 043).** ``prompt_templates``
rows now require a non-empty ``description``. Existing empty rows
get backfilled with a ``"Skill: <name>"`` placeholder on upgrade.
The installer (``admin_skill_discover``) and MCP prompt sync both
synthesise a placeholder when the upstream description is blank
so non-admin write paths satisfy the invariant.
- **Skill kind classifier (migration 044).** New
``prompt_templates.kind`` column (``interactive`` / ``coordinator``
/ ``any``; defaults to ``any``). New
``turnstone.core.skill_kind.SkillKind`` StrEnum is the single
source of truth; Pydantic schemas type ``kind`` as ``SkillKind``
(OpenAPI advertises the enum) and the handler validator catches
the ValueError. ``list_skills_filtered`` gains a
``kinds: list[str] | None = None`` SQL filter.
``CoordinatorClient.list_skills`` defaults to
``kinds=["coordinator", "any"]`` so interactive-only skills are
hidden from the orchestrator.
- **``scan_status`` → ``risk_level`` rename (migration 045).**
Lossless column rename to align with ``IntentVerdict.risk_level``
terminology. Swept storage (both backends + schema + protocol),
handlers, API schemas, tool JSON, generated OpenAPI specs,
TypeScript SDK types, frontend (``governance.js``), tests, and
English prose in ``docs/judge.md`` + ``docs/tools.md``. The
user-facing on-load warning now reads ``has risk level:
{risk_tier}``. Tool JSON's ``risk_level`` enum corrected to the
scanner's actual taxonomy (``safe / low / medium / high /
critical``; was the never-shipped ``clean / flagged / unscanned /
pending``). Historical migration 021 left untouched.
## Migrations
042 (``coordinator.trust.send`` perm — PR A)
043 (description backfill — PR B)
044 (``kind`` column add — PR B)
045 (``scan_status`` → ``risk_level`` rename — PR B)
All four use position-anchored permission strings / host-side
parse-filter-rejoin on downgrade where SQL ``REPLACE`` could
corrupt prefix-overlapping values.
## Verification
- ``ruff check turnstone tests`` clean.
- ``mypy turnstone`` clean on 165 source files.
- ``pytest -m "not live"``: 4431 passed (+85 over the phase-6
baseline). Includes +32 tests in ``tests/test_service_auth_boundary.py``
and +38 in ``tests/test_coordinator_governance.py``; shared fixtures
extracted to ``tests/_coord_test_helpers.py``.
- Generated OpenAPI JSON (``sdk/typescript/openapi-{console,server}.json``)
regenerated via ``sdk/typescript/scripts/generate-types.py``; zero
``scan_status`` occurrences remaining outside the historical
migration 021 and the rename migration 045.
## Security reviews
Both reviews flagged by the phase-7 plan (items 1 + 5, plus 0a's
refuse-to-serve gate) ran through the multi-stage ``/review``
pipeline twice per sub-PR; all confirmed findings landed in-branch.
* fixup(phase-7): CI lint + PR #383 review fixups
Addresses the lint CI failure (ruff format) plus 12 findings from the
two automated PR reviewers.
Copilot:
- ``_sqlite.list_installed_skill_urls`` / ``_postgresql.list_installed_skill_urls``
used positional row indexing (``r[0]``/``r[1]``/``r[2]``) while this
same PR's ``StorageBackend`` class docstring forbids it. Switched
both to ``r._mapping["..."]`` access.
- ``list_skills.json`` previously advertised ``risk_level=""`` as a
filter for unscanned skills, but the implementation treats empty
strings as "no filter". Clarified the tool description to say
omit the filter entirely to include unscanned rows, and added an
explicit ``enum`` on the parameter restricting it to the scanner
tiers. ``_prepare_list_skills`` keeps the ``strip() or None``
normalisation — unscanned filtering now has an unambiguous contract.
- ``test_storage_skills_filtered.test_risk_level_filter`` used the
legacy ``clean`` / ``flagged`` values from the pre-rename column.
Rewritten with the scanner's actual taxonomy (``safe`` / ``high``).
github-code-quality (CodeQL):
- ``test_deny_sentinel_is_singleton`` previously asserted
``cs.DENY_EMPTY_SUB is cs.DENY_EMPTY_SUB`` — an identical-expression
comparison. Rewritten as two separate ``from ... import ... as`` aliases
(``FIRST_READ`` / ``SECOND_READ``) so the identity check is between
distinct bindings.
- ``test_restrict_empty_revoke_is_noop_but_audits`` unpacked ``state``
without using it. Renamed to ``_state``.
- Mixed import styles in ``test_service_auth_boundary.py`` — the
file previously used both ``import turnstone.console.server as cs``
and ``from turnstone.console.server import ...`` for the same
module (same story for ``turnstone.core.auth`` and
``turnstone.server``). Consolidated to the ``from X import Y`` style
used elsewhere in the file; the ``_fetch_live_block`` test now
patches via pytest's ``monkeypatch`` fixture instead of a manual
rebind through a module alias.
CI:
- ``ruff format`` reformatted one line in
``tests/test_coordinator_endpoints.py``.
Verification: ruff check + mypy clean (166 files); 4459 non-live
pytest pass.
* fix(tests): swap asyncio marker for anyio in service-auth boundary tests
PR #383 CI caught that the 13 ``@pytest.mark.asyncio`` decorators I
added in ``test_service_auth_boundary.py`` are an off-convention
choice — the rest of the repo uses ``@pytest.mark.anyio`` (148 sites
vs my 13). The CI environment pulls in ``anyio`` but not
``pytest-asyncio``, so every async test in this one file was failing
with "async def functions are not natively supported". It passed
locally by accident — my dev venv happens to have pytest-asyncio
installed ambiently.
Swapped all 13 marker sites to ``@pytest.mark.anyio``. No functional
change; the tests run under the same default asyncio backend anyio
provides.
Verification: ruff + mypy clean (166 files); 4459 non-live pytest
pass.
* refactor(channels): backfill review of Slack/Discord adapters
Retrospective multi-stage review of the Slack (PR #355) and Discord
channel adapters — they shipped before the review pipeline existed,
so this pass goes back and fixes everything the pipeline would have
caught plus a follow-up round of ultrareview findings.
## Security (8 fixes)
- Adapter-side owner checks on all interactive flows: Discord
ApprovalView / PlanReviewView encode the owner Discord user ID in
the embed footer (`{ws_id}|{corr_id}|{owner_id}`) and reject
non-owner clicks; Slack plan-approve / request-changes /
feedback-modal gain owner tracking in `_pending_plan_review_ts`
and a shared `_ensure_plan_review_owner` gate. These closed the
two critical authz gaps where the gateway's service-scoped JWT
bypassed server-side ownership checks.
- Discord thread-message gate: only the registered invoker can
drive the workstream (prevents a linked user posting in another
user's public thread from injecting into their assistant).
Invoker recorded explicitly so `/ask` follow-ups survive the
`channel.create_thread` bot-as-owner quirk.
- Slack /link flow + per-user identity gate: unlinked Slack users
see an ephemeral `/turnstone link <token>` prompt on every
message instead of silently creating workstreams under the
shared gateway identity. Rate-limited (5/hour) to block online
token enumeration.
- Gateway `/v1/api/notify` requires `write` scope on the validated
JWT; low-scope tokens get 403 + audit.
- Thumbnail URL validator DNS-resolves the hostname before fetch
and rejects any resolved IP that's loopback / link-local /
multicast / reserved, plus an explicit deny-list for IPv6 cloud
metadata (`fd00:ec2::/32` — AWS Nitro IMDS + ECS task metadata)
that would otherwise slip past the `is_private` allowance.
- Per-user rate limit (10 msgs / 60s) + 8 KiB inbound size cap on
Slack DMs / channels / notification-reply threads so one user
can't exhaust the shared LLM budget.
- Discord /link rate limit (5/hour) for token-enumeration defense.
## Bug fixes (9 correctness issues)
- Slack DM routing: each top-level DM no longer spawns a fresh
workstream (was using per-message `ts` as the route key).
- Multi-chunk Slack responses thread correctly under the first
chunk's ts instead of fragmenting as independent top-level
messages.
- Finalize the outgoing StreamingMessage before swapping channel /
thread_ts mid-stream, so buffered tokens still land on the old
thread.
- Redundant `chat_update` on approve/deny eliminated by popping
`_pending_approval[ws_id]` after local resolution.
- Notification reply tracking on Discord only registers for DMs
(guild-channel targets were storing channel IDs where user IDs
were expected, so legitimate replies were always rejected).
- `get_channel_default_alias` rolls `_channel_default_ts` back on
`list_models()` failure so the next caller retries instead of
serving an empty alias for the full TTL.
- Slack `subscribe_ws` purges dead SSE tasks before the
membership short-circuit (previously an unhandled exception left
the ws_id in `_subscribed_ws` forever, silently no-opping
subsequent subscribes).
- ChannelRouter `_create_locks` is now an LRU-bounded OrderedDict
that evicts only unheld locks (original dict grew unbounded;
naive LRU could evict a held lock and let a second caller race
through the critical section, creating duplicate workstreams).
- Slack `_parse_ts` pads the fractional field to 6 digits so
`"1.2"` and `"1.000002"` stop colliding as `(1, 2)` in the
latest-session tiebreaker.
## Performance (6 fixes)
- StreamingMessage keeps a rolling truncated display string capped
at `max_length` so per-flush cost is O(max_length) instead of
O(total_streamed_chars) — long streaming responses no longer do
quadratic work every edit interval.
- `StreamingMessage.finalize()` caches the joined content so the
Discord stream-end DM-forward path doesn't re-join a multi-MB
buffer twice.
- `PendingApproval` stores the Block Kit payload posted to Slack;
`IntentVerdictEvent` appends the verdict in-place and
`chat_update`s, skipping an extra `conversations_history`
round-trip.
- ChannelRouter `lookup_ws_id()` TTL-caches the channel →
ws_id resolution (30s TTL, 4096-entry LRU); hot inbound paths
skip storage on every message.
- Service-discovery startup retry uses exponential backoff
(1s → 8s cap) with a 30s deadline instead of 30 × 1s fixed
sleep.
- `_archive_session` now calls `router.close_workstream` so the
`_node_urls` cache entry is dropped (was leaking one entry per
archived session).
## Quality / refactors (19 improvements)
- `cli.main()` extracted from a 365-line function into focused
helpers; imports carefully kept lazy where test patches target
source-module paths.
- `_run_gateway` finally block now awaits `adapter.stop()` on
every adapter so SSE tasks, httpx clients, and the Slack socket
handler close cleanly on shutdown.
- Shared SSE reconnect loop extracted to `turnstone/channels/_sse.py`
(`run_sse_stream` with `on_event` + `on_stale` callbacks); both
adapters' `_sse_listener` methods just wire up callbacks. The
"404 stops reconnect" invariant is enforced inside the helper
so a broken `on_stale` can't livelock.
- `_on_ws_event` god-dispatchers split into per-event `_handle_*`
methods with a thin isinstance dispatcher at the top.
- Slack `_on_approve` / `_on_deny` collapsed into a single
`_resolve_approval(*, approved: bool)`.
- `ApproveRequestEvent` policy evaluation hoisted into
`ChannelRouter.evaluate_tool_policies` returning a
`PolicyVerdict`; adapters switch on the verdict kind.
- `ChannelAdapter` protocol trimmed to the four methods adapters
actually implement; unused `ChannelEvent` dataclass removed.
- Shared constants lifted to `turnstone/channels/_config.py`.
- `_cleanup_stale_route` and `unsubscribe_ws` share a
`_clear_ws_state` helper.
- `StreamingMessage` private attrs promoted to `message` /
`message_ts` / `accumulated_text` properties so callers don't
reach past the `_`-prefix.
- Various cleanups: dead var, noqa'd lambdas, renamed
`_policy_handled` → `policy_handled`, inlined single-use
helpers, added module docstrings, documented
`SlackRoute.parse` edge cases.
- `chunk_message` plain-text fast path (no backticks → skip
fence bookkeeping).
## Test coverage
Added 45 tests (178 → 223):
- `tests/test_channel_sse.py` (new) — SSE reconnect / backoff /
404-stale-route / on-stale-exception / invalid-JSON-skip /
on-event-exception-doesn't-kill-stream / per-connection token
refresh / ConnectError retry.
- ApprovalView + PlanReviewView owner-check regression tests
(owner allowed, non-owner rejected, legacy 2-pipe footer fails
closed, modal path rejected for non-owner, `/ask`
bot-as-thread-owner follow-up allowed).
- Slack `_recover_routes` latest-ts-wins, `_archive_session`
drops route + closes workstream.
- SSRF tests: DNS rebinding rejected, IPv4 link-local metadata
rejected, IPv6 ULA metadata (fd00:ec2::254 / fd00:ec2::23)
rejected.
- Slack link prefix match (natural-language prompts don't
hijack), link rate-limit ceiling.
- SlackRoute round-trip across all three shapes + lax-parse
behaviour.
Lint (ruff) + mypy clean; 210 channel-focused tests pass.
* chore(channels): address PR #382 review-bot feedback
Three line-level findings from github-code-quality on the backfill
review PR. Copilot had no line-level comments.
- _sse.py:132 — the `except httpx.HTTPStatusError: pass` branch was
flagged as an empty except. The original status was already logged
at WARNING inside the try block (we re-raise ourselves after
logging), so the handler has real intent. Added a debug log of the
exception text + a comment explaining the control flow, so the
empty-except lint stops firing and the next reader sees why we
fall through to backoff.
- discord/bot.py:430, cli.py:354, slack/bot.py:1127 — `await task`
inside `contextlib.suppress` was flagged as "statement has no
effect". It's a false positive (await is an effect) and the
alternative try/except/pass triggers ruff SIM105. Kept the
contextlib.suppress pattern and added an explanatory comment above
each call so the intent (await CancelledError propagation before
state cleanup) is obvious; will reply on the PR thread noting the
false positive.
No behavior change. Lint + mypy clean; 210 channel tests pass.
* feat(coordinator): phase 6 — polish, observability, active-coords via SSE, frontend cleanup
Squashed from two working commits:
1. phase-6 backend polish + active-coords SSE
2. phase-6 frontend cleanup (legacy chat-view classes + designer nits)
Both tier-A/B observability items and tier-C frontend consolidation
ship together — the shared-vocabulary migration touches surfaces the
backend polish already had its hands in, so one combined commit keeps
the diff reviewable as a coherent phase.
Observability
-------------
- **Coordinator-side wait dashboard** — `_exec_wait_for_workstream`
emits `wait_started` / `wait_progress` / `wait_ended` SSE events
via a new `progress_callback` hook on
`CoordinatorClient.wait_for_workstream`; coordinator.js renders a
"⧗ waiting · N ws · Ts" header indicator keyed by call_id so
overlapping waits coexist. Progress throttled to emit only on
snapshot-diff or 5s heartbeat; full results dict attached only on
transitions so a 600s wait doesn't flood SSE listener queues.
Indicator only attaches when a proper header host exists (no
floating document.body fallback) and is cleared on SSE reconnect
so a dropped `wait_ended` can't pin the badge.
- **`cancel_workstream` forensics** — `server.cancel_generation`
captures `ui._pending_approval` tool names +
`session._queued_messages` count / preview before invoking
`session.cancel`, returning the snapshot as `dropped`; routing
proxy passes it through to the tool result. Preview runs through
`redact_credentials` before the 120-char truncate so pasted
secrets / connection strings don't land verbatim in the
coordinator's conversation history.
- **Per-coordinator metrics** — `GET /v1/api/coordinator/{ws_id}/metrics`
returns `spawns_total` / `spawns_last_hour` / `child_state_counts`
/ `judge_fallback_rate` (substring match on verdict.tier) plus
zero placeholders for wait_* pending dedicated instrumentation.
Derived from new `storage.count_workstreams_by_state` +
`count_workstreams_since` aggregate helpers — no 10k-row
hydrated-select to compute a histogram. Ownership 404-mask
matches `coordinator_detail`.
- **Coordinator skill in inspect** — `CoordinatorManager.create`
resolves `skill` → `template_id` / `applied_version` via
`get_skill_by_name` + new `storage.count_skill_versions`
(replacing the SELECT-all-for-COUNT anti-pattern) and persists
them on the workstreams row. `/new` handler dispatches via
`asyncio.to_thread` so blocking storage calls don't stall the
event loop.
- **`wait_for_workstream(since=…)`** — optional prior-snapshot
hint; when supplied, the wait loop diffs each polled ws_id that
IS in `since_map` and exits on any change, independent of mode.
ws_ids absent from `since_map` fall through to the normal mode
condition — a disjoint since dict no longer silently exits the
wait on tick one.
- **`task_list.child_ws_id` referential cleanup** —
`CoordinatorClient.cleanup_dead_task_child_refs(ws_id)` holds the
same per-ws `_task_lock` as `task_list_*` so a close racing a
task_list write can't lose the mutation. `CoordinatorManager.close`
delegates. Final save-failure logs at `warning` instead of
`debug`.
Home view live-updates
----------------------
- **Active-coordinators via SSE** instead of a 5s poll —
`ClusterCollector.ensure_console_pseudo_node` +
`emit_console_ws_created / _closed / _state / _rename` plumbing;
`CoordinatorManager.create / open / close / eviction` +
`ConsoleCoordinatorUI.on_state_change / on_rename` all fan out
through the collector. The pseudo-node is exempt from the
discovery-loop eviction; rehydrate-path eviction now also emits
`console_ws_closed` for the evicted row so other tabs drop it
live. `app.js` reads coordinators from
`clusterState.nodes["console"]`; poller + back-compat shims
deleted (9 call sites). Overview / nodes list skip the
pseudo-node so it doesn't inflate cluster totals. Tenant-filtering
preserved by excluding the pseudo-node from
`collector.get_workstreams` so `/v1/api/cluster/workstreams` still
uses the existing tenant-filtered `_coordinator_rows` path.
`CoordinatorManager.NODE_ID` bound from
`ClusterCollector.CONSOLE_PSEUDO_NODE_ID` so the two literals
can't drift.
Frontend perf
-------------
- **Bulk cluster-ws live endpoint** — `GET /v1/api/cluster/ws/live?ids=`
returns `{results, denied, truncated}` (cap 50); coordinator.js
batches visible-row live-badge fetches into one bulk request per
~250ms window (replaces per-row /detail polling). Ownership
check routes through the empty-string-safe pattern (non-admin
with empty `caller_uid` doesn't match empty-owner rows).
Legacy chat-view class cleanup
------------------------------
- Drop the `.msg` / `.msg-user` / `.msg-assistant` / `.msg-tool` /
`.msg-error` / `.msg-info` / `.approval-block` / `.approval-tool`
/ `.approval-btn` / `.approval-badge` / `.approval-prompt` /
`.approval-feedback-input` / `.approval-actions` / `.pane-input` /
`.pane-input-area` / `.pane-input-row` / `.pane-attach` /
`.pane-attach-chip` / `.coord-msg` / `.coord-body` / `role-*` /
`btn-approve` / `btn-deny` / `btn-always` / `verdict-glow-*`
legacy dual-class names left over from the phase-4 migration.
Every JS className concatenation + querySelector + CSS selector
now uses the `ts-*` vocabulary from `shared_static/chat.css` (and
`ui/static/style.css` where the interactive-page extensions
live). Feature-specific class names that don't map to `ts-*`
stay — `msg-queued` / `msg-editing` / `msg-actions` / `msg-edit-*`
/ `msg-user-attach*` / `msg-user-text` / `queued-badge` /
`queued-dismiss` / `tool-name` / `tool-cmd` / `tool-diff` /
`tool-header` / `tool-preview`.
Designer nits
-------------
- **`.ui-btn--icon:focus-visible`** — new rule matching `.ui-btn`'s
`outline: 2px solid var(--accent); outline-offset: 1px` so the
compact icon variant gets the accent ring instead of the
browser-default outline.
- **Dropped speculative 701-880px composer wrap rule** — the flex
math at ≥701px fits comfortably in every desktop viewport, so the
mid-zone break rule was forcing a 2-line layout where the browser
wouldn't have wrapped naturally. The existing `<700px` full-stack
covers the original wrap observation.
- **`.verdict-badge` border-top** + **`.ch-row.highlight`
prefers-reduced-motion** — confirmed already in main; no
additional code change needed for phase 6.
Follow-up designer-review findings
----------------------------------
- Dropped `border-top` from `.ts-approval-badge` + `.ts-approval-body`
(chat.css's max-content width / flex-gap made them read as
truncated / floating lines).
- `var(--muted)` → `var(--fg-dim)` on denied tool names (undefined
token was silently failing).
- Dropped 3 dead `.ts-approval-badge.badge-*` rules + duplicated
`.ts-approval-btn:focus-visible` + dead `.reasoning` CSS rule +
`contains("reasoning")` JS guard.
- Dropped `tool-row` / `approval-header` / `btn-row` / `label` dead
legacy classes in the coordinator.
- Added `:focus-visible` to `.ch-row a.ws-link` + `.task-row` so
keyboard users get the accent ring on sidebar rows.
Cleanups
--------
- `_WAIT_REAL_TERMINAL_STATES` / `_WAIT_TERMINAL_STATES` /
`_WAIT_MAX_*` / `_WAIT_POLL_INTERVAL` hoisted to module level on
`coordinator_client` so `session.py` no longer reads a class
internal; ClassVar aliases kept for back-compat.
- `ConsoleCoordinatorUI` state/rename observers typed as
`Callable[[str], None] | None` instead of `Any`.
- SSE error renderer verified end-to-end (coordinator.js already
handles `case "error"` → `appendText`; no code change).
Tests
-----
- 26 new test cases: 11 for `_diff_since` + `cleanup_dead_task_child_refs`,
15 for `cluster_ws_live_bulk` + `coordinator_metrics`. Full suite
4345 passing (4319 base + 26 phase-6).
Gate: ruff + mypy + pytest -m "not live" (4345 passed) all clean.
* fix(coordinator): address PR #381 review feedback
Copilot comments:
- Cross-tenant aggregate leak in coordinator_metrics — the new
count_workstreams_by_state / count_workstreams_since aggregates
took parent_ws_id but not user_id, so a non-admin caller could
observe drifted / forged child rows that share parent_ws_id with
their coord but whose user_id drifted to another tenant. The
404-mask on coord ownership (_resolve_coordinator_or_404) is the
primary defense; this is defense-in-depth inside the aggregate
queries. Pass filter_user_id (None for admin, caller_uid for
non-admin) — matches coordinator_children's tenant-push-into-SQL
pattern.
- wait_for_workstream(since=…) docstring + tool schema were stale —
claimed "A missing entry counts as changed on first observation"
but the implementation ignores ws_ids absent from since_map to
prevent a disjoint since dict from silently exiting on tick one.
Rewrote both doc sites to match the actual semantics: only ws_ids
present in `since` participate in the diff-exit check; others fall
back to the normal mode-based completion condition.
- WAIT_TERMINAL_STATES comment drift — the comment claimed it was
"used by the resolved-count summary" but the summary counts only
WAIT_REAL_TERMINAL_STATES (denied is a rejection, not a
resolution). Rewrote the comment to describe the real usage:
mode='any' pure-denied short-circuit + mode='all' settle check.
github-code-quality (CodeQL):
- coordinator.js — dropped the dead typeof _renderWaitIndicator
guard + the typeof activeWaits guard around the reconnect clear.
Both symbols are defined in the same IIFE; the onopen handler
fires strictly AFTER IIFE execution finishes, so the guards
always evaluated to true. Removing the dead branching also
removes a CodeQL nit.
- Protocol-method `...` statements — the bot flagged the three new
methods (count_workstreams_by_state / count_workstreams_since /
count_skill_versions) with "statement has no effect". Left as
`...` to match the file's universal convention (216 `...` bodies
/ 0 `pass` bodies pre-change); swapping just the new methods to
`pass` would introduce inconsistency with every other Protocol
method. Resolved as non-actionable.
Test: new test_metrics_tenant_filter_excludes_forged_cross_tenant_child
covering the aggregate-query tenant filter with both a legitimate
alice child and a forged bob child sharing parent_ws_id. Non-admin
alice sees 1; admin sees 2.
Gate: ruff + mypy + pytest -m "not live" (4346 passed) all clean.
* fix(server,console): kind filter on saved-workstreams + closed coords on landing
Two independent bugs folded into one hotfix:
1. Coordinators leaking into the interactive UI's "saved workstreams"
sidebar. ``list_workstreams_with_history`` (SQLite + postgres) was
kind-agnostic — every coordinator row with conversation history came
back alongside interactive rows, and ``list_saved_workstreams``
serialized them uniformly with no kind field so the interactive UI
rendered coordinators as regular interactive entries.
Fix: add optional ``kind: WorkstreamKind | str | None = None`` kwarg
on ``list_workstreams_with_history`` (storage protocol + both
backends + the ``turnstone.core.memory`` helper). Pass
``kind=WorkstreamKind.INTERACTIVE`` from the /v1/api/workstreams/saved
handler so the interactive surface only sees interactive rows.
Default ``None`` preserves legacy all-kinds behaviour for any
other caller that wants both.
2. Closed coordinators vanish from the console landing page.
``_coordinator_rows`` in console/server.py built dashboard rows
exclusively from the in-memory ``CoordinatorManager`` registry,
which pops rows on ``close()``. The persisted storage row stays
(state='closed') but never reached the landing-page poller at
/v1/api/cluster/workstreams?node=console.
Fix: two-lane merge in ``_coordinator_rows``. The in-memory lane
(manager) stays authoritative for live session state (model /
model_alias / current state / tokens). A new persisted lane queries
``storage.list_workstreams(kind=COORDINATOR, user_id=uid, limit=200)``
and appends rows NOT already in the in-memory set — surfacing
closed / error / deleted coordinators so the operator can still
see them on the landing page. Ownership semantics unchanged —
non-admin callers only see their own tenant, admin-bypass via
admin.users/admin.roles honored on both lanes, empty-string
defense-in-depth matches _check_row_owner_or_404.
Tests:
- tests/test_storage_sqlite.py — two new tests: kind filter excludes
coordinators from the history list; string form of kind accepted
(matches the memory.py forwarding shape).
- tests/test_coordinator_endpoints.py — four new tests:
- closed coordinators from storage surface alongside active ones.
- in-memory row wins on ws_id dedup (live state authoritative).
- persisted rows respect tenant filter (non-admin, admin bypass).
- orphan rows (empty user_id) never leak to empty-sub callers.
Gate: ruff + mypy + pytest -m "not live" (4315 passed) all clean.
* fix(server,console): address Copilot review on PR #380
Three review comments folded in:
1. Tenancy leak in /v1/api/workstreams/saved — the handler called
list_workstreams_with_history without a user_id filter, so any
authenticated user could see every other user's saved workstream
aliases / titles / names. Fix:
- Add ``user_id: str | None = None`` kwarg to
list_workstreams_with_history on the protocol + both backends
(SQLite + postgres). Pushes the filter into SQL.
- memory.py helper forwards the kwarg.
- /v1/api/workstreams/saved reads ``_auth_scopes(request)``: a
service-scoped caller gets cluster-wide visibility (None), a
non-service caller with a blank ``sub`` returns an empty list,
otherwise the SQL filter is scoped to the caller's uid. Matches
the _visible_workstreams pattern used on /workstreams and
/dashboard.
2. Loose type annotation on the memory.py helper — ``kind: Any``
tightened to ``WorkstreamKind | str | None`` so mypy catches
invalid callers. WorkstreamKind was already imported in the
module.
3. Brittle positional indexing in _coordinator_rows persisted-rows
lane — ``row[10]`` for user_id encoded a column offset that would
silently corrupt the projection on any future SELECT reorder.
Drop the test-double fallback entirely; the storage-protocol
contract already requires SQLAlchemy Row with _mapping, and every
real caller (SQLite + postgres) provides it.
Tests:
- test_server_authz.py TestSavedWorkstreamsTenantScoping — four new
regression tests covering: non-service caller sees only own rows,
service scope sees cluster-wide, blank-sub non-service returns
empty, and coordinator rows excluded even for service callers.
Gate: ruff + mypy + pytest -m "not live" (4319 passed) all clean.
* fix(console): service scope on collector token + surface upstream 4xx
CRITICAL: the console's ClusterCollector ServiceTokenManager was
configured with only frozenset({"read"}) scope, but every upstream
node's /v1/api/events/global hard-gates on "service" scope (added in
PR #375 for cross-tenant authz hardening). Every console→upstream
SSE connect 403'd, the collector never populated node state, and the
failure was silent — node health, idle workstreams, and interactive-
kind workstream rows all disappeared from the console dashboard with
no user-visible error. The only surface was a log.debug line in the
collector's _node_sse_task that operators had to opt into via DEBUG
logging or browser DevTools.
Fix:
- Add "service" to the collector_token_mgr scopes
(turnstone/console/server.py). Matches the proxy_token_mgr (which
already has it) and the existing cli / admin / channel-gateway
service tokens. Restores /v1/api/events/global SSE subscription
and /v1/api/dashboard visibility (which silently tenant-filters
non-service callers to zero rows).
- Upgrade the 4xx path in _node_sse_task to log.warning with the
status code + 200-char body preview, so configuration-level
failures (scope misconfig, JWT secret mismatch, expired token)
show up in operator logs instead of being masked by the generic
except-block debug line. Keep transient network errors
(CancelledError, ConnectError) at debug so the log doesn't flood
during brief node restarts.
- Add reachable_reason field to NodeSnapshot + surface via
get_nodes / get_node_detail / get_snapshot (and the browser's
buildNodeInfoFromSnapshot). Operators now see the failure cause
on the cluster node list without tailing the log. Cleared on
successful reconnect in _apply_snapshot.
- Test coverage: test_server_authz.py TestGlobalEventsServiceGate
gains a positive-path test asserting that a token with exactly
the collector's scope set ({"read", "service"}) is accepted by
/v1/api/events/global. Locks in the scope contract so any future
rename breaks the test before it breaks the dashboard.
Gate: ruff + mypy + pytest -m "not live" (4309 passed) all clean.
* fix(console): address Copilot review on PR #379
Two review comments folded in:
- collector.py — bounded body read for 4xx SSE error previews. The
prior ``await source.response.aread()`` buffered the entire
upstream error body into memory just to log a 200-char preview; a
malicious / oversized upstream response (HTML error page, proxy-
generated body) could have forced the collector to download an
arbitrary amount of bytes. Iterate ``aiter_bytes()`` and stop once
the preview cap (256 bytes, ~200 chars after UTF-8 decode) is
satisfied.
- test_server_authz.py — tighten the service-scope positive test.
The prior ``assert resp.status_code != 403`` could pass on
unrelated 500s AND left an SSE stream open indefinitely. Send
``?expected_node_id=definitely-wrong-node-id`` so the handler
passes the scope gate, hits the post-auth node-identity check, and
returns 409. Now ``assert resp.status_code == 409`` proves the
scope contract precisely and terminates the request immediately.
Gate: ruff + mypy + pytest -m "not live" (4309 passed) all clean.
* feat(coordinator): phase 5 — harness-test polish + wait_for_workstream + judge fix
Closes the bug list surfaced by the 2026-04-17 coordinator harness test
plus the post-phase-4 wait_for_workstream ask, and folds in three
adjacent cleanups that landed in the same window. Tightens defense-in-
depth on the model-invoked mutating ops, fixes the LLM judge silent
no-op, kills the inspect-poll token burn, and rounds out a handful of
observability / docstring / spec gaps.
The session-factory pre-resolve at console/session_factory.py and
server.py was rewriting `judge.model` from an alias (e.g. `judge-mini`)
to the resolved underlying id (e.g. `gpt-5-mini`). IntentJudge then
checked `model_registry.has_alias(config.model)`, found nothing, and
fell back to the SESSION's provider/client with that bare model id —
silent `llm_fallback / "did not return a verdict"` whenever the
coordinator and judge alias resolved to different providers.
Pass the alias through unchanged; IntentJudge's existing alias-
resolution path picks up the matching client + provider. Validate
the alias exists so an obvious typo still surfaces, but don't replace
the model field.
Regression: `test_alias_uses_registry_provider_not_session_provider`
constructs an alias whose provider differs from the session's and
asserts the judge picks up the alias's provider/client/model;
`test_coordinator_tool_call_returns_llm_verdict_not_fallback` asserts
the verdict tier is `llm` (not `llm_fallback`) on the happy path.
New `cancel_workstream` tool (approval required, primary_key=ws_id) —
cancels in-flight generation, unblocks any pending approval / plan,
moves the child to idle, leaves the row in storage so a fresh
send_to_workstream lands cleanly. Re-uses the existing
`/v1/api/route/cancel` route + `route.cancel` audit namespace; no
new server endpoint.
`CoordinatorClient.cancel/close_workstream/delete/send` now enforce
a tenant guard inline (`_is_own_subtree`) — only the coordinator
itself or one of its own children is targetable. Foreign ids return
the same 404-shape inspect/wait_for_workstream use, so the model
can't distinguish foreign from missing (no existence oracle).
Defense-in-depth — the upstream node enforcement is the perimeter,
this is the second line.
`list_workstreams` advertised `state="deleted"` and an
`include_closed=true` that surfaced deleted rows. Hard-deletes
cascade the workstream + conversation rows out of storage, so
deleted is unreachable in normal operation. Doc-only fix; the
synthetic-test path that registers `state="deleted"` rows still
works (terminal-state filter still excludes them via
`_terminal_states = {"closed", "deleted"}` in list_children).
Documented that the 120s service-registry heartbeat window means a
node returned by list_nodes can drop out before a follow-up
`spawn_workstream(target_node=…)` lands — the spawn fails with "No
available node for routing" rather than falling back. Two-line
clarification on each tool. No code change (a code fallback is a
bigger discussion deferred to 1.6).
`close_workstream` accepts `reason`; the upstream server handler now
persists it to `workstream_config.close_reason` (capped at 512 BYTES,
sliced on UTF-8 not code points so a CJK / emoji-heavy payload can't
4× the documented budget). `CoordinatorClient.inspect()` reads it
and surfaces as `close_reason` in the result dict — only for
terminal-state children (closed/error/deleted) so the live-child hot
path doesn't pay a per-inspect DB round-trip.
Tests: server-side persistence covers success / no-reason /
length-cap / non-string / storage-failure / multi-byte-utf8 paths;
client-side surface covers terminal vs. live workstreams.
For idle children whose node-dashboard live counter is 0 (the live
block only surfaces in-flight token counters), fall back to
`SUM(prompt_tokens + completion_tokens)` from `usage_events` so the
inspect output reflects cumulative spend.
New `storage.sum_workstream_tokens(ws_id) -> int` on the protocol +
both backends. The fallback is folded INTO `_fetch_cluster_live` so
the merged live block (with persisted total applied) is what gets
cached — back-to-back inspects of an idle child amortize through
the existing 2s LRU cache instead of each firing a fresh aggregation.
`CoordinatorClient.list_skills()` now projects `allowed_tools` per
skill — capped at 20 with a `+N more` sentinel so a skill that
whitelists a wide MCP surface doesn't bloat the per-row payload.
Reads the existing `prompt_templates.allowed_tools` column; no
storage change. Coordinators no longer have to guess what tools a
skill brings.
`route_create` now sets `routing_strategy: "hash_ring" | "target_node"
| "resume"` on the spawn response so the coordinator's spawn
response (and the `spawn_workstream` tool output) carries why a
given node was chosen. 3 lines + 3 covering tests in
test_console_routing_proxy.py.
New coordinator tool `wait_for_workstream(ws_ids, timeout=60,
mode='any'|'all')` that absorbs the wait into a single tool call —
the model sees one call + one result regardless of how long the
children take. Kills the busy-poll inspect loop that burned 20+
turns on a 3-child fan-out.
Storage-poll loop with batched primitives —
`get_workstreams_batch` + `sum_workstream_tokens_batch` issue exactly
two storage calls per tick regardless of N. At the cap (32 ws_ids /
600s / 0.5s tick) that's ~2400 round-trips for a full wait, down
from ~38k under the naive per-id shape.
Validation single-source-of-truth: the client owns mode whitelist,
ws_ids dedup + cap, timeout coerce + clamp. The session preparer
is a thin pass-through that builds the header + dispatches; bad
input surfaces at exec time as a tool error via `result.get("error")`.
Tenant-isolation collapse: missing-row and cross-tenant cases both
return `state="denied"` so wait can't be used as an existence oracle
(matches the 404-mask contract `inspect` uses).
Prompt-side: tools_coordinator.md adds a `wait_for_workstream`
pattern + an explicit "PREFER wait_for_workstream OVER a loop of
inspect_workstream" line in the workflow-shape section.
Replaces the quote-bracketed substring LIKE/ILIKE pattern with proper
JSON-array containment. The previous shape effectively did
`LOWER(tags) LIKE '%"<lower-tag>"%'`, which broke for tag values
containing `"` (the JSON encoder escapes it to `\"` and the literal-
substring search misses), `\` (encoded as `\\`), or non-ASCII
characters that the encoder rendered as `\uXXXX`. Also exposed a
small spoofing surface — `tags=["foo\","bar"]` would have matched a
query for `bar`. Real-world tag values are alphanumeric+dash today
so it hadn't fired in production, but the fix is small.
- SQLite: `EXISTS (SELECT 1 FROM json_each(prompt_templates.tags)
WHERE lower(value) = lower(:tag))` (JSON1 extension; SQLite 3.38+).
- PostgreSQL: `EXISTS (SELECT 1 FROM jsonb_array_elements_text(
prompt_templates.tags::jsonb) AS jat(elem) WHERE lower(jat.elem) =
lower(:tag))`.
Three new tests prove the substring pattern was broken for
quoted / backslash / unicode tag values; the existing case-fold +
wildcard tests continue to pin the contract.
Phase 1 added the coordinator workstream API; phase 2 added only
`/open` to the OpenAPI catalog and missed every other coordinator
endpoint plus phase 3's `/children`, `/tasks`, and the
`/cluster/ws/{ws_id}/detail` aggregator. SDK consumers + operators
browsing `/docs` couldn't discover the surface. Doc-only addition:
12 endpoints + 9 new Pydantic models, all under the `Coordinator`
OpenAPI tag so /docs groups them together.
Sidebar re-fetches `GET /tasks` on every `task_list` `tool_result`
SSE event. A model that runs `add → list` (or any back-to-back
mutation pair) double-fetches the same envelope. Coalesced into
one fetch per 150ms window via a new `loadTasksDebounced` wrapper;
direct UI actions (refresh button, page load) keep calling
`loadTasks` directly so user clicks aren't delayed.
- `ruff check turnstone tests` — clean
- `mypy turnstone` — clean (157 source files)
- `pytest -m "not live"` — 4284 passed, 3 deselected (was 4226 on
main; +58 new tests across coordinator client, tools, judge,
storage, console routing proxy, server close-handler,
storage_skills_filtered, OpenAPI catalog, server close-reason
persistence)
- New tools added: 2 (cancel_workstream, wait_for_workstream) —
TOOLS count 28 → 30; coordinator subset 9 → 11; auto_approve adds
wait_for_workstream; primary_key adds cancel_workstream
- New OpenAPI endpoints: 12 (every phase-1/2/3 coordinator route +
the cluster-inspect aggregator)
- New storage protocol methods: 3 (sum_workstream_tokens,
sum_workstream_tokens_batch, get_workstreams_batch)
All phase 1 / 2 / 3 / 4 invariants preserved: COORDINATOR_TOOLS /
INTERACTIVE_TOOLS disjoint; coordinator sessions have no MCP surface;
list-style tools return {items, truncated}; route-proxy emits
route.<action> audit on 2xx; 404-mask on ownership failures; tenant
filters pushed into SQL; per-coordinator JWT carries scope context.
* fix(coordinator): address Copilot review on PR #378
Three valid Copilot findings on the wait_for_workstream surface:
1. ``wait_for_workstream.json`` description claimed the tool returns a
top-level mapping ``ws_id -> {state, tokens, updated}`` plus
elapsed/complete/mode at the same level, but the actual shape is
``{results: {ws_id: {...}}, elapsed, complete, mode}``. Description
now matches the implementation. Also adds ``deleted`` to the
advertised terminal-state list (it's in ``_WAIT_REAL_TERMINAL_STATES``;
the doc and runtime now agree).
2. ``CoordinatorClient.wait_for_workstream`` docstring listed
``idle / error / closed`` as the real terminal set but the constant
includes ``deleted``. Same fix — list ``deleted`` with a parenthetical
noting it's unreachable in normal operation (hard-delete cascades the
row).
3. Storage protocol docstring math: ``sum_workstream_tokens_batch``
claimed "from ~38k to ~1200" round-trips per wait at the cap, but
``wait_for_workstream`` issues TWO storage calls per tick
(``get_workstreams_batch`` + this one), so 1200 ticks × 2 = ~2400.
Updated to "~2400" with the math spelled out.
Also a clean rebase onto today's main (PR #377 — the rebalancer node_id
snapshot doc — landed since phase 5's last push). Single conflict in
``inspect_workstream.json`` resolved by keeping both notes (rebalancer
node_id binding semantics + the new ``close_reason`` surface from phase
5); ``spawn_workstream.json`` auto-merged.
The github-code-quality bot also flagged three items on
``_protocol.py`` asking to replace ``...`` with ``pass`` in Protocol
method bodies. Refuted: ``...`` is the canonical PEP 544 idiom for
Protocol method bodies and the rest of the file uses it consistently.
The bot's lint rule misfires for ``Protocol`` classes.
Verification:
- ``ruff check turnstone tests`` clean
- ``mypy turnstone`` clean (158 source files)
- ``pytest -m "not live"`` — 4308 passed, 3 deselected (no test count
change; pure doc/comment edits)
Phase 3 fixed spawn_workstream's response to return the storage-
authoritative node_id at spawn time, but neither tool description
mentioned that the cluster rebalancer can migrate the workstream to
a different node afterwards. A coordinator that cached the
spawn-time node_id for a long-running callback would silently dispatch
to a node that no longer owns the workstream.
- spawn_workstream: ``node_id`` is a POINT-IN-TIME snapshot at spawn;
re-read with inspect_workstream when you need the current binding.
- inspect_workstream: ``node_id`` is the CURRENT (storage-authoritative)
binding; reflects any rebalancer migration that happened since spawn.
Pure description edit — no schema or runtime change.
Third and final PR of the retrospective-review series. Addresses the
remaining bug / perf / doc findings from the original multi-stage review
plus the three inline comments left on #374 and #375.
From the original review:
- bug-3: delete_workstream now nulls out parent_ws_id on every child
row before dropping the target — previously, deleting a coordinator
left orphaned parent_ws_id pointers and list_workstreams(parent_ws_id=
<deleted>) kept returning ghost-parented rows. Fix lives at the
storage edge so both SQLite and PostgreSQL benefit without a schema
migration.
- perf-1 / perf-2 / perf-3: new migration 041 drops the low-cardinality
idx_workstreams_kind outright, rebuilds idx_workstreams_parent as a
partial index (WHERE parent_ws_id IS NOT NULL) to halve its btree,
and uses CREATE INDEX CONCURRENTLY on postgres so the rebuild
doesn't take ACCESS EXCLUSIVE on populated tables. Dialect-guarded;
sqlite path is a straight partial CREATE INDEX.
- perf-5: _rebuild_children_from_storage bumps its limit sentinel to
10_000 and logs a warning when the cap is hit instead of silently
truncating the tail on every console cold-start.
- q-2: turnstone.core.memory.list_workstreams wrapper deleted (zero
live callers; PR #374 kept it forward-compatible with the new
kwargs as a stepping stone).
- q-5: migration 039's docstring now warns operators that downgrade
drops parent_ws_id irreversibly and notes the 041 dependency.
- q-7: GET /v1/api/workstreams row shape now includes kind +
parent_ws_id to match /v1/api/dashboard; the Pydantic
WorkstreamInfo schema follows so SDK consumers see the same fields.
Inline review comments:
- #374 (copilot): console/server.py::coordinator_children now pushes
user_id into the SQL filter for non-admin callers, so forged /
migration-era rows with matching parent_ws_id but a different
owner can't leak through. Admins bypass the filter — they're
expected to see the full subtree.
- #375 (copilot, delete handler): storage.get_workstream(ws_id) for
the audit snapshot moved inside the try: block so a transient DB
error surfaces through the endpoint's redacted 500 handler instead
of an unhandled exception.
- #375 (copilot, _require_ws_access): added optional mgr= kwarg —
when the workstream is live in the in-memory manager, trust its
cached user_id instead of round-tripping storage. In-memory-only
handlers (approve / plan / cancel / command / close / events_sse /
refresh-title / set-title) pass mgr= so they stay functional
during transient DB outages and skip one query on the hot path.
Storage-backed handlers (/delete, /open) omit mgr= and keep the
storage path for persisted-but-not-loaded rows.
Tests:
- tests/test_workstream_kind.py adds regression tests for the cascade
null-out on delete and the new user_id SQL filter.
- tests/test_workstream_endpoints.py updated so the title-handler
tests exercise the in-memory fast path (MagicMock manager returning
None falls through to storage; explicit ws.user_id set where the
mock ws is used).
Lint (ruff), typecheck (strict mypy), pytest -m 'not live' all green
(4209 passing).
Second of three PRs addressing the retrospective review of the
turnstone-server interactive-kind feature. The first (PR #374) put
the structural pieces in place — WorkstreamKind enum + user_id
kwarg on the storage protocol. This PR uses them to close the
handler-level ownership gaps that shipped under the prior design.
- sec-1: approve / plan_feedback / cancel_generation / command now
call _require_ws_access before touching the target UI. Previously
any authenticated user could resolve pending tool-approvals on
another tenant's workstream — RCE-adjacent because the attacker
could approve destructive operations the victim would have denied.
- sec-2: /v1/api/workstreams/{ws_id}/delete now gates on ownership
AND writes a workstream.deleted audit event. Previously any
authenticated user could destroy any other tenant's workstream,
conversations, and attachments in one call with no tamper-evident
trail.
- sec-3: /v1/api/events (per-ws SSE) gates before _register_listener
so non-owners can't subscribe to another tenant's message / tool /
approval stream.
- sec-4 / sec-5: /v1/api/workstreams and /v1/api/dashboard filter
to the caller's tenant view via a new _visible_workstreams helper;
service-scoped tokens (cluster / routing proxy) keep the full view.
- sec-6: /v1/api/events/global requires service scope. The global
snapshot carries cross-tenant workstream inventory and was never
intended for end-user browsers.
- sec-7: /v1/api/workstreams/{ws_id}/open verifies the caller is
the stored owner (or holds service scope) before rehydrating.
Returns 404 on mismatch — existence isn't enumerable by response
code.
- sec-8 / sec-9: /workstreams/close, /refresh-title, /title all gate
on ownership. Cross-tenant close aborts the victim's running
generation; cross-tenant rename is a phishing / denial-of-use
vector in list / dashboard responses.
- sec-11: workstream.created / .deleted / .closed / .opened now
land in the audit_events table with kind + parent_ws_id detail,
so forensic review can reconstruct lifecycle even after the row
is gone.
- q-4: new tests/test_server_authz.py covers every gate above via
TestClient, plus the PR #1 HTTP-boundary kind-validation branches
that had no regression coverage (coordinator / unknown-kind / 400,
cross-tenant parent_ws_id / 403, non-interactive open / 400).
- q-3: test_workstream_kind.py now uses the conftest storage fixture
so it runs against both SQLite and PostgreSQL under
--storage-backend=postgresql, closing the sqlite↔postgres drift
risk the prior review flagged. Added storage-edge ValueError and
user_id SQL filter tests alongside.
Tests, lint (ruff), typecheck (strict mypy) all green. Stacked on
PR #374 — merges after that lands.
Foundation PR for the multi-stage-review follow-up. Introduces a
single source of truth for workstream kind values and pushes tenant
scoping into the storage protocol so list callers can't forget to
filter client-side.
- WorkstreamKind(StrEnum) replaces bare "interactive" / "coordinator"
literals across 17 production modules. Strict mypy narrows every
internal call site; raw strings still work at wide boundaries
(HTTP body, DB row) via WorkstreamKind(raw) parse at the edge.
- StorageBackend.list_workstreams(..., user_id=None) adds a SQL-level
WHERE user_id = :user_id gate on both sqlite and postgres impls.
Memory wrapper forwards the new filters.
- register_workstream now validates kind at the storage edge so SDK /
restore / internal callers can't silently corrupt the NOT NULL
column with empty / mis-cased / unknown values.
- WebUI.__init__ normalizes empty-string parent_ws_id to None, matching
the storage-edge and WorkstreamManager invariants.
- POST /v1/api/workstreams/new parses body["kind"] through the enum
and returns 400 on unknown kinds instead of silent coercion.
Absorbs bug-1, bug-2, bug-4/q-6, q-1, q-8, and partial q-2 (wrapper
signature forwards the new filters; full deletion of the unused
wrapper stays in the cleanup PR).
* feat(ui): phase 4 — chat-UX unification + coordinator-first console landing
Phase 4 unifies the three turnstone UIs (server-node chat, console
dashboard, coordinator page) around a shared design-system layer,
promotes coordinator sessions to first-class citizens on the console
landing, and folds the chat-view itself onto a shared vocabulary so
the two chat pages no longer reinvent messages / approvals / composer /
header / sidebar chrome from scratch.
## Shared static consolidation
- turnstone/shared_static/renderer.js — consolidates the two copies
(ui/static/ + console/static/coordinator/) into one. Adds
streamingRender / streamingRenderFinalize helpers with
requestAnimationFrame coalescing + per-element buffer cache so both
chat views re-render the streamed markdown smoothly without the
prior "plain-text → final pop" on the coordinator page and without
thrashing renderMarkdown + DOM replacement faster than the paint
cycle. renderMarkdown stays the trust boundary for innerHTML
assignment (escapeHtml internal); postRenderMarkdown (syntax
highlighting, mermaid, KaTeX) is deferred to finalize.
- turnstone/shared_static/ui-base.css — flat form-control + button +
state-glyph + pill + panel vocabulary on top of base.css. Sizes in
px to match the 11/12/13px scale used elsewhere. Namespace rubric
documented inline (.ui-* shared controls; .dash-* dashboard legacy;
.ts-* chat vocabulary; page-local stays unprefixed).
- turnstone/shared_static/chat.css (new) — chat-view component
vocabulary: .ts-msg (user / assistant / reasoning / tool / error /
info), .ts-msg-actions floating toolbar, .ts-approval (inline +
batch layout hooks sharing a visual language), .ts-verdict-badge,
.ts-composer shell, .ts-header shell, .ts-sidebar shell. Mobile +
reduced-motion covered.
## Interactive server UI migration
- turnstone/ui/static/app.js — dual-class adoption of .ts-msg + .ts-
approval + .ts-composer + .ts-msg-actions alongside existing class
names so feature-specific rules (.msg-user-text, .msg-queued,
.msg-editing, .msg-action-btn toolbar, attachment chips, verdict
details, media embeds, plan inline) keep working while the shared
chat.css baseline takes over padding / border / typography.
- Left-aligned user messages: .msg-user loses align-self: flex-end
and .msg-assistant loses align-self: flex-start. Both roles now
render as single-column blocks distinguished by left-border colour
(amber for user, neutral for assistant, dashed for reasoning,
mono + code-bg for tool, red for error) per the locked design.
- style.css trimmed: .msg / .msg-info / .msg-error baseline rules
dropped (chat.css provides); all other feature rules intact.
- index.html links /shared/chat.css and tags the header with
.ts-header + .ts-header-title.
## Coordinator page migration
- coordinator.js appendMsg drops the visible .role-label <div> per
the hybrid no-labels design, preserving the role text on
data-ts-role + aria-label so screen readers and SSE dedup-by-call-
id continue to see meaningful labels. Adds .ts-msg + .ts-msg--*
variants + .ts-msg-body onto the existing .coord-msg / .coord-body
elements.
- coordinator/index.html adopts .ts-header, .ts-header-title,
.ts-header-spacer, .ts-header-status on the header; .ts-approval +
.ts-approval--batch on the pinned bar with .ts-approval-btn
variants on the buttons; .ts-composer + .ts-composer-input +
.ts-composer-send on the composer; .ts-sidebar + .ts-sidebar-
section + .ts-sidebar-section-heading on the children + tasks
sidebar. Inline <style> pared from ~300 to ~130 lines — only
genuinely coordinator-specific layout (flex wiring, sidebar list
rows, mobile accordion breakpoint) remains.
- Nits fixed along the way: .ch-row .glyph-thinking recoloured cyan
to match the shared .ui-glyph vocabulary; .task-row .status-done
lost its 0.7 opacity (colour already signals done; opacity
reduced contrast for no gain).
## Console landing + admin panel redesign (phase 4 scope-expansion)
Replaces the node-list-first console landing with a coordinator-first
layout on a new #view-home pane:
- #coord-composer-panel — persistent "Start a new coordinator task"
composer (textarea + optional name + skill dropdown + submit).
Permission-gated on admin.coordinator (same rule the +coordinator
header button uses). Pre-probes GET /v1/api/coordinator on init
and after login (bug-1 fix) so a 503 (no coordinator.model_alias
resolvable) surfaces as a remediation banner linking to Admin →
Models instead of failing on submit. Probe gates on r.ok instead
of r.status !== 503 (bug-2 fix) so auth/permission errors don't
incorrectly flip the banner to ready. Composer and modal share a
_createCoordinator helper (q-1 fix) — POST + redirect + error-
handling tail is not forked.
- #active-coordinators — SSE-driven list of kind=="coordinator"
workstreams, rendered through the shared _renderWsRow helper so
state glyphs + child-count badges match the existing tree view.
- #cluster-summary-compact — one-line aggregate. Clicking expands
into the legacy #view-overview via showOverview() so deep-link
callers of ?view=overview / ?view=node / ?view=filtered keep
working unchanged.
View switching consolidated into a _setLandingView helper so every
show* / drillDown* function toggles the four landing panes through
one call path. Default currentView flipped from "overview" to
"home"; popstate + init history.replaceState land on {view: "home"}.
patchClusterState preserves kind / parent_ws_id / user_id on
ws_created events (phase 3 invariant) so the active-coordinators
list picks up new coordinators immediately without a snapshot
refetch.
Header H1 is now a home link so operators have a single-click path
back to the coordinator landing from any drill-down / admin view.
## Design polish (review pipeline fixes)
- .home-panel-title dropped from 13px/accent to 11px/fg-dim so it
sits in the same heading tier as .ui-section-heading / .dash-
header-title / .home-section-title instead of outweighing them
(dsn-3).
- .home-composer-banner recoloured from amber-on-amber-glow to
bg-surface + 1px yellow border + fg-bright text + accent link
with thicker underline (dsn-1).
- .ui-pill--done dropped the 0.75 opacity — colour signals done,
opacity reduced contrast for no gain (dsn-6).
- .ui-heading fleshed out with --sm/--md/--lg tiers so the utility
actually conveys size (dsn-12).
- ui-base.css size scale moved from rem to px matching 11/12/13px
(dsn-2).
* fixup: address Copilot feedback on PR #373
- admin.js: drop the stale `#view-overview` display:none mutation in
showAdmin. #view-overview is now nested inside #view-home and
toggled via the `hidden` attribute; inline display:none here would
stick after returning to home and suppress the cluster-details
expand.
- app.js: reword the _renderHomeView token-bucket fingerprint comment
to match the actual `Math.floor(tokens / 100)` bucketing — the
prior comment said "thousands / sub-thousand drift".
* feat(coordinator): tree-view UI, cluster-wide live inspect, dashboard grouping — phase 3
Closes out the 1.5 coordinator UX surface: a right-sidebar tree view at
/coordinator/{ws_id} showing spawned children + task list, a new
cluster-wide live inspect endpoint that powers the tree's live badges,
and 2-level dashboard tree grouping that nests spawned children under
their coordinator parent.
## Cluster-wide live `inspect_workstream`
New `GET /v1/api/cluster/ws/{ws_id}/detail` on the console, gated by a
new `admin.cluster.inspect` permission (unassigned to any builtin role;
operators opt in). Aggregates `storage.get_workstream` with a
short-timeout (2s) HTTP fetch against the owning node's
`/v1/api/dashboard`. Coordinator-hosted workstreams get their `live`
block from the in-process `CoordinatorManager` instead of a proxy hop.
Response shape `{persisted, live, messages}` — `live: null` on node
unreachability / 5xx / missing-entry with status 200 so the UI can
degrade gracefully without an error state. Correlation-id masks
unexpected exceptions. 404-masks cross-tenant reads (non-admin
callers see only their own workstreams).
`CoordinatorClient.inspect()` best-effort merges the `live` block onto
its storage snapshot so the model-facing `inspect_workstream` tool
gains a `live` key without any schema change. Model-facing tool
schema stays identical.
## Tree-view UI
New right sidebar at `/coordinator/{ws_id}` with a 2-level children
tree + the phase-2 task list.
Backend:
- New `GET /v1/api/coordinator/{ws_id}/children` returns
`{items, truncated}` — identical row shape to the `list_children`
tool — filtered via `storage.list_workstreams(parent_ws_id=..., kind=None)`.
- New `GET /v1/api/coordinator/{ws_id}/tasks` returns the
`{version, tasks}` envelope via the shared module-level
`load_task_envelope` decoder (extracted from `CoordinatorClient`
so both the tool path and the UI read share corruption semantics).
Corrupt envelopes return an empty list for UI resilience — the
`task_list` tool remains the authoritative write + error path.
- `CoordinatorManager` subscribes to the `ClusterCollector`'s
listener channel from the console lifespan and dispatches filtered
`child_ws_created / child_ws_state / child_ws_closed / child_ws_rename`
events onto each coordinator's SSE stream. Filter authoritative
on the server via a per-coordinator child-ws_id registry populated
lazily on `open()` from storage and incrementally on `ws_created`
events; cleared on `close()` / eviction. One SSE connection per
client, no client-side filtering.
Frontend:
- DOM-method-only child-row rendering (no innerHTML of user content).
- State glyph vocabulary (● running / ◐ thinking / ⚠ attention /
✗ error / ○ idle) plus text labels — WCAG 1.4.1 carries info in
both glyph and label.
- Live badges (tokens + pending-approval pip) fetched via
`/cluster/ws/{ws_id}/detail` with a 5s TTL cache and 250ms debounce
per child. One request per state change, not per second.
- SSE child events update in place; renderChildren() re-sorts.
- Mobile (<700px) sidebar collapses to an accordion above the chat
with a toggle button flipping aria-expanded; a `.highlight` flash
marks task→child scroll targets; `prefers-reduced-motion` respected.
- Deep-link child rows to `/node/{node_id}/?ws_id=<child>` via
`<a target="_blank" rel="noopener">` with encodeURIComponent on
regex-validated ids.
## Dashboard tree grouping
Cluster dashboard rows now group by `parent_ws_id`. Coordinator rows
(`kind == "coordinator"` or children present) get an expand/collapse
caret (button with `aria-expanded`); collapsed shows "(N children)".
Expanded renders children indented as sibling rows with a left-border
gutter. Orphaned children (parent missing or closed) render at top
level with a muted "orphan" badge. Expansion state persisted in
`localStorage` keyed per coordinator ws_id so operator preference
survives reloads. Coordinator rows deep-link to `/coordinator/{id}`;
node-backed workstreams keep their existing proxy deep-link.
Per-node `ws_created / ws_state / ws_activity` SSE event payloads
gained `parent_ws_id` + `kind` so the collector can propagate them
through its fan-out to browser clients without a second lookup;
`_build_node_snapshot` and `/v1/api/dashboard` rows include the
same. Coordinators (which don't live on cluster nodes) merge into
`/cluster/workstreams` via a new `_coordinator_rows` helper that
threads them through the collector's `get_workstreams(extra_rows=...)`
parameter — extras share the filter / sort / paginate pipeline with
node-backed rows.
## Tests
- `tests/test_coordinator_endpoints.py` — 19 new cases covering
children (empty / populated / ownership 404 / admin bypass /
invalid ws_id / truncation), tasks (empty / round-trip / corrupt /
ownership), and cluster-inspect (auth gates / 400 / 404 / ownership /
coordinator self-path / unloaded-live-null / message-limit clamp).
- `tests/test_coordinator_manager.py` — 8 new cases covering registry
bootstrap on create + open, dispatch for each event type,
unrelated-parent filtering, shutdown idempotency.
- `tests/test_console.py` — existing `cluster_workstreams` assert
updated for the new `extra_rows` kwarg.
## Verification
- `ruff check turnstone tests` clean.
- `mypy turnstone` clean.
- `pytest -m "not live"` — 4184 passed, 3 deselected.
* fix(coordinator): race in dispatch + ui_factory kwarg filtering — PR #370 review
Addresses feedback from the GitHub Copilot + code-quality bot review
passes on PR #370.
## Race in _dispatch_child_event ws_created branch
Copilot flagged a TOCTOU where the lock-free read of
``self._active_coords`` (line 912) could see the parent coordinator,
then ``close()`` / eviction pops ``_children[parent]`` + drops the
coord from ``_active_coords`` before we acquire ``_children_lock``,
and then ``setdefault(parent, set())`` resurrects the entry —
leaking the registry key forever and fanning events to a closed UI.
Fix: re-check ``parent in self._active_coords`` inside
``_children_lock``. The reference swap is still atomic; holding
``_children_lock`` and re-reading the snapshot catches the race
without serializing back through ``self._lock``.
Regression test: create → close → dispatch a ws_created → assert
neither ``_children`` nor ``_active_coords`` regained the entry.
## ui_factory kwarg filtering via inspect.signature
code-quality bot flagged that the previous ``try ui_factory(…, kind=,
parent_ws_id=) except TypeError`` dance fired on every call with
legacy test factories (``lambda wid: WebUI(ws_id=wid)``) — wasteful
and masks real signature mismatches.
Fix: inspect the factory's signature and only pass kwargs it
actually accepts (explicit param name OR ``**kwargs`` absorber).
Keep a conservative ``except TypeError`` fallback for C-callables
and odd signatures ``inspect`` can't introspect.
Copilot also flagged a comment mismatch (the old comment said
"KeyError on **kwargs" — it's ``TypeError``, which is what the code
caught). The rewritten comment is correct.
## Nit: side-effect in assert
code-quality bot flagged ``assert mgr.close(ws.id)`` in
test_coordinator_manager.py. Split into two statements.
## Verification
- ``ruff check`` clean.
- ``mypy turnstone`` clean.
- ``pytest -m "not live"`` — 4223 passed, 3 deselected, 0 failed.
* feat(coordinator): audit middleware on routing proxy — phase 2
Adds per-tool-call audit attribution to the multi-node routing proxy
handlers so coordinator → server hops land observable rows in
``audit_events``. Phase 1 preserved the ``src="coordinator"`` claim
through ``_proxy_auth_headers``'s upstream re-mint; this commit
closes the recording side. Was the last real security gap from
phase 1 — an enterprise deployment with ``admin.coordinator``
granted got only the three console-side
``coordinator.{create,close,cancel}`` rows; per-tool-call
attribution was missing.
## Action-naming scheme
route.workstream.create POST /v1/api/route/workstreams/new
route.workstream.send POST /v1/api/route/send
route.workstream.close POST /v1/api/route/workstreams/close
route.workstream.delete POST /v1/api/route/workstreams/delete
route.approve POST /v1/api/route/approve
route.cancel POST /v1/api/route/cancel
route.command POST /v1/api/route/command
route.plan POST /v1/api/route/plan
Action-name conventions documented in ``turnstone/core/audit.py``
module docstring alongside the existing namespaces — the docstring
is now ``<resource>.<verb>`` shaped (non-exhaustive) rather than
trying to enumerate every prefix.
## Recording rules
- ``record_audit()`` fires only on a 2xx upstream response. 4xx/5xx
are observable via ``_record_route``'s metrics path; doubling the
audit-events table size for failure rows would dilute signal
without giving operators much extra value.
- ``detail`` JSON carries ``{src, node_id, coord_ws_id?}`` — ``src``
lands verbatim from ``auth.token_source`` so non-coordinator
origins (``"jwt"``, ``"console-proxy"``) also get attribution;
``coord_ws_id`` only appears when the inbound JWT carried it.
- Wrapped in ``try/except`` + ``log.debug("route.audit_failed", ...)``
defence-in-depth. ``record_audit`` itself is fire-and-forget;
the outer try guards against a programmer error in the call site.
## Routing-proxy specifics
- ``route_create``: emits at the post-multipart/JSON convergence
``if resp.status_code == 200`` block. Both branches set
``audit_ws_id`` correctly — multipart from the query-string ws_id,
JSON from ``body["ws_id"]`` (post-503-retry) or ``body["resume_ws"]``.
- ``route_proxy``: emits the URL-method-mapped action. ``ref`` is
reassigned to ``new_ref`` after a successful 404→cache-refresh
retry so audit attribution uses the retried node, not the failed
first node.
- ``route_workstream_delete``: emits on 2xx using the ws_id from
the request body.
- ``route_attachment_proxy``: out of scope (upstream attachment
endpoints emit their own ``workstream.attachment.*`` rows;
auditing here would double-count).
## Tests
16 new tests in ``tests/test_route_proxy_audit.py`` covering:
- Coordinator-origin emission with full detail payload.
- 502 / 400 / 503-retry-final-node-id paths.
- Parametrised method→action mapping for the 6 ``route_proxy`` URLs.
- Plain-JWT origin (no ``coord_ws_id`` in detail).
- Delete handler 2xx + 502.
- Audit-storage exception swallowed (proxied response unchanged).
- ``auth_storage`` absent → no-op (existing route-handler tests
unaffected).
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` 4087 passed,
3 deselected (live-backend), 0 failed.
* feat(coordinator): discovery tools and /open parity — list_nodes, list_skills, POST /coordinator/{ws_id}/open
Adds the read-side surface coordinators need to make informed
orchestration decisions plus an explicit rehydration endpoint
matching the server's ``POST /v1/api/workstreams/{ws_id}/open``.
## list_nodes (auto-approved)
``list_nodes(filters={key: value, ...})`` reads ``node_metadata`` via
``storage.filter_nodes_by_metadata`` + ``get_all_node_metadata`` —
one query each, no N+1. Each row carries its full metadata dict so
the coordinator has both auto keys (``arch`` / ``cpu_count`` /
``fqdn`` / ``hostname`` / ``os`` / ``os_release`` / ``python``;
always present) and operator-supplied user keys (``capability`` /
``region`` / ``tenant`` / ``role``) without a second round-trip.
Tool description enumerates the auto keys explicitly so the model
knows what's always available vs deployment-specific.
Storage stores metadata values as JSON-encoded strings (the write
path in ``server.py`` / ``admin.py`` / ``console/server.py`` all go
through ``json.dumps``). The client re-encodes filter values
before the stored-text comparison and decodes stored values before
returning them to the model — so ``{"capability": "gpu"}`` is the
natural form the model uses, not ``{"capability": "\"gpu\""}``.
Ints round-trip as ints.
Returns ``{nodes, truncated}``; ``truncated=True`` when the page
was full.
## list_skills (auto-approved)
``list_skills(category?, tag?, scan_status?, enabled_only?, limit?)``
surfaces the skill registry so coordinators can discover worker
profiles. New storage protocol method ``list_skills_filtered(...)``
on both SQLite and PostgreSQL backends pushes filters into SQL.
``tag`` filter matches against the JSON-array ``tags`` column with
quote-bracketed substring (``%"tag"%``) — quote-safe against
``foo`` vs ``foobar`` collisions on both backends.
Returns ``{skills, truncated}`` with ``name`` / ``category`` /
``tags`` (decoded to list) / ``version`` / ``description`` /
``model`` / ``enabled`` / ``scan_status`` / ``activation`` — the
discovery projection, not the full row.
## POST /v1/api/coordinator/{ws_id}/open
Explicit rehydration endpoint. Lazy ``GET`` rehydration works for
the UI; this gives SDK callers and operators a way to warm a
coordinator without browsing to it. Same ownership / 404-on-
mismatch / correlation-id-masked error semantics as
``coordinator_detail``. Returns ``{ws_id, name, already_loaded?}``.
Registered in ``turnstone/api/console_spec.py`` with a dedicated
``CoordinatorOpenResponse`` Pydantic model so the OpenAPI schema
matches the wire shape.
## Tests
- ``tests/test_storage_skills_filtered.py`` — 8 cases validated on
BOTH SQLite and PostgreSQL backends (``pytest --storage-backend
postgresql``). Covers no-filter ordering, category exact-match,
tag quote-safety (``"foo"`` matches ``["foo","bar"]`` but not
``["foobar"]``), scan_status, enabled_only, limit, AND semantics,
empty result.
- ``tests/test_coordinator_client.py`` — 11 new cases covering
node/skill shape decoding, JSON-encoded filter round-trip (the
``"gpu"`` vs ``'"gpu"'`` case), int filter encoding, truncation,
no-match empty, no N+1 (``get_prompt_template`` /
``get_node_metadata`` call counts asserted zero).
- ``tests/test_coordinator_tools.py`` — 11 new cases for
``_prepare``/``_exec`` dispatch, filter type-drop, limit clamping
(``limit=0`` falls back to 100, negatives clamp to 1),
truncation-signal summary.
- ``tests/test_coordinator_endpoints.py`` — 8 new cases for
``/open``: ``already_loaded`` on in-memory hit, 404 on ownership
mismatch, lazy rehydrate on miss, admin bypass, unknown ws_id,
503 on ``coord_mgr`` unavailable, 500 with correlation-id mask on
factory failure, 503 passthrough on ``ValueError``.
- ``tests/test_workstream_kind.py`` / ``test_tools_schema.py``
updated to include ``list_nodes`` and ``list_skills`` in the
disjoint-namespace regression guard and the tool-count check.
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` 4122 passed, 3
deselected (live-backend), 0 failed. Postgres backend storage
tests green (``pytest --storage-backend postgresql
tests/test_storage_skills_filtered.py`` 8 passed).
* feat(coordinator): task_list tool — persistent planning state
Adds a coordinator-only ``task_list`` tool persisted on the
coordinator's own ``workstream_config`` row. Gives coordinators a
scratch surface for work decomposition that survives restarts so the
UI can render planned-vs-done state once the tree view lands.
## Tool surface
``task_list(action, ...)`` with five actions:
- ``list`` auto-approved read. Returns ``{tasks, truncated}``;
truncated=True when the list exceeded the 200-row
page cap.
- ``add`` needs approval. ``title`` required; optional
``status`` and ``child_ws_id``. Title clamped at
200 chars. Capacity cap at 500 tasks — hitting the
cap is an explicit signal to prune done/blocked rows.
- ``update`` needs approval. Mutate by ``task_id``; fields
``title`` / ``status`` / ``child_ws_id`` optional.
- ``remove`` needs approval. Drop by ``task_id``.
- ``reorder`` needs approval. Pass ``task_ids``; validated as an
exact permutation of the current set (rejects
partial, extra, or substituted ids — prevents silent
task loss).
Status enum: ``pending`` / ``in_progress`` / ``done`` / ``blocked``.
``child_ws_id`` links a task to the child workstream spawned for it
(no enforcement; the coordinator owns the relationship).
## Persistence
Stored as a single JSON-envelope value on ``workstream_config`` —
``{"version": 1, "tasks": [...]}``. No new table; the kanban v2
work will supersede this row via a format migration keyed on
``version``. ``_save_task_list`` writes only the ``tasks`` key so
concurrent writers to other ``workstream_config`` keys (e.g. the
admin Settings UI updating ``reasoning_effort``) aren't clobbered
by a read-modify-write on the full row.
## Corrupt-envelope safety
A hand-edited or legacy config row that doesn't parse as the
expected shape logs a warning and returns an empty envelope from
``task_list_get``. Mutators refuse to overwrite corrupt data —
they detect the sentinel and return a clear error so the operator
can inspect or clear the row rather than losing work silently.
## Concurrency
Per-(ws) ``threading.Lock`` cached on the client. The worker
thread is single-threaded for tool execs so this is mostly
defence-in-depth against future maintenance-script call sites.
Cache never grows beyond one entry per coordinator session because
the scope guard short-circuits foreign ``ws_id`` before the lock
is acquired.
## Malformed-JSON recovery
``_prepare_tool`` fallback-1 regex-extract allowlist extended with
``action`` / ``status`` / ``task_id`` / ``title`` (alphabetized) so
slightly-malformed ``task_list`` calls get the same
self-correction behaviour as the other coordinator tools.
## Tests
- ``tests/test_coordinator_client.py`` — 15 new cases covering:
fresh-envelope shape, add/get roundtrip, empty-title + invalid-
status rejection, 200-char title clamp, update by id + missing
id, remove semantics, reorder permutation validation (partial +
extra + wrong id + valid), cross-ws scope violation, corrupt-
JSON read recovery, corrupt-envelope write refusal (all four
mutators), 500-task capacity cap, workstream_config key
preservation across ``_save_task_list``.
- ``tests/test_coordinator_tools.py`` — 12 new cases covering the
dispatch layer: list auto-approved, each mutating action needs
approval, unknown-action / missing-required-arg errors, list
returns tasks, page-cap at 200 with truncated signal, add
dispatches to client, reorder surfaces permutation error,
remove-not-found.
- ``tests/test_tools_schema.py`` / ``tests/test_workstream_kind.py``
extend the tool-count + disjoint-namespace + primary-key
regression guards with ``task_list``.
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` 4148 passed,
3 deselected (live-backend), 0 failed.
* feat(coordinator): coordinator workstream kind — phase 1
Adds a new ``kind="coordinator"`` workstream that runs inside the
``turnstone-console`` process (first ChatSession hosted on the console)
with a dedicated tool set for spawning and driving child workstreams.
Supersedes the external ``turnstone-coordinator`` MCP side-car for new
installs; the extension is marked deprecated in
``examples/mcp-cluster-ops/README.md`` but still works on 1.4-and-earlier
clusters.
Phase 1 ships: the workstream class, 6 lifecycle tools, console hosting,
9 HTTP endpoints, per-user audit attribution, and a one-pane web UI at
``/coordinator/{ws_id}``. Node/skill discovery tools, task-list tool,
tree-view UI, and routing-proxy audit middleware follow in a later PR.
## Schema
Migration 039 adds ``kind`` / ``parent_ws_id`` columns + indexes to
``workstreams``. Both SQLite and PostgreSQL backends take the new
kwargs on ``register_workstream``; empty-string ``parent_ws_id``
normalises to ``NULL`` at the storage edge. PostgreSQL uses
``INSERT ... ON CONFLICT DO NOTHING`` to match SQLite's ``OR IGNORE``
and close a pre-existing SELECT-then-INSERT TOCTOU window.
``list_workstreams`` gains optional ``parent_ws_id`` / ``kind`` filters;
new ``get_workstream(ws_id)`` returns the full row (the existing
``get_workstream_metadata`` stays untouched for back-compat).
## Core session + kind routing
- ``ChatSession.__init__`` accepts ``kind`` / ``parent_ws_id`` /
``coord_client``. On ``kind="coordinator"`` it swaps
``_tools = COORDINATOR_TOOLS`` and zeros sub-agent tool lists.
- ``Workstream`` dataclass extended with ``user_id`` / ``kind`` /
``parent_ws_id``. Both ``WorkstreamManager`` and the new
``CoordinatorManager`` use the same type — no parallel hierarchy.
- ``_SessionFactory`` Protocol + server / cli factory closures thread
the new kwargs. ``POST /v1/api/workstreams/new`` rejects
``kind != "interactive"`` with 400; ``POST
/v1/api/workstreams/{ws_id}/open`` refuses coordinator rows so a
server node can't accidentally rehydrate one.
## Coordinator tool set
Six tools (``spawn``, ``inspect``, ``send``, ``close``, ``delete``,
``list_workstreams``) with a ``coordinator: true`` metadata flag,
scoped to coordinator-kind sessions only. ``inspect`` and ``list`` are
auto-approved reads; the four mutators need approval. ``list`` returns
``{"children": [...], "truncated": bool}`` so the model can detect
post-filter under-fill and paginate.
## CoordinatorClient (in-process, sync)
Mutating ops HTTP-POST to the console's own ``/v1/api/route/*`` on the
local bind URL so every existing middleware (auth, rate-limit) runs.
Read ops hit ``storage.list_workstreams`` / ``get_workstream`` /
``load_messages`` directly — the routing proxy doesn't expose
list/inspect paths. URL paths are a validated constant table (avoids
an httpx ``base_url``-merge trap). A new
``/v1/api/route/workstreams/delete`` proxy handler joins the existing
route-proxy endpoints.
## Per-session coordinator JWT
``CoordinatorTokenManager`` mints short-lived JWTs with ``sub=<real
user>`` (attribution preserved), ``src="coordinator"``,
``aud="turnstone-console"``, ``coord_ws_id=<ws>`` custom claim.
``_proxy_auth_headers`` preserves ``src`` + ``coord_ws_id`` across the
upstream re-mint so server-side middleware sees coordinator-origin,
not ``console-proxy``. ``AuthResult.extra_claims`` carries
non-reserved claims through validate→remint; ``create_jwt``'s
reserved-claim set (now including ``nbf`` / ``jti``) is symmetric with
``validate_jwt``.
## Console hosts the ChatSession
- New ConfigStore settings: ``coordinator.model_alias`` (required),
``reasoning_effort``, ``max_active`` (default 5),
``session_jwt_ttl_seconds``.
- Console lifespan builds a ``ModelRegistry`` +
``CoordinatorManager``. Missing / unresolvable alias returns **503**
with remediation text — never 500.
- ``CoordinatorManager``: placeholder-slot reservation under lock,
rollback on factory failure, per-ws_id rehydration lock to serialise
concurrent lazy-opens, ``max_active`` enforced via ``close_idle``
eviction semantics.
- ``ConsoleCoordinatorUI`` is a thin ``SessionUI`` implementation — no
global broadcast, no per-node metrics, shared
``_APPROVAL_WAIT_TIMEOUT`` constant across approval + plan paths.
- No eager startup rehydration: persisted coordinator rows load lazily
on first ``GET /v1/api/coordinator/{ws_id}``.
## Console coordinator API
Nine endpoints under ``/v1/api/coordinator/*`` gated by ``approve``
scope + new **``admin.coordinator``** permission (added to
``_VALID_PERMISSIONS``; not in any builtin role — operators opt in
explicitly). Ownership failures return **404, not 403** and use
strict equality so empty-owner rows don't leak across tenants.
Correlation-id masking on every factory-raising path
(``coordinator_create`` + ``coordinator_detail`` lazy rehydrate) — no
stack traces to the client.
## Audit attribution
Three console-side events (``coordinator.create`` / ``.close`` /
``.cancel``) with the real creator's ``user_id`` plus
``detail={coord_ws_id, src="coordinator"}``. No schema migration
required. Per-tool-call audit across the routing proxy is deferred
(needs either a ``source`` column on ``audit_events`` or
``record_audit`` calls wired into the route-proxy handlers).
## Web UI (``/coordinator/{ws_id}``)
One-pane chat served by the console. Reuses ``shared_static``
(``base.css``, ``auth.js``, ``theme.js``, ``toast.js``, ``utils.js``,
``kb.js``) and the server UI's ``renderer.js`` pipeline (KaTeX, Mermaid,
highlight.js already bundled).
- SSE to ``/v1/api/coordinator/{ws_id}/events`` with exponential-
backoff reconnect; status line carries a leading glyph
(● / ○ / ⚠) so state isn't conveyed by colour alone.
- Renders content, reasoning (dimmed italic
``.role-reasoning``), tool_result, approve_request, intent_verdict,
output_warning.
- Child ws_id references auto-wrap to
``/node/{node_id}/?ws_id={child}`` links — both ids regex-validated
before interpolation, everything else HTML-escaped.
- Non-modal approval bar (``role="region"``) with a batch header
("Approve N tool calls"), initial focus on the approve button,
buttons disabled during the in-flight POST, red-bordered deny.
``aria-live`` flips to ``off`` during streaming.
- "New coordinator" button on the dashboard header — permission-gated
on the UI side, matching the backend 403.
- Mobile composer capped under ``@media (max-width: 700px)``.
## Tests
~120 new tests across 8 files: workstream-kind storage + dataclass
semantics, CoordinatorClient URL map + token minting + storage reads +
truncation signalling, tool prepare/exec dispatch and approval gating,
CoordinatorManager create / rollback / eviction / lazy rehydration +
concurrency, HTTP endpoint auth + 404-on-ownership + 503-on-misconfig,
proxy-auth ``src`` preservation, full lifecycle end-to-end, coordinator
page HTML-injection guard. ``test_tools_schema.py`` widened to 25
tools (19 existing + 6 coordinator).
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 files), ``pytest`` 4054 passed (5 pre-existing failures unrelated
to this change — confirmed against ``main``).
* polish(coordinator): address PR review + CI + tool-namespace isolation
CI:
- `ruff format`: two files reformatted, matches the in-repo pre-commit config.
- `wheel-completeness`: add `turnstone/console/static/coordinator/*.html` +
`*.js` to the hatch wheel-include list. Without this the coordinator UI
was missing from published wheels.
- `test (3.11/3.12/3.13)` + `test-postgres`: three `TestExecReadImage`
tests were masking a real bug — my 6 new tool JSONs pushed tool count
19→25, crossing the default `tool_search.auto` threshold (20), which
made `ChatSession.__init__` construct a `ToolSearchManager` and cache
`_cached_capabilities` during init. Tests that later patched
`session._provider.get_capabilities` saw the cached value instead.
Root-cause fix: the tool-search threshold code path now reads
capabilities through `_resolve_capabilities(...)` directly — no cache
populate — so the patch takes.
Tool-namespace isolation (bigger fix than CI symptoms suggested):
- `TOOLS` was the union of all loaded tool JSONs including the 6 new
coordinator tools. Interactive sessions were getting coordinator
tools in their function-calling surface (which is nonsense — they
require a console-hosted `coord_client`), and coordinator sessions
counted against the interactive tool-search threshold. Fix:
- New `INTERACTIVE_TOOLS` / `INTERACTIVE_TOOL_NAMES` in
`turnstone/core/tools.py` exclude anything with `coordinator: true`
metadata. `TOOLS` stays as the union for schema introspection +
eval catalog.
- `ChatSession.__init__` selects tool set by kind: coordinator gets
fixed `COORDINATOR_TOOLS` (no MCP merge, no listeners registered);
interactive gets `INTERACTIVE_TOOLS` (+ MCP if configured).
Coordinators are meta-orchestrators that spawn child workstreams;
MCP tools / resources / prompts live on the children, not on the
coordinator's own surface.
- `_on_mcp_tools_changed` no-ops for coordinator sessions
(defence-in-depth in case listeners were registered).
- `always_on_names` on `ToolSearchManager` is now the set of builtin
tools actually present in the session (kind-aware) rather than the
full `BUILTIN_TOOL_NAMES` frozenset.
- `turnstone/eval.py` uses `INTERACTIVE_TOOLS` (coordinator tools
aren't in scope for the eval harness which tests interactive agent
behaviour).
- Regression tests in `tests/test_workstream_kind.py`:
- `INTERACTIVE_TOOLS ∩ COORDINATOR_TOOLS == ∅` and their union is
`TOOLS`.
- Interactive `ChatSession._tools` does not include any
coordinator tool name.
- Coordinator `ChatSession._tools` contains `spawn_workstream` but
not `bash` / `edit_file` / `memory`; sub-agent lists are empty.
- Coordinator `ChatSession` with an MCP client attached does NOT
merge MCP tools and does NOT register any MCP listeners.
PR review findings:
- **#10 / #11** (Copilot): coordinator UI claimed to reuse the server
renderer pipeline but loaded none of its JS. Mirrored
`turnstone/ui/static/renderer.js` into
`turnstone/console/static/coordinator/renderer.js` (flagged in-file
as a cleanup candidate to promote into `shared_static/`), added
`katex.min.js` / `highlight.min.js` / `renderer.js` script tags to
`coordinator/index.html`. `coordinator.js` now buffers raw markdown
via `textContent` during streaming, then swaps to `renderMarkdown` +
`postRenderMarkdown` on `stream_end`.
- **#7** (Copilot): N+1 query pattern in
`CoordinatorClient.list_children()` — per-row `storage.get_workstream`
just to read `skill_id`. Pushed `skill_id` + `skill_version` into
the `list_workstreams` SELECT projection on both backends; the
client reads them from `row._mapping` directly. New
`test_list_children_skill_filter_avoids_n_plus_one` pins the
behaviour (asserts `storage.get_workstream` call count is 0).
- **#8 / #9** (Copilot): `spawn_workstream` tool JSON said "if empty,
the workstream is created idle" but the prepare method rejected
empty and the field was marked required. Resolved by allowing
empty end-to-end: removed from `required`, prepare builds a
"spawn idle workstream" header + empty preview when empty,
updated `test_spawn_prepare_allows_empty_initial_message`.
- **#1–#5** (github-code-quality): five asserts with side-effecting
method calls in `test_coordinator_manager.py` (`mgr.close`,
`mgr.open`, `mgr.create` in a dead `_c = ...`). Extracted each
call to a local variable so `python -O` can't strip the side
effect.
Verification:
- `ruff check turnstone tests` clean.
- `mypy turnstone` clean (156 source files).
- `pytest -m "not live"` — 4063 passed, 3 deselected (live-backend
tests), 0 failed. The 3 image tests that were failing on this
branch now pass; wheel + lint both green locally.
* polish(coordinator): address Copilot re-review findings
Two findings from the re-review of #368 after the first polish commit.
**user_id wired into `mgr.create()` at the server handlers.** Phase 1
added ``user_id`` to the ``Workstream`` dataclass and
``WorkstreamManager.create()`` signature, but the two call sites in
``turnstone/server.py`` forgot to pass the authenticated caller
through. Result: interactive workstreams created via
``POST /v1/api/workstreams/new`` (including coordinator-spawned
children, which route through this handler) were landing with blank
``user_id``, defeating ownership-based access control on subsequent
sends / approvals / closes (``_require_ws_access`` treats blank
owners as legacy/allowed). Two changes:
- ``server.py:create_workstream`` forwards ``user_id=uid`` — the same
``uid`` already resolved from the auth result (with trusted-service
forwarding preserved).
- ``server.py:open_workstream`` prefers the persisted owner on the
workstream row over the rehydrating caller so reloading someone
else's workstream doesn't silently re-parent it. Falls back to
the authenticated caller when the stored row has no owner
recorded (pre-phase-1 rows).
Regression test in ``tests/test_workstream.py`` pins
``WorkstreamManager.create(user_id=X)`` → ``ws.user_id == X`` so the
manager seam can't regress silently on a future refactor.
**Malformed-JSON recovery allowlist expanded for coordinator args.**
``_prepare_tool()`` has a two-stage salvage path for models that
emit malformed JSON: a regex-extract (fallback 1) and a bare-string
→ primary_key wrap (fallback 2). The fallback-1 key list didn't
include coordinator argument names, so a slightly malformed
``spawn_workstream`` / ``send_to_workstream`` / etc. call would
hard-fail instead of salvaging into a minimal-args dict for retry.
Added ``ws_id`` / ``message`` / ``initial_message`` / ``parent_ws_id``
to the allowlist (kept alphabetised) so the coordinator tools get
the same model-self-correction behaviour as the interactive tools.
Fallback 2 already covers the ``ws_id``-primary-key tools via
``PRIMARY_KEY_MAP``; the regex path matters when the model emits
``{"ws_id": "abc", "message": "..."}`` with a trailing syntax error.
Verification: ``ruff check`` clean, ``mypy turnstone`` clean
(156 source files), ``pytest -m "not live"`` → 4065 passed, 3
deselected (live-backend), 0 failed.
* fix(coordinator): address ultrareview findings on coordinator workstream kind
Security
- Cross-tenant leak: CoordinatorClient.inspect/list_children now constrain
to the coordinator's own ws_id + direct children; an LLM coerced via
prompt injection can no longer exfiltrate other tenants' workstreams.
- Empty-owner short-circuit bypass: strict equality at coordinator.py
ownership gate and at the storage-fallback branch in coordinator_history;
orphan/system-owned coordinator rows can no longer be rehydrated by
arbitrary holders of admin.coordinator (DoS + history disclosure vector).
- Closed coordinators no longer silently resurrect on subsequent GET —
the Close button is now actually durable across URL revisits and tab
refreshes; rows with state in {closed, deleted} refuse rehydration.
Correctness
- ChatSession.close() now releases the CoordinatorClient httpx.Client
pool; previously every closed/evicted coordinator dropped a connection
pool on the floor until non-deterministic GC.
- open_workstream rehydration now forwards parent_ws_id + kind, so
coordinator-spawned children survive node restart / idle eviction
with their parent link intact instead of becoming silent orphans.
- list_children truncated flag now signals whenever the SQL fetch hit
the page cap (previously permanently False in the no-filter case,
causing confident-but-incomplete summaries from the coordinator).
- ConsoleCoordinatorUI.approve_tools: per-tool auto-approve now checks
auto_approve_tools independently of the blanket auto_approve flag,
so 'Always approve this tool' actually works on the next invocation.
Concurrency
- _spawn_worker no longer falls through to start a second concurrent
worker thread on the same ChatSession when queue.Full fires; instead
send() returns False and the endpoint surfaces HTTP 429.
- _open_locks entries are now refcounted under self._lock and only
popped when the last waiter releases — eliminates the race where a
rehydration-failure path lets two threads serialize on different lock
instances for the same ws_id and trip the "already tracked" guard.
Tests: +6 regression cases covering closed-coordinator refusal,
empty-owner non-admin refusal, queue.Full no-duplicate-worker,
inspect/list_children cross-tenant rejection, and truncated semantics.
All eight suggestions verified against source before applying:
- docs/settings.md — ConfigStore key names are `model.plan_alias` /
`model.task_alias` (not `plan_model` / `task_model`); updated in
both the overview list and the plan/task overrides table.
- docs/security.md — `src` claim values now reflect what actually
gets minted: `password`, `database` (from API-token exchange),
`oidc`, plus service origins `console`, `cli`, `channel`.
- docs/sdk.md — `upload_attachment(ws_id, filename, data, *,
mime_type=...)` matches the real SDK signature; `bytes`-returning
helper is `get_attachment_content` (not `download_attachment`);
code example reordered so it doesn't collide on `filename=` kwarg.
- docs/architecture.md — "prior `plan` tool call" → "prior
`plan_agent` tool call" so wording stays consistent with the
renamed tool.
- docs/tools.md — `plan_agent` `primary_key` is `goal`, not
`prompt`, in both the primary-key table and the summary table
(matches the JSON schema in turnstone/tools/plan_agent.json).
Systematic pass over every doc under docs/, the root-level README /
QUICKSTART / CONTRIBUTING, and the PlantUML diagrams. Memory and docs
had drifted against the code since 1.2 — this catches them up to the
1.4.0 release and the 1.5.0a1 experimental line.
User-facing fixes
- README: fix broken docs/mcp.md link (→ mcp-registry.md); channel
gateway entry reflects shipped Discord + Slack adapters instead of
"Slack/Teams planned"; diagrams table mentions both.
- QUICKSTART: docs/*.md relative links were wrong from the repo root;
wizard version bumped from 0.5.4.
- CONTRIBUTING: add dev extra plus the ruff / mypy / pytest commands
we actually expect before push.
Reference docs
- architecture.md: 19 tool schemas (was 15), 18 admin tabs (was 14),
turnstone-bootstrap added to entry-points table, OpenAI provider
file split (chat/responses/common) documented, 38 SDK event
dataclasses (was 27 and referenced deleted mq/protocol.py), Slack
adapter + multi-adapter gateway, plan_agent/task_agent naming,
governance admin-panel rewrite.
- api-reference.md: full attachment endpoints (POST/GET/content/
DELETE on /v1/api/workstreams/{ws_id}/attachments) plus the
multipart mode on POST /v1/api/workstreams/new.
- channels.md: Slack Setup section (Socket Mode app creation, OAuth
scopes, tokens), Slack CLI/env reference in config table, combined-
adapter architecture diagram.
- console.md: 18-tab listing (was 13) with Channels/Models/Nodes/TLS
descriptions and ConfigStore live-edit note.
- docker.md: Slack env vars block; image entry-point list now
includes turnstone / turnstone-bootstrap.
- sdk.md: attachments methods on the server client, attachments
example (upload-then-send and at-creation), event count fixed.
- releasing.md: four-track table (stable/1.0, 1.3, 1.4 + main 1.5);
promotion workflow uses 1.5 / 1.6 numbering.
- settings.md: plan_model / task_model / plan_effort / task_effort
overrides section.
- governance.md: skill naming (/skill, `skill` field — not /template),
Prompts/Judge tabs called out.
- security.md: two-token-types wording; src claim values match the
AuthResult source strings actually emitted.
- mcp-registry.md: SDK package name is @turnstone/sdk.
- tools.md: plan / task renamed to plan_agent / task_agent in the
section headings and summary table; primary-key table matched.
- design/consistent-hash-ring.md: dead direct-http-transport.md
pointer redirected to architecture.md.
Diagrams
- 02-package-structure: drop phantom chat.py entry point, add admin
and bootstrap, add slack/bot.py, rename channels/gateway.py →
channels/cli.py.
- 16-channel-architecture: Slack is no longer "(future)", add a
SlackBot class and the slack-bolt Socket Mode edges; wire the new
bot into ChannelService. PNGs regenerated from both puml sources.
Audit pass against the actual commit messages between v1.3.0 and
v1.4.0 turned up several substantive items the initial CHANGELOG
under-described or omitted entirely. Fix-forward expansion plus a
Contributors section recognizing external contributors.
Added detail / coverage:
- New "Server compatibility layer for local model servers" entry —
the vLLM / llama.cpp profiles + admin UI fields shipped in #352
alongside the capabilities passthrough; previously buried under one
bullet.
- Per-call plan/task model selection split into three sub-bullets
(backend split, runtime configurability without restart via
ConfigStore admin tab, per-call override) — three PRs that build on
each other deserve to be discoverable independently.
- Opus 4.7 entry expanded with 1M ctx / 128K output, the new
thinking_display capability field, xhigh effort level, and admin
dropdown updates.
- Dashboard composer note: tab-bar `+` modal also gained the paperclip
+ chip strip + first-message field.
- Slack adapter: explicit "session recovery via persisted recoverable
route keys" — ops-relevant promise for restart behaviour.
- pgbouncer swap: helm chart link + ports updates noted.
- Provider capabilities entry: defensive shallow-copy + chat_template
deep-merge follow-ups.
New Fixed entries:
- Cross-user attachment-fetch hardening (get_attachment_content
scopes by user_id).
- Attachment-list DoS guard on /v1/api/send.
- Bounded LRU for upload locks.
- 3.12 CI deadlock root-cause writeup (asyncio.Lock vs Starlette
TestClient loop teardown).
New SDK entry:
- PlanResolvedEvent type + guard, dispatched cross-client when one
client resolves a plan so others dismiss in sync.
New Operational subsection:
- vendor-js workflow now auto-downloads hls.js for future Renovate
bumps so they're merge-ready without manual file fetches.
Contributors:
- Recognise @daoxley (Slack adapter, #355) and @pizzaandcheese
(pgbouncer swap, #353) — the two external contributors with
meaningful net-new work in this release — plus the Renovate bot.
- Pointer to channel-attachment ingest as the headline 1.4.1 feature
for would-be contributors.
Repo previously had no CHANGELOG. Establishes the file with full
1.4.0 coverage (attachments end-to-end, dashboard composer refactor,
Slack adapter, per-call plan/task model, provider capability
passthrough, Opus 4.7) plus a one-line 1.3.1 entry for the Opus 4.7
backport. Format follows Keep a Changelog 1.1.0; release-track
guidance up top covers the three stable branches + main.
Operator-relevant call-out at the top of [1.4.0]: migrations 037 +
038 must be applied before starting 1.4.0 against an existing 1.3.x
database. Both are additive and idempotent.
* feat(ui): dashboard composer polish from PR #362 designer review
Three deferred items from the prior designer pass on the unified
dashboard composer. Pure UX affordances; no server change.
- **Persist Options open/closed in localStorage.** Power users who
routinely set non-default model/skill don't have to click "Options"
on every page load. Key: `turnstone.dashboard.options_open`.
Defaults closed for first-time users. Falls back gracefully when
localStorage is unavailable (private mode, quota).
- **Active-options summary chip.** Renders the non-default model /
judge / skill values inline next to the Options button (mono, dim,
separated by middots). Hidden via `[hidden]` when everything is at
default — no chrome cost in the common case. Updates on any select
change via a single delegated handler on the panel. Hidden on
narrow viewports (the action row stacks vertically there and the
chip would push the layout further).
- **"Drop to attach" overlay during drag.** CSS pseudo-element on
`.dashboard-composer-drop` overlays a centered "Drop to attach"
label so dragging a file makes the action explicit instead of just
showing the dashed-border highlight. pointer-events: none keeps
the underlying composer controls reachable; visual only.
* fix(ui): address Copilot review on dashboard composer polish
- _restoreDashboardOptionsState() forced the panel closed every time
showDashboard() ran when localStorage was unavailable (private mode,
storage quota), contradicting the comment that promised a per-session
fallback. Add a module-scoped _dashOptionsOpenSession variable
updated by _setDashboardOptionsOpen / _toggleDashboardOptions, and
only override the visible state from localStorage when the read
genuinely succeeded. The session value now preserves the user's
choice across hide/show cycles in environments where localStorage
throws.
- Fold the duplicated `.dashboard-composer { position: relative; }`
block into the existing rule above. The position context is needed
for the .dashboard-composer-drop::before overlay; the comment now
says so.
PR #355 added the Slack adapter on the server but missed the console
admin surfaces that talk to channel_type. Three concrete gaps + a
designer-review polish pass.
Functional bug + UI parity:
- _collectNotifyTargets() in admin.js hardcoded `channel_type: "discord"`
— even on a Slack-only deployment the skill notify-on-complete form
always wrote Discord targets, sending notifications to the wrong
adapter (or nowhere). Add a per-row channel-type <select> driven by
a small _NOTIFY_CHANNEL_TYPES table that's the one place to register
a new platform; collector and populator both read from the dropdown.
ID-input placeholder updates dynamically when the platform changes.
- The "Link Channel Account" modal only offered Discord — users
couldn't link a Slack account through the UI at all. Add a Slack
<option> and reuse the same dynamic-placeholder helper. Drop the
static Discord-shaped HTML placeholder so the JS-driven hint doesn't
flash a Discord example before the dropdown initializes.
- Skill create/edit modals only showed Discord in the notify-on-complete
placeholder example. Show both adapters.
- Per-platform .scope-discord / .scope-slack badge classes so the
linked-accounts list distinguishes platforms visually instead of all
rendering as the generic .scope-channel magenta. Falls back to
.scope-channel for any future channel_type the stylesheet doesn't
yet know about.
Designer review polish:
- Theme-aware --discord / --slack / --discord-glow / --slack-glow
tokens in base.css. The first pass shipped raw hex (#818cf8 /
#f472b6) that fails WCAG AA on light theme (1.8:1 and 2.4:1); the
light variants (#4f46e5 indigo, #be185d rose) pass. Badge classes
now reference tokens, matching every other .scope-* rule.
- Notify-row mobile layout: three controls in a row left ~80px for
the ID input at 360px viewport, truncating snowflakes. Tighten
platform select to 76px (labels are short), add flex-wrap, and at
≤700px drop the ID input to its own row so it gets full width.
- Per-platform classes apply alone (not co-classed with scope-channel)
so winning the cascade doesn't depend on stylesheet source order.
- Replace "Discord snowflake" jargon with "Discord ID"; give Slack
ids concrete examples (C01234567 / U01234567) instead of an
ambiguous "C0…".
Combines the substantive bot.py fixes flagged in both review trails on
PR #355. Discord parity items grouped here too since they're the same
surface (slack/bot.py).
From Copilot:
- _notify_reply_routes was read on StreamEndEvent but never popped on
the success path. Result: one notification reply pinned every later
response for that ws_id to the notification thread until the bot
restarted. Pop after read; combine the surrounding ifs (SIM102).
- PlanReviewEvent embedded raw event.content inside a triple-backtick
mrkdwn fence without escaping. A plan with ``` (very common — plans
often quote code) would break the fence and let later content render
as live markup, including unintended Slack mentions/links. Rewrite
_sanitize_slack_preview to splice a zero-width space inside any ```
sequence (Slack stops recognizing it as a delimiter) instead of
escaping every single backtick — keeps single-backtick code snippets
readable while still protecting the fence. Apply to plan-review.
- _send_approval_request joined unbounded tool_lines into one mrkdwn
section, but Slack section.text caps at 3000 chars. Multi-tool
batches with large previews silently failed chat_postMessage,
leaving the user unable to approve/deny. Cap each preview to 600
chars under a 2700-char total budget; append "+N more" when truncated.
From eous (parity with Discord):
- Pass `client_type="chat"` from both `get_or_create_workstream` call
sites (slash-command session + DM). Without it Slack-routed
workstreams loaded the web-default prompt; the chat-specific
system prompt now applies as it does for Discord.
- Add `exc_info=True` to the eleven `log.debug(...)` exception handlers
so underlying tracebacks are available when debug logging is on
instead of being silently dropped. Level stays debug — these are
benign-by-default sites (chat_update on a deleted message, etc.) so
only the visibility changes. Typed-exception handlers
(RemoteProtocolError, etc.) keep their bare debug log.
- Module docstring on slack/__init__.py so pydoc / import errors have
human-readable context.
Tests: rewrite the sanitizer test to match the new (more permissive)
single-backtick behaviour; add coverage for the triple-backtick
neutralization + short-input passthrough; patch httpx.AsyncClient at
all five TurnstoneSlackBot construction sites so each test doesn't
leak an unclosed real client.
- cli.py: ChannelAdapter import is annotation-only; move into
TYPE_CHECKING block and switch the two cast() calls to string-form
so the runtime import isn't required (TC001).
- slack/{config,routes}.py: ruff format fixes (whitespace + drop
redundant string-form annotation now that __future__ annotations
is in effect).
- pyproject.toml: drop the unused `tests.*` mypy override — `mypy
turnstone` (the only invocation in CI + local) never matches it,
so it was pure noise in the "unused section(s)" report. Other
optional-dep overrides stay; they're real safety nets when running
mypy without the [all] extras (e.g. on the test job).
- uv.lock: regenerate to match the slack-bolt + transitive deps the
pyproject changes resolve to (lock-check was failing on stale hash).
* fix(ui): rehydrate chip strip after queued-message dequeue not_found
The dequeue handler only refreshed the per-pane chip strip when the
DELETE returned status="removed". On status="not_found" (the queued
message already dispatched), chips stayed stale: any reservations that
raced the dispatch could leave the UI showing a different pending set
than the server actually had.
Re-fetch on both paths so the chip strip always reflects the
authoritative server state. The queued-message bubble itself stays
visible on not_found, same as before — the promote loop strips the
queued styling on idle.
* feat: sweep orphan attachment reservations periodically
Process crashes between reserve_attachments and consume/unreserve can
leave attachment rows soft-locked forever (reserved_for_msg_id NOT NULL
with no consumer ever coming back). The worker-thread exception path
in /v1/api/send already handles in-process failures, but a hard kill
or oom mid-send escapes that.
Add sweep_orphan_reservations(older_than_seconds) to the storage
protocol — clears reserved_for_msg_id on rows where message_id IS NULL
and created < now() - threshold. Implemented for SQLite + PostgreSQL
using the same string-comparison form (created is ISO-8601 text in
both backends, lexicographic order matches chronological).
Wire into the server lifespan: run once at startup (catches anything
left over from the previous process), then every 30 minutes as
defense-in-depth. Threshold is 4 hours so we don't race a long-running
dispatch and unreserve rows the worker is still about to consume.
Tests cover sweep semantics: clears old reserved rows, leaves fresh
ones alone, skips already-consumed rows, no-ops on zero/negative
threshold.
* fix: track reserved_at for orphan-reservation sweep
Copilot review on PR #363 flagged a real correctness bug: the sweep
used the attachment row's `created` timestamp (upload time) as the
staleness signal. An attachment uploaded hours ago but reserved fresh
could be unreserved mid-send, after which mark_attachments_consumed
silently drops the row because reserved_for_msg_id no longer matches
the send_id.
Add a dedicated `reserved_at` column (migration 038) set on
reserve_attachments and cleared on mark_attachments_consumed /
unreserve_attachments. The sweep now scopes by `reserved_at < cutoff`,
so reservation age is what's measured, not upload age. Backed by a
partial index `(reserved_at) WHERE reserved_at IS NOT NULL` so the
periodic scan stays cheap as the consumed-history grows.
Threshold dropped from 4h to 1h since it now means "longest realistic
single send" rather than "longest plausible time between upload and
send" — a tighter, more defensible bound.
Tests cover the regression (uploaded long ago + reserved fresh must
not be swept), plus reserved_at clearing on both consume and unreserve.
* feat: workstream attachments at creation time + SDK + UI parity
Closes the two big deferred items from PR #356: attaching files as part
of the initial workstream-creation request, and full SDK coverage of the
attachment surface.
Server: POST /v1/api/workstreams/new now accepts multipart/form-data
(meta JSON + 0..N file parts). Files are validated and saved as pending
under the new ws; when initial_message is also set the create handler
reserves them onto that turn before the dispatch worker fires, mirroring
the /v1/api/send pattern. Validation failure rolls back the workstream
via delete_workstream so we don't leak orphan rows or emit a phantom
ws_created/ws_closed pair on SSE. JSON path is unchanged.
Console routing: route_create accepts multipart with ?ws_id=<hex> as a
query parameter (the console hashes the id before the body lands).
Added /v1/api/route/workstreams/{ws_id}/attachments POST/GET/DELETE +
.../{attachment_id}/content GET proxies that forward raw bytes and
preserve upstream headers (Content-Disposition, X-Content-Type-Options,
CSP sandbox).
Python + TypeScript SDKs: AttachmentUpload type, upload_attachment,
list_attachments, get_attachment_content, delete_attachment, and
send(attachment_ids=...). create_workstream(attachments=...) sends
multipart and pre-generates a ws_id client-side so cluster routing
works. SDKs reject attachments+target_node combinations since the
multipart route doesn't honor target_node.
Web UI: dashboard composer refactored to a single unified create flow.
Replaced the inconsistent split (Enter created+sent raw, "New Chat"
opened a modal) with one rich composer carrying a textarea, paperclip
+ chip strip, drag-drop, paste-image, and a collapsible Options panel
for model/judge_model/skill. Submit button dynamically labels Create
vs Send. New-workstream modal also gained the same paperclip + chip
strip + first-message field for the tab-bar + entry point.
Tests: 30 new tests across server multipart create, console route
multipart + attachment proxies, Python + TS SDK attachment surfaces,
plus regressions for the three review-flagged bugs (Content-Type
boundary preservation, attachments+target_node rejection, no phantom
ws_created on validation failure).
* fix: address Copilot review feedback on PR #362
- web_helpers: docstring now matches behaviour — read_multipart_create_or_400
does enforce the optional max_per_file_bytes cap as defense-in-depth.
- app.js: drop the duplicated _formatAttachSize definition (one already
exists earlier for pane chips); add a shared _isAttachmentAllowed helper
that mirrors the server's classifier (png/jpeg/gif/webp images, text/*
MIMEs, allowlisted application/* MIMEs, known text extensions) and call
it from both _newWsAddFiles and _addDashboardFiles so unsupported files
fail fast client-side instead of after a server roundtrip.
- app.js: dashboardSubmit catch now suppresses the redundant error toast
on authFetch's "auth" Error and falls back to a generic message when
err.message is undefined, instead of rendering "Connection error: undefined".
- SendResponse (Pydantic + TS): document and expose attached_ids,
dropped_attachment_ids, priority, and msg_id so attachment-aware SDK
callers can detect partial reservations and dequeue queued messages.
- test_server_attachments_on_create: drop the dual `import turnstone.server`
+ `from turnstone.server import` style — use monkeypatch.setattr by
dotted path for module-level mutation and `from … import …` for the
helpers, keeping a single import style.
* feat: per-call model selection on plan_agent / task_agent
The calling LLM can now pass `model="<alias>"` to plan_agent or
task_agent to override the operator-configured per-kind model for
that one invocation. Useful when subtask difficulty varies within a
session: the model can downgrade to a cheap alias for trivial work
and reach for a stronger one when the problem is hard.
Tool descriptions list the live registered aliases (refreshed when
the operator hits "sync to nodes" / internal_model_reload), so the
calling LLM always sees the current options. Bad aliases return a
corrective error dict with the available choices so the LLM retries
cleanly rather than failing silently.
No whitelist — any alias the registry knows is acceptable; cost
control is intentionally ceded to the model. No per-call effort
override (out of scope; effort stays operator-configured).
Resolution precedence in _run_agent: explicit per-call agent_alias
override > registry per-kind (plan_model/task_model) > legacy
agent_model > session model. The plan retry path (when
_validate_plan fails) reuses the same alias so coaching reflects
real model behaviour rather than a different model masking the
signal.
Implementation:
- plan_agent.json / task_agent.json: optional `model` parameter.
- ChatSession._validate_agent_model_override extracts and validates
the arg; mirrors the existing empty-prompt error pattern.
- _prepare_plan / _prepare_task stash the override in
item["model_override"]; _exec_* pass it through.
- _run_agent gains agent_alias kwarg with defence-in-depth
ValueError on unknown alias.
- _render_agent_tool_descriptions deep-copies plan/task entries
before mutating description so the module-level TOOLS constant
stays untouched across sessions; rebuilds the BM25 tool-search
index when active so its text matches what the LLM sees.
- server._broadcast_agent_tool_schema_refresh walks active
workstreams on internal_model_reload so descriptions update
without restart.
* fix: clarify no-registry placeholder + avoid double BM25 rebuild
Addresses Copilot feedback on PR #361.
1. plan_agent.json / task_agent.json placeholder said the parameter
falls back to the "operator-configured plan/task model". That
text is what no-registry sessions see (registry-bearing sessions
get the templated description with the live alias list); for
those single-model sessions, omitting the param falls back to
the current session model, not an operator-configured one.
Reword so the no-registry user gets accurate guidance.
2. _on_mcp_tools_changed already calls _rebuild_tool_search after
merging MCP tools. _render_agent_tool_descriptions also
rebuilt the BM25 index when active, so the MCP refresh path
was rebuilding twice per refresh. Move the BM25 rebuild out
of the private render helper into the public
refresh_agent_tool_schemas wrapper — _on_mcp_tools_changed
keeps calling the render helper directly (no double rebuild),
and registry-reload callers go through the wrapper which
still keeps the index in sync.
* feat: ConfigStore + admin UI for plan/task agent model and effort
Per-kind sub-agent routing was added in #359 but only via config.toml.
Operators can now switch the plan_agent / task_agent model and reasoning
effort at runtime from the admin Model tab without restarting.
Adds four ConfigStore-backed settings:
model.plan_alias — alias for plan_agent
model.task_alias — alias for task_agent
model.plan_effort — reasoning effort for plan_agent
model.task_effort — reasoning effort for task_agent
Server startup and internal_model_reload both apply these as overrides
on top of the registry's config.toml-loaded values; the new logic
computes "effective" values for all five model-routing fields and only
calls registry.reload() when at least one differs.
Admin UI: extracts ALIAS_SETTING_KEYS to a const used by both the
dynamic-alias-choice injection and the empty-option label rendering.
Adds INHERIT_EMPTY_LABEL_KEYS so plan_effort / task_effort show
"(inherit)" for empty — distinct from the literal "none" choice (which
actually disables reasoning, very different from leaving unset).
Also fixes Copilot review feedback from #359:
- _validate_effort treats empty / whitespace as unset rather than
warning on benign explicit-empty configs (with .strip().lower()
normalisation; "HIGH" and " low " now parse correctly)
- turnstone.example.toml's reasoning_effort comment lists the full
set of accepted values (none, minimal, low, medium, high, xhigh, max)
* fix: apply routing overrides on config-reload + skip no-op model-reload
Addresses Copilot feedback on PR #360.
1. Admin settings updates fan out via /_internal/config-reload, which
only reloaded the ConfigStore — plan/task routing changes weren't
visible until a model-reload or restart, defeating the runtime
configurability this PR is meant to add.
2. /_internal/model-reload always called registry.reload(), churning
cached clients even when nothing changed. Risky when fanned out
across nodes (could close in-flight clients).
Extracts two helpers in server.py:
- _effective_routing(cs, ...) pure function: overlay CS values on base
- _apply_routing_overrides(reg, cs) reload only when something differs
Used by the startup path, config_reload (new), and model_reload (now
short-circuits with a noop response when models + routing are unchanged).
plan_agent and task_agent previously shared a single agent_model knob and
plan_agent hardcoded reasoning_effort="high" in three call sites. They
have different cost/latency profiles — plan is rare and benefits from a
stronger model, task is frequent and benefits from a cheaper one — so
sharing the knob undertunes both.
ModelRegistry gains plan_model, task_model, plan_effort, task_effort.
Per-kind overrides win over the legacy agent_model, which still works
as the single-knob fallback for both. resolve_agent_alias(kind) and
resolve_agent_effort(kind) centralise the resolution; PLAN_DEFAULT_EFFORT
captures the back-compat "high" default in one place rather than at
every call site.
session._run_agent delegates resolution by label ("plan" vs "task").
The three hardcoded reasoning_effort="high" arguments are removed —
behaviour is identical when no plan_effort is configured.
Loader validates effort against {none,minimal,low,medium,high,xhigh,max}
and warns + drops typos rather than passing them to the provider.
ConfigStore parity and admin UI for the new knobs are deferred to a
follow-up — config.toml-only is enough for the backend split.
Previously, resolving a plan on one client (e.g. phone) cleared the
server's pending state and unblocked the worker, but emitted no event
to other connected clients. Their plan-approval modal stayed stuck.
resolve_plan() now enqueues a plan_resolved frame (mirroring the
approval_resolved pattern in resolve_approval) before clearing
_pending_plan_review, so a reconnecting client cannot receive both
the replayed plan_review and the live plan_resolved. Skips the frame
on the cancel-with-no-plan path.
Client adds a plan_resolved handler that dismisses the modal without
re-firing /v1/api/plan, restores keyboard context (skipped on touch
to avoid soft-keyboard pop on mobile), labels the inline plan summary
"(synced)" so remote dismissal is unambiguous, announces via the
existing aria-live #toast for screen-reader parity, and falls back
to an info message if plan_resolved races ahead of plan_review.
Adds PlanResolvedEvent to the Python and TypeScript SDKs with
deserialization and type-guard tests.
- Add claude-opus-4-7 capability entry (1M ctx, 128K output, adaptive
thinking, supports_temperature=False, thinking_display=summarized)
- Suppress temperature param for Opus 4.7 (API returns 400)
- Add thinking display opt-in via new ModelCapabilities.thinking_display
field - Opus 4.7 omits thinking by default, always send summarized
- Add xhigh effort level to mapping and Opus 4.7 effort_levels
- Add xhigh/max options to skill template dropdowns in admin console
- Align reasoning effort label capitalization across all console dropdowns
- Update example config to reference claude-opus-4-7
- 10 new tests with regression guards for Opus 4.6 backward compat
Verified against live API: streaming and completion calls succeed.
Trivy flags two HIGH CVEs in jq/libjq1 1.7.1-6+deb13u1 with no fixed
version yet from Debian:
- CVE-2026-39979: out-of-bounds read in jv_parse_sized() on non-NUL-
terminated buffers
- CVE-2026-40164: DoS via crafted JSON causing hash collisions
jq is invoked only on trusted CLI/admin paths against
process-controlled JSON input in turnstone — never on untrusted
network bytes — so the NUL-terminated invariant holds and the DoS
vector is not reachable.
Will revisit when Debian publishes a patched libjq1.
* feat: workstream attachments (images + text documents)
Adds end-to-end support for attaching images (png/jpeg/gif/webp) and
plain-text documents (markdown, source, JSON, etc.) to a workstream's
next user turn via the web UI.
Storage: new workstream_attachments table (migration 037) with a
three-state lifecycle — pending → reserved → consumed — scoped by
(ws_id, user_id) and linked to conversations.id on consume. Rewind/
truncation cascades attachment rows; delete_workstream does too.
Session: ChatSession.send(attachments, send_id) builds multipart user
content (text + image_url + document parts) and persists text-only to
conversations with attachments joined on load via message_id. Queue
path carries ordered attachment_ids plus a reservation token so
queued multimodal turns can't lose files to overlapping sends.
Providers: internal document content parts translate at the API
boundary — Anthropic emits native document blocks (text/plain
coerced, original MIME folded into title); OpenAI Chat Completions
and the Google OpenAI-compat endpoint inline them as escaped
<document> text blocks (XML-attr escape + </document> neutralization);
Responses API emits input_text with the same wrapper.
Server: POST/GET/DELETE /v1/api/workstreams/{ws_id}/attachments with
multipart upload (magic-byte image sniffing, UTF-8 enforcement for
text, per-kind size caps, Content-Length pre-check, per-(ws,user)
pending cap + TOCTOU lock). /v1/api/send reserves before dispatch
using a full-UUID token, threads it into session.send / queue_message,
releases on worker-thread failure, and reports attached/dropped ids
so the UI can reflect partial reservations. GET /content sets
X-Content-Type-Options, CSP sandbox, inline Content-Disposition, and
forces text/plain for text kinds. Ownership failures mask as 404.
UI: paperclip button, hidden file input with accept allowlist, chip
strip above textarea, drag/drop + paste-image handlers. Chips
rehydrate on ws switch and on queued-message dequeue; send clears
only attached ids and shows a toast when some dropped. Historical
user messages render filename pills via a _attachments_meta sibling
populated on both live-send and reconstruct paths.
530 tests covering CRUD, reservation lifecycle, races (TOCTOU cap,
reserve-then-dispatch overlap), provider translation, XSS headers,
cascade delete, history round-trip, and service-scoped actor flow.
* fix(attachments): address PR review feedback
- get_attachment_content now scopes the row by user_id too, so an
unowned workstream can't be a vector for cross-user blob fetches
via attachment_id guessing (Copilot, server.py:2676)
- send_message rejects attachment_ids lists longer than the pending
cap with 400 — prevents hostile clients from blowing up the
storage IN (...) clause (Copilot, server.py:1515)
- _attachment_upload_locks switched to a bounded LRU OrderedDict;
evicts the oldest unlocked entries past the soft cap so the map
can't grow unboundedly on long-running nodes (Copilot, server.py:2417)
- Pane.dragleave handler uses relatedTarget instead of target so the
drop-zone styling clears correctly when the cursor moves through
child elements; dragend listener added as a fallback for cancelled
drags (Copilot, app.js:297)
- uploadAttachment always cleans up the placeholder chip on failure,
including auth errors — no more stuck "uploading..." chips after
re-auth (Copilot, app.js:427)
- New _swapPlaceholderChip / _removeAttachmentChip helpers preserve
user-selection order through the placeholder→real-id swap; the
pendingAttachments Map is rebuilt in place rather than naïvely
delete+set, which would have moved the entry to iteration end
(Copilot, app.js:420)
- Drop unused `var self = this;` in removeAttachment (github-code-quality)
- Two regression tests: cross-user fetch on an unowned workstream,
and oversized attachment_ids list rejection
* fix(attachments): switch upload-lock to threading.Lock to avoid 3.12 CI hang
The per-(ws, user) upload lock was a module-cached asyncio.Lock.
Starlette's TestClient runs each request on a fresh anyio task /
event loop, so the cached lock's internal _waiters bind to the first
loop that acquired it. When a later request runs in a different
loop, await lock.acquire() blocks on a Future from a closed loop —
silent deadlock.
This surfaced as test (3.12) hanging indefinitely in CI on one push
while the same suite passed on 3.11/3.13 and on the next push. Same
root cause is reproducible against any Starlette TestClient harness
on 3.10+; 3.12 just happens to surface it more often given changes
in how anyio + asyncio.Future interact across loop teardown.
Switched to threading.Lock — loop-agnostic, and the critical section
is one COUNT + one INSERT, short enough that briefly blocking the
event loop is fine. Updated the LRU-eviction probe accordingly
(threading.Lock has no public .locked(), so use a non-blocking
acquire+release as the "is it free?" probe).
TOCTOU pending-cap test still passes; full attachment suite passes
on both 3.12 and 3.13.
* replace bitnami pgbouncer wit edoburu
replaced bitnami pgbouncer with edoburu pgbouncer container and updated environment variables to fit
* updated ports & Kubernetes
Updated ports to fit existing documentation. Also updated the Kubernetes Helm Chart link to use the same container.
* chore(deps): update dependency hls.js to v1.6.16
* chore: download vendored hls.js files + add hls to workflow detection loop
The wheel-completeness check failed on the Renovate bump because
vendor-js.yml only iterated katex/hljs/mermaid — so hls.js PRs
never got their files auto-downloaded. Adding hls to the loop so
future Renovate bumps are merge-ready without manual intervention.
Also running the update now to fix this specific PR.
---------
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: Patrick Buckley <buckleypm@gmail.com>
* feat: pass resolved capabilities through to providers, add server compat layer
The LLMProvider protocol previously forced providers to re-derive
capabilities from static lookup tables, ignoring config overrides set
via the admin UI or config.toml (e.g. thinking_mode, token_param).
This adds an optional capabilities parameter to create_streaming and
create_completion so the session can pass its config-merged
ModelCapabilities through to providers.
On top of this, adds a server compatibility layer for local model
servers (vLLM, llama.cpp). Profiles suggest thinking mode and server
workarounds (skip_special_tokens for vLLM, reasoning_format for
llama.cpp) during model detection, with structured admin UI fields
for server type, thinking mode, and extra body params.
Verified against real vLLM (Gemma 4 31B) and llama.cpp (Gemma 4 E4B)
servers.
* fix: defensive copy in _finalize_extra_body, expose thinking_param in UI
Shallow-copy extra_params and its chat_template_kwargs in the provider
before _apply_thinking_mode mutates them, so callers that reuse the
same dict across models are safe.
Replace the hidden thinking_param input with a visible text field
that appears when thinking mode is enabled. Shows the default
"enable_thinking" and hints that Granite/DeepSeek use "thinking".
* fix: address Copilot review feedback on admin UI and server compat
- Preserve unrepresentable thinking_mode values (e.g. "adaptive") in
raw capabilities JSON instead of silently dropping on edit round-trip
- Validate capabilities and extra body JSON are plain objects, not
arrays or primitives
- Deep-merge chat_template_kwargs from extra_body instead of silently
dropping, so operators can extend/override template kwargs
* fix: hide server compat section for non-local providers
The Server Compatibility fields (server type, thinking mode, extra
body) only apply to openai-compatible (local model servers). Hide
the entire section when the provider is openai, anthropic, or google.
* fix: normalize capsObj to plain object on edit load
Defend against DB rows where capabilities is a JSON literal null,
an array, or a primitive — previous code would crash on the
capsObj.server_compat / capsObj.thinking_mode reads. Same defensive
check also applied to the server_compat nested value.
* refactor: extract _isPlainObject helper for JSON type checks
Consolidates the null/array/typeof check that was inlined at three
different call sites into a single helper. Keeps the intent obvious
at each use site and avoids the awkward multi-condition ternary.
* feat: per-model sampling parameters (temperature, max_tokens, reasoning_effort)
Model sampling parameters were global-only settings applied uniformly to
all models. Different models have fundamentally different requirements
(o-series needs no temperature, Anthropic needs temp=1.0 with thinking,
local models may need different max_tokens). This adds per-model overrides
with global fallback so each model definition can specify its own defaults.
Migration 036 adds nullable temperature, max_tokens, reasoning_effort
columns to model_definitions. NULL inherits the global default from
ConfigStore. The session factory and /model switch command both resolve
per-model override → global fallback consistently.
The admin UI model create/edit modal now has dedicated form fields for
these parameters with client-side validation, a visual section divider,
and per-model override hints in the model table rows.
Removes vestigial model.name and model.context_window global settings
(now handled per-model by the model registry) with startup warnings for
existing config.toml users.
* fix: defensive parsing for config.toml per-model sampling params
Wrap temperature/max_tokens conversions in try/except with range
validation. Invalid values log a warning and fall back to None
(inherit global default) instead of aborting registry load.
* fix: use gethostname() instead of getfqdn() for advertise URLs
socket.getfqdn() does a reverse DNS lookup that often returns a
truncated hostname (e.g. "flat" instead of "flat-blck-io"). Use
gethostname() for advertise URLs in both server and console. For TLS
SANs, include both names so certs cover all variations.
* docs: clarify advertise URL comment re Docker/k8s
* fix: standardize database env vars on TURNSTONE_DB_* naming
compose.yaml used DB_BACKEND/DATABASE_URL in .env which got mapped to
TURNSTONE_DB_BACKEND/TURNSTONE_DB_URL inside containers. Running bare-
metal required the TURNSTONE_ prefix, but docs didn't explain this.
Eliminate the indirection — use TURNSTONE_DB_BACKEND and TURNSTONE_DB_URL
everywhere (compose, .env, bare-metal, docs, bootstrap wizard).
* fix: update .env.example to use TURNSTONE_DB_* naming
* fix: universal tool_call/tool_result orphan detection for OpenAI-compat providers
The Anthropic provider had orphan detection for mismatched tool_call ↔
tool_result pairs, but OpenAI-compatible providers (Chat Completions,
Google, Responses API) had none. When an Anthropic model runs behind
an OpenAI-compat API (e.g. Azure) or cancellation creates orphans,
the API rejects the malformed request.
- Rewrite sanitize_messages() with orphan detection: synthesize error
tool results for unmatched tool_calls, drop tool results with no
matching tool_call, fill empty tool_call IDs with positional remap
- Call sanitize_messages() from Responses API _convert_messages()
* fix: address review feedback on orphan detection
- Track answered IDs per-turn (local_answered) instead of scanning
all of out, preventing false matches from reused IDs across turns
- Drop empty-ID tool results that have no remap entry instead of
passing them through with invalid empty tool_call_id
- Increment empty_result_idx for every empty result, not just remapped
- Remove dead result_ids peek-ahead code
- Add test for repeated tool_call IDs across turns
* fix: accurate token usage tracking for compaction across all providers
Anthropic's input_tokens excluded cached tokens, causing massive
under-reporting (e.g. 327 vs 9000 actual) when prompt caching was
active. This prevented auto-compaction from triggering.
- Normalize Anthropic prompt_tokens to total input (input_tokens +
cache_creation + cache_read), matching OpenAI semantics
- Reset _last_usage per API call so tool-chain iterations get fresh
usage instead of max()-merging with stale values
- Add mid-turn compaction check during tool chains to prevent context
overflow before end-of-turn
- Anchor _remaining_token_budget() on provider-reported prompt_tokens
with local estimates only for the delta since last API call
- Improve _msg_char_count() to include structural overhead (role,
tool_call_id, tool call IDs) and handle image tokens in calibration
- Emit status after every API call, not just end of turn
* fix: defensive null coercion and index clamping from review feedback
- Add `or 0` to all getattr calls for input_tokens/output_tokens in
Anthropic provider (streaming + non-streaming) to handle SDK nulls
- Use getattr for non-streaming input_tokens/output_tokens instead of
direct attribute access for consistency
- Clamp _calibrated_msg_count with min() in _remaining_token_budget()
to prevent stale state from over-slicing after compaction
The ws-check-hint animation clobbered the fadein's forwards fill,
making the checkbox invisible for 0.6s on card-body click — appearing
as a deselect-then-reselect. Remove the hint, the unused role=checkbox
on the card, and restore the original symmetric toggle behavior.
* fix(ui): improve delete workstream UX and accessibility
Card body click no longer deselects (prevents confusing red border loss);
checkbox pulse hint guides users to deselect affordance. Adds keyboard
navigation, aria-labels, hover feedback, animations, and neutral Close
button styling after deletion.
* fix(ui): remove duplicate a11y checkbox from delete-mode cards
Hide the visual checkbox from the a11y tree and tab order so the card
(role=checkbox) is the sole keyboard/screen-reader target. Addresses
Copilot review feedback about nested interactive elements.
* perf: reduce initial rebalance from ~1.5s to ~50ms on PostgreSQL
Increase seed_ring_buckets chunk sizes (PG 500→16k, SQLite 500→8k) to
cut network round-trips from 131 to 5. Add ConsoleRouter.populate_from_assignments()
to build the routing cache directly from computed assignments, eliminating the
65 536-row DB read-back. Router becomes ready in <1ms; DB persistence follows.
* fix: address review — populate after seed write, sync router version
Move router cache population after seed_ring_buckets() so the router
is never "ready" with an unpersisted ring. Pass the new rebalancer
version to populate_from_assignments() so check_version() on the
collector thread does not trigger a redundant 65 536-row refresh.
- Remove role="status" and aria-label during _promoteQueuedMessages
so screen readers don't announce stale "queued" context
- Mark element with pendingDismiss when user dismisses before msg_id
arrives; send deferred DELETE when the send response provides the ID
If the model responds without tool calls, the main loop exits
immediately — no tool-result seam exists for advisory injection.
Queued messages were silently orphaned in the OrderedDict. Now
flushed as regular user messages before emitting idle state.
Bug 1: Extract _promoteQueuedMessages() — removes badge, dismiss
button, queued classes, and data-msgId. Called from setBusy(false)
on state_change: idle.
Bug 2: _dequeueMessage no longer removes the DOM element when server
returns not_found (message already injected). Only removes on
"removed" (actually dequeued). Network errors also preserve the
element. The promote loop handles cleanup on idle instead.
* feat: tool result advisory system with user message queuing
General-purpose advisory injection for tool results — when advisories
are present, tool output is wrapped in <tool_output> tags with
<system-reminder> blocks appended. Two initial producers:
- Output guard advisories: model sees why content was flagged/redacted
- User message interjections: users can queue messages mid-execution
via the web UI, injected at the next tool-call seam
Queued messages use !!! prefix for important priority. Advisory
injection is gated by ModelCapabilities.supports_tool_advisories
(default true for commercial models, false for local/vLLM).
On cancel/error, queued messages are flushed as regular user messages
so nothing is silently lost. Raw tool output (pre-wrap) is persisted
to the DB to keep history clean of ephemeral advisory XML.
* fix: frontend UX for queued messages — rollback, discoverability, a11y
- Send button changes to "Queue" (outline style) during busy state,
visually distinct from filled red Stop button
- Placeholder updates to hint at !!! priority convention
- addQueuedMessage returns element ref for optimistic UI rollback
- Remove queued element on queue_full, busy, or connection error
- Add role="status" and aria-label to queued message elements
- Promote queued messages to normal appearance when generation ends
* feat: queued message removal via dismiss button
Switch backing store from queue.Queue to OrderedDict + Lock for O(1)
removal by ID. Each queued message gets a UUID, returned to the
frontend and stored as data-msg-id on the DOM element.
Dismiss button (x) on queued messages calls DELETE /v1/api/send with
the msg_id. If the message was already injected (race), server returns
not_found and the UI removes the element anyway.
No new endpoint — DELETE method added to the existing /v1/api/send
route. dequeue_message() on ChatSession is O(1) under the lock.
* fix: address PR review — escaping, types, list output, message cap
- Escape </tool_output> and <system-reminder> in tool output to prevent
wrapper tag injection from untrusted tool results
- Change _collect_advisories return type from list[Any] to list[ToolAdvisory]
- Drain queued messages on list/structured output (append as text part)
so they aren't silently stuck until a str result appears
- Cap queued message length at 2000 chars to prevent context bloat
- Remove unused var in _dequeueMessage
* feat: replace workstream action buttons with per-tab dropdown menu
Move refresh-title, edit-title, fork, close, and delete actions from
the header toolbar into a dropdown menu on each workstream tab,
triggered by a ▾ chevron that replaces the × close button.
Dropdown follows the existing pane context menu pattern: keyboard
navigation, mutual exclusion, click-outside/Escape dismiss, toggle
on re-click, aria-expanded + aria-haspopup, and focus restoration.
Delete is visually distinct (red text + wash + red focus ring, 6px
separator). Mobile hides "Refresh title" and sizes the chevron to
36px touch targets.
Removes updateWsActionButtons(), _applyTitleButtonState(), and
_wsTitleState tracking (dead code after button removal).
* fix: remove Ctrl+Shift+R shortcut that overrides browser hard refresh
Refresh title is a low-frequency action accessible from the tab
dropdown; no replacement keybind needed.
* fix: address tab dropdown review findings
- Pass wsId through dropdown actions so they target the correct
workstream even when opened on a non-active tab
- Fix setTimeout race where closeTabDropdown before timeout fires
could leave stale listeners
- Guard Close and Delete on last workstream (dropdown, keyboard
shortcuts, and defense-in-depth in confirmDeleteWorkstream)
- Use aria-disabled instead of disabled so screen reader users can
discover unavailable items via arrow keys
- Enlarge chevron hit target, add hover affordance with subtle
background highlight
- Add 0.1s dropdown open animation (respects prefers-reduced-motion)
- server.py: catch (Exception, GenerationCancelled) instead of
BaseException so KeyboardInterrupt/SystemExit propagate normally
- judge.py: log client close failures instead of bare pass
Gemini's OpenAI-compat endpoint requires thought_signature to survive
the tool-call round-trip. Previously dropped because the Chat Completions
provider cherry-picks only standard fields (id, type, function).
Fix: GoogleProvider now captures raw tool-call dicts (including
thought_signature) via provider_blocks — the same fidelity lane the
Anthropic provider uses for signature round-tripping. On the next turn,
_prepare_messages reconstructs tool_calls from the stored raw data and
strips _provider_content so it never reaches the wire.
Changes:
- _openai_chat.py: add _prepare_messages and _extract_tool_calls hooks
- _google.py: override hooks + tap-pattern _iter_stream for streaming
- model_registry.py: auto-detect .googleapis.com → google provider
- session.py: read cancel_on_approval from ConfigStore
- console/server.py: add PUT/DELETE to proxy route methods
- server.py: fix fork naming (don't inherit source display name)
The channel gateway registers with its Docker-internal hostname
(e.g. http://channel:8091) which is unreachable from a host-side
server. Publish port 8091 and set TURNSTONE_CHANNEL_ADVERTISE_URL
to localhost so the server can reach it for schedule notifications.
GenerationCancelled extends BaseException, not Exception, so it bypassed
the except handler in _run_initial. The finally block ran but
_extract_last_assistant_content returned "" (response never appended to
messages), and _fire_notify_targets bailed on the empty content guard.
Fixes:
- Catch BaseException (not just Exception) in _run_initial so
GenerationCancelled is handled and the UI state is cleaned up
- Remove the empty-content suppression in _fire_notify_targets —
scheduled tasks should always deliver, even with a fallback message
when no output was captured
- Move action buttons (refresh/edit/fork/delete) from header to tab bar,
grouped in #ws-action-group with separators. Contextually adjacent to
the workstream tabs they operate on.
- Toggle group visibility via CSS class (.hidden) instead of per-button
inline style.display — makes media query overrides reliable.
- Call updateWsActionButtons() from renderTabBar() so buttons appear on
initial load and ws_created, not just on tab switch.
- Fix theme loss between nodes: loadInterfaceSettings no longer overwrites
localStorage with server defaults — preserves user's theme choice when
switching nodes via console proxy.
- Add flex-shrink:0 on +/split buttons to prevent squeeze with many tabs.
The model-reload handler read model.default_alias from ConfigStore's
in-memory cache, which could be stale if the earlier best-effort
config-reload notification failed or hadn't arrived yet. Force a
cs.reload() from DB before reading the alias. Also publish config
changes from the console before dispatching model-reload, and
downgrade the misleading "No 'default' model alias" log to debug.
Ctrl+Shift+R Refresh title (regenerate via LLM)
Ctrl+Shift+E Edit title
Ctrl+Shift+F Fork workstream
Ctrl+Shift+X Delete workstream (X not D — avoids Chrome DevTools conflict)
Shortcuts are blocked when any modal is open (edit-title, delete-ws,
batch-delete, new-ws). Help dialog (?) updated with the new bindings.
* feat: add per-node metadata with auto-collection, admin API, and console UI
Adds a normalized node_metadata table for structured per-node key/value
metadata with source tracking (auto/user/config). Auto-populated fields
(hostname, OS, arch, interfaces, cpu_count) are collected at server startup
via stdlib; user-defined fields are managed through the admin API, CLI, or
config.toml [metadata] section.
Storage: migration 035, 7 new protocol methods (get, get_all, set,
set_bulk, delete, delete_by_source, filter), both SQLite and PostgreSQL
backends. Filtering uses single-query GROUP BY/HAVING for efficiency.
Console API: GET/PUT/DELETE endpoints under /admin/nodes/{node_id}/metadata
with auto-source protection. cluster_nodes gains meta.* query param
filtering; cluster_node_detail attaches metadata to responses.
Frontend: new Nodes admin tab with collapsible per-node sections, inline
add form, delete with confirmation. Read-only metadata panel in node
detail drill-down. Proper design token usage, accessibility (ARIA,
keyboard nav, screen reader labels), and mobile responsiveness.
CLI: turnstone-admin list-node-metadata, set-node-metadata, and
delete-node-metadata subcommands.
64 tests (25 storage, 19 node_info, 20 existing unaffected).
* fix: resolve CI typecheck and test failures
- Fix mypy error: use %-style format string instead of structlog kwargs
for standard Logger.warning() in console server
- Fix test_get_nodes assertion to include new node_ids=None parameter
- Add debug logging to _collect_interfaces empty except block
* fix: address Copilot review feedback on node metadata
- Clear stale auto/config metadata before upserting on startup
- Wrap metadata filter in try/except with graceful fallback
- Add metadata field to NodeDetailResponse schema
- Use _VALID_NODE_ID regex for consistent node_id validation
- Defensive JSON decode in admin_get_node_metadata
- Switch to read_json_or_400 and require_storage_or_503 helpers
- Add SetNodeMetadataValueRequest for single-key PUT endpoint
- Add bulk GET /admin/node-metadata endpoint (replaces N+1 fetches)
- Update frontend to use single bulk metadata fetch
* feat: add admin.nodes permission scope for node metadata
- Add admin.nodes to builtin-admin role via migration 035
- Switch all node metadata handlers from admin.settings to admin.nodes
- Register admin.nodes in the admin panel permission set
- Node detail metadata panel fetches from cluster endpoint (no admin
permission needed) instead of admin endpoint
* fix: address second round of Copilot feedback
- Replace inline onclick handlers with data-* attributes and event
delegation to prevent JS string context XSS
- Move NodeMetadataEntry before NodeDetailResponse and use it as the
typed metadata field (was list[dict[str, Any]])
- Clean up config metadata on shutdown (was only cleaning auto)
Add save_messages_bulk() to StorageBackend protocol and both backends.
Fork path now inserts all messages in a single transaction instead of
N individual save_message() calls — for a 200-message workstream this
goes from 200 connection/insert/commit cycles to 1.
FTS5 indexing is intentionally skipped for bulk fork data (historical
messages indexed on rebuild). Ordering preserved via auto-increment id
with a shared timestamp across all rows in the batch.
Also adds 22 endpoint tests covering the 6 new workstream management
endpoints (delete, open, title, refresh-title, list/update interface
settings) and 4 storage-level tests for the bulk insert path.
- Replace inline-style console banner with CSS classes + light/dark theme
- Node ID in banner is now a clickable link back to the node UI
- Add judge_model parameter to create_workstream flow
- Add Google to model provider list with default URL
- Provider-specific placeholder hints in model editor
- Detect results populate model name suggestions datalist
- Theme changes in admin settings apply immediately
- Persist theme selection to server via settings API
- Use workstream title field (with name fallback) in collector SSE events
- Add judge model dropdown to new-workstream modal
Add workstream forking (resume with fork=True keeps new ws_id), custom
naming via aliases, title refresh via LLM, and workstream deletion.
New server endpoints: delete, refresh-title, set-title, open-workstream,
list/update interface settings. Verdict caching with SSE replay on
reconnect, display name fallback (alias→title→name) across all
endpoints, judge_model override per workstream, and settings_changed
broadcast on config reload.
New settings: judge.cancel_on_approval, interface.close_tab_action,
interface.theme. Storage backends updated with name in
list_workstreams_with_history and new get_workstream_metadata method.
Add workstream action buttons in header (refresh title, edit title, fork,
delete) with supporting modals and keyboard shortcuts.
Workstream tabs: always-visible close button, ws_id badge, configurable
close-tab-action (last_used/nearest/dashboard) via interface settings.
Dashboard: batch delete mode with multi-select, saved workstream cards
with ws_id badge, open endpoint for resuming sessions.
Judge display: late-arriving verdict toast when DOM element is gone,
worst-case verdict glow across all tool calls in approval block.
Theme: server-persisted via admin settings API, real-time sync across
clients via SSE settings_changed events.
New workstream modal: judge model dropdown for per-workstream judge
model selection.
- Create fresh HTTP client per evaluation run to avoid stale connections
- Store client factory args instead of client instance for on-demand creation
- Add cancel_on_approval config: when True, abort remaining items on user
approval; when False (default), run all evaluations to completion
- Always deliver LLM verdicts via callback (or fallback when LLM returns None)
- Add _deliver_fallbacks helper for cancelled/incomplete evaluations
- Skip read-only tools for Google provider (requires thought_signature)
- Flatten conversation history to plaintext transcript in _prepare_context
to avoid multi-turn role sequence errors with strict providers like Google
- Use per-turn timeout instead of shared budget so slow turns don't starve
later ones
- Add empty-response retry logic (up to 3 retries without consuming turns)
- Enhanced structured logging throughout judge pipeline
- Update tests to match new signatures and behavioral changes
Add GoogleProvider that extends OpenAIChatCompletionsProvider for
Gemini models via the OpenAI-compatible /v1beta/openai/ endpoint.
- New _google.py with 2M context window defaults and vision support
- Lazy-initialized singleton in create_provider() (thread-safe)
- Route 'google' through OpenAI SDK in create_client()
- Return empty list from list_known_models() (Google models change frequently)
* feat: reconcile judge admin rule UX with edit, disable, and reset actions
Replace the misleading "Customize" button on built-in rules with a
logically consistent 4-state action model: pure built-in (Disable/Edit),
overridden built-in (Disable/Edit/Reset), disabled built-in
(Enable/Edit/Reset), and custom rule (Enable-Disable/Edit/Delete).
Add edit modals for both heuristic rules and output guard patterns,
reusing the existing create modal form structure. Introduce amber
"Reset" button styling to visually distinguish reversible resets from
permanent deletes. Fix source badge redundancy (disabled built-ins now
show grey "built-in" in SOURCE, red "disabled" in STATUS only). Add
aria-labels and role="listitem" for screen reader support.
* fix: preserve built-in pattern_flags and priority on override
Derive pattern_flags from compiled regex for built-in output guard
patterns in the list API so IGNORECASE and other flags survive the
disable/edit/override round-trip. Carry priority through edit modals
via hidden fields so built-in evaluation order is preserved.
* feat: auto-invalidate JWT and static assets on version upgrade
Add a `ver` claim (major.minor) to user-facing JWTs so tokens from
previous versions are rejected after upgrade, triggering re-login.
Service tokens are excluded for rolling-deployment safety. Tokens
without a `ver` claim (pre-upgrade) are accepted for backward compat.
Inject `?v={__version__}` query strings into static asset URLs at
startup so browsers fetch fresh JS/CSS after any release. Vendored
libraries (KaTeX, Highlight.js, etc.) are skipped since they already
carry version numbers in directory paths. HTML responses now include
`Cache-Control: no-cache` to ensure browsers always revalidate.
Frontend detects upgrade-specific 401s and shows a contextual subtitle
("The server was updated — please sign in again"), then performs a full
page reload after re-auth to load the new versioned assets.
* refactor: address PR review — public API name, single decode, idempotent regex
Rename _version_slot() → jwt_version_slot() to make the cross-module
import explicit rather than relying on a private name.
Move version gating from validate_jwt() into check_request() via a new
AuthResult.token_version field. This eliminates the double JWT decode
that occurred on version-mismatch detection — the token is now decoded
once and the version compared afterward.
Guard version_html() regex against double-apply by excluding URLs that
already contain a query string ([^"?]+ instead of [^"]+).
* feat: structured version_mismatch code, ETag, cross-tab auth sync
Add structured "code": "version_mismatch" field to the 401 response
so the frontend detects upgrade-triggered re-auth without string
matching on the error message.
Add ETag headers to HTML index responses (server, console, and proxied
node UI). Combined with Cache-Control: no-cache, browsers send
conditional GETs and receive 304 between upgrades, saving bandwidth.
Add BroadcastChannel-based cross-tab auth sync so logging in on one
tab dismisses the login modal on all other tabs (and vice-versa for
logout).
Add a reminder to the vendored JS update script about the
version_html() regex lookahead.
* fix: remove unused import in test_web_helpers
* feat: Discord /ask model alias, channel default setting, admin UX
Add optional 'model' parameter to Discord /ask command with
autocomplete from available aliases. Model precedence:
explicit > channels.default_model_alias > CLI --model > server default.
- Add channels.default_model_alias to settings registry
- Extend /v1/api/models response with default_alias and
channel_default_alias fields (both server and console)
- Add list_models() to async + sync SDK clients and ChannelRouter
- TTL-cached channel default in ChannelRouter (5min, fail-open)
- @mention path also respects channel default
- Admin Settings tab: model alias settings render as dropdowns
populated from enabled model definitions
- Admin Settings tab: is_secret settings render as write-only
password inputs with save button (replaces static label)
- Update OpenAPI schemas for new response fields
- Validate alias defaults against enabled models on both endpoints
* fix: address PR #306 review feedback
- Move TTL timestamp update before await in get_channel_default_alias
to prevent concurrent duplicate fetches
- Add 30s TTL cache for list_models() to avoid per-keystroke HTTP
traffic during Discord autocomplete
- Type SDK list_models() with ListAvailableModelsResponse instead
of raw dict (both server and console, async + sync)
- Fix IntentJudge.__init__() control flow: model override block was
dangling inside try/except instead of being a separate branch
- Remove provider/base_url/api_key kwargs from server.py and cli.py
JudgeConfig construction (fields removed in prior commit)
- Remove stale TOML mapping entries from config.py
- Remove --judge-provider CLI argument
- Fix Judge settings font sizes to match Settings tab (12px keys,
11px descriptions, tighter spacing, --fg instead of --accent)
Judge model config now uses model aliases exclusively via ModelRegistry.
The separate provider, base_url, and api_key fields on JudgeConfig were
redundant with what's already stored in model definitions. Removes the
fields from JudgeConfig, the explicit-provider resolution path from
IntentJudge.__init__(), and the 3 settings from the registry.
- Replace all raw fetch() + _adminToken with authFetch() helper
- Fix URL paths to use /v1/api/admin/judge/ prefix
- Add r.ok checks on all GET fetches (match existing tab pattern)
- Load model definitions before settings to fix picker race condition
- Escape secret input values with escapeHtml
- Use Mapping type for evaluate_output patterns param (mypy)
- Clean up stale blank lines and comment references
* feat: configurable judge rules with dedicated admin tab
Externalize heuristic intent validation rules and output guard patterns
from hard-coded module constants into the storage abstraction with full
admin UI CRUD. Introduces a dedicated Judge tab in the admin panel that
consolidates all judge configuration (scalar settings, heuristic rules,
output guard patterns) under a single admin.judge permission scope.
- Add heuristic_rules and output_guard_patterns tables (migration 033)
- Add RuleRegistry with thread-safe merge of built-in + DB rules
- Refactor output_guard.py patterns into structured OutputGuardPatternDef
- evaluate_heuristic() and evaluate_output() accept optional rules/patterns
- IntentJudge resolves model aliases via ModelRegistry
- 15 admin API endpoints under /api/admin/judge/ with regex validation
- Judge tab with Settings, Heuristic Rules, and Output Guard sub-panels
- Filter judge.* settings from generic Settings tab
- ConfigStore.storage public property for backend access
* fix: align Judge tab with admin panel design system
- Replace raw <table> with grid-based admin-row/admin-colheaders pattern
- Replace dynamic innerHTML modals with static overlays using focus traps
- Replace confirm() with styled showConfirmModal()
- Replace inline badge styles with scope-badge classes
- Add mobile responsive breakpoints for Judge tab grids
* fix: Judge tab accessibility and polish
- Extract sub-section switcher inline styles to CSS classes
- Add focus-visible outline and reduced-motion support
- Add tab button IDs and fix aria-labelledby on tabpanels
- Add tabindex roving and arrow key navigation for sub-tabs
- Add role=list and aria-live to table containers
- Replace status text with scope-badge classes for scannability
* fix: address CodeQL and Copilot review feedback
- Remove unused validation constants from rule_registry.py (CodeQL)
- Return MappingProxyType from output_patterns for immutability
- Fix ThreadPoolExecutor shutdown(wait=False) to prevent hangs
- Use separate _VALID_OG_RISK_LEVELS (no "critical") for output guard
- Pass pattern_flags to regex validation in update endpoint
- Chain redactions in configurable mode (compose pattern + complex)
- Initialize RuleRegistry on console app.state
- Fix test fixtures to use valid enum values (approve/review/deny)
* fix: use Mapping type for evaluate_output patterns param (mypy)
* feat: multi-model health tracking with runtime default and DB-only startup
Replace active-probe circuit breaker with passive per-backend health
tracking. Backends are marked degraded after consecutive failures and
recover when a request succeeds — requests are never blocked.
- Add model.default_alias ConfigStore setting for runtime default model
- Make load_model_registry CLI args optional for DB-only startup
- Per-(provider, base_url) health trackers via HealthTrackerRegistry
- Two-pass fallback: prefer healthy backends, then try degraded
- Remove BackendHealthMonitor, CircuitState, probe threads, cooldown
- Remove circuit_state from API schema, SDK events, metrics, frontends
* feat: add "Set Default" button to Model Definitions admin panel
Show a "default" badge on the current default model alias and a
"set default" action button on all other models. Clicking it writes
model.default_alias via the settings API. The list endpoint now
includes default_alias in the response so the UI can highlight it.
* fix: address review feedback — metric scoping, effective default, session alias
- Move turnstone_backend_up metric out of BackendHealthTracker into
server callback; only the effective default backend drives the gauge
- _build_health_dict resolves effective default via ConfigStore override
- session_factory computes selected_alias once before registry.resolve
- admin model-definitions endpoint returns effective default (not just
override) so UI shows correct badge when ConfigStore is empty
- Rename circuitTitle → healthTitle in console JS
- Fix ruff SIM117 lint in test
* fix: validate effective default against enabled models, degraded label, log normalization
- admin model-definitions endpoint validates default_alias against
enabled models using same fallback rules as load_model_registry
- UI text "backend down" → "backend degraded" to match advisory semantics
- Health tracker log uses normalized base_url from key, not raw argument
Migration 031 created the prompt_policies table but never registered
admin.prompt_policies in _VALID_PERMISSIONS or granted it to the
builtin-admin role, causing 403 on all prompt-policy admin endpoints.
* fix: capacity-aware tool output truncation and context overflow recovery
Large tool results (e.g. 593K-char search output) could overflow the
context window in a single turn when the conversation was already
partially full. The fixed 50%-of-context truncation limit didn't
account for current usage.
Changes:
- _truncate_output() now accepts remaining token budget and uses
min(tool_truncation, remaining_budget_chars) as the effective limit
- _remaining_token_budget() helper calculates available capacity with
reserves for max_tokens response and 5% safety margin
- Safety truncation at tool-result append: every string tool result is
clamped to remaining budget before entering the message array
- _exec_web_search() now calls _truncate_output() (was missing)
- Context overflow recovery: catches provider errors indicating context
length exceeded (OpenAI + Anthropic patterns), auto-compacts, retries
once. Falls back to original error if compact-and-retry fails.
* fix: address review — zero-budget floor, nested spinner, Anthropic patterns, tests
- Remove 256-char floor from budget truncation — zero budget now returns
a placeholder instead of allowing 256 chars through
- Stop thinking spinner before compact to avoid nested start/stop
- Add Anthropic error patterns (prompt is too long, input tokens)
- Wrap compact-and-retry so failures re-raise the original error
- Add 15 tests covering budget calculation, capacity-aware truncation,
and overflow recovery for both providers
* fix: cap response reservation at 25% of context window
Reserving the full max_tokens in _remaining_token_budget() zeroed the
budget for common configs like max_tokens=32768 on a 32K context,
collapsing all tool output to a placeholder. max_tokens is a ceiling,
not guaranteed consumption — cap the reserve at context_window // 4.
Adds regression test for max_tokens >= context_window.
* fix: skip chat_template_kwargs for commercial OpenAI API
OpenAI rejects chat_template_kwargs as an unknown parameter — it's only
meaningful for local model servers (vLLM, llama.cpp, SGLang).
Split OpenAIProvider into separate singletons for "openai" vs
"openai-compatible" so _provider_extra_params can gate on provider_name
instead of inspecting base_url. Also deduplicates agent inline code into
the same method and fixes pre-existing test pollution where
get_capabilities was mutated on the singleton without cleanup.
* feat: add OpenAI Responses API provider for commercial models
Split the OpenAI provider into three concrete implementations behind the
LLMProvider protocol:
- _openai_chat.py: Chat Completions API for local model servers
(vLLM, llama.cpp, SGLang)
- _openai_responses.py: Responses API for commercial OpenAI
(GPT-5.x, O-series)
- _openai_common.py: shared capability table, temperature/reasoning
gating, cache retention, citations, usage extraction
The Responses API handles reasoning_effort as a {"effort": value} dict,
system messages as an instructions field, and tool format translation at
the provider boundary. ChatSession is unchanged — the provider abstracts
the API difference.
Also fixes diff_file direction when comparing against provided content.
* fix: Responses API input format and local model provider routing
- Assistant input messages use plain string content (not output_text)
- Tool call argument deltas match on item_id, not call_id
- Auto-detect openai-compatible provider for non-api.openai.com URLs
- Fix diff_file direction when comparing against provided content
* fix: resolve env vars before provider auto-detection in config.toml models
Config-file model entries using ${ENV_VAR} placeholders in base_url were
not resolving env vars before _resolve_openai_provider(), causing
commercial OpenAI configs to be misclassified as openai-compatible.
* fix: replace empty except blocks with diagnostic logging
Add log.debug/warning to 7 bare except-pass blocks that silenced
failures in security-relevant or operationally-important paths:
- Channel route lookup, CLI policy evaluation, OIDC JWKS fetch,
prompt policy loading, plan file write, routing override, username
resolution.
Plan write now reports failure to user instead of falsely claiming
"Plan saved."
* fix: replace assert-with-side-effect and narrow BaseException catch
- Convert 4 assert isinstance() to explicit TypeError raises — assertions
are stripped under python -O, removing runtime type checks
- Narrow except BaseException to except Exception in fallback handler —
KeyboardInterrupt/SystemExit should not record as health failures
- Plan write failure now reports error to user instead of "Plan saved"
* fix: wire up toast error type and remove useless conditional
- showToast() now accepts optional type param ("error") with red border
styling — 3 call sites were passing "error" that was silently ignored
- Remove always-true if (q) guard after early-return on empty query
* fix: remove unreachable return None after return self._judge
* fix: parenthesize multi-line string concatenations in dev_parts list
Explicit parens make intentional concatenation unambiguous to static
analysis (CodeQL implicit-string-concatenation-in-list rule).
* fix: remove constant-true filter in test mock — return list directly
* fix: extract side-effecting calls from assert in tests
store.delete() and mgr.close() have side effects that would be
stripped under python -O. Assign to variable first, then assert.
* fix: remove unused local variables in tests
Drop assignments to unused workstream/variable references created
solely for side effects. Use _ for unused tuple unpacking.
* fix: use admin.prompt_policies permission for prompt policy endpoints
All 5 prompt-policy endpoints (list, create, get, update, delete)
were checking admin.policies (the tool-policy permission) instead of
admin.prompt_policies. This caused a mismatch with the admin UI which
gates the tab on admin.prompt_policies — users could see the tab but
get 403, or reach the endpoint but never see the tab.
* fix: use caplog instead of capsys for structlog warning assertion
structlog output goes through the logging system, not stdout/stderr.
* fix: address review — remove dead isinstance, module-level import, unnecessary lambdas
- session.py: remove unreachable isinstance check (has_batch already
validates raw_edits is a list)
- cli.py: move logging import to module level
- test_workstream.py: replace lambda wid: FakeUI(wid) with FakeUI
* fix: harden MCP client against misbehaving servers
Misbehaving/failed/misconfigured MCP servers could peg CPU at 100% due
to anyio cancel-scope busy-loops (SDK #2147), uncancelled orphaned
futures, and missing application-layer resilience.
Five fixes:
1. Cancel orphaned futures on timeout — future.cancel() in all sync
bridge methods prevents coroutine accumulation on the event loop
2. Per-server circuit breaker — 3-failure threshold with exponential
cooldown (30s–5min), per-server jitter, auto-reconnect on half-open
probe, McpError excluded (protocol errors from healthy servers)
3. Safe transport stream pre-close — store stream refs and close them
before stack teardown in all error/shutdown paths, preventing the
anyio zero-buffer CPU busy-loop
4. Notification debounce — 5s per-server rate limit on list_changed
refresh storms from buggy servers
5. Periodic refresh backoff with auto-reconnect — disconnected servers
get reconnection attempts with exponential backoff (60s–1hr) instead
of being silently skipped forever
* docs: add MCP resilience section to architecture docs and diagram
Document the circuit breaker, future cancellation, stream pre-close,
notification debounce, and periodic refresh backoff in the architecture
guide and the MCP architecture PlantUML diagram.
* fix: address review — stack leak on transport error, half-open comment
- Widen _connect_one guard to check _per_server_stacks too, not just
_sessions. Transport errors in sync dispatch methods evict the session
but left the stack behind, leaking anyio tasks on reconnect.
- Clarify half-open design: multiple callers are intentionally allowed
through (reconnects serialize on the event loop, first failure re-trips).
* fix: mobile UX for console sidebar drawer and server chat input
Console admin sidebar: add box-shadow elevation, close button with
focus return, 44px touch targets, focus-into-drawer on open, flip
active indicator to left border, cubic-bezier easing, aria-expanded,
fix resize handler state desync, guard toggle injection for panels
without toolbars.
Server chat input: on touch devices Enter inserts newline (tap Send
button to send), hide Shift+Enter hint from placeholder.
* fix: preserve first group label spacing when close header is injected
Add sibling combinator selector so the first sidebar group keeps its
reduced top padding regardless of whether the close header div is
present as first-child.
* feat: render rich media embeds for MCP tool results
Detect structured media JSON (stream_url, results, sessions) in MCP
tool output and render interactive cards instead of plain text.
Web UI: media cards with thumbnail, title, metadata, and click-to-play
video/audio. HLS via lazy-loaded hls.js with direct-stream preference.
Collapsed raw JSON (API keys redacted) for inspection.
Discord: rich embeds with proxied thumbnail images (fetched by the bot
since Discord CDN cannot reach private media servers). Search results
as numbered lists, session state as "Now Playing" cards. Stream URLs
never exposed in embeds — web_url used for safe clickable links.
CI: vendor hls.js 1.6.15 with renovate tracking and update script.
* fix: address PR #292 review — SSRF guards, streaming fetch, tests
- URL validation: reject non-http(s) schemes and userinfo in thumbnail
URLs. Private IPs intentionally allowed (media servers are on LAN).
- Streaming fetch: use http.stream() with aiter_bytes() and a running
byte count to enforce the 2MB cap without buffering the full response.
Validate content-type is image/* before downloading.
- Resilience: wrap try_build_media_embed in try/except in bot.py so a
media embed failure falls through to the code-block path.
- LICENSE: download hls.js LICENSE from npm on update instead of only
copying from old dir.
- Tests: add 19 new tests — try_parse_media (8 cases), _is_safe_image_url
(7 cases), embed builders (4 cases including stream_url exclusion and
string season/episode safety).
* chore: add LICENSE file for vendored hls.js
* fix: remove ANSI escape codes from tool preview fields
Preview text (tool args, URLs, queries) was wrapped in DIM/RESET ANSI
codes at the source in session.py, which leaked into SSE events and
rendered as raw escape sequences in Discord and the web UI.
Move ANSI styling to the CLI consumer (cli.py) where it belongs. Also
escape markdown in Discord tool name titles to prevent __ from being
interpreted as underline formatting.
* fix: drop [MCP: server] prefix from tool descriptions
The prefix made MCP tools look second-class compared to builtins,
causing models to hesitate using them. The server name is already
encoded in the tool name (mcp__server__tool).
* feat: pretty-print JSON tool output, player error state, broader key redaction
- JSON tool results are detected and pretty-printed with 2-space indent
instead of rendering as a wall of text
- API key redaction extended to cover api_key, apiKey, api-key, and
token query params across all tool output (not just media embeds)
- Video/audio player shows styled error message when stream fails to
load instead of leaving a broken player element
- Both appendToolOutput and replayHistory use shared renderToolOutput()
* fix: designer review — player error retry, contrast, tool-cmd cap
- Player error: role="alert" for screen readers, retry button that
reuses existing play handler, includes media title in error message
- Light theme: darken --red from #dc2626 to #b91c1c (5.7:1 contrast
on --code-bg, was 4.3:1 failing WCAG AA at 12px)
- Pretty-print collapsed raw JSON in media embeds (was missed earlier)
- Cap .tool-cmd at 120px to prevent tools with many args from making
approval blocks disproportionately tall in history replay
- Dedicated .media-player-error class instead of reusing .tool-output
* fix: Discord tool info name matching regression, suppress deprecation warning
The escape_markdown call on tool names was stored for matching against
ToolResultEvent.name, but event.name is raw/unescaped. The escaped name
never matched, so the "Running → Done" transition silently failed and
previews disappeared from the status embed.
Fix: store raw name for matching, use escaped name only for display.
Also suppress discord.py's re.sub count deprecation warning (Python
3.13+ issue, fixed upstream).
* fix: update MCP tool description tests to match prefix removal
* fix: address PR #292 review round 2
- Retry button: handle missing span children in click handler so retry
buttons from player error state don't throw
- Footer count: use len(lines) instead of min(len(results), 10) to
reflect actual rendered count after char budget truncation
- Null display: use "null" instead of "None" in JS tool arg preview
- Broader redaction: also redact JSON "api_key": "..." patterns
- SSRF hardening: block loopback and link-local IPs plus cloud metadata
hostnames in thumbnail fetch (private LAN IPs still allowed)
* fix: bundle production compose.yaml for pipx users (#293)
Users who install via pipx don't have a git clone, so there's no
compose.yaml or Dockerfile. Bootstrap now extracts a bundled production
compose file that uses pre-built ghcr.io images instead of local builds.
- Add turnstone/deploy/compose.yaml (ghcr.io images, no build blocks,
single-node production profile only)
- Add write_compose tool to bootstrap wizard
- Update bootstrap system prompt to check for and write compose.yaml
- Remove stale ddgCluster profile references from system prompt
- Include turnstone/deploy/*.yaml in wheel
* fix: use postgresql+psycopg:// DSN scheme in compose fallbacks
The Docker image ships psycopg3, not psycopg2, so the bare
postgresql:// scheme fails. Also clarify PG usage comment in
production compose.
* fix: improve web_fetch reliability — strip scripts, dynamic truncation, more tokens
- strip_html() now removes <script>, <style>, <template>, <noscript>
element content instead of just their tags
- Truncation budget scales with context window (75% in chars, 50k floor)
and takes from the beginning only instead of head+tail splice
- max_tokens bumped from 2000 to 8192 so thinking models don't starve
the visible extraction answer
- reasoning_effort="low" on summarization call to avoid wasting tokens
- Empty responses and empty extractions now report as tool errors
* refactor: extract _utility_completion to fix reasoning_effort duplication
Callers previously had to pass reasoning_effort both as a direct keyword
(for commercial providers) and via _provider_extra_params (for local
model servers). This duplication was easy to get wrong — web_fetch was
already missing the direct keyword.
_utility_completion threads it through both paths from a single call,
used by title generation, compaction, and web_fetch extraction.
* fix: disable thinking when max_tokens too small, cap extraction at 500k
_reasoning_params now returns empty dict when max_tokens can't fit a
thinking budget (e.g. title gen with max_tokens=200). Previously
produced budget_tokens >= max_tokens which is an API error on
manual-thinking Anthropic models.
Also caps web_fetch content truncation at 500k chars — the dynamic
context-window calc was producing 3M chars on 1M-context models.
* fix: clamp utility max_tokens to model output limit, add strip_html tests
_utility_completion now clamps max_tokens to the model's advertised
max_output_tokens so small/local models don't reject 8192-token
requests.
Adds 8 tests for invisible element stripping (script, style, template,
noscript) including multiline, case-insensitive, and attribute cases.
* fix: mock get_capabilities in title retry tests for _utility_completion
_utility_completion calls _get_capabilities to clamp max_tokens. The
existing title tests mocked _provider as a bare MagicMock, so
caps.max_output_tokens was a truthy MagicMock instead of an int. Set
get_capabilities to return a real ModelCapabilities instance.
Build the image once via the profileless console service and reference
it as turnstone:local from server/channel. Prevents stale images when
users run docker compose build without --profile.
* fix: include prompt .md files in wheel, add wheel-completeness CI (#289)
Prompt markdown files were missing from PyPI wheels since the modular
prompts refactor, causing FileNotFoundError on startup for pip-installed
users. Add the missing include pattern and a new CI job that diffs
source-tree data files against wheel contents so omissions are caught
before merge.
* fix: sanitise ALLOW patterns in wheel-completeness check
Strip blank lines and leading whitespace from the allowlist before
passing to grep -vFxf so empty patterns cannot silently match all lines.
* fix: log clean one-liner when PostgreSQL becomes unavailable
Wrap all 174 connection sites in PostgreSQLBackend through a _conn()
context manager that catches OperationalError, emits a single
database.unavailable log line (with connection URL), and suppresses
repeats until the connection is restored (database.connection_restored).
* fix: add StorageUnavailableError and cover all heartbeat loops
Address review feedback:
- Separate connect-phase from execution-phase in _conn() so that
OperationalError during caller code (e.g. BEGIN IMMEDIATE lock
contention) is not misclassified as a connectivity failure.
- Add StorageUnavailableError exception class so callers can
distinguish transient DB outages without redundant tracebacks.
- Apply the same _conn() wrapper to SQLiteBackend for consistency.
- Catch StorageUnavailableError in all 7 periodic loops: watch
runner, server heartbeat, channel heartbeat, console heartbeat,
collector discovery, rebalancer, and scheduler.
- Guard dedup flag with threading.Lock.
- Add tests for dedup logging and PostgreSQL path.
* fix: chunk IN clauses to stay within DB parameter limits
psycopg caps query parameters at 65 535 and SQLite defaults to 999.
assign_buckets, prune_workstreams, and count_skill_resources_bulk were
passing unbounded lists into single IN(...) clauses, causing
OperationalError during rebalancer runs on full-size hash rings.
Chunk sizes: 10 000 (PostgreSQL), 500 (SQLite).
* fix: deduplicate assign_buckets input, add chunking regression tests
Address review feedback: deduplicate bucket list before chunking to
prevent inflated rowcount from cross-chunk duplicates. Add tests that
exercise the multi-chunk path (1200 buckets > SQLite chunk_size of 500)
and verify dedup preserves accurate counts.
Multiple CI completions for the same commit (tag push + branch push)
caused duplicate publish and docker runs. Concurrency group keyed on
head_sha ensures only one publish runs per commit.
* chore: release infrastructure for dual-track stable/experimental
CI/CD changes for the 1.0 release:
- Gate PyPI publish and Docker publish on CI success via workflow_run
- Add docker-publish.yml: builds and pushes to GHCR with smart tagging
(stable gets :X.Y.Z/:X.Y/:stable/:latest, pre-release gets :experimental)
- Add stable/* and v* tags to CI and docker-scan triggers
- Remove stale [mq] extra and types-redis from CI (Redis MQ deleted)
- Remove stale redis from Renovate package rules
Release tooling:
- scripts/release.sh: bump version, uv lock, commit, tag (with --push)
- docs/releasing.md: documents stable/experimental workflow
Docker:
- Add /workspace mount point (WORKSPACE_MOUNT env var, defaults to empty volume)
- Update .env.example: remove stale Redis/auth-token refs, add workspace/model/discord
README:
- Remove beta warning, add hero image and release tracks table
* fix: derive release tag from git instead of workflow_run.head_branch
Use git tag --points-at HEAD after checkout to resolve the release
tag instead of relying on workflow_run.head_branch, which may not
reliably be the tag name for tag-triggered CI runs. Both publish
and docker-publish workflows now skip cleanly when no v* tag exists
at the checked-out commit.
* fix: replay plan review prompt on SSE reconnection
Plan approval prompts were lost when a user navigated to the server
web UI from the console dashboard (triggering a new SSE connection).
Tool approvals stored pending state in _pending_approval and replayed
it on reconnection, but plan reviews used fire-and-forget _enqueue
with no persistent state.
Mirror the _pending_approval pattern: store _pending_plan_review
before blocking, replay it in events_sse for new SSE clients, and
clear it on resolution. Without this fix, plan reviews silently
timed out after 1 hour and were treated as approval.
* test: add plan review SSE replay regression tests
Covers pending state lifecycle: stored during on_plan_review, cleared
on resolve_plan, available for SSE reconnection replay.
The async variants accepted token_factory for auto-rotating JWTs via
ServiceTokenManager, but the sync wrappers did not expose or forward
the parameter. External SDK users calling the sync clients with
token_factory got a TypeError.
Settings rows and admin table rows used align-items: center, which
caused inputs and source badges to drift away from their labels when
descriptions wrapped to multiple lines. Switch to align-items: start
so controls stay next to their label names regardless of row height.
Add 2px top margin on settings toggles to pixel-align with text input
top padding in start-aligned rows.
* fix(examples): rewrite mcp-cluster-ops to use console SDK for cluster routing
The example MCP server was broken after the direct HTTP transport
refactor — it used TurnstoneServer (single-node) for cluster ops that
require TurnstoneConsole (cluster gateway). Rewrites dispatch flow to:
route via console → SSE stream from node → cleanup via console.
- Switch from TurnstoneServer to TurnstoneConsole for node listing and
workstream routing (TURNSTONE_CONSOLE_URL replaces TURNSTONE_SERVER_URL)
- Add proper workstream lifecycle: create via routing proxy, stream from
node, close in finally block with leak-safe ws_id guard
- Catch dispatch exceptions in run_on_node for structured JSON errors
- Extract _extract_node_ids helper, remove dead n.get("id") fallback
- Normalise _console_kwargs to always include token key
- Rewrite tests against Console+Server mocks (36 → 44 tests)
* fix(examples): paginate node listing and clarify auth in README
Address Copilot review feedback on #278:
- _list_nodes_sync now paginates via offset/limit loop so clusters
with >100 nodes are fully discovered
- README step 2 now mentions token passthrough for authenticated clusters
- New test_paginates_large_clusters verifies multi-page fetch (45 tests)
* fix: remove non-auth support from bootstrap wizard
Auth is now mandatory for all deployments. Remove the
TURNSTONE_AUTH_ENABLED toggle and make JWT_SECRET and AUTH_TOKEN
required in the wizard's system prompt.
* fix: remove auth disable support from runtime and infra
Remove AuthConfig.enabled field — auth is always on. Drop
TURNSTONE_AUTH_ENABLED env var, config toggle, and the
check_request bypass. Update compose.yaml, Helm chart,
Terraform, docs, and tests to match.
* feat: deprecate config tokens, require JWT secret, prefer JWT auth
Phase 1 of config-token removal:
- load_jwt_secret() now exits with error if no secret is configured
(was: silently auto-generated ephemeral secret)
- _authenticate_token() logs deprecation warning on config token use
- CLI /cluster commands use ServiceTokenManager when JWT secret is set
- turnstone-admin tls-list uses ServiceTokenManager when JWT secret is set
- Update bootstrap wizard, docker.md, security.md to mark
TURNSTONE_AUTH_TOKEN as deprecated and JWT_SECRET as required
- Console test fixtures use auth token + headers (auth always enforced)
* feat: add service scope for inter-service JWT auth
Add "service" to VALID_SCOPES and SCOPE_HIERARCHY. Service tokens
bypass require_permission() RBAC checks, replacing the old
empty-user-id bypass that config tokens relied on.
All ServiceTokenManager instances that need admin access now include
"service" in their scopes (console proxy, channel gateway, CLI,
admin CLI). Read-only services (collector, notification) unchanged.
* feat: phase 2 config token deprecation
- SDK doc examples now show API tokens (ts_) instead of config tokens
- Remove _get_config_token() from admin CLI (dead code)
- Block config token exchange in handle_auth_login — only password
and API token login allowed
- Update login tests to use password-based auth instead of config
token exchange
* feat: phase 3 — remove config tokens entirely
Complete removal of config-file token authentication:
- Delete AuthConfig.tokens, check(), _ROLE_TO_SCOPES, hmac dispatch
branch, and config token loading from load_auth_config()
- Remove auth_config parameter from _authenticate_token() and
check_request() — callers updated throughout
- Remove TURNSTONE_AUTH_TOKEN from compose.yaml, Helm charts,
Terraform, turnstone.example.toml
- Remove --auth-token CLI flags from turnstone, turnstone-admin,
and turnstone-console
- Simplify console main() — always use ServiceTokenManager
(no fallback to static tokens)
- Delete config-token-specific tests, rewrite check_request and
integration tests to use JWT auth with proper audience claims
- Remove all config token references from docs (security.md,
docker.md, sdk.md, console.md, architecture.md, bootstrap prompt)
* fix: address code review findings
- Fix 33 broken tests: add JWT auth to test_api_versioning,
test_console_routing_proxy, test_tls_admin, test_tls_manager,
test_server_live (jwt_secret + audience-scoped auth headers)
- Add TestRequirePermissionServiceScope: 4 tests covering the
service scope RBAC bypass path
- Remove stale comments referencing config tokens in auth.py and
console/server.py
- Remove dead proxy_auth_token parameter from console create_app()
and static token fallback in _proxy_auth_headers()
- Remove TURNSTONE_AUTH_TOKEN from env.py scrub list
* fix: address Copilot review — JWT audience, compose require secret
- CLI /cluster: add audience=JWT_AUD_CONSOLE to ServiceTokenManager
(console validates audience, JWTs without it were rejected)
- Admin CLI tls-list: same audience fix
- compose.yaml: TURNSTONE_JWT_SECRET now uses :? to fail fast if unset
- SDK console: fix default port from 8081 to 8090
* test: add auth enforcement tests for TLS admin endpoints
5 new tests: unauthenticated requests return 401 (list, renew,
delete), read-only-scoped requests return 403 (renew, delete).
Closes the TLS auth enforcement test gap noted in PROGRESS.md.
* fix: address remaining Copilot review feedback
- Fix token_source="config" → "test" in TLS test fixtures
- Fix AuthResult.token_source docstring to include service origins
- Require TURNSTONE_JWT_SECRET in cluster compose profile (:?)
- Helm: add auth.jwtSecret + auth.existingSecret values, wire
TURNSTONE_JWT_SECRET into secret.yaml and both deployments
- Terraform: replace auth_token with jwt_secret variable + secret,
remove orphaned auth_token resources and IAM reference
- Remove [[auth.tokens]] from security.md config example
* fix: address full code review — 10 findings
Critical:
- Terraform: replace concat(common_env, auth_env) with common_env
(auth_env local was removed but still referenced)
- Channel gateway: remove hmac static token auth from _check_auth(),
use JWT-only validation. Remove --auth-token CLI arg from channel
- Rebalancer: add token_manager support so migration requests carry
JWT auth (was sending unauthenticated POST to /internal/migrate)
Major:
- Guard _permissions_to_scopes() against "service" privilege
escalation from DB role permissions
- Remove dead AuthConfig class, load_auth_config(), and all
auth_config parameters from create_app() signatures
- Helm: inject JWT secret for both inline and existingSecret paths
Minor:
- Remove dead auth_token param from ClusterCollector
- Remove empty TestLoadAuthConfig class
- Short JWT secret now exits instead of warning
- Compose: add generation command comment above JWT_SECRET
- Clean stale config token references from 6 doc files
- Clean stale AUTH_TOKEN reference from bootstrap wizard prompt
* fix: remove remaining stale config token references from docs
- channels.md: remove --auth-token from options table
- oidc.md: remove "config-file tokens still work" claim
- security.md: remove config token section, fix JWT secret docs
(now required/exits, no ephemeral fallback), remove hmac from
ASCII diagram, remove --auth-token reference
* fix: populate model in _last_usage so usage-by-model records correctly
_last_usage was built purely from UsageInfo token counts, never
including a "model" key. server.py's on_status() fell back to
model="" for every record_usage_event call, so GROUP BY model
collapsed all rows into a single empty-key bucket.
* fix: inject model at emission time, preserve dict[str, int] typing
Address Copilot review: keep _last_usage as dict[str, int] for type
safety, inject "model" from self.model when passing to on_status().
This also fixes stale model after /model switch since the value is
read fresh each time.
* fix: MCP tools not surfacing after Sync to Nodes, update Anthropic tool search
Three fixes:
1. session_factory closure captured mcp_client=None when no --mcp-config
was passed at startup. internal_mcp_reload created a new MCPClientManager
on app.state but the factory never saw it. New workstreams got 0 MCP tools.
Fix: mutable _mcp_ref list shared between factory and reload handler.
2. Anthropic dropped the date suffix from tool_search_tool_bm25_20251119
and now requires name == type. Updated constant and tool definition.
3. Add diagnostic logging around API errors (provider, model, base_url,
message counts, full exception chain) and workstream resume (pre/post
provider state, alias resolution warnings).
Also adds Node.js 24 LTS to Dockerfile via multi-stage copy for npx-based
MCP servers.
* fix: address Copilot review — set_storage on reload, sanitize log output
- Call mcp_mgr.set_storage(storage) when internal_mcp_reload creates a
new MCPClientManager so prompt sync works for post-startup servers
- Strip query params from base_url before logging (may contain API keys
in some vLLM deployments)
- Split API error logging: concise warning (type names only) + separate
debug with exc_info=True for full traceback when needed
* chore: remove DDG MCP sidecar, web_search uses built-in ddgs client
The DuckDuckGo MCP server container is redundant — the built-in
DuckDuckGoClient (via ddgs package, included in all extras) auto-detects
when no Tavily key is configured. Removes the ddg-search service,
ddgCluster profile, and mcp-ddg.json config file.
* fix: materialize skill resources to disk for subprocess access
Skill-bundled scripts stored in skill_resources were loaded into memory
but never written to disk, causing FileNotFoundError when the model
tried to execute them. Write resources to a per-workstream temp directory
on skill load, expose via SKILL_RESOURCES_DIR env var and PATH, clean up
on skill change or session close.
* fix: pre-flight validation warns when skill references missing resources
Scan rendered skill content for path references (scripts/foo.py, etc.)
and compare against bundled skill_resources. Warn via on_info if any
referenced paths are not bundled, so operators see the gap at skill
activation rather than at runtime FileNotFoundError.
* fix: address PR #271 review feedback
- Fix trailing colon in PATH when $PATH is empty (cwd-on-PATH risk)
- Move try/except inside per-resource loop so one bad write doesn't
abort all resources
- Explicit encoding="utf-8" for deterministic writes across locales
* fix: address PR #271 review round 2
- Normalize available paths in _validate_skill_resources() to match
referenced paths (both sides use os.path.normpath now)
- Fix flaky traversal test: assert inside base dir, not escaped path
Multiple browser tabs open to the same server could see workstream
names, states, and content mixed up between workstreams when creating,
closing, and switching tabs rapidly.
Root causes and fixes:
- Global SSE ws_created events were never handled — other tabs never
learned about new workstreams, causing blank names and stale tab bars
- SSE reconnection assigned all stale panes to the first workstream
instead of deduplicating; now uses two-pass assignment with tracking
- switchTab left the old EventSource open while reassigning pane.wsId,
creating a window for events to leak; now disconnects SSE first
- Per-workstream events carried no ws_id — server now stamps ws_id on
all events via _enqueue (shallow copy); client handleEvent drops
events with mismatched ws_id as defense-in-depth
- Plan dialog used pane.wsId at resolve time (could drift after tab
switch); now captures ws_id when the dialog opens
- Global ws_closed could reassign panes before per-ws SSE finished
draining; now disconnects per-ws SSE immediately on close
The 5 prompt policy admin endpoints required "admin.prompt_policies"
but the builtin-admin role only grants "admin.policies". Changed to
match the existing permission used by tool policy endpoints.
* fix: harden Discord bot against gateway disconnects and SSE failures
- Isolate Discord API failures from SSE stream — _on_ws_event exceptions
no longer kill the SSE connection and cause missed events
- Fix broken exponential backoff on 4xx/5xx (delay was reset on every
attempt); skip aiter_sse() on error responses
- Add read timeout (90s) to SSE httpx client so half-open TCP
connections are detected and recovered
- Re-resolve node URL on each SSE reconnect attempt
- Add on_resumed handler to recover SSE tasks that died during brief
gateway disconnects (on_ready is not called on session resume)
- Sync slash commands only on first on_ready to avoid Discord rate limits
* fix: SSE backoff on 4xx/5xx and retrieve dead task exceptions
- Replace `continue` with raise+catch so 4xx/5xx errors hit the
exponential backoff path instead of tight-looping
- Retrieve task exceptions in _purge_dead_sse_tasks to suppress
"Task exception was never retrieved" warnings and log the cause
* feat: modular system message composition with admin prompt policies (#267)
Replace the monolithic persona+tools block in _init_system_messages()
with a modular composition harness (turnstone/prompts/). System messages
are now assembled from five typed layers: BASE (persona), ENV (client
surface — web/cli/chat), CONTEXT (datetime, timezone, username), TOOLS
(usage patterns), and POLICIES (behavioral rules with tool gating).
Prompt policies are admin-managed via a new Prompts tab in the Governance
group (CRUD with modal forms, tool gating, priority ordering, enable/disable).
DB policies override file-based defaults by name; file-based policies serve
as deployment defaults. Migration 031 adds the prompt_policies table.
ClientType is threaded end-to-end from channel adapters through the SDK,
HTTP API, WorkstreamManager, and session factory to ChatSession. Discord
sessions now receive chat-optimized formatting (no tables, no Mermaid,
concise output) instead of the web UI's rich markdown instructions.
* fix: address CI failures and Copilot review feedback
- Add client_type param to CLI session_factory (mypy protocol match)
- Add prompt policy CRUD to PostgreSQL backend (test-postgres CI)
- Fix ClientType resolution: compare against enum values, not members
- Fix null client_type coercion (body.get returns None, not "")
- Use local time with astimezone() instead of UTC with local tz name
- Sanitize tool_gate in update endpoint (coerce null to empty string)
* feat: replace console HTTP polling with persistent SSE streams
Console collector now subscribes to each server node's /v1/api/events/global
SSE stream for real-time state updates instead of polling /v1/api/dashboard
and /health every 15 seconds.
Server changes:
- Emit ws_created/ws_closed events on global queue from create/close handlers
- Add node_snapshot on SSE connect (workstreams, health, aggregate)
- Add ?expected_node_id= identity verification (409 on mismatch)
- Add health_changed callback to BackendHealthMonitor circuit breaker
- Add periodic aggregate emitter thread (10s)
Console collector changes:
- Single asyncio event loop on one thread multiplexes all SSE connections
(scales to 1000+ nodes vs thread-per-node)
- Discovery loop spawns/cancels async SSE tasks per node
- Snapshot reconciliation on connect, delta application for live events
- Fix ws_state→cluster_state event type mismatch
- Remove polling code (poll_interval, max_poll_workers, --poll-interval CLI)
SDK changes:
- Add NodeSnapshotEvent, HealthChangedEvent, AggregateEvent dataclasses
- Add stream_node_events() method (async + sync)
* fix: address review feedback on node event streams
- Fix stop() to let SSE manager exit naturally instead of force-stopping
the event loop (ensures finally cleanup runs)
- Guard against empty/invalid SSE data from ping frames
- Treat missing node_id as identity mismatch (409) when expected_node_id
is provided
- Fix stale docstring on _update_metrics
* feat: show thinking indicators, tool calls, and results in Discord threads
Discord threads now surface real-time activity during multi-tool chains
instead of appearing idle. ThinkingStart/Stop events display a transient
italic status message. ToolInfoEvent sends a per-tool "running" embed
that ToolResultEvent edits in-place with the result (FIFO matching by
tool name, fallback to new message). Includes backtick-injection escaping
in tool output. Visibility respects auto-approve config so tool calls
always appear somewhere.
* fix: address review feedback on Discord action visibility
Delete thinking messages in unsubscribe/stale-route cleanup (not just
pop state). Sanitize tool-call previews (escape backticks, strip
mentions). Fix format_tool_result docstring re ellipsis line count.
Add regression test for triple-backtick escaping.
* fix: disable approval buttons on server-side resolution (timeout)
Handle ApprovalResolvedEvent in _on_ws_event to disable buttons and
grey out the approval embed when the server resolves the approval
externally (timeout, auto-approve from another client). Extract
disable_message_buttons helper from views.py so it works on a plain
Message (not just an Interaction).
* fix: reply with guidance when user DMs the bot directly
Non-reply DMs were silently ignored. Now sends a message directing the
user to /ask or @mention in a server channel.
* fix: address round 2 review feedback
- Show error items (policy-denied) in ToolInfoEvent unconditionally
- Match ToolResultEvent to ToolInfoEvent by call_id (deterministic),
fall back to name-based FIFO when call_id is absent
- Escape triple backticks before truncating in format_tool_result so
the 500-char limit holds after expansion
* fix: edit thinking message in-place instead of delete-and-recreate
ThinkingStopEvent now preserves the message for the next event to reuse.
ContentEvent seeds StreamingMessage with the thinking message so the
first flush edits it. ToolInfoEvent edits the thinking message into the
first tool embed. Eliminates the visible delete → gap → new message
flicker during thinking → tool call transitions.
* feat: separate tool call and result into distinct Discord messages
ToolInfoEvent sends a "running" embed (light grey, tool name + preview).
ToolResultEvent marks it "Done"/"Error" (color + title update) and sends
the result as a separate message. This gives clear lifecycle tracing in
chat-style threads where verbosity aids readability.
* fix: show running embed for all tools and remove redundant name prefix
ToolInfoEvent now shows a running embed for every tool regardless of
needs_approval — the running indicator and approval dialog serve
different purposes. Removes the needs_approval/auto_approve filter
that caused missing running embeds when tools were approved via
"Always Approve" or server-side auto-approve.
Also drops the redundant **name** prefix from format_tool_result since
the embed title already carries the tool name.
* fix: concise logging for SSE connection failures
Catch httpx.ConnectError/ConnectTimeout separately from the generic
exception handler. Logs url and error string instead of the full
httpx/httpcore stack trace, which is noise for expected transient
connection failures during node restarts.
* fix: address round 3 review feedback
- Pop _pending_approval_msgs on button click so ApprovalResolvedEvent
doesn't double-update the embed title (e.g. "Approved - Approved")
- Remove unused name/is_error params from format_tool_result — embed
title carries the name, embed color carries the error status
- Fix _disable_buttons docstring to mention title update
- Fix missing /v1 prefix on SSE endpoint URL (caused all SSE connections
to get text/plain 404 responses instead of event streams)
- Stop treating StreamEndEvent as session-terminal (it fires per-segment,
not per-workstream) so multi-turn conversations work in Discord
- Bail on 404 instead of retrying forever for gone workstreams, and clean
up stale routes from storage
- Check response status before iterating SSE events to avoid retrying
non-retryable upstream errors
- Default rebalancer.enabled to True so hash ring routing works without
manual ConfigStore setup
- Add one-shot cache refresh fallback on route endpoints to handle the
startup race between rebalancer and first routed request
- Remove stale "message queues" language from tagline
- Remove duplicated content covered by docs (governance details,
judge config, config.toml reference, health/rate-limit details,
monitoring metrics, tool table, multi-model config)
- Replace tool table with summary + link to docs/tools.md
- Add documentation index table linking to all doc pages
- Add architecture summary (single-node vs multi-node routing)
- Add component table for entry points
- Trim diagram table to most useful subset
- Consolidate quickstart section
README is now a concise landing page that directs to docs for
details, not a duplicated reference manual.
When TURNSTONE_ADVERTISE_URL is set (Docker deployments), the
_advertise_host variable was never assigned. The TLS upgrade path
tried to use it to construct the https:// URL, causing an
UnboundLocalError that made TLS init fail silently.
Fix: derive the TLS URL from _advertise_url (replace http → https)
instead of reconstructing from _advertise_host.
Address Copilot PR feedback:
- Exit with clear error if neither console_url nor server_url is
available after discovery (prevents cryptic failures downstream)
- Fix log field names: console → console_url, server → server_url
for consistency with other channel log events
Add token_factory parameter to SDK clients (_BaseClient, server,
console) — a Callable[[], str] invoked before each request to get
the current auth token. Supports ServiceTokenManager for auto-rotating
JWTs that re-mint transparently before expiry.
Channel gateway creates dual token managers:
- console-audience JWT for routing proxy calls (via AsyncTurnstoneConsole)
- server-audience JWT for direct SSE connections to server nodes
Both _request() and _stream_sse() inject the factory header per-call,
so long-lived connections get fresh tokens on reconnect.
Also adds TURNSTONE_CONSOLE_URL to console compose service for
DNS-resolvable service discovery.
- Default --server-url is now empty (not localhost:8080) to avoid
unreachable fallback URLs inside Docker containers
- Auto-discovery retries for up to 30s waiting for console or server
to register in the services table (handles startup ordering)
- Log discovery progress (discovering, discovered_console, discovered_server)
and warn on timeout or failure
- Wrap discovery in try/except so storage init failures don't crash startup
- Default --server-url is now empty (not localhost:8080) to avoid
unreachable fallback URLs inside Docker containers
- Auto-discovery retries for up to 30s waiting for console or server
to register in the services table (handles startup ordering)
- Both console_url and server_url are discovered from DB when not
explicitly set via CLI flags or env vars
- 404 retry: use blocking lock acquire so retry waits for cache refresh
to complete instead of skipping on contention
- 404 retry: surface httpx.HTTPError as 502 instead of suppressing it
and returning the original 404
- channel router: pass auto_approve_tools to create_workstream calls
(was silently dropped for console-routed creates)
- api-reference.md: document all /v1/api/route/* console routing proxy
endpoints and console /metrics
- router.route(): validate ws_id length and hex format before bucket
extraction, raise NoAvailableNodeError instead of ValueError
- router: expose version as public property, collector uses it instead
of accessing _version directly
- memory.py: deduplicate _bucket_of with canonical bucket_of from
hash_ring module
- architecture SVG: reroute direct/SSE lines below console to avoid
crossing over the console box
The collector's _apply_poll only detected workstream additions and
removals (set diff on ws_ids). State changes within existing
workstreams (idle → running, running → attention, etc.) were not
emitted to the SSE stream, so the dashboard only updated on manual
page refresh.
Now _apply_poll compares state and name fields between old and new
poll snapshots and emits ws_state and ws_rename events for any
changes. These flow through _fanout to the browser SSE stream,
giving real-time dashboard updates without page refresh.
Bug: Server nodes registered with container ID hostnames (e.g.,
http://a236323a92f6:8080) which aren't DNS-resolvable by other
containers. The console collector failed to poll nodes, causing
stale health/error status on the dashboard.
Fix: Add TURNSTONE_ADVERTISE_URL env var support. In compose, each
server sets it to the Docker service name (http://server-1:8080 etc).
Falls back to socket.getfqdn() when not set.
Also: remove the 100-node stress cluster (ddgStressCluster profile)
from compose.yaml. It was 720 lines of boilerplate from the old
simulator era. The simulator is being rebuilt separately (task #5).
Compose goes from 1028 to 304 lines.
Bug 1: Server's create_workstream handler ignored initial_message from
the request body. The old bridge sent it as a follow-up SendMessage
via Redis, but with direct HTTP nobody was sending it. Now the server
spawns a worker thread to send the initial message after creation,
matching the bridge's behavior.
Bug 2: Channel gateway compose config used --server-url=http://server:8080
which doesn't exist in cluster/ddgCluster profiles. Removed the hardcoded
URL — the channel gateway auto-discovers the console from the services
table via shared PostgreSQL. Added TURNSTONE_DB_URL and auth token to
the channel environment so DB-based service discovery works.
Server SDK create_workstream: add initial_message, auto_approve_tools,
user_id, ws_id params (all optional, omitted when empty).
Console SDK: add auto_approve, auto_approve_tools, user_id to
create_workstream. Add 8 route_* methods for the routing proxy path
(/api/route/*): route_create_workstream, route_send, route_approve,
route_plan_feedback, route_close, route_cancel, route_command,
route_lookup. Sync mirrors for all.
Prepares for channel gateway and scheduler to use SDK clients instead
of raw httpx calls.
Move the consistent hash ring implementation (FNV-1a, virtual nodes,
bisect lookup) from code to docs/design/consistent-hash-ring.md as a
forward-looking reference for future scalability work.
The current rebalancer uses weight-proportional distribution (simpler,
exact splits, no hash variance). The ring algorithm is documented with
test vectors, stability properties, and a comparison table for when
the ring approach becomes advantageous (large clusters, decentralized
routing, cross-language determinism).
hash_ring.py retains: RING_SIZE, bucket_of(), RingNode, NoAvailableNodeError
(all actively used by router and rebalancer).
Replace the full-rehash algorithm (diff ideal vs current across all
65536 buckets) with a donor/recipient algorithm that only moves
buckets from overloaded nodes to underloaded nodes.
Key improvements:
- Adding node C to {A, B} only moves buckets TO C, never between
A and B. Previously the HashRing rehash could shuffle between
existing nodes.
- Seeding uses weight-proportional distribution instead of HashRing
virtual nodes, producing an exact split that doesn't trigger
immediate correction on the next cycle.
- Dead-node buckets are redistributed to the most underloaded
survivors, not rehashed across the whole ring.
- HashRing class is no longer used by the rebalancer (still
available for other uses like the Go rewrite reference).
The threshold check still gates live-to-live moves. Dead-node
recovery remains unconditional.
set_bucket_stat: single-upsert storage method replacing the N-loop
reconciliation in the rebalancer. Reduces DB round-trips from
|ws_delta| per bucket to exactly 1.
Console metrics: /metrics endpoint on the console exposing 6 routing
and ring metrics in Prometheus text format:
- turnstone_router_requests_total (method, status)
- turnstone_router_request_duration_seconds (method)
- turnstone_ring_membership_size
- turnstone_ring_version
- turnstone_ring_rebalance_total (result)
- turnstone_ring_migrations_total
Instrumented in route_create, route_proxy, route_lookup handlers.
Ring gauges updated on collector discovery loop. Rebalance/migration
counters recorded after each rebalancer pass.
When rebalancer.eager_migrate is enabled, the rebalancer POSTs
/_internal/migrate to source nodes after reassigning buckets,
triggering immediate workstream eviction instead of waiting for
lazy resume on the next request.
Only idle workstreams are eagerly migrated — active ones (running,
thinking, attention) are left alone to avoid disrupting in-flight
work. Failed migrations are logged and skipped (the lazy path
handles them eventually).
Rebalancer: daemon thread in the console process that maintains
bucket-to-node assignments in hash_ring_buckets. Seeds the ring on
first run (empty table → 65536 rows via consistent hash). Periodically
checks for membership changes and rebalances: moves cheapest buckets
first (empty > idle > active), respects imbalance threshold, reconciles
bucket_stats against actual workstream counts before each pass.
Uses DB-based leader election (rebalancer_lock in system_settings) for
multi-console deployments. Increments rebalancer_version after writes
so console routers refresh their caches.
Add 6 settings: ring.vnodes_per_unit, rebalancer.enabled/interval/
threshold/eager_migrate, node.weight.
Add /_internal/migrate endpoint on server for eager workstream eviction.
Add routing proxy endpoints to the console server:
- POST /v1/api/route/workstreams/new — hash-ring-routed create with
503 retry, target_node pinning, and node_url injection
- POST /v1/api/route/{send,approve,cancel,command,close} — generic
proxy to workstream owner via O(1) bucket lookup
- GET /v1/api/route?ws_id=X — node URL lookup for direct SSE
Wire ConsoleRouter into console lifespan (cache refresh on startup)
and collector discovery loop (version-based cache invalidation).
Add --console-url to channel gateway CLI for multi-node routing.
ChannelRouter routes control-plane through console when set, SSE
connections go direct to server nodes via node_url from create response.
HashRing: FNV-1a virtual nodes, immutable, computes ideal bucket-to-node
distribution. Used by the rebalancer (next commit) to seed and maintain
the assignment table.
ConsoleRouter: in-memory flat array of 65536 NodeRef entries loaded from
hash_ring_buckets table. O(1) routing via ws_id prefix. Supports
per-workstream overrides, version-based cache refresh, and targeted
ws_id generation.
Both are pure library code with no server integration yet.
Delete the entire turnstone/mq/ package (broker, bridge, protocol,
client) and turnstone/sim/ package. Remove Redis as a dependency.
Channel gateway and console now communicate with server nodes via
direct HTTP (httpx + httpx-sse) instead of Redis pub/sub and queues.
Single-node deployments work with zero infrastructure beyond the
database.
Key changes:
- Channel adapters use httpx POST for create/send/approve/close
and httpx-sse for per-workstream event streaming
- Console collector discovers nodes via services table instead of
Redis SCAN
- Console scheduler dispatches tasks via HTTP POST with DB-based
leader election
- Server registers in services table with 30s heartbeat
- Server accepts optional ws_id in create request (for Phase 2
console-generated routing)
- SDK events gain IntentVerdictEvent and OutputWarningEvent types
- All docs, examples, bootstrap wizard updated
63 files changed, -5968 net lines (Redis transport fully removed)
* fix: sync actual TLS state to ConfigStore on console startup
The console writes tls.enabled to the DB but never clears it when TLS
init fails or isn't configured. Server nodes read the stale DB value
and attempt TLS negotiation with a non-TLS console, producing noisy
SSL errors on every startup.
Console now syncs the actual TLS state after init: if TLS succeeded,
tls.enabled=true; if it failed or wasn't attempted, tls.enabled=false.
Server TLS failure log reduced from full traceback to one-line warning.
* fix: sync TLS state to ConfigStore on console startup
Console now writes the definitive TLS state to ConfigStore so server
nodes don't attempt TLS against a non-TLS console:
- TLS init succeeded → write true
- TLS not configured (DB false/unset) → write false (definitive)
- TLS configured (DB true) but init failed → don't overwrite
(transient failure shouldn't permanently disable)
Server TLS warning reduced to one line with exception type, full
traceback available at debug level.
* feat: auto-detect model changes when LLM backend swaps models
The BackendHealthMonitor already probes /v1/models every 30s but
discarded the response. Now compares the detected model against the
last known one and triggers a registry reload when it changes.
- Extract _extract_context_window() helper for reuse across
detect_model, probe_model_endpoint, and the health monitor
- BackendHealthMonitor: new provider/initial_model/on_model_changed
params; _check_model_change() fires callback on model swap
- Server: wire _handle_model_change callback that updates cli_model_args
and calls registry.reload(); guarded by _user_specified_model flag
so --model overrides are never auto-replaced
- Session: _refresh_model_from_registry() called at top of send();
two string compares when nothing changed, full re-resolve on swap
- 7 new tests for _extract_context_window and model change detection
* fix: address Copilot review on model re-detection
- server: update cli_model_args only after successful reload (not
before), add finally block for new_reg.shutdown(), guard against
cli_model_args not yet initialized
- session: wrap registry lookup in try/except for concurrent reload
race, reset judge on model change, recompute tool_truncation when
context_window changes in auto mode
Replace "You are an expert software engineer" with a grounded
narrative persona: a resident engineer on a focused team with real
tools, real code, and real consequences. Sets expectations about
boundaries, judgment calls, and working within constraints.
* fix: scope-filter memory list/search to current workstream and user
Unscoped memory(action='list') and memory(action='search') returned all
memories across all workstreams. Now applies the same 3-query pattern
(global + current workstream + current user) used by system prompt
injection.
* fix: validate user scope on memory search/list for unauthenticated sessions
Adds _validate_scope guard to search and list prepare paths, matching
save/get/delete. Prevents explicit scope='user' from returning all
user-scoped memories when session is unauthenticated.
* fix: update _get_visible_memories references to _list_visible_memories
* fix: defense-in-depth guard for empty scope_id on search/list
Copilot review: if scope is 'user' or 'workstream' with empty
scope_id, the storage query returns all memories in that scope
across all users/workstreams. The prepare step already validates
via _validate_scope, but add exec-level guard to reject scoped
queries with empty scope_id as defense-in-depth.
* fix: detect context window from vLLM max_model_len field
vLLM exposes the context window as max_model_len on the model object,
not meta.n_ctx_train (llama.cpp format). Both detect_model() and
probe_model_endpoint() now check max_model_len first, falling back
to meta.n_ctx_train for llama.cpp. Fixes 32768 fallback on vLLM
servers that report 262144+ token context windows.
* test: add vLLM max_model_len detection tests
Copilot review: new vLLM context window path had no test coverage.
Add tests for probe_model_endpoint (max_model_len detected, preferred
over meta.n_ctx_train) and detect_model (vLLM model object with
max_model_len).
The first message sent from Discord was silently dropped because the
cog delegated the initial message to the bridge via CreateWorkstream-
Message, but the bridge published response events to the per-workstream
Redis pub/sub channel before the Discord bot had subscribed to it.
Redis pub/sub is fire-and-forget — events with no subscribers are lost.
Fix: create the workstream with initial_message="" (no delegation),
subscribe to the per-workstream event channel, then send the message
through router.send_message() — the same path the second message
already uses successfully.
Applied to both @mention handler and /ask slash command.
* feat: per-workstream status bar above input
Move the global token counter and model name from the header into a
per-pane telemetry strip between messages and the text input. Each
workstream pane now independently shows model name, token usage with
context percentage, tool calls this turn, and turn count.
Backend: add _ws_turn_tool_calls counter (reset per user turn, emitted
in SSE status event alongside turn_count). MQ bridge forwards the new
fields. SDK and TypeScript types updated.
Frontend: build .ws-status-bar DOM in _createDOM, rewrite updateStatus
to target per-pane elements, update SSE connect/disconnect handlers.
Remove #model-name and #status-bar from global header. Restore console
#status-bar CSS in its own stylesheet.
Accessibility: aria-atomic, aria-labels on each field, warning symbols
(▲/⚠) at 80%/95% context for color-blind users, placeholder text
before first status event. Disconnect state uses 2px red border with
dimmed stale fields.
* fix: emit status event on SSE connect so status bar populates on resume
When resuming a workstream, the event_generator only sent connected +
history events. The status bar stayed at placeholder values until the
next LLM response. Now replays session._last_usage as a synthetic
status event right after connected, so token count, tool calls, and
turn count render immediately.
* fix: address Copilot review — remove dead function, clarify locals
Remove updateHeaderForFocusedPane() and its call site (no-op since
status moved per-pane). Rename ambiguous ttc/tc locals to
turn_tool_calls/turn_count in the status replay block.
* feat: add memory get action, reduce search/list preview to 200 chars
search and list truncated memory content to 500 chars with no way to
read the full value. Two changes:
- New 'get' action retrieves a single memory by name with complete
untruncated content. Searches scopes narrowest-first (workstream
→ user → global).
- search/list previews reduced from 500 to 200 chars now that get
exists for full content. Both append a hint:
"Use memory(action='get', name='...') for full content."
Includes get_structured_memory_by_name wrapper in memory.py and
4 tests.
* Update turnstone/tools/memory.json
* fix: include 'get' in _prepare_memory docstring and invalid-action error
PostgreSQL text fields cannot store NUL (0x00) bytes, and SQLite
stores them but they cause downstream issues (API payloads, web UI).
Add sanitize_text() to _utils.py and apply it in both backends'
save_message to content and provider_data fields.
* fix: drop orphaned tool_results with no matching tool_use in _convert_messages
The context window increase from 200K to 1M for Claude 4.6 means
conversations that previously triggered auto-compaction now send their
full history. Older messages with orphaned tool_results (from
pre-fix cancels or compaction boundaries) are now visible to the API,
causing "unexpected tool_use_id in tool_result blocks" errors.
The existing repair code handles orphaned tool_use (synthesizes
missing results), but not the reverse. Now validates each
tool_result against the preceding assistant message's tool_use IDs
and silently drops results with no match.
* fix: filter empty IDs from prev_tool_use_ids, document pass-through
Code review: empty-ID tool_use blocks were added to the filter set,
and the intentional pass-through when prev_tool_use_ids is empty
needed documentation.
* fix: block math sandbox escape via getattr/setattr/type reflection
getattr() with runtime-constructed strings bypassed the AST validator,
allowing full os/subprocess access from the sandboxed math tool via
module.__builtins__['__import__']('os').
Three-layer fix:
- Block getattr, setattr, delattr, type, __import__ in
_MATH_BLOCKED_BUILTINS (prevents direct calls)
- Add AST validation for getattr/setattr/delattr call nodes
(catches them even if builtins dict is bypassed)
- Strip __builtins__ from all pre-imported modules in the execution
namespace (runtime defense — even if AST is somehow bypassed,
module.__builtins__ returns empty dict)
Normal math, sympy, numpy, scipy operations unaffected.
* fix: harden _safe_import to strip __builtins__ from runtime imports
Copilot review: modules imported at runtime via _safe_import still
had their original __builtins__ dict, accessible via
operator.attrgetter('__builtins__'). Now _safe_import strips
__builtins__ from every module it returns. Also blocks
operator.attrgetter/itemgetter at the AST level, and removes the
redundant duplicate getattr check in visit_Call.
* fix: add type ignore for module __builtins__ assignment
* fix: block /proc/*/environ access in bash filter and judge heuristic
/proc/1/environ leaks the full server environment including DB
credentials, API keys, and JWT secrets. Env scrubbing in env.py
only affects subprocess calls, not procfs reads.
- Add /proc/1/environ and /proc/self/environ to BLOCKED_PATTERNS
in safety.py (hard block)
- Add proc-environ-exfil heuristic rule at critical severity with
deny recommendation (catches /proc/<pid>/environ patterns)
* fix: move proc-environ-exfil rule to _CRITICAL_RULES list
Copilot review: rule had risk_level=critical but was placed in
_HIGH_RULES. Move to _CRITICAL_RULES for consistency with the
first-match-wins severity ordering.
The intent judge was receiving up to 50% of the context window in
conversation history (FIFO from end), which grows linearly with
conversation length and causes increasing latency. The judge only
needs the immediate request context to evaluate a tool call's safety.
Now trims to messages from the last user message onward before
applying the FIFO budget cap. Keeps the user's request, the
assistant's response with tool calls, and any recent tool results
while discarding earlier conversation that isn't relevant to the
current intent evaluation.
Claude 4.6 (Opus + Sonnet) unified on 1M token context windows.
Update capabilities table from 200K to 1M for both models. Remove
claude-opus-4 and claude-sonnet-4 entries (end of life). 4.5 models
remain at 200K. Default fallback stays at 200K for unknown models.
* fix: distinguish user cancel from crash in bash tool results
When a user cancels a running bash command, the process is killed
with SIGKILL (exit code -9). Previously this showed as an error,
causing the model to retry. Now checks cancel.is_set() after proc
exit and returns "Cancelled by user." as a non-error result so the
model knows to stop rather than retry.
* fix: use -signal.SIGKILL instead of magic -9
Copilot review: replace hard-coded -9 with -signal.SIGKILL for
clarity. Popen.returncode is negative of signal number when killed.
* feat: add stop_on_error param to bash tool for set -e behavior
New boolean parameter enables 'set -e' in the bash preamble so
multi-step scripts exit on the first command failure instead of
silently continuing. Default false (existing behavior preserved).
pipefail remains always-on.
* fix: strict bool parsing for stop_on_error, treat exit 1 as error with set -e
Copilot review: bool("false") is True — use `is True` for strict
JSON boolean parsing. Also, with stop_on_error enabled, any non-zero
exit code is now treated as an error (set -e means the script halted
on failure), whereas without it exit code 1 remains benign.
* fix: synthesize cancelled tool results instead of stripping turns
When a user cancels during tool execution, the model previously lost
all context about what was attempted (assistant message + tool_calls
stripped entirely). Now synthesizes tool_result messages with
is_error=true and "Cancelled by user." content for any tool_calls
that lack matching results. This keeps the conversation valid for
both providers while preserving the full tool call structure so the
model knows what was tried.
Also applies to KeyboardInterrupt with "Interrupted by user." text.
* fix: persist synthesized cancel results to DB, assert is_error in test
Copilot review: synthesized tool messages were in-memory only,
creating a mismatch with DB that could break rewind/retry. Now
calls save_message() for each synthesized result. Also adds
is_error=True assertion to the cancel test.
* feat: add pagination and longer content to recall tool
- New offset parameter for paginating through recall results
- Content preview increased from 500 to 2000 chars per match with
total length indicator when truncated
- Output passed through _truncate_output for consistency
- OFFSET clause added to SQLite (FTS5 + LIKE) and PostgreSQL
(tsvector + ILIKE) search queries
* fix: defensive int coercion for recall offset/limit
Copilot review: offset/limit could arrive as null, float, or other
non-int types from JSON. Coerce with int() + try/except in prepare,
and int() at the storage layer before binding into SQL OFFSET/LIMIT.
* feat: add diff_file tool for comparing files and content
New read-only tool that shows unified diffs between two files or
between a file and provided content. Useful for verifying edit_file
changes and comparing file versions. Auto-approved (no side effects).
Configurable context lines (default 3). Available to task agents.
* refactor: extract _read_text_lines helper, share across read_file and diff_file
Copilot review: diff_file duplicated file-loading and lacked binary
detection. Extract _read_text_lines() that handles realpath
resolution, null-byte binary detection, and error handling. Used by
both _exec_read_file and _exec_diff for consistent behavior.
* fix: address code review — agent flag, resolved shadowing, read_files
- Add agent: true to diff_file schema so plan agents can use it
- Fix resolved variable shadowing in _exec_read_file (use _ for
unused return from _read_text_lines)
- Register diffed files in _read_files so edit_file read guard
is satisfied after diff_file
- Move difflib import to module level (stdlib, no lazy-load needed)
- Fix description wording ("provided string" not "previous version")
* fix: stream diff with early cutoff, expand paths before header
- Stream difflib output and stop collecting after tool_truncation
chars to avoid large intermediate allocations on big diffs
- Expand paths with expanduser before building the approval header
so display matches actual execution paths
* docs: tool descriptions, bash timeout param, multi-line preview
- task_agent/plan_agent: document the tool subset limitation (no
memory, recall, watch, skill, or further delegation)
- bash: add per-call timeout parameter (1-600s, defaults to 120s),
shown in approval header when specified
- bash: show full command in preview for multi-line scripts so the
approval flow displays the complete command, not just the first line
- bash: document 256KB output cap and stderr prefix in description
* fix: address Copilot review on tool descriptions
- bash: say "truncated" not "256KB" (limit is configurable), document
timeout clamping range (1-600) and global fallback
- bash preview: fix "1 more lines" → "1 more line" singular
- plan_agent: remove bash from listed tools (not in AGENT_TOOLS)
* fix: improve memory save error message, narrow dd command filter
Two minor fixes from harness shakedown:
- memory save: split "both name and content required" into separate
errors for missing name vs empty content
- bash safety: replace blanket "dd if=" block with targeted patterns
for writes to block devices (of=/dev/sd*, /dev/nvme*, /dev/disk/,
etc.) and redirects to the same. Legitimate dd use like generating
test data or benchmarking reads is no longer blocked.
* fix: generalize > /dev/sda redirect pattern to > /dev/sd
Copilot review: only /dev/sda was blocked for redirects while
/dev/sdb, /dev/sdc etc were not. Generalize to match any /dev/sd*
device, consistent with the of= patterns.
* feat: edit_file replace_all, write_file append mode, search match count
Three tool enhancements from harness shakedown feedback:
- edit_file: new replace_all parameter replaces all occurrences of
old_string instead of requiring a unique match. Cannot combine with
near_line or edits array.
- write_file: new mode parameter with "append" option. Appends
content to end of file instead of truncating.
- search: output now includes a summary footer showing total match
count and file count (e.g. "47 matches across 12 files").
* fix: address Copilot review on tool enhancements
- replace_all: skip multi-occurrence rejection in pre-validation so
the feature actually works; show occurrence count in preview
- write_file mode: coerce non-string types safely via str()
- search footer: append before truncation to respect output limits
- edit_file error: mention replace_all as alternative to near_line
read_file silently converted null bytes to spaces, showing corrupted
content with no warning. Now samples the first 8KB for null bytes and
returns a clear error directing the user to bash for binary inspection.
* fix: memory delete searches all scopes when scope not specified
Previously delete defaulted to scope=global, so deleting a
workstream-scoped memory without explicitly passing scope=workstream
silently failed. Now tries narrowest scope first (workstream → user
→ global) and deletes the first match. Explicit scope still honored
when provided.
* fix: reject invalid scope on memory delete instead of silent fallback
Copilot review: invalid scope values were silently treated as
unspecified, which could cause accidental deletion from the wrong
scope. Now returns a clear error listing valid scopes.
* fix: exclude build/vendor/VCS directories from search tool
grep -rn recursed into .git, node_modules, target, __pycache__, etc.
producing hundreds of noise hits from generated content. Add
--exclude-dir flags for common directories that should never appear
in search results.
* fix: glob egg-info pattern and add vendor exclude
Copilot review: .egg-info misses turnstone.egg-info (named dirs),
use *.egg-info glob. Also add vendor to the exclude list.
Agent workflows need git for version control, curl for raw HTTP
requests, jq for JSON processing, and man/info for documentation
lookup. All were missing from the slim base image, leaving the man
tool non-functional and standard dev workflows broken.
* fix: block IPv6 loopback/link-local/private in SSRF filter
check_ssrf used gethostbyname which only resolves IPv4. IPv6 addresses
like ::1, fe80::, fd00:: bypassed the filter entirely. Switch to
getaddrinfo which resolves both address families and check all results.
* fix: handle IPv4-mapped IPv6 and zone IDs in SSRF filter
Copilot review caught two bypasses: ::ffff:127.0.0.1 (IPv4-mapped
IPv6) wasn't normalized before private/loopback checks, and fe80::1%lo0
(zone ID suffix) caused a ValueError that was silently swallowed.
Now normalizes IPv4-mapped addresses and strips zone IDs before parsing.
* fix: resolve symlinks before file I/O to prevent path-based bypass
write_file and edit_file followed symlinks silently — a symlink at
/data/link → /etc/passwd would show the /data path in the approval
header while writing to the real target. Three changes:
- open() calls in _exec_write_file, _exec_edit_file, _exec_read_file
now use the resolved (realpath) path instead of the raw symlink
- Approval headers show both paths when a symlink is detected
(e.g. "⚙ write_file: /data/link → /etc/passwd")
- Judge _get_arg_text includes the resolved path so heuristic rules
like write-system-path fire even through symlinks
* fix: address Copilot review — expanduser in fallback, pre-read, image paths
- edit_file exec fallback: add expanduser before realpath (tilde bypass)
- judge _get_arg_text: compare resolved against abspath(expanduser(path))
so ~/ paths don't false-positive as symlinks
- edit_file pre-read: use resolved path instead of raw symlink path
- _exec_read_image: use resolved path for getsize and binary open
* fix: clear dedup sigs after write tools to avoid false repeat warnings
The read→edit→read workflow triggered "identical repeat" warnings
because the dedup tracker compared (tool_name, args) without
considering intervening state changes. Now clears the signature set
when write_file, edit_file, or bash executes successfully, so
subsequent reads of the same file are not flagged.
* fix: use shared error prefixes for write-success detection in dedup
Copilot review: the error detection for write tools only checked
"Error" prefix, missing "Command timed out", "Blocked:", "Denied",
etc. Now shares the same _error_prefixes tuple used by the repeat
detection below, ensuring consistent classification.
The judge pre-converted tool schemas via convert_tools() before
passing them to create_completion(), which internally calls
convert_tools() again. The second conversion tried to extract
function.name from already-converted Anthropic-format tools,
producing empty tool names that the API rejected with
"tools.0.custom.name: String should have at least 1 character".
Fix: pass raw OpenAI-format schemas directly — create_completion
handles the provider-specific conversion.
* feat: add /retry and /rewind commands for conversation history navigation
Allow users to re-send the last message for a new response (/retry) or
drop the last N turns to restore an earlier conversation state (/rewind N).
Both operations sync in-memory state with the persistent database.
Server path includes conversation.modify permission gate, audit trail
(conversation.rewind / conversation.retry events), and thread-safe retry
dispatch. Migration 029 grants the permission to admin and operator roles.
* feat: add message action controls for retry, edit, and rewind in web UI
Hover toolbar on messages with CSS-only icons matching instrument panel
aesthetic. User messages get edit (pencil) and rewind (chevrons) buttons;
last assistant message gets retry (circular arrow). Edit flow uses
event-driven coordination — rewind completes via SSE history event before
send fires. Includes ARIA labels, keyboard nav, touch device support,
reduced motion, and busy-state gating.
Addresses Copilot review feedback on #219:
1. Anthropic _convert_messages: collect tool_use IDs in order (list
not set), filter empty IDs, defer synthetic results until after
real tool results so _merge_consecutive produces correct ordering.
2. Universal repair in reconstruct_messages: synthesize tool results
for mid-conversation orphaned tool calls on DB load. Benefits all
providers (OpenAI is lenient today but may tighten).
3. Test improvements: assert on is_error flag instead of "cancelled"
substring, verify real-before-synthetic ordering in partial results.
When a cancel interrupts tool execution, the assistant message with
tool_use blocks is saved to DB before tools run, but
GenerationCancelled prevents tool results from being created. The
in-memory rollback removes the orphaned message, but the DB row
persists. On resume, Anthropic rejects the conversation with
"tool_use ids were found without tool_result blocks".
Fix: _convert_messages now peeks ahead after each assistant message
with tool_use blocks. If any tool_use IDs lack matching tool_result
messages, synthetic error results are injected (is_error: true,
"Tool execution was cancelled."). Transparent to all callers,
provider-specific (OpenAI is lenient about this).
5 new tests covering: single orphan, multiple orphans, partial
results, complete results (no synthesis), and trailing orphan.
plan_agent and task_agent fail on Anthropic models with "Streaming is
required for operations that may take longer than 10 minutes" from
the SDK. The non-streaming create_completion path used
client.messages.create() which the SDK rejects for thinking-enabled
models.
Fix: use client.messages.stream() internally and call
get_final_message() to get the same Message object. Transparent to
all callers — fixes sub-agents, title generation, summarization,
web fetch, and judge create_completion calls.
Five improvements from Opus self-evaluation of the turnstone harness:
1. Batch edit_file: edits array parameter for atomic multi-edit in a
single tool call. Overlap detection, reverse-order application,
mutual exclusivity with single-edit params.
2. Sandbox packages: new [sandbox] extras group with sympy, numpy,
scipy, pytest — the sandbox already had graceful ImportError
fallbacks, now the packages are actually installed.
3. Stderr labeling: bash tool output prefixes stderr lines with
[stderr] so the model can distinguish errors from stdout.
4. JSON secret redaction: output guard now detects and redacts secrets
in JSON format ("api_key": "...", "password": "...", etc.) with
18 key patterns and 8-char minimum value length.
5. Model persisted on resume: workstream config now saves model and
model_alias. Resume restores the original model via registry
(same path as /model command), falling back to raw model name
if the alias is no longer available.
24 new tests (23 in test_edit_file.py, 1 in test_sessions.py).
stream_end fires per-segment (between tool calls), not per-turn.
The UI was using stream_end to transition to idle, causing a window
where the Send button appeared but the server worker thread was still
alive. User messages submitted during this window were silently
dropped. No Stop button was visible, so the user had no cancel path.
Root cause: state_change events (idle/thinking/running/error) were
only broadcast to the global SSE stream (console dashboard), never
to the per-workstream SSE that the browser UI listens to.
Fix: (1) server.py: on_state_change now also enqueues to the
per-workstream SSE listeners. (2) app.js: stream_end no longer
calls setBusy(false) — it only finalizes markdown rendering.
New state_change handler manages busy transitions: idle/error
set busy=false, thinking/running set busy=true.
* feat: model detect button, capabilities API, and model dropdowns
Admin Models tab: add Detect button that probes a model endpoint to
verify reachability, list available models, detect context_window, and
identify server type (llama.cpp/vLLM/SGLang/OpenAI/Anthropic). Add
static capability lookup endpoint for auto-filling form fields when
a known model name is entered. Add known-models endpoint for datalist
autocomplete suggestions.
Add "openai-compatible" as a third provider option for local servers,
keeping the OpenAI SDK under the hood but suppressing capability
auto-fill and known-model suggestions.
Replace free-text model input with a select dropdown in both console
and server new-workstream modals, populated from a new lightweight
GET /v1/api/models endpoint.
New endpoints:
- POST /v1/api/admin/model-definitions/detect
- GET /v1/api/admin/model-capabilities
- GET /v1/api/admin/model-capabilities/known
- GET /v1/api/models (both console and server)
* fix: accumulate signature_delta for Anthropic thinking blocks (#214)
The streaming path captured thinking_delta events but not
signature_delta, leaving the signature empty on round-trip and
causing 400 errors on multi-turn conversations with thinking enabled.
* fix: address PR review — empty base_url, capability leak, response schemas
- Don't pass empty base_url to OpenAI client (falls back to SDK default)
- Return None from lookup_model_capabilities for openai-compatible provider
- Only use static capability table for known models in _detect_openai_compat,
avoiding misleading 200k default for unknown local models
- Add AvailableModelInfo + ListAvailableModelsResponse schemas to both
console_spec and server_spec
- Regenerate TypeScript SDK OpenAPI snapshots
* fix: apply same known-model guard to Anthropic context_window detection
Only report context_window from the static capability table when the
Anthropic model is actually known, matching the OpenAI path fix.
* ui: add autocomplete hint to Model ID label in admin modal
* fix: use explicit kwargs for OpenAI() to satisfy strict mypy
The streaming path captured thinking_delta events but not
signature_delta, leaving the signature empty on round-trip and
causing 400 errors on multi-turn conversations with thinking enabled.
When a local model generates tool calls that are silently dropped
(truncation, missing tool-call-parser), there was zero server-side
logging — making it impossible to diagnose from docker compose logs.
- OpenAI provider: log request params and response summary (debug)
- Session: log stream completion, tool call presence (info), and
tool call discard with names when truncated (warning)
CLI unaffected — log level is WARNING there.
Add model_definitions table (migration 028) enabling model management
via the admin console without SSH access or server restarts. Models
defined in the database coexist with config.toml models through a
per-node merge strategy — config.toml overrides DB for the same alias,
DB-only models coexist alongside, no cross-node contamination.
Storage layer:
- model_definitions table with CRUD (SQLite + PostgreSQL)
- MODEL_DEFINITION_MUTABLE allowlist, admin.models permission
ModelRegistry integration:
- load_model_registry() merges DB + config.toml + CLI models
- context_window=0 auto-detects from provider capability table
or inherits CLI-detected value (same as config.toml behavior)
- ModelRegistry.reload() with validation, TOCTOU-safe accessors
- internal_model_reload + internal_model_status server endpoints
Admin API + UI:
- 6 console endpoints (list, create, get, update, delete, reload)
with admin.models permission, audit trail, provider validation
- Models tab in System group with sky blue (--blue) accent color
- Provider badges (openai/anthropic), source badges (config/db)
- Write-only API keys (never readable, "***" sentinel on update)
- Sync-pending indicator, mobile responsive, focus-trapped modal
Also changes is_secret settings from write-blocked (403) to write-only
across all settings, making judge.api_key configurable via admin UI.
* fix: watch dispatch error handler missing stream_end and state cleanup
The watch dispatch run() closure was missing GenerationCancelled
handling, stream_end emission, on_state_change calls, and the
worker_thread identity guard that the send_message path has. This
left the web UI in a stale state when watch-dispatched sends failed.
* fix: address review feedback - on_stream_end, put_nowait, ws._lock, tests
* fix: ruff lint (unused pytest import)
* fix: send_message() use on_stream_end() instead of raw _enqueue
* refactor: add is_error to on_tool_result protocol, remove text heuristics
Add is_error keyword arg to SessionUI.on_tool_result() so tools
report errors structurally. Server and JS client no longer guess
from output text prefixes — each tool sets the flag at the source.
Bash tool: exit code >= 2 is error, exit code 1 is ambiguous (grep
no-match). History reconstruction keeps text heuristic as fallback
for pre-migration data.
Update SDKs (Python + TypeScript), test mocks, docs, and diagrams.
* fix: infinite recursion in _report_tool_result, signal exits, stale docs
* fix: add _tool_error_flags to test_load_skill ChatSession stubs
Experimental multi-node AI orchestration platform. Deploy tool-using AI agents across a cluster of servers, driven by message queues or interactive interfaces.
Multi-node AI orchestration platform. Deploy tool-using AI agents across a cluster of servers with direct HTTP routing, interactive interfaces, and enterprise governance.
> **Beta — Use at your own risk.** Turnstone is under active development and has not reached a stable release. APIs, configuration formats, and database schemas may change between versions without migration paths. We make no guarantees of determinism, reliability, or backward compatibility. Evaluate thoroughly before deploying to any environment where these properties matter.
<p align="center">
<img src="docs/assets/hero.png" alt="Turnstone coordinator — parallel tool batches with judge-graded approval and child workstream tracking" width="960"/>
</p>
Named after the [Ruddy Turnstone](https://en.wikipedia.org/wiki/Ruddy_turnstone) (*Arenaria interpres*) — a shorebird that flips stones to discover what's hiding underneath.
| **Experimental** | `pip install turnstone --pre` | `ghcr.io/turnstonelabs/turnstone:experimental` | New features. May have rough edges. |
See [docs/releasing.md](docs/releasing.md) for the full release process.
## What it does
Turnstone gives LLMs tools — shell, files, search, web, planning — and orchestrates multi-turn conversations where the model investigates, acts, and reports. It runs as:
Turnstone gives LLMs tools — shell, files, search, web, planning — and orchestrates multi-turn conversations where the model investigates, acts, and reports.
- **Interactive sessions** — terminal CLI or browser UI with parallel workstreams
- **Queue-driven agents** — trigger workstreams via message queue, stream progress, approve or auto-approve tool use
- **Multi-node clusters** — generic work load-balances across nodes, directed work routes to a specific server
- **Cluster dashboard** — real-time view of all nodes and workstreams, reverse proxy for server UIs
- **Intent validation** — an LLM judge evaluates every tool call before approval, presenting risk assessments and evidence-based recommendations so users can make informed decisions instead of blindly approving raw tool calls
- **Cluster simulator** — test the stack at scale (up to 1000 nodes) without an LLM backend
Works with any OpenAI-compatible API (vLLM, llama.cpp, NVIDIA NIM) or Anthropic's native Messages API. Supports [MCP](https://modelcontextprotocol.io/) for external tool servers with native deferred tool loading on Anthropic and OpenAI APIs (BM25 fallback for local models).
- **Cluster dashboard** — real-time view of all nodes and workstreams with console routing proxy
- **Intent validation** — LLM judge evaluates every tool call with risk assessments and evidence
- **Multi-provider** — OpenAI-compatible APIs (vLLM, llama.cpp, NIM), Anthropic Messages API, and Google Gemini
- **MCP support** — external tool servers with native deferred loading (Anthropic/OpenAI) or BM25 fallback
<p align="center">
<img src="docs/diagrams/architecture-overview.svg" alt="Turnstone system architecture — data flow from clients through gateways, Redis MQ, cluster nodes, to LLM providers" width="960"/>
<img src="docs/diagrams/architecture-overview.svg" alt="Turnstone system architecture" width="960"/>
Then open `http://localhost:8090` for the cluster-wide dashboard. Create workstreams from the console and interact with any node's server UI through the built-in reverse proxy — no direct server port access required.
### Docker
```bash
cp .env.example .env # edit LLM_BASE_URL, OPENAI_API_KEY, etc.
docker compose up # starts redis + server + bridge + console (SQLite)
docker compose --profile production up
```
For production with PostgreSQL:
See [QUICKSTART.md](QUICKSTART.md) for the bootstrap wizard and [docs/docker.md](docs/docker.md) for Docker configuration and profiles.
```bash
# Requires POSTGRES_PASSWORD and DB_BACKEND=postgresql in .env (or exported)
docker compose --profile production up # adds PostgreSQL, uses it as database
Turnstone includes a built-in governance layer for enterprise deployments — manage who can do what, which tools run unattended, and where every token goes.
- **OIDC SSO** — single sign-on via any OpenID Connect provider (Okta, Azure AD, Google, Keycloak); Authorization Code Flow with PKCE, auto-provisioning, claim-based role mapping with demotion propagation; see [docs/oidc.md](docs/oidc.md)
- **Tool policies** — glob-pattern rules (`allow` / `deny` / `ask`) with priority ordering; automate approvals or lock down dangerous tools
- **Skills** — reusable behavioral profiles with system prompts, `{{variable}}` substitution, session config (model, temperature, token budget), install-time security scanning, version history, external discovery (skills.sh / GitHub), and runtime `skill` tool for model-driven skill activation
- **Usage tracking** — per-request token and tool metrics, aggregation by day / model / user, automatic 90-day pruning
- **Audit logging** — append-only event trail for all admin mutations, IP-aware, 365-day retention
All governance features are managed through the console admin panel (13 tabs) and the full REST API. Runtime settings (model, tools, rate limiting, health, judge, memory) are configurable via the admin Settings tab — no config file edits or restarts needed for most changes. See [docs/governance.md](docs/governance.md) for setup and [docs/settings.md](docs/settings.md) for the settings reference.
### Intent Validation (LLM Judge)
Every tool call that requires human approval is evaluated by an intent validation judge that provides a structured risk assessment alongside the approval prompt — so instead of "approve this bash command?", users see a verdict with risk level, confidence, recommendation, and reasoning.
The system uses a two-tier evaluation pipeline:
1.**Heuristic tier** (instant, free) — 36 pattern-based rules classify tool calls by severity. Catches destructive commands (`rm -rf /`, `DROP TABLE`), privilege escalation (`sudo`), credential access, supply chain risks, browser data export, cloud infrastructure mutations, and more. Results appear immediately.
2.**LLM judge tier** (async) — A full LLM evaluation runs in the background with access to `read_file` and `list_directory` for evidence gathering. The judge can inspect files that a write would overwrite, check directory contents before a delete, and cite specific evidence in its reasoning. Results update the UI progressively when ready.
The judge defaults to the same model as the session (self-consistency) but can be configured to use a separate model — useful when running a small local model for tasks but wanting a commercial model for safety evaluation.
```toml
[judge]
enabled=true# on by default
model=""# empty = same as session model
provider=""# empty = same as session provider
timeout=60.0# generous for local models
```
Verdicts are persisted for audit and exposed via Prometheus metrics (`turnstone_judge_verdicts_total`, `turnstone_judge_llm_latency_seconds`).
Skills are also scanned at install time — the scanner evaluates content, supply chain, vulnerability, and declared capability risk across four independent axes. Results populate `scan_status` (tier) and `scan_report` (structured JSON breakdown) on the skill record so administrators can assess risk before enabling a skill.
Tool execution results are evaluated by an output guard before entering the conversation — detecting prompt injection payloads in fetched content, credential leakage in command output, and encoded payloads. Detected credentials are automatically redacted.
See [docs/judge.md](docs/judge.md) for the full guide.
## Multi-node routing
Each Turnstone server runs a bridge process. Bridges share a Redis instance for coordination:
| Redis Key | Purpose |
|-----------|---------|
| `turnstone:inbound` | Shared work queue — generic tasks, any node |
| `turnstone:events:global` | Global event pub/sub |
| `turnstone:events:cluster` | Cluster-wide state changes (for turnstone-console) |
**Routing rules:**
1. Message has `target_node` → routes to that node's queue
2. Message has `ws_id` → looks up owner, routes to owning node
3. Neither → shared queue, next available bridge picks it up
Bridges BLPOP from their per-node queue (priority) then the shared queue. Directed work always takes precedence.
## Tools
15 built-in tools, 2 agent tools, plus external tools via MCP:
Built-in tools for shell, files, search, web, memory, notifications, and autonomous sub-agents — plus external tools via [MCP](https://modelcontextprotocol.io/) with native deferred loading. See [docs/tools.md](docs/tools.md) for the full reference and [docs/mcp-registry.md](docs/mcp-registry.md) for MCP configuration.
| Tool | Description | Auto-approved |
|------|-------------|:---:|
| `bash` | Execute shell commands | |
| `read_file` | Read file contents (text or images with vision models) | yes |
| `write_file` | Write/create files | |
| `edit_file` | Fuzzy-match file editing | |
| `search` | Search files by name/content | yes |
| `math` | Sandboxed Python evaluation | |
| `man` | Read man pages | yes |
| `web_fetch` | Fetch URL content | |
| `web_search` | Web search (provider-native or Tavily) | |
| `watch` | Periodic command polling with conditions | |
| `task` | Spawn autonomous sub-agent | |
| `plan` | Explore codebase, write .plan.md | |
| `mcp__*` | External tools from MCP servers | |
## Architecture
When the total tool count exceeds a configurable threshold (default 20), MCP tools are automatically deferred using native `defer_loading` on Anthropic and OpenAI APIs, or a transparent client-side BM25 search for local models. The LLM discovers deferred tools on demand via a `tool_search` capability — no configuration needed beyond `--tool-search auto` (the default).
**Single-node**: Client → Server (direct HTTP + SSE). No external dependencies beyond the database.
### MCP Tool Servers
**Multi-node**: Client → Console (rendezvous routing proxy) → Server nodes. The console picks the target node for each workstream via rendezvous (HRW) hashing over the live service registry — pure function of `(ws_id, live_nodes)`, no stored bucket state, deterministic across readers. A node join or drop only re-routes the keys that score highest on the affected node.
Turnstone supports the [Model Context Protocol](https://modelcontextprotocol.io/) (MCP) for connecting external tool servers. MCP tools are discovered at startup, converted to OpenAI function-calling format, and merged with built-in tools. Each MCP tool is prefixed with `mcp__{server}__{tool}` to avoid name collisions. Tool lists stay fresh via push notifications (`tools.listChanged`), periodic polling for servers without push, and manual `/mcp refresh`.
| Component | Purpose |
|-----------|---------|
| `turnstone` | Terminal CLI (REPL) |
| `turnstone-server` | Web UI + REST API + SSE events |
Use `/mcp` in the REPL to list connected tools, `/mcp refresh` to re-fetch tool lists from servers. MCP tools require user approval by default (overridden by `--skip-permissions` or UI auto-approve).
### Multi-Model and Multi-Provider Support
Turnstone supports multiple model backends per server instance, including different LLM providers. `ChatSession` delegates all API communication to pluggable `LLMProvider` adapters — the internal message format stays OpenAI-like, and each provider translates at the API boundary. Define named models in `config.toml` and select per-workstream or switch mid-session with `/model <alias>`.
```toml
[models.local]
base_url="http://localhost:8000/v1"
model="qwen3-32b"
# provider defaults to "openai" (works with vLLM, llama.cpp, etc.)
[models.claude]
provider="anthropic"
api_key="sk-ant-..."
model="claude-opus-4-6"
context_window=200000
[models.openai]
base_url="https://api.openai.com/v1"
api_key="sk-..."
model="gpt-5"
context_window=400000
[model]
default="local"# which model to use by default
fallback=["claude","openai"]# try these if the primary is unreachable
agent_model="claude"# optional: separate model for plan/task sub-agents
```
Supported providers: `"openai"` (default -- OpenAI, vLLM, llama.cpp, any OpenAI-compatible API) and `"anthropic"` (Anthropic Messages API, requires `pip install turnstone[anthropic]`).
Use `/model` to show available models, `/model claude` to switch. Workstreams created via the API accept an optional `model` parameter.
## Configuration
All entry points read `~/.config/turnstone/config.toml`. CLI flags override config values.
```toml
[api]
base_url="http://localhost:8000/v1"
api_key=""
# tavily_key = "" # only needed for local/vLLM models without native search
[model]
name=""# empty = auto-detect
temperature=0.5
reasoning_effort="medium"
default="default"# model alias for new workstreams
fallback=[]# ordered list of fallback model aliases
agent_model=""# model alias for plan/task sub-agents
[tools]
timeout=30
skip_permissions=false
search="auto"# "auto" (enable when >threshold tools), "on", "off"
search_threshold=20# min tools before tool search activates
search_max_results=5# max tools returned per search query
[server]
host="0.0.0.0"
port=8080
max_workstreams=50# auto-evicts oldest idle when full
[redis]
host="localhost"
port=6379
password=""
[bridge]
server_url="http://localhost:8080"
node_id=""# empty = hostname_xxxx
[console]
host="0.0.0.0"
port=8090
url="http://localhost:8090"# used by CLI /cluster commands
poll_interval=10
[health]
backend_probe_interval=30
backend_probe_timeout=5
circuit_breaker_threshold=5
circuit_breaker_cooldown=60
[ratelimit]
enabled=true
requests_per_second=10.0
burst=20
[database]
backend="sqlite"# "sqlite" (default) or "postgresql"
path=".turnstone.db"# SQLite file path (relative to working directory)
Parallel independent conversations, each with its own session and state:
| Symbol | State | Meaning |
|--------|-------|---------|
| `·` | idle | Waiting for input |
| `◌` | thinking | Model is generating |
| `▸` | running | Tool execution in progress |
| `◆` | attention | Waiting for approval |
| `✖` | error | Something went wrong |
Idle workstreams are automatically cleaned up after 2 hours (configurable). In multi-node deployments, workstream ownership is tracked in Redis — follow-up messages auto-route to the owning node.
-`turnstone_judge_enabled` — whether the intent validation judge is active (0/1)
Per-workstream metrics are labeled by `ws_id` (bounded by `[server].max_workstreams`).
### Health & Rate Limiting
**Health degradation.** A background `BackendHealthMonitor` probes the LLM backend every `backend_probe_interval` seconds. When the backend is unreachable, `/health` reports `"status": "degraded"` (HTTP 200) and the `turnstone_backend_up` gauge drops to 0.
**Circuit breaker.** After `circuit_breaker_threshold` consecutive probe failures the circuit opens (CLOSED -> OPEN). While open, `ChatSession._create_stream_with_retry` skips the backend entirely and returns an error. After `circuit_breaker_cooldown` seconds the circuit enters HALF_OPEN, allowing a single probe. A successful probe closes the circuit; a failure re-opens it.
**Per-IP rate limiting.** When `[ratelimit].enabled` is true, each client IP is tracked with a token-bucket limiter (`requests_per_second` / `burst`). Rate limiting is applied in `do_GET`/`do_POST` after authentication but before route dispatch. `/health` and `/metrics` are exempt. Requests that exceed the limit receive HTTP 429 with a `Retry-After` header.
**Workstream eviction.** When `WorkstreamManager.create()` would exceed `max_workstreams`, the oldest IDLE workstream is automatically evicted and the `turnstone_workstreams_evicted_total` counter is incremented. Configure via `[server].max_workstreams` (default 50).
- An OpenAI-compatible API endpoint ([vLLM](https://github.com/vllm-project/vllm), [NVIDIA NIM](https://build.nvidia.com/), [llama.cpp](https://github.com/ggml-org/llama.cpp), etc.) or an Anthropic API key
JWTs are the recommended credential for browser sessions. API tokens are suitable for programmatic access and CI/CD. Config tokens are a simple option for single-node deployments.
JWTs are the recommended credential for browser sessions. API tokens are suitable for programmatic access and CI/CD.
### `POST /v1/api/auth/login`
@@ -232,12 +229,12 @@ below.
---
### `GET /v1/api/events?ws_id=<id>`
### `GET /v1/api/workstreams/{ws_id}/events`
Opens a Server-Sent Events stream scoped to a single workstream. The connection
remains open indefinitely; the server pushes events as they occur.
@@ -284,6 +281,7 @@ Each message in the `messages` array has:
| `role` | string | `"user"`, `"assistant"`, or `"tool"` |
| `content` | string or null | Text content of the message |
| `tool_calls` | array or null | Present only on assistant messages with calls |
| `reasoning` | string (optional) | Concatenated reasoning / chain-of-thought text on assistant turns whose `provider_data` carried reasoning-bearing blocks (Anthropic `thinking`, OpenAI Responses `reasoning`, or synthetic `reasoning_text` from local-model servers). Present only when the active model's `surface_persisted_reasoning` flag is True. |
Each entry in `tool_calls`:
@@ -328,6 +326,44 @@ finalize any in-progress assistant message.
{"type":"stream_end"}
```
**`state_change`** -- the worker thread transitioned to a new state. Drives
the client's busy-mode (composer in send vs. stop, spinner indicators,
auto-focus on idle). Sent live during normal operation AND on every fresh
SSE subscribe (so a mid-stream page refresh restores the correct composer
state without waiting for the next live transition).
**`tool_result`** -- final output from a completed tool execution. The `call_id` matches the corresponding `tool_info`/`approve_request` item and any preceding `tool_output_chunk` events. For bash tools, this arrives after all streaming chunks and includes both stdout and stderr.
**`tool_result`** -- final output from a completed tool execution. The `call_id` matches the corresponding `tool_info`/`approve_request` item and any preceding `tool_output_chunk` events. For bash tools, this arrives after all streaming chunks and includes both stdout and stderr. The `is_error` field is `true` when the tool execution failed (e.g. bash exit code >= 2 or signal, file not found, timeout). Exit code 1 is ambiguous (e.g. `grep` no-match) and is not flagged. User denials are tracked separately via a `denied` flag. Clients should use `is_error` instead of text-prefix heuristics.
Creates a new workstream. The server supports up to 10 concurrent workstreams.
The endpoint accepts **either**`application/json` (legacy shape) **or**
`multipart/form-data` when you want to upload attachments at creation
time. Multipart requests carry one `meta` field containing the JSON body
shown below plus zero-or-more `file` parts; each file is validated and
reserved onto the new workstream's first turn before the dispatch worker
runs, so queued multimodal turns cannot lose files to racing sends. If
validation fails the fresh workstream is rolled back so no orphan rows
leak.
**Request body:**
```json
@@ -860,6 +926,7 @@ All fields are optional. The body can be empty or an empty JSON object.
| `auto_approve` | bool | false | Auto-approve all tool calls for this workstream |
| `resume_ws` | string | "" | Workstream ID to resume atomically during creation (empty = fresh)|
| `skill` | string | "" | Skill name. Applies content (system prompt), model, temperature, reasoning effort, max tokens, auto-approve policy, token budget, and other session config from the skill. Returns 400 if not found or disabled. Ignored when `resume_ws` is set (resumed sessions restore their own skill). |
| `judge_model` | string | "" | Optional model alias for the judge (overrides default judge model for this workstream) |
> **Skill behavior:** When `skill` is specified, the skill's content is injected as a system message and its session config fields (model, temperature, auto-approve, token budget, etc.) override system defaults for the new workstream.
@@ -886,20 +953,32 @@ Status code: `400`
---
### `POST /v1/api/workstreams/close`
### `POST /v1/api/workstreams/{ws_id}/close`
Closes and removes a workstream. The last remaining workstream cannot be
*.json 15 tool schemas (OpenAI function-calling format + turnstone metadata)
*.json 19 tool schemas (OpenAI function-calling format + turnstone metadata)
```
Both UIs share a common design system extracted into `turnstone/shared_static/`: design tokens, login overlay, toast notifications, theme toggle, keyboard shortcuts, and utility functions. Each UI imports `base.css` and the shared JS modules at `/shared/`, then adds only page-specific code at `/static/`.
@@ -231,18 +231,20 @@ The engine emits state changes via `_emit_state()` which calls
> See also: [Core Engine Classes diagram](diagrams/png/03-core-engine-classes.png)
Defined in `turnstone.core.session.SessionUI` as a `typing.Protocol` with 14
Defined in `turnstone.core.session.SessionUI` as a `typing.Protocol` with 16
methods. Every frontend must implement all of them.
`on_rename` is called by the `/name` command (on success) and after a successful `/resume` (if the resumed session has an alias or title). `WebUI.on_rename` broadcasts a `ws_rename` event on the global SSE channel and updates the in-memory `Workstream.name`; `TerminalUI.on_rename` is a no-op.
| `WebUI` | `turnstone.server` | SSE event queue per workstream, `threading.Event` for blocking on approval/plan |
| `WebUI` | `turnstone.server` | SSE event queue per workstream + global broadcast, `threading.Event` for blocking on approval/plan. `on_state_change` sends to both per-workstream and global SSE (the browser UI uses per-workstream `state_change` events to manage busy/idle transitions; `stream_end` only finalizes markdown rendering). |
`turnstone-console` is a cluster management service that provides cluster-wide visibility and control across all turnstone nodes. It connects to the shared Redis broker, discovers nodes via heartbeat keys, polls each node's HTTP API for workstream data, and subscribes to a cluster event channel for real-time state changes.
`turnstone-console` is a cluster management service that provides cluster-wide visibility and control across all turnstone nodes. It discovers nodes via the`services` database table and subscribes to each node's SSE event stream for real-time workstream, health, and metric updates.
The console also supports **workstream creation** (dispatched via MQ to target nodes) and a **reverse proxy** that serves each node's server UI through the console port — so users only need network access to the console, not to individual server nodes.
The console also supports **workstream creation** (dispatched via HTTP proxy to target nodes) and a **reverse proxy** that serves each node's server UI through the console port — so users only need network access to the console, not to individual server nodes.
## Architecture
> See also: [Console Data Flow diagram](diagrams/png/11-console-data-flow.png)
- **Inbound (monitoring):**Bridges publish state changes to `{prefix}:events:cluster` on Redis pub/sub. The console subscribes for real-time updates and periodically polls each node's `GET /v1/api/dashboard` for full workstream snapshots.
- **Outbound (control):** The console pushes `CreateWorkstreamMessage` to Redis inbound queues targeting specific nodes. Bridges pick up these messages and create workstreams on their local servers.
- **Inbound (monitoring):**The console discovers nodes via the `services` database table (nodes register on startup and send periodic heartbeats). It opens a persistent SSE connection to each node's `GET /v1/api/events/global` endpoint, receiving a full snapshot on connect followed by real-time delta events (state changes, health transitions, aggregate metrics).
- **Outbound (control):** The console proxies workstream creation requests to target nodes via HTTP.
- **Proxy (pass-through):** The console reverse-proxies each node's server UI at `/node/{node_id}/`, forwarding HTTP and SSE traffic so the browser never contacts server nodes directly.
| Node HTTP API | `GET/POST {server_url}/*` | Proxy | Server UI, API requests, SSE streams |
### Redis Key: Cluster Event Channel
Bridges publish to `{prefix}:events:cluster` whenever a workstream state change, creation, closure, or rename occurs. Events include `node_id` so the console can attribute them to the correct node.
| `ws_created` | ws_id, name, node_id | New workstream created |
| `ws_closed` | ws_id | Workstream closed |
| `ws_rename` | ws_id, name | Workstream renamed |
---
## ClusterCollector
The collector (`turnstone/console/collector.py`) maintains an in-memory snapshot of all nodes and workstreams. Three daemon threads handle data acquisition:
The collector (`turnstone/console/collector.py`) maintains an in-memory snapshot of all nodes and workstreams. Two daemon threads handle data acquisition:
1. **Event subscriber** — subscribes to`{prefix}:events:cluster` via `RedisBroker.subscribe_cluster()`. Applies state changes, creates, closes, and renames to the in-memory model immediately.
1. **Node discovery** — queries the`services` database table every 60 seconds. Adds newly discovered nodes, removes expired ones (stale heartbeats), emits `node_joined` / `node_lost` events to SSE listeners, and spawns/cancels SSE tasks for new/lost nodes.
2. **Node discovery** — scans heartbeat keys every 15 seconds via `broker.list_nodes()`. Adds newly discovered nodes, removes expired ones, emits `node_joined` / `node_lost` events to SSE listeners.
3. **Poll loop** — fetches `GET /v1/api/dashboard` and `GET /health` from each known node every 10 seconds. Uses `ThreadPoolExecutor(max_workers=50)` for parallelism. Each poll replaces the node's workstream list with the authoritative server data.
2. **SSE manager** — a single asyncio event loop on one thread multiplexes persistent SSE connections to all discovered nodes via `GET /v1/api/events/global`. Each connection receives a `node_snapshot` on connect (workstreams, health, aggregate) followed by real-time delta events (`ws_state`, `ws_created`, `ws_closed`, `ws_rename`, `health_changed`, `aggregate`). On disconnect, the node is marked unreachable and the connection is retried with exponential backoff (1s–30s). An `?expected_node_id=` query parameter provides identity verification against IP reuse (server returns 409 on mismatch).
A `get_snapshot()` method builds the full cluster state under a single lock acquisition — overview aggregates and per-node workstream lists in one atomic read. This is served both as a REST endpoint and as the initial SSE event on client connect.
@@ -70,7 +53,7 @@ All reads and writes to the node/workstream map are protected by a single `threa
### Scale Considerations
- **50,000 workstreams** (1,000 nodes × 50 per node) at ~500 bytes each = ~25 MB in memory
- **1,000 nodes**polled in parallel — fan-out concurrency is configurable via `cluster.node_fan_out_limit` (default 200), yielding 5 batches at ~100ms each = ~0.5 second poll cycle
- **1,000 nodes**connected via persistent SSE — a single asyncio event loop multiplexes all connections with negligible overhead. Ensure `ulimit -n` >= 4096 for fd headroom
- **Filtering and pagination** run in-memory on the full workstream list — sub-millisecond at this scale
- **SSE fan-out** uses per-client queues (2,000 events) — backed-up clients get events dropped, not blocking
- **Database** — for clusters sharing PostgreSQL, use [PgBouncer](pgbouncer.md) in transaction pooling mode
@@ -183,7 +166,7 @@ Full cluster state in a single response — all nodes with their workstreams plu
### `POST /v1/api/cluster/workstreams/new`
Create a new workstream on a target node. Dispatches a `CreateWorkstreamMessage` through the Redis MQ pipeline — the bridge on the target node picks it up and creates the workstream on the server. Requires `write` scope.
Create a new workstream on a target node. The console proxies the creation request to the target node's HTTP API. Requires `write` scope.
Request:
@@ -197,9 +180,9 @@ Request:
All fields are optional:
- `node_id` — targeting mode:
- **omitted or `"auto"`** — console picks the reachable node with the most available capacity (max_ws - ws_total) and pushes to its directed queue.
- **`"pool"`** — pushes to the shared inbound queue; the next available bridge picks it up (true general-pool dispatch).
- **specific node ID** — pushes to that node's directed queue.
- **omitted or `"auto"`** — console picks the reachable node with the most available capacity (max_ws - ws_total) and proxies the request to it.
- **`"pool"`** — console picks a reachable node with available capacity using round-robin selection.
- **specific node ID** — proxies the request to that node directly.
- `name` — workstream display name. Auto-generated if omitted.
- `model` — model alias from the target node's registry. Uses the node's default model if omitted.
@@ -213,7 +196,7 @@ Response:
}
```
Creation is asynchronous — the response confirms the MQ message was dispatched. A `ws_created` event on the cluster SSE stream confirms the workstream was actually created.
The response confirms the workstream creation request was proxied to the target node. A `ws_created` event on the cluster SSE stream confirms the workstream was actually created.
### `GET /v1/api/cluster/events`
@@ -351,7 +334,7 @@ The console reverse-proxies each node's server UI at `/node/{node_id}/`. This al
### URL Rewriting
The server UI uses root-relative URLs (`/v1/api/send`, `/static/app.js`, `/shared/base.css`, etc.). Since `<base>` tags cannot rewrite root-relative URLs, the console uses a JS shim approach:
The server UI uses root-relative URLs (`/v1/api/workstreams/{ws_id}/send`, `/static/app.js`, `/shared/base.css`, etc.). Since `<base>` tags cannot rewrite root-relative URLs, the console uses a JS shim approach:
1. **HTML rewriting** — when serving `index.html`, replaces `href=` and `src=` references to both `/static/` and `/shared/` with the proxy prefix (`/node/{node_id}/static/` and `/node/{node_id}/shared/` respectively).
@@ -361,7 +344,7 @@ The server UI uses root-relative URLs (`/v1/api/send`, `/static/app.js`, `/share
### SSE Proxy
SSE streams (`/v1/api/events`, `/v1/api/events/global`) are proxied as raw byte passthrough — the console opens an `httpx.AsyncClient.stream()` to the upstream server (with `read=None` and `pool=None` timeouts since SSE connections are long-lived) and relays every byte via `StreamingResponse`. This preserves server-side ping comments, event framing, and keepalives verbatim without parsing or re-encoding.
SSE streams (`/v1/api/workstreams/{ws_id}/events`, `/v1/api/events/global`) are proxied as raw byte passthrough — the console opens an `httpx.AsyncClient.stream()` to the upstream server (with `read=None` and `pool=None` timeouts since SSE connections are long-lived) and relays every byte via `StreamingResponse`. This preserves server-side ping comments, event framing, and keepalives verbatim without parsing or re-encoding.
Triggered by the "+ new" header button. A modal dialog with:
- **Node selector** — dropdown with three targeting modes: "Auto (best available)" picks the node with the most headroom, "General pool (any node)" pushes to the shared queue for any bridge to pick up, or a specific node from the list (showing capacity).
- **Node selector** — dropdown with three targeting modes: "Auto (best available)" picks the node with the most headroom, "General pool (any node)" picks a node with available capacity using round-robin, or a specific node from the list (showing capacity).
- **Profile** — optional dropdown listing enabled skills. Applies the skill's model, auto-approve policy, token budget, and other behavioral settings at creation time.
- **Name** — optional text input. Auto-generated if left empty.
- **Model** — optional text input for a model alias from the target node's registry.
- **Judge Model** — optional text input for the judge model alias (overrides the default judge model for this workstream).
Keyboard shortcuts: Ctrl+Shift+R (refresh title), Ctrl+Shift+E (edit title), Ctrl+Shift+F (fork), Ctrl+Shift+X (delete). Press ? for full shortcut help.
On submit, `POST /v1/api/cluster/workstreams/new` dispatches the creation request. A toast confirms success; the SSE stream delivers the `ws_created` event to update the dashboard.
@@ -410,10 +396,19 @@ The browser maintains a local `clusterState` object that mirrors the cluster sna
Accessed via the "admin" button in the header (visible when authenticated
with `approve` scope). Provides user, API token, channel link, MCP server,
and skill management with 13 tabs (see also
[Governance](governance.md) for
the Roles, Policies, Skills, Usage, and Audit tabs, and
[Settings](settings.md) for the database-backed configuration editor):
and skill management with 18 tabs (Users, API Tokens, Channels, Schedules,
Audit, Memories, Models, Nodes, Settings, TLS). See also
[Governance](governance.md) for the Roles, Policies, Skills, Usage, and
Audit tabs, and [Settings](settings.md) for the database-backed
configuration editor.
The **Channels** tab links users to either a Discord or Slack account
via a per-row channel-type selector. The **Models** tab is a CRUD
editor for `model_definitions`, the **Nodes** tab edits per-node
metadata, and the **TLS** tab manages CA and leaf certificates for the
internal mTLS fabric. The **Settings** tab edits ConfigStore values
live; edits apply without restart.
**Users tab:**
@@ -482,17 +477,17 @@ to create the initial admin user and receive a JWT in one step. See
## Scheduled Tasks
The console includes a background **TaskScheduler** daemon that creates workstreams on a timed basis via the MQ broker. It supports cron-based recurring schedules and one-shot `at` schedules.
The console includes a background **TaskScheduler** daemon that creates workstreams on a timed basis via HTTP proxy to target nodes. It supports cron-based recurring schedules and one-shot `at` schedules.
### Architecture
The scheduler runs as a daemon thread inside the console process. Every `check_interval` seconds (default 15) it:
1. Acquires a distributed lock via Redis `SET NX EX` (prevents duplicate dispatch in multi-console deployments)
1. Acquires a distributed lock via the `system_settings` table (prevents duplicate dispatch in multi-console deployments)
2. Queries the storage backend for tasks whose `next_run <= now` and `enabled = true`
3. Dispatches each due task as one or more `CreateWorkstreamMessage` via MQ
3. Dispatches each due task as one or more workstream creation requests via HTTP proxy
4. Updates `last_run` and computes the next `next_run` (or disables one-shot `at` tasks)
5. Releases the lock via Lua script (safe conditional delete)
5. Releases the lock
Run history is automatically pruned (runs older than 90 days) approximately once per hour.
@@ -508,7 +503,7 @@ Run history is automatically pruned (runs older than 90 days) approximately once
| Mode | Behavior |
|------|----------|
| `auto` | Picks the reachable node with the most available capacity |
| `pool` | Pushes to the shared inbound queue (any bridge picks it up) |
| `pool` | Picks a reachable node with available capacity using round-robin |
| `all` | Fan-out to all reachable nodes (capped at `max_fan_out`, default 20) |
| `<node_id>` | Targets a specific node by ID |
@@ -645,12 +640,6 @@ CLI flags for `turnstone-console`:
Open `http://localhost:8090` for the cluster dashboard. Create workstreams via the "+ new" button. Click any workstream to open the proxied server UI — no direct access to server ports required.
| `approve_request` | One or more tool calls need operator approval | `items: [{call_id, header, preview, func_name, approval_label, needs_approval}]` |
| `state_change` | Worker-thread state transition (also re-emitted with the current state on every fresh subscribe so refresh-mid-stream restores composer mode) | `state` ∈ `running`, `thinking`, `attention`, `idle`, `error` |
| `in_progress_snapshot` | One-shot replay of the in-progress turn's content + reasoning when this client connects mid-stream | `content`, `reasoning` |
| `rename` | Session's display name changed | `name` |
| `intent_verdict` | Intent judge produced a verdict on a pending tool call | `risk_level`, `recommendation`, `reasons` |
| `output_warning` | Output guard flagged a tool result | `call_id`, `risk_level`, `flags` |
| `child_ws_created` | A direct child of this coord was just created (fan-out from the cluster bus) | `child_ws_id`, `node_id`, `name`, `parent_ws_id` (`ws_id` in the envelope is always the coord's own id) |
| `child_ws_state` | A direct child transitioned state | `child_ws_id`, `state` |
| `child_ws_closed` | A direct child closed | `child_ws_id` |
| `child_ws_rename` | A direct child's name changed | `child_ws_id`, `name` |
component [**turnstone:node:{node_id}**\n\nNode heartbeat + metadata.\nValue: JSON {server_url, started, ...}\nTTL: 60s (refreshed every 30s)\n\nOps: SET with EX, GET, SCAN] as node_hb <<STRING>>
}
package "Event Channels (Redis PUBSUB)" #FFF3E0 {
component [**turnstone:events:global**\n\nGlobal event broadcast.\nAll state changes, ws lifecycle.\n\nOps: PUBLISH, SUBSCRIBE] as evt_global <<PUBSUB>>
component [**turnstone:events:{ws_id}**\n\nPer-workstream events.\nContent, tools, status.\n\nOps: PUBLISH, SUBSCRIBE] as evt_ws <<PUBSUB>>
component [**turnstone:events:cluster**\n\nCluster-wide state changes.\nUsed by Console dashboard.\n\nOps: PUBLISH, SUBSCRIBE] as evt_cluster <<PUBSUB>>
**Default** (no flag) — starts `redis`, `server`, `bridge`,`console`. Requires an OpenAI-compatible LLM API running on the host (default: `http://localhost:8000/v1`).
**Default** (no flag) — starts `server` and`console`. Requires an OpenAI-compatible LLM API running on the host (default: `http://localhost:8000/v1`).
```bash
docker compose up
@@ -46,22 +39,12 @@ docker compose up
docker compose --profile production up
```
**Cluster** — 10-node server/bridge fleet sharing PostgreSQL and Redis. Access all nodes via the console at `:8090`. Requires `POSTGRES_PASSWORD`:
**Cluster** — 10-node server fleet sharing PostgreSQL. Access all nodes via the console at `:8090`. Requires `POSTGRES_PASSWORD`:
```bash
docker compose --profile cluster up
```
**Sim** — adds the simulator. Can run alongside the full stack or standalone with just Redis and the console:
```bash
# Sim + console (no LLM needed)
docker compose --profile sim up redis console sim
# Everything including sim
docker compose --profile sim up
```
## Configuration
All configuration is via environment variables in `.env` (copy from `.env.example`):
@@ -74,13 +57,6 @@ All configuration is via environment variables in `.env` (copy from `.env.exampl
| `OPENAI_API_KEY` | `dummy` | API key (`dummy` for local servers) |
| `TAVILY_API_KEY` | — | Web search API key (only needed for local/vLLM models; Anthropic and OpenAI search models use native search) |
| `TURNSTONE_DB_URL` | — | Database URL (e.g. `postgresql://user:pass@db:5432/turnstone`). For SQLite, defaults to `/data/.turnstone.db` |
| `TURNSTONE_DB_URL` | — | Database URL (e.g. `postgresql+psycopg://user:pass@postgres:5432/turnstone`). For SQLite, defaults to `/data/.turnstone.db` |
| `TURNSTONE_DB_LISTEN_URL` | (falls back to `TURNSTONE_DB_URL`) | Direct-to-PostgreSQL URL for the console's dedicated `LISTEN` connection. Set this when `TURNSTONE_DB_URL` points at PgBouncer in transaction pooling mode — LISTEN is session state and the transaction-pooled connection can't hold it. See [pgbouncer.md](pgbouncer.md). |
| `TURNSTONE_DB_POOL_SIZE` | `2` | PostgreSQL connection pool size per process (default: 2 base + 3 overflow = 5 max) |
| `POSTGRES_USER` | `turnstone` | PostgreSQL container username (used in default `TURNSTONE_DB_URL` for cluster/channel) |
| `POSTGRES_PASSWORD` | — | PostgreSQL container password (required for production and cluster profiles) |
The database stores workstream history, user accounts, and API tokens. When using JWT auth, a database backend is required for user storage.
> **Upgrading from <1.3.0a4:** Earlier versions used `DB_BACKEND` and `DATABASE_URL` in `.env`, which `compose.yaml` mapped to the `TURNSTONE_`-prefixed names internally. These short aliases have been removed. Rename `DB_BACKEND` → `TURNSTONE_DB_BACKEND` and `DATABASE_URL` → `TURNSTONE_DB_URL` in your `.env` file.
> **Large clusters:** Each turnstone process maintains a small connection pool (5 max). At hundreds of nodes this adds up — use [PgBouncer](pgbouncer.md) in transaction pooling mode between turnstone and PostgreSQL.
> **First-time setup:** After deploying with auth enabled, create an initial admin user by running `turnstone-admin create-user` inside the container:
@@ -129,30 +109,26 @@ The database stores workstream history, user accounts, and API tokens. When usin
| `TURNSTONE_SLACK_SLASH_COMMAND` | `/turnstone` | Slash command registered in the Slack app |
The channel service runs in the `production` profile. When`TURNSTONE_DISCORD_TOKEN` is set, the Discord adapter connects to the Discord Gateway and routes messages through Redis MQ to the bridge and server. See [Channel Integrations](channels.md) for full setup instructions including Discord application creation and user account linking.
### Simulator
| Variable | Default | Description |
|----------|---------|-------------|
| `SIM_NODES` | `100` | Number of simulated nodes |
The channel service runs in the `production` profile. When
`TURNSTONE_DISCORD_TOKEN` or the Slack pair is set the gateway starts the
corresponding adapter; both can run in one process. See
[Channel Integrations](channels.md) for platform app setup and user
account linking.
## Scaling
For multi-node testing, use the `cluster` profile which provides 10 dedicated server+bridge pairs with unique node IDs (`node-1` through `node-10`), resource limits, and shared PostgreSQL:
For multi-node testing, use the `cluster` profile which provides 10 server instances with unique node IDs (`node-1` through `node-10`), resource limits, and shared PostgreSQL:
```bash
POSTGRES_PASSWORD=secret docker compose --profile cluster up
```
The default `server` and `bridge` also run alongside the cluster nodes (11 total). All nodes are accessible via the console dashboard at `:8090`.
The default `server` also runs alongside the cluster nodes (11 total). All nodes are accessible via the console dashboard at `:8090`.
For production clusters beyond ~50 nodes, add PgBouncer between turnstone services and PostgreSQL. See [PgBouncer Connection Pooling](pgbouncer.md) for Docker Compose and Helm configuration.
@@ -160,7 +136,6 @@ For production clusters beyond ~50 nodes, add PgBouncer between turnstone servic
All entry points are installed in a single image: `turnstone-server`, `turnstone-bridge`, `turnstone-console`, `turnstone-channel`, `turnstone-admin`, `turnstone-sim`, `turnstone-eval`.
All entry points are installed in a single image: `turnstone`,
# MCP OAuth — per-user authorization for MCP servers
Turnstone supports **per-(user, MCP server) OAuth 2.1 + PKCE** delegation so each Turnstone user authorizes a remote MCP server with their own identity, rather than sharing a single bearer token across the deployment. This is the right shape for MCP servers that expose user-specific data (a personal CRM, an email inbox, a calendar) and for MCP servers that want per-user audit attribution.
Per-user OAuth is opt-in per `mcp_servers` row. Local-auth Turnstone installs with no `oauth_user` rows exercise zero new code paths — the entire feature is dark by default.
> **Note**: This is a separate authorization layer from Turnstone's own user authentication. A user who logs into Turnstone with a local username + password can still authorize a per-server OAuth MCP server. OIDC SSO and per-server OAuth are orthogonal.
---
## When to use which `auth_type`
The MCP server admin form exposes three authorization modes ("Multitenant Authorization"):
| `auth_type` | What it means | When to use |
|---|---|---|
| `none` | No headers attached. Open MCP server (or one gated by network policy only). | Internal MCP servers on a trusted network. |
| `static` | One static bearer token, configured per server, sent on every request from every user. | Service-to-service MCP servers where per-user attribution doesn't matter, or single-tenant deployments. |
| `oauth_user`*(recommended for user-data servers)* | Each user authorizes separately via OAuth 2.1 + PKCE; Turnstone stores per-user tokens encrypted at rest. | MCP servers that expose user-specific data or that want per-user audit attribution. |
Switching `auth_type` away from `oauth_user` orphans existing per-user tokens. Use the admin **bulk-revoke** affordance on the server row (Phase 9) to clear them, or let them expire naturally — they're inert without the matching `auth_type` value.
---
## Prerequisites for `auth_type=oauth_user`
1. **Encryption key**. Tokens are stored encrypted with Fernet. Set `[security] mcp_token_encryption_key` in `config.toml` (Turnstone won't start with an `oauth_user` row configured but no key installed). Rotate via `MultiFernet` — add the new key first, then later remove the old one once all rows have been re-encrypted.
2. **MCP server publishes RFC 9728 PRM and RFC 8414 AS metadata***or* you configure the AS URL override on the server row. PKCE S256 is mandatory; Turnstone refuses to connect to authorization servers that don't advertise `code_challenge_methods_supported: ["S256"]`.
3. **OAuth client registration**. Two paths:
- **Pre-registered** (most common): you create an OAuth client at the authorization server (manually, via admin console, or via Terraform), then paste the `client_id` / `client_secret` into the Turnstone admin form.
- **Dynamic client registration** (RFC 7591): if the AS supports it and you select that mode in the admin form, Turnstone registers a client at first use and persists the `client_id` automatically.
4. **Redirect URI** registered at the authorization server: `https://your-turnstone-host/v1/api/mcp/oauth/callback`.
---
## Configuration
### Per-server fields (admin UI)
| Field | Required | Description |
|---|---|---|
| Server URL | Yes | The MCP server's `streamable-http` base URL. |
| Authorization Server URL | No | Override for RFC 9728 PRM discovery. Set when your AS endpoint differs from the MCP server URL (e.g., corporate AS protecting a third-party MCP). When unset, Turnstone falls back to PRM discovery against the MCP server itself. |
| Client Secret | Optional (write-only) | OAuth 2.0 client secret (confidential client). Encrypted at rest. Written but never re-read by the API; field stays masked. |
| Scopes | No | Space-separated default scope set requested at the authorize endpoint. Per-tool step-up may union additional scopes from a server's `insufficient_scope` response. |
| Audience | No | RFC 8707 `resource=` parameter sent on every authorize and token request. Defaults to the MCP server URL when unset. Validate against the `aud` claim in returned JWT tokens. |
### Encryption key
```toml
[security]
mcp_token_encryption_key = "base64-fernet-key"
# For rotation, list the keys in priority order — first is used for new
Keep this in `config.toml` rather than environment variables. An in-process LLM with shell-tool access can read the server's environment via `env` / `os.environ` and exfiltrate any secret stored there; secrets in `config.toml` are only loaded into the server at startup and never re-read on a tool-driven path, so a prompt-injection attack against the agent cannot reach them.
---
## Lifecycle
1. **First tool call** for a user against an `oauth_user` MCP server: pool dispatch finds no stored token, returns `mcp_consent_required` to the agent. Dashboard renders an inline "Connect" action card.
2. **User clicks Connect**: opens `/v1/api/mcp/oauth/start?server=<name>` in a popup. Browser redirects through the AS authorize endpoint, user grants consent, AS redirects back to `/v1/api/mcp/oauth/callback`. Turnstone exchanges code → tokens via PKCE, validates audience, encrypts, persists in `mcp_user_tokens`, redirects user back to the originating URL.
3. **Subsequent tool calls** by the same user against the same server reuse the persisted token via the per-(user, server) session pool. Tokens auto-refresh via the refresh-token grant when expired; failed refresh emits `mcp_consent_required` to drive re-consent.
4. **Step-up scope**: when a tool call hits `403` with `WWW-Authenticate: error="insufficient_scope"`, Turnstone emits `mcp_insufficient_scope` with the parsed scope set; the dashboard offers a "Connect with additional scopes" affordance that opens `/v1/api/mcp/oauth/start?server=<name>&scopes=<extra>` so the union of original + new scopes flows into the AS authorize request.
5. **User revoke** (settings modal): `DELETE /v1/api/mcp/oauth/connections/{server_name}` runs the authoritative local delete + best-effort RFC 7009 upstream revoke (fire-and-forget, capped at 256 concurrent in-flight tasks).
6. **Admin bulk-revoke** (Phase 9): `POST /v1/api/admin/mcp-servers/{name}/bulk-revoke` drops every user's token for the server. Upstream RFC 7009 revoke is intentionally **not** attempted in bulk (avoids N upstream HTTP calls per admin click); tokens at the AS expire naturally. Use the per-user revoke endpoint if you need guaranteed upstream invalidation.
---
## Admin status indicators
The MCP Servers admin tab shows per-server status pills (Phase 9):
- **Consented users count** — distinct users with a non-expired token for this server. Surfaced as a `bulk-revoke (N)` button when ≥1; clicking it opens a confirmation dialog. Hidden when 0.
- **Last refresh** — timestamp + outcome (`ok` / `error:ClassName`) of the most recent manual or auto-reconnect refresh. Per node. Absent until at least one refresh has occurred (renders as "never" in the admin UI).
Additional indicators (circuit-breaker state, encryption-key mismatch) are exposed via `get_server_status` on the API but do not yet have a dedicated admin pill — operators see them today via the per-server status text + error tooltip and in audit logs. A future phase may surface these as discrete pills.
---
## Auth-type transitions
| From | To | What happens |
|---|---|---|
| `none` / `static` → `oauth_user` | — | New code path activates for this server. Existing static headers (if any) are no longer sent. Users must authorize on first use. |
| `oauth_user` → `none` / `static` | — | Existing `mcp_user_tokens` rows are **orphaned** — inert without a matching `auth_type`. Use admin bulk-revoke to drop them, or let them expire. Switching back to `oauth_user` later re-activates the orphaned rows if they haven't been deleted. |
| OAuth `client_id` or `client_secret` rotated | — | Existing tokens may stop refreshing if the AS treats them as bound to the previous client. Bulk-revoke after rotation. |
The orphan-by-default behavior is chosen so switching back to `oauth_user` is non-destructive. Bulk-revoke is the explicit cleanup path.
---
## Troubleshooting
| Symptom | Likely cause | Action |
|---|---|---|
| `mcp_consent_required` even after consenting | Token persistence failed, or refresh-token rejected by AS | Check audit log for `mcp_server.oauth.persist_failed` or `mcp_server.oauth.token_revoked`. Re-consent via settings modal. |
| `mcp_token_undecryptable_key_unknown` | Encryption key rotated without keeping the previous key in the keyring | Add the previous key back to `mcp_token_encryption_keys` until all rows have been re-encrypted, then drop. |
| `mcp_oauth_url_insecure` | MCP server URL is `http://` (not `https://`) on a non-loopback host | Use `https://`. Per-user bearers must not transit cleartext. |
| Tools fail in scheduled / Discord / Slack runs | OAuth-MCP requires browser-based consent | Users must pre-consent via the web UI. Phase 9 dashboard badge surfaces deferred consents from these runs on next login. |
| Circuit breaker open repeatedly | Transport-level errors on the MCP server (DNS, TLS, 5xx) | Check the per-server error pill; auth errors do not trip the breaker. |
See also: `docs/operations/mcp-oauth-headless.md` for the cron / channel-driven run caveat.
| `TURNSTONE_OIDC_PROVIDER_NAME` | No | `SSO` | Display name for the login button (e.g. "Google", "Okta") |
| `TURNSTONE_OIDC_ROLE_CLAIM` | No | — | ID token claim containing role/group values (see [Role Mapping](#role-mapping)) |
| `TURNSTONE_OIDC_ROLE_MAP` | No | — | Mapping from claim values to Turnstone role IDs (see [Role Mapping](#role-mapping)) |
| `TURNSTONE_OIDC_PASSWORD_ENABLED` | No | `true` | Set to `false` to hide the password form and block all username/password logins (including admin). API tokens and config-file tokens still work. |
| `TURNSTONE_OIDC_REDIRECT_BASE` | No | — | Externally-reachable origin for the OIDC redirect URI (e.g. `https://app.example.com`). Recommended when running behind a reverse proxy. When unset, derived from the request Host header. |
| `TURNSTONE_OIDC_PASSWORD_ENABLED` | No | `true` | Set to `false` to hide the password form and block all username/password logins (including admin). API tokens continue to work. |
| `TURNSTONE_OIDC_REDIRECT_BASE` | Yes | — | Externally-reachable origin for the OIDC redirect URI (e.g. `https://app.example.com`). Without this, OIDC will refuse to start. The previous Host-header fallback was unsafe under permissive reverse proxies. |
| `TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS` | No | — | Comma-separated list of additional hostnames whose endpoints the IdP discovery document is allowed to reference. See [Cross-host endpoints](#cross-host-endpoints). |
OIDC is enabled when all three required fields (issuer, client ID, client
secret) are non-empty. If any is missing, OIDC is silently disabled and
the login screen shows only the password form.
All four required fields — issuer, client ID, client secret, and
`TURNSTONE_OIDC_REDIRECT_BASE` — must be set. If any are missing OIDC
is disabled at startup (an error is logged when only `redirect_base`
is missing) and the login screen shows only the password form.
### Reverse Proxy / Load Balancer
### Redirect base (required)
When Turnstone runs behind a reverse proxy, the internal `Host` header may
not match the externally-reachable URL. Set `TURNSTONE_OIDC_REDIRECT_BASE`
to the public origin so the redirect URI sent to the identity provider is
correct:
`TURNSTONE_OIDC_REDIRECT_BASE` pins the redirect URI sent to the identity
provider to a known externally-visible origin. Set it to the public origin
# MCP OAuth in headless / scheduled / channel-driven runs
**Constraint**: OAuth-MCP servers (`auth_type=oauth_user`) require browser-based user consent. Users must pre-consent via the web UI before any run that cannot drive a browser redirect.
- Any future channel adapter without an interactive browser session.
**What happens when consent is missing**:
A tool call against an `oauth_user` server returns a structured `mcp_consent_required` error to the agent. The agent surfaces the deferred work in its output. Turnstone persists a record to `mcp_pending_consent` so the dashboard badge surfaces the deferred consent need to the user on next login.
**Recovery**:
The user opens the dashboard, sees the gear-icon badge counting pending consents, opens the settings modal, clicks Connect for each affected server, and completes the OAuth dance. The pending-consent record is cleared by the OAuth callback handler on success. Subsequent scheduled / channel runs use the freshly-stored token.
**Pre-consent recipe**:
Before scheduling a workstream that depends on an `oauth_user` MCP server, the user should:
1. Open the dashboard.
2. Open the settings modal (gear icon).
3. Click Connect on each MCP server the schedule will use.
4. Confirm consent in the popup.
This stores tokens that the scheduled run will reuse. Refresh-token rotation is handled transparently on the run side; only the first consent requires browser interaction.
In `values.yaml`, point the database at PgBouncer:
@@ -106,7 +108,7 @@ pgbouncer:
maxClientConn: 5000
maxDbConnections: 80
```
:
---
## Configuration reference
@@ -197,4 +199,28 @@ does not support prepared statements. Turnstone's SQLAlchemy layer does
not use server-side prepared statements by default, so this is not an
issue.
**LISTEN / NOTIFY not supported in transaction mode** — PgBouncer's
transaction pooling assigns a real server connection only for the
duration of each transaction, then returns it to the pool. PostgreSQL
`LISTEN` is session state — a transaction-pooled client can't hold the
multi-statement session a long-lived `LISTEN` needs. The console's
`NotifyDispatcher` (reactive node discovery via the `services` channel)
therefore opens a **dedicated, direct-to-Postgres** connection that
bypasses PgBouncer.
Configure via `config.toml``[database] listen_url` (preferred —
co-located with the main `url`) or the `TURNSTONE_DB_LISTEN_URL` env var
(config.toml wins when both are set). Defaults to the main DB URL when
unset.
| Setting | Behaviour |
|---|---|
| unset | Listener uses `TURNSTONE_DB_URL` as-is. Fine when PgBouncer is in **session** mode, or when there's no pooler in front of Postgres. With transaction-mode PgBouncer the listener's `LISTEN` will fail and the dispatcher retries with exponential backoff (1 s → 30 s cap) without ever succeeding. Reactive NOTIFY-driven node discovery is silently lost; the cluster collector's 60 s `_discovery_loop` is the only remaining backstop. |
| set to direct-to-PG URL (e.g. `postgresql://…/turnstone`) | Listener bypasses PgBouncer for its one dedicated connection. Reactive discovery latency drops from up-to-60 s to ~500 ms. The rest of the storage layer continues to go through PgBouncer in transaction mode. |
Set this whenever PgBouncer is in transaction mode (the recommended
setting per this doc). The override only adds one long-lived PG
connection per console process — sized into the cluster's
`max_connections` budget alongside the pool.
See also: [Docker deployment](docker.md) · [Security](security.md)
This bumps `pyproject.toml` + `turnstone/__init__.py`, regenerates `uv.lock`, commits, tags `v1.5.0a2`, and pushes. CI runs, then publish + Docker workflows fire automatically.
## Releasing a Stable Patch (from stable/X.Y)
```bash
git checkout stable/1.4
git cherry-pick <commit-hash> # bugfix from main
scripts/release.sh 1.4.1 --push
```
## Promoting Experimental to Stable
When `main` is ready for a stable release:
```bash
# 1. Tag the stable release on main
scripts/release.sh 1.5.0 --push
# 2. Create the stable maintenance branch from that tag
git branch stable/1.5 v1.5.0
git push origin stable/1.5
# 3. Start the next experimental cycle on main
scripts/release.sh 1.6.0a1 --push
```
The previous stable branch (`stable/1.4`) continues to receive
security-only patches; older tracks (`stable/1.0`, `stable/1.3`) are
retired when they fall out of support.
## CI/CD Pipeline
All releases are gated on CI success:
1. `git push` with `v*` tag triggers **CI** (lint, typecheck, test, test-postgres, lock-check, security audit)
2. On CI success, **Publish to PyPI** fires via `workflow_run`
3. On CI success, **Publish Docker Image** fires via `workflow_run`
- `client.logout()` clears the stored JWT from the client.
- If a request returns 401, the SDK raises `TurnstoneAPIError` — the caller is responsible for re-authenticating.
### Backward Compatibility
### Token Types
The config-file token (`TURNSTONE_AUTH_TOKEN`) still works as a simple Bearer token for environments that do not use the user/JWT system. When the server receives a non-JWT Bearer token, it falls back to the legacy token check.
The SDK accepts any Bearer token — JWTs (from `ServiceTokenManager` or login) and API tokens (`ts_` prefix) are both supported. Use `token_factory` for auto-rotating JWTs or a static `token` for API tokens.
When a per-model override is `NULL` (empty in the UI), the global default is
used. Switching models via `/model <alias>` re-resolves sampling parameters
from the new model's overrides or global defaults.
**Removed settings:** `model.name` and `model.context_window` have been removed
from ConfigStore. Model names and context windows are now configured per-model
in the Models tab. A startup warning is logged if these keys appear in
`config.toml`.
### Reasoning persistence (per-model)
Two boolean flags on `model_definitions` (migration 052) control how
reasoning text round-trips per model:
| Flag | Default | Effect |
|------|---------|--------|
| `surface_persisted_reasoning` | `True` | Surface stored reasoning text on `/history` payloads so a page reload re-renders the reasoning bubble. **Storage of reasoning bytes is independent of this flag** — they ride in `provider_data` regardless. |
| `replay_reasoning_to_model` | `False` | Send stored reasoning blocks back to the provider on subsequent turns. Capability-gated: only takes effect when the model's `ModelCapabilities.supports_reasoning_replay` is also `True`. Set on canonical OpenAI gpt-5*/o-series and Anthropic Claude entries; unknown / local-server models default to `False` so an operator who flips the flag on a model whose API doesn't understand reasoning replay silently no-ops rather than 400-ing. |
Edit both via the admin Models tab. See the architecture doc for the
`reasoning_text` for Chat Completions / vLLM / llama.cpp / Gemini-compat).
### Plan / task agent overrides
`plan_agent` and `task_agent` sub-sessions resolve independently from the
conversation model so operators can pick a cheaper/faster model for
autonomous loops:
| Setting | Purpose |
|---------|---------|
| `model.plan_alias` | Alias used for `plan_agent` sub-sessions. Falls back to `[model].plan_model` in config.toml, then `[model].agent_model`, then the session's active model. |
| `model.task_alias` | Alias used for `task_agent` sub-sessions. Same fallback chain as `plan_alias`. |
The simulator (`turnstone-sim`) creates lightweight simulated nodes that talk to a real Redis instance using the standard turnstone protocol. External observers — `TurnstoneClient`, `turnstone-console`, real bridges — see identical behavior. No LLM backend is needed.
Use `--metrics-file report.json` to write the full report as JSON.
## Console Integration
The simulator's nodes appear in `turnstone-console` exactly like real nodes. Run them together to see the dashboard populate with simulated workstreams:
```bash
# Terminal 1: start Redis and console
docker compose up redis console
# Terminal 2: run simulator
docker compose --profile sim up sim
```
Or all at once:
```bash
SIM_NODES=50 SIM_DURATION=120 docker compose --profile sim up redis console sim
```
Open http://localhost:8090 to see simulated nodes, workstream states, token counts, and load bars updating in real time.
## Architecture
> See also: [Simulator Architecture diagram](diagrams/png/10-simulator-architecture.png)
```
turnstone/sim/
├── __init__.py # Public API: SimCluster, SimConfig
├── config.py # SimConfig — all simulation parameters
**Key design:** The `InboundDispatcher` batches ~50 node queues into a single Redis `BLPOP` call, keeping connection count bounded at ~20 regardless of node count. All nodes share a single `ConnectionPool(max_connections=64)`.
description: Use this skill when the user wants to import or migrate conversation history from another LLM chat or coding tool (e.g. ChatGPT, Claude.ai, Cursor, Copilot Chat, Aider, Gemini, a custom JSON export) into Turnstone. The skill teaches Turnstone's destination contracts — workstream identity, the OpenAI-shaped message rows, tool-call/result pairing, provider-fidelity blobs, attachments, and archive-vs-resumable choice — so the agent can map any source format onto them. Trigger phrases: "import my chats", "migrate this transcript into Turnstone", "bring my Claude.ai history over", "load this export as a workstream".
version: 1.0.0
---
# Importing Conversation History into Turnstone
## Overview
Source formats vary; the destination does not. Your job is to translate whatever the user hands you (JSON dump, ZIP export, scraped HTML, screenshot OCR, raw transcript) into Turnstone's internal shape: **one workstream row** plus an ordered sequence of **conversation rows** in OpenAI message format. This skill documents the destination so you can write a correct mapper for any source.
Two questions to settle with the user before writing anything:
1. **Archive or resumable?** An archive ("saved" workstream — `state="closed"`) is read-only history. A resumable workstream (`state="idle"`) lets the user continue the conversation; this only works cleanly when the source LLM matches a Turnstone-supported provider/model and tool definitions still resolve.
2. **One workstream per source thread, or merge?** Default to one-to-one unless the user explicitly asks to merge.
Default to **archive** when in doubt — resuming a foreign transcript with mismatched tool schemas or stale provider signatures will fail at the next turn.
## Turnstone Data Model (the destination)
Two tables carry the conversation:
### `workstreams` (one row per imported thread)
| Column | Required | Notes |
|---|---|---|
| `ws_id` | yes | 32-char lowercase hex. Auto-generate with `secrets.token_hex(16)` if you don't already have one. **First 4 hex chars are the routing bucket** — see "Identity & Routing" below. |
| `name` | yes | Short title. Pull from source thread title; fall back to first ~60 chars of first user message. |
| `state` | yes | `"closed"` for archive, `"idle"` for resumable. Never set `"running"` on import. |
| `kind` | yes | `"interactive"` for normal threads. Do NOT use `"coordinator"` for imports — that's reserved for cluster-spawned coordinator workstreams. |
| `parent_ws_id` | no | Leave NULL. Only set if you're importing a coordinator-spawned subtree and re-parenting it; rare. |
| `user_id` | yes | Owner. Must exist in `users`; importer must know which Turnstone user owns the imported history. |
| `node_id` | yes (multi-node) | Denormalized cache of the node that owns this `ws_id`'s bucket. Single-node deployments can leave it NULL or set it to the only node. |
| `alias` | no | Human-typeable short name. Optional; must be unique cluster-wide if set. |
| `title` | no | Auto-titled later by the LLM; safe to leave NULL on import. |
| `skill_id`, `skill_version` | yes | Default `""` and `0` unless the source thread was scoped to a Turnstone skill. |
| `created`, `updated` | yes | ISO8601 strings. Use the source's first/last message timestamps when available. |
### `conversations` (many rows per thread, ordered by `id`/`timestamp`)
| Column | Notes |
|---|---|
| `ws_id` | The workstream this row belongs to. |
| `timestamp` | ISO8601 string. Preserve source timestamps; fall back to monotonically increasing values if unknown. **Order is canonical via `id` (autoincrement), not `timestamp`** — but always insert in conversational order so both agree. |
| `role` | One of `system`, `user`, `assistant`, `tool`, `developer`. See role mapping below. |
| `content` | Text. May be NULL for assistant rows that are *only* tool calls. |
| `tool_name` | Set on `role="tool"` rows (the tool whose result this is). NULL otherwise. |
| `tool_call_id` | Set on `role="tool"` rows (matches the assistant row's `tool_calls[].id`). NULL otherwise. |
| `tool_calls` | JSON-encoded list, on `role="assistant"` rows that issued tool calls. OpenAI shape — see "Tool Calls" below. |
| `provider_data` | JSON blob preserving provider-native content blocks (Anthropic `signature`, Gemini `thought_signature`, etc.). Optional; only matters for **resumable** imports against the same provider. Skip for archives. |
The internal format is **OpenAI-shaped**, even when the source was Anthropic or Gemini. Providers translate at their own API boundary; storage stays uniform.
## Identity & Routing (`ws_id`)
- `ws_id` is **32-char lowercase hex** (i.e. `secrets.token_hex(16)`).
- The **routing bucket** is `int(ws_id[:4], 16)` — the first 4 hex chars place this workstream on a specific node via the consistent hash ring.
- For multi-node imports: either insert through the console's routing proxy (which forwards to the owning node), or generate `ws_id`s and write directly to each node's database in batches grouped by bucket.
- For single-node imports: bucket math is irrelevant; any `ws_id` works.
- **Do not reuse the source platform's IDs as `ws_id`** unless they happen to be 32-char hex. Generate fresh; if you need the old ID for traceability, store it in `workstream_config` under a key like `import.source_id`.
## Recommended Import Path
Three options, in order of preference:
### 1. Storage protocol (recommended for full history)
Use `turnstone.core.storage.Storage.save_messages_bulk(rows)`. This is the canonical bulk-insert primitive and bypasses the LLM round-trip entirely.
```python
from turnstone.core.storage import get_storage # construct via the same path the server uses
storage = get_storage(...) # see turnstone.core.storage.__init__ for the project's wiring
storage.create_workstream( # or whatever the project's exposed creator is — check turnstone/core/storage/_protocol.py
`save_messages_bulk` handles `timestamp` and the workstream's `updated` column internally, so you don't need to compute them per row. **Verify the exact creator signature** by reading `turnstone/core/storage/_protocol.py` — table layout has shifted across migrations and the Storage protocol is the source of truth.
### 2. SDK `create_workstream(resume_ws=...)` (when the source is already a Turnstone workstream)
Only useful for *Turnstone → Turnstone* re-parenting. Not relevant for foreign sources.
### 3. SDK `create_workstream(initial_message=...)` + `send()` per turn (last resort)
Only fits archives where the source had **no tool calls** and you don't care about preserving assistant turns verbatim. Each `send()` triggers a real LLM round-trip, which is expensive and rewrites assistant content. Don't use this for full history.
## Role Mapping
Common source-role conventions and how they map to Turnstone:
| `system` | `system` | Preserve only if it's content the user wrote (custom instructions). Drop boilerplate provider preambles — Turnstone composes its own system message. |
| `tool`, `function`, `tool_result` | `tool` | Must carry `tool_name` and `tool_call_id` matching the prior assistant row's `tool_calls[].id`. |
| `tool_use` (Anthropic) | `assistant` with `tool_calls` | Anthropic emits tool calls *inside* an assistant message; flatten to OpenAI shape. |
| `human_feedback`, `revision` | `user` | Treat as a follow-up user turn. |
## Tool Calls (the most error-prone part)
Turnstone stores tool calls in OpenAI's nested-function shape on the assistant row, and matches them with `role="tool"` result rows by `tool_call_id`.
### Assistant row with tool calls
```json
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "search_web",
"arguments": "{\"query\":\"turnstone import\"}"
}
}
]
}
```
`tool_calls[].function.arguments` is **a JSON-encoded string**, not an object. Source formats commonly get this wrong — Anthropic stores arguments as a parsed object, Gemini as a struct. Always re-serialize to a string.
### Tool result row
```json
{
"role": "tool",
"tool_name": "search_web",
"tool_call_id": "call_abc123",
"content": "..."
}
```
Pairing rules:
- Every assistant `tool_calls[].id` MUST be followed by exactly one `role="tool"` row with the matching `tool_call_id`, before the next user/assistant turn.
- If the source dropped the tool result (cut-off transcript), insert a synthetic `role="tool"` row with `content="[tool result missing in source]"` to keep the chain valid. An assistant row with an unanswered `tool_calls[].id` will break replay and any LLM round-trip.
- Multi-tool assistant turns: one `role="tool"` row per call, in any order, all before the next non-tool row.
### Tool ID generation
If the source used opaque tool IDs that aren't unique within a thread (some platforms reuse them), regenerate with a stable scheme like `f"call_{i}"` where `i` is a per-thread counter. Update both the assistant and tool rows together.
## Provider Fidelity (`provider_data`)
Skip this entirely for **archive** imports.
For **resumable** imports against the same provider, populate `provider_data` to preserve provider-specific tool-call metadata that the next API round-trip will require:
- **Anthropic**: `signature` field on thinking blocks; required for round-tripping extended-thinking responses.
- **Gemini**: `thought_signature` on tool calls; required for fidelity.
- **OpenAI**: typically nothing to preserve.
The runtime-side dict key is `_provider_content` (a list of provider-native blocks); the persisted column is `provider_data` (the same list, JSON-encoded). If you don't have provider-native blocks from the source — and you usually won't, because a foreign export won't include them — leave `provider_data` NULL. The first new turn will succeed without it, but the previous assistant turn's reasoning won't replay back to the model.
## Attachments
If the source thread had image or file attachments:
- **Size limits**: images ≤ 4 MiB, text documents ≤ 512 KiB. Reject or downsample anything bigger.
- **Allowed types**: server validates magic bytes for images and UTF-8-decodes for text. Binary blobs that aren't images won't pass.
- **Lifecycle**: pending → reserved → consumed. For imports, the cleanest path is to upload as pending and immediately consume by attaching to the relevant `conversations.id`.
Two import paths:
1. **Bulk-insert + post-attach**: insert messages first, get back the assistant/user `conversations.id`, then write `workstream_attachments` rows linking the file to `message_id`.
2. **SDK multipart create**: `create_workstream(attachments=[...], initial_message=...)` for the *first* turn only — the server reserves and consumes them onto that turn. Doesn't help for mid-thread attachments.
For full-history imports with multiple attachments at different turns, path (1) is the only option.
## Validation Checklist
Before declaring success, verify:
- [ ] `ws_id` is 32-char lowercase hex.
- [ ] `workstreams` row exists with the right `user_id`, `state`, `kind`.
- [ ] Conversation rows are inserted **in order** (autoincrement `id` will reflect insert order).
- [ ] Every assistant `tool_calls[].id` has a matching `role="tool"` row with the same `tool_call_id`.
- [ ] `tool_calls[].function.arguments` is a JSON-encoded **string**, not a parsed object.
- [ ] First message is typically `role="user"` (not `system`) — Turnstone composes its own system prompt at runtime.
- [ ] No empty assistant rows (`content=NULL` AND `tool_calls=NULL` is invalid).
- [ ] If multi-node: the `ws_id`'s bucket maps to a node that exists; `workstreams.node_id` matches.
- [ ] Round-trip test: run `Storage.load_messages(ws_id)` and confirm the reconstructed list matches what you inserted (modulo timestamps).
## Anti-patterns
- **Don't import the source provider's system prompt verbatim.** Provider boilerplate ("You are Claude...", "You are ChatGPT...") will conflict with Turnstone's composed system message and confuse the model on resume. Drop it; preserve only user-authored custom instructions.
- **Don't preserve foreign tool definitions as Turnstone tools.** If the source had custom tools that don't exist in Turnstone, the assistant rows that called them are still valid history (archive), but the workstream is **not resumable** — mark `state="closed"`.
- **Don't fabricate `tool_call_id`s without re-pairing.** Mismatched ids silently break the replay chain on the next turn.
- **Don't skip the `tool_name` field on `role="tool"` rows.** Some load paths use it for display and audit; NULL there will render as "unknown tool".
- **Don't write through the LLM (`send()` per turn) for full history.** It's expensive, rewrites assistant turns, and rate-limits will bite long imports.
- `turnstone/core/session.py` (around the message-save section) — how the runtime constructs in-memory message dicts; mirror this shape on import to round-trip cleanly.
- `turnstone/api/server_schemas.py` — Pydantic shapes for the SDK paths if you go through HTTP.
- Special post-execution gate for `plan`: the plan output is shown to the user
for review, and the user can reject or annotate it.
@@ -164,8 +169,8 @@ Every tool defines a `primary_key`. The mapping is:
| `man` | `page` |
| `web_fetch` | `url` |
| `web_search` | `query` |
| `task` | `prompt` |
| `plan` | `prompt` |
| `task_agent` | `prompt` |
| `plan_agent` | `goal` |
| `memory` | `name` |
| `recall` | `query` |
| `notify` | `message` |
@@ -183,8 +188,11 @@ Execute a bash command and return stdout + stderr.
| Parameter | Type | Required | Description |
|-----------|--------|----------|-------------|
| `command` | string | yes | The bash command to execute. |
| `timeout` | integer | no | Timeout in seconds (1-600). Omit to use the global `tools.timeout` setting (typically 120s). |
| `stop_on_error` | boolean | no | Enable `set -e` so the script exits on the first command failure. Default false. |
- **What it does**: Runs the command in a subprocess with a configurable timeout. Commands are sanitized and checked against a blocklist (e.g. `rm -rf /`).
- **What it does**: Runs the command in a subprocess with a configurable timeout. Commands are sanitized and checked against a blocklist (e.g. `rm -rf /`). Environment variables containing secrets are scrubbed (`*_KEY`, `*_SECRET`, `*_TOKEN`, etc.).
- **Output format**: Stdout is returned directly. Stderr lines are prefixed with `[stderr]` so the model can distinguish them. When the command itself redirects stderr to stdout (`2>&1`), no prefix is added. Output exceeding 256KB is truncated (head + tail preserved, middle replaced with a truncation notice).
- **Auto-approve**: No -- requires user confirmation.
- **Agent availability**: `task_agent` only (not available to plan sub-agents).
@@ -216,8 +224,9 @@ Write content to a file, creating it if needed.
| `near_line` | integer | no | Disambiguate when `old_string` matches multiple locations. |
| `edits` | array | no* | Multiple replacements to apply atomically (see below). |
| `replace_all` | boolean | no | Replace ALL occurrences of `old_string`. Cannot combine with `near_line` or `edits`. |
- **What it does**: Finds `old_string` in the file and replaces it with `new_string`. Fails if the string is not found or matches multiple locations (unless `near_line` is provided to pick the nearest match). Requires a prior `read_file` call on the same path.
\* Provide either `old_string`+`new_string` (single edit) or `edits` array (batch), not both.
- **What it does**: Finds `old_string` in the file and replaces it with `new_string`. Fails if the string is not found or matches multiple locations (unless `near_line` or `replace_all` is provided). Requires a prior `read_file` or `diff_file` call on the same path.
- **Batch mode**: The `edits` array accepts multiple `{old_string, new_string, near_line?}` entries applied atomically. All edits are validated before any are applied. Overlapping edits (two entries targeting the same text region) are rejected. Edits are applied in reverse file-position order so character offsets stay stable.
- **Replace-all mode**: When `replace_all` is true, all occurrences are replaced via `str.replace()`. The approval preview shows the occurrence count.
- **Auto-approve**: No -- requires user confirmation.
- **Agent availability**: `task_agent` only.
---
### diff_file
Show a unified diff between two files, or between a file and a provided string.
| `path_a` | string | yes | Path to the first file. |
| `path_b` | string | no | Path to the second file. Mutually exclusive with `content_b`. |
| `content_b` | string | no | String content to compare against `path_a`. Mutually exclusive with `path_b`. |
| `context_lines` | integer | no | Number of context lines around changes (default 3, max 20). |
- **What it does**: Returns unified diff output using Python's `difflib`. Binary files (containing null bytes) are rejected with a clear error. Files read through `diff_file` satisfy `edit_file`'s read guard — you can diff then edit without a separate `read_file` call. Large diffs are streamed with early cutoff at the tool truncation limit.
- **Auto-approve**: Yes (read-only).
- **Agent availability**: `agent` and `task_agent`.
---
### search
Search file contents for a regex pattern.
@@ -265,8 +297,9 @@ Execute Python code for math and computation in a sandbox.
|-----------|--------|----------|-------------|
| `code` | string | yes | Python code to execute. Must use `print()` for output. |
- **What it does**: Runs Python code in a sandboxed environment with pre-imported libraries: `sympy`, `numpy`, `scipy`, `math`, `fractions`, `itertools`, `functools`, `collections`, `decimal`, `operator`, `random`, `re`, `string`. Common sympy names (`symbols`, `solve`, `simplify`, `sqrt`, `Matrix`, etc.) are pre-imported.
- **Auto-approve**: No -- requires user confirmation.
- **What it does**: Runs Python code in a sandboxed environment with pre-imported libraries: `sympy`, `numpy`, `scipy`, `math`, `fractions`, `itertools`, `functools`, `collections`, `decimal`, `operator`, `random`, `re`, `string`. Common sympy names (`symbols`, `solve`, `simplify`, `sqrt`, `Matrix`, etc.) are pre-imported.`pytest` is also available for import.
- **Installation**: `sympy`, `numpy`, `scipy`, and `pytest` require the `[sandbox]` extras group: `pip install turnstone[sandbox]` (included in `[all]`).
- **Auto-approve**: Yes.
- **Agent availability**: `agent` and `task_agent`.
---
@@ -324,7 +357,10 @@ Search the web using a text query.
## Agent
### task
Tool names use the `_agent` suffix — bare `plan` / `task` collide with
chat-template channel names on some local models.
### task_agent
Delegate a general-purpose task to an autonomous sub-agent.
@@ -338,7 +374,7 @@ Delegate a general-purpose task to an autonomous sub-agent.
---
### plan
### plan_agent
Plan before implementing -- an autonomous agent explores the codebase and writes a structured plan.
@@ -510,11 +546,11 @@ pre-configure skills at workstream creation.
- `load` — Activate a skill by name. Calls `set_skill()` which handles content
rendering with `{{model}}`/`{{ws_id}}`/`{{node_id}}` variables, system message
reinitialization, and config persistence. Returns the skill name, description,
and security scan tier. Warns on high/critical scan status.
and security risk level. Warns on high/critical risk level.
- `search` — Find available skills by query. Uses BM25 relevance ranking over
name, description, tags, and category (same `BM25Index` used by memory
relevance and tool search). Returns up to 10 results with name, description,
@@ -596,7 +632,7 @@ CLI flags override the config file:
directly.
2. **Partitioning**: When active, tools are split into two sets:
- **Always-on** -- the 17 built-in tools (members of `BUILTIN_TOOL_NAMES`).
- **Always-on** -- the 19 built-in tools (members of `BUILTIN_TOOL_NAMES`).
These are always visible to the model.
- **Deferred** -- all MCP tools. These are not sent in the tool list unless
the model searches for them.
@@ -639,7 +675,7 @@ MCP-compatible service.
3. **Schema conversion**: Each MCP tool's `inputSchema` is converted to OpenAI
function-calling format. The tool name is prefixed: `mcp__{server}__{tool}`.
4. **Merging**: MCP tools are appended after the 17 built-in tools via
4. **Merging**: MCP tools are appended after the 19 built-in tools via
`merge_mcp_tools()`. Built-in tools appear first, giving them natural LLM priority.
When dynamic tool search is active, MCP tools are deferred rather than directly
visible -- the model discovers them via search as needed (see
@@ -657,7 +693,7 @@ that external tools are read-only. However, global overrides such as
`--skip-permissions` will auto-approve all tools, including MCP tools. The
interactive "Always" button adds specific tool types to the per-tool auto-approve
set. The web UI and server use `approval_label` for MCP tools, giving
per-prompt/per-resource granularity. The CLI and bridge use `func_name`, which
per-prompt/per-resource granularity. The CLI uses`func_name`, which
gives per-tool-type granularity (e.g., all `use_prompt` calls).
### Sub-agent availability
@@ -722,22 +758,18 @@ MCP tools (3):
### Dynamic tool refresh
MCP tool lists stay up-to-date without restart through three mechanisms:
MCP tool lists stay up-to-date without restart through two mechanisms:
1. **Push notifications** -- MCP servers that declare `tools.listChanged: true` in
their capabilities send `notifications/tools/list_changed` when their tool list
changes. `MCPClientManager` registers a `message_handler` on each `ClientSession`
that triggers an immediate refresh for that server.
2. **Periodic timer** -- Servers that do *not* support push notifications are polled
on a configurable interval (default 4 hours). The timer is staggered using a
launch-time seed (`monotonic_ns ^ pid`) so cluster nodes don't all hit MCP
servers simultaneously. Configure via `[mcp] refresh_interval` in `config.toml`
or `--mcp-refresh-interval SECONDS` on the CLI. Set to `0` to disable.
3. **Manual** -- `/mcp refresh` re-fetches tools from all servers immediately.
2. **Manual** -- `/mcp refresh` re-fetches tools from all servers immediately.
`/mcp refresh <server>` targets a single server. If a server has disconnected,
manual refresh attempts reconnection.
manual refresh attempts reconnection. The console admin panel exposes the
same controls (refresh / reconnect buttons per server) for cluster-wide
fan-out.
When tools change, `MCPClientManager` rebuilds its merged tool list using copy-on-write
(new list/dict objects assigned atomically) and notifies all active `ChatSession`
@@ -745,11 +777,6 @@ instances via registered listener callbacks. Each session rebuilds its `_tools`,
`_task_tools`, `_agent_tools`, and reconstructs its `ToolSearchManager` (if active),
preserving the set of previously expanded (discovered) tools.
```toml
[mcp]
refresh_interval = 14400 # seconds (default 4h), 0 to disable
```
```
/mcp refresh
MCP refresh complete:
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.