Four Copilot findings on c6041c6 — all confirmed valid, all bounded
to authenticated-user prompt-injection scenarios but worth closing
before merge.
Wrapper-detect bypass (string + list branches of
``_apply_reminders_for_provider``):
The round-2 fix used ``content.startswith("<tool_output>\\n")`` to
detect already-wrapped content and skip ``escape_wrapper_tags``. A
tool whose RAW output starts with that prefix (e.g. ``echo
'<tool_output>'``) would match and have its escape skipped, letting
literal ``<tool_output>`` / ``<system-reminder>`` tags reach the model
and impersonate a system envelope. Replace the prefix check with
``extract_advisories_from_tool_envelope(content) is not None`` —
parsing requires the open AND matching close tags AND a structurally
valid envelope, raising the bypass bar significantly.
Mirror fix in the list-content branch so a tool emitting an unmatched
envelope as a text part can't bypass the per-text-part escape.
``_build_history`` legitimate-envelope drop:
The list-content drop path previously removed any text part starting
with ``<tool_output>\\n``. A tool that legitimately outputs a
well-formed envelope (documentation viewer, code analyzer demoing the
wrapper, an echo tool) would have that part silently disappear on
replay. Tighten the drop heuristic to require BOTH ``cleaned_text ==
""`` AND at least one extracted advisory — the structural signature of
the injected ``wrap_tool_result("", advisories)`` carrier we produce
in ``session.py`` for list-typed tool output. A legitimate envelope
has non-empty inner body or no advisory blocks and survives the
projection.
Empty advisory body:
``queue_message`` accepts any non-None text including ``""`` and
whitespace-only strings. ``_classify_advisory`` would return a
``user_interjection`` advisory with empty / whitespace body, which
``replayAdvisoriesAfterTool`` then renders as a featureless empty user
bubble. Filter empty / whitespace-only bodies at classification time
so the wire-shape contract is uniform: no empty advisories ever ride
the wire.
Tests:
* ``test_apply_reminders_escapes_tool_output_starting_with_envelope_prefix``
pins the structural-parser bypass close: a string starting with the
envelope prefix but lacking a close tag still gets escaped.
* ``test_apply_reminders_escapes_list_text_part_with_unmatched_envelope_prefix``
mirrors for the list-content branch.
* ``test_build_history_keeps_legitimate_envelope_text_part_with_body``
pins that legitimate envelope output stays in the projected list.
* ``test_decorate_suppresses_empty_advisory_body`` and
``test_decorate_suppresses_whitespace_only_advisory_body`` pin the
empty-body filter in ``_classify_advisory``.
Tests: 5923 passed, 3 deselected. Lint + format + mypy clean.
(cherry picked from commit c2cb6a7ea5)
Reverses the seam-2-only design from the prior commits on this branch.
Queued user messages arriving DURING a tool batch (Seam 1) splice into
the last tool result's envelope as ``UserInterjection`` advisories via
``wrap_tool_result``. Messages arriving BETWEEN turns (Seam 2) drain
as a single trailing user row via ``_flush_queued_messages`` with
``user_feedback`` (operator text alongside an approval, e.g. "y, use
full path") folded in as a prefix. Cancel/exception drains (Seam 3)
keep the existing ``_flush_queued_messages()`` call unchanged.
Why all three seams:
* Strict-template providers (Mistral, Llama via vLLM with stock chat
templates) reject role-alternation violations. A literal ``user``
row mid-tool-batch breaks ``assistant(tool_calls) → tool → ... →
assistant``; back-to-back ``user → user`` rows on the wire also fail.
* The seam-2-only design produced back-to-back ``user`` whenever
``user_feedback`` and queued items both fired — bug-1 from the round-1
review. Folding ``user_feedback`` as a prefix to the queue-drain
collapses the two into one row.
* During-batch arrivals couldn't ride seam 2 — the splice was the only
way to deliver same-turn without violating role alternation.
Storage symmetry:
Tool DB rows now store the wrapped ``output`` (envelope + advisories)
unconditionally — ``self.messages[i]['content']`` and
``conversations.content`` match exactly. List-typed output (image /
structured MCP results) uses ``wrap_tool_result(raw_joined_text,
advisories)`` at save time so the persisted string is anchored on
``<tool_output>\n`` for the replay parser. ``TOOL_RESULT_STORAGE_CAP``
is removed entirely; tools are responsible for bounding their own
output, storage faithfully represents in-memory. Removing the cap
also simplifies the parser — no truncated-envelope edge case.
Replay extraction:
``decorate_history_messages`` (REST ``/history``) and ``_build_history``
(SSE replay, resume, rewind, retry, post-load, rename re-replay) both
call the public ``extract_advisories_from_tool_envelope`` helper to
pull the envelope back into structured ``advisories`` for JS replay.
Both string content and list-typed content (image+queued-message
combo) covered. JS renders extracted advisories as normal user
bubbles after the tool block via the shared ``replayAdvisoriesAfterTool``
helper in ``shared_static/utils.js``.
Wrapper-tag escape and provider splice:
``escape_wrapper_tags`` now encodes pre-existing ``&`` first using an
``&`` sentinel so tool output containing literal entity strings
(documentation viewers, code analyzers, web scrapers returning entity-
encoded markup) round-trips correctly. Both encode and decode helpers
short-circuit on absence of ``<`` / ``&``.
``_apply_reminders_for_provider`` detects already-wrapped content
(string body and list text-part) by ``startswith("<tool_output>\n")``
and skips re-escape so existing envelopes survive intact when a tool
message also carries ``_reminders`` (the queued-message + tool-error
co-occurrence case is now common).
``decorate_history_messages`` runs in ``asyncio.to_thread`` to keep
MB-scale string work off the event loop.
Other cleanup:
* ``_collect_advisories`` delegates the queue drain to a named helper
``_drain_queued_messages_to_advisories`` so the swap-and-clear pattern
lives next to ``_flush_queued_messages``'s identical pattern and the
side-effect is documented at the call site.
* Preamble strings + body marker for ``UserInterjection`` round-trip
detection moved to module-level constants in ``tool_advisory.py``;
imported by ``history_decoration.py`` so a producer-side rephrase
can't silently desync the parser.
* ``_send_with_mocks`` ctxmgr extracted in ``test_session.py`` — the
six new send-driven tests share an 8-deep ``patch.object`` block.
* ``replayAdvisoriesAfterTool`` shared helper in
``shared_static/utils.js``; ``app.js`` and ``coordinator.js`` both
invoke it.
* Dead truncation-pill CSS removed (``.tool-output-truncated`` and
``.coord-tool-truncated``); the JS that added these elements went
away with ``TOOL_RESULT_STORAGE_CAP``.
* Tautological tests (``TestBuildHistoryAdvisoryPropagation``)
replaced with production-realistic round-trip tests built from
``wrap_tool_result(...)`` envelopes — REST and SSE-replay surfaces
pinned to the same wire shape; full DB round-trip pinned end-to-end.
Negative-tested:
* Reverting the prefix-merge in ``_flush_queued_messages`` produces
back-to-back ``user`` rows, breaking
``test_user_feedback_and_queued_coexistence_single_row_with_prefix``.
* Reverting the ``extract_advisories_from_tool_envelope`` call in
``_build_history``'s tool branch leaves the envelope verbatim in
wire content, breaking the round-trip tests.
* Reverting the wrapper-detection in ``_apply_reminders_for_provider``
entity-encodes the existing envelope's literal tags, breaking both
the string-content and list-content envelope-preservation tests.
* Reverting the ``wrap_tool_result(raw_text, advisories)`` projection
at the DB save site produces a string starting with the original
raw text, breaking
``test_tool_db_row_round_trips_list_output_with_advisories``.
Tests: 5918 passed, 3 deselected. Lint + format + mypy clean on
touched files.
(cherry picked from commit eca4bb79e4)
Round-1 ``/review`` apply-pass. Drops stale ``UserInterjection``
references from comments and docstrings that no longer describe the
post-PR drain shape, asserts the two-stream invariant in the new
queued-message persistence test, and pins the ``content.trim()`` +
``renderAssistantToolBatch`` invariants on coord-side so a future
refactor can't silently regress the Qwen3 phantom-card fix or the
chronological-order render fix.
Deferred:
* **bug-1** (back-to-back ``user`` row when ``user_feedback`` from the
approval-prompt UI callback coexists with a queued-message drain).
Reachable on strict OpenAI-compatible local templates (Anthropic and
Anthropic-via-merge-consecutive collapse fine; vLLM-hosted Mistral /
Llama enforcing role alternation can reject). The pre-PR splice
guarded against this case by riding queued items inside the tool
result envelope; that guard is what motivated the original
UserInterjection design, so the fix lane needs a deliberate decision
rather than a quick patch. Sleeping on it.
* **q-1** (delete dead ``UserInterjection`` class + tests). Held for
the bug-1 decision — if the chosen fix is to resume the splice for
the ``user_feedback``+queue coexistence case, the advisory shape
stays load-bearing. Class now carries a docstring note marking it
retained-pending-decision so a passing reader doesn't grep for
producers and assume it's actually dead.
Apply-pass content:
* ``q-2``: drop "queued user interjections" from the persistent-
advisory parenthetical in ``send``'s tool-result loop comment;
rewrite to point at ``_flush_queued_messages`` for the queue path.
* ``q-3``: ``__init__`` channel-routing comment loses "and
``UserInterjection``" — only ``GuardAdvisory`` remains.
* ``q-4``: ``_queue_tool_advisory`` docstring + the tool-error nudge
comment lose the user-interjection mentions; the docstring also now
describes the side-channel + ``_apply_reminders_for_provider``
splice path (the actual mechanism).
* ``q-5``: ``AttachmentsNotQueueableError`` docstring rewritten to
describe the post-PR ``_flush_queued_messages`` flow — the
single-combined-turn ``\n\n``-join shape can't carry image / file
blocks, and per-item separate user turns would expand the strict-
template role-ordering surface that the post-batch drain already
balances.
* ``q-6``: the new ``test_queued_message_persists_as_user_row_after_tool_batch``
in ``test_session.py`` now asserts ``stream_idx == 2`` so a future
regression where the post-batch flush runs but the send-loop short-
circuits before the next iteration surfaces in CI rather than
manual repro.
* ``q-7``: ``test_coordinator_page.py`` gets two new string-grep pins
mirroring the existing ``test_app_js.py`` shape — ``content.trim()``
on coord's assistant-replay branch and ``renderAssistantToolBatch``
for the hoisted helper that orders content card before tool batch.
## Test plan
- [x] ``ruff check`` clean
- [x] ``mypy turnstone/`` clean (189 source files)
- [x] Affected test surface (``test_session.py`` +
``test_tool_advisory.py`` + ``test_app_js.py`` +
``test_coordinator_page.py``) — 240 passed
(cherry picked from commit a032e71ff3)
Three independent rehydrate / replay regressions reported on long
multi-turn conversations after the pull-model wake stack landed.
**1. coord history replay rendered tool_calls above the assistant
narration that announced them.**
In ``coordinator.js``'s loadHistory loop, the ``role === "assistant"``
``tool_calls`` branch sat above the role switch — every assistant turn
with both narration AND tool dispatch produced ``[tool batch][content
card]`` in the DOM, even though chronological order is content first.
On a parallel fan-out (e.g. four ``close_workstream`` calls in one
turn) operators saw the assistant text "Let me close them out and
summarize" with NO tool batch between it and the next assistant
message — the four-row batch had been rendered above the announcing
text and was scrolled out of view.
Hoisted the ``tool_calls`` synthesis into a local
``renderAssistantToolBatch(m)``, called from inside the assistant
branch AFTER the content card. Live SSE order (text → dispatch →
results) now matches replay order.
**2. Whitespace-only assistant content rendered as a blank card on
replay.**
Models with vLLM's ``--reasoning-parser`` (Qwen3 in production)
strip ``<think>…</think>`` and emit only the trailing ``"\n\n"`` as
``content`` before a tool call. ``content_parts = ["\n\n"]`` saves
``content = "\n\n"`` to the conversations row. Live the user only
sees ``.msg.reasoning`` (the thinking content) — the empty
``.msg.assistant`` card lives next to it but reads as a thin
divider. On rehydrate the reasoning bubble is gone (not persisted)
and the empty assistant card is the only thing left, surfacing as
"blank cards where the assistant message was."
Both UIs now check ``content && content.trim()`` before rendering
the body — whitespace-only content skips the card entirely instead
of showing a phantom row. Live render unchanged.
**3. Queued user messages disappeared on reconnect.**
PR #474 routed queued user messages into the tool-result envelope
via ``UserInterjection`` advisories — same-turn delivery, but no
persisted user row. On page reload / cross-tab replay the
optimistic ``.msg-queued`` bubble vanished: there was no DB row to
rehydrate it.
Dropped the ``UserInterjection`` splice in ``_collect_advisories``;
the queue drains through ``_flush_queued_messages`` AFTER the tool
batch completes instead. Sequence becomes
``assistant(tool_calls) → tool … tool → user(drained)``, which is
valid for Mistral and Anthropic strict role validators (the only
forbidden shape was user injected mid-batch BEFORE the tool result,
which this still avoids). Persists a real user row → bubble survives
reconnect, and stays in the session's wire-side context window on
the next turn.
## Test plan
- [x] ``ruff check`` clean
- [x] ``mypy turnstone/`` clean (189 source files)
- [x] ``pytest -m "not live"`` — 5798 passed, 3 deselected
- [x] Updated ``test_collect_advisories_does_not_drain_queued_messages``
(was pinning the old UserInterjection shape)
- [x] Added ``test_queued_message_persists_as_user_row_after_tool_batch``
(drives ``send`` end-to-end with a queued message arriving during
the tool batch; asserts the user row lands in self.messages AND
hits ``save_message``)
- [x] Updated ``test_replay_history_renders_content_before_tool_block``
to tolerate the new ``msg.content && msg.content.trim()`` guard
- [ ] Live browser pass on coord (close_workstream parallel fan-out
rehydrates with the 4-row batch BETWEEN the announcing assistant
text and the summary) and interactive (Qwen3 ``"\n\n"`` rows no
longer paint blank cards on reload; queued bubble survives a tab
refresh)
(cherry picked from commit c11692b327)
PR #489 review feedback (Copilot + github-code-quality):
- closeSettingsPanel now closes nested revoke modal first on close-button
path (Escape was already handled by the parent keydown trap deferring
to the inner trap; missing-modal-on-close-button was an orphan-modal
hazard).
- _refreshConsentBadge now updates the settings button's aria-label +
title dynamically with the pending-consent count for screen readers
(badge stays aria-hidden — the count is in the label).
- _MAX_INSUFFICIENT_SCOPE_REPORTED promoted to public
MAX_INSUFFICIENT_SCOPE_REPORTED in mcp_http_parsers; drops cross-module
private import in mcp_oauth's /start handler.
- Stale test comment in test_session_mcp_dispatch_error.py corrected:
_exec_read_resource does not log with exc_info=True (bearer-leak
invariant).
- Rejected the protocol-method ellipsis warning: rest of _protocol.py
uses ... consistently per Protocol convention.
Lint:
- ruff format applied to test_mcp_pool_auth_integration.py and
test_mcp_pool_auth_resource_integration.py (combined `with` grammar —
pure formatting).
Flake fix — test_integration_pool_reuse_401_refresh_and_retry_succeeds
on Python 3.11 / resource-constrained CI:
Same cross-task scope hazard f6a3b66 fixed at the close side, surfacing
at the connect side. asyncio.wait_for at mcp_client.py:1206 wraps
streamablehttp_client.__aenter__ in a fresh asyncio.Task. That fresh
task enters anyio cancel scopes, completes, and dies. The eventual
stack.aclose() during eviction or auth_401 retry runs from a different
task and tries to exit scopes whose entering task is dead — anyio
raises RuntimeError, the wedged anyio state blocks the retry's stack
teardown + reconnect, and the call exceeds the 15s budget on slow
workers.
Fix: replace asyncio.wait_for with `async with asyncio.timeout(...)` so
the streamablehttp_client.__aenter__ runs in the dispatch task itself,
no fresh-task scope ownership. Aligns with invariant 18 (asyncio.timeout
not asyncio.wait_for for any SDK / AS / pool-loop await crossing anyio
scopes).
Static path (_connect_one) at lines 905 and 1000 deliberately retains
asyncio.wait_for — auth_type ∈ {none, static} is byte-identical
(invariant 1) and the narrow connect-once / no-eviction-then-reuse
pattern doesn't trigger the cross-task hazard. Anchor comments pin
both directions: a future migration there would break invariant 1; a
future revert at 1206 would re-introduce the flake.
The cited test is the symptom (non-deterministically times out under
load), not a structural gate (no deterministic asyncio.timeout
assertion exists). The comment block at line 1206 records this so a
maintainer who reverts and finds green on a fast machine doesn't
conclude the fix is unneeded.
Verified on Python 3.11.14 (/tmp/venv311) and 3.13.7 (.venv): ruff
format clean, ruff check clean, mypy clean. 368 unit tests + 30 pool
integration tests pass on both interpreters; the previously-flaky test
passed 20× in isolation on 3.11.
Multi-stage /review (4 finders × verify × dedupe): bug/security/perf
returned zero findings; quality returned 3 confirmed minor/nit items
all of which are applied here (q-1 anchor comments at 905+1000, q-2
symptom-vs-gate clarification at 1206, q-3 module-docstring sentence
in mcp_http_parsers).
(cherry picked from commit 4a3e3607be)
Wires the structured-error envelopes produced by Phase 7b's pool
dispatcher (mcp_consent_required / mcp_insufficient_scope /
mcp_*_forbidden / mcp_token_undecryptable_key_unknown /
mcp_oauth_url_insecure) through to the user-facing dashboard, and
adds a per-user settings panel for managing MCP server consents.
Changes
- ``_dispatch_pool_sync`` and ``_dispatch_pool_resource_sync`` wrap
structured-error string returns as ``RuntimeError(json_str)`` via
``_is_structured_error()`` so the session-layer ``except Exception``
branch fires uniformly across tool / resource / prompt dispatchers
(the prompt path's ``isinstance(result, str)`` shortcut works only
because prompts return ``list[dict]`` on success). Without this,
the consent UX silently does not render for tool / resource calls.
- ``_structured_error`` extended with an optional ``consent_url``
field; ``_build_consent_url`` produces ``/v1/api/mcp/oauth/start``
query strings (path-relative; the dashboard appends ``return_url``
at click time). Wired to all 12 ``mcp_consent_required`` and the
``mcp_insufficient_scope`` emit sites.
- New endpoints ``GET /v1/api/mcp/oauth/connections`` and
``DELETE /v1/api/mcp/oauth/connections/{server_name}`` registered
on both ``turnstone-server`` and ``turnstone-console``. The DELETE
handler runs local delete + audit + 204 first, then schedules the
RFC 7009 upstream revoke as a fire-and-forget ``asyncio.create_task``
with strong-ref tracking via ``_revoke_upstream_tasks`` (mirrors
the ``_pg_refresh_drain_tasks`` pattern). Soft cap of 256 concurrent
in-flight revokes prevents pile-up under coordinated mass-revoke;
the audit detail records ``upstream_revoke_outcome`` as
``scheduled | no_refresh_token | no_http_client | shed_by_cap``.
- ``ASMetadata`` extended with ``revocation_endpoint`` parsed from
RFC 8414 metadata. ``revoke_token_at_as`` helper posts the form
body under ``asyncio.timeout`` (not ``asyncio.wait_for``) and
never raises; ``_attempt_upstream_revoke`` is wrapped in an outer
``try/except Exception`` so unhandled exceptions don't surface as
``Task exception was never retrieved``.
- ``/v1/api/mcp/oauth/start`` accepts an optional ``scopes=`` query
param; tokens are validated against RFC 6749 §3.3 grammar via
``is_valid_scope_token`` (promoted to ``mcp_http_parsers``),
capped at ``_MAX_INSUFFICIENT_SCOPE_REPORTED`` (32), and unioned
with the configured server scopes for the step-up consent flow.
- Storage primitive ``list_mcp_user_token_metadata_by_user`` projects
the metadata columns at the SQL boundary so ciphertext blobs never
cross the wire on the settings-list path. New
``MCPUserTokenMetadataRow`` TypedDict in ``_protocol.py``;
``MCPTokenStore.list_user_token_metadata`` re-types to the existing
``MCPUserTokenMetadata`` shape.
- Dashboard renderer (``app.js``): ``tryParseMcpError`` detects the
envelope shape on ``tool_result`` SSE events with ``is_error=True``
and ``buildMcpErrorEmbed`` renders an action card mirroring the
existing ``buildMediaEmbed`` pattern. Three categories: actionable
(consent_required / insufficient_scope) with a ``Connect`` button
that opens ``/v1/api/mcp/oauth/start`` in a popup with a scheme
guard, forbidden (mcp_*_forbidden) with a static notice, operator
(key-mismatch / url-insecure) with an operator-action notice.
- New gear button in the appbar opens an MCP-connections settings
modal driven by ``loadMcpConnections`` / ``confirmRevokeMcp``
(two-step revoke confirmation matching the existing delete-ws
pattern). Pending-consent badge tracks unresolved consent prompts
in this tab; cleared after the connections list returns. Console
proxy collision-checked: the IIFE only prepends a node-id pill to
``header.firstChild``, so the right-anchored gear button is safe.
Bearer-leak invariant
- No ``exc_info=True`` on any new path that can carry a chained
``httpx.Request`` (revoke handler, dispatch sites, exec sites).
The two pre-existing ``exc_info=True`` calls in
``_exec_read_resource`` / ``_exec_use_prompt`` were replaced with
structured-field logs as a Phase 8 sibling fix.
Tests
- 440 pytest passes on both Python 3.13 (.venv) and 3.11
(/tmp/venv311); ruff + mypy clean.
- 5 new test files: ``test_mcp_consent_url_sibling_audit`` (structural
gate that every ``code="mcp_consent_required"`` / ``mcp_insufficient_scope``
site carries ``consent_url=``), ``test_mcp_oauth_connections``,
``test_mcp_oauth_revoke``, ``test_mcp_token_store_metadata``,
``test_session_mcp_dispatch_error``.
- End-to-end regression coverage for the bug-1 sibling pattern:
``test_call_tool_sync_raises_on_structured_error_envelope``,
``test_read_resource_sync_raises_on_structured_error_envelope``,
``test_get_prompt_sync_raises_on_structured_error_envelope``, plus
``test_call_tool_sync_does_not_wrap_non_structured_string`` as the
defensive gate (only ``mcp_*`` envelopes are wrapped).
Hard invariants honored
- Static path byte-identical for ``auth_type ∈ {none, static}``: the
wrap fires only when the dispatcher returns a structured-mcp-error
string, which only happens on the oauth_user pool path.
- ``asyncio.timeout`` (not ``asyncio.wait_for``) on every new
AS / SDK / pool-loop await per Python 3.11 anyio cancel-scope
hazard.
- Scope cap ``_MAX_INSUFFICIENT_SCOPE_REPORTED = 32`` enforced at
every output / merge site.
- Cross-user isolation on the revoke endpoint: a non-owner DELETE
returns 404 with the same body shape as a never-existed row;
``http_client_mock.post.assert_not_called()`` pins this in 3 tests.
Deferred (not Phase 8 blockers)
- perf-2 (``asyncio.gather`` parallelisation in revoke handler) —
superseded by perf-1's fire-and-forget pattern.
- q-4 (prompt-path ``isinstance(str)`` vs sibling ``_is_structured_error``
asymmetry) — already documented in the function docstring.
- q-9 (``_pendingConsentServers`` → ``_serversNeedingConsent``
rename) — pure naming taste.
(cherry picked from commit 5a3f46a1fa)
Apply sanitize_text() to the new _source and _reminders columns in
both save_message and save_messages_bulk on SQLite + PostgreSQL,
mirroring the existing pattern used for content and provider_data.
Producers (sanitize_payload on the watch dispatch path,
format_nudge constants on the standard nudge path) already strip
NUL bytes today so nothing in production reaches this clamp — but
the storage layer is opaque to those invariants, and PostgreSQL
TEXT columns reject NUL outright. Without this clamp, a future
producer that forgets sanitize_payload (or hand-builds the column
string) hard-fails the chat-loop persist path on PostgreSQL.
Cost is negligible — sanitize_text early-exits on the common
no-NUL case via 'if value and "\x00" in value'.
Surfaced by Copilot's PR #486 review.
(cherry picked from commit fc8bd6ca33)
Closes round-2 review finding q-7 (nit).
The kwarg was added to close round-1 perf-2 cosmetically — the
storage backend's signature already accepted ``limit``, but the
single in-tree caller (``ChatSession.resume``) doesn't pass it and
other tail-load consumers go direct to ``storage.load_messages``.
Adding signature surface to mark a perf finding closed without an
actual consumer is API-surface bloat.
When a tail-load consumer is written (e.g. a heuristic in
``session.resume`` to skip ancient wake rows), the kwarg can come
back — at that point with a real caller driving the contract.
(cherry picked from commit 14af6f464e)
Closes round-2 review findings q-6 (nit) and perf-1 (nit).
* **q-6:** ``_WATCH_REMINDER_OPTIONAL_KEYS`` carried a leading
underscore (Python's module-private convention) but was imported
from two other modules — clearly a public contract between
``build_watch_reminder`` and its consumers
(``ChatSession._dispatch`` + ``server._build_history``). Drop the
underscore so the import sites match the constant's documented
cross-module role.
* **perf-1:** The dispatch closure imported the constant inside its
body, paying ``IMPORT_NAME`` + ``IMPORT_FROM`` bytecode on every
watch fire. ``server.py`` already imports at module scope; hoist
the same way in ``session.py``. Microsecond savings per dispatch,
but the in-closure form was just an oversight from the apply-pass.
(cherry picked from commit 668da26dce)
Closes round-2 review findings q-1 (minor), q-3 (nit), q-4 (nit), q-5
(nit).
* **q-1:** Drop the ``post-migration 050`` clause from the fork-block
comment — the apply-pass relocated rather than removed the
tombstone-style temporal reference round-1 q-2 was supposed to fix.
The bulk-row dict shape and ``_encode_reminders`` are
self-explanatory; the WHY is pinned by
``test_fork_preserves_source_and_reminders``.
* **q-3:** Replace ``DOES persist now`` framing on the wake-row save
comment with a present-tense invariant. The ``now`` implies the
reader knows the prior state, same family as the temporal
tombstones.
* **q-4:** Trim the 12-line WHAT-narration block above the
resume-time ``_reminders_delivered = True`` loop to two lines
stating the WHY only. The new regression test pins the contract.
* **q-5:** Reframe ``test_fork_preserves_source_and_reminders``
docstring as a forward-looking invariant; drop the
``Dropping them was the original bug`` and ``post-migration 050``
fix-narration.
Project convention: invariant statements, present tense; don't
reference the current task / fix / migration number.
(cherry picked from commit b120ee2fd7)
Closes round-2 review findings bug-1 (minor) and q-2 (minor).
* **bug-1:** ``_encode_reminders`` clamped each entry's ``text`` field
with Python ``str`` slicing, which counts codepoints. Multi-byte
UTF-8 input (CJK, emoji) could land 4 bytes per character past the
cap, defeating the row-width / FTS5-index protection by up to 4x.
Switch to UTF-8 byte clamping with ``errors="ignore"`` on the
decode boundary so a slice mid-codepoint drops the partial
character cleanly.
* **q-2:** Both the constant block-comment and the ``_encode_reminders``
docstring referenced ``docs/design/watch-card-ux-briefing.md`` —
local-only per project convention (``feedback_no_design_doc_commits``)
so the canonical repo reads as a dead reference. The cap value
stands by itself; the row-width / FTS5 WHY is enough.
(cherry picked from commit 779ec638a5)
Closes round-1 review findings q-2 (minor), q-5 (minor), q-6 (nit), q-7
(nit), sec-1 (nit), perf-4 (nit).
* **q-5:** Export ``_WATCH_REMINDER_OPTIONAL_KEYS`` from
``turnstone/core/watch.py`` and import in the dispatch closure
(session.py) and the replay filter (server.py:_build_history). The
three-place duplication of the literal tuple
``("watch_name", "command", "poll_count", "max_polls", "is_final")``
is gone; future field adds touch one constant.
* **sec-1:** Run ``sanitize_payload`` over string-typed metadata fields
(``watch_name`` / ``command``) before they enter the queue. Today's
consumers all use ``textContent``, but the asymmetry — sanitised
``text`` alongside unsanitised metadata — would survive forever in
DB rows and resurface if a future consumer used a non-textContent
sink (aria-label, copy-to-clipboard, markdown render).
* **q-7:** Drop the per-iteration ``isinstance(reminder, dict)`` from
the dispatch closure's metadata comprehension. By the time the
block runs, ``text = reminder.get("text", "") if isinstance(...)``
+ the ``if not sanitized: return`` guard above already established
``reminder`` is a non-empty dict.
* **q-2:** Strip tombstone-style references — "post-#482", "post-#484",
"Step 7 of the watch-card UX plan", "Post-Step-7 dispatch surface",
and the brittle line-anchor "session.py:2685-2686" — across
``session.py``, ``test_session.py``, ``test_watch.py``,
``test_watch_dispatch.py``, ``test_watch_integration.py``. Comment
intent preserved; historical anchors gone.
* **q-6:** Drop the ``del source`` line in ``cli.py``'s
``on_user_reminder``; the parallel ``on_tool_reminder`` ignores
``tool_call_id`` without ``del`` and the comment alone is enough.
* **perf-4:** Document the SQLite ``render_as_batch=True`` recreate
cost in migration 050's docstring — first deployment after upgrade
copies the conversations table twice (one per ``add_column``).
PostgreSQL is unaffected.
5734 non-live tests pass; ruff + mypy clean.
(cherry picked from commit 7e35050b68)
Closes round-1 review findings q-3 + q-4 (minor, merged) and bug-3 + bug-4
(nit, merged).
* **q-3 + q-4:** The new ``.msg.user-reminder .msg-body { white-space:
pre-wrap }`` rule was a no-op on the interactive UI because that
frontend's ``_buildDefaultReminderBubble`` appended label + text spans
directly to the outer ``.msg.user-reminder`` element with no
``.msg-body`` wrapper. Coord rendered the same shape with a wrapper.
The two implementations diverging on DOM structure also meant a
shared-helper extraction was harder than necessary. Reconciled by
wrapping interactive's spans in ``.msg-body`` to match coord; the CSS
rule now applies to both UIs and the shared-extraction follow-up to
``shared_static/cards.js`` is mechanical (deferred per the review
report — out of scope for this commit).
* **bug-3 + bug-4:** The reminder anchor lookup ``.msg.user`` also
matched ``.msg.user.system-nudge`` markers because the marker carries
both classes. A non-wake reminder fired between a wake marker and
the next real user message would anchor below the wake marker rather
than the previous real user message. Edge case (``/history`` reload
corrects), but the fix is mechanical: change the selector to
``.msg.user:not(.system-nudge)`` in both files.
(cherry picked from commit 869135d97a)
Closes round-1 review finding perf-2 (minor).
Storage backends accept ``*, limit: int | None = None`` (see
:meth:`StorageBackend.load_messages` at storage/_protocol.py:146) but
the in-memory wrapper at memory.py:82-85 dropped the kwarg, so
callers that wanted to tail-load (e.g. ``session.resume`` against a
long-running coord with hundreds of wake rows + persisted reminder
JSON) were forced to pull every row through the wrapper anyway.
Wraparound is mechanical: signature widens, default leaves existing
callers unaffected.
(cherry picked from commit 885f6a9185)
Closes round-1 review finding q-1 (major).
The comment block above ``self._attach_pending_user_reminders(user_msg)``
asserted that reminders "stay in-memory only and don't persist across
reloads" — directly contradicted by the comment block immediately below
(at the save_message call site) that explains the new persistence
semantics, plus the actual code that now writes ``_source`` and
``_reminders`` to the conversations row. Future readers hitting both
blocks would lose trust in the surrounding comments.
The lower block already documents the persistence contract, so the
upper block is just deleted rather than rewritten.
(cherry picked from commit 81502c962f)
Closes round-1 review findings bug-2 (major), perf-1 (minor), perf-6 (nit).
* **bug-2:** ``ChatSession.resume(..., fork=True)``'s bulk-row builder
silently dropped the ``_source`` and ``_reminders`` side-channel
data the source workstream had persisted via ``_append_user_turn``.
Both backends' ``save_messages_bulk`` already accept these keys
(the columns exist post-migration 050) — the bulk builder just
didn't supply them. The fork's resumed transcript would then look
like the assistant turn answered out of nowhere: every wake marker
and every reminder bubble that survived to disk on the source got
dropped on the fork. New regression test
``test_fork_preserves_source_and_reminders`` pins the contract.
* **perf-6:** Extracts ``_encode_reminders(reminders) -> str | None``
near ``_apply_reminders_for_provider`` so the user-turn save path,
the tool-turn save path, and the new fork bulk builder share one
encoder. Eliminates the drift risk between three near-identical
``json.dumps(..., separators=(",", ":")) if X else None`` patterns.
* **perf-1:** The new helper clamps each entry's ``text`` field at
``REMINDER_TEXT_STORAGE_CAP = 8192`` characters before encoding so
a single rogue producer (a watch streaming unbounded shell output,
a corruption-class steering payload) can't blow the conversations
row width or the FTS5 index. The in-memory side-channel keeps the
full body — only the persisted JSON is clamped. Mirrors
``TOOL_RESULT_STORAGE_CAP`` on tool result rows.
5734 non-live tests pass; ruff + mypy clean.
(cherry picked from commit 91e7f2daca)
Persisted ``_reminders`` survive ``load_messages`` but the in-memory
``_reminders_delivered`` flag does not (it's session-scoped — set by
``_mark_reminders_delivered`` after each successful provider stream,
never persisted alongside the JSON column). Without a re-splice
guard at resume time, ``_apply_reminders_for_provider`` would walk
every loaded message, see ``_reminders`` set + the flag falsy, and
splice every historical ``<system-reminder>`` envelope onto the wire
on the very next user turn — leaking each reminder a second time, the
turn after it had already advised.
Mirror the post-stream hook in ``resume()``: every loaded message
that carries reminders has already been delivered (it survived to
disk), so flag it accordingly so ``_apply_reminders_for_provider``
short-circuits on the pass-through path.
Test pins the contract end-to-end — stage a workstream with a
persisted reminder, resume into a fresh session, append a live user
turn, run the wire transform, and assert the historical reminder
body does NOT land in the rendered output.
(cherry picked from commit f1466ca7e3)
User-visible slice of the watch-card UX workstream — combines the
replay-path widening, both frontend renderers, the CSS, and the
cross-cutting Python tests.
server._build_history widens the reminder filter from {type, text} to
project on a known set of optional fields (watch_name, command,
poll_count, max_polls, is_final) and surfaces _source as
entry["source"] when set. The known-key filter narrows the blast
radius if a future producer accidentally stuffs sensitive fields
into the dict.
SessionUIBase.on_user_reminder takes a new source: str | None kwarg
that rides on the SSE event when set. _attach_pending_user_reminders
forwards user_msg["_source"] so non-originating tabs see the wake's
"system_nudge" tag and render the thin marker. Protocol + cli + eval
implementations widen accordingly.
Frontend (coordinator.js + app.js — touched in lockstep per project
memory's "logic that lands in BOTH UIs must touch both files"):
* Branch on r.type === "watch_triggered" for a structured
.msg.watch-result card with header / $ command / <pre> body /
poll N/M [· final] footer.
* New addSystemNudgeMarker (interactive) + appendSystemNudgeMarker
(coord) renders a thin .msg.user.system-nudge anchor for
wake-driven reminders, both live (source === "system_nudge" on the
SSE event) and replay (msg.source === "system_nudge").
* Default .msg.user-reminder rendering preserved for every other
metacog nudge type.
CSS (shared_static/chat.css):
* New .msg.watch-result rules — full-width treatment, cyan accent,
monospace body with word-break: break-word for mobile.
* New .msg.user.system-nudge rule — thin yellow marker.
* Bonus newline-collapse fix: .msg.user-reminder .msg-body now sets
white-space: pre-wrap so multi-line shell output / bulleted lists
stay readable inside the advisory bubble.
Plan reference: docs/design/watch-card-ux.md §4 Steps 9-12 + bonus
CSS §11 (Commit 4).
(cherry picked from commit 6ae6877acc)
WatchRunner._dispatch_result now takes a structured reminder dict
produced by build_watch_reminder() — text matches format_watch_message
verbatim (so compaction / channel adapters / wire splice keep their
behaviour), and watch_name / command / poll_count / max_polls /
is_final ride alongside as queue-entry metadata.
The dispatch closure registered in ChatSession.set_watch_runner pulls
the optional fields out of the dict and passes them to enqueue via
the new metadata kwarg. Drain seams already merge metadata into the
rendered reminder dict (Commit 2), so the SSE event for a watch fire
now carries the structured fields without further plumbing.
* turnstone/core/watch.py — new build_watch_reminder() helper, _poll_watch
switches from format_watch_message + dispatch(str) to build_watch_reminder
+ dispatch(dict). set_dispatch_fn / get_dispatch_fn / restore_fn
signatures widen from Callable[[str, str], None] to
Callable[[dict[str, Any], str], None].
* turnstone/core/session.py — dispatch closure builds the metadata dict
via {k: reminder[k] for k in ("watch_name", "command", ...) if k in reminder}
and passes it to nudge_queue.enqueue.
* tests/test_watch.py — new TestBuildWatchReminder class pinning the
builder shape; existing dispatch_fn_registry / restore_fn tests
updated to dict shape.
* tests/test_watch_dispatch.py — every dispatch(...) call updated to
pass a structured reminder dict via _reminder() helper; new
TestMetadataPropagation class pins the metadata-on-enqueue contract.
* tests/test_watch_integration.py — _dispatch_result calls updated to
dict shape.
Plan reference: docs/design/watch-card-ux.md §4 Step 7 + Step 8 watch-test
subset (Commit 3).
(cherry picked from commit 13db19905a)
Producers (today only watch_triggered) can now attach a metadata dict
to a queued nudge so the rendered reminder dict on the user/tool side
carries fields beyond {type, text}. Wire shape stays additive: the
SSE event picks up the optional fields when present, and producers
without metadata leave it None.
* _Entry grows from 4 fields to 5 — metadata: dict[str, Any] | None.
* enqueue accepts metadata=... as a kwarg.
* drain returns list[tuple[str, str, dict | None]] (was 2-tuples).
* pending stays narrow at (type, text) for legacy callers; new
pending_with_metadata projects the third slot for tests that need
to assert producer-specific fields.
* Three drain consumers in session.py — _collect_advisories,
_attach_pending_user_reminders, deliver_wake_nudge_from_queue —
unpack the new 3-tuple shape and merge metadata into each
reminder dict.
* on_user_reminder / on_tool_reminder protocol signatures widen
from list[dict[str, str]] to list[dict[str, Any]] across
ChatSession.UI, SessionUIBase, CLI, eval harness.
Plan reference: docs/design/watch-card-ux.md §4 Step 6 + Step 8 _Entry
subset (Commit 2).
(cherry picked from commit 30b7e4dd24)
Adds two TEXT-NULL columns to the conversations table so multi-tab /
multi-device replay sees the same metacognitive bubble shape the
originating tab saw live. Until now, reminders lived only on the
in-memory ChatSession.messages dict, and the wake-driven empty user
turn was not persisted at all (skip at session.py:2685-2686) — a
second tab connecting via /history saw the assistant turn with no
preceding wake context, and missed every other tab's reminder
bubbles besides.
Single Alembic revision 050 (head was 049) adds:
* conversations._source — today only "system_nudge" for wake rows
* conversations._reminders — JSON-encoded reminder list
Both backends (sqlite + postgresql) thread the columns through
save_message / save_messages_bulk / load_messages. reconstruct_messages
unpacks the row tuple as 9 elements (was 7), JSON-decoding _reminders
on the user AND tool branches with the same contextlib.suppress guard
the existing provider_data / tool_calls decode uses. Tool-row
reminders ride the same column so tool_error / repeat replay shape
matches user-channel parity.
session.py:2685-2686 wake-row persist skip is dropped; _append_user_turn
JSON-encodes user_msg["_reminders"] and passes both source + reminders
to save_message. The tool-message save site at session.py:3014-3020
mirrors with metacog_reminders.
Plan reference: docs/design/watch-card-ux.md §4 Steps 1-5 (Commit 1).
(cherry picked from commit f64c3e7b10)
Address Copilot review feedback on PR #487:
1. **Atomic commit invariant**: ``_bootstrap_coord_subsystem`` previously
stamped ``coord_mgr`` ~50 lines before the final ``coord_registry``
commit, and started threads + subscriptions in between. A concurrent
dashboard request running through ``_require_coord_mgr`` during the
runtime-bootstrap window could observe ``coord_mgr`` set with
``coord_registry`` still ``None`` and surface the misleading
"Restart the console after adding a model definition" 503.
Refactored to two phases: (a) build everything as locals, (b) start
side-effects (StateWriter / observer / nudge watcher / child fan-out
/ cleanup thread), then atomic commit at the end with ``coord_mgr``
stamped LAST. The build-phase ``try/except`` rolls back any started
side-effects from local handles before re-raising — no daemon thread
or subscription leaks across retries, and ``app.state`` is never
stamped on a partial failure.
2. **Class-attr cleanup symmetry**: ``_teardown_partial_coord_subsystem``
now also clears ``ConsoleCoordinatorUI._coord_mgr`` /
``_collector`` / ``_console_metrics`` to match the lifespan shutdown
path (server.py ~line 4629). A failed bootstrap (or test teardown
reuse) no longer leaks process-global pointers at a half-built
subsystem.
3. **Lifespan startup offload**: the lifespan startup error path used
to call ``_teardown_partial_coord_subsystem`` synchronously, which
in turn calls ``StateWriter.shutdown(timeout=2.0)`` — a thread-join
+ sync DB writes that could block the event loop for up to 2s
while the console is still coming up. Wrapped the whole
load-and-bootstrap in ``asyncio.to_thread`` via the new
``_load_and_bootstrap_coord_subsystem`` synchronous helper, so all
blocking work (including any rollback) runs on a worker thread.
Mirrors the pattern the regular lifespan shutdown (line ~4620) and
the runtime CRUD-triggered path already use.
Tests:
- ``test_bootstrap_atomic_commit_no_partial_visibility``: a polling
thread in tight loop watches ``coord_mgr`` / ``coord_registry``
during a real bootstrap and asserts no observation has ``coord_mgr``
set with ``coord_registry`` still ``None``.
- ``test_real_bootstrap_rolls_back_partial_state_on_side_effect_failure``:
monkeypatches ``install_idle_nudge_watcher`` to raise mid-build,
asserts ``app.state`` shows the clean fresh-install state and the
builder-failure error string surfaces ``RuntimeError`` (not the
stale "no models" boot-time message).
(cherry picked from commit c6b4dc26be)
A freshly-installed console with no model rows in the DB at boot
caught the ``ValueError`` from ``load_model_registry()`` in the
lifespan and skipped the entire coord subsystem build, leaving
``coord_mgr`` ``None``. ``_refresh_coord_registry`` then bailed
out at ``existing is None`` rather than building the subsystem on
first model add — operators had to restart the console after
configuring their first model in the admin panel for the
"Coordinator subsystem not initialized" banner to clear.
Extract the lifespan's coord build into a reusable
``_bootstrap_coord_subsystem`` and add ``_maybe_bootstrap_coord_subsystem``
that runs as an ``asyncio.to_thread`` follow-on after every admin
model-CRUD endpoint (create/update/delete/reload). The helper:
- fast-paths to a no-op when ``coord_mgr`` is already set;
- guards concurrent first-install attempts with
``_COORD_BOOTSTRAP_LOCK`` + double-checked re-test inside the lock;
- pre-computes config-derived integers BEFORE any thread starts so
``int(config_store.get(...))`` failures don't strand a started
``StateWriter`` daemon;
- stamps ``coord_state_writer`` to ``app.state`` immediately after
``.start()`` so the new ``_teardown_partial_coord_subsystem`` can
shut it down on a partial failure (no thread leaks across retries);
- atomically commits ``coord_registry`` + clears
``coord_registry_error`` as the final step so callers can rely on
the invariant ``coord_registry`` is set iff ``coord_mgr`` is set;
- replaces the stale boot-time "no model definitions" message with
a builder-failure-specific diagnosis (carrying ``type(exc).__name__``)
on construction failure so the dashboard's 503 banner reflects the
actual cause.
Both the lifespan path and the runtime-bootstrap path now route
through the same helper and the same teardown on failure.
Tests: 12 new tests covering the helper-level wiring (idempotent
fast-path, missing-prereq parametrised over ``config_store`` /
``collector`` / ``console_metrics``, no-rows error recording, builder
failure error replacement, partial-state teardown), the endpoint
integration, the deterministic concurrent-call lock test (uses an
instrumented lock wrapper that signals when a second acquirer arrives,
so the test fails fast on slow CI rather than depending on a
wall-clock sleep), and a real-builder end-to-end case constructing a
working ``SessionManager`` against a real ``ConfigStore`` + real
``ClusterCollector``.
(cherry picked from commit 3143965e00)
Two of five Copilot comments on PR #485 were valid; this commit applies
both. The other three (one duplicate of comment 1, plus the INFO-logging
and `_pending`-naming nits) get rationale on-thread and resolution.
1. emit_oauth_failure_audit action now derived from `code` (#485 bug-1)
The Phase 7b refactor generalized `emit_insufficient_scope_audit` →
`emit_oauth_failure_audit`, routing both `mcp_insufficient_scope` AND
generic-403 (`mcp_*_forbidden`) through the same helper. The audit
`action` field stayed hardcoded as
`"mcp_server.oauth.insufficient_scope_emitted"`, mislabeling generic
forbidden events under the insufficient_scope bucket — downstream
alerting / analytics filtering on `action` would silently fold both
categories together.
The action is now selected from `code`:
* `mcp_insufficient_scope` →
`mcp_server.oauth.insufficient_scope_emitted` (preserves existing
alerting consumers)
* `mcp_tool_call_forbidden` / `mcp_resource_read_forbidden` /
`mcp_prompt_get_forbidden` →
`mcp_server.oauth.forbidden_emitted` (new, distinct label)
Detail row continues to carry both `code` and `kind` so operators get
sub-bucket distinction within either action.
2. Resource-listener docstrings cite RFC §3.2 (#485 doc-1)
Per the codebase convention established in Phase 7b round-1 q-1
(`_rebuild_user_prompt_map` corrected §3.2 → §3.3 because prompts are
§3.3 in the MCP spec), resource-related docstrings should cite §3.2.
The three resource-listener docstrings were citing §3.3, and the
"Mirrors `_notify_listeners` for tools (RFC §3.3)" parenthetical in
both `_notify_resource_listeners` and `_notify_prompt_listeners` read
as "tools are at §3.3" — confusing twice over. All four sites now
carry the correct catalog-kind citation explicitly:
* resource-listener docstrings → "RFC §3.2 (resources)"
* prompt-listener docstrings → "RFC §3.3 (prompts)"
Tests / lint:
* 119 passed on 3.13 + 3.11 (targeted MCP OAuth pool tests)
* ruff + mypy clean on both files
(cherry picked from commit 12cc052bca)
Extends the Phase 7 per-(user, server) ClientSession pool to cover
RFC §3.2 (resources/read) and §3.3 (prompts/get) on the same shape
already proven for tools/call. Pool discovery is capability-gated so
servers without resources/ or prompts/ stay free of extra round-trips.
API additions / widenings (MCPClientManager):
- ``read_resource_sync(uri, *, user_id=None, timeout=120)`` —
per-user-first dispatch; falls through to the byte-identical static
path when ``user_id`` is None or the URI doesn't resolve to an
``oauth_user`` pool entry.
- ``get_prompt_sync(prefixed_name, arguments=None, *, user_id=None,
timeout=30)`` — same dispatch shape; structured-error responses
surface via ``RuntimeError`` so the agent-loop's ``except Exception``
block renders the JSON without polluting the prompt-protocol return
shape.
- ``get_resources(user_id=None)`` / ``get_prompts(user_id=None)`` —
per-user merged catalogs (admin/global call still passes None).
- ``add_{resource,prompt}_listener`` /
``remove_{resource,prompt}_listener`` — ``user_id`` keyword scopes
the listener so a pool-only catalog change for one user does not
wake another user's session.
- ``resource_count_for_user(user_id=None)`` /
``prompt_count_for_user(user_id=None)`` — method-form variants used
by ChatSession's ``read_resource`` / ``use_prompt`` tool gating; the
legacy ``resource_count`` / ``prompt_count`` properties remain
static-only for admin paths.
- ``_dispatch_pool_resource`` / ``_dispatch_pool_prompt`` async coros
— mirror ``_dispatch_pool`` for the new SDK calls; share the
carrier-race-and-cancel core via ``_dispatch_pool_with_entry_call``.
- ``_handle_auth_403`` extended with ``kind=Literal["tool",
"resource", "prompt"]`` so the per-operation ``mcp_*_forbidden``
code surfaces (kind="tool" remains the default for back-compat).
- Pool notification handler now refreshes resources / prompts on
``ResourceListChangedNotification`` / ``PromptListChangedNotification``
via ``_refresh_pool_server_resources`` / ``_refresh_pool_server_prompts``.
ChatSession (``turnstone/core/session.py``) call-site updates:
- 12 sites threaded the session-bound ``user_id`` through
``add_*_listener`` / ``remove_*_listener``, ``get_resources`` /
``get_prompts``, gating, ``read_resource_sync`` /
``get_prompt_sync``, and ``is_mcp_prompt`` so the per-user merged
catalog drives both the visible-tool set and dispatch.
- ``/mcp`` slash command now lists this user's pool resources and
prompts alongside tools (Phase 7 already scoped tools).
Scope decisions:
- Per-user-first URI ordering (decision 0.1): the dispatcher attempts
the user's pool catalog first, falling back to the static catalog
only when no pool entry resolves the URI / prefixed name. Pool-only
users never see the static catalog leak into their resolution.
- Method-form ``*_count_for_user`` (vs property) keeps the legacy
``resource_count`` / ``prompt_count`` properties intact for admin
endpoints whose contract is "static catalog size only".
- Shared ``_dispatch_pool_with_entry_call`` helper accepts an
``sdk_call: Callable[[ClientSession], Awaitable[Any]]`` closure,
keeping the entry-locked carrier-race / classification / retry
plumbing single-source instead of a 3x copy across tool / resource
/ prompt paths.
R6 (anyio uniformity): every pool-side list / read / get path uses
``async with asyncio.timeout(...)`` — ``asyncio.wait_for`` is
forbidden in those paths because it wraps the inner awaitable in a
fresh task and surfaces ``CancelledError`` from inside
``streamablehttp_client``'s anyio TaskGroup on Python 3.11
(per ``feedback_asyncio_timeout_vs_wait_for.md``).
Tests:
- ``test_mcp_pool_auth_resource_integration.py`` — 9 real-transport
resource tests (FastMCP upstream + ``BehaviorMiddleware``):
401-refresh-retry success, persistent 401 -> consent_required,
403+insufficient_scope, 403 generic -> mcp_resource_read_forbidden,
breaker-isolation under repeated auth failures, missing-token,
decrypt-failure, http:// URL guard, unknown-URI ValueError.
- ``test_mcp_pool_auth_prompt_integration.py`` — 9 mirror tests for
the prompt path; structured-error responses verified via
``RuntimeError`` payload shape.
- ``test_mcp_user_catalog.py`` — extended unit coverage for per-user
resource / prompt rebuild + collision policy + symmetric eviction.
- ``test_sessions.py::TestMCPToolGating`` — pool-only-user canary
asserts ``read_resource`` / ``use_prompt`` stay visible when the
static catalog is empty but the user has pool entries.
Round-1 review fixes (4-finder review applied, no push yet):
- bug-1: ``_exec_use_prompt`` was hardcoding ``"MCP prompt error: failed
to invoke prompt"`` — discarding the structured-error JSON that
``_dispatch_pool_prompt_sync`` raises via ``RuntimeError``. Now uses
``f"MCP prompt error: {e}"`` mirroring ``_exec_mcp_tool``; pool-prompt
consent_required / insufficient_scope / forbidden errors now reach
the LLM as intended.
- bug-2 + bug-3: resource template discovery was uncapped —
``_cap_server_resources`` covered ``res_result.resources`` but the
separate ``tmpl_result.resourceTemplates`` loop appended every
template a server returned. Added ``_MAX_RESOURCE_TEMPLATES_PER_SERVER``
(1000) + ``_cap_server_resource_templates`` helper, applied at both
the initial discovery site (``_connect_one_pool``) and the refresh
site (``_refresh_pool_server_resources``). Mirrors the existing
``_MAX_TOOLS_PER_SERVER`` / ``_MAX_PROMPTS_PER_SERVER`` defensive
ceilings.
- sec-1 + sec-2: ``emit_insufficient_scope_audit`` generalized to
``emit_oauth_failure_audit(kind, code, ...)``, called from both the
insufficient_scope branch AND the previously-silent generic 403
branch. Audit detail now records ``{"kind": kind, "code": code,
"scopes_required": [...]}`` so operators can distinguish tool-call
vs resource-read vs prompt-get 403s in audit logs and so cross-
tenant probing on the generic 403 path leaves a trail. The Phase 7
inherited gap (``mcp_tool_call_forbidden`` had the same silence) is
closed in the same refactor.
- perf-1: pool resource discovery now uses ``asyncio.gather(
list_resources, list_resource_templates)`` inside the existing
``async with asyncio.timeout(...)`` budget — disjoint catalogs, no
ordering dependency. Typical-case 2-RTT cold-connect resource block
collapses to 1-RTT. Same change applied at ``_refresh_pool_server_resources``.
- q-1: ``_rebuild_user_prompt_map`` docstring corrected RFC §3.2 →
§3.3 (resources are §3.2; prompts are §3.3).
- q-2: ``_refresh_pool_server_prompts`` docstring now carries the
R6 / mcp-loop note that the resource sibling already had — both
refresh paths now declare the asyncio.timeout invariant explicitly.
- q-5: added the ``_user_resource_map`` / DB-mismatch guard to
``read_resource_sync`` for parity with ``get_prompt_sync``. A stale
per-user map entry with no matching oauth_user row now raises a
specific ValueError instead of silently falling through to a
generic ``Unknown MCP resource``.
- q-6: ``_dispatch_pool_with_entry`` (now a single-caller wrapper
after the ``_dispatch_pool_with_entry_call`` extraction) gains a
one-line docstring explaining why the wrapper is preserved
(tool-decode localization + stack-trace identity for debugging).
- q-7: added 1 resource + 1 prompt end-to-end integration test that
drive REAL discovery + dispatch in the same connect (no
``_seed_pool_*_map`` shortcuts), mirroring the tool path's
``test_integration_pool_reuse_401_refresh_and_retry_succeeds``.
The seeded-map tests stay (faster, focused on dispatch); the new
e2e tests cover the connect-discover-dispatch composition that
caught Phase 6's carrier-on-entry bug.
Pre-push round-1 review fixes (3-finder review on the final state —
the lesson from Phase 7 round-3's q-1 regression: round-2 catches
what the round-1 apply pass missed):
- q-1 (MAJOR): the bug-1 sibling that round-1 missed —
``_exec_read_resource`` was hardcoding ``"MCP resource error: failed
to read resource"`` while ``_exec_use_prompt`` (post-bug-1) preserved
the structured-error JSON via ``f"... error: {e}"``. The round-1
apply pass patched the prompt side but not the resource side. q-5's
per-user-map / DB-mismatch ValueError was being swallowed at the
agent loop boundary, defeating the operator-diagnostic intent. Now
``_exec_read_resource`` mirrors ``_exec_mcp_tool`` and ``_exec_use_prompt``.
- q-6 (nit): defensive-cap comment block at module-level cited
"(RFC §3.2)" while covering both resource and prompt list paths;
prompts are §3.3. Now reads "(RFC §3.2 for resources, §3.3 for
prompts)" matching the convention the q-1 apply established.
- q-5 (rejected with better justification): the reviewer flagged
``_dispatch_pool_with_entry`` as a single-caller wrapper that should
be inlined. After examination — the autouse fixture
``tests/test_mcp_pool_auth_introspection.py::_install_capture_intercept``
monkeypatches this method to stash ``entry.auth_capture`` for the
fake call_tool stubs in dispatcher-asserting tests. Inlining would
redirect the patch to ``_dispatch_pool_with_entry_call`` (different
kwargs shape) and require re-validating every test that depends on
the interception. The wrapper IS load-bearing; q-6 docstring updated
to cite the test-fixture rationale instead of the thin "stack-trace
identity" claim.
Deferred to follow-up (documented rationale):
- perf-2: single-pass partition for system-message resource list
(concrete vs templates). Sub-microsecond at expected scale;
opportunistic-only.
- q-2 (pre-push): ~200 lines of fixture infrastructure
(``BehaviorMiddleware``, ``_build_server``, ``_seed_oauth_server``,
``running_loop_mgr``, etc.) duplicated across three pool-integration
test files. Real maintenance cost, but a 200-line conftest extraction
is a focused refactor that earns its own commit / PR. Tracking as
follow-up rather than balloon Phase 7b's diff further.
- q-3 / q-4 (refactor): extract shared dispatcher / scheduler
helpers to compress three near-identical 90-line bodies (round-1
q-3 was the same root cause; the pre-push q-3/q-4 reviewer
reaffirmed it concretely). Three named methods preserve readability
for the codebase's hottest correctness path; follow-up if
duplication grows further or if a per-path divergence ships.
- q-4 (round-1, distinct from pre-push q-4): split pool concerns
into ``mcp_pool.py``. Out-of-scope per finder; future refactor as
the file approaches the navigation/merge-conflict threshold.
3.13: 5590 passed (5541 baseline -> +49 net; pre-review +47, q-7
e2e tests added +2). Existing audit-detail tests updated in-place
to expect the new ``kind`` and ``code`` fields.
3.11: 5590 passed (parity gate per ``feedback_pytest_env_parity.md``).
(cherry picked from commit 124615cce0)
Closes PR #484 review findings (Copilot): the soft-cap pattern in
``ChatSession.set_watch_runner``'s dispatch closure was a non-atomic
two-call pair (``count_by_type`` then ``drop_oldest_by_type``) with
two separate lock acquisitions. A concurrent drain on the worker
thread (``USER_DRAIN`` / ``TOOL_DRAIN`` consuming ``"watch_triggered"``
entries via the ``"any"`` channel) could slip between the two calls,
making the drop a no-op. The dispatch closure also discarded
``drop_oldest_by_type``'s return value and unconditionally logged
``dropped_oldest=True``, so a no-op drop got reported as a successful
drop.
* New ``NudgeQueue.cap_at_or_drop_oldest(nudge_type, max_depth,
channel=None) -> bool`` does the count+drop in a single critical
section. Returns the actual outcome.
* Dispatch closure (``session.py:1410-1416``) now calls the helper and
uses its return value to gate the WARNING log line, so the log is
accurate when a drop did NOT happen.
* ``drop_oldest_by_type``'s docstring no longer overstates the
per-call lock as covering a count+drop pair — it points readers
to ``cap_at_or_drop_oldest`` for that contract.
7 new tests in ``TestCapAtOrDropOldest`` cover: below-cap no-op,
at-cap drop-oldest, above-cap drop-only-one (per-call), channel
filter, other-type isolation, ``max_depth <= 0`` defensive no-op,
no-match.
5708 non-live tests pass; ruff + mypy clean.
The github-code-quality bot finding ("Statement has no effect" on
``_protocol.py:939``'s ``...`` body) is a false positive — every
Protocol method in ``_protocol.py`` uses ``...`` as its body, which
is the canonical Python Protocol pattern. Replacing with ``pass``
would diverge from the file's existing style. No code change.
(cherry picked from commit c757c22f55)
Closes round-2 review findings q-3, q-4, q-5, q-7.
* **q-4:** ``_NAME_CONTROL_CHARS`` and ``_PAYLOAD_CONTROL_CHARS`` shared
7 lines of Unicode-steering character classes (zero-width / bidi /
separators / BOM / tag chars above BMP). Factored into a single
``_CONTROL_CHARS_TAIL`` constant; each regex now differs only in its
leading ASCII range. Future bidi or zero-width additions edit one
place.
Side effect: this corrects a latent bug where ``_NAME_CONTROL_CHARS``
had two literal ASCII spaces in place of U+2028 / U+2029 (line and
paragraph separators) — visible as ``r" "`` in source but rendered
as the actual codepoints in ``_PAYLOAD_CONTROL_CHARS``. After the
factoring both regexes correctly include U+2028 / U+2029, closing
the gap that would have let a workstream name with embedded line
separators forge a sibling bullet (the same vector ``\n`` was
blocked for in the original bug-1 fix).
Switched to ``\u`` escapes for readability (and to keep future Edit
tool runs against this block reliable).
* **q-3:** Tombstone clause "standing in for the deleted
``_watch_pending`` maxsize bound" survived in
``ChatSession.set_watch_runner``'s docstring after the apply-pass
trim cleaned the inline soft-cap comment. Dropped.
* **q-5:** ``test_newline_in_name_does_not_forge_extra_bullet`` carried
five WHAT-narration comments restating what the immediately-following
asserts already say. Dropped — the docstring carries the security
invariant; the assertions speak for themselves.
* **q-7:** ``patch_session_storage`` had a 14-line docstring including
fallback-guidance and self-justification ("accumulated 7 near-duplicate
sites"). Trimmed to a 3-line contract.
(cherry picked from commit 39e0f930c1)
Closes round-2 review findings q-1, q-2, q-6.
* **q-1:** ``test_valid_until_drops_when_watch_missing`` collapsed to the
same code path as ``test_valid_until_drops_when_watch_inactive`` after
the apply-pass switched the predicate from ``get_watch[active]`` to
``is_watch_active`` (both stubbed via ``patch_session_storage(active=False)``).
The "missing" case has no distinguishable branch at the dispatch
layer, so dropping it removes a tautological duplicate. The
missing-row mapping moves to the storage layer (q-2 below) where it
IS distinguishable.
* **q-2:** ``is_watch_active`` was a new public storage primitive with
zero direct backend coverage — only via-session-via-stub coverage.
New ``TestIsWatchActive`` in ``tests/test_watch_storage.py`` covers
active row → True, inactive row → False, missing row → False.
Pinned at the storage boundary so future backend changes fail loudly
there instead of in the dispatch tests.
* **q-6:** Concurrency test had ``n_threads = 2`` alongside two literal
Thread objects and a tautological ``assert len(threads) == n_threads``.
Threads are now built from a labels tuple, so ``len(threads)`` drives
the slack bound; the redundant assertion is gone.
(cherry picked from commit 751ed9c85f)
Closes review findings bug-4 and q-6.
bug-4 — the watch dispatch concurrency test bounded depth at
``_WATCH_QUEUE_SOFT_CAP + 2 * per_thread`` (= 250) which is
tautologically true: two threads × 100 fires can append at most 200
entries above the cap, so the bound asserted nothing more than what
``depth <= 2 * per_thread`` already says. Tighten to
``_WATCH_QUEUE_SOFT_CAP + N_THREADS`` (= 52): the count-then-drop window
admits at most one slip per concurrent thread.
q-6 — 7 near-duplicate ``monkeypatch.setattr(session_mod, "get_storage",
lambda: _StubStorage())`` sites across ``test_watch_dispatch.py`` +
``test_watch_integration.py`` (4 different stub shapes, mostly trivial
variations on the active flag). Lift a ``patch_session_storage``
helper into the existing ``tests/_helpers.py`` with kwargs for the
common cases (``active``, ``raise_on_is_active``), returns the call list
so call-shape assertions still work. Tests collapse from ~10-line
inline-class blocks to one-line helper calls.
(cherry picked from commit 20c4dfaca6)
Closes review findings q-2 and q-5.
q-2 — ``bound_watch_id = watch_id`` rebind was unnecessary. ``_dispatch``
is constructed fresh per fire (not in a loop), so ``_still_active``
closes over the function parameter directly without any
loop-variable-capture risk. Drop the rebind.
q-5 — the inline soft-cap comment restated rationale already covered by
the ``_WATCH_QUEUE_SOFT_CAP`` block-comment at module scope and dragged
in a tombstone reference to the deleted ``_watch_pending`` path. Trim
to one line stating only the WHY (drop-oldest because latest output is
most useful). Leave the ``set_watch_runner`` docstring's operational
detail at lines 1356-1378 alone — trimming further risks losing the
``valid_until`` predicate semantics.
(cherry picked from commit 28d9bb4802)
Closes review finding q-4.
The closure built inside ``server.py``'s ``_watch_restore_fn`` is the
new contract surface introduced by the switchover — it constructs a
fresh ChatSession, calls ``session.resume(ws_id)`` to adopt the
original ws_id, re-registers the dispatch closure via
``set_watch_runner``, and returns ``WatchRunner.get_dispatch_fn`` for
the runner to invoke directly. No automated coverage exists today;
a future refactor (e.g. swapping ``manager.create + session.resume``
for ``manager.open``) could silently break the watch-restore pipeline.
Adds ``test_watch_dispatch_through_restore_fn_lands_on_rehydrated_session``
to ``tests/test_watch_integration.py`` — drives the full restore path:
persists a kickoff message for the original ws_id, fires
``_dispatch_result`` against a runner with no registered dispatch fn,
asserts the restore_fn ran exactly once, the rehydrated session is a
distinct object that adopted the original ws_id, and the watch payload
landed on the rehydrated session's NudgeQueue (not on the original).
(cherry picked from commit ed1eaee216)
Closes review finding perf-1.
The watch dispatch closure's ``valid_until`` predicate fires once per
watch entry at every drain seam — on the chat-loop hot path. It only
needs the ``active`` flag, but ``storage.get_watch`` runs a full-row
``SELECT *`` and marshals the result into a dict. At the typical drain
depth (cap-50 + a busy chat loop) that's ~50 throwaway dict allocations
per drain pass for one boolean.
Adds ``StorageProtocol.is_watch_active(watch_id) -> bool`` plus
SQLite + Postgres implementations doing a single-column
``SELECT active FROM watches WHERE watch_id = ?`` (returns False on
missing row). ``_still_active`` in ``ChatSession.set_watch_runner``
now calls that instead of indexing into the full row.
Test stubs that mocked ``get_watch`` for the predicate are converted
to mock ``is_watch_active`` directly. Bulk variant deferred — single-row
fix is sufficient at typical drain depths.
(cherry picked from commit 3b495eba15)
Closes review findings perf-2, q-3, bug-3.
The watch dispatch closure's soft-cap pre-check materialised the whole
queue snapshot via ``pending(channel="any")`` only to throw away the
text and count the type — wasteful at typical drain depths (cap-50 +
mixed producers means a 50-tuple allocation per fire just to read a
length). The other half of the cap pair (``drop_oldest_by_type``)
walked the *whole* queue regardless of channel, so a future producer
that enqueued ``"watch_triggered"`` on a different channel could be
dropped by the watch cap, and vice versa — silently surprising once
that producer existed.
Adds ``NudgeQueue.count_by_type(nudge_type, channel=None) -> int`` that
walks ``_items`` once under the queue lock without materialising
tuples; extends ``drop_oldest_by_type`` to take an optional ``channel``
filter so both halves can agree on the entry set being capped. The
watch dispatch closure now passes ``channel="any"`` to both —
consistent with where the closure enqueues — so a future channel split
can't bleed across producers.
Adds ``TestCountByType`` mirroring the existing ``TestDropOldestByType``
shape, plus a ``test_drop_oldest_by_type_channel_filter`` case pinning
the new optional argument's behaviour.
(cherry picked from commit e5e6e13307)
Closes review finding q-1.
The live-marker scaffold in ``tests/test_watch_live.py`` couldn't actually
run as written: the ``live_client`` / ``live_model_id`` fixtures it
referenced live in ``tests/test_server_live.py`` at ``scope="module"``,
not on a shared ``conftest.py``, so the file would have ImportError'd
at collection if anyone ever tried ``pytest -m live`` against it.
Lifting the fixtures into a shared conftest is a larger refactor
than R9 justifies — the deterministic envelope-arrival contract is
already pinned end-to-end by ``test_watch_fires_then_user_send_drains_envelope``
and ``test_three_back_to_back_watch_fires_drain_into_one_turn`` in
``test_watch_integration.py`` (real ChatSession + real WatchRunner +
real chat-loop drain). The model-quality-of-response leg is genuinely
manual; the plan doc's R9 entry is updated locally to reflect that
deferral.
(cherry picked from commit 68a44cc7e2)
Closes review finding bug-1.
The shared ``sanitize_payload`` regex preserved TAB/LF/CR so multi-line
watch shell output kept its layout — necessary for the watch path, but a
correctness gap for the idle_children formatter, which renders the
user-controlled ``name`` field as a single bullet item. A child name
with an embedded ``\n`` would split the bullet across two rendered rows
and let a hostile name forge a fake sibling entry in the listing.
Splits the regex in two: ``_NAME_CONTROL_CHARS`` strips TAB/LF/CR
(used by the new ``sanitize_name`` helper for single-line name fields),
``_PAYLOAD_CONTROL_CHARS`` keeps the existing permissive shape (used by
``sanitize_payload`` for multi-line watch payloads).
``format_idle_children_nudge`` now calls ``sanitize_name``.
Adds ``test_newline_in_name_does_not_forge_extra_bullet`` — feeds a
hostile name with embedded ``\n`` + bullet-shaped continuation, asserts
the rendered listing still has exactly N bullet rows for N children
(no forged sibling), and the hostile newline got flattened to an inline
space. Adds a ``TestSanitizeName`` class mirroring the existing
``TestSanitizePayload`` shape for the new strict variant.
(cherry picked from commit e596650a5c)
The deleted comment claimed the closure may be registered "under the
rehydrated workstream's id, which may differ from the original ws_id we
restored against" — but ``ChatSession.resume(ws_id, fork=False)`` adopts
the parameter as the session's id at session.py:1682, so they match
exactly post-resume. The lookup works because the ids are equal, not
because they may differ.
The accessor name ``get_dispatch_fn`` is self-explanatory; no replacement
comment is needed (per the project's "default to no comments" rule).
(cherry picked from commit d2028aa4f7)
Adds two boundary-crossing integration tests and one live-marker
scaffold for the watch switchover landed in the previous commits:
tests/test_watch_integration.py — drives a real ChatSession + real
WatchRunner end-to-end (LLM stubbed) through the unified pull-model
chat-loop drain seam. Pins:
- test_watch_fires_then_user_send_drains_envelope: a synchronous
WatchRunner.dispatch fire enqueues "watch_triggered" on "any";
session.send drains the entry into the user message's _reminders
side-channel — confirms the envelope splice path.
- test_three_back_to_back_watch_fires_drain_into_one_turn: pins the
intentional behavioural delta from the plan section 3.4 / risk
register R3 — N back-to-back fires now produce ONE assistant turn
with N _reminders entries, not N successive turns.
tests/test_watch_live.py (new file, single test, marked @pytest.mark.live):
risk register R9 verification recipe — confirm a real LLM handles a
<system-reminder>-framed watch payload sensibly. Collects under the
regular -m "not live" run; the user runs it on demand against an
Anthropic-backed config.
Implements watch-switchover plan section 5.2 (integration) and step 11
(live scaffold).
(cherry picked from commit 17c62f7ef3)
Replaces the deleted tests/test_watch_dispatch.py with a focused
14-test suite exercising the closure that ChatSession.set_watch_runner
now constructs (per the previous commit's switchover). Each test
pins one assertion:
- enqueue shape: ("watch_triggered", text, "any") on the per-session
NudgeQueue; not on user / tool channels
- producer-side sanitisation strips control / bidi / zero-width chars
and angle-bracket tag breakers; preserves TAB/LF/CR so multi-line
shell output keeps its layout (R8); empty-after-strip → no enqueue
- soft-cap drop-oldest at _WATCH_QUEUE_SOFT_CAP with a queue_full
WARNING log; non-watch entries on the same queue are not collateral
damage
- valid_until predicate drops on inactive / missing / storage-raises;
delivers when active (counter-test)
- concurrent enqueues across two threads stay bounded under the
3-acquisition count-then-drop window
Implements watch-switchover plan section 5.1 / step 9. No production
changes — pure test rewrite.
(cherry picked from commit 7ca00b564c)
Replaces the bespoke _make_watch_dispatch / _watch_pending /
_dispatch_pending_watch / _MAX_WATCH_CHAIN machinery with a single
NudgeQueue.enqueue("watch_triggered", ...) call inside
ChatSession.set_watch_runner. Watch results now drain at the same
<system-reminder> envelope seams as every other metacog nudge
(USER_DRAIN, TOOL_DRAIN, IdleNudgeWatcher IDLE wake) — no separate
worker-spawn, no recursive watch chain, no per-session queue.Queue.
The dispatch closure built inside set_watch_runner carries:
- producer-side sanitize_payload over the whole formatted message
before enqueue, so steering-vector / control-char shell output
can't tamper with the envelope at interpolation time
- a soft cap of 50 entries on per-session "watch_triggered" depth
via the new NudgeQueue.drop_oldest_by_type, replacing the prior
_watch_pending maxsize=20 + _MAX_WATCH_CHAIN=5 bounds; drop policy
is drop-oldest (latest output most useful), logged at WARNING
- a valid_until predicate that re-checks
storage.get_watch(watch_id)["active"] at drain time so a cancelled
watch's last splat doesn't ride out a future wake
Behavioural delta documented in the plan section 3.4: N back-to-back
watch fires now drain into ONE assistant turn responding to all N
(via the envelope splice) instead of N separate send turns. This is
intentional — fewer model invocations for noisy watches, and uniform
with the rest of the metacog pull-model surface introduced by #482.
Implements watch-switchover plan steps 5-8. Server-side simplifications
let the previously-load-bearing _make_watch_dispatch (47 lines), its
session_worker.send import, and the chat-loop _dispatch_pending_watch
seam at the no-tools IDLE branch all disappear. The obsolete
tests/test_watch_dispatch.py and the wake-tag test in test_session.py
(both pinning contracts that no longer exist) are removed; the
NudgeQueue-based replacement plus an integration test land in the
following commit.
(cherry picked from commit 94ed79d488)
Widens the per-workstream dispatch fn signature from ``(message,)``
to ``(message, watch_id)``. The runner now passes the originating
``watch_id`` through ``_dispatch_result`` so dispatch closures can
capture per-watch metadata at fire time — the upcoming switchover
needs this for the ``valid_until`` predicate that re-checks
``storage.get_watch(watch_id)["active"]`` before a stale entry rides
out a wake.
Also adds ``WatchRunner.get_dispatch_fn(ws_id)`` as the public
accessor used by the server-side restore path to retrieve the
closure that ``set_watch_runner`` constructed during workstream
rehydrate (avoiding private-attr access into ``_dispatch_fns``).
Implements watch-switchover plan step 4 plus risk register R4.
The pre-existing single-arg callers (``_make_watch_dispatch`` and
``set_watch_runner``'s ``dispatch_fn=`` fallback) get replaced
in the next commit; their mypy types are ``Any`` today so the
type mismatch isn't caught at this step.
(cherry picked from commit 195ff985cc)
Renames _sanitize_child_name to sanitize_payload and widens it to be
the shared producer-side sanitiser for both idle_children and the
incoming watch_triggered nudges. The regex now skips TAB / LF / CR
so multi-line shell output rendered into a watch payload keeps its
line structure when sanitised as a whole formatted message — the
pre-switchover code path collapsed multi-line output to one line.
Adds the watch_triggered entry to _NUDGE_MAP alongside idle_children
so ``_NUDGE_MAP``-as-registry consumers (should_nudge gating, future
audit / UI tagging) recognise the type. Body is empty — payload
comes from the producer (the watch dispatch closure), same shape as
idle_children.
Implements watch-switchover plan section 3.2 plus risk register R8
(TAB/LF/CR exclusion) and step 3 (_NUDGE_MAP registration).
(cherry picked from commit 78ae7ae6b5)
Adds an atomic drop-oldest-by-type operation to NudgeQueue used by
producers that need a per-type soft cap on their own queue depth.
The watch dispatcher (next commit in this stack) is the first user:
when "watch_triggered" saturates, the dispatch closure drops its
oldest entry under the queue lock so the count snapshot and drop
can't interleave with a concurrent enqueue from the same producer.
Implements watch-switchover plan section 3.1 — the producer-side soft
cap takes the place of the deleted _watch_pending maxsize=20 bound.
Other producers (idle_children, advisories) have natural rate limiters
already, so the helper is opt-in per producer rather than a global cap
in enqueue itself.
(cherry picked from commit 74f1958e47)
Three Copilot findings on PR #483 (commit dad98c0); one rejected as a
false positive.
- mcp_client.py:1189 — pool notification handler's exception path
used ``log.warning(..., exc_info=True)`` which serializes the
chained ``httpx.Request.headers`` carrying ``Authorization: Bearer
<token>`` into Sentry / faulthandler frame captures. Same threat
model as the round-1 sec-1 dispatch-path fix, applied to a site
the original review missed. Now logs structured fields only
(server, user, exc type) without ``exc_info``.
- mcp_client.py:1202 — ``_connect_one_pool``'s handshake step used
``asyncio.wait_for(session.initialize(), ...)``, the same Python
3.11 + anyio cross-task-cancel-scope anti-pattern that the
Phase 7 round-3 q-1 fix removed from the discovery step (and that
f6a3b66 originally addressed for ``_safe_close_stack``). Pre-
existing Phase 5 code, but the same latent bug class — a 401
during initialize() under 3.11 would surface ``RuntimeError:
Attempted to exit cancel scope in a different task`` as the
SDK's TaskGroup unwinds. Switched to ``async with asyncio.timeout(...)``
matching the discovery step's pattern.
- mcp_client.py:1522 — renamed loop tuple-unpack variable
``_server_name`` → ``server_name`` in ``_rebuild_user_tool_map``.
The leading underscore conventionally signals "intentionally
unused", but the variable is read at the assignment a few lines
below. Two other ``_server_name`` unpacks in this file (1410,
3111) genuinely don't use the value and keep the underscore.
Rejected as false positive:
- test_mcp_user_catalog.py:58 (github-code-quality bot, "Statement
has no effect"): ``await task`` inside ``contextlib.suppress(
BaseException)`` is the standard pattern for cleanly draining a
cancelled task. The bot's static analysis treats ``await`` of a
result that's discarded as a no-op statement, but ``await`` here
triggers cancellation propagation and waits for the task to
finish — load-bearing in the fixture's teardown. No change.
Verified on Python 3.11 (``/tmp/venv311``) and 3.13 (``.venv``):
ruff + mypy clean, full test suite green.
(cherry picked from commit 62909d402c)
Light up production reachability of pool dispatch (RFC §3, invariant 8)
by widening the public catalog API to optionally take a ``user_id``:
- ``MCPClientManager.get_tools(user_id=None)`` returns the merged
static + per-user pool view when ``user_id`` is supplied; the default
preserves the legacy global-only contract.
- ``is_mcp_tool(name, *, user_id=None)`` extends the lookup to the
per-user ``_user_tool_map``. Pool tools become reachable from
``ChatSession._prepare_tool`` only when the session-bound user_id
flows through — flipping invariant 8 from "must hold" to "satisfied".
- Listener identity becomes ``(user_id, callback)``. Static-path
changes fire ALL listeners (admin + every user); pool-entry
changes fire only matching-user + admin (``None``) listeners.
RFC §3.3.
- Pool sessions discover their tool list on first connect
(``_connect_one_pool`` → ``await session.list_tools()``); the
notification closure binds to ``(user_id, server_name)`` so
push-driven ``list_changed`` updates target the correct user's
catalog. R6 verified empirically: ``list_tools()`` 401 propagates
through anyio TaskGroup unwinding, no hang — plain ``await`` is
fine, no carrier-race shape needed for discovery.
- ``_evict_session`` drops ``entry.tools`` and rebuilds the user's
index so an evicted-then-reconnected session doesn't carry
stale catalog state.
- ``web_search.resolve_web_search_client`` refuses
``auth_type=oauth_user`` backends (per-node web search can't
carry per-user tokens).
Resources / prompts pool dispatch deferred to Phase 7b — invariant 8
is satisfied by the tool path alone, and the resource/prompt path
needs sibling ``_dispatch_pool_resource_sync`` /
``_dispatch_pool_prompt_sync`` helpers each with their own
carrier-race plumbing (~400 LOC). Phase 7b will follow the patterns
established here.
CLI sessions default ``user_id=""`` and so cannot use oauth_user
MCP servers — documented limitation; users must use the web UI.
Round-1 review fixes (4-finder review applied, no push yet):
- bug-1: get_tools(user_id) was iterating _user_pool_entries from sync
threads while the mcp-loop concurrently mutated it (RuntimeError:
dictionary changed size during iteration). Now reads from a sibling
_user_tools dict updated atomically by _rebuild_user_tool_map.
- bug-2: _close_pool_entry_if_idle (LRU/TTL eviction) skipped the
catalog cleanup that _evict_session does — stale tools persisted
in _user_tool_map and ChatSession's tool list never rebuilt. Now
mirrors _evict_session.
- perf-1: _last_pool_notification_refresh debounce dict was never
pruned in either eviction path. Now popped alongside the entry.
- perf-3: web_search resolver was issuing a sync SQL query per LLM
turn to gate oauth_user backends. Now reads from the cached
in-memory config.
- sec-1: bearer token could leak into exc_info-rendered tracebacks
via Sentry/faulthandler. log.debug now uses structured fields,
not exc_info.
- sec-2: tools-per-server response now capped at 1000 (defensive,
mirrors _MAX_ERROR_LEN / _MAX_INSUFFICIENT_SCOPE_REPORTED).
- Test cleanup: dropped two listener fan-out tests duplicating
test_mcp_client.py coverage; renamed test_pool_session_notification_handler
to match its actual scope (_refresh_pool_server_tools); removed
stale comments referencing /tmp/r6-spike*.py scratchpads and a
misleading "copy-on-write" comment.
Round-2 pre-push review fixes (focused single-pass review applied):
- round2-1: bug-2's catalog-cleanup block in _close_pool_entry_if_idle
had no integration test (exactly the failure mode flagged in
feedback_tests_through_boundaries.md). Added
test_close_pool_entry_if_idle_clears_catalog_and_fires_listener
driving the LRU/TTL eviction path through real streamablehttp_client +
MockTransport. Negative-test verified: reverting the
_rebuild_user_tool_map / _notify_user_tool_listeners calls makes
the new test fail.
- round2-3: documented the _oauth_user_server_names cache invariant
in add_server_sync / remove_server_sync docstrings. Cache is
reconcile_sync's sole owner — direct callers leave it stale, but
_db_servers_to_config strips oauth_user rows so production paths
are unaffected. Static→oauth_user transitions correctly leave the
name in the cache because remove_server_sync drops the static
connection, not the cache identity.
- round2-6: strengthened test_rebuild_user_tool_map_populates and
test_rebuild_user_tool_map_drops_empty_user to assert on the
_user_tools sibling cache (bug-1 fix). Without this, a future
revert dropping the sibling write would still pass the unit
tests because get_tools coverage lives in separate tests.
Round-3 full-stack review fixes (multi-stage review on the final
state caught what the layered apply passes missed):
- q-1 REGRESSION: pool tool-discovery used asyncio.wait_for around
session.list_tools(), the exact pattern the f6a3b66 fix (and
feedback_asyncio_timeout_vs_wait_for.md) put in place to avoid.
Python 3.11's asyncio.wait_for wraps the inner coroutine in a
fresh task → cross-task scope-exit when the SDK's anyio TaskGroup
unwinds on a 401. Switched to `async with asyncio.timeout(...):`
pattern used by _safe_close_stack.
- sec-2: TOCTOU in _connect_one_pool — entry.tools was published
(via _rebuild_user_tool_map + listener fan-out) BEFORE entry.session
was assigned. A sync-thread reader could observe a tool whose
backing entry has session=None. Defence-in-depth — dispatch
re-fetches its own token and lazy-reconnects on session=None — but
reordering catches the race at the source. entry.session now
publishes BEFORE catalog visibility.
- bug-1: _close_pool_entry_if_idle's _user_pool_locks.pop ran
unconditionally after the try/finally, but the early-return
branches (entry None on re-check, in_flight > 0 under lock) skip
it via Python's return-through-finally semantics. The lock was
never popped on those paths. Now gated behind an `evicted` flag
set only on the success path; in_flight > 0 leaves the lock for
the active dispatcher to reuse, entry-None races leave the lock
for re-allocation by _ensure_pool_entry. Comment now describes
the actual semantics, not the original promise.
- bug-2: softened the _rebuild_user_tool_map docstring's atomicity
claim. The two-dict write is technically non-atomic across Python
statements; in practice the window is sub-microsecond on the
mcp-loop with no awaits between writes, and the listener fan-out
fires AFTER both writes complete. Docstring now says "back-to-back
on the mcp-loop" instead of "atomically alongside".
- q-3: dropped `hasattr(mcp_client, "server_auth_type")` defensive
check in web_search.py. The method ships in this commit; the
hasattr created a silent fallthrough that would let a future
rename silently re-enable oauth_user backends.
- q-4: surfaced the CLI / empty-user_id limitation in a docstring
comment at ChatSession.__init__'s self._user_id assignment. The
note previously lived only inside is_mcp_tool's docstring — a
future maintainer wiring CLI features against MCP pool servers
wouldn't think to read is_mcp_tool to find the constraint.
- q-2 + q-5: deleted a tautological duplicate test in
test_mcp_user_catalog.py whose docstring claimed to test
ChatSession.close but never instantiated a ChatSession (the
manager-level identity semantics are already covered by
test_listener_identity_includes_user_id in the same file and by
test_session_close_removes_listener_with_same_user_id in
test_mcp_client.py which DOES drive a ChatSession). Reworded a
misleading "fixture provides only 5s" comment to point at the
actual `_run_on_loop(..., timeout=5)` site.
- q-6: the `self._user_id or None` collapse repeated at 8 sites
across session.py. Cached once at __init__ as
``self._mcp_user_id`` (since ``_user_id`` is set once and never
mutated); 8 call sites now read the cached value. The empty-
string-is-CLI-sentinel invariant is documented at the assignment
site, not re-asserted at each consumer.
Deferred to follow-up:
- sec-1: a hostile MCP server bound to user-A could craft a
tool.name containing `__` to synthesize a prefixed-name collision
in user-A's own catalog. Bounded impact: cross-tenant dispatch is
prevented by the per-tenant token gate in _dispatch_pool, and
user-B's get_tools(user_id="B") never includes user-A's pool
entries. The fix needs policy decisions (reject vs. sanitize)
and touches _mcp_to_openai which is shared between static and
pool paths; better discussed in its own follow-up where the
policy applies uniformly to static-path servers too. The threat
model already requires user-A to have consented to a malicious
server, who has many more dangerous vectors than tool-name
shenanigans.
Test count delta: +31 tests (5435 → 5466, ``-m "not live"``; one
test deleted in round-3 apply per q-2):
- ``tests/test_mcp_client.py`` +20 (per-user catalog state, listener
identity, session thread-through)
- ``tests/test_mcp_user_catalog.py`` +9 NEW (integration tests
driving real ``streamablehttp_client`` + ``httpx.MockTransport`` per
invariant 14: discovery on connect, user isolation, eviction +
reconnect, LRU/TTL eviction (round2-1), R6 401-propagation
regression, static byte-identical canonical regression; review
passes dropped duplicate listener fan-out tests from earlier
drafts whose coverage lived in test_mcp_client.py)
- ``tests/test_web_search.py`` +2 (oauth_user backend rejection +
static backend acceptance regression; updated to use the new
``server_auth_type`` in-memory accessor)
(cherry picked from commit a8b34bfe54)
Three confirmed findings from the PR #482 bot review pass.
* **Copilot (idle_nudge_watcher.py)**: ``IdleNudgeWatcher`` was gating
wake dispatch on ``len(_nudge_queue) == 0`` (any channel), but
``deliver_wake_nudge_from_queue`` only drains ``USER_DRAIN``. A
``"tool"``-channel entry queued by ``_queue_tool_advisory`` would
pass the gate, spawn a wake daemon, and immediately no-op at the
drain guard — repeating on every IDLE event for as long as the
tool entry sat unconsumed. No correctness bug (the no-op return
prevents bad state) but a wasted thread spawn per IDLE. Fixed by
gating on ``has_pending(USER_DRAIN)``; tool-only queues no longer
trigger the wake path.
* **Copilot (coordinator_idle_observer.py)**: docstring referenced
the old module path ``turnstone.core.metacognition.IdleNudgeWatcher``;
the class moved to ``turnstone.core.idle_nudge_watcher`` in q-3 of
the apply-pass.
* **Copilot (nudge_queue.py)**: ``has_pending`` docstring cited
``ChatSession.deliver_wake_nudge_from_queue`` as its caller, but
that method calls ``drain(USER_DRAIN)`` directly — no production
caller used ``has_pending`` until this commit. Updated to point
at the now-actual caller (``IdleNudgeWatcher``).
* **github-code-quality (test_nudge_queue.py)**: false positive on
``test_channel_is_required`` — the no-channel ``q.enqueue("a", "1")``
call is wrapped in ``pytest.raises(TypeError)`` to verify the
validation contract. No code change.
5571 non-live tests pass; ruff + mypy clean.
(cherry picked from commit 0fbf31e713)
Round-2 review caught 11 confirmed findings on the 3-commit metacog stack;
this commit applies them.
* **bug-1 (major)**: Wake source tag was leaking onto real user messages
flushed during a wake send. ``_append_user_turn`` and ``send`` now
take an explicit ``from_wake: bool`` parameter — only the wake's
synthesized first turn passes True, so ``_flush_queued_messages``'s
real user input no longer inherits the audit tag. Regression test
pins the contract.
* **perf-1 (major)**: ``CoordinatorIdleObserver._maybe_enqueue`` was
issuing list_workstreams + visible_memory_count storage queries
before the cheap cooldown gate could short-circuit. New
``_cooldown_allows`` read-only peek runs first; storage queries only
fire when cooldown actually allows the nudge.
* **q-1 (major)**: Added the missing coord-side integration test that
exercises ``CoordinatorIdleObserver`` + ``IdleNudgeWatcher`` together
in the production install order against a real ``SessionManager``,
protecting the subscription-order contract from silent regression.
* **perf-2/3 (minor)**: Cap check moved above ``_last_assistant_used_wait``;
``_fire_counts`` restructured as ``dict[str, dict[str, int]]`` keyed by
ws_id so the leave-IDLE existence check is O(1).
* **perf-4 (minor)**: ``NudgeQueue.drain`` fast-paths the all-match
case (the common one for chat-loop drain seams) by swapping
``self._items`` directly instead of allocating a fresh ``kept``
deque + per-entry append.
* **perf-5 (minor)**: Wake's synthesized empty user turn no longer
writes a content-empty row to the conversations table — the
``_source`` audit tag isn't column-backed and the side-channel
reminder is stripped before persist, so the row would carry nothing.
* **q-3 (minor)**: Split ``IdleNudgeWatcher`` + ``install_*`` /
``shutdown_*`` helpers out of ``metacognition.py`` into the new
``turnstone/core/idle_nudge_watcher.py``; metacog stays a
static-template module.
* **sec-1 (nit)**: Widened ``_sanitize_child_name``'s control-char
regex to cover Unicode bidi-overrides, zero-width chars,
line/paragraph separators, BOM, and tag chars.
* **q-4/q-5 (nits)**: Docstring referenced the wrong peek primitive
(``has_pending`` → ``len()``); ``_last_assistant_used_wait``'s
``session`` parameter now typed ``ChatSession``.
5571 non-live tests pass; ruff + mypy clean.
(cherry picked from commit 3f106f98b2)
Adds the first concrete consumer of the wake trigger: when a coordinator
goes IDLE while interactive children are still running, a
``CoordinatorIdleObserver`` enqueues an ``idle_children`` nudge that the
``IdleNudgeWatcher`` then dispatches as a synthetic empty-user-turn
``send``. The model receives a system-reminder body listing the active
children (capped at 6 inline + 32 in the suggested ``wait_for_workstream``
call) and a nudge to block on them rather than reply prematurely.
Observer gates (in order): coord-only filter, skip if last assistant
turn used ``wait_for_workstream``, per-(ws, nudge_type) hard cap (3)
that resets only on non-wake leave-IDLE, active-children query,
``should_nudge`` cooldown. Console lifespan registers the observer
BEFORE the watcher so subscriber-fire order has the observer
enqueueing first on the same IDLE event.
Adds an opt-in ``valid_until`` predicate on ``NudgeQueue.enqueue``
(R9 from the design risk register) — drain re-checks the predicate
outside the queue lock; falsy / raising drops the entry without
delivering it. ``deliver_wake_nudge_from_queue`` now drains inline
before synthesizing the empty user turn so a stale predicate-drop
doesn't leave the wake send with empty content; ``_attach_pending_user_reminders``
consumes the pre-drained reminders via ``_wake_drained_reminders``.
The observer's ``valid_until`` uses ``count_workstreams_by_state``
(boolean check, no row fetch) instead of full ``list_workstreams``,
keeping the chat-loop user-attach path off the heavy query.
User-controlled child workstream names are sanitized
(``_sanitize_child_name``) before interpolation so a name like
``</thinking>...`` can't steer the model's reasoning channels through
the rendered body — the wire-boundary ``escape_wrapper_tags`` only
covers ``<system-reminder>`` / ``<tool_output>`` envelopes.
(cherry picked from commit 908e67fe4f)
Adds the third metacog channel: an out-of-band wake that converts a
workstream's IDLE transition into a synthetic empty-user-turn ``send``
when the session has any-channel nudges queued. The ``IdleNudgeWatcher``
subscribes to ``SessionManager.subscribe_to_state``; on IDLE it dispatches
via ``session_worker.send`` with a no-op ``enqueue`` callback so a
busy-worker race silently drops without spawning a competing worker.
Wake-source-tag plumbing on ``ChatSession`` short-circuits metacog
detection on the synthetic empty input, suppresses queue producers
during the wake's own tool dispatch, and stamps ``_source = "system_nudge"``
on the synthetic user-message for audit / replay distinction. The tag
is saved / restored across ``_dispatch_pending_watch`` so watch chains
recursing off the wake are processed as normal user turns rather than
inheriting the wake's guards.
Generic ``install_idle_nudge_watcher`` / ``shutdown_idle_nudge_watchers``
helpers wire the watcher into both the interactive and coord lifespans
via a single ``app.state`` registry so both surfaces share the same
teardown contract.
Foundation for PR 3 (CoordinatorIdleObserver + idle_children formatter)
and PR 4 (watch dispatcher switchover).
(cherry picked from commit f0e7fea549)
Replaces the dual `_pending_user_advisories` / `_pending_tool_advisories`
list pair with a single channel-tagged `NudgeQueue` per session.
Producers tag entries with a channel ("user", "tool", or "any");
consumers drain by channel filter at their existing seams. Foundation
for the wake trigger (PR 2) and coordinator idle-children nudge (PR 3).
Existing nudges (start, correction, completion, denial, resume,
tool_error, repeat) keep their wire shape and drain timing — zero
behavior change. Cancel paths now `clear()` the unified queue.
(cherry picked from commit 94b3720916)
Python 3.11's ``asyncio.wait_for`` wraps its inner coroutine in a fresh
``asyncio.Task`` via ``ensure_future``. When the inner is
``stack.aclose()`` on an ``AsyncExitStack`` containing
``streamablehttp_client(...)`` (anyio cancel scopes entered in the
calling task), the fresh task's attempt to exit those scopes raises
``RuntimeError('Attempted to exit cancel scope in a different task
than it was entered in')``. Python 3.12+ rewrote ``wait_for`` to use
``asyncio.timeout`` internally — runs in the current task — so 3.13
ran the same code path successfully.
Symptom on 3.11: integration tests where ``session.initialize()``
returns 4xx (e.g., 403 insufficient_scope tests) hit
``_connect_one_pool``'s ``except Exception:`` handler →
``_safe_teardown_on_connect_failure`` → ``_safe_close_stack`` → cross-
task RuntimeError. The ``concurrent.futures._base.CancelledError``
that surfaces in ``future.result(timeout=...)`` is the cascade
fallout from the asyncio loop's exception handler reacting to the
unretrieved-task-exception.
Fix: use ``asyncio.timeout`` instead of ``asyncio.wait_for`` for the
5s aclose bound. Equivalent semantics, current-task execution, works
on 3.11+. The 5s guard against ``aclose()`` hanging on a broken stack
is preserved.
Verified on Python 3.11.14 (full suite 5427 passed) and 3.13.7 (full
suite 5427 passed); all 9 integration tests pass on both.
Pre-existing bug — surfaced only after the marker fix in 5c9850c
let CI's test (3.11) actually run the 4xx tests.
(cherry picked from commit f6a3b66ea4)
Two pre-existing defects in the Phase 6 pool dispatch path that only
manifest when a pooled session is reused for a second dispatch:
1. The per-dispatch _AuthCapture allocated in _dispatch_pool was wired
into the httpx response hook only at first connect (via
_connect_one_pool). On a reused session no fresh connect runs, so
the hook continues writing to the original-connect's carrier while
the new dispatch inspects an empty carrier — auth_401/403 silently
misclassified to "other", refresh-and-retry never fires.
2. Even with the carrier on the entry (so the hook writes to a stable
reachable object), session.call_tool itself hangs forever on
upstream 4xx for reused sessions. Trace: SDK's spawned
handle_request_async raises HTTPStatusError, the outer
streamablehttp_client TaskGroup cancels post_writer, post_writer's
finally aclose's read_stream_writer, BaseSession's _receive_loop
exits and enters its CONNECTION_CLOSED-fanout finally. anyio's
send_nowait skips waiting receivers with pending_cancellation; the
dispatch task (created by run_coroutine_threadsafe for the reuse
case) is NOT in any cancel-scope chain, so the send "delivers" but
the receiver's Event is set on stale state — receive() never
wakes. Test 21 doesn't hit this because its 401 happens during
initialize, in the same task that opens streamablehttp_client, so
the cancel scope DOES propagate.
Fix:
- Move _AuthCapture ownership to PoolEntryState (and asyncio.Event
alongside, allocated lazily on the mcp-loop). The hook closes over
entry.auth_capture at first connect and stays valid across
dispatches; reset under open_lock before each call_tool.
- Race session.call_tool against the carrier's fired_event in
_dispatch_pool_with_entry. If the event wins (hook captured 4xx
before SDK propagated), cancel call_tool and raise an internal
_CarrierAuthSignal — _classify_failure resolves to auth_401/403
via the carrier's status, the dispatcher evicts the broken
session, and the cross-task retry handshake reconnects on a fresh
bearer.
Adds tests/test_mcp_pool_auth_integration.py::test_integration_pool_reuse_401_refresh_and_retry_succeeds
which drives the reuse path through real upstream + real SDK and is
the structural gate against this class regressing. Negative-tested
twice: revert PoolEntryState.auth_capture → test fails (carrier
empty); revert the race → test times out (SDK hang).
Also drops the @pytest.mark.asyncio decorator (replaced with
@pytest.mark.anyio) on four tests in test_mcp_pool_auth_introspection.py.
The project depends on anyio's pytest plugin (anyio is in deps);
pytest-asyncio is NOT a project dep and CI's test (3.13) failed on
those four. Local pytest happened to pick it up via system Python.
Found via Copilot review on PR #481.
(cherry picked from commit 97086fc617)
Phase 6 of OAuth-MCP. Recovers upstream 401/403 from MCP servers via a
capturing httpx_client_factory: an async response hook records 4xx
status + WWW-Authenticate header into a per-dispatch carrier before
the SDK's post_writer swallows the underlying httpx.HTTPStatusError.
Splits _classify_failure into auth_401 (refresh-and-retry once) vs
auth_403 (parse insufficient_scope, emit mcp_insufficient_scope with
parsed scope set). The 401 retry runs on a fresh asyncio.Task via
run_coroutine_threadsafe in _dispatch_pool_sync, escaping the anyio
cancel-scope state of the prior dispatch's TaskGroup.
WWW-Authenticate parsing extracted to a new mcp_http_parsers module
with an RFC 7235 challenge tokenizer (replaces hand-rolled substring
scanners). Two-layer defense against multi-Bearer-challenge injection:
the hook uses get_list("www-authenticate")[0] to drop attacker's
second challenge, the parser truncates at challenge boundary as
belt-and-braces. Scope set capped at 32 entries before hitting the
audit row or the LLM-visible structured-error JSON.
Auth failures (401/403) never trip the per-server circuit breaker
(server-only breaker invariant). Static path remains byte-identical.
_PgRefreshLock untouched. Pool dispatch still reachable from the
agent loop only via Phase 7 catalog scoping; Phase 6 behaviour is
testable via direct call_tool_sync.
5557 tests pass. 33 tokenizer unit tests in tests/test_mcp_http_parsers
cover the RFC 7235 grammar + the scope/error wrappers + the 4 KB input
cap. 7 integration tests in tests/test_mcp_pool_auth_integration drive
real upstream 401/403 through streamablehttp_client + a FastMCP
subprocess fixture — the structural exit gate that makes
HTTPStatusError-injection-only unit tests insufficient.
(cherry picked from commit db9260d8c4)
Models often emit page references in the standard man-page form
(``printf(3)``, ``open(2)``, ``perlfunc(3pm)``) rather than splitting
them into ``page`` + ``section`` args. The page-name sanitizer was
rejecting the parens as invalid input, killing the call. Parse the
section out of the page string before sanitization (explicit
``section`` arg still wins) and widen the section validator to accept
multi-letter suffixes like ``3pm`` / ``3perl`` that already appear on
real systems.
(cherry picked from commit 39a6b7b447)
Phase 5 PR #479 review fix-up. Three review rounds (bot + two internal
multi-stage /review) caught:
- _PgRefreshLock now allocates a per-instance ThreadPoolExecutor instead of
a module-global single-worker one. The global shape preserved psycopg2
thread-affinity but serialized every advisory-lock acquire on the node
behind one thread, even for unrelated (user, server) keys.
- get_user_access_token_classified flips to `async with lock, pg_lock:` so
concurrent same-key callers serialize on the in-process asyncio.Lock
before allocating the pg_lock's per-instance executor + spin loop. N
concurrent same-key callers collapse to one executor allocation.
- _drain_orphan_pg_lock no longer re-awaits the cancelled asyncio Future
from `__aenter__`. It receives the underlying concurrent.futures.Future
and re-wraps it via asyncio.wrap_future, getting an independent asyncio
Future tied to the worker outcome. This way cancellation of the awaiter
doesn't poison the drain's wait, and the drain genuinely waits for the
worker to settle before deciding whether to call cm.__exit__.
- Module-level _pg_refresh_drain_tasks set holds strong refs to in-flight
drains (asyncio's task set is weak — fire-and-forget tasks could be GC'd
mid-cleanup; RUF006 hazard).
- Drain narrows except clauses to Exception so a drain-task cancellation
records as cancelled instead of being silently logged as 'completed
normally with no acquire'.
Test integrity (was a major finding in round 2 — old generator-based cm
let the test pass via GC finalization timing rather than drain logic):
- New _ObservableLockCm class-based context manager whose __exit__ is a real
observable method (records call args + thread). Distinguishable from
GeneratorExit thrown by GC of a generator-based cm.
- Strong external ref to the cm via created_cms list — keeps cm alive past
the test's awaits, so a no-op drain genuinely fails the assertion rather
than papering over via GC timing.
- Deterministic drain wait via _pg_refresh_drain_tasks gather — no
fixed-duration sleeps.
- _run_cancel_scenario helper drops the duplicated setup between the two
cancellation tests.
Negative-test verified: replacing _drain_orphan_pg_lock body with `return`
makes test_pg_refresh_lock_cancellation_releases_on_same_thread fail with
'drain did NOT call cm.__exit__ — orphan Postgres lock + open transaction'.
Other fixes: protocol docstring corrected to describe pg_try_advisory_xact_lock
spin + retry (was claiming pg_advisory_xact_lock blocking acquire);
get_user_access_token_classified docstring rewritten for new lock order;
narrow `except BaseException` -> `except Exception` in
test_mcp_user_pool.py concurrent-dispatch helper.
882 tests pass (MCP + auth + storage). ruff + mypy clean.
(cherry picked from commit 3eb9d22ad5)
Phase 5 of OAuth-MCP — adds a per-(user, MCP-server) ClientSession
pool to MCPClientManager alongside the existing static-server path,
gated entirely on the per-server `auth_type='oauth_user'` config.
Pool architecture:
- `_user_pool_entries: dict[(user_id, server_name), PoolEntryState]`
with lazy connect on first dispatch, per-key asyncio.Lock allocated
on the mcp-loop, idle eviction coroutine (default 600s TTL, LRU cap
200), and an `in_flight` counter as the eviction interlock so live
calls can never be torn down mid-flight.
- `_dispatch_pool` runs the token-state machine: missing token →
`mcp_consent_required`; key-rotation decrypt failure →
`mcp_token_undecryptable_key_unknown` with NO consent prompt and NO
auto-delete; expired token → silent refresh under per-(user, server)
advisory lock; refresh failure → revoke + consent.
- `_classify_failure` separates transport (trips breaker) from auth
401/403 (does NOT trip breaker — server-only invariant) from
protocol (no breaker change).
- `entry.open_lock` held only across connect-or-reuse and released
before the `await session.call_tool` so concurrent calls from one
user against one server overlap (validated by Spike 1 scenario 2).
Auth-class failures are fail-soft in Phase 5: any 401/403 surfaced by
the SDK propagates to the agent as a tool error and the next dispatch
reconnects on a fresh refresh. Real introspection of upstream 401/403
is a Phase 6 concern — the MCP SDK's `streamable_http` post_writer
swallows `httpx.HTTPStatusError` upstream, so detecting status from
the response chain requires `McpError(CONNECTION_CLOSED)` payload
parsing or a custom httpx middleware around `streamablehttp_client`.
The mid-flight 401 refresh-retry path and the `mcp_insufficient_scope`
structured error for 403 step-up land together in Phase 6, gated by
an integration test that drives a real upstream 401/403 (the unit-
test injection of `HTTPStatusError` is what masked the production gap
on the first apply-findings pass — the integration test is the
structural gate so the gap can't reopen). RFC §1.5 steps 4-5 and the
phase table in §Implementation phases reflect this scope split.
Multi-node refresh contention:
- New `StorageBackend.acquire_advisory_lock_sync` Protocol method.
SQLite returns nullcontext (single-node, in-process asyncio.Lock
is sufficient). Postgres uses `pg_try_advisory_xact_lock` with
retry on a fresh per-attempt connection, so waiters don't pin pool
connections during the AS roundtrip. Inner try/except + nested
finally ensures conn is always returned to the pool, even when
begin / execute / yield / commit raises mid-body.
- Lock ordering: pg_advisory outer, asyncio.Lock inner. Re-read after
lock collapses cluster-wide contention to one HTTP roundtrip per
(user, server) per refresh window.
- `_PgRefreshLock` enter/exit pinned to a single-worker
ThreadPoolExecutor so SQLAlchemy connection state stays
thread-affine across cancellations.
Token storage refactor:
- `get_user_access_token_classified` returns a tagged TokenLookupResult
(Token / MissingToken / DecryptFailure / RefreshFailed) so the
dispatcher maps each state to the right user-facing error.
- `get_user_access_token` is now a thin wrapper around the classified
variant; the previous duplicated state machine is gone.
Security:
- Pool dispatch + admin endpoints reject `http://` URLs for
`auth_type='oauth_user'` servers (only exact loopback hostnames are
exempt — `*.localhost` is intentionally NOT honored because RFC 6761
localhost-zone resolution is configuration-dependent and could route
bearers to non-loopback IPs via custom resolvers / hosts file /
Docker overlays). Validated at three layers:
`_dispatch_pool` (structured `mcp_oauth_url_insecure` error),
`_connect_one_pool` (defensive ValueError), and
`admin_create_mcp_server` / `admin_update_mcp_server` (400 before
storage write).
- Admin URL change on an oauth_user row purges per-user OAuth tokens
bound to the old URL: bearers are bound (via OAuth resource /
audience) to the URL active at consent time, so silently rebinding
them to a new URL is a token-binding violation. Re-consent forces
fresh issuance for the new resource.
- Encryption-key fingerprints stay in audit logs only; no longer
surfaced in agent-facing error payloads.
User_id thread-through:
- `MCPClientManager.call_tool_sync(..., user_id=None)` (additive;
default None preserves the static path byte-identically).
- `ChatSession._exec_mcp_tool` passes `self._user_id or None`.
- `set_app_state(app_state)` setter wires OAuth state at lifespan
startup, called from both turnstone-server and turnstone-console.
Performance:
- LRU cap eviction iterates `_user_pool_entries` (not
`_user_pool_last_used`) so pre-dispatch entries are eligible.
- Eviction batch closes via `asyncio.gather` instead of serial await.
- `_resolve_pool_target` returns the resolved server row to
`_dispatch_pool` to eliminate the second DB lookup.
- Production reachability of pool dispatch is gated on Phase 7
(catalog scoping) wiring pool tools into `_tool_map`; until then
pool dispatch is reachable only via direct `call_tool_sync` with a
prefixed name (the path the new pool tests exercise).
Hardening parity preserved:
- Static path (auth_type ∈ {none, static}) byte-identical; PR #296
hardening (SDK #2147 mitigations, anyio cancel-scope, stale-session-
and-stack guard, server-only circuit breaker) intact.
- `test_reconnect_preserves_static_state_identity` unchanged + green.
- `MCPTokenStore.get_user_token` does not auto-delete on
MCPTokenDecryptError (key-rotation safety).
- Notification debounce stays manager-level.
- Connect-failure cleanup factored into
`_safe_teardown_on_connect_failure` shared by both connect paths.
Tests: 5475 → 5493 (+18). New file `tests/test_mcp_user_pool.py`
plus additions to test_mcp_oauth_refresh.py, test_mcp_admin_api.py,
and test_mcp_client.py covering: pool data structures, lazy connect,
eviction TTL + LRU + lock interlock, dispatch state machine (token
states), failure classification, http-rejection at dispatch and
admin layers, URL-change-purges-tokens (sec), concurrent dispatch on
one (user, server), pg_advisory lock parity, and user_id threading.
Phase exit criterion (synthetic load test 50 users × 3 servers × LRU
30 × 1000 calls × 200 evictions) deferred to a post-Phase-5 fitness
spike that runs against a staging deployment with real FDs and real
network behaviour, not a CI mock — same shape as Spike 1's
pre-Phase-0 SDK validation.
Out-of-scope for Phase 5 (Phase 6+): SDK-level 401 refresh-retry +
403 `mcp_insufficient_scope` (Phase 6), per-user catalog scoping
(Phase 7), consent UX SSE event + dashboard renderer (Phase 8),
admin UI status indicators (Phase 9).
(cherry picked from commit 4db7d9c6cf)
Spike artifact validating MCP SDK behavior before Phase 5 builds the
per-(user, MCP-server) ClientSession pool. Three scenarios, all pass:
1. N=20 concurrent ClientSession instances against the same URL — no
FD blow-up, no shared transport state, each session's tools/list
returns independently.
2. Two concurrent tools/call on a shared ClientSession with
interleaving payloads — request_id demux works under contention.
3. Per-session Authorization header isolation across 5 sessions —
httpx connection pooling does not cross headers between sessions,
so per-session bearer tokens reach the server unmixed.
Outcome gates the Phase 5 architecture (lazy dict[(user_id,
server_name), ClientSession] + per-key asyncio.Lock + LRU eviction).
Had any scenario failed, the fallback was per-call header injection
(Alternative F in the OAuth-MCP RFC).
Spike-only — not collected by pytest. Run manually:
uv run python tests/spike_sdk_concurrency.py
(cherry picked from commit e695a98c54)
Addresses ten findings on the Phase 4 OAuth-MCP commit: four from the
PR #478 review surface, plus six surfaced by a follow-up multi-stage
review of the first round of fixes. Two of the latter were genuine
security regressions in the very code that claimed to close those
holes.
Security
--------
- _validate_return_url now pins return_url same-origin against the
configured oidc_config.redirect_base instead of request.url. Behind
a permissive front proxy that did not normalise Host, an attacker
could spoof Host and provide a matching absolute return_url to mint
an open redirect off /api/mcp/oauth/start. Same fix pattern as
PR #476 OIDC.
- Reject return_url values containing literal backslashes or starting
with `//` up front. urlparse leaves backslashes inside `path`, so a
value like `/\evil.example/foo` slipped through the path-only branch
and became the protocol-relative `//evil.example/foo` after WHATWG-
conformant browsers normalised the backslash — re-introducing the
open redirect the same-origin pin was meant to close.
- internal_mcp_status (read-scoped) projects through a new
_strip_server_status_for_read helper that drops the verbose `error`
text and replaces it with a coarse `has_error` boolean. The error
string is built as `f"{type(exc).__name__}: {exc}"` and so carries
stdio binary paths (FileNotFoundError) or internal MCP URLs
(httpx.ConnectError) — equivalent to leaking command/url, which
this same patch deliberately strips. Approve-scoped refresh and
reconnect callers continue to receive the full `error` text via
the existing _strip_server_status helper.
- internal_mcp_status now returns the projected (sanitised) entries
for every server in mcp_mgr.get_all_server_status() instead of
emitting the un-sanitised dict that included `command` (stdio argv)
and `url` (remote MCP endpoint). Sibling refresh/reconnect endpoints
already used _public_server_status to strip these.
- internal_mcp_status docstring documents the trust boundary — server
enumeration to read scope is intentional so dashboards can render
per-server indicators; verbose error detail and command/url remain
approve-scoped.
Correctness / UX
----------------
- _validate_return_url comparison normalises (scheme, host, port)
before equality. Lowercases hostname and collapses the scheme's
default port, so `https://App.Example.COM/x` and
`https://app.example.com:443/x` are recognised as same-origin
with `redirect_base = https://app.example.com` instead of being
silently downgraded to the `/` fallback.
- mcp_crypto startup-gate error message now names both
`mcp_token_encryption_keys` (rotation list) and
`mcp_token_encryption_key` (single) so an operator using rotation
isn't misled into thinking only the singular form is valid.
Cleanup
-------
- Delete the unused _KNOWN_TRUSTED_ENDPOINT_HOSTS legacy re-export
shim in oidc.py (zero callers — a no-op that survived the Phase 4
oauth_ssrf extraction). Sphinx :data: docstring reference at
validate_discovered_endpoint updated to point at
turnstone.core.oauth_ssrf.KNOWN_TRUSTED_OAUTH_ENDPOINT_HOSTS
directly. The Google multi-origin allowlist is unaffected — it
lives at the canonical name and is read from oauth_ssrf.py:164.
- test_mcp_oauth_handlers TestValidateReturnUrl imports
_validate_return_url at module level instead of repeating the
import inside each test method.
- test_server_lifespan_mcp_crypto replaces a fragile
`messages.count("mcp_token_encryption_key") >= 2` substring trick
with `re.search(r"mcp_token_encryption_key(?!s)", messages)` —
asserts the singular form directly via negative lookahead.
Tests
-----
5448 pass (+13 vs the prior tip):
- TestValidateReturnUrl gains backslash-bypass, protocol-relative,
default-port, uppercase-host, and explicit-port-mismatch cases
alongside the original same-origin / cross-origin / scheme-
mismatch / path-only cases.
- TestInternalMcpStatusEndpoint asserts the `error` text never
reaches the read-scope wire (binary-path FileNotFoundError no
longer appears anywhere in the rendered response) and that the
coarse `has_error` boolean lights up correctly on the failed
server.
- TestInternalMcpStatusEndpoint also pins the no-mcp-client path to
`{"servers": {}}`.
- _routes_with_internal extended to include the
/api/_internal/mcp-status route so the new tests can exercise it
through TestClient.
- Existing test_startup_aborts_with_oauth_user_row_and_no_key
strengthened to require both singular and plural key names appear
in the error log.
(cherry picked from commit 62bbc332af)
Lands the OAuth flow that uses the token-at-rest store from the prior
commit: discovery (RFC 9728 PRM + RFC 8414 AS metadata with operator-
override precedence), PKCE S256 (mandatory — refuse AS without it),
RFC 8707 resource indicator on every authorize and token request,
RFC 7591 minimal one-shot dynamic client registration, authorization-
code exchange, refresh-token grant with re-read-after-acquire single-
flight lock, and the /v1/api/mcp/oauth/{start,callback} endpoints
mounted on both server and console.
Refactored:
- validate_url_no_ssrf, validate_discovered_endpoint, is_localhost,
effective_port, sanitize_log_text moved out of oidc.py into a shared
oauth_ssrf module; oidc.py re-exports for compatibility. The shared
helpers also expose async wrappers (validate_url_no_ssrf_async,
validate_discovered_endpoint_async) so OAuth-MCP discovery — invoked
from async handlers — does not block the event loop on the
synchronous socket.getaddrinfo call.
- MCPTokenStore.get_oauth_client_secret reader path added (the prior
commit was write-only)
- Storage protocol gains create/pop/cleanup_*_mcp_oauth_pending_state
and get_mcp_oauth_client_secret_ct (mirror OIDC pending-state
pattern: SQLite BEGIN IMMEDIATE select-then-delete, Postgres atomic
DELETE...RETURNING)
Refresh-grant correctness:
- When the AS omits refresh_token (RFC 6749 §6 — MAY rotate), the
existing refresh value is preserved at the OAuth-flow layer rather
than cleared, so production ASes (Google, Auth0 default, Okta) don't
force re-consent every hour
- expires_in accepts int, float, str-with-decimal — earlier int-coerce
through str() failed on float and silently dropped expiry tracking
- The refresh-grant `resource=` parameter (RFC 8707) is the canonical
MCP server URL, not the audience. Audience and resource are distinct
concepts; using audience as resource would mismatch the AS RS
allowlist.
Audience handling:
- _validate_token_audience accepts str or tuple; the callback resolves
accepted_audiences = {server_url, oauth_audience} and validates
against the set, so Auth0-style ASes that honor `audience=` (not
RFC 8707 `resource=`) issue tokens that pass audience-bound
validation
- build_authorize_url emits both `resource=` (RFC 8707) and
`audience=` (Auth0-style) per server config; comment documents which
AS implementations need which form
Security hardening:
- redirect_uri pinned to oidc_config.redirect_base instead of the
request Host header — closes the same Host-header injection PR #476
fixed for OIDC. Both /start and /callback return 503 with operator-
actionable hint when redirect_base is unset
- DCR registration runs under per-server asyncio.Lock with re-fetch
inside the lock, so concurrent /start callers don't both register
and overwrite each other's client_id (the second user's code is no
longer rejected on callback)
- /callback error branch pops the pending state row before redirecting
so a leaked state can't be replayed against a separately-obtained
code in the 60s cleanup window
- WWW-Authenticate Bearer parser handles RFC 7235 quoted-string
escapes (\" and \\) instead of the naive [^"]+ regex
- AS-controlled response bodies and error_description query params go
through sanitize_log_text before reaching exception messages or
audit details. AS error responses are parsed for the standard
RFC 6749 fields (error, error_description, error_uri), each
capped at 80 chars and run through redact_credentials to defend
against ASes that echo the request body back into their error
payload.
- oauth_as_issuer_cached is re-validated against the SSRF guard on
read; on rejection the column is cleared and PRM rediscovery runs
- DCR / token-endpoint / refresh-endpoint response bodies cap at 64
KiB (PRM/AS metadata cap stays at 256 KiB) so a hostile or
malfunctioning AS can't exhaust client memory.
- oauth_client_secret operator input capped at 1024 chars at the
admin-form boundary; longer plaintext rejected with 400.
- /start and /callback responses stamp `X-Frame-Options: DENY` so the
redirected pages can't be framed by attacker sites.
- delete_user cascades to mcp_user_tokens and mcp_oauth_pending so
user deletion no longer leaves dangling per-user OAuth state.
- Renaming or deleting an oauth_user MCP server purges per-user
tokens and pending OAuth state for the previous server name
(delete_mcp_oauth_rows_by_server_name). The OAuth tables key on the
mutable server_name; without this purge, a future server with the
same name (and an attacker-controlled URL) would silently rebind
prior user tokens. A future schema migration will replace the
server_name key with a server_id FK + ON DELETE CASCADE.
- get_user_access_token catches MCPTokenDecryptError (raised when no
installed key can decrypt the row, e.g. after key rotation) and
falls through to None so dispatch surfaces a re-consent rather than
crashing.
- oauth_user MCP server rows are skipped in the static auto-connect
path. Auto-connecting them at startup with empty headers fails the
AS check and trips the circuit breaker; per-user tokens come online
lazily once the user has consented.
Audit (mcp_server.oauth.* prefix):
- consent_started, consent_completed, consent_failed, token_refreshed,
token_revoked, dcr_registered. _audit_event is async and wraps
record_audit in asyncio.to_thread so the audit write doesn't block
the event loop. resource_id on the audit row is the immutable
server_id (PK UUID) so admin-driven server renames don't break
event correlation; server_name is exposed in detail for cross-
reference. dcr_registered detail.has_secret reflects whether the
DCR-issued secret was actually persisted (the prior code reported
has_secret=true even on persistence failure).
- _admin_mcp_action audits the immutable server_id, not the mutable
server_name (which is what the column is — the table's PK was
always server_id).
- All OAuth-flow log keys use the mcp_server.oauth.* prefix to match
the audit-action taxonomy.
Lifespan close-order in turnstone.server and turnstone.console.server
is reversed (LIFO) — mcp_oauth → mcp_crypto → oidc — to match init
order.
Deferred until the upcoming per-user pool integration:
- Multi-node refresh-lock contention via pg_advisory_lock
- DCR re-register on token-endpoint 401 (the dispatch path surfaces
those 401s)
- TTL-LRU caching of decrypted plaintext access tokens
- DNS-rebinding hardening (httpx Transport pin) — documented as
limitation in oauth_ssrf module docstring
Tests: 7 new test files / ~85 new tests covering discovery precedence
+ PRM quoted-string parsing, PKCE round-trip, SSRF helper extraction,
authorize/callback handlers including 503-on-no-redirect-base + DCR
concurrency + JWT audience polymorphism + callback-error-pops-pending,
refresh single-flight lock, refresh resource-vs-audience regression,
decrypt-error fallthrough, _db_servers_to_config skipping oauth_user,
pending-state CRUD round-trip.
(cherry picked from commit 29c42c1427)
Phase 3 of docs/design/oauth-mcp.md. Adds the Fernet/MultiFernet wrapper,
[security] config loader with rotation support, MCPTokenStore CRUD facade,
typed MCPTokenDecryptError that maps to the RFC's mcp_token_undecryptable_
key_unknown class, and a startup gate that fails loud when auth_type=
'oauth_user' rows exist without a configured encryption key.
Crypto module (turnstone/core/mcp_crypto.py):
- MCPTokenCipher wraps cryptography.fernet.Fernet + MultiFernet for
rotation; encrypt with first key, decrypt by trying each in order
- load_mcp_token_cipher_config reads [security] mcp_token_encryption_keys
(plural list) or mcp_token_encryption_key (singular), validates each
key is base64-decodable to exactly 32 bytes
- MCPTokenCipherConfig is repr=False with custom __repr__ that redacts
raw key bytes (defense in depth against accidental log/traceback leak)
- _key_fingerprint produces an 8-hex-char SHA-256 prefix for audit
attribution without exposing the key
- MCPTokenStore handles encrypt-on-write / decrypt-on-read for
mcp_user_tokens and mcp_servers.oauth_client_secret_ct
- get_user_token MUST NOT auto-delete the row on MCPTokenDecryptError
(test_get_user_token_with_wrong_key_raises_decrypt_error verifies
the row stays intact across a key-mismatch read)
- initialize_mcp_crypto_state / close_mcp_crypto_state lifespan helpers
shared between server and console
Storage protocol (5 new ciphertext-only methods):
- set_mcp_oauth_client_secret_ct (dedicated writer; deliberately NOT
added to MCP_SERVER_MUTABLE so generic update_mcp_server cannot write
the secret column)
- create_mcp_user_token, get_mcp_user_token,
update_mcp_user_token_after_refresh, delete_mcp_user_token
Server + console lifespans (turnstone/server.py + console/server.py):
- after OIDC init, count auth_type='oauth_user' rows; if any exist and
no encryption key is configured, log an actionable error and
raise SystemExit(1)
- without oauth_user rows, missing key is fine (lazy validation; admin
flip without restart returns 503 from the admin handler)
- app.state.mcp_token_cipher / .mcp_token_store populated when key
configured; None otherwise
Admin handlers:
- _require_token_store_for_oauth_secret pre-mutation gate validates
token_store availability and oauth_client_secret type BEFORE
storage.create_mcp_server / update_mcp_server runs, so a 503 from a
missing key never leaves an orphan row or partial-update state
- _apply_oauth_client_secret encapsulates the encrypt + audit write
used after the storage mutation; rolled out across both create and
update handlers
- 503 message references both mcp_token_encryption_key (singular) and
mcp_token_encryption_keys (plural for rotation)
- non-string oauth_client_secret payloads (false / 0 / lists / dicts)
are rejected with 400 instead of being str()-coerced
- when auth_type transitions away from oauth_user, the encrypted
secret column is cleared in the same admin call (with audit), so
flipping back doesn't silently resurrect a stale credential
Audit events (mcp_server.oauth.* per audit.py taxonomy; RFC's
mcp.oauth.* renamed for consistency):
- mcp_server.oauth.client_secret_set fired from admin handlers with
cleared:bool and key_fingerprint
- mcp_server.oauth.token_decrypt_failure fired from MCPTokenStore
.get_user_token when no installed key can decrypt; carries
key_fingerprints_attempted
Tests: 35 new tests across test_mcp_crypto, test_mcp_token_store,
test_server_lifespan_mcp_crypto, plus 6 admin-API tests covering the
no-orphan-row, no-partial-update, secret-clear-on-transition, and
non-string-secret-rejection invariants. Suite at 5337 (Phase 3 added
~50 tests including the rebase-imported skill suite).
cryptography>=42 promoted from transitive (lacme[tls]) to direct dep
since the encryption layer is now core, not optional.
Phase 4 (OAuth flow) wires the actual callers; Phase 3 adds only the
crypto layer and is exercised entirely by tests.
(cherry picked from commit 7f132e7230)
Adds the data model and admin UI surface required by the OAuth-MCP flow.
Phase 2 of the per-user delegation initiative.
Schema:
- migration 049 creates mcp_user_tokens (PK user_id, server_name) and
mcp_oauth_pending (PK state, indexed by created_at)
- eight new columns on mcp_servers: auth_type ('none' / 'static' /
'oauth_user', NOT NULL DEFAULT 'static') plus six oauth_* config
fields and oauth_as_issuer_cached
- post-upgrade UPDATE normalises auth_type to 'none' for streamable-http
rows whose headers are NULL/empty/'{}'; stdio rows are left at the
'static' default (auth_type is HTTP-auth-only)
- _schema.py kept in lockstep with the migration so metadata.create_all
and alembic upgrade produce identical shapes
- mcp_user_tokens / mcp_oauth_pending TypedDicts in _protocol.py for
Phase 3/4 use (no CRUD methods yet)
Storage / API:
- create_mcp_server gains the eight kwargs across protocol + sqlite +
postgresql
- MCP_SERVER_MUTABLE picks up auth_type and the six text oauth_* fields;
oauth_client_secret_ct is intentionally NOT in the whitelist — Phase 3
will own ciphertext writes via a dedicated method
- McpServerInfo + Create/Update Pydantic schemas extended; oauth_client_secret
accepted as plaintext input but discarded (Phase 3 wires encryption)
Admin handlers:
- _parse_auth_type validates against {'none', 'static', 'oauth_user'} and
rejects empty / unknown values; shared between create and update
- when auth_type changes away from 'oauth_user', the oauth_* config
columns are explicitly nulled in the same UPDATE so the row stays
consistent
- _clean_oauth_text caps text fields at 512 chars (URLs at 2048) to bound
admin write surface
- _mask_mcp_secrets now masks oauth_client_secret_ct to '***' regardless
of reveal=true (write-only field)
- audit detail dict redacts oauth_client_secret if present
Frontend:
- new "Multitenant Authorization" fieldset on the MCP-server modal with
three radio buttons (None / Shared / Per-user OAuth 2.1)
- conditional OAuth subform: AS URL, registration mode (preregistered /
dcr; cimd is future), client ID, client secret, scopes, audience
- secret input is autocomplete=off and never round-trips on edit
- audience auto-populates from the MCP server URL on blur
- headers textarea hidden and submitted as {} when auth_type is 'none' or
'oauth_user' so flipping the radio cleans up server-side state
Tests: storage round-trip for the new columns, oauth_pending table smoke,
migration 049 upgrade/downgrade with stdio-vs-http normalisation, four
admin-API tests for auth_type validation and oauth_*-clear-on-flip-away.
Suite passes 5284 (matched pre-Phase-2 baseline 5267 + 17 new).
Stacks on Phase 0; no behavioural change for existing rows.
(cherry picked from commit d675b237a3)
Phase 0 of the OAuth-MCP RFC: prepare MCPClientManager for the per-(user,
server) session pool that lands in Phase 5, without changing static-path
behavior.
Two changes:
1. Hardening helpers _pre_close_streams and _tcp_probe rename their first
parameter from `name` to `key`. Type stays `str` for now; widening to
`str | tuple[str, str]` happens in Phase 5 when callers actually pass
tuples. _safe_close_stack takes the stack directly and is unchanged.
2. The eleven parallel name-keyed dicts (_sessions, _per_server_stacks,
_per_server_tools, _per_server_resources, _per_server_prompts,
_supports_list_changed, _supports_resources, _supports_resource_list_changed,
_supports_prompts, _supports_prompt_list_changed, _server_streams) are
consolidated into _static_servers: dict[str, StaticServerState]. Server-
level state (circuit breaker, notification debounce, last-error,
db-managed, merged catalog maps, listener lists) stays on the manager,
unchanged.
PoolEntryState is defined for Phase 5 use but no code instantiates it. The
typed map declarations (dict[str, StaticServerState] vs dict[tuple[str, str],
PoolEntryState]) make accidental cross-keying lookups easier to catch.
PR #296 hardening preserved exactly:
- pre-close-streams atomic take-and-clear before stack teardown
- stale-session-and-stack guard at _connect_one top: both state.session and
state.stack checked, cleared independently, entry preserved (not popped)
- transport-error session-eviction in dispatch sets state.session=None only,
leaving stack/streams for the next connect-time guard sweep
- _safe_close_stack CancelledError suppression unchanged
- TCP probe before streamablehttp_client unchanged
- future.cancel() after TimeoutError in all sync bridges unchanged
- notification debounce stays manager-level (not migrated into the dataclass)
Refresh helpers (_refresh_server_tools/_resources/_prompts) snapshot
state.session into a local immediately after the None guard so concurrent
transport-error eviction during await cannot null the session reference
mid-call.
Tests: shared _seed_static_state helper in tests/conftest.py replaces eleven
direct dict mutations; new test_reconnect_preserves_static_state_identity
guards the entry-preservation invariant. Pass count rises 5266 → 5267.
(cherry picked from commit be0950bb98)
Deletes the _periodic_refresh task and its supporting state
(_refresh_task, _refresh_failures, _refresh_backoff_until,
_REFRESH_BACKOFF_BASE/MAX, _DEFAULT_REFRESH_INTERVAL, refresh_interval
kwarg) from MCPClientManager. Push notifications and operator-driven
manual refresh now cover all catalog-update needs; the long-running
4-hour timer was dead complexity that obscured the per-user pool
work to come.
Catalog freshness on auto-reconnect is preserved by scheduling an
unblocking _refresh_server task on the mcp-loop after _connect_one
succeeds; the calling thread returns immediately so half-open
recovery latency does not double. Adds MCPClientManager.reconnect_sync
(clears the circuit, closes any existing session, calls _connect_one,
clears stale catalog on failure).
Wires a new pair of operator endpoints —
POST /v1/api/admin/mcp-servers/{name}/refresh and
/v1/api/admin/mcp-servers/{name}/reconnect — that fan out to all
nodes through the existing _internal route family, with per-row
"Refresh" and "Reconnect" buttons in the MCP Servers admin tab.
The new node-internal paths /api/_internal/mcp-{refresh,reconnect}/
are gated to the approve scope to prevent direct unprivileged
reconnects bypassing the console's admin.mcp gate. Internal
endpoints return generic error messages and a filtered status
payload (no command/url) to keep transport details admin-gated.
Drops the [mcp] refresh_interval setting, the
--mcp-refresh-interval CLI flag, and the matching config-mapping
entry; updates docs/architecture.md, docs/tools.md,
docs/settings.md, and the three PlantUML diagrams that referenced
the periodic loop.
Tradeoffs (intentional):
- Idle nodes will not auto-rejoin a recovered MCP server until
traffic arrives or an operator clicks Reconnect. The previous
background reconnection loop is gone by design — push
notifications + operator controls replace it.
- Console fan-out blocks on the slowest node (existing pattern);
not changed here.
This is Phase 1 of the OAuth-MCP series — feature subtraction
ahead of per-user state.
(cherry picked from commit eb2a119da9)
* feat(skills): paste SKILL.md to auto-fill the Create Skill modal
When a user pastes an Anthropic-style SKILL.md (YAML frontmatter +
markdown body) into the Create Skill content textarea, the frontend
sniffs the leading ``---``, posts the raw text to a new backend parse
endpoint, and populates name / description / tags / author / version /
license / compatibility / allowed_tools from the parsed fields. The
textarea is left with the body only (frontmatter stripped), and a toast
reports how many fields were set vs. kept (already-typed values are
preserved).
Backend
- ``POST /v1/api/admin/skills/parse`` (admin.skills permission) wraps
the existing ``turnstone.core.skill_parser.parse_skill_md`` so admin
imports and external installs share one parser. ``ParseSkillRequest``
/ ``ParseSkillResponse`` schemas added; OpenAPI spec + sync/async
console SDK methods updated.
- Hardening: 32 KiB cap on ``raw`` (Pydantic ``max_length`` + handler
enforcement); ``Content-Length`` pre-check returns 413 before any body
buffering; parse offloaded via ``asyncio.to_thread`` so deeply-nested
YAML cannot stall the event loop.
Frontend (turnstone/console/static)
- New paste handler with optimistic paint (raw text shown immediately,
textarea disabled + ``aria-busy`` flipped, hint switches to
"Parsing...") so the round-trip is visible on slow networks.
- ``AbortController`` + generation guard (``_ctmPasteController``) so a
fresh paste or modal close cancels a stale fetch — the previous
handler's callbacks see the controller has been replaced and bail
before touching the DOM.
- Non-destructive overwrite: ``_setSkillFormField`` returns "filled" /
"skipped" / "absent" and refuses to clobber non-empty values. Toast
reports counts.
- Bumps ``#toast`` z-index above modal overlays (was 200 vs. modal 600
— toasts fired while a modal was open were invisible). Console-wide
fix exposed by this being the first feature to fire toasts mid-modal.
HTML / CSS
- New ``.skill-paste-hint`` line above the textarea announcing the
affordance, sized to match surrounding ``.label-hint`` text.
- ``aria-describedby`` ties the hint to the textarea; ``aria-live=
"polite"`` announces the busy-state transition to screen readers.
- "Skill Content" heading hint reworded "system message — ..." →
"available: ..." and the variables row label "Variables" → "Used"
to disambiguate available vs. in-use template variables.
Tests
- 11 new cases in ``tests/test_skill_parse_api.py``: happy paths
(full / minimal / nested-metadata / unquoted-colon recovery),
malformed YAML 400, missing/blank/missing-name 400, RBAC 403, raw
body 32 KiB cap (Content-Length pre-check), chunked-encoding bypass
forces the application-layer cap. Test pins ``raw_frontmatter``
omission so a future ``dataclasses.asdict`` refactor can't silently
leak the full YAML dict back to clients.
Validation
- 5146 / 5146 ``pytest -k "not live"`` pass.
- ``ruff`` + ``mypy`` clean on changed sources.
- ``node -c`` clean on governance.js.
- Two-stage code review (full pipeline + bug+quality re-review of the
fix patches) applied; all confirmed findings addressed.
* fix(skills): Copilot PR #477 review fixes (cumulative bug-1, bug-2, q-1)
bug-1 (server.py): Content-Length pre-check was clamped to 32 KiB —
the same number as the per-string char cap on ``raw``. A legitimate
``raw`` of exactly 32 KiB produces a JSON body well above 32 KiB once
the ``{"raw":"..."}`` wrapper and any escaping is added, so valid
near-max requests were 413'd. New constant
``_PARSE_SKILL_MAX_BODY_BYTES = _PARSE_SKILL_MAX_CHARS * 4`` admits the
wrapper + multibyte expansion while still refusing obviously oversized
payloads early; the per-string ``len(raw)`` check stays authoritative.
bug-2 (governance.js): hideCreateTemplateModal aborted the inflight
paste controller and nulled the global, but the handler's ``.catch``
and ``.finally`` guard each DOM mutation behind ``_isCurrent()`` —
both bail when the controller has been nulled, leaving the textarea
``disabled`` + ``aria-busy`` and the hint stuck on "Parsing…".
Reopening the modal landed on a poisoned state. The second-pass
review's q-2 cleanup that dropped the show-side defensive reset
missed this scenario — the verifier's reachability argument confused
"controller is null" with "UI state is reset"; the two are
independent. Hide now resets the paste-induced visible state
alongside the abort.
q-1 (console_spec.py): error_codes for the parse endpoint listed only
400; handler also returns 413 for oversized bodies. Added 413; kept
403 implicit per the convention sibling admin endpoints follow.
Test fixup: bumped the Content-Length test payload to 200 KB so it
clearly exceeds the new 128 KB pre-check threshold; otherwise it was
falling through to the per-string check and duplicating
test_oversized_raw_chunked_returns_413's coverage.
(cherry picked from commit 0a8083e6d5)
PR #476 review feedback (Copilot, oidc.py:584,616):
1. initialize_oidc_state's docstring claimed "on any failure
enabled is False" but the JWKS-prefetch failure branch
intentionally keeps enabled=True so the callback's lazy-fetch
retry can recover from a transient IdP issue at startup.
Docstring rewritten to spell out the three post-conditions:
disable, JWKS-failure-keeps-enabled, success.
2. The long-lived httpx.AsyncClient was created up front, then
three disable branches (discovery exception, discovery-returned-
disabled, missing redirect_base) returned without closing it,
leaving sockets held until shutdown.
Restructured: discovery now uses a transient AsyncClient inside
a context manager (closed at exit). The long-lived client is
only created after the disable checks pass. The JWKS-failure
branch still legitimately keeps the client open because the
lazy-retry path needs it.
The pre-existing single-client-passthrough test was replaced
with three more specific tests: long-lived client only goes to
fetch_jwks (not discover_oidc); discovery-exception path leaves
http_client=None; missing-redirect_base path leaves
http_client=None.
(cherry picked from commit b2153d907f)
q-4: tests/test_oidc.py's _make_config and tests/test_oidc_handlers.py's
_make_oidc_config built the same OIDCConfig with sensible defaults but
had drifted — only the handlers helper set redirect_base. After b3
made redirect_base operationally required, every test_oidc.py test
that exercised redirect_base had to override it explicitly. A future
test could omit redirect_base and silently exercise the wrong
production path.
Moves make_oidc_test_config to tests/conftest.py with the more
complete handler-version defaults (including redirect_base). Both
test files import it under their existing local alias
(_make_config / _make_oidc_config) so the 60+ call sites in
test_oidc.py and the handler tests don't have to change.
q-5: section banner '# Exception' (singular) at oidc.py:79 became
inconsistent after b5 (callback robustness) added OIDCKeyNotFoundError.
Renamed to '# Exceptions'.
(cherry picked from commit 5d4a50d2cd)
The OIDC perf batch added storage.count_users() and migrated the two
OIDC handlers (handle_oidc_authorize, handle_oidc_callback) but missed
handle_auth_status — which still ran storage.list_users() then
len(users) > 0 for the same has-any-users gate.
count_users() is one COUNT(*) round-trip vs list_users() rehydrating
every row dict. Wrapped in asyncio.to_thread to match the OIDC handler
pattern; the async handler no longer blocks the event loop on storage
I/O for what's effectively an existence probe.
(cherry picked from commit 7c6bc22d02)
bug-2 (Postgres) — replace_oidc_roles read existing rows under default
READ COMMITTED with no row lock. Two concurrent OIDC callbacks for the
same user_id (racing token refreshes with differing claim sets) could
both observe the same baseline and produce a final role state matching
neither caller's intent. Adds .with_for_update() to the SELECT so the
existing rows for this user are locked for the duration of the
transaction.
The lock is per-user_id, not table-wide; unrelated user writes are
unaffected. Empty result sets acquire no locks, so a brand-new user
with no rows yet still allows two callers to proceed and merge via
ON CONFLICT DO NOTHING — that's a permissive race that self-heals on
the next reconciliation cycle, documented in code.
perf-1 (SQLite) — replace_oidc_roles took the SQLite global write
lock unconditionally via BEGIN IMMEDIATE before reading. Steady-state
re-logins (claims unchanged, no INSERT/DELETE needed) paid the lock
cost for nothing and serialised against unrelated writers.
Replaces with a double-check pattern: phase 1 reads under the default
deferred transaction (no write lock), computes the diff, and returns
(set(), set()) on no-op. Phase 2, only when mutation is needed,
commits the read txn, escalates to BEGIN IMMEDIATE, RE-READS, and
re-computes the diff under the lock before writing. The returned
(added, removed) reflects what was actually written, so caller logging
in apply_role_mapping stays truthful even when concurrent writers
shifted state between the two reads.
The OR IGNORE on insert is now defense-in-depth (the lock makes it
unnecessary) but kept as a safety net.
(cherry picked from commit d5087ef3b9)
The 8-commit OIDC stack added TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS
(operator allow-list for cross-host IdP discovery endpoints) and
promoted TURNSTONE_OIDC_REDIRECT_BASE to required, but the docs drifted
in two places:
q-1 — Troubleshooting > "OIDC not configured" still listed three
required env vars. An operator hitting the missing-redirect-base
startup error landed on a debugging entry that didn't mention the
variable they were missing. Fixed; added a separate troubleshooting
entry naming the exact log message produced by initialize_oidc_state
when redirect_base is unset.
q-2 — TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS was undocumented entirely.
Added a row to the env-var table and a new "Cross-host endpoints"
section explaining when the knob is needed (Google is the canonical
multi-origin IdP, but it's auto-handled; the env var is for any other
IdP whose discovery doc legitimately references hosts beyond the
issuer's origin). Added a troubleshooting entry pointing at the new
section.
(cherry picked from commit 3cf87628d2)
If apply_role_mapping raised after create_oidc_user committed (transient
storage failure, race with role deletion, etc.), provision_oidc_user's
inline safety-net was skipped — and on retry the existing-identity
branch never reached the safety-net code, leaving the user permanently
stranded with zero roles.
Extracts _ensure_default_role(storage, user_id, desired_role_ids=None)
helper. Calls it on BOTH the new-user and existing-identity paths so a
user stranded by a transient failure recovers on next login.
desired_role_ids is a hint that lets the helper skip list_user_roles
when claim-driven mapping populated at least one role; the new-user
path was already paying that query, the existing-identity path now
pays it only when claim mapping returned an empty desired set.
Documents the admin-strip behavior in the helper docstring: stripping
all roles from an OIDC user no longer locks them out, since the next
login will re-grant builtin-viewer (assigned_by='oidc-default'). The
documented way to deny an OIDC user is to unlink their OIDC identity
via the admin endpoint, not to strip roles. The pre-fix behavior
(stripped user actually locked out) was the bug.
The 'oidc-default' vs 'oidc' assigned_by distinction is preserved:
apply_role_mapping's revocation lane only touches 'oidc' rows, so the
safety-net role survives every subsequent login regardless of claims.
Six new tests cover both paths, the hint short-circuit, the
list_user_roles fallback, the missing-builtin-viewer no-op, and the
self-heal regression case for already-stranded users.
(cherry picked from commit 1c41212f15)
q-5: _derive_username's UUID-retry tier (oidc.py:923-933) was untested.
After perf-6 collapsed tier-2 to a single find_existing_usernames call,
the only remaining tail was the 3-attempt UUID-retry loop and the final
raise. New TestDeriveUsername class covers:
- falls into UUID retry when all 10 suffix candidates are taken
- UUID retry succeeds on the second attempt after one collision
- UUID retry exhausted -> raises OIDCError
q-8: filled the unit-level coverage holes the multi-stage review flagged:
- test_validate_id_token_retry_after_kid_rotation — direct unit test of
the OIDCKeyNotFoundError path with real RS256 keys + JWKS rotation
(previously only exercised end-to-end through the handler).
- test_callback_uses_pending_audience_not_handler_audience — pins down
the bug-3 fix by decoding the issued JWT cookie and asserting aud
matches the audience stored at /authorize time, not the handler param.
- test_apply_role_mapping_int_claim / _dict_claim — exercises the
else: values = [str(claim_value)] branch for non-string non-list
claim shapes.
- TestFetchJWKS — non-200 status, non-dict body, dict-missing-keys,
keys-not-list, transport network error.
- TestExchangeCode network/4xx/5xx error tests (the non-dict-body case
already shipped in batch 5).
Also a small production hardening that fell out of writing the
TestFetchJWKS::test_fetch_jwks_non_dict_body_raises test: fetch_jwks now
guards isinstance(result, dict) before result.get("keys"), matching the
shape-check pattern that discover_oidc and exchange_code already use.
A list/null body now surfaces as OIDCError("...not a JSON object") rather
than AttributeError leaking up to the lifespan.
(cherry picked from commit 5c11ab985f)
Eleven small maintenance fixes; no behavior change beyond bug-3.
bug-3: pending.get('audience', audience) couldn't fall back because
pop_oidc_pending_state always returns a dict with the audience key
set verbatim from a non-null TEXT column. Replaced with
pending.get('audience') or audience to cover the empty-string case
defensively. Comment explains the security rationale.
q-1: extract _env_or_cfg_str / _env_or_cfg_bool helpers in oidc.py;
load_oidc_config's six near-identical env-or-config blocks collapse
to one-liners. role_map / trusted_endpoint_hosts / redirect_base
retain bespoke parsing.
q-3: discover_oidc narrows except (httpx.HTTPError, ValueError, KeyError)
with exc_info=True.
q-4: OIDC_STATE_TTL_SECONDS = 300 constant in oidc.py; auth.py imports
and passes it explicitly. Storage signatures keep the literal default
(storage layer doesn't know OIDC TTL semantics).
q-6: hoist runtime imports (OIDCError, OIDCKeyNotFoundError, exchange_code,
fetch_jwks, provision_oidc_user, validate_id_token, build_authorize_url,
generate_pkce_verifier) to module scope in auth.py. The genuine cycle
is only oidc._derive_username -> auth.is_valid_username, kept
function-scoped. test_oidc_handlers.py mock targets repointed to
turnstone.core.auth.X to match the new binding.
q-7: comment + docs explain the 'oidc' vs 'oidc-default' assigned_by
marker distinction.
q-9: OIDCIdentity / OIDCPendingState TypedDicts in storage protocol.
Implementations construct via TypedDict syntax so mypy structurally
verifies all required fields.
q-10: fetch_jwks narrows except (httpx.HTTPError, ValueError); docstring
matches.
q-11: rename generate_pkce_pair -> generate_pkce_verifier; return only
the verifier (build_authorize_url already recomputes the challenge).
q-12: extract _buildOidcRow helper in admin.js so future field additions
go in one place.
q-13: OIDCConfig docstring lists startup-config vs discovery-derived
field groups.
(cherry picked from commit bae4adca12)
Eight independent perf wins on the OIDC hot path:
perf-1: list_users() full-scan setup-gate replaced with new count_users()
on both authorize and callback. Saves a full users-table fetch per login.
perf-2: handle_oidc_callback's sync DB chain wrapped in asyncio.to_thread
for cleanup, pop_oidc_pending_state, count_users, and provision_oidc_user.
handle_oidc_authorize gets the same treatment for count_users and
create_oidc_pending_state. Event loop no longer blocks for the full
callback duration on Postgres deployments.
perf-3: apply_role_mapping N+1 collapsed via new replace_oidc_roles
storage method. One transaction handles the diff + insert + delete
instead of 2N+1 commits per login. Returns (added, removed) so the
caller can still emit per-role audit logs.
The diff respects the documented invariant "manually-assigned roles
are never touched" — desired_role_ids is filtered against rows where
assigned_by != 'oidc' before computing added/removed. This prevents a
PK conflict (Postgres lockout) or silent OR-IGNORE no-op (SQLite lying
return) when admin-ui or oidc-default already holds the same role_id.
perf-4: provision_oidc_user no longer re-queries list_user_roles after
apply_role_mapping. The new-user builtin-viewer fallback is gated on
desired_role_ids being empty, which is information apply_role_mapping
already returned.
perf-5: JWKS refetch dedup via asyncio.Lock on app.state. Both lazy-fetch
(cold-start recovery) and rotation paths share the same lock with a
double-check pattern: re-resolve kid against the current cache before
issuing a new GET. N concurrent callbacks during rotation now produce
at most 1 fetch.
perf-6: _derive_username's 9-suffix loop collapsed via new
find_existing_usernames(candidates) -> set query. Worst case drops
from 13 sequential queries to 1 + up-to-3 UUID-retry queries.
perf-7: cleanup_expired_oidc_states gated to once-per-60s per process
via app.state.oidc_last_cleanup_monotonic. The pop already deletes
the consumed row; the bulk cleanup is only relevant for abandoned
authorize flows, so frequency was overkill.
perf-8: Long-lived httpx.AsyncClient stashed on app.state.oidc_http_client
by initialize_oidc_state. discover_oidc/fetch_jwks/exchange_code accept
an optional client= kwarg; when set, skip the per-call AsyncClient
context-manager. New close_oidc_state lifespan teardown closes it.
Tests pass client=None to keep the transient-client legacy path.
New storage methods (sqlite + postgresql):
- count_users() -> int
- find_existing_usernames(candidates) -> set[str]
- replace_oidc_roles(user_id, desired) -> (added, removed)
(cherry picked from commit 39a647f39c)
Four small hardening fixes on the OIDC callback hot path:
bug-4: JWKS rotation retry was matching the substring 'not found in JWKS'
inside an OIDCError message. A future rephrasing would silently break
key rotation. Adds OIDCKeyNotFoundError(OIDCError); validate_id_token
raises the subclass at the kid-not-found site; handle_oidc_callback
catches it explicitly. Other 'not found' errors in validate_id_token
remain as plain OIDCError.
bug-5: tokens['id_token'] raised KeyError if the IdP returned 200 without
id_token. exchange_code now rejects non-dict response bodies; the
callback validates id_token shape (must be non-empty str) before
passing to validate_id_token. Both raise OIDCError, surfaced as the
standard 'Authentication failed' redirect.
bug-6: shared_static/auth.js — the OIDC error display raced showLogin's
/v1/api/auth/status fetch via a 300ms setTimeout. showLogin now takes
an optional oidcError parameter and paints it after _switchMode clears
the error, in both the success and catch branches of the fetch.
sec-4: oidc.py exchange_code's non-200 OIDCError interpolated up to 500
bytes of attacker-controlled IdP body, which then went to log.warning
via 'OIDC callback failed: %s'. CRLF in resp.text could forge log
lines. New _sanitize_log_text helper escapes control chars via
unicode_escape and caps at the rendered length.
(cherry picked from commit 0af3adae1d)
provision_oidc_user previously called create_user (INSERT OR IGNORE
on SQLite — silent no-op on UNIQUE conflict), then create_oidc_identity
(also INSERT OR IGNORE), then apply_role_mapping which writes user_role
rows for the supposedly-new user_id. On a username TOCTOU race or
concurrent (issuer, sub) double-create, both inserts no-opped but
user_role rows were already written — leaving orphan rows pointing
at a user_id that doesn't exist.
PostgreSQL's create_user raised IntegrityError instead of silently
no-opping so it produced a misleading 'Authentication failed' error
without orphans, but the user-facing UX was equally poor.
Adds StorageConflictError to the storage protocol and create_oidc_user
that does both inserts in one transaction. Username collision and
(issuer, subject) collision both raise StorageConflictError, mapped
to OIDCError by provision_oidc_user. Crucially the new code does not
silently bind a colliding-username new identity to the existing user
— that would be an account-takeover vector. It raises.
SQLite uses BEGIN IMMEDIATE inside the try block so lock-contention
errors surface as StorageConflictError instead of leaking the raw
sqlalchemy OperationalError.
PostgreSQL relies on SQLAlchemy 2.x begin-on-demand semantics; the
explicit conn.commit()/rollback() in the catch block is the only
materialization path. Discrimination on PG uses
exc.orig.diag.constraint_name with message-substring fallback.
(cherry picked from commit 11618bb1d7)
_build_oidc_redirect_uri previously fell back to the request Host
header when redirect_base was unset. With a permissive reverse proxy
or direct backend access, a spoofed Host minted an authorize URL
pointing to attacker-controlled host — combined with a permissive
IdP redirect_uri allowlist this enables auth-code interception.
There is no production scenario where a Host-derived redirect_uri is
correct, so this fails closed:
- initialize_oidc_state checks redirect_base after discovery succeeds
and disables OIDC (with an explicit error log naming the env var)
if it's empty. Runs before fetch_jwks so a misconfigured deploy
doesn't make a wasted JWKS call.
- _build_oidc_redirect_uri simplifies to f"{redirect_base}/v1/api/auth/oidc/callback".
request parameter dropped; both call sites (handle_oidc_authorize,
handle_oidc_callback) updated.
- docs/oidc.md promotes TURNSTONE_OIDC_REDIRECT_BASE from "Recommended"
to "Required" with the security rationale.
(cherry picked from commit 52aba17740)
The OIDC discovery + JWKS prefetch block was duplicated byte-for-byte
between turnstone/server.py and turnstone/console/server.py. The bare
except branch in that block also left app.state.oidc_config unchanged
on unexpected exceptions — leaving the runtime with enabled=True and
empty endpoints, producing malformed authorize URLs.
Extracts initialize_oidc_state(app_state) into turnstone/core/oidc.py
which guarantees a coherent post-condition on every code path:
- discovery exception -> oidc_config replaced with enabled=False, jwks_data=None
- discovery returns enabled=False -> jwks_data=None
- JWKS prefetch fails -> jwks_data=None but enabled=True preserved (the
callback's lazy-fetch retry path remains the recovery)
- success -> oidc_config + jwks_data both populated
Also hardens discover_oidc against non-dict discovery responses
(list/null/string/int) — previously these raised AttributeError out
of doc.get and propagated past the lifespan's bare except.
server.py and console/server.py lifespan blocks collapse to a single
await initialize_oidc_state(app.state) call.
(cherry picked from commit 6f9e140a41)
OIDC discovery-document endpoints (token_endpoint, jwks_uri,
userinfo_endpoint) were stored verbatim in OIDCConfig and later passed
to httpx without revalidation. Only the issuer URL was checked. A
hostile or compromised IdP could return token_endpoint pointing to an
internal IP (169.254.169.254, 10.0.0.0/8, etc.) and Turnstone would
POST the client_secret there.
Extracts the existing scheme/userinfo/SSRF check into
_validate_url_no_ssrf, adds validate_discovered_endpoint that runs the
same checks plus an issuer-binding check, and wires it into
discover_oidc for authorization_endpoint, token_endpoint, jwks_uri,
and userinfo_endpoint (when present).
Issuer binding accepts:
- Same (scheme, hostname, effective port) as the issuer.
- A hostname in _KNOWN_TRUSTED_ENDPOINT_HOSTS for the issuer (Google's
multi-origin discovery is in the allow-map by default).
- A hostname in OIDCConfig.trusted_endpoint_hosts, settable via
TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS env var or config.toml, for
IdPs not in the static map.
Effective port comparison treats https://host and https://host:443 as
the same origin (urllib.parse.urlparse leaves the explicit form's port
as 443 and the implicit form's as None).
24 new tests cover the validator, the Google known-hosts path, the
operator allow-list, default-port equivalence, foreign-host
rejection, private-IP rejection, embedded credentials, and DNS
rotation between issuer check and endpoint use.
(cherry picked from commit 0df7dc026b)
* feat(console): inline node picker replaces back-to-console banner
Drops the 32px banner the console proxy used to inject above proxied
server-UI pages and replaces it with an inline node-id pill in the
existing #ui-header. Click the pill to open a dropdown that lists
healthy nodes (health dot, ws count, reachable/degraded/unreachable
text) plus a top-row link back to the console.
Reuses the .ws-tab-dropdown shell from ui/static/style.css for
animation, shadow, theme override, and item layout, so the picker
visually matches the workstream-tab chevron menu it sits next to.
Keyboard nav (ArrowDown/Up/Home/End/Tab/Escape) mirrors the chevron
menu's handler with cross-reference comments at both sites.
Lazy-fetches /v1/api/cluster/nodes against the console origin
(bypassing the prefix shim) on first open.
Reclaims 32px of vertical space, consolidates three separate
"you're on node X via console" indicators into one, and turns the
wayfinding chrome into a real cluster-nav primitive.
* fix(console): address Copilot review on node picker
- Request /v1/api/cluster/nodes?limit=1000 (collector's hard cap)
instead of relying on the default 100 — clusters with more than
100 nodes were silently dropping rows from the picker.
- Hand off focus to the first menu item after the async fetch
resolves: openMenu()'s deferred focus hook ran while only the
skeleton was in the DOM, so first-open keyboard users were
stranded on the trigger until they pressed an arrow key.
- Tab now closes the menu without preventDefault, so focus moves
to the next focusable element on the first press (ARIA APG menu
pattern). Escape still preventDefault + returns to the pill.
- Cap pill max-width at 240px and ellipsize the id span; node ids
are accepted up to 256 chars upstream and could otherwise push
the title and right-side controls off the appbar. Pill carries
a title attribute so the full id is still legible on hover.
* fix(session): properly inject queued user messages mid-loop
Two queued-user-message bugs in ``ChatSession.send()``.
**Mid-tool-call: ``Unexpected role 'tool' after role 'user'`` on Mistral.**
The ``supports_tool_advisories`` capability flag (default False for
unknown openai-compatible models) routed cap-off providers down a
short-circuit branch in ``_collect_advisories`` that called
``_flush_queued_messages`` directly. That appended a ``user`` turn
between ``assistant(tool_calls)`` and ``tool``, which mistral-common's
``_validate_message_order`` rejects with a 400.
Drop the flag. All providers now run the unified path: queued user
messages become ``UserInterjection`` advisories that ride inside the
tool result envelope via ``wrap_tool_result``, splicing
``<system-reminder>`` text into the tool message's content. Role
sequence stays ``assistant → tool``. Live-confirmed on Mistral
medium and Qwen3 — both correctly distinguish system-reminder from
tool stdout in their reasoning.
**Mid-stream: queued message orphaned until next user send.**
After a no-tool assistant turn, ``_flush_queued_messages`` would
append the queued user message to history and the loop would
``break``, leaving the message at the tail of history with no
model response. Visible as "two sends to get one reply".
``_flush_queued_messages`` now returns ``bool``. The no-tool branch
``continue``s on drain instead of ``break``ing, so the model gets a
turn over the extended history.
Tests:
- ``test_collect_advisories_drains_text_queued_messages_to_persistent``
pins the unified-path drain (text-only queue → ``UserInterjection``,
no separate user turn appended to ``self.messages``).
- ``test_send_continues_when_messages_queued_during_streaming`` pins
the loop-continue behavior (fails with 1 stream call pre-fix,
passes with 2 post-fix).
* fix(session,ui): reject queued attachments + paperclip busy state
Copilot pointed out that the attachment-bearing branch in
``_collect_advisories`` had the same role-ordering bug as the
text-only path that 802658f fixed: an attachment-bearing queued
item would still call ``_append_user_turn`` mid-tool-call,
injecting ``user`` between ``assistant(tool_calls)`` and ``tool``.
Pragmatic fix: don't allow attachments to be queued at all.
**Backend.** ``ChatSession.queue_message`` raises a new
``AttachmentsNotQueueableError`` when called with non-empty
``attachment_ids``. The interactive ``/send`` route catches it,
releases reservations via the existing ``_release_reservation_on_fail``
hook, and surfaces ``status: "attachments_busy"`` to the caller
with the IDs in ``dropped_attachment_ids``. The coord adapter
mirrors the cleanup (releases the soft-locked reservation taken
for ``_send_id``) so the create-with-attachments path can't leak.
Now that the queue can never carry attachments, the per-item
``att_ids`` slot is gone:
- Queue tuple slimmed ``(cleaned, priority, att_ids)`` →
``(cleaned, priority)``.
- ``_flush_queued_messages`` collapses to a single combined-text
user turn (no attachment branch).
- ``_collect_advisories`` queue-drain pushes ``UserInterjection``
advisories only (no ``attachment_items`` list).
- ``dequeue_message`` no longer unreserves (queue can't reserve).
- ``_resolve_attachment_ids`` had no remaining production callers
and is deleted along with the tests that exercised it in
isolation.
**Frontend.** ``Composer.setBusy`` disables the paperclip whenever
busy (regardless of ``queueWhileBusy``) — text still queues,
attachments don't. ``chat.css`` gains a ``.composer-attach:disabled``
rule (mirrors the existing ``.composer-send:disabled`` treatment)
so the affordance actually looks unclickable instead of falling
through to the UA default. ``title`` and ``aria-label`` are kept in
sync for AT users (WCAG 4.1.2).
Both interactive and coordinator UIs handle the new
``attachments_busy`` response with a chat-surface error bubble:
> Attachments can't be sent while the assistant is working.
> Send a text-only message now, or wait and resend with attachments.
Chips stay in the composer so the user can retry once idle.
**Tests.** Replaced the now-impossible ``TestQueuedWithAttachments``
class with a rejection-coverage class. Rewrote the
``_queue_with_attachment`` route-test fixture to reserve directly
via ``reserve_attachments`` (the queue path no longer reaches the
reserved state). Added a route-level test for the new
``attachments_busy`` contract.
* Bound search tool output against pathological inputs
Replaces the per-line truncation with a fully bounded pipeline so the
search tool can no longer overflow the LLM context — or OOM the parent —
on minified bundles, multi-GB JSONL records, or huge result sets.
Backend:
- Prefer ripgrep when on PATH; grep is the fallback. Detection is
cached via functools.cache.
- ripgrep flags do most of the bounding natively: --max-columns 1024
+ --max-columns-preview, --max-filesize 10M, --max-count 100,
--no-config, --no-messages, plus negative globs for the same
noisy directories grep has been excluding.
- ripgrep added to the Dockerfile.
Streaming subprocess (_search_capture):
- subprocess.Popen with a streaming, byte-capped stdout read (4 MB).
Defends against single-line files (training data, minified bundles)
that would have OOM'd the previous subprocess.run capture.
- threading.Timer watchdog enforces tool_timeout even when the
pipe read is blocked in the kernel — proc.wait(timeout=…) alone
was insufficient because the read sat ahead of it.
- Stderr drained in a daemon thread to avoid pipe-deadlock when the
child writes to stderr while we're still reading stdout. Cap on
captured stderr keeps a hostile child from growing the buffer.
Tier-based formatter (_format_search_results):
- Tier 1: full path:line:content output, stream-emitted with a
running-cost short-circuit so we never materialize past the budget.
- Tier 2: K samples per file with overflow notes; K is computed
analytically from budget / file_count / avg-line-length so we hit
the right ladder rung in a single pass.
- Tier 3: per-file counts only, also budget-bounded with a tail line
reporting the omitted files. Sorted by descending count.
- Total output budget (32 KB) is well under tool_truncation, so the
head+tail _truncate_output strategy never silently drops middle
files in a search result.
Argument injection fix:
- The ripgrep arg list was missing the `--` separator that the grep
branch already had. With auto_approve on the search tool, that was
exploitable: path='--pre=COMMAND' would have made ripgrep run the
script as a per-file preprocessor and surface its stdout. Added
`--` and a regression test.
State-machine cleanup in _exec_search:
- rc < 0 (signal-killed by something other than us) now surfaces a
dedicated 'killed by signal N' message instead of being parsed as
success.
- capped + zero parsed records (e.g. one multi-MB line with no \n)
now returns a dedicated byte-cap message instead of the malformed-
output message that previously masked the real cause.
- _report_tool_result descriptions now match the returned payload
(no more 'no matches' tag on a 'malformed' payload).
Defence-in-depth on env scrub:
- RIPGREP_CONFIG_PATH, GIT_CONFIG, GIT_CONFIG_GLOBAL, GIT_CONFIG_SYSTEM
added to _EXPLICIT_SCRUB. We pass --no-config on the rg CLI today,
but if a future caller forgets the flag, an attacker who can set
one of these env vars could plant a config containing --pre=… and
recreate the same RCE shape.
Tests:
- TestSearchLineTruncation rewritten to mock _search_capture instead
of subprocess.run (the previous tests passed ChatSession kwargs
that no longer satisfy the constructor).
- TestSearchBackendSelection covers rg/grep detection and arg
construction, including the --pre flag-injection regression.
- TestSearchOutputBudget exercises Tier 1/2/3 directly.
- TestSearchCaptureStreaming spawns real Python subprocess writers
to exercise the byte-cap trim, mega-line-no-newline edge case, the
watchdog timeout when the child writes nothing, and the stderr
drain under load.
- test_env_scrub picks up the new tool-config keys.
* Address Copilot review on #473
- Budget the Tier 2/3 header up front so the formatter's emission stays
strictly within _SEARCH_OUTPUT_BUDGET. Previously the fit checks only
counted body bytes, letting the final string overflow by ~120 chars
(header + separator) and triggering _truncate_output's head+tail
dropout — exactly the shape this code was trying to avoid.
- Restore the (5, 3, 1) ladder in Tier 2: the analytical K from perf-2
is kept as a starting estimate, but if that K's actual emission
doesn't fit (the estimate ignores the header and overweights shared-
path compression) we step down through the ladder before falling
through to Tier 3. The previous one-shot K could collapse to counts-
only when 3/file or 1/file would have fit.
- Only normalise rc to 0 in the capped-output path when rc < 0 (our
SIGKILL). There's a narrow race where the child can exit naturally
between our read and our kill; preserving a non-negative rc means
rg's rc=2 ('matches found but some files had errors') no longer
silently turns into a clean success when the byte cap also fires.
- Clarify _MAX_SEARCH_LINE_LENGTH doc: the cap applies to the content
portion (after path:lineno:), not the whole emitted line.
- Add explanatory comments on the two intentional `except Exception:
pass` blocks in _search_capture (stderr drain, pipe close in the
cleanup finally) so static analysis and future readers can see the
silence is deliberate.
- Tighten the budget tests: now assert strict `<= _SEARCH_OUTPUT_BUDGET`
instead of the +512-char slack that was masking the header overflow.
- New regression tests:
- Tier 2 ladder step-down (K=5 over budget, K=3 fits, no Tier 3 fall-through)
- capped + rc=2 surfaces stderr instead of being normalised to success
- capped + rc<0 (our SIGKILL) flows through as a partial-result success
* chore(search): post-review cleanup
Follow-up to the Copilot-review fixes in 39d2aa2 — these are all small
quality items (no behaviour change, no new tests).
- q-1: collapse the Tier 2 candidates filter to a single expression.
Drops the redundant inner ``max(estimated_k, 1)`` and the unreachable
``if not candidates`` branch (the ladder ends in 1 and ``estimated_k``
is already floored at 1, so the comprehension always yields ≥ ``[1]``).
``or [...]`` is kept as defence against future ladder changes.
- q-2: update _format_search_results docstring to match the new ladder
semantics (analytical seed → step down through (5, 3, 1) from the
highest rung ≤ the estimate). The previous wording suggested every
Tier 2 attempt started at 5.
- q-3: combine the two ``from turnstone.core.session import ...``
statements in test_tier2_steps_down_ladder_before_falling_to_tier3
into a single top-of-function import (matches the surrounding tests).
- q-4: shorten the explanatory comments on the two best-effort cleanup
paths in _search_capture to one line each. Both sites now read with
the same shape ("# best-effort: pipe may be torn down by ...").
- q-5: trim the _MAX_SEARCH_LINE_LENGTH comment from 7 lines back to 3.
Keeps the load-bearing semantic (cap is on the content portion only)
and the pathological-line defence; drops the paths-aren't-bounded
parenthetical, which was background reading rather than WHY.
* feat(providers): api_surface toggle + mistral medium reasoning fix
Mistral medium open-weights served by vLLM expects reasoning_effort via
the Responses API (`reasoning.effort`), not as a `chat_template_kwargs`
entry on Chat Completions. The session was unconditionally injecting
`{"reasoning_effort": ...}` into `chat_template_kwargs` for every
openai-compatible request, which corrupted the prompt rendering for any
backend whose chat template didn't consume that key (Mistral medium,
Mistral cloud, Groq, OpenRouter).
Changes:
- Add `api_surface` ("chat" | "responses") to `ModelConfig.server_compat`
and thread it through `create_provider` / `model_registry.get_provider`.
`openai-compatible` defaults to Chat Completions; operators can flip
individual aliases to Responses for endpoints that support it.
- New `vllm-mistral-medium` profile that pre-fills api_surface=responses
on Detect for known Mistral medium model ids.
- Drop the unconditional `reasoning_effort` injection into
`chat_template_kwargs`. Operators running gpt-oss-style local
templates that consume `reasoning_effort` from the chat template now
opt in via `server_compat.extra_body.chat_template_kwargs`.
- New "API Surface" select in the Models admin tab; allowlist-validated
server-side at create/update time; pre-filled by Detect via the
profile suggestion.
- Evict the cached provider singleton in `ModelRegistry.reload()` when
api_surface changes (previously only cfg.provider triggered eviction).
- Fix `_run_agent` fallback path to inherit the session's primary alias
for capability and server_compat resolution; previously the fallback
passed `alias=None`, which silently dropped per-model caps on the
agent path.
Tests: 5117 passed (-m "not live"); ruff + mypy clean.
* fix(providers): don't auto-suggest Responses for Mistral medium
vLLM's Responses API surface for Mistral medium open-weights doesn't
wire up the Mistral tool-call parser as of vLLM 0.x — tool calls leak
into the response as ``[TOOL_CALLS]<name>{...}`` text instead of
structured tool_calls. Chat Completions on the same engine handles
tools cleanly via ``--tool-call-parser mistral``, and reasoning can be
turned on via the vLLM CLI ``--reasoning-parser`` flag.
Drop the auto-suggest mapping so Detect falls back to the generic
``vllm`` profile. Keep the ``vllm-mistral-medium`` profile definition
in place so an operator who specifically wants per-request effort and
accepts the tool-calling limitation can still pick "Responses API"
manually in the admin UI.
* fix(providers): address Copilot review on PR #469
- providers/__init__.py: drop the redundant *_responses_provider /
*_chat_provider names; have create_provider use _openai_provider and
_openai_compat_provider directly so they're not flagged as unused
globals.
- console/server.py: tighten _validate_api_surface to a strict equality
match against the canonical {"chat", "responses"} set. The previous
strip().lower() membership check accepted ' Responses '/'CHAT' but
stored the raw string verbatim, which then failed to round-trip
through the admin <select>.
- console/static/admin.js: gate the entire server_compat block (server
type, api_surface, extra_body) on provider == "openai-compatible" at
save time so toggling provider away can't leave a stale hidden surface
selection in the persisted capabilities JSON.
- tests/test_session.py: splat the bad kwarg via **dict so CodeQL no
longer flags the call as a wrong-name keyword (the point of the test
is the runtime contract, not the static type).
- tests/test_admin_model_registry_refresh.py: add endpoint-level tests
for the api_surface validation on both create and update — covers the
bogus-value rejection, non-canonical-string rejection, and the happy
path persisting through to the refreshed registry.
* fix(memory): query-aware candidate selection + OR-of-terms search
The system-message memory composition path used a recency-ordered
candidate set (`_list_visible_memories(limit=fetch_limit)`). On
deployments with more than `fetch_limit` (default 50) visible
memories, BM25 only ever ranked the 50 most-recently-touched memories
— a relevant memory written months ago was silently invisible
regardless of how well it matched the recent context. Multi-word
search at the SQL layer used AND-of-terms, killing recall on any
multi-word query without an exact field overlap.
## Functional changes
- `_init_system_messages` (`turnstone/core/session.py`): extract
recent context first, then `_search_visible_memories(context)` to
pull query-aware candidates. Search hits below `fetch_limit` union
with the recency list (deduped by memory_id) so the BM25 candidate
pool is always a SUPERSET of the prior recency-only pool — even on
noisy queries where the cap fills with stopwords, the recency-50
the original bug surfaced still reaches BM25. Empty context skips
search entirely. Candidate-selection logic extracted into
`_select_memory_candidates`.
- `search_structured_memories` (PostgreSQL + SQLite): per-term
clauses join with OR instead of AND. A row matches if ANY term
matches ANY of name/description/content. Downstream BM25 narrows
back down by relevance.
## Perf hardening
- Collapse the 1-3 fanned scope queries into a single SQL. New
backend methods `list_visible_structured_memories` /
`search_visible_structured_memories` union the visibility scopes
into one WHERE OR-group, so a composition rebuild now hits the DB
at most twice (search + recency) instead of up to six times.
- Cap and normalize search terms. Composition can hand a multi-KB
pasted message to ILIKE-based search; without a cap, every distinct
token would emit one unindexable predicate per scope-fanned query.
`normalize_search_terms` (`storage/_utils.py`) de-dupes
case-insensitively, drops <2-char tokens, and hard-caps at 16.
- Per-turn search cache. `_init_system_messages` fires from many
call sites within one turn (state transitions, MCP refresh, tool
results) and the recent-context query is identical across them.
Session-instance cache keyed by (query, mem_type, limit) absorbs
the duplicates; invalidated in `_append_user_turn` and after
memory save/delete tool actions.
- Stable secondary sort by `memory_id`. `updated` is second-precision
and `touch_structured_memories` can land a batch on identical
timestamps; without a tie-breaker SQL returns rows in
implementation-defined order, BM25 input shuffles, and the
LLM-side prompt cache misses across calls. All four backend ORDER
BYs now break ties on `memory_id ASC`.
## Quality cleanups
- Coalesce `memory.search.term_count` + `memory.search.zero_results`
into a single `memory.search` log carrying both `term_count` and
`result_count`.
- New `memory.composition` log: source / candidates / injected.
- Promote a shared `make_chat_session` factory to `tests/_helpers.py`.
- Rename SQL builder local `extra` -> `scope_filters` for clarity.
- Add docstrings on `search_structured_memories` so the AND->OR flip
survives future readers.
## Tests
Adds 20 tests across `tests/test_structured_memory.py`,
`tests/test_structured_memory_storage.py`, and
`tests/test_memory_relevance.py`: recency-ceiling regression,
empty-query fallback, sparse-match union, recency-preserved-when-
search-returns-noise (locks in the pool-superset invariant),
OR-of-terms on both backends, scope filtering preserved,
search-facade multi-word behavior, term-cap normalization, the new
visible-scope helpers (list + search + empty-scopes guard),
coord-scope composition isolation, end-to-end
`memory(action='search')` tool execution, per-turn cache hit +
invalidation, and stable ordering under tied `updated` timestamps.
Memory test sweep: 102/102. Broader regression
(session, storage, coordinator, load_skill): 411/411.
* fix(memory): address Copilot review on PR #468
Three follow-ups from Copilot's inline review:
1. SUPERSET invariant violation (Copilot, session.py:5510).
`(search_hits + extra)[:fetch_limit]` capped the union back down to
fetch_limit, evicting the recency tail when search added distinct
hits. Recency tail is exactly where ancient-but-recently-touched
memories live — the recall this PR is supposed to improve — so
tail eviction recreated the bug for the narrow case where a query
term fell off the 16-cap and the matching memory sat in
recency[40-49]. Drop the cap; both halves are already SQL-capped
at fetch_limit, so the union is at most 2 × fetch_limit (~100 with
defaults). BM25 over 100 candidates in pure Python is sub-ms;
irrelevant recency fillers get score=0 and don't pollute ranking.
Updates the docstring to actually be honest about the invariant.
Adds `test_recency_tail_preserved_when_search_adds_distinct_hits`
that locks the behavior in: 5 search hits + 10 recency = 15-item
pool, every recency item present, source="union".
2. Unbounded `query.split()` in normalize_search_terms (Copilot,
_utils.py:74). `str.split()` allocates the full token list before
the cap-after-16 break, so a 100KB pasted query did MB of throwaway
work even though only 16 tokens entered SQL. Switch to
`re.finditer(r'\S+', query)` — streaming iterator, stops scanning
at the first 16 normalized terms regardless of input size.
3. Misleading + unbounded log term_count (Copilot, session.py:8571).
`len(item["query"].split())` had two problems: same unbounded
split as #2, and the value reported the raw input token count
rather than the normalized term count that actually hit the SQL
WHERE clause — misleading metric for an operator trying to
understand storage-side behavior. Switch to
`len(normalize_search_terms(item["query"]))` — accurate count, and
bounded for free via #2.
Refuted: github-code-quality flagged `...` bodies in the new Protocol
methods as "statement has no effect." False positive — `...` is the
canonical Protocol body convention, used 213 other times in the same
file.
Memory test sweep: 103/103. Broader regression: 411/411.
CI failure on main: test_publish_records_metric_outcome saw an empty
calls list — its monkeypatch was patching a different metrics
instance from the one `_publish_models_metadata` reads.
Two changes:
- test_close_reason_persistence.py: replace the bare
`srv_mod._metrics = MetricsCollector()` assignment in `_make_app`
with an autouse `monkeypatch.setattr(srv_mod, "_metrics", ...)`
fixture so the test's metrics swap auto-restores. Other test
files (test_auth.py, test_server_attachments_endpoints.py) carry
the same anti-pattern; left for a follow-up since they're not on
the critical path here.
- test_server_node_models_metadata.py: switch the publish-helper
metric test to a string-form `monkeypatch.setattr("turnstone.
server._metrics", FakeMetrics())` so it replaces whatever binding
the live module currently holds, regardless of what other tests
did to it. Robust against future leaks of the same shape.
* feat(coord): expose healthy model aliases per node on list_nodes
Surfaces a `model_aliases` field on each `list_nodes` row so a
coordinator can discover which model aliases each cluster node will
accept on `spawn_workstream(model=...)` without an HTTP fan-out.
Each server projects its registry into a `models` entry on
`node_metadata` (`{alias, provider, healthy}` per alias) at lifespan
startup, on every 30s heartbeat tick, and after `internal_model_reload`.
The publish helper short-circuits on a payload-equality cache so a
stable cluster doesn't pay UPSERT churn — exposed via the new
`turnstone_node_models_publish_total{outcome="written|skipped"}`
Prometheus counter so operators can graph cache hit-rate.
Coord client filters the per-alias rows to healthy aliases only and
drops the provider-side model identifier (`cfg.model`) — coords kept
reaching for it when they should pass the local alias.
* fix(coord): address Copilot+CodeQL feedback on list_nodes models work
- internal_model_reload: reuse a single get_storage() local across the
registry load and the metadata publish (Copilot:3047)
- _collect_node_models_metadata: iterate sorted aliases so two
structurally identical registries built in different insertion orders
serialize to the same JSON — directly improves the publish-cache hit
rate exposed via turnstone_node_models_publish_total (Copilot:3105)
- tests: drop mixed turnstone.server import style flagged by CodeQL —
hoist _metrics into the from-import block, and use sys.modules in
the shutdown-race regression test instead of `import as srv`
Address Copilot feedback on PR #465:
1. The has_alias fallback in both session_factories silently rewrote
any unknown caller-supplied alias to the default, including on the
fresh-create path where the create handler maps the factory's
ValueError to a 503 with operator-friendly text. A typo in
body.model would now silently start a workstream on the default
instead of telling the caller their requested model could not be
resolved. Move the fallback out of the factories: each factory
raises again on unknown aliases, and SessionManager filters stale
aliases out of the rehydrate path via a new ``model_validator``
constructor kwarg (production wiring passes ``registry.has_alias``
on both interactive and coordinator).
2. ChatSession.resume()'s elif branch flipped self.model to the
persisted model name even when the alias was unresolvable, leaving
the session paired with the constructor's default provider/client
but a removed model name — a broken state whose next API call
fails. Drop the model copy: keep the constructor's coherent
default (provider + model + capabilities) and just log the
unreachable saved values so the missing alias is auditable.
Tests:
- Move stale-alias coverage from the factory level into
SessionManager (tests/test_session_manager.py): validator drops
stale aliases before reaching build_session; live aliases pass
through unchanged.
- tests/test_sessions.py renamed test_resume_restores_model →
test_resume_keeps_defaults_when_alias_unresolvable to match the new
contract.
SessionManager.open() was calling build_session(ws) without a model
arg on the rehydrate path. The session_factory then resolved the
*current* default alias, ChatSession.__init__'s _save_config() (INSERT
OR REPLACE per-key) clobbered the persisted workstream_config with
those defaults, and the subsequent resume() "restored" what was now
the default — silently resetting model_alias, model, temperature,
reasoning_effort, max_tokens, skill, creative_mode, instructions,
token_budget, and notify_on_complete on every reopen and every
service restart, for both interactive and coordinator workstreams.
Three layers:
1. SessionManager.open() now reads workstream_config via
self._storage.load_workstream_config(ws_id) and threads the saved
model_alias into build_session(ws, model=saved_alias).
2. ChatSession.__init__ now skips its initial _save_config() when a
workstream_config row already exists for self._ws_id — protects
every other persisted knob without having to plumb each one
through the adapter signature, and catches any future construction
path that forgets to thread model through build_session.
3. Both session_factories (server.py interactive, console
session_factory.py coordinator) now treat an unknown caller-
supplied alias the same as an unset alias: fall back to the
runtime default rather than raising. Without this, a workstream
pinned to an alias an operator has since removed from the registry
would 500 on every reopen — defeating the "best effort restore,
default if the original is gone" contract this fix is meant to
deliver. Mirrors _effective_default_alias's existing has_alias
guard against a stale ConfigStore default.
Three changes from PR review:
- Permission gating: hide the Roles sub-tab button when the user
lacks ``admin.settings``. The sub-tab loads/saves through
``/v1/api/admin/settings``, so an admin with ``admin.models`` but
no ``admin.settings`` would otherwise see a perpetual 403 loader.
When Roles is the active sub-tab and the permission check fails,
snap the panel back to Definitions so the user lands somewhere
usable.
- Drop the redundant ``/v1/api/admin/model-definitions`` fetch from
``loadAdminModelRoles``. Both entry points (initial Models-tab
open + ``models_changed`` SSE refresh) flow through
``loadAdminModels`` first, which already populates ``_modelDefs``
+ ``_modelDefaultAlias``; ``_saveModelRole`` doesn't touch model
definitions, so the cached snapshot stays accurate when the save
chains back here. Halves the per-render request count and
removes a wasted round-trip on every cluster-wide model edit.
- Add ``test_models_changed_event.py`` covering the SSE fanout the
prior commit introduced: each model-definition CRUD endpoint
emits exactly one ``models_changed``, settings PUT/DELETE only
emit for keys in ``_MODEL_AFFECTING_SETTING_KEYS`` (parametrised
over all eight), and unrelated settings (e.g.
``session.retention_days``) don't trigger spurious refreshes.
The expected key set is pinned in the test so a stray addition
to the allowlist doesn't silently bypass coverage.
Same shape as the coordinator/judge rows already there: alias dropdown
+ reasoning_effort dropdown sourced from the existing
``model.plan_alias`` / ``model.plan_effort`` and
``model.task_alias`` / ``model.task_effort`` settings. Adds the four
keys to the SSE ``models_changed`` allowlist so changes from the
Settings API also trigger a live dropdown refresh, and filters them
out of the Settings tab so they only render in one place.
Lifts judge and coordinator model assignments out of their respective
admin tabs and into a new Models → Roles sub-tab so role overrides live
next to the model definitions they reference. Forward-looking shape for
the upcoming perception.{audio,image,video} model settings — adding a
new role is one entry in the declarative MODEL_ROLES array.
Also drops the misleading "Coordinator subsystem not configured" home
banner. The session factory already falls back to the registry's
default model when coordinator.model_alias is unset, so the banner was
nagging on fresh installs where the system was actually working. The
related _probeCoordSubsystem / _homeCoordReady plumbing went with it.
Wires SSE-driven live refresh: the console now emits a models_changed
event when a model definition is created/updated/deleted/reloaded, or
when a model-affecting setting (model.default_alias, judge.model,
coordinator.model_alias, coordinator.reasoning_effort) changes.
Connected browsers refetch /v1/api/models on receipt so the home
composer's model dropdown and the Roles sub-tab stay accurate without
a manual reload — fixes the case where editing the underlying model
for an existing alias left the dropdown showing the old model id.
Companion cleanups:
- Renamed .judge-section-* CSS classes to .admin-subtab-* and shared
them with the Models sub-tab switcher (same a11y attrs, arrow-key
nav). Old names had no other callers.
- Filtered judge.model out of the Judge Settings sub-tab and
coordinator.model_alias / coordinator.reasoning_effort out of the
Settings tab — they live exclusively under Models → Roles now.
- Reworded the _require_coord_mgr 503 messages to point operators at
the Models tab instead of suggesting they set coordinator.model_alias.
* fix(console): home composer attachments + coord chat user-message pills
Two parity gaps in the console's coordinator surface:
- The embedded creator on the home page accepted only text — the
paperclip / paste / drop pipeline that the in-coord composer and the
interactive new-ws modal both expose was missing, so a user couldn't
attach files at create time. Stage Files in memory (no ws_id yet) and
ship them multipart on Start; the coord create endpoint already accepts
multipart via create_supports_attachments=True.
- User messages with attachments rendered as plain text on both live
send and history replay — no chip cluster like the interactive pane.
Added appendUserMessageWithAttachments and a structured userAttachments
list built from _attachments_meta (preferred) or the multipart parts
themselves, then rendered the same .msg-user-attach pill strip the
interactive pane uses.
Polish from a designer pass:
- Pill background was --panel-2, equal to the .msg bubble background in
both themes (border contrast ≈1.4:1, below WCAG 1.4.11). Switched to
--panel so the pill sits on a different surface than the bubble.
- Capped chip filename width inside the home composer (max-width 200px +
ellipsis) so a long filename doesn't push the strip past the textarea.
- aria-live="assertive" → "polite" on #home-coord-error; client-side
validation isn't an interrupt-level event.
- Reserved min-height on .home-composer-error and dropped the
display: none/block toggling so validation messages no longer reflow
the active-coordinators list below.
* fix(console): address PR #462 review feedback
- Block home-composer submit when files are staged but the task field is
empty. Server's _coord_create_post_install short-circuits on an empty
initial_message, so the multipart upload would create pending
attachment rows that never reserve onto a turn — orphaned until the
GC sweep. Fail in the browser instead.
- Drop the redundant `part &&` guard in coordinator.js's history-replay
multipart loop; the earlier `if (!part || ...) continue` already
filtered.
- Rewrite the home-mount .composer-chip-name CSS comment. shared/chat.css
defines .composer-chip{,-size,-remove} but no .composer-chip-name rule
— the span inherits the parent chip font with no width cap.
- Add smoke-guard string assertions in test_coordinator_page.py for
appendUserMessageWithAttachments and msg-user-attach so a future
rename can't silently regress the attachment affordance.
* fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration
Loading a saved workstream silently dropped tool results and missed
verdict / output-guard / truncation signals on replay. Root cause was
in `Pane.prototype.replayHistory`: an assistant message carrying both
content and tool_calls cleared the `lastToolBlock` anchor before the
following tool-result iteration could attach. The fix reorders content
to render before the tool block (matching live SSE order) and
restructures the tool-result branch to anchor by `data-call-id` so
multi-tool batches render `[hdr A][out A][hdr B][out B]` rather than
bunching outputs at the bottom.
Beyond the bug, replay now reaches near-parity with the live UX:
- Persisted intent verdicts and output_assessments flow through both
the SSE replay (`_build_history`) and the `/history` REST endpoint
used by coord. Single shared helper module owns the wire shape.
- Memory/recall calls persist instead of being filtered at storage
time — full audit trail; UI dims them by default with hover-reveal
so heavy memory usage doesn't crowd the narrative.
- Truncation indicator surfaces as a sibling pill (consistent across
interactive + coord) when a tool result hit the 2000-char cap.
- `replayHistory` wraps DOM work in `aria-busy` so screen readers
don't get a chatty announce-flood on long replays.
- `_build_history`'s storage I/O moves off the event loop via a new
`events_replay_prepare` async hook for the SSE path; other async
callers wrap in `asyncio.to_thread`.
Coord parity:
- `/history` REST endpoint decorates tool_calls with verdict +
output_assessment + truncation flag (was previously raw
`load_messages` output).
- Coord JS stamps `judge_verdict` / `heuristic_verdict` from
history-loaded `tc.verdict` so the existing batch render paints
the persisted pill, seeds the verdict cache to dedupe later live
SSE events, and emits an inline `.coord-tool-row-warning` chip
per call instead of a generic chat line.
- Memory/recall dim rule mirrored on `.coord-tool-row[data-tool-name=...]`.
* fix(replay): address PR #461 review feedback + raise tool-result storage cap
Copilot review feedback:
- Sibling-chain dim rule (memory/recall) now adds :focus-within
alongside :hover for .tool-output / .media-embed / .output-warning
/ .tool-output-truncated — keyboard users tabbing into a faded
subtree now get full opacity.
- ``cfg.open_post_load`` is now invoked via ``await asyncio.to_thread``
so its sync ``_build_history`` call (storage I/O for verdict
indexes + message reconstruction) doesn't block the event loop on
every workstream open. Mirrors the SSE replay path that's already
protected via ``events_replay_prepare``.
- Replaced the hardcoded ``2000`` literal in server.py and session.py
with ``TOOL_RESULT_STORAGE_CAP`` from the shared decoration module
so the UI truncation-pill detection can't silently desync from the
storage write side.
While here:
- Raised ``TOOL_RESULT_STORAGE_CAP`` from 2000 → 10000. A 2000-char
clip routinely cut grep / file-read bodies mid-line, leaving the
audit trail useless for retrospective debugging. FTS5 + row size
grow proportionally; the per-tool upper bound is still bounded
upstream by ``_truncate_output``'s context-budget clamp.
- Updated the user-visible truncation-pill tooltip on both
interactive and coord to reflect the new cap.
- ``test_decorates_tool_calls_and_marks_truncated`` now references
the constant instead of a literal so it stays correct on future
cap changes.
Speculative reliability machinery from the Stage 3 push that turned
out not to address any user-visible bug. The actual fixes (state /
activity disjunction in handleChildState, bulk-fetch race fix in
_fetch_live_block, push approve_request via cluster bus) are what
resolved the wedged-row issues. Manual testing showed the per-tab
SSE listener queue depth never climbed past single digits even when
rows were stuck — overflow was never the cause.
Removed
- ``_CRITICAL_EVENT_TYPES`` + ``_put_with_priority`` helper.
- Per-tab listener queue selective drop (back to plain
``contextlib.suppress(queue.Full)`` everywhere).
- ``ClusterCollector._fanout`` reverts to the same.
- WebUI ``_broadcast_intent_verdict`` / ``_broadcast_approval_resolved``
/ ``_broadcast_approve_request`` revert to plain ``put_nowait``.
- ``_queue_stats`` periodic SSE emit + frontend status-bar indicator
+ the supporting CSS rules.
- Broken ``.approval-block`` ``transition: max-height`` /
``max-height: 80vh`` / ``overflow: hidden`` rules — the transition
never fired (nothing toggled max-height) and ``overflow: hidden``
clipped long verdict reasoning. Layout-shift on auto-expand jumps
again, which is preferable to clipped content (Copilot review).
Tidied
- ``_CollectorProtocol`` / ``_ManagerProtocol`` method bodies switch
from ``...`` ellipsis to docstring-only bodies, silencing four
CodeQL "statement has no effect" warnings without changing the
Protocol contract.
5024 passed, ruff + mypy clean.
Lift the Children primitive out of CoordinatorAdapter into universal
SessionManager core primitives, replace the fragile poll + state-event
piggyback paths with first-class cluster bus event types for inline
approval delivery, and clean up the resulting frontend reducer.
Architecture
- New `turnstone/core/children_registry.py` — universal parent → children
+ reverse-lookup primitive with atomic `add_child` (returns parent UI
for race-free dispatch). Lifted from `CoordinatorAdapter`.
- New `turnstone/core/child_source.py` — `ChildSource` Protocol with
`SameNodeChildSource` (in-process via SessionManager state observer)
and `ClusterChildSource` (cross-node via ClusterCollector listener).
- `SessionManager._on_state_change` upgraded to multi-subscriber
(`subscribe_to_state` / `unsubscribe_from_state`) under a dedicated
lock; CLI consumer migrated.
- `CoordinatorAdapter` shrunk: 731 → ~640 LOC. Children data lives in
the registry; fan-out lives in ClusterChildSource. Backward-compat
property facades dropped; tests updated to use the registry surface.
Cluster bus event vocabulary
- New event types `intent_verdict`, `approval_resolved`,
`approve_request` flow through both `ClusterCollector._apply_delta`
(translation from node SSE) and `emit_console_ws_*` (synthesis on
console pseudo-node).
- `CoordinatorAdapter._dispatch_child_event` re-emits as
`child_ws_intent_verdict` / `child_ws_approval_resolved` /
`child_ws_approve_request` on the parent coord's SSE stream.
- New `_broadcast_intent_verdict` / `_broadcast_approval_resolved` /
`_broadcast_approve_request` no-op hooks on `SessionUIBase`. WebUI
pushes to the global queue; ConsoleCoordinatorUI pushes to the
collector. `approve_tools` calls `_broadcast_approve_request` right
after setting `_pending_approval` so the items reach the coord tree
immediately, eliminating the bulk-fetch race.
Cleanups
- `pending_approval_detail` piggyback on `ws_state` / `cluster_state`
removed end-to-end. Bulk fetch + explicit verdict / approve-request
push are the canonical carriers.
- Browser `_judgePollTick` 90-second poll loop deleted; push path is
authoritative.
- `urgent` flag on `scheduleLiveFetch` deleted (only caller was 409
retry; replaced with `invalidateLiveBadge` + standard schedule).
- Console `_fetch_live_block` derives `pending_approval` from a
disjunction (`activity_state="approval"` OR `state="attention"`
OR detail present) so the bulk fetch can't return false during the
state-transition race window.
- Coord-side merge guard in `flushLiveFetches` no longer clobbered:
`handleChildState` only stamps `sseUpdatedAt` when authoritatively
clearing detail.
- `child_locality` capability flag removed (was inert dead code).
Reliability
- Selective drop on listener queue overflow: critical event types
(verdicts, approvals, ws_closed, child_ws_*) evict one oldest item
to make room rather than dropping themselves on a full queue.
Best-effort events (state ticks, content tokens, status, activity)
drop as before. Applied to `SessionUIBase._enqueue`,
`ClusterCollector._fanout`, and the `WebUI._global_queue` puts in
the new broadcast hooks.
- `_state_subscribers` snapshot under a dedicated lock so concurrent
subscribe / unsubscribe during dispatch can't shift the iterator.
UX / a11y
- Loading placeholder in renderChildRow keeps row height stable while
the bulk fetch is in-flight (sr-friendly aria-label).
- Focus preservation across `_renderChildrenNow` (capture +
restore by row + marker) and across targeted `_updateChildRow` swaps.
- Layout-shift transition on the approval block max-height; respects
`prefers-reduced-motion`.
- Sidebar pending count: `(N children · M pending)`.
- Risk pill `aria-label` spells out level + confidence for SR users.
- Per-coord SSE listener queue depth surfaced in the status bar
(`queue N/500`) with color escalation (warn at >50%, danger at >80%).
Tests
- 305+ test changes across 8 files. New unit tests for
`ChildrenRegistry`, `ChildSource` (both impls + multi-subscriber
observer), the new collector emit + apply_delta cases, the dispatch
cases for new event types, the broadcast hook overrides on both
WebUI and ConsoleCoordinatorUI, and the focus / placeholder /
pending-count frontend assertions in `test_coordinator_page.py`.
5024 passed, ruff + mypy clean.
* feat(console): multi-select delete UX for Saved Coordinators
Mirror the per-server "Saved Workstreams" multi-select delete onto the
console's "Saved Coordinators" section. Coordinator deletes go through
the existing routing proxy at POST /v1/api/route/workstreams/delete
(body-keyed by ws_id, since coordinators live on the node that owns
them) — no backend change required.
Pagination caps the visible page (and therefore the Select-All fan-out)
at 24. Without it, a Select-All on a busy cluster would pin the
console proxy pool with hundreds of parallel deletes through the
fan-out router. While in delete mode the saved-coordinators list is
frozen against SSE re-renders so visible cards don't shuffle out from
under the user's selections (drained on cancel / post-delete close).
Refactor: shared logic now lives in turnstone/shared_static/cards.{css,js}.
* .ws-delete-* CSS moved out of ui/static/style.css into the shared
sheet alongside .dashboard-card; the existing ui/static modal
markup picks up class hooks instead of id-scoped rules.
* createSavedCardsController() owns mode state, checkbox decoration,
toolbar wiring, focus trap, modal lifecycle, and batch fan-out.
Both ui/static (Saved Workstreams) and console/static (Saved
Coordinators) instantiate one controller; ui/static is now ~300
LOC lighter as a result.
* Internalises stale-selection prune across SSE re-renders, the
wsId->item lookup map (was O(selected x N)), and the aria-hidden
wrap on the toggle button's emoji glyph.
Designer review tightened the affordance:
* Modal close restores focus to the toggle button (was landing on
<body>) — WCAG 2.4.3.
* Modal [role="alert"] gets a red-chip treatment when populated,
stays invisible at rest via :not(:empty).
* Pagination consolidated onto the existing .pagination control
(terse "X / Y" label + arrow-glyph buttons) instead of a parallel
.coord-pagination treatment.
* Filled destructive buttons darkened to #dc2626 in dark theme so
the white label clears WCAG AA contrast (was 3.0:1 on --red).
Light theme keeps --red unchanged (5.9:1 already passes).
* Toolbar wraps below 700px viewport — Delete Selected drops to its
own full-width row underneath count + Cancel + Select All for
thumb-target separation.
* .ws-card-check:focus-visible outline + word-break on
.ws-delete-item for narrow-modal long aliases.
* fix(cards): address Copilot review feedback on PR #458
* closeModal focus restore now falls back to the section toggle button
(opts.buttonId) when prevFocus is hidden or detached. The post-delete
Close path runs cancel() before closeModal(), which puts the bar at
display:none — so the captured prevFocus (the bar's "Delete Selected"
button) is no longer focusable and focus would land on <body>,
defeating the WCAG 2.4.3 fix. Esc / Cancel paths still land on the
original focus owner because the bar stays visible in those flows.
* Saved Coordinators onClose drains _savedCoordsRetry before reloading.
Without it, SSE events that arrived during the delete-mode freeze
leave the retry flag true, so loadSavedCoordinators's .finally()
re-fires a second fetch immediately after the first resolves. Mirrors
the same idiom in cancelCoordDeleteMode.
Three issues from the Copilot review on PR #457:
1. SQLite race in bulk_close_stale_orphans (Copilot): the SELECT-then-
UPDATE flow doesn't re-apply the eligibility predicates on the
UPDATE, so a row that gets touch_workstream-bumped (or set_state-
transitioned) between the two statements would still be flipped
to closed. Postgres dodges this via UPDATE...RETURNING (one atomic
statement); SQLite needs the explicit re-application. Fix: rebuild
the WHERE conditions list once, apply on both SELECT and UPDATE,
then SELECT-back by ``state='closed' AND updated=now`` to get the
accurate closed-id list. A row that became fresh between the two
statements skips the UPDATE entirely.
2. SQLite IN-clause bind-parameter limit (Copilot): default 999 cap
could be exceeded on a backlog reap (e.g. after a long outage).
Chunked the candidate id list at 500 — same chunk size
prune_workstreams (line 453) uses for the same reason.
3. Wall-clock-dependent test asserts (Copilot, two locations): the
tests asserted ``updated > '2024-01-01T00:00:00'`` which is fragile
on systems with skewed clocks or pre-2024 dates. Replaced with
``updated != stale_seed`` — captures the same intent (the value
was bumped) without depending on wall-clock date.
Two ``...``-as-no-op flags from github-code-quality were false
positives — ``...`` is the standard Python idiom for Protocol method
bodies and matches every other method in _protocol.py. No code change.
Replaces the ``node_id == self_node_id`` orphan-scoping heuristic from
earlier on this branch with liveness-based scoping using
``services.last_heartbeat``. The heuristic was wrong for the post-#384
world: PR #384 (refactor: replace hash-ring rebalancer with rendezvous
hashing) deleted the rebalancer that used to keep workstreams.node_id
pointing at a live node. Without it, ``workstreams.node_id`` is now
stamped at create time and never updated, so in containerized
deployments with dynamic hostnames a dead pod's rows have ``node_id``
matching no surviving service — they'd accumulate forever under the old
heuristic.
services.last_heartbeat is the same primitive the rendezvous router
uses for routing. Reusing it here keeps reap scoping aligned with
routing: dead pods' rows fall out of the live set after the heartbeat
window and become reapable; alive pods' rows stay protected as long as
they heartbeat.
Mechanics:
- ``bulk_close_stale_orphans`` parameter renamed
``node_id: str | None`` → ``live_node_ids: list[str] | None``. The
WHERE clause becomes ``(node_id IS NULL OR node_id NOT IN
live_node_ids)``. ``None`` skips the filter entirely (single-process
/ tests / operator backfill). ``[]`` treats every row as
unprotected.
- ``SessionManager.close_idle`` pass 2 calls
``storage.list_services(self._service_type)`` to enumerate live
peers, passes their service_ids as ``live_node_ids``. ``_service_type``
is derived from ``self.kind`` (INTERACTIVE→"server",
COORDINATOR→"console") via a module-level mapping — no constructor
param, so production wiring can't miswire the kind/service_type
pairing.
- list_services failure → pass 2 is skipped this tick (conservative;
never reap when liveness state is unknown). Pass 1 still runs.
- ``workstreams.node_id`` with NULL value is always eligible — defends
against ANSI ``NULL NOT IN (...)`` evaluating to NULL (not TRUE) and
silently protecting orphans forever.
- Migration 048 simplified to ``(kind, updated)``; the new query's
``NOT IN (small list)`` predicate against an unbounded-cardinality
column doesn't index well, so leading ``node_id`` would just add
write cost.
Tests cover the live-services protection (own/dead/null cases), the
empty-peers reap-all case, the list_services-failure conservative
fallback, both kind/service_type pairings (interactive→"server",
coordinator→"console"), and the combined live_node_ids +
exclude_ws_ids filter matrix.
bulk_close_stale_orphans runs every min(300s, idle_timeout/4) on
every server and console process. Its WHERE shape is:
WHERE kind = ?
AND state IN ('idle','thinking','attention','running')
AND updated < ?
AND node_id = ? -- multi-node interactive only
At current scale the existing single-column indexes are sufficient —
idx_workstreams_state prunes to non-closed and the planner filters the
rest sequentially. At 100k+ rows that filter becomes a tablescan-
shaped cost.
A partial index covering only BULK_CLOSE_STATE_VALUES rows matches the
reaper's query exactly while staying tiny — closed rows (typically
95%+ of the table) and error rows are excluded, so the index is
roughly 5% the size a full multi-column index would be. Write
amplification only kicks in for transitions touching one of the four
covered states.
Column order (node_id, kind, updated): node_id is the most selective
filter for multi-node interactive (each server prunes to its own
node's rows), kind second so coord-only and interactive-only queries
within a node still get index-only scans, updated last so the range
comparison rides the trailing column.
Postgres uses CREATE INDEX CONCURRENTLY so the build is non-blocking
on a live system; SQLite has no concurrent concept and the table-
level write lock already serializes, so a plain CREATE INDEX is fine.
The console's coord SessionManager had no idle thread — close_idle was
never called for coordinator workstreams. This is the worse half of
the lifecycle leak: the dashboard filters via the in-memory pool, so
DB-only orphan coords were invisible. At empirical diagnosis,
coord closure was 16% (10 closed / 64 total) vs interactive 63%.
Adds _coord_idle_cleanup_thread mirroring turnstone/server.py's
_idle_cleanup_thread but skipping the rate-limiter / global-queue arms
the console doesn't have. Started from the lifespan when coord_mgr is
constructed and server.workstream_idle_timeout > 0 (reuses the
existing setting — same cadence works for both kinds).
Initial sweep runs INSIDE the thread before the first sleep, not
synchronously in the lifespan: cold-start orphans are reaped without
blocking Starlette boot. Important because cold start with many DB
orphans (the precise condition this code targets) is exactly when the
UPDATE is most likely to be slow.
Helper takes an optional stop_event parameter purely for tests —
production callers pass None and the daemon runs for process lifetime.
This avoids the SystemExit-from-stub + module-wide filterwarnings
fragility a previous iteration relied on.
Four tests: initial sweep runs before first sleep, ticks fire each
loop, exceptions don't kill the thread, stop_event exits cleanly.
Real bug: workstream rows accumulate in non-closed states (idle,
thinking, attention, running) when their owning process restarts or
crashes. Empirical diagnosis on a live deployment found ~60 stuck
coord rows in DB invisible to the in-memory-keyed dashboard, plus
100+ interactive rows older than the 2h timeout (one stuck "thinking"
for 2 weeks — impossible across a process restart).
Root cause: close_idle iterates self._workstreams.values() — only the
loaded subset. Anything left behind by a prior process incarnation
sits in DB forever because nothing ever re-loads it.
This commit gives close_idle a second pass.
Pass 1 (existing, unchanged): close loaded IDLE rows whose
ws.last_active (monotonic) is past timeout. IDLE-only so legitimately-
attentive rows (waiting for user response) stay live.
Pass 2 (new): bulk-close DB rows of this manager's kind whose updated
is past the wall-clock cutoff and which aren't currently loaded.
Closes the broader BULK_CLOSE_STATE_VALUES set — any matching row is
by definition not loaded by any process and cannot be in a live
interaction. Scoped by self._node_id so a sibling node can't reap
rows we own (multi-node interactive correctness). No emit_closed —
never-loaded rows have no SSE listeners expecting them.
Lock invariant: pass 1 holds self._lock briefly to snapshot victims
and pop them (existing behavior). Pass 2 holds self._lock briefly to
snapshot the loaded keys, then releases before the DB UPDATE so a slow
reaper query can't block create/get/set_state.
Also fixes a same-process race in open(): the rehydrate path read DB,
released the manager lock, then re-acquired to install — a concurrent
pass 2 between the two acquisitions snapshots loaded keys without the
in-flight ws_id, and could clobber its DB row to closed. open() now
calls touch_workstream(ws_id) on rehydrate so the row's updated is
fresh against any pass-2 cutoff. Pure timestamp write is safe against
concurrent close() (close still wins on the state column).
Three new tests cover the DB orphan pass (basic, exclude-loaded, kind
filter) plus node_id scoping (own/foreign rows, None-skips-filter) and
the open() rehydrate touch.
Two new methods on the StorageBackend Protocol, with implementations on
both Postgres (UPDATE ... RETURNING) and SQLite (SELECT-then-UPDATE in
one transaction). No callers yet — wiring lands in subsequent commits.
bulk_close_stale_orphans(kind, cutoff, exclude_ws_ids, node_id=None)
flips rows in BULK_CLOSE_STATE_VALUES (idle/thinking/attention/running)
to closed when their updated timestamp is lex-older than cutoff. The
node_id filter scopes the reap to a single node's partition — required
for multi-node interactive deployments where each node only has
authority over its own workstreams.node_id rows. Excludes loaded ids
so the in-memory pass owns those.
touch_workstream(ws_id) bumps updated without changing state. Used by
the open() rehydrate path to defend against the orphan reaper clobbering
a freshly-loaded row whose DB updated is older than the cutoff. Pure
timestamp write is safe against concurrent close() because close still
wins on the state column.
BULK_CLOSE_STATE_VALUES is centralized in workstream.py so the two
backend implementations and FakeStorage all agree; if a new transient
state is added to WorkstreamState, deciding whether it joins this set
is part of the change rather than an after-the-fact audit across three
files.
Storage tests (run against both backends via the conftest fixture) cover
the kind/state/cutoff/exclude/node_id matrix plus touch_workstream.
The themed ``tool_reminder`` bubble below the tool block already
shows the metacog text, and the tool block immediately above it
carries the tool name — so a separate gray ``[repeat: list_workstreams()
called with same arguments]`` info line was just duplicate visual
noise (operator-visible in the screenshot below the bubble).
Drop the ``ui.on_info`` call inside ``_apply_post_execute_advisories``
that emitted the diagnostic line. Update the docstring to reflect
that the bubble is the canonical signal. Rename
``test_emit_repeat_ui_line_on_streak_fire`` →
``test_no_legacy_repeat_info_line_on_streak_fire`` and invert the
assertion.
CI typecheck failed because ``WorkstreamTerminalUI(TerminalUI)``
inherits from ``SessionUI`` (the Protocol), and the Protocol's
``on_user_reminder`` / ``on_tool_reminder`` declarations have empty
bodies — mypy treats those as implicitly abstract, so the subclass
became un-instantiable.
Add real implementations on ``TerminalUI`` that render reminders as
``[metacognition · type] text`` lines in yellow. This also restores
the metacog signal on the CLI surface (the legacy
``[metacognition: nudge injected — …]`` info-line went away with
``_emit_nudge_ping``; without this commit the CLI showed no signal
at all for metacog nudges). Tool-channel and user-channel render
identically because terminal output is anchored by stdout flow
rather than by DOM anchor — the line lands directly after the
message it advises.
Address Copilot's review feedback on PR #456 — the docstrings and
inline comments hadn't all caught up with the architectural shift
across the branch:
- ``_apply_reminders_for_provider`` docstring: "every user message"
→ role-agnostic, since tool messages also carry ``_reminders``
(tool_error / repeat).
- ``_mark_reminders_delivered`` docstring: same role-agnostic
update; explicitly note both channels.
- ``_append_user_turn`` callsite comment near
``_attach_pending_user_reminders``: still described splicing
``<system-reminder>`` blocks into user content; updated to
reflect the side-channel attach + transient-copy splice at the
provider boundary.
- ``_build_history`` block comment: was user-message-only; now
mentions tool messages and both ``user_reminder`` /
``tool_reminder`` SSE events.
- ``_build_history`` propagation comment: same role-agnostic note
on the per-entry surface.
- ``app.js`` ``user_reminder`` SSE handler comment: said the
bubble renders "above" the user message, but
``insertAdjacentElement('afterend', el)`` drops it BELOW.
- ``app.js`` ``replayHistory`` comment: said "insertBefore drops
the reminder directly above the just-rendered user bubble";
same fix — bubble lands BELOW.
No behaviour change.
The repeat-detection block in ``_apply_post_execute_advisories`` had
a leftover "clear streak when a write tool succeeded" branch from
when ``RepeatDetector`` tracked cumulative counts. With the
consecutive-streak semantics introduced earlier in the branch the
branch became:
1. Redundant — any different (name, args) signature already resets
the streak via ``RepeatDetector.record``, so an intervening
read/write naturally breaks the streak.
2. Actively wrong — the clear runs ONCE at the top of each
``_apply_post_execute_advisories`` call, before the per-result
loop records sigs. In a single parallel batch
``[bash, bash, bash]`` the clear runs once and then three
``record`` calls accumulate to count=3 in the same call → fires.
But across three sequential turns, each turn calls
``_apply_post_execute_advisories`` fresh, the clear runs at the
top of each call, and only one ``record`` per call follows — so
the count never gets above 1 and the canonical
"small local model stuck on ``bash('echo test')``" pattern
never triggered the nudge.
The asymmetry only existed for successful calls — failures don't
satisfy the ``not _tool_error_flags.get(tc["id"])`` predicate, so
the clear didn't fire and sequential failures already worked. The
fix is to drop the clear entirely; ``RepeatDetector``'s
consecutive-streak semantics handle every case uniformly.
Tests:
- ``test_successful_write_clears_streak`` →
``test_intervening_different_call_resets_streak`` —
rewords the assertion to reflect the actual mechanism (any
different sig resets, write-or-otherwise) since "writes clear"
was the bug, not the contract.
- ``test_failed_write_does_not_clear_streak`` →
``test_sequential_bash_failures_fire_repeat`` — same shape, just
framing fixed.
- New ``test_sequential_bash_same_command_fires_repeat`` —
regression for the bug user hit (three sequential successful
``bash('echo test')`` calls now correctly fire the nudge).
The yellow themed reminder card introduced for user-channel nudges
(correction / denial / resume / start / completion) now also fronts
tool-channel nudges (tool_error / repeat). Pre-fix the tool channel
shipped its reminders inside the tool-result envelope via
``wrap_tool_result``, leaking the ``<system-reminder>`` block into
``self.messages`` content (same problem the user channel had before
the side-channel refactor) and surfacing the legacy gray
``[metacognition: nudge injected — …]`` info line as the only
operator-visible signal — duplicated alongside the new themed bubble
for user-channel nudges.
Tool-channel parity:
- ``_collect_advisories`` now returns
``(persistent_advisories, metacog_reminders)``. Persistent
advisories (``GuardAdvisory`` / ``UserInterjection``) keep
riding ``wrap_tool_result`` because they ARE conversation
history. Metacognitive reminders extract to the second tuple
element; the caller attaches them to the tool message dict's
``_reminders`` side-channel and emits ``on_tool_reminder``.
- ``_apply_reminders_for_provider`` already handles ``_reminders``
on any role, so the tool-channel splice into wire content is
free. ``_build_history`` also already propagates
``entry["reminders"]`` regardless of role, so reload renders the
bubble too.
- ``SessionUI`` Protocol gains ``on_tool_reminder(reminders,
tool_call_id)``; ``SessionUIBase`` enqueues a ``tool_reminder``
SSE event with the ``tool_call_id`` anchor.
- ``_emit_nudge_ping`` had no remaining callers and was removed —
the themed bubble (live SSE + ``/history`` reload) is the
canonical operator signal for both channels now.
UI polish (the four fixes the screenshot caught for the user
channel + their tool-channel mirror):
- Bubble renders BELOW the message it advises (semantically: a
hint to the model right before its turn). ``addUserReminder``
swaps ``insertBefore`` for ``insertAdjacentElement('afterend',
el)``; ``addToolReminder`` anchors below the ``.ts-approval``
block whose tool result triggered the batch's reminder.
- Label uses the full feature name ``metacognition`` (was the
``metacog`` shorthand).
- Card width / alignment inherits from the base ``.msg`` rule —
``align-self: flex-end`` and the explicit ``max-width`` are
gone, so the card matches the user / assistant column instead
of pinning right-aligned narrow.
- The legacy ``[metacognition: nudge injected — …]`` gray info
line is gone for both channels.
Frontend additions:
- ``Pane.prototype.addToolReminder(reminders, toolCallId)``
anchors below the ``.ts-approval`` block (live: by
``data-call-id``; replay: by "last block in messagesEl"
fallback, which is correct because messages render in order).
- SSE switch case ``"tool_reminder"`` calls ``addToolReminder``.
- ``replayHistory``'s tool-message branch now calls
``addToolReminder`` when ``msg.reminders`` is present.
- ``addUserReminder`` advances its anchor on each loop iteration
so multiple reminders stack in queued order rather than
reversed.
Coord console parity:
- ``coordinator.js`` gains ``appendReminderBubble`` /
``appendUserReminderLive`` / ``appendToolReminderLive`` mirroring
the interactive UI. The tool-channel anchor walks
``toolRows[callId].batch`` to attach below the
``.coord-tool-batch`` construct (one bubble per dispatch turn,
matching the "one nudge per batch even with many failing tools"
drain).
- SSE switch handles ``user_reminder`` and ``tool_reminder`` on
the coord conversation surface.
- ``/history`` replay propagates ``msg.reminders`` for user and
tool messages — same wire shape as the interactive pane.
- ``.msg.user-reminder`` styles moved to
``shared_static/chat.css`` so both surfaces inherit the same
yellow themed bubble from the shared base.
Defensive read on ``_apply_reminders_for_provider`` (per Copilot
review on the closed PR): a malformed ``_reminders`` entry (string,
None, etc. — corruption / partial state) used to abort ``send`` via
AttributeError on the ``.get("text", "")`` call. Filter to dicts
before building the block, mirroring the same filter
``_build_history`` already applies on the wire-out side; an
all-malformed list passes through as no-reminders.
Tests:
- ``test_collect_advisories_drains_tool_buffer_on_last_result``
rewritten to assert the ``(persistent, metacog)`` tuple shape
and that ``MetacognitiveAdvisory`` no longer appears in the
persistent list.
- ``test_collect_advisories_holds_*`` and ``_drops_*`` updated for
tuple return.
- ``test_attach_emits_visibility_ping`` /
``test_collect_advisories_emits_visibility_ping`` inverted to
assert the legacy gray line is gone on both channels.
- ``TestSessionUIBaseToolReminderHook`` covers the new SSE event
shape with the ``tool_call_id`` anchor.
- ``test_malformed_reminders_filtered_out`` and
``test_all_malformed_reminders_passes_through`` cover the
Copilot-flagged defensive filter.
User-channel metacognitive nudges (correction, denial, resume, start,
completion) used to be spliced into ``user_msg["content"]`` permanently,
which leaked the ``<system-reminder>`` envelope into every consumer of
``self.messages`` — UI replay (mitigated by a regex strip in /history),
compaction, title generation, and any future channel adapter that
echoes conversation context. The /history strip was a band-aid;
compaction and title-gen still saw the raw spliced text.
Switch to a side-channel: ``_attach_pending_user_reminders`` writes the
rendered reminder list to ``user_msg["_reminders"]`` (sibling key,
leading-underscore convention shared with ``_attachments_meta`` /
``_provider_content``). At the provider boundary, a new
``_apply_reminders_for_provider`` builds a transient shallow-copy with
the reminder spliced into ``content``; the original message dict
stays clean. ``sanitize_messages`` drops the sibling key on the wire.
Once-per-session-not-per-turn semantics for the wire: after stream
success the loop calls ``_mark_reminders_delivered``, which flips a
``_reminders_delivered`` flag on every user message that carried
reminders into that call. ``_apply_reminders_for_provider`` skips
already-delivered messages so the model sees each reminder exactly
once (the turn it advised). ``_build_history`` ignores the delivered
flag entirely, so reconnecting tabs render the same nudge bubble the
originating tab saw via the live ``user_reminder`` SSE event.
UI surface:
- ``SessionUIBase.on_user_reminder`` enqueues a
``{type: "user_reminder", reminders: [...]}`` SSE event with the
same shape ``_build_history`` surfaces.
- ``app.js`` renders a ``.msg.user-reminder`` bubble (yellow accent,
pill-styled) anchored above the user message it advises, both
live and on history replay.
- ``replayHistory`` renders ``addUserMessage`` before
``addUserReminder`` so the anchor lookup finds the just-rendered
turn (not a prior one).
- Multi-tab caveat documented inline: non-originating tabs receive
no ``user_message`` SSE event today, so a reminder may anchor to
a stale prior bubble until ``/history`` reload corrects it.
Pre-existing bug surfaced by the audit: cancel handlers
(``GenerationCancelled`` / ``KeyboardInterrupt`` / generic
``Exception``) in ``ChatSession.send`` cleared
``_pending_tool_advisories`` but not the user-channel buffer. Both
now drain through a shared ``_drain_pending_advisories`` helper.
Removed the ``/history`` regex strip — the side-channel approach
makes it redundant. Hoisted ``escape_wrapper_tags`` +
``render_system_reminder`` imports to module top (called 2-3× per
turn).
Tests:
- ``TestApplyRemindersForProvider`` — pass-through-by-reference,
string + list content splice, escape on user-typed wrapper tags,
multi-reminder ordering, source-untouched invariant, delivered
flag skip path, fallback for unexpected content shape.
- ``TestMarkRemindersDelivered`` — flag idempotency, no-reminders
no-flag, only marks user messages with reminders.
- ``TestUpdateTokenTableMsgsParam`` — calibration uses pre-built
msgs when provided, falls back when not.
- ``TestUserAdvisoryCancelClear`` — all three cancel branches drain
the user buffer.
- ``TestReminderSidechannelIsolation`` — compaction's
``_format_messages_for_summary`` and the title-gen extraction
loop cannot see reminders by construction.
- ``TestSessionUIBaseUserReminderHook`` — ``on_user_reminder``
enqueues the right SSE shape.
- ``TestBuildHistoryReminderPropagation`` — ``entry["reminders"]``
propagation, absent / empty / multi / coexist-with-attachments
cases, malformed input filtering, all-malformed elision.
- ``test_sanitize_messages_strips_underscore_sibling_keys`` covers
``_reminders`` and ``_reminders_delivered``.
Cleanup pass on the metacognitive nudge stack — restores pre-split
errored-counts-toward-repeat behaviour and tightens the is_error
plumbing through the per-batch advisory hook.
The per-batch hook in ``_run_loop`` was duplicating the is_error
signal: ``self._tool_error_flags`` (set by ``_report_tool_result``)
and a string-prefix tuple (``Error`` / ``JSON parse error`` / …).
Two truth sources is what got us here — bash commands that exit
non-zero with normal stdout matched the flag but not the prefix,
the deny path matched the prefix but not the flag, and the result
was that stuck-loop detection silently broke for the most common
failure mode (the model bashing the same broken command).
Single source of truth now:
- ``_execute_tools.run_one`` deny branch routes through
``_report_tool_result(is_error=True)`` so denied calls populate
``_tool_error_flags`` like every other error path.
- The error-prefix tuple is gone; the write-success-clear gate and
the tool-error-nudge gate both read ``_tool_error_flags`` only.
Repeat-detection state moves from a ``set[str]`` (fired on the second
identical call, ignored errors entirely) to a ``RepeatDetector``
helper in ``metacognition.py`` with consecutive-streak semantics:
- Threshold raised from 2 to 3 — two-in-a-row was noisy on
legitimate transient retries; three is the cheapest stuck-loop
signal.
- Recording a different signature resets the count, so [A, A, B, A]
is two short streaks of 2 and not a streak of 4. Bounded by O(1)
state regardless of session length.
- Errored calls now count toward the streak (the split into a
separate metacog module unintentionally introduced a "skip errors"
branch — restored).
While there:
- ``metacognition._COOLDOWN_SECS`` default aligned to 300s (matches
``MemoryConfig.nudge_cooldown`` and the ``memory.nudge_cooldown``
config-store default; was set to 30 by an earlier investigation).
- The per-batch advisory block (~80 lines of mixed orchestration
inside ``_run_loop``) is extracted to
``ChatSession._apply_post_execute_advisories`` so the wired
behaviour is testable without driving ``_run_loop`` end-to-end.
Producer extraction to a dedicated module is deferred to a
follow-up; advisory producers all live on ``ChatSession`` for
now per existing convention.
- Frontend ``appendToolOutput`` (turnstone/ui/static/app.js) now
skips rendering when the parent approval block is denied or
the output starts with ``Denied by user`` / ``Blocked``,
mirroring the history-replay guard at ``_build_history``.
Previously the live SSE path didn't need this guard because
the deny path never emitted a ``tool_result`` event; the
is_error routing change above means it does now, so without
this guard the badge from ``resolveApproval`` and the SSE
output would both render.
Tests: 8 unit tests for ``RepeatDetector`` covering streak,
threshold, clear, and intervening-sig reset; 9 integration tests
for ``_apply_post_execute_advisories`` covering the wired
behaviour (3-identical fires warning + advisory + UI line, errored
calls count toward streak as a regression guard, intervening sig
resets streak, successful write clears, failed write does not,
JSON outputs tracked but not inline-warned, tool_error nudge gates
on memory_count, repeat UI line emitted on streak fire).
The pre-existing comment said pending_approval_detail "rides on
every ws_state event" — that overstated the case. The node-side
emit is gated on ``_pending_approval is not None`` so the field is
absent on the steady-state broadcast and possibly null on a node
mid-rolling-upgrade. The handleChildState fallback already
handles both cases; only the comment was wrong.
Inline child approve/deny in the coord tree UI was rendering downstream
of the bulk-live cache (``GET /v1/api/cluster/ws/live``), not the SSE
stream. ``child_ws_state`` events were tiny notifications that fired
an urgent live-bulk fetch on every activity_state transition into/out
of "approval", just to pick up the rich ``pending_approval_detail``
payload. With multiple coord tabs and multi-child workstreams, that
urgent-fetch pattern compounded the SSE-executor pressure Shape A
is unwinding.
Thread the field through every layer so the SSE event itself carries
the rich payload — browser mutates ``liveBadgeCache`` directly,
no urgent fetch:
1. Node ``WebUI._broadcast_state`` emits ``pending_approval_detail``
on ``ws_state`` events. Gated on ``_pending_approval is not None``
so the per-broadcast verdict-cache deepcopy only runs when there
is actually an approval pending. ``_build_node_snapshot`` also
projects the field so the console's reconnect-via-snapshot
resync path delivers it (without this the new collector
forwarding would never see the field on a snapshot row).
2. Console ``ClusterCollector._apply_delta`` (live ``ws_state``
forwarding) and ``_reconcile_node`` (snapshot resync diff) both
forward the field on the emitted ``cluster_state`` event, AND
``_apply_delta`` persists it on the cached ``ws`` dict so the
``get_node_detail`` / ``get_snapshot`` endpoints between
reconciliations don't render stale approve/deny buttons.
3. ``CoordinatorAdapter._dispatch_child_event`` re-emits the field
on the ``child_ws_state`` event sent to coord listener queues.
4. Frontend ``handleChildState`` reads ``ev.pending_approval_detail``
and writes it directly into ``liveBadgeCache``, tagging the
entry with ``sseUpdatedAt``. ``flushLiveFetches`` honors that
tag for ``SSE_AUTHORITATIVE_MS`` (3s) — the upstream
``/dashboard`` cache has its own ~2s TTL, so a bulk-poll
landing right after a transition can otherwise clobber the
fresh SSE-set state with pre-transition data.
The pre-fix ``enteredApproval`` / ``leftApproval`` urgent-fetch
branch is removed. The 409 stale-call_id retry path keeps its own
urgent fetch — that's a different scenario.
Tests cover the forwarding contract at every layer, the broadcast
gate (event includes the field when an approval is pending,
omits it otherwise, and clears after resolution), and the
``flushLiveFetches`` merge-guard structural shape so a refactor
that keeps the symbols but inverts the comparison or drops the
``prev.live`` check can't pass silently.
``coordinator_children`` was calling ``storage.list_workstreams``
directly on the event loop, ``coordinator_tasks`` did the same with
``load_task_envelope``, and ``_resolve_coordinator_or_404`` (called
from both handlers, plus ``coordinator_history`` and
``_resolve_coord_session``) did the same with
``storage.get_workstream`` on its cold-cache path.
The cold-cache resolver path is hit on every console restart,
coordinator eviction, and console proxy hop — exactly when the
event loop is most contended. Three coord tabs reconnecting after a
brief network blip = three serial event-loop blocks per call site.
Other lifted handlers in this file already use
``asyncio.to_thread``; bring all four call sites onto the same
pattern.
Convert ``_resolve_coordinator_or_404`` to ``async def`` and update
its four call sites to ``await``. Exception flow is unchanged.
Each coord ``events`` SSE listener parks a thread on
``client_queue.get(timeout=5)`` for the connection lifetime. The
console's coord endpoint was wiring no ``sse_executor_lookup`` on
``coord_endpoint_config``, so those parks landed on Python's default
ThreadPoolExecutor (~min(32, cpu_count+4)) and competed with every
other ``asyncio.to_thread`` caller (storage, router, audit). A few
coord tabs against a multi-child workstream would stall new request
handlers waiting for a worker thread.
Mirror the interactive-side precedent (the ``sse_executor`` /
``sse_executor_lookup`` pattern in ``turnstone/server.py``) — build a
dedicated 200-thread ``coord_sse_executor`` in the console lifespan
and wire ``sse_executor_lookup`` onto ``coord_endpoint_config``.
Drain order matters: shut the pool down AFTER ``coord_adapter.shutdown()``
so no new listeners arrive at a dying pool. ``cancel_futures=True``
discards queued-but-not-started futures during teardown.
Update the stale comment on the interactive-side wiring that claimed
"coord wires None and falls back to the default executor" — it now
points at the console's matching wire.
Three follow-ups from Copilot's round-2 review on #453.
ValueError logging surfaced the wrong reason
The catch-all ``except ValueError:`` logged ``reason=no_enabled_rows``
unconditionally, but ``ModelRegistry.__init__`` raises ValueError for
five distinct config issues (empty models, default / fallback / agent /
plan / task alias not present). Operator looking at logs for a
config.toml typo would see the wrong cause. Switch to
``log.warning("...reason=%s", exc)`` so the actual error message
threads through. Behavior unchanged — existing registry still
preserved on every ValueError path.
Misleading shutdown() comment
The ``finally`` comment claimed shutdown() was closing clients the
throwaway registry created during DB load. ``load_model_registry`` only
constructs ModelConfigs and the bare ``ModelRegistry(...)``;
``ModelRegistry.__init__`` leaves ``_clients`` / ``_providers`` empty
and they populate lazily on first resolve. Today shutdown() iterates
empty dicts. Comment now says so explicitly while keeping the call
(and its try/except) for forward-compat against an eager-init future.
Stale "probe" wording in test docstring
``test_helper_preserves_registry_when_db_probe_fails`` →
``test_helper_preserves_registry_when_strict_load_fails``. The
explicit probe was removed in commit 1ba17ed when the helper switched
to ``load_model_registry(..., strict=True)``; the test name and
docstring still talked about a probe. Updated wording reflects that
the loader's strict-mode re-raise is what the helper catches now.
132 tests pass.
Hygiene follow-ups from the multi-stage code review on #453.
perf-1 — sync helper called from async route handlers
``_refresh_coord_registry`` runs two sync DB reads and a registry reload
that takes ``_client_lock``; calling it directly from an async handler
held the event loop for the duration. All four call sites now
``await asyncio.to_thread(_refresh_coord_registry, ...)``, matching the
pattern from commit ``1f7d6ad`` (offloaded ``tenant_check``).
perf-3 — ModelRegistry.reload() tore down all clients unconditionally
The reload always closed every cached client and provider, even when
the changed fields (``model``, ``temperature``, ``context_window``)
didn't touch the connection target. Now selective: clients drop only
when alias removed or ``(base_url, api_key, provider)`` differs;
providers drop only when alias removed or ``provider`` string differs.
Keeps connection pools warm across the common admin-edit case where
only metadata changed. Two new ``test_model_registry`` cases lock the
keep-warm vs drop-on-change behaviour, and the existing
``test_reload_clears_clients`` was updated (it asserted the old
overly-aggressive contract) into
``test_reload_keeps_clients_when_connection_target_unchanged``.
q-5 — helper rename
``_refresh_console_coord_registry`` → ``_refresh_coord_registry``. The
``console_`` prefix was redundant given the function lives in
``turnstone/console/server.py`` and sibling helpers there
(``_notify_nodes_model_reload``, ``_publish_config_change``,
``_collect_model_status``) all omit it.
q-1 — shared test middleware
``tests/test_admin_model_registry_refresh`` now imports the
header-driven ``_AuthMiddleware`` from ``tests/_coord_test_helpers``
and sets default ``X-Test-User`` / ``X-Test-Perms`` headers on the
``TestClient``. The local hardcoded variant duplicated infrastructure
the helper module exists to centralise.
q-3 — multi-alias test registry
``_make_registry`` extracted a ``_make_config`` helper and gained an
``extras={alias: model}`` param so multi-alias scenarios stop
hand-building ``ModelConfig`` literals.
``test_delete_endpoint_refreshes_registry`` now uses the helper.
310 tests pass across the related coordinator + model surfaces.
bug-3 / q-2 from the multi-stage review on #453: the previous test
``test_update_endpoint_with_empty_body_does_not_blow_up`` asserted only
that the registry's model name was unchanged after an empty PUT, which
holds whether or not the refresh ran (DB row matches registry → refresh
is idempotent). A regression that always called
``_refresh_console_coord_registry`` — exactly the gate this test was
meant to lock — would have left the assertion green.
Rename to ``test_update_endpoint_skips_refresh_on_empty_body`` and spy
on the helper via ``monkeypatch.setattr``. Empty-body PUT must register
zero calls; any future change that drops the ``if updates:`` gate now
fails loudly.
Two correctness follow-ups from the multi-stage code review on #453.
bug-2 / perf-2 (DB probe was theatre + double scan)
The previous probe defended nothing the loader didn't already swallow
on the next line: ``load_model_registry``'s row-loop catches Exception
internally, so a transient DB error after the probe still degrades to
a config.toml-only registry that ``existing.reload()`` would apply,
silently dropping every DB-sourced alias. And on the happy path each
CRUD paid for two scans of ``model_definitions``.
Add a ``strict: bool = False`` flag to ``load_model_registry``. When
strict, the row-loop's except re-raises instead of swallowing. The
helper passes ``strict=True`` and drops the probe — single DB scan,
real failure isolation, the loader's silent fallback can no longer
mask a partial-result regression. Default ``strict=False`` so CLI /
lifespan callers keep their boot-with-config-fallback behaviour.
bug-1 (shutdown could escape after a successful reload)
``ModelRegistry.shutdown()`` calls ``client.close()`` unguarded, and the
helper's ``finally`` block ran it outside the try/except. A raising
close() after a successful ``existing.reload()`` would surface as 500
with the registry already mutated and the audit row already recording
success. Wrap ``new_registry.shutdown()`` in its own try/except that
matches the helper's belt-and-suspenders error policy elsewhere.
The helper's docstring also drops the obsolete probe paragraph; the
``if existing is None: return`` branch gets a one-line inline comment
about the boot-from-empty case (the multi-paragraph version restated
behaviour the line itself documents).
129 tests pass (test_admin_model_registry_refresh + test_model_registry).
Two follow-ups from Copilot review of #453:
1. ``load_model_registry`` swallows storage read errors internally
(logs + continues with config.toml-only models). Without a strict
probe in the helper, a transient DB outage on an admin CRUD would
apply a truncated registry that drops every DB-sourced alias —
silently, since the loader returns a non-empty registry built from
``[models.*]`` config.toml entries. Add an explicit
``storage.list_model_definitions(enabled_only=True)`` probe before
the loader call so the failure is visible here and the existing
registry is preserved on outage.
2. The previous docstring claimed ``admin_model_reload`` "has its own
boot-from-empty story." It doesn't — it just calls this helper,
which no-ops when ``coord_registry`` is None. When no model rows
existed at boot, lifespan leaves the entire coord subsystem
uninitialized (no ``coord_mgr``, no ``coord_adapter``, no
``session_factory``), and a console restart remains required after
the operator adds the first row. Tighten the docstring to admit
that limitation rather than overstating the helper's reach.
New test ``test_helper_preserves_registry_when_db_probe_fails``
monkeypatches ``list_model_definitions`` to raise and asserts the
existing registry stays intact.
The console builds ``app.state.coord_registry`` once at lifespan startup
and the coordinator session factory closes over that exact instance.
Until now, the model-definition admin endpoints (create/update/delete)
wrote to the DB but never touched the in-process registry — and the
explicit reload button only fanned out to nodes via HTTP, also leaving
the console's own registry stale.
Symptom: an operator who changed the underlying model name behind a
local-LLM alias (same alias, same endpoint) saw the DB row update
immediately, but coordinator sessions kept calling the prior model
name until the console process was restarted.
Fix: a new helper ``_refresh_console_coord_registry`` rebuilds a fresh
ModelRegistry from DB and applies it to ``app.state.coord_registry``
via the existing thread-safe ``ModelRegistry.reload()`` — in-place
mutation preserves object identity so the factory closure keeps
working, and active coord sessions auto-pick up the swap on their
next ``send()`` via ``ChatSession._refresh_model_from_registry``.
Wired into four endpoints in ``console/server.py``:
- ``admin_create_model_definition`` — after the DB write
- ``admin_update_model_definition`` — after the DB write, gated on
``if updates:`` so a no-op PUT skips the rebuild
- ``admin_delete_model_definition`` — after the DB write
- ``admin_model_reload`` — between ``_publish_config_change`` and
``_notify_nodes_model_reload`` so the console mirrors what the
reload broadcasts to nodes
Failure isolation: a load or reload error leaves the existing registry
intact (logged + swallowed). Coord stays usable while the operator
investigates; the explicit reload remains the user-facing recovery path.
No node fan-out on CRUD — the explicit reload button continues to gate
cluster-wide HTTP propagation, preserving today's UX semantics on shared
clusters.
Tests in ``tests/test_admin_model_registry_refresh.py`` cover:
- helper-level: rebuild from DB, identity preservation, no-op when
registry is None, preservation on load failure / no-enabled-rows /
reload validation error
- endpoint-level: create / update / delete / explicit-reload all
refresh the registry; an empty PUT skips the rebuild
Production fan-outs are frequently hitting the 6 KiB per-child cap by
just 1-2 KiB, forcing the coordinator into a follow-up inspect_workstream
round-trip per truncated child to recover the tail. Bumping the cap to
10 KiB absorbs the common overshoot without changing the truncation
semantics — truncated=True still fires for genuinely oversized messages,
and inspect_workstream remains the unbounded follow-up.
Worst-case context impact: a 32-child fan-out at the cap is now ~320 KiB
(was ~192 KiB), still well within commercial model context windows.
Typical fan-outs of 1-5 children land at 10-50 KiB.
LAST_ERROR_MAX_LEN (1 KiB) is unchanged — it's intentionally smaller
than the wait cap so error truncation happens at write time, and
1 KiB still sits well below 10 KiB.
WAIT_MESSAGE_MAX_BYTES is referenced by name (not literal 6144) in the
truncation test, so no test value needs updating.
The coordinator system message was descriptive about parallelism rather
than prescriptive — "while multiple children run in parallel" framed
fan-out as incidental, and "a tasks entry, a child to own it" primed
singular delegation. The spawn_batch example (benchmark A, benchmark B,
prototype the winner) showed dependent work under a fan-out framing,
teaching the wrong shape.
In practice the coordinator failed to decompose enumerable requests
("top stories on HN, Lobsters, /r/programming, …") without explicit
"please fan this out" instructions, on both GPT-5.5 and Claude Opus.
base_coordinator.md
- Replace singular "a tasks entry, a child to own it" with plural
"enumerate the independent units of work, spawn one child per unit,
run them in parallel by default. Sequential only when one child's
output feeds the next."
- Tighten the delegation paragraph.
tools_coordinator.md
- Drop the persona repetition that duplicated base_coordinator.md.
- Drop the prescriptive "## Workflow shape" section (the cost note is
already in wait_for_workstream's tool description; the edit-X
redirect is already in the persona).
- Drop "in one approval" / "single approval" mentions to avoid
surfacing approval mechanics to the model.
- Replace the misleading spawn_batch example with truly independent
items; drop "(up to 10)" which overstated the cap (it's per-call,
not global, and is documented in the tool schema).
- Add a course-correction example to send_to_workstream — the pattern
coordinators most often replace with cancel-and-respawn.
- Drop the read action from the tasks examples to keep the lifecycle
(add → update → remove) coherent.
Coord system message ~16% shorter (4440 → 3722 chars). Both GPT-5.5
and Claude Opus now naturally decompose the news-board prompt without
explicit fan-out instructions. 29 prompt-composition tests pass.
The server's --help epilog and compose.yaml both reference
--skip-permissions, but the argparser never defined it, so any
container started with SKIP_PERMISSIONS=1 exited with
"unrecognized arguments: --skip-permissions".
Wire the flag through to app.state.skip_permissions, OR-ing it
with the existing tools.skip_permissions config-store setting so
the stored value still works on its own.
2026-04-29 20:20:38 -07:00
497 changed files with 38781 additions and 94542 deletions
Self-hosted, local-first orchestration for tool-using AI agents. Give LLMs real tools — shell, files, search, web — and run them across your own cluster with direct HTTP routing and interactive interfaces. Your code, your models, your data stay on hardware you control: no telemetry, no phone-home.
Multi-node AI orchestration platform. Deploy tool-using AI agents across a cluster of servers with direct HTTP routing, interactive interfaces, and enterprise governance.
<p align="center">
<img src="docs/assets/hero.png" alt="Turnstone coordinator — parallel tool batches with judge-graded approval and child workstream tracking" width="960"/>
@@ -27,13 +26,12 @@ See [docs/releasing.md](docs/releasing.md) for the full release process.
Turnstone gives LLMs tools — shell, files, search, web, planning — and orchestrates multi-turn conversations where the model investigates, acts, and reports.
- **Local-first & private** — runs entirely on hardware you control, with no telemetry and no phone-home. Point it at local models (vLLM, llama.cpp, Ollama) or commercial APIs you hold the keys to — your prompts and data never transit a third party you didn't choose.
- **Bring your own models** — OpenAI-compatible APIs (vLLM, llama.cpp, NIM), the Anthropic Messages API, and Google Gemini, mixed freely per role
- **Interactive sessions** — terminal CLI or browser UI with parallel workstreams
- **Cluster dashboard** — real-time view of every node and workstream, with a rendezvous routing proxy
- **Intent validation** — an LLM judge (your model) grades every tool call with a risk assessment and evidence before it runs
- **Cluster dashboard** — real-time view of all nodes and workstreams with console routing proxy
- **Intent validation** — LLM judge evaluates every tool call with risk assessments and evidence
@@ -281,7 +281,6 @@ Each message in the `messages` array has:
| `role` | string | `"user"`, `"assistant"`, or `"tool"` |
| `content` | string or null | Text content of the message |
| `tool_calls` | array or null | Present only on assistant messages with calls |
| `reasoning` | string (optional) | Concatenated reasoning / chain-of-thought text on assistant turns whose `provider_data` carried reasoning-bearing blocks (Anthropic `thinking`, OpenAI Responses `reasoning`, or synthetic `reasoning_text` from local-model servers). Present only when the active model's `surface_persisted_reasoning` flag is True. |
Each entry in `tool_calls`:
@@ -326,44 +325,6 @@ finalize any in-progress assistant message.
{"type":"stream_end"}
```
**`state_change`** -- the worker thread transitioned to a new state. Drives
the client's busy-mode (composer in send vs. stop, spinner indicators,
auto-focus on idle). Sent live during normal operation AND on every fresh
SSE subscribe (so a mid-stream page refresh restores the correct composer
state without waiting for the next live transition).
*.json 16 tool schemas (OpenAI function-calling format + turnstone metadata)
*.json 19 tool schemas (OpenAI function-calling format + turnstone metadata)
```
Both UIs share a common design system extracted into `turnstone/shared_static/`: design tokens, login overlay, toast notifications, theme toggle, keyboard shortcuts, and utility functions. Each UI imports `base.css` and the shared JS modules at `/shared/`, then adds only page-specific code at `/static/`.
@@ -189,6 +190,7 @@ Phase 3: EXECUTE (parallel)
(cancel_event also checked per line — kills process group on cancel)
Final output (stdout + stderr) delivered via ui.on_tool_result(call_id, name, output)
call_id links tool_info items → streaming chunks → final result
For plan tool: post-execution gate via ui.on_plan_review()
```
### State Transitions
@@ -207,7 +209,7 @@ The engine emits state changes via `_emit_state()` which calls
"running" ---> tool execution
|
v
"attention" ---> waiting for user approval
"attention" ---> waiting for user approval / plan review
|
v
"running" ---> executing approved tools
@@ -229,13 +231,11 @@ The engine emits state changes via `_emit_state()` which calls
> See also: [Core Engine Classes diagram](diagrams/png/03-core-engine-classes.png)
Defined in `turnstone.core.session.SessionUI` as a `typing.Protocol` with 15
Defined in `turnstone.core.session.SessionUI` as a `typing.Protocol` with 14
methods. Every frontend must implement all of them.
`on_rename` is called by the `/name` command (on success) and after a successful `/resume` (if the resumed session has an alias or title). `WebUI.on_rename` broadcasts a `ws_rename` event on the global SSE channel and updates the in-memory `Workstream.name`; `TerminalUI.on_rename` is a no-op.
### Three Implementations
@@ -266,7 +259,7 @@ the per-workstream events stream in
| `WebUI` | `turnstone.server` | SSE event queue per workstream + global broadcast, `threading.Event` for blocking on approval. `on_state_change` sends to both per-workstream and global SSE (the browser UI uses per-workstream `state_change` events to manage busy/idle transitions; `stream_end` only finalizes markdown rendering). |
| `WebUI` | `turnstone.server` | SSE event queue per workstream + global broadcast, `threading.Event` for blocking on approval/plan. `on_state_change` sends to both per-workstream and global SSE (the browser UI uses per-workstream `state_change` events to manage busy/idle transitions; `stream_end` only finalizes markdown rendering). |
| `approve_request` | One or more tool calls need operator approval | `items: [{call_id, header, preview, func_name, approval_label, needs_approval}]` |
| `state_change` | Worker-thread state transition (also re-emitted with the current state on every fresh subscribe so refresh-mid-stream restores composer mode) | `state` ∈ `running`, `thinking`, `attention`, `idle`, `error` |
| `in_progress_snapshot` | One-shot replay of the in-progress turn's content + reasoning when this client connects mid-stream | `content`, `reasoning` |
| `tasks` | plan | Orchestrator-only scratchpad. Children don't see it. |
Explicitly **not** in the coordinator set:
@@ -86,8 +72,8 @@ Explicitly **not** in the coordinator set:
- `bash` / `edit_file` / `write_file` / `append_file` / `diff_file` — no local FS.
- `read_file` / `search` — no local FS reads.
- `web_fetch` / `web_search` — no direct web access.
- `task_agent` — sub-agent tool is zeroed on coord sessions.
- `recall` / `watch` / `read_resource` / `use_prompt`— UX / persistence tools that belong to interactive sessions. The dual-kind `memory` / `skills` / `notify` tools are available on both kinds (see the table above).
- `task_agent` / `plan_agent` — sub-agent tools are zeroed on coord sessions.
- `memory` / `recall` / `notify` / `watch` / `read_resource` / `use_prompt`/ `skill` — the orchestrator's "memory" is its children's outputs; these UX / persistence tools belong to interactive sessions.
If your skill needs a coordinator to "run a command" or "read a
file", write the delegate pattern instead: spawn a child with an
@@ -168,50 +154,29 @@ Every ws_id returned by `spawn_workstream` / `spawn_batch` is a
invent ws_ids — a model that hallucinates `"child-1"` or `"ws-abc"`
hits the tenant guard in `CoordinatorClient._is_own_subtree`, which
validates ws_id against `parent_ws_id=coord_ws_id` AND
`user_id=owner` in storage. The rejection shape is uniform and
recovery-oriented:
`user_id=owner` in storage. The rejection shape varies by tool:
**Default** (no flag) — starts `server` and `console`. Requires an OpenAI-compatible LLM API running on the host (default: `http://localhost:8000/v1`).
The dashboard is at **https://localhost:8443** (Caddy, same as the dev stack);
the console's HTTP port isn't published. For a real domain and a publicly
trusted cert, edit `turnstone/deploy/Caddyfile` to point Caddy at Let's Encrypt
(see [tls.md](tls.md)). Pin the image with `TURNSTONE_IMAGE_TAG` (default:
`latest`).
### mTLS
Layer the TLS overlay on the production stack to enable mutual TLS between
services. A bootstrap container creates a CA and every service auto-provisions
certs via the console's ACME endpoint:
```bash
docker compose -f turnstone/deploy/compose.yaml -f deploy/docker-compose.tls.yml up
```
See [tls.md](tls.md) for details.
## Configuration
Everything is configured with environment variables in `.env` (copy from
[`.env.example`](../.env.example)). The dev stack needs none of them — they're
overrides.
All configuration is via environment variables in `.env` (copy from`.env.example`):
### LLM backend
### LLM Backend
| Variable | Default | Description |
|----------|---------|-------------|
| `LLM_BASE_URL` | `http://host.docker.internal:8000/v1` | Bootstrap OpenAI-compatible API URL (real backends go in the UI) |
| `LLM_BASE_URL` | `http://host.docker.internal:8000/v1` | OpenAI-compatible API URL |
| `OPENAI_API_KEY` | `dummy` | API key (`dummy` for local servers) |
| `TURNSTONE_SEARXNG_URL` | `http://searxng:8080` | SearxNG URL for the `web_search` tool (local/vLLM models only; Anthropic/OpenAI use native search). Defaults to the bundled `searxng` service; set to an external instance's URL. To turn web search off, clear `tools.searxng_url` in the admin Settings tab. |
| `SEARXNG_IMAGE_TAG` | `latest` | Tag for the bundled `searxng/searxng` image |
| `MODEL` | — | Override the default model alias |
| `TAVILY_API_KEY` | — | Web search API key (only needed for local/vLLM models; Anthropic and OpenAI search models use native search) |
### Auth & database
| Variable | Default (dev / prod) | Description |
|----------|----------------------|-------------|
| `TURNSTONE_JWT_SECRET` | insecure default / **required** | JWT signing secret. Every service must share one value. |
| `TURNSTONE_DB_URL` | — | Database URL (e.g. `postgresql+psycopg://user:pass@postgres:5432/turnstone`). For SQLite, defaults to `/data/.turnstone.db` |
| `TURNSTONE_DB_POOL_SIZE` | `2` | PostgreSQL connection pool size per process (default: 2 base + 3 overflow = 5 max) |
| `POSTGRES_USER` | `turnstone` | PostgreSQL container username (used in default `TURNSTONE_DB_URL` for cluster/channel) |
| `POSTGRES_PASSWORD` | — | PostgreSQL container password (required for production and cluster profiles) |
The database stores workstream history, user accounts, and API tokens. When using JWT auth, a database backend is required for user storage.
> **Upgrading from <1.3.0a4:** Earlier versions used `DB_BACKEND` and `DATABASE_URL` in `.env`, which `compose.yaml` mapped to the `TURNSTONE_`-prefixed names internally. These short aliases have been removed. Rename `DB_BACKEND` → `TURNSTONE_DB_BACKEND` and `DATABASE_URL` → `TURNSTONE_DB_URL` in your `.env` file.
> **Large clusters:** Each turnstone process maintains a small connection pool (5 max). At hundreds of nodes this adds up — use [PgBouncer](pgbouncer.md) in transaction pooling mode between turnstone and PostgreSQL.
> **First-time setup:** After deploying with auth enabled, create an initial admin user by running `turnstone-admin create-user` inside the container:
> You will be prompted to set a password. Use it to log in via the UI or SDK, then create additional users through the admin API. Pass `--token --scopes read,write,approve` to also generate an initial API token.
| `TURNSTONE_SLACK_SLASH_COMMAND` | `/turnstone` | Slash command registered in the Slack app |
The channel service runs in the `production` profile. When
`TURNSTONE_DISCORD_TOKEN` or the Slack pair is set the gateway starts the
corresponding adapter; both can run in one process. See
[Channel Integrations](channels.md) for platform app setup and user
account linking.
## Scaling
For multi-node testing, use the `cluster` profile which provides 10 server instances with unique node IDs (`node-1` through `node-10`), resource limits, and shared PostgreSQL:
```bash
docker compose build # build the dev image
docker compose build --no-cache # rebuild from scratch
POSTGRES_PASSWORD=secret docker compose --profile cluster up
```
The default `server` also runs alongside the cluster nodes (11 total). All nodes are accessible via the console dashboard at `:8090`.
For production clusters beyond ~50 nodes, add PgBouncer between turnstone services and PostgreSQL. See [PgBouncer Connection Pooling](pgbouncer.md) for Docker Compose and Helm configuration.
## Volumes
| Volume | Purpose |
|--------|---------|
| `postgres-data` | PostgreSQL data directory |
| `turnstone-data` | `/data` per node (SQLite fallback, local state) |
| `workspace` | `/workspace` (unless `WORKSPACE_MOUNT` is set) |
| `caddy-data` / `caddy-config` | Caddy's local CA and config (dev stack) |
# MCP OAuth — per-user authorization for MCP servers
Turnstone supports **per-(user, MCP server) OAuth 2.1 + PKCE** delegation so each Turnstone user authorizes a remote MCP server with their own identity, rather than sharing a single bearer token across the deployment. This is the right shape for MCP servers that expose user-specific data (a personal CRM, an email inbox, a calendar) and for MCP servers that want per-user audit attribution.
Per-user OAuth is opt-in per `mcp_servers` row. Local-auth Turnstone installs with no `oauth_user` rows exercise zero new code paths — the entire feature is dark by default.
> **Note**: This is a separate authorization layer from Turnstone's own user authentication. A user who logs into Turnstone with a local username + password can still authorize a per-server OAuth MCP server. OIDC SSO and per-server OAuth are orthogonal.
---
## When to use which `auth_type`
The MCP server admin form exposes three authorization modes ("Multitenant Authorization"):
| `auth_type` | What it means | When to use |
|---|---|---|
| `none` | No headers attached. Open MCP server (or one gated by network policy only). | Internal MCP servers on a trusted network. |
| `static` | One static bearer token, configured per server, sent on every request from every user. | Service-to-service MCP servers where per-user attribution doesn't matter, or single-tenant deployments. |
| `oauth_user`*(recommended for user-data servers)* | Each user authorizes separately via OAuth 2.1 + PKCE; Turnstone stores per-user tokens encrypted at rest. | MCP servers that expose user-specific data or that want per-user audit attribution. |
Switching `auth_type` away from `oauth_user` orphans existing per-user tokens. Use the admin **bulk-revoke** affordance on the server row (Phase 9) to clear them, or let them expire naturally — they're inert without the matching `auth_type` value.
---
## Prerequisites for `auth_type=oauth_user`
1. **Encryption key**. Tokens are stored encrypted with Fernet. Set `[security] mcp_token_encryption_key` in `config.toml` (Turnstone won't start with an `oauth_user` row configured but no key installed). Rotate via `MultiFernet` — add the new key first, then later remove the old one once all rows have been re-encrypted.
2. **MCP server publishes RFC 9728 PRM and RFC 8414 AS metadata***or* you configure the AS URL override on the server row. PKCE S256 is mandatory; Turnstone refuses to connect to authorization servers that don't advertise `code_challenge_methods_supported: ["S256"]`.
3. **OAuth client registration**. Two paths:
- **Pre-registered** (most common): you create an OAuth client at the authorization server (manually, via admin console, or via Terraform), then paste the `client_id` / `client_secret` into the Turnstone admin form.
- **Dynamic client registration** (RFC 7591): if the AS supports it and you select that mode in the admin form, Turnstone registers a client at first use and persists the `client_id` automatically.
4. **Redirect URI** registered at the authorization server: `https://your-turnstone-host/v1/api/mcp/oauth/callback`.
---
## Configuration
### Per-server fields (admin UI)
| Field | Required | Description |
|---|---|---|
| Server URL | Yes | The MCP server's `streamable-http` base URL. |
| Authorization Server URL | No | Override for RFC 9728 PRM discovery. Set when your AS endpoint differs from the MCP server URL (e.g., corporate AS protecting a third-party MCP). When unset, Turnstone falls back to PRM discovery against the MCP server itself. |
| Client Secret | Optional (write-only) | OAuth 2.0 client secret (confidential client). Encrypted at rest. Written but never re-read by the API; field stays masked. |
| Scopes | No | Space-separated default scope set requested at the authorize endpoint. Per-tool step-up may union additional scopes from a server's `insufficient_scope` response. |
| Audience | No | RFC 8707 `resource=` parameter sent on every authorize and token request. Defaults to the MCP server URL when unset. Validate against the `aud` claim in returned JWT tokens. |
### Encryption key
```toml
[security]
mcp_token_encryption_key = "base64-fernet-key"
# For rotation, list the keys in priority order — first is used for new
Keep this in `config.toml` rather than environment variables. An in-process LLM with shell-tool access can read the server's environment via `env` / `os.environ` and exfiltrate any secret stored there; secrets in `config.toml` are only loaded into the server at startup and never re-read on a tool-driven path, so a prompt-injection attack against the agent cannot reach them.
---
## Lifecycle
1. **First tool call** for a user against an `oauth_user` MCP server: pool dispatch finds no stored token, returns `mcp_consent_required` to the agent. Dashboard renders an inline "Connect" action card.
2. **User clicks Connect**: opens `/v1/api/mcp/oauth/start?server=<name>` in a popup. Browser redirects through the AS authorize endpoint, user grants consent, AS redirects back to `/v1/api/mcp/oauth/callback`. Turnstone exchanges code → tokens via PKCE, validates audience, encrypts, persists in `mcp_user_tokens`, redirects user back to the originating URL.
3. **Subsequent tool calls** by the same user against the same server reuse the persisted token via the per-(user, server) session pool. Tokens auto-refresh via the refresh-token grant when expired; failed refresh emits `mcp_consent_required` to drive re-consent.
4. **Step-up scope**: when a tool call hits `403` with `WWW-Authenticate: error="insufficient_scope"`, Turnstone emits `mcp_insufficient_scope` with the parsed scope set; the dashboard offers a "Connect with additional scopes" affordance that opens `/v1/api/mcp/oauth/start?server=<name>&scopes=<extra>` so the union of original + new scopes flows into the AS authorize request.
5. **User revoke** (settings modal): `DELETE /v1/api/mcp/oauth/connections/{server_name}` runs the authoritative local delete + best-effort RFC 7009 upstream revoke (fire-and-forget, capped at 256 concurrent in-flight tasks).
6. **Admin bulk-revoke** (Phase 9): `POST /v1/api/admin/mcp-servers/{name}/bulk-revoke` drops every user's token for the server. Upstream RFC 7009 revoke is intentionally **not** attempted in bulk (avoids N upstream HTTP calls per admin click); tokens at the AS expire naturally. Use the per-user revoke endpoint if you need guaranteed upstream invalidation.
---
## Admin status indicators
The MCP Servers admin tab shows per-server status pills (Phase 9):
- **Consented users count** — distinct users with a non-expired token for this server. Surfaced as a `bulk-revoke (N)` button when ≥1; clicking it opens a confirmation dialog. Hidden when 0.
- **Last refresh** — timestamp + outcome (`ok` / `error:ClassName`) of the most recent manual or auto-reconnect refresh. Per node. Absent until at least one refresh has occurred (renders as "never" in the admin UI).
Additional indicators (circuit-breaker state, encryption-key mismatch) are exposed via `get_server_status` on the API but do not yet have a dedicated admin pill — operators see them today via the per-server status text + error tooltip and in audit logs. A future phase may surface these as discrete pills.
---
## Auth-type transitions
| From | To | What happens |
|---|---|---|
| `none` / `static` → `oauth_user` | — | New code path activates for this server. Existing static headers (if any) are no longer sent. Users must authorize on first use. |
| `oauth_user` → `none` / `static` | — | Existing `mcp_user_tokens` rows are **orphaned** — inert without a matching `auth_type`. Use admin bulk-revoke to drop them, or let them expire. Switching back to `oauth_user` later re-activates the orphaned rows if they haven't been deleted. |
| OAuth `client_id` or `client_secret` rotated | — | Existing tokens may stop refreshing if the AS treats them as bound to the previous client. Bulk-revoke after rotation. |
The orphan-by-default behavior is chosen so switching back to `oauth_user` is non-destructive. Bulk-revoke is the explicit cleanup path.
---
## Troubleshooting
| Symptom | Likely cause | Action |
|---|---|---|
| `mcp_consent_required` even after consenting | Token persistence failed, or refresh-token rejected by AS | Check audit log for `mcp_server.oauth.persist_failed` or `mcp_server.oauth.token_revoked`. Re-consent via settings modal. |
| `mcp_token_undecryptable_key_unknown` | Encryption key rotated without keeping the previous key in the keyring | Add the previous key back to `mcp_token_encryption_keys` until all rows have been re-encrypted, then drop. |
| `mcp_oauth_url_insecure` | MCP server URL is `http://` (not `https://`) on a non-loopback host | Use `https://`. Per-user bearers must not transit cleartext. |
| Tools fail in scheduled / Discord / Slack runs | OAuth-MCP requires browser-based consent | Users must pre-consent via the web UI. Phase 9 dashboard badge surfaces deferred consents from these runs on next login. |
| Circuit breaker open repeatedly | Transport-level errors on the MCP server (DNS, TLS, 5xx) | Check the per-server error pill; auth errors do not trip the breaker. |
See also: `docs/operations/mcp-oauth-headless.md` for the cron / channel-driven run caveat.
# MCP OAuth in headless / scheduled / channel-driven runs
**Constraint**: OAuth-MCP servers (`auth_type=oauth_user`) require browser-based user consent. Users must pre-consent via the web UI before any run that cannot drive a browser redirect.
- Any future channel adapter without an interactive browser session.
**What happens when consent is missing**:
A tool call against an `oauth_user` server returns a structured `mcp_consent_required` error to the agent. The agent surfaces the deferred work in its output. Turnstone persists a record to `mcp_pending_consent` so the dashboard badge surfaces the deferred consent need to the user on next login.
**Recovery**:
The user opens the dashboard, sees the gear-icon badge counting pending consents, opens the settings modal, clicks Connect for each affected server, and completes the OAuth dance. The pending-consent record is cleared by the OAuth callback handler on success. Subsequent scheduled / channel runs use the freshly-stored token.
**Pre-consent recipe**:
Before scheduling a workstream that depends on an `oauth_user` MCP server, the user should:
1. Open the dashboard.
2. Open the settings modal (gear icon).
3. Click Connect on each MCP server the schedule will use.
4. Confirm consent in the popup.
This stores tokens that the scheduled run will reuse. Refresh-token rotation is handled transparently on the run side; only the first consent requires browser interaction.
@@ -199,28 +199,4 @@ does not support prepared statements. Turnstone's SQLAlchemy layer does
not use server-side prepared statements by default, so this is not an
issue.
**LISTEN / NOTIFY not supported in transaction mode** — PgBouncer's
transaction pooling assigns a real server connection only for the
duration of each transaction, then returns it to the pool. PostgreSQL
`LISTEN` is session state — a transaction-pooled client can't hold the
multi-statement session a long-lived `LISTEN` needs. The console's
`NotifyDispatcher` (reactive node discovery via the `services` channel)
therefore opens a **dedicated, direct-to-Postgres** connection that
bypasses PgBouncer.
Configure via `config.toml``[database] listen_url` (preferred —
co-located with the main `url`) or the `TURNSTONE_DB_LISTEN_URL` env var
(config.toml wins when both are set). Defaults to the main DB URL when
unset.
| Setting | Behaviour |
|---|---|
| unset | Listener uses `TURNSTONE_DB_URL` as-is. Fine when PgBouncer is in **session** mode, or when there's no pooler in front of Postgres. With transaction-mode PgBouncer the listener's `LISTEN` will fail and the dispatcher retries with exponential backoff (1 s → 30 s cap) without ever succeeding. Reactive NOTIFY-driven node discovery is silently lost; the cluster collector's 60 s `_discovery_loop` is the only remaining backstop. |
| set to direct-to-PG URL (e.g. `postgresql://…/turnstone`) | Listener bypasses PgBouncer for its one dedicated connection. Reactive discovery latency drops from up-to-60 s to ~500 ms. The rest of the storage layer continues to go through PgBouncer in transaction mode. |
Set this whenever PgBouncer is in transaction mode (the recommended
setting per this doc). The override only adds one long-lived PG
connection per console process — sized into the cluster's
`max_connections` budget alongside the pool.
See also: [Docker deployment](docker.md) · [Security](security.md)
- **Stable** tracks receive bugfixes only. The most-recent stable minor
owns the `:stable` / `:latest` Docker tags and the default PyPI
@@ -16,10 +17,8 @@ Turnstone ships several parallel release tracks from a single PyPI package.
- **Experimental** (always on `main`) receives new features. May be
rough around the edges.
- When experimental matures, it is promoted to a new stable minor via
a `stable/X.Y` branch. One prior stable track is maintained alongside
the current one; at each promotion the oldest track is retired — its
branch is deleted, while its tags and released artifacts remain
available.
a `stable/X.Y` branch; older stable branches continue to receive
security fixes until explicitly retired.
## Version Scheme
@@ -34,17 +33,17 @@ Turnstone ships several parallel release tracks from a single PyPI package.
## Releasing an Experimental Version (from main)
```bash
scripts/release.sh 1.7.0a2 --push
scripts/release.sh 1.5.0a2 --push
```
This bumps `pyproject.toml` + `turnstone/__init__.py`, regenerates `uv.lock`, commits, tags `v1.7.0a2`, and pushes. CI runs, then publish + Docker workflows fire automatically.
This bumps `pyproject.toml` + `turnstone/__init__.py`, regenerates `uv.lock`, commits, tags `v1.5.0a2`, and pushes. CI runs, then publish + Docker workflows fire automatically.
## Releasing a Stable Patch (from stable/X.Y)
```bash
git checkout stable/1.6
git checkout stable/1.4
git cherry-pick <commit-hash> # bugfix from main
scripts/release.sh 1.6.1 --push
scripts/release.sh 1.4.1 --push
```
## Promoting Experimental to Stable
@@ -53,19 +52,19 @@ When `main` is ready for a stable release:
```bash
# 1. Tag the stable release on main
scripts/release.sh 1.6.0 --push
scripts/release.sh 1.5.0 --push
# 2. Create the stable maintenance branch from that tag
git branch stable/1.6 v1.6.0
git push origin stable/1.6
git branch stable/1.5 v1.5.0
git push origin stable/1.5
# 3. Start the next experimental cycle on main
scripts/release.sh 1.7.0a1 --push
scripts/release.sh 1.6.0a1 --push
```
The previous stable branch continues to receive security-only patches;
the track before it is retired at each promotion (at 1.6.0:
`stable/1.5` stays maintained, `stable/1.4` is retired).
The previous stable branch (`stable/1.4`) continues to receive
security-only patches; older tracks (`stable/1.0`, `stable/1.3`) are
@@ -59,32 +59,20 @@ from ConfigStore. Model names and context windows are now configured per-model
in the Models tab. A startup warning is logged if these keys appear in
`config.toml`.
### Reasoning persistence (per-model)
### Plan / task agent overrides
Two boolean flags on `model_definitions` (migration 052) control how
reasoning text round-trips per model:
| Flag | Default | Effect |
|------|---------|--------|
| `surface_persisted_reasoning` | `True` | Surface stored reasoning text on `/history` payloads so a page reload re-renders the reasoning bubble. **Storage of reasoning bytes is independent of this flag** — they ride in `provider_data` regardless. |
| `replay_reasoning_to_model` | `False` | Send stored reasoning blocks back to the provider on subsequent turns. Capability-gated: only takes effect when the model's `ModelCapabilities.supports_reasoning_replay` is also `True`. Set on canonical OpenAI gpt-5*/o-series and Anthropic Claude entries; unknown / local-server models default to `False` so an operator who flips the flag on a model whose API doesn't understand reasoning replay silently no-ops rather than 400-ing. |
Edit both via the admin Models tab. See the architecture doc for the
`reasoning_text` for Chat Completions / vLLM / llama.cpp / Gemini-compat).
### Task agent overrides
`task_agent` sub-sessions resolve independently from the conversation model
so operators can pick a cheaper/faster model for autonomous loops:
`plan_agent` and `task_agent` sub-sessions resolve independently from the
conversation model so operators can pick a cheaper/faster model for
autonomous loops:
| Setting | Purpose |
|---------|---------|
| `model.task_alias` | Alias used for `task_agent` sub-sessions. Falls back to `[model].agent_model` in config.toml, then the session's active model. |
| `model.plan_alias` | Alias used for `plan_agent` sub-sessions. Falls back to `[model].plan_model` in config.toml, then `[model].agent_model`, then the session's active model. |
| `model.task_alias` | Alias used for `task_agent` sub-sessions. Same fallback chain as `plan_alias`. |
turnstone exposes 16 built-in tools plus any number of external MCP tools to the
turnstone exposes 19 built-in tools plus any number of external MCP tools to the
LLM via the OpenAI function-calling interface. Built-in tools are defined as JSON
files under `turnstone/tools/` and loaded at startup by `turnstone/core/tools.py`.
MCP tools are discovered from configured MCP servers at startup by
@@ -22,6 +22,7 @@ schema plus turnstone-specific metadata keys:
"properties": { ... },
"required": ["param1"]
},
"agent": true,
"task_agent": true,
"auto_approve": true,
"primary_key": "param1"
@@ -32,7 +33,8 @@ schema plus turnstone-specific metadata keys:
| Key | Type | Meaning |
|----------------|------|---------|
| `task_agent` | bool | Tool is available to task sub-agents. |
| `agent` | bool | Tool is available to plan/task sub-agents (read-only subset). |
| `task_agent` | bool | Tool is available to task sub-agents (broader subset). |
| `auto_approve` | bool | Tool runs without user confirmation (read-only, safe operations). |
| `primary_key` | str | When the model sends a bare string instead of JSON args, map it to this parameter name. |
@@ -44,10 +46,12 @@ schema plus turnstone-specific metadata keys:
| Name | Description |
|---------------------|-------------|
| `TOOLS` | All 28 loaded built-in tool definitions (interactive + coordinator union). Sessions send a kind-specific subset (`INTERACTIVE_TOOLS` or `COORDINATOR_TOOLS`). |
| `TOOLS` | All 19 tool definitions (sent to the model). |
| `AGENT_TOOLS` | Tools with `agent: true` -- available to plan sub-agents. Read-only tools. |
| `TASK_AGENT_TOOLS` | Tools with `task_agent: true` -- available to task sub-agents. Includes write operations. |
| `TASK_AUTO_TOOLS`| Set of all tool names with `auto_approve: true` -- used by task-agent sub-sessions to skip confirmation for matching available tools. |
| `BUILTIN_TOOL_NAMES`| Frozenset of all 28 built-in tool names (interactive + coordinator union). Used by tool search to distinguish always-on tools from deferrable MCP tools. |
| `AGENT_AUTO_TOOLS` | Set of tool names with `auto_approve: true` -- no user confirmation needed. |
| `TASK_AUTO_TOOLS` | Same as `AGENT_AUTO_TOOLS` (identical filter). |
| `BUILTIN_TOOL_NAMES`| Frozenset of all 19 built-in tool names. Used by tool search to distinguish always-on tools from deferrable MCP tools. |
| `PRIMARY_KEY_MAP` | Dict mapping tool name to its `primary_key` parameter name. |
- `notify` -- sends notifications to linked channels (time-sensitive, auto-approved for urgency)
@@ -122,14 +130,16 @@ Each item's `execute` callable is invoked:
- `bash` -- arbitrary command execution
- `write_file` -- creates or overwrites files
- `edit_file` -- modifies file content
- `math` -- sandboxed computation (confirmation required despite being sandboxed)
- `web_fetch` -- fetches a URL (SSRF-protected, but makes network requests)
- `web_search` -- web search via self-hosted SearxNG (makes network requests)
- `task_agent` -- spawns an autonomous sub-agent
- `web_search` -- web search via Tavily API (makes network requests)
- `task` -- spawns an autonomous sub-agent
- `plan` -- spawns a planning sub-agent, plus post-execution review gate
Note: The JSON schema metadata key `auto_approve` controls membership in
`TASK_AUTO_TOOLS` (used for task agent sub-sessions). The actual runtime
approval behavior is determined by the `needs_approval` field set in each
`_prepare_*` method on `ChatSession`. These two mechanisms can differ.
`AGENT_AUTO_TOOLS`/`TASK_AUTO_TOOLS` (used for agent sub-sessions). The actual
runtime approval behavior is determined by the `needs_approval` field set in
each `_prepare_*` method on `ChatSession`. These two mechanisms can differ.
---
@@ -155,9 +165,12 @@ Every tool defines a `primary_key`. The mapping is:
| `write_file` | `content` |
| `edit_file` | `old_string`|
| `search` | `query` |
| `math` | `code` |
| `man` | `page` |
| `web_fetch` | `url` |
| `web_search` | `query` |
| `task_agent` | `prompt` |
| `plan_agent` | `goal` |
| `memory` | `name` |
| `recall` | `query` |
| `notify` | `message` |
@@ -181,7 +194,7 @@ Execute a bash command and return stdout + stderr.
- **What it does**: Runs the command in a subprocess with a configurable timeout. Commands are sanitized and checked against a blocklist (e.g. `rm -rf /`). Environment variables containing secrets are scrubbed (`*_KEY`, `*_SECRET`, `*_TOKEN`, etc.).
- **Output format**: Stdout is returned directly. Stderr lines are prefixed with `[stderr]` so the model can distinguish them. When the command itself redirects stderr to stdout (`2>&1`), no prefix is added. Output exceeding 256KB is truncated (head + tail preserved, middle replaced with a truncation notice).
- **Auto-approve**: No -- requires user confirmation.
- **Agent availability**: `task_agent` only.
- **Agent availability**: `task_agent` only (not available to plan sub-agents).
---
@@ -199,7 +212,7 @@ base64-encoded image data for supported image formats.
- **What it does**: For text files, reads and returns content with line numbers. For image files (PNG, JPEG, GIF, WebP, BMP, TIFF, ICO), returns image data as multi-part content when the model supports vision, or a text description when it does not. SVG files are read as text. Images larger than 4 MB are rejected. Must be called before `edit_file` on the same path (the session tracks which files have been read).
- **Vision support**: Controlled by `ModelCapabilities.supports_vision`. All commercial OpenAI and Anthropic models have vision enabled. Local models (vLLM, llama.cpp, NIM) default to off — enable via `[models.*.capabilities] supports_vision = true` in config.toml.
- **Auto-approve**: Yes.
- **Agent availability**: `task_agent`.
- **Agent availability**: `agent` and `task_agent`.
---
@@ -255,7 +268,7 @@ Show a unified diff between two files, or between a file and a provided string.
- **What it does**: Returns unified diff output using Python's `difflib`. Binary files (containing null bytes) are rejected with a clear error. Files read through `diff_file` satisfy `edit_file`'s read guard — you can diff then edit without a separate `read_file` call. Large diffs are streamed with early cutoff at the tool truncation limit.
- **Auto-approve**: Yes (read-only).
- **Agent availability**: `task_agent`.
- **Agent availability**: `agent` and `task_agent`.
---
@@ -270,12 +283,44 @@ Search file contents for a regex pattern.
- **What it does**: Recursively searches for the pattern using `grep -rn`. Returns matching lines with file paths and line numbers.
- **Auto-approve**: Yes.
- **Agent availability**: `task_agent`.
- **Agent availability**: `agent` and `task_agent`.
---
## Computation
### math
Execute Python code for math and computation in a sandbox.
| Parameter | Type | Required | Description |
|-----------|--------|----------|-------------|
| `code` | string | yes | Python code to execute. Must use `print()` for output. |
- **What it does**: Runs Python code in a sandboxed environment with pre-imported libraries: `sympy`, `numpy`, `scipy`, `math`, `fractions`, `itertools`, `functools`, `collections`, `decimal`, `operator`, `random`, `re`, `string`. Common sympy names (`symbols`, `solve`, `simplify`, `sqrt`, `Matrix`, etc.) are pre-imported. `pytest` is also available for import.
- **Installation**: `sympy`, `numpy`, `scipy`, and `pytest` require the `[sandbox]` extras group: `pip install turnstone[sandbox]` (included in `[all]`).
- **Auto-approve**: Yes.
- **Agent availability**: `agent` and `task_agent`.
---
## Information
### man
Read a man page.
| Parameter | Type | Required | Description |
|-----------|--------|----------|-------------|
| `page` | string | yes | The man page name (e.g. `grep`, `socket`, `printf`). |
- **What it does**: Returns the full formatted manual entry. Preferred over `bash('man ...')` or `web_search` for command/API documentation.
- **Auto-approve**: Yes.
- **Agent availability**: `agent` and `task_agent`.
---
### web_fetch
Fetch a URL and extract specific information from it.
@@ -287,7 +332,7 @@ Fetch a URL and extract specific information from it.
- **What it does**: Fetches the URL, strips HTML to plain text, and uses the LLM to extract the answer to the question from the page content. Protected against SSRF (blocks private/internal IPs).
- **Auto-approve**: No -- requires user confirmation (makes network requests).
- **Agent availability**: `task_agent`.
- **Agent availability**: `agent` and `task_agent`.
---
@@ -299,55 +344,20 @@ Search the web using a text query.
| `max_results` | integer | no | Max results to return (default 5, max 20). |
| `category` | string | no | Search category: `general` (default), `news`,`it` (code/tech), or `science`. Maps to SearxNG categories; the model picks per query. |
| `topic` | string | no | Search topic: `general`, `news`, or `finance` (default `general`). |
- **What it does**: Searches the web and returns ranked results with titles, URLs, and content snippets. Uses provider-native search when available:
- **Anthropic**: Replaced at the API boundary with Anthropic's `web_search_20250305` server-side tool. Claude decides when to search; the API executes it and returns results with citations inline. No backend needed.
- **Anthropic**: Replaced at the API boundary with Anthropic's `web_search_20250305` server-side tool. Claude decides when to search; the API executes it and returns results with citations inline. No Tavily key needed.
- **OpenAI search models** (`gpt-5-search-api`): Replaced with `web_search_options` parameter. The model always searches and returns `url_citation` annotations.
- **Local/vLLM models**: Falls back to a self-hosted [SearxNG](https://searxng.org) instance. Set `searxng_url` in `config.toml``[tools]` or `$TURNSTONE_SEARXNG_URL` (the docker-compose stack bundles a `searxng` service and points at it by default). Operators with a custom MCP search server can instead set `web_search_backend = "mcp:server:tool"`.
- **Local/vLLM models**: Falls back to the Tavily API. Requires `tavily_key` in `config.toml` or `$TAVILY_API_KEY`.
- **Auto-approve**: Yes (auto-approved for all tool dispatch paths).
- **Agent availability**: `task_agent`.
---
### Reranking (optional)
`web_search` can use an external **reranker** to re-order the backend's result pool by relevance to the query before returning the top hits. Turnstone runs no reranker model itself; it POSTs to a Cohere/Jina-compatible `/rerank` endpoint (self-hosted [vLLM](https://docs.vllm.ai) / [TEI](https://github.com/huggingface/text-embeddings-inference) / llama.cpp, or hosted Cohere/Jina/Voyage).
**Disabled by default.** In the console **Models** tab, add a model definition whose `base_url` is a Cohere/Jina-compatible `/rerank` endpoint and whose capabilities include `{"supports_rerank": true}`, then select it under **Models → Roles → Reranker**. It's managed like every other model (write-only key, enable/disable, calibration). The reranker is purely this per-model definition — there is no global `rerank_url`-style endpoint setting.
The `rerank_web_search` toggle defaults on once a reranker is selected. If the endpoint is unreachable or errors, web_search falls back silently to the backend's native result order — reranking never makes a search fail.
When `rerank_bm25` is enabled, the candidate text for memory, tool, and skill retrieval (memory name/description/content and tool/skill names + descriptions) is also sent to the rerank endpoint — a self-hosted endpoint (vLLM/TEI/llama.cpp) keeps it on your infrastructure, a hosted provider (Cohere/Jina/Voyage) sends it off-box.
**Serving a Qwen3-Reranker with vLLM.** The model is instruction-aware, so vLLM **must** apply its chat template — pass `--chat-template` explicitly. Without it the bare query produces near-random scores and reranking actively *hurts* retrieval (verified: an irrelevant passage outscored the correct one):
Then add a reranker model in the **Models** tab with `base_url``http://vllm:8000/rerank` (model name `qwen3-reranker`) and select it under **Models → Roles → Reranker**.
For an endpoint that does *not* apply the model's template, set `rerank_instruction` instead — Turnstone then wraps each query as `<Instruct>: {instruction}` / `<Query>: {query}` (Qwen3's own default is `Given a web search query, retrieve relevant passages that answer the query`). Use the chat template **or** the instruction, not both (they double-wrap).
**Picking `rerank_bm25_threshold`.** The relevance floor that gates proactive memory injection is a probability in `[0, 1]`, but the right value differs per model (a sharp 0.6B reranker may want ~0.95; a broader 4B ~0.33). Calibrate it against your endpoint:
```bash
turnstone-admin rerank-calibrate # probe the endpoint, recommend a floor
It reports the score scale, whether the endpoint cleanly separates relevant from irrelevant probes (a **"no clean separation"** result flags a mis-served or weak reranker), and the suggested floor. Leave the threshold at `0` to rerank-without-filtering.
- **Agent availability**: `agent` and `task_agent`.
---
## Agent
The tool name uses the `_agent` suffix — bare `task` collides with
Tool names use the `_agent` suffix — bare `plan` / `task` collide with
chat-template channel names on some local models.
### task_agent
@@ -358,9 +368,23 @@ Delegate a general-purpose task to an autonomous sub-agent.
|-----------|--------|----------|-------------|
| `prompt` | string | yes | Complete task description for the sub-agent. |
- **What it does**: Spawns a sub-agent that inherits the `TASK_AGENT_TOOLS` set (read, write, edit, search, bash, web tools, memory tools). The sub-agent runs autonomously to completion. Use for work that requires file modifications or command execution.
- **What it does**: Spawns a sub-agent that inherits the `TASK_AGENT_TOOLS` set (read, write, edit, search, bash, math, man, web tools, memory tools). The sub-agent runs autonomously to completion. Use for work that requires file modifications or command execution.
- **Auto-approve**: No -- requires user confirmation.
- **Agent availability**: Top-level only.
- **Agent availability**: Not available to sub-agents (top-level only).
---
### plan_agent
Plan before implementing -- an autonomous agent explores the codebase and writes a structured plan.
| Parameter | Type | Required | Description |
|-----------|--------|----------|-------------|
| `prompt` | string | yes | What to plan -- the goal, constraints, and scope. |
- **What it does**: Spawns a planning sub-agent with `AGENT_TOOLS` (read-only tools: `read_file`, `search`, `math`, `man`, `web_fetch`, `web_search`). The agent explores the codebase and writes a structured plan to `.plan-<ws_id>.md` (unique per workstream, so concurrent workstreams never collide). If the `plan` tool has been called before in the same session, the prior plan is passed to the agent as context so it refines rather than restarts. After completion, the user is prompted to review and can accept, reject, or annotate the plan.
- **Auto-approve**: No -- requires user confirmation, plus post-execution review gate.
- **Agent availability**: Not available to sub-agents (top-level only).
---
@@ -421,7 +445,7 @@ Provide either `username` for user-based targeting or `channel_type` +
- **What it does**: Sends a notification via the channel gateway's HTTP endpoint (`POST /v1/api/notify`). The server queries the `services` table for healthy channel gateways, authenticates with a service JWT (`aud: turnstone-channel`), and delivers to the first healthy gateway. On failure, retries up to 2 additional times with backoff (1s, 3s). Rate-limited to 5 notifications per turn (counter only increments on success).
- **Auto-approve**: Yes — notifications are time-sensitive and auto-approved so the model can alert users urgently.
- **Agent availability**: `task_agent`.
- **Agent availability**: `agent` and `task_agent`.
> See [Channel Integrations: Notifications](channels.md#notifications)
> for the full delivery flow, service registry details, and security
@@ -497,7 +521,7 @@ data.get("mergedAt") is not None
- Duplicate names rejected within the same workstream.
- **Auto-approve**: `create` requires approval; `list` and `cancel` are auto-approved.
- **Agent availability**: Main session only — not available to task sub-agents.
- **Agent availability**: Main session only — not available to plan/task sub-agents.
> See [Watch Architecture](diagrams/png/18-watch-architecture.png) for the
> full poll → evaluate → dispatch flow.
@@ -530,30 +554,33 @@ pre-configure skills at workstream creation.
- **Task sub-agents** — via `self._task_tools` (merged list)
- **Plan sub-agents** — via `self._agent_tools` (merged list)
### Naming convention
@@ -747,7 +774,7 @@ MCP tool lists stay up-to-date without restart through two mechanisms:
When tools change, `MCPClientManager` rebuilds its merged tool list using copy-on-write
(new list/dict objects assigned atomically) and notifies all active `ChatSession`
instances via registered listener callbacks. Each session rebuilds its `_tools`,
`_task_tools`, and reconstructs its `ToolSearchManager` (if active),
`_task_tools`,`_agent_tools`, and reconstructs its `ToolSearchManager` (if active),
preserving the set of previously expanded (discovered) tools.
```
@@ -803,7 +830,7 @@ Use read_resource(uri='...') to access the resources listed above.
- **What it does**: Reads the resource from its MCP server via `MCPClientManager.read_resource_sync()`. Returns text content for text resources or base64-encoded data for binary resources. Output is truncated by the standard tool output limiter.
- **Auto-approve**: No -- requires user confirmation (reads external data).
- **Agent availability**: `task_agent`.
- **Agent availability**: `agent` and `task_agent`.
### Capability guards
@@ -845,7 +872,7 @@ the `initialize` handshake. Each prompt is stored with its prefixed name
- **What it does**: Invokes an MCP prompt by name via `MCPClientManager.get_prompt_sync()`, expanding it into messages. Returns the expanded prompt content formatted as `[role]: content` blocks joined with blank lines. The prompt catalog is listed in the system message so the model knows which prompts are available. Output is truncated by the standard tool output limiter.
- **Auto-approve**: No -- requires user confirmation (invokes external prompt servers).
- **Agent availability**: `task_agent`.
- **Agent availability**: `agent` and `task_agent`.
# Resolve how to invoke docker (direct / via sudo) and make sure the daemon runs.
ensure_docker_usable(){
if ! have docker;then
warn "Docker is not installed."
ask "Install Docker now?" y || die "Docker is required."
install_docker
have docker || die "Docker installation failed."
fi
docker info >/dev/null 2>&1&&{DOCKER="docker";return;}
# Maybe the daemon just isn't running yet.
start_docker_daemon
docker info >/dev/null 2>&1&&{DOCKER="docker";return;}
# Maybe we lack permission (not in the docker group yet, group change not
# active in this shell) — fall back to sudo for this run.
if["$(id -u)" -ne 0]&& have sudo && sudo docker info >/dev/null 2>&1;then
DOCKER="sudo docker"
warn "Using 'sudo docker' for this run. To drop the sudo: 'sudo usermod -aG docker $USER' then log out and back in."
return
fi
if["$IS_WSL" -eq 1];then
die "Docker isn't usable inside WSL. Install Docker Desktop on Windows and enable WSL integration for this distro (Settings -> Resources -> WSL integration), then re-run."
fi
die "Docker is installed but not usable — the daemon may be stopped or you lack permission. Try: 'sudo systemctl start docker', then re-run."
}
ensure_compose(){
$DOCKER compose version >/dev/null 2>&1&&return
warn "The Docker Compose v2 plugin is missing."
case"$PKG" in
apt|dnf|yum) ask "Install docker-compose-plugin?" y && pkg_install docker-compose-plugin ||true;;
pacman) ask "Install docker-compose?" y && pkg_install docker-compose ||true;;
esac
$DOCKER compose version >/dev/null 2>&1\
|| die "'docker compose' is unavailable. Install Docker Compose v2 and re-run."
$DOCKER info 2>/dev/null | grep -qi 'rootless'&&ROOTLESS=1
pick_node_count
info "Building the image (first run pulls dependencies — this can take a few minutes)…"
if ! (cd"$INSTALL_DIR"&&$DOCKER compose build );then
die "image build failed (see output above). Common causes: low memory or disk, or the Docker daemon stopped. Free up resources and re-run — it resumes."
fi
prepare_env
info "Ports — dashboard (Caddy): ${CADDY_PORT}, PostgreSQL: 127.0.0.1:${PG_PORT}"
write_node_override
info "Starting the stack…"
# Plain `up -d` (honoring compose.override.yaml) + --remove-orphans so a
# re-run that lowers the count also stops the now-excluded nodes.
if ! (cd"$INSTALL_DIR"&&$DOCKER compose up -d --remove-orphans );then
die "the stack failed to start (see output above). Inspect logs: cd $INSTALL_DIR && $DOCKER compose logs"
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.