Compare commits

...

273 Commits

Author SHA1 Message Date
Patrick Buckley d76c57e687 chore: bump version to 1.5.12 2026-05-10 18:28:46 -07:00
Patrick Buckley 58bf811a0e fix(server): always emit history event on /rewind to unblock edit-and-resend
Editing the first message in a workstream sends /rewind N where N is the
total user turns, leaving session.messages empty. The handler guarded the
history event with `if history:`, so only clear_ui was emitted. The
frontend dispatches the queued edit-and-resend from the history event
handler (app.js _pendingEditSend), so an empty history orphaned the
pending text and left the composer stuck in busy.

replayHistory already handles the empty case via showEmptyState(), so
emitting the event unconditionally is safe and unblocks the dispatch.
2026-05-10 18:27:23 -07:00
Patrick Buckley df942e375e feat(session): enriched backend error messages with provider + URL
A bare ``httpx.ReadTimeout`` previously surfaced as ``ReadTimeout: timed
out`` — no provider, no base URL, no model — leaving the user with no
signal to tell whether a model server hung, the URL was wrong, or the
model isn't loaded on the backend.

``ChatSession._format_backend_error`` now rewrites known boundary
exceptions (httpx ``ReadTimeout`` / ``ConnectError`` / etc. and OpenAI /
Anthropic SDK ``APITimeoutError`` / ``APIConnectionError`` /
``NotFoundError`` / ``AuthenticationError`` / ``RateLimitError``) into
operator-actionable text that names the provider, base URL (query
string stripped before ``sanitize_error_text`` redacts credentials),
and model.  Matching is by class name so the helper carries no SDK
imports.  Unrecognised exceptions fall through to the legacy
``f"{type(exc).__name__}: {exc}"`` shape, preserving existing grep
targets.
2026-05-10 18:27:14 -07:00
Patrick Buckley 84fd5dc859 chore: bump version to 1.5.11 2026-05-09 17:29:46 -07:00
Patrick Buckley ba8b1d9126 fix(session): AND-gate replay_reasoning_to_model with model capability
The Anthropic call sites in session.py passed the operator-side
`replay_reasoning_to_model` flag through without checking the
model's static `supports_reasoning_replay` capability. The OpenAI
Responses path AND-gated both flags in `_build_kwargs` so a model
without a reasoning lane (gpt-4o, etc.) silently skipped replay even
when the operator flag was set. The Anthropic path had no such gate.

For all current Claude entries this was a no-op asymmetry - every
`_ANTHROPIC_CAPABILITIES` row sets `supports_reasoning_replay=True`,
so `True AND op == op`. But:

- The capability flag was dead code on the Anthropic path
- A future Claude entry (or any Anthropic-shaped surface) shipping
  with the cap left at its False default would have replay fire
  anyway, against the cap declaration
- The asymmetry made `supports_reasoning_replay` an unreliable
  signal - readers couldn't tell if it gated anything per-provider

Move the AND-gate into `_resolve_replay_reasoning_to_model` via a
new optional `caps=` kwarg. When caps is provided, the resolver
returns `operator_on AND caps.supports_reasoning_replay`; when
omitted (back-compat for any caller not yet updated), it returns
the operator flag unchanged.

Thread caps through the three call sites: `_utility_completion`
(non-streaming), `_try_stream` (streaming, hoisted resolution out
of the retry loop since caps are attempt-invariant), and the
agent `_api_call` closure in `_run_agent`.

With the AND-gate now living at the session resolver, the redundant
in-provider gate in `OpenAIResponsesProvider._build_kwargs` is
removed. The provider now trusts the resolved bool it receives,
matching the AnthropicProvider shape and giving the cap a single
source of truth across providers. The two provider-level tests
that pinned the in-provider gate
(`test_include_omitted_when_capability_false`,
`test_include_omitted_by_default`) drop out; the session-level
boundary test
`TestSessionToOpenAIResponsesBoundaryIntegration::test_capability_false_omits_include_even_when_flag_true`
already covers the same end-to-end invariant.

Tests added:
- 4 resolver-level tests pinning the AND-gate semantics +
  back-compat when caps is omitted
- 1 wire-boundary integration test mirroring the OpenAI Responses
  `test_capability_false_omits_include_even_when_flag_true` -
  drives session._try_stream through the real AnthropicProvider
  with operator flag True + capability False and asserts the
  thinking block does NOT reach the SDK boundary

Existing `TestUtilityCompletionPassesFlag` test had its caps mock
upgraded from `SimpleNamespace` to a real `ModelCapabilities`
instance to satisfy the new attribute read and stay robust to
future capability fields.
2026-05-09 17:23:56 -07:00
Patrick Buckley 685b1e3d9b fix(console): preserve cs=None fallback in /v1/api/models placeholder
Copilot review feedback on #500.  The original
``list_available_models`` had an implicit cs=None branch where the
placeholder still advertised ``registry.default`` (filtered against
enabled rows) when ``app.state.config_store`` was None but
``coord_registry`` was bound — useful in the rare degraded state
where lifespan wired the registry but the ConfigStore failed to
initialise.  The PR #500 refactor accidentally dropped that branch:
the helper requires a config_store, so the cs=None case fell out as
"blank coordinator default".

Add an explicit ``elif coord_registry is not None`` branch that
mirrors the helper's tier 3 with the placeholder's enabled-rows
filter applied.  New test exercises this path by passing
``config_store=False`` to the test fixture.
2026-05-09 17:23:56 -07:00
Patrick Buckley 91300060fa fix(console): unify coordinator alias resolution across placeholder + factory
Previously /v1/api/models (home composer placeholder) and
console/session_factory.py walked separate two-/three-tier chains for
the coordinator alias.  session_factory was missing the
``model.default_alias`` tier, so admins who set the system default in
the Models tab would see it advertised but new coordinator sessions
would silently keep launching on ``registry.default``.

This commit:

- Extracts the chain into ``turnstone/console/coordinator_alias.py``.
  ``resolve_coordinator_alias`` returns the effective alias under a
  shared three-tier policy: explicit pin → ``model.default_alias`` →
  ``registry.default``.  Tier 2 is validated against
  ``registry.has_alias`` and falls through to tier 3 with a logged
  warning if unknown.  Tier 1 is intentionally passed through
  unvalidated so an explicit operator pin surfaces as 503 at
  ``registry.resolve`` rather than being silently swapped out.
- Wires both call sites through the helper.  The placeholder supplies
  an ``alias_filter`` that restricts every tier to enabled DB rows so
  the home composer never advertises a model the workstream picker
  can't actually offer; the session factory uses no filter (matches
  prior 503-on-typo behaviour for explicit pins).
- Adds direct integration tests for the session factory's chain
  (``tests/test_console_session_factory.py``) and updates the
  placeholder tests' fixture to provide a stub coord_registry, since
  the helper now requires one.
2026-05-09 17:23:56 -07:00
Patrick Buckley 688f047ce1 docs(console-ui): clarify coordinator placeholder fallback comment
Light-review followup on 389400c8.

The "mirrors session_factory.py:109-110" claim was inaccurate —
session_factory's chain is two tiers (coordinator.model_alias →
registry.default) and skips model.default_alias entirely.  The
placeholder handler extends that chain with model.default_alias as
tier 2 so admins who set the default in the Models tab see it
advertised in the home composer.  Comment now lists the three tiers
explicitly and flags the session_factory-vs-placeholder drift case
(where model.default_alias ≠ registry.default) as a separate issue
to track.

Also lifts the ``from types import SimpleNamespace`` import in the
test fixture to module level — minor readability cleanup.
2026-05-09 17:23:56 -07:00
Patrick Buckley 87893aa4ab fix(console-ui): align coordinator placeholder fallback with session_factory
Two Copilot-review followups on /v1/api/models default resolution.

- console/server.py: coordinator_default_alias now mirrors the full
  fallback chain in console/session_factory.py:109-110 — explicit
  coordinator.model_alias → model.default_alias → registry.default.
  The registry tier was missing, so the home composer placeholder went
  blank whenever an operator never set model.default_alias in the admin
  UI even though new coordinator sessions still launch on
  registry.default (loaded from config.toml [model].default by
  load_model_registry).  Two new tests cover the registry-default
  branch and the disabled-alias guard.
- console/static/app.js: _resolveModelLabel returns "" (not the bare
  alias) when the alias isn't found in the dropdown's model list, so
  callers can rely on the documented "fall back to neutral placeholder"
  contract.  Matches the existing doc comment.
2026-05-09 17:23:56 -07:00
Patrick Buckley 0988142303 feat(console-ui): home composer placeholders, toggle component, admin polish
Bundles the click-around polish on the console admin UX.

Home composer + schedule modals
- /v1/api/models now exposes coordinator_default_alias + judge_default_alias,
  resolved through the same chain console/session_factory.py uses.  Both the
  home composer's MODEL / JUDGE MODEL placeholders and the schedule create /
  edit modal model placeholders rewrite to "Default — alias (model)" once
  the API responds.  The `models_changed` SSE refresh keeps placeholders
  current as operators edit per-role assignments.
- Composer.setOptionPlaceholder added so callers can update just the first
  option's text without disturbing the rest of the choice list.

Admin → Models → Roles
- Channel adapter row added (channels.default_model_alias) — the migration
  to the Roles sub-tab missed it.  Key added to
  _MODEL_AFFECTING_SETTING_KEYS so edits fire the SSE refresh, and to the
  settings-tab roleKeys skip-list so it only renders in one place.
- Plan/Task agent rows now display "(inherit)" instead of the misleading
  "(default — <alias>)" — those roles cascade through plan_model →
  agent_model → session model, not a single concrete default.
- coordinator.reasoning_effort accepts "" (inherit), matching
  model.plan_effort / model.task_effort.
- Blank options in each role's MODEL select now match the "alias (model)"
  shape used by the other rows.

Toggle-switch component
- New .toggle-switch component (visually-hidden native checkbox + styled
  track + label).  40×22 hit target meets WCAG 2.5.5 (AAA), inset ring on
  the off state for ≥1.5:1 contrast against the modal surface.
- .toggle-stack groups toggles in a column with .toggle-group-divider for
  conceptual grouping (used in the Add Model modal between "Active" and the
  paired Reasoning toggles).
- .toggle--flush modifier zeroes the default top margin for toggles that
  sit flush against a heading or a dynamically-rendered row.

Sweep — every admin-modal boolean checkbox is now a toggle:
schedule (cs/es-autoapprove, es-enabled), policy (ep/epp-enabled),
tool-mode (ctm/etm-default), skill (csk/esk-auto-approve, csk/esk-enabled),
MCP (mcp-auto-approve, mcp-enabled), Add Model (Active, surface-persisted-
reasoning, replay-reasoning), judge bool settings (cancel_on_approval et
al.), and the user-roles-modal role assignment list.  The two
ogp-cred / eogp-cred inline credential checkboxes stay as compact inline
boxes since they sit beside text inputs in tight horizontal rows.

Add Model modal — the "Enabled" toggle promoted to "Active" and moved to
the very top of the form.  Tooltip explains it gates dropdown visibility
without removing the definition.

MCP authorization — the three radio buttons replaced with a vertical
.segmented-control option list.  Selected row paints --accent-dim plus a
filled .segmented-indicator; focus ring uses --accent so it stays visible
on the currently-selected option.

Role permissions modal — the 19 permission checkboxes are now
.toggle-switch.perm-toggle (monospace lowercase identifiers preserved).
The permissions are split into Scopes / Admin / Workstreams & Tools
sections under caps-styled section headers so the row-flow grid no longer
slices `admin.*` mid-column.

Judge bool toggles use a static "Enabled" caption rather than flipping
text on `.checked`; flipping lagged 50–300 ms behind the slider position
because the caption was sourced from the post-save reload.

CSS cleanup — dead `.admin-checkbox` / `.perm-checkbox` rules removed.
Specificity audit (scripts/css_specificity_audit.py) returns no conflicts
on any new component class.

Tests — 525 pass on the affected slices; new tests/test_console_available_
models.py pins each branch of the resolution chain in /v1/api/models so the
home composer placeholder stays correct as precedence rules evolve.
2026-05-09 17:23:56 -07:00
Patrick Buckley 40ecebf012 refactor(judge): require alias for judge.model, drop session-provider raw-model fallback
`IntentJudge.__init__` previously had a 3-way resolution chain: registered
alias → raw model id pinned onto the session provider → session model.  The
middle branch was a footgun documented in `console/session_factory.py:130-137`
— pinning the literal `judge.model` string onto the coordinator's session
provider silently broke every verdict whenever that provider didn't recognise
the model id (e.g. coordinator on Anthropic, `judge.model = "gpt-5-mini"` →
uniform `llm_fallback`).

Tightens to alias-only, matching `coordinator.model_alias` /
`model.plan_alias` / `model.task_alias`.  An unknown `config.model` now logs
a warning and inherits the session model — same path as empty.  Help text on
`judge.model` updated to clarify the contract.

Adds two regression tests in `TestModelAliasResolution` covering the
session-model inheritance for unknown values and the empty-model self-
consistency case.
2026-05-09 17:23:56 -07:00
Patrick Buckley c91869c7e5 fix(reasoning): synthesize reasoning_text alongside non-reasoning provider_blocks
GoogleProvider attaches raw tool_call dicts as ``provider_blocks`` on
the finish chunk for ``thought_signature`` round-trip
(``_google.py:_iter_stream``).  When the same turn streamed Gemini's
``reasoning_content`` as ``reasoning_delta`` chunks, the prior
synthesizer bailed out the moment ``provider_blocks`` was non-empty
— so the captured reasoning was visible live but lost on page reload.

Replace the early-return-if-non-empty check with a reasoning-bearing
type test (``thinking`` / ``redacted_thinking`` / ``reasoning`` /
``reasoning_text``).  When none of those types appear, append the
synthetic ``reasoning_text`` block to the existing list rather than
replacing it — preserving Google's tool-call fidelity blocks.

Also addresses two doc-accuracy review findings:
- ``LLMProvider.extract_reasoning_text`` docstring no longer claims
  OpenAI Chat / Responses are unwired (Phase 3+4 shipped extractors).
- Add the method to the Protocol methods table in
  ``docs/architecture.md`` (was missing alongside the class diagram).
2026-05-09 17:23:56 -07:00
Patrick Buckley 53b52092f9 fix(reasoning): per-block ANTHROPIC_VALID_BLOCK_TYPES filter + review fixes
The earlier all-or-nothing shape check on ``_provider_content`` discarded
every valid Anthropic block in a message the moment a single foreign
block (OpenAI ``reasoning``, Gemini thought parts, the synthetic
``reasoning_text`` from path-3 capture) appeared.  In the cross-model
resumption edge case that meant ``server_tool_use`` /
``web_search_tool_result`` blocks lost their ``encrypted_content``
silently, breaking web-search round-trip continuity on subsequent turns.

Replaced with a per-block walk: foreign blocks are dropped individually,
valid blocks ride the verbatim path, and an identity-preserving fast
path reuses the source list reference when nothing was filtered or
stripped (pinned by the ``is`` assertions in test_providers.py).

Also addresses validation-pass review findings:
- Document the single-tier vs three-tier ``surface_persisted_reasoning``
  resolution divergence between server.py:_build_history and
  session_routes.make_history_handler.
- Document why OpenAIResponsesProvider._convert_messages defaults
  ``replay_reasoning_to_model=False`` while Anthropic's defaults True.
- Document the ``source`` metadata field on synthetic ``reasoning_text``
  blocks as reserved-for-future-use, not dead code.
- Add edge tests for non-dict / missing-type-key blocks in
  _provider_content (defensive branches in the per-block walk).
2026-05-09 17:23:56 -07:00
Patrick Buckley 0d1a009a4c test(reasoning): skip wire-boundary tests when anthropic extra missing
CI test job installs `[test]` extras, which omits `anthropic`. The two
TestSessionToWireBoundaryIntegration cases drive the real
AnthropicProvider.create_streaming, which calls _ensure_anthropic() and
raises ImportError. Match the repo convention (test_channel_discord,
test_channel_slack, test_tls_*) by gating the helper with
pytest.importorskip("anthropic").
2026-05-09 17:23:56 -07:00
Patrick Buckley e8352bd8e5 fix(reasoning): apply Copilot review feedback + docs sync
PR #498 round-robin review surfaced 5 findings.  4 applied; 1 rejected
with rationale.

Applied

* **Copilot finding 5** (history_decoration.py:341): dispatcher
  inspected only ``provider_content[0]['type']``.  OpenAI Responses
  captures EVERY ``output_item.done`` event into ``provider_blocks``
  (not just reasoning) — in practice the order is
  ``[reasoning, message, ...]`` but the API doesn't guarantee that;
  a hypothetical ``[message, reasoning]`` ordering would silently
  drop the reasoning under an index-only check.  Now walks the list
  for the first block whose type is in ``_BLOCK_TYPE_PROVIDER_FACTORY``,
  then dispatches the WHOLE list to that provider's extractor.  Each
  provider's extractor already filters internally by its own block
  type, so passing the full list is correct.  Regression test added
  (``test_dispatcher_scans_past_unrecognized_first_blocks``).

* **Copilot finding 3** (migration 052 docstring): the previous
  review-fix wave used sed to rename ``persist_reasoning`` →
  ``surface_persisted_reasoning`` everywhere, which mangled a
  historical reference in the migration docstring ("The earlier name
  ``surface_persisted_reasoning`` was renamed...").  Restored to
  point at the actual pre-rename name (``persist_reasoning``).

* **Copilot finding 4** (sdk/typescript/src/events.ts:26):
  ``HistoryEvent`` JSDoc still referenced ``persist_reasoning`` —
  the sed rename only walked ``turnstone/`` and ``tests/``, missing
  the TypeScript SDK.  Updated to ``surface_persisted_reasoning``.
  Also widened the comment to cover all three reasoning-bearing
  block types (Anthropic ``thinking``, OpenAI Responses ``reasoning``,
  synthetic ``reasoning_text``) instead of mentioning only Anthropic.

* **github-code-quality finding** (session.py:1120): ``_resolve_server_type``
  had a bare ``except Exception: pass``.  Replaced with a
  ``log.debug(..., exc_info=True)`` + explanatory comment.  Behaviour
  unchanged (still returns ``""`` on any lookup failure); failures
  are now observable under DEBUG triage.

Rejected (with rationale)

* **github-code-quality finding** (_protocol.py:265):
  ``extract_reasoning_text``'s body is ``...`` per ``LLMProvider``
  Protocol convention.  Every method in the file uses ``...`` (PEP
  544 idiomatic Protocol style).  Changing only this one to
  ``raise NotImplementedError`` would be inconsistent with the rest
  of the file.  CodeQL's "statement has no effect" warning is
  technically correct for ``...`` as a standalone expression but
  ignores the documented Python Protocol convention.  No fix.

Docs sync

* docs/api-reference.md: ``history`` SSE event message-shape table
  gains the optional ``reasoning`` field.
* docs/architecture.md: ``ModelCapabilities`` row in the type table
  gains ``supports_reasoning_replay``; ``StreamChunk`` and
  ``CompletionResult`` rows gain the existing ``provider_blocks``
  field (was missing pre-PR).  New "Per-model reasoning persistence"
  subsection under the Models config section, documenting the two
  flags + capability gate + three reasoning paths + cross-provider
  shape filter.
* docs/settings.md: new "Reasoning persistence (per-model)"
  subsection with the two-flag table and capability-gate note.
* docs/diagrams/03-core-engine-classes.puml: ``LLMProvider`` interface
  adds ``extract_reasoning_text`` + the new ``replay_reasoning_to_model``
  kwarg; ``ModelCapabilities`` class adds ``supports_reasoning_replay``.
  PNG regenerated.

Lint + test gate

* ruff check + ruff format clean.
* mypy clean (191 source files).
* pytest -m 'not live' — 6116 passed (3 deselected), +1 net new test
  (``test_dispatcher_scans_past_unrecognized_first_blocks``).
2026-05-09 17:23:56 -07:00
Patrick Buckley de4cc568c4 fix(reasoning): apply full-stack review findings
Multi-stage /review on the full Phase 1+2+3+4 stack surfaced 9 findings
(0 critical, 3 major, 5 minor, 1 nit, 1 uncertain).  All applied.

Major

* perf-1 (session_routes.py:2402): make_history_handler ran sync
  storage.load_workstream_config inside async def history on the cold-
  workstream path, blocking the event loop on every dashboard /history
  request for non-resident workstreams.  Every other storage call in
  the same handler correctly used asyncio.to_thread.  Wrap the sync
  call in asyncio.to_thread (preserving the existing try/except so a
  DB failure still degrades to the conservative-default branch instead
  of bubbling out).

* q-2 (test_reasoning_audit_log_discipline.py): the security-sensitive
  test (reasoning text never lands at INFO+ severity) only covered the
  4 Phase 1 surfaces.  Phase 2 added the strip predicate in
  AnthropicProvider._convert_messages and Phase 3 added 3 more code
  paths that touch reasoning text — none guarded.  Added 4 parallel
  tests using the existing capture-and-walk infrastructure:
  OpenAIResponsesProvider.extract_reasoning_text,
  OpenAIChatCompletionsProvider.extract_reasoning_text,
  ChatSession._stream_response (drives the synth-block stamp via a
  fake reasoning-emitting stream), AnthropicProvider._convert_messages
  with replay_reasoning_to_model=False (drives the Phase 2 strip
  predicate).

* q-1 (model_registry.py:42): the persist_reasoning flag name implied
  storage-control but actually gates UI rehydration only — operators
  flipping it could reasonably expect "stop persisting reasoning" but
  storage of reasoning bytes happens in provider_data regardless.
  Renamed everywhere to surface_persisted_reasoning: ModelConfig
  field, migration 052 column (renaming in-place since 052 is not yet
  on main), schema, MODEL_DEFINITION_MUTABLE allowlist, _postgresql.py
  + _sqlite.py CRUD impls, _protocol.py create_model_definition
  signature, 3 console_schemas Pydantic models, console/server.py
  admin POST + PUT, model_registry row mapper, history_decoration.py
  helper parameter, server.py _build_history local var,
  session_routes.py make_history_handler local var, sdk/events.py
  HistoryEvent docstring, admin.js form id + override pill label,
  index.html form input id + UI label + tooltip, coordinator.js (none
  needed), and every test that referenced the old field name.  The
  admin tooltip now reads "Storage of reasoning bytes is unaffected
  by this flag — they ride in provider_data regardless" so the
  decoupling stays explicit at the operator surface.

Minor

* bug-1 (history_decoration.py:336): dispatcher discriminated on
  provider_content[0]["type"] only.  Anthropic's redacted_thinking
  blocks (sealed by the safety system) can appear before, after, or
  interleaved with regular thinking blocks per the API docs.  When a
  redacted block lands first, the dispatcher returned "" and the UI
  silently lost the surrounding thinking text.  Registered
  "redacted_thinking" as a second key in _BLOCK_TYPE_PROVIDER_FACTORY
  pointing at the same AnthropicProvider factory — the existing
  extractor's type=="thinking" filter already correctly skips redacted
  blocks while walking the full list.  Regression test added.

* q-3 (_protocol.py:155): replay_reasoning_to_model defaults split
  across 9 sites — operator-side defaults to False (matches DB
  server_default), provider-API defaults to True (back-compat with
  direct callers).  Original "pick False everywhere" fix would have
  silently flipped behaviour for any direct provider caller.  Instead
  documented the intentional bifurcation in the Protocol's
  create_streaming docstring.

* q-4+q-5 (_protocol.py:107 + 3 providers): MAX_REASONING_DISPLAY_BYTES
  was enforced via Python str slicing which counts code points, not
  UTF-8 bytes — 4-byte CJK/emoji glyphs would blow past the byte
  ceiling.  Renamed to MAX_REASONING_DISPLAY_CHARS to match actual
  behaviour.  Hoisted the 4-line truncation pattern into a shared
  _join_reasoning_with_cap helper in _protocol.py; each provider's
  extractor becomes a single line at the tail.

* q-6 (tests/_session_helpers.py): _NullUI + _make_session were
  duplicated verbatim between test_session_replay_reasoning.py and
  test_session_synth_reasoning_block.py.  Hoisted to a shared
  tests/_session_helpers.py module (importable, leading underscore so
  pytest doesn't try to collect it).  test_model_registry.py's
  _make_session has a different signature (registry/model_alias args
  + _FakeUI) and is not a candidate for sharing.

Nit

* q-7 (history_decoration.py:286): _make_provider_factory used a
  dict-as-cell workaround for closure read-only scope.  Replaced with
  the more idiomatic nonlocal pattern.

Lint + test gate

* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6115 passed (3 deselected).  Net +5 tests
  (4 audit-log discipline + 1 redacted_thinking dispatcher).

Refinements vs the dedupe output (caught during sanity rendering
the report)

* perf-1 fix preserved the try/except wrapper.  The original "wrap in
  to_thread" one-liner would have let an OperationalError bubble out
  instead of degrading to the fallback branch.

* q-3 fix explicitly documented the bifurcation rather than
  collapsing both sides to False.  "Pick False everywhere" would
  silently flip back-compat behaviour for direct provider callers.

* q-1 fix included the admin.js:5292 fallback site
  (m.persist_reasoning !== false) that the original threaded-change
  list missed.

* q-6 fix verified the third _make_session in test_model_registry.py
  is structurally different (different signature + different UI
  helper) and intentionally NOT a dedupe target.
2026-05-09 17:23:55 -07:00
Patrick Buckley b477c85ddc feat(reasoning): OpenAI Responses + Chat Completions capture/replay (Phase 3+4)
Wire reasoning capture and (where the API supports it) replay for the
two remaining provider paths.  Phase 3 was originally scoped as
"OpenAI Responses + Gemini" but a spike against the OpenAI SDK source
revealed that Gemini routes through the OpenAI-compatible endpoint
(``/v1beta/openai/``), which is structurally identical to vLLM /
llama.cpp / any other Chat-Completions-shaped local model.  Phase 3
and Phase 4 collapse into one feature with two distinct sub-paths:

* **Path 2 (OpenAI Responses)** — full capture+replay.  ``include=
  ["reasoning.encrypted_content"]`` on the request makes the API
  surface ``encrypted_content`` on reasoning items in
  ``provider_blocks``; ``_convert_messages`` round-trips them as
  ``ResponseReasoningItemParam`` input items on subsequent turns.
  Verified against the OpenAI Python SDK 2.33.0 source
  (``response_reasoning_item.py:31-62``,
  ``response_reasoning_item_param.py:33-37``,
  ``response_create_params.py:70-74``).  Even with ``store=False``,
  ``encrypted_content`` round-trips correctly per the SDK's own
  documentation.

* **Path 3 (Chat Completions / vLLM / llama.cpp / Gemini-compat)** —
  persist-only.  Canonical OpenAI Chat Completions has no reasoning
  field on the wire, but several local-model servers tack on
  ``delta.reasoning_content`` as Pydantic extras.  ``ChatSession.
  _maybe_synth_reasoning_block`` stamps a synthetic ``{type:
  "reasoning_text", text, source?}`` block onto ``_provider_content``
  at end-of-stream when no native ``provider_blocks`` were emitted but
  ``reasoning_parts`` accumulated text.  The ``source`` field carries
  ``server_compat.server_type`` (vllm, llama.cpp, sglang, …) for
  diagnostic value — informational only, doesn't gate behaviour.
  Reasoning text NEVER replays back to the model on this path; it
  rides ``_provider_content`` only for ``/history`` UI rehydration
  and gets stripped from the wire by the existing
  ``sanitize_messages`` underscore-prefix strip on every request.

What this change does

* ``ModelCapabilities.supports_reasoning_replay: bool = False`` added
  to the dataclass.  Set True on every OpenAI reasoning model
  (gpt-5* + o-series via the Responses API) and every Anthropic
  Claude entry (default + 6 model-specific).  Path-2 wire-build does
  ``replay_active = bool(replay_reasoning_to_model and caps.supports_
  reasoning_replay)`` so an operator who flips the flag on a
  non-reasoning model (gpt-4o via Responses) silently no-ops rather
  than emit a malformed ``include=`` request.

* ``OpenAIResponsesProvider`` gains:
  - ``_build_kwargs`` accepts ``replay_reasoning_to_model: bool``
    (threaded from ``create_streaming``/``create_completion``);
    adds ``include=["reasoning.encrypted_content"]`` when active.
  - ``_convert_messages`` accepts the same flag, captures
    ``_provider_content`` reasoning items pre-sanitization, and
    emits them as input items immediately before the assistant
    message they belong to.  Position is tracked by ASSISTANT
    ORDINAL (not raw index) — ``sanitize_messages`` drops orphan
    tool results and inserts synthesized error tool messages, but
    NEVER drops or duplicates assistant messages, so the n-th
    assistant in the original list is invariably the n-th in the
    sanitized list.  Index-based lookup would have silently
    misrouted reasoning attachments after any tool-message repair.
  - ``extract_reasoning_text`` walks ``type=="reasoning"`` items and
    returns ``summary[*].text`` + ``content[*].text`` concatenation.
  - ``_reasoning_item_for_input`` projects a stored item into
    ``ResponseReasoningItemParam`` shape (drops server-only
    ``status``).  Returns ``None`` when ``id`` is missing or non-
    string per the SDK ``Required[str]`` schema; caller skips
    appending, preventing malformed input items from reaching the API.

* ``OpenAIChatCompletionsProvider`` gains:
  - ``extract_reasoning_text`` walks synthetic
    ``type=="reasoning_text"`` blocks and returns the concatenated
    text directly (no underlying provider semantics — the synth
    block IS the surface).

* ``ChatSession`` gains:
  - ``_resolve_server_type(alias)`` reads ``server_compat.server_type``
    from the active model's capabilities dict.
  - ``_maybe_synth_reasoning_block(provider_blocks, reasoning_parts)``
    creates the synthetic ``reasoning_text`` block when no native
    blocks were emitted but reasoning was captured.  Wired at the
    end of ``_stream_response`` immediately before the
    ``_provider_content`` stamp.

* ``history_decoration.py`` dispatcher collapses three near-identical
  lazy-init singleton getters (one per recognised block type) into a
  single ``_BLOCK_TYPE_PROVIDER_FACTORY`` dict + helper.  Adding a
  fourth provider becomes a one-line dict entry.

* Constants hoist: ``MAX_REASONING_DISPLAY_BYTES = 64 * 1024`` moved
  from three sibling provider modules into ``_protocol.py`` so a
  tuning change propagates uniformly to every provider's display path.

Cross-provider safety

The synthetic ``reasoning_text`` block type is intentionally NOT in
``ANTHROPIC_VALID_BLOCK_TYPES`` (Phase 2 constant).  Cross-model
resumption (operator switches from a local model to Anthropic mid-
workstream) falls through Phase 2's shape filter cleanly to the
text+tool_calls rebuild path rather than reaching Anthropic with a
malformed block.  Pinned by ``test_synthetic_block_falls_through_
anthropic_shape_filter``.

Same protection applies in reverse: OpenAI Responses
``type=="reasoning"`` items reaching Anthropic mid-workstream fail
the shape filter and rebuild from text+tool_calls.

Tests (49 net new tests)

* ``tests/test_provider_openai_responses_reasoning.py`` (21 tests):
  - Extractor unit tests: empty/none/no-reasoning/single/mixed/
    truncation/malformed/non-list (8).
  - ``_reasoning_item_for_input`` projection (4 tests including the
    new None-on-missing-id guard).
  - ``_build_kwargs`` include= gating: flag+capability/flag-false/
    capability-false/default-omits (4).
  - ``_convert_messages`` reasoning round-trip: emit-before-assistant/
    drop-on-replay-false/foreign-shape-skipped/default-replay-false (5).

* ``tests/test_session_synth_reasoning_block.py`` (23 tests):
  - ``_maybe_synth_reasoning_block`` direct unit tests (6).
  - Cross-provider safety regression — synthetic block falls through
    Anthropic shape filter (2).
  - ``OpenAIChatCompletionsProvider.extract_reasoning_text`` for the
    new synthetic block type (6).
  - ``_resolve_server_type`` direct unit tests (5).
  - ``_stream_response`` integration tests driving fake reasoning-
    emitting streams through the actual session method (3 tests
    — added in response to a code-review finding that pinned the
    wire-up at session.py needs an integration test).

* ``tests/test_session_replay_reasoning.py`` extended with 4
  ``TestSessionToOpenAIResponsesBoundaryIntegration`` tests driving
  ``session._try_stream`` -> real ``OpenAIResponsesProvider`` ->
  captured ``client.responses.create`` SDK boundary call.  Negative-
  tested: temporarily reverting the ``include=`` step in
  ``_build_kwargs`` makes ``test_replay_true_adds_include_to_
  responses_request`` fail; restoring makes it pass.

* ``tests/test_history_decoration.py`` extended with the new
  ``reasoning_text`` dispatcher branch test, and the Phase 1 stub
  test for the OpenAI Responses dispatcher branch was tightened
  (it now asserts real text extraction instead of the empty-string
  stub).

* ``tests/test_provider_anthropic_reasoning.py`` had its Phase 1
  ``OpenAIResponses returns "" for reasoning blocks`` stub test
  retitled and updated to assert the real Phase 3 behaviour.

Code-review pass

Multi-stage ``/review`` pipeline (4 finders + verify + dedupe) ran
on this diff.  6 findings (1 major, 3 minor, 2 nit), 0 critical, 0
security, 0 performance.  All applied:

* MAJOR (bug-1+bug-4+q-1): ``_convert_messages`` enumerate-index
  lookup was unsound under ``sanitize_messages`` length changes.
  Fixed by switching to assistant-ordinal-keyed lookup.
* MINOR (q-2+q-3): ``_MAX_REASONING_DISPLAY_BYTES`` duplicated
  across three provider modules + declared after first use.
  Fixed by hoisting to ``_protocol.py``.
* MINOR (q-4): three near-identical singleton getters in dispatcher.
  Fixed by collapsing to ``_BLOCK_TYPE_PROVIDER_FACTORY`` dict.
* MINOR (q-5): ``_maybe_synth_reasoning_block`` wire-up not pinned
  by integration test.  Fixed by adding three
  ``TestStreamResponseSynthBlockIntegration`` tests.
* NIT (bug-2): ``_reasoning_item_for_input`` fell back to ``id=""``;
  fixed to return ``None`` on missing/non-string id.
* NIT (q-6): four naming variants for the same concept; renamed
  ``_convert_messages`` kwarg to match the operator-flag name.
* REFUTED (bug-3): SDK distinguishes summary vs content as separate
  fields; no double-counting concern.

Briefing departures

The briefing's Phase 3 plan grouped Gemini with OpenAI Responses on
the assumption that Gemini reasoning had its own native shape (like
Anthropic's ``thinking``).  The spike confirmed Gemini-via-OpenAI-
compat is path-3 (Chat Completions shape, no native reasoning
items).  Phase 3+4 merger handles Gemini for free via the synthetic
``reasoning_text`` block — same mechanism used for vLLM and
llama.cpp.  Whether Gemini's specific endpoint actually emits
``reasoning_content`` deltas is server-dependent and not yet
empirically verified; capture is best-effort (server-emission-driven,
no flag gate).

The briefing's Phase 4 plan stamped reasoning as Anthropic-shaped
``thinking`` blocks ``{type: "thinking", thinking: <text>}``.  This
PR uses a distinct ``{type: "reasoning_text", text, source?}`` shape
to avoid a cross-model resumption hazard the briefing missed: an
unsigned synthetic Anthropic-shape block reaching Anthropic's wire
would 400 the API.  The distinct shape falls through Phase 2's shape
filter cleanly without needing signature validation in the filter.

Lint + test gate

* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6110 passed (3 deselected).  Phase 3+4
  added 49 net new tests.
2026-05-09 17:23:55 -07:00
Patrick Buckley 00bd80a658 feat(reasoning): wire-build shape filter + replay flag (Phase 2)
Make ``replay_reasoning_to_model=False`` actually suppress prior-turn
thinking blocks on the Anthropic wire (Phase 1 stored the operator
flag but the wire path always re-sent ``_provider_content``
verbatim).  As a side benefit, close a pre-existing latent bug where
foreign-shaped ``_provider_content`` (e.g. an OpenAI Responses
``type="reasoning"`` block reaching Anthropic on a mid-workstream
model switch, post-Phase-3) would have 400'd the API.

Why now: Phase 1 shipped the operator knob and UI rehydration but
the wire payload still always carried thinking blocks for
Anthropic-with-thinking turns.  Operators flipping replay=False saw
no behaviour change on the actual API call -- the flag only affected
``/history`` rendering.  Phase 2 closes that gap.

What this change does

* ``ANTHROPIC_VALID_BLOCK_TYPES`` (frozenset of 8 block types
  Anthropic's input boundary accepts) and
  ``ANTHROPIC_REASONING_BLOCK_TYPES`` (the strip subset) added at
  the top of ``_anthropic.py``.  The strip set is intentionally
  narrow: ``{"thinking", "redacted_thinking"}`` -- ``tool_use`` /
  ``server_tool_use`` / ``web_search_tool_result`` (which carry
  web-search ``encrypted_content``) MUST survive for round-trip
  continuity, and a regression test pins this.
* ``_convert_messages`` signature gains
  ``replay_reasoning_to_model: bool = True`` (back-compat default
  -- production call sites pass the resolved value explicitly).
  The verbatim ``_provider_content`` replay path is now wrapped by
  a shape-validity check using ``ANTHROPIC_VALID_BLOCK_TYPES``;
  foreign-shaped payloads fall through to the existing text+
  tool_calls rebuild path rather than reaching the API.  When
  shape is valid AND replay=False, a list comprehension drops
  thinking blocks from ``wire_blocks`` while preserving
  tool_use / web_search blocks.  When all blocks are stripped
  (message had only thinking, no text or tool_calls), the message
  also falls through to the rebuild path -- which silently skips
  if both content and tool_calls are empty (correct: stripped
  reasoning has nothing to replay).
* Orphan-tool detection still walks the ORIGINAL ``provider_content``
  (not ``wire_blocks``) so the strip cannot accidentally lose the
  source-of-truth tool_use IDs.  The implementation comment pins
  this invariant.
* Protocol surface grows the kwarg on both ``create_streaming`` and
  ``create_completion``.  ``OpenAIChatCompletionsProvider``,
  ``OpenAIResponsesProvider``, and ``GoogleProvider`` (via
  inheritance) accept the kwarg and ignore it -- they have no
  first-class reasoning shape on the wire today.  Phase 3 will use
  it on the OpenAI Responses adapter to gate
  ``include=["reasoning.encrypted_content"]``.
* ``ChatSession._resolve_replay_reasoning_to_model(alias)`` reads
  ``ModelConfig.replay_reasoning_to_model`` from the registry,
  defaulting to ``False`` on lookup failure (the conservative
  miss-fallback: replaying reasoning text against an unknown
  operator preference is worse than missing the strip).  Threaded
  into the three production call sites:
  ``ChatSession._try_stream`` (streaming), ``_utility_completion``
  (title gen / compaction / extraction), and the agent provider
  call site (plan / task agents).

Token calibration deferred to Phase 4

The briefing's optional Phase 2 step (extending ``_msg_text_chars``
to count ``_provider_content`` bytes that survive the strip)
required either invasive flag-threading through every call site
of the static method or a lossy approximation that picked the wrong
direction for the default case.  Per the briefing's ``pick a
phase'' guidance, this is bumped to Phase 4.  The pre-existing
silent under-count on Anthropic-thinking turns persists when
replay=True.  Strip-when-False naturally fixes the under-count by
keeping the bytes off the wire entirely; the residual case is the
opt-in replay path.

Tests (28 new, all driving through real boundary objects)

* ``tests/test_provider_anthropic_replay.py`` (19 tests):
  - Strip vs preserve under both flag values (3 tests including
    redacted_thinking).
  - Default-kwarg back-compat preserves verbatim replay (1 test).
  - Web-search tool_use + server_tool_use + web_search_tool_result
    survive strip with encrypted_content intact (2 tests, edge 14).
  - Orphan-tool synthesis after strip -- pins the
    ``provider_content`` source-of-truth read at lines 397-433
    (1 test).
  - Foreign-shape fallthrough: OpenAI ``type="reasoning"`` block
    rebuilds via text+tool_calls (1 test).
  - Mixed-shape fallthrough: even one foreign block forces
    rebuild (1 test).
  - Empty / None / non-list ``_provider_content`` fallthrough
    (3 tests).
  - Legacy Anthropic-thinking row pre-Phase-2 stays in verbatim
    path -- no regression on existing conversations (2 tests).
  - All-blocks-stripped fallthrough behaviour: rebuild from text
    if available, silently skip if not (2 tests).
  - Constants pinning: strip set is narrow, valid set includes
    web search, strip is subset of valid (3 tests).
* ``tests/test_session_replay_reasoning.py`` (12 tests):
  - Resolver: 6 tests covering miss / default / set / explicit /
    fallback alias / exception.
  - Streaming call site: 3 tests pinning the kwarg propagates
    through ``_try_stream`` to a stub provider.
  - Non-streaming call site: 1 test pinning
    ``_utility_completion`` propagates the flag.
  - End-to-end boundary integration: 2 tests driving
    ``_try_stream`` -> real ``AnthropicProvider`` -> captured
    Anthropic SDK ``client.messages.stream`` boundary, asserting
    on the ACTUAL wire payload shape.  Negative-tested:
    temporarily reverting the kwarg-thread at
    ``_anthropic.py:create_streaming`` makes the wire test fail
    with ``Strip predicate did not fire at wire boundary``;
    restoring makes it pass.

The boundary integration tests were added in response to a code
review finding that the bare-stub call-site tests would not catch
a regression where the provider stops reading the kwarg or
``_convert_messages`` silently drops the strip.  The integration
tests close that gap by inspecting what reaches the (mocked) SDK,
not just what the provider was called with.

Lint + test gate

* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6061 passed (3 deselected).  Phase 2
  added 28 net new tests.
2026-05-09 17:23:55 -07:00
Patrick Buckley 47df9d23c5 feat(reasoning): persist reasoning text on history payload (Phase 1)
Surface stored Anthropic thinking blocks on /history responses so
refreshing the page rehydrates the reasoning bubble. Wire payloads
unchanged. Per-model operator knobs added to model_definitions for
both UI rehydration and (Phase 2) wire-build replay.

Why now: reasoning is already round-tripped via _provider_content for
Anthropic-with-thinking turns, but never surfaces on the history wire,
so a tab reload showed only the final answer with no rationale.
Operators also have no per-model lever to opt out of UI display or to
opt in to replay-to-model on subsequent calls.

What this change does

* Migration 052 adds two boolean columns to model_definitions:
  persist_reasoning (default 1) controls UI rehydration; replay_
  reasoning_to_model (default 0) reserved for Phase 2's wire-build
  shape filter. Mirrors the enabled column pattern (NOT NULL +
  integer server_default).
* LLMProvider Protocol gains extract_reasoning_text(provider_blocks)
  with concrete impls on AnthropicProvider (walks type=='thinking'
  blocks, joins with newline, caps at 64 KiB) and no-op stubs on
  OpenAIChatCompletionsProvider + OpenAIResponsesProvider. Google
  inherits the no-op via OpenAIChat. Phase 3 will wire the OpenAI
  Responses extractor once include=['reasoning.encrypted_content']
  is requested.
* turnstone.core.history_decoration gains a structural dispatcher
  extract_reasoning_text_from_provider_content keyed off the first
  block's type field (Anthropic 'thinking' / OpenAI Responses
  'reasoning' / Gemini 'thought' are non-overlapping by API design).
  Both history surfaces use it: _build_history calls the dispatcher
  directly (the SSE-replay path builds entry dicts from scratch),
  and the lifted make_history_handler runs the list-helper variant
  in the existing to_thread block.
* make_history_handler resolves persist_reasoning via three tiers:
  live session -> workstream_config.model_alias (the same key
  SessionManager uses to rehydrate the original model after process
  restart) -> conservative True default. Operator flag-flip takes
  effect uniformly on both warm and cold workstreams.
* Frontend: app.js replayHistory and coordinator.js role==='assistant'
  branch each call the existing reasoning-bubble construction (for
  app.js, the document.createElement pattern from the live SSE
  handler; for coord, the appendMsg('reasoning') helper) when
  msg.reasoning is non-empty. Reasoning bubbles render before the
  content bubble, matching live SSE order.
* Admin UI: two checkboxes ('Persist reasoning', 'Replay reasoning
  to model') in the model edit modal, plus override-pill display in
  the model row when set to non-default values.

What is intentionally out of scope

* Phase 2 -- ANTHROPIC_VALID_BLOCK_TYPES shape filter at
  _anthropic.py:312-316, _convert_messages replay_reasoning_to_model
  parameter, thinking-strip branch, _msg_text_chars token-calibration
  extension. The replay flag is stored but not consumed on the wire.
* Phase 3 -- OpenAI Responses include=['reasoning.encrypted_content'],
  Gemini include_thoughts spike, ModelCapabilities.supports_
  reasoning_replay.
* Phase 4 -- Local-model / chat-template reasoning persistence
  (session.py:3486 reasoning_parts accumulator).

Tests

* AnthropicProvider.extract_reasoning_text -- 13 unit tests covering
  None / empty / mixed / multi-block / cap / malformed / non-list
  inputs plus other-provider no-op verification (real provider
  instances, no mocks).
* extract_reasoning_for_history -- 10 dispatcher tests including
  block-type discriminator routing (thinking vs reasoning vs
  unknown), strip-when-flag-false, empty / non-dict guards, and
  cross-role isolation.
* _build_history -- 6 boundary tests through the real Anthropic
  extractor with stub sessions, including the registry-lookup
  failure default-True branch.
* make_history_handler -- 5 round-trip tests through real storage:
  the storage layer's reconstruct_messages decodes provider_data
  into _provider_content, and the helper extracts through the real
  AnthropicProvider. Includes the live-session flag honoring path,
  the cold-workstream workstream_config lookup path, and the
  no-alias default-True fallback path.
* Audit-log discipline -- 4 structural mock-and-assert tests that
  capture every Logger.info / warning / error call across the
  pipeline (extractor, dispatcher, list-helper, _build_history)
  and assert no captured payload contains a marker reasoning string.
* model_definitions storage -- 6 round-trip tests: default flags,
  explicit create with both flags, individual update of each flag,
  and list-includes-flags assertion.
* model_registry -- 4 tests: dataclass defaults, dataclass with
  explicit flags, DB-row-mapping with both flags, and pre-052
  legacy-row default-fallback.

Edge cases pinned by the test suite

* Pre-052 DB rows missing the new columns degrade to dataclass
  defaults (test_db_reasoning_flags_default_when_absent).
* Live session in memory has its flag honored (test_history_handler_
  with_persist_flag_false_via_live_session).
* Cold workstream resolves the flag via workstream_config +
  app.state.registry (test_history_handler_cold_workstream_resolves_
  via_workstream_config) -- this closes the gap where a process
  restart would have silently un-honored an operator flag-flip.
* Cold workstream without persisted model_alias falls through to
  default True (test_history_handler_cold_workstream_no_alias_
  defaults_true).
* Foreign / unknown / missing block types degrade silently to no
  reasoning field rather than misroute or crash.

Lint + test gate

* ruff check + ruff format -- clean.
* mypy -- no issues across all 191 source files.
* pytest -m 'not live' -- 6030 passed (3 deselected).
2026-05-09 17:23:55 -07:00
Patrick Buckley 16fc7efce2 style(sse): align comments with always-advance seq invariant
Doc-debt cleanup flagged by /review on 9dc29db7. The cap+seq fix
flipped the seq-advance rule but left two doc sites describing the
old "incremented only on actual append" shape — exactly the buggy
invariant the previous commit removed. Future readers trusting the
stale docs would be one wrong assumption away from re-introducing
the silent-drop bug.

Updates the field-init comment block and the docstring on
register_listener_with_in_progress_snapshot (which sits at the
snap_seq capture site, so its contract is consumer-facing).

Also drops the now-dead `seq: int = 0` initializer in
on_reasoning_token and on_content_token — under the new shape, the
unconditional `seq = self._ws_inflight_seq` inside the lock makes
the initializer unreachable. Was load-bearing under the old
else-branch; harmless now but signals "some path leaves seq at 0"
to a reader.
2026-05-09 17:23:55 -07:00
Patrick Buckley e8eca2ec9b fix(sse): always advance _ws_inflight_seq on emit, even past cap
Copilot caught a real bug in the cap+seq interaction: the previous
shape only advanced ``_ws_inflight_seq`` when the buffer actually
appended, on the theory that "every _seq corresponds to a buffered
fragment" was a useful invariant. It wasn't — once the buffer hit
its cap, seq stalled at the high-water-pre-cap, so a subscriber that
registered AFTER the cap was hit would capture
``snap_seq == stalled_seq``, and every subsequent live token (also
tagged with the stalled seq) would be filter-dropped by the events
handler's ``seq <= snap_seq`` dedup. Silent loss of the entire
post-cap stream for refresh-past-cap tabs.

Fix: advance seq on every emit, regardless of buffer cap. The cap
is a buffer-size limit, not a stop-streaming signal. Past-cap tokens
are absent from the snapshot's text payload (the buffer was
truncated at cap) but the live stream past them is now correctly
delivered — refresh-after-cap renders snapshot-up-to-cap then live
tokens past it, with a visual gap equal to the past-cap chunk and
no silent drop of subsequent tokens.

Test ``test_inflight_seq_increments_only_on_actual_append`` enforced
the buggy invariant and is renamed/flipped to
``test_inflight_seq_advances_on_every_emit_even_at_cap``. Added
``test_subscriber_after_cap_hit_receives_subsequent_tokens`` (and
the reasoning equivalent) as direct regressions for the
silent-token-loss scenario.
2026-05-09 17:23:55 -07:00
Patrick Buckley 57563b0c12 docs(sse): document state_change + in_progress_snapshot events
Updates the docs that describe the per-workstream SSE event stream and
the SessionUI lifecycle to match the refresh-resume changes:

- api-reference.md: documented the `state_change` event (previously
  undocumented despite already being a live event) and the new
  `in_progress_snapshot` event; rewrote the multi-consumer fan-out
  paragraph to mention the kind-specific replay tail (state_change +
  optional in_progress_snapshot) so the "no catch-up needed" claim
  is no longer misleading.
- architecture.md: bumped the SessionUI Protocol stub to 16 methods
  (added `on_turn_start` / `on_turn_committed`) and pointed at the
  in_progress_snapshot section in the API reference.
- sdk.md: added rows for `state_change`, `in_progress_snapshot`, and
  `approval_resolved` (preexisting gap) to the per-workstream event
  table.
- coordinator-api-tour.md: added an `in_progress_snapshot` row to the
  event table and rewrote the reconnection-contract paragraph to
  cover mid-stream content/reasoning restoration.
- diagrams/04-conversation-turn.puml: added `on_turn_start()` before
  the thinking-start emit and `on_turn_committed()` immediately after
  `messages.append(assistant_msg)`, with notes explaining the inflight-
  buffer reset semantics. PNG regenerated.
2026-05-09 17:23:55 -07:00
Patrick Buckley 43e622840b feat(sse): refresh-resume for mid-stream page reloads
Refreshing a coordinator or interactive workstream pane while the LLM
is mid-stream now restores the partial assistant text + reasoning
immediately and flips the composer back to stop-mode, instead of
showing nothing until the response completes.

Per-turn inflight buffers (`_ws_inflight_content`, `_ws_inflight_reasoning`,
`_ws_inflight_seq`) on `SessionUIBase` are kept separate from the
existing multi-turn `_ws_turn_content` buffer that drives the
dashboard's IDLE-piggyback payload. New `on_turn_start` (top of
send-loop, defensive) and `on_turn_committed` (right after
`messages.append(assistant_msg)`, primary) lifecycle hooks reset
inflight at turn boundaries. The seq counter is monotonic across
turns so a long-lived subscriber's `snap_seq` cutoff stays valid for
the lifetime of the connection — resetting per-turn would silently
drop turn N+1's first M tokens (M = whatever was streamed pre-snapshot
in turn N).

`snapshot_and_consume_state_payload` also drains inflight at idle/error
so cancel and exception paths don't leak stale text. New
`register_listener_with_in_progress_snapshot` atomically registers a
listener and snapshots the inflight buffers; `make_events_handler`
emits a `state_change` event (so the JS busy machine flips to
stop-mode) followed by a one-shot `in_progress_snapshot` after the
kind-specific replay, then strips the internal `_seq` field from
yielded live events while filtering against `snap_seq`. A per-listener
shallow `dict` copy in the live drain prevents the multi-tab race
where one listener's `del event["_seq"]` would corrupt another
listener's filter view.

`_synthesize_cancelled_results` now emits synthetic `on_tool_result`
events for each cancelled tool so live coord tabs can drop the
newly-additive `coord-tool-batch--running` indicator cleanly. The
indicator now coexists with `--auto`/`--approved` (applied on
`tool_info` and `approval_resolved` approved; removed when every row
in the batch has a result), making live tool execution visually
parallel to the replay-time orphan rendering.

Frontend handlers in `app.js` (interactive) and `coordinator.js` (coord)
absorb EventSource auto-reconnect re-replays via a length-based
prefix check on the in-progress buffer. New `InProgressSnapshotEvent`
+ `StateChangeEvent` dataclasses in the Python and TypeScript SDKs
with type guards.

`_MAX_TURN_CONTENT_CHARS` lifted 256 KiB → 512 KiB (single constant
for both buffers — headroom for current commercial models).

Regression tests cover race-free composition under concurrent writers,
seq-filter dedup invariants, the cross-turn seq monotonic invariant,
idle/error inflight drain, synthesized `on_tool_result` on cancel
(including UI-hook failure isolation), and the multi-listener
shared-dict invariant.
2026-05-09 17:23:55 -07:00
Patrick Buckley 584437b98e chore: bump version to 1.5.10 2026-05-08 15:18:46 -07:00
Patrick Buckley 3abe1c0058 fix(skills): apply Copilot review feedback on PR #495
Two findings, both confirmed against the source:

1. Migration 051's downgrade rewrote every '[]' row back to '{}',
   which would (a) destroy operator-written empty arrays and
   (b) reintroduce the known-invalid sentinel that every consumer
   rejects. Pre-migration '{}' rows and operator-authored '[]' rows
   are indistinguishable after upgrade — there is no clean inverse
   for the data state. Made downgrade an explicit no-op with the
   rationale documented inline; '[]' is the correct shape under any
   consumer's interpretation, so leaving the data untouched on
   downgrade is strictly safer than reversing it. Updated the
   module docstring to call this out.

2. admin_update_skill's notify_on_complete validator short-circuited
   on empty string: `if nc and nc != "[]":` skipped the JSON-parse
   branch when nc=="" and persisted the empty string straight to
   storage, leaving a non-JSON value behind. Folded the empty case
   into the existing "{}" coercion so any blank/whitespace/legacy
   value normalises to "[]" before the array-validation gate.

Tests: three new regressions in TestSkillAPI — empty-string
normalises, "{}" sentinel coerces, non-array JSON 400s. The third
locks in the array-only validator that the previous "valid JSON"
gate would have accepted.
2026-05-08 15:18:10 -07:00
Patrick Buckley e24d8b9597 fix(skills): notify_on_complete default is "[]" not "{}"
Every consumer of prompt_templates.notify_on_complete treats it as a
JSON-array string (the admin form's array editor, the JSON.isArray
validator in submitEditTemplate, _validate_notify_targets in
server.py, the documented "list of channel/contact identifiers"
shape). But the column's server_default — set in migration 011 and
inherited through 021's lift into prompt_templates — has been "{}"
(an empty JSON object) since day one.

Newly-installed remote skills inherit the schema default, so every
unlock-then-edit flow trips the array validator on the inherited
"{}" and the request never leaves the browser. The user-visible
symptom was "click Save, nothing happens"; the latent symptom was
silent shape divergence between every install and every operator-
authored skill.

Migration 051: rewrites every legacy "{}" row to "[]". Operator-
edited values (anything that's neither "{}" nor NULL) are left
intact. Downgrade restores "{}" only on rows still holding the
post-migration "[]" so any later operator edits stick.

Server-side defaults flipped to "[]" in the same PR so new rows
land correct without depending on the column's server_default:

- _schema.py prompt_templates.notify_on_complete server_default
- StorageBackend protocol create_prompt_template kwarg
- sqlite + postgres create_prompt_template kwargs
- console_schemas.py SkillCreateRequest / SkillUpdateRequest /
  SkillInfo Pydantic defaults
- core/session.py ChatSession._notify_on_complete initial value
- server.py initial-message worker fallback when skill_data omits
  the field

admin_update_skill validator now also rejects non-array JSON (was
"valid JSON" only — would have accepted "{}" or "{\"a\": 1}").
_skill_to_response coerces legacy "{}" rows to "[]" on read so the
admin UI sees a consistent shape even before migration 051 runs.
The frontend's `tmpl.notify_on_complete || "[]"` fallback already
handled empty-string but not "{}" — the read-side coercion makes
it moot.
2026-05-08 15:18:10 -07:00
Patrick Buckley e11e6f6b70 fix(ui): aria-atomic on modal error elements + drop stale inline display
Designer-review follow-up to the .is-visible sweep. With role=alert
+ aria-live=assertive, AT engines re-announce when the element's
text content changes — but without aria-atomic some engines only
read the diff between old and new content. With aria-atomic=true
the entire updated message is read each time, which matters when a
validation error is replaced by a server error on retry (or
vice-versa).

Added aria-atomic=true to all 24 modal error elements (every
role=alert with aria-live=assertive). Same accessibility uplift
across the board — no per-modal exceptions.

Also dropped the stale `style="display: none"` attribute from the
three MCP error elements (mcp-create-error, mcp-import-error,
mcp-install-error). The CSS rule

  .admin-modal [role="alert"] { display: none; }

already hides them by default — the inline attribute was redundant
and would have overridden the .is-visible toggle if the class-based
contract is ever changed.
2026-05-08 15:18:10 -07:00
Patrick Buckley b8d728fd79 fix(ui): convert remaining modal-error toggles to .is-visible class
Sweep of the latent bug PR #494 fixed for the skill modals: the
project's CSS contract for modal errors is

  .admin-modal [role="alert"]            { display: none; }
  .admin-modal [role="alert"].is-visible { display: block; }

…but ~30 sites across governance.js and admin.js were toggling
`style.display = ""` instead of the .is-visible class. The "show"
side broke silently — clearing the inline style fell back to the
CSS `display: none` so the error never rendered, and any
validation failure looked like an unresponsive button.

Mechanical conversion of every show/hide site for these modal
error elements:

  governance.js
    create-role-error, edit-role-error
    create-policy-error, edit-policy-error
    github-import-error
    cpp-error, epp-error  (custom + eval prompt policies)
    create-hr-error, edit-hr-error  (heuristic rules)
    create-ogp-error, edit-ogp-error  (output-guard patterns)

  admin.js
    mcp-create-error, mcp-import-error, mcp-install-error

Plus the global `_showModalError` helper in admin.js — its
`style.display = "block"` happened to work today (inline display
beats the CSS rule), but normalising it to .is-visible keeps every
modal on a single canonical path. The five modals that route their
show side through that helper (create-user, create-token,
create-channel, create-schedule, edit-schedule) had their hide
sides converted in lockstep.

Added a comment on `_showModalError` documenting the contract so
the next contributor doesn't reintroduce the bug.

Out of scope: model-create-error (already canonical), home-coord-error
(not in .admin-modal), edit/create-template-error (fixed in #494).
No CSS or HTML changes; behaviour-equivalent for hide sides; show
sides go from broken-silent-no-render to correct-render-with-AT-
announcement.
2026-05-08 15:18:10 -07:00
Patrick Buckley 2138c19821 fix(skills-ui): clear prior error at submit-start so it doesn't go stale
Once edit-template-error is actually visible (the visibility fix in
this same PR), a stale error now persists across resubmit cycles:
the user sees a red message, fixes the input, clicks Save, the
validator passes, the PUT goes out — and the previous error stays
on-screen the whole time, only clearing when the modal closes on
success.

Fix at the start of submitEditTemplate / submitCreateTemplate:
clear .is-visible AND empty textContent. Cheaper than tracking
every validator branch and every .catch path; a fresh submit is a
clean slate.
2026-05-08 15:18:10 -07:00
Patrick Buckley 627bf06ced fix(skills-ui): show validation errors via .is-visible, not style.display
Smoke-testing the unlock flow surfaced a latent bug: clicking Save
on the edit-skill modal silently no-op'd whenever the
notify-on-complete field had non-JSON content. The error div was
DOM-correct (text content set, role=alert, aria-live=assertive),
but invisible — because the project's modal-error CSS contract is:

  .admin-modal [role="alert"]              { display: none; }
  .admin-modal [role="alert"].is-visible   { display: block; }

…and the JS in submitEditTemplate / submitCreateTemplate was
clearing the inline `display: none` via `el.style.display = ""`.
That falls back to the CSS rule, which still says `display: none`,
so the error never rendered. The user saw no error and the click
felt unresponsive (compounded by the early-return before the
disabled-state reset, which also made Save look broken).

Fixed both skill-modal flows (create + edit) by toggling the
canonical `.is-visible` class instead. Six sites in governance.js:
the two early-return show paths, the two .catch show paths, and
the two modal-open hide-resets.

Scope note: this same bug pattern exists in ~20 other modal error
sites across governance.js and admin.js (create-role, edit-role,
create-policy, edit-policy, github-import, cpp, epp, create-hr,
edit-hr, create-ogp, edit-ogp, mcp-create, mcp-import, mcp-install,
plus admin.js sites that don't go through _showModalError). All
pre-existing, broken silently for who knows how long. Out of scope
for this PR — recommend a follow-up sweep that also normalises
_showModalError's `style.display = "block"` to the same convention.
2026-05-08 15:18:10 -07:00
Patrick Buckley b67da0f48a fix(skills): apply designer review on lock-icon UX
Designer review of the cb5fa1b lock-icon iteration flagged five
items; four are addressed here, one was a deliberate trade-off
documented below.

- Glyph hardening (#2): the lock character is now 🔒︎ — U+1F512 with
  the U+FE0E text variation selector — paired with the existing
  font-variant-emoji: text rule. font-variant-emoji shipped late
  and isn't universal yet (Chrome 131+, Safari 16.4+, Firefox 132+);
  the explicit text VS is belt-and-braces so older Chromium / most
  Linux don't fall back to a coloured emoji that would clash with
  the monochrome instrument-panel aesthetic.
- Accent-line de-conflict (#3): top:14px → 18px so the lock button
  sits below the modal's ::before accent-line decoration's visual
  band rather than competing with it horizontally. h2's
  padding-right reservation (44px) still gives the title clearance.
- Mobile touch target (#4): @media (max-width: 700px) bumps the
  button to 44×44 (WCAG 2.5.5 / Apple HIG / Material minimum) and
  shifts it to top:8px right:8px, with h2 padding-right widened to
  56px to match.
- Keyboard discoverability (#6): on readonly open, focus lands on
  the lock button instead of Cancel. Keyboard users hit the unlock
  affordance immediately instead of having to Tab past every
  disabled spec input to reach it. Cancel is one Shift-Tab away.

Deferred:
- (#1) Reviewer flagged top-right placement as risking confusion
  with the universal × close-button convention. Keeping the
  icon-only design per product direction; the bordered chip styling
  + accent-coloured hover make it visually distinct from the
  thin-stroke unbordered × pattern, and the confirm dialog catches
  any misclick safely.
- (#5) Optional empty-corner indicator after unlock — the
  "Customized from upstream" badge text already carries the signal;
  not adding new chrome.
2026-05-08 15:18:10 -07:00
Patrick Buckley e05b6adc67 fix(skills): unlock UX — lock icon top-right, save reset, confirm z-index
Three issues from manual smoke-testing the unlock flow:

1. Confirm dialog rendered behind the edit-skill modal. Both
   overlays sat at z-index 600, and confirm-overlay is earlier in
   the DOM than edit-template-overlay — so DOM order put the parent
   modal on top of its own confirm. Bumped confirm-overlay to 650
   (still below toasts at 700) since confirm dialogs are launched
   FROM other overlays and need to sit above them.

2. Save button stayed disabled (or non-functional) after unlock.
   submitEditTemplate disables etm-submit on click and re-enables in
   .finally, but a stale disabled=true survives the mutate-in-place
   re-render that runs after unlock. Always reset
   submitBtn.disabled = false in showEditTemplateModal so the
   re-render path can never inherit a stuck disabled state.

3. UX redesign — moved the unlock affordance from a "Customize…"
   button at the bottom of the footer to a 🔒 icon button at the
   top-right of the modal. The lock glyph is the universal "this is
   locked, click to unlock" affordance and reads more clearly than
   a footer button next to Cancel/Save. font-variant-emoji: text
   keeps it monochrome on browsers that support it (instrument-panel
   aesthetic) with graceful fallback to coloured emoji elsewhere.
   admin-modal-skill h2 reserves padding-right so a long title can
   never collide with the absolute-positioned button.

Cleanup: removed the now-unused .modal-secondary and
.modal-buttons-spacer rules; the bottom etm-unlock button + flex
spacer are gone from the modal footer.
2026-05-08 15:18:10 -07:00
Patrick Buckley 40355d8303 fix(skills): match readonly column int idiom in postgres unlock_skill
Copilot caught that prompt_templates.readonly is an Integer column
(_schema.py: sa.Column("readonly", sa.Integer, nullable=False,
server_default="0")) and create_prompt_template stores it as 1/0,
but unlock_skill in the postgres backend was passing a Python bool
(readonly=False). The sqlite impl already uses 0; this aligns the
two backends and matches the 0/1 idiom used for the sibling flag
columns (is_default, auto_approve, enabled).

The other Copilot findings on this PR (loadGovSkills race, NBSP
double-space, list_skill_versions O(history_size), ignored
set_skill_readonly return value + None re-read) were all closed by
the prior review-feedback commit (eea795d): the snapshot+flip is
now an atomic unlock_skill() that uses SELECT MAX(version)+1
internally, the handler guards both the unlock_skill return and the
post-flip get_prompt_template re-read, the JS chains
showEditTemplateModal off loadGovSkills's promise, and the badge
NBSP matches the sibling pattern.
2026-05-08 15:18:10 -07:00
Patrick Buckley 2cbd926b9d fix(skills): apply review feedback on unlock action
Code review caught a race + a missing None guard; designer review
caught a window.confirm regression and a button-hierarchy issue.

Backend:
- Race fix (bug-2): replace set_skill_readonly+create_skill_version
  with a single atomic unlock_skill(template_id, snapshot, changed_by)
  -> int|None on the storage protocol (sqlite + postgres). Snapshot
  insert + readonly flip happen in one transaction; the next version
  number is computed via SELECT MAX(version)+1 inside the txn rather
  than len(list)+1 outside, closing the (skill_id, version)
  collision window where two concurrent admin actions could both pick
  the same version.
- None guard (bug-3): check the post-flip get_prompt_template re-read;
  return 404 instead of letting _skill_to_response(None) raise.
- Audit body: also record snapshot_version, and harden None-vs-empty
  with `or ""` on the existing.get(...) calls.

Frontend:
- D-1: replace window.confirm with the existing showConfirmModal
  (admin.js:2350) — themed dialog, focus-trap, can render the source
  URL with consistent typography. The native dialog could collapse
  the multi-paragraph copy depending on browser.
- D-2: mutate-in-place on success rather than hide → reload → reopen.
  loadGovSkills now returns its fetch promise so unlockSkill can
  chain showEditTemplateModal after the cache refresh — no flicker,
  no focus bounce, and it kills bug-1 (the reopen was reading stale
  _govSkills before loadGovSkills resolved). showEditTemplateModal
  is idempotent when already open: it skips the trigger-element
  capture and the focus-trap reinstall.
- D-3: button hierarchy. Drop flex:1 from .modal-secondary so the
  Save button keeps a stable width whether or not Customize is
  rendered; insert a flex-spacer between Customize and Save so the
  destructive-ish detach groups left next to Cancel and the primary
  action floats right.
- D-4: NBSP normalized to match the existing   escape pattern
  on the sibling badge line (was an actual NBSP byte).
- D-5: success toast now reads "Skill unlocked — fields are now
  editable" so the operator gets a positive affirmation that the
  edit affordance is live.
- D-10: aria-describedby="etm-origin-badge" on disabled spec inputs
  so screen-reader users get the same "this came from upstream"
  context that sighted users see in the cyan badge.

Tests: + test_unlock_skill_versions_after_existing_history seeds an
out-of-order version (3) and asserts unlock picks 4, defending
against the len()-based version computation regressing.
2026-05-08 15:18:10 -07:00
Patrick Buckley 40c0ab5a58 feat(skills): unlock action lets operators customize installed skills
skills.sh / GitHub installs land with readonly=True so admins can only
tune runtime config (model, temperature, etc.); the SKILL.md spec is
locked. In practice, upstream skills aren't always tuned for turnstone,
so locking the spec adds friction without a real safety win — every
edit is audited and version-snapshotted regardless.

This adds an explicit unlock so the boundary stays visible (multi-user
audit trail benefits from a discrete event, vs. silently dropping the
gate). Behaviour:

- POST /v1/api/admin/skills/{id}/unlock — flips readonly=False on a
  readonly row. Snapshots the pre-unlock state into skill_versions so
  the upstream-pristine version is recoverable from the History tab.
  Records skill.unlock audit with {name, source_url, origin}. 400 on
  already-unlocked, 404 on missing.
- origin stays "source" after unlock so the UI keeps a "Customized
  from upstream" provenance badge — the readonly flag is the gate, the
  origin field is the lineage.
- Storage: dedicated set_skill_readonly writer on the protocol +
  sqlite + postgres backends. readonly is intentionally absent from
  SKILL_MUTABLE so the generic update path can't piggyback on a
  provenance flip — the dedicated writer pattern matches what's
  already used for set_mcp_oauth_client_secret_ct.
- Frontend: "Customize…" button in the edit modal (visible only when
  readonly), with a confirm dialog explaining the upstream-detach.
  Once unlocked the existing edit-skill flow handles spec edits with
  no other changes. Origin badge updates to show "Customized from"
  the upstream URL when a source-origin row is unlocked.

Tests cover: unlock flips readonly + persists, pre-unlock snapshot
written to skill_versions, 400 on already-unlocked, 404 on missing,
post-unlock PUT can edit name/content/description (the readonly gate
no longer fires).
2026-05-08 15:18:10 -07:00
Patrick Buckley a5e8b8b17e fix(skills): apply PR #491 review feedback (size cap + dedup + conflict mapping)
Three issues caught by Copilot on the initial PR:

1. SKILL.md size cap was measured in code points, not UTF-8 bytes.
   `len(str)` is a *lower* bound on encoded byte length — multi-byte
   chars (emoji, CJK) inflate up to 4×, so a 100k-emoji SKILL.md
   (400KB encoded) would slip past the 256KB cap. Switch to
   `len(contents.encode("utf-8"))` and surface lone-surrogate failures
   as SkillSourceError instead of dropping them silently. New
   regression test feeds emoji content.

2. _skills_sh_source_url did not normalize the skill_id, so a sloppy
   id from `/api/search` (whitespace, surrounding slashes) would pass
   `_split_skills_sh_id`'s charset check (which strips first) and
   produce a malformed persisted source_url that broke the
   discover-UI dedup contract. Strip the id inside the helper, and
   reconstruct the canonical id from validated parts in
   download_skill's listing so downstream callers never see the raw
   input.

3. The catch-all `except Exception:` around create_prompt_template
   relabeled every storage failure (DB connection, disk full,
   permission errors) as "conflict", masking operational issues.
   Translate IntegrityError → StorageConflictError at the storage
   shim (matching the pattern already used for OIDC user
   provisioning) in both sqlite and postgres backends, then catch
   StorageConflictError specifically in the install handler. Real
   conflicts → "conflict" + warning; other exceptions → new
   "internal error" reason + log.exception.

Tests: +3 (oversized multibyte SKILL.md, source_url normalization,
storage-layer conflict translation). 226 passing.
2026-05-08 15:18:10 -07:00
Patrick Buckley bed691c62a fix(skills): switch skills.sh install to /api/download endpoint
The skills.sh install path was failing with 404s because their public
API surface changed: /api/skills/{id} is gone, replaced by
/api/skill/[owner]/[repo]/[skill] (auth-walled) and
/api/download/[owner]/[repo]/[skill] (unauthenticated, returns the
SKILL.md + bundled resources inline as JSON). The error was not
surfacing in logs because admin_skill_install had a silent
`except Exception:` around create_prompt_template that relabeled every
storage failure as "conflict" with no log entry.

- Replace SkillsShClient.resolve_github_url with download_skill that
  hits /api/download/{owner}/{repo}/{skill} and returns a SkillPackage
  directly. No GitHub round-trip; no rate-limit surface.
- Add _split_skills_sh_id with strict per-segment charset validation
  ([A-Za-z0-9._-]+) so URL-hostile content can't produce a malformed
  request or divergent persisted source_url.
- Use len(contents) instead of len(contents.encode("utf-8",
  errors="ignore")) for the SKILL.md size cap — errors='ignore' was
  silently dropping invalid units, making the cap bypassable.
- Extract _accept_resource(rel_path, byte_size) gate predicate; share
  it between download_skill and the GitHub _find_resource_files helper.
- Have search() derive a deterministic source_url from the skill id
  when /api/search omits one (which it currently always does), so the
  discover-UI "already installed" check matches what download_skill
  persists.
- Add structured logging across admin_skill_install and
  admin_skill_discover: a shared _log_install_failure helper for the
  four except branches (was four near-duplicate log calls with one
  drift), plus per-resource failure tallying — partial-resource
  installs now surface failed_resources in the response and audit
  record instead of silently committing the skill row with missing
  assets.

Tests: 7 new — empty/non-list files, oversized SKILL.md, resource
cap, non-text extension filtering, plus _split_skills_sh_id charset
rejection (whitespace, query chars). Verified end-to-end against
live skills.sh with tavily-search.
2026-05-08 15:18:10 -07:00
Patrick Buckley c930078f3d chore: bump version to 1.5.9 2026-05-07 22:56:06 -07:00
Patrick Buckley f148c4b423 fix: apply repair=False to all display-read load_messages call sites 2026-05-07 22:48:13 -07:00
Patrick Buckley 8ce4c8e737 chore: bump version to 1.5.8 2026-05-07 17:36:12 -07:00
Patrick Buckley 95e67dc768 fix(replay): apply PR #488 review findings
Four Copilot findings on c6041c6 — all confirmed valid, all bounded
to authenticated-user prompt-injection scenarios but worth closing
before merge.

Wrapper-detect bypass (string + list branches of
``_apply_reminders_for_provider``):

The round-2 fix used ``content.startswith("<tool_output>\\n")`` to
detect already-wrapped content and skip ``escape_wrapper_tags``.  A
tool whose RAW output starts with that prefix (e.g. ``echo
'<tool_output>'``) would match and have its escape skipped, letting
literal ``<tool_output>`` / ``<system-reminder>`` tags reach the model
and impersonate a system envelope.  Replace the prefix check with
``extract_advisories_from_tool_envelope(content) is not None`` —
parsing requires the open AND matching close tags AND a structurally
valid envelope, raising the bypass bar significantly.

Mirror fix in the list-content branch so a tool emitting an unmatched
envelope as a text part can't bypass the per-text-part escape.

``_build_history`` legitimate-envelope drop:

The list-content drop path previously removed any text part starting
with ``<tool_output>\\n``.  A tool that legitimately outputs a
well-formed envelope (documentation viewer, code analyzer demoing the
wrapper, an echo tool) would have that part silently disappear on
replay.  Tighten the drop heuristic to require BOTH ``cleaned_text ==
""`` AND at least one extracted advisory — the structural signature of
the injected ``wrap_tool_result("", advisories)`` carrier we produce
in ``session.py`` for list-typed tool output.  A legitimate envelope
has non-empty inner body or no advisory blocks and survives the
projection.

Empty advisory body:

``queue_message`` accepts any non-None text including ``""`` and
whitespace-only strings.  ``_classify_advisory`` would return a
``user_interjection`` advisory with empty / whitespace body, which
``replayAdvisoriesAfterTool`` then renders as a featureless empty user
bubble.  Filter empty / whitespace-only bodies at classification time
so the wire-shape contract is uniform: no empty advisories ever ride
the wire.

Tests:

* ``test_apply_reminders_escapes_tool_output_starting_with_envelope_prefix``
  pins the structural-parser bypass close: a string starting with the
  envelope prefix but lacking a close tag still gets escaped.
* ``test_apply_reminders_escapes_list_text_part_with_unmatched_envelope_prefix``
  mirrors for the list-content branch.
* ``test_build_history_keeps_legitimate_envelope_text_part_with_body``
  pins that legitimate envelope output stays in the projected list.
* ``test_decorate_suppresses_empty_advisory_body`` and
  ``test_decorate_suppresses_whitespace_only_advisory_body`` pin the
  empty-body filter in ``_classify_advisory``.

Tests: 5923 passed, 3 deselected.  Lint + format + mypy clean.
(cherry picked from commit c2cb6a7ea5)
2026-05-07 17:35:23 -07:00
Patrick Buckley dc35cbc7bf fix(replay): seam 1 splice + storage symmetry for queued user messages
Reverses the seam-2-only design from the prior commits on this branch.
Queued user messages arriving DURING a tool batch (Seam 1) splice into
the last tool result's envelope as ``UserInterjection`` advisories via
``wrap_tool_result``.  Messages arriving BETWEEN turns (Seam 2) drain
as a single trailing user row via ``_flush_queued_messages`` with
``user_feedback`` (operator text alongside an approval, e.g. "y, use
full path") folded in as a prefix.  Cancel/exception drains (Seam 3)
keep the existing ``_flush_queued_messages()`` call unchanged.

Why all three seams:

* Strict-template providers (Mistral, Llama via vLLM with stock chat
  templates) reject role-alternation violations.  A literal ``user``
  row mid-tool-batch breaks ``assistant(tool_calls) → tool → ... →
  assistant``; back-to-back ``user → user`` rows on the wire also fail.
* The seam-2-only design produced back-to-back ``user`` whenever
  ``user_feedback`` and queued items both fired — bug-1 from the round-1
  review.  Folding ``user_feedback`` as a prefix to the queue-drain
  collapses the two into one row.
* During-batch arrivals couldn't ride seam 2 — the splice was the only
  way to deliver same-turn without violating role alternation.

Storage symmetry:

Tool DB rows now store the wrapped ``output`` (envelope + advisories)
unconditionally — ``self.messages[i]['content']`` and
``conversations.content`` match exactly.  List-typed output (image /
structured MCP results) uses ``wrap_tool_result(raw_joined_text,
advisories)`` at save time so the persisted string is anchored on
``<tool_output>\n`` for the replay parser.  ``TOOL_RESULT_STORAGE_CAP``
is removed entirely; tools are responsible for bounding their own
output, storage faithfully represents in-memory.  Removing the cap
also simplifies the parser — no truncated-envelope edge case.

Replay extraction:

``decorate_history_messages`` (REST ``/history``) and ``_build_history``
(SSE replay, resume, rewind, retry, post-load, rename re-replay) both
call the public ``extract_advisories_from_tool_envelope`` helper to
pull the envelope back into structured ``advisories`` for JS replay.
Both string content and list-typed content (image+queued-message
combo) covered.  JS renders extracted advisories as normal user
bubbles after the tool block via the shared ``replayAdvisoriesAfterTool``
helper in ``shared_static/utils.js``.

Wrapper-tag escape and provider splice:

``escape_wrapper_tags`` now encodes pre-existing ``&`` first using an
``&amp;`` sentinel so tool output containing literal entity strings
(documentation viewers, code analyzers, web scrapers returning entity-
encoded markup) round-trips correctly.  Both encode and decode helpers
short-circuit on absence of ``<`` / ``&``.

``_apply_reminders_for_provider`` detects already-wrapped content
(string body and list text-part) by ``startswith("<tool_output>\n")``
and skips re-escape so existing envelopes survive intact when a tool
message also carries ``_reminders`` (the queued-message + tool-error
co-occurrence case is now common).

``decorate_history_messages`` runs in ``asyncio.to_thread`` to keep
MB-scale string work off the event loop.

Other cleanup:

* ``_collect_advisories`` delegates the queue drain to a named helper
  ``_drain_queued_messages_to_advisories`` so the swap-and-clear pattern
  lives next to ``_flush_queued_messages``'s identical pattern and the
  side-effect is documented at the call site.
* Preamble strings + body marker for ``UserInterjection`` round-trip
  detection moved to module-level constants in ``tool_advisory.py``;
  imported by ``history_decoration.py`` so a producer-side rephrase
  can't silently desync the parser.
* ``_send_with_mocks`` ctxmgr extracted in ``test_session.py`` — the
  six new send-driven tests share an 8-deep ``patch.object`` block.
* ``replayAdvisoriesAfterTool`` shared helper in
  ``shared_static/utils.js``; ``app.js`` and ``coordinator.js`` both
  invoke it.
* Dead truncation-pill CSS removed (``.tool-output-truncated`` and
  ``.coord-tool-truncated``); the JS that added these elements went
  away with ``TOOL_RESULT_STORAGE_CAP``.
* Tautological tests (``TestBuildHistoryAdvisoryPropagation``)
  replaced with production-realistic round-trip tests built from
  ``wrap_tool_result(...)`` envelopes — REST and SSE-replay surfaces
  pinned to the same wire shape; full DB round-trip pinned end-to-end.

Negative-tested:

* Reverting the prefix-merge in ``_flush_queued_messages`` produces
  back-to-back ``user`` rows, breaking
  ``test_user_feedback_and_queued_coexistence_single_row_with_prefix``.
* Reverting the ``extract_advisories_from_tool_envelope`` call in
  ``_build_history``'s tool branch leaves the envelope verbatim in
  wire content, breaking the round-trip tests.
* Reverting the wrapper-detection in ``_apply_reminders_for_provider``
  entity-encodes the existing envelope's literal tags, breaking both
  the string-content and list-content envelope-preservation tests.
* Reverting the ``wrap_tool_result(raw_text, advisories)`` projection
  at the DB save site produces a string starting with the original
  raw text, breaking
  ``test_tool_db_row_round_trips_list_output_with_advisories``.

Tests: 5918 passed, 3 deselected.  Lint + format + mypy clean on
touched files.

(cherry picked from commit eca4bb79e4)
2026-05-07 17:35:23 -07:00
Patrick Buckley 79f4d0030d fix(replay): apply review findings q-2 through q-7
Round-1 ``/review`` apply-pass.  Drops stale ``UserInterjection``
references from comments and docstrings that no longer describe the
post-PR drain shape, asserts the two-stream invariant in the new
queued-message persistence test, and pins the ``content.trim()`` +
``renderAssistantToolBatch`` invariants on coord-side so a future
refactor can't silently regress the Qwen3 phantom-card fix or the
chronological-order render fix.

Deferred:

* **bug-1** (back-to-back ``user`` row when ``user_feedback`` from the
  approval-prompt UI callback coexists with a queued-message drain).
  Reachable on strict OpenAI-compatible local templates (Anthropic and
  Anthropic-via-merge-consecutive collapse fine; vLLM-hosted Mistral /
  Llama enforcing role alternation can reject).  The pre-PR splice
  guarded against this case by riding queued items inside the tool
  result envelope; that guard is what motivated the original
  UserInterjection design, so the fix lane needs a deliberate decision
  rather than a quick patch.  Sleeping on it.

* **q-1** (delete dead ``UserInterjection`` class + tests).  Held for
  the bug-1 decision — if the chosen fix is to resume the splice for
  the ``user_feedback``+queue coexistence case, the advisory shape
  stays load-bearing.  Class now carries a docstring note marking it
  retained-pending-decision so a passing reader doesn't grep for
  producers and assume it's actually dead.

Apply-pass content:

* ``q-2``: drop "queued user interjections" from the persistent-
  advisory parenthetical in ``send``'s tool-result loop comment;
  rewrite to point at ``_flush_queued_messages`` for the queue path.
* ``q-3``: ``__init__`` channel-routing comment loses "and
  ``UserInterjection``" — only ``GuardAdvisory`` remains.
* ``q-4``: ``_queue_tool_advisory`` docstring + the tool-error nudge
  comment lose the user-interjection mentions; the docstring also now
  describes the side-channel + ``_apply_reminders_for_provider``
  splice path (the actual mechanism).
* ``q-5``: ``AttachmentsNotQueueableError`` docstring rewritten to
  describe the post-PR ``_flush_queued_messages`` flow — the
  single-combined-turn ``\n\n``-join shape can't carry image / file
  blocks, and per-item separate user turns would expand the strict-
  template role-ordering surface that the post-batch drain already
  balances.
* ``q-6``: the new ``test_queued_message_persists_as_user_row_after_tool_batch``
  in ``test_session.py`` now asserts ``stream_idx == 2`` so a future
  regression where the post-batch flush runs but the send-loop short-
  circuits before the next iteration surfaces in CI rather than
  manual repro.
* ``q-7``: ``test_coordinator_page.py`` gets two new string-grep pins
  mirroring the existing ``test_app_js.py`` shape — ``content.trim()``
  on coord's assistant-replay branch and ``renderAssistantToolBatch``
  for the hoisted helper that orders content card before tool batch.

## Test plan

- [x] ``ruff check`` clean
- [x] ``mypy turnstone/`` clean (189 source files)
- [x] Affected test surface (``test_session.py`` +
  ``test_tool_advisory.py`` + ``test_app_js.py`` +
  ``test_coordinator_page.py``) — 240 passed

(cherry picked from commit a032e71ff3)
2026-05-07 17:35:23 -07:00
Patrick Buckley 14e504db1f fix(replay): coord render order + blank assistant cards + queued message persistence
Three independent rehydrate / replay regressions reported on long
multi-turn conversations after the pull-model wake stack landed.

**1. coord history replay rendered tool_calls above the assistant
narration that announced them.**

In ``coordinator.js``'s loadHistory loop, the ``role === "assistant"``
``tool_calls`` branch sat above the role switch — every assistant turn
with both narration AND tool dispatch produced ``[tool batch][content
card]`` in the DOM, even though chronological order is content first.
On a parallel fan-out (e.g. four ``close_workstream`` calls in one
turn) operators saw the assistant text "Let me close them out and
summarize" with NO tool batch between it and the next assistant
message — the four-row batch had been rendered above the announcing
text and was scrolled out of view.

Hoisted the ``tool_calls`` synthesis into a local
``renderAssistantToolBatch(m)``, called from inside the assistant
branch AFTER the content card.  Live SSE order (text → dispatch →
results) now matches replay order.

**2. Whitespace-only assistant content rendered as a blank card on
replay.**

Models with vLLM's ``--reasoning-parser`` (Qwen3 in production)
strip ``<think>…</think>`` and emit only the trailing ``"\n\n"`` as
``content`` before a tool call.  ``content_parts = ["\n\n"]`` saves
``content = "\n\n"`` to the conversations row.  Live the user only
sees ``.msg.reasoning`` (the thinking content) — the empty
``.msg.assistant`` card lives next to it but reads as a thin
divider.  On rehydrate the reasoning bubble is gone (not persisted)
and the empty assistant card is the only thing left, surfacing as
"blank cards where the assistant message was."

Both UIs now check ``content && content.trim()`` before rendering
the body — whitespace-only content skips the card entirely instead
of showing a phantom row.  Live render unchanged.

**3. Queued user messages disappeared on reconnect.**

PR #474 routed queued user messages into the tool-result envelope
via ``UserInterjection`` advisories — same-turn delivery, but no
persisted user row.  On page reload / cross-tab replay the
optimistic ``.msg-queued`` bubble vanished: there was no DB row to
rehydrate it.

Dropped the ``UserInterjection`` splice in ``_collect_advisories``;
the queue drains through ``_flush_queued_messages`` AFTER the tool
batch completes instead.  Sequence becomes
``assistant(tool_calls) → tool … tool → user(drained)``, which is
valid for Mistral and Anthropic strict role validators (the only
forbidden shape was user injected mid-batch BEFORE the tool result,
which this still avoids).  Persists a real user row → bubble survives
reconnect, and stays in the session's wire-side context window on
the next turn.

## Test plan

- [x] ``ruff check`` clean
- [x] ``mypy turnstone/`` clean (189 source files)
- [x] ``pytest -m "not live"`` — 5798 passed, 3 deselected
- [x] Updated ``test_collect_advisories_does_not_drain_queued_messages``
  (was pinning the old UserInterjection shape)
- [x] Added ``test_queued_message_persists_as_user_row_after_tool_batch``
  (drives ``send`` end-to-end with a queued message arriving during
  the tool batch; asserts the user row lands in self.messages AND
  hits ``save_message``)
- [x] Updated ``test_replay_history_renders_content_before_tool_block``
  to tolerate the new ``msg.content && msg.content.trim()`` guard
- [ ] Live browser pass on coord (close_workstream parallel fan-out
  rehydrates with the 4-row batch BETWEEN the announcing assistant
  text and the summary) and interactive (Qwen3 ``"\n\n"`` rows no
  longer paint blank cards on reload; queued bubble survives a tab
  refresh)

(cherry picked from commit c11692b327)
2026-05-07 17:35:23 -07:00
Patrick Buckley 0abe0cb77d fix(mcp): apply PR #489 review feedback + de-flake pool reuse 401 retry
PR #489 review feedback (Copilot + github-code-quality):
- closeSettingsPanel now closes nested revoke modal first on close-button
  path (Escape was already handled by the parent keydown trap deferring
  to the inner trap; missing-modal-on-close-button was an orphan-modal
  hazard).
- _refreshConsentBadge now updates the settings button's aria-label +
  title dynamically with the pending-consent count for screen readers
  (badge stays aria-hidden — the count is in the label).
- _MAX_INSUFFICIENT_SCOPE_REPORTED promoted to public
  MAX_INSUFFICIENT_SCOPE_REPORTED in mcp_http_parsers; drops cross-module
  private import in mcp_oauth's /start handler.
- Stale test comment in test_session_mcp_dispatch_error.py corrected:
  _exec_read_resource does not log with exc_info=True (bearer-leak
  invariant).
- Rejected the protocol-method ellipsis warning: rest of _protocol.py
  uses ... consistently per Protocol convention.

Lint:
- ruff format applied to test_mcp_pool_auth_integration.py and
  test_mcp_pool_auth_resource_integration.py (combined `with` grammar —
  pure formatting).

Flake fix — test_integration_pool_reuse_401_refresh_and_retry_succeeds
on Python 3.11 / resource-constrained CI:

Same cross-task scope hazard f6a3b66 fixed at the close side, surfacing
at the connect side. asyncio.wait_for at mcp_client.py:1206 wraps
streamablehttp_client.__aenter__ in a fresh asyncio.Task. That fresh
task enters anyio cancel scopes, completes, and dies. The eventual
stack.aclose() during eviction or auth_401 retry runs from a different
task and tries to exit scopes whose entering task is dead — anyio
raises RuntimeError, the wedged anyio state blocks the retry's stack
teardown + reconnect, and the call exceeds the 15s budget on slow
workers.

Fix: replace asyncio.wait_for with `async with asyncio.timeout(...)` so
the streamablehttp_client.__aenter__ runs in the dispatch task itself,
no fresh-task scope ownership. Aligns with invariant 18 (asyncio.timeout
not asyncio.wait_for for any SDK / AS / pool-loop await crossing anyio
scopes).

Static path (_connect_one) at lines 905 and 1000 deliberately retains
asyncio.wait_for — auth_type ∈ {none, static} is byte-identical
(invariant 1) and the narrow connect-once / no-eviction-then-reuse
pattern doesn't trigger the cross-task hazard. Anchor comments pin
both directions: a future migration there would break invariant 1; a
future revert at 1206 would re-introduce the flake.

The cited test is the symptom (non-deterministically times out under
load), not a structural gate (no deterministic asyncio.timeout
assertion exists). The comment block at line 1206 records this so a
maintainer who reverts and finds green on a fast machine doesn't
conclude the fix is unneeded.

Verified on Python 3.11.14 (/tmp/venv311) and 3.13.7 (.venv): ruff
format clean, ruff check clean, mypy clean. 368 unit tests + 30 pool
integration tests pass on both interpreters; the previously-flaky test
passed 20× in isolation on 3.11.

Multi-stage /review (4 finders × verify × dedupe): bug/security/perf
returned zero findings; quality returned 3 confirmed minor/nit items
all of which are applied here (q-1 anchor comments at 905+1000, q-2
symptom-vs-gate clarification at 1206, q-3 module-docstring sentence
in mcp_http_parsers).

(cherry picked from commit 4a3e3607be)
2026-05-07 17:35:23 -07:00
Patrick Buckley 610513398b feat(mcp): per-user MCP server consent UX (Phase 8)
Wires the structured-error envelopes produced by Phase 7b's pool
dispatcher (mcp_consent_required / mcp_insufficient_scope /
mcp_*_forbidden / mcp_token_undecryptable_key_unknown /
mcp_oauth_url_insecure) through to the user-facing dashboard, and
adds a per-user settings panel for managing MCP server consents.

Changes
- ``_dispatch_pool_sync`` and ``_dispatch_pool_resource_sync`` wrap
  structured-error string returns as ``RuntimeError(json_str)`` via
  ``_is_structured_error()`` so the session-layer ``except Exception``
  branch fires uniformly across tool / resource / prompt dispatchers
  (the prompt path's ``isinstance(result, str)`` shortcut works only
  because prompts return ``list[dict]`` on success). Without this,
  the consent UX silently does not render for tool / resource calls.
- ``_structured_error`` extended with an optional ``consent_url``
  field; ``_build_consent_url`` produces ``/v1/api/mcp/oauth/start``
  query strings (path-relative; the dashboard appends ``return_url``
  at click time). Wired to all 12 ``mcp_consent_required`` and the
  ``mcp_insufficient_scope`` emit sites.
- New endpoints ``GET /v1/api/mcp/oauth/connections`` and
  ``DELETE /v1/api/mcp/oauth/connections/{server_name}`` registered
  on both ``turnstone-server`` and ``turnstone-console``. The DELETE
  handler runs local delete + audit + 204 first, then schedules the
  RFC 7009 upstream revoke as a fire-and-forget ``asyncio.create_task``
  with strong-ref tracking via ``_revoke_upstream_tasks`` (mirrors
  the ``_pg_refresh_drain_tasks`` pattern). Soft cap of 256 concurrent
  in-flight revokes prevents pile-up under coordinated mass-revoke;
  the audit detail records ``upstream_revoke_outcome`` as
  ``scheduled | no_refresh_token | no_http_client | shed_by_cap``.
- ``ASMetadata`` extended with ``revocation_endpoint`` parsed from
  RFC 8414 metadata. ``revoke_token_at_as`` helper posts the form
  body under ``asyncio.timeout`` (not ``asyncio.wait_for``) and
  never raises; ``_attempt_upstream_revoke`` is wrapped in an outer
  ``try/except Exception`` so unhandled exceptions don't surface as
  ``Task exception was never retrieved``.
- ``/v1/api/mcp/oauth/start`` accepts an optional ``scopes=`` query
  param; tokens are validated against RFC 6749 §3.3 grammar via
  ``is_valid_scope_token`` (promoted to ``mcp_http_parsers``),
  capped at ``_MAX_INSUFFICIENT_SCOPE_REPORTED`` (32), and unioned
  with the configured server scopes for the step-up consent flow.
- Storage primitive ``list_mcp_user_token_metadata_by_user`` projects
  the metadata columns at the SQL boundary so ciphertext blobs never
  cross the wire on the settings-list path. New
  ``MCPUserTokenMetadataRow`` TypedDict in ``_protocol.py``;
  ``MCPTokenStore.list_user_token_metadata`` re-types to the existing
  ``MCPUserTokenMetadata`` shape.
- Dashboard renderer (``app.js``): ``tryParseMcpError`` detects the
  envelope shape on ``tool_result`` SSE events with ``is_error=True``
  and ``buildMcpErrorEmbed`` renders an action card mirroring the
  existing ``buildMediaEmbed`` pattern. Three categories: actionable
  (consent_required / insufficient_scope) with a ``Connect`` button
  that opens ``/v1/api/mcp/oauth/start`` in a popup with a scheme
  guard, forbidden (mcp_*_forbidden) with a static notice, operator
  (key-mismatch / url-insecure) with an operator-action notice.
- New gear button in the appbar opens an MCP-connections settings
  modal driven by ``loadMcpConnections`` / ``confirmRevokeMcp``
  (two-step revoke confirmation matching the existing delete-ws
  pattern). Pending-consent badge tracks unresolved consent prompts
  in this tab; cleared after the connections list returns. Console
  proxy collision-checked: the IIFE only prepends a node-id pill to
  ``header.firstChild``, so the right-anchored gear button is safe.

Bearer-leak invariant
- No ``exc_info=True`` on any new path that can carry a chained
  ``httpx.Request`` (revoke handler, dispatch sites, exec sites).
  The two pre-existing ``exc_info=True`` calls in
  ``_exec_read_resource`` / ``_exec_use_prompt`` were replaced with
  structured-field logs as a Phase 8 sibling fix.

Tests
- 440 pytest passes on both Python 3.13 (.venv) and 3.11
  (/tmp/venv311); ruff + mypy clean.
- 5 new test files: ``test_mcp_consent_url_sibling_audit`` (structural
  gate that every ``code="mcp_consent_required"`` / ``mcp_insufficient_scope``
  site carries ``consent_url=``), ``test_mcp_oauth_connections``,
  ``test_mcp_oauth_revoke``, ``test_mcp_token_store_metadata``,
  ``test_session_mcp_dispatch_error``.
- End-to-end regression coverage for the bug-1 sibling pattern:
  ``test_call_tool_sync_raises_on_structured_error_envelope``,
  ``test_read_resource_sync_raises_on_structured_error_envelope``,
  ``test_get_prompt_sync_raises_on_structured_error_envelope``, plus
  ``test_call_tool_sync_does_not_wrap_non_structured_string`` as the
  defensive gate (only ``mcp_*`` envelopes are wrapped).

Hard invariants honored
- Static path byte-identical for ``auth_type ∈ {none, static}``: the
  wrap fires only when the dispatcher returns a structured-mcp-error
  string, which only happens on the oauth_user pool path.
- ``asyncio.timeout`` (not ``asyncio.wait_for``) on every new
  AS / SDK / pool-loop await per Python 3.11 anyio cancel-scope
  hazard.
- Scope cap ``_MAX_INSUFFICIENT_SCOPE_REPORTED = 32`` enforced at
  every output / merge site.
- Cross-user isolation on the revoke endpoint: a non-owner DELETE
  returns 404 with the same body shape as a never-existed row;
  ``http_client_mock.post.assert_not_called()`` pins this in 3 tests.

Deferred (not Phase 8 blockers)
- perf-2 (``asyncio.gather`` parallelisation in revoke handler) —
  superseded by perf-1's fire-and-forget pattern.
- q-4 (prompt-path ``isinstance(str)`` vs sibling ``_is_structured_error``
  asymmetry) — already documented in the function docstring.
- q-9 (``_pendingConsentServers`` → ``_serversNeedingConsent``
  rename) — pure naming taste.

(cherry picked from commit 5a3f46a1fa)
2026-05-07 17:35:23 -07:00
Patrick Buckley 2d6519f9a8 fix(storage): sanitize NUL bytes on _source + _reminders columns
Apply sanitize_text() to the new _source and _reminders columns in
both save_message and save_messages_bulk on SQLite + PostgreSQL,
mirroring the existing pattern used for content and provider_data.

Producers (sanitize_payload on the watch dispatch path,
format_nudge constants on the standard nudge path) already strip
NUL bytes today so nothing in production reaches this clamp — but
the storage layer is opaque to those invariants, and PostgreSQL
TEXT columns reject NUL outright.  Without this clamp, a future
producer that forgets sanitize_payload (or hand-builds the column
string) hard-fails the chat-loop persist path on PostgreSQL.

Cost is negligible — sanitize_text early-exits on the common
no-NUL case via 'if value and "\x00" in value'.

Surfaced by Copilot's PR #486 review.

(cherry picked from commit fc8bd6ca33)
2026-05-07 17:35:23 -07:00
Patrick Buckley a99ce49311 revert(memory): drop dormant limit kwarg from load_messages
Closes round-2 review finding q-7 (nit).

The kwarg was added to close round-1 perf-2 cosmetically — the
storage backend's signature already accepted ``limit``, but the
single in-tree caller (``ChatSession.resume``) doesn't pass it and
other tail-load consumers go direct to ``storage.load_messages``.
Adding signature surface to mark a perf finding closed without an
actual consumer is API-surface bloat.

When a tail-load consumer is written (e.g. a heuristic in
``session.resume`` to skip ancient wake rows), the kwarg can come
back — at that point with a real caller driving the contract.

(cherry picked from commit 14af6f464e)
2026-05-07 17:35:22 -07:00
Patrick Buckley 46f3571c93 refactor(watch): rename _WATCH_REMINDER_OPTIONAL_KEYS public + hoist import
Closes round-2 review findings q-6 (nit) and perf-1 (nit).

* **q-6:** ``_WATCH_REMINDER_OPTIONAL_KEYS`` carried a leading
  underscore (Python's module-private convention) but was imported
  from two other modules — clearly a public contract between
  ``build_watch_reminder`` and its consumers
  (``ChatSession._dispatch`` + ``server._build_history``).  Drop the
  underscore so the import sites match the constant's documented
  cross-module role.

* **perf-1:** The dispatch closure imported the constant inside its
  body, paying ``IMPORT_NAME`` + ``IMPORT_FROM`` bytecode on every
  watch fire.  ``server.py`` already imports at module scope; hoist
  the same way in ``session.py``.  Microsecond savings per dispatch,
  but the in-closure form was just an oversight from the apply-pass.

(cherry picked from commit 668da26dce)
2026-05-07 17:35:22 -07:00
Patrick Buckley 53cabe7e20 fix(session): trim tombstone refs + WHAT-narration in apply-pass comments
Closes round-2 review findings q-1 (minor), q-3 (nit), q-4 (nit), q-5
(nit).

* **q-1:** Drop the ``post-migration 050`` clause from the fork-block
  comment — the apply-pass relocated rather than removed the
  tombstone-style temporal reference round-1 q-2 was supposed to fix.
  The bulk-row dict shape and ``_encode_reminders`` are
  self-explanatory; the WHY is pinned by
  ``test_fork_preserves_source_and_reminders``.

* **q-3:** Replace ``DOES persist now`` framing on the wake-row save
  comment with a present-tense invariant.  The ``now`` implies the
  reader knows the prior state, same family as the temporal
  tombstones.

* **q-4:** Trim the 12-line WHAT-narration block above the
  resume-time ``_reminders_delivered = True`` loop to two lines
  stating the WHY only.  The new regression test pins the contract.

* **q-5:** Reframe ``test_fork_preserves_source_and_reminders``
  docstring as a forward-looking invariant; drop the
  ``Dropping them was the original bug`` and ``post-migration 050``
  fix-narration.

Project convention: invariant statements, present tense; don't
reference the current task / fix / migration number.

(cherry picked from commit b120ee2fd7)
2026-05-07 17:35:22 -07:00
Patrick Buckley 1d9fd94e23 fix(session): byte-clamp REMINDER_TEXT_STORAGE_CAP + drop local-only doc citation
Closes round-2 review findings bug-1 (minor) and q-2 (minor).

* **bug-1:** ``_encode_reminders`` clamped each entry's ``text`` field
  with Python ``str`` slicing, which counts codepoints.  Multi-byte
  UTF-8 input (CJK, emoji) could land 4 bytes per character past the
  cap, defeating the row-width / FTS5-index protection by up to 4x.
  Switch to UTF-8 byte clamping with ``errors="ignore"`` on the
  decode boundary so a slice mid-codepoint drops the partial
  character cleanly.

* **q-2:** Both the constant block-comment and the ``_encode_reminders``
  docstring referenced ``docs/design/watch-card-ux-briefing.md`` —
  local-only per project convention (``feedback_no_design_doc_commits``)
  so the canonical repo reads as a dead reference.  The cap value
  stands by itself; the row-width / FTS5 WHY is enough.

(cherry picked from commit 779ec638a5)
2026-05-07 17:35:22 -07:00
Patrick Buckley eb89ddab1e fix(metacog): cleanup batch — share watch-key constant, sanitize metadata, drop tombstones
Closes round-1 review findings q-2 (minor), q-5 (minor), q-6 (nit), q-7
(nit), sec-1 (nit), perf-4 (nit).

* **q-5:** Export ``_WATCH_REMINDER_OPTIONAL_KEYS`` from
  ``turnstone/core/watch.py`` and import in the dispatch closure
  (session.py) and the replay filter (server.py:_build_history).  The
  three-place duplication of the literal tuple
  ``("watch_name", "command", "poll_count", "max_polls", "is_final")``
  is gone; future field adds touch one constant.

* **sec-1:** Run ``sanitize_payload`` over string-typed metadata fields
  (``watch_name`` / ``command``) before they enter the queue.  Today's
  consumers all use ``textContent``, but the asymmetry — sanitised
  ``text`` alongside unsanitised metadata — would survive forever in
  DB rows and resurface if a future consumer used a non-textContent
  sink (aria-label, copy-to-clipboard, markdown render).

* **q-7:** Drop the per-iteration ``isinstance(reminder, dict)`` from
  the dispatch closure's metadata comprehension.  By the time the
  block runs, ``text = reminder.get("text", "") if isinstance(...)``
  + the ``if not sanitized: return`` guard above already established
  ``reminder`` is a non-empty dict.

* **q-2:** Strip tombstone-style references — "post-#482", "post-#484",
  "Step 7 of the watch-card UX plan", "Post-Step-7 dispatch surface",
  and the brittle line-anchor "session.py:2685-2686" — across
  ``session.py``, ``test_session.py``, ``test_watch.py``,
  ``test_watch_dispatch.py``, ``test_watch_integration.py``.  Comment
  intent preserved; historical anchors gone.

* **q-6:** Drop the ``del source`` line in ``cli.py``'s
  ``on_user_reminder``; the parallel ``on_tool_reminder`` ignores
  ``tool_call_id`` without ``del`` and the comment alone is enough.

* **perf-4:** Document the SQLite ``render_as_batch=True`` recreate
  cost in migration 050's docstring — first deployment after upgrade
  copies the conversations table twice (one per ``add_column``).
  PostgreSQL is unaffected.

5734 non-live tests pass; ruff + mypy clean.

(cherry picked from commit 7e35050b68)
2026-05-07 17:35:22 -07:00
Patrick Buckley eb92e61755 fix(ui): wrap interactive reminder spans in .msg-body + exclude system-nudge from anchor lookup
Closes round-1 review findings q-3 + q-4 (minor, merged) and bug-3 + bug-4
(nit, merged).

* **q-3 + q-4:** The new ``.msg.user-reminder .msg-body { white-space:
  pre-wrap }`` rule was a no-op on the interactive UI because that
  frontend's ``_buildDefaultReminderBubble`` appended label + text spans
  directly to the outer ``.msg.user-reminder`` element with no
  ``.msg-body`` wrapper.  Coord rendered the same shape with a wrapper.
  The two implementations diverging on DOM structure also meant a
  shared-helper extraction was harder than necessary.  Reconciled by
  wrapping interactive's spans in ``.msg-body`` to match coord; the CSS
  rule now applies to both UIs and the shared-extraction follow-up to
  ``shared_static/cards.js`` is mechanical (deferred per the review
  report — out of scope for this commit).

* **bug-3 + bug-4:** The reminder anchor lookup ``.msg.user`` also
  matched ``.msg.user.system-nudge`` markers because the marker carries
  both classes.  A non-wake reminder fired between a wake marker and
  the next real user message would anchor below the wake marker rather
  than the previous real user message.  Edge case (``/history`` reload
  corrects), but the fix is mechanical: change the selector to
  ``.msg.user:not(.system-nudge)`` in both files.

(cherry picked from commit 869135d97a)
2026-05-07 17:35:22 -07:00
Patrick Buckley 0e2ea122eb fix(memory): wire limit kwarg through load_messages
Closes round-1 review finding perf-2 (minor).

Storage backends accept ``*, limit: int | None = None`` (see
:meth:`StorageBackend.load_messages` at storage/_protocol.py:146) but
the in-memory wrapper at memory.py:82-85 dropped the kwarg, so
callers that wanted to tail-load (e.g. ``session.resume`` against a
long-running coord with hundreds of wake rows + persisted reminder
JSON) were forced to pull every row through the wrapper anyway.

Wraparound is mechanical: signature widens, default leaves existing
callers unaffected.

(cherry picked from commit 885f6a9185)
2026-05-07 17:35:22 -07:00
Patrick Buckley eb9dd2402a fix(session): delete stale 'reminders stay in-memory' comment
Closes round-1 review finding q-1 (major).

The comment block above ``self._attach_pending_user_reminders(user_msg)``
asserted that reminders "stay in-memory only and don't persist across
reloads" — directly contradicted by the comment block immediately below
(at the save_message call site) that explains the new persistence
semantics, plus the actual code that now writes ``_source`` and
``_reminders`` to the conversations row.  Future readers hitting both
blocks would lose trust in the surrounding comments.

The lower block already documents the persistence contract, so the
upper block is just deleted rather than rewritten.

(cherry picked from commit 81502c962f)
2026-05-07 17:35:22 -07:00
Patrick Buckley 0c58910c4b fix(session): preserve _source/_reminders on fork + cap persisted reminder text
Closes round-1 review findings bug-2 (major), perf-1 (minor), perf-6 (nit).

* **bug-2:** ``ChatSession.resume(..., fork=True)``'s bulk-row builder
  silently dropped the ``_source`` and ``_reminders`` side-channel
  data the source workstream had persisted via ``_append_user_turn``.
  Both backends' ``save_messages_bulk`` already accept these keys
  (the columns exist post-migration 050) — the bulk builder just
  didn't supply them.  The fork's resumed transcript would then look
  like the assistant turn answered out of nowhere: every wake marker
  and every reminder bubble that survived to disk on the source got
  dropped on the fork.  New regression test
  ``test_fork_preserves_source_and_reminders`` pins the contract.

* **perf-6:** Extracts ``_encode_reminders(reminders) -> str | None``
  near ``_apply_reminders_for_provider`` so the user-turn save path,
  the tool-turn save path, and the new fork bulk builder share one
  encoder.  Eliminates the drift risk between three near-identical
  ``json.dumps(..., separators=(",", ":")) if X else None`` patterns.

* **perf-1:** The new helper clamps each entry's ``text`` field at
  ``REMINDER_TEXT_STORAGE_CAP = 8192`` characters before encoding so
  a single rogue producer (a watch streaming unbounded shell output,
  a corruption-class steering payload) can't blow the conversations
  row width or the FTS5 index.  The in-memory side-channel keeps the
  full body — only the persisted JSON is clamped.  Mirrors
  ``TOOL_RESULT_STORAGE_CAP`` on tool result rows.

5734 non-live tests pass; ruff + mypy clean.

(cherry picked from commit 91e7f2daca)
2026-05-07 17:35:22 -07:00
Patrick Buckley 2e393d76b4 fix(session): flag persisted reminders delivered on resume
Persisted ``_reminders`` survive ``load_messages`` but the in-memory
``_reminders_delivered`` flag does not (it's session-scoped — set by
``_mark_reminders_delivered`` after each successful provider stream,
never persisted alongside the JSON column).  Without a re-splice
guard at resume time, ``_apply_reminders_for_provider`` would walk
every loaded message, see ``_reminders`` set + the flag falsy, and
splice every historical ``<system-reminder>`` envelope onto the wire
on the very next user turn — leaking each reminder a second time, the
turn after it had already advised.

Mirror the post-stream hook in ``resume()``: every loaded message
that carries reminders has already been delivered (it survived to
disk), so flag it accordingly so ``_apply_reminders_for_provider``
short-circuits on the pass-through path.

Test pins the contract end-to-end — stage a workstream with a
persisted reminder, resume into a fresh session, append a live user
turn, run the wire transform, and assert the historical reminder
body does NOT land in the rendered output.

(cherry picked from commit f1466ca7e3)
2026-05-07 17:35:22 -07:00
Patrick Buckley dec175f176 feat(ui): structured watch-result card + system-nudge marker on replay
User-visible slice of the watch-card UX workstream — combines the
replay-path widening, both frontend renderers, the CSS, and the
cross-cutting Python tests.

server._build_history widens the reminder filter from {type, text} to
project on a known set of optional fields (watch_name, command,
poll_count, max_polls, is_final) and surfaces _source as
entry["source"] when set.  The known-key filter narrows the blast
radius if a future producer accidentally stuffs sensitive fields
into the dict.

SessionUIBase.on_user_reminder takes a new source: str | None kwarg
that rides on the SSE event when set.  _attach_pending_user_reminders
forwards user_msg["_source"] so non-originating tabs see the wake's
"system_nudge" tag and render the thin marker.  Protocol + cli + eval
implementations widen accordingly.

Frontend (coordinator.js + app.js — touched in lockstep per project
memory's "logic that lands in BOTH UIs must touch both files"):
* Branch on r.type === "watch_triggered" for a structured
  .msg.watch-result card with header / $ command / <pre> body /
  poll N/M [· final] footer.
* New addSystemNudgeMarker (interactive) + appendSystemNudgeMarker
  (coord) renders a thin .msg.user.system-nudge anchor for
  wake-driven reminders, both live (source === "system_nudge" on the
  SSE event) and replay (msg.source === "system_nudge").
* Default .msg.user-reminder rendering preserved for every other
  metacog nudge type.

CSS (shared_static/chat.css):
* New .msg.watch-result rules — full-width treatment, cyan accent,
  monospace body with word-break: break-word for mobile.
* New .msg.user.system-nudge rule — thin yellow marker.
* Bonus newline-collapse fix: .msg.user-reminder .msg-body now sets
  white-space: pre-wrap so multi-line shell output / bulleted lists
  stay readable inside the advisory bubble.

Plan reference: docs/design/watch-card-ux.md §4 Steps 9-12 + bonus
CSS §11 (Commit 4).

(cherry picked from commit 6ae6877acc)
2026-05-07 17:35:22 -07:00
Patrick Buckley 592433b46d feat(metacog): structured watch reminders carry watch metadata onto NudgeQueue
WatchRunner._dispatch_result now takes a structured reminder dict
produced by build_watch_reminder() — text matches format_watch_message
verbatim (so compaction / channel adapters / wire splice keep their
behaviour), and watch_name / command / poll_count / max_polls /
is_final ride alongside as queue-entry metadata.

The dispatch closure registered in ChatSession.set_watch_runner pulls
the optional fields out of the dict and passes them to enqueue via
the new metadata kwarg.  Drain seams already merge metadata into the
rendered reminder dict (Commit 2), so the SSE event for a watch fire
now carries the structured fields without further plumbing.

* turnstone/core/watch.py — new build_watch_reminder() helper, _poll_watch
  switches from format_watch_message + dispatch(str) to build_watch_reminder
  + dispatch(dict).  set_dispatch_fn / get_dispatch_fn / restore_fn
  signatures widen from Callable[[str, str], None] to
  Callable[[dict[str, Any], str], None].
* turnstone/core/session.py — dispatch closure builds the metadata dict
  via {k: reminder[k] for k in ("watch_name", "command", ...) if k in reminder}
  and passes it to nudge_queue.enqueue.
* tests/test_watch.py — new TestBuildWatchReminder class pinning the
  builder shape; existing dispatch_fn_registry / restore_fn tests
  updated to dict shape.
* tests/test_watch_dispatch.py — every dispatch(...) call updated to
  pass a structured reminder dict via _reminder() helper; new
  TestMetadataPropagation class pins the metadata-on-enqueue contract.
* tests/test_watch_integration.py — _dispatch_result calls updated to
  dict shape.

Plan reference: docs/design/watch-card-ux.md §4 Step 7 + Step 8 watch-test
subset (Commit 3).

(cherry picked from commit 13db19905a)
2026-05-07 17:35:22 -07:00
Patrick Buckley da5321eb88 refactor(metacog): widen NudgeQueue._Entry with optional metadata field
Producers (today only watch_triggered) can now attach a metadata dict
to a queued nudge so the rendered reminder dict on the user/tool side
carries fields beyond {type, text}.  Wire shape stays additive: the
SSE event picks up the optional fields when present, and producers
without metadata leave it None.

* _Entry grows from 4 fields to 5 — metadata: dict[str, Any] | None.
* enqueue accepts metadata=... as a kwarg.
* drain returns list[tuple[str, str, dict | None]] (was 2-tuples).
* pending stays narrow at (type, text) for legacy callers; new
  pending_with_metadata projects the third slot for tests that need
  to assert producer-specific fields.
* Three drain consumers in session.py — _collect_advisories,
  _attach_pending_user_reminders, deliver_wake_nudge_from_queue —
  unpack the new 3-tuple shape and merge metadata into each
  reminder dict.
* on_user_reminder / on_tool_reminder protocol signatures widen
  from list[dict[str, str]] to list[dict[str, Any]] across
  ChatSession.UI, SessionUIBase, CLI, eval harness.

Plan reference: docs/design/watch-card-ux.md §4 Step 6 + Step 8 _Entry
subset (Commit 2).

(cherry picked from commit 30b7e4dd24)
2026-05-07 17:35:22 -07:00
Patrick Buckley baa2214f96 feat(storage): persist _source + _reminders side-channels on conversations
Adds two TEXT-NULL columns to the conversations table so multi-tab /
multi-device replay sees the same metacognitive bubble shape the
originating tab saw live.  Until now, reminders lived only on the
in-memory ChatSession.messages dict, and the wake-driven empty user
turn was not persisted at all (skip at session.py:2685-2686) — a
second tab connecting via /history saw the assistant turn with no
preceding wake context, and missed every other tab's reminder
bubbles besides.

Single Alembic revision 050 (head was 049) adds:
  * conversations._source — today only "system_nudge" for wake rows
  * conversations._reminders — JSON-encoded reminder list

Both backends (sqlite + postgresql) thread the columns through
save_message / save_messages_bulk / load_messages.  reconstruct_messages
unpacks the row tuple as 9 elements (was 7), JSON-decoding _reminders
on the user AND tool branches with the same contextlib.suppress guard
the existing provider_data / tool_calls decode uses.  Tool-row
reminders ride the same column so tool_error / repeat replay shape
matches user-channel parity.

session.py:2685-2686 wake-row persist skip is dropped; _append_user_turn
JSON-encodes user_msg["_reminders"] and passes both source + reminders
to save_message.  The tool-message save site at session.py:3014-3020
mirrors with metacog_reminders.

Plan reference: docs/design/watch-card-ux.md §4 Steps 1-5 (Commit 1).

(cherry picked from commit f64c3e7b10)
2026-05-07 17:35:22 -07:00
Patrick Buckley 3b60a69e4f fix(console): atomic coord-subsystem commit + offload startup teardown
Address Copilot review feedback on PR #487:

1. **Atomic commit invariant**: ``_bootstrap_coord_subsystem`` previously
   stamped ``coord_mgr`` ~50 lines before the final ``coord_registry``
   commit, and started threads + subscriptions in between.  A concurrent
   dashboard request running through ``_require_coord_mgr`` during the
   runtime-bootstrap window could observe ``coord_mgr`` set with
   ``coord_registry`` still ``None`` and surface the misleading
   "Restart the console after adding a model definition" 503.

   Refactored to two phases: (a) build everything as locals, (b) start
   side-effects (StateWriter / observer / nudge watcher / child fan-out
   / cleanup thread), then atomic commit at the end with ``coord_mgr``
   stamped LAST.  The build-phase ``try/except`` rolls back any started
   side-effects from local handles before re-raising — no daemon thread
   or subscription leaks across retries, and ``app.state`` is never
   stamped on a partial failure.

2. **Class-attr cleanup symmetry**: ``_teardown_partial_coord_subsystem``
   now also clears ``ConsoleCoordinatorUI._coord_mgr`` /
   ``_collector`` / ``_console_metrics`` to match the lifespan shutdown
   path (server.py ~line 4629).  A failed bootstrap (or test teardown
   reuse) no longer leaks process-global pointers at a half-built
   subsystem.

3. **Lifespan startup offload**: the lifespan startup error path used
   to call ``_teardown_partial_coord_subsystem`` synchronously, which
   in turn calls ``StateWriter.shutdown(timeout=2.0)`` — a thread-join
   + sync DB writes that could block the event loop for up to 2s
   while the console is still coming up.  Wrapped the whole
   load-and-bootstrap in ``asyncio.to_thread`` via the new
   ``_load_and_bootstrap_coord_subsystem`` synchronous helper, so all
   blocking work (including any rollback) runs on a worker thread.
   Mirrors the pattern the regular lifespan shutdown (line ~4620) and
   the runtime CRUD-triggered path already use.

Tests:
- ``test_bootstrap_atomic_commit_no_partial_visibility``: a polling
  thread in tight loop watches ``coord_mgr`` / ``coord_registry``
  during a real bootstrap and asserts no observation has ``coord_mgr``
  set with ``coord_registry`` still ``None``.
- ``test_real_bootstrap_rolls_back_partial_state_on_side_effect_failure``:
  monkeypatches ``install_idle_nudge_watcher`` to raise mid-build,
  asserts ``app.state`` shows the clean fresh-install state and the
  builder-failure error string surfaces ``RuntimeError`` (not the
  stale "no models" boot-time message).

(cherry picked from commit c6b4dc26be)
2026-05-07 17:35:22 -07:00
Patrick Buckley 5d1213d3dc fix(console): bootstrap coord subsystem on first model add
A freshly-installed console with no model rows in the DB at boot
caught the ``ValueError`` from ``load_model_registry()`` in the
lifespan and skipped the entire coord subsystem build, leaving
``coord_mgr`` ``None``.  ``_refresh_coord_registry`` then bailed
out at ``existing is None`` rather than building the subsystem on
first model add — operators had to restart the console after
configuring their first model in the admin panel for the
"Coordinator subsystem not initialized" banner to clear.

Extract the lifespan's coord build into a reusable
``_bootstrap_coord_subsystem`` and add ``_maybe_bootstrap_coord_subsystem``
that runs as an ``asyncio.to_thread`` follow-on after every admin
model-CRUD endpoint (create/update/delete/reload).  The helper:

- fast-paths to a no-op when ``coord_mgr`` is already set;
- guards concurrent first-install attempts with
  ``_COORD_BOOTSTRAP_LOCK`` + double-checked re-test inside the lock;
- pre-computes config-derived integers BEFORE any thread starts so
  ``int(config_store.get(...))`` failures don't strand a started
  ``StateWriter`` daemon;
- stamps ``coord_state_writer`` to ``app.state`` immediately after
  ``.start()`` so the new ``_teardown_partial_coord_subsystem`` can
  shut it down on a partial failure (no thread leaks across retries);
- atomically commits ``coord_registry`` + clears
  ``coord_registry_error`` as the final step so callers can rely on
  the invariant ``coord_registry`` is set iff ``coord_mgr`` is set;
- replaces the stale boot-time "no model definitions" message with
  a builder-failure-specific diagnosis (carrying ``type(exc).__name__``)
  on construction failure so the dashboard's 503 banner reflects the
  actual cause.

Both the lifespan path and the runtime-bootstrap path now route
through the same helper and the same teardown on failure.

Tests: 12 new tests covering the helper-level wiring (idempotent
fast-path, missing-prereq parametrised over ``config_store`` /
``collector`` / ``console_metrics``, no-rows error recording, builder
failure error replacement, partial-state teardown), the endpoint
integration, the deterministic concurrent-call lock test (uses an
instrumented lock wrapper that signals when a second acquirer arrives,
so the test fails fast on slow CI rather than depending on a
wall-clock sleep), and a real-builder end-to-end case constructing a
working ``SessionManager`` against a real ``ConfigStore`` + real
``ClusterCollector``.

(cherry picked from commit 3143965e00)
2026-05-07 17:35:22 -07:00
Patrick Buckley 9ae2b376c7 fix(mcp): apply Phase 7b PR #485 review feedback
Two of five Copilot comments on PR #485 were valid; this commit applies
both. The other three (one duplicate of comment 1, plus the INFO-logging
and `_pending`-naming nits) get rationale on-thread and resolution.

1. emit_oauth_failure_audit action now derived from `code` (#485 bug-1)

The Phase 7b refactor generalized `emit_insufficient_scope_audit` →
`emit_oauth_failure_audit`, routing both `mcp_insufficient_scope` AND
generic-403 (`mcp_*_forbidden`) through the same helper. The audit
`action` field stayed hardcoded as
`"mcp_server.oauth.insufficient_scope_emitted"`, mislabeling generic
forbidden events under the insufficient_scope bucket — downstream
alerting / analytics filtering on `action` would silently fold both
categories together.

The action is now selected from `code`:
  * `mcp_insufficient_scope` →
    `mcp_server.oauth.insufficient_scope_emitted` (preserves existing
    alerting consumers)
  * `mcp_tool_call_forbidden` / `mcp_resource_read_forbidden` /
    `mcp_prompt_get_forbidden` →
    `mcp_server.oauth.forbidden_emitted` (new, distinct label)

Detail row continues to carry both `code` and `kind` so operators get
sub-bucket distinction within either action.

2. Resource-listener docstrings cite RFC §3.2 (#485 doc-1)

Per the codebase convention established in Phase 7b round-1 q-1
(`_rebuild_user_prompt_map` corrected §3.2 → §3.3 because prompts are
§3.3 in the MCP spec), resource-related docstrings should cite §3.2.
The three resource-listener docstrings were citing §3.3, and the
"Mirrors `_notify_listeners` for tools (RFC §3.3)" parenthetical in
both `_notify_resource_listeners` and `_notify_prompt_listeners` read
as "tools are at §3.3" — confusing twice over. All four sites now
carry the correct catalog-kind citation explicitly:
  * resource-listener docstrings → "RFC §3.2 (resources)"
  * prompt-listener docstrings → "RFC §3.3 (prompts)"

Tests / lint:
  * 119 passed on 3.13 + 3.11 (targeted MCP OAuth pool tests)
  * ruff + mypy clean on both files

(cherry picked from commit 12cc052bca)
2026-05-07 17:35:22 -07:00
Patrick Buckley b368bdeecc feat(mcp): per-user resource + prompt pool dispatch (Phase 7b)
Extends the Phase 7 per-(user, server) ClientSession pool to cover
RFC §3.2 (resources/read) and §3.3 (prompts/get) on the same shape
already proven for tools/call. Pool discovery is capability-gated so
servers without resources/ or prompts/ stay free of extra round-trips.

API additions / widenings (MCPClientManager):
- ``read_resource_sync(uri, *, user_id=None, timeout=120)`` —
  per-user-first dispatch; falls through to the byte-identical static
  path when ``user_id`` is None or the URI doesn't resolve to an
  ``oauth_user`` pool entry.
- ``get_prompt_sync(prefixed_name, arguments=None, *, user_id=None,
  timeout=30)`` — same dispatch shape; structured-error responses
  surface via ``RuntimeError`` so the agent-loop's ``except Exception``
  block renders the JSON without polluting the prompt-protocol return
  shape.
- ``get_resources(user_id=None)`` / ``get_prompts(user_id=None)`` —
  per-user merged catalogs (admin/global call still passes None).
- ``add_{resource,prompt}_listener`` /
  ``remove_{resource,prompt}_listener`` —  ``user_id`` keyword scopes
  the listener so a pool-only catalog change for one user does not
  wake another user's session.
- ``resource_count_for_user(user_id=None)`` /
  ``prompt_count_for_user(user_id=None)`` — method-form variants used
  by ChatSession's ``read_resource`` / ``use_prompt`` tool gating; the
  legacy ``resource_count`` / ``prompt_count`` properties remain
  static-only for admin paths.
- ``_dispatch_pool_resource`` / ``_dispatch_pool_prompt`` async coros
  — mirror ``_dispatch_pool`` for the new SDK calls; share the
  carrier-race-and-cancel core via ``_dispatch_pool_with_entry_call``.
- ``_handle_auth_403`` extended with ``kind=Literal["tool",
  "resource", "prompt"]`` so the per-operation ``mcp_*_forbidden``
  code surfaces (kind="tool" remains the default for back-compat).
- Pool notification handler now refreshes resources / prompts on
  ``ResourceListChangedNotification`` / ``PromptListChangedNotification``
  via ``_refresh_pool_server_resources`` / ``_refresh_pool_server_prompts``.

ChatSession (``turnstone/core/session.py``) call-site updates:
- 12 sites threaded the session-bound ``user_id`` through
  ``add_*_listener`` / ``remove_*_listener``, ``get_resources`` /
  ``get_prompts``, gating, ``read_resource_sync`` /
  ``get_prompt_sync``, and ``is_mcp_prompt`` so the per-user merged
  catalog drives both the visible-tool set and dispatch.
- ``/mcp`` slash command now lists this user's pool resources and
  prompts alongside tools (Phase 7 already scoped tools).

Scope decisions:
- Per-user-first URI ordering (decision 0.1): the dispatcher attempts
  the user's pool catalog first, falling back to the static catalog
  only when no pool entry resolves the URI / prefixed name. Pool-only
  users never see the static catalog leak into their resolution.
- Method-form ``*_count_for_user`` (vs property) keeps the legacy
  ``resource_count`` / ``prompt_count`` properties intact for admin
  endpoints whose contract is "static catalog size only".
- Shared ``_dispatch_pool_with_entry_call`` helper accepts an
  ``sdk_call: Callable[[ClientSession], Awaitable[Any]]`` closure,
  keeping the entry-locked carrier-race / classification / retry
  plumbing single-source instead of a 3x copy across tool / resource
  / prompt paths.

R6 (anyio uniformity): every pool-side list / read / get path uses
``async with asyncio.timeout(...)`` — ``asyncio.wait_for`` is
forbidden in those paths because it wraps the inner awaitable in a
fresh task and surfaces ``CancelledError`` from inside
``streamablehttp_client``'s anyio TaskGroup on Python 3.11
(per ``feedback_asyncio_timeout_vs_wait_for.md``).

Tests:
- ``test_mcp_pool_auth_resource_integration.py`` — 9 real-transport
  resource tests (FastMCP upstream + ``BehaviorMiddleware``):
  401-refresh-retry success, persistent 401 -> consent_required,
  403+insufficient_scope, 403 generic -> mcp_resource_read_forbidden,
  breaker-isolation under repeated auth failures, missing-token,
  decrypt-failure, http:// URL guard, unknown-URI ValueError.
- ``test_mcp_pool_auth_prompt_integration.py`` — 9 mirror tests for
  the prompt path; structured-error responses verified via
  ``RuntimeError`` payload shape.
- ``test_mcp_user_catalog.py`` — extended unit coverage for per-user
  resource / prompt rebuild + collision policy + symmetric eviction.
- ``test_sessions.py::TestMCPToolGating`` — pool-only-user canary
  asserts ``read_resource`` / ``use_prompt`` stay visible when the
  static catalog is empty but the user has pool entries.

Round-1 review fixes (4-finder review applied, no push yet):
- bug-1: ``_exec_use_prompt`` was hardcoding ``"MCP prompt error: failed
  to invoke prompt"`` — discarding the structured-error JSON that
  ``_dispatch_pool_prompt_sync`` raises via ``RuntimeError``. Now uses
  ``f"MCP prompt error: {e}"`` mirroring ``_exec_mcp_tool``; pool-prompt
  consent_required / insufficient_scope / forbidden errors now reach
  the LLM as intended.
- bug-2 + bug-3: resource template discovery was uncapped —
  ``_cap_server_resources`` covered ``res_result.resources`` but the
  separate ``tmpl_result.resourceTemplates`` loop appended every
  template a server returned. Added ``_MAX_RESOURCE_TEMPLATES_PER_SERVER``
  (1000) + ``_cap_server_resource_templates`` helper, applied at both
  the initial discovery site (``_connect_one_pool``) and the refresh
  site (``_refresh_pool_server_resources``). Mirrors the existing
  ``_MAX_TOOLS_PER_SERVER`` / ``_MAX_PROMPTS_PER_SERVER`` defensive
  ceilings.
- sec-1 + sec-2: ``emit_insufficient_scope_audit`` generalized to
  ``emit_oauth_failure_audit(kind, code, ...)``, called from both the
  insufficient_scope branch AND the previously-silent generic 403
  branch. Audit detail now records ``{"kind": kind, "code": code,
  "scopes_required": [...]}`` so operators can distinguish tool-call
  vs resource-read vs prompt-get 403s in audit logs and so cross-
  tenant probing on the generic 403 path leaves a trail. The Phase 7
  inherited gap (``mcp_tool_call_forbidden`` had the same silence) is
  closed in the same refactor.
- perf-1: pool resource discovery now uses ``asyncio.gather(
  list_resources, list_resource_templates)`` inside the existing
  ``async with asyncio.timeout(...)`` budget — disjoint catalogs, no
  ordering dependency. Typical-case 2-RTT cold-connect resource block
  collapses to 1-RTT. Same change applied at ``_refresh_pool_server_resources``.
- q-1: ``_rebuild_user_prompt_map`` docstring corrected RFC §3.2 →
  §3.3 (resources are §3.2; prompts are §3.3).
- q-2: ``_refresh_pool_server_prompts`` docstring now carries the
  R6 / mcp-loop note that the resource sibling already had — both
  refresh paths now declare the asyncio.timeout invariant explicitly.
- q-5: added the ``_user_resource_map`` / DB-mismatch guard to
  ``read_resource_sync`` for parity with ``get_prompt_sync``. A stale
  per-user map entry with no matching oauth_user row now raises a
  specific ValueError instead of silently falling through to a
  generic ``Unknown MCP resource``.
- q-6: ``_dispatch_pool_with_entry`` (now a single-caller wrapper
  after the ``_dispatch_pool_with_entry_call`` extraction) gains a
  one-line docstring explaining why the wrapper is preserved
  (tool-decode localization + stack-trace identity for debugging).
- q-7: added 1 resource + 1 prompt end-to-end integration test that
  drive REAL discovery + dispatch in the same connect (no
  ``_seed_pool_*_map`` shortcuts), mirroring the tool path's
  ``test_integration_pool_reuse_401_refresh_and_retry_succeeds``.
  The seeded-map tests stay (faster, focused on dispatch); the new
  e2e tests cover the connect-discover-dispatch composition that
  caught Phase 6's carrier-on-entry bug.

Pre-push round-1 review fixes (3-finder review on the final state —
the lesson from Phase 7 round-3's q-1 regression: round-2 catches
what the round-1 apply pass missed):
- q-1 (MAJOR): the bug-1 sibling that round-1 missed —
  ``_exec_read_resource`` was hardcoding ``"MCP resource error: failed
  to read resource"`` while ``_exec_use_prompt`` (post-bug-1) preserved
  the structured-error JSON via ``f"... error: {e}"``. The round-1
  apply pass patched the prompt side but not the resource side. q-5's
  per-user-map / DB-mismatch ValueError was being swallowed at the
  agent loop boundary, defeating the operator-diagnostic intent. Now
  ``_exec_read_resource`` mirrors ``_exec_mcp_tool`` and ``_exec_use_prompt``.
- q-6 (nit): defensive-cap comment block at module-level cited
  "(RFC §3.2)" while covering both resource and prompt list paths;
  prompts are §3.3. Now reads "(RFC §3.2 for resources, §3.3 for
  prompts)" matching the convention the q-1 apply established.
- q-5 (rejected with better justification): the reviewer flagged
  ``_dispatch_pool_with_entry`` as a single-caller wrapper that should
  be inlined. After examination — the autouse fixture
  ``tests/test_mcp_pool_auth_introspection.py::_install_capture_intercept``
  monkeypatches this method to stash ``entry.auth_capture`` for the
  fake call_tool stubs in dispatcher-asserting tests. Inlining would
  redirect the patch to ``_dispatch_pool_with_entry_call`` (different
  kwargs shape) and require re-validating every test that depends on
  the interception. The wrapper IS load-bearing; q-6 docstring updated
  to cite the test-fixture rationale instead of the thin "stack-trace
  identity" claim.

Deferred to follow-up (documented rationale):
- perf-2: single-pass partition for system-message resource list
  (concrete vs templates). Sub-microsecond at expected scale;
  opportunistic-only.
- q-2 (pre-push): ~200 lines of fixture infrastructure
  (``BehaviorMiddleware``, ``_build_server``, ``_seed_oauth_server``,
  ``running_loop_mgr``, etc.) duplicated across three pool-integration
  test files. Real maintenance cost, but a 200-line conftest extraction
  is a focused refactor that earns its own commit / PR. Tracking as
  follow-up rather than balloon Phase 7b's diff further.
- q-3 / q-4 (refactor): extract shared dispatcher / scheduler
  helpers to compress three near-identical 90-line bodies (round-1
  q-3 was the same root cause; the pre-push q-3/q-4 reviewer
  reaffirmed it concretely). Three named methods preserve readability
  for the codebase's hottest correctness path; follow-up if
  duplication grows further or if a per-path divergence ships.
- q-4 (round-1, distinct from pre-push q-4): split pool concerns
  into ``mcp_pool.py``. Out-of-scope per finder; future refactor as
  the file approaches the navigation/merge-conflict threshold.

3.13: 5590 passed (5541 baseline -> +49 net; pre-review +47, q-7
e2e tests added +2). Existing audit-detail tests updated in-place
to expect the new ``kind`` and ``code`` fields.
3.11: 5590 passed (parity gate per ``feedback_pytest_env_parity.md``).

(cherry picked from commit 124615cce0)
2026-05-07 17:35:22 -07:00
Patrick Buckley d767aca784 fix(metacog): atomic cap-and-drop helper for soft-cap producers
Closes PR #484 review findings (Copilot): the soft-cap pattern in
``ChatSession.set_watch_runner``'s dispatch closure was a non-atomic
two-call pair (``count_by_type`` then ``drop_oldest_by_type``) with
two separate lock acquisitions.  A concurrent drain on the worker
thread (``USER_DRAIN`` / ``TOOL_DRAIN`` consuming ``"watch_triggered"``
entries via the ``"any"`` channel) could slip between the two calls,
making the drop a no-op.  The dispatch closure also discarded
``drop_oldest_by_type``'s return value and unconditionally logged
``dropped_oldest=True``, so a no-op drop got reported as a successful
drop.

* New ``NudgeQueue.cap_at_or_drop_oldest(nudge_type, max_depth,
  channel=None) -> bool`` does the count+drop in a single critical
  section.  Returns the actual outcome.

* Dispatch closure (``session.py:1410-1416``) now calls the helper and
  uses its return value to gate the WARNING log line, so the log is
  accurate when a drop did NOT happen.

* ``drop_oldest_by_type``'s docstring no longer overstates the
  per-call lock as covering a count+drop pair — it points readers
  to ``cap_at_or_drop_oldest`` for that contract.

7 new tests in ``TestCapAtOrDropOldest`` cover: below-cap no-op,
at-cap drop-oldest, above-cap drop-only-one (per-call), channel
filter, other-type isolation, ``max_depth <= 0`` defensive no-op,
no-match.

5708 non-live tests pass; ruff + mypy clean.

The github-code-quality bot finding ("Statement has no effect" on
``_protocol.py:939``'s ``...`` body) is a false positive — every
Protocol method in ``_protocol.py`` uses ``...`` as its body, which
is the canonical Python Protocol pattern.  Replacing with ``pass``
would diverge from the file's existing style.  No code change.

(cherry picked from commit c757c22f55)
2026-05-07 17:35:22 -07:00
Patrick Buckley 12bc580dee fix(metacog): factor sanitiser regex tail + trim docstrings + drop tombstone
Closes round-2 review findings q-3, q-4, q-5, q-7.

* **q-4:** ``_NAME_CONTROL_CHARS`` and ``_PAYLOAD_CONTROL_CHARS`` shared
  7 lines of Unicode-steering character classes (zero-width / bidi /
  separators / BOM / tag chars above BMP).  Factored into a single
  ``_CONTROL_CHARS_TAIL`` constant; each regex now differs only in its
  leading ASCII range.  Future bidi or zero-width additions edit one
  place.

  Side effect: this corrects a latent bug where ``_NAME_CONTROL_CHARS``
  had two literal ASCII spaces in place of U+2028 / U+2029 (line and
  paragraph separators) — visible as ``r"  "`` in source but rendered
  as the actual codepoints in ``_PAYLOAD_CONTROL_CHARS``.  After the
  factoring both regexes correctly include U+2028 / U+2029, closing
  the gap that would have let a workstream name with embedded line
  separators forge a sibling bullet (the same vector ``\n`` was
  blocked for in the original bug-1 fix).

  Switched to ``\u`` escapes for readability (and to keep future Edit
  tool runs against this block reliable).

* **q-3:** Tombstone clause "standing in for the deleted
  ``_watch_pending`` maxsize bound" survived in
  ``ChatSession.set_watch_runner``'s docstring after the apply-pass
  trim cleaned the inline soft-cap comment.  Dropped.

* **q-5:** ``test_newline_in_name_does_not_forge_extra_bullet`` carried
  five WHAT-narration comments restating what the immediately-following
  asserts already say.  Dropped — the docstring carries the security
  invariant; the assertions speak for themselves.

* **q-7:** ``patch_session_storage`` had a 14-line docstring including
  fallback-guidance and self-justification ("accumulated 7 near-duplicate
  sites").  Trimmed to a 3-line contract.

(cherry picked from commit 39e0f930c1)
2026-05-07 17:35:22 -07:00
Patrick Buckley aa1446364b test(metacog): drop redundant valid_until test + tighten concurrency bound + cover is_watch_active
Closes round-2 review findings q-1, q-2, q-6.

* **q-1:** ``test_valid_until_drops_when_watch_missing`` collapsed to the
  same code path as ``test_valid_until_drops_when_watch_inactive`` after
  the apply-pass switched the predicate from ``get_watch[active]`` to
  ``is_watch_active`` (both stubbed via ``patch_session_storage(active=False)``).
  The "missing" case has no distinguishable branch at the dispatch
  layer, so dropping it removes a tautological duplicate.  The
  missing-row mapping moves to the storage layer (q-2 below) where it
  IS distinguishable.

* **q-2:** ``is_watch_active`` was a new public storage primitive with
  zero direct backend coverage — only via-session-via-stub coverage.
  New ``TestIsWatchActive`` in ``tests/test_watch_storage.py`` covers
  active row → True, inactive row → False, missing row → False.
  Pinned at the storage boundary so future backend changes fail loudly
  there instead of in the dispatch tests.

* **q-6:** Concurrency test had ``n_threads = 2`` alongside two literal
  Thread objects and a tautological ``assert len(threads) == n_threads``.
  Threads are now built from a labels tuple, so ``len(threads)`` drives
  the slack bound; the redundant assertion is gone.

(cherry picked from commit 751ed9c85f)
2026-05-07 17:35:22 -07:00
Patrick Buckley 21507dc02e fix(metacog): tighten concurrency bound + lift storage-patch helper
Closes review findings bug-4 and q-6.

bug-4 — the watch dispatch concurrency test bounded depth at
``_WATCH_QUEUE_SOFT_CAP + 2 * per_thread`` (= 250) which is
tautologically true: two threads × 100 fires can append at most 200
entries above the cap, so the bound asserted nothing more than what
``depth <= 2 * per_thread`` already says.  Tighten to
``_WATCH_QUEUE_SOFT_CAP + N_THREADS`` (= 52): the count-then-drop window
admits at most one slip per concurrent thread.

q-6 — 7 near-duplicate ``monkeypatch.setattr(session_mod, "get_storage",
lambda: _StubStorage())`` sites across ``test_watch_dispatch.py`` +
``test_watch_integration.py`` (4 different stub shapes, mostly trivial
variations on the active flag).  Lift a ``patch_session_storage``
helper into the existing ``tests/_helpers.py`` with kwargs for the
common cases (``active``, ``raise_on_is_active``), returns the call list
so call-shape assertions still work.  Tests collapse from ~10-line
inline-class blocks to one-line helper calls.

(cherry picked from commit 20c4dfaca6)
2026-05-07 17:35:22 -07:00
Patrick Buckley a0ed4b9897 fix(metacog): drop watch_id rebind + trim soft-cap inline comment
Closes review findings q-2 and q-5.

q-2 — ``bound_watch_id = watch_id`` rebind was unnecessary.  ``_dispatch``
is constructed fresh per fire (not in a loop), so ``_still_active``
closes over the function parameter directly without any
loop-variable-capture risk.  Drop the rebind.

q-5 — the inline soft-cap comment restated rationale already covered by
the ``_WATCH_QUEUE_SOFT_CAP`` block-comment at module scope and dragged
in a tombstone reference to the deleted ``_watch_pending`` path.  Trim
to one line stating only the WHY (drop-oldest because latest output is
most useful).  Leave the ``set_watch_runner`` docstring's operational
detail at lines 1356-1378 alone — trimming further risks losing the
``valid_until`` predicate semantics.

(cherry picked from commit 28d9bb4802)
2026-05-07 17:35:22 -07:00
Patrick Buckley ab8ee0d759 test(metacog): integration coverage for _watch_restore_fn closure
Closes review finding q-4.

The closure built inside ``server.py``'s ``_watch_restore_fn`` is the
new contract surface introduced by the switchover — it constructs a
fresh ChatSession, calls ``session.resume(ws_id)`` to adopt the
original ws_id, re-registers the dispatch closure via
``set_watch_runner``, and returns ``WatchRunner.get_dispatch_fn`` for
the runner to invoke directly.  No automated coverage exists today;
a future refactor (e.g. swapping ``manager.create + session.resume``
for ``manager.open``) could silently break the watch-restore pipeline.

Adds ``test_watch_dispatch_through_restore_fn_lands_on_rehydrated_session``
to ``tests/test_watch_integration.py`` — drives the full restore path:
persists a kickoff message for the original ws_id, fires
``_dispatch_result`` against a runner with no registered dispatch fn,
asserts the restore_fn ran exactly once, the rehydrated session is a
distinct object that adopted the original ws_id, and the watch payload
landed on the rehydrated session's NudgeQueue (not on the original).

(cherry picked from commit ed1eaee216)
2026-05-07 17:35:22 -07:00
Patrick Buckley d34f6cd0b1 fix(metacog): is_watch_active storage primitive for hot-path valid_until
Closes review finding perf-1.

The watch dispatch closure's ``valid_until`` predicate fires once per
watch entry at every drain seam — on the chat-loop hot path.  It only
needs the ``active`` flag, but ``storage.get_watch`` runs a full-row
``SELECT *`` and marshals the result into a dict.  At the typical drain
depth (cap-50 + a busy chat loop) that's ~50 throwaway dict allocations
per drain pass for one boolean.

Adds ``StorageProtocol.is_watch_active(watch_id) -> bool`` plus
SQLite + Postgres implementations doing a single-column
``SELECT active FROM watches WHERE watch_id = ?`` (returns False on
missing row).  ``_still_active`` in ``ChatSession.set_watch_runner``
now calls that instead of indexing into the full row.

Test stubs that mocked ``get_watch`` for the predicate are converted
to mock ``is_watch_active`` directly.  Bulk variant deferred — single-row
fix is sufficient at typical drain depths.

(cherry picked from commit 3b495eba15)
2026-05-07 17:35:22 -07:00
Patrick Buckley b219c47ba8 fix(metacog): NudgeQueue.count_by_type primitive + channel-aligned soft cap
Closes review findings perf-2, q-3, bug-3.

The watch dispatch closure's soft-cap pre-check materialised the whole
queue snapshot via ``pending(channel="any")`` only to throw away the
text and count the type — wasteful at typical drain depths (cap-50 +
mixed producers means a 50-tuple allocation per fire just to read a
length).  The other half of the cap pair (``drop_oldest_by_type``)
walked the *whole* queue regardless of channel, so a future producer
that enqueued ``"watch_triggered"`` on a different channel could be
dropped by the watch cap, and vice versa — silently surprising once
that producer existed.

Adds ``NudgeQueue.count_by_type(nudge_type, channel=None) -> int`` that
walks ``_items`` once under the queue lock without materialising
tuples; extends ``drop_oldest_by_type`` to take an optional ``channel``
filter so both halves can agree on the entry set being capped.  The
watch dispatch closure now passes ``channel="any"`` to both —
consistent with where the closure enqueues — so a future channel split
can't bleed across producers.

Adds ``TestCountByType`` mirroring the existing ``TestDropOldestByType``
shape, plus a ``test_drop_oldest_by_type_channel_filter`` case pinning
the new optional argument's behaviour.

(cherry picked from commit e5e6e13307)
2026-05-07 17:35:21 -07:00
Patrick Buckley d770a811a8 fix(metacog): drop test_watch_live.py — defer R9 to operator-driven verification
Closes review finding q-1.

The live-marker scaffold in ``tests/test_watch_live.py`` couldn't actually
run as written: the ``live_client`` / ``live_model_id`` fixtures it
referenced live in ``tests/test_server_live.py`` at ``scope="module"``,
not on a shared ``conftest.py``, so the file would have ImportError'd
at collection if anyone ever tried ``pytest -m live`` against it.

Lifting the fixtures into a shared conftest is a larger refactor
than R9 justifies — the deterministic envelope-arrival contract is
already pinned end-to-end by ``test_watch_fires_then_user_send_drains_envelope``
and ``test_three_back_to_back_watch_fires_drain_into_one_turn`` in
``test_watch_integration.py`` (real ChatSession + real WatchRunner +
real chat-loop drain).  The model-quality-of-response leg is genuinely
manual; the plan doc's R9 entry is updated locally to reflect that
deferral.

(cherry picked from commit 68a44cc7e2)
2026-05-07 17:35:21 -07:00
Patrick Buckley 912e9c57b0 fix(metacog): split sanitiser regex — strict for names, permissive for payloads
Closes review finding bug-1.

The shared ``sanitize_payload`` regex preserved TAB/LF/CR so multi-line
watch shell output kept its layout — necessary for the watch path, but a
correctness gap for the idle_children formatter, which renders the
user-controlled ``name`` field as a single bullet item.  A child name
with an embedded ``\n`` would split the bullet across two rendered rows
and let a hostile name forge a fake sibling entry in the listing.

Splits the regex in two: ``_NAME_CONTROL_CHARS`` strips TAB/LF/CR
(used by the new ``sanitize_name`` helper for single-line name fields),
``_PAYLOAD_CONTROL_CHARS`` keeps the existing permissive shape (used by
``sanitize_payload`` for multi-line watch payloads).
``format_idle_children_nudge`` now calls ``sanitize_name``.

Adds ``test_newline_in_name_does_not_forge_extra_bullet`` — feeds a
hostile name with embedded ``\n`` + bullet-shaped continuation, asserts
the rendered listing still has exactly N bullet rows for N children
(no forged sibling), and the hostile newline got flattened to an inline
space.  Adds a ``TestSanitizeName`` class mirroring the existing
``TestSanitizePayload`` shape for the new strict variant.

(cherry picked from commit e596650a5c)
2026-05-07 17:35:21 -07:00
Patrick Buckley d7c6053441 fix(metacog): drop misleading _watch_restore_fn comment
The deleted comment claimed the closure may be registered "under the
rehydrated workstream's id, which may differ from the original ws_id we
restored against" — but ``ChatSession.resume(ws_id, fork=False)`` adopts
the parameter as the session's id at session.py:1682, so they match
exactly post-resume.  The lookup works because the ids are equal, not
because they may differ.

The accessor name ``get_dispatch_fn`` is self-explanatory; no replacement
comment is needed (per the project's "default to no comments" rule).

(cherry picked from commit d2028aa4f7)
2026-05-07 17:35:21 -07:00
Patrick Buckley e7a17a20b0 test(metacog): watch switchover boundary integration + live scaffold
Adds two boundary-crossing integration tests and one live-marker
scaffold for the watch switchover landed in the previous commits:

tests/test_watch_integration.py — drives a real ChatSession + real
WatchRunner end-to-end (LLM stubbed) through the unified pull-model
chat-loop drain seam.  Pins:

- test_watch_fires_then_user_send_drains_envelope: a synchronous
  WatchRunner.dispatch fire enqueues "watch_triggered" on "any";
  session.send drains the entry into the user message's _reminders
  side-channel — confirms the envelope splice path.
- test_three_back_to_back_watch_fires_drain_into_one_turn: pins the
  intentional behavioural delta from the plan section 3.4 / risk
  register R3 — N back-to-back fires now produce ONE assistant turn
  with N _reminders entries, not N successive turns.

tests/test_watch_live.py (new file, single test, marked @pytest.mark.live):
risk register R9 verification recipe — confirm a real LLM handles a
<system-reminder>-framed watch payload sensibly.  Collects under the
regular -m "not live" run; the user runs it on demand against an
Anthropic-backed config.

Implements watch-switchover plan section 5.2 (integration) and step 11
(live scaffold).

(cherry picked from commit 17c62f7ef3)
2026-05-07 17:35:21 -07:00
Patrick Buckley 931a1eca9d test(metacog): NudgeQueue-based dispatch tests for watch closure
Replaces the deleted tests/test_watch_dispatch.py with a focused
14-test suite exercising the closure that ChatSession.set_watch_runner
now constructs (per the previous commit's switchover).  Each test
pins one assertion:

- enqueue shape: ("watch_triggered", text, "any") on the per-session
  NudgeQueue; not on user / tool channels
- producer-side sanitisation strips control / bidi / zero-width chars
  and angle-bracket tag breakers; preserves TAB/LF/CR so multi-line
  shell output keeps its layout (R8); empty-after-strip → no enqueue
- soft-cap drop-oldest at _WATCH_QUEUE_SOFT_CAP with a queue_full
  WARNING log; non-watch entries on the same queue are not collateral
  damage
- valid_until predicate drops on inactive / missing / storage-raises;
  delivers when active (counter-test)
- concurrent enqueues across two threads stay bounded under the
  3-acquisition count-then-drop window

Implements watch-switchover plan section 5.1 / step 9.  No production
changes — pure test rewrite.

(cherry picked from commit 7ca00b564c)
2026-05-07 17:35:21 -07:00
Patrick Buckley 048285a423 feat(metacog): switchover — watches enqueue onto NudgeQueue not _watch_pending
Replaces the bespoke _make_watch_dispatch / _watch_pending /
_dispatch_pending_watch / _MAX_WATCH_CHAIN machinery with a single
NudgeQueue.enqueue("watch_triggered", ...) call inside
ChatSession.set_watch_runner.  Watch results now drain at the same
<system-reminder> envelope seams as every other metacog nudge
(USER_DRAIN, TOOL_DRAIN, IdleNudgeWatcher IDLE wake) — no separate
worker-spawn, no recursive watch chain, no per-session queue.Queue.

The dispatch closure built inside set_watch_runner carries:
- producer-side sanitize_payload over the whole formatted message
  before enqueue, so steering-vector / control-char shell output
  can't tamper with the envelope at interpolation time
- a soft cap of 50 entries on per-session "watch_triggered" depth
  via the new NudgeQueue.drop_oldest_by_type, replacing the prior
  _watch_pending maxsize=20 + _MAX_WATCH_CHAIN=5 bounds; drop policy
  is drop-oldest (latest output most useful), logged at WARNING
- a valid_until predicate that re-checks
  storage.get_watch(watch_id)["active"] at drain time so a cancelled
  watch's last splat doesn't ride out a future wake

Behavioural delta documented in the plan section 3.4: N back-to-back
watch fires now drain into ONE assistant turn responding to all N
(via the envelope splice) instead of N separate send turns.  This is
intentional — fewer model invocations for noisy watches, and uniform
with the rest of the metacog pull-model surface introduced by #482.

Implements watch-switchover plan steps 5-8.  Server-side simplifications
let the previously-load-bearing _make_watch_dispatch (47 lines), its
session_worker.send import, and the chat-loop _dispatch_pending_watch
seam at the no-tools IDLE branch all disappear.  The obsolete
tests/test_watch_dispatch.py and the wake-tag test in test_session.py
(both pinning contracts that no longer exist) are removed; the
NudgeQueue-based replacement plus an integration test land in the
following commit.

(cherry picked from commit 94ed79d488)
2026-05-07 17:35:21 -07:00
Patrick Buckley 481347eb17 refactor(metacog): widen WatchRunner dispatch_fn signature to (msg, watch_id)
Widens the per-workstream dispatch fn signature from ``(message,)``
to ``(message, watch_id)``.  The runner now passes the originating
``watch_id`` through ``_dispatch_result`` so dispatch closures can
capture per-watch metadata at fire time — the upcoming switchover
needs this for the ``valid_until`` predicate that re-checks
``storage.get_watch(watch_id)["active"]`` before a stale entry rides
out a wake.

Also adds ``WatchRunner.get_dispatch_fn(ws_id)`` as the public
accessor used by the server-side restore path to retrieve the
closure that ``set_watch_runner`` constructed during workstream
rehydrate (avoiding private-attr access into ``_dispatch_fns``).

Implements watch-switchover plan step 4 plus risk register R4.
The pre-existing single-arg callers (``_make_watch_dispatch`` and
``set_watch_runner``'s ``dispatch_fn=`` fallback) get replaced
in the next commit; their mypy types are ``Any`` today so the
type mismatch isn't caught at this step.

(cherry picked from commit 195ff985cc)
2026-05-07 17:35:21 -07:00
Patrick Buckley 31a554a4bd refactor(metacog): shared sanitize_payload + watch_triggered nudge type
Renames _sanitize_child_name to sanitize_payload and widens it to be
the shared producer-side sanitiser for both idle_children and the
incoming watch_triggered nudges.  The regex now skips TAB / LF / CR
so multi-line shell output rendered into a watch payload keeps its
line structure when sanitised as a whole formatted message — the
pre-switchover code path collapsed multi-line output to one line.

Adds the watch_triggered entry to _NUDGE_MAP alongside idle_children
so ``_NUDGE_MAP``-as-registry consumers (should_nudge gating, future
audit / UI tagging) recognise the type.  Body is empty — payload
comes from the producer (the watch dispatch closure), same shape as
idle_children.

Implements watch-switchover plan section 3.2 plus risk register R8
(TAB/LF/CR exclusion) and step 3 (_NUDGE_MAP registration).

(cherry picked from commit 78ae7ae6b5)
2026-05-07 17:35:21 -07:00
Patrick Buckley af2c0ae13a feat(metacog): NudgeQueue.drop_oldest_by_type helper for soft-cap producers
Adds an atomic drop-oldest-by-type operation to NudgeQueue used by
producers that need a per-type soft cap on their own queue depth.
The watch dispatcher (next commit in this stack) is the first user:
when "watch_triggered" saturates, the dispatch closure drops its
oldest entry under the queue lock so the count snapshot and drop
can't interleave with a concurrent enqueue from the same producer.

Implements watch-switchover plan section 3.1 — the producer-side soft
cap takes the place of the deleted _watch_pending maxsize=20 bound.
Other producers (idle_children, advisories) have natural rate limiters
already, so the helper is opt-in per producer rather than a global cap
in enqueue itself.

(cherry picked from commit 74f1958e47)
2026-05-07 17:35:21 -07:00
Patrick Buckley 0808dc0af0 fix(mcp): apply Phase 7 PR review feedback
Three Copilot findings on PR #483 (commit dad98c0); one rejected as a
false positive.

- mcp_client.py:1189 — pool notification handler's exception path
  used ``log.warning(..., exc_info=True)`` which serializes the
  chained ``httpx.Request.headers`` carrying ``Authorization: Bearer
  <token>`` into Sentry / faulthandler frame captures. Same threat
  model as the round-1 sec-1 dispatch-path fix, applied to a site
  the original review missed. Now logs structured fields only
  (server, user, exc type) without ``exc_info``.

- mcp_client.py:1202 — ``_connect_one_pool``'s handshake step used
  ``asyncio.wait_for(session.initialize(), ...)``, the same Python
  3.11 + anyio cross-task-cancel-scope anti-pattern that the
  Phase 7 round-3 q-1 fix removed from the discovery step (and that
  f6a3b66 originally addressed for ``_safe_close_stack``). Pre-
  existing Phase 5 code, but the same latent bug class — a 401
  during initialize() under 3.11 would surface ``RuntimeError:
  Attempted to exit cancel scope in a different task`` as the
  SDK's TaskGroup unwinds. Switched to ``async with asyncio.timeout(...)``
  matching the discovery step's pattern.

- mcp_client.py:1522 — renamed loop tuple-unpack variable
  ``_server_name`` → ``server_name`` in ``_rebuild_user_tool_map``.
  The leading underscore conventionally signals "intentionally
  unused", but the variable is read at the assignment a few lines
  below. Two other ``_server_name`` unpacks in this file (1410,
  3111) genuinely don't use the value and keep the underscore.

Rejected as false positive:
- test_mcp_user_catalog.py:58 (github-code-quality bot, "Statement
  has no effect"): ``await task`` inside ``contextlib.suppress(
  BaseException)`` is the standard pattern for cleanly draining a
  cancelled task. The bot's static analysis treats ``await`` of a
  result that's discarded as a no-op statement, but ``await`` here
  triggers cancellation propagation and waits for the task to
  finish — load-bearing in the fixture's teardown. No change.

Verified on Python 3.11 (``/tmp/venv311``) and 3.13 (``.venv``):
ruff + mypy clean, full test suite green.

(cherry picked from commit 62909d402c)
2026-05-07 17:35:21 -07:00
Patrick Buckley cfc8a6c8c0 feat(mcp): per-user catalog scoping (Phase 7 — tools)
Light up production reachability of pool dispatch (RFC §3, invariant 8)
by widening the public catalog API to optionally take a ``user_id``:

- ``MCPClientManager.get_tools(user_id=None)`` returns the merged
  static + per-user pool view when ``user_id`` is supplied; the default
  preserves the legacy global-only contract.
- ``is_mcp_tool(name, *, user_id=None)`` extends the lookup to the
  per-user ``_user_tool_map``. Pool tools become reachable from
  ``ChatSession._prepare_tool`` only when the session-bound user_id
  flows through — flipping invariant 8 from "must hold" to "satisfied".
- Listener identity becomes ``(user_id, callback)``. Static-path
  changes fire ALL listeners (admin + every user); pool-entry
  changes fire only matching-user + admin (``None``) listeners.
  RFC §3.3.
- Pool sessions discover their tool list on first connect
  (``_connect_one_pool`` → ``await session.list_tools()``); the
  notification closure binds to ``(user_id, server_name)`` so
  push-driven ``list_changed`` updates target the correct user's
  catalog. R6 verified empirically: ``list_tools()`` 401 propagates
  through anyio TaskGroup unwinding, no hang — plain ``await`` is
  fine, no carrier-race shape needed for discovery.
- ``_evict_session`` drops ``entry.tools`` and rebuilds the user's
  index so an evicted-then-reconnected session doesn't carry
  stale catalog state.
- ``web_search.resolve_web_search_client`` refuses
  ``auth_type=oauth_user`` backends (per-node web search can't
  carry per-user tokens).

Resources / prompts pool dispatch deferred to Phase 7b — invariant 8
is satisfied by the tool path alone, and the resource/prompt path
needs sibling ``_dispatch_pool_resource_sync`` /
``_dispatch_pool_prompt_sync`` helpers each with their own
carrier-race plumbing (~400 LOC). Phase 7b will follow the patterns
established here.

CLI sessions default ``user_id=""`` and so cannot use oauth_user
MCP servers — documented limitation; users must use the web UI.

Round-1 review fixes (4-finder review applied, no push yet):
- bug-1: get_tools(user_id) was iterating _user_pool_entries from sync
  threads while the mcp-loop concurrently mutated it (RuntimeError:
  dictionary changed size during iteration). Now reads from a sibling
  _user_tools dict updated atomically by _rebuild_user_tool_map.
- bug-2: _close_pool_entry_if_idle (LRU/TTL eviction) skipped the
  catalog cleanup that _evict_session does — stale tools persisted
  in _user_tool_map and ChatSession's tool list never rebuilt. Now
  mirrors _evict_session.
- perf-1: _last_pool_notification_refresh debounce dict was never
  pruned in either eviction path. Now popped alongside the entry.
- perf-3: web_search resolver was issuing a sync SQL query per LLM
  turn to gate oauth_user backends. Now reads from the cached
  in-memory config.
- sec-1: bearer token could leak into exc_info-rendered tracebacks
  via Sentry/faulthandler. log.debug now uses structured fields,
  not exc_info.
- sec-2: tools-per-server response now capped at 1000 (defensive,
  mirrors _MAX_ERROR_LEN / _MAX_INSUFFICIENT_SCOPE_REPORTED).
- Test cleanup: dropped two listener fan-out tests duplicating
  test_mcp_client.py coverage; renamed test_pool_session_notification_handler
  to match its actual scope (_refresh_pool_server_tools); removed
  stale comments referencing /tmp/r6-spike*.py scratchpads and a
  misleading "copy-on-write" comment.

Round-2 pre-push review fixes (focused single-pass review applied):
- round2-1: bug-2's catalog-cleanup block in _close_pool_entry_if_idle
  had no integration test (exactly the failure mode flagged in
  feedback_tests_through_boundaries.md). Added
  test_close_pool_entry_if_idle_clears_catalog_and_fires_listener
  driving the LRU/TTL eviction path through real streamablehttp_client +
  MockTransport. Negative-test verified: reverting the
  _rebuild_user_tool_map / _notify_user_tool_listeners calls makes
  the new test fail.
- round2-3: documented the _oauth_user_server_names cache invariant
  in add_server_sync / remove_server_sync docstrings. Cache is
  reconcile_sync's sole owner — direct callers leave it stale, but
  _db_servers_to_config strips oauth_user rows so production paths
  are unaffected. Static→oauth_user transitions correctly leave the
  name in the cache because remove_server_sync drops the static
  connection, not the cache identity.
- round2-6: strengthened test_rebuild_user_tool_map_populates and
  test_rebuild_user_tool_map_drops_empty_user to assert on the
  _user_tools sibling cache (bug-1 fix). Without this, a future
  revert dropping the sibling write would still pass the unit
  tests because get_tools coverage lives in separate tests.

Round-3 full-stack review fixes (multi-stage review on the final
state caught what the layered apply passes missed):
- q-1 REGRESSION: pool tool-discovery used asyncio.wait_for around
  session.list_tools(), the exact pattern the f6a3b66 fix (and
  feedback_asyncio_timeout_vs_wait_for.md) put in place to avoid.
  Python 3.11's asyncio.wait_for wraps the inner coroutine in a
  fresh task → cross-task scope-exit when the SDK's anyio TaskGroup
  unwinds on a 401. Switched to `async with asyncio.timeout(...):`
  pattern used by _safe_close_stack.
- sec-2: TOCTOU in _connect_one_pool — entry.tools was published
  (via _rebuild_user_tool_map + listener fan-out) BEFORE entry.session
  was assigned. A sync-thread reader could observe a tool whose
  backing entry has session=None. Defence-in-depth — dispatch
  re-fetches its own token and lazy-reconnects on session=None — but
  reordering catches the race at the source. entry.session now
  publishes BEFORE catalog visibility.
- bug-1: _close_pool_entry_if_idle's _user_pool_locks.pop ran
  unconditionally after the try/finally, but the early-return
  branches (entry None on re-check, in_flight > 0 under lock) skip
  it via Python's return-through-finally semantics. The lock was
  never popped on those paths. Now gated behind an `evicted` flag
  set only on the success path; in_flight > 0 leaves the lock for
  the active dispatcher to reuse, entry-None races leave the lock
  for re-allocation by _ensure_pool_entry. Comment now describes
  the actual semantics, not the original promise.
- bug-2: softened the _rebuild_user_tool_map docstring's atomicity
  claim. The two-dict write is technically non-atomic across Python
  statements; in practice the window is sub-microsecond on the
  mcp-loop with no awaits between writes, and the listener fan-out
  fires AFTER both writes complete. Docstring now says "back-to-back
  on the mcp-loop" instead of "atomically alongside".
- q-3: dropped `hasattr(mcp_client, "server_auth_type")` defensive
  check in web_search.py. The method ships in this commit; the
  hasattr created a silent fallthrough that would let a future
  rename silently re-enable oauth_user backends.
- q-4: surfaced the CLI / empty-user_id limitation in a docstring
  comment at ChatSession.__init__'s self._user_id assignment. The
  note previously lived only inside is_mcp_tool's docstring — a
  future maintainer wiring CLI features against MCP pool servers
  wouldn't think to read is_mcp_tool to find the constraint.
- q-2 + q-5: deleted a tautological duplicate test in
  test_mcp_user_catalog.py whose docstring claimed to test
  ChatSession.close but never instantiated a ChatSession (the
  manager-level identity semantics are already covered by
  test_listener_identity_includes_user_id in the same file and by
  test_session_close_removes_listener_with_same_user_id in
  test_mcp_client.py which DOES drive a ChatSession). Reworded a
  misleading "fixture provides only 5s" comment to point at the
  actual `_run_on_loop(..., timeout=5)` site.
- q-6: the `self._user_id or None` collapse repeated at 8 sites
  across session.py. Cached once at __init__ as
  ``self._mcp_user_id`` (since ``_user_id`` is set once and never
  mutated); 8 call sites now read the cached value. The empty-
  string-is-CLI-sentinel invariant is documented at the assignment
  site, not re-asserted at each consumer.

Deferred to follow-up:
- sec-1: a hostile MCP server bound to user-A could craft a
  tool.name containing `__` to synthesize a prefixed-name collision
  in user-A's own catalog. Bounded impact: cross-tenant dispatch is
  prevented by the per-tenant token gate in _dispatch_pool, and
  user-B's get_tools(user_id="B") never includes user-A's pool
  entries. The fix needs policy decisions (reject vs. sanitize)
  and touches _mcp_to_openai which is shared between static and
  pool paths; better discussed in its own follow-up where the
  policy applies uniformly to static-path servers too. The threat
  model already requires user-A to have consented to a malicious
  server, who has many more dangerous vectors than tool-name
  shenanigans.

Test count delta: +31 tests (5435 → 5466, ``-m "not live"``; one
test deleted in round-3 apply per q-2):
- ``tests/test_mcp_client.py`` +20 (per-user catalog state, listener
  identity, session thread-through)
- ``tests/test_mcp_user_catalog.py`` +9 NEW (integration tests
  driving real ``streamablehttp_client`` + ``httpx.MockTransport`` per
  invariant 14: discovery on connect, user isolation, eviction +
  reconnect, LRU/TTL eviction (round2-1), R6 401-propagation
  regression, static byte-identical canonical regression; review
  passes dropped duplicate listener fan-out tests from earlier
  drafts whose coverage lived in test_mcp_client.py)
- ``tests/test_web_search.py`` +2 (oauth_user backend rejection +
  static backend acceptance regression; updated to use the new
  ``server_auth_type`` in-memory accessor)

(cherry picked from commit a8b34bfe54)
2026-05-07 17:35:21 -07:00
Patrick Buckley 266e3536aa fix(metacog): bot-review fixes — watcher gate + two stale docstrings
Three confirmed findings from the PR #482 bot review pass.

* **Copilot (idle_nudge_watcher.py)**: ``IdleNudgeWatcher`` was gating
  wake dispatch on ``len(_nudge_queue) == 0`` (any channel), but
  ``deliver_wake_nudge_from_queue`` only drains ``USER_DRAIN``.  A
  ``"tool"``-channel entry queued by ``_queue_tool_advisory`` would
  pass the gate, spawn a wake daemon, and immediately no-op at the
  drain guard — repeating on every IDLE event for as long as the
  tool entry sat unconsumed.  No correctness bug (the no-op return
  prevents bad state) but a wasted thread spawn per IDLE.  Fixed by
  gating on ``has_pending(USER_DRAIN)``; tool-only queues no longer
  trigger the wake path.

* **Copilot (coordinator_idle_observer.py)**: docstring referenced
  the old module path ``turnstone.core.metacognition.IdleNudgeWatcher``;
  the class moved to ``turnstone.core.idle_nudge_watcher`` in q-3 of
  the apply-pass.

* **Copilot (nudge_queue.py)**: ``has_pending`` docstring cited
  ``ChatSession.deliver_wake_nudge_from_queue`` as its caller, but
  that method calls ``drain(USER_DRAIN)`` directly — no production
  caller used ``has_pending`` until this commit.  Updated to point
  at the now-actual caller (``IdleNudgeWatcher``).

* **github-code-quality (test_nudge_queue.py)**: false positive on
  ``test_channel_is_required`` — the no-channel ``q.enqueue("a", "1")``
  call is wrapped in ``pytest.raises(TypeError)`` to verify the
  validation contract.  No code change.

5571 non-live tests pass; ruff + mypy clean.

(cherry picked from commit 0fbf31e713)
2026-05-07 17:35:21 -07:00
Patrick Buckley 42bf9aecaf fix(metacog): apply-pass fixes from pre-push full-stack review
Round-2 review caught 11 confirmed findings on the 3-commit metacog stack;
this commit applies them.

* **bug-1 (major)**: Wake source tag was leaking onto real user messages
  flushed during a wake send.  ``_append_user_turn`` and ``send`` now
  take an explicit ``from_wake: bool`` parameter — only the wake's
  synthesized first turn passes True, so ``_flush_queued_messages``'s
  real user input no longer inherits the audit tag.  Regression test
  pins the contract.

* **perf-1 (major)**: ``CoordinatorIdleObserver._maybe_enqueue`` was
  issuing list_workstreams + visible_memory_count storage queries
  before the cheap cooldown gate could short-circuit.  New
  ``_cooldown_allows`` read-only peek runs first; storage queries only
  fire when cooldown actually allows the nudge.

* **q-1 (major)**: Added the missing coord-side integration test that
  exercises ``CoordinatorIdleObserver`` + ``IdleNudgeWatcher`` together
  in the production install order against a real ``SessionManager``,
  protecting the subscription-order contract from silent regression.

* **perf-2/3 (minor)**: Cap check moved above ``_last_assistant_used_wait``;
  ``_fire_counts`` restructured as ``dict[str, dict[str, int]]`` keyed by
  ws_id so the leave-IDLE existence check is O(1).

* **perf-4 (minor)**: ``NudgeQueue.drain`` fast-paths the all-match
  case (the common one for chat-loop drain seams) by swapping
  ``self._items`` directly instead of allocating a fresh ``kept``
  deque + per-entry append.

* **perf-5 (minor)**: Wake's synthesized empty user turn no longer
  writes a content-empty row to the conversations table — the
  ``_source`` audit tag isn't column-backed and the side-channel
  reminder is stripped before persist, so the row would carry nothing.

* **q-3 (minor)**: Split ``IdleNudgeWatcher`` + ``install_*`` /
  ``shutdown_*`` helpers out of ``metacognition.py`` into the new
  ``turnstone/core/idle_nudge_watcher.py``; metacog stays a
  static-template module.

* **sec-1 (nit)**: Widened ``_sanitize_child_name``'s control-char
  regex to cover Unicode bidi-overrides, zero-width chars,
  line/paragraph separators, BOM, and tag chars.

* **q-4/q-5 (nits)**: Docstring referenced the wrong peek primitive
  (``has_pending`` → ``len()``); ``_last_assistant_used_wait``'s
  ``session`` parameter now typed ``ChatSession``.

5571 non-live tests pass; ruff + mypy clean.

(cherry picked from commit 3f106f98b2)
2026-05-07 17:35:21 -07:00
Patrick Buckley 191775dd7e feat(metacog): coord idle-children nudge — observer + valid_until predicates
Adds the first concrete consumer of the wake trigger: when a coordinator
goes IDLE while interactive children are still running, a
``CoordinatorIdleObserver`` enqueues an ``idle_children`` nudge that the
``IdleNudgeWatcher`` then dispatches as a synthetic empty-user-turn
``send``.  The model receives a system-reminder body listing the active
children (capped at 6 inline + 32 in the suggested ``wait_for_workstream``
call) and a nudge to block on them rather than reply prematurely.

Observer gates (in order): coord-only filter, skip if last assistant
turn used ``wait_for_workstream``, per-(ws, nudge_type) hard cap (3)
that resets only on non-wake leave-IDLE, active-children query,
``should_nudge`` cooldown.  Console lifespan registers the observer
BEFORE the watcher so subscriber-fire order has the observer
enqueueing first on the same IDLE event.

Adds an opt-in ``valid_until`` predicate on ``NudgeQueue.enqueue``
(R9 from the design risk register) — drain re-checks the predicate
outside the queue lock; falsy / raising drops the entry without
delivering it.  ``deliver_wake_nudge_from_queue`` now drains inline
before synthesizing the empty user turn so a stale predicate-drop
doesn't leave the wake send with empty content; ``_attach_pending_user_reminders``
consumes the pre-drained reminders via ``_wake_drained_reminders``.

The observer's ``valid_until`` uses ``count_workstreams_by_state``
(boolean check, no row fetch) instead of full ``list_workstreams``,
keeping the chat-loop user-attach path off the heavy query.

User-controlled child workstream names are sanitized
(``_sanitize_child_name``) before interpolation so a name like
``</thinking>...`` can't steer the model's reasoning channels through
the rendered body — the wire-boundary ``escape_wrapper_tags`` only
covers ``<system-reminder>`` / ``<tool_output>`` envelopes.

(cherry picked from commit 908e67fe4f)
2026-05-07 17:35:21 -07:00
Patrick Buckley c41fd2be2e feat(metacog): wake trigger — IdleNudgeWatcher + ChatSession.deliver_wake_nudge_from_queue
Adds the third metacog channel: an out-of-band wake that converts a
workstream's IDLE transition into a synthetic empty-user-turn ``send``
when the session has any-channel nudges queued.  The ``IdleNudgeWatcher``
subscribes to ``SessionManager.subscribe_to_state``; on IDLE it dispatches
via ``session_worker.send`` with a no-op ``enqueue`` callback so a
busy-worker race silently drops without spawning a competing worker.

Wake-source-tag plumbing on ``ChatSession`` short-circuits metacog
detection on the synthetic empty input, suppresses queue producers
during the wake's own tool dispatch, and stamps ``_source = "system_nudge"``
on the synthetic user-message for audit / replay distinction.  The tag
is saved / restored across ``_dispatch_pending_watch`` so watch chains
recursing off the wake are processed as normal user turns rather than
inheriting the wake's guards.

Generic ``install_idle_nudge_watcher`` / ``shutdown_idle_nudge_watchers``
helpers wire the watcher into both the interactive and coord lifespans
via a single ``app.state`` registry so both surfaces share the same
teardown contract.

Foundation for PR 3 (CoordinatorIdleObserver + idle_children formatter)
and PR 4 (watch dispatcher switchover).

(cherry picked from commit f0e7fea549)
2026-05-07 17:35:21 -07:00
Patrick Buckley 1787fb5c11 refactor(metacog): unify advisory channels into pull-model NudgeQueue
Replaces the dual `_pending_user_advisories` / `_pending_tool_advisories`
list pair with a single channel-tagged `NudgeQueue` per session.
Producers tag entries with a channel ("user", "tool", or "any");
consumers drain by channel filter at their existing seams. Foundation
for the wake trigger (PR 2) and coordinator idle-children nudge (PR 3).

Existing nudges (start, correction, completion, denial, resume,
tool_error, repeat) keep their wire shape and drain timing — zero
behavior change. Cancel paths now `clear()` the unified queue.

(cherry picked from commit 94b3720916)
2026-05-07 17:35:21 -07:00
Patrick Buckley 814c42763d fix(mcp): asyncio.timeout (not wait_for) for safe-close-stack on Python 3.11
Python 3.11's ``asyncio.wait_for`` wraps its inner coroutine in a fresh
``asyncio.Task`` via ``ensure_future``. When the inner is
``stack.aclose()`` on an ``AsyncExitStack`` containing
``streamablehttp_client(...)`` (anyio cancel scopes entered in the
calling task), the fresh task's attempt to exit those scopes raises
``RuntimeError('Attempted to exit cancel scope in a different task
than it was entered in')``. Python 3.12+ rewrote ``wait_for`` to use
``asyncio.timeout`` internally — runs in the current task — so 3.13
ran the same code path successfully.

Symptom on 3.11: integration tests where ``session.initialize()``
returns 4xx (e.g., 403 insufficient_scope tests) hit
``_connect_one_pool``'s ``except Exception:`` handler →
``_safe_teardown_on_connect_failure`` → ``_safe_close_stack`` → cross-
task RuntimeError. The ``concurrent.futures._base.CancelledError``
that surfaces in ``future.result(timeout=...)`` is the cascade
fallout from the asyncio loop's exception handler reacting to the
unretrieved-task-exception.

Fix: use ``asyncio.timeout`` instead of ``asyncio.wait_for`` for the
5s aclose bound. Equivalent semantics, current-task execution, works
on 3.11+. The 5s guard against ``aclose()`` hanging on a broken stack
is preserved.

Verified on Python 3.11.14 (full suite 5427 passed) and 3.13.7 (full
suite 5427 passed); all 9 integration tests pass on both.

Pre-existing bug — surfaced only after the marker fix in 5c9850c
let CI's test (3.11) actually run the 4xx tests.

(cherry picked from commit f6a3b66ea4)
2026-05-07 17:35:21 -07:00
Patrick Buckley 242596ced3 fix(mcp): pool-reuse 401 — entry-owned carrier + race-and-cancel
Two pre-existing defects in the Phase 6 pool dispatch path that only
manifest when a pooled session is reused for a second dispatch:

1. The per-dispatch _AuthCapture allocated in _dispatch_pool was wired
   into the httpx response hook only at first connect (via
   _connect_one_pool). On a reused session no fresh connect runs, so
   the hook continues writing to the original-connect's carrier while
   the new dispatch inspects an empty carrier — auth_401/403 silently
   misclassified to "other", refresh-and-retry never fires.

2. Even with the carrier on the entry (so the hook writes to a stable
   reachable object), session.call_tool itself hangs forever on
   upstream 4xx for reused sessions. Trace: SDK's spawned
   handle_request_async raises HTTPStatusError, the outer
   streamablehttp_client TaskGroup cancels post_writer, post_writer's
   finally aclose's read_stream_writer, BaseSession's _receive_loop
   exits and enters its CONNECTION_CLOSED-fanout finally. anyio's
   send_nowait skips waiting receivers with pending_cancellation; the
   dispatch task (created by run_coroutine_threadsafe for the reuse
   case) is NOT in any cancel-scope chain, so the send "delivers" but
   the receiver's Event is set on stale state — receive() never
   wakes. Test 21 doesn't hit this because its 401 happens during
   initialize, in the same task that opens streamablehttp_client, so
   the cancel scope DOES propagate.

Fix:
- Move _AuthCapture ownership to PoolEntryState (and asyncio.Event
  alongside, allocated lazily on the mcp-loop). The hook closes over
  entry.auth_capture at first connect and stays valid across
  dispatches; reset under open_lock before each call_tool.
- Race session.call_tool against the carrier's fired_event in
  _dispatch_pool_with_entry. If the event wins (hook captured 4xx
  before SDK propagated), cancel call_tool and raise an internal
  _CarrierAuthSignal — _classify_failure resolves to auth_401/403
  via the carrier's status, the dispatcher evicts the broken
  session, and the cross-task retry handshake reconnects on a fresh
  bearer.

Adds tests/test_mcp_pool_auth_integration.py::test_integration_pool_reuse_401_refresh_and_retry_succeeds
which drives the reuse path through real upstream + real SDK and is
the structural gate against this class regressing. Negative-tested
twice: revert PoolEntryState.auth_capture → test fails (carrier
empty); revert the race → test times out (SDK hang).

Also drops the @pytest.mark.asyncio decorator (replaced with
@pytest.mark.anyio) on four tests in test_mcp_pool_auth_introspection.py.
The project depends on anyio's pytest plugin (anyio is in deps);
pytest-asyncio is NOT a project dep and CI's test (3.13) failed on
those four. Local pytest happened to pick it up via system Python.

Found via Copilot review on PR #481.

(cherry picked from commit 97086fc617)
2026-05-07 17:35:21 -07:00
Patrick Buckley bde0913442 feat(mcp): SDK 401/403 introspection via httpx response hook
Phase 6 of OAuth-MCP. Recovers upstream 401/403 from MCP servers via a
capturing httpx_client_factory: an async response hook records 4xx
status + WWW-Authenticate header into a per-dispatch carrier before
the SDK's post_writer swallows the underlying httpx.HTTPStatusError.

Splits _classify_failure into auth_401 (refresh-and-retry once) vs
auth_403 (parse insufficient_scope, emit mcp_insufficient_scope with
parsed scope set). The 401 retry runs on a fresh asyncio.Task via
run_coroutine_threadsafe in _dispatch_pool_sync, escaping the anyio
cancel-scope state of the prior dispatch's TaskGroup.

WWW-Authenticate parsing extracted to a new mcp_http_parsers module
with an RFC 7235 challenge tokenizer (replaces hand-rolled substring
scanners). Two-layer defense against multi-Bearer-challenge injection:
the hook uses get_list("www-authenticate")[0] to drop attacker's
second challenge, the parser truncates at challenge boundary as
belt-and-braces. Scope set capped at 32 entries before hitting the
audit row or the LLM-visible structured-error JSON.

Auth failures (401/403) never trip the per-server circuit breaker
(server-only breaker invariant). Static path remains byte-identical.
_PgRefreshLock untouched. Pool dispatch still reachable from the
agent loop only via Phase 7 catalog scoping; Phase 6 behaviour is
testable via direct call_tool_sync.

5557 tests pass. 33 tokenizer unit tests in tests/test_mcp_http_parsers
cover the RFC 7235 grammar + the scope/error wrappers + the 4 KB input
cap. 7 integration tests in tests/test_mcp_pool_auth_integration drive
real upstream 401/403 through streamablehttp_client + a FastMCP
subprocess fixture — the structural exit gate that makes
HTTPStatusError-injection-only unit tests insufficient.

(cherry picked from commit db9260d8c4)
2026-05-07 17:35:21 -07:00
Patrick Buckley 570b198f1b fix(man): accept canonical name(section) page notation
Models often emit page references in the standard man-page form
(``printf(3)``, ``open(2)``, ``perlfunc(3pm)``) rather than splitting
them into ``page`` + ``section`` args. The page-name sanitizer was
rejecting the parens as invalid input, killing the call. Parse the
section out of the page string before sanitization (explicit
``section`` arg still wins) and widen the section validator to accept
multi-letter suffixes like ``3pm`` / ``3perl`` that already appear on
real systems.

(cherry picked from commit 39a6b7b447)
2026-05-07 17:35:21 -07:00
Patrick Buckley 96d935f1f7 fix(mcp): cancellation-safe orphan-lock drain + lock-reorder + test integrity
Phase 5 PR #479 review fix-up. Three review rounds (bot + two internal
multi-stage /review) caught:

- _PgRefreshLock now allocates a per-instance ThreadPoolExecutor instead of
  a module-global single-worker one. The global shape preserved psycopg2
  thread-affinity but serialized every advisory-lock acquire on the node
  behind one thread, even for unrelated (user, server) keys.
- get_user_access_token_classified flips to `async with lock, pg_lock:` so
  concurrent same-key callers serialize on the in-process asyncio.Lock
  before allocating the pg_lock's per-instance executor + spin loop. N
  concurrent same-key callers collapse to one executor allocation.
- _drain_orphan_pg_lock no longer re-awaits the cancelled asyncio Future
  from `__aenter__`. It receives the underlying concurrent.futures.Future
  and re-wraps it via asyncio.wrap_future, getting an independent asyncio
  Future tied to the worker outcome. This way cancellation of the awaiter
  doesn't poison the drain's wait, and the drain genuinely waits for the
  worker to settle before deciding whether to call cm.__exit__.
- Module-level _pg_refresh_drain_tasks set holds strong refs to in-flight
  drains (asyncio's task set is weak — fire-and-forget tasks could be GC'd
  mid-cleanup; RUF006 hazard).
- Drain narrows except clauses to Exception so a drain-task cancellation
  records as cancelled instead of being silently logged as 'completed
  normally with no acquire'.

Test integrity (was a major finding in round 2 — old generator-based cm
let the test pass via GC finalization timing rather than drain logic):

- New _ObservableLockCm class-based context manager whose __exit__ is a real
  observable method (records call args + thread). Distinguishable from
  GeneratorExit thrown by GC of a generator-based cm.
- Strong external ref to the cm via created_cms list — keeps cm alive past
  the test's awaits, so a no-op drain genuinely fails the assertion rather
  than papering over via GC timing.
- Deterministic drain wait via _pg_refresh_drain_tasks gather — no
  fixed-duration sleeps.
- _run_cancel_scenario helper drops the duplicated setup between the two
  cancellation tests.

Negative-test verified: replacing _drain_orphan_pg_lock body with `return`
makes test_pg_refresh_lock_cancellation_releases_on_same_thread fail with
'drain did NOT call cm.__exit__ — orphan Postgres lock + open transaction'.

Other fixes: protocol docstring corrected to describe pg_try_advisory_xact_lock
spin + retry (was claiming pg_advisory_xact_lock blocking acquire);
get_user_access_token_classified docstring rewritten for new lock order;
narrow `except BaseException` -> `except Exception` in
test_mcp_user_pool.py concurrent-dispatch helper.

882 tests pass (MCP + auth + storage). ruff + mypy clean.

(cherry picked from commit 3eb9d22ad5)
2026-05-07 17:35:21 -07:00
Patrick Buckley 1a1043c4df feat(mcp): per-(user, server) ClientSession pool with OAuth dispatch
Phase 5 of OAuth-MCP — adds a per-(user, MCP-server) ClientSession
pool to MCPClientManager alongside the existing static-server path,
gated entirely on the per-server `auth_type='oauth_user'` config.

Pool architecture:
- `_user_pool_entries: dict[(user_id, server_name), PoolEntryState]`
  with lazy connect on first dispatch, per-key asyncio.Lock allocated
  on the mcp-loop, idle eviction coroutine (default 600s TTL, LRU cap
  200), and an `in_flight` counter as the eviction interlock so live
  calls can never be torn down mid-flight.
- `_dispatch_pool` runs the token-state machine: missing token →
  `mcp_consent_required`; key-rotation decrypt failure →
  `mcp_token_undecryptable_key_unknown` with NO consent prompt and NO
  auto-delete; expired token → silent refresh under per-(user, server)
  advisory lock; refresh failure → revoke + consent.
- `_classify_failure` separates transport (trips breaker) from auth
  401/403 (does NOT trip breaker — server-only invariant) from
  protocol (no breaker change).
- `entry.open_lock` held only across connect-or-reuse and released
  before the `await session.call_tool` so concurrent calls from one
  user against one server overlap (validated by Spike 1 scenario 2).

Auth-class failures are fail-soft in Phase 5: any 401/403 surfaced by
the SDK propagates to the agent as a tool error and the next dispatch
reconnects on a fresh refresh. Real introspection of upstream 401/403
is a Phase 6 concern — the MCP SDK's `streamable_http` post_writer
swallows `httpx.HTTPStatusError` upstream, so detecting status from
the response chain requires `McpError(CONNECTION_CLOSED)` payload
parsing or a custom httpx middleware around `streamablehttp_client`.
The mid-flight 401 refresh-retry path and the `mcp_insufficient_scope`
structured error for 403 step-up land together in Phase 6, gated by
an integration test that drives a real upstream 401/403 (the unit-
test injection of `HTTPStatusError` is what masked the production gap
on the first apply-findings pass — the integration test is the
structural gate so the gap can't reopen). RFC §1.5 steps 4-5 and the
phase table in §Implementation phases reflect this scope split.

Multi-node refresh contention:
- New `StorageBackend.acquire_advisory_lock_sync` Protocol method.
  SQLite returns nullcontext (single-node, in-process asyncio.Lock
  is sufficient). Postgres uses `pg_try_advisory_xact_lock` with
  retry on a fresh per-attempt connection, so waiters don't pin pool
  connections during the AS roundtrip. Inner try/except + nested
  finally ensures conn is always returned to the pool, even when
  begin / execute / yield / commit raises mid-body.
- Lock ordering: pg_advisory outer, asyncio.Lock inner. Re-read after
  lock collapses cluster-wide contention to one HTTP roundtrip per
  (user, server) per refresh window.
- `_PgRefreshLock` enter/exit pinned to a single-worker
  ThreadPoolExecutor so SQLAlchemy connection state stays
  thread-affine across cancellations.

Token storage refactor:
- `get_user_access_token_classified` returns a tagged TokenLookupResult
  (Token / MissingToken / DecryptFailure / RefreshFailed) so the
  dispatcher maps each state to the right user-facing error.
- `get_user_access_token` is now a thin wrapper around the classified
  variant; the previous duplicated state machine is gone.

Security:
- Pool dispatch + admin endpoints reject `http://` URLs for
  `auth_type='oauth_user'` servers (only exact loopback hostnames are
  exempt — `*.localhost` is intentionally NOT honored because RFC 6761
  localhost-zone resolution is configuration-dependent and could route
  bearers to non-loopback IPs via custom resolvers / hosts file /
  Docker overlays). Validated at three layers:
  `_dispatch_pool` (structured `mcp_oauth_url_insecure` error),
  `_connect_one_pool` (defensive ValueError), and
  `admin_create_mcp_server` / `admin_update_mcp_server` (400 before
  storage write).
- Admin URL change on an oauth_user row purges per-user OAuth tokens
  bound to the old URL: bearers are bound (via OAuth resource /
  audience) to the URL active at consent time, so silently rebinding
  them to a new URL is a token-binding violation. Re-consent forces
  fresh issuance for the new resource.
- Encryption-key fingerprints stay in audit logs only; no longer
  surfaced in agent-facing error payloads.

User_id thread-through:
- `MCPClientManager.call_tool_sync(..., user_id=None)` (additive;
  default None preserves the static path byte-identically).
- `ChatSession._exec_mcp_tool` passes `self._user_id or None`.
- `set_app_state(app_state)` setter wires OAuth state at lifespan
  startup, called from both turnstone-server and turnstone-console.

Performance:
- LRU cap eviction iterates `_user_pool_entries` (not
  `_user_pool_last_used`) so pre-dispatch entries are eligible.
- Eviction batch closes via `asyncio.gather` instead of serial await.
- `_resolve_pool_target` returns the resolved server row to
  `_dispatch_pool` to eliminate the second DB lookup.
- Production reachability of pool dispatch is gated on Phase 7
  (catalog scoping) wiring pool tools into `_tool_map`; until then
  pool dispatch is reachable only via direct `call_tool_sync` with a
  prefixed name (the path the new pool tests exercise).

Hardening parity preserved:
- Static path (auth_type ∈ {none, static}) byte-identical; PR #296
  hardening (SDK #2147 mitigations, anyio cancel-scope, stale-session-
  and-stack guard, server-only circuit breaker) intact.
- `test_reconnect_preserves_static_state_identity` unchanged + green.
- `MCPTokenStore.get_user_token` does not auto-delete on
  MCPTokenDecryptError (key-rotation safety).
- Notification debounce stays manager-level.
- Connect-failure cleanup factored into
  `_safe_teardown_on_connect_failure` shared by both connect paths.

Tests: 5475 → 5493 (+18). New file `tests/test_mcp_user_pool.py`
plus additions to test_mcp_oauth_refresh.py, test_mcp_admin_api.py,
and test_mcp_client.py covering: pool data structures, lazy connect,
eviction TTL + LRU + lock interlock, dispatch state machine (token
states), failure classification, http-rejection at dispatch and
admin layers, URL-change-purges-tokens (sec), concurrent dispatch on
one (user, server), pg_advisory lock parity, and user_id threading.

Phase exit criterion (synthetic load test 50 users × 3 servers × LRU
30 × 1000 calls × 200 evictions) deferred to a post-Phase-5 fitness
spike that runs against a staging deployment with real FDs and real
network behaviour, not a CI mock — same shape as Spike 1's
pre-Phase-0 SDK validation.

Out-of-scope for Phase 5 (Phase 6+): SDK-level 401 refresh-retry +
403 `mcp_insufficient_scope` (Phase 6), per-user catalog scoping
(Phase 7), consent UX SSE event + dashboard renderer (Phase 8),
admin UI status indicators (Phase 9).

(cherry picked from commit 4db7d9c6cf)
2026-05-07 17:35:21 -07:00
Patrick Buckley 55aab54774 test(mcp): SDK 1.27 concurrency spike for per-(user, server) pool
Spike artifact validating MCP SDK behavior before Phase 5 builds the
per-(user, MCP-server) ClientSession pool. Three scenarios, all pass:

1. N=20 concurrent ClientSession instances against the same URL — no
   FD blow-up, no shared transport state, each session's tools/list
   returns independently.

2. Two concurrent tools/call on a shared ClientSession with
   interleaving payloads — request_id demux works under contention.

3. Per-session Authorization header isolation across 5 sessions —
   httpx connection pooling does not cross headers between sessions,
   so per-session bearer tokens reach the server unmixed.

Outcome gates the Phase 5 architecture (lazy dict[(user_id,
server_name), ClientSession] + per-key asyncio.Lock + LRU eviction).
Had any scenario failed, the fallback was per-call header injection
(Alternative F in the OAuth-MCP RFC).

Spike-only — not collected by pytest. Run manually:

  uv run python tests/spike_sdk_concurrency.py

(cherry picked from commit e695a98c54)
2026-05-07 17:35:20 -07:00
Patrick Buckley 0f8c8b38a3 fix(mcp): pin OAuth return_url + sanitise read-scope status
Addresses ten findings on the Phase 4 OAuth-MCP commit: four from the
PR #478 review surface, plus six surfaced by a follow-up multi-stage
review of the first round of fixes. Two of the latter were genuine
security regressions in the very code that claimed to close those
holes.

Security
--------

- _validate_return_url now pins return_url same-origin against the
  configured oidc_config.redirect_base instead of request.url. Behind
  a permissive front proxy that did not normalise Host, an attacker
  could spoof Host and provide a matching absolute return_url to mint
  an open redirect off /api/mcp/oauth/start. Same fix pattern as
  PR #476 OIDC.
- Reject return_url values containing literal backslashes or starting
  with `//` up front. urlparse leaves backslashes inside `path`, so a
  value like `/\evil.example/foo` slipped through the path-only branch
  and became the protocol-relative `//evil.example/foo` after WHATWG-
  conformant browsers normalised the backslash — re-introducing the
  open redirect the same-origin pin was meant to close.
- internal_mcp_status (read-scoped) projects through a new
  _strip_server_status_for_read helper that drops the verbose `error`
  text and replaces it with a coarse `has_error` boolean. The error
  string is built as `f"{type(exc).__name__}: {exc}"` and so carries
  stdio binary paths (FileNotFoundError) or internal MCP URLs
  (httpx.ConnectError) — equivalent to leaking command/url, which
  this same patch deliberately strips. Approve-scoped refresh and
  reconnect callers continue to receive the full `error` text via
  the existing _strip_server_status helper.
- internal_mcp_status now returns the projected (sanitised) entries
  for every server in mcp_mgr.get_all_server_status() instead of
  emitting the un-sanitised dict that included `command` (stdio argv)
  and `url` (remote MCP endpoint). Sibling refresh/reconnect endpoints
  already used _public_server_status to strip these.
- internal_mcp_status docstring documents the trust boundary — server
  enumeration to read scope is intentional so dashboards can render
  per-server indicators; verbose error detail and command/url remain
  approve-scoped.

Correctness / UX
----------------

- _validate_return_url comparison normalises (scheme, host, port)
  before equality. Lowercases hostname and collapses the scheme's
  default port, so `https://App.Example.COM/x` and
  `https://app.example.com:443/x` are recognised as same-origin
  with `redirect_base = https://app.example.com` instead of being
  silently downgraded to the `/` fallback.
- mcp_crypto startup-gate error message now names both
  `mcp_token_encryption_keys` (rotation list) and
  `mcp_token_encryption_key` (single) so an operator using rotation
  isn't misled into thinking only the singular form is valid.

Cleanup
-------

- Delete the unused _KNOWN_TRUSTED_ENDPOINT_HOSTS legacy re-export
  shim in oidc.py (zero callers — a no-op that survived the Phase 4
  oauth_ssrf extraction). Sphinx :data: docstring reference at
  validate_discovered_endpoint updated to point at
  turnstone.core.oauth_ssrf.KNOWN_TRUSTED_OAUTH_ENDPOINT_HOSTS
  directly. The Google multi-origin allowlist is unaffected — it
  lives at the canonical name and is read from oauth_ssrf.py:164.
- test_mcp_oauth_handlers TestValidateReturnUrl imports
  _validate_return_url at module level instead of repeating the
  import inside each test method.
- test_server_lifespan_mcp_crypto replaces a fragile
  `messages.count("mcp_token_encryption_key") >= 2` substring trick
  with `re.search(r"mcp_token_encryption_key(?!s)", messages)` —
  asserts the singular form directly via negative lookahead.

Tests
-----

5448 pass (+13 vs the prior tip):

- TestValidateReturnUrl gains backslash-bypass, protocol-relative,
  default-port, uppercase-host, and explicit-port-mismatch cases
  alongside the original same-origin / cross-origin / scheme-
  mismatch / path-only cases.
- TestInternalMcpStatusEndpoint asserts the `error` text never
  reaches the read-scope wire (binary-path FileNotFoundError no
  longer appears anywhere in the rendered response) and that the
  coarse `has_error` boolean lights up correctly on the failed
  server.
- TestInternalMcpStatusEndpoint also pins the no-mcp-client path to
  `{"servers": {}}`.
- _routes_with_internal extended to include the
  /api/_internal/mcp-status route so the new tests can exercise it
  through TestClient.
- Existing test_startup_aborts_with_oauth_user_row_and_no_key
  strengthened to require both singular and plural key names appear
  in the error log.

(cherry picked from commit 62bbc332af)
2026-05-07 17:35:20 -07:00
Patrick Buckley b0f7029ff1 feat(mcp): per-(user, server) OAuth 2.1 + PKCE flow
Lands the OAuth flow that uses the token-at-rest store from the prior
commit: discovery (RFC 9728 PRM + RFC 8414 AS metadata with operator-
override precedence), PKCE S256 (mandatory — refuse AS without it),
RFC 8707 resource indicator on every authorize and token request,
RFC 7591 minimal one-shot dynamic client registration, authorization-
code exchange, refresh-token grant with re-read-after-acquire single-
flight lock, and the /v1/api/mcp/oauth/{start,callback} endpoints
mounted on both server and console.

Refactored:
- validate_url_no_ssrf, validate_discovered_endpoint, is_localhost,
  effective_port, sanitize_log_text moved out of oidc.py into a shared
  oauth_ssrf module; oidc.py re-exports for compatibility. The shared
  helpers also expose async wrappers (validate_url_no_ssrf_async,
  validate_discovered_endpoint_async) so OAuth-MCP discovery — invoked
  from async handlers — does not block the event loop on the
  synchronous socket.getaddrinfo call.
- MCPTokenStore.get_oauth_client_secret reader path added (the prior
  commit was write-only)
- Storage protocol gains create/pop/cleanup_*_mcp_oauth_pending_state
  and get_mcp_oauth_client_secret_ct (mirror OIDC pending-state
  pattern: SQLite BEGIN IMMEDIATE select-then-delete, Postgres atomic
  DELETE...RETURNING)

Refresh-grant correctness:
- When the AS omits refresh_token (RFC 6749 §6 — MAY rotate), the
  existing refresh value is preserved at the OAuth-flow layer rather
  than cleared, so production ASes (Google, Auth0 default, Okta) don't
  force re-consent every hour
- expires_in accepts int, float, str-with-decimal — earlier int-coerce
  through str() failed on float and silently dropped expiry tracking
- The refresh-grant `resource=` parameter (RFC 8707) is the canonical
  MCP server URL, not the audience. Audience and resource are distinct
  concepts; using audience as resource would mismatch the AS RS
  allowlist.

Audience handling:
- _validate_token_audience accepts str or tuple; the callback resolves
  accepted_audiences = {server_url, oauth_audience} and validates
  against the set, so Auth0-style ASes that honor `audience=` (not
  RFC 8707 `resource=`) issue tokens that pass audience-bound
  validation
- build_authorize_url emits both `resource=` (RFC 8707) and
  `audience=` (Auth0-style) per server config; comment documents which
  AS implementations need which form

Security hardening:
- redirect_uri pinned to oidc_config.redirect_base instead of the
  request Host header — closes the same Host-header injection PR #476
  fixed for OIDC. Both /start and /callback return 503 with operator-
  actionable hint when redirect_base is unset
- DCR registration runs under per-server asyncio.Lock with re-fetch
  inside the lock, so concurrent /start callers don't both register
  and overwrite each other's client_id (the second user's code is no
  longer rejected on callback)
- /callback error branch pops the pending state row before redirecting
  so a leaked state can't be replayed against a separately-obtained
  code in the 60s cleanup window
- WWW-Authenticate Bearer parser handles RFC 7235 quoted-string
  escapes (\" and \\) instead of the naive [^"]+ regex
- AS-controlled response bodies and error_description query params go
  through sanitize_log_text before reaching exception messages or
  audit details. AS error responses are parsed for the standard
  RFC 6749 fields (error, error_description, error_uri), each
  capped at 80 chars and run through redact_credentials to defend
  against ASes that echo the request body back into their error
  payload.
- oauth_as_issuer_cached is re-validated against the SSRF guard on
  read; on rejection the column is cleared and PRM rediscovery runs
- DCR / token-endpoint / refresh-endpoint response bodies cap at 64
  KiB (PRM/AS metadata cap stays at 256 KiB) so a hostile or
  malfunctioning AS can't exhaust client memory.
- oauth_client_secret operator input capped at 1024 chars at the
  admin-form boundary; longer plaintext rejected with 400.
- /start and /callback responses stamp `X-Frame-Options: DENY` so the
  redirected pages can't be framed by attacker sites.
- delete_user cascades to mcp_user_tokens and mcp_oauth_pending so
  user deletion no longer leaves dangling per-user OAuth state.
- Renaming or deleting an oauth_user MCP server purges per-user
  tokens and pending OAuth state for the previous server name
  (delete_mcp_oauth_rows_by_server_name). The OAuth tables key on the
  mutable server_name; without this purge, a future server with the
  same name (and an attacker-controlled URL) would silently rebind
  prior user tokens. A future schema migration will replace the
  server_name key with a server_id FK + ON DELETE CASCADE.
- get_user_access_token catches MCPTokenDecryptError (raised when no
  installed key can decrypt the row, e.g. after key rotation) and
  falls through to None so dispatch surfaces a re-consent rather than
  crashing.
- oauth_user MCP server rows are skipped in the static auto-connect
  path. Auto-connecting them at startup with empty headers fails the
  AS check and trips the circuit breaker; per-user tokens come online
  lazily once the user has consented.

Audit (mcp_server.oauth.* prefix):
- consent_started, consent_completed, consent_failed, token_refreshed,
  token_revoked, dcr_registered. _audit_event is async and wraps
  record_audit in asyncio.to_thread so the audit write doesn't block
  the event loop. resource_id on the audit row is the immutable
  server_id (PK UUID) so admin-driven server renames don't break
  event correlation; server_name is exposed in detail for cross-
  reference. dcr_registered detail.has_secret reflects whether the
  DCR-issued secret was actually persisted (the prior code reported
  has_secret=true even on persistence failure).
- _admin_mcp_action audits the immutable server_id, not the mutable
  server_name (which is what the column is — the table's PK was
  always server_id).
- All OAuth-flow log keys use the mcp_server.oauth.* prefix to match
  the audit-action taxonomy.

Lifespan close-order in turnstone.server and turnstone.console.server
is reversed (LIFO) — mcp_oauth → mcp_crypto → oidc — to match init
order.

Deferred until the upcoming per-user pool integration:
- Multi-node refresh-lock contention via pg_advisory_lock
- DCR re-register on token-endpoint 401 (the dispatch path surfaces
  those 401s)
- TTL-LRU caching of decrypted plaintext access tokens
- DNS-rebinding hardening (httpx Transport pin) — documented as
  limitation in oauth_ssrf module docstring

Tests: 7 new test files / ~85 new tests covering discovery precedence
+ PRM quoted-string parsing, PKCE round-trip, SSRF helper extraction,
authorize/callback handlers including 503-on-no-redirect-base + DCR
concurrency + JWT audience polymorphism + callback-error-pops-pending,
refresh single-flight lock, refresh resource-vs-audience regression,
decrypt-error fallthrough, _db_servers_to_config skipping oauth_user,
pending-state CRUD round-trip.

(cherry picked from commit 29c42c1427)
2026-05-07 17:35:20 -07:00
Patrick Buckley a4c335d7bf feat(mcp): token-at-rest encryption layer for OAuth-MCP
Phase 3 of docs/design/oauth-mcp.md. Adds the Fernet/MultiFernet wrapper,
[security] config loader with rotation support, MCPTokenStore CRUD facade,
typed MCPTokenDecryptError that maps to the RFC's mcp_token_undecryptable_
key_unknown class, and a startup gate that fails loud when auth_type=
'oauth_user' rows exist without a configured encryption key.

Crypto module (turnstone/core/mcp_crypto.py):
- MCPTokenCipher wraps cryptography.fernet.Fernet + MultiFernet for
  rotation; encrypt with first key, decrypt by trying each in order
- load_mcp_token_cipher_config reads [security] mcp_token_encryption_keys
  (plural list) or mcp_token_encryption_key (singular), validates each
  key is base64-decodable to exactly 32 bytes
- MCPTokenCipherConfig is repr=False with custom __repr__ that redacts
  raw key bytes (defense in depth against accidental log/traceback leak)
- _key_fingerprint produces an 8-hex-char SHA-256 prefix for audit
  attribution without exposing the key
- MCPTokenStore handles encrypt-on-write / decrypt-on-read for
  mcp_user_tokens and mcp_servers.oauth_client_secret_ct
- get_user_token MUST NOT auto-delete the row on MCPTokenDecryptError
  (test_get_user_token_with_wrong_key_raises_decrypt_error verifies
  the row stays intact across a key-mismatch read)
- initialize_mcp_crypto_state / close_mcp_crypto_state lifespan helpers
  shared between server and console

Storage protocol (5 new ciphertext-only methods):
- set_mcp_oauth_client_secret_ct (dedicated writer; deliberately NOT
  added to MCP_SERVER_MUTABLE so generic update_mcp_server cannot write
  the secret column)
- create_mcp_user_token, get_mcp_user_token,
  update_mcp_user_token_after_refresh, delete_mcp_user_token

Server + console lifespans (turnstone/server.py + console/server.py):
- after OIDC init, count auth_type='oauth_user' rows; if any exist and
  no encryption key is configured, log an actionable error and
  raise SystemExit(1)
- without oauth_user rows, missing key is fine (lazy validation; admin
  flip without restart returns 503 from the admin handler)
- app.state.mcp_token_cipher / .mcp_token_store populated when key
  configured; None otherwise

Admin handlers:
- _require_token_store_for_oauth_secret pre-mutation gate validates
  token_store availability and oauth_client_secret type BEFORE
  storage.create_mcp_server / update_mcp_server runs, so a 503 from a
  missing key never leaves an orphan row or partial-update state
- _apply_oauth_client_secret encapsulates the encrypt + audit write
  used after the storage mutation; rolled out across both create and
  update handlers
- 503 message references both mcp_token_encryption_key (singular) and
  mcp_token_encryption_keys (plural for rotation)
- non-string oauth_client_secret payloads (false / 0 / lists / dicts)
  are rejected with 400 instead of being str()-coerced
- when auth_type transitions away from oauth_user, the encrypted
  secret column is cleared in the same admin call (with audit), so
  flipping back doesn't silently resurrect a stale credential

Audit events (mcp_server.oauth.* per audit.py taxonomy; RFC's
mcp.oauth.* renamed for consistency):
- mcp_server.oauth.client_secret_set fired from admin handlers with
  cleared:bool and key_fingerprint
- mcp_server.oauth.token_decrypt_failure fired from MCPTokenStore
  .get_user_token when no installed key can decrypt; carries
  key_fingerprints_attempted

Tests: 35 new tests across test_mcp_crypto, test_mcp_token_store,
test_server_lifespan_mcp_crypto, plus 6 admin-API tests covering the
no-orphan-row, no-partial-update, secret-clear-on-transition, and
non-string-secret-rejection invariants. Suite at 5337 (Phase 3 added
~50 tests including the rebase-imported skill suite).

cryptography>=42 promoted from transitive (lacme[tls]) to direct dep
since the encryption layer is now core, not optional.

Phase 4 (OAuth flow) wires the actual callers; Phase 3 adds only the
crypto layer and is exercised entirely by tests.

(cherry picked from commit 7f132e7230)
2026-05-07 17:35:20 -07:00
Patrick Buckley 21663d1567 feat(mcp): oauth schema + minimum admin form
Adds the data model and admin UI surface required by the OAuth-MCP flow.
Phase 2 of the per-user delegation initiative.

Schema:
- migration 049 creates mcp_user_tokens (PK user_id, server_name) and
  mcp_oauth_pending (PK state, indexed by created_at)
- eight new columns on mcp_servers: auth_type ('none' / 'static' /
  'oauth_user', NOT NULL DEFAULT 'static') plus six oauth_* config
  fields and oauth_as_issuer_cached
- post-upgrade UPDATE normalises auth_type to 'none' for streamable-http
  rows whose headers are NULL/empty/'{}'; stdio rows are left at the
  'static' default (auth_type is HTTP-auth-only)
- _schema.py kept in lockstep with the migration so metadata.create_all
  and alembic upgrade produce identical shapes
- mcp_user_tokens / mcp_oauth_pending TypedDicts in _protocol.py for
  Phase 3/4 use (no CRUD methods yet)

Storage / API:
- create_mcp_server gains the eight kwargs across protocol + sqlite +
  postgresql
- MCP_SERVER_MUTABLE picks up auth_type and the six text oauth_* fields;
  oauth_client_secret_ct is intentionally NOT in the whitelist — Phase 3
  will own ciphertext writes via a dedicated method
- McpServerInfo + Create/Update Pydantic schemas extended; oauth_client_secret
  accepted as plaintext input but discarded (Phase 3 wires encryption)

Admin handlers:
- _parse_auth_type validates against {'none', 'static', 'oauth_user'} and
  rejects empty / unknown values; shared between create and update
- when auth_type changes away from 'oauth_user', the oauth_* config
  columns are explicitly nulled in the same UPDATE so the row stays
  consistent
- _clean_oauth_text caps text fields at 512 chars (URLs at 2048) to bound
  admin write surface
- _mask_mcp_secrets now masks oauth_client_secret_ct to '***' regardless
  of reveal=true (write-only field)
- audit detail dict redacts oauth_client_secret if present

Frontend:
- new "Multitenant Authorization" fieldset on the MCP-server modal with
  three radio buttons (None / Shared / Per-user OAuth 2.1)
- conditional OAuth subform: AS URL, registration mode (preregistered /
  dcr; cimd is future), client ID, client secret, scopes, audience
- secret input is autocomplete=off and never round-trips on edit
- audience auto-populates from the MCP server URL on blur
- headers textarea hidden and submitted as {} when auth_type is 'none' or
  'oauth_user' so flipping the radio cleans up server-side state

Tests: storage round-trip for the new columns, oauth_pending table smoke,
migration 049 upgrade/downgrade with stdio-vs-http normalisation, four
admin-API tests for auth_type validation and oauth_*-clear-on-flip-away.
Suite passes 5284 (matched pre-Phase-2 baseline 5267 + 17 new).

Stacks on Phase 0; no behavioural change for existing rows.

(cherry picked from commit d675b237a3)
2026-05-07 17:35:20 -07:00
Patrick Buckley c823156af5 refactor(mcp): consolidate per-server state into StaticServerState dataclass
Phase 0 of the OAuth-MCP RFC: prepare MCPClientManager for the per-(user,
server) session pool that lands in Phase 5, without changing static-path
behavior.

Two changes:

1. Hardening helpers _pre_close_streams and _tcp_probe rename their first
   parameter from `name` to `key`.  Type stays `str` for now; widening to
   `str | tuple[str, str]` happens in Phase 5 when callers actually pass
   tuples.  _safe_close_stack takes the stack directly and is unchanged.

2. The eleven parallel name-keyed dicts (_sessions, _per_server_stacks,
   _per_server_tools, _per_server_resources, _per_server_prompts,
   _supports_list_changed, _supports_resources, _supports_resource_list_changed,
   _supports_prompts, _supports_prompt_list_changed, _server_streams) are
   consolidated into _static_servers: dict[str, StaticServerState].  Server-
   level state (circuit breaker, notification debounce, last-error,
   db-managed, merged catalog maps, listener lists) stays on the manager,
   unchanged.

PoolEntryState is defined for Phase 5 use but no code instantiates it.  The
typed map declarations (dict[str, StaticServerState] vs dict[tuple[str, str],
PoolEntryState]) make accidental cross-keying lookups easier to catch.

PR #296 hardening preserved exactly:
- pre-close-streams atomic take-and-clear before stack teardown
- stale-session-and-stack guard at _connect_one top: both state.session and
  state.stack checked, cleared independently, entry preserved (not popped)
- transport-error session-eviction in dispatch sets state.session=None only,
  leaving stack/streams for the next connect-time guard sweep
- _safe_close_stack CancelledError suppression unchanged
- TCP probe before streamablehttp_client unchanged
- future.cancel() after TimeoutError in all sync bridges unchanged
- notification debounce stays manager-level (not migrated into the dataclass)

Refresh helpers (_refresh_server_tools/_resources/_prompts) snapshot
state.session into a local immediately after the None guard so concurrent
transport-error eviction during await cannot null the session reference
mid-call.

Tests: shared _seed_static_state helper in tests/conftest.py replaces eleven
direct dict mutations; new test_reconnect_preserves_static_state_identity
guards the entry-preservation invariant.  Pass count rises 5266 → 5267.

(cherry picked from commit be0950bb98)
2026-05-07 17:35:20 -07:00
Patrick Buckley bace928477 refactor(mcp): remove periodic refresh, add manual refresh/reconnect controls
Deletes the _periodic_refresh task and its supporting state
(_refresh_task, _refresh_failures, _refresh_backoff_until,
_REFRESH_BACKOFF_BASE/MAX, _DEFAULT_REFRESH_INTERVAL, refresh_interval
kwarg) from MCPClientManager. Push notifications and operator-driven
manual refresh now cover all catalog-update needs; the long-running
4-hour timer was dead complexity that obscured the per-user pool
work to come.

Catalog freshness on auto-reconnect is preserved by scheduling an
unblocking _refresh_server task on the mcp-loop after _connect_one
succeeds; the calling thread returns immediately so half-open
recovery latency does not double. Adds MCPClientManager.reconnect_sync
(clears the circuit, closes any existing session, calls _connect_one,
clears stale catalog on failure).

Wires a new pair of operator endpoints —
POST /v1/api/admin/mcp-servers/{name}/refresh and
/v1/api/admin/mcp-servers/{name}/reconnect — that fan out to all
nodes through the existing _internal route family, with per-row
"Refresh" and "Reconnect" buttons in the MCP Servers admin tab.
The new node-internal paths /api/_internal/mcp-{refresh,reconnect}/
are gated to the approve scope to prevent direct unprivileged
reconnects bypassing the console's admin.mcp gate. Internal
endpoints return generic error messages and a filtered status
payload (no command/url) to keep transport details admin-gated.

Drops the [mcp] refresh_interval setting, the
--mcp-refresh-interval CLI flag, and the matching config-mapping
entry; updates docs/architecture.md, docs/tools.md,
docs/settings.md, and the three PlantUML diagrams that referenced
the periodic loop.

Tradeoffs (intentional):
- Idle nodes will not auto-rejoin a recovered MCP server until
  traffic arrives or an operator clicks Reconnect. The previous
  background reconnection loop is gone by design — push
  notifications + operator controls replace it.
- Console fan-out blocks on the slowest node (existing pattern);
  not changed here.

This is Phase 1 of the OAuth-MCP series — feature subtraction
ahead of per-user state.

(cherry picked from commit eb2a119da9)
2026-05-07 17:35:20 -07:00
Patrick Buckley d16c911750 feat(skills): paste SKILL.md to auto-fill the Create Skill modal (#477)
* feat(skills): paste SKILL.md to auto-fill the Create Skill modal

When a user pastes an Anthropic-style SKILL.md (YAML frontmatter +
markdown body) into the Create Skill content textarea, the frontend
sniffs the leading ``---``, posts the raw text to a new backend parse
endpoint, and populates name / description / tags / author / version /
license / compatibility / allowed_tools from the parsed fields.  The
textarea is left with the body only (frontmatter stripped), and a toast
reports how many fields were set vs. kept (already-typed values are
preserved).

Backend
- ``POST /v1/api/admin/skills/parse`` (admin.skills permission) wraps
  the existing ``turnstone.core.skill_parser.parse_skill_md`` so admin
  imports and external installs share one parser.  ``ParseSkillRequest``
  / ``ParseSkillResponse`` schemas added; OpenAPI spec + sync/async
  console SDK methods updated.
- Hardening: 32 KiB cap on ``raw`` (Pydantic ``max_length`` + handler
  enforcement); ``Content-Length`` pre-check returns 413 before any body
  buffering; parse offloaded via ``asyncio.to_thread`` so deeply-nested
  YAML cannot stall the event loop.

Frontend (turnstone/console/static)
- New paste handler with optimistic paint (raw text shown immediately,
  textarea disabled + ``aria-busy`` flipped, hint switches to
  "Parsing...") so the round-trip is visible on slow networks.
- ``AbortController`` + generation guard (``_ctmPasteController``) so a
  fresh paste or modal close cancels a stale fetch — the previous
  handler's callbacks see the controller has been replaced and bail
  before touching the DOM.
- Non-destructive overwrite: ``_setSkillFormField`` returns "filled" /
  "skipped" / "absent" and refuses to clobber non-empty values.  Toast
  reports counts.
- Bumps ``#toast`` z-index above modal overlays (was 200 vs. modal 600
  — toasts fired while a modal was open were invisible).  Console-wide
  fix exposed by this being the first feature to fire toasts mid-modal.

HTML / CSS
- New ``.skill-paste-hint`` line above the textarea announcing the
  affordance, sized to match surrounding ``.label-hint`` text.
- ``aria-describedby`` ties the hint to the textarea; ``aria-live=
  "polite"`` announces the busy-state transition to screen readers.
- "Skill Content" heading hint reworded "system message — ..." →
  "available: ..." and the variables row label "Variables" → "Used"
  to disambiguate available vs. in-use template variables.

Tests
- 11 new cases in ``tests/test_skill_parse_api.py``: happy paths
  (full / minimal / nested-metadata / unquoted-colon recovery),
  malformed YAML 400, missing/blank/missing-name 400, RBAC 403, raw
  body 32 KiB cap (Content-Length pre-check), chunked-encoding bypass
  forces the application-layer cap.  Test pins ``raw_frontmatter``
  omission so a future ``dataclasses.asdict`` refactor can't silently
  leak the full YAML dict back to clients.

Validation
- 5146 / 5146 ``pytest -k "not live"`` pass.
- ``ruff`` + ``mypy`` clean on changed sources.
- ``node -c`` clean on governance.js.
- Two-stage code review (full pipeline + bug+quality re-review of the
  fix patches) applied; all confirmed findings addressed.

* fix(skills): Copilot PR #477 review fixes (cumulative bug-1, bug-2, q-1)

bug-1 (server.py): Content-Length pre-check was clamped to 32 KiB —
the same number as the per-string char cap on ``raw``.  A legitimate
``raw`` of exactly 32 KiB produces a JSON body well above 32 KiB once
the ``{"raw":"..."}`` wrapper and any escaping is added, so valid
near-max requests were 413'd.  New constant
``_PARSE_SKILL_MAX_BODY_BYTES = _PARSE_SKILL_MAX_CHARS * 4`` admits the
wrapper + multibyte expansion while still refusing obviously oversized
payloads early; the per-string ``len(raw)`` check stays authoritative.

bug-2 (governance.js): hideCreateTemplateModal aborted the inflight
paste controller and nulled the global, but the handler's ``.catch``
and ``.finally`` guard each DOM mutation behind ``_isCurrent()`` —
both bail when the controller has been nulled, leaving the textarea
``disabled`` + ``aria-busy`` and the hint stuck on "Parsing…".
Reopening the modal landed on a poisoned state.  The second-pass
review's q-2 cleanup that dropped the show-side defensive reset
missed this scenario — the verifier's reachability argument confused
"controller is null" with "UI state is reset"; the two are
independent.  Hide now resets the paste-induced visible state
alongside the abort.

q-1 (console_spec.py): error_codes for the parse endpoint listed only
400; handler also returns 413 for oversized bodies.  Added 413; kept
403 implicit per the convention sibling admin endpoints follow.

Test fixup: bumped the Content-Length test payload to 200 KB so it
clearly exceeds the new 128 KB pre-check threshold; otherwise it was
falling through to the per-string check and duplicating
test_oversized_raw_chunked_returns_413's coverage.

(cherry picked from commit 0a8083e6d5)
2026-05-07 17:35:20 -07:00
Patrick Buckley b8fadad94f fix(oidc): close transient client on disable paths + correct docstring
PR #476 review feedback (Copilot, oidc.py:584,616):

1. initialize_oidc_state's docstring claimed "on any failure
   enabled is False" but the JWKS-prefetch failure branch
   intentionally keeps enabled=True so the callback's lazy-fetch
   retry can recover from a transient IdP issue at startup.
   Docstring rewritten to spell out the three post-conditions:
   disable, JWKS-failure-keeps-enabled, success.

2. The long-lived httpx.AsyncClient was created up front, then
   three disable branches (discovery exception, discovery-returned-
   disabled, missing redirect_base) returned without closing it,
   leaving sockets held until shutdown.

   Restructured: discovery now uses a transient AsyncClient inside
   a context manager (closed at exit). The long-lived client is
   only created after the disable checks pass. The JWKS-failure
   branch still legitimately keeps the client open because the
   lazy-retry path needs it.

   The pre-existing single-client-passthrough test was replaced
   with three more specific tests: long-lived client only goes to
   fetch_jwks (not discover_oidc); discovery-exception path leaves
   http_client=None; missing-redirect_base path leaves
   http_client=None.

(cherry picked from commit b2153d907f)
2026-05-07 17:35:20 -07:00
Patrick Buckley 63aecdf2fa chore(oidc): consolidate test OIDCConfig helper + fix exceptions banner (cumulative q-4, q-5)
q-4: tests/test_oidc.py's _make_config and tests/test_oidc_handlers.py's
_make_oidc_config built the same OIDCConfig with sensible defaults but
had drifted — only the handlers helper set redirect_base. After b3
made redirect_base operationally required, every test_oidc.py test
that exercised redirect_base had to override it explicitly. A future
test could omit redirect_base and silently exercise the wrong
production path.

Moves make_oidc_test_config to tests/conftest.py with the more
complete handler-version defaults (including redirect_base). Both
test files import it under their existing local alias
(_make_config / _make_oidc_config) so the 60+ call sites in
test_oidc.py and the handler tests don't have to change.

q-5: section banner '# Exception' (singular) at oidc.py:79 became
inconsistent after b5 (callback robustness) added OIDCKeyNotFoundError.
Renamed to '# Exceptions'.

(cherry picked from commit 5d4a50d2cd)
2026-05-07 17:35:20 -07:00
Patrick Buckley cbe8940b30 perf(auth): migrate handle_auth_status to count_users (cumulative q-3)
The OIDC perf batch added storage.count_users() and migrated the two
OIDC handlers (handle_oidc_authorize, handle_oidc_callback) but missed
handle_auth_status — which still ran storage.list_users() then
len(users) > 0 for the same has-any-users gate.

count_users() is one COUNT(*) round-trip vs list_users() rehydrating
every row dict. Wrapped in asyncio.to_thread to match the OIDC handler
pattern; the async handler no longer blocks the event loop on storage
I/O for what's effectively an existence probe.

(cherry picked from commit 7c6bc22d02)
2026-05-07 17:35:20 -07:00
Patrick Buckley 1dcd1e2ec4 fix(oidc): serialise role-mapping concurrency + skip no-op write lock (cumulative bug-2, perf-1)
bug-2 (Postgres) — replace_oidc_roles read existing rows under default
READ COMMITTED with no row lock. Two concurrent OIDC callbacks for the
same user_id (racing token refreshes with differing claim sets) could
both observe the same baseline and produce a final role state matching
neither caller's intent. Adds .with_for_update() to the SELECT so the
existing rows for this user are locked for the duration of the
transaction.

The lock is per-user_id, not table-wide; unrelated user writes are
unaffected. Empty result sets acquire no locks, so a brand-new user
with no rows yet still allows two callers to proceed and merge via
ON CONFLICT DO NOTHING — that's a permissive race that self-heals on
the next reconciliation cycle, documented in code.

perf-1 (SQLite) — replace_oidc_roles took the SQLite global write
lock unconditionally via BEGIN IMMEDIATE before reading. Steady-state
re-logins (claims unchanged, no INSERT/DELETE needed) paid the lock
cost for nothing and serialised against unrelated writers.

Replaces with a double-check pattern: phase 1 reads under the default
deferred transaction (no write lock), computes the diff, and returns
(set(), set()) on no-op. Phase 2, only when mutation is needed,
commits the read txn, escalates to BEGIN IMMEDIATE, RE-READS, and
re-computes the diff under the lock before writing. The returned
(added, removed) reflects what was actually written, so caller logging
in apply_role_mapping stays truthful even when concurrent writers
shifted state between the two reads.

The OR IGNORE on insert is now defense-in-depth (the lock makes it
unnecessary) but kept as a safety net.

(cherry picked from commit d5087ef3b9)
2026-05-07 17:35:20 -07:00
Patrick Buckley 32e29ff255 docs(oidc): document TRUSTED_ENDPOINT_HOSTS + fix three-vs-four required drift (cumulative q-1, q-2)
The 8-commit OIDC stack added TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS
(operator allow-list for cross-host IdP discovery endpoints) and
promoted TURNSTONE_OIDC_REDIRECT_BASE to required, but the docs drifted
in two places:

q-1 — Troubleshooting > "OIDC not configured" still listed three
required env vars. An operator hitting the missing-redirect-base
startup error landed on a debugging entry that didn't mention the
variable they were missing. Fixed; added a separate troubleshooting
entry naming the exact log message produced by initialize_oidc_state
when redirect_base is unset.

q-2 — TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS was undocumented entirely.
Added a row to the env-var table and a new "Cross-host endpoints"
section explaining when the knob is needed (Google is the canonical
multi-origin IdP, but it's auto-handled; the env var is for any other
IdP whose discovery doc legitimately references hosts beyond the
issuer's origin). Added a troubleshooting entry pointing at the new
section.

(cherry picked from commit 3cf87628d2)
2026-05-07 17:35:20 -07:00
Patrick Buckley 3e2fe0bc9d fix(oidc): self-heal stranded user when role mapping fails post-create (cumulative bug-1)
If apply_role_mapping raised after create_oidc_user committed (transient
storage failure, race with role deletion, etc.), provision_oidc_user's
inline safety-net was skipped — and on retry the existing-identity
branch never reached the safety-net code, leaving the user permanently
stranded with zero roles.

Extracts _ensure_default_role(storage, user_id, desired_role_ids=None)
helper. Calls it on BOTH the new-user and existing-identity paths so a
user stranded by a transient failure recovers on next login.
desired_role_ids is a hint that lets the helper skip list_user_roles
when claim-driven mapping populated at least one role; the new-user
path was already paying that query, the existing-identity path now
pays it only when claim mapping returned an empty desired set.

Documents the admin-strip behavior in the helper docstring: stripping
all roles from an OIDC user no longer locks them out, since the next
login will re-grant builtin-viewer (assigned_by='oidc-default'). The
documented way to deny an OIDC user is to unlink their OIDC identity
via the admin endpoint, not to strip roles. The pre-fix behavior
(stripped user actually locked out) was the bug.

The 'oidc-default' vs 'oidc' assigned_by distinction is preserved:
apply_role_mapping's revocation lane only touches 'oidc' rows, so the
safety-net role survives every subsequent login regardless of claims.

Six new tests cover both paths, the hint short-circuit, the
list_user_roles fallback, the missing-builtin-viewer no-op, and the
self-heal regression case for already-stranded users.

(cherry picked from commit 1c41212f15)
2026-05-07 17:35:20 -07:00
Patrick Buckley 5f5eee4aab test(oidc): close coverage gaps + tighten fetch_jwks shape check (q-5, q-8)
q-5: _derive_username's UUID-retry tier (oidc.py:923-933) was untested.
  After perf-6 collapsed tier-2 to a single find_existing_usernames call,
  the only remaining tail was the 3-attempt UUID-retry loop and the final
  raise. New TestDeriveUsername class covers:
  - falls into UUID retry when all 10 suffix candidates are taken
  - UUID retry succeeds on the second attempt after one collision
  - UUID retry exhausted -> raises OIDCError

q-8: filled the unit-level coverage holes the multi-stage review flagged:
  - test_validate_id_token_retry_after_kid_rotation — direct unit test of
    the OIDCKeyNotFoundError path with real RS256 keys + JWKS rotation
    (previously only exercised end-to-end through the handler).
  - test_callback_uses_pending_audience_not_handler_audience — pins down
    the bug-3 fix by decoding the issued JWT cookie and asserting aud
    matches the audience stored at /authorize time, not the handler param.
  - test_apply_role_mapping_int_claim / _dict_claim — exercises the
    else: values = [str(claim_value)] branch for non-string non-list
    claim shapes.
  - TestFetchJWKS — non-200 status, non-dict body, dict-missing-keys,
    keys-not-list, transport network error.
  - TestExchangeCode network/4xx/5xx error tests (the non-dict-body case
    already shipped in batch 5).

Also a small production hardening that fell out of writing the
TestFetchJWKS::test_fetch_jwks_non_dict_body_raises test: fetch_jwks now
guards isinstance(result, dict) before result.get("keys"), matching the
shape-check pattern that discover_oidc and exchange_code already use.
A list/null body now surfaces as OIDCError("...not a JSON object") rather
than AttributeError leaking up to the lifespan.

(cherry picked from commit 5c11ab985f)
2026-05-07 17:35:20 -07:00
Patrick Buckley 6d532ed776 refactor(oidc): quality cleanup (bug-3, q-1/3/4/6/7/9/10/11/12/13)
Eleven small maintenance fixes; no behavior change beyond bug-3.

bug-3: pending.get('audience', audience) couldn't fall back because
  pop_oidc_pending_state always returns a dict with the audience key
  set verbatim from a non-null TEXT column. Replaced with
  pending.get('audience') or audience to cover the empty-string case
  defensively. Comment explains the security rationale.

q-1: extract _env_or_cfg_str / _env_or_cfg_bool helpers in oidc.py;
  load_oidc_config's six near-identical env-or-config blocks collapse
  to one-liners. role_map / trusted_endpoint_hosts / redirect_base
  retain bespoke parsing.

q-3: discover_oidc narrows except (httpx.HTTPError, ValueError, KeyError)
  with exc_info=True.

q-4: OIDC_STATE_TTL_SECONDS = 300 constant in oidc.py; auth.py imports
  and passes it explicitly. Storage signatures keep the literal default
  (storage layer doesn't know OIDC TTL semantics).

q-6: hoist runtime imports (OIDCError, OIDCKeyNotFoundError, exchange_code,
  fetch_jwks, provision_oidc_user, validate_id_token, build_authorize_url,
  generate_pkce_verifier) to module scope in auth.py. The genuine cycle
  is only oidc._derive_username -> auth.is_valid_username, kept
  function-scoped. test_oidc_handlers.py mock targets repointed to
  turnstone.core.auth.X to match the new binding.

q-7: comment + docs explain the 'oidc' vs 'oidc-default' assigned_by
  marker distinction.

q-9: OIDCIdentity / OIDCPendingState TypedDicts in storage protocol.
  Implementations construct via TypedDict syntax so mypy structurally
  verifies all required fields.

q-10: fetch_jwks narrows except (httpx.HTTPError, ValueError); docstring
  matches.

q-11: rename generate_pkce_pair -> generate_pkce_verifier; return only
  the verifier (build_authorize_url already recomputes the challenge).

q-12: extract _buildOidcRow helper in admin.js so future field additions
  go in one place.

q-13: OIDCConfig docstring lists startup-config vs discovery-derived
  field groups.
(cherry picked from commit bae4adca12)
2026-05-07 17:35:20 -07:00
Patrick Buckley c3d9cdae82 perf(oidc): batch perf hardening (perf-1..8)
Eight independent perf wins on the OIDC hot path:

perf-1: list_users() full-scan setup-gate replaced with new count_users()
  on both authorize and callback. Saves a full users-table fetch per login.

perf-2: handle_oidc_callback's sync DB chain wrapped in asyncio.to_thread
  for cleanup, pop_oidc_pending_state, count_users, and provision_oidc_user.
  handle_oidc_authorize gets the same treatment for count_users and
  create_oidc_pending_state. Event loop no longer blocks for the full
  callback duration on Postgres deployments.

perf-3: apply_role_mapping N+1 collapsed via new replace_oidc_roles
  storage method. One transaction handles the diff + insert + delete
  instead of 2N+1 commits per login. Returns (added, removed) so the
  caller can still emit per-role audit logs.

  The diff respects the documented invariant "manually-assigned roles
  are never touched" — desired_role_ids is filtered against rows where
  assigned_by != 'oidc' before computing added/removed. This prevents a
  PK conflict (Postgres lockout) or silent OR-IGNORE no-op (SQLite lying
  return) when admin-ui or oidc-default already holds the same role_id.

perf-4: provision_oidc_user no longer re-queries list_user_roles after
  apply_role_mapping. The new-user builtin-viewer fallback is gated on
  desired_role_ids being empty, which is information apply_role_mapping
  already returned.

perf-5: JWKS refetch dedup via asyncio.Lock on app.state. Both lazy-fetch
  (cold-start recovery) and rotation paths share the same lock with a
  double-check pattern: re-resolve kid against the current cache before
  issuing a new GET. N concurrent callbacks during rotation now produce
  at most 1 fetch.

perf-6: _derive_username's 9-suffix loop collapsed via new
  find_existing_usernames(candidates) -> set query. Worst case drops
  from 13 sequential queries to 1 + up-to-3 UUID-retry queries.

perf-7: cleanup_expired_oidc_states gated to once-per-60s per process
  via app.state.oidc_last_cleanup_monotonic. The pop already deletes
  the consumed row; the bulk cleanup is only relevant for abandoned
  authorize flows, so frequency was overkill.

perf-8: Long-lived httpx.AsyncClient stashed on app.state.oidc_http_client
  by initialize_oidc_state. discover_oidc/fetch_jwks/exchange_code accept
  an optional client= kwarg; when set, skip the per-call AsyncClient
  context-manager. New close_oidc_state lifespan teardown closes it.
  Tests pass client=None to keep the transient-client legacy path.

New storage methods (sqlite + postgresql):
- count_users() -> int
- find_existing_usernames(candidates) -> set[str]
- replace_oidc_roles(user_id, desired) -> (added, removed)

(cherry picked from commit 39a647f39c)
2026-05-07 17:35:20 -07:00
Patrick Buckley 366d316941 fix(oidc): callback robustness — typed exceptions, shape checks, log sanitize, JS race (bug-4, bug-5, bug-6, sec-4)
Four small hardening fixes on the OIDC callback hot path:

bug-4: JWKS rotation retry was matching the substring 'not found in JWKS'
  inside an OIDCError message. A future rephrasing would silently break
  key rotation. Adds OIDCKeyNotFoundError(OIDCError); validate_id_token
  raises the subclass at the kid-not-found site; handle_oidc_callback
  catches it explicitly. Other 'not found' errors in validate_id_token
  remain as plain OIDCError.

bug-5: tokens['id_token'] raised KeyError if the IdP returned 200 without
  id_token. exchange_code now rejects non-dict response bodies; the
  callback validates id_token shape (must be non-empty str) before
  passing to validate_id_token. Both raise OIDCError, surfaced as the
  standard 'Authentication failed' redirect.

bug-6: shared_static/auth.js — the OIDC error display raced showLogin's
  /v1/api/auth/status fetch via a 300ms setTimeout. showLogin now takes
  an optional oidcError parameter and paints it after _switchMode clears
  the error, in both the success and catch branches of the fetch.

sec-4: oidc.py exchange_code's non-200 OIDCError interpolated up to 500
  bytes of attacker-controlled IdP body, which then went to log.warning
  via 'OIDC callback failed: %s'. CRLF in resp.text could forge log
  lines. New _sanitize_log_text helper escapes control chars via
  unicode_escape and caps at the rendered length.
(cherry picked from commit 0af3adae1d)
2026-05-07 17:35:20 -07:00
Patrick Buckley 3e87f4262e fix(oidc): atomic user + identity provisioning to prevent orphan rows (bug-1)
provision_oidc_user previously called create_user (INSERT OR IGNORE
on SQLite — silent no-op on UNIQUE conflict), then create_oidc_identity
(also INSERT OR IGNORE), then apply_role_mapping which writes user_role
rows for the supposedly-new user_id. On a username TOCTOU race or
concurrent (issuer, sub) double-create, both inserts no-opped but
user_role rows were already written — leaving orphan rows pointing
at a user_id that doesn't exist.

PostgreSQL's create_user raised IntegrityError instead of silently
no-opping so it produced a misleading 'Authentication failed' error
without orphans, but the user-facing UX was equally poor.

Adds StorageConflictError to the storage protocol and create_oidc_user
that does both inserts in one transaction. Username collision and
(issuer, subject) collision both raise StorageConflictError, mapped
to OIDCError by provision_oidc_user. Crucially the new code does not
silently bind a colliding-username new identity to the existing user
— that would be an account-takeover vector. It raises.

SQLite uses BEGIN IMMEDIATE inside the try block so lock-contention
errors surface as StorageConflictError instead of leaking the raw
sqlalchemy OperationalError.

PostgreSQL relies on SQLAlchemy 2.x begin-on-demand semantics; the
explicit conn.commit()/rollback() in the catch block is the only
materialization path. Discrimination on PG uses
exc.orig.diag.constraint_name with message-substring fallback.

(cherry picked from commit 11618bb1d7)
2026-05-07 17:35:20 -07:00
Patrick Buckley f50b559792 fix(oidc): require TURNSTONE_OIDC_REDIRECT_BASE; drop Host-header fallback (sec-2)
_build_oidc_redirect_uri previously fell back to the request Host
header when redirect_base was unset. With a permissive reverse proxy
or direct backend access, a spoofed Host minted an authorize URL
pointing to attacker-controlled host — combined with a permissive
IdP redirect_uri allowlist this enables auth-code interception.

There is no production scenario where a Host-derived redirect_uri is
correct, so this fails closed:

- initialize_oidc_state checks redirect_base after discovery succeeds
  and disables OIDC (with an explicit error log naming the env var)
  if it's empty. Runs before fetch_jwks so a misconfigured deploy
  doesn't make a wasted JWKS call.
- _build_oidc_redirect_uri simplifies to f"{redirect_base}/v1/api/auth/oidc/callback".
  request parameter dropped; both call sites (handle_oidc_authorize,
  handle_oidc_callback) updated.
- docs/oidc.md promotes TURNSTONE_OIDC_REDIRECT_BASE from "Recommended"
  to "Required" with the security rationale.

(cherry picked from commit 52aba17740)
2026-05-07 17:35:20 -07:00
Patrick Buckley c6b3c0bc5f refactor(oidc): unify server+console lifespan via initialize_oidc_state (q-2, bug-2)
The OIDC discovery + JWKS prefetch block was duplicated byte-for-byte
between turnstone/server.py and turnstone/console/server.py. The bare
except branch in that block also left app.state.oidc_config unchanged
on unexpected exceptions — leaving the runtime with enabled=True and
empty endpoints, producing malformed authorize URLs.

Extracts initialize_oidc_state(app_state) into turnstone/core/oidc.py
which guarantees a coherent post-condition on every code path:
- discovery exception -> oidc_config replaced with enabled=False, jwks_data=None
- discovery returns enabled=False -> jwks_data=None
- JWKS prefetch fails -> jwks_data=None but enabled=True preserved (the
  callback's lazy-fetch retry path remains the recovery)
- success -> oidc_config + jwks_data both populated

Also hardens discover_oidc against non-dict discovery responses
(list/null/string/int) — previously these raised AttributeError out
of doc.get and propagated past the lifespan's bare except.

server.py and console/server.py lifespan blocks collapse to a single
await initialize_oidc_state(app.state) call.

(cherry picked from commit 6f9e140a41)
2026-05-07 17:35:20 -07:00
Patrick Buckley cefb74a226 fix(oidc): SSRF + plaintext credential exfil via discovery doc (sec-1, sec-3)
OIDC discovery-document endpoints (token_endpoint, jwks_uri,
userinfo_endpoint) were stored verbatim in OIDCConfig and later passed
to httpx without revalidation. Only the issuer URL was checked. A
hostile or compromised IdP could return token_endpoint pointing to an
internal IP (169.254.169.254, 10.0.0.0/8, etc.) and Turnstone would
POST the client_secret there.

Extracts the existing scheme/userinfo/SSRF check into
_validate_url_no_ssrf, adds validate_discovered_endpoint that runs the
same checks plus an issuer-binding check, and wires it into
discover_oidc for authorization_endpoint, token_endpoint, jwks_uri,
and userinfo_endpoint (when present).

Issuer binding accepts:
- Same (scheme, hostname, effective port) as the issuer.
- A hostname in _KNOWN_TRUSTED_ENDPOINT_HOSTS for the issuer (Google's
  multi-origin discovery is in the allow-map by default).
- A hostname in OIDCConfig.trusted_endpoint_hosts, settable via
  TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS env var or config.toml, for
  IdPs not in the static map.

Effective port comparison treats https://host and https://host:443 as
the same origin (urllib.parse.urlparse leaves the explicit form's port
as 443 and the implicit form's as None).

24 new tests cover the validator, the Google known-hosts path, the
operator allow-list, default-port equivalence, foreign-host
rejection, private-IP rejection, embedded credentials, and DNS
rotation between issuer check and endpoint use.

(cherry picked from commit 0df7dc026b)
2026-05-07 17:35:20 -07:00
Patrick Buckley 273d547f4e chore: bump version to 1.5.7 2026-05-04 03:04:09 -07:00
Patrick Buckley bbb404c363 feat(console): inline node picker replaces back-to-console banner (#475)
* feat(console): inline node picker replaces back-to-console banner

Drops the 32px banner the console proxy used to inject above proxied
server-UI pages and replaces it with an inline node-id pill in the
existing #ui-header.  Click the pill to open a dropdown that lists
healthy nodes (health dot, ws count, reachable/degraded/unreachable
text) plus a top-row link back to the console.

Reuses the .ws-tab-dropdown shell from ui/static/style.css for
animation, shadow, theme override, and item layout, so the picker
visually matches the workstream-tab chevron menu it sits next to.
Keyboard nav (ArrowDown/Up/Home/End/Tab/Escape) mirrors the chevron
menu's handler with cross-reference comments at both sites.
Lazy-fetches /v1/api/cluster/nodes against the console origin
(bypassing the prefix shim) on first open.

Reclaims 32px of vertical space, consolidates three separate
"you're on node X via console" indicators into one, and turns the
wayfinding chrome into a real cluster-nav primitive.

* fix(console): address Copilot review on node picker

- Request /v1/api/cluster/nodes?limit=1000 (collector's hard cap)
  instead of relying on the default 100 — clusters with more than
  100 nodes were silently dropping rows from the picker.
- Hand off focus to the first menu item after the async fetch
  resolves: openMenu()'s deferred focus hook ran while only the
  skeleton was in the DOM, so first-open keyboard users were
  stranded on the trigger until they pressed an arrow key.
- Tab now closes the menu without preventDefault, so focus moves
  to the next focusable element on the first press (ARIA APG menu
  pattern).  Escape still preventDefault + returns to the pill.
- Cap pill max-width at 240px and ellipsize the id span; node ids
  are accepted up to 256 chars upstream and could otherwise push
  the title and right-side controls off the appbar.  Pill carries
  a title attribute so the full id is still legible on hover.
2026-05-04 02:59:35 -07:00
Patrick Buckley 072113f7ca fix(session): properly inject queued user messages mid-loop (#474)
* fix(session): properly inject queued user messages mid-loop

Two queued-user-message bugs in ``ChatSession.send()``.

**Mid-tool-call: ``Unexpected role 'tool' after role 'user'`` on Mistral.**
The ``supports_tool_advisories`` capability flag (default False for
unknown openai-compatible models) routed cap-off providers down a
short-circuit branch in ``_collect_advisories`` that called
``_flush_queued_messages`` directly. That appended a ``user`` turn
between ``assistant(tool_calls)`` and ``tool``, which mistral-common's
``_validate_message_order`` rejects with a 400.

Drop the flag. All providers now run the unified path: queued user
messages become ``UserInterjection`` advisories that ride inside the
tool result envelope via ``wrap_tool_result``, splicing
``<system-reminder>`` text into the tool message's content. Role
sequence stays ``assistant → tool``. Live-confirmed on Mistral
medium and Qwen3 — both correctly distinguish system-reminder from
tool stdout in their reasoning.

**Mid-stream: queued message orphaned until next user send.**
After a no-tool assistant turn, ``_flush_queued_messages`` would
append the queued user message to history and the loop would
``break``, leaving the message at the tail of history with no
model response. Visible as "two sends to get one reply".

``_flush_queued_messages`` now returns ``bool``. The no-tool branch
``continue``s on drain instead of ``break``ing, so the model gets a
turn over the extended history.

Tests:
- ``test_collect_advisories_drains_text_queued_messages_to_persistent``
  pins the unified-path drain (text-only queue → ``UserInterjection``,
  no separate user turn appended to ``self.messages``).
- ``test_send_continues_when_messages_queued_during_streaming`` pins
  the loop-continue behavior (fails with 1 stream call pre-fix,
  passes with 2 post-fix).

* fix(session,ui): reject queued attachments + paperclip busy state

Copilot pointed out that the attachment-bearing branch in
``_collect_advisories`` had the same role-ordering bug as the
text-only path that 802658f fixed: an attachment-bearing queued
item would still call ``_append_user_turn`` mid-tool-call,
injecting ``user`` between ``assistant(tool_calls)`` and ``tool``.

Pragmatic fix: don't allow attachments to be queued at all.

**Backend.** ``ChatSession.queue_message`` raises a new
``AttachmentsNotQueueableError`` when called with non-empty
``attachment_ids``. The interactive ``/send`` route catches it,
releases reservations via the existing ``_release_reservation_on_fail``
hook, and surfaces ``status: "attachments_busy"`` to the caller
with the IDs in ``dropped_attachment_ids``. The coord adapter
mirrors the cleanup (releases the soft-locked reservation taken
for ``_send_id``) so the create-with-attachments path can't leak.

Now that the queue can never carry attachments, the per-item
``att_ids`` slot is gone:

- Queue tuple slimmed ``(cleaned, priority, att_ids)`` →
  ``(cleaned, priority)``.
- ``_flush_queued_messages`` collapses to a single combined-text
  user turn (no attachment branch).
- ``_collect_advisories`` queue-drain pushes ``UserInterjection``
  advisories only (no ``attachment_items`` list).
- ``dequeue_message`` no longer unreserves (queue can't reserve).
- ``_resolve_attachment_ids`` had no remaining production callers
  and is deleted along with the tests that exercised it in
  isolation.

**Frontend.** ``Composer.setBusy`` disables the paperclip whenever
busy (regardless of ``queueWhileBusy``) — text still queues,
attachments don't. ``chat.css`` gains a ``.composer-attach:disabled``
rule (mirrors the existing ``.composer-send:disabled`` treatment)
so the affordance actually looks unclickable instead of falling
through to the UA default. ``title`` and ``aria-label`` are kept in
sync for AT users (WCAG 4.1.2).

Both interactive and coordinator UIs handle the new
``attachments_busy`` response with a chat-surface error bubble:

> Attachments can't be sent while the assistant is working.
> Send a text-only message now, or wait and resend with attachments.

Chips stay in the composer so the user can retry once idle.

**Tests.** Replaced the now-impossible ``TestQueuedWithAttachments``
class with a rejection-coverage class. Rewrote the
``_queue_with_attachment`` route-test fixture to reserve directly
via ``reserve_attachments`` (the queue path no longer reaches the
reserved state). Added a route-level test for the new
``attachments_busy`` contract.
2026-05-04 02:59:35 -07:00
Patrick Buckley 7f1b0acf7a Bound search tool output against pathological inputs (#473)
* Bound search tool output against pathological inputs

Replaces the per-line truncation with a fully bounded pipeline so the
search tool can no longer overflow the LLM context — or OOM the parent —
on minified bundles, multi-GB JSONL records, or huge result sets.

Backend:
- Prefer ripgrep when on PATH; grep is the fallback. Detection is
  cached via functools.cache.
- ripgrep flags do most of the bounding natively: --max-columns 1024
  + --max-columns-preview, --max-filesize 10M, --max-count 100,
  --no-config, --no-messages, plus negative globs for the same
  noisy directories grep has been excluding.
- ripgrep added to the Dockerfile.

Streaming subprocess (_search_capture):
- subprocess.Popen with a streaming, byte-capped stdout read (4 MB).
  Defends against single-line files (training data, minified bundles)
  that would have OOM'd the previous subprocess.run capture.
- threading.Timer watchdog enforces tool_timeout even when the
  pipe read is blocked in the kernel — proc.wait(timeout=…) alone
  was insufficient because the read sat ahead of it.
- Stderr drained in a daemon thread to avoid pipe-deadlock when the
  child writes to stderr while we're still reading stdout. Cap on
  captured stderr keeps a hostile child from growing the buffer.

Tier-based formatter (_format_search_results):
- Tier 1: full path:line:content output, stream-emitted with a
  running-cost short-circuit so we never materialize past the budget.
- Tier 2: K samples per file with overflow notes; K is computed
  analytically from budget / file_count / avg-line-length so we hit
  the right ladder rung in a single pass.
- Tier 3: per-file counts only, also budget-bounded with a tail line
  reporting the omitted files. Sorted by descending count.
- Total output budget (32 KB) is well under tool_truncation, so the
  head+tail _truncate_output strategy never silently drops middle
  files in a search result.

Argument injection fix:
- The ripgrep arg list was missing the `--` separator that the grep
  branch already had. With auto_approve on the search tool, that was
  exploitable: path='--pre=COMMAND' would have made ripgrep run the
  script as a per-file preprocessor and surface its stdout. Added
  `--` and a regression test.

State-machine cleanup in _exec_search:
- rc < 0 (signal-killed by something other than us) now surfaces a
  dedicated 'killed by signal N' message instead of being parsed as
  success.
- capped + zero parsed records (e.g. one multi-MB line with no \n)
  now returns a dedicated byte-cap message instead of the malformed-
  output message that previously masked the real cause.
- _report_tool_result descriptions now match the returned payload
  (no more 'no matches' tag on a 'malformed' payload).

Defence-in-depth on env scrub:
- RIPGREP_CONFIG_PATH, GIT_CONFIG, GIT_CONFIG_GLOBAL, GIT_CONFIG_SYSTEM
  added to _EXPLICIT_SCRUB. We pass --no-config on the rg CLI today,
  but if a future caller forgets the flag, an attacker who can set
  one of these env vars could plant a config containing --pre=… and
  recreate the same RCE shape.

Tests:
- TestSearchLineTruncation rewritten to mock _search_capture instead
  of subprocess.run (the previous tests passed ChatSession kwargs
  that no longer satisfy the constructor).
- TestSearchBackendSelection covers rg/grep detection and arg
  construction, including the --pre flag-injection regression.
- TestSearchOutputBudget exercises Tier 1/2/3 directly.
- TestSearchCaptureStreaming spawns real Python subprocess writers
  to exercise the byte-cap trim, mega-line-no-newline edge case, the
  watchdog timeout when the child writes nothing, and the stderr
  drain under load.
- test_env_scrub picks up the new tool-config keys.

* Address Copilot review on #473

- Budget the Tier 2/3 header up front so the formatter's emission stays
  strictly within _SEARCH_OUTPUT_BUDGET. Previously the fit checks only
  counted body bytes, letting the final string overflow by ~120 chars
  (header + separator) and triggering _truncate_output's head+tail
  dropout — exactly the shape this code was trying to avoid.
- Restore the (5, 3, 1) ladder in Tier 2: the analytical K from perf-2
  is kept as a starting estimate, but if that K's actual emission
  doesn't fit (the estimate ignores the header and overweights shared-
  path compression) we step down through the ladder before falling
  through to Tier 3. The previous one-shot K could collapse to counts-
  only when 3/file or 1/file would have fit.
- Only normalise rc to 0 in the capped-output path when rc < 0 (our
  SIGKILL). There's a narrow race where the child can exit naturally
  between our read and our kill; preserving a non-negative rc means
  rg's rc=2 ('matches found but some files had errors') no longer
  silently turns into a clean success when the byte cap also fires.
- Clarify _MAX_SEARCH_LINE_LENGTH doc: the cap applies to the content
  portion (after path:lineno:), not the whole emitted line.
- Add explanatory comments on the two intentional `except Exception:
  pass` blocks in _search_capture (stderr drain, pipe close in the
  cleanup finally) so static analysis and future readers can see the
  silence is deliberate.
- Tighten the budget tests: now assert strict `<= _SEARCH_OUTPUT_BUDGET`
  instead of the +512-char slack that was masking the header overflow.
- New regression tests:
  - Tier 2 ladder step-down (K=5 over budget, K=3 fits, no Tier 3 fall-through)
  - capped + rc=2 surfaces stderr instead of being normalised to success
  - capped + rc<0 (our SIGKILL) flows through as a partial-result success

* chore(search): post-review cleanup

Follow-up to the Copilot-review fixes in 39d2aa2 — these are all small
quality items (no behaviour change, no new tests).

- q-1: collapse the Tier 2 candidates filter to a single expression.
  Drops the redundant inner ``max(estimated_k, 1)`` and the unreachable
  ``if not candidates`` branch (the ladder ends in 1 and ``estimated_k``
  is already floored at 1, so the comprehension always yields ≥ ``[1]``).
  ``or [...]`` is kept as defence against future ladder changes.
- q-2: update _format_search_results docstring to match the new ladder
  semantics (analytical seed → step down through (5, 3, 1) from the
  highest rung ≤ the estimate). The previous wording suggested every
  Tier 2 attempt started at 5.
- q-3: combine the two ``from turnstone.core.session import ...``
  statements in test_tier2_steps_down_ladder_before_falling_to_tier3
  into a single top-of-function import (matches the surrounding tests).
- q-4: shorten the explanatory comments on the two best-effort cleanup
  paths in _search_capture to one line each. Both sites now read with
  the same shape ("# best-effort: pipe may be torn down by ...").
- q-5: trim the _MAX_SEARCH_LINE_LENGTH comment from 7 lines back to 3.
  Keeps the load-bearing semantic (cap is on the content portion only)
  and the pathological-line defence; drops the paths-aren't-bounded
  parenthetical, which was background reading rather than WHY.
2026-05-04 02:59:35 -07:00
renovate[bot] 6904bd8f39 chore(deps): lock file maintenance 2026-05-04 02:59:35 -07:00
renovate[bot] 4d677d1ebf chore(deps): update github actions 2026-05-04 02:59:35 -07:00
Patrick Buckley 35f462a46d chore: bump version to 1.5.6 2026-05-03 13:44:50 -07:00
Patrick Buckley ec74334e74 feat(providers): api_surface toggle + mistral medium reasoning fix (#469)
* feat(providers): api_surface toggle + mistral medium reasoning fix

Mistral medium open-weights served by vLLM expects reasoning_effort via
the Responses API (`reasoning.effort`), not as a `chat_template_kwargs`
entry on Chat Completions.  The session was unconditionally injecting
`{"reasoning_effort": ...}` into `chat_template_kwargs` for every
openai-compatible request, which corrupted the prompt rendering for any
backend whose chat template didn't consume that key (Mistral medium,
Mistral cloud, Groq, OpenRouter).

Changes:
- Add `api_surface` ("chat" | "responses") to `ModelConfig.server_compat`
  and thread it through `create_provider` / `model_registry.get_provider`.
  `openai-compatible` defaults to Chat Completions; operators can flip
  individual aliases to Responses for endpoints that support it.
- New `vllm-mistral-medium` profile that pre-fills api_surface=responses
  on Detect for known Mistral medium model ids.
- Drop the unconditional `reasoning_effort` injection into
  `chat_template_kwargs`.  Operators running gpt-oss-style local
  templates that consume `reasoning_effort` from the chat template now
  opt in via `server_compat.extra_body.chat_template_kwargs`.
- New "API Surface" select in the Models admin tab; allowlist-validated
  server-side at create/update time; pre-filled by Detect via the
  profile suggestion.
- Evict the cached provider singleton in `ModelRegistry.reload()` when
  api_surface changes (previously only cfg.provider triggered eviction).
- Fix `_run_agent` fallback path to inherit the session's primary alias
  for capability and server_compat resolution; previously the fallback
  passed `alias=None`, which silently dropped per-model caps on the
  agent path.

Tests: 5117 passed (-m "not live"); ruff + mypy clean.

* fix(providers): don't auto-suggest Responses for Mistral medium

vLLM's Responses API surface for Mistral medium open-weights doesn't
wire up the Mistral tool-call parser as of vLLM 0.x — tool calls leak
into the response as ``[TOOL_CALLS]<name>{...}`` text instead of
structured tool_calls.  Chat Completions on the same engine handles
tools cleanly via ``--tool-call-parser mistral``, and reasoning can be
turned on via the vLLM CLI ``--reasoning-parser`` flag.

Drop the auto-suggest mapping so Detect falls back to the generic
``vllm`` profile.  Keep the ``vllm-mistral-medium`` profile definition
in place so an operator who specifically wants per-request effort and
accepts the tool-calling limitation can still pick "Responses API"
manually in the admin UI.

* fix(providers): address Copilot review on PR #469

- providers/__init__.py: drop the redundant *_responses_provider /
  *_chat_provider names; have create_provider use _openai_provider and
  _openai_compat_provider directly so they're not flagged as unused
  globals.
- console/server.py: tighten _validate_api_surface to a strict equality
  match against the canonical {"chat", "responses"} set.  The previous
  strip().lower() membership check accepted ' Responses '/'CHAT' but
  stored the raw string verbatim, which then failed to round-trip
  through the admin <select>.
- console/static/admin.js: gate the entire server_compat block (server
  type, api_surface, extra_body) on provider == "openai-compatible" at
  save time so toggling provider away can't leave a stale hidden surface
  selection in the persisted capabilities JSON.
- tests/test_session.py: splat the bad kwarg via **dict so CodeQL no
  longer flags the call as a wrong-name keyword (the point of the test
  is the runtime contract, not the static type).
- tests/test_admin_model_registry_refresh.py: add endpoint-level tests
  for the api_surface validation on both create and update — covers the
  bogus-value rejection, non-canonical-string rejection, and the happy
  path persisting through to the refreshed registry.
2026-05-03 13:40:30 -07:00
Patrick Buckley 2fd0c29a92 fix(memory): query-aware candidate selection + OR-of-terms search (#468)
* fix(memory): query-aware candidate selection + OR-of-terms search

The system-message memory composition path used a recency-ordered
candidate set (`_list_visible_memories(limit=fetch_limit)`).  On
deployments with more than `fetch_limit` (default 50) visible
memories, BM25 only ever ranked the 50 most-recently-touched memories
— a relevant memory written months ago was silently invisible
regardless of how well it matched the recent context.  Multi-word
search at the SQL layer used AND-of-terms, killing recall on any
multi-word query without an exact field overlap.

## Functional changes

- `_init_system_messages` (`turnstone/core/session.py`): extract
  recent context first, then `_search_visible_memories(context)` to
  pull query-aware candidates.  Search hits below `fetch_limit` union
  with the recency list (deduped by memory_id) so the BM25 candidate
  pool is always a SUPERSET of the prior recency-only pool — even on
  noisy queries where the cap fills with stopwords, the recency-50
  the original bug surfaced still reaches BM25.  Empty context skips
  search entirely.  Candidate-selection logic extracted into
  `_select_memory_candidates`.

- `search_structured_memories` (PostgreSQL + SQLite): per-term
  clauses join with OR instead of AND.  A row matches if ANY term
  matches ANY of name/description/content.  Downstream BM25 narrows
  back down by relevance.

## Perf hardening

- Collapse the 1-3 fanned scope queries into a single SQL.  New
  backend methods `list_visible_structured_memories` /
  `search_visible_structured_memories` union the visibility scopes
  into one WHERE OR-group, so a composition rebuild now hits the DB
  at most twice (search + recency) instead of up to six times.

- Cap and normalize search terms.  Composition can hand a multi-KB
  pasted message to ILIKE-based search; without a cap, every distinct
  token would emit one unindexable predicate per scope-fanned query.
  `normalize_search_terms` (`storage/_utils.py`) de-dupes
  case-insensitively, drops <2-char tokens, and hard-caps at 16.

- Per-turn search cache.  `_init_system_messages` fires from many
  call sites within one turn (state transitions, MCP refresh, tool
  results) and the recent-context query is identical across them.
  Session-instance cache keyed by (query, mem_type, limit) absorbs
  the duplicates; invalidated in `_append_user_turn` and after
  memory save/delete tool actions.

- Stable secondary sort by `memory_id`.  `updated` is second-precision
  and `touch_structured_memories` can land a batch on identical
  timestamps; without a tie-breaker SQL returns rows in
  implementation-defined order, BM25 input shuffles, and the
  LLM-side prompt cache misses across calls.  All four backend ORDER
  BYs now break ties on `memory_id ASC`.

## Quality cleanups

- Coalesce `memory.search.term_count` + `memory.search.zero_results`
  into a single `memory.search` log carrying both `term_count` and
  `result_count`.
- New `memory.composition` log: source / candidates / injected.
- Promote a shared `make_chat_session` factory to `tests/_helpers.py`.
- Rename SQL builder local `extra` -> `scope_filters` for clarity.
- Add docstrings on `search_structured_memories` so the AND->OR flip
  survives future readers.

## Tests

Adds 20 tests across `tests/test_structured_memory.py`,
`tests/test_structured_memory_storage.py`, and
`tests/test_memory_relevance.py`: recency-ceiling regression,
empty-query fallback, sparse-match union, recency-preserved-when-
search-returns-noise (locks in the pool-superset invariant),
OR-of-terms on both backends, scope filtering preserved,
search-facade multi-word behavior, term-cap normalization, the new
visible-scope helpers (list + search + empty-scopes guard),
coord-scope composition isolation, end-to-end
`memory(action='search')` tool execution, per-turn cache hit +
invalidation, and stable ordering under tied `updated` timestamps.

Memory test sweep: 102/102.  Broader regression
(session, storage, coordinator, load_skill): 411/411.

* fix(memory): address Copilot review on PR #468

Three follow-ups from Copilot's inline review:

1. SUPERSET invariant violation (Copilot, session.py:5510).
   `(search_hits + extra)[:fetch_limit]` capped the union back down to
   fetch_limit, evicting the recency tail when search added distinct
   hits.  Recency tail is exactly where ancient-but-recently-touched
   memories live — the recall this PR is supposed to improve — so
   tail eviction recreated the bug for the narrow case where a query
   term fell off the 16-cap and the matching memory sat in
   recency[40-49].  Drop the cap; both halves are already SQL-capped
   at fetch_limit, so the union is at most 2 × fetch_limit (~100 with
   defaults).  BM25 over 100 candidates in pure Python is sub-ms;
   irrelevant recency fillers get score=0 and don't pollute ranking.
   Updates the docstring to actually be honest about the invariant.
   Adds `test_recency_tail_preserved_when_search_adds_distinct_hits`
   that locks the behavior in: 5 search hits + 10 recency = 15-item
   pool, every recency item present, source="union".

2. Unbounded `query.split()` in normalize_search_terms (Copilot,
   _utils.py:74).  `str.split()` allocates the full token list before
   the cap-after-16 break, so a 100KB pasted query did MB of throwaway
   work even though only 16 tokens entered SQL.  Switch to
   `re.finditer(r'\S+', query)` — streaming iterator, stops scanning
   at the first 16 normalized terms regardless of input size.

3. Misleading + unbounded log term_count (Copilot, session.py:8571).
   `len(item["query"].split())` had two problems: same unbounded
   split as #2, and the value reported the raw input token count
   rather than the normalized term count that actually hit the SQL
   WHERE clause — misleading metric for an operator trying to
   understand storage-side behavior.  Switch to
   `len(normalize_search_terms(item["query"]))` — accurate count, and
   bounded for free via #2.

Refuted: github-code-quality flagged `...` bodies in the new Protocol
methods as "statement has no effect."  False positive — `...` is the
canonical Protocol body convention, used 213 other times in the same
file.

Memory test sweep: 103/103.  Broader regression: 411/411.
2026-05-03 13:40:30 -07:00
Patrick Buckley 1207d27363 fix(tests): isolate metrics-singleton swaps so they don't leak across files
CI failure on main: test_publish_records_metric_outcome saw an empty
calls list — its monkeypatch was patching a different metrics
instance from the one `_publish_models_metadata` reads.

Two changes:

- test_close_reason_persistence.py: replace the bare
  `srv_mod._metrics = MetricsCollector()` assignment in `_make_app`
  with an autouse `monkeypatch.setattr(srv_mod, "_metrics", ...)`
  fixture so the test's metrics swap auto-restores. Other test
  files (test_auth.py, test_server_attachments_endpoints.py) carry
  the same anti-pattern; left for a follow-up since they're not on
  the critical path here.

- test_server_node_models_metadata.py: switch the publish-helper
  metric test to a string-form `monkeypatch.setattr("turnstone.
  server._metrics", FakeMetrics())` so it replaces whatever binding
  the live module currently holds, regardless of what other tests
  did to it. Robust against future leaks of the same shape.
2026-05-03 13:40:29 -07:00
Patrick Buckley 733c9818d4 feat(coord): expose healthy model aliases per node on list_nodes (#466)
* feat(coord): expose healthy model aliases per node on list_nodes

Surfaces a `model_aliases` field on each `list_nodes` row so a
coordinator can discover which model aliases each cluster node will
accept on `spawn_workstream(model=...)` without an HTTP fan-out.

Each server projects its registry into a `models` entry on
`node_metadata` (`{alias, provider, healthy}` per alias) at lifespan
startup, on every 30s heartbeat tick, and after `internal_model_reload`.
The publish helper short-circuits on a payload-equality cache so a
stable cluster doesn't pay UPSERT churn — exposed via the new
`turnstone_node_models_publish_total{outcome="written|skipped"}`
Prometheus counter so operators can graph cache hit-rate.

Coord client filters the per-alias rows to healthy aliases only and
drops the provider-side model identifier (`cfg.model`) — coords kept
reaching for it when they should pass the local alias.

* fix(coord): address Copilot+CodeQL feedback on list_nodes models work

- internal_model_reload: reuse a single get_storage() local across the
  registry load and the metadata publish (Copilot:3047)
- _collect_node_models_metadata: iterate sorted aliases so two
  structurally identical registries built in different insertion orders
  serialize to the same JSON — directly improves the publish-cache hit
  rate exposed via turnstone_node_models_publish_total (Copilot:3105)
- tests: drop mixed turnstone.server import style flagged by CodeQL —
  hoist _metrics into the from-import block, and use sys.modules in
  the shutdown-race regression test instead of `import as srv`
2026-05-03 13:40:29 -07:00
Patrick Buckley 28a2779c10 fix(core): scope rehydrate fallback to manager, fix resume orphan
Address Copilot feedback on PR #465:

1. The has_alias fallback in both session_factories silently rewrote
   any unknown caller-supplied alias to the default, including on the
   fresh-create path where the create handler maps the factory's
   ValueError to a 503 with operator-friendly text. A typo in
   body.model would now silently start a workstream on the default
   instead of telling the caller their requested model could not be
   resolved. Move the fallback out of the factories: each factory
   raises again on unknown aliases, and SessionManager filters stale
   aliases out of the rehydrate path via a new ``model_validator``
   constructor kwarg (production wiring passes ``registry.has_alias``
   on both interactive and coordinator).

2. ChatSession.resume()'s elif branch flipped self.model to the
   persisted model name even when the alias was unresolvable, leaving
   the session paired with the constructor's default provider/client
   but a removed model name — a broken state whose next API call
   fails. Drop the model copy: keep the constructor's coherent
   default (provider + model + capabilities) and just log the
   unreachable saved values so the missing alias is auditable.

Tests:
- Move stale-alias coverage from the factory level into
  SessionManager (tests/test_session_manager.py): validator drops
  stale aliases before reaching build_session; live aliases pass
  through unchanged.
- tests/test_sessions.py renamed test_resume_restores_model →
  test_resume_keeps_defaults_when_alias_unresolvable to match the new
  contract.
2026-05-03 13:40:29 -07:00
Patrick Buckley 56364b0b5b fix(core): preserve workstream model + config on rehydrate
SessionManager.open() was calling build_session(ws) without a model
arg on the rehydrate path. The session_factory then resolved the
*current* default alias, ChatSession.__init__'s _save_config() (INSERT
OR REPLACE per-key) clobbered the persisted workstream_config with
those defaults, and the subsequent resume() "restored" what was now
the default — silently resetting model_alias, model, temperature,
reasoning_effort, max_tokens, skill, creative_mode, instructions,
token_budget, and notify_on_complete on every reopen and every
service restart, for both interactive and coordinator workstreams.

Three layers:

1. SessionManager.open() now reads workstream_config via
   self._storage.load_workstream_config(ws_id) and threads the saved
   model_alias into build_session(ws, model=saved_alias).

2. ChatSession.__init__ now skips its initial _save_config() when a
   workstream_config row already exists for self._ws_id — protects
   every other persisted knob without having to plumb each one
   through the adapter signature, and catches any future construction
   path that forgets to thread model through build_session.

3. Both session_factories (server.py interactive, console
   session_factory.py coordinator) now treat an unknown caller-
   supplied alias the same as an unset alias: fall back to the
   runtime default rather than raising. Without this, a workstream
   pinned to an alias an operator has since removed from the registry
   would 500 on every reopen — defeating the "best effort restore,
   default if the original is gone" contract this fix is meant to
   deliver. Mirrors _effective_default_alias's existing has_alias
   guard against a stale ConfigStore default.
2026-05-03 13:40:29 -07:00
Patrick Buckley afb5804a7c fix(console): address Copilot feedback on Models → Roles sub-tab
Three changes from PR review:

- Permission gating: hide the Roles sub-tab button when the user
  lacks ``admin.settings``.  The sub-tab loads/saves through
  ``/v1/api/admin/settings``, so an admin with ``admin.models`` but
  no ``admin.settings`` would otherwise see a perpetual 403 loader.
  When Roles is the active sub-tab and the permission check fails,
  snap the panel back to Definitions so the user lands somewhere
  usable.

- Drop the redundant ``/v1/api/admin/model-definitions`` fetch from
  ``loadAdminModelRoles``.  Both entry points (initial Models-tab
  open + ``models_changed`` SSE refresh) flow through
  ``loadAdminModels`` first, which already populates ``_modelDefs``
  + ``_modelDefaultAlias``; ``_saveModelRole`` doesn't touch model
  definitions, so the cached snapshot stays accurate when the save
  chains back here.  Halves the per-render request count and
  removes a wasted round-trip on every cluster-wide model edit.

- Add ``test_models_changed_event.py`` covering the SSE fanout the
  prior commit introduced: each model-definition CRUD endpoint
  emits exactly one ``models_changed``, settings PUT/DELETE only
  emit for keys in ``_MODEL_AFFECTING_SETTING_KEYS`` (parametrised
  over all eight), and unrelated settings (e.g.
  ``session.retention_days``) don't trigger spurious refreshes.
  The expected key set is pinned in the test so a stray addition
  to the allowlist doesn't silently bypass coverage.
2026-05-03 13:40:29 -07:00
Patrick Buckley 4b508a1319 feat(console): add plan_agent + task_agent to Models → Roles
Same shape as the coordinator/judge rows already there: alias dropdown
+ reasoning_effort dropdown sourced from the existing
``model.plan_alias`` / ``model.plan_effort`` and
``model.task_alias`` / ``model.task_effort`` settings.  Adds the four
keys to the SSE ``models_changed`` allowlist so changes from the
Settings API also trigger a live dropdown refresh, and filters them
out of the Settings tab so they only render in one place.
2026-05-03 13:40:29 -07:00
Patrick Buckley 9c2cb185e1 feat(console): consolidate role-model settings + live-refresh dropdowns
Lifts judge and coordinator model assignments out of their respective
admin tabs and into a new Models → Roles sub-tab so role overrides live
next to the model definitions they reference. Forward-looking shape for
the upcoming perception.{audio,image,video} model settings — adding a
new role is one entry in the declarative MODEL_ROLES array.

Also drops the misleading "Coordinator subsystem not configured" home
banner. The session factory already falls back to the registry's
default model when coordinator.model_alias is unset, so the banner was
nagging on fresh installs where the system was actually working. The
related _probeCoordSubsystem / _homeCoordReady plumbing went with it.

Wires SSE-driven live refresh: the console now emits a models_changed
event when a model definition is created/updated/deleted/reloaded, or
when a model-affecting setting (model.default_alias, judge.model,
coordinator.model_alias, coordinator.reasoning_effort) changes.
Connected browsers refetch /v1/api/models on receipt so the home
composer's model dropdown and the Roles sub-tab stay accurate without
a manual reload — fixes the case where editing the underlying model
for an existing alias left the dropdown showing the old model id.

Companion cleanups:
- Renamed .judge-section-* CSS classes to .admin-subtab-* and shared
  them with the Models sub-tab switcher (same a11y attrs, arrow-key
  nav). Old names had no other callers.
- Filtered judge.model out of the Judge Settings sub-tab and
  coordinator.model_alias / coordinator.reasoning_effort out of the
  Settings tab — they live exclusively under Models → Roles now.
- Reworded the _require_coord_mgr 503 messages to point operators at
  the Models tab instead of suggesting they set coordinator.model_alias.
2026-05-03 13:40:29 -07:00
Patrick Buckley 0519b847bd docs(skills): add import-conversation-history SKILL.md
Source-agnostic guide that teaches an agent Turnstone's destination
contracts (workstream + conversations schema, ws_id routing, OpenAI
message shape, tool-call/result pairing, provider_data fidelity blob,
attachment lifecycle) so it can map any external chat export onto them.
Validated against turnstone.core.skill_parser.
2026-05-03 13:40:29 -07:00
Patrick Buckley bbc8b99a9f fix(console): home composer attachments + coord chat user-message pills (#462)
* fix(console): home composer attachments + coord chat user-message pills

Two parity gaps in the console's coordinator surface:

- The embedded creator on the home page accepted only text — the
  paperclip / paste / drop pipeline that the in-coord composer and the
  interactive new-ws modal both expose was missing, so a user couldn't
  attach files at create time. Stage Files in memory (no ws_id yet) and
  ship them multipart on Start; the coord create endpoint already accepts
  multipart via create_supports_attachments=True.

- User messages with attachments rendered as plain text on both live
  send and history replay — no chip cluster like the interactive pane.
  Added appendUserMessageWithAttachments and a structured userAttachments
  list built from _attachments_meta (preferred) or the multipart parts
  themselves, then rendered the same .msg-user-attach pill strip the
  interactive pane uses.

Polish from a designer pass:

- Pill background was --panel-2, equal to the .msg bubble background in
  both themes (border contrast ≈1.4:1, below WCAG 1.4.11). Switched to
  --panel so the pill sits on a different surface than the bubble.
- Capped chip filename width inside the home composer (max-width 200px +
  ellipsis) so a long filename doesn't push the strip past the textarea.
- aria-live="assertive" → "polite" on #home-coord-error; client-side
  validation isn't an interrupt-level event.
- Reserved min-height on .home-composer-error and dropped the
  display: none/block toggling so validation messages no longer reflow
  the active-coordinators list below.

* fix(console): address PR #462 review feedback

- Block home-composer submit when files are staged but the task field is
  empty.  Server's _coord_create_post_install short-circuits on an empty
  initial_message, so the multipart upload would create pending
  attachment rows that never reserve onto a turn — orphaned until the
  GC sweep.  Fail in the browser instead.
- Drop the redundant `part &&` guard in coordinator.js's history-replay
  multipart loop; the earlier `if (!part || ...) continue` already
  filtered.
- Rewrite the home-mount .composer-chip-name CSS comment.  shared/chat.css
  defines .composer-chip{,-size,-remove} but no .composer-chip-name rule
  — the span inherits the parent chip font with no width cap.
- Add smoke-guard string assertions in test_coordinator_page.py for
  appendUserMessageWithAttachments and msg-user-attach so a future
  rename can't silently regress the attachment affordance.
2026-05-03 13:40:29 -07:00
Patrick Buckley bd9f780b21 chore: bump version to 1.5.5 2026-05-01 14:09:38 -07:00
Patrick Buckley 5d14b5f675 fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration (#461)
* fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration

Loading a saved workstream silently dropped tool results and missed
verdict / output-guard / truncation signals on replay. Root cause was
in `Pane.prototype.replayHistory`: an assistant message carrying both
content and tool_calls cleared the `lastToolBlock` anchor before the
following tool-result iteration could attach. The fix reorders content
to render before the tool block (matching live SSE order) and
restructures the tool-result branch to anchor by `data-call-id` so
multi-tool batches render `[hdr A][out A][hdr B][out B]` rather than
bunching outputs at the bottom.

Beyond the bug, replay now reaches near-parity with the live UX:

- Persisted intent verdicts and output_assessments flow through both
  the SSE replay (`_build_history`) and the `/history` REST endpoint
  used by coord. Single shared helper module owns the wire shape.
- Memory/recall calls persist instead of being filtered at storage
  time — full audit trail; UI dims them by default with hover-reveal
  so heavy memory usage doesn't crowd the narrative.
- Truncation indicator surfaces as a sibling pill (consistent across
  interactive + coord) when a tool result hit the 2000-char cap.
- `replayHistory` wraps DOM work in `aria-busy` so screen readers
  don't get a chatty announce-flood on long replays.
- `_build_history`'s storage I/O moves off the event loop via a new
  `events_replay_prepare` async hook for the SSE path; other async
  callers wrap in `asyncio.to_thread`.

Coord parity:

- `/history` REST endpoint decorates tool_calls with verdict +
  output_assessment + truncation flag (was previously raw
  `load_messages` output).
- Coord JS stamps `judge_verdict` / `heuristic_verdict` from
  history-loaded `tc.verdict` so the existing batch render paints
  the persisted pill, seeds the verdict cache to dedupe later live
  SSE events, and emits an inline `.coord-tool-row-warning` chip
  per call instead of a generic chat line.
- Memory/recall dim rule mirrored on `.coord-tool-row[data-tool-name=...]`.

* fix(replay): address PR #461 review feedback + raise tool-result storage cap

Copilot review feedback:

- Sibling-chain dim rule (memory/recall) now adds :focus-within
  alongside :hover for .tool-output / .media-embed / .output-warning
  / .tool-output-truncated — keyboard users tabbing into a faded
  subtree now get full opacity.
- ``cfg.open_post_load`` is now invoked via ``await asyncio.to_thread``
  so its sync ``_build_history`` call (storage I/O for verdict
  indexes + message reconstruction) doesn't block the event loop on
  every workstream open. Mirrors the SSE replay path that's already
  protected via ``events_replay_prepare``.
- Replaced the hardcoded ``2000`` literal in server.py and session.py
  with ``TOOL_RESULT_STORAGE_CAP`` from the shared decoration module
  so the UI truncation-pill detection can't silently desync from the
  storage write side.

While here:

- Raised ``TOOL_RESULT_STORAGE_CAP`` from 2000 → 10000. A 2000-char
  clip routinely cut grep / file-read bodies mid-line, leaving the
  audit trail useless for retrospective debugging. FTS5 + row size
  grow proportionally; the per-tool upper bound is still bounded
  upstream by ``_truncate_output``'s context-budget clamp.
- Updated the user-visible truncation-pill tooltip on both
  interactive and coord to reflect the new cap.
- ``test_decorates_tool_calls_and_marks_truncated`` now references
  the constant instead of a literal so it stays correct on future
  cap changes.
2026-05-01 14:06:08 -07:00
Patrick Buckley 4693fa95f1 chore: bump version to 1.5.4 2026-04-30 23:51:06 -07:00
renovate[bot] c3423d6606 chore(deps): update ghcr.io/astral-sh/uv docker tag to v0.11.8 2026-04-30 23:48:53 -07:00
Patrick Buckley 53f1222c22 refactor(coord): remove priority queue + queue depth indicator + broken CSS
Speculative reliability machinery from the Stage 3 push that turned
out not to address any user-visible bug. The actual fixes (state /
activity disjunction in handleChildState, bulk-fetch race fix in
_fetch_live_block, push approve_request via cluster bus) are what
resolved the wedged-row issues. Manual testing showed the per-tab
SSE listener queue depth never climbed past single digits even when
rows were stuck — overflow was never the cause.

Removed
- ``_CRITICAL_EVENT_TYPES`` + ``_put_with_priority`` helper.
- Per-tab listener queue selective drop (back to plain
  ``contextlib.suppress(queue.Full)`` everywhere).
- ``ClusterCollector._fanout`` reverts to the same.
- WebUI ``_broadcast_intent_verdict`` / ``_broadcast_approval_resolved``
  / ``_broadcast_approve_request`` revert to plain ``put_nowait``.
- ``_queue_stats`` periodic SSE emit + frontend status-bar indicator
  + the supporting CSS rules.
- Broken ``.approval-block`` ``transition: max-height`` /
  ``max-height: 80vh`` / ``overflow: hidden`` rules — the transition
  never fired (nothing toggled max-height) and ``overflow: hidden``
  clipped long verdict reasoning. Layout-shift on auto-expand jumps
  again, which is preferable to clipped content (Copilot review).

Tidied
- ``_CollectorProtocol`` / ``_ManagerProtocol`` method bodies switch
  from ``...`` ellipsis to docstring-only bodies, silencing four
  CodeQL "statement has no effect" warnings without changing the
  Protocol contract.

5024 passed, ruff + mypy clean.
2026-04-30 23:48:53 -07:00
Patrick Buckley 802d87a57f feat(coord): Stage 3 SessionManager Children primitive lift + cluster bus push paths
Lift the Children primitive out of CoordinatorAdapter into universal
SessionManager core primitives, replace the fragile poll + state-event
piggyback paths with first-class cluster bus event types for inline
approval delivery, and clean up the resulting frontend reducer.

Architecture
- New `turnstone/core/children_registry.py` — universal parent → children
  + reverse-lookup primitive with atomic `add_child` (returns parent UI
  for race-free dispatch). Lifted from `CoordinatorAdapter`.
- New `turnstone/core/child_source.py` — `ChildSource` Protocol with
  `SameNodeChildSource` (in-process via SessionManager state observer)
  and `ClusterChildSource` (cross-node via ClusterCollector listener).
- `SessionManager._on_state_change` upgraded to multi-subscriber
  (`subscribe_to_state` / `unsubscribe_from_state`) under a dedicated
  lock; CLI consumer migrated.
- `CoordinatorAdapter` shrunk: 731 → ~640 LOC. Children data lives in
  the registry; fan-out lives in ClusterChildSource. Backward-compat
  property facades dropped; tests updated to use the registry surface.

Cluster bus event vocabulary
- New event types `intent_verdict`, `approval_resolved`,
  `approve_request` flow through both `ClusterCollector._apply_delta`
  (translation from node SSE) and `emit_console_ws_*` (synthesis on
  console pseudo-node).
- `CoordinatorAdapter._dispatch_child_event` re-emits as
  `child_ws_intent_verdict` / `child_ws_approval_resolved` /
  `child_ws_approve_request` on the parent coord's SSE stream.
- New `_broadcast_intent_verdict` / `_broadcast_approval_resolved` /
  `_broadcast_approve_request` no-op hooks on `SessionUIBase`. WebUI
  pushes to the global queue; ConsoleCoordinatorUI pushes to the
  collector. `approve_tools` calls `_broadcast_approve_request` right
  after setting `_pending_approval` so the items reach the coord tree
  immediately, eliminating the bulk-fetch race.

Cleanups
- `pending_approval_detail` piggyback on `ws_state` / `cluster_state`
  removed end-to-end. Bulk fetch + explicit verdict / approve-request
  push are the canonical carriers.
- Browser `_judgePollTick` 90-second poll loop deleted; push path is
  authoritative.
- `urgent` flag on `scheduleLiveFetch` deleted (only caller was 409
  retry; replaced with `invalidateLiveBadge` + standard schedule).
- Console `_fetch_live_block` derives `pending_approval` from a
  disjunction (`activity_state="approval"` OR `state="attention"`
  OR detail present) so the bulk fetch can't return false during the
  state-transition race window.
- Coord-side merge guard in `flushLiveFetches` no longer clobbered:
  `handleChildState` only stamps `sseUpdatedAt` when authoritatively
  clearing detail.
- `child_locality` capability flag removed (was inert dead code).

Reliability
- Selective drop on listener queue overflow: critical event types
  (verdicts, approvals, ws_closed, child_ws_*) evict one oldest item
  to make room rather than dropping themselves on a full queue.
  Best-effort events (state ticks, content tokens, status, activity)
  drop as before. Applied to `SessionUIBase._enqueue`,
  `ClusterCollector._fanout`, and the `WebUI._global_queue` puts in
  the new broadcast hooks.
- `_state_subscribers` snapshot under a dedicated lock so concurrent
  subscribe / unsubscribe during dispatch can't shift the iterator.

UX / a11y
- Loading placeholder in renderChildRow keeps row height stable while
  the bulk fetch is in-flight (sr-friendly aria-label).
- Focus preservation across `_renderChildrenNow` (capture +
  restore by row + marker) and across targeted `_updateChildRow` swaps.
- Layout-shift transition on the approval block max-height; respects
  `prefers-reduced-motion`.
- Sidebar pending count: `(N children · M pending)`.
- Risk pill `aria-label` spells out level + confidence for SR users.
- Per-coord SSE listener queue depth surfaced in the status bar
  (`queue N/500`) with color escalation (warn at >50%, danger at >80%).

Tests
- 305+ test changes across 8 files. New unit tests for
  `ChildrenRegistry`, `ChildSource` (both impls + multi-subscriber
  observer), the new collector emit + apply_delta cases, the dispatch
  cases for new event types, the broadcast hook overrides on both
  WebUI and ConsoleCoordinatorUI, and the focus / placeholder /
  pending-count frontend assertions in `test_coordinator_page.py`.

5024 passed, ruff + mypy clean.
2026-04-30 23:48:53 -07:00
Patrick Buckley 8349d9994d feat(console): multi-select delete UX for Saved Coordinators (#458)
* feat(console): multi-select delete UX for Saved Coordinators

Mirror the per-server "Saved Workstreams" multi-select delete onto the
console's "Saved Coordinators" section.  Coordinator deletes go through
the existing routing proxy at POST /v1/api/route/workstreams/delete
(body-keyed by ws_id, since coordinators live on the node that owns
them) — no backend change required.

Pagination caps the visible page (and therefore the Select-All fan-out)
at 24.  Without it, a Select-All on a busy cluster would pin the
console proxy pool with hundreds of parallel deletes through the
fan-out router.  While in delete mode the saved-coordinators list is
frozen against SSE re-renders so visible cards don't shuffle out from
under the user's selections (drained on cancel / post-delete close).

Refactor: shared logic now lives in turnstone/shared_static/cards.{css,js}.

  * .ws-delete-* CSS moved out of ui/static/style.css into the shared
    sheet alongside .dashboard-card; the existing ui/static modal
    markup picks up class hooks instead of id-scoped rules.
  * createSavedCardsController() owns mode state, checkbox decoration,
    toolbar wiring, focus trap, modal lifecycle, and batch fan-out.
    Both ui/static (Saved Workstreams) and console/static (Saved
    Coordinators) instantiate one controller; ui/static is now ~300
    LOC lighter as a result.
  * Internalises stale-selection prune across SSE re-renders, the
    wsId->item lookup map (was O(selected x N)), and the aria-hidden
    wrap on the toggle button's emoji glyph.

Designer review tightened the affordance:

  * Modal close restores focus to the toggle button (was landing on
    <body>) — WCAG 2.4.3.
  * Modal [role="alert"] gets a red-chip treatment when populated,
    stays invisible at rest via :not(:empty).
  * Pagination consolidated onto the existing .pagination control
    (terse "X / Y" label + arrow-glyph buttons) instead of a parallel
    .coord-pagination treatment.
  * Filled destructive buttons darkened to #dc2626 in dark theme so
    the white label clears WCAG AA contrast (was 3.0:1 on --red).
    Light theme keeps --red unchanged (5.9:1 already passes).
  * Toolbar wraps below 700px viewport — Delete Selected drops to its
    own full-width row underneath count + Cancel + Select All for
    thumb-target separation.
  * .ws-card-check:focus-visible outline + word-break on
    .ws-delete-item for narrow-modal long aliases.

* fix(cards): address Copilot review feedback on PR #458

* closeModal focus restore now falls back to the section toggle button
  (opts.buttonId) when prevFocus is hidden or detached.  The post-delete
  Close path runs cancel() before closeModal(), which puts the bar at
  display:none — so the captured prevFocus (the bar's "Delete Selected"
  button) is no longer focusable and focus would land on <body>,
  defeating the WCAG 2.4.3 fix.  Esc / Cancel paths still land on the
  original focus owner because the bar stays visible in those flows.

* Saved Coordinators onClose drains _savedCoordsRetry before reloading.
  Without it, SSE events that arrived during the delete-mode freeze
  leave the retry flag true, so loadSavedCoordinators's .finally()
  re-fires a second fetch immediately after the first resolves.  Mirrors
  the same idiom in cancelCoordDeleteMode.
2026-04-30 23:48:53 -07:00
Patrick Buckley ac1fd67137 chore: bump version to 1.5.3 2026-04-30 13:34:35 -07:00
Patrick Buckley 1b40ae79f9 fix(storage): address PR #457 review feedback
Three issues from the Copilot review on PR #457:

1. SQLite race in bulk_close_stale_orphans (Copilot): the SELECT-then-
   UPDATE flow doesn't re-apply the eligibility predicates on the
   UPDATE, so a row that gets touch_workstream-bumped (or set_state-
   transitioned) between the two statements would still be flipped
   to closed.  Postgres dodges this via UPDATE...RETURNING (one atomic
   statement); SQLite needs the explicit re-application.  Fix: rebuild
   the WHERE conditions list once, apply on both SELECT and UPDATE,
   then SELECT-back by ``state='closed' AND updated=now`` to get the
   accurate closed-id list.  A row that became fresh between the two
   statements skips the UPDATE entirely.

2. SQLite IN-clause bind-parameter limit (Copilot): default 999 cap
   could be exceeded on a backlog reap (e.g. after a long outage).
   Chunked the candidate id list at 500 — same chunk size
   prune_workstreams (line 453) uses for the same reason.

3. Wall-clock-dependent test asserts (Copilot, two locations): the
   tests asserted ``updated > '2024-01-01T00:00:00'`` which is fragile
   on systems with skewed clocks or pre-2024 dates.  Replaced with
   ``updated != stale_seed`` — captures the same intent (the value
   was bumped) without depending on wall-clock date.

Two ``...``-as-no-op flags from github-code-quality were false
positives — ``...`` is the standard Python idiom for Protocol method
bodies and matches every other method in _protocol.py.  No code change.
2026-04-30 13:34:15 -07:00
Patrick Buckley b078ddccf0 fix(session_manager): scope orphan reaper by services.last_heartbeat
Replaces the ``node_id == self_node_id`` orphan-scoping heuristic from
earlier on this branch with liveness-based scoping using
``services.last_heartbeat``.  The heuristic was wrong for the post-#384
world: PR #384 (refactor: replace hash-ring rebalancer with rendezvous
hashing) deleted the rebalancer that used to keep workstreams.node_id
pointing at a live node.  Without it, ``workstreams.node_id`` is now
stamped at create time and never updated, so in containerized
deployments with dynamic hostnames a dead pod's rows have ``node_id``
matching no surviving service — they'd accumulate forever under the old
heuristic.

services.last_heartbeat is the same primitive the rendezvous router
uses for routing.  Reusing it here keeps reap scoping aligned with
routing: dead pods' rows fall out of the live set after the heartbeat
window and become reapable; alive pods' rows stay protected as long as
they heartbeat.

Mechanics:

- ``bulk_close_stale_orphans`` parameter renamed
  ``node_id: str | None`` → ``live_node_ids: list[str] | None``.  The
  WHERE clause becomes ``(node_id IS NULL OR node_id NOT IN
  live_node_ids)``.  ``None`` skips the filter entirely (single-process
  / tests / operator backfill).  ``[]`` treats every row as
  unprotected.
- ``SessionManager.close_idle`` pass 2 calls
  ``storage.list_services(self._service_type)`` to enumerate live
  peers, passes their service_ids as ``live_node_ids``.  ``_service_type``
  is derived from ``self.kind`` (INTERACTIVE→"server",
  COORDINATOR→"console") via a module-level mapping — no constructor
  param, so production wiring can't miswire the kind/service_type
  pairing.
- list_services failure → pass 2 is skipped this tick (conservative;
  never reap when liveness state is unknown).  Pass 1 still runs.
- ``workstreams.node_id`` with NULL value is always eligible — defends
  against ANSI ``NULL NOT IN (...)`` evaluating to NULL (not TRUE) and
  silently protecting orphans forever.
- Migration 048 simplified to ``(kind, updated)``; the new query's
  ``NOT IN (small list)`` predicate against an unbounded-cardinality
  column doesn't index well, so leading ``node_id`` would just add
  write cost.

Tests cover the live-services protection (own/dead/null cases), the
empty-peers reap-all case, the list_services-failure conservative
fallback, both kind/service_type pairings (interactive→"server",
coordinator→"console"), and the combined live_node_ids +
exclude_ws_ids filter matrix.
2026-04-30 13:34:15 -07:00
Patrick Buckley 4b6c93a0e9 perf(storage): partial composite index for the orphan reaper query
bulk_close_stale_orphans runs every min(300s, idle_timeout/4) on
every server and console process.  Its WHERE shape is:

    WHERE kind = ?
      AND state IN ('idle','thinking','attention','running')
      AND updated < ?
      AND node_id = ?           -- multi-node interactive only

At current scale the existing single-column indexes are sufficient —
idx_workstreams_state prunes to non-closed and the planner filters the
rest sequentially.  At 100k+ rows that filter becomes a tablescan-
shaped cost.

A partial index covering only BULK_CLOSE_STATE_VALUES rows matches the
reaper's query exactly while staying tiny — closed rows (typically
95%+ of the table) and error rows are excluded, so the index is
roughly 5% the size a full multi-column index would be.  Write
amplification only kicks in for transitions touching one of the four
covered states.

Column order (node_id, kind, updated): node_id is the most selective
filter for multi-node interactive (each server prunes to its own
node's rows), kind second so coord-only and interactive-only queries
within a node still get index-only scans, updated last so the range
comparison rides the trailing column.

Postgres uses CREATE INDEX CONCURRENTLY so the build is non-blocking
on a live system; SQLite has no concurrent concept and the table-
level write lock already serializes, so a plain CREATE INDEX is fine.
2026-04-30 13:34:15 -07:00
Patrick Buckley 9d283e951f fix(console): periodic idle cleanup for the coordinator pool
The console's coord SessionManager had no idle thread — close_idle was
never called for coordinator workstreams.  This is the worse half of
the lifecycle leak: the dashboard filters via the in-memory pool, so
DB-only orphan coords were invisible.  At empirical diagnosis,
coord closure was 16% (10 closed / 64 total) vs interactive 63%.

Adds _coord_idle_cleanup_thread mirroring turnstone/server.py's
_idle_cleanup_thread but skipping the rate-limiter / global-queue arms
the console doesn't have.  Started from the lifespan when coord_mgr is
constructed and server.workstream_idle_timeout > 0 (reuses the
existing setting — same cadence works for both kinds).

Initial sweep runs INSIDE the thread before the first sleep, not
synchronously in the lifespan: cold-start orphans are reaped without
blocking Starlette boot.  Important because cold start with many DB
orphans (the precise condition this code targets) is exactly when the
UPDATE is most likely to be slow.

Helper takes an optional stop_event parameter purely for tests —
production callers pass None and the daemon runs for process lifetime.
This avoids the SystemExit-from-stub + module-wide filterwarnings
fragility a previous iteration relied on.

Four tests: initial sweep runs before first sleep, ticks fire each
loop, exceptions don't kill the thread, stop_event exits cleanly.
2026-04-30 13:34:15 -07:00
Patrick Buckley 4e407e7d4f fix(session_manager): close DB-orphan workstreams in close_idle
Real bug: workstream rows accumulate in non-closed states (idle,
thinking, attention, running) when their owning process restarts or
crashes.  Empirical diagnosis on a live deployment found ~60 stuck
coord rows in DB invisible to the in-memory-keyed dashboard, plus
100+ interactive rows older than the 2h timeout (one stuck "thinking"
for 2 weeks — impossible across a process restart).

Root cause: close_idle iterates self._workstreams.values() — only the
loaded subset.  Anything left behind by a prior process incarnation
sits in DB forever because nothing ever re-loads it.

This commit gives close_idle a second pass.

Pass 1 (existing, unchanged): close loaded IDLE rows whose
ws.last_active (monotonic) is past timeout.  IDLE-only so legitimately-
attentive rows (waiting for user response) stay live.

Pass 2 (new): bulk-close DB rows of this manager's kind whose updated
is past the wall-clock cutoff and which aren't currently loaded.
Closes the broader BULK_CLOSE_STATE_VALUES set — any matching row is
by definition not loaded by any process and cannot be in a live
interaction.  Scoped by self._node_id so a sibling node can't reap
rows we own (multi-node interactive correctness).  No emit_closed —
never-loaded rows have no SSE listeners expecting them.

Lock invariant: pass 1 holds self._lock briefly to snapshot victims
and pop them (existing behavior).  Pass 2 holds self._lock briefly to
snapshot the loaded keys, then releases before the DB UPDATE so a slow
reaper query can't block create/get/set_state.

Also fixes a same-process race in open(): the rehydrate path read DB,
released the manager lock, then re-acquired to install — a concurrent
pass 2 between the two acquisitions snapshots loaded keys without the
in-flight ws_id, and could clobber its DB row to closed.  open() now
calls touch_workstream(ws_id) on rehydrate so the row's updated is
fresh against any pass-2 cutoff.  Pure timestamp write is safe against
concurrent close() (close still wins on the state column).

Three new tests cover the DB orphan pass (basic, exclude-loaded, kind
filter) plus node_id scoping (own/foreign rows, None-skips-filter) and
the open() rehydrate touch.
2026-04-30 13:34:15 -07:00
Patrick Buckley 7ab24e500b fix(storage): add bulk_close_stale_orphans + touch_workstream primitives
Two new methods on the StorageBackend Protocol, with implementations on
both Postgres (UPDATE ... RETURNING) and SQLite (SELECT-then-UPDATE in
one transaction).  No callers yet — wiring lands in subsequent commits.

bulk_close_stale_orphans(kind, cutoff, exclude_ws_ids, node_id=None)
flips rows in BULK_CLOSE_STATE_VALUES (idle/thinking/attention/running)
to closed when their updated timestamp is lex-older than cutoff.  The
node_id filter scopes the reap to a single node's partition — required
for multi-node interactive deployments where each node only has
authority over its own workstreams.node_id rows.  Excludes loaded ids
so the in-memory pass owns those.

touch_workstream(ws_id) bumps updated without changing state.  Used by
the open() rehydrate path to defend against the orphan reaper clobbering
a freshly-loaded row whose DB updated is older than the cutoff.  Pure
timestamp write is safe against concurrent close() because close still
wins on the state column.

BULK_CLOSE_STATE_VALUES is centralized in workstream.py so the two
backend implementations and FakeStorage all agree; if a new transient
state is added to WorkstreamState, deciding whether it joins this set
is part of the change rather than an after-the-fact audit across three
files.

Storage tests (run against both backends via the conftest fixture) cover
the kind/state/cutoff/exclude/node_id matrix plus touch_workstream.
2026-04-30 13:34:15 -07:00
Patrick Buckley 5bcbcb73b9 chore: bump version to 1.5.2 2026-04-30 03:15:59 -07:00
Patrick Buckley af6749421a fix(metacog): drop duplicate [repeat: tool()] info line
The themed ``tool_reminder`` bubble below the tool block already
shows the metacog text, and the tool block immediately above it
carries the tool name — so a separate gray ``[repeat: list_workstreams()
called with same arguments]`` info line was just duplicate visual
noise (operator-visible in the screenshot below the bubble).

Drop the ``ui.on_info`` call inside ``_apply_post_execute_advisories``
that emitted the diagnostic line.  Update the docstring to reflect
that the bubble is the canonical signal.  Rename
``test_emit_repeat_ui_line_on_streak_fire`` →
``test_no_legacy_repeat_info_line_on_streak_fire`` and invert the
assertion.
2026-04-30 03:15:22 -07:00
Patrick Buckley 4d6cb77075 fix(cli): add on_user_reminder + on_tool_reminder to TerminalUI
CI typecheck failed because ``WorkstreamTerminalUI(TerminalUI)``
inherits from ``SessionUI`` (the Protocol), and the Protocol's
``on_user_reminder`` / ``on_tool_reminder`` declarations have empty
bodies — mypy treats those as implicitly abstract, so the subclass
became un-instantiable.

Add real implementations on ``TerminalUI`` that render reminders as
``[metacognition · type] text`` lines in yellow.  This also restores
the metacog signal on the CLI surface (the legacy
``[metacognition: nudge injected — …]`` info-line went away with
``_emit_nudge_ping``; without this commit the CLI showed no signal
at all for metacog nudges).  Tool-channel and user-channel render
identically because terminal output is anchored by stdout flow
rather than by DOM anchor — the line lands directly after the
message it advises.
2026-04-30 03:15:22 -07:00
Patrick Buckley 5c225ef39b docs(metacog): align comments with side-channel + tool-channel scope
Address Copilot's review feedback on PR #456 — the docstrings and
inline comments hadn't all caught up with the architectural shift
across the branch:

  - ``_apply_reminders_for_provider`` docstring: "every user message"
    → role-agnostic, since tool messages also carry ``_reminders``
    (tool_error / repeat).
  - ``_mark_reminders_delivered`` docstring: same role-agnostic
    update; explicitly note both channels.
  - ``_append_user_turn`` callsite comment near
    ``_attach_pending_user_reminders``: still described splicing
    ``<system-reminder>`` blocks into user content; updated to
    reflect the side-channel attach + transient-copy splice at the
    provider boundary.
  - ``_build_history`` block comment: was user-message-only; now
    mentions tool messages and both ``user_reminder`` /
    ``tool_reminder`` SSE events.
  - ``_build_history`` propagation comment: same role-agnostic note
    on the per-entry surface.
  - ``app.js`` ``user_reminder`` SSE handler comment: said the
    bubble renders "above" the user message, but
    ``insertAdjacentElement('afterend', el)`` drops it BELOW.
  - ``app.js`` ``replayHistory`` comment: said "insertBefore drops
    the reminder directly above the just-rendered user bubble";
    same fix — bubble lands BELOW.

No behaviour change.
2026-04-30 03:15:22 -07:00
Patrick Buckley f5a843f44a fix(metacog): drop write-success-clear so sequential same-call streaks fire
The repeat-detection block in ``_apply_post_execute_advisories`` had
a leftover "clear streak when a write tool succeeded" branch from
when ``RepeatDetector`` tracked cumulative counts.  With the
consecutive-streak semantics introduced earlier in the branch the
branch became:

  1. Redundant — any different (name, args) signature already resets
     the streak via ``RepeatDetector.record``, so an intervening
     read/write naturally breaks the streak.
  2. Actively wrong — the clear runs ONCE at the top of each
     ``_apply_post_execute_advisories`` call, before the per-result
     loop records sigs.  In a single parallel batch
     ``[bash, bash, bash]`` the clear runs once and then three
     ``record`` calls accumulate to count=3 in the same call → fires.
     But across three sequential turns, each turn calls
     ``_apply_post_execute_advisories`` fresh, the clear runs at the
     top of each call, and only one ``record`` per call follows — so
     the count never gets above 1 and the canonical
     "small local model stuck on ``bash('echo test')``" pattern
     never triggered the nudge.

The asymmetry only existed for successful calls — failures don't
satisfy the ``not _tool_error_flags.get(tc["id"])`` predicate, so
the clear didn't fire and sequential failures already worked.  The
fix is to drop the clear entirely; ``RepeatDetector``'s
consecutive-streak semantics handle every case uniformly.

Tests:

  - ``test_successful_write_clears_streak`` →
    ``test_intervening_different_call_resets_streak`` —
    rewords the assertion to reflect the actual mechanism (any
    different sig resets, write-or-otherwise) since "writes clear"
    was the bug, not the contract.
  - ``test_failed_write_does_not_clear_streak`` →
    ``test_sequential_bash_failures_fire_repeat`` — same shape, just
    framing fixed.
  - New ``test_sequential_bash_same_command_fires_repeat`` —
    regression for the bug user hit (three sequential successful
    ``bash('echo test')`` calls now correctly fire the nudge).
2026-04-30 03:15:22 -07:00
Patrick Buckley cf44841624 feat(metacog): themed reminder bubble unifies user + tool channels
The yellow themed reminder card introduced for user-channel nudges
(correction / denial / resume / start / completion) now also fronts
tool-channel nudges (tool_error / repeat).  Pre-fix the tool channel
shipped its reminders inside the tool-result envelope via
``wrap_tool_result``, leaking the ``<system-reminder>`` block into
``self.messages`` content (same problem the user channel had before
the side-channel refactor) and surfacing the legacy gray
``[metacognition: nudge injected — …]`` info line as the only
operator-visible signal — duplicated alongside the new themed bubble
for user-channel nudges.

Tool-channel parity:

  - ``_collect_advisories`` now returns
    ``(persistent_advisories, metacog_reminders)``.  Persistent
    advisories (``GuardAdvisory`` / ``UserInterjection``) keep
    riding ``wrap_tool_result`` because they ARE conversation
    history.  Metacognitive reminders extract to the second tuple
    element; the caller attaches them to the tool message dict's
    ``_reminders`` side-channel and emits ``on_tool_reminder``.
  - ``_apply_reminders_for_provider`` already handles ``_reminders``
    on any role, so the tool-channel splice into wire content is
    free.  ``_build_history`` also already propagates
    ``entry["reminders"]`` regardless of role, so reload renders the
    bubble too.
  - ``SessionUI`` Protocol gains ``on_tool_reminder(reminders,
    tool_call_id)``; ``SessionUIBase`` enqueues a ``tool_reminder``
    SSE event with the ``tool_call_id`` anchor.
  - ``_emit_nudge_ping`` had no remaining callers and was removed —
    the themed bubble (live SSE + ``/history`` reload) is the
    canonical operator signal for both channels now.

UI polish (the four fixes the screenshot caught for the user
channel + their tool-channel mirror):

  - Bubble renders BELOW the message it advises (semantically: a
    hint to the model right before its turn).  ``addUserReminder``
    swaps ``insertBefore`` for ``insertAdjacentElement('afterend',
    el)``; ``addToolReminder`` anchors below the ``.ts-approval``
    block whose tool result triggered the batch's reminder.
  - Label uses the full feature name ``metacognition`` (was the
    ``metacog`` shorthand).
  - Card width / alignment inherits from the base ``.msg`` rule —
    ``align-self: flex-end`` and the explicit ``max-width`` are
    gone, so the card matches the user / assistant column instead
    of pinning right-aligned narrow.
  - The legacy ``[metacognition: nudge injected — …]`` gray info
    line is gone for both channels.

Frontend additions:

  - ``Pane.prototype.addToolReminder(reminders, toolCallId)``
    anchors below the ``.ts-approval`` block (live: by
    ``data-call-id``; replay: by "last block in messagesEl"
    fallback, which is correct because messages render in order).
  - SSE switch case ``"tool_reminder"`` calls ``addToolReminder``.
  - ``replayHistory``'s tool-message branch now calls
    ``addToolReminder`` when ``msg.reminders`` is present.
  - ``addUserReminder`` advances its anchor on each loop iteration
    so multiple reminders stack in queued order rather than
    reversed.

Coord console parity:

  - ``coordinator.js`` gains ``appendReminderBubble`` /
    ``appendUserReminderLive`` / ``appendToolReminderLive`` mirroring
    the interactive UI.  The tool-channel anchor walks
    ``toolRows[callId].batch`` to attach below the
    ``.coord-tool-batch`` construct (one bubble per dispatch turn,
    matching the "one nudge per batch even with many failing tools"
    drain).
  - SSE switch handles ``user_reminder`` and ``tool_reminder`` on
    the coord conversation surface.
  - ``/history`` replay propagates ``msg.reminders`` for user and
    tool messages — same wire shape as the interactive pane.
  - ``.msg.user-reminder`` styles moved to
    ``shared_static/chat.css`` so both surfaces inherit the same
    yellow themed bubble from the shared base.

Defensive read on ``_apply_reminders_for_provider`` (per Copilot
review on the closed PR): a malformed ``_reminders`` entry (string,
None, etc. — corruption / partial state) used to abort ``send`` via
AttributeError on the ``.get("text", "")`` call.  Filter to dicts
before building the block, mirroring the same filter
``_build_history`` already applies on the wire-out side; an
all-malformed list passes through as no-reminders.

Tests:

  - ``test_collect_advisories_drains_tool_buffer_on_last_result``
    rewritten to assert the ``(persistent, metacog)`` tuple shape
    and that ``MetacognitiveAdvisory`` no longer appears in the
    persistent list.
  - ``test_collect_advisories_holds_*`` and ``_drops_*`` updated for
    tuple return.
  - ``test_attach_emits_visibility_ping`` /
    ``test_collect_advisories_emits_visibility_ping`` inverted to
    assert the legacy gray line is gone on both channels.
  - ``TestSessionUIBaseToolReminderHook`` covers the new SSE event
    shape with the ``tool_call_id`` anchor.
  - ``test_malformed_reminders_filtered_out`` and
    ``test_all_malformed_reminders_passes_through`` cover the
    Copilot-flagged defensive filter.
2026-04-30 03:15:22 -07:00
Patrick Buckley 7ffab6a272 fix(session): metacog reminders ride a side-channel, not user content
User-channel metacognitive nudges (correction, denial, resume, start,
completion) used to be spliced into ``user_msg["content"]`` permanently,
which leaked the ``<system-reminder>`` envelope into every consumer of
``self.messages`` — UI replay (mitigated by a regex strip in /history),
compaction, title generation, and any future channel adapter that
echoes conversation context.  The /history strip was a band-aid;
compaction and title-gen still saw the raw spliced text.

Switch to a side-channel: ``_attach_pending_user_reminders`` writes the
rendered reminder list to ``user_msg["_reminders"]`` (sibling key,
leading-underscore convention shared with ``_attachments_meta`` /
``_provider_content``).  At the provider boundary, a new
``_apply_reminders_for_provider`` builds a transient shallow-copy with
the reminder spliced into ``content``; the original message dict
stays clean.  ``sanitize_messages`` drops the sibling key on the wire.

Once-per-session-not-per-turn semantics for the wire: after stream
success the loop calls ``_mark_reminders_delivered``, which flips a
``_reminders_delivered`` flag on every user message that carried
reminders into that call.  ``_apply_reminders_for_provider`` skips
already-delivered messages so the model sees each reminder exactly
once (the turn it advised).  ``_build_history`` ignores the delivered
flag entirely, so reconnecting tabs render the same nudge bubble the
originating tab saw via the live ``user_reminder`` SSE event.

UI surface:

  - ``SessionUIBase.on_user_reminder`` enqueues a
    ``{type: "user_reminder", reminders: [...]}`` SSE event with the
    same shape ``_build_history`` surfaces.
  - ``app.js`` renders a ``.msg.user-reminder`` bubble (yellow accent,
    pill-styled) anchored above the user message it advises, both
    live and on history replay.
  - ``replayHistory`` renders ``addUserMessage`` before
    ``addUserReminder`` so the anchor lookup finds the just-rendered
    turn (not a prior one).
  - Multi-tab caveat documented inline: non-originating tabs receive
    no ``user_message`` SSE event today, so a reminder may anchor to
    a stale prior bubble until ``/history`` reload corrects it.

Pre-existing bug surfaced by the audit: cancel handlers
(``GenerationCancelled`` / ``KeyboardInterrupt`` / generic
``Exception``) in ``ChatSession.send`` cleared
``_pending_tool_advisories`` but not the user-channel buffer.  Both
now drain through a shared ``_drain_pending_advisories`` helper.

Removed the ``/history`` regex strip — the side-channel approach
makes it redundant.  Hoisted ``escape_wrapper_tags`` +
``render_system_reminder`` imports to module top (called 2-3× per
turn).

Tests:

  - ``TestApplyRemindersForProvider`` — pass-through-by-reference,
    string + list content splice, escape on user-typed wrapper tags,
    multi-reminder ordering, source-untouched invariant, delivered
    flag skip path, fallback for unexpected content shape.
  - ``TestMarkRemindersDelivered`` — flag idempotency, no-reminders
    no-flag, only marks user messages with reminders.
  - ``TestUpdateTokenTableMsgsParam`` — calibration uses pre-built
    msgs when provided, falls back when not.
  - ``TestUserAdvisoryCancelClear`` — all three cancel branches drain
    the user buffer.
  - ``TestReminderSidechannelIsolation`` — compaction's
    ``_format_messages_for_summary`` and the title-gen extraction
    loop cannot see reminders by construction.
  - ``TestSessionUIBaseUserReminderHook`` — ``on_user_reminder``
    enqueues the right SSE shape.
  - ``TestBuildHistoryReminderPropagation`` — ``entry["reminders"]``
    propagation, absent / empty / multi / coexist-with-attachments
    cases, malformed input filtering, all-malformed elision.
  - ``test_sanitize_messages_strips_underscore_sibling_keys`` covers
    ``_reminders`` and ``_reminders_delivered``.
2026-04-30 03:15:22 -07:00
Patrick Buckley ba3bc9d989 fix(metacog): N>=3 streak detector + drop redundant error-prefix list
Cleanup pass on the metacognitive nudge stack — restores pre-split
errored-counts-toward-repeat behaviour and tightens the is_error
plumbing through the per-batch advisory hook.

The per-batch hook in ``_run_loop`` was duplicating the is_error
signal: ``self._tool_error_flags`` (set by ``_report_tool_result``)
and a string-prefix tuple (``Error`` / ``JSON parse error`` / …).
Two truth sources is what got us here — bash commands that exit
non-zero with normal stdout matched the flag but not the prefix,
the deny path matched the prefix but not the flag, and the result
was that stuck-loop detection silently broke for the most common
failure mode (the model bashing the same broken command).

Single source of truth now:

- ``_execute_tools.run_one`` deny branch routes through
  ``_report_tool_result(is_error=True)`` so denied calls populate
  ``_tool_error_flags`` like every other error path.
- The error-prefix tuple is gone; the write-success-clear gate and
  the tool-error-nudge gate both read ``_tool_error_flags`` only.

Repeat-detection state moves from a ``set[str]`` (fired on the second
identical call, ignored errors entirely) to a ``RepeatDetector``
helper in ``metacognition.py`` with consecutive-streak semantics:

- Threshold raised from 2 to 3 — two-in-a-row was noisy on
  legitimate transient retries; three is the cheapest stuck-loop
  signal.
- Recording a different signature resets the count, so [A, A, B, A]
  is two short streaks of 2 and not a streak of 4. Bounded by O(1)
  state regardless of session length.
- Errored calls now count toward the streak (the split into a
  separate metacog module unintentionally introduced a "skip errors"
  branch — restored).

While there:

- ``metacognition._COOLDOWN_SECS`` default aligned to 300s (matches
  ``MemoryConfig.nudge_cooldown`` and the ``memory.nudge_cooldown``
  config-store default; was set to 30 by an earlier investigation).
- The per-batch advisory block (~80 lines of mixed orchestration
  inside ``_run_loop``) is extracted to
  ``ChatSession._apply_post_execute_advisories`` so the wired
  behaviour is testable without driving ``_run_loop`` end-to-end.
  Producer extraction to a dedicated module is deferred to a
  follow-up; advisory producers all live on ``ChatSession`` for
  now per existing convention.
- Frontend ``appendToolOutput`` (turnstone/ui/static/app.js) now
  skips rendering when the parent approval block is denied or
  the output starts with ``Denied by user`` / ``Blocked``,
  mirroring the history-replay guard at ``_build_history``.
  Previously the live SSE path didn't need this guard because
  the deny path never emitted a ``tool_result`` event; the
  is_error routing change above means it does now, so without
  this guard the badge from ``resolveApproval`` and the SSE
  output would both render.

Tests: 8 unit tests for ``RepeatDetector`` covering streak,
threshold, clear, and intervening-sig reset; 9 integration tests
for ``_apply_post_execute_advisories`` covering the wired
behaviour (3-identical fires warning + advisory + UI line, errored
calls count toward streak as a regression guard, intervening sig
resets streak, successful write clears, failed write does not,
JSON outputs tracked but not inline-warned, tool_error nudge gates
on memory_count, repeat UI line emitted on streak fire).
2026-04-30 03:15:22 -07:00
Patrick Buckley dbe023b4dd chore: bump version to 1.5.1 2026-04-29 20:21:12 -07:00
Patrick Buckley 3d3a8b7367 docs(coord): tighten handleChildState comment per Copilot review
The pre-existing comment said pending_approval_detail "rides on
every ws_state event" — that overstated the case.  The node-side
emit is gated on ``_pending_approval is not None`` so the field is
absent on the steady-state broadcast and possibly null on a node
mid-rolling-upgrade.  The handleChildState fallback already
handles both cases; only the comment was wrong.
2026-04-29 20:20:38 -07:00
Patrick Buckley 99eff73a97 feat(coord): pass pending_approval_detail on child_ws_state SSE events
Inline child approve/deny in the coord tree UI was rendering downstream
of the bulk-live cache (``GET /v1/api/cluster/ws/live``), not the SSE
stream. ``child_ws_state`` events were tiny notifications that fired
an urgent live-bulk fetch on every activity_state transition into/out
of "approval", just to pick up the rich ``pending_approval_detail``
payload. With multiple coord tabs and multi-child workstreams, that
urgent-fetch pattern compounded the SSE-executor pressure Shape A
is unwinding.

Thread the field through every layer so the SSE event itself carries
the rich payload — browser mutates ``liveBadgeCache`` directly,
no urgent fetch:

  1. Node ``WebUI._broadcast_state`` emits ``pending_approval_detail``
     on ``ws_state`` events. Gated on ``_pending_approval is not None``
     so the per-broadcast verdict-cache deepcopy only runs when there
     is actually an approval pending. ``_build_node_snapshot`` also
     projects the field so the console's reconnect-via-snapshot
     resync path delivers it (without this the new collector
     forwarding would never see the field on a snapshot row).

  2. Console ``ClusterCollector._apply_delta`` (live ``ws_state``
     forwarding) and ``_reconcile_node`` (snapshot resync diff) both
     forward the field on the emitted ``cluster_state`` event, AND
     ``_apply_delta`` persists it on the cached ``ws`` dict so the
     ``get_node_detail`` / ``get_snapshot`` endpoints between
     reconciliations don't render stale approve/deny buttons.

  3. ``CoordinatorAdapter._dispatch_child_event`` re-emits the field
     on the ``child_ws_state`` event sent to coord listener queues.

  4. Frontend ``handleChildState`` reads ``ev.pending_approval_detail``
     and writes it directly into ``liveBadgeCache``, tagging the
     entry with ``sseUpdatedAt``. ``flushLiveFetches`` honors that
     tag for ``SSE_AUTHORITATIVE_MS`` (3s) — the upstream
     ``/dashboard`` cache has its own ~2s TTL, so a bulk-poll
     landing right after a transition can otherwise clobber the
     fresh SSE-set state with pre-transition data.

The pre-fix ``enteredApproval`` / ``leftApproval`` urgent-fetch
branch is removed. The 409 stale-call_id retry path keeps its own
urgent fetch — that's a different scenario.

Tests cover the forwarding contract at every layer, the broadcast
gate (event includes the field when an approval is pending,
omits it otherwise, and clears after resolution), and the
``flushLiveFetches`` merge-guard structural shape so a refactor
that keeps the symbols but inverts the comparison or drops the
``prev.live`` check can't pass silently.
2026-04-29 20:20:38 -07:00
Patrick Buckley 99fcd30299 fix(console): offload sync DB calls in coord children/tasks handlers
``coordinator_children`` was calling ``storage.list_workstreams``
directly on the event loop, ``coordinator_tasks`` did the same with
``load_task_envelope``, and ``_resolve_coordinator_or_404`` (called
from both handlers, plus ``coordinator_history`` and
``_resolve_coord_session``) did the same with
``storage.get_workstream`` on its cold-cache path.

The cold-cache resolver path is hit on every console restart,
coordinator eviction, and console proxy hop — exactly when the
event loop is most contended. Three coord tabs reconnecting after a
brief network blip = three serial event-loop blocks per call site.
Other lifted handlers in this file already use
``asyncio.to_thread``; bring all four call sites onto the same
pattern.

Convert ``_resolve_coordinator_or_404`` to ``async def`` and update
its four call sites to ``await``. Exception flow is unchanged.
2026-04-29 20:20:38 -07:00
Patrick Buckley 423c2e80b7 fix(console): isolate coord SSE polling on a dedicated 200-thread pool
Each coord ``events`` SSE listener parks a thread on
``client_queue.get(timeout=5)`` for the connection lifetime. The
console's coord endpoint was wiring no ``sse_executor_lookup`` on
``coord_endpoint_config``, so those parks landed on Python's default
ThreadPoolExecutor (~min(32, cpu_count+4)) and competed with every
other ``asyncio.to_thread`` caller (storage, router, audit). A few
coord tabs against a multi-child workstream would stall new request
handlers waiting for a worker thread.

Mirror the interactive-side precedent (the ``sse_executor`` /
``sse_executor_lookup`` pattern in ``turnstone/server.py``) — build a
dedicated 200-thread ``coord_sse_executor`` in the console lifespan
and wire ``sse_executor_lookup`` onto ``coord_endpoint_config``.
Drain order matters: shut the pool down AFTER ``coord_adapter.shutdown()``
so no new listeners arrive at a dying pool. ``cancel_futures=True``
discards queued-but-not-started futures during teardown.

Update the stale comment on the interactive-side wiring that claimed
"coord wires None and falls back to the default executor" — it now
points at the console's matching wire.
2026-04-29 20:20:38 -07:00
Patrick Buckley a0eb77360d fix(coord): tighten coord_registry refresh logging + comments per round-2 review
Three follow-ups from Copilot's round-2 review on #453.

ValueError logging surfaced the wrong reason
The catch-all ``except ValueError:`` logged ``reason=no_enabled_rows``
unconditionally, but ``ModelRegistry.__init__`` raises ValueError for
five distinct config issues (empty models, default / fallback / agent /
plan / task alias not present).  Operator looking at logs for a
config.toml typo would see the wrong cause.  Switch to
``log.warning("...reason=%s", exc)`` so the actual error message
threads through.  Behavior unchanged — existing registry still
preserved on every ValueError path.

Misleading shutdown() comment
The ``finally`` comment claimed shutdown() was closing clients the
throwaway registry created during DB load.  ``load_model_registry`` only
constructs ModelConfigs and the bare ``ModelRegistry(...)``;
``ModelRegistry.__init__`` leaves ``_clients`` / ``_providers`` empty
and they populate lazily on first resolve.  Today shutdown() iterates
empty dicts.  Comment now says so explicitly while keeping the call
(and its try/except) for forward-compat against an eager-init future.

Stale "probe" wording in test docstring
``test_helper_preserves_registry_when_db_probe_fails`` →
``test_helper_preserves_registry_when_strict_load_fails``.  The
explicit probe was removed in commit 1ba17ed when the helper switched
to ``load_model_registry(..., strict=True)``; the test name and
docstring still talked about a probe.  Updated wording reflects that
the loader's strict-mode re-raise is what the helper catches now.

132 tests pass.
2026-04-29 20:20:38 -07:00
Patrick Buckley 3dd0e196fe refactor(coord): hygiene pass on coord_registry refresh — async + selective teardown + test cleanup
Hygiene follow-ups from the multi-stage code review on #453.

perf-1 — sync helper called from async route handlers
``_refresh_coord_registry`` runs two sync DB reads and a registry reload
that takes ``_client_lock``; calling it directly from an async handler
held the event loop for the duration.  All four call sites now
``await asyncio.to_thread(_refresh_coord_registry, ...)``, matching the
pattern from commit ``1f7d6ad`` (offloaded ``tenant_check``).

perf-3 — ModelRegistry.reload() tore down all clients unconditionally
The reload always closed every cached client and provider, even when
the changed fields (``model``, ``temperature``, ``context_window``)
didn't touch the connection target.  Now selective: clients drop only
when alias removed or ``(base_url, api_key, provider)`` differs;
providers drop only when alias removed or ``provider`` string differs.
Keeps connection pools warm across the common admin-edit case where
only metadata changed.  Two new ``test_model_registry`` cases lock the
keep-warm vs drop-on-change behaviour, and the existing
``test_reload_clears_clients`` was updated (it asserted the old
overly-aggressive contract) into
``test_reload_keeps_clients_when_connection_target_unchanged``.

q-5 — helper rename
``_refresh_console_coord_registry`` → ``_refresh_coord_registry``.  The
``console_`` prefix was redundant given the function lives in
``turnstone/console/server.py`` and sibling helpers there
(``_notify_nodes_model_reload``, ``_publish_config_change``,
``_collect_model_status``) all omit it.

q-1 — shared test middleware
``tests/test_admin_model_registry_refresh`` now imports the
header-driven ``_AuthMiddleware`` from ``tests/_coord_test_helpers``
and sets default ``X-Test-User`` / ``X-Test-Perms`` headers on the
``TestClient``.  The local hardcoded variant duplicated infrastructure
the helper module exists to centralise.

q-3 — multi-alias test registry
``_make_registry`` extracted a ``_make_config`` helper and gained an
``extras={alias: model}`` param so multi-alias scenarios stop
hand-building ``ModelConfig`` literals.
``test_delete_endpoint_refreshes_registry`` now uses the helper.

310 tests pass across the related coordinator + model surfaces.
2026-04-29 20:20:38 -07:00
Patrick Buckley 19c3db5329 test(coord): lock the empty-body gate with a refresh-call spy
bug-3 / q-2 from the multi-stage review on #453: the previous test
``test_update_endpoint_with_empty_body_does_not_blow_up`` asserted only
that the registry's model name was unchanged after an empty PUT, which
holds whether or not the refresh ran (DB row matches registry → refresh
is idempotent).  A regression that always called
``_refresh_console_coord_registry`` — exactly the gate this test was
meant to lock — would have left the assertion green.

Rename to ``test_update_endpoint_skips_refresh_on_empty_body`` and spy
on the helper via ``monkeypatch.setattr``.  Empty-body PUT must register
zero calls; any future change that drops the ``if updates:`` gate now
fails loudly.
2026-04-29 20:20:38 -07:00
Patrick Buckley 0bea72019e fix(coord): strict-mode loader + guarded shutdown for coord_registry refresh
Two correctness follow-ups from the multi-stage code review on #453.

bug-2 / perf-2 (DB probe was theatre + double scan)
The previous probe defended nothing the loader didn't already swallow
on the next line: ``load_model_registry``'s row-loop catches Exception
internally, so a transient DB error after the probe still degrades to
a config.toml-only registry that ``existing.reload()`` would apply,
silently dropping every DB-sourced alias.  And on the happy path each
CRUD paid for two scans of ``model_definitions``.

Add a ``strict: bool = False`` flag to ``load_model_registry``.  When
strict, the row-loop's except re-raises instead of swallowing.  The
helper passes ``strict=True`` and drops the probe — single DB scan,
real failure isolation, the loader's silent fallback can no longer
mask a partial-result regression.  Default ``strict=False`` so CLI /
lifespan callers keep their boot-with-config-fallback behaviour.

bug-1 (shutdown could escape after a successful reload)
``ModelRegistry.shutdown()`` calls ``client.close()`` unguarded, and the
helper's ``finally`` block ran it outside the try/except.  A raising
close() after a successful ``existing.reload()`` would surface as 500
with the registry already mutated and the audit row already recording
success.  Wrap ``new_registry.shutdown()`` in its own try/except that
matches the helper's belt-and-suspenders error policy elsewhere.

The helper's docstring also drops the obsolete probe paragraph; the
``if existing is None: return`` branch gets a one-line inline comment
about the boot-from-empty case (the multi-paragraph version restated
behaviour the line itself documents).

129 tests pass (test_admin_model_registry_refresh + test_model_registry).
2026-04-29 20:20:38 -07:00
Patrick Buckley b9ff52d582 fix(coord): tighten coord_registry refresh — DB probe + accurate boot-from-empty docstring
Two follow-ups from Copilot review of #453:

1. ``load_model_registry`` swallows storage read errors internally
   (logs + continues with config.toml-only models).  Without a strict
   probe in the helper, a transient DB outage on an admin CRUD would
   apply a truncated registry that drops every DB-sourced alias —
   silently, since the loader returns a non-empty registry built from
   ``[models.*]`` config.toml entries.  Add an explicit
   ``storage.list_model_definitions(enabled_only=True)`` probe before
   the loader call so the failure is visible here and the existing
   registry is preserved on outage.

2. The previous docstring claimed ``admin_model_reload`` "has its own
   boot-from-empty story."  It doesn't — it just calls this helper,
   which no-ops when ``coord_registry`` is None.  When no model rows
   existed at boot, lifespan leaves the entire coord subsystem
   uninitialized (no ``coord_mgr``, no ``coord_adapter``, no
   ``session_factory``), and a console restart remains required after
   the operator adds the first row.  Tighten the docstring to admit
   that limitation rather than overstating the helper's reach.

New test ``test_helper_preserves_registry_when_db_probe_fails``
monkeypatches ``list_model_definitions`` to raise and asserts the
existing registry stays intact.
2026-04-29 20:20:38 -07:00
Patrick Buckley 961f999c93 fix(coord): auto-refresh console coord_registry on model-definition changes
The console builds ``app.state.coord_registry`` once at lifespan startup
and the coordinator session factory closes over that exact instance.
Until now, the model-definition admin endpoints (create/update/delete)
wrote to the DB but never touched the in-process registry — and the
explicit reload button only fanned out to nodes via HTTP, also leaving
the console's own registry stale.

Symptom: an operator who changed the underlying model name behind a
local-LLM alias (same alias, same endpoint) saw the DB row update
immediately, but coordinator sessions kept calling the prior model
name until the console process was restarted.

Fix: a new helper ``_refresh_console_coord_registry`` rebuilds a fresh
ModelRegistry from DB and applies it to ``app.state.coord_registry``
via the existing thread-safe ``ModelRegistry.reload()`` — in-place
mutation preserves object identity so the factory closure keeps
working, and active coord sessions auto-pick up the swap on their
next ``send()`` via ``ChatSession._refresh_model_from_registry``.

Wired into four endpoints in ``console/server.py``:

- ``admin_create_model_definition`` — after the DB write
- ``admin_update_model_definition`` — after the DB write, gated on
  ``if updates:`` so a no-op PUT skips the rebuild
- ``admin_delete_model_definition`` — after the DB write
- ``admin_model_reload`` — between ``_publish_config_change`` and
  ``_notify_nodes_model_reload`` so the console mirrors what the
  reload broadcasts to nodes

Failure isolation: a load or reload error leaves the existing registry
intact (logged + swallowed). Coord stays usable while the operator
investigates; the explicit reload remains the user-facing recovery path.

No node fan-out on CRUD — the explicit reload button continues to gate
cluster-wide HTTP propagation, preserving today's UX semantics on shared
clusters.

Tests in ``tests/test_admin_model_registry_refresh.py`` cover:

- helper-level: rebuild from DB, identity preservation, no-op when
  registry is None, preservation on load failure / no-enabled-rows /
  reload validation error
- endpoint-level: create / update / delete / explicit-reload all
  refresh the registry; an empty PUT skips the rebuild
2026-04-29 20:20:38 -07:00
Patrick Buckley 8bdb916064 fix(coord): raise wait_for_workstream message cap to 10 KiB
Production fan-outs are frequently hitting the 6 KiB per-child cap by
just 1-2 KiB, forcing the coordinator into a follow-up inspect_workstream
round-trip per truncated child to recover the tail. Bumping the cap to
10 KiB absorbs the common overshoot without changing the truncation
semantics — truncated=True still fires for genuinely oversized messages,
and inspect_workstream remains the unbounded follow-up.

Worst-case context impact: a 32-child fan-out at the cap is now ~320 KiB
(was ~192 KiB), still well within commercial model context windows.
Typical fan-outs of 1-5 children land at 10-50 KiB.

LAST_ERROR_MAX_LEN (1 KiB) is unchanged — it's intentionally smaller
than the wait cap so error truncation happens at write time, and
1 KiB still sits well below 10 KiB.

WAIT_MESSAGE_MAX_BYTES is referenced by name (not literal 6144) in the
truncation test, so no test value needs updating.
2026-04-29 20:20:38 -07:00
Patrick Buckley 25fe4e728a fix(coord): make coordinator fan out independent work by default
The coordinator system message was descriptive about parallelism rather
than prescriptive — "while multiple children run in parallel" framed
fan-out as incidental, and "a tasks entry, a child to own it" primed
singular delegation. The spawn_batch example (benchmark A, benchmark B,
prototype the winner) showed dependent work under a fan-out framing,
teaching the wrong shape.

In practice the coordinator failed to decompose enumerable requests
("top stories on HN, Lobsters, /r/programming, …") without explicit
"please fan this out" instructions, on both GPT-5.5 and Claude Opus.

base_coordinator.md
- Replace singular "a tasks entry, a child to own it" with plural
  "enumerate the independent units of work, spawn one child per unit,
  run them in parallel by default. Sequential only when one child's
  output feeds the next."
- Tighten the delegation paragraph.

tools_coordinator.md
- Drop the persona repetition that duplicated base_coordinator.md.
- Drop the prescriptive "## Workflow shape" section (the cost note is
  already in wait_for_workstream's tool description; the edit-X
  redirect is already in the persona).
- Drop "in one approval" / "single approval" mentions to avoid
  surfacing approval mechanics to the model.
- Replace the misleading spawn_batch example with truly independent
  items; drop "(up to 10)" which overstated the cap (it's per-call,
  not global, and is documented in the tool schema).
- Add a course-correction example to send_to_workstream — the pattern
  coordinators most often replace with cancel-and-respawn.
- Drop the read action from the tasks examples to keep the lifecycle
  (add → update → remove) coherent.

Coord system message ~16% shorter (4440 → 3722 chars). Both GPT-5.5
and Claude Opus now naturally decompose the news-board prompt without
explicit fan-out instructions. 29 prompt-composition tests pass.
2026-04-29 20:20:38 -07:00
Robert DeAngelis 0d1a32ff65 fix(server): accept --skip-permissions CLI flag (#450)
The server's --help epilog and compose.yaml both reference
--skip-permissions, but the argparser never defined it, so any
container started with SKIP_PERMISSIONS=1 exited with
"unrecognized arguments: --skip-permissions".

Wire the flag through to app.state.skip_permissions, OR-ing it
with the existing tools.skip_permissions config-store setting so
the stored value still works on its own.
2026-04-29 20:20:38 -07:00
Patrick Buckley a8f6348f51 chore: bump version to 1.5.0 2026-04-29 00:16:57 -07:00
Patrick Buckley f24c6d6c73 docs(readme): refresh hero image to coordinator UX shot
Replaces the old mermaid-rendering shot with a coordinator session
mid-attention — parallel tool batches, judge-graded approval,
children + tasks side panels — which more accurately represents
what the platform does today.
2026-04-29 00:13:48 -07:00
Patrick Buckley 1f7d6ad23b perf(api): offload tenant_check to thread on lifted session handlers (#449)
* perf(api): offload tenant_check to thread on lifted session handlers

Every make_*_handler factory in turnstone/core/session_routes.py invoked
cfg.tenant_check(request, ws_id, mgr) synchronously inside its async
handler. For the interactive surface tenant_check chains through
_interactive_tenant_check → _require_ws_access → resolve_workstream_owner,
which short-circuits on mgr.get(ws_id) for warm cache but falls through
to a synchronous get_workstream_owner SQL call on a cold cache,
blocking the event loop for the duration of the storage round-trip.

Wrap each of the 8 call sites (approve, close, cancel, events, history,
detail, send, dequeue) in await asyncio.to_thread(...) — mirroring the
existing storage-offload pattern at make_history_handler's other call
sites. Coord wires tenant_check=None and is unaffected. Five handlers
gain a local import asyncio (matching the per-handler lazy-import
convention in this module). Centralizes the offload rationale on
SessionEndpointConfig.tenant_check's field docstring.

Adds two regression tests in TestTenantCheckOnReadEndpoints that wire
the real resolve_workstream_owner as tenant_check and force the
storage fall-through path the existing class only stubbed past with
fake allow/deny callables.

* test(api): spy asyncio.to_thread to pin tenant_check offload

Copilot flagged the cold-cache regression tests for asserting the
response shape but not the offload itself: reverting
await asyncio.to_thread(cfg.tenant_check, ...) to the sync call shape
would still leave the storage fall-through working and the tests
green. Patch asyncio.to_thread inside both tests with an async spy
that records every offloaded callable, then assert cold_check is in
the call list — sanity-checked by reverting the history wrap locally
and watching the assertion bite (offloaded only contained
storage.get_workstream + storage.load_messages, missing cold_check).
2026-04-28 23:56:43 -07:00
Patrick Buckley 353ff4d18b feat(coord): inline tool-batch construct replaces approval dock (#447)
* feat(coord): inline tool-batch construct replaces approval dock

The pinned bottom approval-dock didn't scale: a 10-call spawn_workstream
fan-out filled the whole pane with a wall of repeated verdict chips,
and the call → approval → result lifecycle was split across three
disconnected surfaces (.msg.tool bubble + dock + .msg.tool result).

Replaces it with one chat-stream construct per dispatch turn that
pairs each tool call with its result and embeds the approval gate:

  - .coord-tool-batch--solo      single-call serial turn
  - .coord-tool-batch--parallel  ≥2 calls; rows share a left rail
                                 + per-row tick so they read as
                                 siblings of one assistant decision

Lifecycle: rows render with optional "judge evaluating…" placeholder,
upgrade in place when intent_verdict arrives, and on tool_result the
output lands paired under the originating row.  When the batch needs
approval, one Approve/Deny/Always action row renders inside the
construct (envelope-level — server semantics resolve siblings
together).  After approval_resolved the action row morphs into a
✓ approved / ✗ denied status pill that stays as a receipt.

Critical bug closed: when a page reload races a pending approval,
pre-scan tool_call_ids in history; turns whose call_ids have no
matching tool result are rendered pending (not resolved-approved).
The SSE approve_request replay then upgrades the existing batch
in place — drops --approved/--denied, adds --pending, swaps the
status pill for actions, and assigns activeBatch.  Without this
the operator was locked out of any approval pending at reload.

Defence-in-depth follow-ups from the same review:

  - approval_resolved falls back to a DOM lookup if activeBatch
    is null (cross-tab resolution where this tab never set it).
  - _appendVerdictLineTo dedupes via a row.dataset.verdictSig so
    SSE reconnect storms + repeat intent_verdict events don't
    tear down + rebuild an unchanged verdict line.
  - judgeVerdicts Map soft-capped at 500 entries (FIFO eviction)
    via _cacheJudgeVerdict.
  - toolRows entries hold {batch, row} only — the originating
    item payload is no longer pinned for the page lifetime.
  - _scheduleScroll coalesces messagesEl.scrollTop writes through
    requestAnimationFrame so history replay doesn't reflow once
    per appended message.
  - Rationale <details> now inserts immediately after the verdict
    line (was tail-appending, breaking ordering once a result
    landed below).
  - .coord-tool-batch--error wired: _appendResultToRow lifts a
    row's error onto the enclosing batch; _renderBatchRow does
    the same for policy-blocked rows at construction.
  - _buildStatusPill extracted; both _morphBatchResolved and the
    appendToolBatch resolved-replay branch route through it.

Removed: ~248 lines of dead .approval-dock CSS, the dock <aside>
element from index.html, and the dead helpers showApproval's
prior body, hideApproval, claimApprovalFocus,
claimApprovalFocusForVerdict, applyJudgeVerdictToRow,
applyJudgePendingToRow, ensureDctxAfterRow, removeRationale,
setApprovalButtonsDisabled, the appendToolCall single-row wrapper,
and window.coordApprove.  Five stale comment blocks referencing
the dock as if live also swept.

Children-tree's renderApprovalBlock is independent and untouched
(different surface, different .approval-block / .approval-pill
vocabulary).

* fix(coord): close four Copilot review gaps on PR 447

Copilot review on caa07e6 flagged four follow-ups:

1. History replay was rendering EVERY orphan tool_calls turn (one
   that lacks a matching tool result message) as `pending: true,
   judgePending: true`.  That paints Approve/Deny on turns that
   could be just running — auto-approved-and-still-in-flight, or
   already-approved-and-still-in-flight — and clicking would 409
   because the call_id isn't in `pending_items`.  Add a new
   `--running` state for the orphan case (no actions, neutral
   accent stripe).  SSE then upgrades in place: `--running` →
   `--pending` when `approve_request` replays, or `--running` →
   `--auto` when `tool_info` replays.  Tool_result events still
   route into the rows for the third case (already-approved + in
   flight) since `toolRows` is populated.  Kicker text reads
   "Running · Parallel N" while ambiguous, so the operator can
   tell the in-flight-replay state apart from a fresh "Parallel ·
   N tools" auto-approved batch.

2. Removing the dock also removed its `aria-live="assertive"`
   region — pending tool-batches now append into the polite
   `#coord-messages` log (which gets flipped to `aria-live="off"`
   during streaming), so a screen reader could miss the
   action-required signal.  Add an off-screen
   `aria-live="assertive"` `#coord-sr-announcer` region and route
   "Approval required: <name> + N more" through it whenever a
   pending batch is created OR an upgrade-in-place promotes a
   running batch to pending.  Also mark pending batches with
   `role="region"` + a matching `aria-label` so SR landmark
   navigation surfaces them; both are dropped on resolve so the
   resolved batch stops claiming the landmark.

3. `_resolveBatchAction` was selecting the first row whose
   `data-call-id` was set and that wasn't `.error` — but
   `approve_request` envelopes carry the FULL items list,
   including auto-approved siblings whose `needs_approval=false`
   means the server's `pending_items` won't recognise their
   call_id (→ 409 on submit, or resolves the wrong gate).  Tag
   rows that are genuinely in `pending_items` with
   `data-needs-approval="1"` at construction (and during
   upgrade-in-place when SSE arrives), and select against that
   selector specifically.  Restores the legacy
   `pendingApprovalCallId` contract that filtered on
   `needs_approval` before the dock was retired.

4. The `.coord-tool-row-result` comment claimed the styles applied
   a click-to-expand "collapsed" affordance like the interactive
   UI's `.tool-output.collapsed`, but the implementation only set
   `max-height: 240px; overflow: auto` (a scroll pane, not a
   collapse with expand control).  Update the comment to describe
   what the rules actually do and explain the deliberate
   divergence from interactive (coord is a diagnostic-leaning
   read-once surface; an internal scroll pane reads with lower
   friction than a click-to-expand control on the operator's
   primary monitoring view).

No Python touched; node --check on coordinator.js clean.

* fix(coord): restore reload-time pending approval gate

Agent-Logs-Url: https://github.com/turnstonelabs/turnstone/sessions/30f630fe-3ded-4abe-991b-b5a95f699127

Co-authored-by: eous <13773563+eous@users.noreply.github.com>

* feat(api): expose pending_approval on workstream detail response

PR 447 / 93cb3d9 (Copilot autonomous follow-up) added a JS path that
reads ``wsSnapshot.pending_approval_detail`` off the
``GET /v1/api/workstreams/{ws_id}`` snapshot in coordinator.js
init() so a freshly-loaded chat tab can paint the inline approval
gate immediately at reload, without waiting for the SSE
approve_request replay (which leaves a brief --running flash on the
inflight orphan placeholder).

But the server's ``WorkstreamDetailResponse`` schema only declared
``{ws_id, name, state, user_id, kind}`` and the lifted
``make_detail_handler`` matched: nothing was populating
``pending_approval`` or ``pending_approval_detail`` on the wire.
The frontend block silently no-op'd at runtime; Copilot's
accompanying assertion only grep'd the JS source for the literal
strings, so it stayed green while the actual contract was missing.

Extend the contract to match the Copilot frontend:

  - Add ``pending_approval: bool`` + ``pending_approval_detail:
    PendingApprovalDetail | None`` to ``WorkstreamDetailResponse``,
    same shape as the dashboard / cluster live projection.
  - ``make_detail_handler`` reads ``ws.ui._pending_approval`` (only
    treats it as live when ``isinstance(_, dict)`` so MagicMock-
    based unit tests don't trip the path) and calls
    ``ui.serialize_pending_approval_detail()`` to fill the detail.
    A serializer raise falls back to ``pending_approval=True`` +
    ``detail=None`` instead of 500ing the whole response — SSE
    replay still carries the authoritative payload.
  - ``test_returns_workstream_fields`` updated for the two extra
    fields (False / None on a MagicMock UI).
  - ``test_pending_approval_fields_propagate_from_ui`` is the new
    behavioural test: stub a UI with a realistic
    ``_pending_approval`` dict + serializer return, assert the JSON
    surfaces ``pending_approval=True`` + the items list.
  - ``test_pending_serializer_failure_falls_back_to_bool_only``
    pins the defensive degradation so a future serializer
    regression can't 500 every reload.

Tests: 4822 pass (3 deselected live).  Ruff + mypy clean.

* fix(coord): three regressions on PR 447 inline tool-batch refactor

Three regressions reported during operator harness shakedown, all
landed by the inline tool-batch refactor in caa07e6:

1. ``stripAnsi`` ReferenceError on every ``tool_result``.
   ``_appendResultToRow`` called ``stripAnsi(output || "")`` but the
   helper only existed in ``ui/static/app.js`` — coord.js never
   imported or defined it.  The thrown ReferenceError propagated up
   through ``appendToolResult``, aborting the SSE handler before
   ``loadTasksDebounced()`` could fire, AND the result block never
   appended to the row, AND history replay's tool-message loop
   bailed out at the first orphan-tool-result.  Three reported
   bugs (tasks pane stops auto-refreshing, tool output missing in
   the modal, reload only rebuilds the conversation up to the first
   tool result), one root cause.

   Fix: hoist a local ``stripAnsi`` mirroring the interactive UI's
   regex.  Keep it local rather than centralised — coord and
   interactive tool-output paths have different rendering
   strategies, and the interactive helper isn't on the shared
   module surface today.

2. JSON tool output rendered as a single unreadable line.  Coord
   tool surfaces (``list_nodes``, ``tasks``, ``spawn_workstream``,
   ...) emit JSON by default, and ``textContent = stripAnsi(raw)``
   showed the whole envelope on one line.  The parent
   ``.coord-tool-row-result`` already has ``white-space: pre-wrap``
   so a ``JSON.stringify(parsed, null, 2)`` body lays out as
   intended without a nested ``<pre>``.  Non-JSON / unparseable
   output falls through to the raw cleaned string.

3. Header tier badge stuck on ``⚙ heuristic`` after the LLM judge
   landed an upgraded verdict.  ``_pickBatchTier(items)`` ran once
   at batch-creation time; later ``intent_verdict`` SSE events
   updated the per-row chip via ``_appendVerdictLineTo`` but never
   refreshed the head.

   Fix: persist the verdict's tier on ``row.dataset.verdictTier``
   (+ ``verdictModel`` when set), add ``_refreshBatchTier(batch)``
   that scans the rows and computes the cross-row best tier (LLM
   beats heuristic), and call it from ``_appendVerdictLineTo``
   whenever a row writes a verdict.  ``_pickBatchTier`` gets the
   same prefer-LLM scan so the initial render is consistent.  The
   ``intent_verdict`` cache entry tags ``tier: "llm"`` so a late
   verdict landing on a previously heuristic-only row escalates
   the badge correctly.

No Python touched; node --check on coordinator.js clean.

* fix(coord): close five Copilot review gaps on PR 447

Five distinct findings from the second Copilot pass on the inline
tool-batch refactor (the sixth — stripAnsi ReferenceError — already
shipped in 77dc24e):

1. CSS rail tucks never matched.  The ``--first / --last`` row trims
   used ``:first-of-type`` / ``:last-of-type``, but the batch
   contains other ``<div>`` siblings (.coord-tool-batch-head,
   .coord-tool-actions / .coord-tool-status) — the
   structural-pseudo-class is type-based (``div``), not class-
   based, so the first .coord-tool-row is not the first ``<div>``
   in the parent.  Selector silently no-op'd, leaving the rail
   butting against the inner top/bottom edges of the batch.  Fix:
   apply explicit ``.coord-tool-row--first`` / ``--last`` markers
   in JS at row-build time and key the CSS off them.

2. Upgrade-in-place left stale ``data-needs-approval`` markers on
   non-pending sibling rows.  The original block only added the
   attribute for items where ``needs_approval=true``, never
   clearing it for rows whose earlier (replay-time) shell tagged
   them.  ``_resolveBatchAction`` could then pick a non-pending
   row's call_id, yielding a 409 stale call_id on approve / deny.

3. Upgrade-in-place left row-level status pills out of sync with
   the SSE-authoritative item shape.  When a ``--running`` orphan
   gained a ``tool_info`` envelope, the ✓ auto pill never
   appeared; when it gained an ``approve_request`` envelope with
   policy-blocked siblings, the ✗ blocked pill / ``.error`` class
   were missed.  Batch-level state classes flipped, but per-row
   visual cues lagged.

   Fix for 2 + 3: extract ``_refreshRowStatus(row, item)`` from
   ``_renderBatchRow``.  It clears prior ``data-needs-approval`` +
   pills and re-applies from the item, preserving runtime
   ``tool_result`` errors via the new
   ``.coord-tool-row-result--error`` marker on the result block.
   Both ``_renderBatchRow`` (initial render) and the
   upgrade-in-place loop now route through it, so the two paths
   can't drift.

4. History replay defaulted ``item.needs_approval = true`` on
   every synthesized tool call.  ``_renderBatchRow`` then tagged
   the row with ``data-needs-approval="1"`` regardless of whether
   the call genuinely needed approval.  Combined with the missing
   clear in finding 2, an SSE upgrade with a mixed envelope kept
   incorrect markers on auto-approved siblings.  Drop the
   replay-time default; let SSE supply the authoritative bit when
   the upgrade fires (``_refreshRowStatus`` reads it from the
   item).

5. Tool result routed into an existing batch row didn't trigger
   ``_scheduleScroll()``.  Result blocks grow ``scrollHeight``;
   without the rAF-coalesced scroll the user pinned at the bottom
   loses their pin when the row inflates.  Add the call after
   ``_appendResultToRow`` in the early-return path so this branch
   matches ``appendMsg``'s pinning behaviour.

Plus comment-only:

6. Detail-handler comment claimed "the JSON omits the section"
   when the UI doesn't expose ``serialize_pending_approval_detail``,
   but the response always includes both keys (with ``False`` /
   ``null`` for the bool / detail).  Updated to match the actual
   shape.

Tests: ``test_workstream_endpoints.TestDetailInteractive`` +
coordinator-detail + page tests pass (14 / 0 failed). Ruff +
mypy clean.  ``node --check`` on coordinator.js clean.

* fix(coord): close 17 review findings on PR 447

Second /review pipeline pass surfaced 16 confirmed findings (1 sec
major, 1 bug major, several minor + nit); operator harness shakedown
+ this commit's stale-comment sweep adds one more.  All addressed
here.

Security:

  sec-1 (major) — make_detail_handler + make_history_handler in
  session_routes.py now invoke ``cfg.tenant_check`` after ws_id
  validation, matching every other lifted session verb (send /
  approve / close / cancel / events / attachments).  Pre-fix the
  detail response carried 5 low-data fields and history exposed
  message rows; PR 447 added pending_approval_detail to detail
  (tool previews + LLM judge reasoning) which made cross-tenant
  reads via the missing gate a real disclosure on the interactive
  surface (coord wires tenant_check=None and is unaffected).  Plus
  4 new regression tests in TestTenantCheckOnReadEndpoints that
  wire a tenant_check function into the test cfg and assert the
  gate fires on detail + history.

Bug fixes:

  bug-1 (major) — history replay used to render every fully-
  resolved tool batch as ``resolved: { approved: true }`` regardless
  of the persisted tool result content.  A denied tool round-trip
  showed the green "✓ approved" pill alongside the persisted
  "Denied by user" result text — directly contradictory state.  Fix:
  pre-scan classifies each tool message via a ``callOutcomes`` Map
  by inspecting content prefix ("Denied by user" / "Blocked by
  tool policy" / "Error:") and ``m.is_error``.  Assistant tool_calls
  render ``resolved.approved=false`` when any call's outcome is
  "denied"; the existing --running fallback covers orphan turns
  (any call lacking an outcome).

  bug-2 — _verdictSig joined recommendation/risk_level/confidence/
  reasoning only.  When a late LLM verdict text-matched the earlier
  heuristic verdict, the dedupe early-return fired before the
  row's dataset.verdictTier was updated, so _refreshBatchTier
  never escalated the header from "⚙ heuristic" to "⚖ llm".
  Fix: include verdict.tier and verdict.judge_model in the
  signature (with a "\x1f" separator instead of the empty join,
  reducing field-boundary collision risk).

  bug-3 — history replay's tool-result rendering hardcoded
  isError=false.  A runtime tool error on reload rendered without
  the .error class, --error stripe, or "✗ error:" lead.  Fix:
  the same callOutcomes pre-scan that drives bug-1's denial path
  also classifies "Error:" prefixes; appendToolResult now receives
  isError=callOutcomes.get(callId) === "error".

  bug-4 — approval_resolved derived ``wasAlways`` exclusively from
  this tab's ``batch.dataset.requestedAlways``; cross-tab "Always"
  click never propagated to peer tabs' status pill.  Fix: server's
  resolve_approval now takes a keyword ``always`` arg and includes
  it on the SSE event body; client prefers ``ev.always`` and falls
  back to the dataset stash for the hot-deploy window where the
  SSE event might briefly omit the field.

  bug-5 (nit) — appendToolBatch's create-new path overwrote
  toolRows entries unconditionally.  A partial-mapped envelope
  (some call_ids previously seen, some new) silently orphaned the
  prior batch's row pointers.  Fix: detect the partial overlap,
  console.warn, unmap the stale entries before the new batch
  claims them.

Performance:

  perf-1 — _refreshBatchTier did a querySelectorAll per verdict
  insertion; for an N-row batch upgrade this was O(N²) DOM walks.
  Coalesce via queueMicrotask + a _tierDirtyBatches Set so a burst
  of N verdict updates collapses into ONE tier scan.  Synchronous
  body extracted to _refreshBatchTierImmediate (called from the
  microtask flush).

  perf-2 — _appendResultToRow pretty-printed JSON via
  JSON.parse + JSON.stringify(parsed, null, 2) on every tool
  result with no size cap.  A 100KB JSON output stalled the main
  thread; 10 parallel tool_result events compounded.  Fix: gate
  on cleaned.length <= 32 KiB AND a first-char check (0x7B / 0x5B)
  so plain text + oversized payloads skip the parse.  Parent CSS
  is white-space: pre-wrap so raw text still wraps.

Quality:

  q-1 — deleted dead row.dataset.funcName write (no readers).

  q-2 — extracted _formatTierLabel(llmModel, hasHeuristic) shared
  by _pickBatchTier (item-driven) and _refreshBatchTierImmediate
  (dataset-driven).  Single source of truth for the tier label
  literals.

  q-3 — extracted _pendingKickerText(items) used by both the
  upgrade-in-place and fresh-build paths in appendToolBatch.

  q-4 — added string-presence assertions to
  test_coordinator_js_exposes_inline_approval_helpers covering
  the new tool-batch helpers (appendToolBatch, _morphBatchResolved,
  _resolveBatchAction, _refreshBatchTier, _refreshRowStatus), the
  --running / --pending state classes, and the callOutcomes
  outcome classifier.

  q-5 — renamed _announcePolitelyAssertive → _announceAssertive.
  Function unconditionally writes into the aria-live="assertive"
  region; "politely assertive" was contradictory.

  q-6 — rescoped the test docstring to acknowledge it covers two
  layers (Chunk 3 children-tree + PR 447 tool-batch).

  q-7 — tightened pending_approval_detail: Any → dict[str, Any]
  | None in make_detail_handler.  Mypy-confirmed.

Plus the third /review pass's q-1 stale-comment sweep:
  _resolveBatchAction's comment still claimed the server doesn't
  echo ``always`` on approval_resolved — wrong post-bug-4-fix.
  Updated to reflect that the dataset stash is now backward-compat
  fallback only, not the primary source.

Tests: 4826 pass (+4 new from TestTenantCheckOnReadEndpoints, plus
expanded assertions in TestDetailInteractive).  Ruff + mypy clean.
``node --check`` on coordinator.js clean.

Verifier confirmed all 16 findings; pass-3 /review on the
addressing-commit surfaced only 0 critical / 0 major / 2 minor /
2 nit, none blocking.  The two pass-3 minor findings are
pre-existing patterns across all lifted verbs (sync tenant_check
inside async handlers) and best addressed in a dedicated follow-up
PR auditing the whole lifted-verb surface.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: eous <13773563+eous@users.noreply.github.com>
2026-04-28 23:01:18 -07:00
Patrick Buckley 64d5205dd6 perf(session): split metacognitive nudges out of the system message
The system-message developer block was rebuilt every turn with two
unstable inputs: minute-precision current_datetime in the middle of
the composed prefix, and _pending_nudge entries appended-then-cleared
at the bottom. Both invalidated prompt-cache reuse on Anthropic /
OpenAI for the entire prefix, every turn.

current_datetime now rounds to the top of the hour. Nudges no longer
ride on the system message at all — they drain through two channels:

- tool_error and repeat ride the existing tool-result <system-reminder>
  envelope via a new MetacognitiveAdvisory ToolAdvisory subtype, drained
  in _collect_advisories alongside GuardAdvisory and UserInterjection.
- correction, denial, resume, start, completion splice as
  <system-reminder> blocks at the trailing edge of the next user
  message via a new _splice_pending_user_advisories helper.

User content passes through escape_wrapper_tags before concatenation
so a user typing literal <system-reminder> tags cannot fabricate an
envelope; the same escape now runs on advisory.render() output inside
wrap_tool_result for defense-in-depth across all advisory types.

Cancel handlers (GenerationCancelled / KeyboardInterrupt / bare
Exception) now clear _pending_tool_advisories alongside the existing
_flush_queued_messages so a queued nudge from an aborted batch
cannot leak into the next generation.

Visibility ping ([metacognition: nudge injected — ...]) preserved at
both new attach points via a single _emit_nudge_ping helper.

Also adds a "Session kind" line (interactive | coordinator) to the
composed Session Context so the model can see which manager hosts
its session.

Tests: 4847 passing (+7 new in TestMetacognitiveBuffers and
test_tool_advisory). ruff + mypy clean.
2026-04-28 19:51:00 -07:00
Patrick Buckley 08c6eeb1e5 fix(coord): close three copilot review gaps on PR 446
Copilot review on c5fe3e7 flagged three follow-ups:

1. tasks-write batch order was scheduler-dependent.  Prior comment
   claimed the result was "deterministic against the input set even
   if the dispatch order isn't" — true for the SET of tasks, but
   ``tasks_add`` appends under a per-ws lock, so the FINAL list
   ordering (and order-derived timestamps) varied with whichever
   thread happened to acquire the lock first.  Fix: when a batch
   contains any tasks-write, the dispatcher runs the WHOLE batch
   serially in input order.  Other batches stay parallel.

2. ``tasks_add`` test stubs were the wrong shape.  ``CoordinatorClient.
   tasks_add()`` returns the task dict directly with top-level
   ``id`` / ``title`` / ``status`` / ``child_ws_id`` / ``created`` /
   ``updated`` — the previous stubs wrapped it as ``{"ok": True,
   "task": {...}}`` and weakened the tests since
   ``_exec_tasks``'s summary path reads ``result.get("id")`` and
   would have seen ``"?"`` against the wrong shape.  Both stubs
   updated to match the real contract.

3. New regression test pins the input-order property on tasks-write
   batches.  ``test_tasks_writes_run_in_input_order`` captures the
   ``tasks_add`` call sequence and asserts it matches the model's
   emit order exactly — pre-fix this would be scheduler-dependent.
   Plus ``test_tasks_writes_serial_when_mixed_with_non_tasks_siblings``
   pins the same property when the batch interleaves a
   ``list_nodes`` call with two ``tasks(add)`` calls.

Tests: 4820 pass (+2 net since the prior PR 446 push).  Ruff +
mypy clean.
2026-04-28 14:23:32 -07:00
Patrick Buckley f6fbf2d85b fix(coord): list_nodes accepts flat-arg filters too
Operator's harness shakedown found list_nodes filters silently
ignored on every call:

    list_nodes(os="Linux")            → returns ALL 10 nodes
    list_nodes(has_gpu=true)          → returns ALL 10 nodes
    list_nodes(memory_gb=751)         → returns ALL 10 nodes (no
                                         node has 751 GiB; should be 0)

Storage filter pipeline is fine (pinned by an existing
test_list_nodes_filter_uses_natural_value_not_quoted).  The bug is
upstream in ``_prepare_list_nodes``: it only honoured
``args["filters"]`` (the canonical nested shape).  Several models
drop the nesting and emit each filter as a top-level kwarg —
``list_nodes(os="Linux")`` instead of ``list_nodes(filters={"os":
"Linux"})`` — and the strict prepare silently degraded those calls
to "no filter" → full-cluster return.

Fix: any top-level kwarg that ISN'T one of the four reserved control
parameters (``filters``, ``limit``, ``include_network_detail``,
``include_inactive``) is now treated as a flat filter.  Nested entries
still win on key collision so the canonical shape stays
deterministic.  Tool description unchanged so well-behaved models
keep using ``filters={...}``; the relaxation is purely receiver-side.

Tests: 4818 pass (+5 net).  Five new tests pin both shapes plus the
collision-precedence rule and the prepare→exec wiring.  Ruff + mypy
clean.
2026-04-28 14:23:32 -07:00
Patrick Buckley fca1ac3736 fix(coord): relax tasks parallel-batch rule to mixed read+write only
Operator observed the prior rule rejecting a natural decompose-the-
plan turn:

    [tasks(add×4), list_nodes, list_skills, list_workstreams]

The 4 tasks(add) calls landed (per-ws lock serialised them) but the
guard blanket-rejected EVERY tasks(...) regardless of what its
siblings actually were.  All-write batches converge under the
per-ws lock; all-read batches can't race.  The only genuinely-
hazardous shape is the read+write mix where tasks(list)
paralleled with tasks(add=...) inside ``run_one``'s
ThreadPoolExecutor has unspecified ordering and the read can land
on either side of the write.

The rule now scopes precisely:

  - All ``tasks`` writes in a batch — permitted.
  - All ``tasks`` reads in a batch — permitted.
  - ``tasks`` paralleled with non-``tasks`` siblings — permitted
    in either direction.  Non-tasks tools don't touch the tasks
    state, so there's no read-after-write surface.
  - ``tasks`` read AND ``tasks`` write in the same batch — REJECTED
    (still, because that IS the actual hazard).

Tests: 4813 pass (+4 net).  Six new tests pin the relaxation
(all-write OK, all-read OK, write+sibling OK, read+sibling OK,
non-tasks-only batch unaffected) and the one tightened rejection
case (read+write mixed in tasks specifically).  Ruff + mypy clean.
2026-04-28 14:23:32 -07:00
Patrick Buckley d0f5f50650 feat(node): auto-detect node capabilities via kernel interfaces (#445)
* feat(node): auto-detect node capabilities via kernel interfaces

Closes the operator-burden gap the harness shakedown surfaced — the
list_nodes capability/region/role filtering surface that nodes were
launching with empty.  Auto-detection runs at server startup and
populates ``node_metadata`` rows with sensible defaults that
operators can still override via the ``[metadata]`` section of
config.toml (operator-config writes win on the per-key upsert).

What's detected, all from kernel interfaces (no userspace binaries
on PATH — works the same way regardless of whether nvidia-smi /
rocm-smi / lspci is installed):

- ``gpu_count`` / ``gpu_vendor`` / ``gpu_vendors`` / ``gpus`` —
  walks ``/sys/class/drm/cardN/device/{vendor,device}`` and decodes
  PCI vendor IDs to friendly names (NVIDIA / AMD / Intel / Apple).
  Heterogeneous-GPU nodes get the first KNOWN vendor in the flat
  ``gpu_vendor`` key — never ``"unknown"`` when known vendors are
  present — so a coord filtering on ``gpu_vendor=nvidia`` matches
  nodes whose first card happened to be exotic.
- ``memory_gb`` — reads ``/proc/meminfo``, rounds GiB down so
  ``filters={"memory_gb": 32}`` doesn't match a 31.5 GiB node.
- ``cpu_model`` — first ``model name`` line from ``/proc/cpuinfo``.
- ``cloud_provider`` / ``cloud_region`` / ``cloud_zone`` /
  ``cloud_instance_type`` / ``cloud_instance_id`` — DMI sysfs
  identifies the cloud provider from BIOS/SMBIOS strings (no
  network call) and only THEN does the IMDS probe fire.  Baremetal
  hosts pay zero startup latency on the cloud path.

Hardening highlights:

- IMDS probes target the link-local IP literal ``169.254.169.254``
  for AWS, GCP, AND Azure — no DNS-resolvable hostname for any
  vendor, so a host with attacker-controlled DNS can't redirect
  the probe even when its DMI claims a cloud provider.
- Response bodies capped at 64 KiB on read; per-field strings
  capped at 256 chars and stripped of control characters before
  persistence.  Stops a hostile IMDS responder from spraying
  multi-megabyte / newline-injected payloads into ``node_metadata``
  and from there into coord-LLM ``list_nodes`` context.
- ``isinstance(doc, dict)`` guards on every JSON IMDS response so
  a non-conformant body (list / scalar / null) returns clean ``{}``
  instead of raising.
- ``collect_node_info()`` runs via ``asyncio.to_thread`` from the
  server's lifespan handler so the IMDS probe latency never blocks
  the event loop.
- GCP fans the three zone/machine-type/id probes concurrently so a
  misidentified host's worst case is one timeout window (~1 s)
  instead of three (~3 s).
- Operator opt-out via ``TURNSTONE_AUTO_CLOUD_METADATA=0`` skips the
  IMDS phase entirely; the DMI-derived ``cloud_provider`` still
  populates because that's a kernel interface.

Tests: 4807 pass (+11 net, 73 in test_node_info.py).  Ruff + mypy
clean on every modified file.  New tests pin the heterogeneous-GPU
flat-key fix, the IMDS hardening (non-dict JSON, control-char
sanitisation, body cap, per-field cap), and the GCP IP-literal
property.

* fix(node): filter synthetic display adapters + per-vendor GPU flags

PR review on 68cb0ab flagged two real issues with the GPU surface:

1. Hyper-V synthetic display adapter (vendor 0x1414, device 0x06)
   registers a /sys/class/drm/cardN entry on Linux but is NOT a
   compute GPU.  A CI runner reproduced this and came back with
   gpu_count=1 on a CPU-only VM.  Same hazard for AWS Nitro VGA,
   QEMU virtio-gpu, and any other hypervisor synthetic display
   adapter.  Fix: ``_detect_gpus`` filters DRM cards by PCI vendor
   against the GPU allow-list (NVIDIA / AMD / Intel / Apple); cards
   from other vendors are skipped entirely.  Operators with exotic
   accelerators that don't match any known vendor can still set
   ``gpu_count`` + the relevant flags via [metadata] config.

2. ``gpu_vendor`` (singular flat key) was sorted-alphabetical-first
   of the unique known vendors.  list_nodes() filtering does
   exact-equality JSON matching, so a mixed AMD+NVIDIA node ended
   up with ``gpu_vendor=amd`` and was invisible to a coord
   filtering ``gpu_vendor=nvidia``.  Fix: drop the singular key
   entirely; emit per-vendor booleans (``gpu_has_nvidia=true``,
   ``gpu_has_amd=true``) so multi-vendor nodes match EITHER vendor.
   Also add ``has_gpu=true`` for "any compute GPU at all" filtering.

Tests: 4809 pass (+2 net).  New / updated tests pin both behaviors:
- _detect_gpus: Hyper-V synthetic + arbitrary unknown-vendor card
  now filter out; mixed-known-and-unknown keeps only the known card.
- collect_node_info integration: multi-vendor node has both
  ``gpu_has_amd`` and ``gpu_has_nvidia`` set; ``gpu_vendor`` (singular)
  is asserted absent so a future regression that re-introduces it
  fails loudly.

Ruff + mypy clean.
2026-04-28 13:49:31 -07:00
Patrick Buckley 7d6b31e18a fix(coord): close gaps an operator's harness shakedown surfaced (#444)
* fix(coord): close gaps an operator's harness shakedown surfaced

Operator-driven shakedown of the coordinator tool surface flagged
five issues; this commit addresses all of them plus the review
findings against the initial fix.

1. Cancelled-mid-stream partial assistant content now carries a
   "[generation cancelled before completion]" marker.  Without it,
   ``inspect_workstream`` / ``wait_for_workstream`` callers and the
   next coord-LLM turn read the truncated text as a complete answer.
   ``_cancelled_partial_msg`` no longer ships ``_provider_content``
   (Anthropic would otherwise read that lane verbatim and bypass the
   marker; partial tool_use blocks could also leak through).

2. ``spawn_workstream`` / ``spawn_batch`` no longer surface the
   routing-proxy ``status`` field (always HTTP 200 on the success
   path).  The tool description claimed it was "lifecycle state at
   creation"; code that did ``if result["status"] == "idle"``
   silently never matched.  Lifecycle state lives on the workstream
   row — ``inspect_workstream`` is the read.  Tool JSON descriptions
   plus docs/coordinator-skills.md and docs/bulk-endpoints.md
   examples updated to match.

3. ``inspect_workstream`` not-found error string is bare ("workstream
   not found"); the structured ``ws_id`` field carries the queried
   id.  Pre-fix the error STRING echoed the id back at the caller
   who just sent it — redundant and out of step with the rest of the
   surface.  Cross-tenant + missing rows still return the same shape,
   preserving the existence-leak guarantee.

4. ``tasks(...)`` is now rejected when called in a parallel tool
   batch.  The prior shape relied on a docstring warning ("a list
   paralleled with writes can reflect pre-write state") that put
   cognitive overhead on every model invocation; turning the silent
   footgun into an explicit error means the model only thinks about
   the rule the moment it actually breaks it.  Warning dropped from
   the tasks tool description.  ``_PARALLEL_INCOMPATIBLE_TOOLS``
   constant in session.py is the extension point for any future
   tool with the same read-after-write hazard.

Plus the multi-stage code review's findings against the initial
fix (q-1 / q-2 docs drift, q-3 idiom, q-4 keys-assertion, q-5
duplicate guard) — all addressed in the same pass.

Tests: 4752 pass, +6 net since the pre-fix baseline.  Ruff + mypy
clean.  Three new tests pin the parallel-batch-rejection behaviour
on tasks (rejected when batched, runs alone, sibling tools
unaffected); existing cancel + spawn + inspect tests updated to
match the new shape.

* fix(coord): close two copilot review gaps on PR 444

Copilot review on PR 444 flagged two follow-ups:

1. Empty-content cancel divergence — when ``GenerationCancelled``
   races BEFORE the first content token, the prior shape skipped
   ``save_message`` and only appended an empty-content msg in
   memory.  In-memory and storage diverged: a rehydrate would see
   nothing in storage but the session would carry an empty
   assistant turn.  Both branches now persist; on the empty-content
   shape the marker becomes the entire message
   ("[generation cancelled before completion]") so storage matches
   the in-memory history.

2. Test stub cleanup — three new tests injected ``ui.approve_tools``
   via ad-hoc ``lambda + type: ignore[attr-defined]``.  Replaced
   with a permissive ``approve_tools`` method on ``_StubUI`` so the
   stub matches the SessionUI surface the dispatcher actually
   reads.  Tests that exercise approval pathways can still override
   per-instance.

Tests: 4752 pass.  Ruff + mypy clean.
2026-04-28 12:52:36 -07:00
Patrick Buckley 9a30530d41 feat(coord): surface child errors, isolate tool exceptions, add memory tool (#443)
* feat(coord): surface child errors, isolate tool exceptions, add memory tool

Closes four coordinator gaps identified during operator triage:

1. Child workstream errors now surface in inspect/wait. Worker-thread
   exception text is sanitized (URL userinfo masked, sk-/Bearer/ghp_/
   github_pat_/AKIA tokens redacted, capped at 1024 chars) and persisted
   to workstream_config.last_error before _emit_state("error") fires, so
   coord polling never sees state=error with a missing cause. The row
   is cleared on recovery transitions (idle/running) so a once-leaked
   exception body doesn't outlive the failure. inspect_workstream and
   wait_for_workstream return last_error for state=error rows; the
   wait surface prefers it over the assistant-tail walk.

2. Tool exceptions now return as tool_results with sibling-aware
   guidance. ChatSession._safe_prepare_tool wraps every per-call
   _prepare_tool invocation; a buggy preparer becomes an error item
   for that call only — sibling parallel tool_calls keep going,
   never orphaning the assistant message's tool_calls block.
   run_one's runtime exception path includes the exception class
   and a short note that other tool calls in the batch completed
   independently so the model can recover.

3. Memory tool exposed to coordinator with a coord-only scope.
   memory.json gains coordinator: true + interactive: true + per-kind
   kind_variants. Coord sessions see scope enum ["coordinator"] and
   an orchestration-flavored description; IC sessions see ["global",
   "workstream", "user"] and the existing flavor. Coord-scope rows
   are private to the coordinator session (children cannot read or
   write them), closing the cross-session prompt-injection lane that
   an adversarially-steered child would otherwise have. Coord
   visibility is also restricted to coord-scope only — coords no
   longer see global / workstream / user memories that belong to the
   user's interactive sessions.

4. Per-call exception isolation in tool batches. _safe_prepare_tool
   was previously the implicit shield; now it's an explicit method
   with documented invariants. KeyboardInterrupt / GenerationCancelled
   re-raise so the cooperative cancel path still works.

Other notable changes:
- LAST_ERROR_CONFIG_KEY + persist_last_error / clear_last_error /
  load_last_error / sanitize_error_text moved to turnstone.core.memory
  (the storage facade hub) — readers in coordinator_client.py import
  the constant.
- Memory scope tuples extracted to module constants
  _VALID_MEMORY_SCOPES and _IMPLICIT_SCOPE_WALK; seven inline
  duplicates collapsed.
- tools.py grows _apply_kind_variant for the per-kind tool surface;
  tools without kind_variants pass through unchanged (no spurious
  deep-copies).
- Session adds _coordinator_scope_id, _default_memory_scope,
  _implicit_scope_walk, and _record_fatal_error chokepoints so the
  worker-thread fatal path is one site rather than three.
- Removed duplicate on_error / on_state_change emits from
  session_routes.py and coordinator_adapter.py — session.send()'s
  _record_fatal_error owns the sequence now.

Tests: 4742 pass (no live), +30 net since the baseline. Ruff + mypy
clean on every modified production file.

* fix(coord): redact secrets in tool error paths via output_guard

Copilot review flagged two paths where ``str(exc)`` flowed back into
the model-facing tool_result without going through the credential-
redaction the new fatal-error path applies:

  - ``ChatSession._safe_prepare_tool``: a preparer-side exception
    becomes an error item whose ``error`` field embedded the raw
    exception text.
  - ``ChatSession._execute_tools.run_one``: a runtime tool exception
    became an ``Error executing X: <e>`` tool_result, again with
    the raw exception text.

Both now route through ``sanitize_error_text`` (sanitised log line +
sanitised tool_result), and ``sanitize_error_text`` itself was
refactored to delegate to ``output_guard.redact_credentials`` instead
of carrying its own parallel regex catalog — the audit log + post-tool
guard already use that pattern set, so the credential definition
stays in one place.

Also extended ``_RE_CONNECTION_STRING`` in ``output_guard`` to cover
``http(s)://user:pass@host`` so a misconfigured ``OPENAI_BASE_URL``
that lands in an httpx ``ConnectError.__str__`` is redacted by every
caller of ``redact_credentials`` (audit details, close-reason
persistence, last_error, the two tool error paths).  The
host (useful for triage) survives; only the password is replaced
with the standard ``[REDACTED:password]`` marker.

Tests: full suite (4745 pass), ruff + mypy clean.  Two new tests pin
the redaction behaviour in both tool error paths so a future refactor
can't drift back to leaking ``str(exc)`` verbatim.
2026-04-28 12:19:34 -07:00
Patrick Buckley 352a27915a feat(coord): per-coordinator status bar + richer history replay
Bring the coord dashboard toward parity with the interactive pane on
two operator-visible surfaces:

- Status bar pinned above the composer.  Same four cells as the
  interactive pane (model, token / context-window usage with effort
  suffix, tool calls this turn, conversation turn) driven by the
  same on_status SSE events.  ws-status-bar CSS hoisted from
  ui/static/style.css to shared_static/chat.css so both UIs read one
  copy.  StatusBar.paint helper extracted to
  shared_static/status_bar.js; both Pane.prototype.updateStatus and
  the new coord updateStatusBar delegate to it so warn/danger
  thresholds, prefix glyphs, and effort-suffix rules can't drift.
  CTX_WARN_PCT / CTX_DANGER_PCT now named constants on a single line.

- _coord_events_replay now yields the connected + status preamble
  via a shared session_replay_preamble helper in
  turnstone/core/session_replay.py.  _interactive_events_replay
  routes through the same helper so a future field add lands once.
  Coord still skips conversation history in the SSE replay (the
  dashboard fetches it via GET /history); only the status preamble
  is shared.

- History replay reconstructs tool calls.  Pre-fix, an assistant
  turn that only dispatched tools rendered as an empty bubble
  followed by raw tool-result text — the call's intent and
  parameters were lost on reload.  synthesizeHistoricalToolCall
  builds an appendToolCall-shaped item from the persisted
  function.name + function.arguments (special-casing bash so the
  shell line shows in the header).  Tool result rows now resolve
  their label from the matching tool_call_id instead of always
  printing "tool".

- onopen restores the tokens placeholder when no prior status was
  seen, so a transient SSE blip on a fresh coord doesn't leave the
  dim "Reconnecting…" copy stuck until the next live tick.

Tests: 4 new tests for the shared replay preamble (connected first,
status only when last_usage present, status payload shape, no-session
fallthrough); existing approval/verdict ordering tests refactored
through a shared make_replay_mocks helper in tests/_replay_helpers.py
that both interactive and coord suites import.
2026-04-28 10:27:26 -07:00
Patrick Buckley b1de1584c6 fix(coord): None-safe slice in _evaluate_intent projection
tasks(update) is the only mutation that allows title to be omitted,
so _prepare_tasks stores ``item["title"] = None`` for an update that
only changes status/child_ws_id. _evaluate_intent then projected via
``it.get("title", "")[:100]`` — but dict.get returns the stored None
(the default kicks in only when the key is absent), and the slice
crashed with ``TypeError: 'NoneType' object is not subscriptable``.

The exception fired before any tool in the parallel batch executed,
so the assistant's tool-call message was already on the wire while
no tool-result entries followed.  Reconstruction/sanitisation later
synthesised "Tool execution was cancelled" for every sibling — the
visible symptom that masked the real None-slice failure.

- Switch tasks/notify/task_agent/plan_agent/spawn_workstream/
  spawn_batch/send_to_workstream/close_workstream/close_all_children
  projections to ``(it.get(x) or "")[:N]`` so absent and explicit-None
  both fall back to the empty string.  The other tools weren't
  observed crashing, but the bug shape is identical at every site;
  hardening the projection layer once costs one extra ``or`` per line
  and removes the foot-gun for any future preparer that stores None.
- Regression tests reproduce the original TypeError on
  ``tasks(update)`` without title both standalone and in a parallel
  batch alongside ``tasks(add)``.
2026-04-28 09:50:48 -07:00
Patrick Buckley 39aa493d76 feat(console): per-call model + judge_model on coord composer (#440)
* feat(console): per-call model + judge_model on coord composer

Brings the landing-page coordinator composer toward parity with the
interactive new-ws modal — operators can now pick a model and judge
model per session without round-tripping through the Models admin tab.

- Add Model + Judge Model selects to the home composer's options
  panel, populated from /v1/api/models. Empty / non-string fields
  collapse to None so the factory falls back to ConfigStore defaults
  (coordinator.model_alias, judge.model).
- _coord_create_build_kwargs threads the body fields onto mgr.create.
- Console session factory accepts judge_model and overrides the
  JudgeConfig via dataclasses.replace, mirroring the server-side
  interactive factory's pattern (alias preserved for IntentJudge's
  provider/client resolution).
- Sanitise the 503 factory-misconfig response across make_open_handler,
  make_create_handler, and make_detail_handler: a new
  _safe_factory_misconfig_message helper strips control characters
  and caps at 200 chars before echoing exc text. Operators still get
  the full alias in the warning log; clients see a bounded printable
  string. Defends the user-controlled body["model"] reflection
  surface on the create path.
- _build_mgr_with_factory test helper extracted from _build_mgr so
  tests that need to capture factory kwargs don't reconstruct the
  CoordinatorAdapter + SessionManager scaffolding inline.
- Tests cover: passthrough of model + judge_model, empty / whitespace
  / non-string body fields collapsing to None, and the 503 sanitiser
  truncating + scrubbing a hostile alias payload.

* fixup: address PR #440 Copilot review

- _safe_factory_misconfig_message: hard-cap return at
  _FACTORY_MISCONFIG_MAX_LEN total (was MAX_LEN+1 because the slice
  was MAX_LEN long with the ellipsis appended on top).  Reserve one
  codepoint for the ellipsis so the cap is honoured.  Update the
  regression test to assert the tighter bound.
- Composer judge_model placeholder: "Default (agent model)" was
  misleading when ConfigStore judge.model is set — the actual fallback
  is judge.model when set, IntentJudge's agent-model fallback when
  not.  Use "Default judge model" instead so the label matches both
  configs.
2026-04-28 09:37:54 -07:00
Patrick Buckley 36f7bd5c80 refactor(console): trim landing-page friction
- Drop the duplicate "N nodes · M workstreams" header span — same data is
  already on the page.
- Drop the "+ new" workstream header button + modal; the coordinator
  composer is now the primary entry point on the landing page.
- Always render the NODES list inline; remove the cluster-summary
  compact toggle since the list already self-collapses same-prefix
  nodes into groups.
- Replace the meta node-detail page (#view-node) with direct navigation
  to /node/{node_id}/. Removes drillDownToNode, loadNodeDetail,
  _loadNodeMetadataPanel, the popstate "node" branch, and the
  currentNodeId/currentServerUrl state.
- popstate now falls back to showHome() for unknown state shapes so a
  back-nav from a tab on an older build doesn't no-op.
- test_index_landing_surfaces guards the removed IDs from
  reintroduction.
2026-04-28 08:52:32 -07:00
Patrick Buckley ea204226ad chore: bump version to 1.5.0a5 2026-04-28 00:32:01 -07:00
Patrick Buckley 6f5cb33923 feat(coord): composer parity with interactive — stop/queue/attach (#438)
* feat(coord): composer parity with interactive — stop/queue/attach

Bring the coordinator one-pane UI to feature parity with the
interactive composer: in-composer Stop button replaces Send during a
turn, queue-while-busy with !!! priority + dismiss, paperclip attach
+ drag/drop/paste. The coord backend already supported all three
(lifted send/cancel/attachment handlers, emit_message_queued=True,
supports_attachments=True); this wires the UI through.

Backend:
- Wire make_dequeue_handler(coord_endpoint_config) so DELETE
  /v1/api/workstreams/{ws_id}/send works for coord-kind workstreams.
- Add the matching OpenAPI EndpointSpec.
- Five new test_dequeue_* tests (success, not_found, missing msg_id,
  unknown ws, scope gate) pin the URL/method/scope contract.

Frontend extraction:
- New shared modules composer_attachments.js (createAttachmentController)
  and composer_queue.js (createQueueController) replace ~300 LOC of
  pre-existing duplication between the interactive Pane and the coord
  IIFE. Both panes now share one source of truth for the chip pipeline,
  optimistic queue bubble, and busy-edge promote sweep.

Coordinator pane:
- Composer constructor adds attachments/stopBtn/queueWhileBusy/
  busyPlaceholder/dragDrop options.
- setBusy now drives off SSE state_change (running/thinking/attention →
  busy; idle/error → idle), with composer.setBusy unconditional and the
  edge-only work (timer cleanup + queue.onIdleEdge) gated on the actual
  transition.
- Cancel uses the in-composer Stop with a 2s "Force Stop" affordance +
  10s safety auto-recover; the legacy header-mounted #coord-cancel-btn
  is removed.
- coordCloseSession suspends SSE before close and re-establishes it on
  any failure path so the UI never goes dark on a still-alive session.
- Race handling: bind() releases the queued slot server-side when the
  bubble was already dismissed or promoted; rehydrate re-checks getWsId
  in its .then so a stale-tab response can't clobber the new tab's
  chips.

Interactive pane:
- Pane class adopts the same controllers via this.attachments /
  this.queue. Pane.prototype.uploadAttachment, _renderAttachmentChip,
  _swapPlaceholderChip, _removeAttachmentChip, removeAttachment,
  rehydrateAttachments wrapper, addQueuedMessage, _dequeueMessage, and
  _promoteQueuedMessages are all gone — the controllers own the state.
- setBusy collapses to the same shape as coord: composer.setBusy +
  edge calc + queue.onIdleEdge on idle.

CSS:
- Move .msg-queued / .queued-badge / .queued-dismiss styles from
  ui/static/style.css into shared_static/chat.css so both panes share
  one rendering.
- Add .coord-drop-target overlay rule so the coord pane shows the
  drag-and-drop affordance.

Tests pass: 160 in the impacted suites (coord endpoints + attachments
+ session routes), including 5 new dequeue tests for coord.

* fix(coord): Copilot review + lint follow-ups

Lint:
- ruff: cast(MagicMock, ...) → cast("MagicMock", ...) under
  ``from __future__ import annotations`` (UP037).

Copilot review (PR #438):
- composer_queue _sendDelete now invokes onAfterDequeue on success
  so a bind() race-DELETE (queued bubble dismissed pre-bind or
  promote sweep raced ahead) still rehydrates the caller's chip pile;
  released attachment reservations no longer linger invisibly until
  the next page load.
- Coord's createQueueController gains onAfterDequeue: attachments.
  rehydrate(). The previous omission was a v2 review carry-over from
  before coord supported attachments — now it does, so the same
  contract as interactive applies.
- Both panes' send-response handler now accepts status:queued without
  a queuedEl (SSE-not-yet-connected race on initial load): flips busy
  so subsequent sends queue correctly. The current message keeps its
  optimistic user bubble — accepted UX gap (no in-UI dismiss for
  THIS message) since flipping a rendered user bubble into a queued
  one mid-stream would be jarring.
- Doc updates: chat.css comment + composer_queue.js module docstring
  refer to the renamed onIdleEdge() instead of the removed
  promote()/promoteQueuedMessages.
2026-04-28 00:25:32 -07:00
Patrick Buckley 5ad5f4d12a chore(compose): raise per-node memory caps to fit current footprint
Cluster nodes were OOM-killing under MCP child-process load with the
old 384M/0.5cpu budget chosen for a leaner, pre-MCP turnstone. Bump
each cluster server to 4G/4cpu and postgres to 4G/4cpu. The single-node
server, console, and channel services remain uncapped.
2026-04-27 23:36:09 -07:00
Patrick Buckley dea2729292 refactor(coordinator): rename task_list → tasks, doc/prompt sweep (#437)
Four themes from a coordinator-feature shakedown:

1. Correctness fixes (return shapes / examples / behavior)

   - tools_coordinator.md: drop fake skill names from spawn examples;
     fix wrong kwarg ``node_id=`` → ``target_node=``.
   - wait_for_workstream.json: document ``message`` + ``truncated``
     per-ws fields (always enriched in the client; the JSON shape
     lagged the docstring).
   - cancel_workstream.json: document the conditional ``dropped``
     payload — ``was_running`` always present when ``dropped`` is,
     ``pending_approval`` and ``queued_messages`` conditional sub-shapes.
   - spawn_workstream.json: document full return shape including
     ``routing_strategy ∈ {rendezvous, target_node, resume}`` and
     ``status``.
   - close_all_children.json: clarify ``skipped`` covers BOTH
     hard-deleted children AND already-closed-and-evicted children
     (wire shape doesn't distinguish); drop incorrect "echoed back
     in response" claim — server returns ``{status, closed, failed,
     skipped}``, never echoes ``reason``.
   - console/server.py: comment in ``_fanout_on_children`` clarifying
     that the 400 "No session" branch fires for cancel-cascade
     callers and is unreachable from close_all_children (close
     handler 404s instead).
   - coordinator_client._utc_now_iso(): switch to bare ISO format
     matching the rest of the storage row format used in the codebase.

2. Tightened the 11 longest tool descriptions (~23% cut on the
   coord set). Removed ALL-CAPS emphasis, normalised em-dashes,
   dropped informal phrasing. No new claims.

3. Removed static approval annotations from descriptions.
   Approval is governed at runtime by the unified ``approve_tools``
   body and admin-defined ``tool_policies`` (#436); static
   "Auto-approved" / "Approval required" / per-action approval
   tags become a stale signal. Field names (``pending_approval``)
   and operational verb behaviour ("cancel unblocks pending
   approvals") stay.

4. Renamed ``task_list`` coord tool → ``tasks``. The previous name
   compounded the bare word ``task`` (which collides with chat-template
   channels on local models — same reason ``task_agent`` carries
   the suffix); the plural form sidesteps the collision and reads
   more accurately, since the tool acts on the whole list rather
   than a single task. Sweep covers tool JSON, Python methods (5
   client methods + 2 session methods + 1 helper + 1 constant),
   audit event name (``task_list.update`` → ``tasks.update``), log
   tag (``task_list.corrupt_envelope`` → ``tasks.corrupt_envelope``),
   frontend SSE event matcher, prompts, docs, and tests. CHANGELOG
   entry added.

Plus: dropped the ENV block (Output Environment / Available
rendering / Formatting principles) from coordinator system
prompts. Coordinators orchestrate rather than render rich output
to the user, so the rendering capability matrix is not actionable
for them. Coord prompt drops ~29% (6309 → 4493 chars).

SDK regeneration via ``generate-types.py`` updates both
``openapi-console.json`` (the rename's downstream change) and
``openapi-server.json`` (PR #436 drift — its merge added
``pending_approval_detail`` + ``recent_auto_approvals`` fields to
the Python schemas but didn't regenerate the JSON artifact).

## Behavior changes (operator-visible)

- Audit event name: ``task_list.update`` → ``tasks.update``.
  Audit dashboards / SIEM filters / log greps that pinned the old
  prefix should update.
- SSE ``tool_result`` events now ship ``name="tasks"`` for the
  scratchpad tool. The bundled coord-tree UI is updated atomically;
  external consumers reading SSE events by tool name need to update.
- Existing task envelopes in production storage have ``+00:00``
  timestamps from the old ``_utc_now_iso``. New writes are bare;
  old rows are not backfilled. Within an envelope you may briefly
  see mixed formats until each row is re-touched. No code path
  string-compares timestamps within an envelope, so this is
  cosmetic.

## Validation

- ``ruff check`` + ``ruff format --check`` clean
- ``mypy turnstone/`` clean (175 source files)
- ``pytest -m "not live"`` — 4679 passed, 3 deselected
2026-04-27 22:51:31 -07:00
Patrick Buckley fb44652850 refactor(core): unify approve_tools across both kinds (#436)
* refactor(core): unify approve_tools across kinds + judge visibility + perf

Lift WebUI.approve_tools to SessionUIBase so both interactive and
coordinator workstreams run the same body. The shared body now owns
tool-policy gating, per-tool auto-approve, blanket carve-out for
__budget_override__, activity tagging, heuristic-verdict persistence,
and the approve_request/approval_event blocking pattern. Subclass
hooks layer kind-specific surfaces on top.

This closes the drift the LLM-judge audit flagged on coord — the
judge (heuristic + LLM tier) now sees actual tool args for every
coord tool call instead of empty func_args. spawn_batch projects
the full children list so a malicious mid-batch entry is no longer
hidden.

= Unification core =
- SessionUIBase.approve_tools: lifted body covering policy / per-tool
  auto-approve / blanket / activity tagging / heuristic-verdict
  persistence / approval gate
- _APPROVAL_WAIT_TIMEOUT class constant + _record_judge_metric hook
- WebUI.approve_tools deleted; _record_judge_metric override fires
  per-node MetricsCollector.record_judge_verdict
- ConsoleCoordinatorUI.approve_tools deleted; _record_judge_metric
  + on_intent_verdict overrides fire ConsoleMetrics.record_judge_verdict
- ConsoleMetrics.record_judge_verdict + turnstone_judge_verdicts_total
  in /metrics text output (cluster PromQL rolls coord+interactive up
  uniformly)
- _console_metrics class attribute wired in console lifespan
- Frontend: coord SSE event tools_auto_approved -> tool_info for parity

= Judge args visibility =
- _evaluate_intent populates func_args for all coord tools that hit
  approval (spawn_workstream / spawn_batch / send_to_workstream /
  close_workstream / close_all_children / cancel_workstream /
  delete_workstream / task_list)
- spawn_batch projects every child's skill / initial_message[:200] /
  target_node so the judge sees the full fan-out (was first child only)
- fire_judge_verdict_metric helper collapses 4 sites of identical
  record_judge_verdict shape across WebUI + ConsoleCoordinatorUI

= Hardening =
- __budget_override__ carve-out reads from pre-filter items list, not
  post-filter pending; policy block skips matching the synthetic
  name entirely so a wildcard `*: allow` cannot strip the override
  before the gate sees it
- _persist_intent_verdict default_tier parameter so heuristic + llm
  paths share the storage write helper

= Performance =
- TTL cache on list_tool_policies in turnstone/core/policy.py
  (60s, keyed by org_id, lock-free hits)
- Storage-layer invalidation: create/update/delete_tool_policy on
  both SQLite and PostgreSQL backends call invalidate_policy_cache
  (covers admin-API path + direct test fixtures + any future caller)
- Admin-API handlers also call invalidate_policy_cache as
  defense-in-depth
- storage.create_intent_verdicts_bulk on both backends: one
  multi-row INSERT + one commit instead of N round-trips. approve_tools
  switches to the bulk path so a fan-out turn no longer pays N x commit
  before the approval prompt enqueues
- _persist_intent_verdicts_bulk helper on SessionUIBase

= Test coverage =
- tests/test_coord_ui_approve_tools.py (NEW, 17 cases): inheritance
  regression, tool-policy deny/allow/mixed on coord, heuristic verdict
  persistence (bulk path), activity tagging on auto-approve and pending,
  judge_pending dynamic flag (true + false), event-name parity,
  per-tool auto-approve, __budget_override__ carve-out under blanket
  + wildcard policy, _record_judge_metric wired/unwired, on_intent_verdict
  llm-tier metric
- tests/test_console_metrics.py: 3 cases for the new
  record_judge_verdict counter
- tests/test_judge_storage.py: 3 cases for create_intent_verdicts_bulk
- tests/test_coordinator_tools.py: 3 cases pinning the spawn_batch
  full-children projection (truncation, mid-batch visibility, empty
  defensive)
- tests/conftest.py: autouse _clear_policy_cache fixture so the
  process-level cache doesn't leak between tests with distinct storage
  instances

= Drift fixes (review feedback) =
- Refresh stale "no-op on coord" comments now that coord overrides
  the hook
- WebUI.on_plan_review timeout uses self._APPROVAL_WAIT_TIMEOUT
  instead of literal 3600
- Drop redundant bool() wrapper around any() in judge_pending
- Rephrase broken docstring grammar in _coord_spawn_metrics
- Hoist redundant get_storage import out of approve_tools per-item loop
  (folded into _persist_intent_verdicts_bulk helper)

= Validation =
- pytest -m "not live": 4679 passed, 3 deselected
- ruff check + ruff format: clean
- mypy: no issues in 175 source files

* fix(approval): apply Copilot feedback on PR #436

- Policy-cache invalidation now drops both the org-scoped slot AND the
  default ``""`` slot on ``create_tool_policy`` for both SQLite and
  PostgreSQL backends. ``list_tool_policies("")`` returns rows from
  every org_id, and the production evaluators (SessionUIBase.approve_tools
  / cli.py) read with the default ``org_id=""``, so an org-scoped insert
  that only invalidated its own slot would leave the default cache slot
  stale until the TTL window expired.
- Cap ``reason`` to 200 chars in ``_evaluate_intent`` for ``close_workstream``
  and ``close_all_children`` — both fields are LLM/user-provided and the
  preparer doesn't size-limit them, so an unbounded reason could bloat
  the persisted verdict row's func_args. Matches the cap applied to other
  free-form coord tool fields (initial_message, message, title).
- Refresh ``_PolicyCache`` docstring: it claimed lock-free reads on
  cache hit but ``get()`` always acquires ``self._lock``. Updated to
  reflect that the lock is held briefly to copy the policies reference.

Validation: targeted suite 201/201, ruff + mypy clean.
2026-04-27 21:52:57 -07:00
Patrick Buckley 1fe800f832 refactor(ui): drop legacy ts-composer prefix on shared composer classes
Follow-up to #434.  That PR unified the chat-message primitive on .msg
and noted that the parallel .ts-composer prefix on the shared composer
widget was still in place; this drops it so the widget sits in the
shared/* vocabulary the same way .msg does.

Mechanical 1:1 rename (`ts-composer` -> `composer`) across:

  shared_static/chat.css       — 48 selectors
  shared_static/composer.js    — 19 className strings
  ui/static/style.css          — 11 per-node UI overrides
  ui/static/app.js             — 7 chip queries / className strings

Pre-rename collision check confirmed clean: the only `composer`-substring
matches in the codebase were IDs (#coord-composer-mount, #coord-composer-
panel, #coord-composer-503, #home-coord-composer-mount — IDs are a
different namespace from classes) and the unrelated console
.home-composer-banner / .home-composer-error pair (different prefix).

CSS specificity audit (scripts/css_specificity_audit.py): 26 findings on
origin/main, 26 on this branch — no new cascade flips.

Tests: 189 affected tests pass (test_app_js, test_webui_content,
test_webui_auto_approve_visibility, test_html, test_web_helpers,
test_coordinator_adapter, test_coordinator_client).

Manual visual verification of composer surfaces (textarea, send button,
stop button, attach button + file picker, chip pills + remove buttons,
options panel toggle, paste-image and drag/drop attach paths, stacked
layout used by creation forms) recommended before merge.
2026-04-27 20:35:38 -07:00
Patrick Buckley 94edd741d3 refactor(ui): drop legacy .ts-msg* dual-classing in chat surfaces (#434)
* refactor(ui): drop legacy .ts-msg* dual-classing, chat surfaces share .msg primitive

Third and final follow-up after #431 stripped the data-design="v1"
gate.  This drops the parallel .ts-msg* family that had been kept as a
transitional bridge during the gated rollout.  Per-node UI now renders
pure .msg classes (previously dual-classed as
"ts-msg ts-msg--user msg user"), matching the coordinator chat view
which already used pure .msg*.

chat.css: deleted the ~210-line legacy .ts-msg* rule block (Messages +
floating action toolbar + mobile + reduced-motion sections); renamed
.ts-msg.ts-approval--inline to .msg.ts-approval--inline; restored the
streaming-markdown rationale (white-space: normal intent + partial-fence
behavior + .msg-user-text path) on .msg-body that previously lived on
the deleted .ts-msg-body, with white-space: normal now declared
explicitly so a future "simplification" can't silently break streaming.

ui/static/app.js: dropped the ts-msg* half of every dual-class string
and updated querySelector callsites (.ts-msg--user -> .msg.user,
.ts-msg--assistant -> .msg.assistant).

ui/static/style.css: renamed all .ts-msg--* selectors to .msg.*; removed
two now-dead override rules (.ts-msg.msg:not(.tool) and
.ts-msg-body.msg-body font-family overrides) that existed solely to
unwind the legacy .ts-msg font-mono default that's now gone.

The .msg.ts-approval--inline selector intentionally keeps the .msg
qualifier (rather than bare .ts-approval--inline) so its (0,2,0)
specificity ties with .ts-approval.approved/.denied/.error and the
later-cascade rule wins; without the qualifier those state classes
would suddenly flip the inline-approval card colour based on state.

Composer rename (.ts-composer-* -> .composer-*) deferred to a follow-up
PR; ~80 occurrences across composer.js + chat.css would have obscured
this verification.

Tests: 538 affected tests pass (test_app_js, test_webui_content,
test_webui_auto_approve_visibility, test_web_helpers, test_html,
test_auth, test_console, test_api_versioning, test_coordinator_*).
Visual verification (message cards, hover toolbar, approval/denial/error
cards in light + dark themes) recommended before merge.

* docs(ui): clarify .msg.reasoning emission comment per Copilot review

The previous wording — ".reasoning as a bare role class is no longer
emitted" — implied .reasoning is never emitted, but the new className
is "msg reasoning" so .reasoning IS emitted, just always alongside .msg.
Reword to make the actual invariant (never on its own) explicit.
2026-04-27 20:19:59 -07:00
Patrick Buckley 4b5edce8c5 fix(css): two cascade-flip bugs found by specificity audit (#433)
* fix(css): two cascade-flip bugs found by specificity audit

PR #431 stripped [data-design="v1"] from ~400 rules, dropping each by a
specificity tier; two cascade flips (#header outranking .appbar, #header h1
outranking .appbar-title) were caught visually during that PR's review and
fixed by renaming id="header" → id="ui-header" on the per-node UI page.
This is the audit follow-up; it found two more:

- textarea.skill-content-area (was .skill-content-area) — bumped to (0,1,1)
  so the rule ties with `.admin-modal textarea` (0,1,1) and wins on source
  order. Without the bump, min-height: 220px was clobbered to 40px by the
  modal default and the spec-content textarea rendered short. The three
  !important markers (font-family/size/line-height) are now redundant
  against the modal's font: inherit shorthand and are dropped.

- h3.skill-spec-heading — removed `font-size: inherit;`. The author wrote
  it to "reset UA defaults" but it locked font-size to the parent's
  (~14-16px) at (0,1,1), silently overriding `.skill-spec-heading`'s 10px
  at (0,1,0). The bare class already beats UA `h3` on specificity (class >
  tag), so no font-size reset was needed; the `margin-block: 0` line stays
  because the bare class's `margin: 14px 0 6px` shorthand may not reset
  the UA's logical margin-block-start/end on every engine.

Adds scripts/css_specificity_audit.py — the audit tool. It parses every
CSS file referenced from the project's three HTML entry points, computes
selector specificity (incl. :not/:is/:has math, attribute selectors, and
!important), and flags every place an unscoped legacy rule could outrank
a bare-class designed primitive. Honours per-page stylesheet manifests,
state-pseudo subset gating (a `:hover` rule overriding a resting-state
base rule is intentional, not a flip), and shorthand→longhand expansion
for font/padding/margin/border/background. Triage of remaining findings
(26 id-tier in default mode, 74 total at --all-tiers) confirmed all are
intentional designer overrides — id-scoped buttons, BEM modifier classes,
contextual ancestor selectors, last-child margin reset, [hidden] toggle.

* fix(css-audit): correct two cascade-resolution bugs flagged by Copilot

1. _parse_declarations dict insertion order didn't update on overwrite, so
   a sequence like `font-size: 13px; font: inherit; font-size: 12px;` would
   iterate as (font-size=12px, font=inherit) and the shorthand expansion
   then clobbered font-size back to `inherit` — wrong.  Delete-then-insert
   on overwrite so the last occurrence lands at the dict's tail and the
   shorthand expansion sees the real source order.

2. The cascade-winner tie-break used `rule.line_no` only, ignoring the
   stylesheet load order.  A rule at line 1000 of `base.css` looked
   "later" than a rule at line 50 of `style.css`, even though the page
   loads `base.css` BEFORE `style.css`.  Sort by `(file_index, line_no)`
   keyed off the element's per-page stylesheet manifest instead.
2026-04-27 20:05:25 -07:00
Patrick Buckley 83a97ba485 fix(ui): preserve approval pill when tool errors (#432)
* fix(ui): preserve approval pill when tool errors

When an approved (or auto-approved) tool subsequently failed during
execution, both `replayHistory` and `appendToolOutput` located the
existing `.ts-approval-badge` and overwrote its className + textContent
with the `--error` variant — losing the record that the user had
approved the call.

Append a separate `--error` pill as a sibling of the existing approval
pill instead. The `.ts-approval` parent is `flex-direction: column` with
a 6px gap, so the two pills stack vertically and read as a small
status timeline ("you approved this, then it errored"). Idempotency
guard via `querySelector(".ts-approval-badge--error")` so duplicate
fires don't stack badges.

CSS classes are unchanged (the `--error` modifier already exists in
both per-node and shared chat stylesheets).

Adds a static-string guard in `tests/test_app_js.py` that pins both
call sites and forbids the mutate-in-place anti-pattern via a regex
that pairs a queried `.ts-approval-badge` handle with an `--error`
className overwrite.

Deferred from #431.

* refactor(ui): extract appendToolErrorBadge helper, broaden test guard

Address Copilot feedback on #432:

- Extract the duplicated 5-line error-pill construction into a single
  module-level `appendToolErrorBadge(blockEl)` helper next to the
  other approval-related helpers (`buildToolDiv`, `renderVerdictBadge`,
  `toggleVerdictDetail`). Reduces drift risk on ARIA / class / text
  string between the two call sites.

- Loosen the affirmative test check from a literal substring keyed on
  the local variable name to a regex matching any
  `querySelector(".ts-approval-badge--error")` lookup, in either
  quote style, in either guard idiom (`if (!q) {...}` at a call site
  or `if (q) return;` inside the helper). A future refactor that
  preserves behaviour shouldn't trip CI on cosmetics.

- Broaden the anti-pattern regex to accept single quotes and to
  catch the `classList.add("ts-approval-badge--error")` form on a
  queried badge handle, not only `className = "..."`.
2026-04-27 19:37:26 -07:00
Patrick Buckley cc20c7008d refactor(css): unify design system, eliminate data-design="v1" gating (#431)
* refactor(css): unify design system, eliminate data-design="v1" gating

Strip the [data-design="v1"] attribute that was wrapping every DS rule
since #389 and never came back out. Result: every styled element on
coord/ui pages had two CSS rules (default + v1-gated), reviewers
couldn't tell which one rendered, and the bundle shipped duplicates.

Changes:
* Strip [data-design="v1"] prefix from ~400 gated rules.  Remove the
  attribute from coordinator/index.html, ui/index.html, preview.html.
* Merge shared_static/design/* into pre-v1 sheets:
    tokens + typography  → base.css (:root, dark default)
    appbar + panel + buttons + pills + field  → ui-base.css
    message primitives   → chat.css
    sidebar + approval-dock → console/static/coordinator/coordinator.css
  (new file linked from coord only — admin no longer ships ~9KB of
  coordinator-only chrome on first load).
* Delete preview.html + 5 preview-only orphan stylesheets (topbar /
  stats / feed / fleet-grid / live-feed).  Drop the empty
  shared_static/design/ directory.
* Drop unused primitives the merge dragged in: .pill / .k-badge /
  .chip / .field / .t-* utilities / .side-item / .shell.  Drop unused
  tokens (--accent-c / --accent-l / --row-h / --density / --gap /
  --font-display alias).  Find-replace var(--font-display) →
  var(--font-ui) across 6 files (103 sites).
* Standardize on the DS font stack: Inter body, JetBrains Mono code.
  Admin's body shifts from IBM Plex Mono → Inter via the alias rename.
* Re-tune legacy --bg/--bg-surface/--bg-highlight/--bg-elevated from
  steel-blue to neutral charcoal so admin and v1 pages share one
  palette.  Drop the cyan radial-gradient overlay on body that added a
  blue tint to the formerly-blue bg.
* Polish:
  - .msg.tool / .ts-msg--tool / .ts-approval / inline-approval all use
    --cyan instead of amber, removing the user/tool colour collision.
  - .msg-action-btn reverts to icon-button styling (transparent,
    28x24) after the merge gave it text-button chrome that dwarfed the
    13-14px icon glyphs inside.
  - Light-mode composer contrast: flip .ts-composer surface roles
    (wrapper recessed, textarea elevated) so the textarea reads
    against its container; bump .dashboard-composer textarea
    border-bottom + options panel surface so they're visible on white.
  - .pane-messages padding 20px -> 16px 12px and gap 14px -> 0 (the
    flexbox gap was stacking with .ts-msg margin-bottom for ~18px
    inter-card spacing); .ts-msg/.msg margin-bottom 8px -> 4px.
  - Restore WCAG 2.5.5 36x36 touch target on .msg-action-btn.
  - Rename ui's <div id="header"> to id="ui-header" so the legacy
    #header chrome no longer outranks the .appbar primitive on the
    per-node page (3 getElementById calls in app.js updated).
  - Drop chat.css link from coord (coord renders pure DS classes; the
    composer.js consumer of chat.css is on admin + ui only).

Verified: ruff + mypy clean (175 files); 4653 non-live tests pass;
zero data-design / --font-display / shared_static/design hits remain.
Net source change: -1885 lines (924 added, 2809 deleted across 28 files).

* build: include coordinator.css in wheel; drop dead design/ glob

The previous commit added turnstone/console/static/coordinator/coordinator.css
(coord-only chrome moved out of console/static/style.css) but didn't
update the [tool.hatch.build.targets.wheel] include list, so CI's
wheel-completeness check failed.

Also drop the now-stale 'turnstone/shared_static/design/**/*' glob —
that directory was deleted in the same v1-elimination commit.

Verified locally: replicating the CI step's source-vs-wheel diff
returns MISSING: none.

* fix(css): address Copilot review feedback on PR #431

* coordinator/index.html — comment now correctly points to the moved
  .sidebar rules at console/static/coordinator/coordinator.css (was
  console/static/style.css before the perf-1 split-out).
* ui/static/index.html — restore <h1 class="appbar-title">; the cascade
  conflict that motivated the h1→div change is gone now that the
  wrapper id was renamed away from #header (the legacy #header h1 rule
  no longer matches).  Page semantics + accessibility regain the
  top-level heading.
* governance.js — drop the inline font-family:var(--font-ui) on the
  config-key <code> elements; let them inherit the global mono default
  from base.css.  The inline style was an artifact of the
  --font-display → --font-ui find-replace; the original Outfit was
  already odd on a <code> tag.
* ui-base.css — typography-helpers comment said "Body text still
  inherits var(--font-mono) at 13px" but base.css now sets var(--font-ui)
  at 14px.  Reword to match current defaults.
2026-04-27 19:04:15 -07:00
Patrick Buckley 9b5096fe3c fix(approve): visibility for child tool calls bypassing operator gate (#430)
* fix(approve): visibility for child tool calls bypassing operator gate

When a coord LLM spawns a child with `skill="X"`, the skill template's
`allowed_tools` JSON list silently populates the child UI's
`auto_approve_tools` set. Tool calls whose names are in that set
short-circuit the approval gate without prompting the operator —
matching the user-reported bug "tool calls of children occasionally
getting approved instead of waiting for approve/deny".

The auto-approve paths themselves are unchanged (Option C — visibility
only). Surfaces:

- Per-item annotations: each pending tool gets `auto_approved=True` +
  `auto_approve_reason` ("skill" / "always" / "policy" / "blanket" /
  "auto_approve_tools") at the four gate-bypass paths.
- Per-ws ring buffer (cap 10) of recent bypasses, exposed via
  `/dashboard` and the cluster live-bulk projection so the coord-
  tree row can render an "auto-approved by ..." pill.
- `tool.auto_approved` audit row per `approve_tools` call —
  forensic durability beyond the in-memory ring buffer.
- Per-ws WebUI page: inline "auto: <reason>" badge next to each
  tool name, so an operator who clicks through from the coord tree
  to the child's page sees the same bypass signal.

Persistence across UI rebuilds:
- The ring buffer is in-memory only; a saved-workstream rehydrate /
  coord→node click-through / process restart all build a fresh UI.
  `replay_recent_auto_approvals_from_audit` runs at the end of
  `SessionUIBase.__init__` and re-seeds the buffer from recent
  `tool.auto_approved` audit rows scoped to this ws_id.
- Adds `resource_id` filter to `list_audit_events` (protocol +
  SQLite + Postgres) so the replay is a single indexed query.

Source provenance:
- `_auto_approve_tools_source: dict[str, str]` per UI tracks which
  writer added each tool name to `auto_approve_tools` ("skill" at
  skill-template setup time, "always" on Approve+Always click).
  Lets the dashboard pill distinguish a skill-driven bypass from
  an explicit operator-Always click — those are very different
  signals that previously rendered the same.

Magic-string drift mitigation:
- `AutoApproveReason` constants in `core/session_ui_base.py` lift
  the five reason strings into a single source of truth.
- `KNOWN_AUTO_APPROVE_REASONS` JS constant + validator render
  unknown reasons as "unknown" with a console.warn instead of
  rendering raw (a typo would otherwise silently desync wire ↔
  pill).

Recording-leak fixes (q-2 from review):
- Policy `allow` partial-resolve now records the policy-tagged
  items at two previously-leaking branches: the early-return-on-
  deny path and the still_pending-non-empty fall-through to the
  prompt path.

Other review fixes:
- Heuristic verdict surfaces consistently as `heuristic_verdict`
  in both `_serialize_approval_items` and the dashboard
  serializer (was inconsistent: one emitted `verdict`, the other
  `heuristic_verdict`). app.js updated to read either key for
  mid-deploy compatibility.
- `_tag_auto_approved` helper on SessionUIBase replaces the
  verbatim tag loops previously copy-pasted across WebUI and
  ConsoleCoordinatorUI.

* fix(approve): apply Copilot review feedback on PR #430

- coordinator_ui: use ``approval_label or func_name`` for the
  ``auto_approve_tools`` subset check, matching WebUI.  Pre-fix
  an "Approve + Always" entry whose approval_label differs from
  func_name (skill__name, mcp_resource__uri) wouldn't match on
  the coord page and the operator would be re-prompted.
- _parse_audit_timestamp: treat naive ISO strings as UTC.  Audit
  rows are written via ``datetime.now(UTC).strftime(...)`` with
  no timezone marker; ``datetime.fromisoformat`` returns a naive
  datetime, and ``.timestamp()`` on a naive datetime interprets
  it in the server's local timezone — wrong on any non-UTC
  server.  Stamp UTC explicitly before converting.
- server.py: drop the dead ``pending = []`` after the blanket
  tag — the function returns inside the same block without
  reading ``pending`` again.
- _protocol.py: fix docstring reference from
  ``_replay_recent_auto_approvals`` to
  ``replay_recent_auto_approvals_from_audit`` (the actual
  method name).
2026-04-27 15:51:01 -07:00
Patrick Buckley d15f182b80 fix(coord): tree UI not updating when LLM deletes workstream (#429)
* fix(coord): tree UI not updating when LLM deletes workstream

The coord LLM's `delete_workstream` tool wiped the storage row but
fired no SSE event, so a long-lived dashboard tab kept the deleted
child visible (with its last-known idle/closed state) until a full
reload. A coordinator that spawns→completes→deletes children would
leave an ever-growing tree.

Fix: add `SessionManager.delete()` that drops the in-memory slot if
present and emits `ws_closed` with `reason="deleted"` (mirrors
`close()`'s shape). Wire `delete_workstream_endpoint` to call it
after the storage delete succeeds, snapshotting the workstream's
name into the event payload before the row is wiped. The cluster
collector → coord adapter chain re-emits as `child_ws_closed`; the
browser's existing `handleChildClosed` already keys on
`reason === "deleted"` to mark the row, so no JS changes needed.

Event emit is best-effort — a fan-out failure logs a warning but
doesn't roll back the storage delete (the row is already gone).

* fix(coord): apply Copilot review feedback on PR #429

- server.py: clarify that ``name`` is forwarded to mgr.delete only
  (not into the audit detail) — comment previously claimed both.
- test_session_manager.py: extract ``mgr.delete(ws_id)`` to a local
  before asserting (CodeQL: no side-effecting calls inside ``assert``,
  which would be stripped under ``python -O``).
- test_workstream_endpoints.py: docstring said "Yield" but the
  fixture ``return``s; switch to "Return".
2026-04-27 14:44:01 -07:00
Patrick Buckley e33519275e docs(coord): fix wait_for_workstream message-field claim re deleted state
Copilot caught a doc/code mismatch from the q-2 cleanup: the
docstring still claimed `closed` / `deleted` / `denied` all return a
sentinel, but the `deleted` branch was dropped (hard deletes cascade
rows out of storage so the state is unreachable). Update the
docstring to align with `_wait_message_for`'s actual behaviour —
`deleted` falls into the same null-message shape as a still-running
entry.
2026-04-27 13:56:16 -07:00
Patrick Buckley 91b07aaf4b feat(coord): bundle child last-message inline in wait_for_workstream
Each per-ws snapshot now carries `message` + `truncated` so the
coordinator LLM doesn't need a follow-up `inspect_workstream`
round-trip per child to read what came back. idle/error states
return the last assistant turn (capped at 6 KiB UTF-8 bytes,
truncated from the end); closed/denied return a sentinel; running
children carry null. Storage reads for idle/error parallelize
across an 8-worker thread pool so a 32-child fan-out lands in 4
batches instead of 32 sequential round-trips.
2026-04-27 13:56:16 -07:00
Patrick Buckley 15d5ddde12 fix(ui): bootstrap pane in switchTab when none exists (#427)
Creating or opening a workstream from the dashboard left the chat
UI blank until the operator refreshed: switchTab early-returned at
``if (!pane) return;`` because getFocusedPane was null on a fresh-
loaded page that had no workstreams. The freshly-created ws was
added to the workstreams dict and the dashboard was hidden, but no
pane was bootstrapped, no SSE connected, and the chat area sat
empty until refresh — at which point initWorkstreams saw the
populated list and bootstrapped the pane via the existing
"if (!Object.keys(panes).length)" branch.

switchTab now mirrors that bootstrap when no focused pane exists:
createPane + splitRoot leaf + setFocusedPane + renderLayout. The
rest of switchTab (disconnectSSE / reset / connectSSE) runs as
before — no-ops on the just-constructed pane up to the connectSSE
call which is exactly what we want.

Subsequent creations on the same node already worked because the
first create populated panes and switchTab found a focused one.

Static smoke test in tests/test_app_js.py guards against the
early-return regressing.
2026-04-27 13:03:56 -07:00
Patrick Buckley 438e6f41ba feat(renderer): progressive mermaid rendering during streaming (#426)
* feat(renderer): progressive mermaid rendering during streaming

Mermaid diagrams used to materialize all-at-once at stream_end via
streamingRenderFinalize, which felt laggy on long responses with
multiple diagrams. Now closed mermaid fences render progressively
as each fence completes during streaming.

The blocker was streamingRender's wholesale `el.innerHTML = html`
on every rAF tick, which destroys any rendered SVG nodes — without
caching, calling postRenderMermaid per tick would re-trigger an
async mermaid.render every time, thrashing the renderer.

Added a source-keyed SVG cache (_mermaidSvgCache, FIFO-bounded at
64 entries):

  - Cache hit on identical source: synchronous innerHTML swap, no
    loading flash, no async work. Mermaid is deterministic for a
    given init, so identical source ⇒ identical SVG, safe to reuse.
  - Cache miss: queue async render, populate cache on success.
  - Errored sources cached separately (_mermaidErrorCache) so a
    syntactically-broken diagram doesn't re-thrash mermaid on every
    tick. The user can fix the diagram and the new source string
    misses the cache, triggering a fresh render.

_streamingRenderApply now calls postRenderMermaid after the
innerHTML replace. Per-stream cost: each unique mermaid source
pays mermaid.render once, then synchronous cache hits for every
subsequent rAF tick. hljs syntax highlighting stays deferred to
streamingRenderFinalize (it's a separate pass and benefits less
from progressive rendering — code blocks tend to be short and
already legible without color).

Tests: built a richer Node-driven harness with a fake DOM that
tracks attributes / classList / parent chain / replaceWith, plus
a stubbed mermaid.render with a call counter. 5 new tests cover:
cache-hit skips render, distinct sources render independently,
errors cache to avoid thrash, FIFO eviction at cap, and a static
guard that _streamingRenderApply actually calls postRenderMermaid.

* fix(renderer): apply Copilot feedback on PR #426

Six review items, all real:

1. _cacheMermaidEntry evicted on overwrite — overwriting an
   existing source unnecessarily dropped the oldest entry.
   Now: only evict when inserting a new key.

2. _initMermaid didn't clear caches — a theme change via
   reRenderAllMermaid (which calls _initMermaid) would serve
   stale SVG keyed by source-only, since rendered output
   depends on themeVariables. Now clears both caches on
   (re-)init.

3. bindFunctions never re-applied on cache hits — mermaid's
   bindFunctions attaches link/click handlers to each rendered
   SVG instance. Pre-fix, only the first render got bindings;
   subsequent cache hits via raw innerHTML left the SVG inert.
   Cache value is now {svg, bindFunctions}; cache hits go
   through _applyMermaidSvg which re-applies bindings on each
   new container instance.

4. Truthiness checks on cache lookups — empty-string SVG / error
   would have masqueraded as a miss. Switched to cache.has()
   (and then .get) so intent is explicit.

5. Concurrent mermaid.render — postRenderMermaid now fires on
   every streaming rAF tick, so multiple ticks could overlap
   while earlier render Promises pend. mermaid.render uses
   module-level state internally — concurrent calls clobber it.
   Two layers of serialization fix this:
   - _mermaidPending: per-source. While a render is in flight
     for source X, additional containers asking for X are queued
     and the single render result fans out to all pending
     containers when it lands.
   - _mermaidRenderChain: across-source. Promises chain so
     mermaid.render runs at most one at a time globally.
   - Detached containers (no longer in the DOM by the time the
     render completes) are skipped via isConnected guard —
     wholesale innerHTML replace during streaming detaches them
     and a later tick is already taking care of the live one.

6. Test brittleness — _streamingRenderApply guard used
   body.index("\\n}\\n", start) which would stop at the first
   inner-block closing brace inside the function. Switched to a
   bounded-window string search (Copilot's suggestion).

Three new tests added: overwrite doesn't evict; _initMermaid
clears caches; cache hit re-applies bindFunctions. Existing tests
updated for the new {svg, bindFunctions} cache shape and the
async serialization (drain via setTimeout hops instead of bare
microtask resolves).

Test harness fix: fake DOM elements now have an isConnected
getter derived from the parent chain, so the new guard
exercises correctly under test.
2026-04-27 12:47:21 -07:00
Patrick Buckley 33d16d19ce fix(renderer): handle LaTeX-style \(...\) and \[...\] math delimiters (#425)
* fix(renderer): handle LaTeX-style \(...\) and \[...\] math delimiters

The browser renderer at turnstone/shared_static/renderer.js only
recognized TeX-style $...$ / $$...$$ delimiters. Most modern LLMs
(GPT-5 / o-series, Claude with reasoning effort) emit LaTeX-style
\(...\) for inline math and \[...\] for display by default — those
slipped through as raw text in the coord + interactive WebUIs,
making KaTeX appear "broken when nested inside a markdown block"
(actually broken everywhere, the surrounding markdown just made
the failure noticeable).

Added a second pass for each delimiter style alongside the
existing $...$ / $$...$$ patterns. Both styles now feed the same
mathBlocks / inlineMaths placeholder pipeline so all the existing
nested-block handling (lists, blockquotes, tables, bold, headings,
details, post-render KaTeX markup) Just Works.

Edge cases verified by the new test_renderer_js.py harness:
- \(...\) inside inline code stays literal
- \(...\) inside fenced code blocks stays literal
- Solo \[ with no closing \] doesn't trigger spurious math
- Markdown links [text](url) untouched (regex uses \[ \], not [ ])
- Mixed TeX + LaTeX delimiters in one message both render

The harness drives renderer.js through Node via vm.runInThisContext
with stubbed document/katex globals — first JS-side regression
guard for the renderer; previously it had no test coverage at all.

* fix(renderer): apply Copilot feedback on PR #425

Three review items from Copilot:

1. Display-math sentinel could leak through inline-code spans.
   The original ordering ran $$...$$ / \[...\] extraction BEFORE
   inline code, so a backtick span around math (e.g. `$$x$$` or
   `\[x\]`) had its delimiters consumed by the math regex and
   replaced with \x00MB…\x00. Inline code then captured the
   sentinel; restore order put MB after IC, leaving the null-byte
   placeholder visible inside the rendered <code>. Reorder: inline
   code first, then display math, then inline math. Code spans
   now seal their content before any math regex sees it. The
   reverse edge case (math containing backticks, e.g. \verb|`x`|)
   is much rarer and KaTeX rejects \verb anyway.

2. Inline LaTeX-style \(...\) regex used [\s\S]+? which allowed
   newlines, so an unterminated \( on one line would eat the
   next paragraph until it found a closing \). Aligned with the
   existing $...$ behavior by switching to [^\n]+? — display
   math (\[...\] / $$...$$) stays multi-line by design.

3. tests/test_renderer_js.py was guarded with a node-availability
   skip, but CI's test + test-postgres jobs didn't explicitly
   install Node, so the suite would have silently no-op'd if the
   runner image dropped Node. Added actions/setup-node@v5 to
   both jobs.

Four new regression tests cover the leak (both delimiter styles
inside backticks must stay literal) and the cross-paragraph span
(both \(...\) and $...$ must not eat newlines).
2026-04-27 12:19:19 -07:00
Patrick Buckley 1f271789b3 fix(approve): global judge poll + Copilot round-2 feedback
Bug: LLM judge verdicts stayed stuck on heuristic-only render.
Root cause: per-row poller called scheduleLiveFetch which
short-circuits on non-visible rows — invalidate cleared the
cache, no fetch fired, the row kept rendering its last-cached
heuristic indefinitely. The 12s attempt cap also gave up before
slow LLM judges (>15s with reasoning effort) could land.

Replaced with a single global poller _maybeStartJudgePoll /
_judgePollTick:
  - Walks the full childrenState (not just visible rows)
  - Bypasses scheduleLiveFetch's visibility + TTL gates by
    adding to pendingLiveIds directly + flushing
  - One bulk request covers every pending row per tick
  - Self-terminates when every verdict lands or 90s elapses
    (operator can hit Refresh to retry on a failed judge)
  - 90s cap is wall-clock, not attempt count, so an LLM that
    takes 60s no longer prematurely gives up

Copilot round-2 feedback:

- _proxy_sse with use_service_auth=True silently fell back to
  empty headers when proxy_token_mgr was None, producing a
  retry-storm 401/403 loop. Fail fast with a 503 + clear log
  so the misconfig surfaces immediately.

- Mobile <700px CSS comment claimed buttons "stretch to full
  row width" but the rule keeps flex-direction: row with
  flex: 1 on each, giving 50/50 side-by-side. Updated the
  comment to match the deliberate side-by-side layout
  (stacking would push the action row below preview/disclosure
  on tall envelopes; 50/50 keeps both verbs reachable).
2026-04-27 11:41:14 -07:00
Patrick Buckley 3b92c96b31 fix(coord): align coord client routes with post-#422 path-keyed mounts
#422's legacy URL adapter removal deleted the body-keyed
/v1/api/route/{verb} endpoints (with ws_id in JSON body) but
turnstone/console/coordinator_client.py still pointed at them.
The coord LLM's close_workstream / close_all_children tools
404'd; send / approve / cancel were equally broken though
exercised less often.

_ROUTE_PATHS now uses {ws_id}-templated path-keyed forms:
  send → /v1/api/route/workstreams/{ws_id}/send
  approve → /v1/api/route/workstreams/{ws_id}/approve
  cancel → /v1/api/route/workstreams/{ws_id}/cancel
  close → /v1/api/route/workstreams/{ws_id}/close

_post() interpolates {ws_id} at call time when the template has
the slot; body-keyed paths (delete, close_all_children) still
work via the same code path. Each affected caller (send, approve,
cancel, close_workstream, close_all_children) was updated to pass
ws_id as the kwarg and drop ws_id from the body.

Added test_route_paths_match_actual_console_mounts: walks the
real Starlette app's routes and asserts every _ROUTE_PATHS entry
corresponds to an actually-mounted route. Catches the next URL
unification drift before runtime. Updated the existing literal
assertions + path-checking tests for the new shape.

Pre-existing bug surfaced while testing inline-child-approvals.
2026-04-27 11:41:14 -07:00
Patrick Buckley 93875ebca5 fix(console): proxy events/global with service auth (not user JWT)
The interactive WebUI's app.js opens an EventSource against
/v1/api/events/global on load (cluster-wide tab indicators,
ws_state for the dashboard). When loaded via the console proxy
at /node/{node_id}/, the JS shim rewrites that to
/node/node-X/v1/api/events/global and the proxy forwards using
the user's re-minted JWT.

Upstream global_events_sse requires `service` scope by design
— the stream carries cross-tenant cluster inventory, intended
for the cluster collector, not browsers. End-user JWTs don't
carry service scope, so every proxied call returned 403, the
browser auto-retried with exponential backoff, and the console
log filled with proxy.sse.non_200 warnings.

_proxy_sse gains a use_service_auth flag. proxy_api flips it on
for events/global only, swapping the user JWT for the console's
proxy_token_mgr bearer token. Per-ws events stay on user auth
(tenant filtering on the upstream still requires user identity).

The upstream-side privacy posture is unchanged — the data on
events/global is the same cluster-wide inventory the console's
own /v1/api/cluster/events endpoint already serves to any
read-scoped caller under the trusted-team posture. The console's
AuthMiddleware on /node/{node_id}/v1/api/ remains the gate that
decides who can use the proxy at all.
2026-04-27 11:41:14 -07:00
Patrick Buckley b0f78ae4c0 fix(console): route per-workstream events to SSE proxy
The console's node-API passthrough at /node/{node_id}/v1/api/{path}
detected SSE only on the bare events / events/global paths. After
#422 removed the legacy /v1/api/events?ws_id= shape and moved
per-workstream SSE under /v1/api/workstreams/{ws_id}/events, the
proxy never got updated to match the new path — per-ws events
fell through to the regular GET branch, the upstream returned a
text/event-stream payload that the regular GET response couldn't
hold open, and Firefox surfaced the failure as "can't establish a
connection to the server".

Extend the SSE detection to also match
``workstreams/{ws_id}/events``. Pre-existing bug surfaced while
testing inline-child-approvals (operator clicks through from the
coord tree to the per-child interactive WebUI) but affects every
caller hitting a node's per-ws events stream via the console
proxy.

Two new tests in TestConsoleProxy: per-ws events route to
_proxy_sse with the correct upstream path; existing
events/global routing still works.
2026-04-27 11:41:14 -07:00
Patrick Buckley ebf562de93 fix(approve): suppress 409 storm from rapid approve/deny clicks
Previously the 409 stale-call_id branch in submitChildApproval
re-enabled both buttons synchronously before kicking off the
urgent live-bulk refresh. That opened a window where rapid clicks
on an already-resolved approval (or a row whose call_id had
rolled) each re-armed the click handler, fired another POST, and
collected another 409. Operators rage-clicking saw a network 409
storm and a stack of warning toasts.

Keep the buttons disabled in the 409 path. The row is about to be
re-rendered wholesale via the urgent refresh — the disabled DOM
gets dropped along with it. If the row's approval truly resolved,
the new render has no buttons. If a new round started, the new
render has fresh enabled buttons. Either way the operator-facing
signal IS the row updating, not the toast.

Drop the toast.warn (noisy on every rapid-click race) in favour
of a single console.warn for diagnostics.

If the urgent refresh fails entirely, the buttons stay disabled
on that row — but the operator can hit the Refresh button on the
children panel to force a full reload. Acceptable degraded state
vs the previous 409 loop.
2026-04-27 11:41:14 -07:00
Patrick Buckley 4d08a19bd5 fix(approve): replay cached LLM verdicts on coord SSE reconnect
The coord's _coord_events_replay re-yielded _pending_approval on
connect but not the cached _llm_verdicts entries. A tab refreshing
mid-approval saw the approve_request prompt without the judge chip
because intent_verdict is a one-shot SSE event with no
late-subscriber push — the chip would only ever land if the operator
re-invoked the tool call.

Mirrored the interactive path at turnstone/server.py:875-878:
after re-injecting the pending_approval prompt, walk
ui._llm_verdicts under _ws_lock and yield each cached verdict as
an intent_verdict event. Pre-existing bug surfaced during the
inline-child-approvals work but the coord-self dock UX was always
affected on reconnect — not introduced by this PR.

Two new tests: cached verdicts replay after pending_approval; stale
verdicts from a prior round don't replay when no approval is pending.
2026-04-27 11:41:14 -07:00
Patrick Buckley 68e1332c59 fix(approve): route child approvals through proxy + poll for late judge verdict
Two bugs reported from local repro on PR #424:

1. Approve/Deny buttons return HTTP 404 on every click. The new
   approveWorkstream helper hit /v1/api/workstreams/{ws_id}/approve
   regardless of target — that path is only mounted for coord
   workstreams (which live on the console process). Child
   workstreams live on cluster nodes and need to round-trip
   through the routing proxy at
   /v1/api/route/workstreams/{ws_id}/approve, which resolves the
   ws_id to its owning node and forwards the body verbatim.
   approveWorkstream now picks the path based on whether targetWsId
   matches the coord's own wsId.

2. LLM judge verdict never populates — rows freeze on the
   heuristic-tier pill ("⚙ heuristic") even after the judge would
   have completed. The judge runs async on the child node via a
   daemon thread and updates _llm_verdicts there, but no signal
   propagates back to the coord — cluster_state events don't fire
   on verdict-only changes, and the live-bulk TTL is 5s with no
   periodic poll.

   Added _maybePollForJudgeVerdict: when renderChildRow encounters
   a pending_approval_detail with judge_pending=true and items
   missing judge_verdict, schedule a recursive 2s urgent
   live-bulk re-fetch. Self-terminates when the verdict lands,
   the row closes, the approval clears, or attempts hit the cap
   (≈12s for a failed/timed-out judge so we don't poll forever).
   Single timer per ws_id; re-renders are no-ops while a timer is
   in flight.

Smoke-test assertions added for both fixes so a regression on
either path surfaces at test-time.
2026-04-27 11:41:14 -07:00
Patrick Buckley a23ef7306c fix(approve): apply Copilot feedback + remove plan doc
Copilot review on PR #424 flagged three items:

1. Schema drift on /v1/api/dashboard — DashboardWorkstream didn't
   declare the new pending_approval_detail field, so generated
   OpenAPI / typed clients were out of sync. Added
   PendingApprovalItem + PendingApprovalDetail Pydantic models
   and referenced PendingApprovalDetail from DashboardWorkstream.

2. deepcopy under _ws_lock in serialize_pending_approval_detail
   could extend lock hold under contention with on_intent_verdict
   (daemon judge thread) and per-token activity writes that also
   take _ws_lock. _llm_verdicts entries are only assigned/cleared,
   never mutated in place, so a snapped reference is stable after
   the lock drops. Snapshot refs under lock; deepcopy after release.

3. Plan doc removed from the branch — design docs are local-only
   working artifacts, same posture as PROGRESS.md.
2026-04-27 11:41:14 -07:00
Patrick Buckley 7e33fc68bb fix(approve): apply /review feedback on inline child approvals
Critical:
- coordinator.js RISK_SEVERITY accepted 'crit' only; production
  emits 'critical' (per turnstone/core/judge.py:1556 + heuristic
  seeds). A risk_level=='critical' verdict ranked as 0 and
  rendered with .risk.low (green) styling, never triggering
  the crit-risk auto-expand. Now accepts both aliases. Unknown
  risk_level falls back to rank 2 ('high') so future schema
  drift fails *safe* (over-alert) instead of silently
  downgrading. Pill ternary handles both 'crit' and 'critical'
  alias to the existing .risk.crit class.

Major:
- Urgent live-badge flush now coalesces N urgent calls in the
  same JS tick into one bulk request via queueMicrotask, instead
  of firing N single-id fetches. The motivating 10-children-
  pending-bash scenario in the design doc now lands on one bulk
  /v1/api/cluster/ws/live request.
- Test coverage gap: added test_session_ui_base.py cases for
  POLICY-BLOCKED (item.error + needs_approval=False) and
  judge-unavailable (no verdict + no judge_pending) matrix rows.
  Added literal-string assertions to the smoke list in
  test_coordinator_page.py so a refactor dropping either branch
  surfaces at test-time.

Minor batch (4 coord.js + 1 CSS + 1 fake-divergence):
- 409 stale-call_id path re-enables both buttons before return
  (urgent fetch is best-effort; could also fail).
- judgePending pill no longer conflicts with a present heuristic
  verdict — guard changed from !judge to !verdict.
- Empty <div class="approval-reasoning"> no longer appended when
  reasoning is absent but evidence is present (evidence still
  renders inside the disclosure).
- Dead .ch-row .approval-pill.rec-* CSS rules removed (JS never
  combines those classes). Recommendation chip in the disclosure
  footer now has its own scoped rules so the chip is actually
  styled.
- _FakeUI.serialize_pending_approval_detail call_id selection
  aligned to the real impl's "first non-empty" semantics.
- liveBadgeCache reconnect cleanup now preserves permanent
  (403/404) entries — denied users no longer pay one wasted
  bulk fetch per denied id per reconnect.

All 4465 non-live tests pass. Ruff + mypy clean. node --check OK.
2026-04-27 11:41:14 -07:00
Patrick Buckley a369d5f0d0 feat(approve): clear live cache on SSE reconnect — chunk 4 reconnect parity
Closes the stale-button window where a sub-5s SSE gap would leave
liveBadgeCache holding pending_approval_detail for a child whose
approval was actually resolved during the gap. Without this clear,
zombie approve/deny buttons render until either the next child_ws_state
event or the natural TTL expiry (whichever comes first).

The clear sits beside the existing activeWaits.clear() in the
reconnect handler — same posture (drop client-only state that the
server's SSE replay doesn't cover) and same blast radius. The 409
race guard in submitChildApproval would catch a stale-call_id POST
even without this, but rendering wrong UI until the operator clicks
is the worse failure mode.

loadChildren's finally block already fires scheduleLiveFetch for
every visible row after the replace-mode refresh, so the cache
repopulates with authoritative pending_approval_detail in one bulk
request within the next debounce window.

Plan: docs/design/inline-child-approvals.md (chunk 4 of 4 — last
required chunk; 5/6 are stretch).
2026-04-27 11:41:14 -07:00
Patrick Buckley 54f04496c3 feat(approve): inline approve/deny buttons + judge verdict pill on coord tree
Chunk 3 of the inline-child-approvals plan + the SSE pipeline plumbing
needed for sub-second urgent fetches.

JS (coordinator.js):
- approveWorkstream(targetWsId, body) — generic POST helper, callable
  for both the coord-self dock and the new per-child inline buttons.
- renderApprovalBlock(child, detail) — risk-level pill (.risk.* per
  the design system primitives), tool-name summary with "+ N more"
  for envelope-level approvals, intent_summary, ↳ judge reasoning
  teaser, ▸ more disclosure carrying the recommendation chip,
  evidence list, and items 2..N stacked sub-blocks. Plus matrix
  coverage: judge_pending / judge unavailable / tool-policy
  blocked / multi-item.
- submitChildApproval — handles the 409 stale call_id race by
  invalidating the live cache + urgent-refetching, optimistically
  clears pending_approval_detail on success.
- scheduleLiveFetch({ urgent: true }) — bypasses the 5s TTL +
  cancels the debounce so attention transitions surface inline UI
  immediately instead of after the next polling window.
- handleChildState fires urgent on activity_state="approval"
  enter/leave; handleChildClosed eagerly invalidates the live
  cache so closed rows can't render stale buttons.

CSS (index.html):
- New .approval-block / pill / preview / actions / disclosure
  styles. Inline .act buttons duplicate the dock's colour treatment
  (the dock-scoped rules don't reach the children-tree). Mobile
  <700px touch targets ≥44px.

Pipeline (collector.py + coordinator_adapter.py):
- All three cluster_state event emitters and the child_ws_state
  re-emit now carry activity_state. The previous omission left
  the urgent-fetch trigger as dead code — discovered in review.

Tests:
- Static smoke test in test_coordinator_page.py asserting the new
  helper names exist + the pending_approval_detail key is read.

Plan: docs/design/inline-child-approvals.md (chunk 3 of 4).
2026-04-27 11:41:14 -07:00
Patrick Buckley 7d2d7db9d2 feat(approve): pass pending_approval_detail through cluster live-bulk
Threads the field added by Chunk 1 through the console's live-bulk
endpoint so coord tree UI can read it without a separate per-child
fetch. Three touchpoints:

- _CLUSTER_WS_LIVE_KEYS gains the new key so _fetch_live_block's
  projection forwards it from the upstream /dashboard response on
  node-backed child rows.
- _coordinator_live_snapshot synthesizes the same shape from
  ConsoleCoordinatorUI._pending_approval for in-process coord
  rows (no upstream /dashboard exists on the console pseudo-node).
- One source of truth: SessionUIBase.serialize_pending_approval_detail.

Both branches now emit the same 12-key live block; coord judge isn't
wired today so coord-self judge_verdict is always None — flagged in
the plan as a stretch follow-up.

Plan: docs/design/inline-child-approvals.md (chunk 2 of 4).
2026-04-27 11:41:14 -07:00
Patrick Buckley fbb9be27f9 feat(approve): expose pending_approval_detail on /dashboard + guard stale call_id
Lays the server-side groundwork for inline approve/deny buttons + judge
verdict on the coordinator children-tree UI. Two surgical changes:

1. SessionUIBase.serialize_pending_approval_detail() merges the active
   _pending_approval items[] with per-call_id verdicts from
   _llm_verdicts. The dashboard handler embeds this on every per-ws
   row so cluster live-bulk callers can render inline UI without an
   extra per-child round-trip.

2. make_approve_handler now returns 409 when the body sends a call_id
   that doesn't match any currently-pending item. Closes the stale
   call_id race where an operator clicks approve on a row showing
   call A while the child has rolled over to call B. Empty/missing
   call_id preserves backwards compatibility with CLI + channel
   adapters that don't track it.

Cross-tenant exposure on /dashboard is consistent with the trusted-team
posture already in place for activity / tokens — documented in the new
method's docstring so the choice survives the next reviewer.

Plan: docs/design/inline-child-approvals.md (chunk 1 of 4).
2026-04-27 11:41:14 -07:00
renovate[bot] 5ebee015d2 chore(deps): lock file maintenance 2026-04-27 07:50:27 -07:00
Patrick Buckley b8e51fa9ed fix(api): add DequeueRequest schema for DELETE /workstreams/{ws_id}/send
Copilot review on PR #422 flagged that the DELETE-on-send (dequeue)
EndpointSpec declared no request_model, so the generated OpenAPI
showed no requestBody for an operation that *requires* a JSON body
with ``msg_id`` and 400s when it's missing.

- Add ``DequeueRequest`` to ``server_schemas.py`` with the single
  required ``msg_id: str`` field.
- Wire ``request_model=DequeueRequest`` and ``response_model=
  StatusResponse`` on the DELETE EndpointSpec; trim the now-redundant
  inline body example from the description.
- Re-import the schema in ``server_spec.py`` and add the entry to
  ``_ALL_MODELS`` so the OpenAPI components list carries it.
- Regenerate ``openapi-server.json``.

Sibling thread on the close EndpointSpec was already addressed in
4000ae2 (request_model=CloseWorkstreamRequest).

4558 tests passing under ``-m "not live"``; ruff + mypy clean.
2026-04-26 22:14:22 -07:00
Patrick Buckley 5874159ffd fix(close): require non-empty body, restore CloseWorkstreamRequest
Copilot caught three real issues in PR #422 review, all clustered
around the close request body contract:

1. The interactive close handler runs with
   ``supports_close_reason=True``, which calls
   ``read_json_or_400(request)`` — an empty / non-JSON body returns
   ``400 {"error": "Invalid JSON body"}``. The previous SDK fix
   sent NO body via ``json_body=None``, which would 400 against a
   real server. The mock-transport test silently masked it because
   the mock answered without inspecting the body.
2. The doc said the body was empty (or ``{}``), with no mention
   of the optional ``reason`` field, its 512-byte cap, or the
   credential-redaction guard.
3. The Pydantic schema for close was deleted outright; OpenAPI
   and SDKs lost their typed shape for the optional ``reason``.

Changes:

- ``turnstone/api/server_schemas.py``: reintroduce
  ``CloseWorkstreamRequest`` with a single optional
  ``reason: str | None = None`` field. Docstring documents the
  must-be-valid-JSON contract and notes that coord ignores the body
  (``supports_close_reason=False``).
- ``turnstone/api/server_spec.py``: re-import the schema, point the
  close ``EndpointSpec`` at it via ``request_model=``, restore the
  ``_ALL_MODELS`` entry. OpenAPI JSON regenerated.
- ``turnstone/sdk/server.py``: ``close_workstream`` (sync + async)
  gains an optional ``reason: str | None = None`` parameter and
  always sends ``json_body={}`` (or ``{"reason": ...}``) so the
  body is never empty. Adds a regression test
  (``test_close_workstream_sends_valid_json_body``) that inspects the
  raw transport content rather than relying on a path-keyed mock —
  the kind of check that would have caught this bug pre-merge.
- ``sdk/typescript/src/server.ts``: ``closeWorkstream`` gains an
  optional ``opts.reason`` parameter; reintroduce
  ``CloseWorkstreamRequest`` interface in ``types.ts`` and re-export
  from ``index.ts``.
- ``docs/api-reference.md``: close section documents the JSON-body
  requirement, the ``reason`` field, the 512-byte cap, the
  multibyte-safe behavior, the credential-redaction guard, and the
  non-string-coercion path.
- ``CHANGELOG.md``: amend the 1.5.0 BREAKING block to reflect the
  schema reintroduction (slim form, ``reason`` optional) instead of
  the prior "removed outright" claim.

4558 tests passing under ``-m "not live"`` (was 4557 — +1 from the
regression test). ruff + mypy clean.
2026-04-26 22:14:22 -07:00
Patrick Buckley d6e615d324 fix: apply /review feedback on legacy URL cleanup
Reviewer caught real misses on the consumer-swap claim:

- TypeScript SDK still defined and re-exported `CloseWorkstreamRequest`
  (types.ts + index.ts) — drop both. Now matches the Python-side
  removal.
- Four `tests/test_auth.py` cases (`test_write_full_token_ok`,
  `test_approve_full_token_ok`, `test_bearer_takes_precedence_over_cookie`,
  `test_cookie_full_on_write_ok`) were tautological after the legacy
  URL removal: they posted to `/api/send` / `/api/approve` and asserted
  `allowed is True`, but those paths now classify as `read` so a read
  token would also pass — they no longer tested the write/approve
  scope enforcement. Swap to path-keyed URLs to restore the original
  intent.
- `is_public_path("/api/send")` test renamed + retargeted to a
  path-keyed URL.

Doc-table drift the previous commit missed:

- `docs/security.md` path-to-scope mapping rewritten for the
  path-keyed verb family (write set, DELETE-on-/send dequeue,
  per-ws_id approve).
- `docs/architecture.md` scope-model row text swap from `/api/send`
  / `/api/approve` to the path-keyed equivalents.
- `docs/diagrams/01-system-context.puml` channel→server edge label
  swap.
- `docs/diagrams/15-auth-architecture.puml` scope class swap.

Cosmetic comment-only stragglers:

- `tests/test_session_worker.py` module docstring URL update.
- `tests/test_ratelimit.py` ~11 `/api/send` fixture-key strings
  retargeted to `/api/workstreams/abc/send` so the URL fixtures
  reflect the post-1.5 surface (rate limiter is path-agnostic; the
  swap is purely cosmetic).

4557 tests still passing under -m "not live"; ruff + mypy clean.
2026-04-26 22:14:22 -07:00
Patrick Buckley ad0e7ce6eb docs: mark 1.5.0 legacy URL surface removal
CHANGELOG [Unreleased] / Removed (BREAKING — 1.5.0) block calling out
the legacy URL family removal with the swap table. Doc passes on
api-reference.md (per-endpoint sections rewritten with path
parameters and slimmer body shapes), architecture.md (handler-list
diagram and console-proxy URL example), console.md (URL-rewriting
JS shim docstring + SSE proxy example), and the two PlantUML
diagrams (11-console-data-flow, 16-channel-architecture).

Also picks up two test-side stragglers from step 5 that referenced
the legacy adapters in a docstring + a stale /v1/api/events SSE
test: turn into path-keyed equivalents. OpenAPI JSON dump regenerated
to reflect the catalog edits from step 3.

After this commit:
- 4557 tests passing under -m "not live"
- ruff + mypy clean on turnstone/ tests/ sdk/
- grep for "/v1/api/send", "/v1/api/approve", "/v1/api/cancel",
  "/v1/api/workstreams/close" returns zero hits across turnstone/
  sdk/ docs/ tests/ (excluding CHANGELOG.md, which intentionally
  documents the old shape).
- grep for make_legacy_body_keyed_adapter, make_legacy_query_keyed_adapter,
  _make_method_dispatch, close_legacy returns zero hits.
2026-04-26 22:14:22 -07:00
Patrick Buckley 1358121d52 chore(tests): refresh fixtures for path-keyed URL family
Mechanical updates across the test suite to swap legacy
/v1/api/{send,approve,cancel,events,workstreams/close} URLs for the
path-keyed equivalents under /v1/api/workstreams/{ws_id}/<verb>, and
to drop ws_id from request bodies (the path provides it now).

Per file:

- test_session_routes.py: deletes test_close_legacy_mounts_when_handler_provided
  (the close_legacy slot is gone); test_send_mounts_post_and_delete_when_dequeue_provided
  (added in PR commit 1) stays.
- test_openapi.py: expected-paths set swaps to path-keyed shape;
  test_send_endpoint_has_request_body now asserts the OpenAPI for
  /v1/api/workstreams/{ws_id}/send.
- test_auth.py / test_auth_identity.py: required_scope and
  check_request fixtures swap to path-keyed shape; new tests cover
  write/approve/read scope assignment for the path-keyed verbs +
  the /node/* proxy mirror.
- test_sdk_server.py / test_sdk_console.py: mock-transport URL keys
  swap; bodies drop ws_id.
- test_server_attachments_endpoints.py: ~17 send sites migrated to
  /v1/api/workstreams/<ws>/send (a small Python script ran the bulk
  rewrite — body ws_id stripped, URL rebuilt).
- test_server_authz.py: cross-tenant approve/close/cancel/events
  tests retargeted to path-keyed URLs;
  test_events_legacy_query_keyed_url_still_resolves_to_404_for_unknown_ws
  renamed to test_events_path_keyed_url_resolves_to_404_for_unknown_ws
  with the docstring updated to note the legacy adapter is gone.
- test_close_reason_persistence.py: 7 close sites all swap.
- test_console_routing_proxy.py: route-proxy tests swap to
  /v1/api/route/workstreams/{ws_id}/<verb>; the upstream-URL
  assertion now reads from .request (route_proxy uses
  client.request(method, url, ...) for method passthrough); _wire_proxy
  helper installs both .post and .request mocks for compatibility.
- test_route_proxy_audit.py: parametrized URLs migrated;
  _make_proxy now also exposes a .request side-effect that delegates
  to .post for the same compatibility surface.
- test_api_versioning.py: openapi.json path assertion swaps to the
  path-keyed shape.

4557 passing under -m "not live"; ruff + mypy clean.
2026-04-26 22:14:22 -07:00
Patrick Buckley 3ea6fb30b4 refactor(consumers): swap UI/SDK/console-proxy/channels to path-keyed URLs
All in-tree consumers of the legacy /v1/api/send | /approve | /cancel |
events?ws_id= | /workstreams/close URLs now hit the path-keyed shape
under /v1/api/workstreams/{ws_id}/<verb>. Bodies drop ws_id (the path
provides it). The SSE event stream URL likewise moves to the path-keyed
form; channel adapters drop the params={"ws_id": ...} kwarg on
aconnect_sse.

Touched:

- turnstone/ui/static/app.js: 7 call sites (send×3, dequeue, approve,
  cancel, close + EventSource SSE URL).
- turnstone/sdk/server.py (Python SDK): close_workstream, send,
  approve, cancel, stream_events, send_and_wait's internal SSE
  consumer.
- sdk/typescript/src/server.ts: closeWorkstream, send, approve,
  cancel, streamEvents + sendAndWait's internal SSE consumer.
- turnstone/sdk/console.py: route_send, route_approve, route_close,
  route_cancel — proxy URLs swap to /v1/api/route/workstreams/{ws_id}/<verb>.
  route_plan_feedback / route_command remain body-keyed (out of scope).
- turnstone/console/server.py:
  - Proxy mount table swaps the four legacy /api/route/{send,approve,
    cancel,workstreams/close} mounts for path-keyed equivalents under
    /api/route/workstreams/{ws_id}/<verb>; /send accepts both POST
    and DELETE for dequeue.
  - route_proxy reads ws_id from path_params (with body-fallback for
    the surviving plan/command body-keyed mounts), uses
    client.request(request.method, ...) so DELETE on /send proxies
    through correctly, and audits DELETE-on-/send as a separate
    "route.workstream.dequeue" action via _ROUTE_PROXY_AUDIT_ACTIONS.
  - Internal `method` variable renamed to `verb` to avoid confusion
    with HTTP method now that the two diverge.
- turnstone/channels/_sse.py: SSE URL builder swaps to path-keyed.
- turnstone/channels/{discord,slack}/bot.py: docstring URL updates.
- turnstone/server.py, turnstone/core/session_worker.py,
  turnstone/sdk/events.py, turnstone/api/server_spec.py: comment /
  docstring URL updates only.

Test fixtures still reference legacy URLs and will be swapped in step
5 of this PR.
2026-04-26 22:14:22 -07:00
Patrick Buckley 41e83f98d6 refactor(auth,api): drop legacy paths from scope tables, slim verb schemas
- WRITE_PATHS / APPROVE_PATHS in turnstone/core/auth.py drop the four
  legacy literal entries (/api/send, /api/cancel, /api/workstreams/close,
  /api/approve). The path-keyed verb match for write expands from
  {delete, open, refresh-title, title, attachments} to also include
  {send, cancel, close}; a sibling branch maps POST /workstreams/{ws_id}/approve
  to the approve scope, and a DELETE branch maps DELETE
  /workstreams/{ws_id}/send (dequeue) to write. The /node/* proxy
  block mirrors all four expansions so the console routing proxy
  stays in lockstep.
- server_schemas.py drops the body-keyed ws_id field from SendRequest,
  ApproveRequest, CancelRequest. CloseWorkstreamRequest deleted in
  full (its only field was ws_id, now provided by the path).
- server_spec.py: drops CloseWorkstreamRequest from imports and
  _ALL_MODELS, swaps the five legacy EndpointSpec entries to their
  path-keyed equivalents (POST/DELETE workstreams/{ws_id}/send, POST
  /approve, POST /cancel, POST /close, GET /events). Catalogue retains
  /api/plan and /api/command unchanged (out of scope).

Tests still reference the legacy URLs and will fail at this commit;
test fixture updates land in step 5 of this PR. Step 4 swaps the
UI / SDK / console proxy / channels callers next.
2026-04-26 22:14:22 -07:00
Patrick Buckley da12c6b268 refactor(routes): drop legacy body-keyed and query-keyed URL adapters
Removes the pre-1.5 interactive URL family that mounted body- and
query-keyed shapes on top of the lifted path-keyed handlers via
make_legacy_body_keyed_adapter / make_legacy_query_keyed_adapter.
Path-keyed equivalents under /v1/api/workstreams/{ws_id}/<verb>
already serve every consumer; coord never used the legacy URLs.

Removed:

- make_legacy_body_keyed_adapter / make_legacy_query_keyed_adapter
  from turnstone/core/session_routes.py.
- _make_method_dispatch from turnstone/server.py (zero callers
  after legacy /api/send POST+DELETE block goes — its only purpose
  was to bridge that single dual-method legacy URL).
- 5 legacy Route mounts in turnstone/server.py:
  /api/events?ws_id, /api/send POST+DELETE, /api/approve, /api/cancel,
  /api/workstreams/close.
- close_legacy field on SharedSessionVerbHandlers and its mount in
  register_session_routes — the only surviving body-keyed slot in
  the registrar, no longer needed.

Tightened make_dequeue_handler to read ws_id from the path only;
the body-fallback existed solely for the legacy DELETE /api/send
path and is now dead.

Test-suite updates and consumer call-site swaps (UI / SDK /
console proxy / channels) follow in subsequent commits in the same
PR — main stays broken across this commit until step 4 lands.
External SDK consumers on stable 1.0/1.3/1.4 calling these URLs
will receive 404s on upgrade to 1.5.0; CHANGELOG breaking-change
call-out lands with the docs commit.
2026-04-26 22:14:22 -07:00
Patrick Buckley 2b435263e3 refactor(routes): wire DELETE on path-keyed workstreams/{ws_id}/send
Pre-flight for the legacy URL adapter removal: the path-keyed
`/v1/api/workstreams/{ws_id}/send` route only mounted POST today;
the dequeue handler was reachable only via the legacy
`DELETE /v1/api/send` body-keyed URL through `_make_method_dispatch`.

Add a new `dequeue: Handler | None = None` slot on
`SharedSessionVerbHandlers` next to `send`, mounted as a second
`Route` on the same path with `methods=["DELETE"]` (two distinct
Routes rather than collapsing methods on one Route — different
handler callables, and collapsing would force the same
method-dispatch wrapper this cleanup is tearing out).

Wire `dequeue=dequeue_handler` in `turnstone/server.py`'s
`SharedSessionVerbHandlers(...)` call so DELETE on the path-keyed
shape works in the same merge as the legacy mount removal.

Adds a regression-locking test covering both the POST+DELETE and
the dequeue-alone cases.
2026-04-26 22:14:22 -07:00
Patrick Buckley fef266dbd9 docs: apply Copilot review feedback on PR #421
Switch fenced-code language tag from `json` to `http` on the seven
example blocks that mix an HTTP request line with a JSON body
(/trust, /restrict, /stop_cascade, /close_all_children, /approve,
/cancel, /close). Pure JSON response blocks stay tagged `json`.

Pre-existing pattern in the doc that Copilot flagged on the lines
this PR touched; fixed across all instances for consistency. No
content / URL changes — only fence-tag adjustment for correct
syntax highlighting.
2026-04-26 20:00:45 -07:00
Patrick Buckley 059bbc3729 docs: update coord URL tree to post-Stage-2 unified /v1/api/workstreams
The Stage 2 verb-shape lift converged coord and interactive on the
unified /v1/api/workstreams/{ws_id}/<verb> URL tree; the
/v1/api/coordinator/* tree was removed in P0. Two docs still
documented the pre-lift surface:

- coordinator-api-tour.md (the integrator's lifecycle walk-through):
  rewrites all 9 step URLs to the post-lift paths, keeps a one-block
  callout noting the historical /v1/api/coordinator/* tree and why
  it converged, and drops the operation-id column (operation ids
  shifted with the URL move and are now best looked up live via
  /openapi.json + Swagger UI rather than baked into prose).
- bulk-endpoints.md (the cascade-mutation shape contract): two table
  rows for stop_cascade / close_all_children fixed.

No code changes. CHANGELOG entry kept implicit since this is doc-only
and the URL convergence itself was already documented under the P0
verb-lift CHANGELOG block.
2026-04-26 20:00:45 -07:00
Patrick Buckley 6572437c5d refactor(server): rename dashboard row id → ws_id for v1 row-shape consistency
The /v1/api/dashboard endpoint was the last workstream-listing surface
keyed on `id` rather than `ws_id`. The Stage 2 list-verb lift converged
the active list (`/v1/api/workstreams`) and saved list
(`/v1/api/workstreams/saved`) on `ws_id` but explicitly left dashboard
alone to keep that PR's diff focused. This lands the same rename on
the remaining endpoint so v1 row shape is consistent across the family.

Scope kept narrow:

- Pydantic `DashboardWorkstream` and TS SDK `DashboardWorkstream`
  interface both rename `id: str/string` → `ws_id`.
- The bundled web UI (`turnstone/ui/static/app.js`) is the only consumer
  reading `dashboard.workstreams[].id` and is updated atomically.
- Console `_fetch_live_block` (cluster-inspect's projection over a
  remote node's dashboard payload at `turnstone/console/server.py`)
  flips its `entry.get("id")` lookup to `entry.get("ws_id")`.
- Drive-by: stale `id` example in `docs/api-reference.md` for the
  earlier `/v1/api/workstreams` rename also fixed.

`_build_node_snapshot` (the global-events SSE node_snapshot payload
consumed by the cluster collector) deliberately stays on `id` — it's
part of a separate cluster-row family (collector → cluster_workstreams
→ console UI) that is internally consistent on `id` and would need its
own coordinated sweep. CHANGELOG documents the bounded blast radius.

Tests: 4554 passing (-m "not live"). ruff + mypy clean.
2026-04-26 20:00:45 -07:00
Patrick Buckley 3abd2c441b feat(console): coord rich ws_state payload + live activity broadcast (#420)
* feat(console): coord rich ws_state payload + live activity broadcast (Stage 2 follow-up)

Pre-lift coord's cluster broadcast was state-only — the dashboard's
coord rows showed the state column flipping but ``tokens`` /
``context_ratio`` / ``activity`` / ``content`` were all hardcoded
to zero / empty. The lift makes coord populate the same per-ws
metric fields interactive does and broadcasts them through the
cluster collector with the rich kwargs.

**Architecture changes:**

- Lift ``on_status`` / ``on_content_token`` / ``on_thinking_start`` /
  ``on_thinking_stop`` / ``on_stream_end`` / ``on_tool_result`` /
  ``on_reasoning_token`` / ``on_tool_output_chunk`` / ``on_info`` /
  ``on_error`` from ``WebUI`` to :class:`SessionUIBase` as base
  implementations. Coord inherits the bodies; the per-ws metric
  fields it had at the base but never populated now flow.
- ``WebUI`` keeps overrides for ``on_status`` / ``on_tool_result`` /
  ``on_error`` to layer Prometheus ``_metrics.record_*`` calls
  on top of ``super()`` (node-only — the console isn't a node).
  ``WebUI._broadcast_state`` now uses the new
  :meth:`SessionUIBase.snapshot_and_consume_state_payload` helper
  for the rich-payload snapshot read.
- ``ConsoleCoordinatorUI`` adds a ``_broadcast_activity`` override
  that calls the new
  :meth:`ClusterCollector.update_console_ws_activity` (in-memory
  pseudo-node row update; named ``update_*`` rather than ``emit_*``
  to flag the no-fanout asymmetry vs. the rest of the
  ``emit_console_ws_*`` family).
- ``coord_adapter.emit_state`` reads ``ws.ui``'s snapshot under
  ``_ws_lock`` and passes the rich kwargs to the extended
  :meth:`ClusterCollector.emit_console_ws_state`. Defensive when
  ``ws.ui is None`` mid-eviction (broadcasts state-only).
- ``coord_endpoint_config`` wires a new ``_coord_spawn_metrics``
  hook so per-spawn ``_ws_messages`` / ``_ws_turn_tool_calls``
  bookkeeping fires on coord too.
- ``_MAX_TURN_CONTENT_CHARS`` moved from ``turnstone.server`` to
  ``turnstone.core.session_ui_base`` so coord enforces the same
  per-turn content cap.

**Three observable behaviour changes** (CHANGELOG-callout-worthy):

- Coord persists ``usage_event`` storage rows on every status
  emission (governance dashboards / token-spend queries gain
  coord visibility).
- Coord broadcasts live activity transitions to the cluster
  collector (dashboard's coord rows show activity ticks between
  state changes the same way interactive does), with last-emitted
  dedup so a tool-heavy turn's repeated ``activity=""`` clears
  don't hammer the collector lock.
- Cluster ``cluster_state`` events for coord rows now carry
  non-zero ``tokens`` / ``content``. Frontend rendering that
  conditionally hid these on coord can drop the branch.

**Tests:** 23 new tests in ``tests/test_coord_rich_ws_state_payload.py``
(per-ws metric writes, snapshot helper drain semantics +
single-lock-acquisition, adapter rich-payload pass-through +
None-UI defensive handling, activity broadcast wire + dedup +
failure swallow + no-op-when-collector-unset, spawn_metrics
hook, concurrent-writes-during-snapshot stress with reader
cycling through running/idle/error so drain branches actually
run, on_stream_end activity-clear pin). Plus WebUI override
regression tests confirming ``_metrics.record_*`` still fires
on top of the lifted bodies. Existing
``tests/test_webui_content.py`` updated to import
``_MAX_TURN_CONTENT_CHARS`` from its new home;
``tests/test_coordinator_adapter.py`` updated to expect the
rich-payload kwargs (default zeros) on
``emit_console_ws_state``. Total: ``4491 → 4514``.
``ruff check`` clean, ``mypy`` clean on touched files.

**/review pipeline** (4 finders → verify → dedupe) caught 14
findings → 12 unique (3 collapsed as duplicates of the lockless
``on_content_token`` writer):

- bug-1 Minor: ``on_status`` regressed coord's defensive
  ``usage.get(...)`` indexing → restored ``.get(..., 0)`` for
  ``prompt_tokens`` / ``completion_tokens`` on both base + WebUI
  override.
- bug-2 Nit: concurrent-snapshot reader only used ``"running"`` →
  cycled through ``("running", "idle", "error")`` so drain
  branches run; also captures + re-raises thread exceptions
  instead of silently passing.
- bug-3 + sec-2 + perf-3 Nit (merged): ``on_content_token``
  mutated ``_ws_turn_content`` lockless while the snapshot drained
  under lock → wrapped the cap-check + append + size-update in
  ``_ws_lock``.
- perf-2 Minor: collector lock contention from per-event activity
  broadcasts → cached last-emitted ``(activity, activity_state)``
  on the UI; subsequent identical ticks return early without
  acquiring the collector lock.
- perf-4 Nit: join-under-lock in snapshot helper → swap-then-join
  pattern (capture list reference under lock, reassign to empty,
  join the captured list outside the lock). Halves the lock
  hold and decouples the join walk from concurrent appenders.
- q-1 Minor: ``emit_console_ws_activity`` was misleading (no
  ``_fanout`` call, unlike the rest of the ``emit_console_ws_*``
  family) → renamed to ``update_console_ws_activity`` + docstring
  call-out for the asymmetry.
- q-2 + q-3 Minor/Nit: stale docstrings on
  ``coordinator_ui.py`` (still claimed "no per-node metrics —
  Phase D") and ``_interactive_spawn_metrics`` (still claimed
  "counters live on WebUI only") → both updated to reflect the
  lifted base class + coord's new hook.
- q-4 Nit: broken Sphinx cross-ref
  ``:meth:\`_snapshot_and_consume_state_payload\``` → dropped
  the leading underscore.
- q-5 Nit: missing ``test_coord_on_stream_end_clears_activity``
  → added.

**Two findings explicitly deferred** (out-of-scope follow-ups,
documented in CHANGELOG):

- perf-1: synchronous ``record_usage_event`` INSERT on coord
  worker thread per status tick. Parity with WebUI is the lift's
  goal; if throughput becomes a concern, batch usage_event writes
  on a background flusher (would apply to both kinds).
- sec-1: coord assistant content now flows on the cluster SSE
  stream, which has no per-user filter today. Pre-existing
  exposure for interactive ``cluster_state`` events; the lift
  extends to coord rows. Proper fix needs SSE auth gating
  (``admin.cluster.inspect``) or per-listener user_id filtering
  — separate security project, doesn't gate this lift.

* fix(console): apply review feedback on PR #420

Three review findings, all confirmed against source:

1. **Copilot — dedup-state-vs-failure race in `_broadcast_activity`**
   (correctness bug): pre-fix ``self._last_broadcast_activity = current``
   was assigned inside the ``_ws_lock`` block BEFORE the collector call.
   If the collector raised mid-broadcast, the exception was swallowed
   but the dedup state was already updated, so subsequent identical
   activity ticks would be deduped and never retried — leaving the
   dashboard's coord row stranded at the pre-failure activity until
   the activity actually changed.

   Fix: move the dedup-state update OUT of the lock and place it AFTER
   a successful collector call. On failure, ``_last_broadcast_activity``
   stays unchanged so the next identical tick retries. Two new
   regression tests pin both the failure-recovery (``test_coord_ui_
   broadcast_activity_failure_does_not_strand_dedup``) and the
   happy-path dedup behavior (``test_coord_ui_broadcast_activity_
   dedup_skips_identical_after_success``).

2. **Copilot — stale `emit_console_ws_activity` reference in
   CHANGELOG**: the method was renamed to ``update_console_ws_activity``
   per /review's q-1 finding before the original commit landed, but the
   CHANGELOG entry was written ahead of the rename. Updated to match
   the actual API + added the no-fanout asymmetry rationale inline so
   readers don't have to chase the method name.

3. **code-quality bot ×2 — `except BaseException` in test workers**:
   the concurrent-snapshot stress test caught thread-worker exceptions
   with ``except BaseException`` (with a noqa to suppress BLE001).
   ``BaseException`` is overkill for a thread worker — ``SystemExit``
   / ``KeyboardInterrupt`` are main-thread signals and ``Exception``
   is the right scope. Narrowed to ``except Exception`` on both
   workers; ``writer_exc`` / ``reader_exc`` types narrowed from
   ``list[BaseException]`` to ``list[Exception]``.

Tests: ``4514 → 4516`` (+2 regression tests for the dedup race fix).
``ruff check`` clean, ``mypy`` clean. No code-path changes outside
the dedup-state placement; the rich-payload broadcast surface is
unchanged.
2026-04-26 17:32:44 -07:00
Patrick Buckley acbe18d5f5 docs: apply Copilot review feedback on PR #419
Server-side history endpoint declared ``error_codes=[404]`` but the
lifted ``make_history_handler`` factory can also return:

- ``400`` on empty ``ws_id`` (defensive — Starlette routing makes
  it unreachable in practice, but the factory has the branch).
- ``500`` on the ``cfg.list_kind is None`` misconfig gate added in
  the /review fix-up (defense-in-depth fail-loud; both production
  cfgs wire ``list_kind`` so the gate doesn't fire today).
- ``503`` via ``cfg.manager_lookup`` when the kind's manager isn't
  available (interactive's lookup never returns 503; coord's can).

Updated ``server_spec.py`` to ``[400, 404, 500, 503]`` per Copilot's
suggestion — matches the existing detail entry's shape so the two
endpoints document the same possible-error envelope.

Caught the parallel asymmetry on ``console_spec.py``: history was
``[403, 404, 503]`` but the lifted factory's misconfig + empty-
ws_id branches reach coord too. Updated to
``[400, 403, 404, 500, 503]`` — same factory body, same possible
responses, plus ``403`` from coord's ``admin.coordinator``
permission gate.

Regenerated ``openapi-{server,console}.json``. No code changes;
spec metadata only. Tests + lint + mypy unchanged.
2026-04-26 15:48:26 -07:00
Patrick Buckley d555816016 refactor(core): lift history + detail verb bodies across both kinds (Stage 2 verb lift)
Last verb-shape lift before v1.5.0 stable can tag. Adds two new
factories to ``turnstone/core/session_routes.py``:

- ``make_history_handler(cfg)`` — body lifted from coord's
  ``coordinator_history`` near-verbatim. ``?limit=`` query param
  defaults to 100, clamps to [1, 500], malformed values fall back
  to 100. Storage operations (``get_workstream`` on the
  storage-fallback path, ``load_messages`` for the row read) now
  run via ``asyncio.to_thread`` (was inline pre-lift on coord).
- ``make_detail_handler(cfg)`` — body lifted from coord's
  ``coordinator_detail``. Lazy-rehydrates a closed/evicted
  workstream via ``mgr.open()`` on miss; mirrors
  :func:`make_open_handler`'s exception envelope (``ValueError``
  → 503 with the session-factory's remediation text; bare
  ``Exception`` → correlation_id'd 500 with the per-kind noun
  via ``cfg.audit_action_prefix``).

NO new ``SessionEndpointConfig`` fields — the factories reuse
``permission_gate``, ``manager_lookup``, ``not_found_label``,
``audit_action_prefix``, and (for history's storage-fallback
kind check) ``list_kind`` — all already wired by both production
lifespans for the list/saved factories.

Coord side: ``coordinator_history`` and ``coordinator_detail``
standalone handler bodies removed from ``console/server.py``;
``register_session_routes`` now wires
``history=make_history_handler(coord_endpoint_config)`` and
``detail=make_detail_handler(coord_endpoint_config)``.

Interactive side: GAINS both endpoints as a feature gain. Pre-lift
interactive had no ``GET /v1/api/workstreams/{ws_id}`` and no
``GET /v1/api/workstreams/{ws_id}/history`` — SDK consumers had to
subscribe to ``/events`` SSE just to read display fields or
message rows. The same lifted factories are wired with the
interactive endpoint config; cross-kind isolation is preserved on
both sides (history via ``cfg.list_kind`` storage-fallback gate
+ fail-loud-on-misconfig 500; detail via ``mgr.open()``'s internal
kind check).

Pydantic schemas: ``CoordinatorDetailResponse`` /
``CoordinatorHistoryResponse`` removed from ``console_schemas.py``;
``WorkstreamDetailResponse`` / ``WorkstreamHistoryResponse`` added
to ``server_schemas.py`` (mirrors the list lift's pattern for
``WorkstreamInfo``). Both server and console OpenAPI specs
reference the unified schemas; ``server_spec.py`` gains
``EndpointSpec`` entries for the new interactive endpoints. TS
SDK gains both interfaces in ``sdk/typescript/src/types.ts``;
``openapi-{server,console}.json`` regenerated.

Tests: 6 new coord regression/parity tests in
``test_coordinator_endpoints.py`` (limit clamping, cross-kind 404
on storage fallback, storage-only history, detail 503 on
session-factory misconfig, detail 500 with correlation_id on
unexpected rehydrate failure, history swallows
``load_messages`` exception → 200 with empty messages). 10 new
interactive parity tests in ``test_workstream_endpoints.py``
(``TestHistoryInteractive`` + ``TestDetailInteractive``). 1 new
openapi spec test pinning the server-side ``?limit=`` query param.
Total: ``4490 → 4491`` after the new exception-swallow
regression test landed. ``ruff check`` clean, ``mypy`` clean on
touched files.

/review pipeline (4 finders → verify → dedupe) caught 1 Minor
defense-in-depth (bug-1/sec-1, merged: ``make_history_handler``
fail-closed gate when ``cfg.list_kind is None``, mirroring
``make_saved_handler``'s same gate) + 1 Minor test-helper rename
(q-1: ``_interactive_history_cfg`` → ``_interactive_endpoint_cfg``)
+ 4 Nits (q-2 unused fixture parameter, q-3 CHANGELOG TS SDK
mention, q-4 missing exception-swallow regression test, q-5
misleading test comment) — all addressed in the same commit.
2026-04-26 15:48:26 -07:00
Patrick Buckley e8a6b0632d docs: apply Copilot review feedback on PR #418
Three docstring + CHANGELOG drift items from the post-review
M3 + Mi1 fixes:

- ``make_list_handler`` docstring referenced ``cfg.list_resolve_title``
  (singular) but the field renamed to ``list_resolve_titles``
  (bulk variant) when the N+1 fix landed. Updated to the plural
  name + a one-line note about the bulk SELECT pattern.
- ``make_saved_handler`` docstring still claimed kind was derived
  from ``cfg.audit_action_prefix`` string-compare. The Mi1 fix
  replaced that with the explicit ``cfg.list_kind`` field +
  fail-loud-on-missing semantic; docstring now describes the
  current contract.
- CHANGELOG ``[Unreleased]`` entry said "Three new
  ``SessionEndpointConfig`` fields" and listed the singular
  ``list_resolve_title`` wired to ``get_workstream_display_name``.
  Updated to "Four" + the bulk plural names + the new
  ``list_kind`` field with its rationale (distinct from
  ``audit_action_prefix``; fail-loud on misconfig).

The fourth review comment — code-quality bot flagging the ``...``
ellipsis body on the new ``get_workstream_display_names`` Protocol
method as "statement has no effect" — is a false positive.
``...`` is the canonical Protocol method body throughout
``turnstone/core/storage/_protocol.py`` (every other method uses
it). Refuting; the file's pattern wins over the bot's per-method
suggestion.

No code changes; docstring + CHANGELOG only. Tests + lint + mypy
unchanged.
2026-04-26 13:11:07 -07:00
Patrick Buckley edf52016ac refactor(core): lift list + saved verb bodies across both kinds (Stage 2 verb lift)
New ``make_list_handler(cfg)`` and ``make_saved_handler(cfg)``
factories in ``turnstone/core/session_routes.py`` replace four
pre-lift bodies (interactive ``list_workstreams`` +
``list_saved_workstreams``; coord ``coordinator_list`` +
``coordinator_saved``). Same factory + capability-flag pattern as
the merged cancel / open / events / create lifts.

Four new ``SessionEndpointConfig`` fields:

- ``list_resolve_titles: ListResolveTitles | None`` — bulk lookup
  ``(ws_ids) -> {ws_id: title-or-None}``. Interactive wires
  ``get_workstream_display_names`` (new bulk helper added on the
  storage layer + memory.py); the lifted body resolves every active
  row in ONE ``SELECT ... WHERE ws_id IN (...)`` instead of the
  pre-lift N+1 (one SELECT per row).
- ``list_kind: WorkstreamKind | None`` — explicit kind classifier
  for the saved-list storage filter. Replaces the initial draft's
  ``audit_action_prefix == "coordinator"`` string compare which
  would have silently leaked INTERACTIVE rows for any future kind
  whose audit prefix didn't match. Required when a kind mounts
  list/saved; misconfig surfaces as a 500 with a clear log line.
- ``saved_state_filter: str | None`` — coord wires ``"closed"``;
  interactive wires ``None``.
- ``saved_loaded_lookup: SavedLoadedLookup | None`` — coord-only
  defence-in-depth filter that excludes ws_ids in the warm pool.

Behaviour changes (all observable in CHANGELOG):

- **Active-list row shape converges on always-include** ``{ws_id,
  name, state, kind, parent_ws_id, user_id}``. Interactive renames
  ``id`` → ``ws_id``; both kinds populate every field (coord adds
  kind + parent_ws_id; interactive adds user_id).
- **Top-level response key converges on ``"workstreams"``** on
  both endpoints. Coord ``coordinators`` key removed — coord is a
  1.5.0aN-only surface (never shipped stable) so the convergence
  has no compat shim; SDK / frontend consumers swap once.
- **Storage + manager-lock work moved off the event loop on
  interactive**. ``list_workstreams_with_history`` runs through
  ``asyncio.to_thread`` on both kinds (matches coord's pre-existing
  perf-2 pattern from the saved-coordinators review); ``mgr.list_all``
  + per-row work also offloaded.
- **N+1 storage round-trips on /v1/api/workstreams eliminated**.
  Pre-lift interactive resolved the alias for every active row in a
  separate SELECT (up to 50 round-trips per dashboard refresh on a
  saturated node). Lifted body issues one bulk SELECT.

Pydantic schemas: ``WorkstreamInfo.id`` renamed → ``ws_id``,
``WorkstreamInfo.user_id`` field added. ``CoordinatorInfo`` and
``CoordinatorListResponse`` removed (folded into the unified
``WorkstreamInfo`` / ``ListWorkstreamsResponse``). OpenAPI spec
snapshots regenerated. TS SDK types updated (``WorkstreamInfo``
interface gains ws_id + the always-include fields); TS test
mock + assertion updated to match.

``GET /v1/api/dashboard`` is intentionally NOT in this PR's scope
and still returns rows keyed on ``id``. Tracked as a separate
cleanup PR (tombstone-note added at the dashboard handler).

/review pipeline run; the four Major findings + one Minor + six
nits all addressed in the same commit:

- M1: TS SDK ``WorkstreamInfo`` interface stale (id: string) →
  renamed + fields added.
- M2: TS SDK test masked the type-mismatch with stale mock → updated.
- M3: N+1 alias resolution on active list → bulk
  ``get_workstream_display_names`` helper + ``list_resolve_titles``
  bulk cfg hook.
- M4: Missing interactive parity regression test for unified row
  shape → mirror of coord's added in test_server_authz.py.
- Mi1: ``audit_action_prefix`` string-compare deriving kind →
  explicit ``cfg.list_kind: WorkstreamKind`` field.
- Six nits: redundant inner asyncio import, forward-ref quotes on
  Awaitable, duplicated frontend comments, dashboard ``id`` field
  has no tombstone-note, empty-coord_mgr short-circuit on
  ``saved_loaded_lookup``.

4512 tests passing; ruff + mypy clean.
2026-04-26 13:11:07 -07:00
Patrick Buckley c77b237033 refactor(core): defer emit_created on SessionManager.create + commit_create / discard pair (#417)
* refactor(core): defer emit_created on SessionManager.create + commit_create / discard pair

Eliminates the phantom create→close pair on coord rollback that was
documented as a known limitation in PR #416. The pair surfaced on the
cluster events stream when a multipart workstream-create request
failed attachment validation: coord's ``mgr.create`` fired
``emit_created`` synchronously, then the rollback called
``mgr.close`` which fired ``emit_closed``. Cluster consumers had to
reconcile via the collector's diff path. Post-fix, a rejected upload
produces zero events.

API changes on ``SessionManager``:

- ``create(..., defer_emit_created: bool = False)`` — when True,
  skip the trailing ``emit_created`` so the caller can run additional
  post-create work (attachment validation in the lifted HTTP handler)
  before advertising the workstream. Default preserves the existing
  "advertise immediately" contract for direct callers (test fixtures,
  CLI REPL, channel adapters).
- ``commit_create(ws)`` — fires the deferred ``emit_created`` event
  after the caller's post-create work confirms the workstream should
  be advertised. Synchronous; the wrapped work is in-memory and
  non-blocking on every kind (interactive: documented no-op stub;
  coord: dict updates under a lock + ``queue.put_nowait`` fan-out).
- ``discard(ws_id)`` — releases the in-memory slot + cleans up the UI
  WITHOUT firing ``emit_closed``. Distinct from ``close`` which
  advertises the transition; ``discard`` is for the rollback case
  where the workstream's existence was never advertised. Storage-row
  deletion stays a separate concern (caller invokes
  ``delete_workstream``), mirroring ``mgr.create``'s split between
  slot reservation and ``register_workstream``.

Caller-bug detection: ``Workstream._emit_created_fired`` is set
inside ``create`` (non-deferred path) and ``commit_create``;
``discard`` logs ``session_mgr.discard.after_emit_created`` warning
when invoked on an already-advertised workstream. Slot is still
released so capacity isn't stranded.

Lifted ``make_create_handler`` updated to use the deferred bracket:
pass ``defer_emit_created=True``, validate uploaded attachments,
then ``mgr.commit_create(ws)`` on success / ``mgr.discard(ws.id)``
on failure. Ordering invariants (``commit_create`` BEFORE
``audit_emit`` and ``post_install`` so any state events the worker
fires reach the cluster collector for an already-known ws_id) are
documented in the handler docstring.

Tests:

- 5 new ``SessionManager`` unit tests (defer skips emit, commit
  fires it, commit no-ops without emitter, discard releases without
  emit_closed, discard returns False on unknown id).
- 2 caller-bug regression tests (commit_create after discard pins
  the silent re-emit behaviour; discard after non-deferred create
  asserts the warning fires + slot still releases).
- 1 coord regression test asserting the cluster collector sees zero
  events when attachment validation fails.

``/review`` pipeline run; M1 (test gap on caller-bug paths) +
Mi1 (no runtime guard for already-advertised) + Mi2
(``_make_manager`` event_emitter override) + Mi3 / N2 (duplicated
comments + ordering invariant) + N1 (drop ``to_thread`` on
``commit_create``) all addressed.

4509 tests passing; ruff + mypy clean.

* fix(core): apply Copilot + code-quality review feedback on PR #417

Copilot review:

- ``Workstream._emit_created_fired`` comment claimed the flag was
  "set under the manager's _lock-protected emit", but the actual
  ordering set it OUTSIDE the lock. Comment updated to describe the
  real synchronization (non-deferred ``create`` sets it immediately
  before ``emit_created``; ``commit_create`` sets it under the
  manager lock alongside the tracked-ws check).
- ``commit_create`` had no guard against duplicate calls,
  post-discard calls, or calls on workstreams not tracked by this
  manager — any of those would have fired duplicate or phantom
  ``ws_created`` events. Added a guard symmetric to ``discard``'s
  after-emit warning: under ``self._lock``, check ``_emit_created_fired``
  + ``_workstreams.get(ws.id) is ws``, no-op + log a warning
  (``session_mgr.commit_create.already_fired`` /
  ``session_mgr.commit_create.untracked``) on either failure. The
  emit itself still runs outside the lock so coord's collector
  fan-out doesn't couple to the manager mutex.
- ``test_commit_create_after_discard_is_caller_bug_no_op`` was
  internally inconsistent — name + docstring said "must not re-emit"
  but the assertion expected the re-emit. Renamed to
  ``test_commit_create_after_discard_is_no_op`` and updated to
  assert the new no-op + warning behaviour.

New test ``test_commit_create_is_idempotent_on_duplicate_call``
pins the second-commit-call code path: exactly one ``ws_created``
event fires, second call short-circuits via the guard with a
``commit_create.already_fired`` warning.

Code-quality bot review (3 findings, identical pattern):

- Three test ``assert`` statements wrapped side-effecting calls
  (``assert mgr.discard(ws_id) is True/False``); under ``python -O``
  the asserts strip and the side-effect strips with them. Refactored
  all three to assign the result to a local first, assert on the
  local. No behaviour change.

4510 tests passing; ruff + mypy clean.
2026-04-26 12:06:59 -07:00
Patrick Buckley 16916dc257 fix(core,console): coord create-time attachments coordination + Copilot review feedback on PR #416
Coord initial-message + create-time-attachments coordination:

- ``CoordinatorAdapter.send`` gains optional ``attachments`` + ``send_id``
  kwargs so the worker dispatched at create time can carry the uploaded
  files onto the first turn. Mirrors interactive's pre-existing
  worker-thread pattern. The ``send_id`` reservation token soft-locks
  the rows; the adapter's failure path unreserves so a worker crash
  returns them to pending.
- ``_coord_create_post_install`` reserves any uploaded ``attachment_ids``
  via the lifted ``reserve_and_resolve_attachments`` helper before
  dispatching through the adapter — closes the parity gap with
  interactive's create-with-attachments+initial_message flow.
- ``_reserve_and_resolve_attachments`` lifted from ``turnstone/server.py``
  to ``turnstone/core/attachments.py`` as ``reserve_and_resolve_attachments``
  so both processes use one kind-agnostic implementation.

Copilot review fixes on PR #416:

- Skill lookup now calls ``storage.get_prompt_template_by_name`` directly
  rather than going through ``turnstone.core.memory.get_skill_by_name``;
  that helper swallows storage exceptions into ``None`` which would have
  masked outages as the 400 "Skill not found" branch. Calling storage
  directly lets exceptions bubble to the lifted body's correlation_id'd
  500 path so operators chasing skill-related reports can distinguish
  real misses from registry outages.
- ``_interactive_create_build_kwargs`` /
  ``_coord_create_build_kwargs`` thread ``skill_data["name"]`` (the
  canonical row name) into ``mgr.create`` instead of the raw
  ``body["skill"]`` value. Pre-fix a whitespace-padded request body
  ``"skill": "  my-skill "`` would have persisted the dirty name even
  though the lookup ran on the stripped key.
- ``make_create_handler`` docstring corrected: audit-emit failures
  return 200 (not 201).
- ``_audit_workstream_created`` docstring corrected: factory keeps the
  successful 200 response on audit-emit failure (was 201).

New regression test:
``test_create_with_multipart_attachments_and_initial_message_reserves``
asserts attachments are reserved (not pending) when both
``initial_message`` and uploads land in the same coord create request.
Updated ``_SendSession`` stub in ``test_coordinator_adapter.py`` to
match the new ``send`` / ``queue_message`` signatures.

4501 tests passing; ruff + mypy clean.
2026-04-26 04:07:39 -07:00
Patrick Buckley 9ed8b1e0b5 refactor(core): lift create verb body across both kinds (Stage 2 verb lift)
New ``make_create_handler(cfg, *, audit_emit=None)`` factory in
``turnstone/core/session_routes.py`` consumes five new ``SessionEndpointConfig``
fields (``create_supports_attachments``, ``create_supports_user_id_override``,
``create_validate_request``, ``create_build_kwargs``, ``create_post_install``)
and replaces both ``create_workstream`` and ``coordinator_create`` bodies.
Same factory + capability-flag pattern as the merged cancel / open / events
lifts. ``_validate_and_save_uploaded_files`` lifted to
``turnstone.core.attachments`` so both processes call one kind-agnostic
implementation.

Coord parity gains (§ Post-P3 reckoning item #1 + carry-forward):
- Create-time attachments: multipart parsing, validate+save+rollback,
  ``attachment_ids`` on the response. Coord adapter ``send`` doesn't yet
  reserve attachments at create time, so the rows save as pending and the
  next ``/send`` picks them up via the standard send-with-attachments path.
- Disabled-skill rejection (matches interactive's pre-lift gate).
- Always-include response shape ``{ws_id, name, resumed, message_count,
  attachment_ids}`` populated with default ``False``/``0``/``[]`` on the
  fields coord doesn't fill.
- 200 status (was 201).
- Audit-emit failures swallow + warning log instead of 500.

Both kinds converge on the manager-at-capacity 429, factory-misconfig 503,
and correlation_id'd 500 for unexpected ``mgr.create`` failure (interactive
lifted up to coord's safer error envelope).

Three /review fixes folded in:
- ``notify_targets`` malformed input gates at the validator (400) instead
  of bubbling out of post_install as a 500 — pre-fix the workstream had
  already been created + audited + broadcast by the time the validation
  raised.
- Skill-lookup storage failures now share the correlation_id'd 500 path
  with ``mgr.create`` (was masquerading as 400 "Skill not found").
- Whitespace-only ``skill`` field treated as empty (matches pre-lift coord).

CHANGELOG entry under [Unreleased] documents every observable behaviour
change. OpenAPI spec regenerated. Three new coord regression tests
(create-time-attachments save pending rows, always-include parity fields,
disabled-skill rejection) plus one interactive regression test
(notify_targets 400). 4500 tests passing.
2026-04-26 04:07:39 -07:00
Patrick Buckley 577ad2824f refactor(core): lift events verb body across both kinds (Stage 2 verb lift) (#415)
* refactor(core): lift events verb body across both kinds (Stage 2 verb lift)

The interactive ``GET /v1/api/events?ws_id=...`` and coord
``GET /v1/api/workstreams/{ws_id}/events`` SSE handlers now share
one body via ``make_events_handler(cfg)``. Per-kind divergence
captured by two new ``SessionEndpointConfig`` fields:

* ``events_replay: EventsReplay | None`` — Protocol-typed callback
  that yields the kind-specific initial replay payload. Interactive
  wires ``_interactive_events_replay`` (connected + status + history
  + pending_approval + cached intent verdicts + pending_plan_review);
  coord wires ``_coord_events_replay`` (just pending_approval +
  pending_plan_review). The lifted body iterates the callback
  before starting the live event loop.
* ``sse_executor_lookup: SseExecutorLookup | None`` — per-kind
  executor for the live loop's blocking ``client_queue.get``.
  Interactive returns the dedicated 200-thread ``sse_executor``
  from app state so SSE polling stays isolated from every other
  ``asyncio.to_thread`` caller in the process; coord returns
  ``None`` and the lifted body falls through to the default executor.

Also adds ``make_legacy_query_keyed_adapter(handler)`` (sister to
``make_legacy_body_keyed_adapter`` from earlier lifts): reads
``ws_id`` from the query string and splices into ``request.path_params``
before delegating to the lifted body. Preserves the
``GET /v1/api/events?ws_id=...`` legacy URL shape so any 1.x SDK
consumer keeps working.

Old ``events_sse`` (server.py) + ``coordinator_events``
(console/server.py) bodies deleted.

Two convergence wins for coord:

* **SSE connect/disconnect metrics** — pre-lift coord didn't record
  per-stream metrics; the lifted body always calls
  ``metrics.record_sse_connect()`` / ``record_sse_disconnect()``,
  giving the cluster dashboard the same per-stream observability
  interactive's had since 1.0.
* **Both kinds now check ``request.is_disconnected()`` AND the
  ``ws_closed`` event** to terminate. Pre-lift interactive relied
  solely on ``ws_closed`` (which never fires if the client just
  goes away without closing the workstream); pre-lift coord relied
  solely on ``is_disconnected``. The lifted body uses both.

One observable shape change for coord callers: the lifted body
returns 409 ``"session has no UI"`` when ``ws.ui`` is missing
(placeholder / build-failed UI), matching pre-lift coord.
Pre-lift interactive returned 404 in this case; the lift converges
on 409 because the workstream EXISTS (404 would imply it doesn't).

Item #2 from § Post-P3 reckoning (rich ``ws_state`` payload parity
for coord) split out during scoping — touches different files
(``coordinator_ui.py`` + ``collector.py`` + ``session_ui_base.py``)
with different reviewer concerns. Tracked as standalone follow-up
``feat/coord-rich-ws-state-payload``.

Two /review fixes folded in:

* **Dedicated SSE thread pool restored.** Initial draft used
  ``asyncio.to_thread`` (default executor, ~32 workers). Pre-lift
  interactive deliberately used a dedicated 200-thread
  ``sse_executor`` to avoid pool starvation; the
  ``sse_executor_lookup`` cfg field above restores that isolation.
* **5s poll timeout restored.** Initial draft shortened to 1s,
  multiplying thread-wakeup rate 5x while the pool was already
  starving. ``is_disconnected()`` between polls covers cancel-
  detection latency.

Plus minor cleanups: stale ``coordinator_events`` comment
references in coordinator.js refreshed; ``TestInteractiveEventsLifted``
gets a ``_make_interactive_replay_mocks`` fixture so per-test
intent stays clear; live-loop coverage gap documented in the
test class docstring.

Lint + mypy clean. 4497 tests passing (+8 new events tests).

* fix(core): stream events replay from inside the generator instead of pre-building

PR #415 review caught that ``make_events_handler`` pre-built the
full replay payload (``connected`` + ``status`` + ``history`` +
pending prompts) into a list before constructing the
``EventSourceResponse``. Two real costs:

* **TTFB delay** — the client saw nothing until the heaviest
  replay event finished serialising (``_build_history`` on a
  long-running interactive workstream can take 10s of ms). With
  pre-build, the ``connected`` event was buried at the end of
  the materialisation pass instead of streaming first.
* **Listener-queue accumulation** — registering the per-UI
  listener BEFORE building the replay let live events queue
  during the build window. On a chatty mid-generation
  workstream that window can fill the 500-slot listener queue
  and drop events before the live loop starts draining.

Fix: iterate ``cfg.events_replay`` inside the async generator
so each event ships as soon as the callback yields it. The
observational-failure swallow semantics are preserved by
wrapping the iteration in the same try/except + log.debug as
before — partial replay is still acceptable; the live loop
continues either way.

Resolves the Copilot review thread on PR #415. Lint + mypy
clean. 4497 tests passing (no test changes — the replay
callbacks themselves are unchanged; only the lifted body's
consumption pattern flipped from eager-build to lazy-stream).
2026-04-26 01:58:41 -07:00
Patrick Buckley f9ed4d3071 refactor(core): lift open verb body across both kinds (Stage 2 verb lift) (#414)
* refactor(core): lift open verb body across both kinds (Stage 2 verb lift)

The interactive ``POST /v1/api/workstreams/{ws_id}/open`` and coord
``POST /v1/api/workstreams/{ws_id}/open`` handlers now share one
body via ``make_open_handler(cfg, *, audit_emit=None)``. Per-kind
divergence captured by two new ``SessionEndpointConfig`` fields:

* ``open_resolve_alias: AliasResolver | None`` — interactive wires
  ``resolve_workstream`` so callers can pass user-friendly aliases
  in the path param. Coord wires ``None``.
* ``open_post_load: OpenPostLoad | None`` — interactive wires
  ``_interactive_open_post_load`` (display-name sync + UI replay
  via ``clear_ui`` + history + handler-side ``ws_created`` enqueue
  onto the global SSE queue). Coord wires ``None`` and relies on
  the cluster collector fan-out from
  ``CoordinatorAdapter.emit_rehydrated``.

Plus an optional ``audit_emit`` parameter (interactive wires
``_audit_workstream_opened``; coord wires ``None`` — coord doesn't
audit open today). Old ``open_workstream`` (server.py) +
``coordinator_open`` (console/server.py) bodies deleted.

**Load-bearing fix** (§ Post-P3 reckoning item #3 from the planning
docs): pre-lift interactive's ``open_workstream`` called
``mgr.create(ws_id=resolved_id)`` + ``ws.session.resume(...)`` to
rehydrate, bypassing ``mgr.open()`` entirely. After the lift both
kinds route through ``mgr.open()`` — which makes
``InteractiveAdapter.emit_rehydrated`` reachable on interactive
(it had been dead-by-routing) and gives the manager a single
rehydrate code path to maintain. ``emit_rehydrated`` stays a
documented no-op stub on the interactive adapter; the handler-side
``ws_created`` enqueue from the post-load callback is the
load-bearing emission for the SSE consumers.

Behaviour changes for interactive callers (documented in CHANGELOG):

* **Cross-kind open returns 404** (was 400 with
  ``"Workstream is not an interactive kind"``). The lift consolidates
  on ``mgr.open()``'s single ``None``-return contract for missing /
  wrong-kind / tombstoned rows. Security boundary unchanged.
* **Already-loaded response uses ``ws.name`` directly** (was
  ``get_workstream_display_name(resolved_id) or resolved_id``).
  The dashboard listing endpoint still resolves aliases on its own
  pass, so the user-visible name in the tab strip isn't affected.

Two /review fixes folded in:

* **Resume failures now return 5xx instead of broken-200.**
  ``SessionManager.open()`` previously caught and ``log.debug``-
  swallowed exceptions from ``ChatSession.resume``. Since
  ``ChatSession.resume`` assigns ``self.messages`` *before* the
  config-restore block, a partial-failure resume (corrupted
  ``workstream_config`` row, model-registry mismatch on a saved
  alias, malformed ``temperature`` / ``max_tokens``) would leave
  the session with history but with default config. Pre-lift the
  interactive open handler called ``ws.session.resume`` directly
  and let exceptions propagate as 500. Restored that behaviour:
  ``mgr.open()`` now re-raises resume exceptions after rolling
  back the slot (``cleanup_ui`` + ``_remove_locked``), so the
  lifted handler returns 500 with a correlation id and the storage
  row stays available for a retry.
* **Bare ``except Exception`` documents intent.** A one-line
  rationale in the handler body explains why the catch is broad
  (no documented exception spec on ``adapter.build_session``;
  resume can propagate via the new contract above). Keeps a future
  contributor from narrowing it incorrectly.

Test scaffolding:

* ``tests/test_workstream_endpoints.py`` — fixture rebuilt to
  use ``make_open_handler`` + a minimal cfg with a lazy alias
  resolver so per-test ``@patch`` calls take effect. Added 5 new
  tests: already-loaded uses ws.name, alias resolution runs first,
  ``mgr.open`` is called (NOT ``mgr.create``), post-load callback
  fires with (request, ws) only on the load-from-storage path
  (not the already-loaded shortcut), post-load exception swallowed
  → 200.
* ``tests/test_coordinator_endpoints.py`` — fixture imports
  updated to ``make_open_handler``.
* ``tests/test_server_authz.py`` — ``TestOpenKindGate`` now expects
  404 (not pre-lift's 400) for cross-kind open attempts. Docstring
  explains the consolidation.

Two nit cleanups: dropped the unnecessary ``import secrets as
_secrets`` aliasing in the exception handler; refreshed the stale
``open_workstream`` reference in the ``AliasResolver`` doc-comment.

Lint + mypy clean. 4488 tests passing (was 4475; +13 new open
tests).

* fix(core): use cfg.audit_action_prefix for the per-kind noun in open's 500 error

PR #414 review caught the hardcoded ``"failed to open workstream"``
in ``make_open_handler``'s 500 path: coord callers got misleading
text (pre-lift coord said ``"failed to open coordinator"``).

The fix derives the noun from ``cfg.audit_action_prefix``
("workstream" interactive, "coordinator" coord) — a field both
production lifespans already construct, and which the previous
/review pipeline (q-5) flagged as dead config (set but read by
no factory). Reusing it here both fixes the wording AND gives
the field its first runtime reader.

Pinned by a new test
(``test_open_500_message_uses_kind_noun_from_cfg``) that wires a
coord-shaped cfg, forces ``mgr.open`` to raise, and asserts the
500 body contains ``"failed to open coordinator"`` + the
correlation id, without echoing the exception text.

Lint + mypy clean. 4489 tests passing (+1 new).
2026-04-26 00:44:14 -07:00
Patrick Buckley 412c99f486 refactor(core): lift cancel verb body across both kinds (Stage 2 verb lift) (#413)
* refactor(core): lift cancel verb body across both kinds (Stage 2 verb lift)

The interactive ``/v1/api/cancel`` (body-keyed ws_id) and coord
``/v1/api/workstreams/{ws_id}/cancel`` (path-keyed) handlers now
share one body via ``make_cancel_handler(cfg, *, audit_emit=None)``
in ``turnstone.core.session_routes``. Per-kind divergence captured
by a new ``cancel_forensics: CancelForensics | None`` field on
``SessionEndpointConfig`` (interactive wires
``_capture_cancel_forensics``; coord wires ``None``) plus an
optional ``audit_emit`` (coord wires ``_audit_cancel_coordinator``;
interactive wires ``None`` — pre-lift interactive didn't audit
cancel).

Same factory + capability-flag pattern as P1.5's ``make_send_handler``
+ make_attachment_handlers. Old ``cancel_generation`` body deleted
from ``server.py``; old ``coordinator_cancel`` body deleted from
``console/server.py``.

Behavior changes (documented in CHANGELOG):

* **Coord gains the ``force`` flag.** Pre-lift coord ignored
  ``force``; the lifted body honours it on both kinds. Stuck-worker
  recovery becomes available on coord (parity gain — coord workers
  hang the same way interactive's can).
* **Coord cancel response always includes ``"dropped"``.** Pre-lift
  returned bare ``{"status": "ok"}``; lifted returns
  ``{"status": "ok", "dropped": {}}``. Always-include parity with
  interactive so SDK consumers don't branch on kind.
* **Coord cancel returns 400 ``"No session"``** on placeholder /
  build-failed workstreams (was a silent 200 no-op pre-lift). Parity
  with interactive's existing 400 branch.
* **Coord ``coordinator.cancel`` audit detail now includes
  ``force``** so operator-driven recovery is distinguishable from
  routine cancels.

Three /review fixes folded in:

* **bug-1**: lifted body's ``resolve_approval`` is now gated on
  ``ui._pending_approval is not None``. Pre-fix, the unconditional
  call leaked a stale ``approval_resolved`` SSE event on every
  idle cancel — listener UIs that key on the event would dismiss
  prompts they didn't have. ``resolve_plan`` keeps its existing
  internal no-pending guard so the unconditional call is still
  safe there.
* **bug-2**: force-cancel now clears ``_worker_running`` alongside
  ``worker_thread`` inside the same ``with ws._lock`` block. Prior
  half-state ``(_worker_running=True, worker_thread=None)`` routed
  follow-up sends through the queue-enqueue path onto the abandoned
  worker (whose cancel flag short-circuits the queue-drain seam,
  leaving messages orphaned until next spawn). Restores the
  ``(worker_thread, _worker_running)`` invariant
  ``session_worker.send`` documents.
* **bug-3**: ``coordinator_stop_cascade._fanout_on_children`` now
  treats child cancel ``400 + "No session"`` as ``skipped`` (was
  ``failed``). Lifted coord cancel returns 400 on placeholder
  children; matches the pre-lift outcome where those children were
  silently no-op'd, so the cascade response's ``failed`` bucket
  stops firing spurious operator alerts.

Test scaffolding:

* ``tests/test_coordinator_endpoints.py`` — replace ``coordinator_cancel``
  fixture with ``make_cancel_handler(...)`` wiring; add 6 new
  tests covering always-include shape, force-flag worker-abandon,
  400-on-null-session, cancel_forensics swallowed-exception,
  audit_emit swallowed-exception, no-stale-approval-resolved-on-idle.
* ``tests/test_server_authz.py`` — new ``TestInteractiveCancelLifted``
  class with HTTP-level coverage of ``/v1/api/cancel`` for the
  dropped shape, force-flag + ``_worker_running`` clearing, and
  400-on-null-session. Pre-lift ``cancel_generation`` had no
  HTTP-level test; this is the first.

One observable change for interactive (pre-existing call site):
``resolve_approval`` / ``resolve_plan`` now run on every cancel
regardless of ``was_running`` (was gated). Lifts coord's
unconditional behaviour onto interactive — a stuck approval-pending
state from a crashed worker can now be cleared via cancel without
requiring close + rehydrate.

Lint + mypy clean. 4484 tests passing (was 4475; +9 new cancel
tests minus the moved one that became part of the new suite).

* docs(core,changelog): correct cancel-lift behaviour description for resolve_approval

Two review comments on PR #413 caught the same drift between the
implementation and its documentation: my bug-1 fix gated
``resolve_approval`` on ``_pending_approval is not None`` (because
it broadcasts ``approval_resolved`` unconditionally), but the
``make_cancel_handler`` docstring and the CHANGELOG entry still
claimed both ``resolve_approval`` and ``resolve_plan`` "run on
every cancel" and "the calls are idempotent and no-op when
nothing is blocked".

Reality:

* ``resolve_plan`` does run on every cancel and its no-op-when-
  nothing-pending behaviour is real (the method has an internal
  ``_pending_plan_review is None`` short-circuit).
* ``resolve_approval`` runs only when ``ui._pending_approval is
  not None``. Without the gate, every idle cancel would broadcast
  a stale ``approval_resolved`` SSE event and overwrite
  ``_approval_result``.

Updated:

* ``make_cancel_handler`` docstring (turnstone/core/session_routes.py
  in the "Behavior changes vs the pre-lift handlers" section) —
  splits the two methods into separate bullets, explains why
  ``resolve_approval`` is gated and ``resolve_plan`` isn't.
* CHANGELOG.md ``[Stage 2 Verb Lift — cancel]`` entry — same
  split + rationale; the asymmetric coord pre-lift parity is
  still flagged as the recovery path that drove the lift.

Docs-only change; lint + mypy clean; cancel test suite (59 tests)
unchanged.

* style(core): replace CancelForensics ellipsis stub with docstring

github-code-quality bot flagged the ``...`` body of
``CancelForensics.__call__`` as "Statement has no effect". The
ellipsis is the canonical Protocol method-body idiom (no real
issue), but switching to a one-line docstring satisfies the bot
AND adds a small piece of method-level documentation. The class-
level rationale (why Protocol-typed instead of a plain Callable
alias) moves from a wall of leading ``#`` comments into a proper
class docstring at the same time.

Style-only change; the Protocol semantics are identical.
2026-04-25 23:43:19 -07:00
Patrick Buckley 48c9ad2a40 refactor(core): split SessionKindAdapter Protocol into construction +… (#412)
* refactor(core): split SessionKindAdapter Protocol into construction + emission (Stage 2 P3)

The single ``SessionKindAdapter`` Protocol that ``SessionManager``
takes is split into two:

* ``SessionKindAdapter`` — kind / build_ui / build_session /
  cleanup_ui. Required for every kind. The shared lifecycle
  manager always delegates here for construction + cleanup.
* ``SessionEventEmitter`` — emit_created / emit_state /
  emit_rehydrated / emit_closed. **Optional**, wired through a new
  ``event_emitter: SessionEventEmitter | None = None`` kwarg on
  ``SessionManager``. Reserved for future kinds whose lifecycle
  transitions don't fan out anywhere; both production kinds wire
  one today.

Both production adapters implement both Protocols. The interactive
lifespan (``server.py``) and console lifespan
(``console/server.py``) pass their adapter as both ``adapter`` and
``event_emitter`` — production behaviour is unchanged. Six lifecycle
sites in ``SessionManager`` (create / open eviction / open rehydrate /
close / set_state / close_idle / _reserve_and_install_locked unwind)
now call ``self._event_emitter.emit_*(...)`` guarded by
``if self._event_emitter is not None``.

InteractiveAdapter asymmetry preserved + documented:

* ``emit_closed`` stays load-bearing — it's the **sole** transport
  path for ``ws_closed`` onto the process-wide global SSE queue
  (Stage 1 consolidated emission from the create handler here so
  there's exactly one emission point; ``name`` powers the
  frontend's eviction toast).
* ``emit_created`` / ``emit_state`` / ``emit_rehydrated`` are
  documented no-op stubs (``del ws[, state]``). Those events fire
  from out-of-band paths — the create HTTP handler enqueues
  ``ws_created`` directly onto ``global_queue`` *after* attachment
  validation (so a rejected upload doesn't surface a phantom
  create→close pair); ``WebUI._broadcast_state`` emits the full
  ``ws_state`` payload (tokens + context_ratio + activity) via the
  ``SessionUI.on_state_change`` callback chain. The stubs exist
  solely to satisfy ``SessionEventEmitter`` Protocol so the
  adapter can be wired as the manager's ``event_emitter`` for the
  ``emit_closed`` path. Each stub has a 1-line inline rationale to
  match the in-repo convention (``coordinator_adapter.py:210``).

Test scaffolding:

* ``tests/test_session_manager.py`` — ``_make_manager`` and
  ``_make_with_writer`` wire ``FakeAdapter`` as both ``adapter``
  and ``event_emitter`` for production parity; the standalone
  ``test_create_uses_configured_node_id`` does the same.
  ``FakeAdapter.emit_rehydrated`` now records as
  ``_Event("rehydrated", ...)`` rather than conflating with
  ``"created"``, and ``test_open_resurrects_closed_state`` asserts
  against ``events_of("rehydrated")`` so a regression where the
  manager fires the wrong call on the open path actually fails.
* ``tests/_coord_test_helpers.py`` and
  ``tests/test_coordinator_end_to_end.py`` — wire
  ``CoordinatorAdapter`` as both args.
* Six interactive test fixtures (``test_skills.py``,
  ``test_prompt_templates_runtime.py`` x2, ``test_model_registry.py``,
  ``test_server_authz.py``, ``test_server_attachments_on_create.py``)
  — wire ``event_emitter=adapter`` so they match the production
  wiring, removing the footgun where a future contributor adds a
  ``gq.get_nowait()`` assertion and silently loses the only
  ``ws_closed`` transport.
* ``tests/test_interactive_adapter.py`` — drops the three
  tautological no-op-emit_* tests (``test_emit_created_is_noop``,
  ``test_emit_state_is_noop``, ``test_emit_rehydrated_is_noop``);
  keeps the four ``emit_closed`` tests (real behaviour).

Lint + mypy clean. 4475 tests passing.

* docs(core): correct SessionKindAdapter + SessionEventEmitter docstrings to match implementation

Two Copilot review threads on PR #412 caught the same real
discrepancy: my P3 docstrings on ``SessionKindAdapter`` and
``SessionEventEmitter`` described an *intent* — "interactive
doesn't implement ``SessionEventEmitter``; the manager skips emit
calls when no emitter is wired" — that doesn't match the actual
wiring. ``InteractiveAdapter`` does implement both Protocols and
``server.py`` does pass it as ``event_emitter``; only the three
no-op stubs (``emit_created`` / ``emit_state`` / ``emit_rehydrated``)
are dead, while ``emit_closed`` is load-bearing.

Updated both docstrings to:

* State that both production adapters implement both Protocols.
* Explain the asymmetry is in *which* emit methods carry real
  bodies (coord: 4; interactive: 1, with 3 documented stubs because
  the out-of-band paths — create handler ``ws_created`` after
  attachment validation, ``WebUI._broadcast_state`` carrying the
  richer ``ws_state`` payload — fire those events).
* Clarify the ``if self._event_emitter is not None`` guard exists
  for the kwarg-omitted case (tests that don't care about events,
  reserved for future kinds whose transitions don't fan out
  anywhere).

Docstring-only change. Lint + mypy clean; the 75 tests in
test_session_manager + test_interactive_adapter + test_coordinator_adapter
pass.

Resolves the two Copilot review threads on PR #412 (commits
PRRC_kwDORcMomM67VyPD, PRRC_kwDORcMomM67VyPI).
2026-04-25 22:52:39 -07:00
Patrick Buckley 02e4a01207 fix(core,server): apply Copilot review feedback on PR #411
Five fixes from Copilot's review of Stage 2 P1.5 — all preserve
behaviour, narrow docstring claims, and round out the response shape:

* **session_routes.py:supports_attachments docstring** — claimed
  the handler "accepts only ``{"message": ...}``" when ``False``,
  but the implementation silently ignores ``attachment_ids``
  rather than rejecting. Updated wording to say the
  attachment-resolution block short-circuits and any
  ``attachment_ids`` are silently ignored. Behaviour unchanged
  (silent-ignore is the right choice for forward compat — clients
  passing ``attachment_ids`` speculatively to a not-yet-lit-up
  kind shouldn't get a 400).

* **session_routes.py:queue_full response shape** — restored the
  always-include guarantee for ``attached_ids`` /
  ``dropped_attachment_ids``. The queue_full path now returns
  ``attached_ids: []`` and ``dropped_attachment_ids: list(requested_ids)``
  so SDK consumers don't have to branch on status.

* **server.py:_interactive_spawn_metrics guard** — added
  ``_ws_turn_tool_calls`` to the ``hasattr`` chain. Previously
  the guard checked ``_ws_lock`` + ``_ws_messages`` and then
  unconditionally assigned ``_ws_turn_tool_calls`` — would
  raise on a SessionUI subclass with the first two but not the
  third.

* **console_spec.py:coord_send error_codes** — added 409
  (the 'session UI not available' branch in
  ``make_send_handler`` returns 409, but the spec didn't list
  it). OpenAPI spec regenerated; TS SDK types refreshed.

* **session_routes.py:tenant_check docstring** — claimed
  interactive uses ``_require_ws_access`` with "404 on owner
  mismatch", but the helper now delegates to
  ``resolve_workstream_owner`` which explicitly does NOT enforce
  row-level ownership (trusted-team semantics; 404s only on
  missing rows). Updated wording to match.
2026-04-25 21:44:40 -07:00
Patrick Buckley ad56192a96 fix(core,console): address /review feedback on Stage 2 P1.5
Six fixes from the local /review pipeline (find-bug + find-security +
find-quality, all confirmed by verify):

* **sec-1 (major)** — coord ``attachment_owner_resolver`` now
  resolves through ``coord_mgr.get(ws_id)`` only and does NOT fall
  back to storage. Without the kind-strict check, an
  ``admin.coordinator``-scoped caller could pass an *interactive*
  workstream ws_id to the new coord attachment endpoints; the
  generic ``get_workstream_owner`` storage call (kind-agnostic)
  would resolve and grant cross-kind read / write access to
  interactive attachments. New regression test
  ``test_coord_attachment_endpoints_404_on_interactive_ws_id``
  pins the surface.

* **bug-1 (minor)** — UI hook calls in the spawn-path ``_run``
  closure are now wrapped per-hook (via ``_emit_ui``) so a failure
  in ``ui.on_error`` doesn't suppress the subsequent
  ``ui.on_stream_end`` / ``ui.on_state_change`` calls. Mirrors the
  pre-P1.5 coord_adapter.send per-hook defense.

* **bug-2 (minor)** — ``make_dequeue_handler`` now 404s when
  ``ws.ui is None`` (preserves the pre-P1.5 ``_get_ws`` contract;
  a partially-constructed or close-window workstream shouldn't
  answer DELETE).

* **bug-3 (minor)** — ``coordinator.js`` gains a
  ``case "message_queued":`` handler that surfaces the queueing
  as an info row. Coord wires ``emit_message_queued=True`` for
  parity with interactive but the dashboard had no router branch
  for these events, silently dropping them.

* **bug-4 (minor)** — error-message format on coord regressed
  from ``f"{type(exc).__name__}: {exc}"`` to ``f"Error: {e}"``
  (lost the exception class name, which coord operators rely on
  to triage failures). Restored.

* **q-1 (major)** — duplicate ``_auth_user_id`` and
  ``_require_ws_access`` helpers in ``server.py`` and
  ``console/server.py`` now delegate to the lifted
  ``turnstone.core.web_helpers.auth_user_id`` /
  ``resolve_workstream_owner``. The lifted versions are the
  canonical implementations; the shims keep existing call sites
  working without a sweeping rename.

CHANGELOG entry adds a Security section noting the kind-strict
resolver fix and a behaviour callout for the cancel-state semantic.
2026-04-25 21:44:40 -07:00
Patrick Buckley e0c78e2aec test,docs: coord attachment + queue parity tests + spec regen + CHANGELOG
Five new TestCoordinatorAttachments tests in
``tests/test_coordinator_endpoints.py`` exercising the lifted
attachment surface end-to-end on coord:

* upload → list round-trip
* get_content returns raw bytes with text/plain forced for text
* delete removes pending entries and clears them from the listing
* send with attachment_ids consumes pending under the send_id token
* send response carries attached_ids / dropped_attachment_ids even
  on plain-text sends (unified shape parity)

The existing ``_coord_endpoint_config`` fixture grew capability
flags to mirror the production console wiring, and ``_make_client``
now mounts the four coord attachment routes via
``make_attachment_handlers``.

OpenAPI specs regenerated; TS SDK bumped to 0.5.0. CHANGELOG entry
under [Unreleased] documents the verb-shape lift, the coord
attachment surface coming online, the response-shape change for
``coordinator_send``, the unification of the three lifted classifier /
lock helpers under ``turnstone.core.attachments``, and the new SDK
helpers.
2026-04-25 21:44:40 -07:00
Patrick Buckley 61fe759b6c refactor(server,console): wire both kinds to lifted send/attachments factories
Replaces per-kind ``send_message`` / ``coordinator_send`` and the
four interactive attachment handlers with calls to the shared
factories from ``turnstone.core.session_routes``. Net deletion of
~660 LOC from ``server.py`` (the lifted body lives in
``session_routes`` and is mounted twice — once interactive, once
coord).

Interactive (``turnstone/server.py``):

* ``SessionEndpointConfig`` now carries ``supports_attachments=True``,
  ``attachment_owner_resolver`` (delegates to ``_require_ws_access``
  via storage path to preserve test fixtures using MagicMock
  managers), ``attachment_helpers`` (the lifted classifiers +
  upload-lock), ``spawn_metrics`` (records the per-conversation
  WebUI counters that coord doesn't have), and
  ``emit_message_queued=True``.
* New ``_make_method_dispatch`` adapter lets the legacy body-keyed
  ``/v1/api/send`` URL serve both POST (send) and DELETE (dequeue)
  via the lifted handlers.
* The four attachment handler bodies (``upload_attachment`` etc.)
  are deleted; the shared registrar mounts them via
  ``make_attachment_handlers(cfg)``.

Coord (``turnstone/console/server.py``):

* Same wiring with coord-specific resolvers
  (``_coord_attachment_owner`` via the lifted
  ``resolve_workstream_owner``). ``spawn_metrics=None`` since the
  coord dashboard doesn't have per-conversation counters; cluster
  metrics fan out via the collector.
* Old ``coordinator_send`` body deleted.
* Console-side coord attachment endpoints come up automatically
  through the shared ``AttachmentHandlers`` slot — no per-kind
  attachment handler bodies needed at all.

Coord dashboard (``coordinator.js``): user messages with
attachments arriving on history replay now extract just the text
portion + a ``📎 N attachment(s)`` count badge instead of
JSON-stringifying the multipart content. Full chip-rendering with
click-to-view stays deferred.

Python SDK adds coord-side helpers on
``AsyncTurnstoneConsole`` + ``TurnstoneConsole``:
``coordinator_send`` (with ``attachment_ids``),
``coordinator_upload_attachment``,
``coordinator_list_attachments``,
``coordinator_get_attachment_content``,
``coordinator_delete_attachment``. URL prefix is direct
``/v1/api/workstreams/`` since coord workstreams live on the
console — no routing-proxy hop needed.

Behaviour change for coord callers:

* Worker-queue-full responses are now ``200 {"status": "queue_full"}``
  for parity with interactive (was ``429 {"error": "..."}``). SDK
  consumers checking for 429 should switch to the status field.
* Send response now always carries ``attached_ids`` /
  ``dropped_attachment_ids`` (empty arrays on plain text sends);
  the live-worker reuse path also surfaces ``priority`` /
  ``msg_id``.
2026-04-25 21:44:40 -07:00
Patrick Buckley 3398c4b6e7 feat(core): lift send + attachments to shared factories with capability flags
Stage 2 P1.5 — verb-shape unification at the HTTP layer for both
``send`` and the four attachment endpoints. New factories in
``turnstone.core.session_routes``:

* ``make_send_handler(cfg)`` — single body covering the
  attachment-resolution dance, dispatcher hand-off, queue/spawn
  outcome surfacing, and metrics increment. Capability flags on
  ``SessionEndpointConfig`` (``supports_attachments``,
  ``attachment_owner_resolver``, ``attachment_helpers``,
  ``spawn_metrics``, ``emit_message_queued``) toggle the per-kind
  bits without forking the body.
* ``make_dequeue_handler(cfg)`` — DELETE branch (cancel a queued
  message by ``msg_id``). Path-keyed; mountable on both new
  ``/v1/api/workstreams/{ws_id}/send`` and the legacy body-keyed
  ``/v1/api/send`` URL via ``make_legacy_body_keyed_adapter``.
* ``make_attachment_handlers(cfg)`` — quartet of upload / list /
  get_content / delete with shared scope checks and 404 masking.
  Per-kind classification + locking comes in via the new
  ``AttachmentUploadHelpers`` bundle so the cfg stays declarative.

Three pure helpers (``sniff_image_mime``,
``classify_text_attachment``, ``upload_lock``) moved from
``turnstone/server.py`` to ``turnstone/core/attachments.py`` so the
console process can wire them into the lifted attachment endpoints
without depending on the node-side server module. Behaviour is
unchanged.

``turnstone.core.web_helpers`` gains ``auth_user_id`` and
``resolve_workstream_owner`` so both kinds share the owner-resolution
helper underpinning attachment scoping. The interactive ``trusted-team``
404-on-missing semantics are preserved; ``not_found_label`` is
parameterised so coord can return ``coordinator not found``.

Console spec adds ``CoordinatorSendResponse`` (parity with interactive
``SendResponse``) and four new endpoint declarations for the coord
attachment surface.
2026-04-25 21:44:40 -07:00
Patrick Buckley a8cd9444b1 fix(server): apply Copilot + code-quality review feedback
PR #410 review pass:

* **session_worker**: ``except BaseException`` → ``except Exception``
  in ``_runner`` (code-quality bot). Daemon threads don't receive
  SystemExit/KeyboardInterrupt, so the wider catch was unjustified
  defensive style. Same defense-in-depth for unexpected ``run()``
  exceptions; doesn't widen scope to runtime signals.

* **session_worker**: ``threading.Thread()`` construction moved
  inside the spawn branch under ``ws._lock`` (Copilot). The
  enqueue path no longer allocates and then discards a Thread
  object on each call against a busy workstream. Thread()
  construction is microsecond-cheap, so the lock-window growth is
  negligible vs. the saved allocation churn.

* **lifespans**: ``state_writer.shutdown()`` (and the console
  equivalent) now run via ``asyncio.to_thread`` so the daemon-
  thread join + sync DB drain don't block the event loop and
  delay other teardown tasks (Copilot, ×2).

* **tests**: five remaining ``writer._flush_once()`` calls
  switched to the public ``writer.flush()`` API across
  test_session_manager.py (4) and test_state_writer.py (1)
  (Copilot, ×5). Tests no longer depend on private internals.
2026-04-25 20:11:47 -07:00
Patrick Buckley 52e09e87d6 fix(core): address /review feedback on Stage 2 P1
Six fixes from the local /review pipeline (find-bug + find-perf +
find-quality, all confirmed by verify):

* **bug-1 (critical)** — ``StateWriter.record(flush_now=True)`` now
  drops any pending buffered transient for the same ws_id AND waits
  on the flush_lock before its sync UPDATE. Without this, an
  earlier buffered 'running' could flush AFTER the sync 'error'
  write and clobber the terminal state — same shape as the
  close-vs-buffered-transient race ``discard`` was already
  guarding. New regression tests cover both the drop and the
  in-flight wait.

* **bug-3 (major)** — ``session_worker.send`` now assigns
  ``ws.worker_thread = t`` AND sets ``ws._worker_running = True``
  under the same ``ws._lock`` acquisition. Previously
  ``worker_thread`` was assigned outside the lock, so a reader
  holding ``ws._lock`` could observe ``_worker_running=True``
  paired with a stale (already-exited) ``worker_thread`` —
  defeating every ``ws.worker_thread is me`` identity check
  downstream.

* **bug-2 (major)** — rewind/retry busy gate in
  ``server.py:command`` now reads ``ws._worker_running`` instead
  of ``ws.worker_thread.is_alive()``. The is_alive() gate could
  see a stale dead thread under ws._lock while a new worker was
  in the middle of starting (post bug-3 fix the window narrows
  but the gate-mismatch was independent — ``_worker_running`` is
  the canonical gate post-Stage-2-P1).

* **perf-2 (major)** — ``StateWriter.discard`` now waits on
  ``_flush_lock`` with a 5s timeout (configurable). Without a
  bound, a stuck Postgres connection inside an in-flight flush
  would block ``close()`` and ``close_idle()`` indefinitely while
  they hold ``ws._lock`` — a system-wide hang on every close
  path. On timeout we log + proceed; the worst-case degrades to
  "buffered transient flushes shortly after sync 'closed'"
  (eventual consistency) rather than process hang.

* **q-1** — inline comments on ``run_retry`` and ``_run_initial``
  now explain why those two spawn sites don't go through
  ``session_worker.send``: retry-when-busy is a hard reject (no
  fallback queue), and init-on-create can't have a pre-existing
  worker by construction (enqueue branch is dead code). Both
  still set ``_worker_running`` + ``ws.worker_thread`` together
  under ws._lock for parity with the dispatcher.

* **q-3** — ``state_writer.discard`` callsite comments in
  ``session_manager.py`` no longer reference 'bug-3' (which lived
  only in untracked working notes). Now describe the invariant
  inline by what it prevents.

* **q-5** — ``StateWriter._flush_once`` promoted to public
  ``flush()``. Tests now drive flushes via the public API.
2026-04-25 20:11:47 -07:00
Patrick Buckley c3d24749f5 docs(changelog): note Stage 2 P1 worker dispatch + write-behind
Two new bullets under [Unreleased]:

* Worker dispatch unified — ``session_worker.send`` shared by
  interactive ``/v1/api/send``, the coord adapter, watches, retry,
  and initial-message paths. Gate is ``_worker_running`` (atomic
  under ws._lock) instead of ``Thread.is_alive()``. Closes a
  parallel-worker race that any concurrent path (watch + /send,
  retry + /send, init + /send) could trigger pre-P1.
* Buffered ``StateWriter`` for set_state — non-terminal transitions
  now show up in storage up to ~1s late (SSE consumers see them
  immediately via the adapter). Terminal ERROR + close still write
  sync; bug-3 invariant preserved via state_writer.discard before
  the sync 'closed' write.

The `/send` HTTP body convergence stays out of scope — interactive's
attachments / reservations / queue-outcome distinctions diverge from
coord's response shape too far for a clean factory split until
coord grows attachments parity (post-1.5.0).
2026-04-25 20:11:47 -07:00
Patrick Buckley 8240e32704 test(core): regression tests for state_writer + close ordering
Five new tests under ``TestSessionManagerWithStateWriter`` exercise
the bug-3 invariant under write-behind:

* set_state buffers via the writer (long flush_interval → no sync
  write until drain).
* set_state(ERROR) flushes synchronously.
* close after a buffered transient writes 'closed' as the final
  state — the buffered 'running' must NOT be flushed to storage
  AFTER close's sync 'closed' write.
* close_idle exhibits the same invariant.
* set_state arriving AFTER close short-circuits on ws._closed and
  never reaches the buffer.
2026-04-25 20:11:47 -07:00
Patrick Buckley 436ae79d19 refactor(core): wire StateWriter into SessionManager + lifespans
``SessionManager.__init__`` accepts an optional ``state_writer``;
when present, ``set_state`` for non-terminal transitions records via
the buffered writer instead of holding ``ws._lock`` across a sync DB
UPDATE. Terminal ERROR transitions still flush sync (error-surfacing
paths need durability before any observer sees the state).

``close()`` and ``close_idle()`` call ``state_writer.discard(ws_id)``
under ws._lock BEFORE their sync 'closed' write — drops any pending
buffered transient and waits on the flush_lock for any in-flight
flush to complete. Without this, a buffered 'running' could land in
storage AFTER the sync 'closed' write and resurrect the closed row
(bug-3 invariant under write-behind).

Lifespan wiring on both servers: build the StateWriter alongside the
SessionManager, ``state_writer.start()`` on enter, ``shutdown()`` on
teardown (drains any pending writes synchronously). Tests can leave
``state_writer=None`` and get the legacy direct-write behaviour.
2026-04-25 20:11:47 -07:00
Patrick Buckley 470a6af6a9 feat(core): add state_writer for buffered set_state persistence
``turnstone.core.state_writer.StateWriter`` buffers non-terminal
``update_workstream_state`` writes (last state per ws_id wins) and
flushes them on a ~1s cadence (configurable). Terminal ERROR
transitions and close()'s 'closed' write bypass the buffer.

Bounded buffer (``max_buffer=10000`` default) evicts the oldest
ws_id on insertion overflow — protects against unbounded growth
when storage is unreachable. ``discard(ws_id)`` drops any pending
buffered transition AND waits on a flush_lock for any in-flight
write to complete; this is the close-path hook that preserves the
bug-3 invariant (a closed row can't be resurrected by a buffered
transient writing AFTER close's sync 'closed').

13 unit tests cover coalescing, flush_now, bounded buffer, the
discard / in-flight-flush wait, lifecycle (start/shutdown
idempotence), wake-on-record latency, and resilience to storage
errors poisoning subsequent flushes.
2026-04-25 20:11:47 -07:00
Patrick Buckley 4e791cfb15 refactor(server): swap interactive workers to session_worker.send
Five spawn sites in turnstone/server.py now share the worker dispatch:

* ``send_message`` (``POST /v1/api/send``) — the main path. Now uses
  ``session_worker.send`` with separate ``_enqueue`` / ``_run``
  closures; the queue-vs-spawn outcome is conveyed via a captured
  ``queue_outcome`` dict so the existing response shapes
  (``status: queued`` vs ``status: ok``) survive.
* ``_make_watch_dispatch`` — watch results dispatch.
* ``run_retry`` (post-rewind) and ``_run_initial`` (initial-message
  on workstream creation) — set ``_worker_running`` directly under
  ws._lock instead of going through session_worker (their structural
  shape doesn't fit a queue-vs-spawn decision) but stay consistent
  with the shared gate so they can't race with /send into parallel
  workers.
* ``cancel_generation``'s ``was_running`` snapshot now reads
  ``_worker_running`` for parity with the dispatcher.

Pre-dispatch cancel-await also gates on ``_worker_running`` for
consistency. The ``busy_error`` /  ``status: busy`` legacy branch
(reached only when worker is alive but ws.session is None) is gone
— the new path checks ws.session up front and returns the same
500 shape.

Test fixtures in test_server_attachments_endpoints.py and
test_watch_dispatch.py updated to set ws._worker_running explicitly
(MagicMock auto-truthifies the field, which would otherwise mis-route
all idle paths into queue mode).
2026-04-25 20:11:47 -07:00
Patrick Buckley 7ffe8d1ca3 feat(core): add session_worker shared dispatch + delegate coord adapter
Introduces ``turnstone.core.session_worker.send`` — the atomic
check-and-(spawn-or-queue) decision both interactive and coordinator
HTTP paths use to drive ``ChatSession.send``. Callers pass no-arg
``enqueue`` / ``run`` closures; the shared module owns only the
``ws._worker_running`` lifecycle.

CoordinatorAdapter.send now delegates to the shared module — its
``_spawn_worker`` body is gone. Workstream._worker_running's
docstring updated to note both kinds use it post-Stage-2-P1.
2026-04-25 20:11:47 -07:00
Patrick Buckley abf7f62301 fix(server): address PR #409 review feedback
PR #409 line-level review feedback. Three of four findings valid;
the fourth (code-quality bot's "unused TYPE_CHECKING imports")
verified as false-positive — removing the imports breaks mypy on
the string-form annotations in ``ManagerLookup`` / ``TenantCheck``
/ ``CloseAuditEmitter``.

CI lint failure (ruff format on ``tests/_coord_test_helpers.py``)
addressed alongside.

Findings addressed:

- **Copilot #1** (``session_routes.py`` SessionEndpointConfig
  docstring): said the config is "stored on
  ``app.state.session_endpoint_config``" and "handler bodies pull
  this config from app.state". Stale after the previous fixup
  switched the factories to capture ``cfg`` via closure. Rewrote
  the class docstring + the lifted-handler comment block + the
  module docstring + the ``create_app`` block comments in both
  ``server.py`` and ``console/server.py``.
- **Copilot #2** (``server.py:_interactive_manager_lookup``
  docstring): referenced ``:data:SessionRouteHandlers`` which was
  renamed to ``SharedSessionVerbHandlers`` AND wasn't the right
  reference anyway — the callable matches
  ``SessionEndpointConfig.manager_lookup``. Fixed.
- **Bonus**: dropped the now-dead
  ``app.state.session_endpoint_config = ...`` assignments in both
  servers (nothing reads them since the closure-capture switch).
- **Bonus**: dropped the stale "close (interactive caps + redacts +
  persists close_reason)" entry from the deferred-verbs comment in
  ``session_routes.py`` — close was lifted in the previous commit
  and is no longer in the deferred set.
- **CI lint**: ``ruff format`` joined the
  ``MockStorage.list_services`` signature in
  ``tests/_coord_test_helpers.py`` to a single line (95 chars,
  fits the 100-char limit).

ruff + ruff format + mypy clean. 88 affected tests pass.
2026-04-24 16:56:06 -07:00
Patrick Buckley 74670cd53e refactor(server): apply 2nd-pass /review fixups
Addresses the eight verified findings from the second review pass on
the body-convergence work (one bug-flagged behavior change, one
defensive-style nit, six quality items). One quality item (q-6,
``request.scope[\"path_params\"]`` mutation in the legacy adapter)
is documented but not refactored — restructuring the lifted handler
signatures to take ``ws_id`` as an explicit param is bigger than
this fixup's scope; the adapter docstring already explains the
choice.

Findings addressed:

- **bug-1 + q-5**: hoist module-level ``log = get_logger(__name__)``
  in ``session_routes.py``; bump audit-failure log from ``debug``
  to ``warning`` (compliance signal). Document the interactive
  500-on-audit-failure → 200+log behavior change in CHANGELOG +
  in ``make_close_handler``'s docstring.
- **bug-2**: switch ``_audit_close_workstream`` to
  ``getattr(request.app.state, \"auth_storage\", None)`` for
  consistency with the upstream gate. Same fix on coord side.
- **q-1**: pass ``SessionEndpointConfig`` into
  ``make_approve_handler(cfg)`` and
  ``make_close_handler(cfg, *, audit_emit, supports_close_reason)``
  via closure capture. Removes the implicit ``app.state`` contract
  and parallels the two factory signatures. Tests + production
  wiring updated.
- **q-2**: promote ``_audit_close_coordinator`` to a module-level
  function in ``turnstone/console/server.py``. Both test fixtures
  import it instead of duplicating the body. The previous three
  near-identical implementations collapse to one.
- **q-3**: lift ``_interactive_tenant_check`` and
  ``_audit_close_workstream`` from nested ``create_app`` closures
  to module-level functions in ``turnstone/server.py``, beside the
  other ``_audit_*`` / ``_require_*`` helpers. Add
  ``_interactive_manager_lookup`` so the config doesn't need a
  lambda. ``create_app`` shrinks accordingly.
- **q-4**: merge the bottom ``if TYPE_CHECKING`` block into the
  one at the top of ``session_routes.py``.
- **q-7**: replace ``assert mgr is not None`` with
  ``mgr = cast(\"SessionManager\", mgr_opt)`` in both lifted
  handlers — survives ``python -O`` and makes the type-checker-only
  intent explicit.
- **q-8**: update ``test_coordinator_endpoints.py`` file docstring
  to mention the lifted-handler wiring.

ruff + mypy + 4366 pytest pass. Live console smoke against the
unified URLs returns 503 (no coord_mgr in smoke env) — proves the
factory-captured config is reachable + manager_lookup fires.

CHANGELOG ``[Unreleased]`` entry expanded to flag the audit-failure
swallow as an interactive behavior change alongside the existing
500→404 standardization.
2026-04-24 16:56:06 -07:00
Patrick Buckley 0ac5c75dcf docs(changelog): expand the [Unreleased] entry with body-convergence + SDK status
Adds two paragraphs:

- Notes the two verbs (``approve``, ``close``) whose bodies were
  successfully lifted into the shared registrar, plus the close-
  failure status-code standardization (500 → 404 on coord). Calls
  out the verbs whose bodies are intentionally NOT lifted, with
  the underlying reason (Priority 1 dependency, response-shape
  unification, etc.) so the next-session reader doesn't re-litigate.
- Notes the TS SDK 0.4.0 bump and the regenerated reference specs.

No code change.
2026-04-24 16:56:06 -07:00
Patrick Buckley 06c91294a4 refactor(server): lift close handler into shared session_routes body
Stage 2 Priority 0 Step 0.2 body-convergence — second verb.
``make_close_handler(audit_emit=..., supports_close_reason=...)``
factory in ``turnstone/core/session_routes.py`` produces the lifted
body; both interactive and coord pass their kind-specific audit
emitter at app construction.

The two body-keyed close URL aliases on the interactive side reach
the same lifted body:

- ``POST /v1/api/workstreams/{ws_id}/close`` (new, path-keyed)
  via ``register_session_routes(handlers.close=...)``.
- ``POST /v1/api/workstreams/close`` (legacy, body-keyed) via
  ``make_legacy_body_keyed_adapter(close_handler)``.

Coord exposes only the path-keyed shape.

Behavior gains:

- ``supports_close_reason=True`` (interactive only) keeps the 512-
  byte UTF-8 cap + credential redaction + ``workstream_config``
  persistence path. Coord stays at ``False``; if coord ever wants
  close-reason metadata, flipping the flag is a one-line change.
- ``audit_emit`` is per-kind so each owns its detail dict shape
  (``{kind, parent_ws_id, reason}`` vs ``{coord_ws_id, src}``) and
  audit action name (``workstream.closed`` vs ``coordinator.close``).
- Standardizes the close-failure status code to 404 across both
  kinds. The coord code previously returned 500 on a
  ``mgr.close()`` race-loss, which was overly pessimistic — the
  semantic is "the ws was popped between .get() and .close()", i.e.
  not-found.

Coord-side test fixtures (``test_coordinator_endpoints``,
``test_coordinator_end_to_end``) swap the imported
``coordinator_close`` for the lifted handler + a local audit_emit
adapter so the tests exercise the same code path the live console
does.

ruff + mypy + 4366 pytest pass. Live console smoke against
``POST /v1/api/workstreams/abc/close`` returns 503 (no coord_mgr
loaded in the smoke env) — proves the lifted handler is reachable
+ the manager_lookup callable fires correctly.

Two verbs converged so far (``approve`` + ``close``); the remaining
pairs (``send``, ``cancel``, ``open``, ``events``, ``create``,
``list``, ``saved``, ``history``, ``detail``) have substantive
behavior divergence that doesn't factor cleanly into the
SessionEndpointConfig + factory-handler pattern — see the
session_routes module docstring for the per-verb status.
2026-04-24 16:56:06 -07:00
Patrick Buckley 6415eeb91e refactor(server): lift approve handler into shared session_routes body
Stage 2 Priority 0 Step 0.2 body-convergence — first verb. Both
interactive ``approve`` and coord ``coordinator_approve`` handler
bodies collapse into ``make_approve_handler()`` in
``turnstone/core/session_routes.py``. Each kind sets a
``SessionEndpointConfig`` on ``app.state`` carrying the kind-
specific policies (auth gate, manager lookup, tenant check, audit
prefix, not-found label) the lifted body consults at request time.

The two interactive URLs converge:

- ``POST /v1/api/workstreams/{ws_id}/approve`` (new, path-keyed)
  reaches the lifted body directly via ``register_session_routes``.
- ``POST /v1/api/approve`` (legacy, body-keyed) keeps shipping;
  ``make_legacy_body_keyed_adapter`` peeks the body for ``ws_id``,
  splices it into ``request.path_params``, and forwards to the same
  lifted body. Frontend can keep using the legacy URL — no caller
  churn.

Coord exposes only the path-keyed shape (its URLs were experimental
in 1.5.0aN; the URL-shape commit already removed the ``coordinator/``
prefix).

Tenant-check is split out from permission-gate so interactive's
``_require_ws_access`` (404 on cross-owner) and coord's
``_require_admin_coordinator`` (cluster-wide scope) coexist without
either kind triggering the wrong gate.

Coord-side test fixture (``test_coordinator_endpoints._make_client``)
swaps the imported ``coordinator_approve`` for the lifted handler
and seeds ``app.state.session_endpoint_config`` so the tests
exercise the same code path the live console does.

Net delta: ~−25 LOC for this verb on top of the SessionEndpointConfig
+ legacy-adapter scaffolding (~80 LOC paid once). Subsequent verb
lifts amortize against that scaffolding.

ruff + mypy + 4366 pytest pass. Live console smoke against the
unified URL returns 503 (no coord_mgr loaded in the smoke env) —
proves the lifted handler is reachable + the manager_lookup callable
fires correctly.

Verbs still kind-specific (deferred — bodies have substantive
behavior divergence, not just naming): ``send`` (Priority 1
worker dispatch), ``cancel`` (interactive forensics + force flag),
``close`` (interactive close-reason cap+redact+persist), ``open``
(interactive resume vs coord rehydrate), ``events`` (different SSE
replay shapes), ``create`` (interactive attachments vs coord
initial_message), ``list`` / ``saved`` (different response keys).
2026-04-24 16:56:06 -07:00
Patrick Buckley 4a72b2ce19 build(sdk): regen openapi specs + bump TS SDK to 0.4.0
Stage 2 Priority 0 Step 0.5 follow-on. The handwritten Python
OpenAPI spec already moved to ``/v1/api/workstreams/`` in the URL
sweep commit; this just regenerates ``sdk/typescript/openapi-{server,console}.json``
from those specs so generated TS callers see the new paths.

Bumps the TS SDK to 0.4.0 to flag the URL-shape break for any
1.5.0aN-era consumer of the experimental coord client. Python SDK
needs no change — it never exposed the coord HTTP surface.

TS typecheck + 32 vitest tests pass.
2026-04-24 16:56:06 -07:00
Patrick Buckley ae8ffd4bad refactor(server): tighten registrar shape per code-review pass
Addresses the eight quality findings the per-priority /review pass
flagged on the registrar refactor. All confirmed by the verifier;
none blocking. Net −186 LOC in this fixup.

q-1, q-9: trim ``session_routes.py`` module docstring + console
``create_app`` comments to the timeless explanation. The Step 0.1 →
0.4 narrative was already stale within the PR that introduced it
(every step had landed by the final commit) and would rot further as
the body-convergence follow-on lands.

q-2, q-5: group the four attachment handlers into an
``AttachmentHandlers`` dataclass exposed as
``handlers.attachments: AttachmentHandlers | None``. The type system
now carries the all-or-none invariant; the parallel four-condition
chain + bare ValueError disappear.

q-3: drop the ``mgr`` and ``adapter`` placeholder kwargs from both
``register_session_routes`` and ``register_coord_verbs``. Pre-
threading them so a future commit avoids "callsite churn" violated
the project's "don't pre-build for the next step" norm — the
body-convergence follow-on will edit the callsites anyway. Drops
``SessionManager.adapter`` for the same reason.

q-4: drop ``SessionRouteConfig`` outright. It existed solely to
carry ``supports_legacy_close``; the registrar now mounts the legacy
close route whenever ``handlers.close_legacy is not None``, matching
the all-Optional convention used for every other handler.

q-6: move ``MockStorage`` from ``tests/test_console.py`` into the
shared ``tests/_coord_test_helpers.py`` and re-import in test_console
+ test_session_routes. No more cross-test-module import.

q-7: trim the two exhaustive route-table set-equality assertions
(``test_coord_shape_mounts_expected_verbs``,
``test_register_coord_verbs_mounts_expected_paths``); replaced with
focused ``test_attachment_routes_mount_when_quartet_provided`` and
``test_close_legacy_mounts_when_handler_provided``. The targeted
ordering tests still catch the actual registrar bugs.

q-8: delete the tombstone comment block where the legacy
``/api/coordinator/`` Routes used to be — per the user's
``feedback_no_tombstone_comments`` norm, deletions don't get
narrated inline.

q-10: rename ``SessionRouteHandlers`` → ``SharedSessionVerbHandlers``
and ``CoordVerbHandlers`` → ``CoordOnlyVerbHandlers`` so the
"shared verbs vs coord-only verbs" symmetry is visible at the type
names. Drop the back-compat aliases since nothing uses them.
2026-04-24 16:56:06 -07:00
Patrick Buckley ffac49d098 docs(changelog): flag the coord URL move under [Unreleased]
The plan called the CHANGELOG callout for the
``/v1/api/coordinator/`` → ``/v1/api/workstreams/`` move
non-negotiable since experimental SDK consumers from 1.5.0aN lose
the URL outright. Adds the path-mapping table under [Unreleased]
so the line lands in the 1.5.0 release notes when the version cuts.

Stable upgraders (1.0 / 1.3 / 1.4) never saw the coord URL prefix
so the change is a no-op for them — the entry says so explicitly.
2026-04-24 16:56:06 -07:00
Patrick Buckley df7c0c2f44 refactor(server): delete legacy /v1/api/coordinator/ URL tree
Stage 2 Priority 0 Steps 0.4–0.7 — collapses the four migration
steps into one commit since they have to land together. The legacy
``/v1/api/coordinator/`` URL prefix never shipped in a stable release
(it appeared in 1.5.0aN experimental), so there's no compat carry-
forward — just rip and replace.

What moves:

- Step 0.4: deletes the eighteen ``Route("/api/coordinator/...")``
  entries from ``console/server.py``. Coord traffic now flows
  exclusively through the unified ``/v1/api/workstreams/`` shape
  mounted via ``register_session_routes`` + ``register_coord_verbs``
  (Steps 0.2 and 0.3).
- Step 0.5: rewrites the OpenAPI spec (``console_spec.py``) and
  schemas (``console_schemas.py``, ``server_schemas.py``) to
  document the new paths. ``test_openapi.py`` parity assertions
  swap with them.
- Step 0.6: mechanical URL sweep across the frontend
  (``console/static/app.js`` — 9 sites; ``coordinator/coordinator.js``
  — 16 sites; ``index.html`` — 1 comment).
- Step 0.7: same sweep across the test suite
  (``test_coordinator_endpoints.py``, ``test_coordinator_end_to_end.py``,
  ``test_coordinator_governance.py``, ``test_coordinator_close_all_children.py``,
  ``test_coordinator_client.py``, ``test_phase6_endpoints.py``).

Also touched:

- Server-side ``CoordinatorClient`` (``coordinator_client.py``) —
  the coord agent's HTTP path for ``close_all_children`` updates
  to the new shape.
- Handler docstrings in ``console/server.py`` say
  ``POST /v1/api/workstreams/...`` not ``/coordinator/...`` so a
  ``grep`` for a verb's URL lands on the right line.
- ``settings_registry.py`` setting descriptions, ``server.py``
  cross-process error message, migration 042 docstring — all
  updated to the unified shape.

The handler functions stay named ``coordinator_*`` until the
body-convergence follow-on lifts them into ``session_routes`` with
kind branching behind ``SessionRouteConfig`` flags. URL surface is
the only thing that changes here.

The ``test_session_routes`` route-walk now asserts the legacy paths
are GONE — previously it asserted both shapes coexisted. A future
accidental remount of ``/api/coordinator/`` would fail that test.
2026-04-24 16:56:06 -07:00
Patrick Buckley 377bd58b67 refactor(server): mount coord-only verbs through register_coord_verbs
Stage 2 Priority 0 Step 0.3 — adds ``CoordVerbHandlers`` +
``register_coord_verbs`` to ``turnstone.core.session_routes`` and
wires the seven coord-only verbs (``children`` / ``tasks`` /
``metrics`` / ``trust`` / ``restrict`` / ``stop_cascade`` /
``close_all_children``) through it on the console.

These verbs are legitimately kind-specific — they read or mutate
state (children registry, parent quota, trust / restrict policy,
cascade controls) that doesn't exist on interactive workstreams —
so they live on a Protocol distinct from
``SessionRouteHandlers``. The unified URL prefix
``/api/workstreams/{ws_id}/`` is shared with the session verbs;
the separate registrar call keeps the kind separation explicit at
the wiring site.

Legacy ``/api/coordinator/{ws_id}/{verb}`` paths stay live during
the transition; both URL shapes resolve to the same handler
function. Step 0.4 deletes the legacy shape.

The route-table walk in ``test_session_routes`` now covers all
eighteen verb pairs (eleven session + seven coord-only) so a
future drift between legacy and unified handlers fails CI.
2026-04-24 16:56:06 -07:00
Patrick Buckley e1ee84af42 refactor(server): mount coord verbs through shared session route registrar
Stage 2 Priority 0 Step 0.2 — extends ``register_session_routes`` to
cover the per-``{ws_id}`` interaction verbs (``send`` / ``approve`` /
``plan`` / ``cancel`` / ``close`` / ``events`` / ``history`` /
``detail``) and wires ``console/server.py`` to mount coord at the
unified ``/v1/api/workstreams/`` shape.

The legacy ``/v1/api/coordinator/`` paths stay live during the
transition; both URL shapes resolve to the same handler functions.
Step 0.4 deletes the legacy shape outright once the frontend
(Step 0.6) and tests (Step 0.7) move off it.

Handler bodies still live in their server modules — body
convergence (kind branching behind ``SessionRouteConfig`` flags) is
the next Step 0.2 follow-on. Splitting "URL surface unified" from
"handler bodies converged" keeps the soak-able diffs small.

``mgr`` and ``adapter`` registrar arguments are now Optional because
the console builds its coord ``SessionManager`` inside the lifespan
(after app construction). They become required again in the
body-convergence follow-on once the lifted handlers read them.

New ``tests/test_session_routes.py`` covers the registrar's mounting
rules (route ordering, attachment-quartet enforcement, legacy-close
gate) and asserts that each unified path on the console points at
the SAME handler object as its legacy counterpart — the transition
is a pure URL alias, not a fork.
2026-04-24 16:56:06 -07:00
Patrick Buckley 2b5e6cb252 refactor(server): scaffold shared session route registrar
Stage 2 Priority 0 Step 0.1 — introduces
``turnstone/core/session_routes.py`` (``SessionRouteConfig`` +
``SessionRouteHandlers`` + ``register_session_routes``) and rewires
``server.py``'s ``/v1/api/workstreams/*`` route table to mount through
it. Pure scaffolding: handler bodies stay where they are, URL surface
is byte-identical, tests pass unchanged.

Sets up Step 0.2 to lift handler bodies into the registrar and have
the console mount the same shape against its coord manager.

Adds ``SessionManager.adapter`` accessor so the registrar can pick
up the kind adapter without callers re-threading it through every
construction layer.
2026-04-24 16:56:06 -07:00
Patrick Buckley c837e3fa6d feat(core): Stage 1 SessionManager unification (#408)
* feat(core): scaffold SessionManager + SessionKindAdapter Protocol

Stage 1 step 1 — pure addition, no production wiring. Defines the
shape later steps will port the shared mechanics onto: slot
accounting, per-ws-id refcounted rehydrate locks, kind-agnostic
lifecycle; kind-specific event transport + session construction on
the adapter.

Pruned from the earlier Protocol draft (see design brief): per-kind
permission_scope (static handler map is simpler), allows_child_spawn /
quota_policy (deleted in #403), on_child_spawned (coordinator tool
owns children registry), allows_active_focus / active_id / switch
(frontend owns the active-tab state).

* feat(core): port shared session-lifecycle mechanics onto SessionManager

Stage 1 step 2. Adds create / open / close / set_state / close_idle /
get / list_all / count on top of the Step 1 scaffolding. Pure
addition — still no production wiring; the new class doesn't replace
any call sites yet.

Concurrency shape is ported from CoordinatorManager (the more-
complete side): single-phase slot reservation under the manager
lock, per-ws refcounted open-lock to serialize concurrent lazy
rehydrate, placeholder workstreams count toward max_active but can't
evict each other. WSM's two-phase eviction outside the lock is not
carried over; it had a window where a burst of creates could silently
exceed max_active.

Deletions (vs. the union of the two old managers):
- "refuse to close last workstream" guard — handled by the
  dashboard; only existed to protect the now-deleted default startup
  workstream.
- active_id / switch / get_active — frontend owns focus; server-side
  duplicate state is gone.
- _active_coords presence cache — defer measurement to Step 4; if it
  pays for itself at realistic cluster sizes, the CoordinatorAdapter
  can maintain it by observing emit_* calls.
- Children registry + reverse index — coordinator tool owns this,
  manager stays kind-agnostic.

Skill resolution (name → template_id + applied_version) is now
shared via SessionManager._resolve_skill, so WSM's pre-resolve-at-
callsite pattern and CM's internal-lookup pattern converge. Callers
pass the skill name; the manager does the lookup once.

26 smoke tests cover create eviction + overflow, concurrent-create
cap, persist/session rollback, open for missing/deleted/wrong-
kind/wrong-user rows, concurrent-open serialization, close unblocks
UI + emits closed, set_state + storage + adapter observer,
close_idle, list_all ordering, count, eviction fires adapter
transport, node_id passthrough.

* feat(core): add InteractiveAdapter for SessionManager

Stage 1 step 3. Adapter that bridges SessionManager to the node's
interactive transport:

- emit_created/state/closed → pushes onto the process-wide SSE
  global_queue (same shape current server.py handlers produce inline)
- cleanup_ui → ports WorkstreamManager._cleanup_ui body: unblock
  _approval_event / _plan_event / _fg_event, broadcast ws_closed to
  per-UI listener queues (with full-queue fallback), cancel + close
  the session
- build_ui/build_session → delegate to injected factories
  (ui_factory builds WebUI, session_factory is the existing closure
  from server.py with judge_model + memory_config captures)

Also extends SessionKindAdapter.build_session with **extra passthrough
so interactive callers can pass judge_model per-call without polluting
the manager API; and adds a reason= kwarg to emit_closed so the
frontend's "evicted" special-case keeps working (frontend doesn't
differentiate "idle" from "closed", so close_idle collapses into
close()).

14 new adapter tests cover wire payload shape, queue.Full tolerance,
cleanup_ui event unblocking + listener broadcast + queue-full
fallback, session cancel+close, graceful handling of stub UIs / None
session, and kwarg passthrough to the session factory.

* feat(console): add CoordinatorAdapter for SessionManager

Stage 1 step 4. Coordinator-side SessionKindAdapter implementation:

- emit_created/state/closed → delegate to the existing
  ClusterCollector.emit_console_ws_* methods (same wire shape the old
  CoordinatorManager emitted inline)
- cleanup_ui → ports the listener-queue + approval/plan event
  unblocks from CoordinatorManager._cleanup, with queue-full
  fallback so an unresponsive browser tab can't wedge close
- build_ui/build_session → delegate to injected factories; session
  factory doesn't accept client_type so we strip it at the adapter
  boundary

Collector emission exceptions are swallowed (same policy as today's
inline fan-out — dashboard lag on one tick is preferable to breaking
the lifecycle path).

Intentionally out of scope: the children registry (_children /
_child_to_coord) stays in the coordinator tool when wired in Step 5;
the _active_coords lock-free presence cache is deferred pending a
measurement at realistic cluster sizes. 10 new tests cover transport
payloads, collector-exception tolerance, cleanup_ui event unblock +
listener broadcast + queue-full eviction, construction passthrough.

* feat(server): wire interactive server.py to SessionManager

Stage 1 step 5a. Production-path swap: WorkstreamManager →
SessionManager(InteractiveAdapter(...)).

- Construction at server startup: build the adapter with the
  process-wide global_queue, a WebUI ui_factory closure, and the
  existing session_factory. SessionManager gets storage + max_active.
- Default startup workstream wiring removed (the CLI-REPL leftover
  flagged in the handoff's "Convergence is also a pruning
  opportunity" section). --resume now lazily creates a workstream
  scoped to the resumed content; no workstream at all if --resume
  isn't given. The dashboard handles the 0-ws state.
- HTTP handler mgr.create() calls switched to the new kw-only
  signature (user_id, name, model, skill, ws_id, client_type,
  judge_model, parent_ws_id). ui_factory/skill_id/skill_version/kind
  no longer threaded through — adapter handles UI construction and
  manager resolves skill internally.
- Dropped the mgr.last_evicted block in the /new handler (adapter
  emits ws_closed:evicted automatically on capacity eviction).
- mgr.max_workstreams → mgr.max_active.
- Added active_id / switch / switch_by_index / get_active / index_of
  / eviction_count to SessionManager because turnstone/cli.py uses
  them extensively; the handoff's "delete unless there's a live
  caller" rule flips here — CLI is a live caller.

Test fixtures across 9 files updated to build SessionManager +
InteractiveAdapter rather than WorkstreamManager. test_workstream.py
stays unchanged (it tests WSM directly; it'll be deleted in step 5d
alongside the class itself).

Full pytest: 4528 passed. Ruff + mypy clean. Next: 5b (console-side
wiring, with the children-registry relocation to the coordinator
tool).

* feat(console): wire console server to SessionManager

Stage 1 step 5b. Production-path swap: CoordinatorManager →
SessionManager(CoordinatorAdapter(...)).

- CoordinatorAdapter now owns the coord-specific bits that were bolted
  onto the old CoordinatorManager: the children registry (forward +
  reverse index), the lock-free active-coords presence cache, the
  cluster-event fan-out thread, and the worker-dispatch path
  (send / _spawn_worker). The shared SessionManager stays kind-agnostic.
- Added CoordinatorAdapter.attach(mgr) for late-binding the owning
  manager (the manager's ctor takes the adapter, so the dependency has
  to break here). Used inside _rebuild_children_registry for the tenant-
  filtered SQL query, inside send/dispatch for mgr.get(ws_id), and
  inside the fan-out seed path for mgr.list_all().
- emit_created now seeds the children registry + active-coords slot AND
  calls _rebuild_children_registry (covers both create — empty query —
  and open/rehydrate, where the subtree is persisted). emit_closed
  drops both entries. Collapses the three old call-sites in
  CoordinatorManager's create/open/close into one per-event hook.
- Console server.py builds the manager via:
      coord_adapter = CoordinatorAdapter(collector=..., ...)
      coord_mgr = SessionManager(coord_adapter, storage=..., max_active=...,
                                 node_id=ClusterCollector.CONSOLE_PSEUDO_NODE_ID)
      coord_adapter.attach(coord_mgr)
      ConsoleCoordinatorUI._coord_mgr = coord_mgr
      app.state.coord_adapter = coord_adapter
- HTTP handler call-site updates:
  - coord_mgr.create drops initial_message; the handler now calls
    coord_adapter.send(ws.id, initial_message) after create so the
    worker spawn stays out of the shared manager.
  - coord_mgr.open_admin(ws_id) → coord_mgr.open(ws_id, user_id="",
    admin=True). Matches SessionManager.open's unified signature.
  - coord_mgr.list_for_user(uid) inlined as a list comp on list_all()
    (SessionManager doesn't expose the filter; two callers).
  - coord_mgr.children_snapshot / send → coord_adapter.*.
  - coord_mgr.cancel stays (now lives on SessionManager from 5a).
- ConsoleCoordinatorUI.on_state_change now flows state transitions
  through ConsoleCoordinatorUI._coord_mgr.set_state, mirroring the
  WebUI pattern. The old _on_state_observer / _on_rename_observer
  closures the manager used to install are dead code now; leaving the
  fields in place for 5d cleanup.
- Lifespan shutdown calls coord_adapter.shutdown() (was coord_mgr.
  shutdown()) and resets ConsoleCoordinatorUI._coord_mgr on teardown.

Test fixture updates in _coord_test_helpers, test_coordinator_end_to_end,
test_coordinator_endpoints, test_phase6_endpoints: build SessionManager
+ CoordinatorAdapter in _build_mgr, set app.state.coord_adapter, switch
mgr.register_children / mgr.children_snapshot tests to mgr._adapter.*,
and rewrite test_open_admin_uses_open_admin to assert the unified
open(user_id="", admin=True) call shape.

Full pytest: 4486 passed. Ruff + mypy clean. Next: 5d (remove
CoordinatorManager + WorkstreamManager class bodies and their test
files).

* feat(core): delete WorkstreamManager + CoordinatorManager classes

Stage 1 step 5c + 5d. Final step of the unification — the legacy
classes and their test files go away now that every production
caller has been ported.

- Delete turnstone/console/coordinator.py entirely (CoordinatorManager
  class + the _enqueue_on_ui helper, which CoordinatorAdapter now hosts
  its own copy of).
- Trim turnstone/core/workstream.py to just the Workstream dataclass +
  WorkstreamKind + WorkstreamState. ~385 lines of WorkstreamManager
  logic gone; the remaining shape is pure data types shared by both
  managers.
- Delete tests/test_workstream.py (WSM-specific) and
  tests/test_coordinator_manager.py (CM-specific).
- Wire turnstone/cli.py to SessionManager + InteractiveAdapter, same
  pattern as turnstone/server.py. The CLI's WorkstreamTerminalUI uses
  manager.set_state + manager.active_id — both preserved on
  SessionManager (CLI is a live caller that keeps the focus API
  honest, per the handoff's "delete unless it pulls its weight" rule).
- Add an optional manager-level ``_on_state_change`` observer hook
  restored for the CLI's background-attention notification (the web
  path uses the adapter's emit_state; this hook covers callers that
  don't consume SSE).
- Drop dead ``_on_state_observer`` / ``_on_rename_observer`` fields
  from ConsoleCoordinatorUI — the old CoordinatorManager installed
  them; SessionManager/CoordinatorAdapter handle fan-out directly.

Vulture @ 80% confidence: zero unused symbols across the new
SessionManager + adapter files. Ruff + mypy clean (170 files).
Full pytest (excluding tests/live): 4414 passed.

Net across the whole Stage 1 branch: one unified SessionManager +
adapter Protocol replaces two ~500-line parallel managers + a
~600-line CoordinatorManager, and the interactive + coordinator
transports stay cleanly separated at the adapter boundary.

* refactor(auth): drop workstream row-level ownership gates

Turnstone is a trusted-team tool (per #400). user_id stays as
metadata for audit + display; it no longer rejects requests. Scope-
level auth via admin.workstreams / admin.coordinator tokens is the
only gate now.

Solves sec-1 (cross-tenant delete via collision on caller-supplied
ws_id, because the gate was half-implemented) and sec-2 (blank-sub
JWT bypass on empty-owner rows). Net: 359 lines of defensive
empty-string comparisons and admin=True bypass plumbing deleted.

* fix(core): serialize set_state vs close + worker spawn

Three concurrency fixes from the multi-stage review:

- bug-3: set_state now looks up ws under self._lock and gates its
  storage write on ws._closed (a new tombstone flag). close() sets
  ws._closed=True and does its storage write under ws._lock. A
  set_state that acquires ws._lock after close sees the tombstone
  and skips its write instead of resurrecting the closed row.

- bug-1: _spawn_worker wraps the check-and-spawn in ws._lock so two
  concurrent send() HTTP requests can't both observe "no live worker"
  and start duplicate worker threads on the same ChatSession.

- bug-2: replaces Thread.is_alive() as the reuse gate with an
  explicit ws._worker_running flag. The flag is set before the worker
  thread starts and cleared in its finally block — both under
  ws._lock. Using is_alive() left a narrow window where the worker
  could exit between the check and a queue_message call, stranding
  the user's message with no consumer.

perf-2 (lock-held-across-DB-write) is accepted as-is: per-ws
serialization of state transitions behind a DB round-trip is real
cost but bounded — a given ws's state flips happen sequentially on
its worker thread anyway. Dropping ws._lock around the DB write
would reintroduce the bug-3 race.

Full pytest: 4401 passed. Ruff + mypy clean.

* refactor(core): drop _resolve_skill from SessionManager

Skill resolution (name → template_id + applied_version) moves out of
the shared manager and back to the HTTP handlers that own the
create request. The interactive handler already resolved skill_data
+ applied_skill_version for other purposes (model override, judge
config, post-create session seed) and was passing the name to
SessionManager which then redundantly re-resolved via
get_skill_by_name + count_skill_versions — two wasted DB round-trips
per create on a user-visible latency path.

- SessionManager.create: accepts skill_id + skill_version as
  already-resolved kwargs; _resolve_skill helper deleted.
- turnstone/server.py create_workstream: passes the skill_id /
  applied_skill_version it already computed.
- turnstone/console/server.py coordinator_create: pre-resolves
  inline (parity with interactive) before calling coord_mgr.create.

Fixes perf-1 (redundant skill queries per create), q-4 (divergent
skill-version computation between manager and handler), q-5
(coordinator-specific lookup on the shared manager surface).

Full pytest: 4401 passed. Ruff + mypy clean.

* refactor(adapters): extract shared cleanup_ui + drop dead child-registry methods

Both InteractiveAdapter.cleanup_ui and CoordinatorAdapter.cleanup_ui
(plus their _broadcast_ws_closed_to_listeners helpers) were byte-identical.
Pull them into turnstone/core/adapters/_ui_cleanup.py:cleanup_session_ui
so the two adapters delegate to one implementation.

Also drop CoordinatorAdapter.register_children (only test callers — now
use _seed_children in tests/_coord_test_helpers.py) and _add_child
(zero callers anywhere).

* refactor(adapters): symmetric attach() + fail-loud on unattached manager

Add InteractiveAdapter.attach(manager) + .manager property mirroring
the coord-side pattern. CLI (cli.py) now uses cli_adapter.attach(manager)
instead of the _mgr_ref list-ref late-binding hack; server.py picks up
the same call for consistency.

CoordinatorAdapter.send / _rebuild_children_registry /
_prime_children_from_snapshot no longer silently return when
self._manager is None — raise RuntimeError so a forgotten attach() at
startup fails loud instead of dropping the whole fan-out.

* docs: replace stale WorkstreamManager / CoordinatorManager references

Both classes were deleted in 965e0b6; prose docstrings across the
codebase still named them. Update to SessionManager (or describe the
collapsed-into-one-class architecture where the distinction matters).

Leaves the 'Ported from …' historical markers in session_manager.py /
coordinator_adapter.py / interactive_adapter.py intact — those are
deliberate pointers back to the pre-unification code.

* fix(core): atomic close_if_idle + batch pop under one lock

bug-5: SessionManager.close_idle re-checked ws.state == IDLE outside
the lock, so a pending tool result could flip state IDLE→RUNNING
between the snapshot and close() acquiring self._lock. Add
_close_if_idle_locked that tests state + pops under self._lock.

perf-5: drop the per-victim self._lock acquisition; collect + pop the
whole batch in one acquisition, then run cleanup_ui / storage write /
emit_closed outside the lock.

* perf(coord): split emit_created / emit_rehydrated to skip storage query on fresh creates

CoordinatorAdapter.emit_created was unconditionally calling
_rebuild_children_registry (storage.list_workstreams with
parent_ws_id=... limit=10001) on every create, even for fresh-create
paths that provably have zero children.

Add emit_rehydrated to the SessionKindAdapter Protocol. SessionManager
.create still calls emit_created; .open (lazy rehydrate) now calls
emit_rehydrated. CoordinatorAdapter.emit_created seeds the registry +
fan-out but skips the rebuild; emit_rehydrated seeds + rebuilds + fans
out. InteractiveAdapter.emit_rehydrated delegates to emit_created (no
children-registry on the interactive transport).

* perf(coord): fold _active_coords into _children_lock + mutate payload in place

perf-4: _active_coords used a copy-on-write dict-swap pattern so the
fan-out dispatch could read it lock-free, but _dispatch_child_event
already re-validates the parent under _children_lock anyway — the
lock-free snapshot was premature. Replace with a plain dict read+write
both under _children_lock; install and remove collapse to one-liners.
Value also drops the user_id half — dead after a46dab1 removed
row-level ownership gates — so _active_coords is now just
coord_ws_id → ui.

perf-6: _enqueue_on_ui was doing {**payload, "ws_id": coord_ws_id} on
every dispatch. The dispatch path owns payload and doesn't reuse it —
mutate in place.

* test(coord): add adapter tests for worker dispatch + children registry + fan-out

Fills the coverage gap on CoordinatorAdapter — the review (q-3) flagged the
coord-specific concurrency paths ported from the deleted CoordinatorManager
as untested. Three new test classes:

- TestCoordinatorAdapterWorkerDispatch: _spawn_worker reuse gate, queue.Full
  backpressure, concurrent-call bug-1 reproducer (two threads → exactly one
  worker via ws._lock + _worker_running), finally-clears-flag.
- TestCoordinatorAdapterChildrenRegistry: registry seed on emit_created vs
  emit_rehydrated rebuild, _pop_coord_registry_locked reverse-index cleanup,
  _merge_child_ids_locked idempotency, _prime_children_from_snapshot merge.
- TestCoordinatorAdapterDispatchChildEvent: unknown-parent drop, ws_created
  fan-out, cluster_state / ws_closed reverse-index routing, perf-6 in-place
  ws_id stamp.

* fix: regressions flagged by ultrareview

Verify stage of the cloud review surfaced 6 confirmed regressions
from Stage 1's adapter layer. Fixing together since they share the
same root cause (plumbing moved into adapters without retiring the
old emission paths).

- Interactive adapter emit_created / emit_state / emit_rehydrated
  become no-ops. The create_workstream HTTP handler still fires
  ws_created (after attachment validation, per the pre-Stage-1
  "no phantom events on rejected upload" contract); WebUI
  _broadcast_state still fires ws_state with the full payload
  (tokens + context_ratio + activity). Firing from the adapter too
  was duplicating both events. Also closes the phantom-ws-created
  regression (adapter fired before attachment validation ran).

- emit_closed Protocol gains a ``name`` kwarg; the adapter is the
  sole emitter for ws_closed on interactive now, and the frontend
  eviction toast needs the name. Manager passes ws.name from
  close() / create()+open() eviction / close_idle paths.

- _idle_cleanup_thread stops firing its own reason="idle" ws_closed
  — close_idle already fires via the adapter with reason="closed",
  and the frontend never differentiated the two anyway.

- close_workstream_endpoint fix: "Cannot close last workstream" 400
  was a stale error (the guard went away with the default-startup
  workstream). Return 404 on close() == False (which now means the
  ws was already closed or unknown). Also switches the audit actor
  from _require_ws_access's stored owner to _auth_user_id — the
  stored owner is metadata post-#400, so attributing actions to it
  misrepresents who actually did them.

- CLI /ws close mirrors the same stale-error fix.

- SessionManager.close now calls storage.delete_workstream_override
  alongside update_workstream_state, same as the old
  WorkstreamManager.close did. Without it overrides leak until
  tombstone cleanup. close_idle does the same.

- SessionManager._reserve_and_install_locked records the eviction
  on turnstone.core.metrics so the global eviction counter keeps
  working. Old WSM did this inline; the unification dropped it.

- ConsoleCoordinatorUI.on_rename now fans out to the cluster
  collector via a new class attribute ``_collector`` (set at
  console startup alongside ``_coord_mgr``). The old
  ``_on_rename_observer`` plumbing went away with
  CoordinatorManager and the "adapter emit_console_ws_rename runs
  from whichever code path renames" comment was aspirational —
  nothing actually did it.

Full pytest: 4375 passed (tests/live + test_server_live.py excluded;
both pre-existing live-backend failures unrelated to this branch).
Ruff + mypy clean.

* refactor(ui): extract SessionUIBase for shared UI scaffolding

Direct response to review feedback that the unification wasn't
merging enough of the two workstream kinds. WebUI (node) and
ConsoleCoordinatorUI (console) both:

- Keep a per-UI list of SSE listener queues guarded by a lock
- Block a worker thread on _approval_event / _plan_event
- Fan enqueued events out with the same ws_id-stamping pattern
- Resolve approvals / plans with the same broadcast-then-signal
  pattern

All of that now lives once in turnstone/core/session_ui_base.py.
Both UIs subclass SessionUIBase; kind-specific bodies (WebUI's
per-UI metrics + _broadcast_state + intent-verdict bookkeeping,
ConsoleCoordinatorUI's collector fan-out) stay in the subclasses.

WebUI.resolve_approval still overrides the base (it adds intent-
verdict updates) but now calls super() for the shared broadcast +
event-set steps. Same shape as the other approval/plan hooks:
subclasses extend, base provides skeleton.

Net file-level: +156 LOC for the base, -144 LOC across the two
subclasses. The raw number is unexciting — but there's now a
single source of truth for the listener + blocking-gate machinery,
and bugs (like the duplicate ws_created / ws_state events that
prompted this refactor) can't arise from the two implementations
drifting.

Full pytest: 4375 passed. Ruff + mypy clean.

* refactor(ui): move metrics + verdict bookkeeping into SessionUIBase

Second pass at unifying the two UIs. Per-workstream metrics
accumulators (token counts, tool-call counts, context ratio,
activity tracking), intent-judge verdict cache + pending-decision
list, and the verdict-persistence path all move to SessionUIBase.

Before: WebUI tracked all of it; ConsoleCoordinatorUI tracked none
of it (a comment on the old on_intent_verdict literally admitted
the deferral — "skip the persistence + late-decision plumbing that
WebUI does"). Coord sessions never got verdict rows in storage, never
had a user_decision stamped, and the dashboard had no way to show
coord token usage because the data wasn't captured.

Now the base class captures the data and persists the rows for
every kind. Kind-specific broadcast (WebUI's _broadcast_state with
rich per-UI payloads) stays on WebUI; prometheus counters on the
node (_metrics.record_judge_verdict) stay on WebUI's on_intent_verdict
override. Everything else shared.

Behaviour change worth flagging: coord sessions now write
intent_verdicts and output_assessments rows for every judge call
and every output-guard warning. Previously silent; the storage rows
now exist and any future coord-dashboard surface can read them.

Shape of the unification:
- resolve_approval: was overridden on WebUI (intent-verdict decision
  propagation); now lives on the base. Both kinds inherit unchanged.
- on_intent_verdict: WebUI overrides only to add _metrics.record_*;
  rest of the body is the base.
- on_output_warning: was on both separately; fully base-shared now.

Full pytest: 4375 passed. Ruff + mypy clean.

* fix: regressions flagged by second-pass review

Three confirmed findings with direct fixes + a dedicated test file
for SessionUIBase (was previously uncovered).

bug-1 — Coord approve_tools didn't reset _last_verdict_decision or
clear _llm_verdicts between approval rounds. WebUI did (inline).
Coord inherited SessionUIBase.on_intent_verdict which stamps via
the decision flag, so after the first resolve every subsequent
round's verdicts were stamped with the prior round's user_decision
before the user had decided the new round.

Fix: add SessionUIBase._reset_approval_cycle() clearing both under
_ws_lock; call from the top of both subclass approve_tools methods.
Single-source invariant — can't drift again.

sec-1, sec-2 — delete_workstream_endpoint and open_workstream's
rehydrate path recorded the audit row under the stored ws.user_id
("owner_uid") rather than the authenticated caller. With row-level
ownership gating gone (a46dab1), any team member acting on a peer's
workstream produced an audit row naming the victim as the actor.
Fix: pass _auth_user_id(request) as the audit actor, matching the
pattern close_workstream already follows.

q-2 — SessionUIBase had no direct tests. The new
tests/test_session_ui_base.py covers listener fan-out, approval +
plan blocking gates, intent-verdict cache + FIFO eviction, verdict
persistence paths, output-guard persistence, the reset-between-rounds
invariant (bug-1 regression test), a cross-subclass test that
verifies BOTH WebUI.approve_tools and ConsoleCoordinatorUI.approve_tools
call _reset_approval_cycle (verified it fails without the fix), and
a concurrent enqueue/register smoke.

Full pytest: 4395 passed (+20 new). Ruff + mypy clean.

* fix: PR #408 review findings from copilot + code-quality

Three substantive fixes + mechanical side-effect-in-assert cleanup.

Copilot findings:

- session_ui_base.py: on_intent_verdict had a race with
  resolve_approval. Previously acquired _ws_lock twice (read decision
  → release → if unset, acquire again to append). resolve_approval
  could interleave between the two acquisitions, swap-and-clear the
  pending list and set the decision — our verdict then got appended
  to the fresh (empty) list and stamped with the NEXT round's
  decision on the following resolve. Fix: decision-check + append
  under ONE acquisition; storage UPDATE (if decision already set)
  runs outside the lock. New regression test counts lock
  acquisitions during on_intent_verdict and fails if the two-phase
  pattern returns.

- server.py close_workstream_endpoint: comment said "treat as
  already-closed success" but handler returned 404. Comment
  rewritten to match the 404 behaviour ("the ws isn't tracked here"
  is the only reachable meaning for close() → False now).

- test_session_ui_base.py concurrency smoke: the test ended with
  ``pytest.assume = lambda ...`` — a leftover that mutates pytest
  globals and can surprise other tests. Replaced with explicit
  ``not is_alive()`` assertions so the "threads completed cleanly"
  intent survives -O optimization stripping.

Code-quality (assert side-effects):

Six ``assert mgr.open(...)`` / ``assert mgr.close(...)`` in
test_session_manager.py stripped under ``python -O``. Mechanical
fix: extract to local before asserting.

Ignored the two "Protocol method body is `...`" flags — that's the
standard Protocol idiom; replacing with ``pass`` or
``NotImplementedError`` changes typing semantics.

Full pytest: 4396 passed.
2026-04-24 14:28:51 -07:00
renovate[bot] e7fd9e53b8 chore(deps): lock file maintenance (#407)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-04-23 21:59:35 -07:00
renovate[bot] 47cd1dbfeb chore(deps): update astral-sh/setup-uv action to v8 (#406)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-04-23 21:15:51 -07:00
renovate[bot] bf36461187 chore(deps): update helm release postgresql to ~18.6.0 (#405)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-04-23 21:15:38 -07:00
renovate[bot] 2ef4243024 chore(deps): update dependency vitest to v4.1.5 (#404)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-04-23 21:15:26 -07:00
Patrick Buckley 58d20f4012 chore(coord): remove spawn-quota subsystem (#403)
* chore(coord): remove spawn-quota subsystem

The quota gate was operator-level safety per its own comments, not a
security boundary, and never fired in a week of heavy use. Runaway
coordinator spawns are already bounded by max_active slot exhaustion,
which surfaces to the coord LLM as a tool error — same operational
shape, one fewer moving part. Precedes the Stage 1 SessionManager
unification so the coord tool doesn't inherit quota bookkeeping.

Upgraded deployments with the three removed settings persisted will
log three "Skipping invalid setting" warnings on startup and
otherwise degrade cleanly; a follow-up migration to delete the rows
would silence that noise.

* chore(migrations): drop stale coord spawn-quota settings rows (047)

Clears persisted rows for the three ConfigStore keys removed in the
previous commit so upgraded deployments don't log "Skipping invalid
setting" warnings on every startup. Downgrade is a no-op — the rows
were operator-set values, and a rollback to pre-1.5.0 code falls back
to the registry defaults for any key not present.
2026-04-23 20:09:12 -07:00
Patrick Buckley f5ec9cd2b7 fix(coord): render markdown on history reload (#402)
* fix(coord): render markdown on history reload

The coordinator chat's history-load path piped assistant content
through ``appendText`` → ``appendMsg(role, esc(text))``, which dumps
escaped raw text into the message body without ever calling the
markdown converter or the post-render hooks (highlight.js, mermaid,
KaTeX).  Live streaming uses ``streamingRender`` /
``streamingRenderFinalize`` which DO render markdown, so a fresh
stream looked correct but a page-reload / reconnect surfaced every
table, code fence, and math block as literal characters.

Now the history loop dispatches by role: assistant + reasoning go
through ``streamingRenderFinalize`` (mirrors what live streaming does
on stream_end); tool messages keep ``appendToolResult``; user / system
stay on ``appendText`` since they're typed verbatim and don't carry
markdown structure.

* fix(coord): keep reasoning role on plain-text path on history replay

The history loop routed reasoning role through streamingRenderFinalize,
but live streaming renders reasoning tokens via textContent
(appendReasoningToken).  History replay would render reasoning as
markdown while a fresh stream rendered it as plain text — inconsistent
look and unnecessary hljs / mermaid / KaTeX work on reasoning content.

Reasoning now uses appendText on replay, matching the live path.

Addresses Copilot review feedback on PR #402.
2026-04-23 19:18:09 -07:00
326 changed files with 94674 additions and 20390 deletions
+12 -2
View File
@@ -43,6 +43,13 @@ jobs:
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: ${{ matrix.python-version }}
# Node is required by tests/test_renderer_js.py — without
# explicit setup, that suite silently skips if the runner
# image happens not to ship Node, masking regressions in
# the browser-side renderer.
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6
with:
node-version: "24"
- run: pip install -e ".[test]"
- run: pytest tests/ -m "not live" --cov=turnstone --cov-report=term-missing --cov-report=xml -q
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
@@ -72,6 +79,9 @@ jobs:
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: "3.14"
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6
with:
node-version: "24"
- run: pip install -e ".[test,postgres]"
- run: pytest tests/ -m "not live" --storage-backend=postgresql -q
env:
@@ -128,7 +138,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6
- uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7
- uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # v8.1.0
with:
uv-version: "0.9.18"
- run: uv lock --check
@@ -137,7 +147,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6
- uses: astral-sh/setup-uv@37802adc94f370d6bfd71619e3f0bf239e1f3b78 # v7
- uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # v8.1.0
with:
uv-version: "0.9.18"
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
+947
View File
@@ -15,6 +15,953 @@ Three release tracks are maintained:
## [Unreleased]
### Removed (BREAKING — 1.5.0)
- **Legacy body-keyed and query-keyed URL family for the workstream
interaction verbs.** Pre-1.5 interactive shipped both a path-keyed
and a body-keyed surface for the same five verbs; this release drops
the body-keyed and query-keyed mounts (and the
``make_legacy_body_keyed_adapter`` /
``make_legacy_query_keyed_adapter`` shims that backed them). External
SDK consumers on stable 1.0/1.3/1.4 must move to the path-keyed
shape:
| Removed (1.0/1.3/1.4) | Use instead |
| ---------------------------------------------- | ------------------------------------------------------ |
| ``GET /v1/api/events?ws_id=X`` | ``GET /v1/api/workstreams/{ws_id}/events`` |
| ``POST /v1/api/send`` (body ``ws_id``) | ``POST /v1/api/workstreams/{ws_id}/send`` |
| ``DELETE /v1/api/send`` (body ``ws_id``) | ``DELETE /v1/api/workstreams/{ws_id}/send`` |
| ``POST /v1/api/approve`` (body ``ws_id``) | ``POST /v1/api/workstreams/{ws_id}/approve`` |
| ``POST /v1/api/cancel`` (body ``ws_id``) | ``POST /v1/api/workstreams/{ws_id}/cancel`` |
| ``POST /v1/api/workstreams/close`` (body) | ``POST /v1/api/workstreams/{ws_id}/close`` |
Calls to the old URLs return **404** on 1.5.0+. Bodies on the new
URLs no longer carry ``ws_id`` (the path provides it); the
``SendRequest`` / ``ApproveRequest`` / ``CancelRequest`` Pydantic
schemas drop the field, and ``CloseWorkstreamRequest`` slims to a
single optional ``reason`` field (the body is still required to be
valid JSON — send ``{}`` when omitting all fields).
``/v1/api/plan`` and ``/v1/api/command`` are unaffected and remain
body-keyed in this release. The bundled web UI, channel adapters,
Python SDK, TypeScript SDK, and console routing-proxy SDK ship the
new URLs automatically; pinning to ≥ 1.5.0 is enough.
The console routing proxy's ``/v1/api/route/...`` family is updated
alongside: ``/v1/api/route/workstreams/{ws_id}/<verb>`` replaces the
pre-1.5 ``/v1/api/route/{send,approve,cancel,workstreams/close}``
mounts. ``DELETE`` is now passed through (``client.request(method,
...)`` instead of ``client.post(...)``) so the new dequeue route
works through the proxy. Audit attribution for ``DELETE`` on
``/send`` is logged as ``route.workstream.dequeue`` rather than
``route.workstream.send``.
Auth scope wiring (``WRITE_PATHS`` / ``APPROVE_PATHS`` literals plus
the path-keyed verb match in ``required_scope``) updated to grant
``write`` for path-keyed ``send/cancel/close``, ``approve`` for
path-keyed ``approve``, and ``write`` for ``DELETE`` on
path-keyed ``/send``. The ``/node/*`` proxy branch mirrors all four.
### Changed
- **Dashboard row shape: ``id`` → ``ws_id``.** The
``GET /v1/api/dashboard`` row dict now keys the workstream
identifier as ``ws_id`` (matching the rest of the v1 workstream
surface — active list, saved list, history, detail). The Stage 2
list-verb lift converged ``/v1/api/workstreams`` and
``/v1/api/workstreams/saved`` on ``ws_id`` but left dashboard
alone to keep that PR's diff focused; this lands the same rename
on the remaining endpoint so the v1 row shape is consistent
across the family. Pydantic ``DashboardWorkstream`` and the
TypeScript SDK ``DashboardWorkstream`` interface both rename the
field accordingly. The bundled web UI is the only consumer that
reads ``dashboard.workstreams[].id`` and is updated atomically;
no external SDK on a stable line reads the field, so the swap is
bounded by normal static-asset reload. Console
``_fetch_live_block`` (cluster-inspect's projection over a
remote node's dashboard payload) is updated to match.
- **Coordinator gains rich `ws_state` payload + live activity broadcast**
([§ Post-P3 reckoning item #2 follow-up]). Pre-lift coord's
cluster broadcast was state-only — the dashboard's coord rows
showed the state column flipping but the ``tokens``,
``context_ratio``, ``activity``, and per-turn ``content`` fields
were all hardcoded to zero / empty. The lift turns
``on_status`` / ``on_content_token`` / ``on_thinking_start`` /
``on_thinking_stop`` / ``on_stream_end`` / ``on_tool_result``
into shared bodies on :class:`SessionUIBase` so coord populates
the same per-ws metric fields interactive does (the fields were
already declared on the base; only the writes were
WebUI-specific). ``coord_adapter.emit_state`` now reads the UI's
snapshot under ``_ws_lock`` via the new
:meth:`SessionUIBase.snapshot_and_consume_state_payload` helper
and passes the rich kwargs through to
``collector.emit_console_ws_state``; the cluster dashboard's
coord rows now render with the same tokens / activity / content /
context_ratio fields interactive rows do.
Three observable behaviour changes (all CHANGELOG-callout-worthy):
- **Coord persists ``usage_event`` storage rows.** Pre-lift only
WebUI did. The lifted ``on_status`` body unifies usage tracking
so governance dashboards / token-spend queries see coordinator
consumption alongside interactive. Operators querying
``usage_event`` by ``ws_id`` will see coord rows for the first
time.
- **Coord broadcasts live activity transitions.** New
``ClusterCollector.update_console_ws_activity(ws_id, *,
activity, activity_state)`` method (named ``update_*`` rather
than ``emit_*`` to flag the no-fan-out asymmetry vs. the rest
of the ``emit_console_ws_*`` family — it updates the in-memory
pseudo-node row but intentionally does NOT fan out a separate
SSE event). The cluster dashboard's per-ws polling reads the
in-memory pseudo-node row, so activity ticks land on the next
snapshot fetch (matches WebUI's behaviour where activity
events are observational; not fanned out through the cluster
SSE stream).
- **Cluster ``cluster_state`` events for coord rows now carry
non-zero ``tokens`` / ``content`` fields.** Frontend rendering
that conditionally hid these on coord rows can drop the
branch.
Architecture changes:
- ``_MAX_TURN_CONTENT_CHARS`` moved from ``turnstone.server`` to
``turnstone.core.session_ui_base`` so coord enforces the same
per-turn content cap interactive does.
- WebUI keeps ``on_status`` / ``on_tool_result`` / ``on_error``
overrides that layer Prometheus ``_metrics.record_*`` calls
(node-only) on top of the shared body via ``super()`` — the
Prometheus surface stays node-scoped (the console isn't a
node and has no /metrics endpoint).
- ``ConsoleCoordinatorUI`` adds a ``_broadcast_activity``
override that fans out via the cluster collector instead of
the global SSE queue (which is node-only on interactive).
- ``coord_endpoint_config`` wires a new ``_coord_spawn_metrics``
hook (mirrors interactive's) so the per-spawn ``_ws_messages``
increment + ``_ws_turn_tool_calls`` reset happen on coord too.
Test additions: 23 new tests in ``tests/test_coord_rich_ws_state_payload.py``
pin the per-ws metric writes (status, content accumulation,
activity tracking, tool-result counters, stream-end activity
clear), the snapshot helper's IDLE/ERROR drain semantics +
single-lock-acquisition guarantee, the adapter's rich-payload
pass-through + defensive None-UI handling, the activity
broadcast (collector wire + failure swallow + no-op-when-
collector-unset + dedup against last-emitted state), the
spawn_metrics hook, and a concurrent-writes-during-snapshot
stress case (cycles through running / idle / error so the
drain branches actually run against a concurrent writer).
Plus WebUI-override regression tests confirming
``_metrics.record_*`` still fires on top of the lifted bodies.
Existing ``tests/test_webui_content.py`` updated to import
``_MAX_TURN_CONTENT_CHARS`` from its new home in
``turnstone.core.session_ui_base``;
``tests/test_coordinator_adapter.py`` updated to expect the
rich-payload kwargs (``tokens=0`` defaults) on
``emit_console_ws_state``.
Two deferred follow-ups (out-of-scope for this lift,
flagged for tracking):
- **Synchronous ``record_usage_event`` INSERT on coord worker
thread.** The lifted ``on_status`` body persists usage rows on
every provider response — same shape WebUI uses, but coord
workers can fire multi-step plan/task agent loops where each
response blocks the worker for a write transaction. Parity
with WebUI is the explicit goal here; if coord throughput
becomes a concern, batch usage_event writes onto a background
flusher thread (one batch INSERT per N events / per K ms) on
both kinds.
- **Coord assistant turn content now flows on the cluster SSE
stream (``/v1/api/cluster/events``).** Pre-lift the broadcast
was ``content=""``; post-lift it carries the joined assistant
output. The cluster SSE stream has no per-user filter today —
extends an existing cross-tenant exposure (interactive
``cluster_state`` events already carry content) to a
previously-empty channel (coord rows). Proper fix needs the
SSE endpoint gated on ``admin.cluster.inspect`` (matching
``/v1/api/cluster/ws/{ws_id}/detail``) or per-listener
user_id filtering. Tracked as a separate security-tightening
project; not gating this lift since it inherits an existing
exposure rather than introducing a new mechanism.
- **`history` / `detail` verb bodies lifted across both kinds**
([Stage 2 Verb Lift — `history` / `detail`]). The coord
``GET /v1/api/workstreams/{ws_id}/history`` and
``GET /v1/api/workstreams/{ws_id}`` handlers now share two factory
bodies via ``make_history_handler(cfg)`` and
``make_detail_handler(cfg)``. The lift adds both endpoints to the
interactive surface as a feature gain (pre-lift only coord exposed
them; interactive consumers had to subscribe to ``/events`` SSE
just to read history rows or display fields). No new
``SessionEndpointConfig`` fields — the factories reuse
``permission_gate``, ``manager_lookup``, ``not_found_label``,
``audit_action_prefix``, and (for history's storage-fallback kind
check) ``list_kind`` — all already wired by both production
lifespans.
Three observable behaviour changes (all documented per kind):
- **Interactive gains ``GET /v1/api/workstreams/{ws_id}``.** Pre-lift
interactive had no detail endpoint — SDK consumers had to read
display fields from the SSE replay on ``/events`` or scrape the
active list. The lifted body lazy-rehydrates a closed/evicted
workstream via ``mgr.open()`` so the response shape is stable
across loaded / persisted-only states. Same
``{ws_id, name, state, user_id, kind}`` shape coord exposed
pre-lift, now available on both surfaces.
- **Interactive gains ``GET /v1/api/workstreams/{ws_id}/history``.**
Same ``?limit=`` query param contract as coord (default 100, max
500, malformed values fall back to 100, out-of-range clamps to
[1, 500]). Persisted-but-not-loaded interactives serve history
without rehydrating — the lifted body falls back to a storage-row
+ kind check (via ``cfg.list_kind``) when ``mgr.get`` returns
``None``, mirroring coord's pre-lift
``_resolve_coordinator_or_404`` ladder.
- **Storage / manager-lock work moved off the event loop on coord.**
The lifted ``history`` body always runs ``storage.get_workstream``
(storage-fallback path) and ``storage.load_messages`` through
``asyncio.to_thread``; pre-lift coord ran them inline on the
event loop. Long-tail message reads on a saturated console no
longer stall every other async handler for the duration of the
SQL.
Pydantic schemas: ``CoordinatorDetailResponse`` and
``CoordinatorHistoryResponse`` removed; both folded into
``WorkstreamDetailResponse`` / ``WorkstreamHistoryResponse`` on
the shared ``server_schemas.py`` (mirrors the list lift's pattern
for ``WorkstreamInfo``). Both server and console OpenAPI specs
reference the unified schemas; ``server_spec.py`` gains
``EndpointSpec`` entries for the new interactive endpoints. TS
SDK gains ``WorkstreamDetailResponse`` / ``WorkstreamHistoryResponse``
interfaces in ``sdk/typescript/src/types.ts``;
``openapi-{server,console}.json`` regenerated.
``GET /v1/api/workstreams/{ws_id}/history`` is the only verb
whose lifted body keeps a kind-aware storage fallback (via
``cfg.list_kind``); ``detail`` defers cross-kind isolation to
``mgr.open()`` itself.
- **`list` / `saved` verb bodies lifted across both kinds** ([Stage 2
Verb Lift — `list` / `saved`]). The interactive
``GET /v1/api/workstreams`` + ``GET /v1/api/workstreams/saved``
and coord ``GET /v1/api/workstreams`` + ``GET /v1/api/workstreams/saved``
handlers now share two factory bodies via
``make_list_handler(cfg)`` and ``make_saved_handler(cfg)``. Four
new ``SessionEndpointConfig`` fields capture the per-kind
divergence:
- ``list_resolve_titles: ListResolveTitles | None`` — interactive
wires :func:`turnstone.core.memory.get_workstream_display_names`
(new bulk helper added on the storage layer + ``memory.py``)
so the active-list endpoint resolves every user-set alias in
ONE ``SELECT ... WHERE ws_id IN (...)`` instead of the pre-lift
per-row N+1. Coord wires ``None`` (no alias surface today).
- ``list_kind: WorkstreamKind | None`` — required storage-side
kind classifier passed to ``list_workstreams_with_history``.
Interactive wires ``WorkstreamKind.INTERACTIVE``; coord wires
``WorkstreamKind.COORDINATOR``. Distinct from
``audit_action_prefix`` (audit-action namespacing) so adding a
third kind doesn't have to overload the audit prefix as a
classifier; missing value surfaces as 500 with a clear log
line rather than silently filtering for the wrong kind.
- ``saved_state_filter: str | None`` — coord wires ``"closed"``
so only explicitly-closed coordinators surface in the
saved-card grid. Interactive wires ``None`` (the storage
layer already excludes ``state='deleted'`` tombstones).
- ``saved_loaded_lookup: SavedLoadedLookup | None`` — coord-only
defence-in-depth filter that excludes ws_ids currently in the
in-memory pool (a row can be ``state='closed'`` for a few
seconds while the close-emit sequence races the in-memory pop).
Interactive wires ``None``.
Five observable behaviour changes (all documented per kind):
- **Active-list top-level key converges on ``"workstreams"``.**
Pre-lift coord returned ``{"coordinators": [...]}``; the lifted
body returns ``{"workstreams": [...]}`` for response-shape
parity with interactive. Coord is a 1.5.0aN-only surface — never
shipped stable — so SDK / frontend consumers swap once and
there's no compat shim or fallback (the convergence MUST land
before v1.5.0 stable per
``project_unification_before_stable.md``).
- **Saved-list top-level key converges on ``"workstreams"``.**
Same shape change as the active list, applied to
``GET /v1/api/workstreams/saved`` on coord. Coord-only surface;
no compat shim.
- **Active-list row key renames ``"id"`` → ``"ws_id"``** on
interactive. Pre-lift interactive used the bare ``id`` field
while every other shared verb on this surface (cancel, open,
events, create, saved-list) uses ``ws_id``. Convergence
eliminates the internal inconsistency. Frontend consumers
reading ``ws.id`` from the active-list response swap to
``ws.ws_id``. Interactive HAS shipped stable across 1.0 / 1.3 /
1.4, but the active-list endpoint is consumed by the bundled
JS only — there's no external SDK on those stable lines reading
the field. Browser-cache staleness is bounded by normal
static-asset reload on next page load.
- **Active-list row gains always-include fields.** ``user_id``
was coord-only; ``kind`` + ``parent_ws_id`` were
interactive-only. Both kinds now populate all three.
``parent_ws_id`` defaults to ``None`` for coord (coordinators
have no parent).
- **Storage / manager-lock work moved off the event loop on
interactive.** The lifted ``saved`` body always uses
``asyncio.to_thread`` for ``list_workstreams_with_history``;
pre-lift interactive ran it inline (correlated COUNT subquery
can stall every other async handler on a cluster with thousands
of saved rows). Coord already used ``to_thread`` (perf-2 from
the saved-coordinators review); convergence lifts interactive
up. The active-list body also moves ``mgr.list_all`` +
per-row title resolution off the event loop on both kinds.
Pydantic schemas: ``WorkstreamInfo.id`` renamed → ``ws_id``,
``WorkstreamInfo.user_id`` field added. ``CoordinatorInfo`` and
``CoordinatorListResponse`` removed (folded into the unified
``WorkstreamInfo`` / ``ListWorkstreamsResponse``); ``console_spec``
active-list endpoint now points at ``ListWorkstreamsResponse``.
OpenAPI spec snapshots regenerated.
``GET /v1/api/dashboard`` is **not** in the lift's scope and
still returns rows keyed on ``id``. A separate cleanup PR will
converge the dashboard row shape with the rest of the v1 surface.
- **`SessionManager.create` gains a deferred-emit option; lifted
``create`` HTTP handler eliminates the phantom create→close
pair on coord rollback.** ``SessionManager.create`` now accepts
``defer_emit_created: bool = False`` (default preserves the
legacy "advertise immediately" contract for direct callers); two
new methods complete the deferred-create bracket:
- ``SessionManager.commit_create(ws)`` fires the deferred
``emit_created`` event after the caller's post-create work
confirms the workstream should be advertised.
- ``SessionManager.discard(ws_id)`` releases the in-memory slot
+ cleans up the UI WITHOUT firing ``emit_closed`` — the
workstream's existence was never advertised, so there's
nothing to advertise on rollback. Storage-row deletion stays
a separate concern (caller invokes ``delete_workstream`` for
a complete rollback), mirroring ``mgr.create``'s split between
slot reservation and ``register_workstream``. Logs a
``warning`` (``session_mgr.discard.after_emit_created``) when
invoked on a workstream that's already been advertised
(non-deferred create or post-``commit_create``); the slot is
still released so capacity isn't stranded, but the warning
surfaces the caller-bug case where ``close`` would have been
the right call.
The lifted ``make_create_handler`` now uses this bracket: pass
``defer_emit_created=True``, validate uploaded attachments, then
``mgr.commit_create(ws)`` on success / ``mgr.discard(ws.id)`` on
failure. Pre-fix, coord's ``mgr.create`` fired ``emit_created``
synchronously — a rollback then called ``mgr.close`` which
fired ``emit_closed``, surfacing a quick create→close pair on
the cluster events stream that the collector's diff-reconcile
had to handle. Post-fix, a rejected upload produces zero
events. Interactive's ``emit_created`` is a documented no-op
stub so the deferral is observably a no-op there; the
``ws_created`` broadcast on the global SSE queue continues to
fire from the kind's post_install callback after attachment
validation passes (unchanged).
Direct callers of ``mgr.create`` (test fixtures, the CLI REPL,
channel adapters) keep the default ``defer_emit_created=False``
and see no behaviour change.
- **Coordinator HTTP surface unified under `/v1/api/workstreams/`**
([Stage 2 Priority 0]). The experimental `/v1/api/coordinator/*`
URL tree from 1.5.0aN is removed; coord verbs now mount at the
same shape as interactive workstreams via a shared route
registrar (`turnstone.core.session_routes`). Path mapping:
| Was (1.5.0aN) | Now |
|--------------------------------------------------|--------------------------------------------------|
| `POST /v1/api/coordinator/new` | `POST /v1/api/workstreams/new` |
| `GET /v1/api/coordinator` | `GET /v1/api/workstreams` |
| `GET /v1/api/coordinator/saved` | `GET /v1/api/workstreams/saved` |
| `GET /v1/api/coordinator/{ws_id}` | `GET /v1/api/workstreams/{ws_id}` |
| `POST /v1/api/coordinator/{ws_id}/{verb}` | `POST /v1/api/workstreams/{ws_id}/{verb}` |
Permission scopes, request / response bodies, and SSE event shapes
are unchanged. Callers on the experimental 1.5.0aN coord SDK must
swap their URL prefix; the legacy paths are gone with no compat
shim. Stable releases (1.0 / 1.3 / 1.4) never exposed
`/v1/api/coordinator/`, so this change is a no-op for anyone
upgrading from a stable line.
Two handler bodies (`approve`, `close`) lifted into the shared
registrar with kind branching behind `SessionEndpointConfig` —
both kinds share one implementation per verb. Two related
behavior changes on the interactive close path:
- `mgr.close()` race-loss returns 404 (was 500 on coord;
"popped between .get() and .close()" is a not-found semantic,
not a server error).
- Audit-write failures (`record_audit` raising on the storage
write) are now caught and logged at `warning` level; the close
still returns 200. Previously the interactive path let the
exception propagate as HTTP 500. Coord previously already
swallowed; convergence is intentional — operators monitor the
`ws.close.audit_failed` log line in both kinds the same way.
Other shared verbs (`send`, `cancel`, `open`, `events`, `create`,
`list`, `saved`, `history`, `detail`) keep their per-kind
handlers — body convergence for those requires SessionManager-
side refactors (e.g. Priority 1's worker-dispatch unification
for `send`) or coordinated frontend changes (response-shape
unification for `list` / `saved`) that fall outside Priority 0
scope.
- **TypeScript SDK bumped to 0.4.0** to flag the URL change for any
1.5.0aN-era consumer of the experimental coord client. The
`openapi-{server,console}.json` reference specs ship with the
unified path tree.
- **Worker dispatch unified across interactive + coordinator**
([Stage 2 Priority 1]). The atomic check-and-(spawn-or-queue)
decision for ``ChatSession.send`` now lives in
``turnstone.core.session_worker.send`` and is shared by both
paths. Interactive ``/v1/api/send``, the coordinator adapter, the
watch-result dispatch, the rewind/retry path, and the
initial-message-on-create path all gate on
``Workstream._worker_running`` (set/cleared atomically under
``ws._lock``) instead of ``Thread.is_alive()`` — closes a race
where two senders could spawn parallel workers on the same
ChatSession.
The ``/send`` HTTP body itself stays per-kind in this PR.
Verb-shape convergence (one shared factory body with capability
flags for attachments / queue priorities / metric increments) is
tracked as P1.5 and MUST land before 1.5.0 stable — letting the
fork ship into the stable line bakes the duplication in for the
lifetime of the 1.5 track.
- **`/send` body lift + coordinator attachments + queue surface
parity** ([Stage 2 Priority 1.5]). The ``/send`` HTTP handler is
now ONE factory body (``make_send_handler(cfg)``) wired with
capability flags on both kinds; the four attachment endpoints
(``upload`` / ``list`` / ``get_content`` / ``delete``) are also
unified via ``make_attachment_handlers(cfg)``. Coord workstreams
light up:
- ``POST/GET /v1/api/workstreams/{ws_id}/attachments``,
``GET .../attachments/{aid}/content``,
``DELETE .../attachments/{aid}`` — same shape, same caps, same
reservation flow as interactive.
- ``POST /v1/api/workstreams/{ws_id}/send`` accepts
``attachment_ids`` (or auto-consumes pending) and returns
``attached_ids`` / ``dropped_attachment_ids`` for surfacing
partial reservations. Live-worker reuse path also returns
``priority`` / ``msg_id`` (parity with the interactive
``status: queued`` shape).
Backend parity is end-to-end: storage layer was already
kind-agnostic; the route registrar's ``AttachmentHandlers`` slot
has been there since Stage 2 P0; the multi-node attachment
routing-proxy on the console (``route_attachment_proxy``) was
already shipping. P1.5 is the wiring + verb-shape lift that lets
these primitives surface on the coord side.
Coord dashboard rendering surfaces an attachment-count badge on
past messages with attachments; full chip rendering with
click-to-view is deferred (the coord dashboard is
diagnostic-leaning and chip parity isn't on the critical path
for the unification thesis). Python SDK adds
``coordinator_send`` / ``coordinator_upload_attachment`` /
``coordinator_list_attachments`` /
``coordinator_get_attachment_content`` /
``coordinator_delete_attachment`` on
``AsyncTurnstoneConsole`` + ``TurnstoneConsole``. TS SDK
regenerated; bumped to 0.5.0.
Three lifted helpers (``sniff_image_mime``,
``classify_text_attachment``, ``upload_lock``) moved from
``turnstone/server.py`` to ``turnstone/core/attachments.py`` so
both processes use the canonical implementation. The interactive
surface keeps the same behaviour; the helpers are simply
imported from their new home.
``coordinator_send`` no longer returns ``429`` on a full worker
queue — the unified body returns ``200 {"status": "queue_full"}``
for parity with interactive. Existing callers checking for ``429``
should switch to the status-code shape.
Coord ``GenerationCancelled`` now emits ``state=idle`` +
``stream_end`` (parity with interactive); pre-P1.5 a cancel-killed
coord worker would have terminated silently with no state event.
Cluster fanout / alerting keyed on ``state=error`` for cancelled
coord workers should switch to monitoring ``stream_end`` /
``state=idle`` together.
- **`SessionKindAdapter` Protocol split into construction +
emission** ([Stage 2 Priority 3]). The adapter Protocol now covers
only what every kind must implement (``kind`` / ``build_ui`` /
``build_session`` / ``cleanup_ui``); the four lifecycle emit
methods (``emit_created`` / ``emit_state`` / ``emit_rehydrated`` /
``emit_closed``) move to a separate ``SessionEventEmitter``
Protocol wired through a new optional
``event_emitter: SessionEventEmitter | None`` kwarg on
``SessionManager``. Both production adapters (interactive on
``server.py``, coordinator on ``console/server.py``) implement
both Protocols and are passed as both ``adapter`` and
``event_emitter`` at lifespan-construction time, so production
behavior is unchanged. The interactive adapter's three
``emit_created`` / ``emit_state`` / ``emit_rehydrated`` methods
remain documented no-op stubs (those events fire from out-of-band
paths — the create handler enqueues ``ws_created`` after
attachment validation, ``WebUI._broadcast_state`` emits
``ws_state``); ``emit_closed`` stays load-bearing as the sole
transport path for ``ws_closed`` onto the global SSE queue.
- **`cancel` verb body lifted across both kinds** ([Stage 2 Verb
Lift — `cancel`]). The interactive ``/v1/api/cancel`` and coord
``/v1/api/workstreams/{ws_id}/cancel`` handlers now share one
body via ``make_cancel_handler(cfg, *, audit_emit=None)``;
per-kind divergence captured by a new
``cancel_forensics: CancelForensics | None`` field on
``SessionEndpointConfig`` (interactive wires
``_capture_cancel_forensics``; coord wires ``None``).
Three observable behaviour changes for coord callers:
- **Coord cancel now accepts a ``force`` flag.** Same shape as
interactive: posting ``{"force": true}`` abandons the worker
thread and emits ``stream_end`` so a stuck coord generation
can be recovered without waiting for the daemon thread to
exit. Pre-lift coord ignored ``force``.
- **Coord cancel response always includes ``"dropped"``.**
Pre-lift coord returned bare ``{"status": "ok"}``; the lifted
body returns ``{"status": "ok", "dropped": {}}`` (always-include
parity with interactive). SDK consumers don't need to branch
on kind to read ``dropped``.
- **Coord cancel returns 400 when the workstream's session is
``None``.** Pre-lift coord called ``coord_mgr.cancel`` which
silently no-op'd on a placeholder/build-failed workstream; the
lifted body 400s with ``{"error": "No session"}`` for parity
with interactive's pre-existing branch.
Two observable changes for interactive (asymmetric — coord
pre-lift already had this behaviour):
- ``resolve_plan`` now runs on every cancel (previously gated
on ``was_running``). ``resolve_plan`` has an internal
``_pending_plan_review is None`` guard, so the call is no-op
when no plan review is pending. Lift gives interactive coord's
pre-lift recovery path: a stuck plan-pending state from a
crashed worker can be cleared via ``cancel`` instead of
requiring a workstream close + rehydrate.
- ``resolve_approval`` runs on every cancel **only when
``ui._pending_approval is not None``** (the lifted body gates
the call). ``resolve_approval`` is not idempotent — it always
broadcasts ``approval_resolved`` and overwrites
``_approval_result`` — so the gate prevents a stale resolution
event from leaking on idle cancels while preserving the recovery
path when an approval really is pending.
Coord ``coordinator.cancel`` audit detail now includes ``force``
so operator-driven recovery is distinguishable from a routine
cancel in the audit log.
Three /review fixes folded into the same commit:
- **No more stale ``approval_resolved`` SSE event on idle cancel.**
The lifted body's ``resolve_approval`` call is now gated on
``ui._pending_approval is not None``. Pre-fix, the unconditional
call would broadcast a phantom ``approval_resolved`` to every
SSE listener even when no prompt was pending — listener UIs
that key on the event would dismiss prompts they didn't have.
- **Force-cancel now clears ``_worker_running`` alongside
``worker_thread``.** Previously the force path left the half-
state ``(_worker_running=True, worker_thread=None)``, which
routed any follow-up ``send`` through the queue-enqueue path
onto the abandoned worker (where the cancel flag short-circuits
the queue-drain seam, leaving the message orphaned until the
next spawn). Restores the
``(worker_thread, _worker_running)`` invariant
``session_worker.send`` documents.
- **``coordinator_stop_cascade`` now treats child cancel
``400 + "No session"`` as ``skipped``** (was previously
``failed``). Lifted coord cancel returns 400 on placeholder /
build-failed children — matching the pre-lift outcome where
those children were silently no-op'd, so the cascade response's
``failed`` bucket no longer fires spurious operator alerts.
- **`open` verb body lifted across both kinds** ([Stage 2 Verb
Lift — `open`]). The interactive
``POST /v1/api/workstreams/{ws_id}/open`` and coord
``POST /v1/api/workstreams/{ws_id}/open`` handlers now share one
body via ``make_open_handler(cfg, *, audit_emit=None)``. Per-kind
divergence captured by two new ``SessionEndpointConfig`` fields:
- ``open_resolve_alias: AliasResolver | None`` — interactive
wires :func:`turnstone.core.memory.resolve_workstream` so
callers can pass user-friendly aliases ("my-debug-ws") in the
path param. Coord wires ``None`` (hex ids only).
- ``open_post_load: OpenPostLoad | None`` — interactive wires the
UI-replay (``clear_ui`` + history) + handler-side ``ws_created``
enqueue onto the global SSE queue. Coord wires ``None`` and
relies on the cluster collector fan-out from
``CoordinatorAdapter.emit_rehydrated``.
**Load-bearing fix** (§ Post-P3 reckoning item #3): interactive
``open_workstream`` previously called
``mgr.create(ws_id=resolved_id)`` + ``ws.session.resume(...)`` to
rehydrate, bypassing ``mgr.open()`` entirely. After the lift both
kinds route through ``mgr.open()`` — which makes
``InteractiveAdapter.emit_rehydrated`` reachable on interactive
(it had been dead-by-routing) and gives the manager a single
rehydrate code path to maintain. ``emit_rehydrated`` stays a
documented no-op stub on the interactive adapter (the
handler-side ``ws_created`` enqueue from ``open_post_load`` is
the load-bearing emission).
Two observable behaviour changes for interactive callers:
- **Cross-kind open returns 404** (was 400). Pre-lift had a
pre-mgr storage probe that returned ``400`` with
``"Workstream is not an interactive kind"`` for coord rows;
the lift consolidates on ``mgr.open()``'s single ``None``-
return contract for missing / wrong-kind / tombstoned rows.
Security boundary unchanged.
- **Already-loaded response uses ``ws.name`` directly** (was
``get_workstream_display_name(resolved_id) or resolved_id``).
A workstream renamed via ``set_workstream_alias`` after being
loaded into memory will surface the storage-row name in the
open response's ``name`` field instead of the latest alias.
The dashboard listing endpoint still resolves aliases on its
own pass, so the user-visible workstream name in the tab strip
isn't affected.
Coord behaviour unchanged.
Two /review fixes folded into the same commit:
- **Resume failures now return 5xx instead of broken-200.**
``SessionManager.open()`` previously caught and ``log.debug``-
swallowed exceptions from ``ChatSession.resume`` (which assigns
``self.messages`` *before* the config-restore block, so a
partial-failure resume — corrupted ``workstream_config`` row,
model-registry mismatch on a saved alias, malformed
``temperature`` / ``max_tokens`` — would leave the session with
history but with default config). Pre-lift, the interactive
open handler called ``ws.session.resume(...)`` directly and let
exceptions propagate as 500. The lift accidentally inherited
the swallow because it routed through ``mgr.open()``. Restored
pre-lift behaviour: ``mgr.open()`` now re-raises resume
exceptions after rolling back the slot (``cleanup_ui`` +
``_remove_locked``), so the lifted handler returns 500 with
a correlation id and the storage row stays available for a
retry instead of silently 200'ing with broken state.
- **``except Exception`` in the lifted body documents intent.**
The bare exception catch around ``mgr.open(ws_id)`` is
intentional — the kind's session factory has no documented
exception spec, and resume can propagate from
``ChatSession.resume``. A one-line rationale comment in the
handler body keeps a future contributor from narrowing it
incorrectly.
- **`events` verb body lifted across both kinds** ([Stage 2 Verb
Lift — `events`]). The interactive
``GET /v1/api/events?ws_id=...`` and coord
``GET /v1/api/workstreams/{ws_id}/events`` SSE handlers now
share one body via ``make_events_handler(cfg)``. Per-kind
divergence captured by a new
``events_replay: EventsReplay | None`` cfg field — a Protocol-
typed callback yielding the kind-specific initial replay
payload that the lifted body iterates and sends as ``data:``
lines before starting the live event loop. Interactive's
``_interactive_events_replay`` yields the pre-lift sequence
(``connected`` + ``status`` + ``history`` + ``pending_approval``
+ cached intent verdicts + ``pending_plan_review``); coord's
``_coord_events_replay`` yields just ``pending_approval`` +
``pending_plan_review`` (matches pre-lift coord behaviour).
The legacy interactive query-keyed URL is preserved via a new
``make_legacy_query_keyed_adapter`` helper (sister to
``make_legacy_body_keyed_adapter`` from earlier lifts) — it
reads ``ws_id`` from the query string and splices into
``request.path_params`` before delegating to the lifted body.
``GET /v1/api/events?ws_id=...`` continues to work for any 1.x
SDK consumer.
Two convergence wins:
- **Coord gains SSE connect/disconnect metrics.** Pre-lift
coord didn't record per-stream metrics; the lifted body
always calls ``metrics.record_sse_connect()`` /
``...disconnect()``, giving the cluster dashboard the same
per-stream observability interactive's had since 1.0.
- **Both kinds now check ``request.is_disconnected()`` AND
the ``ws_closed`` event** to terminate. Pre-lift interactive
relied solely on ``ws_closed`` (which never fires if the
client just goes away without closing the workstream);
pre-lift coord relied solely on ``is_disconnected``. The
lifted body uses both — whichever fires first wins.
One observable shape change for coord callers: the lifted body
returns 409 ``"session has no UI"`` when ``ws.ui`` is missing
(placeholder / build-failed UI), matching pre-lift coord.
Pre-lift interactive returned 404 in this case; the lift
converges on 409 across kinds because the workstream EXISTS
(404 would imply it doesn't).
**Item #2 from § Post-P3 reckoning split out** of this lift
during scoping (rich ``ws_state`` payload parity for coord —
lifting coord's ``ConsoleCoordinatorUI`` to broadcast
``tokens + context_ratio + activity + content`` like
``WebUI._broadcast_state`` does). The body lift touches
``session_routes.py`` + ``server.py`` + ``console/server.py``;
the rich-payload work touches ``coordinator_ui.py`` +
``collector.py`` + ``session_ui_base.py`` (different files,
different reviewer concern). Tracked as standalone follow-up
``feat/coord-rich-ws-state-payload``.
Two /review fixes folded into the same commit:
- **Restored interactive's dedicated SSE thread pool.** The
initial draft of ``make_events_handler`` used
``asyncio.to_thread`` (default executor, capped at
``min(32, cpu_count + 4)``) for the per-connection
``client_queue.get`` blocking wait. Pre-lift interactive used
a dedicated 200-thread ``sse_executor`` (created in the
lifespan with ``thread_name_prefix="sse"``) precisely to
avoid this — under high concurrent SSE counts the default
pool starves and SSE polling contends with every other
``asyncio.to_thread`` caller in the process (storage, router,
audit). Restored isolation via a new
``sse_executor_lookup: SseExecutorLookup | None`` cfg field;
interactive returns ``request.app.state.sse_executor``, coord
wires ``None`` and falls through to the default executor.
- **Restored 5s queue.get poll** (was shortened to 1s in the
initial draft). The 5x wakeup-rate bump compounded the thread-
pool starvation; the ``request.is_disconnected()`` probe
between polls already covers cancel-detection latency the
timeout would otherwise gate.
- **Replay phase streams events directly from the generator
instead of pre-building into a list.** The initial draft
materialised the entire kind-specific replay payload
(``connected`` + ``status`` + ``history`` + pending prompts)
into a list before constructing the ``EventSourceResponse``,
delaying time-to-first-byte until the heaviest replay event
(``_build_history`` for long-running interactive workstreams)
finished serialising AND letting the per-UI listener queue
accumulate over its 500-slot cap on a chatty mid-generation
workstream. The lifted body now iterates ``cfg.events_replay``
inside the async generator so each event ships as soon as the
callback yields it; the existing observational-failure swallow
semantics are preserved by wrapping the iteration in the same
try/except.
- **`create` verb body lifted across both kinds** ([Stage 2 Verb
Lift — `create`]). The interactive
``POST /v1/api/workstreams/new`` and coord
``POST /v1/api/workstreams/new`` handlers now share one body via
``make_create_handler(cfg, *, audit_emit=None)``. Per-kind
divergence captured by five new ``SessionEndpointConfig`` fields:
- ``create_supports_attachments: bool`` — multipart body parsing
+ attachment validation+save+rollback. Both kinds wire ``True``.
- ``create_supports_user_id_override: bool`` — trusted-source
body ``user_id`` override (interactive ``True`` for console-
proxied creates; coord ``False``).
- ``create_validate_request: CreateRequestValidator | None`` —
per-kind pre-create gates (interactive: ws_id format, kind,
parent ownership, attachments+resume_ws combo; coord: 401-on-
empty-uid).
- ``create_build_kwargs: CreateKwargsBuilder | None`` — per-kind
kwargs dict for ``mgr.create``.
- ``create_post_install: CreatePostInstall | None`` — per-kind
tail end (interactive: WebUI auto_approve + watch_runner +
``ws_created`` global broadcast + atomic resume + skill session
config + notify_targets + routing override + initial-message
worker thread; coord: ``coord_adapter.send`` for the optional
initial_message).
The pure helper ``_validate_and_save_uploaded_files`` lifted from
``turnstone.server`` to ``turnstone.core.attachments`` as
``validate_and_save_uploaded_files`` so both processes can call
the same kind-agnostic implementation.
**§ Post-P3 reckoning item #1 done — coord gains create-time
attachments.** Pre-lift ``coordinator_create`` accepted JSON only
and ignored uploads; the lifted body parses ``multipart/form-data``
on coord and saves attachments through the kind-agnostic storage
layer. ``CoordinatorAdapter.send`` gained optional
``attachments`` + ``send_id`` kwargs so when a create request
carries both ``initial_message`` and uploads, the attachments
are reserved onto the dispatched first turn — the worker's
``ChatSession.send(..., send_id=...)`` consumes them on dequeue
exactly the way interactive's create-with-attachments worker
thread does. The ``send_id`` reservation token soft-locks the
rows, and the adapter's failure path unreserves so a worker
crash returns them to pending. The pure helper
``_reserve_and_resolve_attachments`` lifted from ``server.py``
to ``turnstone.core.attachments`` as
``reserve_and_resolve_attachments`` so both kinds call one
kind-agnostic implementation.
Note on broadcast timing: coord's ``mgr.create`` fires
``emit_created`` (cluster collector fan-out) BEFORE the lifted
body runs attachment validation. If validation fails on coord and
the rollback (``mgr.close`` → ``emit_closed``) fires, the cluster
events stream sees a phantom create→close pair. Cluster consumers
handle this gracefully (same shape as any quick-create-close);
decoupling ``emit_created`` from ``mgr.create`` would be a bigger
refactor that doesn't belong in the verb lift. Interactive's
broadcast (``gq.put_nowait("ws_created")``) is held until after
attachment validation by the post-install callback, so interactive
never sees the phantom pair.
Five observable behaviour changes on the create response:
- **Both kinds converge on 200 OK.** Pre-lift interactive
returned 200 (default JSONResponse status); pre-lift coord
returned 201. Picked 200 over 201 for response-shape parity
with every other shared verb at the cost of REST-strict
correctness — a one-time release note rather than ongoing
client churn (the rest of the v1 SDK already uses
``response.ok`` per ``feedback_test_frontend_locally.md``).
SDK consumers that branched on ``status == 201`` for coord
must switch to ``response.ok``.
- **Always-include response shape.** Pre-lift interactive
returned ``{ws_id, name, resumed, message_count, attachment_ids}``
(5 fields); pre-lift coord returned ``{ws_id, name}`` (2). The
lifted body always returns the full shape, with ``resumed=False``
/ ``message_count=0`` / ``attachment_ids=[]`` on kinds whose
post-install doesn't populate them. Coord callers will see the
parity fields appear with default values.
- **Both kinds converge on the manager-at-capacity 429
semantic.** Pre-lift interactive translated ``mgr.create``'s
``RuntimeError`` to 400; coord already translated to 429. The
documented contract on ``SessionManager.create`` is "raises
RuntimeError when the manager is at capacity" — 429 (rate-
limit / try-later) is the correct shape.
- **Both kinds converge on the factory-misconfig 503 semantic.**
Pre-lift interactive let ``ValueError`` propagate as 500 with
a stack trace; coord already translated to 503 with the
factory's remediation text. Operators get the actionable
message instead of the trace.
- **Both kinds get a correlation_id'd 500 on unexpected
``mgr.create`` failure.** Pre-lift interactive let unexpected
exceptions propagate as 500 with a stack trace (potential
information leak via frame names / file paths); coord already
returned a correlation_id'd 500 with the message redacted. The
lifted body adopts coord's safer pattern on both kinds.
Two coord-specific parity gains:
- **Coord rejects disabled skills.** Pre-lift
``coordinator_create`` silently allowed disabled skills to
flow through to ``mgr.create`` — the row would create with a
skill the operator had marked inert, surprising both the
operator and the next user. The lifted body returns 400
"Skill not found or disabled" matching interactive's
behaviour.
- **Coord audit-emit failures no longer 500.** Pre-lift
``coordinator_create`` already swallowed; pre-lift interactive
let the failure propagate as 500. The lifted body wraps
``audit_emit`` in try/except + ``warning`` log, returning the
successful 200 to the caller. Mirrors the close / cancel /
open / events lift contracts.
No legacy adapter is needed for create — both kinds already
mounted ``POST {prefix}/new`` pre-lift; the lifted handler slots
in at the same path on each kind.
Three /review fixes folded into the same commit:
- **Pre-lift's 400 on malformed ``notify_targets`` preserved.** The
initial draft surfaced ``notify_targets`` validation errors from
inside the interactive ``post_install`` callback, which the
factory had no return-the-400 channel for — the only signal was
to ``raise``, which the factory's generic exception handler
turned into a redacted 500. Worse, by the time ``post_install``
ran the workstream was fully built (audit row written,
``ws_created`` broadcast emitted), so a malformed-input request
surfaced as "create failed" with the workstream actually live.
Fixed by moving the ``notify_targets`` validation into
:func:`_interactive_create_validate_request` (the pre-create
gate), which returns the 400 before ``mgr.create`` runs and
keeps storage clean. New regression test:
``test_create_lift_400s_on_malformed_notify_targets``.
- **Skill-lookup storage failure now correlation_id'd.** The
initial draft swallowed ``get_skill_by_name`` exceptions into
``skill_data = None`` and returned a 400 "Skill not found or
disabled" — masking storage outages as user-input misses and
making operator triage of skill-related reports impossible. The
lifted body now lets the storage exception propagate to the
same correlation_id'd 500 path that ``mgr.create`` failures
use; the skill-lookup + version count + ``mgr.create`` all live
inside one ``try / except`` so storage outages anywhere in the
create-prelude get the redacted-message-with-correlation-id
treatment instead of a stack-traced 500 leak.
- **Whitespace-only ``skill`` field treated as empty.** The
initial draft took ``body.get("skill") or ""`` literally — a
payload with ``"skill": " "`` would have hit
``get_skill_by_name(" ")`` and 400'd as "Skill not found".
Pre-lift coord stripped via ``(body.get("skill") or "").strip()
or None``; the lifted body now strips for both kinds (interactive
never received whitespace-only skills from the web UI but the
convergence is the safer default).
- **Canonical skill name persisted to ``mgr.create``.** The initial
draft's ``_interactive_create_build_kwargs`` /
``_coord_create_build_kwargs`` passed the raw ``body["skill"]``
through, so a whitespace-padded request would have persisted
``" my-skill "`` even though the lookup was done on the
stripped name. The build_kwargs callbacks now thread
``skill_data["name"]`` (the resolved row's canonical name) so
the persisted ``Workstream.skill`` matches the row that was
actually applied — keeps later session-side ``skill`` lookups
working regardless of how dirty the inbound payload was.
- **Coordinator scratchpad tool renamed: ``task_list`` → ``tasks``.**
The tool name on the LLM-facing schema, the audit event name
(``task_list.update`` → ``tasks.update``), the SSE
``tool_result`` event name (the coord-tree UI keys
``ev.name === "tasks"`` for /tasks-refetch debounce), and the
log tag (``task_list.corrupt_envelope`` → ``tasks.corrupt_envelope``)
all switch together. Operators with audit dashboards / SIEM filters
/ log greps that pinned the old prefix should update; the rename
is observable on the wire, not just internal. Internal Python
surface follows: ``CoordinatorClient.task_list_*`` → ``tasks_*``,
``ChatSession._prepare_task_list`` / ``_exec_task_list`` →
``_prepare_tasks`` / ``_exec_tasks``, ``_TASK_LIST_MAX`` →
``_TASKS_MAX``. The previous name compounded the bare word
``task`` (which collides with chat-template channels on local
models — the same reason ``task_agent`` carries the suffix); the
plural form sidesteps the collision and is more accurate, since
the tool acts on the whole list rather than a single task.
### Security
- **Coord attachment endpoints are now kind-strict**
([Stage 2 P1.5]). The coord ``attachment_owner_resolver``
resolves through the in-memory ``coord_mgr`` only — it does NOT
fall back to storage. Without this, an
``admin.coordinator``-scoped caller could pass an *interactive*
workstream ws_id to the new coord attachment endpoints; the
generic ``get_workstream_owner`` storage call (kind-agnostic)
would resolve cleanly and grant cross-kind read / write access
to interactive attachments. The kind-strict resolver returns
404 for any ws_id not currently held by the coord manager,
closing the cross-kind path. Persisted-but-not-loaded
coordinators must be ``open``ed before their attachment endpoints
respond. Caught by /review pre-merge; no exploit observed.
- **Workstream state writes are now buffered through ``StateWriter``.**
``SessionManager.set_state`` no longer holds ``ws._lock`` across a
synchronous Postgres ``UPDATE`` for non-terminal transitions;
instead a ``StateWriter`` (constructed at app startup, started /
shutdown by the lifespan) coalesces transient transitions per
ws_id and flushes every ~1s. **Observable behavior change**:
transient state (``thinking`` / ``running`` / ``idle`` /
``attention``) shows up in storage up to ~1s late; SSE consumers
see it immediately via the adapter's ``emit_state``. Terminal
``ERROR`` transitions and ``close()`` write synchronously and
remain durable on return. The bug-3 invariant — a closed row
can't be resurrected by a buffered transient — is preserved by
``close()`` calling ``state_writer.discard(ws_id)`` (drops
pending + waits for any in-flight flush) before its sync
``state='closed'`` write.
## [1.4.0]
User-visible additions: a full attachment system (images + text documents,
+6 -3
View File
@@ -8,14 +8,17 @@ FROM python:3.14-slim
LABEL org.opencontainers.image.title="turnstone" \
org.opencontainers.image.description="Multi-node AI orchestration platform"
COPY --from=ghcr.io/astral-sh/uv:0.11.7 /uv /usr/local/bin/uv
COPY --from=ghcr.io/astral-sh/uv:0.11.8 /uv /usr/local/bin/uv
# Remove the slim image's man page exclusion so man-db has actual content
RUN rm -f /etc/dpkg/dpkg.cfg.d/docker
# System dependencies: psycopg (libpq5), developer tooling for agent workflows
# System dependencies: psycopg (libpq5), developer tooling for agent workflows.
# ripgrep is the preferred backend for the search tool — natively bounds
# per-line, per-file, and per-filesize so pathological inputs (minified
# bundles, training-data JSONL with multi-MB single records) can't OOM us.
RUN apt-get update && apt-get upgrade -y && apt-get install -y --no-install-recommends \
libpq5 git curl jq man-db manpages procps file \
libpq5 git curl jq man-db manpages procps file ripgrep \
&& rm -rf /var/lib/apt/lists/*
# Node.js LTS (for npx-based MCP servers like @modelcontextprotocol/server-github)
+1 -1
View File
@@ -8,7 +8,7 @@
Multi-node AI orchestration platform. Deploy tool-using AI agents across a cluster of servers with direct HTTP routing, interactive interfaces, and enterprise governance.
<p align="center">
<img src="docs/assets/hero.png" alt="Turnstone console — multi-workstream AI orchestration with mermaid diagrams" width="960"/>
<img src="docs/assets/hero.png" alt="Turnstone coordinator — parallel tool batches with judge-graded approval and child workstream tracking" width="960"/>
</p>
Named after the [Ruddy Turnstone](https://en.wikipedia.org/wiki/Ruddy_turnstone) (*Arenaria interpres*) — a shorebird that flips stones to discover what's hiding underneath.
+3 -3
View File
@@ -57,8 +57,8 @@ services:
deploy:
resources:
limits:
memory: 1G
cpus: '1.0'
memory: 4G
cpus: '4.0'
restart: unless-stopped
# -------------------------------------------------------------------
@@ -231,7 +231,7 @@ services:
start_period: 60s
deploy:
resources:
limits: { memory: 384M, cpus: '0.5' }
limits: { memory: 4G, cpus: '4' }
restart: unless-stopped
server-2:
+1 -1
View File
@@ -7,6 +7,6 @@ appVersion: "0.3.0"
dependencies:
- name: postgresql
version: ~18.5.0
version: ~18.6.0
repository: https://charts.bitnami.com/bitnami
condition: postgresql.enabled
+101 -29
View File
@@ -229,12 +229,12 @@ below.
---
### `GET /v1/api/events?ws_id=<id>`
### `GET /v1/api/workstreams/{ws_id}/events`
Opens a Server-Sent Events stream scoped to a single workstream. The connection
remains open indefinitely; the server pushes events as they occur.
**Query parameters:**
**Path parameters:**
| Parameter | Type | Required | Description |
|-----------|--------|----------|----------------------------|
@@ -281,6 +281,7 @@ Each message in the `messages` array has:
| `role` | string | `"user"`, `"assistant"`, or `"tool"` |
| `content` | string or null | Text content of the message |
| `tool_calls` | array or null | Present only on assistant messages with calls |
| `reasoning` | string (optional) | Concatenated reasoning / chain-of-thought text on assistant turns whose `provider_data` carried reasoning-bearing blocks (Anthropic `thinking`, OpenAI Responses `reasoning`, or synthetic `reasoning_text` from local-model servers). Present only when the active model's `surface_persisted_reasoning` flag is True. |
Each entry in `tool_calls`:
@@ -325,6 +326,44 @@ finalize any in-progress assistant message.
{"type": "stream_end"}
```
**`state_change`** -- the worker thread transitioned to a new state. Drives
the client's busy-mode (composer in send vs. stop, spinner indicators,
auto-focus on idle). Sent live during normal operation AND on every fresh
SSE subscribe (so a mid-stream page refresh restores the correct composer
state without waiting for the next live transition).
```json
{"type": "state_change", "state": "running"}
```
| Field | Type | Description |
|----------|--------|----------------------------------------------------------------------|
| `state` | string | One of `"running"`, `"thinking"`, `"attention"`, `"idle"`, `"error"` |
**`in_progress_snapshot`** -- one-shot replay of the in-progress turn's
content + reasoning text-so-far when this client connects mid-stream.
Lets a refreshing browser tab restore partial assistant text immediately
instead of waiting for the response to complete. Yielded once after the
kind-specific replay phase (history + pending), only when at least one
of `content` / `reasoning` is non-empty. Both halves render into the same
assistant bubble the live `content` / `reasoning` events would target;
clients should treat the snapshot as idempotent (skip overwrite if the
current local buffer is already a superset prefix — covers EventSource
auto-reconnect re-replays).
```json
{
"type": "in_progress_snapshot",
"content": "Here is the answer so far: it depends on ",
"reasoning": "The user is asking about a comparison; let me think about..."
}
```
| Field | Type | Description |
|--------------|--------|------------------------------------------------------------|
| `content` | string | Joined assistant content text accumulated this turn |
| `reasoning` | string | Joined reasoning / chain-of-thought text accumulated |
**`tool_info`** -- one or more tool calls that were auto-approved (no user
action required).
@@ -346,7 +385,7 @@ action required).
```
**`approve_request`** -- one or more tool calls that require user approval. The
client must respond via `POST /v1/api/approve`.
client must respond via `POST /v1/api/workstreams/{ws_id}/approve`.
```json
{
@@ -450,7 +489,7 @@ after `/clear` or `/new` commands).
```
**`cancelled`** -- a cancel request was acknowledged (via the Stop button or
`POST /v1/api/cancel`). This signals that cancellation is in progress, not
`POST /v1/api/workstreams/{ws_id}/cancel`). This signals that cancellation is in progress, not
that it is complete. The worker thread may still be finishing — wait for
`stream_end` before transitioning to a ready state. The client should clear
any in-progress assistant rendering but not re-enable the send button until
@@ -522,7 +561,13 @@ Each SSE connection to a workstream receives its own delivery queue. Events
produced by the worker thread are fanned out to all registered listener queues,
so multiple consumers (browser, console proxy, SDK) can connect
simultaneously and each receives every event. On reconnect the client receives
a full history replay, so no catch-up mechanism is needed.
the kind-specific replay (`connected` + `status` + `history` + pending
approval / plan for interactive; `connected` + `status` + pending for coord)
followed by a `state_change` carrying the current worker state and an
optional `in_progress_snapshot` carrying any partial content / reasoning
buffered for the in-progress turn — so a mid-stream refresh restores both
the busy-mode UI and the partial assistant text without waiting for the
response to complete.
---
@@ -558,7 +603,7 @@ Possible `state` values:
and copies each event to every client queue. If a client queue is full, the
event is silently dropped for that client.
**Keepalive:** Same as `/v1/api/events` -- an SSE comment every 5 seconds.
**Keepalive:** Same as `/v1/api/workstreams/{ws_id}/events` -- an SSE comment every 5 seconds.
---
@@ -571,8 +616,8 @@ Returns a list of all active workstreams.
```json
{
"workstreams": [
{"id": "abc123", "name": "default", "state": "idle"},
{"id": "def456", "name": "hacker-news", "state": "thinking"}
{"ws_id": "abc123", "name": "default", "state": "idle"},
{"ws_id": "def456", "name": "hacker-news", "state": "thinking"}
]
}
```
@@ -581,7 +626,7 @@ Each workstream object:
| Field | Type | Description |
|--------------|-------------|--------------------------------------------------------|
| `id` | string | Unique workstream routing identifier |
| `ws_id` | string | Unique workstream routing identifier |
| `name` | string | Display name (alias if set, otherwise `ws-xxxx`) |
| `state` | string | Current state (see state values above) |
@@ -654,21 +699,26 @@ Each skill summary:
---
### `POST /v1/api/send`
### `POST /v1/api/workstreams/{ws_id}/send`
Sends a user message to a workstream. Spawns a daemon worker thread that calls
`session.send()` and streams results back via the SSE channel.
**Path parameters:**
| Parameter | Type | Required | Description |
|-----------|--------|----------|----------------------|
| `ws_id` | string | yes | Target workstream ID |
**Request body:**
```json
{"message": "Explain how the server works", "ws_id": "abc123"}
{"message": "Explain how the server works"}
```
| Field | Type | Required | Description |
|-----------|--------|----------|-------------------------|
| `message` | string | yes | The user's message text |
| `ws_id` | string | yes | Target workstream ID |
**Response (success):**
@@ -692,15 +742,21 @@ from a previous request. Also pushes a `busy_error` event to the SSE stream.
---
### `POST /v1/api/approve`
### `POST /v1/api/workstreams/{ws_id}/approve`
Responds to a tool approval request. The SSE stream must have previously sent
an `approve_request` event for the given workstream.
**Path parameters:**
| Parameter | Type | Required | Description |
|-----------|--------|----------|----------------------|
| `ws_id` | string | yes | Target workstream ID |
**Request body:**
```json
{"approved": true, "feedback": null, "always": false, "ws_id": "abc123"}
{"approved": true, "feedback": null, "always": false}
```
| Field | Type | Required | Description |
@@ -708,7 +764,6 @@ an `approve_request` event for the given workstream.
| `approved` | bool | yes | `true` to approve, `false` to deny |
| `feedback` | string/null | no | Optional feedback text (sent as denial reason) |
| `always` | bool | no | If `true` and `approved`, enables auto-approve |
| `ws_id` | string | yes | Target workstream ID |
When `always` is `true` and `approved` is `true`, the workstream's WebUI
instance sets `auto_approve = True`, causing all subsequent tool calls to be
@@ -789,7 +844,7 @@ containing the resumed session's messages.
---
### `POST /v1/api/cancel`
### `POST /v1/api/workstreams/{ws_id}/cancel`
Cancels the active generation in a workstream. Sets a cooperative cancellation
flag that is checked at multiple points in the generation loop (per streaming
@@ -812,15 +867,20 @@ for the orphaned thread. Use force cancel when cooperative cancel has not
resolved within a few seconds — the web UI offers this as a "Force Stop"
button automatically.
**Path parameters:**
| Parameter | Type | Required | Description |
|-----------|--------|----------|----------------------|
| `ws_id` | string | yes | Target workstream ID |
**Request body:**
```json
{"ws_id": "abc123", "force": false}
{"force": false}
```
| Field | Type | Required | Description |
|--------|--------|----------|----------------------|
| `ws_id`| string | yes | Target workstream ID |
| `force`| bool | no | Abandon stuck worker immediately (default: `false`) |
**Response:**
@@ -893,20 +953,32 @@ Status code: `400`
---
### `POST /v1/api/workstreams/close`
### `POST /v1/api/workstreams/{ws_id}/close`
Closes and removes a workstream. The last remaining workstream cannot be
closed.
**Path parameters:**
| Parameter | Type | Required | Description |
|-----------|--------|----------|------------------------|
| `ws_id` | string | yes | Workstream ID to close |
**Request body:**
```json
{"ws_id": "abc123"}
```
The body must be valid JSON. If you are not supplying any optional
fields, send `{}` — an empty / non-JSON body is rejected with a
`400`.
| Field | Type | Required | Description |
|---------|--------|----------|---------------------------|
| `ws_id` | string | yes | Workstream ID to close |
| Field | Type | Required | Description |
|----------|--------|----------|----------------------------------------------------------|
| `reason` | string | no | Optional close reason persisted to `workstream_config`. |
The `reason` is capped at **512 UTF-8 bytes** (multibyte-safe — the
cap holds for CJK and emoji payloads), and the output guard's
credential-redaction pass strips secrets before the value is
persisted. A non-string `reason` is silently coerced to empty and
the close proceeds without writing the field.
**Response (success):**
@@ -937,7 +1009,7 @@ turn on this workstream.
The attachment moves through three states: `pending → reserved →
consumed`. Reservation tokens are threaded through
`POST /v1/api/send` so a queued multimodal turn cannot lose its file to
`POST /v1/api/workstreams/{ws_id}/send` so a queued multimodal turn cannot lose its file to
an overlapping send.
Ownership failures are masked as `404` so non-owners cannot enumerate
@@ -1952,7 +2024,7 @@ Status code: `200` with an empty body.
| Malformed or unparseable JSON body | Treated as an empty dict `{}`; missing fields use defaults |
| Unknown `ws_id` | `404` with `{"error": "Unknown workstream"}` |
| Unknown path (GET or POST) | `404` with plain-text body `Not found` |
| Empty `message` on `/v1/api/send` | `400` with `{"error": "Empty message"}` |
| Empty `message` on `/v1/api/workstreams/{ws_id}/send` | `400` with `{"error": "Empty message"}` |
| Empty `command` on `/v1/api/command` | `400` with `{"error": "Empty command"}` |
| Rate limit exceeded | `429` with `Retry-After` header (see below) |
@@ -1995,7 +2067,7 @@ reconnection:
On reconnect, the server replays the full conversation history via the
`history` event, so the client can rebuild its UI state without data loss. The
same reconnection strategy applies to both the per-workstream SSE stream
(`/v1/api/events`) and the global state stream (`/v1/api/events/global`).
(`/v1/api/workstreams/{ws_id}/events`) and the global state stream (`/v1/api/events/global`).
---
@@ -2105,7 +2177,7 @@ turnstone_workstreams_active_total 1
# TYPE turnstone_http_requests_total counter
turnstone_http_requests_total{method="GET",endpoint="/health",status_code="200"} 42
turnstone_http_requests_total{method="GET",endpoint="/metrics",status_code="200"} 7
turnstone_http_requests_total{method="POST",endpoint="/v1/api/send",status_code="200"} 18
turnstone_http_requests_total{method="POST",endpoint="/v1/api/workstreams/{ws_id}/send",status_code="200"} 18
# HELP turnstone_tokens_total Total tokens consumed
# TYPE turnstone_tokens_total counter
turnstone_tokens_total{type="prompt"} 84320
+64 -25
View File
@@ -231,11 +231,13 @@ The engine emits state changes via `_emit_state()` which calls
> See also: [Core Engine Classes diagram](diagrams/png/03-core-engine-classes.png)
Defined in `turnstone.core.session.SessionUI` as a `typing.Protocol` with 14
Defined in `turnstone.core.session.SessionUI` as a `typing.Protocol` with 16
methods. Every frontend must implement all of them.
```python
class SessionUI(Protocol):
def on_turn_start(self) -> None: ...
def on_turn_committed(self) -> None: ...
def on_thinking_start(self) -> None: ...
def on_thinking_stop(self) -> None: ...
def on_reasoning_token(self, text: str) -> None: ...
@@ -252,6 +254,14 @@ class SessionUI(Protocol):
def on_rename(self, name: str) -> None: ... # propagate alias to tab/UI label
```
`on_turn_start` fires at the top of each iteration of the send-loop;
`on_turn_committed` fires immediately after `messages.append(assistant_msg)`.
`SessionUIBase` uses both to reset the per-turn inflight buffers
(`_ws_inflight_content` / `_ws_inflight_reasoning` / `_ws_inflight_seq`)
that fuel the SSE refresh-resume `in_progress_snapshot` event — see
the per-workstream events stream in
[`docs/api-reference.md`](api-reference.md#get-v1apiworkstreamsws_idevents).
`on_rename` is called by the `/name` command (on success) and after a successful `/resume` (if the resumed session has an alias or title). `WebUI.on_rename` broadcasts a `ws_rename` event on the global SSE channel and updates the in-memory `Workstream.name`; `TerminalUI.on_rename` is a no-op.
### Three Implementations
@@ -384,11 +394,11 @@ non-idle background workstreams above the input prompt.
(`Ctrl+\`, `Ctrl+Shift+\`). Max 6 panes; no duplicate workstreams across panes.
Layout persisted to `localStorage`.
- **Per-pane SSE**: `Pane.connectSSE(wsId)` opens
`/v1/api/events?ws_id=<id>` for each pane's event stream independently.
`/v1/api/workstreams/{ws_id}/events` for each pane's event stream independently.
- **Global SSE**: `connectGlobalSSE()` opens `/v1/api/events/global` which
receives `ws_state` broadcasts from all workstreams, used to update tab
indicators and pane headers without switching.
- **New tab / close**: POST `/v1/api/workstreams/new`, POST `/v1/api/workstreams/close`.
- **New tab / close**: POST `/v1/api/workstreams/new`, POST `/v1/api/workstreams/{ws_id}/close`.
### Thread Safety
@@ -546,11 +556,9 @@ adds, removes, or reconnects servers as needed.
6. `_exec_mcp_tool()` calls `call_tool_sync()` which dispatches to the async loop
via `asyncio.run_coroutine_threadsafe()`
**Tool refresh:** Three mechanisms keep tools up-to-date without restart:
**Tool refresh:** Two mechanisms keep tools up-to-date without restart:
- **Push:** Servers declaring `tools.listChanged` send `ToolListChangedNotification`;
the registered `message_handler` triggers immediate single-server refresh.
- **Periodic:** Servers without push support are polled on a staggered interval
(default 4 h, configurable via `[mcp] refresh_interval` or `--mcp-refresh-interval`).
- **Manual:** `/mcp refresh [server]` calls `refresh_sync()` for on-demand refresh
(also attempts reconnection for disconnected servers).
@@ -572,10 +580,11 @@ from a healthy connection do not trip the breaker. When the cooldown expires
(`call_tool_sync`, `read_resource_sync`, `get_prompt_sync`, `refresh_sync`)
cancel orphaned futures on timeout to prevent coroutine accumulation on the
background event loop. Push notification refreshes are debounced (5 s per
server) to protect against notification storms. The periodic refresh loop
attempts reconnection for disconnected servers with exponential backoff
(60 s1 h). Transport stream references are pre-closed before stack teardown to
work around the MCP SDK's anyio cancel-scope CPU busy-loop (SDK #2147).
server) to protect against notification storms. Operators can force a
catalog refresh or full reconnect from the admin panel; reconnects clear
the circuit breaker and run a fresh handshake. Transport stream references
are pre-closed before stack teardown to work around the MCP SDK's anyio
cancel-scope CPU busy-loop (SDK #2147).
**Error isolation:** Per-server connection/refresh failures are caught and logged; other
servers are unaffected. Tool execution errors return error strings to the LLM
@@ -620,14 +629,15 @@ LLMProvider (protocol)
| `get_capabilities()` | Per-model flags (`ModelCapabilities`) |
| `convert_tools()` | Translate OpenAI tool schemas to provider format |
| `retryable_error_names` | Exception class names that trigger retry |
| `extract_reasoning_text()` | Walk stored `provider_blocks`, return concatenated reasoning text for UI rehydration (per-provider block-type knowledge: Anthropic `thinking`, OpenAI Responses `reasoning`, OpenAI Chat synthetic `reasoning_text`) |
**Normalized data types:**
| Type | Fields |
|------|--------|
| `StreamChunk` | `content_delta`, `reasoning_delta`, `tool_call_deltas`, `info_delta`, `usage`, `finish_reason` |
| `CompletionResult` | `content`, `tool_calls`, `finish_reason`, `usage` |
| `ModelCapabilities` | `context_window`, `max_output_tokens`, `supports_temperature`, `token_param`, `thinking_mode`, `supports_effort`, `supports_web_search`, `supports_tool_search`, `supports_vision` |
| `StreamChunk` | `content_delta`, `reasoning_delta`, `tool_call_deltas`, `info_delta`, `usage`, `finish_reason`, `provider_blocks` |
| `CompletionResult` | `content`, `tool_calls`, `finish_reason`, `usage`, `provider_blocks` |
| `ModelCapabilities` | `context_window`, `max_output_tokens`, `supports_temperature`, `token_param`, `thinking_mode`, `supports_effort`, `supports_web_search`, `supports_tool_search`, `supports_vision`, `supports_reasoning_replay` |
| `UsageInfo` | `prompt_tokens`, `completion_tokens`, `total_tokens`, `cache_creation_tokens`, `cache_read_tokens` |
**OpenAIProvider** (`_openai.py`): passes messages through unchanged (they are
@@ -715,6 +725,35 @@ and `"openai-compatible"`.
`max_tokens`, and `reasoning_effort` to override the global defaults from
ConfigStore. When unset (`NULL`), the global default is used.
**Per-model reasoning persistence:** Two booleans on `model_definitions`
(migration 052) control how reasoning text round-trips:
* `surface_persisted_reasoning` (default `True`) — gates whether stored
reasoning text is surfaced on `/history` payloads for UI rehydration.
**Storage of reasoning bytes happens regardless of this flag** — they
ride in `provider_data` independently. Phase-1 admin UI label "Surface
persisted reasoning."
* `replay_reasoning_to_model` (default `False`) — gates whether stored
reasoning blocks are sent back to the provider on subsequent turns.
Capability-gated: `ModelCapabilities.supports_reasoning_replay` must
also be `True` for the wire path to actually replay (canonical OpenAI
gpt-5*/o-series and Anthropic Claude entries set it; unknown / local-
server models default to `False`).
Three reasoning paths are recognised:
| Path | Provider | Capture | Persist | Replay |
|------|----------|---------|---------|--------|
| 1 | Anthropic Messages API | `thinking_delta` | `provider_blocks` (`type="thinking"`) | Verbatim via `_provider_content` |
| 2 | OpenAI Responses (gpt-5*, o-series) | `response.reasoning_text.delta` events | `provider_blocks` (`type="reasoning"`) — only when `include=["reasoning.encrypted_content"]` | `ResponseReasoningItemParam` input items |
| 3 | OpenAI Chat Completions (vLLM, llama.cpp, Gemini-compat) | `delta.reasoning_content` Pydantic extras | Synthetic `{type: "reasoning_text", text, source}` block stamped at end-of-stream | None — no API surface for replay on Chat Completions |
Cross-provider safety is enforced by `ANTHROPIC_VALID_BLOCK_TYPES` (a
shape filter in `_anthropic.py:_convert_messages`): foreign blocks
(OpenAI `reasoning`, synthetic `reasoning_text`) fall through to the
text+tool_calls rebuild path rather than reaching Anthropic's input
boundary as malformed content.
```toml
[models.local]
base_url = "http://localhost:8000/v1"
@@ -1100,8 +1139,8 @@ Three hierarchical scopes control endpoint access:
| Scope | Grants | Endpoints |
|-------|--------|-----------|
| `read` | SSE streams, workstream listing, history | GET endpoints |
| `write` | `read` + send, command, workstream create/close | POST to `/api/send`, `/api/command`, etc. |
| `approve` | `write` + tool approval, admin operations | POST to `/api/approve`, `/api/admin/*` |
| `write` | `read` + send, command, workstream create/close | POST to `/api/workstreams/{ws_id}/send`, `/api/command`, etc. |
| `approve` | `write` + tool approval, admin operations | POST to `/api/workstreams/{ws_id}/approve`, `/api/admin/*` |
### Middleware Flow
@@ -1198,12 +1237,12 @@ stderr so it does not interfere with readline. Tool execution may use a
Starlette ASGI app (served by uvicorn)
|
+-- Async request handlers (all under /v1/ prefix)
| POST /v1/api/send -> starts worker thread per workstream
| POST /v1/api/approve -> unblocks WebUI._approval_event
| POST /v1/api/plan -> unblocks WebUI._plan_event
| POST /v1/api/workstreams/new -> creates workstream + worker
| GET /v1/api/events -> SSE via EventSourceResponse (per workstream)
| GET /v1/api/events/global -> SSE via EventSourceResponse (fan-out)
| POST /v1/api/workstreams/{ws_id}/send -> starts worker thread per workstream
| POST /v1/api/workstreams/{ws_id}/approve -> unblocks WebUI._approval_event
| POST /v1/api/plan -> unblocks WebUI._plan_event
| POST /v1/api/workstreams/new -> creates workstream + worker
| GET /v1/api/workstreams/{ws_id}/events -> SSE via EventSourceResponse (per workstream)
| GET /v1/api/events/global -> SSE via EventSourceResponse (fan-out)
|
+-- ASGI middleware stack
| MetricsMiddleware -> CORSMiddleware -> AuthMiddleware -> RateLimitMiddleware
@@ -1277,10 +1316,10 @@ Monitoring (2 daemon threads) Control + Proxy (async Starlette)
| SSE manager | | GET /node/{node_id}/ |
| asyncio loop | | → httpx.AsyncClient |
| 1 task per node | | proxy to server_url |
| /events/global | | GET /node/{id}/v1/api/events |
| snapshot+deltas | | → SSE stream proxy |
+------------------+ | POST /node/{id}/v1/api/send |
| → forwarded to server |
| /events/global | | GET /node/{id}/v1/api/workstreams/{ws_id}/events |
| snapshot+deltas | | → SSE stream proxy |
+------------------+ | POST /node/{id}/v1/api/workstreams/{ws_id}/send |
| → forwarded to server |
+----------------------------+
```
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:75c1832b6079e8628f4bbf4ce98d37880c4de133636b7555e3869990b046ddc6
size 567704
oid sha256:5d500479d3be2363d4f594042a27e2ef5e2974750f580f6c4037a1fe85868ed9
size 251904
+4 -4
View File
@@ -12,8 +12,8 @@ Existing bulk endpoints at time of writing:
|---------------------------------------------------------|--------------------------|------------------------------------------|
| `GET /v1/api/cluster/ws/live?ids=a,b,c` | bulk read | `{results, denied, truncated}` |
| model tool `spawn_batch` | bulk create (per-item) | `{results, denied}` |
| `POST /v1/api/coordinator/{ws_id}/stop_cascade` | cascade mutation | `{cancelled, failed, skipped}` |
| `POST /v1/api/coordinator/{ws_id}/close_all_children` | cascade mutation | `{closed, failed, skipped}` |
| `POST /v1/api/workstreams/{ws_id}/stop_cascade` | cascade mutation | `{cancelled, failed, skipped}` |
| `POST /v1/api/workstreams/{ws_id}/close_all_children` | cascade mutation | `{closed, failed, skipped}` |
---
@@ -113,8 +113,8 @@ owns it; the node is just currently unreachable.
```json
{
"results": {
"0": {"ws_id": "d4e5f6...", "name": "csrf-audit", "node_id": "gpu-3", "status": 200},
"2": {"ws_id": "f1a2b3...", "name": "xss-audit", "node_id": "gpu-1", "status": 200}
"0": {"ws_id": "d4e5f6...", "name": "csrf-audit", "node_id": "gpu-3"},
"2": {"ws_id": "f1a2b3...", "name": "xss-audit", "node_id": "gpu-1"}
},
"denied": [
{"idx": 1, "reason": "skill not found: nonexistent-skill"}
+2 -2
View File
@@ -334,7 +334,7 @@ The console reverse-proxies each node's server UI at `/node/{node_id}/`. This al
### URL Rewriting
The server UI uses root-relative URLs (`/v1/api/send`, `/static/app.js`, `/shared/base.css`, etc.). Since `<base>` tags cannot rewrite root-relative URLs, the console uses a JS shim approach:
The server UI uses root-relative URLs (`/v1/api/workstreams/{ws_id}/send`, `/static/app.js`, `/shared/base.css`, etc.). Since `<base>` tags cannot rewrite root-relative URLs, the console uses a JS shim approach:
1. **HTML rewriting** — when serving `index.html`, replaces `href=` and `src=` references to both `/static/` and `/shared/` with the proxy prefix (`/node/{node_id}/static/` and `/node/{node_id}/shared/` respectively).
@@ -344,7 +344,7 @@ The server UI uses root-relative URLs (`/v1/api/send`, `/static/app.js`, `/share
### SSE Proxy
SSE streams (`/v1/api/events`, `/v1/api/events/global`) are proxied as raw byte passthrough — the console opens an `httpx.AsyncClient.stream()` to the upstream server (with `read=None` and `pool=None` timeouts since SSE connections are long-lived) and relays every byte via `StreamingResponse`. This preserves server-side ping comments, event framing, and keepalives verbatim without parsing or re-encoding.
SSE streams (`/v1/api/workstreams/{ws_id}/events`, `/v1/api/events/global`) are proxied as raw byte passthrough — the console opens an `httpx.AsyncClient.stream()` to the upstream server (with `read=None` and `pool=None` timeouts since SSE connections are long-lived) and relays every byte via `StreamingResponse`. This preserves server-side ping comments, event framing, and keepalives verbatim without parsing or re-encoding.
### Authentication
+58 -38
View File
@@ -26,28 +26,43 @@ schema changes.
## The 9 steps
| # | Action | Operation | Operation id |
|---|------------------------------|-------------------------------------------------------------|-------------------------------------------------------------|
| 1 | Create | `POST /v1/api/coordinator/new` | `v1_api_coordinator_new_post` |
| 2 | Subscribe to events | `GET /v1/api/coordinator/{ws_id}/events` (SSE) | `v1_api_coordinator_{ws_id}_events_get` |
| 3 | Send a user message | `POST /v1/api/coordinator/{ws_id}/send` | `v1_api_coordinator_{ws_id}_send_post` |
| 4 | Inspect children | `GET /v1/api/coordinator/{ws_id}/children` | `v1_api_coordinator_{ws_id}_children_get` |
| 5 | Inspect one workstream | `GET /v1/api/cluster/ws/{ws_id}/detail` | `v1_api_cluster_ws_{ws_id}_detail_get` |
| 6 | Wait for fan-out | model-side tool `wait_for_workstream` | — (tool call, not HTTP) |
| 7 | Govern | `POST /v1/api/coordinator/{ws_id}/trust` | `v1_api_coordinator_{ws_id}_trust_post` |
| | | `POST /v1/api/coordinator/{ws_id}/restrict` | `v1_api_coordinator_{ws_id}_restrict_post` |
| | | `POST /v1/api/coordinator/{ws_id}/stop_cascade` | `v1_api_coordinator_{ws_id}_stop_cascade_post` |
| | | `POST /v1/api/coordinator/{ws_id}/close_all_children` | `v1_api_coordinator_{ws_id}_close_all_children_post` |
| 8 | Approve / cancel | `POST /v1/api/coordinator/{ws_id}/approve` | `v1_api_coordinator_{ws_id}_approve_post` |
| | | `POST /v1/api/coordinator/{ws_id}/cancel` | `v1_api_coordinator_{ws_id}_cancel_post` |
| 9 | Close | `POST /v1/api/coordinator/{ws_id}/close` | `v1_api_coordinator_{ws_id}_close_post` |
> **URL convergence (1.5.0).** Pre-1.5 coord-only endpoints lived
> under `/v1/api/coordinator/...`. The Stage 2 verb-shape lift
> consolidated coord and interactive onto the unified
> `/v1/api/workstreams/{ws_id}/<verb>` tree; coord still distinguishes
> itself via the `kind=coordinator` row classifier rather than a
> separate URL space. The endpoints below reflect the post-lift
> surface served by `turnstone-console`.
| # | Action | Operation |
|---|------------------------------|-------------------------------------------------------------|
| 1 | Create | `POST /v1/api/workstreams/new` |
| 2 | Subscribe to events | `GET /v1/api/workstreams/{ws_id}/events` (SSE) |
| 3 | Send a user message | `POST /v1/api/workstreams/{ws_id}/send` |
| 4 | Inspect children | `GET /v1/api/workstreams/{ws_id}/children` |
| 5 | Inspect one workstream | `GET /v1/api/cluster/ws/{ws_id}/detail` |
| 6 | Wait for fan-out | model-side tool `wait_for_workstream` |
| 7 | Govern | `POST /v1/api/workstreams/{ws_id}/trust` |
| | | `POST /v1/api/workstreams/{ws_id}/restrict` |
| | | `POST /v1/api/workstreams/{ws_id}/stop_cascade` |
| | | `POST /v1/api/workstreams/{ws_id}/close_all_children` |
| 8 | Approve / cancel | `POST /v1/api/workstreams/{ws_id}/approve` |
| | | `POST /v1/api/workstreams/{ws_id}/cancel` |
| 9 | Close | `POST /v1/api/workstreams/{ws_id}/close` |
Refer to `/openapi.json` (Swagger UI at `/docs`) on any
`turnstone-console` process for the authoritative operation ids and
schemas. Coordinator-only verbs (`/children`, `/trust`, `/restrict`,
`/stop_cascade`, `/close_all_children`) 404 against `kind=interactive`
rows; the shared verbs (`/send`, `/approve`, `/cancel`, `/events`,
`/history`, `/open`, `/close`, etc.) work on both kinds.
---
## 1. Create a coordinator
```http
POST /v1/api/coordinator/new
POST /v1/api/workstreams/new
Content-Type: application/json
Authorization: Bearer <token>
@@ -80,7 +95,7 @@ subscribers (step 2) see the session warm up as token traffic starts.
## 2. Subscribe to the per-coordinator event stream
```http
GET /v1/api/coordinator/{ws_id}/events HTTP/1.1
GET /v1/api/workstreams/{ws_id}/events HTTP/1.1
Accept: text/event-stream
Authorization: Bearer <token>
```
@@ -100,7 +115,8 @@ with a `type` field. The recurring shapes a UI has to handle:
| `tool_output_chunk` | Streaming tool output (e.g. long bash command) | `call_id`, `chunk` |
| `approve_request` | One or more tool calls need operator approval | `items: [{call_id, header, preview, func_name, approval_label, needs_approval}]` |
| `approval_resolved` | Operator answered the approval prompt | `approved`, `feedback` |
| `state_change` | Worker-thread state transition | `state``running`, `thinking`, `attention`, `idle`, `error` |
| `state_change` | Worker-thread state transition (also re-emitted with the current state on every fresh subscribe so refresh-mid-stream restores composer mode) | `state``running`, `thinking`, `attention`, `idle`, `error` |
| `in_progress_snapshot` | One-shot replay of the in-progress turn's content + reasoning when this client connects mid-stream | `content`, `reasoning` |
| `status` | Token usage + context-window snapshot (fires on every streaming tick) | `prompt_tokens`, `completion_tokens`, `total_tokens`, `context_window`, `pct`, `effort`, `cache_creation_tokens`, `cache_read_tokens` |
| `rename` | Session's display name changed | `name` |
| `intent_verdict` | Intent judge produced a verdict on a pending tool call | `risk_level`, `recommendation`, `reasons` |
@@ -115,16 +131,20 @@ with a `type` field. The recurring shapes a UI has to handle:
**Reconnection contract:** a freshly-opened SSE connection receives
the current snapshot of any pending tool approval (`approve_request`
is re-sent if unresolved) and any in-flight `wait_*` / `batch_*`
indicator — so a tab refresh mid-approval doesn't strand the
operator.
is re-sent if unresolved), any in-flight `wait_*` / `batch_*`
indicator, the worker's current `state_change`, and an
`in_progress_snapshot` carrying any partial content / reasoning the
model has produced for the in-progress turn — so a tab refresh
mid-approval, mid-tool-execution, or mid-stream restores both the
correct composer mode and the partial assistant text without waiting
for the response to complete.
---
## 3. Send the first user message
```http
POST /v1/api/coordinator/{ws_id}/send
POST /v1/api/workstreams/{ws_id}/send
Content-Type: application/json
{"message": "audit /auth for CSRF handling across all active routes"}
@@ -147,7 +167,7 @@ events, finishing with `state_change → idle` or an
## 4. Inspect direct children
```http
GET /v1/api/coordinator/{ws_id}/children HTTP/1.1
GET /v1/api/workstreams/{ws_id}/children HTTP/1.1
```
```json
@@ -243,8 +263,8 @@ burst can't starve audit writes.
### `POST /trust` — auto-approve own-subtree sends
```json
POST /v1/api/coordinator/{ws_id}/trust
```http
POST /v1/api/workstreams/{ws_id}/trust
{"send": true}
```
@@ -257,8 +277,8 @@ second grants a service token the opt-in it otherwise wouldn't get).
### `POST /restrict` — revoke tool access mid-session
```json
POST /v1/api/coordinator/{ws_id}/restrict
```http
POST /v1/api/workstreams/{ws_id}/restrict
{"revoke": ["spawn_workstream", "delete_workstream"]}
```
@@ -269,8 +289,8 @@ opt in per session. Cap 256 tool names per request, 128 chars each.
### `POST /stop_cascade` — cancel the subtree
```json
POST /v1/api/coordinator/{ws_id}/stop_cascade
```http
POST /v1/api/workstreams/{ws_id}/stop_cascade
{}
```
@@ -291,8 +311,8 @@ propagate via the child's SSE stream.
### `POST /close_all_children` — soft-close the direct fan-out
```json
POST /v1/api/coordinator/{ws_id}/close_all_children
```http
POST /v1/api/workstreams/{ws_id}/close_all_children
{"reason": "audit round complete"}
```
@@ -322,8 +342,8 @@ The `approve` endpoint is what resolves an `approve_request` SSE
event. The coordinator's worker thread is blocked inside
`ui.approve_tools` waiting for this POST.
```json
POST /v1/api/coordinator/{ws_id}/approve
```http
POST /v1/api/workstreams/{ws_id}/approve
{"approved": true, "feedback": null, "always": false}
{"approved": false, "feedback": "spawn count looks too high — try 3 not 10"}
{"approved": true, "feedback": null, "always": true} // always-approve this tool name
@@ -332,8 +352,8 @@ POST /v1/api/coordinator/{ws_id}/approve
`cancel` drops the in-flight generation but leaves the coordinator
idle and open for a fresh `send`:
```json
POST /v1/api/coordinator/{ws_id}/cancel
```http
POST /v1/api/workstreams/{ws_id}/cancel
{}
```
@@ -341,8 +361,8 @@ POST /v1/api/coordinator/{ws_id}/cancel
## 9. Close
```json
POST /v1/api/coordinator/{ws_id}/close
```http
POST /v1/api/workstreams/{ws_id}/close
{}
```
@@ -350,7 +370,7 @@ Soft-closes the session — state persists, children keep running (use
`close_all_children` or `stop_cascade` first to wind them down), the
worker thread exits, SSE streams send a final `stream_end` and
disconnect. The row is reopenable via
`POST /v1/api/coordinator/{ws_id}/open` so long as it hasn't been
`POST /v1/api/workstreams/{ws_id}/open` so long as it hasn't been
deleted.
---
+21 -19
View File
@@ -65,7 +65,7 @@ or MCP config can do adds to it. Current members:
| `delete_workstream` | wind-down | Hard-delete one child. Requires approval. |
| `list_nodes` | discover | Enumerate live cluster nodes + capabilities. |
| `list_skills` | discover | Coordinator-visible skills only (SkillKind filter above). |
| `task_list` | plan | Orchestrator-only scratchpad. Children don't see it. |
| `tasks` | plan | Orchestrator-only scratchpad. Children don't see it. |
Explicitly **not** in the coordinator set:
@@ -108,9 +108,9 @@ the skill should end on.
---
## `task_list` integration
## `tasks` integration
`task_list` is the coordinator's scratchpad — a persisted, ordered
`tasks` is the coordinator's scratchpad — a persisted, ordered
list of rows with fields `{id, title, status, child_ws_id, created,
updated}` that only this coordinator sees. Children don't see it;
the user does via the sidebar. Five actions: `add`, `update`,
@@ -125,13 +125,13 @@ a skill can set it to a placeholder before `spawn_workstream`
returns or keep it pointing at a closed child for later audit.
A skill's initial prompt can seed the task list by calling
`task_list(action="add", title=...)` as its very first tool calls —
`tasks(action="add", title=...)` as its very first tool calls —
the user gets a visible plan before any child is spawned, and the
coordinator's future self has something concrete to iterate on.
Status transitions (`pending``in_progress``done` / `blocked`)
are the skill's main feedback loop: mutate the task when the child
covering it finishes, not when the child starts. Use
`task_list(action="update", task_id=..., child_ws_id=<ws_id>)` to
`tasks(action="update", task_id=..., child_ws_id=<ws_id>)` to
link a task to the child that owns it once spawn returns.
A final gotcha: parallel tool dispatch does NOT serialise reads
@@ -160,7 +160,7 @@ validates ws_id against `parent_ws_id=coord_ws_id` AND
`cancel_workstream`, `delete_workstream`) return
`{"error": "workstream not in coordinator subtree: <ws_id>", "status": 404}`
— the skill should treat this as a tool error, not an empty result.
- **`inspect_workstream`** returns `{"error": "workstream not found: <ws_id>"}`
- **`inspect_workstream`** returns `{"error": "workstream not found", "ws_id": "<ws_id>"}`
(same shape as a genuinely missing row, so the guard can't be
used as an existence oracle).
- **`wait_for_workstream`** reports the offending id with
@@ -170,9 +170,10 @@ validates ws_id against `parent_ws_id=coord_ws_id` AND
Pattern: capture each spawn result in the next tool call's input.
The JSON tool-result carries `{"ws_id": "...", "name": "...",
"node_id": "...", "status": 200}`; the model should extract the
ws_id and pass it to `inspect_workstream` / `wait_for_workstream` /
`send_to_workstream` / `close_workstream` verbatim.
"node_id": "...", "routing_strategy": "..."}`; the model should
extract the ws_id and pass it to `inspect_workstream` /
`wait_for_workstream` / `send_to_workstream` / `close_workstream`
verbatim.
A UI that wants human-readable identifiers should render the `name`
field and keep the ws_id as the click-through key.
@@ -215,12 +216,12 @@ to the user. Appropriate when the user's request is "run the thing
and tell me what happened" and the work fits in one workstream.
```
task_list(action='add', title='audit /auth for CSRF')
tasks(action='add', title='audit /auth for CSRF')
spawn_workstream(skill='engineer', initial_message='audit /auth ...')
wait_for_workstream(ws_ids=[<child>], timeout=300)
inspect_workstream(ws_id=<child>)
→ synthesise the final message into a user-facing response
task_list(action='update', task_id='t_01', status='done')
tasks(action='update', task_id='t_01', status='done')
close_workstream(ws_id=<child>, reason='audit complete')
```
@@ -231,7 +232,7 @@ waited-on together, then synthesised. Appropriate when the user's
request naturally decomposes into independent subtasks.
```
task_list seeds:
tasks seeds:
t_01 benchmark Anthropic 4.7 latency on summarisation
t_02 benchmark OpenAI GPT-5.2 latency on summarisation
t_03 benchmark Gemini 2.5 latency on summarisation
@@ -239,7 +240,7 @@ spawn_batch(children=[...3 briefs...])
wait_for_workstream(ws_ids=[c1, c2, c3], mode='all', timeout=600)
inspect_workstream(ws_id=c1); ...(c2); ...(c3)
→ synthesise head-to-head comparison
task_list → all done
tasks → all done
close_all_children(reason='benchmark complete')
```
@@ -252,20 +253,20 @@ approval each.
### Pattern 3 — plan-then-delegate
The coordinator first uses its own reasoning to carve the plan,
records it in `task_list`, then spawns children that each own one
records it in `tasks`, then spawns children that each own one
task. Appropriate when the user's request is "figure out how to X"
and the coordinator's planning step is itself valuable.
```
→ coord reasons about the shape of the work
task_list(action='add', title='...') × N # the plan, visible in the sidebar
tasks(action='add', title='...') × N # the plan, visible in the sidebar
for task in tasks:
spawn_workstream(skill=..., initial_message=task.brief)
task_list(action='update', task_id=task.id, notes='ws=<child_ws_id>')
tasks(action='update', task_id=task.id, notes='ws=<child_ws_id>')
wait_for_workstream(ws_ids=[...], mode='all', timeout=...)
for child in children:
inspect_workstream(ws_id=child)
task_list(action='update', task_id=..., status='done', notes='result summary')
tasks(action='update', task_id=..., status='done', notes='result summary')
→ synthesise
```
@@ -290,13 +291,14 @@ For a new coordinator skill:
`coord_session` fixture's `skill=` kwarg (see
`tests/test_coordinator_tools.py` for the pattern).
2. Build a small fake cluster: one node + two children via
`mgr.register_children(coord.id, ["child-1", "child-2"])`.
the `_seed_children` helper in `tests/_coord_test_helpers.py`
(``_seed_children(mgr._adapter, coord.id, ["child-1", "child-2"])``).
3. Drive the session with seeded tool_call dicts matching the
provider layer's shape. The unit-level tests in
`tests/test_coordinator_tools.py` show the helper (`_tc(name,
args, call_id)`).
4. Assert the skill's decision shape — which tools fire in what
order, what the task_list looks like at the end, which
order, what the tasks looks like at the end, which
`_error` reasons appear on the denied-path.
A full end-to-end test isn't required for every skill; a
+1 -1
View File
@@ -43,7 +43,7 @@ eval --> sqlite : SQLite
console --> server : HTTP proxy\n(hash-ring bucket lookup,\nproxy /node/{id}/* traffic)
channel --> server : HTTP + SSE\n(POST /v1/api/send,\nGET /v1/api/events)
channel --> server : HTTP + SSE\n(POST /v1/api/workstreams/{ws_id}/send,\nGET /v1/api/workstreams/{ws_id}/events)
' Notes
note right of console
+1 -1
View File
@@ -40,7 +40,7 @@ package "turnstone/core/" <<Rectangle>> {
component [auth.py\nAuthentication] as auth <<core>>
component [healthcheck.py\nBackendHealthMonitor] as healthcheck <<core>>
component [ratelimit.py\nRateLimiter] as ratelimit <<core>>
component [mcp_client.py\nMCPClientManager\n(push + periodic refresh)] as mcp <<core>>
component [mcp_client.py\nMCPClientManager\n(push + manual refresh)] as mcp <<core>>
component [tool_search.py\nToolSearchManager, BM25] as toolsearch <<core>>
component [model_registry.py\nModelRegistry] as registry <<core>>
}
+5 -3
View File
@@ -69,9 +69,10 @@ class "NullUI" as NullUI {
interface "LLMProvider" as LLMProvider <<Protocol>> {
+ provider_name: str {property}
+ get_capabilities(model) → ModelCapabilities
+ create_streaming(client, model, messages, ...) → Iterator[StreamChunk]
+ create_completion(client, model, messages, ...) → CompletionResult
+ create_streaming(client, model, messages, ..., replay_reasoning_to_model) → Iterator[StreamChunk]
+ create_completion(client, model, messages, ..., replay_reasoning_to_model) → CompletionResult
+ convert_tools(tools) → list[dict]
+ extract_reasoning_text(provider_blocks) → str
+ retryable_error_names: frozenset[str] {property}
--
core/providers/_protocol.py
@@ -126,6 +127,7 @@ class "ModelCapabilities" as ModelCaps <<frozen>> {
+ supports_web_search: bool
+ supports_tool_search: bool
+ supports_vision: bool
+ supports_reasoning_replay: bool
}
' ChatSession
@@ -253,7 +255,7 @@ class "MCPClientManager" as MCPMgr {
Background asyncio event loop
bridges async MCP SDK to
sync ChatSession dispatch.
Push + periodic + manual refresh.
Push + manual refresh.
Resources + prompts discovered
alongside tools at startup.
--
+16
View File
@@ -24,6 +24,14 @@ CS -> DB : save_message(ws_id, "user", input)
group loop [while tool_calls present]
CS -> UI : on_turn_start()
note right of UI
SessionUIBase resets the per-turn inflight
buffers (_ws_inflight_content / reasoning /
seq) that fuel the SSE in_progress_snapshot
event for mid-stream refresh resume.
end note
CS -> UI : on_state_change("thinking")
CS -> UI : on_thinking_start()
@@ -73,6 +81,14 @@ group loop [while tool_calls present]
CS -> CS : _update_token_table()\ncalibrate chars_per_token ratio
CS -> CS : messages.append(assistant_msg)
CS -> UI : on_turn_committed()
note right of UI
Drops the per-turn inflight buffers — the
assistant message is now in the history
list, so the in_progress_snapshot must
not re-render it during the next tool-
execution window or the next streaming turn.
end note
CS -> DB : save_message(ws_id, "assistant", content)
CS -> DB : save_message(ws_id, "tool_call", ...) ×N
+6 -6
View File
@@ -170,15 +170,15 @@ Server --> Browser : Shimmed app.js
deactivate Server
note right of Browser
All fetch("/v1/api/send") calls in the
server UI now become fetch("/node/nodeA/v1/api/send"),
All fetch("/v1/api/workstreams/{ws_id}/send") calls in the
server UI now become fetch("/node/nodeA/v1/api/workstreams/{ws_id}/send"),
routed through the console proxy.
end note
Browser -> Server : GET /node/nodeA/v1/api/events?ws_id=ws789
Browser -> Server : GET /node/nodeA/v1/api/workstreams/ws789/events
activate Server #FFF9C4
Server -> NodeA : GET http://10.0.1.1:8080/v1/api/events?ws_id=ws789\n(SSE stream via httpx.AsyncClient timeout=None)
Server -> NodeA : GET http://10.0.1.1:8080/v1/api/workstreams/ws789/events\n(SSE stream via httpx.AsyncClient timeout=None)
activate NodeA
loop SSE streaming
@@ -189,10 +189,10 @@ end
deactivate NodeA
deactivate Server
Browser -> Server : POST /node/nodeA/v1/api/send\n{message:"hello", ws_id:"ws789"}
Browser -> Server : POST /node/nodeA/v1/api/workstreams/ws789/send\n{message:"hello"}
activate Server #FFF9C4
Server -> NodeA : POST http://10.0.1.1:8080/v1/api/send\n(body forwarded)
Server -> NodeA : POST http://10.0.1.1:8080/v1/api/workstreams/ws789/send\n(body forwarded)
activate NodeA
NodeA --> Server : {status:"ok"}
deactivate NodeA
+1 -1
View File
@@ -79,7 +79,7 @@ class "Scope Hierarchy" as SH <<scope>> {
--
GET → read
POST write paths → write
POST /api/approve → approve
POST /api/workstreams/{ws_id}/approve → approve
/api/admin/* → approve
}
+9 -9
View File
@@ -95,10 +95,10 @@ class "ChannelRouter" as Router <<service>> {
' -- Server --
class "turnstone-server" as Server <<server>> {
POST /v1/api/send
POST /v1/api/approve
POST /v1/api/workstreams/{ws_id}/send
POST /v1/api/workstreams/{ws_id}/approve
POST /v1/api/workstreams/new
GET /v1/api/events?ws_id=
GET /v1/api/workstreams/{ws_id}/events
--
LLM execution + tool use
SSE event stream
@@ -148,15 +148,15 @@ Bot --> Router : on_message\non_interaction
Router --> CU : resolve identity
Router --> CR : resolve / register route
Router --> Server : POST /v1/api/send\nPOST /v1/api/approve\nPOST /v1/api/workstreams/new
Bot --> Server : GET /v1/api/events?ws_id=\n(SSE via httpx-sse)
Router --> Server : POST /v1/api/workstreams/{ws_id}/send\nPOST /v1/api/workstreams/{ws_id}/approve\nPOST /v1/api/workstreams/new
Bot --> Server : GET /v1/api/workstreams/{ws_id}/events\n(SSE via httpx-sse)
Server --> Bot : SSE event stream
Bot --> Discord : reply / embed\nbutton callback
Slack --> SlackBot : socket-mode\nevents
SlackBot --> Router : on_message / on_action
SlackBot --> Server : POST /v1/api/send\nGET /v1/api/events?ws_id=
SlackBot --> Server : POST /v1/api/workstreams/{ws_id}/send\nGET /v1/api/workstreams/{ws_id}/events
SlackBot --> Slack : post / update\nBlock Kit button callbacks
Teams .[hidden]. Slack
@@ -179,7 +179,7 @@ note right of Bot
(or creates new workstream)
4. ChannelRouter resolves platform user -> user_id
via channel_users table
5. Router sends POST /v1/api/send to server
5. Router sends POST /v1/api/workstreams/{ws_id}/send to server
**Workstream Resume (evicted workstreams)**
1. Stale route detected (no active SSE listener)
@@ -193,7 +193,7 @@ end note
note right of Server
**Outbound Flow**
1. Server emits SSE events on
GET /v1/api/events?ws_id=
GET /v1/api/workstreams/{ws_id}/events
2. Bot subscribes via httpx-sse
3. Bot formats and sends to Discord thread
end note
@@ -204,7 +204,7 @@ note bottom of CR
2. Bot renders Discord buttons (Approve / Deny)
3. User clicks button -> on_interaction()
4. Router builds ApproveMessage
5. Router sends POST /v1/api/approve to server
5. Router sends POST /v1/api/workstreams/{ws_id}/approve to server
end note
note bottom of CU
+15 -11
View File
@@ -190,21 +190,25 @@ group Push Notifications (debounced 5s per server)
MCPMgr -> Storage : sync_prompts_to_storage()
end
group Periodic Polling (default 4h)
MCPMgr -> MCPMgr : _periodic_refresh()
group Manual Refresh
Session -> MCPMgr : refresh_sync()
note right
Only polls capabilities
without push support.
Staggered per-server.
Disconnected servers get
reconnect attempts with
exponential backoff (60s-1h).
/mcp refresh [server] —
re-fetches catalog and
attempts reconnect for
disconnected servers.
end note
end
group Manual Refresh
Session -> MCPMgr : refresh_sync()
note right: /mcp refresh [server]
group Manual Reconnect
Session -> MCPMgr : reconnect_sync(name)
note right
Operator-driven via the
console admin panel —
tears down session, clears
circuit breaker, runs a
fresh handshake.
end note
end
== Policy Evaluation ==
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a3b5c59403a6febd81667fc8fd2a7d22bc59da6130eba0dea5449c42668d0ede
size 387044
oid sha256:95dd5ebc899a1261d516686a5aa3319a7f45015d411302825fa28afbfc82e1ce
size 326766
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:474b900448ec04d1117b48a2b55614524721b2f04ac4bda66170bd0a06aae0f2
size 624573
oid sha256:9857db23fe3c4316d492073aac69c7e7558b1abe3b95ad7756d4a5933bd0ece7
size 620214
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3aa8d972bba40d78152f9f0c762b9f5ec616d8052c45fa52b7dd1c679ed81d61
size 325245
oid sha256:c14dfbb2db8dcb22cd332b2cf0e53ba75141adb213dae47dd9dbfd389ed482fe
size 354799
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7623df33be9baf7647ca1c2450640df57e1cd73e8be1f8168aae16e546ad683c
size 459941
oid sha256:d6aff446a062aa08f316985d00c2183148694f786d7f22172bc50b30046c728b
size 379259
+82 -13
View File
@@ -39,18 +39,19 @@ are set.
| `TURNSTONE_OIDC_ROLE_CLAIM` | No | — | ID token claim containing role/group values (see [Role Mapping](#role-mapping)) |
| `TURNSTONE_OIDC_ROLE_MAP` | No | — | Mapping from claim values to Turnstone role IDs (see [Role Mapping](#role-mapping)) |
| `TURNSTONE_OIDC_PASSWORD_ENABLED` | No | `true` | Set to `false` to hide the password form and block all username/password logins (including admin). API tokens continue to work. |
| `TURNSTONE_OIDC_REDIRECT_BASE` | No | — | Externally-reachable origin for the OIDC redirect URI (e.g. `https://app.example.com`). Recommended when running behind a reverse proxy. When unset, derived from the request Host header. |
| `TURNSTONE_OIDC_REDIRECT_BASE` | Yes | — | Externally-reachable origin for the OIDC redirect URI (e.g. `https://app.example.com`). Without this, OIDC will refuse to start. The previous Host-header fallback was unsafe under permissive reverse proxies. |
| `TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS` | No | — | Comma-separated list of additional hostnames whose endpoints the IdP discovery document is allowed to reference. See [Cross-host endpoints](#cross-host-endpoints). |
OIDC is enabled when all three required fields (issuer, client ID, client
secret) are non-empty. If any is missing, OIDC is silently disabled and
the login screen shows only the password form.
All four required fields issuer, client ID, client secret, and
`TURNSTONE_OIDC_REDIRECT_BASE` — must be set. If any are missing OIDC
is disabled at startup (an error is logged when only `redirect_base`
is missing) and the login screen shows only the password form.
### Reverse Proxy / Load Balancer
### Redirect base (required)
When Turnstone runs behind a reverse proxy, the internal `Host` header may
not match the externally-reachable URL. Set `TURNSTONE_OIDC_REDIRECT_BASE`
to the public origin so the redirect URI sent to the identity provider is
correct:
`TURNSTONE_OIDC_REDIRECT_BASE` pins the redirect URI sent to the identity
provider to a known externally-visible origin. Set it to the public origin
of your Turnstone deployment:
```bash
TURNSTONE_OIDC_REDIRECT_BASE=https://app.example.com
@@ -60,6 +61,44 @@ The resulting callback URL will be
`https://app.example.com/v1/api/auth/oidc/callback` — register this as the
authorized redirect URI in your identity provider.
OIDC will refuse to start when this variable is unset. There is no
Host-header fallback: a permissive reverse proxy or direct backend access
would otherwise let an attacker spoof `Host` and steer the IdP redirect
to a callback origin they control.
### Cross-host endpoints
By default, every endpoint in the IdP discovery document
(`token_endpoint`, `jwks_uri`, `userinfo_endpoint`) must share the
issuer's `(scheme, host, port)`. This prevents a hostile or compromised
IdP from redirecting the token-exchange POST (which carries
`client_secret`) to an arbitrary host, and prevents JWKS fetches from
being aimed at internal services.
A few public IdPs legitimately split endpoints across hostnames. Google
is the canonical example:
| Field | Hostname |
|-------|----------|
| issuer | `accounts.google.com` |
| token_endpoint | `oauth2.googleapis.com` |
| jwks_uri | `www.googleapis.com` |
| userinfo_endpoint | `openidconnect.googleapis.com` |
Google's set is built in — operators using `https://accounts.google.com`
need no extra configuration.
For other IdPs whose discovery document references a non-issuer host,
extend the allow-list explicitly:
```bash
TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS=token.example.com,keys.example.com
```
The same scheme / no-userinfo / SSRF rules apply to allow-listed hosts —
this knob only relaxes the same-origin check, not the security gates.
Each entry is a hostname (no scheme, no path).
### config.toml alternative
```toml
@@ -198,6 +237,19 @@ TURNSTONE_OIDC_ROLE_MAP="admin:builtin-admin,engineering:builtin-operator,viewer
the user authenticates via OIDC, so new group memberships are picked
up on the next login.
### `assigned_by` markers
Role assignments record an `assigned_by` value that controls how the
sync logic treats them. OIDC-driven flows use two distinct markers:
- `oidc` — set by claim-driven role mapping; revoked automatically on
the next login when the corresponding claim value is no longer
present.
- `oidc-default` — applied to brand-new OIDC users who have no
claim-mapped roles, as a safety net so they still get
`builtin-viewer` access on first login. Survives subsequent logins
regardless of claim contents and is never revoked by `apply_role_mapping`.
### Built-in Roles
| Role ID | Permissions |
@@ -375,10 +427,27 @@ callback validation. Entries are automatically cleaned up after 5 minutes.
### "OIDC not configured"
All three required environment variables must be set:
`TURNSTONE_OIDC_ISSUER`, `TURNSTONE_OIDC_CLIENT_ID`, and
`TURNSTONE_OIDC_CLIENT_SECRET`. Check that none are empty or
whitespace-only.
All four required environment variables must be set:
`TURNSTONE_OIDC_ISSUER`, `TURNSTONE_OIDC_CLIENT_ID`,
`TURNSTONE_OIDC_CLIENT_SECRET`, and `TURNSTONE_OIDC_REDIRECT_BASE`.
Check that none are empty or whitespace-only.
### "OIDC enabled but TURNSTONE_OIDC_REDIRECT_BASE is unset"
This error is logged when the three credential variables are set but
`TURNSTONE_OIDC_REDIRECT_BASE` is missing. OIDC is disabled at startup
to prevent Host-header-derived redirect URI spoofing. Set the variable
to your service's externally-visible origin (e.g.
`https://app.example.com`) and restart the server. See
[Redirect base](#redirect-base-required) for the rationale.
### Discovery silently disables OIDC with "host does not match issuer"
The IdP discovery document points `token_endpoint`, `jwks_uri`, or
`userinfo_endpoint` at a hostname that doesn't share the issuer's
origin. If the IdP is legitimate, add the additional hostname(s) to
`TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS`. Google is allow-listed
automatically; see [Cross-host endpoints](#cross-host-endpoints).
### "Login session expired"
+3
View File
@@ -138,6 +138,9 @@ SSE events are deserialized into typed dataclasses. Use `event.type` to discrimi
| `error` | `ErrorEvent` | `message` |
| `info` | `InfoEvent` | `message` |
| `stream_end` | `StreamEndEvent` | — |
| `state_change` | `StateChangeEvent` | `state``running`/`thinking`/`attention`/`idle`/`error` |
| `in_progress_snapshot` | `InProgressSnapshotEvent` | `content`, `reasoning` (one-shot mid-stream refresh resume) |
| `approval_resolved` | `ApprovalResolvedEvent` | `approved`, `feedback` |
| `cancelled` | `CancelledEvent` | — |
**Global events** (from `stream_global_events()`):
+5 -4
View File
@@ -67,10 +67,11 @@ Scopes are hierarchical — higher scopes imply all lower ones.
| Method | Path pattern | Required scope |
|--------|-------------|----------------|
| GET | Any protected path | `read` |
| POST | `/api/send`, `/api/plan`, `/api/command` | `write` |
| POST | `/api/workstreams/new`, `/api/workstreams/close` | `write` |
| POST | `/api/cluster/workstreams/new` | `write` |
| POST | `/api/approve` | `approve` |
| POST | `/api/plan`, `/api/command` | `write` |
| POST | `/api/workstreams/new`, `/api/cluster/workstreams/new` | `write` |
| POST | `/api/workstreams/{ws_id}/{send,cancel,close,delete,open,refresh-title,title,attachments}` | `write` |
| DELETE | `/api/workstreams/{ws_id}/send` (dequeue), `/api/workstreams/{ws_id}/attachments/{attachment_id}` | `write` |
| POST | `/api/workstreams/{ws_id}/approve` | `approve` |
| Any | `/api/admin/*` | `approve` |
Public paths bypass authentication entirely: `/`, `/health`, `/metrics`,
+16 -1
View File
@@ -59,6 +59,21 @@ from ConfigStore. Model names and context windows are now configured per-model
in the Models tab. A startup warning is logged if these keys appear in
`config.toml`.
### Reasoning persistence (per-model)
Two boolean flags on `model_definitions` (migration 052) control how
reasoning text round-trips per model:
| Flag | Default | Effect |
|------|---------|--------|
| `surface_persisted_reasoning` | `True` | Surface stored reasoning text on `/history` payloads so a page reload re-renders the reasoning bubble. **Storage of reasoning bytes is independent of this flag** — they ride in `provider_data` regardless. |
| `replay_reasoning_to_model` | `False` | Send stored reasoning blocks back to the provider on subsequent turns. Capability-gated: only takes effect when the model's `ModelCapabilities.supports_reasoning_replay` is also `True`. Set on canonical OpenAI gpt-5*/o-series and Anthropic Claude entries; unknown / local-server models default to `False` so an operator who flips the flag on a model whose API doesn't understand reasoning replay silently no-ops rather than 400-ing. |
Edit both via the admin Models tab. See the architecture doc for the
provider-side mechanics (Anthropic `thinking`, OpenAI Responses
`reasoning` + `include=["reasoning.encrypted_content"]`, synthetic
`reasoning_text` for Chat Completions / vLLM / llama.cpp / Gemini-compat).
### Plan / task agent overrides
`plan_agent` and `task_agent` sub-sessions resolve independently from the
@@ -100,7 +115,7 @@ initialization:
| `tools` | timeout, truncation, agent_max_turns, skip_permissions, search, search_threshold, search_max_results |
| `server` | workstream_idle_timeout, max_workstreams |
| `cluster` | node_fan_out_limit, mcp_max_servers |
| `mcp` | config_path, refresh_interval, registry_url |
| `mcp` | config_path, registry_url |
| `ratelimit` | enabled, requests_per_second, burst, trusted_proxies |
| `health` | backend_probe_interval, backend_probe_timeout, circuit_breaker_threshold, circuit_breaker_cooldown |
| `judge` | enabled, model, provider, base_url, api_key, confidence_threshold, max_context_ratio, timeout, read_only_tools, output_guard, redact_secrets, cancel_on_approval |
@@ -0,0 +1,233 @@
---
name: import-conversation-history
description: Use this skill when the user wants to import or migrate conversation history from another LLM chat or coding tool (e.g. ChatGPT, Claude.ai, Cursor, Copilot Chat, Aider, Gemini, a custom JSON export) into Turnstone. The skill teaches Turnstone's destination contracts — workstream identity, the OpenAI-shaped message rows, tool-call/result pairing, provider-fidelity blobs, attachments, and archive-vs-resumable choice — so the agent can map any source format onto them. Trigger phrases: "import my chats", "migrate this transcript into Turnstone", "bring my Claude.ai history over", "load this export as a workstream".
version: 1.0.0
---
# Importing Conversation History into Turnstone
## Overview
Source formats vary; the destination does not. Your job is to translate whatever the user hands you (JSON dump, ZIP export, scraped HTML, screenshot OCR, raw transcript) into Turnstone's internal shape: **one workstream row** plus an ordered sequence of **conversation rows** in OpenAI message format. This skill documents the destination so you can write a correct mapper for any source.
Two questions to settle with the user before writing anything:
1. **Archive or resumable?** An archive ("saved" workstream — `state="closed"`) is read-only history. A resumable workstream (`state="idle"`) lets the user continue the conversation; this only works cleanly when the source LLM matches a Turnstone-supported provider/model and tool definitions still resolve.
2. **One workstream per source thread, or merge?** Default to one-to-one unless the user explicitly asks to merge.
Default to **archive** when in doubt — resuming a foreign transcript with mismatched tool schemas or stale provider signatures will fail at the next turn.
## Turnstone Data Model (the destination)
Two tables carry the conversation:
### `workstreams` (one row per imported thread)
| Column | Required | Notes |
|---|---|---|
| `ws_id` | yes | 32-char lowercase hex. Auto-generate with `secrets.token_hex(16)` if you don't already have one. **First 4 hex chars are the routing bucket** — see "Identity & Routing" below. |
| `name` | yes | Short title. Pull from source thread title; fall back to first ~60 chars of first user message. |
| `state` | yes | `"closed"` for archive, `"idle"` for resumable. Never set `"running"` on import. |
| `kind` | yes | `"interactive"` for normal threads. Do NOT use `"coordinator"` for imports — that's reserved for cluster-spawned coordinator workstreams. |
| `parent_ws_id` | no | Leave NULL. Only set if you're importing a coordinator-spawned subtree and re-parenting it; rare. |
| `user_id` | yes | Owner. Must exist in `users`; importer must know which Turnstone user owns the imported history. |
| `node_id` | yes (multi-node) | Denormalized cache of the node that owns this `ws_id`'s bucket. Single-node deployments can leave it NULL or set it to the only node. |
| `alias` | no | Human-typeable short name. Optional; must be unique cluster-wide if set. |
| `title` | no | Auto-titled later by the LLM; safe to leave NULL on import. |
| `skill_id`, `skill_version` | yes | Default `""` and `0` unless the source thread was scoped to a Turnstone skill. |
| `created`, `updated` | yes | ISO8601 strings. Use the source's first/last message timestamps when available. |
### `conversations` (many rows per thread, ordered by `id`/`timestamp`)
| Column | Notes |
|---|---|
| `ws_id` | The workstream this row belongs to. |
| `timestamp` | ISO8601 string. Preserve source timestamps; fall back to monotonically increasing values if unknown. **Order is canonical via `id` (autoincrement), not `timestamp`** — but always insert in conversational order so both agree. |
| `role` | One of `system`, `user`, `assistant`, `tool`, `developer`. See role mapping below. |
| `content` | Text. May be NULL for assistant rows that are *only* tool calls. |
| `tool_name` | Set on `role="tool"` rows (the tool whose result this is). NULL otherwise. |
| `tool_call_id` | Set on `role="tool"` rows (matches the assistant row's `tool_calls[].id`). NULL otherwise. |
| `tool_calls` | JSON-encoded list, on `role="assistant"` rows that issued tool calls. OpenAI shape — see "Tool Calls" below. |
| `provider_data` | JSON blob preserving provider-native content blocks (Anthropic `signature`, Gemini `thought_signature`, etc.). Optional; only matters for **resumable** imports against the same provider. Skip for archives. |
The internal format is **OpenAI-shaped**, even when the source was Anthropic or Gemini. Providers translate at their own API boundary; storage stays uniform.
## Identity & Routing (`ws_id`)
- `ws_id` is **32-char lowercase hex** (i.e. `secrets.token_hex(16)`).
- The **routing bucket** is `int(ws_id[:4], 16)` — the first 4 hex chars place this workstream on a specific node via the consistent hash ring.
- For multi-node imports: either insert through the console's routing proxy (which forwards to the owning node), or generate `ws_id`s and write directly to each node's database in batches grouped by bucket.
- For single-node imports: bucket math is irrelevant; any `ws_id` works.
- **Do not reuse the source platform's IDs as `ws_id`** unless they happen to be 32-char hex. Generate fresh; if you need the old ID for traceability, store it in `workstream_config` under a key like `import.source_id`.
## Recommended Import Path
Three options, in order of preference:
### 1. Storage protocol (recommended for full history)
Use `turnstone.core.storage.Storage.save_messages_bulk(rows)`. This is the canonical bulk-insert primitive and bypasses the LLM round-trip entirely.
```python
from turnstone.core.storage import get_storage # construct via the same path the server uses
storage = get_storage(...) # see turnstone.core.storage.__init__ for the project's wiring
storage.create_workstream( # or whatever the project's exposed creator is — check turnstone/core/storage/_protocol.py
ws_id=ws_id,
user_id=user_id,
name=name,
state="closed",
kind="interactive",
...
)
storage.save_messages_bulk([
{"ws_id": ws_id, "role": "user", "content": "Hello"},
{"ws_id": ws_id, "role": "assistant", "content": "Hi! What can I help with?"},
{"ws_id": ws_id, "role": "assistant", "content": None,
"tool_calls": json.dumps([{"id": "call_1", "type": "function",
"function": {"name": "search", "arguments": "{\"q\":\"x\"}"}}])},
{"ws_id": ws_id, "role": "tool", "tool_name": "search", "tool_call_id": "call_1",
"content": "result text"},
# ...
])
```
`save_messages_bulk` handles `timestamp` and the workstream's `updated` column internally, so you don't need to compute them per row. **Verify the exact creator signature** by reading `turnstone/core/storage/_protocol.py` — table layout has shifted across migrations and the Storage protocol is the source of truth.
### 2. SDK `create_workstream(resume_ws=...)` (when the source is already a Turnstone workstream)
Only useful for *Turnstone → Turnstone* re-parenting. Not relevant for foreign sources.
### 3. SDK `create_workstream(initial_message=...)` + `send()` per turn (last resort)
Only fits archives where the source had **no tool calls** and you don't care about preserving assistant turns verbatim. Each `send()` triggers a real LLM round-trip, which is expensive and rewrites assistant content. Don't use this for full history.
## Role Mapping
Common source-role conventions and how they map to Turnstone:
| Source role | Turnstone `role` | Notes |
|---|---|---|
| `user`, `human` | `user` | Direct map. |
| `assistant`, `ai`, `model`, `bot` | `assistant` | Direct map. |
| `system` | `system` | Preserve only if it's content the user wrote (custom instructions). Drop boilerplate provider preambles — Turnstone composes its own system message. |
| `developer` (OpenAI o-series) | `developer` | Preserve. |
| `tool`, `function`, `tool_result` | `tool` | Must carry `tool_name` and `tool_call_id` matching the prior assistant row's `tool_calls[].id`. |
| `tool_use` (Anthropic) | `assistant` with `tool_calls` | Anthropic emits tool calls *inside* an assistant message; flatten to OpenAI shape. |
| `human_feedback`, `revision` | `user` | Treat as a follow-up user turn. |
## Tool Calls (the most error-prone part)
Turnstone stores tool calls in OpenAI's nested-function shape on the assistant row, and matches them with `role="tool"` result rows by `tool_call_id`.
### Assistant row with tool calls
```json
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "search_web",
"arguments": "{\"query\":\"turnstone import\"}"
}
}
]
}
```
`tool_calls[].function.arguments` is **a JSON-encoded string**, not an object. Source formats commonly get this wrong — Anthropic stores arguments as a parsed object, Gemini as a struct. Always re-serialize to a string.
### Tool result row
```json
{
"role": "tool",
"tool_name": "search_web",
"tool_call_id": "call_abc123",
"content": "..."
}
```
Pairing rules:
- Every assistant `tool_calls[].id` MUST be followed by exactly one `role="tool"` row with the matching `tool_call_id`, before the next user/assistant turn.
- If the source dropped the tool result (cut-off transcript), insert a synthetic `role="tool"` row with `content="[tool result missing in source]"` to keep the chain valid. An assistant row with an unanswered `tool_calls[].id` will break replay and any LLM round-trip.
- Multi-tool assistant turns: one `role="tool"` row per call, in any order, all before the next non-tool row.
### Tool ID generation
If the source used opaque tool IDs that aren't unique within a thread (some platforms reuse them), regenerate with a stable scheme like `f"call_{i}"` where `i` is a per-thread counter. Update both the assistant and tool rows together.
## Provider Fidelity (`provider_data`)
Skip this entirely for **archive** imports.
For **resumable** imports against the same provider, populate `provider_data` to preserve provider-specific tool-call metadata that the next API round-trip will require:
- **Anthropic**: `signature` field on thinking blocks; required for round-tripping extended-thinking responses.
- **Gemini**: `thought_signature` on tool calls; required for fidelity.
- **OpenAI**: typically nothing to preserve.
The runtime-side dict key is `_provider_content` (a list of provider-native blocks); the persisted column is `provider_data` (the same list, JSON-encoded). If you don't have provider-native blocks from the source — and you usually won't, because a foreign export won't include them — leave `provider_data` NULL. The first new turn will succeed without it, but the previous assistant turn's reasoning won't replay back to the model.
## Attachments
If the source thread had image or file attachments:
- **Size limits**: images ≤ 4 MiB, text documents ≤ 512 KiB. Reject or downsample anything bigger.
- **Allowed types**: server validates magic bytes for images and UTF-8-decodes for text. Binary blobs that aren't images won't pass.
- **Lifecycle**: pending → reserved → consumed. For imports, the cleanest path is to upload as pending and immediately consume by attaching to the relevant `conversations.id`.
Two import paths:
1. **Bulk-insert + post-attach**: insert messages first, get back the assistant/user `conversations.id`, then write `workstream_attachments` rows linking the file to `message_id`.
2. **SDK multipart create**: `create_workstream(attachments=[...], initial_message=...)` for the *first* turn only — the server reserves and consumes them onto that turn. Doesn't help for mid-thread attachments.
For full-history imports with multiple attachments at different turns, path (1) is the only option.
## Validation Checklist
Before declaring success, verify:
- [ ] `ws_id` is 32-char lowercase hex.
- [ ] `workstreams` row exists with the right `user_id`, `state`, `kind`.
- [ ] Conversation rows are inserted **in order** (autoincrement `id` will reflect insert order).
- [ ] Every assistant `tool_calls[].id` has a matching `role="tool"` row with the same `tool_call_id`.
- [ ] `tool_calls[].function.arguments` is a JSON-encoded **string**, not a parsed object.
- [ ] First message is typically `role="user"` (not `system`) — Turnstone composes its own system prompt at runtime.
- [ ] No empty assistant rows (`content=NULL` AND `tool_calls=NULL` is invalid).
- [ ] If multi-node: the `ws_id`'s bucket maps to a node that exists; `workstreams.node_id` matches.
- [ ] Round-trip test: run `Storage.load_messages(ws_id)` and confirm the reconstructed list matches what you inserted (modulo timestamps).
## Anti-patterns
- **Don't import the source provider's system prompt verbatim.** Provider boilerplate ("You are Claude...", "You are ChatGPT...") will conflict with Turnstone's composed system message and confuse the model on resume. Drop it; preserve only user-authored custom instructions.
- **Don't preserve foreign tool definitions as Turnstone tools.** If the source had custom tools that don't exist in Turnstone, the assistant rows that called them are still valid history (archive), but the workstream is **not resumable** — mark `state="closed"`.
- **Don't fabricate `tool_call_id`s without re-pairing.** Mismatched ids silently break the replay chain on the next turn.
- **Don't skip the `tool_name` field on `role="tool"` rows.** Some load paths use it for display and audit; NULL there will render as "unknown tool".
- **Don't write through the LLM (`send()` per turn) for full history.** It's expensive, rewrites assistant turns, and rate-limits will bite long imports.
## Quick Reference
| Task | Path |
|---|---|
| Generate ws_id | `secrets.token_hex(16)` |
| Bulk insert messages | `Storage.save_messages_bulk(rows)` |
| Archive (read-only) | `state="closed"`, skip `provider_data` |
| Resumable | `state="idle"`, populate `provider_data` if same provider |
| Tool call id | OpenAI shape: `{"id": ..., "type": "function", "function": {"name": ..., "arguments": "<json string>"}}` |
| Tool result row | `role="tool"`, `tool_name`, `tool_call_id`, `content` |
| Source role → Turnstone role | See "Role Mapping" table |
| Per-thread metadata | Store source IDs in `workstream_config` under `import.*` keys |
## Files to read before writing the importer
- `turnstone/core/storage/_schema.py` — authoritative table definitions.
- `turnstone/core/storage/_protocol.py``save_message`, `save_messages_bulk`, `load_messages` signatures.
- `turnstone/core/session.py` (around the message-save section) — how the runtime constructs in-memory message dicts; mirror this shape on import to round-trip cleanly.
- `turnstone/api/server_schemas.py` — Pydantic shapes for the SDK paths if you go through HTTP.
+5 -14
View File
@@ -758,22 +758,18 @@ MCP tools (3):
### Dynamic tool refresh
MCP tool lists stay up-to-date without restart through three mechanisms:
MCP tool lists stay up-to-date without restart through two mechanisms:
1. **Push notifications** -- MCP servers that declare `tools.listChanged: true` in
their capabilities send `notifications/tools/list_changed` when their tool list
changes. `MCPClientManager` registers a `message_handler` on each `ClientSession`
that triggers an immediate refresh for that server.
2. **Periodic timer** -- Servers that do *not* support push notifications are polled
on a configurable interval (default 4 hours). The timer is staggered using a
launch-time seed (`monotonic_ns ^ pid`) so cluster nodes don't all hit MCP
servers simultaneously. Configure via `[mcp] refresh_interval` in `config.toml`
or `--mcp-refresh-interval SECONDS` on the CLI. Set to `0` to disable.
3. **Manual** -- `/mcp refresh` re-fetches tools from all servers immediately.
2. **Manual** -- `/mcp refresh` re-fetches tools from all servers immediately.
`/mcp refresh <server>` targets a single server. If a server has disconnected,
manual refresh attempts reconnection.
manual refresh attempts reconnection. The console admin panel exposes the
same controls (refresh / reconnect buttons per server) for cluster-wide
fan-out.
When tools change, `MCPClientManager` rebuilds its merged tool list using copy-on-write
(new list/dict objects assigned atomically) and notifies all active `ChatSession`
@@ -781,11 +777,6 @@ instances via registered listener callbacks. Each session rebuilds its `_tools`,
`_task_tools`, `_agent_tools`, and reconstructs its `ToolSearchManager` (if active),
preserving the set of previously expanded (discovered) tools.
```toml
[mcp]
refresh_interval = 14400 # seconds (default 4h), 0 to disable
```
```
/mcp refresh
MCP refresh complete:
+4 -3
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "turnstone"
version = "1.5.0a4"
version = "1.5.12"
description = "Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation."
readme = "README.md"
license = "BUSL-1.1"
@@ -24,7 +24,7 @@ classifiers = [
dependencies = [
"openai>=2.24",
"httpx>=0.28",
"mcp>=1.6",
"mcp>=1.27",
"starlette>=0.45",
"uvicorn>=0.34",
"sse-starlette>=2.0",
@@ -35,6 +35,7 @@ dependencies = [
"structlog>=24.1",
"PyJWT>=2.8",
"bcrypt>=4.0",
"cryptography>=42",
"python-frontmatter>=1.0",
]
@@ -77,10 +78,10 @@ include = [
"turnstone/console/static/*.css",
"turnstone/console/static/*.js",
"turnstone/console/static/coordinator/*.html",
"turnstone/console/static/coordinator/*.css",
"turnstone/console/static/coordinator/*.js",
"turnstone/shared_static/*.css",
"turnstone/shared_static/*.js",
"turnstone/shared_static/design/**/*",
"turnstone/shared_static/katex-0.16.45/**/*",
"turnstone/shared_static/hljs-11.11.1/**/*",
"turnstone/shared_static/mermaid-11.14.0/**/*",
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+539 -56
View File
@@ -2,7 +2,7 @@
"openapi": "3.1.0",
"info": {
"title": "turnstone Server API",
"version": "1.5.0a2",
"version": "1.5.0a4",
"description": "Single-node workstream management, chat interaction, and real-time streaming."
},
"paths": {
@@ -55,7 +55,7 @@
"tags": [
"Workstreams"
],
"description": "Accepts two content types. Default is `application/json` with a `CreateWorkstreamRequest` body. Alternatively, `multipart/form-data` with one `meta` field (JSON-encoded `CreateWorkstreamRequest` shape) plus zero-or-more `file` parts saves each file as an attachment under the new workstream. When `initial_message` is also set, attachments are reserved onto that turn before the worker thread dispatches; otherwise they remain pending for a follow-up `POST /v1/api/send`.",
"description": "Accepts two content types. Default is `application/json` with a `CreateWorkstreamRequest` body. Alternatively, `multipart/form-data` with one `meta` field (JSON-encoded `CreateWorkstreamRequest` shape) plus zero-or-more `file` parts saves each file as an attachment under the new workstream. When `initial_message` is also set, attachments are reserved onto that turn before the worker thread dispatches; otherwise they remain pending for a follow-up `POST /v1/api/workstreams/{ws_id}/send`.",
"requestBody": {
"required": true,
"content": {
@@ -110,13 +110,23 @@
}
}
},
"/v1/api/workstreams/close": {
"/v1/api/workstreams/{ws_id}/close": {
"post": {
"summary": "Close a workstream",
"operationId": "v1_api_workstreams_close_post",
"operationId": "v1_api_workstreams_{ws_id}_close_post",
"tags": [
"Workstreams"
],
"parameters": [
{
"name": "ws_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
}
}
],
"requestBody": {
"required": true,
"content": {
@@ -147,17 +157,37 @@
}
}
}
},
"404": {
"description": "Error 404",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
}
}
}
},
"/v1/api/send": {
"/v1/api/workstreams/{ws_id}/send": {
"post": {
"summary": "Send a user message",
"operationId": "v1_api_send_post",
"operationId": "v1_api_workstreams_{ws_id}_send_post",
"tags": [
"Chat"
],
"parameters": [
{
"name": "ws_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
}
}
],
"requestBody": {
"required": true,
"content": {
@@ -200,15 +230,85 @@
}
}
}
}
},
"/v1/api/approve": {
"post": {
"summary": "Approve or deny a tool call",
"operationId": "v1_api_approve_post",
},
"delete": {
"summary": "Cancel a queued message",
"operationId": "v1_api_workstreams_{ws_id}_send_delete",
"tags": [
"Chat"
],
"description": "Removes a previously-queued message from the workstream's pending queue. Returns ``status: removed`` when the queue had the entry, ``status: not_found`` otherwise.",
"parameters": [
{
"name": "ws_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
}
}
],
"requestBody": {
"required": true,
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/DequeueRequest"
}
}
}
},
"responses": {
"200": {
"description": "Success",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/StatusResponse"
}
}
}
},
"400": {
"description": "Error 400",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"404": {
"description": "Error 404",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
}
}
}
},
"/v1/api/workstreams/{ws_id}/approve": {
"post": {
"summary": "Approve or deny a tool call",
"operationId": "v1_api_workstreams_{ws_id}_approve_post",
"tags": [
"Chat"
],
"parameters": [
{
"name": "ws_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
}
}
],
"requestBody": {
"required": true,
"content": {
@@ -335,13 +435,23 @@
}
}
},
"/v1/api/cancel": {
"/v1/api/workstreams/{ws_id}/cancel": {
"post": {
"summary": "Cancel the active generation in a workstream",
"operationId": "v1_api_cancel_post",
"operationId": "v1_api_workstreams_{ws_id}_cancel_post",
"tags": [
"Chat"
],
"parameters": [
{
"name": "ws_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
}
}
],
"requestBody": {
"required": true,
"content": {
@@ -386,10 +496,10 @@
}
}
},
"/v1/api/events": {
"/v1/api/workstreams/{ws_id}/events": {
"get": {
"summary": "Per-workstream SSE event stream",
"operationId": "v1_api_events_get",
"operationId": "v1_api_workstreams_{ws_id}_events_get",
"tags": [
"Streaming"
],
@@ -397,12 +507,11 @@
"parameters": [
{
"name": "ws_id",
"in": "query",
"in": "path",
"required": true,
"schema": {
"type": "string"
},
"description": "Workstream identifier"
}
}
],
"responses": {
@@ -623,6 +732,160 @@
}
}
},
"/v1/api/workstreams/{ws_id}": {
"get": {
"summary": "Get workstream detail (rehydrates lazily on miss)",
"operationId": "v1_api_workstreams_{ws_id}_get",
"tags": [
"Workstreams"
],
"description": "Returns the persisted workstream's display fields. If the session isn't currently in memory the manager rehydrates it before responding; ``500`` on rehydrate failure carries a correlation id matching the server log line. Lifted from the coord-only surface in the Stage 2 history/detail verb lift \u2014 interactive previously had no detail endpoint.",
"parameters": [
{
"name": "ws_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
}
}
],
"responses": {
"200": {
"description": "Success",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/WorkstreamDetailResponse"
}
}
}
},
"400": {
"description": "Error 400",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"404": {
"description": "Error 404",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"500": {
"description": "Error 500",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"503": {
"description": "Error 503",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
}
}
}
},
"/v1/api/workstreams/{ws_id}/history": {
"get": {
"summary": "Read the workstream's reconstructed message history",
"operationId": "v1_api_workstreams_{ws_id}_history_get",
"tags": [
"Workstreams"
],
"description": "Returns the tail of the conversation in OpenAI-like message format. Persisted-but-not-loaded workstreams (closed / evicted) serve history without rehydrating. Lifted from the coord-only surface in the Stage 2 history/detail verb lift \u2014 interactive previously only exposed history through the SSE replay on ``/events``.",
"parameters": [
{
"name": "ws_id",
"in": "path",
"required": true,
"schema": {
"type": "string"
}
},
{
"name": "limit",
"in": "query",
"required": false,
"schema": {
"type": "integer",
"default": 100
},
"description": "Max conversation rows to fetch from storage (default 100, max 500)."
}
],
"responses": {
"200": {
"description": "Success",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/WorkstreamHistoryResponse"
}
}
}
},
"400": {
"description": "Error 400",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"404": {
"description": "Error 404",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"500": {
"description": "Error 500",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"503": {
"description": "Error 503",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
}
}
}
},
"/v1/api/workstreams/{ws_id}/attachments": {
"post": {
"summary": "Upload a file (multipart/form-data, field 'file') and attach it to the caller's next user turn on this workstream. Validates size, MIME, and UTF-8 for text; magic-byte sniff for images. Ownership failures are masked as 404 so non-owners cannot enumerate workstream existence; a 403 indicates a scope/auth failure from the middleware layer.",
@@ -1669,11 +1932,6 @@
"title": "Message",
"type": "string"
},
"ws_id": {
"description": "Target workstream ID",
"title": "Ws Id",
"type": "string"
},
"attachment_ids": {
"anyOf": [
{
@@ -1692,8 +1950,7 @@
}
},
"required": [
"message",
"ws_id"
"message"
],
"title": "SendRequest",
"type": "object"
@@ -1760,6 +2017,21 @@
"title": "SendResponse",
"type": "object"
},
"DequeueRequest": {
"description": "Body for ``DELETE /v1/api/workstreams/{ws_id}/send``.\n\nRemoves a previously-queued message from the workstream's pending\nqueue. ``msg_id`` is the id returned in a prior ``send`` response\nwhen the workstream was busy and the message was queued.",
"properties": {
"msg_id": {
"description": "Id of the queued message to remove",
"title": "Msg Id",
"type": "string"
}
},
"required": [
"msg_id"
],
"title": "DequeueRequest",
"type": "object"
},
"ApproveRequest": {
"properties": {
"approved": {
@@ -1785,16 +2057,10 @@
"description": "Auto-approve the tools in this batch going forward",
"title": "Always",
"type": "boolean"
},
"ws_id": {
"description": "Target workstream ID",
"title": "Ws Id",
"type": "string"
}
},
"required": [
"approved",
"ws_id"
"approved"
],
"title": "ApproveRequest",
"type": "object"
@@ -1841,11 +2107,6 @@
},
"CancelRequest": {
"properties": {
"ws_id": {
"description": "Target workstream ID",
"title": "Ws Id",
"type": "string"
},
"force": {
"default": false,
"description": "Force cancel: abandon the stuck worker thread immediately. Use when cooperative cancel has not resolved within a few seconds.",
@@ -1853,9 +2114,6 @@
"type": "boolean"
}
},
"required": [
"ws_id"
],
"title": "CancelRequest",
"type": "object"
},
@@ -1931,7 +2189,7 @@
"kind": {
"$ref": "#/components/schemas/WorkstreamKind",
"default": "interactive",
"description": "Workstream kind \u2014 'interactive' (default) or 'coordinator'. Coordinator workstreams are created by the console's own /v1/api/coordinator/new endpoint; clients hitting /v1/api/workstreams/new should leave this at the default."
"description": "Workstream kind \u2014 'interactive' (default) or 'coordinator'. Coordinator workstreams are created by the console's own /v1/api/workstreams/new endpoint; clients hitting /v1/api/workstreams/new should leave this at the default."
},
"parent_ws_id": {
"anyOf": [
@@ -1984,7 +2242,7 @@
"type": "integer"
},
"attachment_ids": {
"description": "Ids of attachments saved by this request (multipart variant only). Already reserved onto the initial_message turn when one was provided; otherwise left pending for a follow-up POST /v1/api/send.",
"description": "Ids of attachments saved by this request (multipart variant only). Already reserved onto the initial_message turn when one was provided; otherwise left pending for a follow-up POST /v1/api/workstreams/{ws_id}/send.",
"items": {
"type": "string"
},
@@ -2000,20 +2258,27 @@
"type": "object"
},
"CloseWorkstreamRequest": {
"description": "Body for ``POST /v1/api/workstreams/{ws_id}/close``.\n\nThe body must be valid JSON; send ``{}`` when omitting all\nfields. Pre-1.5 the model also carried a body-keyed ``ws_id``;\n1.5 moved that to the path so the body shrinks to the optional\n``reason``. Coord ignores the body entirely (its close handler\nis wired ``supports_close_reason=False``).",
"properties": {
"ws_id": {
"description": "Workstream ID to close",
"title": "Ws Id",
"type": "string"
"reason": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Optional close reason persisted to ``workstream_config`` for postmortem. Capped at 512 UTF-8 bytes server-side; credential-redaction is applied via the output guard.",
"title": "Reason"
}
},
"required": [
"ws_id"
],
"title": "CloseWorkstreamRequest",
"type": "object"
},
"ListWorkstreamsResponse": {
"description": "Response body for ``GET /v1/api/workstreams`` on either kind.\n\nTop-level key is ``workstreams`` regardless of the kind serving\nthe request \u2014 pre-lift coord returned ``{\"coordinators\": [...]}``;\nconvergence lifted both kinds onto the same shape. Coord SDK /\nfrontend consumers branching on ``data.coordinators`` swap to\n``data.workstreams``.",
"properties": {
"workstreams": {
"items": {
@@ -2030,9 +2295,10 @@
"type": "object"
},
"WorkstreamInfo": {
"description": "Active-list row shape, shared across both kinds.\n\nRenamed ``id`` \u2192 ``ws_id`` and added ``user_id`` in the Stage 2\n``list``/``saved`` verb lift so the active-list response shape\nmatches the rest of the v1 surface (every other shared verb's\npayload uses ``ws_id``). ``user_id`` was previously coord-only;\ninteractive now populates it too. SDK consumers reading\n``row.id`` should swap to ``row.ws_id``.",
"properties": {
"id": {
"title": "Id",
"ws_id": {
"title": "Ws Id",
"type": "string"
},
"name": {
@@ -2058,16 +2324,77 @@
],
"default": null,
"title": "Parent Ws Id"
},
"user_id": {
"default": "",
"title": "User Id",
"type": "string"
}
},
"required": [
"id",
"ws_id",
"name",
"state"
],
"title": "WorkstreamInfo",
"type": "object"
},
"WorkstreamDetailResponse": {
"description": "Response body for ``GET /v1/api/workstreams/{ws_id}``.\n\nRenamed and relocated from ``CoordinatorDetailResponse`` in the\nStage 2 history/detail verb lift. Both kinds populate every field;\nSDK consumers don't branch on kind to read them. The lift adds the\nendpoint to interactive as a feature gain (pre-lift only coord\nexposed it).",
"properties": {
"ws_id": {
"title": "Ws Id",
"type": "string"
},
"name": {
"title": "Name",
"type": "string"
},
"state": {
"title": "State",
"type": "string"
},
"user_id": {
"title": "User Id",
"type": "string"
},
"kind": {
"$ref": "#/components/schemas/WorkstreamKind",
"default": "interactive"
}
},
"required": [
"ws_id",
"name",
"state",
"user_id"
],
"title": "WorkstreamDetailResponse",
"type": "object"
},
"WorkstreamHistoryResponse": {
"description": "Response body for ``GET /v1/api/workstreams/{ws_id}/history``.\n\nRenamed and relocated from ``CoordinatorHistoryResponse`` in the\nStage 2 history/detail verb lift. Same OpenAI-like message-row\nshape on both kinds; the lift adds the endpoint to interactive as\na feature gain (pre-lift interactive only exposed history through\nthe SSE replay on ``/events``).",
"properties": {
"ws_id": {
"title": "Ws Id",
"type": "string"
},
"messages": {
"description": "Tail of the workstream's reconstructed message history (provider-fidelity OpenAI-like shape). Bounded by the ``limit`` query parameter (default 100, max 500).",
"items": {
"additionalProperties": true,
"type": "object"
},
"title": "Messages",
"type": "array"
}
},
"required": [
"ws_id"
],
"title": "WorkstreamHistoryResponse",
"type": "object"
},
"DashboardResponse": {
"properties": {
"workstreams": {
@@ -2125,9 +2452,10 @@
"type": "object"
},
"DashboardWorkstream": {
"description": "Dashboard row shape for ``GET /v1/api/dashboard``.\n\nRenamed ``id`` \u2192 ``ws_id`` for v1 row-shape consistency with\nthe rest of the workstream surface (active list, saved list,\nhistory, detail, etc.). Frontend consumers reading\n``dashboard.workstreams[].id`` swap to ``.ws_id``.",
"properties": {
"id": {
"title": "Id",
"ws_id": {
"title": "Ws Id",
"type": "string"
},
"name": {
@@ -2203,16 +2531,171 @@
"default": "",
"title": "User Id",
"type": "string"
},
"pending_approval_detail": {
"anyOf": [
{
"$ref": "#/components/schemas/PendingApprovalDetail"
},
{
"type": "null"
}
],
"default": null,
"description": "Inline approval payload for the coordinator children-tree UI. Carries the merged ``_pending_approval`` items list + per-call_id LLM verdict cache so a coord can render approve/deny buttons + judge pill without a separate per-child round-trip. ``None`` when no approval is pending. Also surfaced (verbatim) on ``GET /v1/api/cluster/ws/live`` via the ``_CLUSTER_WS_LIVE_KEYS`` projection."
},
"recent_auto_approvals": {
"description": "Per-ws ring buffer (cap 10) of recent tool calls that bypassed the operator approval gate. Surfaces ``WebUI._recent_auto_approvals`` so the coord-tree row can render an 'auto-approved by ...' pill when the child's skill / blanket / admin-policy rules silently let a tool through. Also projected onto ``GET /v1/api/cluster/ws/live`` via ``_CLUSTER_WS_LIVE_KEYS``.",
"items": {
"$ref": "#/components/schemas/RecentAutoApproval"
},
"title": "Recent Auto Approvals",
"type": "array"
}
},
"required": [
"id",
"ws_id",
"name",
"state"
],
"title": "DashboardWorkstream",
"type": "object"
},
"PendingApprovalDetail": {
"description": "Inline approval payload merged into ``DashboardWorkstream``.\n\nSet when a workstream's ``approve_tools`` is parked on\n``_approval_event``; ``None`` (omitted) otherwise. Cross-tenant\nexposure here follows the same trusted-team posture as\n``activity`` / ``tokens`` \u2014 see ``server.py``'s ``dashboard``\nhandler comment.",
"properties": {
"call_id": {
"default": "",
"description": "Primary call_id \u2014 first non-empty call_id in items list order. Matches the 409 ``current_call_id`` response from ``POST /v1/api/workstreams/{ws_id}/approve`` so the UI can render the same identifier the server reports as current.",
"title": "Call Id",
"type": "string"
},
"judge_pending": {
"default": false,
"description": "LLM judge tier still running; heuristic verdicts may already be present on items.",
"title": "Judge Pending",
"type": "boolean"
},
"items": {
"items": {
"$ref": "#/components/schemas/PendingApprovalItem"
},
"title": "Items",
"type": "array"
}
},
"title": "PendingApprovalDetail",
"type": "object"
},
"PendingApprovalItem": {
"description": "One pending tool-call inside a ``PendingApprovalDetail`` envelope.\n\nMirrors the dict ``SessionUIBase.serialize_pending_approval_detail``\nemits per item. ``heuristic_verdict`` / ``judge_verdict`` are kept\nloosely-typed because the underlying verdict shape varies by tier;\nconsumers that want the full structure can decode against\n:class:`turnstone.sdk.events.IntentVerdictEvent`.",
"properties": {
"call_id": {
"default": "",
"title": "Call Id",
"type": "string"
},
"header": {
"default": "",
"title": "Header",
"type": "string"
},
"preview": {
"default": "",
"title": "Preview",
"type": "string"
},
"func_name": {
"default": "",
"title": "Func Name",
"type": "string"
},
"approval_label": {
"default": "",
"title": "Approval Label",
"type": "string"
},
"needs_approval": {
"default": false,
"title": "Needs Approval",
"type": "boolean"
},
"error": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Error"
},
"heuristic_verdict": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Heuristic Verdict"
},
"judge_verdict": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Judge Verdict"
}
},
"title": "PendingApprovalItem",
"type": "object"
},
"RecentAutoApproval": {
"description": "One ring-buffer entry for ``DashboardWorkstream.recent_auto_approvals``.\n\nRecords a tool call that bypassed the operator approval gate\n(admin tool policy / skill ``allowed_tools`` allowlist / blanket\n``auto_approve`` / \"Approve + Always\" memory). The coord-tree\npill reads this list to surface \"auto-approved by skill X\" so\nthe operator can see WHICH calls bypassed and WHY.",
"properties": {
"call_id": {
"default": "",
"title": "Call Id",
"type": "string"
},
"func_name": {
"default": "",
"title": "Func Name",
"type": "string"
},
"approval_label": {
"default": "",
"title": "Approval Label",
"type": "string"
},
"auto_approve_reason": {
"default": "",
"description": "Source that fired the bypass. ``skill`` (skill template's ``allowed_tools``), ``always`` (user 'Approve + Always' click), ``policy`` (admin tool-policy ``allow`` rule), ``blanket`` (workstream-level ``auto_approve=True``), or ``auto_approve_tools`` (legacy / unknown writer).",
"title": "Auto Approve Reason",
"type": "string"
},
"ts": {
"default": 0.0,
"description": "Unix epoch seconds when the auto-approve fired.",
"title": "Ts",
"type": "number"
}
},
"title": "RecentAutoApproval",
"type": "object"
},
"ListSavedWorkstreamsResponse": {
"properties": {
"workstreams": {
+148 -148
View File
@@ -1,12 +1,12 @@
{
"name": "@turnstone/sdk",
"version": "0.3.0",
"version": "0.4.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "@turnstone/sdk",
"version": "0.3.0",
"version": "0.4.0",
"license": "BUSL-1.1",
"devDependencies": {
"typescript": "^6.0.0",
@@ -14,9 +14,9 @@
}
},
"node_modules/@emnapi/core": {
"version": "1.9.2",
"resolved": "https://registry.npmjs.org/@emnapi/core/-/core-1.9.2.tgz",
"integrity": "sha512-UC+ZhH3XtczQYfOlu3lNEkdW/p4dsJ1r/bP7H8+rhao3TTTMO1ATq/4DdIi23XuGoFY+Cz0JmCbdVl0hz9jZcA==",
"version": "1.10.0",
"resolved": "https://registry.npmjs.org/@emnapi/core/-/core-1.10.0.tgz",
"integrity": "sha512-yq6OkJ4p82CAfPl0u9mQebQHKPJkY7WrIuk205cTYnYe+k2Z8YBh11FrbRG/H6ihirqcacOgl2BIO8oyMQLeXw==",
"dev": true,
"license": "MIT",
"optional": true,
@@ -26,9 +26,9 @@
}
},
"node_modules/@emnapi/runtime": {
"version": "1.9.2",
"resolved": "https://registry.npmjs.org/@emnapi/runtime/-/runtime-1.9.2.tgz",
"integrity": "sha512-3U4+MIWHImeyu1wnmVygh5WlgfYDtyf0k8AbLhMFxOipihf6nrWC4syIm/SwEeec0mNSafiiNnMJwbza/Is6Lw==",
"version": "1.10.0",
"resolved": "https://registry.npmjs.org/@emnapi/runtime/-/runtime-1.10.0.tgz",
"integrity": "sha512-ewvYlk86xUoGI0zQRNq/mC+16R1QeDlKQy21Ki3oSYXNgLb45GV1P6A0M+/s6nyCuNDqe5VpaY84BzXGwVbwFA==",
"dev": true,
"license": "MIT",
"optional": true,
@@ -74,9 +74,9 @@
}
},
"node_modules/@oxc-project/types": {
"version": "0.124.0",
"resolved": "https://registry.npmjs.org/@oxc-project/types/-/types-0.124.0.tgz",
"integrity": "sha512-VBFWMTBvHxS11Z5Lvlr3IWgrwhMTXV+Md+EQF0Xf60+wAdsGFTBx7X7K/hP4pi8N7dcm1RvcHwDxZ16Qx8keUg==",
"version": "0.127.0",
"resolved": "https://registry.npmjs.org/@oxc-project/types/-/types-0.127.0.tgz",
"integrity": "sha512-aIYXQBo4lCbO4z0R3FHeucQHpF46l2LbMdxRvqvuRuW2OxdnSkcng5B8+K12spgLDj93rtN3+J2Vac/TIO+ciQ==",
"dev": true,
"license": "MIT",
"funding": {
@@ -84,9 +84,9 @@
}
},
"node_modules/@rolldown/binding-android-arm64": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-android-arm64/-/binding-android-arm64-1.0.0-rc.15.tgz",
"integrity": "sha512-YYe6aWruPZDtHNpwu7+qAHEMbQ/yRl6atqb/AhznLTnD3UY99Q1jE7ihLSahNWkF4EqRPVC4SiR4O0UkLK02tA==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-android-arm64/-/binding-android-arm64-1.0.0-rc.17.tgz",
"integrity": "sha512-s70pVGhw4zqGeFnXWvAzJDlvxhlRollagdCCKRgOsgUOH3N1l0LIxf83AtGzmb5SiVM4Hjl5HyarMRfdfj3DaQ==",
"cpu": [
"arm64"
],
@@ -101,9 +101,9 @@
}
},
"node_modules/@rolldown/binding-darwin-arm64": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-darwin-arm64/-/binding-darwin-arm64-1.0.0-rc.15.tgz",
"integrity": "sha512-oArR/ig8wNTPYsXL+Mzhs0oxhxfuHRfG7Ikw7jXsw8mYOtk71W0OkF2VEVh699pdmzjPQsTjlD1JIOoHkLP1Fg==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-darwin-arm64/-/binding-darwin-arm64-1.0.0-rc.17.tgz",
"integrity": "sha512-4ksWc9n0mhlZpZ9PMZgTGjeOPRu8MB1Z3Tz0Mo02eWfWCHMW1zN82Qz/pL/rC+yQa+8ZnutMF0JjJe7PjwasYw==",
"cpu": [
"arm64"
],
@@ -118,9 +118,9 @@
}
},
"node_modules/@rolldown/binding-darwin-x64": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-darwin-x64/-/binding-darwin-x64-1.0.0-rc.15.tgz",
"integrity": "sha512-YzeVqOqjPYvUbJSWJ4EDL8ahbmsIXQpgL3JVipmN+MX0XnXMeWomLN3Fb+nwCmP/jfyqte5I3XRSm7OfQrbyxw==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-darwin-x64/-/binding-darwin-x64-1.0.0-rc.17.tgz",
"integrity": "sha512-SUSDOI6WwUVNcWxd02QEBjLdY1VPHvlEkw6T/8nYG322iYWCTxRb1vzk4E+mWWYehTp7ERibq54LSJGjmouOsw==",
"cpu": [
"x64"
],
@@ -135,9 +135,9 @@
}
},
"node_modules/@rolldown/binding-freebsd-x64": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-freebsd-x64/-/binding-freebsd-x64-1.0.0-rc.15.tgz",
"integrity": "sha512-9Erhx956jeQ0nNTyif1+QWAXDRD38ZNjr//bSHrt6wDwB+QkAfl2q6Mn1k6OBPerznjRmbM10lgRb1Pli4xZPw==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-freebsd-x64/-/binding-freebsd-x64-1.0.0-rc.17.tgz",
"integrity": "sha512-hwnz3nw9dbJ05EDO/PvcjaaewqqDy7Y1rn1UO81l8iIK1GjenME75dl16ajbvSSMfv66WXSRCYKIqfgq2KCfxw==",
"cpu": [
"x64"
],
@@ -152,9 +152,9 @@
}
},
"node_modules/@rolldown/binding-linux-arm-gnueabihf": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-arm-gnueabihf/-/binding-linux-arm-gnueabihf-1.0.0-rc.15.tgz",
"integrity": "sha512-cVwk0w8QbZJGTnP/AHQBs5yNwmpgGYStL88t4UIaqcvYJWBfS0s3oqVLZPwsPU6M0zlW4GqjP0Zq5MnAGwFeGA==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-arm-gnueabihf/-/binding-linux-arm-gnueabihf-1.0.0-rc.17.tgz",
"integrity": "sha512-IS+W7epTcwANmFSQFrS1SivEXHtl1JtuQA9wlxrZTcNi6mx+FDOYrakGevvvTwgj2JvWiK8B29/qD9BELZPyXQ==",
"cpu": [
"arm"
],
@@ -169,9 +169,9 @@
}
},
"node_modules/@rolldown/binding-linux-arm64-gnu": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-arm64-gnu/-/binding-linux-arm64-gnu-1.0.0-rc.15.tgz",
"integrity": "sha512-eBZ/u8iAK9SoHGanqe/jrPnY0JvBN6iXbVOsbO38mbz+ZJsaobExAm1Iu+rxa4S1l2FjG0qEZn4Rc6X8n+9M+w==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-arm64-gnu/-/binding-linux-arm64-gnu-1.0.0-rc.17.tgz",
"integrity": "sha512-e6usGaHKW5BMNZOymS1UcEYGowQMWcgZ71Z17Sl/h2+ZziNJ1a9n3Zvcz6LdRyIW5572wBCTH/Z+bKuZouGk9Q==",
"cpu": [
"arm64"
],
@@ -189,9 +189,9 @@
}
},
"node_modules/@rolldown/binding-linux-arm64-musl": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-arm64-musl/-/binding-linux-arm64-musl-1.0.0-rc.15.tgz",
"integrity": "sha512-ZvRYMGrAklV9PEkgt4LQM6MjQX2P58HPAuecwYObY2DhS2t35R0I810bKi0wmaYORt6m/2Sm+Z+nFgb0WhXNcQ==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-arm64-musl/-/binding-linux-arm64-musl-1.0.0-rc.17.tgz",
"integrity": "sha512-b/CgbwAJpmrRLp02RPfhbudf5tZnN9nsPWK82znefso832etkem8H7FSZwxrOI9djcdTP7U6YfNhbRnh7djErg==",
"cpu": [
"arm64"
],
@@ -209,9 +209,9 @@
}
},
"node_modules/@rolldown/binding-linux-ppc64-gnu": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-ppc64-gnu/-/binding-linux-ppc64-gnu-1.0.0-rc.15.tgz",
"integrity": "sha512-VDpgGBzgfg5hLg+uBpCLoFG5kVvEyafmfxGUV0UHLcL5irxAK7PKNeC2MwClgk6ZAiNhmo9FLhRYgvMmedLtnQ==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-ppc64-gnu/-/binding-linux-ppc64-gnu-1.0.0-rc.17.tgz",
"integrity": "sha512-4EII1iNGRUN5WwGbF/kOh/EIkoDN9HsupgLQoXfY+D1oyJm7/F4t5PYU5n8SWZgG0FEwakyM8pGgwcBYruGTlA==",
"cpu": [
"ppc64"
],
@@ -229,9 +229,9 @@
}
},
"node_modules/@rolldown/binding-linux-s390x-gnu": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-s390x-gnu/-/binding-linux-s390x-gnu-1.0.0-rc.15.tgz",
"integrity": "sha512-y1uXY3qQWCzcPgRJATPSOUP4tCemh4uBdY7e3EZbVwCJTY3gLJWnQABgeUetvED+bt1FQ01OeZwvhLS2bpNrAQ==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-s390x-gnu/-/binding-linux-s390x-gnu-1.0.0-rc.17.tgz",
"integrity": "sha512-AH8oq3XqQo4IibpVXvPeLDI5pzkpYn0WiZAfT05kFzoJ6tQNzwRdDYQ45M8I/gslbodRZwW8uxLhbSBbkv96rA==",
"cpu": [
"s390x"
],
@@ -249,9 +249,9 @@
}
},
"node_modules/@rolldown/binding-linux-x64-gnu": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-x64-gnu/-/binding-linux-x64-gnu-1.0.0-rc.15.tgz",
"integrity": "sha512-023bTPBod7J3Y/4fzAN6QtpkSABR0rigtrwaP+qSEabUh5zf6ELr9Nc7GujaROuPY3uwdSIXWrvhn1KxOvurWA==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-x64-gnu/-/binding-linux-x64-gnu-1.0.0-rc.17.tgz",
"integrity": "sha512-cLnjV3xfo7KslbU41Z7z8BH/E1y5mzUYzAqih1d1MDaIGZRCMqTijqLv76/P7fyHuvUcfGsIpqCdddbxLLK9rA==",
"cpu": [
"x64"
],
@@ -269,9 +269,9 @@
}
},
"node_modules/@rolldown/binding-linux-x64-musl": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-x64-musl/-/binding-linux-x64-musl-1.0.0-rc.15.tgz",
"integrity": "sha512-witB2O0/hU4CgfOOKUoeFgQ4GktPi1eEbAhaLAIpgD6+ZnhcPkUtPsoKKHRzmOoWPZue46IThdSgdo4XneOLYw==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-linux-x64-musl/-/binding-linux-x64-musl-1.0.0-rc.17.tgz",
"integrity": "sha512-0phclDw1spsL7dUB37sIARuis2tAgomCJXAHZlpt8PXZ4Ba0dRP1e+66lsRqrfhISeN9bEGNjQs+T/Fbd7oYGw==",
"cpu": [
"x64"
],
@@ -289,9 +289,9 @@
}
},
"node_modules/@rolldown/binding-openharmony-arm64": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-openharmony-arm64/-/binding-openharmony-arm64-1.0.0-rc.15.tgz",
"integrity": "sha512-UCL68NJ0Ud5zRipXZE9dF5PmirzJE4E4BCIOOssEnM7wLDsxjc6Qb0sGDxTNRTP53I6MZpygyCpY8Aa8sPfKPg==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-openharmony-arm64/-/binding-openharmony-arm64-1.0.0-rc.17.tgz",
"integrity": "sha512-0ag/hEgXOwgw4t8QyQvUCxvEg+V0KBcA6YuOx9g0r02MprutRF5dyljgm3EmR02O292UX7UeS6HzWHAl6KgyhA==",
"cpu": [
"arm64"
],
@@ -306,9 +306,9 @@
}
},
"node_modules/@rolldown/binding-wasm32-wasi": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-wasm32-wasi/-/binding-wasm32-wasi-1.0.0-rc.15.tgz",
"integrity": "sha512-ApLruZq/ig+nhaE7OJm4lDjayUnOHVUa77zGeqnqZ9pn0ovdVbbNPerVibLXDmWeUZXjIYIT8V3xkT58Rm9u5Q==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-wasm32-wasi/-/binding-wasm32-wasi-1.0.0-rc.17.tgz",
"integrity": "sha512-LEXei6vo0E5wTGwpkJ4KoT3OZJRnglwldt5ziLzOlc6qqb55z4tWNq2A+PFqCJuvWWdP53CVhG1Z9NtToDPJrA==",
"cpu": [
"wasm32"
],
@@ -316,18 +316,18 @@
"license": "MIT",
"optional": true,
"dependencies": {
"@emnapi/core": "1.9.2",
"@emnapi/runtime": "1.9.2",
"@napi-rs/wasm-runtime": "^1.1.3"
"@emnapi/core": "1.10.0",
"@emnapi/runtime": "1.10.0",
"@napi-rs/wasm-runtime": "^1.1.4"
},
"engines": {
"node": ">=14.0.0"
"node": "^20.19.0 || >=22.12.0"
}
},
"node_modules/@rolldown/binding-win32-arm64-msvc": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-win32-arm64-msvc/-/binding-win32-arm64-msvc-1.0.0-rc.15.tgz",
"integrity": "sha512-KmoUoU7HnN+Si5YWJigfTws1jz1bKBYDQKdbLspz0UaqjjFkddHsqorgiW1mxcAj88lYUE6NC/zJNwT+SloqtA==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-win32-arm64-msvc/-/binding-win32-arm64-msvc-1.0.0-rc.17.tgz",
"integrity": "sha512-gUmyzBl3SPMa6hrqFUth9sVfcLBlYsbMzBx5PlexMroZStgzGqlZ26pYG89rBb45Mnia+oil6YAIFeEWGWhoZA==",
"cpu": [
"arm64"
],
@@ -342,9 +342,9 @@
}
},
"node_modules/@rolldown/binding-win32-x64-msvc": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/binding-win32-x64-msvc/-/binding-win32-x64-msvc-1.0.0-rc.15.tgz",
"integrity": "sha512-3P2A8L+x75qavWLe/Dll3EYBJLQmtkJN8rfh+U/eR3MqMgL/h98PhYI+JFfXuDPgPeCB7iZAKiqii5vqOvnA0g==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/binding-win32-x64-msvc/-/binding-win32-x64-msvc-1.0.0-rc.17.tgz",
"integrity": "sha512-3hkiolcUAvPB9FLb3UZdfjVVNWherN1f/skkGWJP/fgSQhYUZpSIRr0/I8ZK9TkF3F7kxvJAk0+IcKvPHk9qQg==",
"cpu": [
"x64"
],
@@ -359,9 +359,9 @@
}
},
"node_modules/@rolldown/pluginutils": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/@rolldown/pluginutils/-/pluginutils-1.0.0-rc.15.tgz",
"integrity": "sha512-UromN0peaE53IaBRe9W7CjrZgXl90fqGpK+mIZbA3qSTeYqg3pqpROBdIPvOG3F5ereDHNwoHBI2e50n1BDr1g==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/@rolldown/pluginutils/-/pluginutils-1.0.0-rc.17.tgz",
"integrity": "sha512-n8iosDOt6Ig1UhJ2AYqoIhHWh/isz0xpicHTzpKBeotdVsTEcxsSA/i3EVM7gQAj0rU27OLAxCjzlj15IWY7bg==",
"dev": true,
"license": "MIT"
},
@@ -373,9 +373,9 @@
"license": "MIT"
},
"node_modules/@tybys/wasm-util": {
"version": "0.10.1",
"resolved": "https://registry.npmjs.org/@tybys/wasm-util/-/wasm-util-0.10.1.tgz",
"integrity": "sha512-9tTaPJLSiejZKx+Bmog4uSubteqTvFrVrURwkmHixBo0G4seD0zUxp98E1DzUBJxLQ3NPwXrGKDiVjwx/DpPsg==",
"version": "0.10.2",
"resolved": "https://registry.npmjs.org/@tybys/wasm-util/-/wasm-util-0.10.2.tgz",
"integrity": "sha512-RoBvJ2X0wuKlWFIjrwffGw1IqZHKQqzIchKaadZZfnNpsAYp2mM0h36JtPCjNDAHGgYez/15uMBpfGwchhiMgg==",
"dev": true,
"license": "MIT",
"optional": true,
@@ -409,16 +409,16 @@
"license": "MIT"
},
"node_modules/@vitest/expect": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/@vitest/expect/-/expect-4.1.4.tgz",
"integrity": "sha512-iPBpra+VDuXmBFI3FMKHSFXp3Gx5HfmSCE8X67Dn+bwephCnQCaB7qWK2ldHa+8ncN8hJU8VTMcxjPpyMkUjww==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/@vitest/expect/-/expect-4.1.5.tgz",
"integrity": "sha512-PWBaRY5JoKuRnHlUHfpV/KohFylaDZTupcXN1H9vYryNLOnitSw60Mw9IAE2r67NbwwzBw/Cc/8q9BK3kIX8Kw==",
"dev": true,
"license": "MIT",
"dependencies": {
"@standard-schema/spec": "^1.1.0",
"@types/chai": "^5.2.2",
"@vitest/spy": "4.1.4",
"@vitest/utils": "4.1.4",
"@vitest/spy": "4.1.5",
"@vitest/utils": "4.1.5",
"chai": "^6.2.2",
"tinyrainbow": "^3.1.0"
},
@@ -427,13 +427,13 @@
}
},
"node_modules/@vitest/mocker": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/@vitest/mocker/-/mocker-4.1.4.tgz",
"integrity": "sha512-R9HTZBhW6yCSGbGQnDnH3QHfJxokKN4KB+Yvk9Q1le7eQNYwiCyKxmLmurSpFy6BzJanSLuEUDrD+j97Q+ZLPg==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/@vitest/mocker/-/mocker-4.1.5.tgz",
"integrity": "sha512-/x2EmFC4mT4NNzqvC3fmesuV97w5FC903KPmey4gsnJiMQ3Be1IlDKVaDaG8iqaLFHqJ2FVEkxZk5VmeLjIItw==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/spy": "4.1.4",
"@vitest/spy": "4.1.5",
"estree-walker": "^3.0.3",
"magic-string": "^0.30.21"
},
@@ -454,9 +454,9 @@
}
},
"node_modules/@vitest/pretty-format": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/@vitest/pretty-format/-/pretty-format-4.1.4.tgz",
"integrity": "sha512-ddmDHU0gjEUyEVLxtZa7xamrpIefdEETu3nZjWtHeZX4QxqJ7tRxSteHVXJOcr8jhiLoGAhkK4WJ3WqBpjx42A==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/@vitest/pretty-format/-/pretty-format-4.1.5.tgz",
"integrity": "sha512-7I3q6l5qr03dVfMX2wCo9FxwSJbPdwKjy2uu/YPpU3wfHvIL4QHwVRp57OfGrDFeUJ8/8QdfBKIV12FTtLn00g==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -467,13 +467,13 @@
}
},
"node_modules/@vitest/runner": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/@vitest/runner/-/runner-4.1.4.tgz",
"integrity": "sha512-xTp7VZ5aXP5ZJrn15UtJUWlx6qXLnGtF6jNxHepdPHpMfz/aVPx+htHtgcAL2mDXJgKhpoo2e9/hVJsIeFbytQ==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/@vitest/runner/-/runner-4.1.5.tgz",
"integrity": "sha512-2D+o7Pr82IEO46YPpoA/YU0neeyr6FTerQb5Ro7BUnBuv6NQtT/kmVnczngiMEBhzgqz2UZYl5gArejsyERDSQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/utils": "4.1.4",
"@vitest/utils": "4.1.5",
"pathe": "^2.0.3"
},
"funding": {
@@ -481,14 +481,14 @@
}
},
"node_modules/@vitest/snapshot": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/@vitest/snapshot/-/snapshot-4.1.4.tgz",
"integrity": "sha512-MCjCFgaS8aZz+m5nTcEcgk/xhWv0rEH4Yl53PPlMXOZ1/Ka2VcZU6CJ+MgYCZbcJvzGhQRjVrGQNZqkGPttIKw==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/@vitest/snapshot/-/snapshot-4.1.5.tgz",
"integrity": "sha512-zypXEt4KH/XgKGPUz4eC2AvErYx0My5hfL8oDb1HzGFpEk1P62bxSohdyOmvz+d9UJwanI68MKwr2EquOaOgMQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/pretty-format": "4.1.4",
"@vitest/utils": "4.1.4",
"@vitest/pretty-format": "4.1.5",
"@vitest/utils": "4.1.5",
"magic-string": "^0.30.21",
"pathe": "^2.0.3"
},
@@ -497,9 +497,9 @@
}
},
"node_modules/@vitest/spy": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/@vitest/spy/-/spy-4.1.4.tgz",
"integrity": "sha512-XxNdAsKW7C+FLydqFJLb5KhJtl3PGCMmYwFRfhvIgxJvLSXhhVI1zM8f1qD3Zg7RCjTSzDVyct6sghs9UEgBEQ==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/@vitest/spy/-/spy-4.1.5.tgz",
"integrity": "sha512-2lNOsh6+R2Idnf1TCZqSwYlKN2E/iDlD8sgU59kYVl+OMDmvldO1VDk39smRfpUNwYpNRVn3w4YfuC7KfbBnkQ==",
"dev": true,
"license": "MIT",
"funding": {
@@ -507,13 +507,13 @@
}
},
"node_modules/@vitest/utils": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/@vitest/utils/-/utils-4.1.4.tgz",
"integrity": "sha512-13QMT+eysM5uVGa1rG4kegGYNp6cnQcsTc67ELFbhNLQO+vgsygtYJx2khvdt4gVQqSSpC/KT5FZZxUpP3Oatw==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/@vitest/utils/-/utils-4.1.5.tgz",
"integrity": "sha512-76wdkrmfXfqGjueGgnb45ITPyUi1ycZ4IHgC2bhPDUfWHklY/q3MdLOAB+TF1e6xfl8NxNY0ZYaPCFNWSsw3Ug==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/pretty-format": "4.1.4",
"@vitest/pretty-format": "4.1.5",
"convert-source-map": "^2.0.0",
"tinyrainbow": "^3.1.0"
},
@@ -559,9 +559,9 @@
}
},
"node_modules/es-module-lexer": {
"version": "2.0.0",
"resolved": "https://registry.npmjs.org/es-module-lexer/-/es-module-lexer-2.0.0.tgz",
"integrity": "sha512-5POEcUuZybH7IdmGsD8wlf0AI55wMecM9rVBTI/qEAy2c1kTOm3DjFYjrBdI2K3BaJjJYfYFeRtM0t9ssnRuxw==",
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/es-module-lexer/-/es-module-lexer-2.1.0.tgz",
"integrity": "sha512-n27zTYMjYu1aj4MjCWzSP7G9r75utsaoc8m61weK+W8JMBGGQybd43GstCXZ3WNmSFtGT9wi59qQTW6mhTR5LQ==",
"dev": true,
"license": "MIT"
},
@@ -902,9 +902,9 @@
}
},
"node_modules/nanoid": {
"version": "3.3.11",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.11.tgz",
"integrity": "sha512-N8SpfPUnUp1bK+PMYW8qSWdl9U+wwNWI4QKxOYDy9JAro3WMX7p2OeVRF9v+347pnakNevPmiHhNmZ2HbFA76w==",
"version": "3.3.12",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.12.tgz",
"integrity": "sha512-ZB9RH/39qpq5Vu6Y+NmUaFhQR6pp+M2Xt76XBnEwDaGcVAqhlvxrl3B2bKS5D3NH3QR76v3aSrKaF/Kiy7lEtQ==",
"dev": true,
"funding": [
{
@@ -959,9 +959,9 @@
}
},
"node_modules/postcss": {
"version": "8.5.10",
"resolved": "https://registry.npmjs.org/postcss/-/postcss-8.5.10.tgz",
"integrity": "sha512-pMMHxBOZKFU6HgAZ4eyGnwXF/EvPGGqUr0MnZ5+99485wwW41kW91A4LOGxSHhgugZmSChL5AlElNdwlNgcnLQ==",
"version": "8.5.13",
"resolved": "https://registry.npmjs.org/postcss/-/postcss-8.5.13.tgz",
"integrity": "sha512-qif0+jGGZoLWdHey3UFHHWP0H7Gbmsk8T5VEqyYFbWqPr1XqvLGBbk/sl8V5exGmcYJklJOhOQq1pV9IcsiFag==",
"dev": true,
"funding": [
{
@@ -988,14 +988,14 @@
}
},
"node_modules/rolldown": {
"version": "1.0.0-rc.15",
"resolved": "https://registry.npmjs.org/rolldown/-/rolldown-1.0.0-rc.15.tgz",
"integrity": "sha512-Ff31guA5zT6WjnGp0SXw76X6hzGRk/OQq2hE+1lcDe+lJdHSgnSX6nK3erbONHyCbpSj9a9E+uX/OvytZoWp2g==",
"version": "1.0.0-rc.17",
"resolved": "https://registry.npmjs.org/rolldown/-/rolldown-1.0.0-rc.17.tgz",
"integrity": "sha512-ZrT53oAKrtA4+YtBWPQbtPOxIbVDbxT0orcYERKd63VJTF13zPcgXTvD4843L8pcsI7M6MErt8QtON6lrB9tyA==",
"dev": true,
"license": "MIT",
"dependencies": {
"@oxc-project/types": "=0.124.0",
"@rolldown/pluginutils": "1.0.0-rc.15"
"@oxc-project/types": "=0.127.0",
"@rolldown/pluginutils": "1.0.0-rc.17"
},
"bin": {
"rolldown": "bin/cli.mjs"
@@ -1004,21 +1004,21 @@
"node": "^20.19.0 || >=22.12.0"
},
"optionalDependencies": {
"@rolldown/binding-android-arm64": "1.0.0-rc.15",
"@rolldown/binding-darwin-arm64": "1.0.0-rc.15",
"@rolldown/binding-darwin-x64": "1.0.0-rc.15",
"@rolldown/binding-freebsd-x64": "1.0.0-rc.15",
"@rolldown/binding-linux-arm-gnueabihf": "1.0.0-rc.15",
"@rolldown/binding-linux-arm64-gnu": "1.0.0-rc.15",
"@rolldown/binding-linux-arm64-musl": "1.0.0-rc.15",
"@rolldown/binding-linux-ppc64-gnu": "1.0.0-rc.15",
"@rolldown/binding-linux-s390x-gnu": "1.0.0-rc.15",
"@rolldown/binding-linux-x64-gnu": "1.0.0-rc.15",
"@rolldown/binding-linux-x64-musl": "1.0.0-rc.15",
"@rolldown/binding-openharmony-arm64": "1.0.0-rc.15",
"@rolldown/binding-wasm32-wasi": "1.0.0-rc.15",
"@rolldown/binding-win32-arm64-msvc": "1.0.0-rc.15",
"@rolldown/binding-win32-x64-msvc": "1.0.0-rc.15"
"@rolldown/binding-android-arm64": "1.0.0-rc.17",
"@rolldown/binding-darwin-arm64": "1.0.0-rc.17",
"@rolldown/binding-darwin-x64": "1.0.0-rc.17",
"@rolldown/binding-freebsd-x64": "1.0.0-rc.17",
"@rolldown/binding-linux-arm-gnueabihf": "1.0.0-rc.17",
"@rolldown/binding-linux-arm64-gnu": "1.0.0-rc.17",
"@rolldown/binding-linux-arm64-musl": "1.0.0-rc.17",
"@rolldown/binding-linux-ppc64-gnu": "1.0.0-rc.17",
"@rolldown/binding-linux-s390x-gnu": "1.0.0-rc.17",
"@rolldown/binding-linux-x64-gnu": "1.0.0-rc.17",
"@rolldown/binding-linux-x64-musl": "1.0.0-rc.17",
"@rolldown/binding-openharmony-arm64": "1.0.0-rc.17",
"@rolldown/binding-wasm32-wasi": "1.0.0-rc.17",
"@rolldown/binding-win32-arm64-msvc": "1.0.0-rc.17",
"@rolldown/binding-win32-x64-msvc": "1.0.0-rc.17"
}
},
"node_modules/siginfo": {
@@ -1060,9 +1060,9 @@
"license": "MIT"
},
"node_modules/tinyexec": {
"version": "1.1.1",
"resolved": "https://registry.npmjs.org/tinyexec/-/tinyexec-1.1.1.tgz",
"integrity": "sha512-VKS/ZaQhhkKFMANmAOhhXVoIfBXblQxGX1myCQ2faQrfmobMftXeJPcZGp0gS07ocvGJWDLZGyOZDadDBqYIJg==",
"version": "1.1.2",
"resolved": "https://registry.npmjs.org/tinyexec/-/tinyexec-1.1.2.tgz",
"integrity": "sha512-dAqSqE/RabpBKI8+h26GfLq6Vb3JVXs30XYQjdMjaj/c2tS8IYYMbIzP599KtRj7c57/wYApb3QjgRgXmrCukA==",
"dev": true,
"license": "MIT",
"engines": {
@@ -1119,17 +1119,17 @@
}
},
"node_modules/vite": {
"version": "8.0.8",
"resolved": "https://registry.npmjs.org/vite/-/vite-8.0.8.tgz",
"integrity": "sha512-dbU7/iLVa8KZALJyLOBOQ88nOXtNG8vxKuOT4I2mD+Ya70KPceF4IAmDsmU0h1Qsn5bPrvsY9HJstCRh3hG6Uw==",
"version": "8.0.10",
"resolved": "https://registry.npmjs.org/vite/-/vite-8.0.10.tgz",
"integrity": "sha512-rZuUu9j6J5uotLDs+cAA4O5H4K1SfPliUlQwqa6YEwSrWDZzP4rhm00oJR5snMewjxF5V/K3D4kctsUTsIU9Mw==",
"dev": true,
"license": "MIT",
"dependencies": {
"lightningcss": "^1.32.0",
"picomatch": "^4.0.4",
"postcss": "^8.5.8",
"rolldown": "1.0.0-rc.15",
"tinyglobby": "^0.2.15"
"postcss": "^8.5.10",
"rolldown": "1.0.0-rc.17",
"tinyglobby": "^0.2.16"
},
"bin": {
"vite": "bin/vite.js"
@@ -1197,19 +1197,19 @@
}
},
"node_modules/vitest": {
"version": "4.1.4",
"resolved": "https://registry.npmjs.org/vitest/-/vitest-4.1.4.tgz",
"integrity": "sha512-tFuJqTxKb8AvfyqMfnavXdzfy3h3sWZRWwfluGbkeR7n0HUev+FmNgZ8SDrRBTVrVCjgH5cA21qGbCffMNtWvg==",
"version": "4.1.5",
"resolved": "https://registry.npmjs.org/vitest/-/vitest-4.1.5.tgz",
"integrity": "sha512-9Xx1v3/ih3m9hN+SbfkUyy0JAs72ap3r7joc87XL6jwF0jGg6mFBvQ1SrwaX+h8BlkX6Hz9shdd1uo6AF+ZGpg==",
"dev": true,
"license": "MIT",
"dependencies": {
"@vitest/expect": "4.1.4",
"@vitest/mocker": "4.1.4",
"@vitest/pretty-format": "4.1.4",
"@vitest/runner": "4.1.4",
"@vitest/snapshot": "4.1.4",
"@vitest/spy": "4.1.4",
"@vitest/utils": "4.1.4",
"@vitest/expect": "4.1.5",
"@vitest/mocker": "4.1.5",
"@vitest/pretty-format": "4.1.5",
"@vitest/runner": "4.1.5",
"@vitest/snapshot": "4.1.5",
"@vitest/spy": "4.1.5",
"@vitest/utils": "4.1.5",
"es-module-lexer": "^2.0.0",
"expect-type": "^1.3.0",
"magic-string": "^0.30.21",
@@ -1237,12 +1237,12 @@
"@edge-runtime/vm": "*",
"@opentelemetry/api": "^1.9.0",
"@types/node": "^20.0.0 || ^22.0.0 || >=24.0.0",
"@vitest/browser-playwright": "4.1.4",
"@vitest/browser-preview": "4.1.4",
"@vitest/browser-webdriverio": "4.1.4",
"@vitest/coverage-istanbul": "4.1.4",
"@vitest/coverage-v8": "4.1.4",
"@vitest/ui": "4.1.4",
"@vitest/browser-playwright": "4.1.5",
"@vitest/browser-preview": "4.1.5",
"@vitest/browser-webdriverio": "4.1.5",
"@vitest/coverage-istanbul": "4.1.5",
"@vitest/coverage-v8": "4.1.5",
"@vitest/ui": "4.1.5",
"happy-dom": "*",
"jsdom": "*",
"vite": "^6.0.0 || ^7.0.0 || ^8.0.0"
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@turnstone/sdk",
"version": "0.3.0",
"version": "0.4.0",
"description": "TypeScript client SDK for the turnstone AI orchestration platform",
"type": "module",
"main": "./dist/index.js",
+29
View File
@@ -13,6 +13,20 @@ export interface ConnectedEvent {
export interface HistoryEvent {
type: "history";
/**
* Per-message dicts the frontend consumes directly. Common optional keys:
* - `role`: "user" | "assistant" | "tool"
* - `content`: string or list (image/document parts)
* - `tool_calls`: assistant turns list of `{id, name, arguments, verdict?, output_assessment?}`
* - `tool_call_id`: tool turns id of the originating call
* - `reminders`: metacognitive nudge bubbles (user/tool channels)
* - `advisories`: extracted `UserInterjection` payloads on tool turns
* - `reasoning`: concatenated reasoning text for assistant turns whose
* `provider_data` carried reasoning-bearing blocks (Anthropic
* `thinking`, OpenAI Responses `reasoning`, or synthetic
* `reasoning_text` from path-3 servers). Present only when the
* active model's `surface_persisted_reasoning` flag is true.
*/
messages: Array<Record<string, unknown>>;
}
@@ -38,6 +52,14 @@ export interface StreamEndEvent {
type: "stream_end";
}
/** One-shot replay of the in-progress turn's content + reasoning emitted
* by the events SSE handler when a fresh subscriber connects mid-stream. */
export interface InProgressSnapshotEvent {
type: "in_progress_snapshot";
content: string;
reasoning: string;
}
export interface StateChangeEvent {
type: "state_change";
state: "idle" | "thinking" | "running" | "attention" | "error";
@@ -162,6 +184,7 @@ export type ServerEvent =
| ContentEvent
| ReasoningEvent
| StreamEndEvent
| InProgressSnapshotEvent
| StateChangeEvent
| ToolInfoEvent
| ApproveRequestEvent
@@ -261,6 +284,12 @@ export function isStreamEndEvent(e: ServerEvent): e is StreamEndEvent {
return e.type === "stream_end";
}
export function isInProgressSnapshotEvent(
e: ServerEvent,
): e is InProgressSnapshotEvent {
return e.type === "in_progress_snapshot";
}
export function isStateChangeEvent(e: ServerEvent): e is StateChangeEvent {
return e.type === "state_change";
}
+38 -18
View File
@@ -93,10 +93,17 @@ export class TurnstoneServer extends BaseClient {
});
}
async closeWorkstream(wsId: string): Promise<StatusResponse> {
return this.request("POST", "/v1/api/workstreams/close", {
json: { ws_id: wsId },
});
async closeWorkstream(
wsId: string,
opts?: { reason?: string },
): Promise<StatusResponse> {
const body: Record<string, unknown> = {};
if (opts?.reason !== undefined) body.reason = opts.reason;
return this.request(
"POST",
`/v1/api/workstreams/${encodeURIComponent(wsId)}/close`,
{ json: body },
);
}
// -- Chat interaction -----------------------------------------------------
@@ -106,11 +113,15 @@ export class TurnstoneServer extends BaseClient {
wsId: string,
opts?: { attachmentIds?: string[] },
): Promise<SendResponse> {
const body: Record<string, unknown> = { message, ws_id: wsId };
const body: Record<string, unknown> = { message };
if (opts?.attachmentIds !== undefined) {
body.attachment_ids = opts.attachmentIds;
}
return this.request("POST", "/v1/api/send", { json: body });
return this.request(
"POST",
`/v1/api/workstreams/${encodeURIComponent(wsId)}/send`,
{ json: body },
);
}
// -- Attachments ----------------------------------------------------------
@@ -156,14 +167,17 @@ export class TurnstoneServer extends BaseClient {
feedback?: string | null;
always?: boolean;
}): Promise<StatusResponse> {
return this.request("POST", "/v1/api/approve", {
json: {
ws_id: opts.wsId,
approved: opts.approved ?? true,
feedback: opts.feedback,
always: opts.always,
return this.request(
"POST",
`/v1/api/workstreams/${encodeURIComponent(opts.wsId)}/approve`,
{
json: {
approved: opts.approved ?? true,
feedback: opts.feedback,
always: opts.always,
},
},
});
);
}
async planFeedback(opts: {
@@ -188,15 +202,21 @@ export class TurnstoneServer extends BaseClient {
wsId: string,
opts?: { force?: boolean },
): Promise<StatusResponse> {
const body: Record<string, unknown> = { ws_id: wsId };
const body: Record<string, unknown> = {};
if (opts?.force) body.force = true;
return this.request("POST", "/v1/api/cancel", { json: body });
return this.request(
"POST",
`/v1/api/workstreams/${encodeURIComponent(wsId)}/cancel`,
{ json: body },
);
}
// -- Streaming ------------------------------------------------------------
async *streamEvents(wsId: string): AsyncIterableIterator<ServerEvent> {
yield* this.streamSSE<ServerEvent>("/v1/api/events", { ws_id: wsId });
yield* this.streamSSE<ServerEvent>(
`/v1/api/workstreams/${encodeURIComponent(wsId)}/events`,
);
}
async *streamGlobalEvents(): AsyncIterableIterator<ServerEvent> {
@@ -236,8 +256,8 @@ export class TurnstoneServer extends BaseClient {
try {
// Start consuming the per-workstream SSE stream first
const events = this.streamSSE<ServerEvent>(
"/v1/api/events",
{ ws_id: wsId },
`/v1/api/workstreams/${encodeURIComponent(wsId)}/events`,
undefined,
controller.signal,
);
+33 -3
View File
@@ -162,21 +162,51 @@ export interface CreateWorkstreamResponse {
}
export interface CloseWorkstreamRequest {
ws_id: string;
/**
* Optional close reason persisted to `workstream_config` for
* postmortem. Capped at 512 UTF-8 bytes server-side; credential
* redaction is applied via the output guard.
*/
reason?: string;
}
export interface WorkstreamInfo {
id: string;
// Renamed `id` → `ws_id` and added kind/parent_ws_id/user_id in
// the Stage 2 list-verb lift. Pre-1.5 readers branching on
// `row.id` should swap to `row.ws_id`.
ws_id: string;
name: string;
state: string;
kind: string;
parent_ws_id: string | null;
user_id: string;
}
export interface ListWorkstreamsResponse {
workstreams: WorkstreamInfo[];
}
export interface WorkstreamDetailResponse {
// Lifted from coord-only into a shared verb in the Stage 2
// history/detail verb lift. Both kinds populate every field; SDK
// consumers don't branch on kind.
ws_id: string;
name: string;
state: string;
user_id: string;
kind: string;
}
export interface WorkstreamHistoryResponse {
ws_id: string;
// Tail of the workstream's reconstructed message history
// (provider-fidelity OpenAI-like shape). Bounded by the ?limit=
// query param (default 100, max 500).
messages: Record<string, unknown>[];
}
export interface DashboardWorkstream {
id: string;
ws_id: string;
name: string;
state: string;
title?: string;
+13 -2
View File
@@ -26,7 +26,16 @@ function mockFetchError(
describe("TurnstoneServer", () => {
it("listWorkstreams returns parsed response", async () => {
const fetchFn = mockFetch({
workstreams: [{ id: "ws1", name: "test", state: "idle" }],
workstreams: [
{
ws_id: "ws1",
name: "test",
state: "idle",
kind: "interactive",
parent_ws_id: null,
user_id: "u1",
},
],
});
const client = new TurnstoneServer({
baseUrl: "http://test",
@@ -34,7 +43,9 @@ describe("TurnstoneServer", () => {
});
const resp = await client.listWorkstreams();
expect(resp.workstreams).toHaveLength(1);
expect(resp.workstreams[0].id).toBe("ws1");
// Row key renamed id → ws_id in the Stage 2 list-verb lift.
expect(resp.workstreams[0].ws_id).toBe("ws1");
expect(resp.workstreams[0].kind).toBe("interactive");
expect(fetchFn).toHaveBeenCalledWith(
"http://test/v1/api/workstreams",
expect.objectContaining({ method: "GET" }),
+64 -10
View File
@@ -12,14 +12,33 @@ list differs per file.
from __future__ import annotations
from typing import Any
from typing import TYPE_CHECKING, Any
from unittest.mock import MagicMock
from starlette.middleware.base import BaseHTTPMiddleware
from turnstone.console.coordinator import CoordinatorManager
from turnstone.console.collector import ClusterCollector
from turnstone.console.coordinator_adapter import CoordinatorAdapter
from turnstone.console.coordinator_ui import ConsoleCoordinatorUI
from turnstone.core.auth import AuthResult
from turnstone.core.session_manager import SessionManager
if TYPE_CHECKING:
from collections.abc import Iterable
def _seed_children(
adapter: CoordinatorAdapter, coord_ws_id: str, child_ws_ids: Iterable[str]
) -> None:
"""Seed the coordinator adapter's children registry directly.
The production path populates the registry via the cluster-event
fan-out thread observing ``ws_created`` events. These tests just
need a known-children set for the endpoint handlers to iterate
inject directly via the registry's bulk-merge surface rather than
spinning up the collector + fan-out plumbing.
"""
adapter._registry.merge_children(coord_ws_id, child_ws_ids)
class _AuthMiddleware(BaseHTTPMiddleware):
@@ -59,17 +78,52 @@ def _fake_registry() -> MagicMock:
return reg
def _build_mgr(storage: Any) -> CoordinatorManager:
"""Build a CoordinatorManager with stub factories (test default)."""
def _build_mgr_with_factory(storage: Any, session_factory: Any) -> SessionManager:
"""Build a SessionManager(CoordinatorAdapter) with a caller-supplied factory.
Used by tests that need to capture or assert factory kwargs (e.g.
per-call ``model`` / ``judge_model`` overrides). Plain :func:`_build_mgr`
is the right entry point when the test doesn't care about the
factory.
"""
adapter = CoordinatorAdapter(
collector=MagicMock(),
ui_factory=lambda ws: ConsoleCoordinatorUI(ws_id=ws.id, user_id=ws.user_id or ""),
session_factory=session_factory,
)
mgr = SessionManager(
adapter,
storage=storage,
max_active=3,
node_id=ClusterCollector.CONSOLE_PSEUDO_NODE_ID,
event_emitter=adapter,
)
adapter.attach(mgr)
return mgr
def _build_mgr(storage: Any) -> SessionManager:
"""Build a SessionManager(CoordinatorAdapter) with stub factories (test default)."""
def _sf(ui, model_alias=None, ws_id=None, **kw): # type: ignore[no-untyped-def]
s = MagicMock()
s.send.return_value = None
return s
return CoordinatorManager(
session_factory=_sf,
ui_factory=lambda w, u: ConsoleCoordinatorUI(ws_id=w, user_id=u),
storage=storage,
max_active=3,
)
return _build_mgr_with_factory(storage, _sf)
class MockStorage:
"""Minimal storage mock that implements ``list_services``.
Used by the collector tests + the console route-walk tests. The
collector calls ``list_services("turnstone-server", ...)`` to
discover nodes; tests that don't care about discovery push an
empty list (the default).
"""
def __init__(self) -> None:
self.services: list[dict[str, str]] = []
def list_services(self, service_type: str, max_age_seconds: int = 120) -> list[dict[str, str]]:
return list(self.services)
+53
View File
@@ -0,0 +1,53 @@
"""Shared test helpers — kept out of conftest.py since these are factories,
not fixtures, and several test files want to import them directly."""
from __future__ import annotations
from typing import Any
from unittest.mock import MagicMock
def make_chat_session(**overrides: Any) -> Any:
"""Build a minimal ``ChatSession`` with sane test defaults.
Caller passes any constructor arg as a kwarg to override the default
e.g. ``make_chat_session(memory_config=MemoryConfig(fetch_limit=5))``.
"""
from turnstone.core.session import ChatSession
defaults: dict[str, Any] = {
"client": MagicMock(),
"model": "test-model",
"ui": MagicMock(),
"instructions": None,
"temperature": 0.5,
"max_tokens": 4096,
"tool_timeout": 30,
}
defaults.update(overrides)
return ChatSession(**defaults)
def patch_session_storage(
monkeypatch: Any,
*,
active: bool = True,
raise_on_is_active: bool = False,
) -> list[str]:
"""Patch ``session.get_storage`` to a stub whose ``is_watch_active``
returns *active* (or raises if *raise_on_is_active*). Returns the
list of ``watch_id``s the predicate was called with.
"""
from turnstone.core import session as session_mod
calls: list[str] = []
class _Stub:
def is_watch_active(self, watch_id: str) -> bool:
calls.append(watch_id)
if raise_on_is_active:
raise RuntimeError("storage down")
return active
monkeypatch.setattr(session_mod, "get_storage", lambda: _Stub())
return calls
+58
View File
@@ -0,0 +1,58 @@
"""Shared mock factory for ``events_replay`` tests.
Both interactive (:func:`turnstone.server._interactive_events_replay`)
and coord (:func:`turnstone.console.server._coord_events_replay`) drive
the same shared preamble at
:func:`turnstone.core.session_replay.session_replay_preamble`. Their
test suites share the underlying mock surface (session.model,
session.model_alias, session._last_usage, ui._pending_*, ui._ws_lock,
counters); this module is the single home for that shape so a future
field add lands once.
"""
from __future__ import annotations
import threading
from typing import Any
from unittest.mock import MagicMock
def make_replay_mocks(
*,
last_usage: dict[str, Any] | None = None,
**ui_overrides: Any,
) -> tuple[Any, Any, Any]:
"""Build ``(ws, ui, request)`` MagicMocks for events-replay tests.
Defaults match a fresh workstream that hasn't completed a turn
(no ``last_usage``, no pending prompts).
Args:
last_usage: Sets ``ws.session._last_usage`` directly so tests
don't have to reach into the nested mock; when ``None``
(default), the status replay branch stays inert.
**ui_overrides: Additional attributes set directly on the ``ui``
mock (e.g. ``_pending_approval``, ``_pending_plan_review``,
``_llm_verdicts``, ``_ws_turn_tool_calls``, ``_ws_messages``).
"""
session = MagicMock()
session.model = "gpt-5"
session.model_alias = "default"
session._last_usage = last_usage
session.context_window = 100000
session.reasoning_effort = "medium"
session.messages = []
ui = MagicMock()
ui.auto_approve = False
ui._pending_approval = None
ui._pending_plan_review = None
ui._llm_verdicts = {}
ui._ws_lock = threading.Lock()
ui._ws_turn_tool_calls = 0
ui._ws_messages = 0
for key, value in ui_overrides.items():
setattr(ui, key, value)
ws = MagicMock()
ws.session = session
request = MagicMock()
return ws, ui, request
+45
View File
@@ -0,0 +1,45 @@
"""Shared session-test helpers.
Two reasoning-test modules (``test_session_replay_reasoning.py`` and
``test_session_synth_reasoning_block.py``) need the same minimal
``ChatSession`` factory + a ``SessionUIBase`` no-op subclass. Hoisting
keeps a future third caller from drifting on the defaults the third
existing ``_make_session`` (``test_model_registry.py``) deliberately
takes a different signature (registry / model_alias / reasoning_effort
+ ``_FakeUI``) and is NOT a candidate for sharing this helper.
Module is named with a leading underscore so pytest doesn't try to
collect it as a test file it's an importable utility, not a test.
"""
from __future__ import annotations
from typing import Any
from unittest.mock import MagicMock
from turnstone.core.session import ChatSession
from turnstone.core.session_ui_base import SessionUIBase
class NullUI(SessionUIBase):
"""Bare-bones UI satisfying the SessionUIBase contract for tests
that don't care about UI side effects."""
def __init__(self) -> None:
super().__init__()
def make_session(**kwargs: Any) -> ChatSession:
"""Build a ChatSession with minimal defaults; tests override
individual fields via kwargs."""
defaults: dict[str, Any] = {
"client": MagicMock(),
"model": "test-model",
"ui": NullUI(),
"instructions": None,
"temperature": 0.5,
"max_tokens": 4096,
"tool_timeout": 30,
}
defaults.update(kwargs)
return ChatSession(**defaults)
+87
View File
@@ -1,10 +1,79 @@
from __future__ import annotations
import os
from typing import TYPE_CHECKING, Any
from unittest.mock import MagicMock
import pytest
if TYPE_CHECKING:
from turnstone.core.mcp_client import MCPClientManager, StaticServerState
from turnstone.core.mcp_crypto import MCPTokenCipher
from turnstone.core.oidc import OIDCConfig
def make_mcp_token_cipher() -> MCPTokenCipher:
"""Build a single-key MCP token cipher for tests.
Used by test files that need to exercise ``MCPTokenStore`` round-
trips without the lifespan-side configuration loader; centralised
here so the key/material defaults stay aligned across files.
"""
import base64
from cryptography.fernet import Fernet
from turnstone.core.mcp_crypto import MCPTokenCipher, MCPTokenCipherConfig
raw = base64.urlsafe_b64decode(Fernet.generate_key())
return MCPTokenCipher(MCPTokenCipherConfig(keys=(raw,)))
def _seed_static_state(mgr: MCPClientManager, name: str, **overrides: Any) -> StaticServerState:
"""Get-or-create a ``StaticServerState`` on ``mgr`` and apply ``overrides``.
Shared across MCP test files so the helper stays in one place. Imported
where needed; ``StaticServerState`` is constructed lazily so non-MCP
tests don't pay the import cost.
"""
from turnstone.core.mcp_client import StaticServerState
state = mgr._static_servers.get(name)
if state is None:
state = StaticServerState(name=name)
mgr._static_servers[name] = state
for k, v in overrides.items():
setattr(state, k, v)
return state
def make_oidc_test_config(**overrides: Any) -> OIDCConfig:
"""Build a test ``OIDCConfig`` with sensible defaults.
Shared between ``test_oidc.py`` and ``test_oidc_handlers.py`` so the
defaults (including the now-required ``redirect_base``) stay aligned.
"""
from turnstone.core.oidc import OIDCConfig
defaults: dict[str, Any] = {
"enabled": True,
"issuer": "https://idp.example.com",
"client_id": "my-client",
"client_secret": "my-secret",
"scopes": "openid email profile",
"provider_name": "TestIDP",
"role_claim": "",
"role_map": {},
"password_enabled": True,
"redirect_base": "https://app.example.com",
"authorization_endpoint": "https://idp.example.com/authorize",
"token_endpoint": "https://idp.example.com/token",
"userinfo_endpoint": "https://idp.example.com/userinfo",
"jwks_uri": "https://idp.example.com/.well-known/jwks.json",
}
defaults.update(overrides)
return OIDCConfig(**defaults)
def pytest_addoption(parser: pytest.Parser) -> None:
parser.addoption(
@@ -95,3 +164,21 @@ def mock_openai_client():
client = MagicMock()
client.models.list.return_value.data = [MagicMock(id="test-model")]
return client
@pytest.fixture(autouse=True)
def _clear_policy_cache():
"""Drop the in-process tool-policy cache between tests.
The cache is keyed by org_id (default ``""``), so without this
autouse hook a policy created in test A would leak into test B's
``evaluate_tool_policy`` call distinct storage instances, same
cache slot. Production singleton storage doesn't see the leak
because there's only one storage instance for the process lifetime;
the test isolation requirement is what motivates the autouse.
"""
from turnstone.core.policy import invalidate_policy_cache
invalidate_policy_cache()
yield
invalidate_policy_cache()
+391
View File
@@ -0,0 +1,391 @@
"""Spike 1 — validate MCP SDK behavior for the per-(user, server) session pool.
Three scenarios:
1. N=20 concurrent ClientSession instances to the same URL.
Verifies: no FD blow-up, no shared transport state, each session's
tools/list returns independently.
2. Two concurrent tools/call on a shared ClientSession with interleaving
payloads. Verifies: request_id demux works under contention.
3. Per-session Authorization header isolation. Verifies: different Bearer
tokens per ClientSession reach the server with the expected
Authorization header i.e. httpx connection pooling does not cross
headers between sessions.
Run: uv run python tests/spike_sdk_concurrency.py
Outcome gates Phase 5's pool architecture; if any scenario fails, fall
back to per-call header injection (Alternative F in the OAuth-MCP RFC).
"""
from __future__ import annotations
import asyncio
import contextlib
import logging
import os
import socket
import sys
import threading
import time
from collections import defaultdict
from typing import TYPE_CHECKING
import uvicorn
from mcp import ClientSession
from mcp.client.streamable_http import streamablehttp_client
from mcp.server.fastmcp import FastMCP
from starlette.middleware.base import BaseHTTPMiddleware
if TYPE_CHECKING:
from collections.abc import Callable
from starlette.requests import Request
from starlette.responses import Response
# Reduce uvicorn / mcp log noise so spike output is readable.
logging.getLogger("uvicorn.error").setLevel(logging.WARNING)
logging.getLogger("uvicorn.access").setLevel(logging.WARNING)
logging.getLogger("mcp").setLevel(logging.WARNING)
# Records (auth_header, tool_name) per request — populated by the
# AuthHeaderRecorder middleware below. Indexed by call sequence.
SERVER_OBSERVATIONS: list[tuple[str | None, str | None]] = []
# Tool-call payloads observed (for request_id demux verification).
TOOL_CALL_PAYLOADS: list[str] = []
class AuthHeaderRecorder(BaseHTTPMiddleware):
"""Records the Authorization header on every request the server sees."""
async def dispatch(self, request: Request, call_next: Callable) -> Response:
auth = request.headers.get("authorization")
# We only record the auth header here; tool name comes from the
# body payload which we can't read non-destructively. The tool
# handler logs the payload it received.
SERVER_OBSERVATIONS.append((auth, None))
return await call_next(request)
def find_free_port() -> int:
"""Bind to port 0, return the assigned port."""
s = socket.socket()
s.bind(("127.0.0.1", 0))
port = s.getsockname()[1]
s.close()
return port
def build_server(port: int) -> uvicorn.Server:
"""Create a minimal FastMCP server with one echo tool."""
mcp = FastMCP(name="spike-target", streamable_http_path="/mcp")
@mcp.tool()
async def echo(payload: str) -> str:
"""Echo the payload back. Records the payload server-side."""
TOOL_CALL_PAYLOADS.append(payload)
# Add a small await so two concurrent calls can interleave
# on the wire if the SDK pools the requests.
await asyncio.sleep(0.05)
return f"echoed:{payload}"
app = mcp.streamable_http_app()
app.add_middleware(AuthHeaderRecorder)
config = uvicorn.Config(
app,
host="127.0.0.1",
port=port,
log_level="warning",
access_log=False,
)
return uvicorn.Server(config)
def run_server_in_thread(server: uvicorn.Server) -> threading.Thread:
"""Boot the server in a background thread on its own asyncio loop."""
def _run() -> None:
loop = asyncio.new_event_loop()
asyncio.set_event_loop(loop)
loop.run_until_complete(server.serve())
t = threading.Thread(target=_run, daemon=True, name="spike-server")
t.start()
return t
async def wait_for_server_ready(url: str, timeout: float = 5.0) -> None:
"""Poll the server until it accepts connections."""
import urllib.parse
parsed = urllib.parse.urlparse(url)
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
try:
reader, writer = await asyncio.open_connection(parsed.hostname, parsed.port)
writer.close()
await writer.wait_closed()
return
except OSError:
await asyncio.sleep(0.05)
raise TimeoutError(f"server at {url} not ready within {timeout}s")
def fd_count() -> int:
"""Count open file descriptors for the current process."""
try:
return len(os.listdir(f"/proc/{os.getpid()}/fd"))
except OSError:
return -1
# ---------------------------------------------------------------------------
# Scenario 1: N=20 concurrent ClientSession instances
# ---------------------------------------------------------------------------
async def scenario_1_concurrent_sessions(url: str, n: int = 20) -> dict:
"""Open N concurrent ClientSession instances and call tools/list on each."""
print(f"\n=== Scenario 1: {n} concurrent ClientSession instances ===")
fd_before = fd_count()
async def one_session(idx: int) -> dict:
headers = {"Authorization": f"Bearer test-token-{idx}"}
async with (
streamablehttp_client(url=url, headers=headers) as (read, write, _),
ClientSession(read, write) as session,
):
await session.initialize()
tools = await session.list_tools()
return {
"idx": idx,
"tool_count": len(tools.tools),
"tool_names": [t.name for t in tools.tools],
}
start = time.monotonic()
results = await asyncio.gather(*[one_session(i) for i in range(n)], return_exceptions=True)
elapsed = time.monotonic() - start
fd_after = fd_count()
# Allow some settling time for FDs to release.
await asyncio.sleep(0.5)
fd_settled = fd_count()
successes = [r for r in results if isinstance(r, dict)]
failures = [r for r in results if isinstance(r, Exception)]
# Verify every session got the same tool catalog.
catalog_consistent = (
len(successes) == n and len({tuple(r["tool_names"]) for r in successes}) == 1
)
return {
"scenario": "concurrent_sessions",
"n": n,
"successes": len(successes),
"failures": len(failures),
"elapsed_seconds": round(elapsed, 3),
"fd_before": fd_before,
"fd_during_peak": fd_after,
"fd_settled": fd_settled,
"fd_growth_during": fd_after - fd_before,
"fd_growth_settled": fd_settled - fd_before,
"catalog_consistent": catalog_consistent,
"first_failure": str(failures[0]) if failures else None,
}
# ---------------------------------------------------------------------------
# Scenario 2: 2 concurrent tools/call on a shared session
# ---------------------------------------------------------------------------
async def scenario_2_concurrent_calls_shared_session(url: str) -> dict:
"""Two concurrent tools/call on one ClientSession with interleaving payloads.
The echo tool sleeps 50ms, so concurrent calls overlap on the wire.
Each call passes a distinct payload (~10KB) to make request bodies
spannable across multiple stream frames.
"""
print("\n=== Scenario 2: 2 concurrent tools/call on shared session ===")
# Generous-size payloads so both bodies live during the await.
payload_a = "A" * 10000
payload_b = "B" * 10000
headers = {"Authorization": "Bearer shared-session-token"}
async with (
streamablehttp_client(url=url, headers=headers) as (read, write, _),
ClientSession(read, write) as session,
):
await session.initialize()
TOOL_CALL_PAYLOADS.clear()
start = time.monotonic()
results = await asyncio.gather(
session.call_tool("echo", {"payload": payload_a}),
session.call_tool("echo", {"payload": payload_b}),
return_exceptions=True,
)
elapsed = time.monotonic() - start
successes = [r for r in results if not isinstance(r, Exception)]
failures = [r for r in results if isinstance(r, Exception)]
# Each result.content[0].text should be "echoed:{payload}".
response_payloads: list[str] = []
if len(successes) == 2:
for r in successes:
text = r.content[0].text if r.content else ""
response_payloads.append(text)
# Order may not match call order — what matters is both payloads echo.
expected = {f"echoed:{payload_a}", f"echoed:{payload_b}"}
received = set(response_payloads)
demux_ok = received == expected
# Did both calls actually overlap? If sequential, elapsed ~= 0.1+s;
# if concurrent, ~0.05s.
concurrent_observed = elapsed < 0.09
return {
"scenario": "concurrent_calls_shared_session",
"successes": len(successes),
"failures": len(failures),
"elapsed_seconds": round(elapsed, 3),
"demux_ok": demux_ok,
"expected_payloads_received": list(received) if demux_ok else None,
"actual_payloads_received_count": len(received),
"appears_concurrent_on_wire": concurrent_observed,
"first_failure": str(failures[0]) if failures else None,
}
# ---------------------------------------------------------------------------
# Scenario 3: per-session header isolation
# ---------------------------------------------------------------------------
async def scenario_3_header_isolation(url: str, n: int = 5) -> dict:
"""Open N sessions with distinct Authorization headers, call echo on each.
Verifies the server sees each session's own header — i.e. httpx
connection pooling does not cross headers between concurrent
ClientSession instances against the same URL.
"""
print(f"\n=== Scenario 3: {n}-session Authorization-header isolation ===")
SERVER_OBSERVATIONS.clear()
async def one_session(idx: int) -> str | None:
headers = {"Authorization": f"Bearer iso-token-{idx}"}
async with (
streamablehttp_client(url=url, headers=headers) as (read, write, _),
ClientSession(read, write) as session,
):
await session.initialize()
# One call per session.
await session.call_tool("echo", {"payload": f"session-{idx}"})
return f"Bearer iso-token-{idx}"
start = time.monotonic()
expected_tokens = await asyncio.gather(*[one_session(i) for i in range(n)])
elapsed = time.monotonic() - start
# Tally observed Authorization headers, ignoring None entries (initial
# handshake sometimes lacks auth).
observed_auth = [auth for auth, _ in SERVER_OBSERVATIONS if auth]
expected_set = set(expected_tokens)
observed_set = set(observed_auth)
# Every expected token must show up at least once on the server.
all_present = expected_set.issubset(observed_set)
# No spurious tokens.
no_extras = observed_set.issubset(expected_set)
# Frequency: at least one observation per token.
counts = defaultdict(int)
for a in observed_auth:
counts[a] += 1
each_seen = all(counts[t] >= 1 for t in expected_tokens)
return {
"scenario": "header_isolation",
"n": n,
"elapsed_seconds": round(elapsed, 3),
"expected_tokens": sorted(expected_set),
"observed_tokens": sorted(observed_set),
"all_expected_present": all_present,
"no_extra_tokens_observed": no_extras,
"each_token_seen_at_least_once": each_seen,
"header_counts_per_token": dict(counts),
"total_requests_observed": len(observed_auth),
}
# ---------------------------------------------------------------------------
# Driver
# ---------------------------------------------------------------------------
async def main() -> None:
port = find_free_port()
url = f"http://127.0.0.1:{port}/mcp"
server = build_server(port)
server_thread = run_server_in_thread(server)
try:
await wait_for_server_ready(url)
print(f"server up at {url}\n")
result_1 = await scenario_1_concurrent_sessions(url, n=20)
print_scenario_result(result_1)
result_2 = await scenario_2_concurrent_calls_shared_session(url)
print_scenario_result(result_2)
result_3 = await scenario_3_header_isolation(url, n=5)
print_scenario_result(result_3)
# Final verdict
verdict_1 = (
result_1["successes"] == result_1["n"]
and result_1["catalog_consistent"]
and result_1["fd_growth_settled"] < 30 # 20 sessions, generous bound
)
verdict_2 = result_2["demux_ok"] and result_2["successes"] == 2
verdict_3 = (
result_3["all_expected_present"]
and result_3["no_extra_tokens_observed"]
and result_3["each_token_seen_at_least_once"]
)
print("\n=== VERDICT ===")
print(f" Scenario 1 (concurrent sessions): {'PASS' if verdict_1 else 'FAIL'}")
print(f" Scenario 2 (concurrent calls shared): {'PASS' if verdict_2 else 'FAIL'}")
print(f" Scenario 3 (header isolation): {'PASS' if verdict_3 else 'FAIL'}")
all_pass = verdict_1 and verdict_2 and verdict_3
print(
f"\n Phase 5 per-(user, server) pool architecture: "
f"{'VIABLE' if all_pass else 'NEEDS REWORK (Alternative F fallback)'}"
)
sys.exit(0 if all_pass else 1)
finally:
server.should_exit = True
server_thread.join(timeout=5)
def print_scenario_result(result: dict) -> None:
print(f"\nresult[{result['scenario']}]:")
for k, v in result.items():
if k == "scenario":
continue
print(f" {k}: {v}")
if __name__ == "__main__":
with contextlib.suppress(KeyboardInterrupt):
asyncio.run(main())
File diff suppressed because it is too large Load Diff
+2 -2
View File
@@ -49,7 +49,7 @@ class TestServerVersioning:
mock_mgr = MagicMock()
mock_mgr.list_all.return_value = []
mock_mgr.max_workstreams = 10
mock_mgr.max_active = 10
app = create_app(
workstreams=mock_mgr,
global_queue=queue.Queue(),
@@ -76,7 +76,7 @@ class TestServerVersioning:
assert resp.status_code == 200
spec = resp.json()
assert spec["openapi"] == "3.1.0"
assert "/v1/api/send" in spec["paths"]
assert "/v1/api/workstreams/{ws_id}/send" in spec["paths"]
def test_docs_page(self, client):
resp = client.get("/docs")
+457
View File
@@ -0,0 +1,457 @@
"""Static smoke guards for ``turnstone/ui/static/app.js``.
The interactive WebUI's app.js has no JS test framework on the
project side. This file holds Python-side string-presence assertions
that catch regressions on critical paths the kind of one-line
deletion or rename that breaks the UI silently and only surfaces in
manual testing.
"""
from __future__ import annotations
import re
from pathlib import Path
_APP_JS = Path(__file__).resolve().parent.parent / "turnstone/ui/static/app.js"
def test_switch_tab_bootstraps_pane_when_none_exists() -> None:
"""``switchTab`` must create a pane when none exists. A fresh-
loaded interactive UI with no workstreams shows the dashboard
and creates no panes (per ``initWorkstreams``); the user's first
``create`` or ``open`` then calls ``switchTab(newWsId)``. Pre-fix,
the early ``if (!pane) return;`` left switchTab with nowhere to
attach the chat UI never connected SSE for the freshly-created
workstream, and only a page refresh fixed it. This test guards
against accidentally re-introducing the early-return."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("function switchTab(wsId) {")
# Bound the search to the function body — switchTab is short.
fn = body[start : start + 2000]
assert "if (!pane) return;" not in fn, (
"switchTab must not early-return when no pane exists — that's "
"the no-chat-after-first-create bug. Bootstrap a pane instead."
)
# Affirmatively check the bootstrap path exists.
assert "createPane(wsId)" in fn, (
"switchTab must call createPane(wsId) to bootstrap the first "
"pane when getFocusedPane returns null"
)
def test_tool_error_does_not_overwrite_approval_badge() -> None:
"""When an approved tool subsequently errors, the existing
`` approved`` (or `` auto-approved``) pill must remain visible
the error indicator is appended as a sibling pill, not by mutating
the approval pill in place. Pre-fix, both ``appendToolOutput``
(live) and ``replayHistory`` (history reconstruction) located the
existing approval badge via ``querySelector(".ts-approval-badge")``
and overwrote its className + textContent with the ``--error``
state, so the user lost the record that they had approved the
call. This test pins the new append-sibling behaviour."""
body = _APP_JS.read_text(encoding="utf-8")
# Affirmatively check that an idempotency guard exists somewhere:
# a ``querySelector(".ts-approval-badge--error")`` lookup is the
# structural marker of the fix. Pre-fix the modifier never appeared
# in app.js at all. Loose on quote style and surrounding form (the
# guard might be a negated ``if (!q) {build...}`` block at a call
# site, or a positive ``if (q) return;`` early-exit inside an
# extracted helper) so a later refactor doesn't trip CI on
# cosmetics.
error_guard_re = re.compile(
r"""querySelector\(\s*['"]\.ts-approval-badge--error['"]\s*\)""",
)
assert error_guard_re.search(body), (
"The error-badge code path must guard creation with a "
"querySelector for .ts-approval-badge--error so duplicate fires "
"(live + history re-render) do not stack badges."
)
# Forbid the mutate-existing-badge sequence: a generic
# ``.ts-approval-badge`` lookup followed within a handful of lines
# by mutating that same handle into the ``--error`` state. Two
# unrelated call sites (history rendering + live tool-output
# insertion) legitimately query ``.ts-approval-badge`` to position
# output above it, so the bare query alone is not the anti-pattern;
# the close pairing with an ``--error`` class mutation is. Accept
# either quote style and catch both ``className = "..."`` and
# ``classList.add("ts-approval-badge--error")`` forms.
overwrite_re = re.compile(
r"""(\w+)\s*=\s*\w+\.querySelector\(\s*(["'])\.ts-approval-badge\2\s*\)\s*;"""
r""".{0,200}?"""
r"""(?:"""
r"""\1\.className\s*=\s*(["'])[^"']*\bts-approval-badge--error\b[^"']*\3"""
r"""|"""
r"""\1\.classList\.add\([^)]*(["'])ts-approval-badge--error\4[^)]*\)"""
r""")""",
re.DOTALL,
)
assert not overwrite_re.search(body), (
"Found the badge-overwrite anti-pattern: a queried "
".ts-approval-badge handle is mutated into the --error variant "
"(via className overwrite or classList.add). Append a sibling "
"badge instead so the approval verdict stays visible alongside "
"the error."
)
def test_replay_history_renders_content_before_tool_block() -> None:
"""In ``replayHistory``'s ``role === "assistant"`` branch, the
``msg.content`` render must precede the ``msg.tool_calls`` render.
Two reasons, both load-bearing:
1. **Structural** the next loop iteration's ``role === "tool"``
message anchors to ``lastToolBlock``. The tool-block branch sets
that anchor; the content branch clears it. If content runs after
the tool block, the clear silently drops the upcoming tool
result. Pre-fix, every interactive tool result was missing from
saved-workstream replays whenever the assistant turn carried
both narration and tool calls (very common output shape).
2. **Visual** the live SSE path renders content first
(``stream_text`` streams before ``tool_info`` /
``approve_request``), so replay should match.
The test pins the order via the offsets of the ``msg.content`` and
``msg.tool_calls`` branch headers inside the function body."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.replayHistory = function")
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
fn = body[start:end]
# Locate the assistant branch and bound the search to its body —
# the function also handles user / tool roles which would otherwise
# confuse the offset comparison.
asst_start = fn.index('msg.role === "assistant"')
asst_end = fn.index('msg.role === "tool"', asst_start)
asst = fn[asst_start:asst_end]
# ``if (msg.content && msg.content.trim())`` guards against a
# whitespace-only content row (Qwen-style "\n\n" left over after a
# reasoning-parser model strips ``<think>…</think>`` and emits
# nothing else before the tool call). Pre-trim guard, those rows
# rendered as a visible-but-empty ``.msg.assistant`` card on
# replay. Match the substring up to ``msg.content`` so the test
# tolerates either guard shape without locking the trim() in.
content_idx = asst.index("if (msg.content")
tool_calls_idx = asst.index("if (msg.tool_calls && msg.tool_calls.length)")
assert content_idx < tool_calls_idx, (
"replayHistory must render msg.content BEFORE msg.tool_calls "
"inside the assistant branch — otherwise the lastToolBlock "
"anchor is clobbered before the next iteration's tool result "
"can attach to it (and the visual order also drifts from the "
"live SSE flow)."
)
def test_replay_history_renders_persisted_verdict_badge() -> None:
"""Saved-workstream replays must paint the persisted intent verdict
next to each tool div, using the same ``renderVerdictBadge`` helper
the live ``showInlineToolBlock`` path uses. Pre-fix the audit trail
was complete in storage (``intent_verdicts`` table) but never
surfaced on replay operators reviewing a saved workstream
couldn't see what the heuristic / LLM judge thought of any tool
call. This test pins the call site so a refactor that drops the
decoration regresses the audit surface."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.replayHistory = function")
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
fn = body[start:end]
# Match a `renderVerdictBadge(<something>.verdict, ...)` call inside
# the replay loop. Loose on whitespace + identifier so a future
# rename of the iteration variable doesn't trip CI.
badge_call_re = re.compile(
r"renderVerdictBadge\(\s*\w+\.verdict\b",
)
assert badge_call_re.search(fn), (
"replayHistory must call renderVerdictBadge(tc.verdict, ...) "
"when a persisted verdict is attached to a tool_call entry — "
"otherwise the audit-trail data persisted to intent_verdicts "
"doesn't surface on saved-workstream replays."
)
def test_shared_utils_defines_replay_advisories_after_tool() -> None:
"""The shared ``replayAdvisoriesAfterTool`` helper in
``shared_static/utils.js`` is the single source of advisory-walk +
type-filter logic for both ``app.js`` (interactive) and
``coordinator.js`` (coord). A refactor that drops the helper
breaks both surfaces, so guard its definition + filter shape here.
"""
utils_js = Path(__file__).resolve().parent.parent / "turnstone/shared_static/utils.js"
body = utils_js.read_text(encoding="utf-8")
assert "function replayAdvisoriesAfterTool" in body, (
"shared/utils.js must define replayAdvisoriesAfterTool — "
"interactive and coord both invoke it."
)
# The type filter — ``adv.type !== 'user_interjection'`` — must
# remain in the helper so a future advisory shape (output_guard,
# metacognitive nudge, etc.) doesn't silently render as a user
# bubble.
assert 'adv.type !== "user_interjection"' in body, (
"replayAdvisoriesAfterTool must filter by advisory type so a "
"future non-user_interjection advisory shape doesn't silently "
"render as a user bubble."
)
def test_replay_renders_user_interjection_advisory_after_tool_block() -> None:
"""Queued user messages spliced into the last tool-result envelope
of a batch (Seam 1) persist on the tool DB row as a wrapped
``<tool_output>`` envelope. ``decorate_history_messages`` extracts
the advisory back out and the wire layer projects it onto
``msg.advisories``; ``replayHistory`` must invoke the shared
``replayAdvisoriesAfterTool`` helper (defined in
``shared/utils.js``) so each ``user_interjection`` renders through
``addUserMessage`` and the bubble looks identical to a Seam 2/3
user row.
This test pins the call site so a refactor that drops the helper
invocation regresses the queued-during-batch replay shape
silently."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.replayHistory = function")
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
fn = body[start:end]
# The replay loop must invoke the shared helper, passing
# ``msg.advisories`` and a renderer that routes through
# ``addUserMessage``. The helper itself filters on
# ``adv.type !== "user_interjection"``; that branch lives in
# ``shared/utils.js`` (test_shared_utils_js or runtime smoke covers
# the helper's body).
assert "replayAdvisoriesAfterTool(msg.advisories" in fn, (
"replayHistory must invoke replayAdvisoriesAfterTool with "
"msg.advisories so queued messages spliced into the tool "
"envelope render as user bubbles after the tool block."
)
assert "addUserMessage(text" in fn, (
"replayHistory's renderer callback must route the extracted "
"advisory text through addUserMessage so the rendered bubble "
"matches a normal user-row replay."
)
# ---------------------------------------------------------------------------
# Phase 8 — Chunk D: MCP error embed + settings panel UX
# ---------------------------------------------------------------------------
_INDEX_HTML = Path(__file__).resolve().parent.parent / "turnstone/ui/static/index.html"
_STYLE_CSS = Path(__file__).resolve().parent.parent / "turnstone/ui/static/style.css"
# The Phase-8 D-chunk pins the absence of an unsafe DOM-write API
# in two regions of app.js. Spell the property name out of literal
# concatenation so the tooling that flags occurrences in code
# strings doesn't false-positive on the test source.
_UNSAFE_DOM_WRITE_RE = re.compile(r"\.inner" + r"HTML\s*=")
def test_phase8_mcp_error_helpers_defined_in_app_js() -> None:
"""The Phase 8 dashboard renderer adds three load-bearing helpers
next to the existing media-embed pattern: ``tryParseMcpError``
(envelope detector), ``buildMcpErrorEmbed`` (interactive card),
and the ``_pendingConsentServers`` set that drives the gear-icon
badge. A regression that drops any of them silently degrades the
OAuth consent UX to a plain JSON dump, so guard their existence
here."""
body = _APP_JS.read_text(encoding="utf-8")
assert "function tryParseMcpError" in body, (
"tryParseMcpError must remain defined — appendToolOutput's "
"error branch depends on it to detect the MCP error envelope."
)
assert "function buildMcpErrorEmbed" in body, (
"buildMcpErrorEmbed must remain defined — it renders the "
"interactive consent / forbidden / operator card."
)
assert "_pendingConsentServers" in body, (
"_pendingConsentServers state must remain — it backs the "
"gear-icon badge so a user who scrolls past a consent prompt "
"still has a stable signal that consent is pending."
)
# The buildMcpErrorEmbed pattern must also wire the "actionable"
# branch (consent_required / insufficient_scope) into the badge
# via _onConsentDetected; pin the helper name.
assert "_onConsentDetected" in body, (
"_onConsentDetected must remain — buildMcpErrorEmbed calls it "
"for the actionable category to surface the gear-icon badge."
)
def test_phase8_settings_panel_handlers_defined() -> None:
"""The settings modal exposes four entry points that the inline
``onclick`` attributes in index.html depend on. Renaming or
deleting any of them breaks the modal silently (the buttons are
still rendered but click-to-action is dead). Catch that here."""
body = _APP_JS.read_text(encoding="utf-8")
for name in [
"function openSettingsPanel",
"function closeSettingsPanel",
"function confirmRevokeMcp",
"function cancelRevokeMcp",
]:
assert name in body, f"Missing required handler: {name}"
# The connections list is fetched against the Phase-7 endpoint —
# pin the URL so a server-side rename forces an explicit UI bump.
assert "/v1/api/mcp/oauth/connections" in body, (
"Settings panel must fetch /v1/api/mcp/oauth/connections — "
"a server-side rename needs an explicit UI update."
)
def test_phase8_appendtooloutput_dispatches_mcp_error_before_renderer() -> None:
"""``appendToolOutput`` must call ``tryParseMcpError`` inside its
``isError`` branch BEFORE falling through to the plain
``renderToolOutput`` path. The ordering is what makes the
interactive consent card replace the JSON dump; reverse the calls
and the user sees the raw error envelope as text again."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.appendToolOutput = function")
end = body.index("Pane.prototype.", start + 10)
fn = body[start:end]
parse_idx = fn.find("tryParseMcpError(")
render_idx = fn.find("renderToolOutput(")
assert parse_idx >= 0, (
"appendToolOutput must call tryParseMcpError on the error path "
"before renderToolOutput, otherwise the consent card never "
"replaces the plain JSON output."
)
assert render_idx >= 0, "renderToolOutput call must remain present"
assert parse_idx < render_idx, (
"tryParseMcpError must run BEFORE renderToolOutput so the "
"interactive card path takes precedence over plain rendering."
)
def test_phase8_no_unsafe_dom_write_in_settings_panel() -> None:
"""Defensive XSS guard: the settings panel renders user-controlled
server names, scope strings, and timestamp values into the DOM.
The whole section MUST go through ``textContent``-style APIs; an
unsafe-DOM-write assignment would be a regression vector. Bound
the check to the section 15 body to avoid false positives
elsewhere."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("// 15. MCP server connections settings panel")
# Bound to the full settings section (terminates at the next
# top-level keydown handler block).
end = body.index('document.addEventListener("keydown"', start)
section = body[start:end]
assert not _UNSAFE_DOM_WRITE_RE.search(section), (
"Section 15 must not assign to the unsafe DOM-write property — "
"server names and scope values flow through here and would be "
"XSS-injectable. Use textContent / DOM APIs instead."
)
def test_phase8_settings_button_in_index_html() -> None:
"""The gear-icon entry-point for the settings panel must remain
in the appbar's actions span. The console proxy IIFE prepends a
node pill to ``header.firstChild`` (turnstone/console/server.py:
202); our button is appended inside ``<span class='appbar-actions'>``
on the right, so they don't collide. Pin both shape constraints
here so a future appbar refactor keeps them disjoint."""
body = _INDEX_HTML.read_text(encoding="utf-8")
assert 'id="settings-btn"' in body, (
"index.html must keep the #settings-btn — onclick handlers "
"and the consent badge target it by id."
)
assert 'onclick="openSettingsPanel()"' in body, (
"settings-btn must wire onclick=openSettingsPanel() — losing "
"the binding leaves the panel unreachable."
)
# The button must live inside <span class="appbar-actions"> so the
# console proxy's header.insertBefore(pill, header.firstChild)
# leaves it untouched.
actions_open = body.index('class="appbar-actions"')
actions_close = body.index("</span>", actions_open)
assert 'id="settings-btn"' in body[actions_open:actions_close], (
"settings-btn must be inside <span class='appbar-actions'> "
"so the console proxy's firstChild prepend doesn't shift it."
)
def test_phase8_settings_modal_in_index_html() -> None:
"""Both the settings overlay and the revoke-confirmation overlay
must remain in the modal area. The Escape-key deferral list in
app.js targets these ids, so removing them silently breaks the
handler chain."""
body = _INDEX_HTML.read_text(encoding="utf-8")
assert 'id="settings-overlay"' in body
assert 'id="revoke-mcp-overlay"' in body
# Each overlay must have role="dialog" + aria-modal="true" so
# screen readers and the existing modal-deferral handlers can
# treat them like the rest of the modal stack.
for overlay_id in ("settings-overlay", "revoke-mcp-overlay"):
idx = body.index(f'id="{overlay_id}"')
# Bound to ~600 chars after the open tag so we only check this
# overlay's attributes.
chunk = body[idx : idx + 600]
assert 'role="dialog"' in chunk, f"{overlay_id} missing role=dialog"
assert 'aria-modal="true"' in chunk, f"{overlay_id} missing aria-modal=true"
def test_phase8_xss_safe_render_in_build_mcp_error_embed() -> None:
"""Adversarial input — the renderer for an MCP error envelope
must use ``textContent`` (not the unsafe DOM-write API) for every
field that flows from the server: ``err.detail``, ``err.server``,
scopes list. The card builder uses createElement + textContent
throughout so a script-tag server name renders harmlessly. Pin
the absence of the unsafe-write inside ``buildMcpErrorEmbed``."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("function buildMcpErrorEmbed(")
# Bound to the function body — find its closing brace at column 0.
rest = body[start:]
# Closing function brace at line start (matches existing functions)
end_match = re.search(r"\n}\n", rest)
assert end_match is not None
fn = rest[: end_match.end()]
assert not _UNSAFE_DOM_WRITE_RE.search(fn), (
"buildMcpErrorEmbed must not use the unsafe-DOM-write API — "
"server names and detail strings flow through here. An "
"adversarial server name must render harmlessly via "
"textContent."
)
def test_phase8_css_classes_present_in_stylesheet() -> None:
"""The card / badge / modal classes referenced from app.js must
have CSS rules. Without them the DOM still works but the visual
treatment is gone, which would silently degrade the consent UX."""
css = _STYLE_CSS.read_text(encoding="utf-8")
for selector in [
".mcp-error-card",
".mcp-error-icon",
".mcp-error-action-btn",
".mcp-scope-pill",
"#settings-overlay",
"#settings-box",
".settings-revoke-btn",
".settings-consent-badge",
"#revoke-mcp-overlay",
]:
assert selector in css, f"Missing CSS rule for {selector}"
def test_phase8_consent_url_prefix_check_in_click_handler() -> None:
"""Defence-in-depth: the consent button's click handler must reject
any ``consent_url`` that doesn't start with the dispatcher's known
prefix (``/v1/api/mcp/oauth/start``). ``_build_consent_url`` always
emits a path-relative URL with that exact prefix; a non-prefix
value implies the producer drifted (or was compromised) and a
``window.open("javascript:...")`` would be catastrophic.
The renderer is the last line of defence before ``window.open`` and
must not rely on the producer-side guarantee alone. Pin the prefix
string and the ``startsWith`` form so a future refactor can't
silently weaken the guard.
"""
body = _APP_JS.read_text(encoding="utf-8")
# Bound the search to the click handler region (between the
# ``buildMcpErrorEmbed`` function and the next top-level helper) to
# avoid false positives from unrelated string occurrences.
start = body.index("function buildMcpErrorEmbed(")
end = body.index("\n}\n", start) + 1
fn = body[start:end]
assert 'consentUrl.startsWith("/v1/api/mcp/oauth/start")' in fn, (
"Click handler must guard window.open with "
'consentUrl.startsWith("/v1/api/mcp/oauth/start"). Without it '
"a future producer drift to a non-path-relative URL (or a "
'"javascript:" injection) would be passed straight to '
"window.open."
)
+1 -1
View File
@@ -87,7 +87,7 @@ def test_record_audit_redacts_nested_strings(storage):
record_audit(
storage,
"u1",
"task_list.update",
"tasks.update",
detail={
"tasks": [
{"title": "normal task"},
+85 -37
View File
@@ -53,8 +53,8 @@ class TestIsPublicPath:
def test_api_workstreams_not_public(self):
assert is_public_path("/api/workstreams") is False
def test_api_send_not_public(self):
assert is_public_path("/api/send") is False
def test_api_workstreams_send_not_public(self):
assert is_public_path("/api/workstreams/abc/send") is False
def test_api_cluster_overview_not_public(self):
assert is_public_path("/api/cluster/overview") is False
@@ -71,8 +71,8 @@ class TestIsPublicPath:
def test_v1_api_workstreams_not_public(self):
assert is_public_path("/v1/api/workstreams") is False
def test_v1_api_send_not_public(self):
assert is_public_path("/v1/api/send") is False
def test_v1_api_workstreams_send_not_public(self):
assert is_public_path("/v1/api/workstreams/abc/send") is False
def test_openapi_json_public(self):
assert is_public_path("/openapi.json") is True
@@ -97,10 +97,22 @@ class TestRequiredScope:
assert required_scope("GET", "/api/events") == "read"
def test_post_send_needs_write(self):
assert required_scope("POST", "/api/send") == "write"
assert required_scope("POST", "/api/workstreams/abc/send") == "write"
def test_delete_send_needs_write(self):
assert required_scope("DELETE", "/api/workstreams/abc/send") == "write"
def test_post_approve_needs_approve(self):
assert required_scope("POST", "/api/approve") == "approve"
assert required_scope("POST", "/api/workstreams/abc/approve") == "approve"
def test_post_cancel_needs_write(self):
assert required_scope("POST", "/api/workstreams/abc/cancel") == "write"
def test_post_close_needs_write(self):
assert required_scope("POST", "/api/workstreams/abc/close") == "write"
def test_get_events_per_ws_needs_read(self):
assert required_scope("GET", "/api/workstreams/abc/events") == "read"
def test_post_plan_needs_write(self):
assert required_scope("POST", "/api/plan") == "write"
@@ -111,9 +123,6 @@ class TestRequiredScope:
def test_post_workstreams_new_needs_write(self):
assert required_scope("POST", "/api/workstreams/new") == "write"
def test_post_workstreams_close_needs_write(self):
assert required_scope("POST", "/api/workstreams/close") == "write"
def test_all_write_paths_need_write(self):
for path in WRITE_PATHS:
scope = required_scope("POST", path)
@@ -123,10 +132,10 @@ class TestRequiredScope:
assert required_scope("POST", "/api/unknown") == "read"
def test_v1_post_send_needs_write(self):
assert required_scope("POST", "/v1/api/send") == "write"
assert required_scope("POST", "/v1/api/workstreams/abc/send") == "write"
def test_v1_post_approve_needs_approve(self):
assert required_scope("POST", "/v1/api/approve") == "approve"
assert required_scope("POST", "/v1/api/workstreams/abc/approve") == "approve"
def test_v1_get_workstreams_needs_read(self):
assert required_scope("GET", "/v1/api/workstreams") == "read"
@@ -135,10 +144,10 @@ class TestRequiredScope:
assert required_scope("POST", "/v1/api/cluster/workstreams/new") == "write"
def test_proxy_v1_send_needs_write(self):
assert required_scope("POST", "/node/node-a/v1/api/send") == "write"
assert required_scope("POST", "/node/node-a/v1/api/workstreams/abc/send") == "write"
def test_proxy_v1_approve_needs_approve(self):
assert required_scope("POST", "/node/node-a/v1/api/approve") == "approve"
assert required_scope("POST", "/node/node-a/v1/api/workstreams/abc/approve") == "approve"
def test_proxy_v1_read_endpoint_needs_read(self):
assert required_scope("GET", "/node/node-a/v1/api/workstreams") == "read"
@@ -198,6 +207,30 @@ class TestRequiredScope:
"""Only POST is elevated — GET falls through to read."""
assert required_scope("GET", "/api/_internal/mcp-reload") == "read"
def test_internal_mcp_refresh_one_needs_approve(self):
assert required_scope("POST", "/api/_internal/mcp-refresh/srv") == "approve"
def test_v1_internal_mcp_refresh_one_needs_approve(self):
assert required_scope("POST", "/v1/api/_internal/mcp-refresh/srv") == "approve"
def test_proxy_internal_mcp_refresh_one_needs_approve(self):
assert required_scope("POST", "/node/n1/v1/api/_internal/mcp-refresh/srv") == "approve"
def test_proxy_no_v1_internal_mcp_refresh_one_needs_approve(self):
assert required_scope("POST", "/node/n1/api/_internal/mcp-refresh/srv") == "approve"
def test_internal_mcp_reconnect_one_needs_approve(self):
assert required_scope("POST", "/api/_internal/mcp-reconnect/srv") == "approve"
def test_v1_internal_mcp_reconnect_one_needs_approve(self):
assert required_scope("POST", "/v1/api/_internal/mcp-reconnect/srv") == "approve"
def test_proxy_internal_mcp_reconnect_one_needs_approve(self):
assert required_scope("POST", "/node/n1/v1/api/_internal/mcp-reconnect/srv") == "approve"
def test_proxy_no_v1_internal_mcp_reconnect_one_needs_approve(self):
assert required_scope("POST", "/node/n1/api/_internal/mcp-reconnect/srv") == "approve"
# Workstream sub-resource mutations (parametric paths)
def test_ws_delete_needs_write(self):
assert required_scope("POST", "/api/workstreams/abc123/delete") == "write"
@@ -402,7 +435,7 @@ class TestCheckRequest:
def test_write_read_token_403(self, read_jwt):
allowed, status, msg, _result = check_request(
"POST", "/api/send", read_jwt, jwt_secret=self._SECRET
"POST", "/api/workstreams/abc/send", read_jwt, jwt_secret=self._SECRET
)
assert allowed is False
assert status == 403
@@ -410,14 +443,14 @@ class TestCheckRequest:
def test_write_full_token_ok(self, full_jwt):
allowed, status, msg, _result = check_request(
"POST", "/api/send", full_jwt, jwt_secret=self._SECRET
"POST", "/api/workstreams/abc/send", full_jwt, jwt_secret=self._SECRET
)
assert allowed is True
assert status == 200
def test_approve_read_token_403(self, read_jwt):
allowed, status, msg, _result = check_request(
"POST", "/api/approve", read_jwt, jwt_secret=self._SECRET
"POST", "/api/workstreams/abc/approve", read_jwt, jwt_secret=self._SECRET
)
assert allowed is False
assert status == 403
@@ -425,7 +458,10 @@ class TestCheckRequest:
def test_proxy_write_read_token_403(self, read_jwt):
"""Read tokens cannot escalate to write ops via proxy routes."""
allowed, status, msg, _result = check_request(
"POST", "/node/node-a/api/send", read_jwt, jwt_secret=self._SECRET
"POST",
"/node/node-a/api/workstreams/abc/send",
read_jwt,
jwt_secret=self._SECRET,
)
assert allowed is False
assert status == 403
@@ -433,7 +469,10 @@ class TestCheckRequest:
def test_proxy_write_trailing_slash_read_token_403(self, read_jwt):
"""Trailing slash must not bypass write-role check on proxy routes."""
allowed, status, msg, _result = check_request(
"POST", "/node/node-a/api/send/", read_jwt, jwt_secret=self._SECRET
"POST",
"/node/node-a/api/workstreams/abc/send/",
read_jwt,
jwt_secret=self._SECRET,
)
assert allowed is False
assert status == 403
@@ -441,7 +480,7 @@ class TestCheckRequest:
def test_direct_write_trailing_slash_read_token_403(self, read_jwt):
"""Trailing slash must not bypass write-role check on direct routes."""
allowed, status, msg, _result = check_request(
"POST", "/api/send/", read_jwt, jwt_secret=self._SECRET
"POST", "/api/workstreams/abc/send/", read_jwt, jwt_secret=self._SECRET
)
assert allowed is False
assert status == 403
@@ -449,14 +488,20 @@ class TestCheckRequest:
def test_proxy_write_full_token_ok(self, full_jwt):
"""Full tokens pass through proxy write routes."""
allowed, status, msg, _result = check_request(
"POST", "/node/node-a/api/send", full_jwt, jwt_secret=self._SECRET
"POST",
"/node/node-a/api/workstreams/abc/send",
full_jwt,
jwt_secret=self._SECRET,
)
assert allowed is True
def test_proxy_v1_write_read_token_403(self, read_jwt):
"""Read tokens cannot escalate to write ops via v1 proxy routes."""
allowed, status, msg, _result = check_request(
"POST", "/node/node-a/v1/api/send", read_jwt, jwt_secret=self._SECRET
"POST",
"/node/node-a/v1/api/workstreams/abc/send",
read_jwt,
jwt_secret=self._SECRET,
)
assert allowed is False
assert status == 403
@@ -464,7 +509,10 @@ class TestCheckRequest:
def test_proxy_v1_write_full_token_ok(self, full_jwt):
"""Full tokens pass through v1 proxy write routes."""
allowed, status, msg, _result = check_request(
"POST", "/node/node-a/v1/api/send", full_jwt, jwt_secret=self._SECRET
"POST",
"/node/node-a/v1/api/workstreams/abc/send",
full_jwt,
jwt_secret=self._SECRET,
)
assert allowed is True
@@ -496,7 +544,7 @@ class TestCheckRequest:
def test_approve_full_token_ok(self, full_jwt):
allowed, status, msg, _result = check_request(
"POST", "/api/approve", full_jwt, jwt_secret=self._SECRET
"POST", "/api/workstreams/abc/approve", full_jwt, jwt_secret=self._SECRET
)
assert allowed is True
@@ -538,7 +586,7 @@ class TestCheckRequestWithCookie:
def test_bearer_takes_precedence_over_cookie(self, read_jwt, full_jwt):
allowed, status, _, _r = check_request(
"POST",
"/api/send",
"/api/workstreams/abc/send",
f"Bearer {full_jwt}",
cookie_header=f"turnstone_auth={read_jwt}",
jwt_secret=self._SECRET,
@@ -559,7 +607,7 @@ class TestCheckRequestWithCookie:
def test_cookie_read_on_write_403(self, read_jwt):
allowed, status, _, _r = check_request(
"POST",
"/api/send",
"/api/workstreams/abc/send",
None,
cookie_header=f"turnstone_auth={read_jwt}",
jwt_secret=self._SECRET,
@@ -570,7 +618,7 @@ class TestCheckRequestWithCookie:
def test_cookie_full_on_write_ok(self, full_jwt):
allowed, status, _, _r = check_request(
"POST",
"/api/send",
"/api/workstreams/abc/send",
None,
cookie_header=f"turnstone_auth={full_jwt}",
jwt_secret=self._SECRET,
@@ -643,7 +691,7 @@ class TestServerAuth:
mock_ws.user_id = "u1"
mock_mgr = MagicMock()
mock_mgr.list_all.return_value = [mock_ws]
mock_mgr.max_workstreams = 10
mock_mgr.max_active = 10
from turnstone.core.auth import JWT_AUD_SERVER
@@ -700,25 +748,25 @@ class TestServerAuth:
def test_api_send_read_token_403(self):
resp = self.client.post(
"/v1/api/send",
"/v1/api/workstreams/x/send",
headers=self._read_hdr,
json={"message": "hello", "ws_id": "x"},
json={"message": "hello"},
)
assert resp.status_code == 403
assert "Forbidden" in resp.json().get("error", "")
def test_api_send_full_token_passes_auth(self):
resp = self.client.post(
"/v1/api/send",
"/v1/api/workstreams/nonexistent/send",
headers=self._full_hdr,
json={"message": "hello", "ws_id": "nonexistent"},
json={"message": "hello"},
)
assert resp.status_code not in (401, 403)
def test_api_send_no_token_401(self):
resp = self.client.post(
"/v1/api/send",
json={"message": "hello", "ws_id": "x"},
"/v1/api/workstreams/x/send",
json={"message": "hello"},
)
assert resp.status_code == 401
@@ -731,7 +779,7 @@ class TestServerAuth:
def test_options_no_auth_required(self):
resp = self.client.options(
"/v1/api/send",
"/v1/api/workstreams/x/send",
headers={
"Origin": "http://example.com",
"Access-Control-Request-Method": "POST",
@@ -865,7 +913,7 @@ class TestServerLogin:
mock_ws.user_id = "u1"
mock_mgr = MagicMock()
mock_mgr.list_all.return_value = [mock_ws]
mock_mgr.max_workstreams = 10
mock_mgr.max_active = 10
# Mock storage with a test user for password login
from turnstone.core.auth import hash_password
@@ -1671,7 +1719,7 @@ class TestCorsConfigurable:
mgr = MagicMock()
mgr.list_all.return_value = []
mgr.max_workstreams = 10
mgr.max_active = 10
app = srv_mod.create_app(
workstreams=mgr,
global_queue=queue.Queue(),
@@ -1692,7 +1740,7 @@ class TestCorsConfigurable:
mgr = MagicMock()
mgr.list_all.return_value = []
mgr.max_workstreams = 10
mgr.max_active = 10
app = srv_mod.create_app(
workstreams=mgr,
global_queue=queue.Queue(),
+11 -11
View File
@@ -175,10 +175,10 @@ class TestRequiredScope:
assert required_scope("GET", "/api/workstreams") == "read"
def test_post_write(self):
assert required_scope("POST", "/api/send") == "write"
assert required_scope("POST", "/api/workstreams/abc/send") == "write"
def test_post_approve(self):
assert required_scope("POST", "/api/approve") == "approve"
assert required_scope("POST", "/api/workstreams/abc/approve") == "approve"
def test_admin_prefix(self):
assert required_scope("GET", "/api/admin/users") == "approve"
@@ -186,14 +186,14 @@ class TestRequiredScope:
assert required_scope("DELETE", "/api/admin/users/abc") == "approve"
def test_versioned_path(self):
assert required_scope("POST", "/v1/api/send") == "write"
assert required_scope("POST", "/v1/api/approve") == "approve"
assert required_scope("POST", "/v1/api/workstreams/abc/send") == "write"
assert required_scope("POST", "/v1/api/workstreams/abc/approve") == "approve"
def test_proxy_write(self):
assert required_scope("POST", "/node/n1/api/send") == "write"
assert required_scope("POST", "/node/n1/api/workstreams/abc/send") == "write"
def test_proxy_approve(self):
assert required_scope("POST", "/node/n1/api/approve") == "approve"
assert required_scope("POST", "/node/n1/api/workstreams/abc/approve") == "approve"
# ---------------------------------------------------------------------------
@@ -270,7 +270,7 @@ class TestCheckRequestScopes:
jwt_tok = create_jwt("u1", frozenset({"read"}), "test", self._SECRET)
allowed, status, msg, _ = check_request(
"POST",
"/api/send",
"/api/workstreams/abc/send",
f"Bearer {jwt_tok}",
jwt_secret=self._SECRET,
)
@@ -282,7 +282,7 @@ class TestCheckRequestScopes:
jwt_tok = create_jwt("u1", frozenset({"read"}), "test", self._SECRET)
allowed, status, msg, _ = check_request(
"POST",
"/api/approve",
"/api/workstreams/abc/approve",
f"Bearer {jwt_tok}",
jwt_secret=self._SECRET,
)
@@ -294,7 +294,7 @@ class TestCheckRequestScopes:
jwt_tok = create_jwt("u1", frozenset({"read", "write", "approve"}), "test", self._SECRET)
allowed, status, msg, result = check_request(
"POST",
"/api/approve",
"/api/workstreams/abc/approve",
f"Bearer {jwt_tok}",
jwt_secret=self._SECRET,
)
@@ -306,7 +306,7 @@ class TestCheckRequestScopes:
jwt_tok = create_jwt("u1", frozenset({"read", "write"}), "db", self._SECRET)
allowed, status, msg, result = check_request(
"POST",
"/api/send",
"/api/workstreams/abc/send",
f"Bearer {jwt_tok}",
jwt_secret=self._SECRET,
)
@@ -318,7 +318,7 @@ class TestCheckRequestScopes:
jwt_tok = create_jwt("u1", frozenset({"read"}), "db", self._SECRET)
allowed, status, msg, _ = check_request(
"POST",
"/api/send",
"/api/workstreams/abc/send",
f"Bearer {jwt_tok}",
jwt_secret=self._SECRET,
)
+266
View File
@@ -0,0 +1,266 @@
"""Tests for ``turnstone.server._build_history`` reminder + source surfacing.
The replay path (``_build_history``) projects the ``_source`` and
``_reminders`` side-channels onto the wire entry the frontend
consumes. Persisted via migration 050 (Commit 1) so multi-tab /
multi-device replay sees the same metacognitive bubble shape the
originating tab saw live.
"""
from __future__ import annotations
from types import SimpleNamespace
from typing import Any
from unittest.mock import patch
from turnstone.server import _build_history
def _make_stub_session(messages: list[dict[str, Any]]) -> Any:
"""Minimal ChatSession-shaped stub. ``_build_history`` only reads
``session.messages`` plus calls ``_load_verdict_indexes(ws_id)``
the latter we patch out below.
"""
return SimpleNamespace(messages=messages, _ws_id="ws-test")
def _build(messages: list[dict[str, Any]]) -> list[dict[str, Any]]:
"""Run ``_build_history`` against a stub session, bypassing the
verdicts / output-assessment storage round-trip (no tool_calls in
these tests, so the indexes are unused anyway).
"""
session = _make_stub_session(messages)
with patch(
"turnstone.server._load_verdict_indexes",
return_value=({}, {}),
):
return _build_history(session)
class TestSourceSurfacing:
def test_source_surfaces_when_set(self) -> None:
msg = {
"role": "user",
"content": "",
"_source": "system_nudge",
}
history = _build([msg])
assert len(history) == 1
assert history[0]["source"] == "system_nudge"
def test_source_absent_when_unset(self) -> None:
msg = {"role": "user", "content": "hello"}
history = _build([msg])
assert "source" not in history[0]
class TestRemindersWidening:
def test_watch_triggered_optional_fields_propagate(self) -> None:
"""The widened payload (Commit 2) carries watch_name / command /
poll_count / max_polls / is_final on each ``watch_triggered``
reminder so the frontend renders ``.msg.watch-result``.
"""
msg = {
"role": "user",
"content": "",
"_source": "system_nudge",
"_reminders": [
{
"type": "watch_triggered",
"text": "$ ls\nfile.txt",
"watch_name": "w1",
"command": "ls",
"poll_count": 2,
"max_polls": 100,
"is_final": False,
}
],
}
history = _build([msg])
assert history[0]["source"] == "system_nudge"
assert history[0]["reminders"] == [
{
"type": "watch_triggered",
"text": "$ ls\nfile.txt",
"watch_name": "w1",
"command": "ls",
"poll_count": 2,
"max_polls": 100,
"is_final": False,
}
]
def test_legacy_two_field_reminders_still_work(self) -> None:
"""Producers without optional fields (correction / denial /
idle_children) keep the legacy ``{type, text}`` shape the
widened filter just doesn't add anything beyond that."""
msg = {
"role": "user",
"content": "noted",
"_reminders": [{"type": "correction", "text": "watch out"}],
}
history = _build([msg])
assert history[0]["reminders"] == [{"type": "correction", "text": "watch out"}]
def test_unknown_keys_are_dropped(self) -> None:
"""The wire-layer filter projects on a known set of keys so a
future producer accidentally stuffing arbitrary fields can't
leak them through replay.
"""
msg = {
"role": "user",
"content": "x",
"_reminders": [
{
"type": "correction",
"text": "hi",
"secret": "leak-me",
"internal_id": 42,
}
],
}
history = _build([msg])
clean = history[0]["reminders"][0]
assert "secret" not in clean
assert "internal_id" not in clean
assert clean == {"type": "correction", "text": "hi"}
def test_malformed_reminder_skipped(self) -> None:
"""A non-dict / empty entry is filtered out instead of breaking
the rest of the list (mirrors the defensive filter in
``_apply_reminders_for_provider``).
"""
msg = {
"role": "user",
"content": "x",
"_reminders": [
"garbage string",
{"type": "", "text": ""}, # empty type + text → drop
{"type": "denial", "text": "ok"},
],
}
history = _build([msg])
assert history[0]["reminders"] == [{"type": "denial", "text": "ok"}]
class _StubRegistry:
"""Minimal model registry — only ``get_config`` is read by
``_build_history``."""
def __init__(self, surface_persisted_reasoning: bool = True) -> None:
self._cfg = SimpleNamespace(surface_persisted_reasoning=surface_persisted_reasoning)
def get_config(self, alias: str) -> Any:
return self._cfg
def _build_with_registry(
messages: list[dict[str, Any]],
surface_persisted_reasoning: bool = True,
) -> list[dict[str, Any]]:
session = SimpleNamespace(
messages=messages,
_ws_id="ws-test",
_registry=_StubRegistry(surface_persisted_reasoning=surface_persisted_reasoning),
_model_alias="claude-opus-4-7",
)
with patch(
"turnstone.server._load_verdict_indexes",
return_value=({}, {}),
):
return _build_history(session)
class TestReasoningSurfacing:
"""Phase 1 — surface stored Anthropic thinking blocks on the
history payload so refresh-the-page rehydrates the reasoning bubble.
Drives through the real ``AnthropicProvider`` extractor (no mock-of-
extractor) only the model registry is stubbed.
"""
def test_reasoning_surfaces_for_anthropic_thinking_msg(self) -> None:
msg = {
"role": "assistant",
"content": "Final answer.",
"_provider_content": [
{"type": "thinking", "thinking": "let me think", "signature": "s"},
{"type": "text", "text": "Final answer."},
],
}
history = _build_with_registry([msg], surface_persisted_reasoning=True)
assert len(history) == 1
assert history[0]["reasoning"] == "let me think"
def test_reasoning_empty_when_persist_flag_false(self) -> None:
msg = {
"role": "assistant",
"content": "Final answer.",
"_provider_content": [
{"type": "thinking", "thinking": "hidden", "signature": "s"},
],
}
history = _build_with_registry([msg], surface_persisted_reasoning=False)
assert "reasoning" not in history[0]
def test_provider_content_never_in_wire_entry(self) -> None:
# The build path does not copy ``_provider_content`` into the
# entry dict regardless of flag — wire payload stays tight.
msg = {
"role": "assistant",
"content": "Final answer.",
"_provider_content": [
{"type": "thinking", "thinking": "x", "signature": "s"},
],
}
history = _build_with_registry([msg], surface_persisted_reasoning=True)
assert "_provider_content" not in history[0]
def test_no_reasoning_field_when_provider_content_missing(self) -> None:
msg = {"role": "assistant", "content": "plain answer"}
history = _build_with_registry([msg], surface_persisted_reasoning=True)
assert "reasoning" not in history[0]
def test_no_reasoning_field_for_non_assistant_messages(self) -> None:
# Defensive — user/tool messages with a stray _provider_content
# do not get the reasoning field stamped.
msgs: list[dict[str, Any]] = [
{"role": "user", "content": "hi"},
{
"role": "tool",
"tool_call_id": "c1",
"content": "out",
"_provider_content": [{"type": "thinking", "thinking": "leak", "signature": "s"}],
},
]
history = _build_with_registry(msgs, surface_persisted_reasoning=True)
assert "reasoning" not in history[0]
assert "reasoning" not in history[1]
def test_default_true_when_registry_lookup_raises(self) -> None:
# Conservative default — Phase 1 spec mandates rehydration on
# refresh. A registry/alias mismatch must not silently kill the
# bubble.
class BrokenRegistry:
def get_config(self, alias: str) -> Any:
raise KeyError(alias)
session = SimpleNamespace(
messages=[
{
"role": "assistant",
"content": "x",
"_provider_content": [
{"type": "thinking", "thinking": "still works", "signature": "s"}
],
}
],
_ws_id="ws-test",
_registry=BrokenRegistry(),
_model_alias="missing-alias",
)
with patch(
"turnstone.server._load_verdict_indexes",
return_value=({}, {}),
):
history = _build_history(session)
assert history[0]["reasoning"] == "still works"
+127 -4
View File
@@ -19,6 +19,12 @@ class NullUI:
self.infos = []
self.stream_ends = 0
def on_turn_start(self):
pass
def on_turn_committed(self):
pass
def on_thinking_start(self):
pass
@@ -168,10 +174,18 @@ class TestCancelDuringStreaming:
assert ui.states[-1] == "idle"
# Check that "[Generation cancelled]" was emitted
assert any("cancelled" in i.lower() for i in ui.infos)
# The partial content should be preserved as an assistant message
# The partial content should be preserved as an assistant
# message AND annotated with a marker that downstream readers
# (inspect_workstream, the next coord turn) can use to
# distinguish a cancelled fragment from a completed turn — the
# raw "Hello world" without a marker would look like the
# final assistant answer to a coord LLM reading the child's
# transcript.
assistant_msgs = [m for m in session.messages if m["role"] == "assistant"]
assert len(assistant_msgs) == 1
assert assistant_msgs[0]["content"] == "Hello world"
content = assistant_msgs[0]["content"]
assert content.startswith("Hello world")
assert "[generation cancelled before completion]" in content
# No tool_calls in the partial message
assert "tool_calls" not in assistant_msgs[0]
@@ -511,10 +525,13 @@ class TestStreamAbort:
# Should complete as cancelled, not error
assert "idle" in ui.states
assert any("cancelled" in i.lower() for i in ui.infos)
# Partial content preserved
# Partial content preserved AND annotated with the
# cancelled-before-completion marker.
assistant_msgs = [m for m in session.messages if m["role"] == "assistant"]
assert len(assistant_msgs) == 1
assert assistant_msgs[0]["content"] == "Hello"
content = assistant_msgs[0]["content"]
assert content.startswith("Hello")
assert "[generation cancelled before completion]" in content
def test_non_cancel_exception_not_swallowed(self, tmp_db):
"""Exceptions during streaming that aren't caused by cancel
@@ -809,3 +826,109 @@ class TestForceCancelThreaded:
assert "idle" in ui.states
assistant_msgs = [m for m in session.messages if m["role"] == "assistant"]
assert any("Fresh response" in m.get("content", "") for m in assistant_msgs)
class TestSynthesizeCancelledResults:
"""Regression coverage for ``_synthesize_cancelled_results`` — must
fire ``on_tool_result`` for each synthesized cancellation so live
SSE listeners (e.g. coord's ``--running`` indicator added by
tool_info) can complete the in-DOM tool batch. Without this, the
coord JS would spin the running indicator forever on cancelled
batches because ``state_change`` doesn't strip ``--running`` from
individual batches."""
def _ui_with_tool_result_tracking(self):
class _TrackingUI(NullUI):
def __init__(self) -> None:
super().__init__()
self.tool_results: list[tuple[str, str, str, bool]] = []
def on_tool_result(self, call_id, name, output, **kwargs):
self.tool_results.append(
(call_id, name, output, bool(kwargs.get("is_error", False))),
)
return _TrackingUI()
def test_synthesizes_tool_result_for_unanswered_calls(self, tmp_db):
ui = self._ui_with_tool_result_tracking()
session = _make_session(ui=ui)
session.messages.append(
{
"role": "assistant",
"content": "calling tools",
"tool_calls": [
{"id": "call_a", "function": {"name": "search", "arguments": "{}"}},
{"id": "call_b", "function": {"name": "compute", "arguments": "{}"}},
],
},
)
session._msg_tokens.append(1)
session._synthesize_cancelled_results("Cancelled by user.")
# Both unanswered calls fired ``on_tool_result``.
assert len(ui.tool_results) == 2
ids = {tr[0] for tr in ui.tool_results}
assert ids == {"call_a", "call_b"}
# All emitted as errors so the live UI renders them as
# ``coord-tool-row-result--error``.
assert all(tr[3] is True for tr in ui.tool_results)
# Reason text propagates as the synthetic tool output.
assert all(tr[2] == "Cancelled by user." for tr in ui.tool_results)
# And the message list has the synthesized tool entries
# (preserves the prior contract).
tool_msgs = [m for m in session.messages if m.get("role") == "tool"]
assert len(tool_msgs) == 2
def test_skips_calls_already_answered(self, tmp_db):
ui = self._ui_with_tool_result_tracking()
session = _make_session(ui=ui)
session.messages.append(
{
"role": "assistant",
"tool_calls": [
{"id": "call_a", "function": {"name": "search", "arguments": "{}"}},
{"id": "call_b", "function": {"name": "compute", "arguments": "{}"}},
],
},
)
session._msg_tokens.append(1)
# call_a already answered.
session.messages.append(
{"role": "tool", "tool_call_id": "call_a", "content": "result"},
)
session._msg_tokens.append(1)
session._synthesize_cancelled_results("Cancelled by user.")
# Only call_b synthesized.
assert len(ui.tool_results) == 1
assert ui.tool_results[0][0] == "call_b"
def test_ui_emit_failure_does_not_break_synthesis(self, tmp_db):
"""The UI hook is wrapped in try/except — a hook failure
during cancel must NOT compound the problem. Synthesis still
appends to messages + storage."""
class _ExplodingUI(NullUI):
def on_tool_result(self, call_id, name, output, **kwargs):
raise RuntimeError("ui hook blew up")
ui = _ExplodingUI()
session = _make_session(ui=ui)
session.messages.append(
{
"role": "assistant",
"tool_calls": [
{"id": "call_a", "function": {"name": "search", "arguments": "{}"}},
],
},
)
session._msg_tokens.append(1)
# Must not raise.
session._synthesize_cancelled_results("Cancelled by user.")
tool_msgs = [m for m in session.messages if m.get("role") == "tool"]
assert len(tool_msgs) == 1
+270
View File
@@ -0,0 +1,270 @@
"""Unit tests for :mod:`turnstone.core.child_source`.
Covers both strategies in isolation against fakes no live collector,
no live SessionManager. Adapter-level integration coverage continues to
live in ``test_coordinator_adapter.py``.
"""
from __future__ import annotations
import contextlib
import time
from typing import TYPE_CHECKING, Any
from turnstone.core.child_source import ClusterChildSource, SameNodeChildSource
from turnstone.core.children_registry import ChildrenRegistry
from turnstone.core.workstream import WorkstreamState
if TYPE_CHECKING:
import queue
# ---------------------------------------------------------------------------
# SameNodeChildSource
# ---------------------------------------------------------------------------
class _FakeManager:
"""Minimal SessionManager stand-in implementing the subscribe API."""
def __init__(self) -> None:
self.subscribers: list[Any] = []
def subscribe_to_state(self, callback: Any) -> None:
self.subscribers.append(callback)
def unsubscribe_from_state(self, callback: Any) -> None:
with contextlib.suppress(ValueError):
self.subscribers.remove(callback)
def fire(self, ws_id: str, state: WorkstreamState) -> None:
for cb in self.subscribers:
cb(ws_id, state)
class TestSameNodeChildSource:
def test_start_subscribes_to_manager(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
sink_calls: list[dict[str, Any]] = []
src.start(sink=sink_calls.append)
assert len(mgr.subscribers) == 1
def test_state_change_for_known_child_pushes_to_sink(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
registry.install("p1", object())
registry.add_child("p1", "c1")
src = SameNodeChildSource(mgr, registry)
sink_calls: list[dict[str, Any]] = []
src.start(sink=sink_calls.append)
mgr.fire("c1", WorkstreamState.RUNNING)
assert len(sink_calls) == 1
ev = sink_calls[0]
assert ev["type"] == "cluster_state"
assert ev["ws_id"] == "c1"
assert ev["state"] == "running"
def test_state_change_for_unknown_workstream_is_dropped(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
sink_calls: list[dict[str, Any]] = []
src.start(sink=sink_calls.append)
# No registry entry — pre-filter drops the event without
# invoking the sink.
mgr.fire("ws-unknown", WorkstreamState.IDLE)
assert sink_calls == []
def test_shutdown_unsubscribes(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
src.start(sink=lambda ev: None)
assert len(mgr.subscribers) == 1
src.shutdown()
assert mgr.subscribers == []
def test_start_is_idempotent(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
src.start(sink=lambda ev: None)
src.start(sink=lambda ev: None)
# Second start is a no-op; only one subscription.
assert len(mgr.subscribers) == 1
def test_sink_exception_does_not_propagate(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
registry.install("p1", object())
registry.add_child("p1", "c1")
src = SameNodeChildSource(mgr, registry)
def bad_sink(ev: dict[str, Any]) -> None:
raise RuntimeError("sink boom")
src.start(sink=bad_sink)
# Should not raise — the strategy catches sink failures and logs.
mgr.fire("c1", WorkstreamState.RUNNING)
# ---------------------------------------------------------------------------
# ClusterChildSource
# ---------------------------------------------------------------------------
class _FakeCollector:
"""Minimal ClusterCollector stand-in providing the listener API."""
def __init__(self, snapshot: dict[str, Any] | None = None) -> None:
self._snapshot = snapshot or {"nodes": []}
self.queues: list[queue.Queue[dict[str, Any]]] = []
self.unregistered: list[queue.Queue[dict[str, Any]]] = []
def get_snapshot_and_register(self, q: queue.Queue[dict[str, Any]]) -> dict[str, Any]:
self.queues.append(q)
return self._snapshot
def unregister_listener(self, q: queue.Queue[dict[str, Any]]) -> None:
self.unregistered.append(q)
def emit(self, event: dict[str, Any]) -> None:
"""Push an event to all registered listener queues."""
for q in self.queues:
q.put(event)
class TestClusterChildSource:
def test_start_subscribes_to_collector(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
try:
src.start(sink=lambda ev: None)
assert len(coll.queues) == 1
finally:
src.shutdown()
def test_start_primes_registry_from_snapshot(self) -> None:
snapshot = {
"nodes": [
{
"workstreams": [
{"id": "c1", "parent_ws_id": "p1"},
{"id": "c2", "parent_ws_id": "p1"},
# Unknown parent — dropped
{"id": "x", "parent_ws_id": "p-unknown"},
],
},
],
}
coll = _FakeCollector(snapshot)
registry = ChildrenRegistry()
registry.install("p1", object())
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=lambda: ["p1"],
)
try:
src.start(sink=lambda ev: None)
assert set(registry.children_of("p1")) == {"c1", "c2"}
assert registry.parent_for("x") is None
finally:
src.shutdown()
def test_event_dispatched_to_sink(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
sink_calls: list[dict[str, Any]] = []
try:
src.start(sink=sink_calls.append)
coll.emit({"type": "cluster_state", "ws_id": "c1", "state": "running"})
# Daemon thread loop has 1.0s queue timeout; poll briefly.
for _ in range(20):
if sink_calls:
break
time.sleep(0.05)
assert len(sink_calls) == 1
assert sink_calls[0]["ws_id"] == "c1"
finally:
src.shutdown()
def test_shutdown_unregisters_and_joins_thread(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
src.start(sink=lambda ev: None)
src.shutdown()
assert coll.unregistered == coll.queues
# Second shutdown is a no-op (idempotent).
src.shutdown()
def test_start_is_idempotent(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
try:
src.start(sink=lambda ev: None)
src.start(sink=lambda ev: None)
assert len(coll.queues) == 1
finally:
src.shutdown()
def test_sink_exception_does_not_kill_thread(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
survived_calls: list[dict[str, Any]] = []
call_count = [0]
def flaky_sink(ev: dict[str, Any]) -> None:
call_count[0] += 1
if call_count[0] == 1:
raise RuntimeError("first one boom")
survived_calls.append(ev)
try:
src.start(sink=flaky_sink)
coll.emit({"type": "cluster_state", "ws_id": "c1", "state": "x"})
coll.emit({"type": "cluster_state", "ws_id": "c2", "state": "y"})
for _ in range(40):
if survived_calls:
break
time.sleep(0.05)
assert len(survived_calls) == 1
assert survived_calls[0]["ws_id"] == "c2"
finally:
src.shutdown()
# Multi-subscriber observer tests for ``SessionManager.subscribe_to_state``
# / ``unsubscribe_from_state`` live in ``test_session_manager.py`` where
# the proper FakeAdapter / FakeStorage construction helpers already exist.
+239
View File
@@ -0,0 +1,239 @@
"""Unit tests for :class:`turnstone.core.children_registry.ChildrenRegistry`.
The registry was lifted from ``CoordinatorAdapter`` in Stage 3 Step 1.
Adapter-level coverage for the integrated behavior already lives in
``test_coordinator_adapter.py``; this file pins the data structure
invariants in isolation so the registry can be reused by future
``ChildSource`` strategies (Step 2) without re-deriving the behavior
from the adapter test surface.
"""
from __future__ import annotations
import threading
import pytest
from turnstone.core.children_registry import ChildrenRegistry
class _Sentinel:
"""Lightweight UI stand-in; identity-comparable, no behavior."""
@pytest.fixture
def registry() -> ChildrenRegistry:
return ChildrenRegistry()
# ---------------------------------------------------------------------------
# install / uninstall
# ---------------------------------------------------------------------------
class TestInstallUninstall:
def test_install_seeds_empty_child_set_and_presence(self, registry: ChildrenRegistry) -> None:
ui = _Sentinel()
registry.install("p1", ui)
assert registry.children_of("p1") == []
assert registry.ui_for("p1") is ui
assert registry.parents() == ["p1"]
def test_install_is_idempotent_repoints_ui_keeps_children(
self, registry: ChildrenRegistry
) -> None:
ui_a = _Sentinel()
ui_b = _Sentinel()
registry.install("p1", ui_a)
registry.merge_children("p1", ["c1", "c2"])
registry.install("p1", ui_b)
assert registry.ui_for("p1") is ui_b
assert set(registry.children_of("p1")) == {"c1", "c2"}
def test_uninstall_clears_forward_reverse_and_presence(
self, registry: ChildrenRegistry
) -> None:
ui = _Sentinel()
registry.install("p1", ui)
registry.merge_children("p1", ["c1", "c2"])
registry.uninstall("p1")
assert registry.children_of("p1") == []
assert registry.ui_for("p1") is None
assert registry.parents() == []
assert registry.parent_for("c1") is None
assert registry.parent_for("c2") is None
def test_uninstall_unknown_parent_is_noop(self, registry: ChildrenRegistry) -> None:
registry.uninstall("never-installed") # must not raise
def test_uninstall_does_not_clobber_other_parents(self, registry: ChildrenRegistry) -> None:
registry.install("p1", _Sentinel())
registry.install("p2", _Sentinel())
registry.merge_children("p1", ["c1"])
registry.merge_children("p2", ["c2"])
registry.uninstall("p1")
assert registry.parent_for("c1") is None
assert registry.parent_for("c2") == "p2"
assert registry.parents() == ["p2"]
# ---------------------------------------------------------------------------
# add_child — atomic check-and-route
# ---------------------------------------------------------------------------
class TestAddChild:
def test_add_child_returns_ui_on_success(self, registry: ChildrenRegistry) -> None:
ui = _Sentinel()
registry.install("p1", ui)
assert registry.add_child("p1", "c1") is ui
assert registry.parent_for("c1") == "p1"
assert registry.children_of("p1") == ["c1"]
def test_add_child_returns_none_when_parent_not_installed(
self, registry: ChildrenRegistry
) -> None:
assert registry.add_child("absent", "c1") is None
assert registry.parent_for("c1") is None
def test_add_child_returns_none_on_duplicate(self, registry: ChildrenRegistry) -> None:
ui = _Sentinel()
registry.install("p1", ui)
assert registry.add_child("p1", "c1") is ui
# second add for same child returns None — caller must not
# double-dispatch.
assert registry.add_child("p1", "c1") is None
assert registry.children_of("p1") == ["c1"]
# ---------------------------------------------------------------------------
# merge_children — bulk seeding
# ---------------------------------------------------------------------------
class TestMergeChildren:
def test_merge_seeds_forward_and_reverse(self, registry: ChildrenRegistry) -> None:
registry.merge_children("p1", ["c1", "c2", "c3"])
assert set(registry.children_of("p1")) == {"c1", "c2", "c3"}
for cid in ("c1", "c2", "c3"):
assert registry.parent_for(cid) == "p1"
def test_merge_is_idempotent(self, registry: ChildrenRegistry) -> None:
registry.merge_children("p1", ["c1"])
registry.merge_children("p1", ["c1"])
assert registry.children_of("p1") == ["c1"]
def test_merge_skips_empty_or_falsy_ids(self, registry: ChildrenRegistry) -> None:
registry.merge_children("p1", ["", "c1", "", "c2"])
assert set(registry.children_of("p1")) == {"c1", "c2"}
def test_merge_does_not_require_install(self, registry: ChildrenRegistry) -> None:
# Snapshot-priming may run before the parent's install fires —
# the merge still seeds the forward set so the install picks
# the children up. (Storage-seeded rebuild relies on this.)
registry.merge_children("p1", ["c1"])
assert registry.children_of("p1") == ["c1"]
# ui_for is still None because install hasn't run
assert registry.ui_for("p1") is None
# ---------------------------------------------------------------------------
# Lookups — return copies, not live refs
# ---------------------------------------------------------------------------
class TestLookups:
def test_children_of_returns_copy(self, registry: ChildrenRegistry) -> None:
registry.install("p1", _Sentinel())
registry.merge_children("p1", ["c1", "c2"])
snap = registry.children_of("p1")
snap.append("c3-injected")
assert "c3-injected" not in registry.children_of("p1")
def test_children_of_unknown_parent_returns_empty(self, registry: ChildrenRegistry) -> None:
assert registry.children_of("absent") == []
def test_parent_for_unknown_child_returns_none(self, registry: ChildrenRegistry) -> None:
assert registry.parent_for("absent") is None
def test_parents_returns_copy(self, registry: ChildrenRegistry) -> None:
registry.install("p1", _Sentinel())
snap = registry.parents()
snap.append("p2-injected")
assert "p2-injected" not in registry.parents()
# ---------------------------------------------------------------------------
# Concurrency — concurrent add_child must not exceed the unique-set
# invariant or leave a half-installed reverse-index entry.
# ---------------------------------------------------------------------------
class TestConcurrency:
def test_concurrent_add_child_returns_ui_exactly_once_per_unique(
self, registry: ChildrenRegistry
) -> None:
ui = _Sentinel()
registry.install("p1", ui)
results: list[object] = []
results_lock = threading.Lock()
def attempt_add(child_id: str) -> None:
r = registry.add_child("p1", child_id)
with results_lock:
results.append(r)
threads = [threading.Thread(target=attempt_add, args=("c1",)) for _ in range(20)]
for t in threads:
t.start()
for t in threads:
t.join()
# Exactly one thread sees the UI; the remaining 19 see None
# (duplicate). The forward + reverse indexes carry exactly one
# entry for c1.
successes = [r for r in results if r is ui]
nones = [r for r in results if r is None]
assert len(successes) == 1
assert len(nones) == 19
assert registry.children_of("p1") == ["c1"]
assert registry.parent_for("c1") == "p1"
def test_concurrent_install_and_add_child_no_resurrect(
self, registry: ChildrenRegistry
) -> None:
# add_child racing with uninstall: either lands first (registry
# populated) or the parent is gone (returns None). Must NOT
# leave a forward-set entry without presence — that would be
# the "resurrected after close" leak the locked dispatch path
# was guarding against.
ui = _Sentinel()
registry.install("p1", ui)
outcomes: list[object] = []
def adder() -> None:
outcomes.append(registry.add_child("p1", "c1"))
def uninstaller() -> None:
registry.uninstall("p1")
threads = [
threading.Thread(target=adder),
threading.Thread(target=uninstaller),
]
for t in threads:
t.start()
for t in threads:
t.join()
# If add_child landed first: c1 is in the forward set, then
# uninstall clears everything. End state: nothing.
# If uninstall landed first: add_child sees no presence,
# returns None, no entry added. End state: nothing.
# Either way, the leak invariant holds: child set is empty or
# parent is gone, never "child set populated but no presence".
children = registry.children_of("p1")
ui_present = registry.ui_for("p1") is not None
if children:
assert ui_present, "registry leaked: children set without presence"
+33 -17
View File
@@ -30,9 +30,25 @@ def _full_hdr() -> dict[str, str]:
}
@pytest.fixture(autouse=True)
def _isolate_metrics(monkeypatch):
"""Swap ``turnstone.server._metrics`` for a fresh collector
per-test, with auto-restore.
Bare ``srv_mod._metrics = MetricsCollector()`` (the prior
pattern) leaks into any test file that already bound the name
via ``from turnstone.server import _metrics`` at import time
those tests' patches then operate on a different instance from
the one the live ``_publish_models_metadata`` reads, and the
monkeypatch silently no-ops. ``monkeypatch.setattr`` restores
after the test, so the leak is contained.
"""
fresh = MetricsCollector()
fresh.model = "test-model"
monkeypatch.setattr(srv_mod, "_metrics", fresh)
def _make_app(storage: Any) -> TestClient:
srv_mod._metrics = MetricsCollector()
srv_mod._metrics.model = "test-model"
mock_session = MagicMock()
mock_ws = MagicMock()
mock_ws.id = "ws-target"
@@ -50,7 +66,7 @@ def _make_app(storage: Any) -> TestClient:
mock_mgr.get.return_value = mock_ws
mock_mgr.close.return_value = True
mock_mgr.list_all.return_value = [mock_ws]
mock_mgr.max_workstreams = 10
mock_mgr.max_active = 10
app = srv_mod.create_app(
workstreams=mock_mgr,
@@ -73,8 +89,8 @@ def storage(tmp_path):
def test_close_with_reason_persists_to_workstream_config(storage):
client = _make_app(storage)
resp = client.post(
"/v1/api/workstreams/close",
json={"ws_id": "ws-target", "reason": "task complete"},
"/v1/api/workstreams/ws-target/close",
json={"reason": "task complete"},
headers=_full_hdr(),
)
assert resp.status_code == 200
@@ -85,8 +101,8 @@ def test_close_with_reason_persists_to_workstream_config(storage):
def test_close_without_reason_does_not_touch_config(storage):
client = _make_app(storage)
resp = client.post(
"/v1/api/workstreams/close",
json={"ws_id": "ws-target"},
"/v1/api/workstreams/ws-target/close",
json={},
headers=_full_hdr(),
)
assert resp.status_code == 200
@@ -102,8 +118,8 @@ def test_close_reason_capped_at_512_bytes(storage):
huge = "x" * 5000
client = _make_app(storage)
resp = client.post(
"/v1/api/workstreams/close",
json={"ws_id": "ws-target", "reason": huge},
"/v1/api/workstreams/ws-target/close",
json={"reason": huge},
headers=_full_hdr(),
)
assert resp.status_code == 200
@@ -120,8 +136,8 @@ def test_close_reason_byte_cap_holds_for_multibyte_utf8(storage):
huge = "\u6f22" * 600 # 3 bytes/char in UTF-8
client = _make_app(storage)
resp = client.post(
"/v1/api/workstreams/close",
json={"ws_id": "ws-target", "reason": huge},
"/v1/api/workstreams/ws-target/close",
json={"reason": huge},
headers=_full_hdr(),
)
assert resp.status_code == 200
@@ -137,8 +153,8 @@ def test_close_with_non_string_reason_drops_silently(storage):
proceeds without writing to workstream_config."""
client = _make_app(storage)
resp = client.post(
"/v1/api/workstreams/close",
json={"ws_id": "ws-target", "reason": {"unexpected": "shape"}},
"/v1/api/workstreams/ws-target/close",
json={"reason": {"unexpected": "shape"}},
headers=_full_hdr(),
)
assert resp.status_code == 200
@@ -154,8 +170,8 @@ def test_close_reason_redacts_credentials(storage):
client = _make_app(storage)
secret = "AKIAIOSFODNN7EXAMPLE" # AWS access key — output guard catches.
resp = client.post(
"/v1/api/workstreams/close",
json={"ws_id": "ws-target", "reason": f"task done; key={secret}"},
"/v1/api/workstreams/ws-target/close",
json={"reason": f"task done; key={secret}"},
headers=_full_hdr(),
)
assert resp.status_code == 200
@@ -177,8 +193,8 @@ def test_close_reason_persistence_failure_does_not_block_close(storage):
storage.save_workstream_config = _boom # type: ignore[method-assign]
resp = client.post(
"/v1/api/workstreams/close",
json={"ws_id": "ws-target", "reason": "task complete"},
"/v1/api/workstreams/ws-target/close",
json={"reason": "task complete"},
headers=_full_hdr(),
)
assert resp.status_code == 200
+382 -37
View File
@@ -31,16 +31,7 @@ _TEST_AUTH_HEADERS = {"Authorization": f"Bearer {_test_jwt()}"}
# Mock storage for collector tests
# ---------------------------------------------------------------------------
class MockStorage:
"""Minimal storage mock that implements list_services for collector tests."""
def __init__(self):
self.services: list[dict[str, str]] = []
def list_services(self, service_type: str, max_age_seconds: int = 120) -> list[dict[str, str]]:
return list(self.services)
from tests._coord_test_helpers import MockStorage # noqa: E402, F401
# ---------------------------------------------------------------------------
# Helpers
@@ -323,6 +314,46 @@ class TestCollectorSnapshot:
assert event["ws_id"] == "ws1"
assert event["state"] == "running"
def test_apply_snapshot_state_change_does_not_carry_pending_approval_detail(self):
"""Stage 3 cleanup — the snapshot-resync cluster_state event no
longer piggybacks ``pending_approval_detail`` (the field is
gone from cluster_state entirely). On reconnect the browser's
bulk fetch triggered by the ``activity_state="approval"``
transition in the reducer pulls the items directly from
``ui.serialize_pending_approval_detail()`` via the dashboard
endpoint."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "same", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_snapshot(
"node-a",
{
"type": "node_snapshot",
"node_id": "node-a",
"workstreams": [
{
"id": "ws1",
"name": "same",
"state": "running",
"activity_state": "approval",
}
],
"health": {},
"aggregate": {},
},
)
event = q.get_nowait()
assert event["type"] == "cluster_state"
assert event["activity_state"] == "approval"
assert "pending_approval_detail" not in event
def test_apply_snapshot_skips_empty_id_workstream(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
@@ -368,6 +399,36 @@ class TestCollectorDelta:
# Verify in-memory state was updated
assert c._nodes["node-a"].workstreams["ws1"]["state"] == "running"
def test_apply_delta_ws_state_does_not_carry_pending_approval_detail(self):
"""Stage 3 cleanup — ``cluster_state`` no longer carries the
``pending_approval_detail`` piggyback. Approval items now arrive
via bulk fetch on activity_state transition; verdicts via the
explicit ``intent_verdict`` event class. Symmetric event flow,
no piggyback to dedupe against."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta(
"node-a",
{
"type": "ws_state",
"ws_id": "ws1",
"state": "running",
"activity_state": "approval",
},
)
event = q.get_nowait()
assert event["type"] == "cluster_state"
assert event["activity_state"] == "approval"
assert "pending_approval_detail" not in event
def test_apply_delta_ws_created(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
@@ -413,6 +474,139 @@ class TestCollectorDelta:
assert event["name"] == "new-name"
assert c._nodes["node-a"].workstreams["ws1"]["name"] == "new-name"
def test_apply_delta_intent_verdict_forwards_verbatim(self):
"""Stage 3 Step 5 — node-emitted intent_verdict events flow
through _apply_delta to cluster fan-out so coord adapters can
re-emit as child_ws_intent_verdict on the parent's SSE."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
verdict = {
"call_id": "c1",
"risk_level": "low",
"confidence": 0.9,
"recommendation": "approve",
}
c._apply_delta(
"node-a",
{"type": "intent_verdict", "ws_id": "ws1", "verdict": verdict},
)
event = q.get_nowait()
assert event["type"] == "intent_verdict"
assert event["ws_id"] == "ws1"
assert event["node_id"] == "node-a"
assert event["verdict"] == verdict
def test_apply_delta_intent_verdict_drops_when_ws_id_missing(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta("node-a", {"type": "intent_verdict", "verdict": {}})
assert q.empty()
def test_apply_delta_approval_resolved_forwards_verbatim(self):
"""Stage 3 Step 5 — paired with intent_verdict; clears the
coord tree's pending-approval pill in lockstep with the
actual decision rather than waiting for the state-change
piggyback."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta(
"node-a",
{
"type": "approval_resolved",
"ws_id": "ws1",
"approved": True,
"feedback": "lgtm",
"always": False,
},
)
event = q.get_nowait()
assert event["type"] == "approval_resolved"
assert event["ws_id"] == "ws1"
assert event["node_id"] == "node-a"
assert event["approved"] is True
assert event["feedback"] == "lgtm"
assert event["always"] is False
def test_apply_delta_approve_request_forwards_detail(self):
"""Push path for the initial approval items — eliminates the
bulk-fetch race that left the coord row stuck on a loading
placeholder when the bulk fetch landed in the gap between
_emit_state(ATTENTION) and approve_tools setting _pending_approval."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
detail = {
"type": "approve_request",
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": True,
}
c._apply_delta(
"node-a",
{"type": "approve_request", "ws_id": "ws1", "detail": detail},
)
event = q.get_nowait()
assert event["type"] == "approve_request"
assert event["ws_id"] == "ws1"
assert event["node_id"] == "node-a"
assert event["detail"] == detail
def test_apply_delta_approve_request_drops_when_ws_id_missing(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta("node-a", {"type": "approve_request", "detail": {}})
assert q.empty()
def test_apply_delta_approval_resolved_coerces_missing_fields(self):
"""Defensive: ``approved`` / ``always`` / ``feedback`` may be
omitted by older nodes mid-rolling-upgrade; collector coerces
to safe defaults."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta("node-a", {"type": "approval_resolved", "ws_id": "ws1"})
event = q.get_nowait()
assert event["approved"] is False
assert event["feedback"] == ""
assert event["always"] is False
def test_apply_delta_health_changed(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
@@ -834,16 +1028,18 @@ class TestConsoleHTTPEndpoints:
resp = client.get("/nonexistent")
assert resp.status_code == 404
def test_index_has_new_ws_button(self, client):
def test_index_landing_surfaces(self, client):
status, body, ct = self._get_raw(client, "/")
assert status == 200
assert 'id="new-ws-btn"' in body
assert "showNewWsModal" in body
def test_index_has_new_ws_modal(self, client):
status, body, ct = self._get_raw(client, "/")
assert 'id="new-ws-overlay"' in body
assert 'id="new-ws-node"' in body
# Coordinator-first landing keeps the node list always-visible.
assert 'id="view-overview"' in body
assert 'id="node-table"' in body
# Removed in the 1.5.0 landing-page cleanup — guard against
# accidental reintroduction.
assert 'id="new-ws-overlay"' not in body
assert 'id="new-ws-btn"' not in body
assert 'id="cluster-summary-compact"' not in body
assert 'id="view-node"' not in body
# ---------------------------------------------------------------------------
@@ -1201,6 +1397,98 @@ class TestConsoleProxy:
)
assert resp.status_code == 404
def test_proxy_api_per_ws_events_routes_to_sse_handler(self, client, mock_collector):
"""``/node/{node_id}/v1/api/workstreams/{ws_id}/events`` is the
per-workstream SSE stream the interactive WebUI subscribes to.
Without explicit detection, the path falls through to the
regular GET branch and the EventSource API can't consume the
one-shot response Firefox surfaces it as "can't establish a
connection". Regression guard for the legacy URL surface
removal (#422) that moved per-ws SSE under
``/workstreams/{ws_id}/events`` without updating the proxy."""
from unittest.mock import AsyncMock, patch
from starlette.responses import Response
mock_collector.get_node_detail.return_value = {
"node_id": "node-a",
"server_url": "http://a:8080",
"reachable": True,
}
ws_id = "a" * 32
with (
patch(
"turnstone.console.server._proxy_sse",
new_callable=AsyncMock,
return_value=Response("ok", status_code=200),
) as sse_mock,
patch(
"turnstone.console.server._proxy_get",
new_callable=AsyncMock,
return_value=Response("ok", status_code=200),
) as get_mock,
):
client.get(f"/node/node-a/v1/api/workstreams/{ws_id}/events")
assert sse_mock.await_count == 1, (
"per-ws events path must route to _proxy_sse, not _proxy_get"
)
assert get_mock.await_count == 0
# Path passed to _proxy_sse must be the workstreams-prefixed
# form so the upstream URL is reconstructed correctly.
sse_args = sse_mock.await_args
assert sse_args.args[2] == f"workstreams/{ws_id}/events"
def test_proxy_api_global_events_still_routes_to_sse(self, client, mock_collector):
"""The bare ``events/global`` path was the only SSE path the
proxy recognized before the per-ws fix. Verify it still routes
correctly so the new branch didn't regress the existing case."""
from unittest.mock import AsyncMock, patch
from starlette.responses import Response
mock_collector.get_node_detail.return_value = {
"node_id": "node-a",
"server_url": "http://a:8080",
"reachable": True,
}
with patch(
"turnstone.console.server._proxy_sse",
new_callable=AsyncMock,
return_value=Response("ok", status_code=200),
) as sse_mock:
client.get("/node/node-a/v1/api/events/global")
assert sse_mock.await_count == 1
# events/global must use the console's service token —
# the upstream gates this path on `service` scope and
# end-user JWTs don't carry it. Without this, the
# browser's interactive UI 403-loops on every retry.
assert sse_mock.await_args.kwargs.get("use_service_auth") is True
def test_proxy_api_per_ws_events_uses_user_auth_not_service(self, client, mock_collector):
"""Per-ws events route uses the user's re-minted JWT, not the
service token the upstream per-ws SSE handler scopes by
user identity for tenant filtering, and a service-scoped
call would bypass that gate. Only ``events/global``
(cross-tenant inventory by design) opts into service auth."""
from unittest.mock import AsyncMock, patch
from starlette.responses import Response
mock_collector.get_node_detail.return_value = {
"node_id": "node-a",
"server_url": "http://a:8080",
"reachable": True,
}
ws_id = "b" * 32
with patch(
"turnstone.console.server._proxy_sse",
new_callable=AsyncMock,
return_value=Response("ok", status_code=200),
) as sse_mock:
client.get(f"/node/node-a/v1/api/workstreams/{ws_id}/events")
assert sse_mock.await_count == 1
assert sse_mock.await_args.kwargs.get("use_service_auth") is False
# ---------------------------------------------------------------------------
# Proxy URL rewriting unit tests (no HTTP needed)
@@ -1224,11 +1512,62 @@ class TestProxyRewriting:
assert "window.fetch" in _JS_PROXY_SHIM
assert "window.EventSource" in _JS_PROXY_SHIM
def test_console_banner_contains_placeholder(self):
from turnstone.console.server import _CONSOLE_BANNER_TEMPLATE
def test_js_shim_carries_node_id_placeholder(self):
"""The picker reads the current node_id from the shim's _nodeId
closure variable; the placeholder must be present and substitutable."""
from turnstone.console.server import _JS_PROXY_SHIM
assert "NODE_ID_PLACEHOLDER" in _CONSOLE_BANNER_TEMPLATE
assert "Console" in _CONSOLE_BANNER_TEMPLATE
assert "NODE_ID_PLACEHOLDER" in _JS_PROXY_SHIM
replaced = _JS_PROXY_SHIM.replace("NODE_ID_PLACEHOLDER", "node-a")
assert "node-a" in replaced
assert "NODE_ID_PLACEHOLDER" not in replaced
def test_js_shim_includes_picker_pieces(self):
"""Picker logic ships in the same IIFE as the prefix shim — verify
the moving parts are present so a future refactor doesn't silently
drop them. /v1/api/cluster/nodes is the lazy-fetch target;
#ui-header is the DOM anchor; console-node-pill is the trigger
class; ws-tab-dropdown is the menu shell we share with the
workstream chevron menu (style + behaviour parity); ArrowDown is
the keyboard-nav primitive that disambiguates this from a plain
click-only menu."""
from turnstone.console.server import _JS_PROXY_SHIM
# limit=1000 matches the collector's hard cap; without it the
# picker would silently drop nodes past the 100-default in
# clusters with >100 nodes.
assert "/v1/api/cluster/nodes?limit=1000" in _JS_PROXY_SHIM
assert "ui-header" in _JS_PROXY_SHIM
assert "console-node-pill" in _JS_PROXY_SHIM
assert "ws-tab-dropdown" in _JS_PROXY_SHIM
assert "ArrowDown" in _JS_PROXY_SHIM
assert "DOMContentLoaded" in _JS_PROXY_SHIM
def test_proxy_style_drops_banner_styles(self):
"""The legacy banner CSS classes (.console-banner, .ts-header-back-link
offsets, .dashboard-overlay top:32px hack) should be gone the new
picker lives inside #ui-header and doesn't need overlay offsets."""
from turnstone.console.server import _CONSOLE_PROXY_STYLE
assert ".console-banner" not in _CONSOLE_PROXY_STYLE
assert "dashboard-overlay" not in _CONSOLE_PROXY_STYLE
assert ".console-node-pill" in _CONSOLE_PROXY_STYLE
assert ".console-node-menu" in _CONSOLE_PROXY_STYLE
def test_proxy_style_uses_canonical_degraded_color(self):
"""Degraded health dot must use --accent (the canonical "needs
attention" token used by the cluster-overview node table at
console/static/style.css:548) and not --yellow. Yellow is reserved
for the dash-state attention dot, a stronger signal."""
from turnstone.console.server import _CONSOLE_PROXY_STYLE
assert "console-node-menu-item-dot--degraded" in _CONSOLE_PROXY_STYLE
# The degraded rule sits on its own line; assert it uses --accent
# by checking the CSS substring has --accent and not --yellow.
idx = _CONSOLE_PROXY_STYLE.find("console-node-menu-item-dot--degraded")
rule = _CONSOLE_PROXY_STYLE[idx : idx + 200]
assert "var(--accent)" in rule
assert "var(--yellow)" not in rule
def test_html_rewriting_changes_static_paths(self):
"""Simulate the proxy_index rewriting logic."""
@@ -1245,16 +1584,24 @@ class TestProxyRewriting:
assert 'href="/static/' not in rewritten
assert 'src="/static/' not in rewritten
def test_banner_injection_after_body(self):
"""Simulate the banner injection logic."""
from turnstone.console.server import _CONSOLE_BANNER_TEMPLATE
def test_shim_injection_after_body(self):
"""Simulate the proxy shim injection — the shim ships the node-id
and prefix as JS literals and renders the picker at runtime, so
we assert the substituted JS literals land in the page."""
from turnstone.console.server import _CONSOLE_PROXY_STYLE, _JS_PROXY_SHIM
sample_html = "<html><body><div>content</div></body></html>"
banner = _CONSOLE_BANNER_TEMPLATE.replace("NODE_ID_PLACEHOLDER", "node-a")
result = sample_html.replace("<body>", "<body>" + banner, 1)
assert "node-a" in result
assert "Console" in result
assert result.startswith("<html><body><div")
prefix = "/node/node-a"
shim_js = _JS_PROXY_SHIM.replace('"PREFIX_PLACEHOLDER"', json.dumps(prefix)).replace(
'"NODE_ID_PLACEHOLDER"', json.dumps("node-a")
)
injection = _CONSOLE_PROXY_STYLE + "<script>" + shim_js + "</script>"
result = sample_html.replace("<body>", "<body>" + injection, 1)
assert '"node-a"' in result
assert '"/node/node-a"' in result
assert "PREFIX_PLACEHOLDER" not in result
assert "NODE_ID_PLACEHOLDER" not in result
assert result.startswith("<html><body><style>")
# ---------------------------------------------------------------------------
@@ -1489,17 +1836,15 @@ class TestProxySharedStatic:
def test_proxy_shim_injected_in_html(self):
"""Verify shim is injected as inline script in proxied HTML."""
from turnstone.console.server import _CONSOLE_BANNER_TEMPLATE, _JS_PROXY_SHIM
from turnstone.console.server import _JS_PROXY_SHIM
sample_html = "<html><body><div>content</div></body></html>"
prefix = "/node/test-node"
banner = _CONSOLE_BANNER_TEMPLATE.replace("NODE_ID_PLACEHOLDER", "test-node")
shim = (
"<script>"
+ _JS_PROXY_SHIM.replace('"PREFIX_PLACEHOLDER"', json.dumps(prefix))
+ "</script>"
shim_js = _JS_PROXY_SHIM.replace('"PREFIX_PLACEHOLDER"', json.dumps(prefix)).replace(
'"NODE_ID_PLACEHOLDER"', json.dumps("test-node")
)
result = sample_html.replace("<body>", "<body>" + banner + shim, 1)
shim = "<script>" + shim_js + "</script>"
result = sample_html.replace("<body>", "<body>" + shim, 1)
assert "<script>" in result
assert "/node/test-node" in result
assert "window.fetch" in result
+345
View File
@@ -0,0 +1,345 @@
"""``GET /v1/api/models`` resolution-chain coverage.
The console handler resolves four defaults from settings + the enabled
model list:
* ``default_alias`` ``model.default_alias``
* ``channel_default_alias`` ``channels.default_model_alias``
* ``coordinator_default_alias`` ``coordinator.model_alias``, falling
back to ``default_alias`` when empty *or* pointing at a disabled /
removed alias (mirrors :mod:`turnstone.console.session_factory`).
* ``judge_default_alias`` ``judge.model``, falling back to the
resolved coordinator alias when empty *or* pointing at a value that
isn't an enabled alias. ``judge.model`` is alias-only — same
contract as the other model roles and
:class:`turnstone.core.judge.IntentJudge` silently inherits the
session model when an unknown value is configured, so the API
surfaces the resolved coordinator alias rather than echoing the
misconfigured string.
These tests pin each branch so the home composer's resolved-alias
placeholder stays correct as the precedence rules evolve.
"""
from __future__ import annotations
from typing import Any
import pytest
from starlette.applications import Starlette
from starlette.middleware import Middleware
from starlette.routing import Route
from starlette.testclient import TestClient
from tests._coord_test_helpers import _AuthMiddleware, _FakeConfigStore
from turnstone.console.server import list_available_models
from turnstone.core.storage._sqlite import SQLiteBackend
@pytest.fixture
def storage(tmp_path: Any) -> SQLiteBackend:
return SQLiteBackend(str(tmp_path / "available_models.db"))
def _seed_model(
storage: SQLiteBackend,
*,
definition_id: str,
alias: str,
model: str = "model-x",
enabled: bool = True,
) -> None:
storage.create_model_definition(
definition_id=definition_id,
alias=alias,
model=model,
provider="openai-compatible",
base_url="http://localhost:8000/v1",
api_key="sk-test",
context_window=8192,
capabilities="{}",
enabled=enabled,
created_by="admin",
)
class _StubRegistry:
"""Mimics the surface ``resolve_coordinator_alias`` reads from
``coord_registry``: ``.default`` and ``.has_alias()``.
Production wires this through ``ModelRegistry``, which in turn
pulls aliases from both DB rows and config.toml. The fixture
mirrors the storage's enabled-row set so ``has_alias()`` agrees
with what the placeholder's enabled-row filter would accept —
without that alignment the helper rejects every tier-2 candidate
and the placeholder goes blank in cases that production handles
fine."""
def __init__(self, *, default: str, known: set[str]) -> None:
self.default = default
self._known = known
def has_alias(self, alias: str) -> bool:
return alias in self._known
def _make_client(
storage: SQLiteBackend,
*,
settings: dict[str, str] | None = None,
registry_default: str = "",
config_store: bool = True,
) -> TestClient:
app = Starlette(
routes=[Route("/v1/api/models", list_available_models)],
middleware=[Middleware(_AuthMiddleware)],
)
app.state.auth_storage = storage
if config_store:
app.state.config_store = _FakeConfigStore(dict(settings or {}))
# ``coord_registry`` is always set in production after lifespan
# startup; mirror that here. ``has_alias`` answers from the same
# enabled-rows set the handler filters against.
enabled = {r["alias"] for r in storage.list_model_definitions(enabled_only=True)}
app.state.coord_registry = _StubRegistry(default=registry_default, known=enabled)
client = TestClient(app)
client.headers.update({"X-Test-User": "admin", "X-Test-Perms": ""})
return client
def _get_models(client: TestClient) -> dict[str, Any]:
resp = client.get("/v1/api/models")
assert resp.status_code == 200, resp.text
return resp.json()
# ---------------------------------------------------------------------------
# Coordinator resolution
# ---------------------------------------------------------------------------
def test_no_settings_leaves_all_defaults_blank(storage: SQLiteBackend) -> None:
"""No model.default_alias, no per-role overrides → every default
field is empty and ``models`` is an empty list."""
body = _get_models(_make_client(storage))
assert body == {
"models": [],
"default_alias": "",
"channel_default_alias": "",
"coordinator_default_alias": "",
"judge_default_alias": "",
}
def test_coordinator_inherits_default_alias_when_unset(
storage: SQLiteBackend,
) -> None:
_seed_model(storage, definition_id="m1", alias="primary")
body = _get_models(_make_client(storage, settings={"model.default_alias": "primary"}))
assert body["default_alias"] == "primary"
assert body["coordinator_default_alias"] == "primary"
def test_coordinator_explicit_enabled_alias_passes_through(
storage: SQLiteBackend,
) -> None:
_seed_model(storage, definition_id="m1", alias="primary")
_seed_model(storage, definition_id="m2", alias="fast")
body = _get_models(
_make_client(
storage,
settings={
"model.default_alias": "primary",
"coordinator.model_alias": "fast",
},
)
)
assert body["coordinator_default_alias"] == "fast"
def test_coordinator_set_to_disabled_alias_falls_back_to_default(
storage: SQLiteBackend,
) -> None:
"""Operator disabled the alias the coordinator was pinned to —
fall back to the registry default rather than advertising a model
that workstream creation would refuse to use."""
_seed_model(storage, definition_id="m1", alias="primary")
_seed_model(storage, definition_id="m2", alias="legacy", enabled=False)
body = _get_models(
_make_client(
storage,
settings={
"model.default_alias": "primary",
"coordinator.model_alias": "legacy",
},
)
)
assert body["coordinator_default_alias"] == "primary"
def test_coordinator_set_to_unknown_alias_falls_back_to_default(
storage: SQLiteBackend,
) -> None:
_seed_model(storage, definition_id="m1", alias="primary")
body = _get_models(
_make_client(
storage,
settings={
"model.default_alias": "primary",
"coordinator.model_alias": "ghost",
},
)
)
assert body["coordinator_default_alias"] == "primary"
def test_coordinator_falls_back_to_registry_default_when_config_store_empty(
storage: SQLiteBackend,
) -> None:
"""Match ``console/session_factory.py:109-110``: when both
``coordinator.model_alias`` and ``model.default_alias`` are unset, new
coordinator sessions run on ``registry.default`` (loaded from
config.toml ``[model].default``). The placeholder must report the
same alias rather than going blank otherwise the home composer
advertises "Default model" while sessions actually launch on a
concrete alias."""
_seed_model(storage, definition_id="m1", alias="primary")
body = _get_models(_make_client(storage, registry_default="primary"))
assert body["default_alias"] == ""
assert body["coordinator_default_alias"] == "primary"
assert body["judge_default_alias"] == "primary"
def test_coordinator_skips_registry_default_when_alias_disabled(
storage: SQLiteBackend,
) -> None:
"""Registry default points at an alias that's been disabled in the DB
the placeholder stays blank rather than advertising a model that
workstream creation would refuse to use."""
_seed_model(storage, definition_id="m1", alias="legacy", enabled=False)
body = _get_models(_make_client(storage, registry_default="legacy"))
assert body["coordinator_default_alias"] == ""
def test_coordinator_falls_back_to_registry_default_when_config_store_missing(
storage: SQLiteBackend,
) -> None:
"""Edge case from PR #500 review: lifespan can leave
``app.state.config_store`` as None (e.g. a startup exception) while
``coord_registry`` still binds successfully. The placeholder must
still advertise ``registry.default`` (filtered against enabled rows)
rather than going blank otherwise the home composer is uselessly
empty in a degraded-but-recoverable state."""
_seed_model(storage, definition_id="m1", alias="primary")
body = _get_models(_make_client(storage, registry_default="primary", config_store=False))
assert body["default_alias"] == ""
assert body["coordinator_default_alias"] == "primary"
assert body["judge_default_alias"] == "primary"
# ---------------------------------------------------------------------------
# Judge resolution
# ---------------------------------------------------------------------------
def test_judge_empty_inherits_resolved_coordinator_alias(
storage: SQLiteBackend,
) -> None:
_seed_model(storage, definition_id="m1", alias="primary")
_seed_model(storage, definition_id="m2", alias="fast")
body = _get_models(
_make_client(
storage,
settings={
"model.default_alias": "primary",
"coordinator.model_alias": "fast",
},
)
)
assert body["coordinator_default_alias"] == "fast"
assert body["judge_default_alias"] == "fast"
def test_judge_explicit_enabled_alias_passes_through(
storage: SQLiteBackend,
) -> None:
_seed_model(storage, definition_id="m1", alias="primary")
_seed_model(storage, definition_id="m2", alias="judge-fast")
body = _get_models(
_make_client(
storage,
settings={
"model.default_alias": "primary",
"judge.model": "judge-fast",
},
)
)
assert body["judge_default_alias"] == "judge-fast"
def test_judge_set_to_unknown_value_inherits_coordinator(
storage: SQLiteBackend,
) -> None:
"""``judge.model`` is alias-only — same contract as the other model
roles. An unknown value silently inherits the session model in
:class:`IntentJudge`, so the API surfaces the resolved coordinator
alias rather than echoing the misconfigured string."""
_seed_model(storage, definition_id="m1", alias="primary")
body = _get_models(
_make_client(
storage,
settings={
"model.default_alias": "primary",
"judge.model": "anthropic/claude-haiku-4-5", # raw, not an alias
},
)
)
assert body["coordinator_default_alias"] == "primary"
assert body["judge_default_alias"] == "primary"
def test_judge_set_to_disabled_alias_inherits_coordinator(
storage: SQLiteBackend,
) -> None:
"""Disabled-alias case is handled identically to the unknown-value
case both trip the alias-not-resolved path."""
_seed_model(storage, definition_id="m1", alias="primary")
_seed_model(storage, definition_id="m2", alias="judge-old", enabled=False)
body = _get_models(
_make_client(
storage,
settings={
"model.default_alias": "primary",
"judge.model": "judge-old",
},
)
)
assert body["judge_default_alias"] == "primary"
# ---------------------------------------------------------------------------
# Pre-existing fields stay correct under the new resolution code
# ---------------------------------------------------------------------------
def test_channel_default_alias_blanked_when_disabled(
storage: SQLiteBackend,
) -> None:
_seed_model(storage, definition_id="m1", alias="primary", enabled=False)
body = _get_models(
_make_client(
storage,
settings={"channels.default_model_alias": "primary"},
)
)
assert body["channel_default_alias"] == ""
def test_models_payload_strips_secret_fields(storage: SQLiteBackend) -> None:
"""Regression guard: only alias/model/provider land in the response,
never api_key / base_url / context_window / capabilities."""
_seed_model(storage, definition_id="m1", alias="primary")
body = _get_models(_make_client(storage))
assert body["models"] == [
{"alias": "primary", "model": "model-x", "provider": "openai-compatible"}
]
+106
View File
@@ -0,0 +1,106 @@
"""Tests for the console's coordinator idle-cleanup thread helper.
The helper itself is a tiny loop wrapping ``mgr.close_idle``; the heavy
lifting is in ``SessionManager.close_idle`` (covered in
``test_session_manager.py``) and ``bulk_close_stale_orphans`` (covered
in ``test_storage_sqlite.py``). These tests verify the glue:
- the helper runs an initial sweep BEFORE its first sleep (cold-start
cleanup without blocking the lifespan),
- the helper swallows exceptions so a transient DB blip can't kill the
daemon thread,
- the helper exits cleanly when ``stop_event`` is set.
The ``stop_event`` parameter is exclusively for tests production
callers pass ``None`` and the daemon runs for process lifetime.
"""
from __future__ import annotations
import threading
from unittest.mock import patch
from turnstone.console.server import _coord_idle_cleanup_thread
class _StubMgr:
def __init__(
self, *, stop_event: threading.Event, expected_calls: int, raise_after: int = -1
) -> None:
self.calls: list[float] = []
self.sleep_calls_at_each_close: list[int] = []
self._stop_event = stop_event
self._expected = expected_calls
self._raise_after = raise_after
self._sleep_count = 0
def close_idle(self, timeout_sec: float) -> list[str]:
# Snapshot how many sleeps preceded this close — lets the
# "initial sweep" test verify the first close_idle ran with
# zero preceding sleeps.
self.sleep_calls_at_each_close.append(self._sleep_count)
self.calls.append(timeout_sec)
try:
if 0 <= self._raise_after < len(self.calls):
raise RuntimeError("simulated DB blip")
finally:
# Set stop after the helper has been exercised enough,
# regardless of whether this call raised.
if len(self.calls) >= self._expected:
self._stop_event.set()
return []
def record_sleep(self, _seconds: float) -> None:
self._sleep_count += 1
def _run_until_done(mgr: _StubMgr, stop_event: threading.Event, timeout_sec: float) -> None:
with patch("turnstone.console.server.time.sleep", mgr.record_sleep):
thread = threading.Thread(
target=_coord_idle_cleanup_thread,
args=(mgr, timeout_sec, stop_event),
daemon=True,
)
thread.start()
thread.join(timeout=2.0)
assert not thread.is_alive(), "helper failed to exit on stop_event"
def test_coord_idle_cleanup_runs_initial_sweep_before_sleep() -> None:
"""The first close_idle call must happen BEFORE the first time.sleep —
otherwise cold-start orphans wait one ``check_every`` interval (~30 min
on default 2h timeout) for the first reap. Crucial because the
lifespan no longer does a synchronous initial sweep."""
stop_event = threading.Event()
mgr = _StubMgr(stop_event=stop_event, expected_calls=1)
_run_until_done(mgr, stop_event, timeout_sec=120.0)
assert mgr.sleep_calls_at_each_close == [0], "first close_idle should run before any sleep"
def test_coord_idle_cleanup_calls_close_idle_each_tick() -> None:
stop_event = threading.Event()
mgr = _StubMgr(stop_event=stop_event, expected_calls=3)
_run_until_done(mgr, stop_event, timeout_sec=120.0)
assert len(mgr.calls) == 3
assert all(t == 120.0 for t in mgr.calls)
def test_coord_idle_cleanup_survives_close_idle_exceptions() -> None:
"""A transient DB error must not kill the daemon thread — the next
tick should still fire close_idle. Without the try/except, a single
blip would silently leak orphans forever."""
stop_event = threading.Event()
mgr = _StubMgr(stop_event=stop_event, expected_calls=4, raise_after=1)
_run_until_done(mgr, stop_event, timeout_sec=120.0)
# All four calls must have fired despite calls 2-4 raising.
assert len(mgr.calls) == 4
def test_coord_idle_cleanup_exits_cleanly_on_stop_event() -> None:
"""The stop_event mechanism is the test contract; verify the thread
actually exits when the event is set, without needing exceptions or
daemon-process termination."""
stop_event = threading.Event()
mgr = _StubMgr(stop_event=stop_event, expected_calls=2)
_run_until_done(mgr, stop_event, timeout_sec=120.0)
assert stop_event.is_set()
+27
View File
@@ -37,6 +37,33 @@ class TestRecordRoute:
assert "turnstone_router_request_duration_seconds_sum" in text
class TestRecordJudgeVerdict:
"""Coord-side intent-judge verdict counter."""
def test_single_verdict(self) -> None:
m = ConsoleMetrics()
m.record_judge_verdict("heuristic", "high", 12)
text = m.generate_text()
assert 'turnstone_judge_verdicts_total{tier="heuristic",risk_level="high"} 1' in text
def test_aggregates_by_tier_and_risk(self) -> None:
m = ConsoleMetrics()
m.record_judge_verdict("heuristic", "low", 5)
m.record_judge_verdict("heuristic", "low", 7)
m.record_judge_verdict("llm", "high", 250)
text = m.generate_text()
assert 'turnstone_judge_verdicts_total{tier="heuristic",risk_level="low"} 2' in text
assert 'turnstone_judge_verdicts_total{tier="llm",risk_level="high"} 1' in text
def test_section_omitted_when_empty(self) -> None:
"""No verdicts recorded → don't emit the empty header block."""
m = ConsoleMetrics()
text = m.generate_text()
assert "turnstone_judge_verdicts_total" not in text
class TestRouterInfo:
"""Live-membership gauge + refresh counter."""
+34 -20
View File
@@ -93,6 +93,15 @@ def _wire_proxy(app: Any, mock_post: MagicMock | None = None) -> None:
mock_post = _make_proxy_post()
mock_proxy = MagicMock(spec=httpx.AsyncClient)
mock_proxy.post = mock_post
# route_proxy uses ``client.request(method, url, ...)`` for path-keyed
# routes (so DELETE on /send proxies through correctly). Wire a
# request-shim that drops the leading method positional and forwards
# to the same mock_post for compatibility.
async def _request_shim(method: str, *args: Any, **kwargs: Any) -> httpx.Response:
return await mock_post(*args, **kwargs)
mock_proxy.request = MagicMock(side_effect=_request_shim)
app.state.proxy_client = mock_proxy
@@ -283,7 +292,8 @@ class TestRouteCreate503Retry:
class TestRouteProxy:
"""POST /v1/api/route/send (and other routed endpoints)."""
"""POST /v1/api/route/workstreams/{ws_id}/<verb> (and the surviving
body-keyed plan/command routes)."""
@pytest.fixture()
def client(self):
@@ -296,29 +306,33 @@ class TestRouteProxy:
def test_route_proxy_send(self, client):
resp = client.post(
"/v1/api/route/send",
json={"ws_id": "abc123", "message": "hello"},
"/v1/api/route/workstreams/abc123/send",
json={"message": "hello"},
headers=_TEST_AUTH_HEADERS,
)
assert resp.status_code == 200
# Verify upstream URL was /v1/api/send (not /v1/api/route/send)
mock_post = client.app.state.proxy_client.post
call_args = mock_post.call_args
assert "/v1/api/send" in call_args[0][0]
assert "/route/" not in call_args[0][0]
# Verify upstream URL was /v1/api/workstreams/abc123/send
# (not /v1/api/route/workstreams/abc123/send).
mock_request = client.app.state.proxy_client.request
call_args = mock_request.call_args
# request is called as ``request(method, url, ...)`` — url is the
# second positional arg.
upstream_url = call_args[0][1]
assert "/v1/api/workstreams/abc123/send" in upstream_url
assert "/route/" not in upstream_url
def test_route_proxy_approve(self, client):
resp = client.post(
"/v1/api/route/approve",
json={"ws_id": "abc123", "approved": True},
"/v1/api/route/workstreams/abc123/approve",
json={"approved": True},
headers=_TEST_AUTH_HEADERS,
)
assert resp.status_code == 200
def test_route_proxy_cancel(self, client):
resp = client.post(
"/v1/api/route/cancel",
json={"ws_id": "abc123"},
"/v1/api/route/workstreams/abc123/cancel",
json={},
headers=_TEST_AUTH_HEADERS,
)
assert resp.status_code == 200
@@ -333,8 +347,8 @@ class TestRouteProxy:
def test_route_proxy_close(self, client):
resp = client.post(
"/v1/api/route/workstreams/close",
json={"ws_id": "abc123"},
"/v1/api/route/workstreams/abc123/close",
json={},
headers=_TEST_AUTH_HEADERS,
)
assert resp.status_code == 200
@@ -413,8 +427,8 @@ class TestRouteNotReady:
def test_route_proxy_no_router_503(self, client_no_router):
resp = client_no_router.post(
"/v1/api/route/send",
json={"ws_id": "abc", "message": "hello"},
"/v1/api/route/workstreams/abc/send",
json={"message": "hello"},
headers=_TEST_AUTH_HEADERS,
)
assert resp.status_code == 503
@@ -425,8 +439,8 @@ class TestRouteNotReady:
def test_route_proxy_empty_cache_503(self, client_empty_cache):
resp = client_empty_cache.post(
"/v1/api/route/send",
json={"ws_id": "abc", "message": "hello"},
"/v1/api/route/workstreams/abc/send",
json={"message": "hello"},
headers=_TEST_AUTH_HEADERS,
)
assert resp.status_code == 503
@@ -465,8 +479,8 @@ class TestRouteNoNode:
def test_route_proxy_no_node_503(self, client):
resp = client.post(
"/v1/api/route/send",
json={"ws_id": "abc", "message": "hello"},
"/v1/api/route/workstreams/abc/send",
json={"message": "hello"},
headers=_TEST_AUTH_HEADERS,
)
assert resp.status_code == 503
+219
View File
@@ -0,0 +1,219 @@
"""``console/session_factory.py`` alias-resolution coverage.
The console session factory resolves the coordinator alias through a
three-tier chain that must stay in lockstep with the placeholder logic
in ``console/server.py:list_available_models`` otherwise the home
composer advertises one alias while sessions launch on another.
Tier order (highest priority first):
1. Per-call ``model_alias`` arg, or the ``coordinator.model_alias``
ConfigStore setting (admin-pinned coordinator-specific override).
2. ``model.default_alias`` ConfigStore setting (admin-managed system
default surfaced in the Models tab).
3. ``registry.default`` (config.toml ``[model].default``, the boot-time
fallback).
These tests pin each branch by intercepting ``registry.resolve``
they short-circuit before ChatSession construction so the test never
has to satisfy ChatSession's full kwarg contract.
"""
from __future__ import annotations
from typing import Any
from unittest.mock import MagicMock
import pytest
from tests._coord_test_helpers import _FakeConfigStore
from turnstone.console.session_factory import build_console_session_factory
class _StopBeforeChatSessionError(Exception):
"""Sentinel raised by the capturing registry to short-circuit
factory execution after alias resolution but before ChatSession is
built. The factory's outer code path is irrelevant to alias
resolution and would force the test to satisfy a long kwarg
contract for no extra coverage."""
class _CapturingRegistry:
"""Records the alias passed to ``resolve()`` and short-circuits.
``has_alias`` answers from the configured known set so the
``model.default_alias`` validation tier behaves realistically.
Mirrors the public surface ``ModelRegistry`` exposes to
session_factory: ``has_alias``, ``resolve``, and ``default``.
"""
def __init__(self, *, default: str, known: set[str]) -> None:
self.default = default
self._known = known
self.captured_alias: str | None = None
def has_alias(self, alias: str) -> bool:
return alias in self._known
def resolve(self, alias: str) -> Any:
self.captured_alias = alias
raise _StopBeforeChatSessionError()
def _build_factory(
*,
registry_default: str = "registry-default",
known_aliases: set[str] | None = None,
settings: dict[str, Any] | None = None,
) -> tuple[Any, _CapturingRegistry]:
"""Construct the factory with stub deps. Returns ``(factory_callable,
registry)`` so tests can read back ``registry.captured_alias``."""
registry = _CapturingRegistry(
default=registry_default,
known=known_aliases if known_aliases is not None else {registry_default},
)
config_store = _FakeConfigStore(dict(settings or {}))
factory = build_console_session_factory(
registry=registry, # type: ignore[arg-type]
config_store=config_store, # type: ignore[arg-type]
node_id="console",
coord_client_factory=lambda ws_id, uid: MagicMock(),
)
return factory, registry
def _invoke(factory: Any, **factory_kwargs: Any) -> None:
"""Call the factory with a stub UI and absorb the sentinel.
Forwards ``factory_kwargs`` to the factory so per-call overrides
(e.g. ``model_alias``) can flow through. Raises if any other
exception comes out the test should fail loudly when alias
resolution itself errors rather than swallowing it.
"""
ui = MagicMock()
ui._user_id = "" # skip storage-backed username lookup branch
with pytest.raises(_StopBeforeChatSessionError):
factory(ui, **factory_kwargs)
# ---------------------------------------------------------------------------
# Tier 1 — explicit pin (per-call arg or coordinator.model_alias)
# ---------------------------------------------------------------------------
def test_per_call_model_alias_arg_wins_over_everything() -> None:
"""The ``model_alias`` kwarg on the factory call (e.g. body field on
POST /workstreams/new) wins over both ConfigStore tiers and the
registry default."""
factory, registry = _build_factory(
known_aliases={"per-call", "coord-pin", "admin-default", "registry-default"},
settings={
"coordinator.model_alias": "coord-pin",
"model.default_alias": "admin-default",
},
)
_invoke(factory, model_alias="per-call")
assert registry.captured_alias == "per-call"
def test_coordinator_model_alias_wins_when_no_per_call_override() -> None:
factory, registry = _build_factory(
known_aliases={"coord-pin", "admin-default", "registry-default"},
settings={
"coordinator.model_alias": "coord-pin",
"model.default_alias": "admin-default",
},
)
_invoke(factory)
assert registry.captured_alias == "coord-pin"
def test_coordinator_model_alias_passed_through_unvalidated() -> None:
"""Tier 1 is an *explicit* operator pin — when it's stale or typoed
we deliberately pass it through to ``registry.resolve`` so the
request layer turns it into a 503 with the alias surfaced in the
error. Falling through silently would mask the misconfiguration."""
factory, registry = _build_factory(
known_aliases={"admin-default", "registry-default"},
settings={
"coordinator.model_alias": "ghost", # unknown
"model.default_alias": "admin-default",
},
)
_invoke(factory)
assert registry.captured_alias == "ghost"
def test_per_call_model_alias_arg_passed_through_unvalidated() -> None:
"""The per-call ``model_alias`` kwarg (POST body field — the more
common production trigger) is the same kind of explicit pin as the
ConfigStore setting, so a stale value passes through to
``registry.resolve`` rather than silently falling through to the
system default."""
factory, registry = _build_factory(
known_aliases={"registry-default"},
settings={"model.default_alias": "registry-default"},
)
_invoke(factory, model_alias="ghost")
assert registry.captured_alias == "ghost"
# ---------------------------------------------------------------------------
# Tier 2 — model.default_alias (admin-managed system default)
# ---------------------------------------------------------------------------
def test_model_default_alias_used_when_coordinator_unset() -> None:
"""Regression for the historical drift: admin sets the system
default in the Models tab, the home composer advertises it, and new
coordinator sessions must launch on the same alias rather than
silently falling through to ``registry.default``."""
factory, registry = _build_factory(
known_aliases={"admin-default", "registry-default"},
settings={"model.default_alias": "admin-default"},
)
_invoke(factory)
assert registry.captured_alias == "admin-default"
def test_unknown_model_default_alias_falls_through_to_registry_default() -> None:
"""Tier 2 is *not* an explicit pin — operators set
``model.default_alias`` once in the UI and forget about it; an alias
that's later disabled or typo'd should not 503 the coordinator,
since tier 3 (``registry.default``) is guaranteed to resolve."""
factory, registry = _build_factory(
known_aliases={"registry-default"}, # admin-default got removed
settings={"model.default_alias": "admin-default"},
)
_invoke(factory)
assert registry.captured_alias == "registry-default"
def test_blank_model_default_alias_falls_through_to_registry_default() -> None:
factory, registry = _build_factory(
settings={"model.default_alias": ""},
)
_invoke(factory)
assert registry.captured_alias == "registry-default"
# ---------------------------------------------------------------------------
# Tier 3 — registry.default (config.toml [model].default)
# ---------------------------------------------------------------------------
def test_no_settings_uses_registry_default() -> None:
factory, registry = _build_factory()
_invoke(factory)
assert registry.captured_alias == "registry-default"
def test_whitespace_only_coord_alias_falls_through() -> None:
"""``" "`` is not an explicit pin — ``.strip()`` reduces it to
"", which the chain should treat as unset."""
factory, registry = _build_factory(
settings={"coordinator.model_alias": " "},
)
_invoke(factory)
assert registry.captured_alias == "registry-default"
+599
View File
@@ -0,0 +1,599 @@
"""Tests for the rich ``ws_state`` payload on coord (Stage 2 follow-up).
Pre-lift coord's ``ConsoleCoordinatorUI`` populated none of the per-ws
metric fields ``SessionUIBase`` defines (``_ws_prompt_tokens`` /
``_ws_context_ratio`` / ``_ws_current_activity`` / ``_ws_turn_content``)
and the ``coord_adapter.emit_state`` broadcast was state-only
``tokens=0`` / ``content=""`` were hardcoded into
``collector.emit_console_ws_state``. The lift turned ``on_status`` /
``on_content_token`` / ``on_thinking_*`` / ``on_tool_result`` into
shared bodies on :class:`SessionUIBase` so coord populates the same
fields, then enriched ``coord_adapter.emit_state`` to read them under
lock and pass through to the cluster collector with the rich kwargs.
The cluster dashboard's coord rows now render with the same
tokens / activity / content / context_ratio fields interactive rows do.
"""
from __future__ import annotations
import threading
from typing import Any
from unittest.mock import MagicMock, patch
import pytest
from turnstone.console.coordinator_ui import ConsoleCoordinatorUI
from turnstone.core.session_ui_base import _MAX_TURN_CONTENT_CHARS
from turnstone.core.workstream import WorkstreamState
# ---------------------------------------------------------------------------
# Per-ws metric writes — lifted to SessionUIBase, both subclasses inherit
# ---------------------------------------------------------------------------
def _patch_get_storage(storage: Any):
return patch("turnstone.core.storage._registry.get_storage", return_value=storage)
def test_coord_on_status_writes_per_ws_metrics() -> None:
"""Pre-lift coord ``on_status`` was an enqueue-only stub — ``_ws_*``
fields stayed at their initial zero values regardless of token usage.
Post-lift coord inherits SessionUIBase's body, so token counters and
context ratio populate just like interactive."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
with _patch_get_storage(MagicMock()):
ui.on_status(
{"prompt_tokens": 100, "completion_tokens": 50},
context_window=1000,
effort="medium",
)
assert ui._ws_prompt_tokens == 100
assert ui._ws_completion_tokens == 50
assert ui._ws_context_ratio == pytest.approx(0.15)
def test_coord_on_status_persists_usage_event() -> None:
"""Pre-lift coord didn't persist usage_event rows — only WebUI did.
Lift extends usage tracking to coord so governance dashboards see
coordinator token consumption alongside interactive."""
storage = MagicMock()
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
with _patch_get_storage(storage):
ui.on_status(
{"prompt_tokens": 7, "completion_tokens": 3, "model": "gpt-x"},
context_window=200,
effort="low",
)
storage.record_usage_event.assert_called_once()
kwargs = storage.record_usage_event.call_args.kwargs
assert kwargs["ws_id"] == "coord-ws"
assert kwargs["user_id"] == "u1"
assert kwargs["model"] == "gpt-x"
assert kwargs["prompt_tokens"] == 7
assert kwargs["completion_tokens"] == 3
def test_coord_on_content_token_accumulates() -> None:
"""Pre-lift coord ``on_content_token`` only enqueued; lift turns it
into the same per-ws accumulator WebUI uses so the collector
broadcast can piggyback the joined turn content on the IDLE
state-change event."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui.on_content_token("Hello ")
ui.on_content_token("world")
assert ui._ws_turn_content == ["Hello ", "world"]
assert ui._ws_turn_content_size == len("Hello world")
def test_coord_on_content_token_caps_at_ceiling() -> None:
"""Same content cap interactive enforces — keeps a runaway turn from
ballooning the cluster broadcast event past listener queue size."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
chunk = "x" * 1024
rounds = (_MAX_TURN_CONTENT_CHARS // 1024) + 50
for _ in range(rounds):
ui.on_content_token(chunk)
# Cap is enforced at the size check; one over-cap chunk still
# gets in (per the original ``< _MAX``-not-``<=`` semantics) but
# nothing past that lands.
assert ui._ws_turn_content_size <= _MAX_TURN_CONTENT_CHARS + 1024
def test_coord_on_thinking_start_sets_activity() -> None:
"""Live activity tracking — coord's dashboard row now flips
``activity_state`` to ``"thinking"`` when the model starts."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui.on_thinking_start()
assert ui._ws_current_activity == "Thinking…"
assert ui._ws_activity_state == "thinking"
def test_coord_on_tool_result_clears_activity_and_increments_counters() -> None:
"""Lifted ``on_tool_result`` body increments ``_ws_tool_calls`` /
``_ws_turn_tool_calls`` and clears the activity. Pre-lift coord
just enqueued without touching counters."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui._ws_current_activity = "⚙ bash: ls -la"
ui._ws_activity_state = "tool"
ui.on_tool_result("call-1", "bash", "output")
assert ui._ws_tool_calls == {"bash": 1}
assert ui._ws_turn_tool_calls == 1
assert ui._ws_current_activity == ""
assert ui._ws_activity_state == ""
# ---------------------------------------------------------------------------
# Snapshot helper — drains turn content on IDLE/ERROR
# ---------------------------------------------------------------------------
def test_snapshot_idle_returns_content_and_clears_accumulator() -> None:
"""IDLE snapshot piggybacks the joined assistant content onto the
state-change broadcast (so the dashboard renders the turn without
a storage round-trip), then clears the accumulator for the next
turn."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui.on_content_token("Here's ")
ui.on_content_token("the result.")
payload = ui.snapshot_and_consume_state_payload("idle")
assert payload["content"] == "Here's the result."
assert ui._ws_turn_content == []
assert ui._ws_turn_content_size == 0
def test_snapshot_error_clears_accumulator_without_emitting_content() -> None:
"""ERROR clears the partial content (the turn's broken; nothing to
render) but the broadcast itself doesn't carry it — the state
transition is what matters."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui.on_content_token("partial...")
payload = ui.snapshot_and_consume_state_payload("error")
assert payload["content"] == ""
assert ui._ws_turn_content == []
def test_snapshot_thinking_does_not_touch_accumulator() -> None:
"""Mid-turn state transitions (running / thinking / attention)
don't drain the accumulator — only IDLE / ERROR are terminal."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui.on_content_token("partial mid-turn")
payload = ui.snapshot_and_consume_state_payload("thinking")
assert payload["content"] == ""
# Accumulator preserved.
assert ui._ws_turn_content == ["partial mid-turn"]
def test_snapshot_carries_token_and_activity_snapshot() -> None:
"""Snapshot reads tokens / context_ratio / activity under one lock
acquisition so concurrent on_status / on_thinking_start writes
don't tear the snapshot."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
with _patch_get_storage(MagicMock()):
ui.on_status(
{"prompt_tokens": 80, "completion_tokens": 20},
context_window=400,
effort="medium",
)
ui.on_thinking_start() # sets activity = "Thinking…"
payload = ui.snapshot_and_consume_state_payload("running")
assert payload["tokens"] == 100
assert payload["context_ratio"] == pytest.approx(0.25)
assert payload["activity"] == "Thinking…"
assert payload["activity_state"] == "thinking"
# ---------------------------------------------------------------------------
# Coord adapter — passes rich payload to collector
# ---------------------------------------------------------------------------
class _FakeCollectorRecorder:
"""Captures emit_console_ws_state calls so we can assert on the
rich kwargs the lifted coord_adapter.emit_state passes through."""
def __init__(self) -> None:
self.state_calls: list[dict[str, Any]] = []
self.activity_calls: list[dict[str, Any]] = []
def emit_console_ws_state(
self,
ws_id: str,
state: str,
*,
tokens: int = 0,
context_ratio: float = 0.0,
activity: str = "",
activity_state: str = "",
content: str = "",
) -> None:
self.state_calls.append(
{
"ws_id": ws_id,
"state": state,
"tokens": tokens,
"context_ratio": context_ratio,
"activity": activity,
"activity_state": activity_state,
"content": content,
}
)
def update_console_ws_activity(self, ws_id: str, *, activity: str, activity_state: str) -> None:
self.activity_calls.append(
{"ws_id": ws_id, "activity": activity, "activity_state": activity_state}
)
def emit_console_ws_created(self, *_a: Any, **_kw: Any) -> None:
pass
def emit_console_ws_closed(self, *_a: Any, **_kw: Any) -> None:
pass
def emit_console_ws_rename(self, *_a: Any, **_kw: Any) -> None:
pass
def ensure_console_pseudo_node(self) -> None:
pass
def _build_adapter_and_ws(ws_id: str = "coord-ws-1") -> tuple[Any, Any, _FakeCollectorRecorder]:
"""Construct a minimal adapter + Workstream + UI for emit_state tests.
Skips the full SessionManager wire-up the adapter's ``emit_state``
only reads ``ws.id`` and ``ws.ui``, so a real ``Workstream`` with
a populated ``ConsoleCoordinatorUI`` is enough.
"""
from turnstone.console.coordinator_adapter import CoordinatorAdapter
from turnstone.core.workstream import Workstream
recorder = _FakeCollectorRecorder()
adapter = CoordinatorAdapter(
collector=recorder, # type: ignore[arg-type]
ui_factory=lambda ws: ConsoleCoordinatorUI(ws_id=ws.id, user_id=ws.user_id),
session_factory=lambda ws: MagicMock(),
)
ws = Workstream(id=ws_id, user_id="u1", name="my-coord")
ws.ui = ConsoleCoordinatorUI(ws_id=ws_id, user_id="u1")
return adapter, ws, recorder
def test_coord_adapter_emit_state_passes_rich_payload_to_collector() -> None:
"""Pre-lift coord_adapter.emit_state called collector with state-only;
post-lift it reads the UI's per-ws snapshot under lock and passes
tokens / context_ratio / activity / content kwargs through."""
adapter, ws, recorder = _build_adapter_and_ws()
with _patch_get_storage(MagicMock()):
ws.ui.on_status(
{"prompt_tokens": 60, "completion_tokens": 40},
context_window=400,
effort="medium",
)
ws.ui.on_content_token("partial answer")
ws.ui.on_thinking_start()
adapter.emit_state(ws, WorkstreamState.RUNNING)
assert len(recorder.state_calls) == 1
call = recorder.state_calls[0]
assert call["ws_id"] == ws.id
assert call["state"] == "running"
assert call["tokens"] == 100
assert call["context_ratio"] == pytest.approx(0.25)
assert call["activity"] == "Thinking…"
assert call["activity_state"] == "thinking"
# Mid-turn (RUNNING) — content stays accumulated for the eventual IDLE drain.
assert call["content"] == ""
def test_coord_adapter_emit_state_idle_drains_content() -> None:
"""IDLE state-change drains the turn-content accumulator and
piggybacks the joined content on the broadcast same shape WebUI
uses on global_queue. Subsequent emit_state must see the
accumulator cleared."""
adapter, ws, recorder = _build_adapter_and_ws()
ws.ui.on_content_token("Here's ")
ws.ui.on_content_token("the result.")
adapter.emit_state(ws, WorkstreamState.IDLE)
assert len(recorder.state_calls) == 1
assert recorder.state_calls[0]["content"] == "Here's the result."
# Accumulator drained — next emit_state sees nothing carried over.
adapter.emit_state(ws, WorkstreamState.IDLE)
assert recorder.state_calls[1]["content"] == ""
def test_coord_adapter_emit_state_handles_missing_ui_defensively() -> None:
"""``ws.ui`` can be ``None`` mid-eviction; emit_state still
broadcasts the state-change with empty rich fields so the
dashboard's coord row still flips state instead of going stale."""
adapter, ws, recorder = _build_adapter_and_ws()
ws.ui = None # simulate teardown race
adapter.emit_state(ws, WorkstreamState.RUNNING)
assert len(recorder.state_calls) == 1
call = recorder.state_calls[0]
assert call["state"] == "running"
assert call["tokens"] == 0
assert call["content"] == ""
# ---------------------------------------------------------------------------
# Coord activity broadcast — UI fans out directly to the collector
# ---------------------------------------------------------------------------
def test_coord_ui_broadcast_activity_calls_collector() -> None:
"""Live activity transitions on coord (between state changes) reach
the cluster collector via the new ``update_console_ws_activity``
method. WebUI's analog goes via the global SSE queue; coord's
UI calls the collector directly since the console isn't a node."""
recorder = _FakeCollectorRecorder()
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ConsoleCoordinatorUI._collector = recorder # type: ignore[assignment]
try:
ui.on_thinking_start() # base impl calls _broadcast_activity
assert len(recorder.activity_calls) == 1
call = recorder.activity_calls[0]
assert call["ws_id"] == "coord-ws"
assert call["activity"] == "Thinking…"
assert call["activity_state"] == "thinking"
finally:
ConsoleCoordinatorUI._collector = None
def test_coord_ui_broadcast_activity_swallows_collector_failure() -> None:
"""A flaky collector must NOT block the worker thread — activity
fan-out is observational, the worker keeps running on collector
failure."""
recorder = MagicMock()
recorder.update_console_ws_activity.side_effect = RuntimeError("collector dead")
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ConsoleCoordinatorUI._collector = recorder
try:
ui.on_thinking_start() # must not raise
recorder.update_console_ws_activity.assert_called_once()
finally:
ConsoleCoordinatorUI._collector = None
def test_coord_ui_broadcast_activity_no_op_when_collector_unset() -> None:
"""Tests / tooling that don't wire a collector shouldn't crash."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ConsoleCoordinatorUI._collector = None
ui.on_thinking_start() # must not raise
def test_coord_ui_broadcast_activity_failure_does_not_strand_dedup() -> None:
"""Regression for the Copilot finding on PR #420: post-fix the
dedup state ``_last_broadcast_activity`` is updated **only after**
a successful collector call. If the collector raises mid-broadcast
on tick #1, tick #2 with the same activity tuple must still
attempt the broadcast (otherwise a transient collector failure
would strand the dashboard's coord row at the pre-failure
activity until the activity actually changes). Pre-fix the
dedup state was assigned inside the lock before the collector
call, so the failed broadcast still updated it and tick #2
silently no-op'd."""
recorder = MagicMock()
# First call fails (transient collector outage); second call succeeds.
recorder.update_console_ws_activity.side_effect = [
RuntimeError("collector dead"),
None,
]
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ConsoleCoordinatorUI._collector = recorder
try:
# Tick #1 — collector raises; dedup state must NOT update.
ui.on_thinking_start()
assert ui._last_broadcast_activity is None, (
"dedup state was updated despite a failed collector call — "
"next identical tick would be silently suppressed"
)
# Tick #2 — same activity tuple. Pre-fix this would no-op
# (because dedup state was already (Thinking…, thinking)).
# Post-fix it retries; collector succeeds; dedup state lands.
ui.on_thinking_start()
assert recorder.update_console_ws_activity.call_count == 2, (
"second tick was deduped despite the first call failing"
)
assert ui._last_broadcast_activity == ("Thinking…", "thinking")
finally:
ConsoleCoordinatorUI._collector = None
def test_coord_ui_broadcast_activity_dedup_skips_identical_after_success() -> None:
"""Happy-path dedup: after a successful broadcast, the next identical
tick is deduped the cluster collector lock is not re-acquired
for a no-op write. This is the perf optimization the dedup is
there for; the regression test above checks the failure-recovery
invariant doesn't break it."""
recorder = MagicMock()
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ConsoleCoordinatorUI._collector = recorder
try:
ui.on_thinking_start() # tick 1 — fires
ui.on_thinking_start() # tick 2 — same tuple, deduped
ui.on_thinking_start() # tick 3 — same tuple, deduped
assert recorder.update_console_ws_activity.call_count == 1
assert ui._last_broadcast_activity == ("Thinking…", "thinking")
finally:
ConsoleCoordinatorUI._collector = None
# ---------------------------------------------------------------------------
# Spawn metrics — coord wires its own hook
# ---------------------------------------------------------------------------
def test_coord_spawn_metrics_increments_messages_and_resets_tool_count() -> None:
"""Coord's ``_coord_spawn_metrics`` mirrors interactive's per-spawn
counter writes (sans the Prometheus call) so the rich ``ws_state``
broadcast renders the same per-turn shape."""
from turnstone.console.server import _coord_spawn_metrics
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui._ws_messages = 5
ui._ws_turn_tool_calls = 3
_coord_spawn_metrics(MagicMock(), ui)
assert ui._ws_messages == 6
assert ui._ws_turn_tool_calls == 0
def test_coord_spawn_metrics_tolerates_ui_without_counters() -> None:
"""A SessionUI subclass without the per-ws counters shouldn't trip
the hook defensive guard mirrors the interactive analog."""
from turnstone.console.server import _coord_spawn_metrics
class _StubUI:
pass
_coord_spawn_metrics(MagicMock(), _StubUI()) # must not raise
# ---------------------------------------------------------------------------
# Snapshot lock — single-acquisition guarantee
# ---------------------------------------------------------------------------
def test_snapshot_acquires_ws_lock_exactly_once() -> None:
"""Snapshot must read all four fields under a single lock acquisition
so concurrent on_status / on_thinking_start writes can't tear the
payload."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
acquire_count = 0
inner = ui._ws_lock
class _CountingLock:
def __enter__(self) -> None:
nonlocal acquire_count
acquire_count += 1
inner.acquire()
def __exit__(self, *a: Any) -> None:
inner.release()
def acquire(self, *a: Any, **kw: Any) -> bool:
return inner.acquire(*a, **kw)
def release(self) -> None:
inner.release()
ui._ws_lock = _CountingLock() # type: ignore[assignment]
ui.snapshot_and_consume_state_payload("idle")
assert acquire_count == 1, (
f"snapshot acquired _ws_lock {acquire_count} times; concurrent "
"writes could tear the rich payload"
)
# ---------------------------------------------------------------------------
# Concurrency — snapshot under load
# ---------------------------------------------------------------------------
def test_snapshot_under_concurrent_writes_does_not_crash() -> None:
"""Sanity stress: snapshot reads while on_status / on_thinking_start /
on_content_token write concurrently. Reader cycles through
``("running", "idle", "error")`` so the IDLE/ERROR drain branches
that mutate ``_ws_turn_content`` actually get exercised against
concurrent appends running-only would only hit the read-only
snapshot path. Each thread's exception (if any) is captured + raised
on join so a silent worker crash can't slip through as a bare
deadlock-check pass."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
writer_exc: list[Exception] = []
reader_exc: list[Exception] = []
def _writer() -> None:
try:
with _patch_get_storage(MagicMock()):
for i in range(50):
ui.on_status(
{"prompt_tokens": i, "completion_tokens": i},
context_window=1000,
effort="low",
)
ui.on_content_token(f"chunk-{i}")
ui.on_thinking_start()
except Exception as exc: # noqa: BLE001 — surface to main thread
writer_exc.append(exc)
def _reader() -> None:
try:
states = ("running", "idle", "error")
for i in range(50):
ui.snapshot_and_consume_state_payload(states[i % len(states)])
except Exception as exc: # noqa: BLE001 — surface to main thread
reader_exc.append(exc)
writer = threading.Thread(target=_writer)
reader = threading.Thread(target=_reader)
writer.start()
reader.start()
writer.join(timeout=5)
reader.join(timeout=5)
assert not writer.is_alive(), "writer thread deadlocked"
assert not reader.is_alive(), "reader thread deadlocked"
assert not writer_exc, f"writer raised: {writer_exc[0]!r}"
assert not reader_exc, f"reader raised: {reader_exc[0]!r}"
def test_coord_on_stream_end_clears_activity() -> None:
"""Lifted ``on_stream_end`` body clears ``_ws_current_activity``
and ``_ws_activity_state`` so the dashboard's coord row stops
showing the stale 'Thinking…' indicator after the stream
finishes. Pre-lift coord just enqueued ``stream_end`` without
touching activity this test pins the new clear path so a
future re-stub doesn't silently re-introduce a stuck activity
indicator."""
ui = ConsoleCoordinatorUI(ws_id="coord-ws", user_id="u1")
ui._ws_current_activity = "Thinking…"
ui._ws_activity_state = "thinking"
ui.on_stream_end()
assert ui._ws_current_activity == ""
assert ui._ws_activity_state == ""
# ---------------------------------------------------------------------------
# WebUI override semantics still preserved
# ---------------------------------------------------------------------------
def test_webui_on_status_still_records_prometheus_metrics() -> None:
"""The lift moves the per-ws writes to SessionUIBase but WebUI's
override must still fire ``_metrics.record_*`` (Prometheus on the
node /metrics endpoint). Regression guard against a future refactor
accidentally dropping the override."""
import queue
from turnstone.server import WebUI
WebUI._global_queue = queue.Queue()
try:
ui = WebUI(ws_id="ws-int", user_id="u1")
with patch("turnstone.server._metrics") as mock_metrics, _patch_get_storage(MagicMock()):
ui.on_status(
{"prompt_tokens": 10, "completion_tokens": 5},
context_window=200,
effort="low",
)
mock_metrics.record_tokens.assert_called_once_with(10, 5)
mock_metrics.record_cache_tokens.assert_called_once()
mock_metrics.record_context_ratio.assert_called_once()
finally:
WebUI._global_queue = None
def test_webui_on_tool_result_still_records_prometheus_tool_call() -> None:
"""Same as above for ``on_tool_result``."""
import queue
from turnstone.server import WebUI
WebUI._global_queue = queue.Queue()
try:
ui = WebUI(ws_id="ws-int", user_id="u1")
with patch("turnstone.server._metrics") as mock_metrics:
ui.on_tool_result("call-1", "bash", "output")
mock_metrics.record_tool_call.assert_called_once_with("bash")
# Per-ws counter writes happened too (inherited from base).
assert ui._ws_tool_calls == {"bash": 1}
assert ui._ws_turn_tool_calls == 1
finally:
WebUI._global_queue = None
+611
View File
@@ -0,0 +1,611 @@
"""Tests for the unified ``approve_tools`` body, viewed from the coord side.
The body itself is exercised by ``test_webui_auto_approve_visibility``;
this file pins down the coord-specific contracts that lifting the body
to ``SessionUIBase`` automatically enables:
- Tool-policy gating now applies to coord tool calls (was interactive-only).
- Heuristic verdicts persist on coord (was interactive-only).
- The activity tag fields populate on coord during pending approval.
- ``judge_pending`` is dynamic on the coord ``approve_request``
(was hardcoded ``False``).
- The auto-approve fall-through emits ``tool_info`` (was
``tools_auto_approved``).
- ``_record_judge_metric`` is a no-op on coord (no Prometheus on console).
"""
from __future__ import annotations
import threading
from typing import Any
from unittest.mock import MagicMock, patch
from turnstone.console.coordinator_ui import ConsoleCoordinatorUI
def _make_items(*specs: tuple[str, str], needs_approval: bool = True) -> list[dict[str, Any]]:
return [
{
"call_id": call_id,
"header": f"Tool: {func}",
"preview": "preview text",
"func_name": func,
"approval_label": func,
"needs_approval": needs_approval,
}
for call_id, func in specs
]
def _patch_storage(storage: Any):
return patch("turnstone.core.storage._registry.get_storage", return_value=storage)
def _patch_policies(verdicts: dict[str, str]):
return patch(
"turnstone.core.policy.evaluate_tool_policies_batch",
return_value=verdicts,
)
# ---------------------------------------------------------------------------
# Inheritance regression — the unification itself
# ---------------------------------------------------------------------------
def test_coord_inherits_approve_tools_from_base() -> None:
"""``ConsoleCoordinatorUI`` must NOT define its own ``approve_tools``;
the shared body lives on :class:`SessionUIBase`. A future drift
adding a coord-only override is exactly the kind of bug this
unification is meant to prevent, so guard it explicitly."""
assert "approve_tools" not in ConsoleCoordinatorUI.__dict__, (
"ConsoleCoordinatorUI shouldn't redefine approve_tools — "
"the shared body on SessionUIBase covers both kinds."
)
assert ConsoleCoordinatorUI.approve_tools.__qualname__ == "SessionUIBase.approve_tools"
# ---------------------------------------------------------------------------
# Tool-policy gating now applies to coord
# ---------------------------------------------------------------------------
def test_coord_tool_policy_deny_blocks_coord_tool() -> None:
"""Admin-defined ``deny`` policies now fire on coord tool calls.
Pre-lift this was interactive-only; an admin who wanted to block
e.g. ``delete_workstream`` on the coord couldn't."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
items = _make_items(("c1", "delete_workstream"))
storage = MagicMock()
with _patch_storage(storage), _patch_policies({"delete_workstream": "deny"}):
approved, err = ui.approve_tools(items)
assert approved is False
assert err == "Blocked by tool policy"
assert items[0].get("denied") is True
def test_coord_tool_policy_allow_tags_with_policy_source() -> None:
"""Admin ``allow`` rule auto-approves the item with
``AutoApproveReason.POLICY``. This was a no-op on coord pre-lift."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
items = _make_items(("c1", "spawn_workstream"))
storage = MagicMock()
with _patch_storage(storage), _patch_policies({"spawn_workstream": "allow"}):
approved, _err = ui.approve_tools(items)
assert approved is True
snapshot = ui.serialize_recent_auto_approvals()
assert len(snapshot) == 1
assert snapshot[0]["func_name"] == "spawn_workstream"
assert snapshot[0]["auto_approve_reason"] == "policy"
def test_coord_tool_policy_mixed_allow_deny_records_allowed_sibling() -> None:
"""Same ``mixed-policy`` audit-leak fix that
``test_webui_auto_approve_visibility`` validates for interactive,
now auto-applies to coord via the lifted body."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
items = _make_items(("c1", "delete_workstream"), ("c2", "list_workstreams"))
storage = MagicMock()
with (
_patch_storage(storage),
_patch_policies({"delete_workstream": "deny", "list_workstreams": "allow"}),
):
approved, _err = ui.approve_tools(items)
assert approved is False
snapshot = ui.serialize_recent_auto_approvals()
assert len(snapshot) == 1
assert snapshot[0]["func_name"] == "list_workstreams"
assert snapshot[0]["auto_approve_reason"] == "policy"
# ---------------------------------------------------------------------------
# Heuristic-verdict persistence + metric hook
# ---------------------------------------------------------------------------
def test_coord_heuristic_verdict_persists_to_storage() -> None:
"""Heuristic verdicts attached to items now flow through to
``storage.create_intent_verdicts_bulk`` on coord. Pre-lift coord
silently dropped them; only LLM-tier verdicts (from the daemon
judge thread via ``on_intent_verdict``) reached storage. Post
perf-2 the path uses bulk INSERT so a fan-out turn pays one commit
instead of N."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
hv = {
"verdict_id": "v1",
"call_id": "c1",
"func_name": "spawn_workstream",
"tier": "heuristic",
"risk_level": "high",
"confidence": 0.75,
"recommendation": "review",
"reasoning": "spawning child with bash skill",
"evidence": ["bash"],
"latency_ms": 12,
}
items = _make_items(("c1", "spawn_workstream"))
items[0]["_heuristic_verdict"] = hv
storage = MagicMock()
timer = threading.Timer(0.05, lambda: ui.resolve_approval(False))
timer.start()
try:
with _patch_storage(storage):
ui.approve_tools(items)
finally:
timer.cancel()
storage.create_intent_verdicts_bulk.assert_called_once()
rows = storage.create_intent_verdicts_bulk.call_args.args[0]
assert len(rows) == 1
assert rows[0]["verdict_id"] == "v1"
assert rows[0]["tier"] == "heuristic"
assert rows[0]["ws_id"] == "coord-1"
def test_coord_record_judge_metric_fires_console_metrics() -> None:
"""``_record_judge_metric`` increments the console's
``ConsoleMetrics`` judge counter when the class attribute is wired,
so coord verdicts surface on the console's /metrics endpoint
alongside the per-node series."""
from turnstone.console.metrics import ConsoleMetrics
cm = ConsoleMetrics()
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
try:
ConsoleCoordinatorUI._console_metrics = cm
ui._record_judge_metric({"tier": "heuristic", "risk_level": "high", "latency_ms": 12})
finally:
ConsoleCoordinatorUI._console_metrics = None
text = cm.generate_text()
assert 'turnstone_judge_verdicts_total{tier="heuristic",risk_level="high"} 1' in text
def test_coord_record_judge_metric_safe_when_unwired() -> None:
"""No /metrics instance set → silent no-op. Test fixtures that
don't spin up a full console app must not crash on judge
verdicts during the shared ``approve_tools`` body."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
# Sanity: class attribute is None at module import time outside
# the lifespan — exactly the test-fixture state.
assert ConsoleCoordinatorUI._console_metrics is None
# Should not raise.
ui._record_judge_metric({"tier": "heuristic", "risk_level": "low"})
def test_coord_on_intent_verdict_fires_metric_for_llm_tier() -> None:
"""Async LLM verdicts from the daemon judge thread land at
``on_intent_verdict``. Coord overrides it to fire the same
``record_judge_verdict`` call WebUI does different tier label,
same cluster-wide histogram."""
from turnstone.console.metrics import ConsoleMetrics
cm = ConsoleMetrics()
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
try:
ConsoleCoordinatorUI._console_metrics = cm
with _patch_storage(MagicMock()):
ui.on_intent_verdict(
{
"verdict_id": "v1",
"call_id": "c1",
"tier": "llm",
"risk_level": "medium",
"latency_ms": 250,
}
)
finally:
ConsoleCoordinatorUI._console_metrics = None
text = cm.generate_text()
assert 'turnstone_judge_verdicts_total{tier="llm",risk_level="medium"} 1' in text
# ---------------------------------------------------------------------------
# Activity tagging during pending approval
# ---------------------------------------------------------------------------
def test_coord_pending_approval_sets_activity_tag() -> None:
"""The shared body tags ``_ws_current_activity`` /
``_ws_activity_state`` so the cluster collector's coord-row
snapshot reflects the approval wait. Pre-lift coord left these
fields empty during pending approval."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
items = _make_items(("c1", "delete_workstream"))
captured: dict[str, str] = {}
def _capture_activity() -> None:
captured["activity"] = ui._ws_current_activity
captured["state"] = ui._ws_activity_state
ui.resolve_approval(False)
timer = threading.Timer(0.05, _capture_activity)
timer.start()
try:
with _patch_storage(MagicMock()):
ui.approve_tools(items)
finally:
timer.cancel()
assert "Awaiting approval" in captured["activity"]
assert "delete_workstream" in captured["activity"]
assert captured["state"] == "approval"
def test_coord_auto_approve_sets_tool_activity_tag() -> None:
"""Blanket auto-approve flips activity to the ``⚙ {tool}: {preview}``
shape WebUI has used; coord row now mirrors it."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
ui.auto_approve = True
items = _make_items(("c1", "spawn_workstream"))
with _patch_storage(MagicMock()):
approved, _err = ui.approve_tools(items)
assert approved is True
assert "spawn_workstream" in ui._ws_current_activity
assert ui._ws_activity_state == "tool"
# ---------------------------------------------------------------------------
# judge_pending flag + event-name parity
# ---------------------------------------------------------------------------
def test_coord_judge_pending_flag_dynamic_when_heuristic_present() -> None:
"""Pre-lift coord hardcoded ``judge_pending=False`` on every
``approve_request``; the unified body computes the bool from the
items, matching WebUI."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
items = _make_items(("c1", "spawn_workstream"))
items[0]["_heuristic_verdict"] = {"verdict_id": "v1", "tier": "heuristic"}
captured_events: list[dict[str, Any]] = []
ui._enqueue = captured_events.append # type: ignore[method-assign]
timer = threading.Timer(0.05, lambda: ui.resolve_approval(False))
timer.start()
try:
with _patch_storage(MagicMock()):
ui.approve_tools(items)
finally:
timer.cancel()
approve_requests = [e for e in captured_events if e.get("type") == "approve_request"]
assert len(approve_requests) == 1
assert approve_requests[0]["judge_pending"] is True
def test_coord_blanket_auto_approve_emits_tool_info() -> None:
"""Event-name parity: the auto-approve fall-through emits
``tool_info`` for both kinds. Pre-lift coord emitted
``tools_auto_approved`` the rename happens implicitly via
inheritance."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
ui.auto_approve = True
items = _make_items(("c1", "spawn_workstream"))
captured_events: list[dict[str, Any]] = []
ui._enqueue = captured_events.append # type: ignore[method-assign]
with _patch_storage(MagicMock()):
ui.approve_tools(items)
types = [e.get("type") for e in captured_events]
assert "tool_info" in types
assert "tools_auto_approved" not in types
def test_coord_judge_pending_false_when_no_heuristic_verdict() -> None:
"""Counterpart to ``test_coord_judge_pending_flag_dynamic_when_heuristic_present``:
items with no ``_heuristic_verdict`` produce ``approve_request`` with
``judge_pending=False``. Without this case pinned, a regression that
hardcodes ``judge_pending=True`` (the inverse of the pre-lift coord
bug) would slip through unnoticed."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
items = _make_items(("c1", "spawn_workstream"))
# Deliberately no _heuristic_verdict on any item.
captured_events: list[dict[str, Any]] = []
ui._enqueue = captured_events.append # type: ignore[method-assign]
timer = threading.Timer(0.05, lambda: ui.resolve_approval(False))
timer.start()
try:
with _patch_storage(MagicMock()):
ui.approve_tools(items)
finally:
timer.cancel()
approve_requests = [e for e in captured_events if e.get("type") == "approve_request"]
assert len(approve_requests) == 1
assert approve_requests[0]["judge_pending"] is False
# ---------------------------------------------------------------------------
# Per-tool auto-approve via auto_approve_tools (set membership)
# ---------------------------------------------------------------------------
def test_coord_per_tool_auto_approve_tags_with_source() -> None:
"""When a coord tool name lands in ``auto_approve_tools`` (e.g. via a
skill template's ``allowed_tools``), the lifted body short-circuits
the prompt and tags the item with ``AutoApproveReason.AUTO_APPROVE_TOOLS``
(or the per-tool source from ``_auto_approve_tools_source``).
Mirrors the WebUI test ``test_auto_approve_tools_skill_source_renders_as_skill``
on the coord side so the unified body gains parity coverage."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
ui.auto_approve_tools = {"spawn_workstream"}
ui._auto_approve_tools_source = {"spawn_workstream": "skill"}
items = _make_items(("c1", "spawn_workstream"))
storage = MagicMock()
with _patch_storage(storage):
approved, _err = ui.approve_tools(items)
assert approved is True
snapshot = ui.serialize_recent_auto_approvals()
assert len(snapshot) == 1
assert snapshot[0]["func_name"] == "spawn_workstream"
assert snapshot[0]["auto_approve_reason"] == "skill"
# ---------------------------------------------------------------------------
# __budget_override__ carve-out — sec-2 hardening
# ---------------------------------------------------------------------------
def test_coord_budget_override_prompts_even_under_blanket_auto_approve() -> None:
"""The carve-out promises ``__budget_override__`` always prompts the
operator. Pin that behavior on the coord side so a future regression
of the post-filter / pre-filter check (sec-2) gets caught.
``__budget_override__`` is interactive-only today (coord workstreams
don't have token budgets), but the synthetic item can be threaded
through ``approve_tools`` directly the same way ``ChatSession.send``
does on the interactive side. The carve-out fires uniformly across
both kinds."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
ui.auto_approve = True # blanket flag — should NOT bypass the carve-out
items = [
{
"call_id": "c1",
"header": "Token budget exhausted",
"preview": "Token budget (200,000) exhausted. Approve to continue.",
"func_name": "__budget_override__",
"approval_label": "__budget_override__",
"needs_approval": True,
}
]
captured_events: list[dict[str, Any]] = []
ui._enqueue = captured_events.append # type: ignore[method-assign]
timer = threading.Timer(0.05, lambda: ui.resolve_approval(True))
timer.start()
try:
with _patch_storage(MagicMock()):
approved, _err = ui.approve_tools(items)
finally:
timer.cancel()
assert approved is True
# The carve-out forces the prompt path, NOT the auto-approve fall-through.
types = [e.get("type") for e in captured_events]
assert "approve_request" in types, (
"Budget override must produce an approve_request even under blanket auto_approve"
)
assert "tool_info" not in types, (
"Auto-approve fall-through must not fire when a budget override is present"
)
def test_coord_budget_override_survives_wildcard_allow_policy() -> None:
"""A wildcard ``*: allow`` policy must not strip ``__budget_override__``
from the gate. Pre-sec-2, the policy block could mark the item
``needs_approval=False`` and remove it from ``pending``, after which
the carve-out (which read ``pending``) would see no override and
blanket auto-approve would silently fire. Post-fix the carve-out
reads from the pre-filter ``items`` list AND the policy block skips
matching the synthetic name entirely."""
ui = ConsoleCoordinatorUI(ws_id="coord-1", user_id="u1")
ui.auto_approve = True
items = [
{
"call_id": "c1",
"header": "Token budget exhausted",
"preview": "Token budget exhausted. Approve to continue.",
"func_name": "__budget_override__",
"approval_label": "__budget_override__",
"needs_approval": True,
}
]
captured_events: list[dict[str, Any]] = []
ui._enqueue = captured_events.append # type: ignore[method-assign]
timer = threading.Timer(0.05, lambda: ui.resolve_approval(True))
timer.start()
try:
with _patch_storage(MagicMock()), _patch_policies({"__budget_override__": "allow"}):
approved, _err = ui.approve_tools(items)
finally:
timer.cancel()
assert approved is True
types = [e.get("type") for e in captured_events]
assert "approve_request" in types, "Wildcard allow must not strip the budget-override prompt"
# ---------------------------------------------------------------------------
# Cluster-bus broadcast hooks — _broadcast_intent_verdict / _approval_resolved
# ---------------------------------------------------------------------------
class TestBroadcastIntentVerdict:
"""``ConsoleCoordinatorUI._broadcast_intent_verdict`` overrides the
no-op base hook to push the verdict onto the cluster bus via
``ClusterCollector.emit_console_ws_intent_verdict``. The far more
common path is the per-node ``WebUI`` override (covered in
test_webui_content.py); this lights up the rare coord-self path
(a coord that runs its own LLM judge).
"""
def test_calls_collector_emit_with_ws_id_and_verdict(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
verdict = {
"call_id": "c1",
"risk_level": "high",
"confidence": 0.91,
}
ui._broadcast_intent_verdict(verdict)
collector.emit_console_ws_intent_verdict.assert_called_once_with(
"coord-a",
verdict,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_no_op_when_collector_unset(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
ConsoleCoordinatorUI._collector = None
# Doesn't raise.
ui._broadcast_intent_verdict({"call_id": "c1"})
def test_collector_exception_swallowed(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
collector.emit_console_ws_intent_verdict.side_effect = RuntimeError("boom")
ConsoleCoordinatorUI._collector = collector
try:
# Doesn't raise — collector failures are observational only.
ui._broadcast_intent_verdict({"call_id": "c1"})
finally:
ConsoleCoordinatorUI._collector = None
class TestBroadcastApprovalResolved:
"""``ConsoleCoordinatorUI._broadcast_approval_resolved`` overrides
the base hook to push the resolution onto the cluster bus via
``ClusterCollector.emit_console_ws_approval_resolved``."""
def test_calls_collector_with_decision_fields(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
ui._broadcast_approval_resolved(True, "lgtm", always=True)
collector.emit_console_ws_approval_resolved.assert_called_once_with(
"coord-a",
approved=True,
feedback="lgtm",
always=True,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_normalises_none_feedback_to_empty_string(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
ui._broadcast_approval_resolved(False, None)
collector.emit_console_ws_approval_resolved.assert_called_once_with(
"coord-a",
approved=False,
feedback="",
always=False,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_no_op_when_collector_unset(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
ConsoleCoordinatorUI._collector = None
# Doesn't raise.
ui._broadcast_approval_resolved(True, None)
def test_collector_exception_swallowed(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
collector.emit_console_ws_approval_resolved.side_effect = RuntimeError("boom")
ConsoleCoordinatorUI._collector = collector
try:
# Doesn't raise.
ui._broadcast_approval_resolved(True, "ok")
finally:
ConsoleCoordinatorUI._collector = None
class TestBroadcastApproveRequest:
"""Coord-side override for the approve_request push. Same rationale
as the WebUI override the coord-self path is rare today, but
parity keeps the override symmetric with the rest of the broadcast
family."""
def test_calls_collector_emit_with_ws_id_and_detail(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
detail = {
"type": "approve_request",
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": True,
}
ui._broadcast_approve_request(detail)
collector.emit_console_ws_approve_request.assert_called_once_with(
"coord-a",
detail,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_no_op_when_collector_unset(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
ConsoleCoordinatorUI._collector = None
# Doesn't raise.
ui._broadcast_approve_request({"items": []})
def test_collector_exception_swallowed(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
collector.emit_console_ws_approve_request.side_effect = RuntimeError("boom")
ConsoleCoordinatorUI._collector = collector
try:
# Doesn't raise.
ui._broadcast_approve_request({"items": []})
finally:
ConsoleCoordinatorUI._collector = None
+718
View File
@@ -0,0 +1,718 @@
"""Tests for CoordinatorAdapter.
Mirrors test_interactive_adapter.py: focuses on the transport contract
(what gets sent to the ClusterCollector) and cleanup_ui behavior
(unblock listener queues, cancel session). The SessionManager-level
tests in test_session_manager.py cover the lifecycle path.
"""
from __future__ import annotations
import queue
import threading
from typing import Any
from unittest.mock import MagicMock
from turnstone.console.coordinator_adapter import CoordinatorAdapter
from turnstone.core.workstream import Workstream, WorkstreamKind, WorkstreamState
class _StubCoordUI:
"""Stub matching the subset of ConsoleCoordinatorUI the adapter touches."""
def __init__(self) -> None:
self._approval_event = threading.Event()
self._approval_result: tuple[bool, str | None] = (True, "initial")
self._plan_event = threading.Event()
self._plan_result: str = "accept"
self._fg_event = threading.Event()
self._listeners_lock = threading.Lock()
self._listeners: list[queue.Queue[dict[str, Any]]] = []
class _StubSession:
def __init__(self) -> None:
self.cancelled = False
self.closed = False
def cancel(self) -> None:
self.cancelled = True
def close(self) -> None:
self.closed = True
def _make_adapter(
collector: Any = None,
*,
ui_factory: Any = None,
session_factory: Any = None,
) -> tuple[CoordinatorAdapter, MagicMock]:
collector = collector or MagicMock()
adapter = CoordinatorAdapter(
collector=collector,
ui_factory=ui_factory or (lambda ws: _StubCoordUI()),
session_factory=session_factory or (lambda *a, **kw: _StubSession()),
)
return adapter, collector
def _make_ws(**overrides: Any) -> Workstream:
ws = Workstream(id="coord-1", name="my-coord")
ws.kind = WorkstreamKind.COORDINATOR
ws.user_id = "u1"
ws.ui = _StubCoordUI()
ws.session = _StubSession()
for k, v in overrides.items():
setattr(ws, k, v)
return ws
# ---------------------------------------------------------------------------
# Transport — emit_created / emit_state / emit_closed
# ---------------------------------------------------------------------------
def test_emit_created_calls_collector_with_coord_fields() -> None:
adapter, collector = _make_adapter()
ws = _make_ws()
adapter.emit_created(ws)
collector.emit_console_ws_created.assert_called_once_with(
"coord-1",
name="my-coord",
user_id="u1",
kind=WorkstreamKind.COORDINATOR.value,
state=WorkstreamState.IDLE.value,
parent_ws_id=None,
)
def test_emit_state_calls_collector_state() -> None:
"""Post-rich-payload, emit_state passes tokens / context_ratio /
activity / activity_state / content kwargs read from ws.ui's
snapshot. Default values (zeros / empty strings) when the UI
hasn't recorded any per-ws metrics yet."""
adapter, collector = _make_adapter()
ws = _make_ws()
adapter.emit_state(ws, WorkstreamState.RUNNING)
collector.emit_console_ws_state.assert_called_once_with(
"coord-1",
WorkstreamState.RUNNING.value,
tokens=0,
context_ratio=0.0,
activity="",
activity_state="",
content="",
)
def test_emit_closed_calls_collector_closed() -> None:
adapter, collector = _make_adapter()
adapter.emit_closed("coord-1")
collector.emit_console_ws_closed.assert_called_once_with("coord-1")
def test_emit_closed_swallows_reason_kwarg() -> None:
"""The console collector doesn't propagate a 'reason' — the console
frontend's evicted special-case only fires for real-node
workstreams. Protocol compatibility only."""
adapter, collector = _make_adapter()
adapter.emit_closed("coord-1", reason="evicted")
collector.emit_console_ws_closed.assert_called_once_with("coord-1")
def test_emit_tolerates_collector_exception() -> None:
collector = MagicMock()
collector.emit_console_ws_created.side_effect = RuntimeError("collector dead")
collector.emit_console_ws_state.side_effect = RuntimeError("collector dead")
collector.emit_console_ws_closed.side_effect = RuntimeError("collector dead")
adapter, _ = _make_adapter(collector=collector)
ws = _make_ws()
# All three must swallow — the session lifecycle must not break
# because the collector had a transient failure.
adapter.emit_created(ws)
adapter.emit_state(ws, WorkstreamState.RUNNING)
adapter.emit_closed("coord-1")
# ---------------------------------------------------------------------------
# cleanup_ui
# ---------------------------------------------------------------------------
def test_cleanup_ui_unblocks_events_and_broadcasts_to_listeners() -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
ws.ui._approval_event.clear() # type: ignore[attr-defined]
ws.ui._plan_event.clear() # type: ignore[attr-defined]
ws.ui._fg_event.clear() # type: ignore[attr-defined]
lq: queue.Queue[dict[str, Any]] = queue.Queue(maxsize=5)
ws.ui._listeners.append(lq) # type: ignore[attr-defined]
adapter.cleanup_ui(ws)
assert ws.ui._approval_event.is_set() # type: ignore[attr-defined]
assert ws.ui._plan_event.is_set() # type: ignore[attr-defined]
assert ws.ui._fg_event.is_set() # type: ignore[attr-defined]
assert ws.ui._approval_result == (False, None) # type: ignore[attr-defined]
assert ws.ui._plan_result == "reject" # type: ignore[attr-defined]
assert lq.get_nowait() == {"type": "ws_closed"}
assert ws.ui._listeners == [] # type: ignore[attr-defined]
assert ws.session.cancelled is True # type: ignore[attr-defined]
assert ws.session.closed is True # type: ignore[attr-defined]
def test_cleanup_ui_listener_full_queue_evicts_head() -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
lq: queue.Queue[dict[str, Any]] = queue.Queue(maxsize=1)
lq.put_nowait({"type": "stale"})
ws.ui._listeners.append(lq) # type: ignore[attr-defined]
adapter.cleanup_ui(ws)
assert lq.get_nowait() == {"type": "ws_closed"}
def test_cleanup_ui_tolerates_missing_session_and_ui() -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
ws.session = None
ws.ui = None
adapter.cleanup_ui(ws) # no crash
# ---------------------------------------------------------------------------
# Construction passthrough
# ---------------------------------------------------------------------------
def test_build_session_forwards_skill_model_kind_parent() -> None:
captured: dict[str, Any] = {}
def _sf(ui: Any, model: str | None, ws_id: str, **kwargs: Any) -> Any:
captured["ui"] = ui
captured["model"] = model
captured["ws_id"] = ws_id
captured.update(kwargs)
return _StubSession()
adapter, _ = _make_adapter(session_factory=_sf)
ws = _make_ws()
ws.parent_ws_id = None
adapter.build_session(ws, skill="coordinator", model="gpt-5")
assert captured["ui"] is ws.ui
assert captured["model"] == "gpt-5"
assert captured["skill"] == "coordinator"
assert captured["kind"] == WorkstreamKind.COORDINATOR
assert captured["parent_ws_id"] is None
# client_type intentionally NOT forwarded — coord session_factory
# doesn't accept it (fixed as 'console').
assert "client_type" not in captured
def test_build_ui_delegates_to_ui_factory() -> None:
captured_ws: list[Workstream] = []
def _ui_factory(ws: Workstream) -> Any:
captured_ws.append(ws)
return _StubCoordUI()
adapter, _ = _make_adapter(ui_factory=_ui_factory)
ws = _make_ws()
result = adapter.build_ui(ws)
assert captured_ws == [ws]
assert isinstance(result, _StubCoordUI)
# ---------------------------------------------------------------------------
# Worker dispatch — _spawn_worker / send
# ---------------------------------------------------------------------------
class _SendSession:
"""ChatSession stub with send / queue_message accounting."""
def __init__(
self,
*,
queue_full: bool = False,
send_gate: threading.Event | None = None,
) -> None:
self.send_calls: list[str] = []
self.queue_calls: list[str] = []
self._queue_full = queue_full
# When set, ``send`` blocks on this event — lets the test pin a
# worker inside session.send while a second thread races through
# _spawn_worker, proving the lock gate (not Thread.is_alive) is
# what serialises them.
self._send_gate = send_gate
self._send_lock = threading.Lock()
self.cancelled = False
self.closed = False
def send(
self,
message: str,
attachments: Any = None,
send_id: str | None = None,
) -> None:
if self._send_gate is not None:
self._send_gate.wait(timeout=2.0)
with self._send_lock:
self.send_calls.append(message)
def queue_message(
self,
message: str,
attachment_ids: Any = None,
queue_msg_id: str | None = None,
) -> None:
if self._queue_full:
raise queue.Full
self.queue_calls.append(message)
def cancel(self) -> None:
self.cancelled = True
def close(self) -> None:
self.closed = True
class _StubManager:
"""Minimal SessionManager stub exposing ``get`` for adapter.send."""
def __init__(self, ws: Workstream | None = None) -> None:
self._ws = ws
def get(self, ws_id: str) -> Workstream | None:
if self._ws is not None and self._ws.id == ws_id:
return self._ws
return None
class TestCoordinatorAdapterWorkerDispatch:
def test_spawn_worker_reuses_when_worker_running(self) -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
session = _SendSession()
ws.session = session # type: ignore[assignment]
ws._worker_running = True # pre-existing worker
adapter.attach(_StubManager(ws)) # type: ignore[arg-type]
assert adapter.send(ws.id, "hello") is True
assert session.queue_calls == ["hello"]
assert session.send_calls == []
# worker_thread not replaced
assert ws.worker_thread is None
def test_spawn_worker_returns_false_on_queue_full(self) -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
session = _SendSession(queue_full=True)
ws.session = session # type: ignore[assignment]
ws._worker_running = True
adapter.attach(_StubManager(ws)) # type: ignore[arg-type]
assert adapter.send(ws.id, "hello") is False
assert session.send_calls == []
def test_spawn_worker_concurrent_calls_produce_one_worker(self) -> None:
"""Bug-1 reproducer: two simultaneous send() calls under ws._lock
must land as exactly one ChatSession.send and one queued message,
not two parallel workers on the same ChatSession."""
adapter, _ = _make_adapter()
ws = _make_ws()
send_gate = threading.Event()
session = _SendSession(send_gate=send_gate)
ws.session = session # type: ignore[assignment]
adapter.attach(_StubManager(ws)) # type: ignore[arg-type]
results: list[bool] = []
start_barrier = threading.Barrier(2)
results_lock = threading.Lock()
def _caller(msg: str) -> None:
start_barrier.wait(timeout=1.0)
r = adapter.send(ws.id, msg)
with results_lock:
results.append(r)
t1 = threading.Thread(target=_caller, args=("first",))
t2 = threading.Thread(target=_caller, args=("second",))
t1.start()
t2.start()
# Both callers return quickly: the winner spawns the worker
# (returns True immediately) and the loser queues (returns True).
t1.join(timeout=3.0)
t2.join(timeout=3.0)
assert not t1.is_alive() and not t2.is_alive()
# At this point session.send is still blocked on send_gate —
# the second caller MUST have taken the queue path.
assert len(session.queue_calls) == 1
# Release the worker and let it finish.
send_gate.set()
if ws.worker_thread is not None:
ws.worker_thread.join(timeout=3.0)
assert results == [True, True]
assert len(session.send_calls) == 1
assert set(session.send_calls + session.queue_calls) == {"first", "second"}
assert ws._worker_running is False
def test_worker_finally_clears_running_flag(self) -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
session = _SendSession()
ws.session = session # type: ignore[assignment]
adapter.attach(_StubManager(ws)) # type: ignore[arg-type]
assert adapter.send(ws.id, "hello") is True
assert ws.worker_thread is not None
ws.worker_thread.join(timeout=2.0)
assert ws._worker_running is False
assert session.send_calls == ["hello"]
# ---------------------------------------------------------------------------
# Children registry
# ---------------------------------------------------------------------------
class TestCoordinatorAdapterChildrenRegistry:
"""Adapter-level integration with :class:`ChildrenRegistry`.
Pure-registry invariants (forward/reverse consistency, idempotent
merge, locking) live in ``test_children_registry.py``. These
tests cover the adapter's wiring: that ``emit_*`` paths drive the
registry correctly and that the snapshot-priming bridge between
a collector snapshot and the registry preserves merge semantics.
"""
def test_emit_created_installs_parent(self) -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
adapter.emit_created(ws)
assert adapter._registry.children_of(ws.id) == []
assert adapter._registry.ui_for(ws.id) is ws.ui
def test_emit_rehydrated_calls_rebuild(self) -> None:
adapter, _ = _make_adapter()
calls: list[str] = []
# Monkeypatch the rebuild hook to count invocations without
# requiring a real storage backend.
adapter._rebuild_children_registry = calls.append # type: ignore[method-assign, assignment]
ws = _make_ws()
adapter.emit_created(ws)
assert calls == []
adapter.emit_rehydrated(ws)
assert calls == [ws.id]
def test_emit_closed_uninstalls_parent_and_clears_children(self) -> None:
adapter, _ = _make_adapter()
adapter._registry.install("coord-a", object())
adapter._registry.install("coord-b", object())
adapter._registry.merge_children("coord-a", ["child-a1", "child-a2"])
adapter._registry.merge_children("coord-b", ["child-b1"])
adapter.emit_closed("coord-a")
assert adapter._registry.ui_for("coord-a") is None
assert adapter._registry.children_of("coord-a") == []
assert adapter._registry.parent_for("child-a1") is None
assert adapter._registry.parent_for("child-a2") is None
# coord-b untouched
assert adapter._registry.parent_for("child-b1") == "coord-b"
assert adapter._registry.ui_for("coord-b") is not None
def test_prime_children_from_snapshot_merges_without_overwriting(self) -> None:
# Snapshot priming now lives on ClusterChildSource (production
# path). The adapter no longer carries its own duplicate copy.
from turnstone.core.child_source import ClusterChildSource
adapter, _ = _make_adapter()
adapter._registry.merge_children("coord-a", ["child-a1"])
source = ClusterChildSource(
collector=MagicMock(),
registry=adapter._registry,
parents_provider=lambda: ["coord-a"],
)
snapshot = {
"nodes": [
{
"workstreams": [
{"id": "child-a2", "parent_ws_id": "coord-a"},
# Unknown parent — skipped
{"id": "child-x", "parent_ws_id": "coord-unknown"},
# Missing fields — skipped
{"id": "", "parent_ws_id": "coord-a"},
],
},
],
}
source._prime_from_snapshot(snapshot)
assert set(adapter._registry.children_of("coord-a")) == {
"child-a1",
"child-a2",
}
assert adapter._registry.parent_for("child-a2") == "coord-a"
assert adapter._registry.parent_for("child-x") is None
# ---------------------------------------------------------------------------
# Dispatch — _dispatch_child_event
# ---------------------------------------------------------------------------
class _UIRecorder:
"""UI stub capturing _enqueue payloads for dispatch assertions."""
def __init__(self) -> None:
self.enqueued: list[dict[str, Any]] = []
def _enqueue(self, payload: dict[str, Any]) -> None:
self.enqueued.append(payload)
class TestCoordinatorAdapterDispatchChildEvent:
def _setup(
self, coord_id: str = "coord-a"
) -> tuple[CoordinatorAdapter, _UIRecorder, Workstream]:
adapter, _ = _make_adapter()
coord_ws = _make_ws()
coord_ws.id = coord_id
recorder = _UIRecorder()
coord_ws.ui = recorder # type: ignore[assignment]
adapter._registry.install(coord_id, recorder)
adapter.attach(_StubManager(coord_ws)) # type: ignore[arg-type]
return adapter, recorder, coord_ws
def test_dispatch_unknown_parent_drops_event(self) -> None:
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{"type": "ws_created", "ws_id": "orphan", "parent_ws_id": "coord-unknown"}
)
adapter._dispatch_child_event({"type": "cluster_state", "ws_id": "orphan"})
adapter._dispatch_child_event({"type": "ws_closed", "ws_id": "orphan"})
assert recorder.enqueued == []
def test_dispatch_ws_created_routes_to_parent_coord_ui(self) -> None:
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{
"type": "ws_created",
"ws_id": "child-a1",
"parent_ws_id": "coord-a",
"name": "kid",
"node_id": "node-1",
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_created"
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
# Reverse index updated for subsequent cluster_state events.
assert adapter._registry.parent_for("child-a1") == "coord-a"
def test_dispatch_cluster_state_routes_via_reverse_index(self) -> None:
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "cluster_state",
"ws_id": "child-a1",
"state": "running",
"tokens": 42,
"node_id": "node-1",
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_state"
assert payload["state"] == "running"
assert payload["tokens"] == 42
def test_dispatch_ws_closed_routes_to_parent_coord(self) -> None:
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{"type": "ws_closed", "ws_id": "child-a1", "reason": "evicted"}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_closed"
assert payload["reason"] == "evicted"
assert payload["parent_ws_id"] == "coord-a"
def test_dispatch_adds_ws_id_in_place(self) -> None:
"""perf-6: _enqueue_on_ui mutates the payload dict in place with
the coord's ws_id so the browser can discriminate child events."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "cluster_state",
"ws_id": "child-a1",
"state": "running",
}
)
assert recorder.enqueued[0]["ws_id"] == "coord-a"
def test_dispatch_cluster_state_does_not_carry_pending_approval_detail(
self,
) -> None:
"""Stage 3 cleanup — the ``pending_approval_detail`` piggyback
on ``cluster_state`` is gone. Approval items now arrive via
bulk fetch (triggered by ``activity_state="approval"`` in the
browser); verdicts via ``child_ws_intent_verdict``; resolution
via ``child_ws_approval_resolved``. The state event carries
only state + activity_state no detail field."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "cluster_state",
"ws_id": "child-a1",
"state": "running",
"activity_state": "approval",
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_state"
assert payload["activity_state"] == "approval"
assert "pending_approval_detail" not in payload
def test_dispatch_intent_verdict_emits_child_ws_intent_verdict(self) -> None:
"""Stage 3 Step 6 — explicit verdict events are re-emitted as
child_ws_intent_verdict on the parent's SSE so the tree UI
renders the risk pill without polling."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
verdict = {
"call_id": "c1",
"risk_level": "low",
"confidence": 0.92,
"recommendation": "approve",
}
adapter._dispatch_child_event(
{
"type": "intent_verdict",
"ws_id": "child-a1",
"node_id": "node-1",
"verdict": verdict,
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_intent_verdict"
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
assert payload["node_id"] == "node-1"
assert payload["verdict"] == verdict
def test_dispatch_intent_verdict_unknown_child_drops(self) -> None:
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{
"type": "intent_verdict",
"ws_id": "ws-orphan",
"verdict": {"call_id": "c1"},
}
)
assert recorder.enqueued == []
def test_dispatch_approval_resolved_emits_child_ws_approval_resolved(
self,
) -> None:
"""Stage 3 Step 6 — paired with intent_verdict; clears the
pending-approval pill on the parent's tree UI in lockstep
with the actual decision."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "approval_resolved",
"ws_id": "child-a1",
"node_id": "node-1",
"approved": True,
"feedback": "lgtm",
"always": False,
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_approval_resolved"
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
assert payload["approved"] is True
assert payload["feedback"] == "lgtm"
assert payload["always"] is False
def test_dispatch_approval_resolved_coerces_missing_fields(self) -> None:
"""Older nodes mid-rolling-upgrade may omit approved / always /
feedback; dispatch coerces to safe defaults."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event({"type": "approval_resolved", "ws_id": "child-a1"})
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["approved"] is False
assert payload["feedback"] == ""
assert payload["always"] is False
def test_dispatch_approval_resolved_unknown_child_drops(self) -> None:
"""Symmetric to the intent_verdict drop test — events for
ws_ids the registry doesn't know about silently drop instead
of fanning out to a parent that has no business seeing them."""
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{
"type": "approval_resolved",
"ws_id": "ws-orphan",
"approved": True,
},
)
assert recorder.enqueued == []
def test_dispatch_approve_request_emits_child_ws_approve_request(
self,
) -> None:
"""Push path for the initial approval items — eliminates the
bulk-fetch race that left the coord row stuck on a loading
placeholder when the bulk fetch landed in the gap between
_emit_state(ATTENTION) and approve_tools setting _pending_approval."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
detail = {
"type": "approve_request",
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": True,
}
adapter._dispatch_child_event(
{
"type": "approve_request",
"ws_id": "child-a1",
"node_id": "node-1",
"detail": detail,
},
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_approve_request"
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
assert payload["node_id"] == "node-1"
assert payload["detail"] == detail
def test_dispatch_approve_request_unknown_child_drops(self) -> None:
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{
"type": "approve_request",
"ws_id": "ws-orphan",
"detail": {"items": []},
},
)
assert recorder.enqueued == []
+636 -134
View File
@@ -171,14 +171,56 @@ def test_route_map_matches_console_routes():
mirrors the shape we expect.
"""
assert _ROUTE_PATHS["spawn"] == "/v1/api/route/workstreams/new"
assert _ROUTE_PATHS["send"] == "/v1/api/route/send"
assert _ROUTE_PATHS["approve"] == "/v1/api/route/approve"
assert _ROUTE_PATHS["cancel"] == "/v1/api/route/cancel"
assert _ROUTE_PATHS["close"] == "/v1/api/route/workstreams/close"
assert _ROUTE_PATHS["send"] == "/v1/api/route/workstreams/{ws_id}/send"
assert _ROUTE_PATHS["approve"] == "/v1/api/route/workstreams/{ws_id}/approve"
assert _ROUTE_PATHS["cancel"] == "/v1/api/route/workstreams/{ws_id}/cancel"
assert _ROUTE_PATHS["close"] == "/v1/api/route/workstreams/{ws_id}/close"
# ``delete`` keeps the body-keyed shape — it has its own
# ``route_workstream_delete`` handler instead of going through
# the generic route_proxy.
assert _ROUTE_PATHS["delete"] == "/v1/api/route/workstreams/delete"
# Cascade endpoint lives on the console itself (not a node), so the
# path slots in the coord ws_id rather than routing through a proxy.
assert _ROUTE_PATHS["close_all_children"] == "/v1/api/coordinator/{ws_id}/close_all_children"
assert _ROUTE_PATHS["close_all_children"] == "/v1/api/workstreams/{ws_id}/close_all_children"
def test_route_paths_match_actual_console_mounts():
"""Every entry in ``_ROUTE_PATHS`` must correspond to an actually
mounted Starlette route on the console app. Catches the kind of
drift that broke close_workstream / close_all_children when the
#422 legacy URL adapter removal deleted the body-keyed
/v1/api/route/{verb} routes without a corresponding update to
the coord client's route table."""
from unittest.mock import MagicMock
from starlette.routing import Mount, Route
from turnstone.console.coordinator_client import _ROUTE_PATHS
from turnstone.console.server import create_app
app = create_app(
collector=MagicMock(),
jwt_secret="x" * 64,
)
def _walk(routes, prefix=""):
for r in routes:
if isinstance(r, Mount):
yield from _walk(r.routes, prefix=prefix + r.path)
elif isinstance(r, Route):
yield prefix + r.path
mounted = set(_walk(app.routes))
for key, template in _ROUTE_PATHS.items():
# Starlette's Route.path uses ``{name}`` placeholders just
# like our templates, so a literal containment check works.
assert template in mounted, (
f"_ROUTE_PATHS[{key!r}] = {template!r} is not a mounted "
f"console route. Mounted routes containing 'route' or "
f"'workstreams': "
f"{sorted(p for p in mounted if 'route' in p or 'workstreams' in p)}"
)
def test_spawn_posts_to_routing_proxy_with_bearer_token():
@@ -220,24 +262,26 @@ def test_spawn_omits_optional_empty_fields():
def test_send_posts_to_send_route():
client, captured = _mock_client(_ok_json({"status": 200}))
client.send("ws-x", "hello")
assert captured[0].url.path == "/v1/api/route/send"
# Path-keyed shape post-#422: ws_id rides in the URL, not the body.
assert captured[0].url.path == "/v1/api/route/workstreams/ws-x/send"
body = json.loads(captured[0].content)
assert body == {"ws_id": "ws-x", "message": "hello"}
assert body == {"message": "hello"}
def test_close_workstream_posts_to_close_route():
client, captured = _mock_client(_ok_json({"status": 200}))
client.close_workstream("ws-x")
assert captured[0].url.path == "/v1/api/route/workstreams/close"
assert captured[0].url.path == "/v1/api/route/workstreams/ws-x/close"
body = json.loads(captured[0].content)
assert body == {"ws_id": "ws-x"} # no reason → omitted
assert body == {} # no reason → omitted; ws_id rides the path
def test_close_workstream_includes_reason_when_provided():
client, captured = _mock_client(_ok_json({"status": 200}))
client.close_workstream("ws-x", reason="done")
assert captured[0].url.path == "/v1/api/route/workstreams/ws-x/close"
body = json.loads(captured[0].content)
assert body == {"ws_id": "ws-x", "reason": "done"}
assert body == {"reason": "done"}
def test_close_all_children_posts_to_console_endpoint():
@@ -256,7 +300,7 @@ def test_close_all_children_posts_to_console_endpoint():
)
result = client.close_all_children(reason="batch done")
assert result["closed"] == ["c-1", "c-2"]
assert captured[0].url.path == "/v1/api/coordinator/coord-1/close_all_children"
assert captured[0].url.path == "/v1/api/workstreams/coord-1/close_all_children"
assert captured[0].headers["Authorization"] == "Bearer test-token"
body = json.loads(captured[0].content)
assert body == {"reason": "batch done"}
@@ -301,11 +345,15 @@ def test_approve_and_cancel_hit_their_routes():
client, captured = _mock_client(_ok_json({"status": 200}))
client.approve("ws-x", call_id="c-1", approved=True, feedback="ok", always=True)
client.cancel("ws-x")
assert captured[0].url.path == "/v1/api/route/approve"
assert captured[1].url.path == "/v1/api/route/cancel"
# Path-keyed shape post-#422: ws_id rides the URL.
assert captured[0].url.path == "/v1/api/route/workstreams/ws-x/approve"
assert captured[1].url.path == "/v1/api/route/workstreams/ws-x/cancel"
approve_body = json.loads(captured[0].content)
assert approve_body["approved"] is True
assert approve_body["always"] is True
assert approve_body["call_id"] == "c-1"
# ws_id moved to the URL — make sure we didn't double-encode it.
assert "ws_id" not in approve_body
def test_http_error_returns_structured_failure():
@@ -365,7 +413,7 @@ def test_mutating_ops_accept_self_ws_id():
client, captured = _mock_client(_ok_json({"status": 200}))
client.send("coord-1", "hi")
assert len(captured) == 1
assert captured[0].url.path == "/v1/api/route/send"
assert captured[0].url.path == "/v1/api/route/workstreams/coord-1/send"
# ---------------------------------------------------------------------------
@@ -481,66 +529,6 @@ def test_list_children_skill_filter_avoids_n_plus_one(populated_storage, monkeyp
assert call_count["n"] == 0
def test_count_active_children_counts_non_terminal_states(populated_storage):
"""Budget count must use an aggregate SQL query so a tail of
recently-closed children can't push live rows past a LIMIT and
silently undercount (Copilot #3 on PR #387).
populated_storage has:
- coord-1 (coordinator, excluded)
- child-a (interactive, idle) counted
- child-b (interactive, running) counted
- child-coord (coordinator child) excluded (kind filter doesn't apply
to count_workstreams_by_state, but
it still matches parent_ws_id+user_id)
- unrelated (no parent) excluded (parent filter)
- cross-tenant-child (user-2) excluded (user_id filter)
child-coord DOES count against count_workstreams_by_state because
the aggregate doesn't filter by kind — the budget is per-coord
across any descendant type. That's fine semantically: a
coordinator that spawns a nested coord still occupies a slot.
"""
client = _make_read_client(populated_storage)
count = client.count_active_children("coord-1")
# child-a (idle) + child-b (running) + child-coord (running/default) = 3
assert count == 3
def test_count_active_children_excludes_closed_and_deleted(populated_storage):
"""A closed tail must not count toward the active-children budget —
this is the whole reason for switching off list_children's
LIMIT-then-filter path.
"""
# Close child-a and mark child-b deleted. child-coord stays active.
populated_storage.update_workstream_state("child-a", "closed")
populated_storage.update_workstream_state("child-b", "deleted")
client = _make_read_client(populated_storage)
count = client.count_active_children("coord-1")
assert count == 1 # only child-coord survives
def test_count_active_children_rejects_foreign_parent(populated_storage):
"""Tenant guard — a crafted parent_ws_id other than the coord's own
returns 0 without hitting storage."""
client = _make_read_client(populated_storage)
# The client's coord_ws_id is "coord-1" (see _make_read_client).
# Counting against a different id must not leak anyone else's count.
assert client.count_active_children("other-coord") == 0
def test_count_active_children_fails_open_on_storage_error(populated_storage, monkeypatch):
"""Budget is operator safety, not a security gate — a broken storage
path must return 0 so the coord still makes progress."""
client = _make_read_client(populated_storage)
def _boom(**_kwargs):
raise RuntimeError("storage broken")
monkeypatch.setattr(populated_storage, "count_workstreams_by_state", _boom)
assert client.count_active_children("coord-1") == 0
def test_list_children_signals_truncation_when_page_full_and_filter_drops(
populated_storage,
):
@@ -564,6 +552,36 @@ def test_inspect_missing_ws_returns_error(populated_storage):
assert "error" in result
def test_inspect_not_found_does_not_echo_ws_id_in_error_string(populated_storage):
"""The error STRING is bare ("workstream not found") — the
structured ``ws_id`` field carries the queried id. Pre-fix the
error message echoed the ws_id back at the caller who just sent
it, which was redundant and a stylistic departure from the rest
of the surface. Echo-in-string is also one more place a
hostile/oversize ws_id could land in operator-facing text."""
client = _make_read_client(populated_storage)
result = client.inspect("does-not-exist-xyz")
assert result["error"] == "workstream not found"
# The structured field still carries the ws_id for context.
assert result["ws_id"] == "does-not-exist-xyz"
def test_inspect_cross_tenant_returns_same_shape_as_missing(populated_storage):
"""The cross-tenant guard MUST return the exact same shape as a
genuinely missing ws_id that's the existence-leak defence the
error-string echo was carrying weight for too. Asserting the
shape match here pins the property going forward."""
# ``unrelated`` exists in storage but is not a coord-1 child.
client = _make_read_client(populated_storage)
cross_tenant = client.inspect("unrelated")
missing = client.inspect("does-not-exist-abc")
# Same key set, same error string, only the ws_id field differs.
assert cross_tenant.keys() == missing.keys()
assert cross_tenant["error"] == missing["error"] == "workstream not found"
assert cross_tenant["ws_id"] == "unrelated"
assert missing["ws_id"] == "does-not-exist-abc"
def test_list_children_excludes_closed_by_default(tmp_path):
"""Default ``list_children`` filters out closed / deleted rows —
the common "what's still running?" query shouldn't have to
@@ -931,6 +949,137 @@ def test_list_nodes_empty_on_no_matching_filters(storage_with_nodes):
assert result["truncated"] is False
def test_list_nodes_surfaces_healthy_model_aliases(tmp_path):
"""The node's heartbeat loop projects its registry into a ``models``
metadata entry shaped like ``[{alias, provider, healthy}, ...]``.
``list_nodes`` flattens that to the healthy-alias list at the top
level (under ``model_aliases``) so a coordinator can pass aliases
straight to ``spawn_workstream(model=)`` without having to
introspect the metadata blob. The provider-side model identifier
(``cfg.model``) is intentionally NOT in the payload coords kept
reaching for it when they should pass the local alias."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-x",
[
("arch", "x86_64", "auto"),
(
"models",
[
{"alias": "gpt5", "provider": "openai", "healthy": True},
{"alias": "claude-opus-47", "provider": "anthropic", "healthy": True},
{"alias": "broken", "provider": "openai", "healthy": False},
],
"auto",
),
],
)
_register_service(st, "node-x")
client = _make_read_client(st)
result = client.list_nodes()
node = result["nodes"][0]
assert node["model_aliases"] == ["gpt5", "claude-opus-47"]
# Full per-alias info still available under metadata for callers
# that want provider / healthy detail (e.g. surfacing degraded
# aliases in a UI).
full = node["metadata"]["models"]["value"]
assert {row["alias"] for row in full} == {"gpt5", "claude-opus-47", "broken"}
# ``model`` (the provider-side identifier) is intentionally absent
# — keep the payload to the three values a coord actually uses.
for row in full:
assert "model" not in row
def test_list_nodes_model_aliases_distinct_from_metadata_models(tmp_path):
"""Pin the naming distinction explicitly: the top-level shortlist
(``model_aliases``, list of strings) and the rich metadata blob
(``metadata.models.value``, list of dicts) live under different
keys so a caller that confuses them gets a clear KeyError rather
than a silent shape mismatch."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-x",
[
(
"models",
[{"alias": "a", "provider": "openai", "healthy": True}],
"auto",
),
],
)
_register_service(st, "node-x")
client = _make_read_client(st)
node = client.list_nodes()["nodes"][0]
# No top-level ``models`` field — only ``model_aliases``.
assert "models" not in node
assert node["model_aliases"] == ["a"]
# Rich shape stays under metadata.
assert isinstance(node["metadata"]["models"]["value"], list)
assert isinstance(node["metadata"]["models"]["value"][0], dict)
def test_list_nodes_model_aliases_empty_when_node_has_not_published(tmp_path):
"""Nodes from older builds — or a node mid-startup before its first
metadata write won't have a ``models`` entry. The top-level
``model_aliases`` field defaults to ``[]`` rather than being
omitted so coordinators can rely on the key being present."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(st, "node-y", [("arch", "x86_64", "auto")])
_register_service(st, "node-y")
client = _make_read_client(st)
result = client.list_nodes()
assert result["nodes"][0]["model_aliases"] == []
def test_list_nodes_models_tolerates_malformed_entries(tmp_path):
"""If a node ever stores a malformed ``models`` entry (wrong outer
type, missing alias, non-bool healthy), the projection drops the
bad rows rather than raising the rest of the response should
still be useful."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-z",
[
(
"models",
[
{"alias": "ok", "provider": "p", "healthy": True},
"not-a-dict",
{"provider": "p", "healthy": True}, # missing alias
{"alias": "", "healthy": True}, # empty alias
{"alias": "degraded", "healthy": False},
{"alias": 42, "healthy": True}, # non-string alias
],
"auto",
),
],
)
_register_service(st, "node-z")
client = _make_read_client(st)
result = client.list_nodes()
assert result["nodes"][0]["model_aliases"] == ["ok"]
def test_list_nodes_models_handles_non_list_payload(tmp_path):
"""A node with a corrupted models entry (dict, scalar, null) shouldn't
blow up the whole list_nodes call. ``model_aliases`` falls back to ``[]``."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-w",
[
("models", {"oops": "not a list"}, "auto"),
],
)
_register_service(st, "node-w")
client = _make_read_client(st)
result = client.list_nodes()
assert result["nodes"][0]["model_aliases"] == []
# ---------------------------------------------------------------------------
# list_skills
# ---------------------------------------------------------------------------
@@ -1142,6 +1291,37 @@ def test_inspect_omits_close_reason_when_absent(populated_storage):
assert "close_reason" not in result
def test_inspect_surfaces_last_error_when_state_is_error(populated_storage):
"""A child that crashed (e.g. provider 4xx after retry exhaustion)
has its exception text persisted to workstream_config.last_error
by the worker-thread error path; inspect surfaces it for terminal
error rows so the coordinator can triage without parsing the
assistant tail."""
populated_storage.update_workstream_state("child-a", "error")
populated_storage.save_workstream_config(
"child-a",
{"last_error": "AuthenticationError: invalid api key"},
)
client = _make_read_client(populated_storage)
result = client.inspect("child-a")
assert result.get("last_error") == "AuthenticationError: invalid api key"
def test_inspect_omits_last_error_for_non_error_terminal_states(populated_storage):
"""A historic last_error from an earlier failed turn that was later
closed cleanly must NOT surface on the close the coord would
misread the close as an error close. Gating on state=='error'
keeps the surface honest."""
populated_storage.update_workstream_state("child-a", "closed")
populated_storage.save_workstream_config(
"child-a",
{"last_error": "stale error from a previous failed turn"},
)
client = _make_read_client(populated_storage)
result = client.inspect("child-a")
assert "last_error" not in result
def test_inspect_skips_workstream_config_read_for_live_workstreams(populated_storage, monkeypatch):
"""Hot-path optimisation: live (non-terminal) workstreams must NOT
pay the per-inspect load_workstream_config round-trip. close_reason
@@ -1402,7 +1582,329 @@ def test_wait_for_workstream_handles_non_string_mode(populated_storage):
# ---------------------------------------------------------------------------
# task_list
# wait_for_workstream — last-message bundling
# ---------------------------------------------------------------------------
#
# Each terminal child's last assistant turn (or a status sentinel) is
# bundled inline so the coord LLM doesn't need a follow-up
# inspect_workstream round-trip per ws. The fields are additive
# (``message`` / ``truncated``), so existing wait tests stay green.
def test_wait_for_workstream_idle_returns_last_assistant_message(populated_storage):
"""A child that finished normally surfaces its final assistant
turn inline so the coord doesn't have to inspect to read it."""
populated_storage.save_message("child-a", "user", "what's the answer?")
populated_storage.save_message("child-a", "assistant", "the answer is 42")
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["state"] == "idle"
assert snap["message"] == "the answer is 42"
assert snap["truncated"] is False
def test_wait_for_workstream_idle_walks_past_trailing_tool_messages(populated_storage):
"""The most recent assistant turn often sits behind a few tool
messages (assistant emits tool_calls tool results land final
assistant content follows). The walk must skip non-assistant
rows when picking the last assistant content."""
populated_storage.save_message("child-a", "user", "do the thing")
populated_storage.save_message("child-a", "assistant", "calling tool")
populated_storage.save_message("child-a", "tool", "tool output", tool_call_id="t1")
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
# The assistant message above is the most recent assistant turn —
# the trailing tool row must not block extraction.
assert result["results"]["child-a"]["message"] == "calling tool"
def test_wait_for_workstream_idle_skips_empty_assistant_with_tool_calls(populated_storage):
"""An assistant message with empty content + only tool_calls isn't
a final answer walk further back for the last assistant message
that actually has text."""
populated_storage.save_message("child-a", "user", "first turn")
populated_storage.save_message("child-a", "assistant", "first assistant reply")
populated_storage.save_message("child-a", "user", "second turn")
populated_storage.save_message(
"child-a", "assistant", "", tool_calls='[{"id": "t1", "name": "x"}]'
)
populated_storage.save_message("child-a", "tool", "tool result", tool_call_id="t1")
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
# Last assistant with non-empty content is the FIRST assistant message
# — the empty-content tool-calls assistant must be skipped.
assert result["results"]["child-a"]["message"] == "first assistant reply"
def test_wait_for_workstream_idle_no_assistant_returns_sentinel(populated_storage):
"""A workstream that reaches idle without an assistant turn in the
tail (rare but possible for a freshly registered ws closed before
generation, or a long-running ws whose final assistant message is
buried beyond the tail window) gets a hedged sentinel rather than
null the model can distinguish 'no recent output' from 'still
running'."""
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["state"] == "idle"
# No messages were saved for child-a in this test — sentinel kicks in.
# Wording is hedged ("recent") because the tail-only walk can't
# actually prove no assistant output exists in the full history.
assert snap["message"] == "(no recent assistant output)"
assert snap["truncated"] is False
def test_wait_for_workstream_error_returns_last_assistant_message(populated_storage):
"""An errored child still gets its last assistant turn surfaced —
that's usually the most useful diagnostic ('I was about to ...
when the error happened')."""
populated_storage.update_workstream_state("child-a", "error")
populated_storage.save_message("child-a", "user", "hi")
populated_storage.save_message("child-a", "assistant", "partial output before crash")
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["state"] == "error"
assert snap["message"] == "partial output before crash"
def test_wait_for_workstream_error_with_no_output_returns_sentinel(populated_storage):
"""When error fires with no assistant content in the tail (e.g. a
pre-flight provider auth failure that crashes before the model
speaks, or a >18-parallel-tool-call burst whose only assistant
row carries empty content), the same hedged sentinel applies.
The wording deliberately doesn't claim 'before producing output'
the tail-only walk can't prove that.
"""
populated_storage.update_workstream_state("child-a", "error")
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["state"] == "error"
assert snap["message"] == "(no recent assistant output)"
assert snap["truncated"] is False
def test_wait_for_workstream_error_prefers_persisted_last_error(populated_storage):
"""When the worker thread persists ``last_error`` on a crash (e.g.
provider 429 after retry exhaustion, model misconfig), the error
text wins over the assistant tail the actual cause is more
actionable than a half-finished prior turn."""
populated_storage.update_workstream_state("child-a", "error")
populated_storage.save_message("child-a", "assistant", "partial output before crash")
populated_storage.save_workstream_config(
"child-a",
{"last_error": "RateLimitError: 429 too many requests after 5 retries"},
)
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["state"] == "error"
assert snap["message"] == "RateLimitError: 429 too many requests after 5 retries"
assert snap["truncated"] is False
def test_wait_for_workstream_error_falls_back_to_assistant_when_no_last_error(populated_storage):
"""Legacy / pre-fix error rows (state=error, no last_error config)
keep the existing assistant-tail behaviour the upgrade is
additive."""
populated_storage.update_workstream_state("child-a", "error")
populated_storage.save_message("child-a", "user", "hi")
populated_storage.save_message("child-a", "assistant", "partial output before crash")
# Note: no save_workstream_config call.
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["message"] == "partial output before crash"
def test_wait_for_workstream_closed_returns_sentinel(populated_storage):
"""Closed children get a status sentinel rather than a partial
last message a half-finished thought from a workstream the
operator explicitly closed isn't useful (and could be misleading)."""
populated_storage.update_workstream_state("child-a", "closed")
populated_storage.save_message("child-a", "assistant", "mid-thought when closed")
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["state"] == "closed"
assert snap["message"] == "(workstream closed)"
assert snap["truncated"] is False
def test_wait_for_workstream_denied_returns_sentinel(populated_storage):
"""Cross-tenant / nonexistent ws_ids surface as denied — the
sentinel lets the coord LLM recognise the rejection without
parsing state strings on its own."""
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["unrelated"], timeout=5, mode="any")
snap = result["results"]["unrelated"]
assert snap["state"] == "denied"
assert snap["message"].startswith("(workstream denied")
assert snap["truncated"] is False
def test_wait_for_workstream_running_child_message_is_null(populated_storage):
"""A still-running child after a timeout must report
``message=None`` anything else would be a partial last message
pretending to be a final answer. The coord uses null to know
'still working, inspect later'."""
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a", "child-b"], timeout=1.0, mode="all")
# mode='all' on (idle, running) hits the timeout — child-b is still
# running and must come back with message=None.
assert result["complete"] is False
assert result["results"]["child-b"]["state"] == "running"
assert result["results"]["child-b"]["message"] is None
assert result["results"]["child-b"]["truncated"] is False
def test_wait_for_workstream_truncates_oversize_message(populated_storage):
"""A message past WAIT_MESSAGE_MAX_BYTES is truncated from the
END (preserve the lead) and ``truncated=True`` so the coord LLM
knows to inspect for the rest if it needs the full text."""
from turnstone.console.coordinator_client import WAIT_MESSAGE_MAX_BYTES
big = "A" * (WAIT_MESSAGE_MAX_BYTES * 2)
populated_storage.save_message("child-a", "user", "hi")
populated_storage.save_message("child-a", "assistant", big)
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
# Truncated — exactly the cap in bytes (single-byte chars), with the
# head preserved.
assert snap["truncated"] is True
assert len(snap["message"].encode("utf-8")) == WAIT_MESSAGE_MAX_BYTES
assert snap["message"].startswith("AAAA")
def test_wait_for_workstream_storage_failure_leaves_message_null(populated_storage, monkeypatch):
"""A transient storage error during the message read must not
fail the wait the coord still gets state/tokens/updated, and
the per-ws ``message`` collapses to None so the model can fall
back to inspect."""
populated_storage.update_workstream_state("child-a", "idle")
def _broken_load(*_a, **_kw):
raise RuntimeError("simulated storage outage")
monkeypatch.setattr(populated_storage, "load_messages", _broken_load)
client = _make_read_client(populated_storage)
result = client.wait_for_workstream(["child-a"], timeout=5, mode="any")
snap = result["results"]["child-a"]
assert snap["state"] == "idle"
assert snap["message"] is None
assert snap["truncated"] is False
def test_wait_for_workstream_does_not_pollute_progress_callback(populated_storage):
"""The wait_progress SSE event shape is documented as separate
from the tool result the per-tick snapshot dicts handed to the
progress callback must NOT carry the new ``message`` /
``truncated`` fields, since enrichment happens after the loop
exits."""
populated_storage.save_message("child-a", "assistant", "ok")
client = _make_read_client(populated_storage)
captured: list[dict[str, dict[str, Any]]] = []
def _cb(snap: dict[str, dict[str, Any]], _elapsed: float) -> None:
# Deep-copy so a later mutation by enrichment can't fool the
# assertion (we want the shape AT CALLBACK TIME, not at end).
import copy
captured.append(copy.deepcopy(snap))
client.wait_for_workstream(["child-a"], timeout=5, mode="any", progress_callback=_cb)
assert captured # at least one tick fired
for tick in captured:
for per_ws in tick.values():
assert "message" not in per_ws
assert "truncated" not in per_ws
# ---------------------------------------------------------------------------
# wait_for_workstream — helper-function unit tests
# ---------------------------------------------------------------------------
def test_truncate_wait_message_below_cap_is_passthrough():
from turnstone.console.coordinator_client import _truncate_wait_message
text, trunc = _truncate_wait_message("hello", 100)
assert text == "hello"
assert trunc is False
def test_truncate_wait_message_exact_cap_is_passthrough():
from turnstone.console.coordinator_client import _truncate_wait_message
text, trunc = _truncate_wait_message("a" * 5, 5)
assert text == "aaaaa"
assert trunc is False
def test_truncate_wait_message_oversize_truncates_to_byte_cap():
from turnstone.console.coordinator_client import _truncate_wait_message
text, trunc = _truncate_wait_message("a" * 10, 5)
assert text == "aaaaa"
assert trunc is True
def test_truncate_wait_message_handles_utf8_boundary():
"""A multi-byte codepoint must never be split — back off to a valid
UTF-8 boundary even if it lands a couple bytes under the cap."""
from turnstone.console.coordinator_client import _truncate_wait_message
# "café" is 5 bytes (c=1, a=1, f=1, é=2). Cap at 4 bytes lands
# mid-codepoint on the é; truncation must back off to 3 bytes.
text, trunc = _truncate_wait_message("café", 4)
assert trunc is True
assert text == "caf"
# And the result must be valid UTF-8 — re-encoding doesn't error.
text.encode("utf-8")
def test_truncate_wait_message_zero_or_negative_cap_returns_empty():
from turnstone.console.coordinator_client import _truncate_wait_message
text, trunc = _truncate_wait_message("anything", 0)
assert text == ""
assert trunc is True
def test_last_assistant_text_returns_content_when_present(populated_storage):
"""Pins the third leg of the tri-state contract: a populated tail
returns the actual assistant content string (not ``""``, not
``None``). Integration tests cover this through enrichment, but a
direct unit test makes the contract harder to break in a refactor."""
from turnstone.console.coordinator_client import _last_assistant_text
populated_storage.save_message("child-a", "user", "hello")
populated_storage.save_message("child-a", "assistant", "hi back")
assert _last_assistant_text(populated_storage, "child-a") == "hi back"
def test_last_assistant_text_returns_empty_when_no_messages(populated_storage):
from turnstone.console.coordinator_client import _last_assistant_text
# child-a has no messages saved.
assert _last_assistant_text(populated_storage, "child-a") == ""
def test_last_assistant_text_returns_none_on_storage_failure(populated_storage, monkeypatch):
from turnstone.console.coordinator_client import _last_assistant_text
def _broken(*_a, **_kw):
raise RuntimeError("boom")
monkeypatch.setattr(populated_storage, "load_messages", _broken)
assert _last_assistant_text(populated_storage, "child-a") is None
# ---------------------------------------------------------------------------
# tasks
# ---------------------------------------------------------------------------
@@ -1412,177 +1914,177 @@ def _task_client(tmp_path) -> CoordinatorClient:
return _make_read_client(st)
def test_task_list_get_empty_envelope_on_fresh_ws(tmp_path):
def test_tasks_get_empty_envelope_on_fresh_ws(tmp_path):
client = _task_client(tmp_path)
env = client.task_list_get("coord-1")
env = client.tasks_get("coord-1")
assert env == {"version": 1, "tasks": []}
def test_task_list_add_then_get_roundtrip(tmp_path):
def test_tasks_add_then_get_roundtrip(tmp_path):
client = _task_client(tmp_path)
task = client.task_list_add("coord-1", title="spawn worker")
task = client.tasks_add("coord-1", title="spawn worker")
assert task["title"] == "spawn worker"
assert task["status"] == "pending"
env = client.task_list_get("coord-1")
env = client.tasks_get("coord-1")
assert len(env["tasks"]) == 1
assert env["tasks"][0]["id"] == task["id"]
def test_task_list_add_rejects_empty_title(tmp_path):
def test_tasks_add_rejects_empty_title(tmp_path):
client = _task_client(tmp_path)
result = client.task_list_add("coord-1", title=" ")
result = client.tasks_add("coord-1", title=" ")
assert "error" in result
def test_task_list_add_rejects_invalid_status(tmp_path):
def test_tasks_add_rejects_invalid_status(tmp_path):
client = _task_client(tmp_path)
result = client.task_list_add("coord-1", title="x", status="nonsense")
result = client.tasks_add("coord-1", title="x", status="nonsense")
assert "error" in result
def test_task_list_add_rejects_title_over_200(tmp_path):
def test_tasks_add_rejects_title_over_200(tmp_path):
"""Silent truncation is a data-integrity footgun: the model may
rely on the title it sent, not the one stored. Reject instead."""
client = _task_client(tmp_path)
long_title = "a" * 201
result = client.task_list_add("coord-1", title=long_title)
result = client.tasks_add("coord-1", title=long_title)
assert "error" in result
assert "too long" in result["error"]
# Exactly 200 chars is the boundary and still accepted.
boundary = "a" * 200
task = client.task_list_add("coord-1", title=boundary)
task = client.tasks_add("coord-1", title=boundary)
assert "error" not in task
assert len(task["title"]) == 200
def test_task_list_update_rejects_title_over_200(tmp_path):
def test_tasks_update_rejects_title_over_200(tmp_path):
client = _task_client(tmp_path)
added = client.task_list_add("coord-1", title="original")
result = client.task_list_update("coord-1", task_id=added["id"], title="b" * 201)
added = client.tasks_add("coord-1", title="original")
result = client.tasks_update("coord-1", task_id=added["id"], title="b" * 201)
assert "error" in result
assert "too long" in result["error"]
# Original title untouched when update rejected.
env = client.task_list_get("coord-1")
env = client.tasks_get("coord-1")
assert env["tasks"][0]["title"] == "original"
def test_task_list_update_by_id(tmp_path):
def test_tasks_update_by_id(tmp_path):
client = _task_client(tmp_path)
added = client.task_list_add("coord-1", title="plan")
updated = client.task_list_update(
added = client.tasks_add("coord-1", title="plan")
updated = client.tasks_update(
"coord-1", task_id=added["id"], status="done", child_ws_id="ws-child"
)
assert updated["status"] == "done"
assert updated["child_ws_id"] == "ws-child"
def test_task_list_update_missing_id(tmp_path):
def test_tasks_update_missing_id(tmp_path):
client = _task_client(tmp_path)
result = client.task_list_update("coord-1", task_id="nope", status="done")
result = client.tasks_update("coord-1", task_id="nope", status="done")
assert "error" in result
def test_task_list_remove(tmp_path):
def test_tasks_remove(tmp_path):
client = _task_client(tmp_path)
added = client.task_list_add("coord-1", title="plan")
first = client.task_list_remove("coord-1", task_id=added["id"])
added = client.tasks_add("coord-1", title="plan")
first = client.tasks_remove("coord-1", task_id=added["id"])
assert first.get("ok") is True
assert first.get("task_id") == added["id"]
# Second remove of the same id returns a distinguishable not-found
# error (NOT a silent False that would mask a corrupt envelope).
second = client.task_list_remove("coord-1", task_id=added["id"])
second = client.tasks_remove("coord-1", task_id=added["id"])
assert "error" in second
assert "not found" in second["error"]
assert client.task_list_get("coord-1")["tasks"] == []
assert client.tasks_get("coord-1")["tasks"] == []
def test_task_list_reorder_requires_permutation(tmp_path):
def test_tasks_reorder_requires_permutation(tmp_path):
client = _task_client(tmp_path)
a = client.task_list_add("coord-1", title="a")
b = client.task_list_add("coord-1", title="b")
a = client.tasks_add("coord-1", title="a")
b = client.tasks_add("coord-1", title="b")
# Partial set — must reject.
bad = client.task_list_reorder("coord-1", task_ids=[a["id"]])
bad = client.tasks_reorder("coord-1", task_ids=[a["id"]])
assert "error" in bad
# Wrong id — reject.
wrong = client.task_list_reorder("coord-1", task_ids=[a["id"], "ghost"])
wrong = client.tasks_reorder("coord-1", task_ids=[a["id"], "ghost"])
assert "error" in wrong
# Valid permutation — accept.
ok = client.task_list_reorder("coord-1", task_ids=[b["id"], a["id"]])
ok = client.tasks_reorder("coord-1", task_ids=[b["id"], a["id"]])
assert ok.get("ok") is True
env = client.task_list_get("coord-1")
env = client.tasks_get("coord-1")
assert [t["id"] for t in env["tasks"]] == [b["id"], a["id"]]
def test_task_list_cross_ws_scope_violation_is_noop(tmp_path):
def test_tasks_cross_ws_scope_violation_is_noop(tmp_path):
client = _task_client(tmp_path)
# Client is bound to coord-1; anything else returns an empty envelope
# or an error without touching storage.
assert client.task_list_get("other-ws") == {"version": 1, "tasks": []}
res_add = client.task_list_add("other-ws", title="sneak")
assert client.tasks_get("other-ws") == {"version": 1, "tasks": []}
res_add = client.tasks_add("other-ws", title="sneak")
assert "error" in res_add
res_remove = client.task_list_remove("other-ws", task_id="x")
res_remove = client.tasks_remove("other-ws", task_id="x")
assert "error" in res_remove
assert "scope violation" in res_remove["error"]
def test_task_list_corrupt_json_returns_empty_envelope(tmp_path):
def test_tasks_corrupt_json_returns_empty_envelope(tmp_path):
"""A hand-edited / corrupt config row must not crash the tool."""
st = SQLiteBackend(str(tmp_path / "tasks.db"))
st.register_workstream("coord-1", kind="coordinator", user_id="user-1")
st.save_workstream_config("coord-1", {"tasks": "{not json"})
client = _make_read_client(st)
env = client.task_list_get("coord-1")
env = client.tasks_get("coord-1")
assert env == {"version": 1, "tasks": []}
def test_task_list_mutations_refuse_corrupt_envelope(tmp_path):
def test_tasks_mutations_refuse_corrupt_envelope(tmp_path):
"""When the envelope is corrupt on disk, mutators must error out
(rather than silently overwrite lost-data safety)."""
st = SQLiteBackend(str(tmp_path / "tasks.db"))
st.register_workstream("coord-1", kind="coordinator", user_id="user-1")
st.save_workstream_config("coord-1", {"tasks": "{not json"})
client = _make_read_client(st)
add_result = client.task_list_add("coord-1", title="new")
add_result = client.tasks_add("coord-1", title="new")
assert "error" in add_result
assert "corrupt" in add_result["error"]
# Also: the corrupt blob is preserved after the refused mutation.
assert st.load_workstream_config("coord-1").get("tasks") == "{not json"
update_result = client.task_list_update("coord-1", task_id="x", status="done")
update_result = client.tasks_update("coord-1", task_id="x", status="done")
assert "error" in update_result
reorder_result = client.task_list_reorder("coord-1", task_ids=[])
reorder_result = client.tasks_reorder("coord-1", task_ids=[])
assert "error" in reorder_result
remove_result = client.task_list_remove("coord-1", task_id="x")
remove_result = client.tasks_remove("coord-1", task_id="x")
assert "error" in remove_result
assert "corrupt" in remove_result["error"]
def test_task_list_add_enforces_capacity_cap(tmp_path, monkeypatch):
def test_tasks_add_enforces_capacity_cap(tmp_path, monkeypatch):
from turnstone.console import coordinator_client as cc_module
monkeypatch.setattr(cc_module, "_TASK_LIST_MAX", 3)
monkeypatch.setattr(cc_module, "_TASKS_MAX", 3)
client = _task_client(tmp_path)
for i in range(3):
client.task_list_add("coord-1", title=f"t{i}")
overflow = client.task_list_add("coord-1", title="no-room")
client.tasks_add("coord-1", title=f"t{i}")
overflow = client.tasks_add("coord-1", title="no-room")
assert "error" in overflow
assert "capacity" in overflow["error"]
# After a remove, add succeeds again.
env = client.task_list_get("coord-1")
client.task_list_remove("coord-1", task_id=env["tasks"][0]["id"])
added = client.task_list_add("coord-1", title="retry")
env = client.tasks_get("coord-1")
client.tasks_remove("coord-1", task_id=env["tasks"][0]["id"])
added = client.tasks_add("coord-1", title="retry")
assert "error" not in added
def test_task_list_save_preserves_other_workstream_config_keys(tmp_path):
"""_save_task_list writes only the 'tasks' key so other keys survive."""
def test_tasks_save_preserves_other_workstream_config_keys(tmp_path):
"""_save_tasks writes only the 'tasks' key so other keys survive."""
st = SQLiteBackend(str(tmp_path / "tasks.db"))
st.register_workstream("coord-1", kind="coordinator", user_id="user-1")
st.save_workstream_config("coord-1", {"reasoning_effort": "high"})
client = _make_read_client(st)
client.task_list_add("coord-1", title="plan")
client.tasks_add("coord-1", title="plan")
config = st.load_workstream_config("coord-1")
assert config.get("reasoning_effort") == "high"
assert config.get("tasks") # task_list wrote its key too
assert config.get("tasks") # tasks wrote its key too
def test_live_cache_lru_eviction_caps_memory(tmp_path):
@@ -1777,7 +2279,7 @@ def test_cleanup_dead_task_child_refs_blanks_dead_links(populated_storage):
)
blanked = client.cleanup_dead_task_child_refs("coord-1")
assert blanked == 1
envelope = client.task_list_get("coord-1")
envelope = client.tasks_get("coord-1")
tasks_by_id = {t["id"]: t for t in envelope["tasks"]}
# Live link preserved.
assert tasks_by_id["t1"]["child_ws_id"] == "child-a"
@@ -1813,9 +2315,9 @@ def test_cleanup_dead_task_child_refs_all_alive_is_noop(populated_storage):
def test_cleanup_dead_task_child_refs_empty_envelope(populated_storage):
"""A coordinator with no task_list persisted returns 0 without
"""A coordinator with no tasks persisted returns 0 without
raising the cleanup runs on every close, including those that
never used the task_list tool."""
never used the tasks tool."""
client = _make_read_client(populated_storage)
blanked = client.cleanup_dead_task_child_refs("coord-1")
assert blanked == 0
@@ -1832,7 +2334,7 @@ def test_cleanup_dead_task_child_refs_corrupt_envelope_skips(populated_storage):
def test_cleanup_dead_task_child_refs_uses_task_lock(populated_storage):
"""The cleanup must acquire the same per-ws _task_lock that
task_list_add/update/remove/reorder hold, so a close racing an
tasks_add/update/remove/reorder hold, so a close racing an
in-flight mutation can't lose writes (#bug-6). Verified by
swapping the cached lock for a stand-in that records acquisition."""
client = _make_read_client(populated_storage)
+14 -12
View File
@@ -21,6 +21,7 @@ from tests._coord_test_helpers import (
_build_mgr,
_fake_registry,
_FakeConfigStore,
_seed_children,
)
from turnstone.console.server import coordinator_close_all_children
from turnstone.core.storage._sqlite import SQLiteBackend
@@ -38,7 +39,7 @@ def _make_client(storage, *, coord_mgr, alias="my-model", registry=None) -> Test
app = Starlette(
routes=[
Route(
"/v1/api/coordinator/{ws_id}/close_all_children",
"/v1/api/workstreams/{ws_id}/close_all_children",
coordinator_close_all_children,
methods=["POST"],
),
@@ -46,6 +47,7 @@ def _make_client(storage, *, coord_mgr, alias="my-model", registry=None) -> Test
middleware=[Middleware(_AuthMiddleware)],
)
app.state.coord_mgr = coord_mgr
app.state.coord_adapter = coord_mgr._adapter if coord_mgr is not None else None
app.state.config_store = _FakeConfigStore({"coordinator.model_alias": alias})
app.state.coord_registry = registry
app.state.coord_registry_error = "" if coord_mgr else "registry missing"
@@ -57,7 +59,7 @@ def _make_client(storage, *, coord_mgr, alias="my-model", registry=None) -> Test
def test_close_all_children_closes_each_child_and_audits(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
mgr.register_children(coord.id, ["child-1", "child-2", "child-3"])
_seed_children(mgr._adapter, coord.id, ["child-1", "child-2", "child-3"])
def _close(wid, reason):
if wid == "child-2":
@@ -71,7 +73,7 @@ def test_close_all_children_closes_each_child_and_audits(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={"reason": "tests done"},
headers=_COORD_HEADERS,
)
@@ -108,7 +110,7 @@ def test_close_all_children_routes_404_to_skipped_bucket(storage):
is 'already gone', not a dispatch failure. Route to skipped."""
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
mgr.register_children(coord.id, ["stale-child"])
_seed_children(mgr._adapter, coord.id, ["stale-child"])
coord_client = MagicMock()
coord_client.close_workstream.return_value = {
@@ -120,7 +122,7 @@ def test_close_all_children_routes_404_to_skipped_bucket(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={},
headers=_COORD_HEADERS,
)
@@ -139,7 +141,7 @@ def test_close_all_children_empty_children_still_audits(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={},
headers=_COORD_HEADERS,
)
@@ -154,13 +156,13 @@ def test_close_all_children_empty_children_still_audits(storage):
def test_close_all_children_without_coord_client_marks_all_failed(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
mgr.register_children(coord.id, ["child-a", "child-b"])
_seed_children(mgr._adapter, coord.id, ["child-a", "child-b"])
coord.session = MagicMock()
coord.session._coord_client = None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={},
headers=_COORD_HEADERS,
)
@@ -179,7 +181,7 @@ def test_close_all_children_rejects_non_string_reason(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={"reason": 123},
headers=_COORD_HEADERS,
)
@@ -194,7 +196,7 @@ def test_close_all_children_rejects_overlong_reason(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={"reason": "x" * 600},
headers=_COORD_HEADERS,
)
@@ -207,7 +209,7 @@ def test_close_all_children_404_when_session_not_loaded(storage):
coord.session = None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={},
headers=_COORD_HEADERS,
)
@@ -227,7 +229,7 @@ def test_close_all_children_service_token_cannot_bypass_admin_coordinator(storag
headers = {"X-Test-User": "user-1", "X-Test-Perms": ""}
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/close_all_children",
f"/v1/api/workstreams/{coord.id}/close_all_children",
json={},
headers=headers,
)
+90 -41
View File
@@ -6,7 +6,7 @@ real in-process components:
1. Create + list + detail round-trip via the Starlette TestClient.
2. CoordinatorClient against a MockTransport "server node" stub.
3. list_children storage read flow (kind filtering, parent scoping).
4. Lazy rehydration via GET /v1/api/coordinator/{ws_id}.
4. Lazy rehydration via GET /v1/api/workstreams/{ws_id}.
Intentionally no real LLM infrastructure session factories return
MagicMock-backed stubs. All four tests run in < 2 s total.
@@ -26,18 +26,45 @@ from starlette.middleware.base import BaseHTTPMiddleware
from starlette.routing import Route
from starlette.testclient import TestClient
from turnstone.console.coordinator import CoordinatorManager
from turnstone.console.collector import ClusterCollector
from turnstone.console.coordinator_adapter import CoordinatorAdapter
from turnstone.console.coordinator_client import CoordinatorClient
from turnstone.console.coordinator_ui import ConsoleCoordinatorUI
from turnstone.console.server import (
coordinator_close,
coordinator_create,
coordinator_detail,
coordinator_list,
_audit_close_coordinator,
_audit_coordinator_create,
_coord_create_build_kwargs,
_coord_create_post_install,
_coord_create_validate_request,
_require_admin_coordinator,
_require_coord_mgr,
)
from turnstone.core.auth import AuthResult
from turnstone.core.session_manager import SessionManager
from turnstone.core.session_routes import (
SessionEndpointConfig,
make_close_handler,
make_create_handler,
make_detail_handler,
make_list_handler,
)
from turnstone.core.storage._sqlite import SQLiteBackend
# Per-kind config the lifted handler factories capture by closure.
_coord_endpoint_config = SessionEndpointConfig(
permission_gate=_require_admin_coordinator,
manager_lookup=_require_coord_mgr,
tenant_check=None,
not_found_label="coordinator not found",
audit_action_prefix="coordinator",
create_supports_attachments=True,
create_supports_user_id_override=False,
create_validate_request=_coord_create_validate_request,
create_build_kwargs=_coord_create_build_kwargs,
create_post_install=_coord_create_post_install,
)
# ---------------------------------------------------------------------------
# Shared auth-injection middleware (mirrors test_coordinator_endpoints.py)
# ---------------------------------------------------------------------------
@@ -81,8 +108,8 @@ def _fake_registry() -> MagicMock:
return reg
def _build_mgr(storage: SQLiteBackend) -> CoordinatorManager:
"""Build a CoordinatorManager backed by stub factories."""
def _build_mgr(storage: SQLiteBackend) -> SessionManager:
"""Build a SessionManager(CoordinatorAdapter) backed by stub factories."""
def _sf(ui, model_alias=None, ws_id=None, **kw):
s = MagicMock()
@@ -90,18 +117,26 @@ def _build_mgr(storage: SQLiteBackend) -> CoordinatorManager:
s.send.return_value = None
return s
return CoordinatorManager(
adapter = CoordinatorAdapter(
collector=MagicMock(),
ui_factory=lambda ws: ConsoleCoordinatorUI(ws_id=ws.id, user_id=ws.user_id or ""),
session_factory=_sf,
ui_factory=lambda w, u: ConsoleCoordinatorUI(ws_id=w, user_id=u),
)
mgr = SessionManager(
adapter,
storage=storage,
max_active=5,
node_id=ClusterCollector.CONSOLE_PSEUDO_NODE_ID,
event_emitter=adapter,
)
adapter.attach(mgr)
return mgr
def _make_client(
storage: SQLiteBackend,
*,
coord_mgr: CoordinatorManager | None = None,
coord_mgr: SessionManager | None = None,
alias: str = "my-model",
registry: Any = None,
) -> TestClient:
@@ -109,25 +144,34 @@ def _make_client(
app = Starlette(
routes=[
Route(
"/v1/api/coordinator/new",
coordinator_create,
methods=["POST"],
),
Route("/v1/api/coordinator", coordinator_list, methods=["GET"]),
Route(
"/v1/api/coordinator/{ws_id}/close",
coordinator_close,
"/v1/api/workstreams/new",
make_create_handler(_coord_endpoint_config, audit_emit=_audit_coordinator_create),
methods=["POST"],
),
Route(
"/v1/api/coordinator/{ws_id}",
coordinator_detail,
"/v1/api/workstreams",
make_list_handler(_coord_endpoint_config),
methods=["GET"],
),
Route(
"/v1/api/workstreams/{ws_id}/close",
make_close_handler(
_coord_endpoint_config,
audit_emit=_audit_close_coordinator,
supports_close_reason=False,
),
methods=["POST"],
),
Route(
"/v1/api/workstreams/{ws_id}",
make_detail_handler(_coord_endpoint_config),
methods=["GET"],
),
],
middleware=[Middleware(_AuthMiddleware)],
)
app.state.coord_mgr = coord_mgr
app.state.coord_adapter = coord_mgr._adapter if coord_mgr is not None else None
app.state.config_store = _FakeConfigStore({"coordinator.model_alias": alias})
app.state.coord_registry = registry
app.state.coord_registry_error = "" if coord_mgr else "registry missing"
@@ -151,32 +195,33 @@ def test_create_list_detail_lifecycle(tmp_path):
# --- Create ---
resp = client.post(
"/v1/api/coordinator/new",
"/v1/api/workstreams/new",
json={"name": "e2e-coord"},
headers=_COORD_HEADERS,
)
assert resp.status_code == 201, resp.text
assert resp.status_code == 200, resp.text
body = resp.json()
ws_id = body["ws_id"]
assert ws_id
assert "e2e-coord" in body["name"]
# --- List: caller sees their own coordinator ---
resp = client.get("/v1/api/coordinator", headers=_COORD_HEADERS)
resp = client.get("/v1/api/workstreams", headers=_COORD_HEADERS)
assert resp.status_code == 200, resp.text
coordinators = resp.json()["coordinators"]
coordinators = resp.json()["workstreams"]
ids = {c["ws_id"] for c in coordinators}
assert ws_id in ids
# Coordinator created by a different user is invisible to our caller.
# Trusted-team visibility: every ``admin.coordinator`` caller sees
# every active coordinator regardless of owner.
mgr.create(user_id="other-user", name="not-mine")
resp = client.get("/v1/api/coordinator", headers=_COORD_HEADERS)
resp = client.get("/v1/api/workstreams", headers=_COORD_HEADERS)
assert resp.status_code == 200
names = {c["name"] for c in resp.json()["coordinators"]}
assert "not-mine" not in names
names = {c["name"] for c in resp.json()["workstreams"]}
assert "not-mine" in names
# --- Detail ---
resp = client.get(f"/v1/api/coordinator/{ws_id}", headers=_COORD_HEADERS)
resp = client.get(f"/v1/api/workstreams/{ws_id}", headers=_COORD_HEADERS)
assert resp.status_code == 200, resp.text
detail = resp.json()
assert detail["ws_id"] == ws_id
@@ -184,7 +229,7 @@ def test_create_list_detail_lifecycle(tmp_path):
assert detail["user_id"] == "user-1"
# --- Close ---
resp = client.post(f"/v1/api/coordinator/{ws_id}/close", headers=_COORD_HEADERS)
resp = client.post(f"/v1/api/workstreams/{ws_id}/close", headers=_COORD_HEADERS)
assert resp.status_code == 200
# Manager no longer tracks it after close.
@@ -272,9 +317,11 @@ def test_coordinator_client_spawn_close_delete(tmp_path):
assert close_result.get("status") in (200, "ok"), close_result
close_req = captured[0]
assert close_req.url.path == "/v1/api/route/workstreams/close"
# Path-keyed shape post-#422: ws_id rides in the URL.
assert close_req.url.path == "/v1/api/route/workstreams/child-99/close"
close_body = json.loads(close_req.content)
assert close_body["ws_id"] == "child-99"
# Body no longer carries ws_id — the path is authoritative.
assert "ws_id" not in close_body
# delete --------------------------------------------------------------
captured.clear()
@@ -380,7 +427,7 @@ def test_list_children_skill_filter(seeded_storage):
# ---------------------------------------------------------------------------
# Test 4 — Lazy rehydration via GET /v1/api/coordinator/{ws_id}
# Test 4 — Lazy rehydration via GET /v1/api/workstreams/{ws_id}
# ---------------------------------------------------------------------------
@@ -389,8 +436,8 @@ def test_lazy_rehydration_on_detail_get(tmp_path):
Sequence:
1. Pre-seed storage with a coordinator row (simulating a previous process).
2. Build a CoordinatorManager that doesn't know about it yet.
3. Hit GET /v1/api/coordinator/{ws_id} expect 200.
2. Build a SessionManager (coordinator kind) that doesn't know about it yet.
3. Hit GET /v1/api/workstreams/{ws_id} expect 200.
4. Manager now tracks the rehydrated session.
5. The response body carries the correct kind / user_id metadata.
"""
@@ -410,7 +457,7 @@ def test_lazy_rehydration_on_detail_get(tmp_path):
assert mgr.get("persisted-coord") is None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.get("/v1/api/coordinator/persisted-coord", headers=_COORD_HEADERS)
resp = client.get("/v1/api/workstreams/persisted-coord", headers=_COORD_HEADERS)
assert resp.status_code == 200, resp.text
body = resp.json()
@@ -421,15 +468,17 @@ def test_lazy_rehydration_on_detail_get(tmp_path):
# The endpoint triggers lazy rehydration — manager now tracks it.
assert mgr.get("persisted-coord") is not None
# Non-owner cannot reach the same endpoint (returns 404 — no existence leak).
# Trusted-team visibility: any admin.coordinator caller can read
# the coordinator's detail, regardless of ``user_id``.
resp_stranger = client.get(
"/v1/api/coordinator/persisted-coord",
"/v1/api/workstreams/persisted-coord",
headers={"X-Test-User": "stranger", "X-Test-Perms": "admin.coordinator"},
)
assert resp_stranger.status_code == 404
assert resp_stranger.status_code == 200
assert resp_stranger.json()["user_id"] == "user-1"
# A workstream with kind='interactive' is not reachable via the coordinator
# endpoint even when it exists in storage.
storage.register_workstream("interactive-ws", kind="interactive", user_id="user-1")
resp_int = client.get("/v1/api/coordinator/interactive-ws", headers=_COORD_HEADERS)
resp_int = client.get("/v1/api/workstreams/interactive-ws", headers=_COORD_HEADERS)
assert resp_int.status_code == 404
File diff suppressed because it is too large Load Diff
+44 -38
View File
@@ -25,6 +25,7 @@ from tests._coord_test_helpers import (
_build_mgr,
_fake_registry,
_FakeConfigStore,
_seed_children,
)
from turnstone.console.server import (
coordinator_restrict,
@@ -45,17 +46,17 @@ def _make_client(storage, *, coord_mgr, alias="my-model", registry=None) -> Test
app = Starlette(
routes=[
Route(
"/v1/api/coordinator/{ws_id}/trust",
"/v1/api/workstreams/{ws_id}/trust",
coordinator_trust,
methods=["POST"],
),
Route(
"/v1/api/coordinator/{ws_id}/restrict",
"/v1/api/workstreams/{ws_id}/restrict",
coordinator_restrict,
methods=["POST"],
),
Route(
"/v1/api/coordinator/{ws_id}/stop_cascade",
"/v1/api/workstreams/{ws_id}/stop_cascade",
coordinator_stop_cascade,
methods=["POST"],
),
@@ -63,6 +64,7 @@ def _make_client(storage, *, coord_mgr, alias="my-model", registry=None) -> Test
middleware=[Middleware(_AuthMiddleware)],
)
app.state.coord_mgr = coord_mgr
app.state.coord_adapter = coord_mgr._adapter if coord_mgr is not None else None
app.state.config_store = _FakeConfigStore({"coordinator.model_alias": alias})
app.state.coord_registry = registry
app.state.coord_registry_error = "" if coord_mgr else "registry missing"
@@ -119,7 +121,7 @@ def test_trust_toggle_requires_trust_send_permission(storage):
coord = mgr.create(user_id="user-1", name="coord-a")
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json={"send": True},
headers=_COORD_HEADERS,
)
@@ -134,7 +136,7 @@ def test_trust_toggle_flips_session_flag_and_audits(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json={"send": True},
headers=_TRUST_HEADERS,
)
@@ -166,23 +168,24 @@ def _service_token_client(
app = Starlette(
routes=[
Route(
"/v1/api/coordinator/{ws_id}/trust",
"/v1/api/workstreams/{ws_id}/trust",
coordinator_trust,
methods=["POST"],
),
Route(
"/v1/api/coordinator/{ws_id}/restrict",
"/v1/api/workstreams/{ws_id}/restrict",
coordinator_restrict,
methods=["POST"],
),
Route(
"/v1/api/coordinator/{ws_id}/stop_cascade",
"/v1/api/workstreams/{ws_id}/stop_cascade",
coordinator_stop_cascade,
methods=["POST"],
),
],
)
app.state.coord_mgr = coord_mgr
app.state.coord_adapter = coord_mgr._adapter if coord_mgr is not None else None
app.state.config_store = _FakeConfigStore({"coordinator.model_alias": "my-model"})
app.state.coord_registry = _fake_registry()
app.state.coord_registry_error = ""
@@ -220,7 +223,7 @@ def test_trust_toggle_service_token_cannot_bypass_permission(storage):
permissions=frozenset({"admin.coordinator"}),
)
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json={"send": True},
)
assert resp.status_code == 403
@@ -243,7 +246,7 @@ def test_trust_toggle_service_token_with_permission_succeeds(storage):
permissions=frozenset({"admin.coordinator", "coordinator.trust.send"}),
)
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json={"send": True},
)
assert resp.status_code == 200
@@ -266,7 +269,7 @@ def test_restrict_service_token_cannot_bypass_admin_coordinator(storage):
permissions=frozenset(), # no admin.coordinator
)
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": ["bash"]},
)
assert resp.status_code == 403
@@ -285,7 +288,7 @@ def test_stop_cascade_service_token_cannot_bypass_admin_coordinator(storage):
permissions=frozenset(),
)
resp = client.post(
f"/v1/api/coordinator/{coord.id}/stop_cascade",
f"/v1/api/workstreams/{coord.id}/stop_cascade",
json={},
)
assert resp.status_code == 403
@@ -297,7 +300,7 @@ def test_trust_toggle_rejects_non_bool(storage):
coord.session, _ = _make_session_mock()
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json={"send": "yes"},
headers=_TRUST_HEADERS,
)
@@ -317,7 +320,7 @@ def test_trust_toggle_rejects_non_object_body(storage):
# only care that none 500.
for body in ([], 42, "string"):
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json=body,
headers=_TRUST_HEADERS,
)
@@ -331,27 +334,30 @@ def test_restrict_rejects_non_object_body(storage):
coord.session, _ = _make_session_mock()
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json=[],
headers=_COORD_HEADERS,
)
assert resp.status_code == 400
def test_trust_toggle_tenant_404_on_foreign_coord(storage):
def test_trust_toggle_cluster_wide_access(storage):
# Trusted-team model: the trust toggle is gated on the scope
# permission, not on row-level ownership. A caller holding
# ``coordinator.trust.send`` may toggle any coord's trust state.
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-owner", name="coord-a")
coord.session, _ = _make_session_mock()
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json={"send": True},
headers={
"X-Test-User": "user-other",
"X-Test-Perms": "admin.coordinator,coordinator.trust.send",
},
)
assert resp.status_code == 404
assert resp.status_code == 200
def test_trust_toggle_404_when_session_not_loaded(storage):
@@ -363,7 +369,7 @@ def test_trust_toggle_404_when_session_not_loaded(storage):
coord.session = None # simulate a closed / lazy-rehydrate coord
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/trust",
f"/v1/api/workstreams/{coord.id}/trust",
json={"send": True},
headers=_TRUST_HEADERS,
)
@@ -466,7 +472,7 @@ def test_restrict_adds_to_revoked_tools_and_audits(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": ["spawn_workstream", "delete_workstream"]},
headers=_COORD_HEADERS,
)
@@ -488,12 +494,12 @@ def test_restrict_is_additive_across_calls(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": ["spawn_workstream"]},
headers=_COORD_HEADERS,
)
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": ["delete_workstream"]},
headers=_COORD_HEADERS,
)
@@ -513,7 +519,7 @@ def test_restrict_empty_revoke_is_noop_but_audits(storage):
coord.session, _state = _make_session_mock(revoked=frozenset({"spawn_workstream"}))
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": []},
headers=_COORD_HEADERS,
)
@@ -532,7 +538,7 @@ def test_restrict_rejects_non_list_body(storage):
coord.session, _ = _make_session_mock()
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": "spawn_workstream"},
headers=_COORD_HEADERS,
)
@@ -547,7 +553,7 @@ def test_restrict_rejects_oversize_list(storage):
coord.session, _ = _make_session_mock()
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": [f"tool_{i}" for i in range(500)]},
headers=_COORD_HEADERS,
)
@@ -560,7 +566,7 @@ def test_restrict_rejects_oversize_name(storage):
coord.session, _ = _make_session_mock()
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": ["x" * 1000]},
headers=_COORD_HEADERS,
)
@@ -573,7 +579,7 @@ def test_restrict_404_when_session_not_loaded(storage):
coord.session = None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/restrict",
f"/v1/api/workstreams/{coord.id}/restrict",
json={"revoke": ["bash"]},
headers=_COORD_HEADERS,
)
@@ -634,7 +640,7 @@ def test_prepare_tool_allows_non_revoked_tool():
def test_stop_cascade_cancels_coord_and_each_child(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
mgr.register_children(coord.id, ["child-1", "child-2", "child-3"])
_seed_children(mgr._adapter, coord.id, ["child-1", "child-2", "child-3"])
def _cancel(wid: str) -> dict:
if wid == "child-2":
@@ -648,7 +654,7 @@ def test_stop_cascade_cancels_coord_and_each_child(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/stop_cascade",
f"/v1/api/workstreams/{coord.id}/stop_cascade",
json={},
headers=_COORD_HEADERS,
)
@@ -683,7 +689,7 @@ def test_stop_cascade_routes_404_to_skipped_bucket(storage):
them apart."""
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
mgr.register_children(coord.id, ["stale-child"])
_seed_children(mgr._adapter, coord.id, ["stale-child"])
coord_client = MagicMock()
coord_client.cancel.return_value = {
@@ -695,7 +701,7 @@ def test_stop_cascade_routes_404_to_skipped_bucket(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/stop_cascade",
f"/v1/api/workstreams/{coord.id}/stop_cascade",
json={},
headers=_COORD_HEADERS,
)
@@ -714,7 +720,7 @@ def test_stop_cascade_empty_children_still_audits(storage):
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/stop_cascade",
f"/v1/api/workstreams/{coord.id}/stop_cascade",
json={},
headers=_COORD_HEADERS,
)
@@ -730,13 +736,13 @@ def test_stop_cascade_without_coord_client_marks_all_failed(storage):
the operator can investigate."""
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
mgr.register_children(coord.id, ["child-a", "child-b"])
_seed_children(mgr._adapter, coord.id, ["child-a", "child-b"])
coord.session = MagicMock()
coord.session._coord_client = None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/stop_cascade",
f"/v1/api/workstreams/{coord.id}/stop_cascade",
json={},
headers=_COORD_HEADERS,
)
@@ -753,7 +759,7 @@ def test_stop_cascade_404_when_session_not_loaded(storage):
coord.session = None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/stop_cascade",
f"/v1/api/workstreams/{coord.id}/stop_cascade",
json={},
headers=_COORD_HEADERS,
)
@@ -763,10 +769,10 @@ def test_stop_cascade_404_when_session_not_loaded(storage):
def test_children_snapshot_returns_copy_not_live_set(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
mgr.register_children(coord.id, ["a", "b", "c"])
snap = mgr.children_snapshot(coord.id)
_seed_children(mgr._adapter, coord.id, ["a", "b", "c"])
snap = mgr._adapter.children_snapshot(coord.id)
assert set(snap) == {"a", "b", "c"}
mgr.register_children(coord.id, ["d"])
_seed_children(mgr._adapter, coord.id, ["d"])
assert set(snap) == {"a", "b", "c"}
+489
View File
@@ -0,0 +1,489 @@
"""Unit tests for :class:`CoordinatorIdleObserver`.
Drives a fake :class:`SessionManager` that mirrors the real one's
``subscribe_to_state`` / ``get`` contract, plus a fake storage with the
``list_workstreams`` slice the observer queries.
"""
from __future__ import annotations
import contextlib
import threading
from typing import Any
from unittest.mock import MagicMock
import pytest
from turnstone.console.coordinator_idle_observer import CoordinatorIdleObserver
from turnstone.core.nudge_queue import NudgeQueue
from turnstone.core.workstream import WorkstreamKind, WorkstreamState
class _FakeRow:
"""SQLAlchemy-Row-like wrapper exposing ``_mapping``."""
def __init__(self, **kwargs: Any) -> None:
self._mapping = kwargs
class _FakeStorage:
def __init__(self) -> None:
self.children: list[dict[str, Any]] = []
self.list_calls: list[dict[str, Any]] = []
self.count_calls: list[dict[str, Any]] = []
self.list_raises: bool = False
self.count_raises: bool = False
def list_workstreams(
self,
node_id: str | None = None,
limit: int = 100,
*,
parent_ws_id: str | None = None,
kind: WorkstreamKind | str | None = None,
user_id: str | None = None,
) -> list[Any]:
self.list_calls.append(
{
"limit": limit,
"parent_ws_id": parent_ws_id,
"kind": kind,
"user_id": user_id,
}
)
if self.list_raises:
raise RuntimeError("storage forced failure")
return [_FakeRow(**c) for c in self.children]
def count_workstreams_by_state(
self,
*,
parent_ws_id: str | None = None,
user_id: str | None = None,
) -> dict[str, int]:
self.count_calls.append({"parent_ws_id": parent_ws_id, "user_id": user_id})
if self.count_raises:
raise RuntimeError("count forced failure")
counts: dict[str, int] = {}
for c in self.children:
counts[c["state"]] = counts.get(c["state"], 0) + 1
return counts
class _FakeSession:
def __init__(self) -> None:
self._nudge_queue = NudgeQueue()
self.messages: list[dict[str, Any]] = []
self._wake_source_tag: str = ""
self._metacog_state: dict[str, float] = {}
self._mem_cfg = MagicMock(nudge_cooldown=300)
def _visible_memory_count(self) -> int:
return 0
class _FakeWorkstream:
def __init__(
self,
ws_id: str = "ws-coord",
kind: WorkstreamKind = WorkstreamKind.COORDINATOR,
user_id: str = "u1",
) -> None:
self.id = ws_id
self.kind = kind
self.user_id = user_id
self.session: _FakeSession | None = _FakeSession()
class _FakeManager:
def __init__(self) -> None:
self._workstreams: dict[str, _FakeWorkstream] = {}
self._subscribers: list[Any] = []
self._lock = threading.Lock()
def add_ws(self, ws: _FakeWorkstream) -> None:
self._workstreams[ws.id] = ws
def remove_ws(self, ws_id: str) -> None:
self._workstreams.pop(ws_id, None)
def get(self, ws_id: str) -> _FakeWorkstream | None:
return self._workstreams.get(ws_id)
def subscribe_to_state(self, callback: Any) -> None:
with self._lock:
self._subscribers.append(callback)
def unsubscribe_from_state(self, callback: Any) -> None:
with self._lock, contextlib.suppress(ValueError):
self._subscribers.remove(callback)
def fire_state(self, ws_id: str, state: WorkstreamState) -> None:
with self._lock:
subs = list(self._subscribers)
for cb in subs:
with contextlib.suppress(Exception):
cb(ws_id, state)
@pytest.fixture
def coord_setup() -> tuple[_FakeManager, _FakeStorage, _FakeWorkstream]:
mgr = _FakeManager()
storage = _FakeStorage()
ws = _FakeWorkstream()
mgr.add_ws(ws)
return mgr, storage, ws
def _add_active_child(storage: _FakeStorage, **overrides: Any) -> None:
storage.children.append(
{
"ws_id": overrides.get("ws_id", "child-1"),
"name": overrides.get("name", "research"),
"state": overrides.get("state", "running"),
}
)
class TestEnqueueOnIdle:
def test_idle_with_active_children_enqueues(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage, ws_id="child-a", state="running")
_add_active_child(storage, ws_id="child-b", state="thinking")
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
snap = ws.session._nudge_queue.pending("any")
assert len(snap) == 1
nudge_type, text = snap[0]
assert nudge_type == "idle_children"
assert "child-a" in text
assert "child-b" in text
def test_idle_with_no_active_children_no_enqueue(self, coord_setup):
mgr, storage, ws = coord_setup
# storage.children is empty
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue) == 0
def test_idle_only_idle_state_children_no_enqueue(self, coord_setup):
mgr, storage, ws = coord_setup
# All children "idle" — terminal-from-coord-perspective; not active.
_add_active_child(storage, state="idle")
_add_active_child(storage, state="closed")
_add_active_child(storage, state="error")
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue) == 0
def test_non_idle_state_no_enqueue(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
for state in (
WorkstreamState.RUNNING,
WorkstreamState.THINKING,
WorkstreamState.ATTENTION,
WorkstreamState.ERROR,
):
mgr.fire_state(ws.id, state)
assert len(ws.session._nudge_queue) == 0
class TestKindFilter:
def test_interactive_workstream_skipped(self):
mgr = _FakeManager()
storage = _FakeStorage()
_add_active_child(storage)
ws = _FakeWorkstream(kind=WorkstreamKind.INTERACTIVE)
mgr.add_ws(ws)
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# Observer ignored the non-coord workstream entirely.
assert len(ws.session._nudge_queue) == 0
# Storage was NOT queried — kind check happens before list_workstreams.
assert storage.list_calls == []
class TestWaitForWorkstreamSkip:
def test_skips_when_last_assistant_used_wait(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
ws.session.messages = [
{"role": "user", "content": "kick off"},
{
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": "call-1",
"function": {"name": "wait_for_workstream", "arguments": "{}"},
}
],
},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# Don't pile on — model is already using the right tool.
assert len(ws.session._nudge_queue) == 0
def test_fires_when_last_assistant_used_different_tool(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
ws.session.messages = [
{"role": "user", "content": "go"},
{
"role": "assistant",
"content": None,
"tool_calls": [
{"id": "call-1", "function": {"name": "spawn_workstream", "arguments": "{}"}}
],
},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue) == 1
class TestHardCap:
def test_hard_cap_blocks_after_n_fires(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
# Bypass cooldown for this test: each call burns a per-type slot
# in ``_metacog_state`` so we need to clear it between fires.
for _ in range(3):
ws.session._metacog_state.clear()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# Cap = 3 fires. Even with cooldown bypassed, the 4th doesn't fire.
ws.session._metacog_state.clear()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# We enqueued 3 entries total; cap blocked the 4th.
snap = ws.session._nudge_queue.pending("any")
assert len(snap) == 3
def test_cap_resets_when_state_leaves_idle_without_wake(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
# Burn the cap.
for _ in range(3):
ws.session._metacog_state.clear()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue.pending("any")) == 3
# Drain the queue (simulate the watcher delivering them).
ws.session._nudge_queue.drain({"any"})
# Real (non-wake) leave-IDLE: tag is empty. Cap resets.
ws.session._wake_source_tag = ""
mgr.fire_state(ws.id, WorkstreamState.RUNNING)
# New IDLE — cap is fresh, fires again.
ws.session._metacog_state.clear()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue.pending("any")) == 1
def test_cap_does_not_reset_during_wake_driven_exit(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
# Burn the cap.
for _ in range(3):
ws.session._metacog_state.clear()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
ws.session._nudge_queue.drain({"any"})
# Wake-driven leave-IDLE: tag is set during the wake send.
ws.session._wake_source_tag = "system_nudge"
mgr.fire_state(ws.id, WorkstreamState.RUNNING)
ws.session._wake_source_tag = "" # tag cleared at end of wake send
# Cap should NOT have reset — re-IDLE shouldn't fire.
ws.session._metacog_state.clear()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue.pending("any")) == 0
class TestCooldown:
def test_cooldown_blocks_within_window(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue.pending("any")) == 1
# Drain so the queue isn't the gate.
ws.session._nudge_queue.drain({"any"})
# Second fire within the cooldown window → should_nudge returns False.
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue.pending("any")) == 0
class TestStorageFailure:
def test_storage_exception_is_swallowed(self, coord_setup):
mgr, storage, ws = coord_setup
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
storage.list_raises = True
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
# Must not raise / propagate.
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue) == 0
class TestValidUntilPredicate:
def test_predicate_drops_when_children_finish_before_drain(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage, ws_id="child-a", state="running")
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue) == 1
# Children now complete (storage shows none active).
storage.children.clear()
# Drain at the user seam — predicate re-queries, finds 0 active,
# drops the entry without delivering.
from turnstone.core.nudge_queue import USER_DRAIN
delivered = ws.session._nudge_queue.drain(USER_DRAIN)
assert delivered == []
assert len(ws.session._nudge_queue) == 0
def test_predicate_delivers_when_children_still_active(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage, ws_id="child-a", state="running")
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# Children still active → predicate returns True → entry delivers.
from turnstone.core.nudge_queue import USER_DRAIN
delivered = ws.session._nudge_queue.drain(USER_DRAIN)
assert len(delivered) == 1
assert delivered[0][0] == "idle_children"
def test_predicate_drops_on_storage_failure(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# Storage failure at drain time. Predicate treats raises as
# "no longer valid" (drop) — see NudgeQueue.drain's predicate
# exception handling.
storage.count_raises = True
from turnstone.core.nudge_queue import USER_DRAIN
delivered = ws.session._nudge_queue.drain(USER_DRAIN)
assert delivered == []
class TestLifecycle:
def test_start_idempotent(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
observer.start() # no-op
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# Double-subscribe would have produced 2 entries.
assert len(ws.session._nudge_queue.pending("any")) == 1
def test_shutdown_unsubscribes(self, coord_setup):
mgr, storage, ws = coord_setup
_add_active_child(storage)
# ≥2 messages so should_nudge's message_count > 1 gate clears.
ws.session.messages = [
{"role": "user", "content": "go"},
{"role": "assistant", "content": "ok"},
]
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
observer.shutdown()
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert len(ws.session._nudge_queue) == 0
def test_shutdown_idempotent(self, coord_setup):
mgr, _storage, _ws = coord_setup
observer = CoordinatorIdleObserver(mgr, _storage)
observer.start()
observer.shutdown()
observer.shutdown() # no error
-955
View File
@@ -1,955 +0,0 @@
"""Tests for :class:`turnstone.console.coordinator.CoordinatorManager`.
Covers the lifecycle semantics without standing up a full ModelRegistry
or ChatSession: a stub session factory returns a MagicMock-backed
session so tests stay fast.
"""
from __future__ import annotations
from typing import Any
from unittest.mock import MagicMock
import pytest
from turnstone.console.coordinator import CoordinatorManager
from turnstone.console.coordinator_ui import ConsoleCoordinatorUI
from turnstone.core.storage._sqlite import SQLiteBackend
from turnstone.core.workstream import WorkstreamState
@pytest.fixture
def storage(tmp_path):
return SQLiteBackend(str(tmp_path / "coord.db"))
@pytest.fixture
def built_mgr(storage):
"""Build a CoordinatorManager with a stub session factory.
The factory records its calls and returns a MagicMock-backed
session so ``_spawn_worker`` can run without hitting real LLM
infrastructure.
"""
call_log: list[dict] = []
def _session_factory(ui, model_alias=None, ws_id=None, **kwargs):
call_log.append(
{
"ui": ui,
"model_alias": model_alias,
"ws_id": ws_id,
**kwargs,
}
)
mock_session = MagicMock()
mock_session.ws_id = ws_id
# send() is the worker thread target; make it a fast no-op.
mock_session.send.return_value = None
return mock_session
def _ui_factory(ws_id, user_id):
return ConsoleCoordinatorUI(ws_id=ws_id, user_id=user_id)
mgr = CoordinatorManager(
session_factory=_session_factory,
ui_factory=_ui_factory,
storage=storage,
max_active=3,
)
return mgr, call_log, storage
# ---------------------------------------------------------------------------
# create
# ---------------------------------------------------------------------------
def test_create_registers_row_with_coordinator_kind(built_mgr):
mgr, _calls, storage = built_mgr
ws = mgr.create(user_id="user-1", name="c1")
row = storage.get_workstream(ws.id)
assert row is not None
assert row["kind"] == "coordinator"
assert row["user_id"] == "user-1"
assert row["node_id"] == "console"
assert row["parent_ws_id"] is None
def test_create_passes_kind_to_factory(built_mgr):
mgr, calls, _s = built_mgr
mgr.create(user_id="user-1")
assert calls[-1]["kind"] == "coordinator"
assert calls[-1]["parent_ws_id"] is None
def test_create_dispatches_initial_message(built_mgr):
import time
mgr, _calls, _s = built_mgr
ws = mgr.create(user_id="user-1", initial_message="hello")
# Give the worker a brief window to run send() on the mock.
for _ in range(20):
if ws.session.send.called:
break
time.sleep(0.01)
ws.session.send.assert_called_once_with("hello")
def test_create_no_initial_message_skips_worker(built_mgr):
mgr, _calls, _s = built_mgr
ws = mgr.create(user_id="user-1")
assert ws.session.send.call_count == 0
# ---------------------------------------------------------------------------
# max_active + eviction
# ---------------------------------------------------------------------------
def test_max_active_enforced_evicts_idle(built_mgr):
mgr, _calls, _s = built_mgr
ws_a = mgr.create(user_id="u1")
ws_b = mgr.create(user_id="u2")
ws_c = mgr.create(user_id="u3")
# All three at capacity. The next create should evict the oldest
# IDLE — ws_a has the oldest last_active.
ws_d = mgr.create(user_id="u4")
# ws_a got evicted from the dict; b/c/d are still present.
assert mgr.get(ws_a.id) is None
for w in (ws_b, ws_c, ws_d):
assert mgr.get(w.id) is not None
def test_max_active_raises_when_all_non_idle(built_mgr):
mgr, _calls, _s = built_mgr
ws_a = mgr.create(user_id="u1")
ws_b = mgr.create(user_id="u2")
ws_c = mgr.create(user_id="u3")
# Force all into a non-idle state so no eviction candidate exists.
for w in (ws_a, ws_b, ws_c):
w.state = WorkstreamState.RUNNING
with pytest.raises(RuntimeError) as exc_info:
mgr.create(user_id="u4")
assert "slots are active" in str(exc_info.value)
def test_rollback_on_factory_failure(storage):
"""If the session factory raises, the slot + persisted row are rolled back."""
def _factory_explodes(*args, **kwargs):
raise RuntimeError("session construction failed")
mgr = CoordinatorManager(
session_factory=_factory_explodes,
ui_factory=lambda w, u: ConsoleCoordinatorUI(ws_id=w, user_id=u),
storage=storage,
max_active=3,
)
with pytest.raises(RuntimeError):
mgr.create(user_id="u1")
# No leaked in-memory workstream.
assert mgr.list_all() == []
# ---------------------------------------------------------------------------
# send / cancel / close
# ---------------------------------------------------------------------------
def test_send_returns_false_when_not_loaded(built_mgr):
mgr, _calls, _s = built_mgr
assert mgr.send("nonexistent", "hello") is False
def test_send_returns_false_on_queue_full_without_spawning_duplicate(storage):
"""If queue_message raises queue.Full, _spawn_worker must NOT fall
through and start a second concurrent worker on the same ChatSession
that would corrupt history / cursors / approvals. Instead, send()
returns False so the endpoint can surface 429."""
import queue
import threading
entered = threading.Event()
block = threading.Event()
def _slow_send(msg):
entered.set()
block.wait(timeout=5.0)
def _session_factory(ui, model_alias=None, ws_id=None, **kwargs):
sess = MagicMock()
sess.send.side_effect = _slow_send
sess.queue_message.side_effect = queue.Full()
return sess
mgr = CoordinatorManager(
session_factory=_session_factory,
ui_factory=lambda w, u: ConsoleCoordinatorUI(ws_id=w, user_id=u),
storage=storage,
max_active=3,
)
ws = mgr.create(user_id="u1", initial_message="first")
try:
assert entered.wait(timeout=2.0), "worker didn't start"
original_thread = ws.worker_thread
assert mgr.send(ws.id, "second") is False
# Must NOT have replaced worker_thread with a fresh second worker.
assert ws.worker_thread is original_thread
finally:
block.set()
if ws.worker_thread:
ws.worker_thread.join(timeout=2.0)
def test_send_enqueues_on_live_worker(storage):
"""When a worker thread is already processing, send() routes through
queue_message instead of spawning a duplicate worker."""
import threading
import time
entered = threading.Event()
block = threading.Event()
def _slow_send(msg):
entered.set()
block.wait(timeout=5.0)
def _session_factory(ui, model_alias=None, ws_id=None, **kwargs):
sess = MagicMock()
sess.send.side_effect = _slow_send
return sess
mgr = CoordinatorManager(
session_factory=_session_factory,
ui_factory=lambda w, u: ConsoleCoordinatorUI(ws_id=w, user_id=u),
storage=storage,
max_active=3,
)
ws = mgr.create(user_id="u1", initial_message="first")
try:
# Wait until the worker is actually inside session.send.
assert entered.wait(timeout=2.0), "worker didn't start"
# Now the worker is alive — mgr.send should route through queue_message.
for _ in range(20):
if ws.worker_thread and ws.worker_thread.is_alive():
break
time.sleep(0.01)
sent = mgr.send(ws.id, "second")
assert sent
ws.session.queue_message.assert_called_with("second")
finally:
block.set()
if ws.worker_thread:
ws.worker_thread.join(timeout=2.0)
def test_cancel_resolves_pending_approval(built_mgr):
mgr, _calls, _s = built_mgr
ws = mgr.create(user_id="u1")
assert ws.ui is not None
assert isinstance(ws.ui, ConsoleCoordinatorUI)
# Put ui into a pending-approval state.
ws.ui._pending_approval = {"type": "approve_request", "items": []}
ws.ui._approval_event.clear()
assert mgr.cancel(ws.id) is True
# resolve_approval should have been called with approved=False.
assert ws.ui._approval_event.is_set()
assert ws.ui._approval_result == (False, "cancelled")
def test_cancel_unblocks_worker_blocked_on_approval(built_mgr):
"""Cancel fires while a worker thread is blocked inside
ui.approve_tools() waiting on _approval_event. The worker must
unblock with approved=False and return."""
import threading
import time
mgr, _calls, _s = built_mgr
ws = mgr.create(user_id="u1")
ui = ws.ui
assert isinstance(ui, ConsoleCoordinatorUI)
# Simulate the session worker entering approve_tools. We call it
# directly on its own thread so the test can observe the unblock.
result_holder: list[tuple[bool, str | None]] = []
def _worker() -> None:
outcome = ui.approve_tools(
[
{
"call_id": "c1",
"func_name": "spawn_workstream",
"approval_label": "spawn_workstream",
"needs_approval": True,
}
]
)
result_holder.append(outcome)
t = threading.Thread(target=_worker, daemon=True)
t.start()
# Give the worker time to enter the approval wait.
for _ in range(50):
if ui._pending_approval is not None:
break
time.sleep(0.01)
assert ui._pending_approval is not None, "worker didn't reach approve_tools"
# Cancel fires — worker should unblock with approved=False.
assert mgr.cancel(ws.id) is True
t.join(timeout=2.0)
assert not t.is_alive()
assert result_holder == [(False, "cancelled")]
def test_close_removes_and_updates_state(built_mgr):
mgr, _calls, storage = built_mgr
ws = mgr.create(user_id="u1")
# Extract side-effectful call from the assert expression so
# python -O (which strips asserts) can't drop the close().
closed = mgr.close(ws.id)
assert closed is True
assert mgr.get(ws.id) is None
row = storage.get_workstream(ws.id)
assert row["state"] == "closed"
# ---------------------------------------------------------------------------
# list_for_user + list_all
# ---------------------------------------------------------------------------
def test_list_for_user_filters_by_owner(built_mgr):
mgr, _calls, _s = built_mgr
a = mgr.create(user_id="user-1")
b = mgr.create(user_id="user-1")
mgr.create(user_id="user-2") # non-owner — existence matters, value doesn't
user1_rows = mgr.list_for_user("user-1")
ids = {r.id for r in user1_rows}
assert ids == {a.id, b.id}
def test_list_all_returns_every_loaded(built_mgr):
mgr, _calls, _s = built_mgr
mgr.create(user_id="u1")
mgr.create(user_id="u2")
assert len(mgr.list_all()) == 2
# ---------------------------------------------------------------------------
# Lazy rehydration
# ---------------------------------------------------------------------------
def test_open_rehydrates_from_storage(built_mgr):
mgr, _calls, storage = built_mgr
# Simulate a coordinator persisted from a previous console process.
storage.register_workstream(
"coord-persisted",
node_id="console",
user_id="user-1",
kind="coordinator",
)
# Initially not loaded in memory.
assert mgr.get("coord-persisted") is None
ws = mgr.open("coord-persisted", "user-1")
assert ws is not None
assert ws.kind == "coordinator"
assert ws.user_id == "user-1"
# Now tracked.
assert mgr.get("coord-persisted") is not None
def test_open_rejects_non_coordinator_kind(built_mgr):
mgr, _calls, storage = built_mgr
storage.register_workstream("interactive-ws", kind="interactive", user_id="user-1")
# open() has side effects (factory call, slot reservation); keep it
# out of the assert expression so python -O can't strip it.
opened = mgr.open("interactive-ws", "user-1")
assert opened is None
def test_open_enforces_ownership(built_mgr):
mgr, _calls, storage = built_mgr
storage.register_workstream("coord-x", kind="coordinator", user_id="owner")
# Non-owner gets None.
stranger_ws = mgr.open("coord-x", "stranger")
assert stranger_ws is None
# Owner gets the row.
owner_ws = mgr.open("coord-x", "owner")
assert owner_ws is not None
def test_open_admin_ignores_ownership(built_mgr):
mgr, _calls, storage = built_mgr
storage.register_workstream("coord-x", kind="coordinator", user_id="owner")
ws = mgr.open_admin("coord-x")
assert ws is not None
def test_open_resurrects_closed_coordinator(built_mgr):
"""A coordinator that was closed (state='closed' in storage) IS now
resurrectable via open(). Restore is an explicit user action via
the Saved Coordinators landing UI; ``_reserve_and_install_locked``
still enforces ``max_active`` (evicts an idle peer or 429s). The
old "URL revisit silently undoes Close" safety lives in the slot
accounting now, not in a flat refusal at the open path."""
mgr, _calls, storage = built_mgr
ws = mgr.create(user_id="u1")
mgr.close(ws.id)
assert storage.get_workstream(ws.id)["state"] == "closed"
reopened = mgr.open(ws.id, "u1")
assert reopened is not None
assert reopened.id == ws.id
# Re-loaded into memory.
assert mgr.get(ws.id) is reopened
# Admin path also resurrects.
mgr.close(ws.id)
assert mgr.open_admin(ws.id) is not None
def test_open_refuses_deleted_coordinator(built_mgr):
"""A coordinator marked state='deleted' is a tombstone — open() must
refuse to resurrect even though closed-state is now resurrectable."""
mgr, _calls, storage = built_mgr
ws = mgr.create(user_id="u1")
mgr.close(ws.id)
storage.update_workstream_state(ws.id, "deleted")
user_open = mgr.open(ws.id, "u1")
assert user_open is None
admin_open = mgr.open_admin(ws.id)
assert admin_open is None
def test_open_refuses_empty_owner_for_non_admin(built_mgr):
"""Empty-owner rows (orphan / pre-002 migrated) must not be
rehydrated by non-admin callers would consume a max_active slot
and let any user evict another tenant's IDLE coordinator."""
mgr, _calls, storage = built_mgr
storage.register_workstream("coord-orphan", kind="coordinator", user_id=None)
# Non-admin caller — empty owner must NOT short-circuit the gate.
assert mgr.open("coord-orphan", "any-user") is None
# Admin path can still rehydrate (e.g. cleanup tooling).
assert mgr.open_admin("coord-orphan") is not None
def test_open_returns_existing_when_loaded(built_mgr):
mgr, _calls, _s = built_mgr
ws1 = mgr.create(user_id="u1")
ws2 = mgr.open(ws1.id, "u1")
assert ws2 is ws1
# ---------------------------------------------------------------------------
# Concurrency regressions — blockers 1 & 2 from review
# ---------------------------------------------------------------------------
def test_concurrent_open_for_same_ws_id_constructs_one_session(storage):
"""Two threads calling open() for the same persisted-but-unloaded
ws_id must not each spin up a session. Per-ws_id serialization
ensures the second thread picks up the first thread's session."""
import threading
import time
construct_count = {"n": 0}
construct_lock = threading.Lock()
first_in = threading.Event()
release_first = threading.Event()
def _slow_factory(ui, model_alias=None, ws_id=None, **kwargs):
with construct_lock:
construct_count["n"] += 1
my_idx = construct_count["n"]
if my_idx == 1:
first_in.set()
# Block so the second thread can race past the storage read.
release_first.wait(timeout=5.0)
sess = MagicMock()
sess.ws_id = ws_id
sess.send.return_value = None
return sess
mgr = CoordinatorManager(
session_factory=_slow_factory,
ui_factory=lambda w, u: ConsoleCoordinatorUI(ws_id=w, user_id=u),
storage=storage,
max_active=5,
)
storage.register_workstream(
"coord-shared",
node_id="console",
user_id="user-1",
kind="coordinator",
)
results: list[Any] = [None, None]
def _open_one(idx: int) -> None:
results[idx] = mgr.open("coord-shared", "user-1")
t1 = threading.Thread(target=_open_one, args=(0,))
t2 = threading.Thread(target=_open_one, args=(1,))
t1.start()
assert first_in.wait(timeout=2.0), "first thread didn't enter factory"
t2.start()
# Give t2 a chance to reach the per-ws lock and block.
time.sleep(0.1)
release_first.set()
t1.join(timeout=5.0)
t2.join(timeout=5.0)
assert construct_count["n"] == 1, (
f"expected exactly 1 session construction, got {construct_count['n']}"
)
assert results[0] is not None
assert results[1] is not None
# Both threads must see the same installed Workstream instance.
assert results[0] is results[1]
# Manager tracks exactly one entry.
assert len(mgr.list_all()) == 1
def test_concurrent_create_respects_max_active(storage):
"""max_active + 2 concurrent creates → exactly max_active succeed
and the overflow raises RuntimeError. Regression for the
check-then-install gap that previously let all creates pass the gate."""
import threading
slow_entered = threading.Event()
release = threading.Event()
def _slow_factory(ui, model_alias=None, ws_id=None, **kwargs):
# Block after construction to widen the race window between
# slot reservation and final install. Only the first N reach
# here — the rest must trip on the capacity gate earlier.
slow_entered.set()
release.wait(timeout=5.0)
sess = MagicMock()
sess.send.return_value = None
return sess
max_active = 3
mgr = CoordinatorManager(
session_factory=_slow_factory,
ui_factory=lambda w, u: ConsoleCoordinatorUI(ws_id=w, user_id=u),
storage=storage,
max_active=max_active,
)
successes: list[bool] = []
failures: list[Exception] = []
successes_lock = threading.Lock()
def _create_one(user_suffix: int) -> None:
try:
mgr.create(user_id=f"u{user_suffix}")
with successes_lock:
successes.append(True)
except RuntimeError as exc:
with successes_lock:
failures.append(exc)
threads = [threading.Thread(target=_create_one, args=(i,)) for i in range(max_active + 2)]
for t in threads:
t.start()
# Wait until at least one creation is blocked inside the factory.
assert slow_entered.wait(timeout=2.0)
release.set()
for t in threads:
t.join(timeout=5.0)
assert len(successes) == max_active, f"expected {max_active} successes, got {len(successes)}"
assert len(failures) == 2
for exc in failures:
assert "slots are active" in str(exc)
assert len(mgr.list_all()) == max_active
# ---------------------------------------------------------------------------
# Cross-tenant leak — blocker 3 from review
# ---------------------------------------------------------------------------
def test_list_for_user_excludes_empty_owner_rows(built_mgr):
"""A coordinator whose user_id is empty (system-created, migration
artifact, or lazily rehydrated from a NULL owner) must NOT appear
in list_for_user() output for other callers doing so would leak
ws_id + name + state across tenants."""
mgr, _calls, storage = built_mgr
# Real user's coordinator.
owned = mgr.create(user_id="alice")
# Simulate a rogue empty-owner session by creating one with
# user_id="" directly. Matches what a rehydrate of a NULL-owner
# row would produce, or a system-created coordinator.
empty_owner = mgr.create(user_id="")
rows = mgr.list_for_user("alice")
ids = {ws.id for ws in rows}
assert owned.id in ids
assert empty_owner.id not in ids, (
"list_for_user must not expose empty-owner coordinators to other callers"
)
# ---------------------------------------------------------------------------
# Phase 3 — child-event fan-out
# ---------------------------------------------------------------------------
def _seed_child_row(storage, *, parent_ws_id: str, ws_id: str, state: str = "idle") -> None:
storage.register_workstream(
ws_id,
node_id="node-a",
user_id="user-1",
name=f"c-{ws_id[:4]}",
kind="interactive",
parent_ws_id=parent_ws_id,
)
if state != "idle":
storage.update_workstream_state(ws_id, state)
def _drain(listener, *, wait: float = 0.5):
"""Drain a ConsoleCoordinatorUI listener queue with a short timeout."""
import queue as _q
items = []
try:
while True:
items.append(listener.get(timeout=wait))
except _q.Empty:
return items
def test_children_registry_bootstrapped_from_storage_on_create(built_mgr):
mgr, _calls, storage = built_mgr
ws = mgr.create(user_id="user-1")
# The registry starts empty — no children yet.
assert mgr._children.get(ws.id, set()) == set()
def test_children_registry_bootstrapped_from_storage_on_open(built_mgr):
mgr, _calls, storage = built_mgr
# Seed a persisted coordinator row + two children directly in storage
# so open() rehydrates them without create() being called.
coord_id = "a" * 32
storage.register_workstream(
coord_id,
node_id="console",
user_id="user-1",
name="persisted",
kind="coordinator",
parent_ws_id=None,
)
_seed_child_row(storage, parent_ws_id=coord_id, ws_id="b" * 32)
_seed_child_row(storage, parent_ws_id=coord_id, ws_id="c" * 32)
ws = mgr.open(coord_id, "user-1")
assert ws is not None
assert mgr._children[coord_id] == {"b" * 32, "c" * 32}
def test_dispatch_ws_created_fans_out_to_parent(built_mgr):
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
listener = ws.ui._register_listener()
mgr._dispatch_child_event(
{
"type": "ws_created",
"ws_id": "d" * 32,
"parent_ws_id": ws.id,
"node_id": "node-a",
"name": "new-child",
"title": "",
"user_id": "user-1",
}
)
events = _drain(listener)
child_created = [e for e in events if e.get("type") == "child_ws_created"]
assert len(child_created) == 1
assert child_created[0]["child_ws_id"] == "d" * 32
assert child_created[0]["parent_ws_id"] == ws.id
assert "d" * 32 in mgr._children[ws.id]
def test_dispatch_ws_created_ignores_unrelated_parent(built_mgr):
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
listener = ws.ui._register_listener()
# A ws_created for a parent this coordinator doesn't own.
mgr._dispatch_child_event(
{
"type": "ws_created",
"ws_id": "e" * 32,
"parent_ws_id": "f" * 32,
"node_id": "node-a",
"name": "stranger-child",
"title": "",
"user_id": "user-1",
}
)
events = _drain(listener, wait=0.1)
assert not any(e.get("type") == "child_ws_created" for e in events)
def test_dispatch_ws_created_cross_tenant_dropped(built_mgr):
"""A ws_created event whose user_id does not match the coordinator's
owner must NOT reach the coordinator's SSE stream — prevents the
cross-tenant info-leak via spoofed parent_ws_id (sec-1)."""
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="alice")
listener = ws.ui._register_listener()
# A mallory-owned workstream claiming alice's coordinator as parent.
mgr._dispatch_child_event(
{
"type": "ws_created",
"ws_id": "d" * 32,
"parent_ws_id": ws.id,
"node_id": "node-a",
"name": "spoofed-child",
"title": "",
"user_id": "mallory",
}
)
events = _drain(listener, wait=0.1)
assert not any(e.get("type") == "child_ws_created" for e in events)
# Registry must not have gained mallory's ws_id either.
assert "d" * 32 not in mgr._children.get(ws.id, set())
def test_dispatch_ws_created_empty_user_id_dropped(built_mgr):
"""An event with empty/missing user_id fails closed — we can't
prove tenancy, so we refuse to route it."""
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="alice")
listener = ws.ui._register_listener()
mgr._dispatch_child_event(
{
"type": "ws_created",
"ws_id": "d" * 32,
"parent_ws_id": ws.id,
"node_id": "node-a",
"name": "no-owner-child",
"title": "",
# user_id intentionally absent
}
)
events = _drain(listener, wait=0.1)
assert not any(e.get("type") == "child_ws_created" for e in events)
assert "d" * 32 not in mgr._children.get(ws.id, set())
def test_dispatch_cluster_state_fans_out_when_child_tracked(built_mgr):
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
child_id = "a" * 32
mgr._add_child(ws.id, child_id)
listener = ws.ui._register_listener()
mgr._dispatch_child_event(
{
"type": "cluster_state",
"ws_id": child_id,
"state": "running",
"tokens": 42,
"node_id": "node-a",
}
)
events = _drain(listener)
state_events = [e for e in events if e.get("type") == "child_ws_state"]
assert len(state_events) == 1
assert state_events[0]["child_ws_id"] == child_id
assert state_events[0]["state"] == "running"
assert state_events[0]["tokens"] == 42
def test_dispatch_ws_closed_fans_out(built_mgr):
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
child_id = "a" * 32
mgr._add_child(ws.id, child_id)
listener = ws.ui._register_listener()
mgr._dispatch_child_event({"type": "ws_closed", "ws_id": child_id, "reason": "closed"})
events = _drain(listener)
close_events = [e for e in events if e.get("type") == "child_ws_closed"]
assert len(close_events) == 1
assert close_events[0]["child_ws_id"] == child_id
assert close_events[0]["reason"] == "closed"
def test_dispatch_unrelated_state_ignored(built_mgr):
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
listener = ws.ui._register_listener()
# No _add_child called — ws_id is not in anyone's registry.
mgr._dispatch_child_event({"type": "cluster_state", "ws_id": "a" * 32, "state": "running"})
events = _drain(listener, wait=0.1)
assert not any(e.get("type", "").startswith("child_ws_") for e in events)
def test_shutdown_is_idempotent(built_mgr):
mgr, _calls, _storage = built_mgr
# No fanout started — shutdown must not raise.
mgr.shutdown()
mgr.shutdown()
# ---------------------------------------------------------------------------
# Phase 3 — review-pass-2 regression tests
# ---------------------------------------------------------------------------
def test_rebuild_registry_unions_with_concurrent_adds(built_mgr):
"""A ws_created event that arrives during open() must survive the
subsequent _rebuild_children_registry call the rebuild must UNION
its storage read with whatever the fan-out thread already added."""
mgr, _calls, storage = built_mgr
coord_id = "a" * 32
# Seed a persisted coordinator row — open() will rehydrate it.
storage.register_workstream(
coord_id,
node_id="console",
user_id="user-1",
name="persisted",
kind="coordinator",
parent_ws_id=None,
)
# Persist one child (will show up in rebuild's storage query).
_seed_child_row(storage, parent_ws_id=coord_id, ws_id="b" * 32)
# Simulate the fan-out thread pre-adding a different child_ws_id
# between the placeholder install and the rebuild call. Calling
# open() in this test runs synchronously, so we emulate the race
# by pre-populating the registry for the coord before open.
mgr._add_child(coord_id, "c" * 32)
ws = mgr.open(coord_id, "user-1")
assert ws is not None
# Both the persisted child (from rebuild) AND the pre-added one
# (from the simulated fan-out race) should be present.
assert "b" * 32 in mgr._children[coord_id]
assert "c" * 32 in mgr._children[coord_id]
def test_dispatch_ws_created_atomic_against_close(built_mgr):
"""Concurrent close() during a ws_created dispatch must not leave
the evicted coordinator's registry entry behind.
Regression for a race where the dispatch reads _active_coords
lock-free, close() runs (pops _children[parent]) between the
snapshot read and the _children_lock acquisition, then setdefault
resurrects the entry leaking the registry key forever."""
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
# Close the coordinator — _children[ws.id] gets popped and
# _active_coords loses the entry.
closed = mgr.close(ws.id)
assert closed
# A ws_created event still arriving for the now-closed parent
# must NOT resurrect the registry entry via setdefault.
mgr._dispatch_child_event(
{
"type": "ws_created",
"ws_id": "d" * 32,
"parent_ws_id": ws.id,
"node_id": "node-a",
"user_id": "user-1",
}
)
assert ws.id not in mgr._children
assert ws.id not in mgr._active_coords
def test_open_impl_eviction_clears_children_registry(built_mgr):
"""When _open_impl evicts an idle coordinator to make room, the
evicted coordinator's _children entry must be popped — matching
the create() eviction path."""
mgr, _calls, storage = built_mgr
# Fill the manager to capacity (max_active=3) with owned coords,
# then pre-seed a 4th as persisted-only so open() triggers eviction.
for i in range(3):
mgr.create(user_id=f"u{i}")
# Record which coord is idlest (oldest create) — it's the eviction
# candidate.
victim_id = mgr._order[0]
# Pre-seed the victim's _children to prove the pop works.
mgr._add_child(victim_id, "z" * 32)
assert victim_id in mgr._children
# Persist a 4th coord row so open() will rehydrate + evict.
fourth_id = "f" * 32
storage.register_workstream(
fourth_id,
node_id="console",
user_id="u3",
name="fourth",
kind="coordinator",
parent_ws_id=None,
)
# Force open() — it must evict the idle victim and clear its
# registry entry in the process.
result = mgr.open_admin(fourth_id)
assert result is not None
assert victim_id not in mgr._workstreams, "victim should have been evicted to make room"
assert victim_id not in mgr._children, (
"_open_impl must pop the evicted coordinator's _children entry "
"(mirrors create() eviction path)"
)
def test_child_to_coord_reverse_index_maintained(built_mgr):
"""_coord_for_child uses the reverse index for O(1) lookup. The
index must stay in sync with the forward set across add/close
paths this test pokes each maintenance point."""
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
# _add_child path — populates both sides.
assert mgr._add_child(ws.id, "child-1")
assert mgr._coord_for_child("child-1") == ws.id
assert mgr._child_to_coord["child-1"] == ws.id
# close() path — pops both sides.
mgr.close(ws.id)
assert mgr._coord_for_child("child-1") is None
assert "child-1" not in mgr._child_to_coord
def test_prime_children_from_snapshot(built_mgr):
"""start_child_event_fanout uses the collector snapshot to prime
the child registry so a just-opened coordinator sees already-live
children without waiting for the next ws_state event. Simulate
by calling the helper directly."""
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
snapshot = {
"nodes": [
{
"node_id": "node-a",
"workstreams": [
{"id": "child-1", "parent_ws_id": ws.id, "state": "running"},
{"id": "child-2", "parent_ws_id": ws.id, "state": "idle"},
# Unrelated — parent isn't a tracked coordinator.
{
"id": "foreign-1",
"parent_ws_id": "some-other-coord",
"state": "idle",
},
],
}
]
}
mgr._prime_children_from_snapshot(snapshot)
assert mgr._children[ws.id] == {"child-1", "child-2"}
assert mgr._coord_for_child("child-1") == ws.id
assert mgr._coord_for_child("child-2") == ws.id
# Foreign children with parents we don't track stay out of the
# registry — we only care about live coordinators.
assert mgr._coord_for_child("foreign-1") is None
def test_prime_children_from_empty_snapshot_noop(built_mgr):
"""No nodes → no state changes. Defensive: snapshot shape can
legitimately be missing the ``nodes`` key right after startup."""
mgr, _calls, _storage = built_mgr
ws = mgr.create(user_id="user-1")
mgr._prime_children_from_snapshot({})
mgr._prime_children_from_snapshot({"nodes": []})
assert mgr._children[ws.id] == set()
+290
View File
@@ -52,3 +52,293 @@ def test_uppercase_hex_rejected(client):
# Our ws_ids are lowercase hex; reject mixed/upper to avoid surprises.
resp = client.get("/coordinator/" + "A" * 32)
assert resp.status_code == 400
def test_coordinator_js_exposes_inline_approval_helpers():
"""Smoke guard for two layers of the coord chat frontend: the
children-tree inline approve/deny block (the original Chunk 3
landing) and the PR #447 tool-batch construct that replaced the
pinned approval dock for the coord-self surface. Both layers'
helper symbols must remain reachable in the served JS so a
refactor that accidentally renames or removes them surfaces here
instead of in production where the affected gates silently stop
rendering. Asserts string presence only no DOM parsing
since coord.js has no JS test framework today (per the plan's
testing notes)."""
from pathlib import Path
coord_js = Path(__file__).resolve().parent.parent / (
"turnstone/console/static/coordinator/coordinator.js"
)
body = coord_js.read_text(encoding="utf-8")
# Approval-block rendering helpers
assert "function renderApprovalBlock" in body
assert "function _maxSeverityItem" in body
assert "function _renderSubItem" in body
# The submit + 409 race-handling path
assert "function submitChildApproval" in body or "submitChildApproval(" in body
# The shared approve POST helper (parameterized for child ws_ids)
assert "function approveWorkstream" in body or "approveWorkstream(" in body
# The 409 stale-call_id retry path uses invalidateLiveBadge +
# scheduleLiveFetch (Stage 3 cleanup removed the urgent flag —
# cache invalidation makes the TTL gate fall through naturally).
assert "invalidateLiveBadge(targetWsId)" in body
# Server-side payload field — drift here means the JS reads stale keys
assert "pending_approval_detail" in body
# Reconnect parity (chunk 4): the SSE re-open handler must drop
# non-permanent entries from the live-badge cache so a stale
# pending_approval_detail (left from before the disconnect)
# can't render zombie approve/deny buttons on a row whose
# approval was resolved during the gap. The implementation
# iterates the cache and deletes only !permanent entries —
# asserting the literal helper call keeps a refactor back to
# _liveBadgeCacheClear() (which would re-pay 403s on every
# reconnect for denied ids) from sneaking in.
assert "_liveBadgeCacheDelete" in body
# Edge-case matrix sentinel labels — POLICY-BLOCKED renders when
# an item has error set + needs_approval=False (server-side
# tool policy already blocked the call); "(judge unavailable)"
# renders when no verdict (judge or heuristic) and no
# judge_pending. Refactors that drop either branch silently
# regress to a buttoned approve UI on the wrong state.
assert "POLICY-BLOCKED" in body
assert "judge unavailable" in body
# Critical-risk handling — bug-1 was that risk_level='critical'
# rendered as low because RISK_SEVERITY only mapped 'crit'.
# Both aliases must remain in the table so a 'critical' verdict
# ranks at 3 and renders with the .risk.crit pill.
assert "critical: 3" in body
# Child approves must round-trip through the routing proxy at
# /v1/api/route/workstreams/{ws_id}/approve — the bare
# /v1/api/workstreams/.../approve path only works for the
# coord-self ws_id (the coord lives on the console process).
# Children live on cluster nodes and 404 without the prefix.
assert "/v1/api/route/workstreams/" in body
# Late-arriving LLM judge verdicts — Stage 3 Step 5 promoted
# ``intent_verdict`` and ``approval_resolved`` to first-class
# cluster-bus event types, so the coord adapter dispatches them
# as ``child_ws_intent_verdict`` / ``child_ws_approval_resolved``
# on the parent's SSE stream. The browser handlers write
# directly to liveBadgeCache (bypassing scheduleLiveFetch's
# visibility gate cleanly) so off-screen rows pick up verdicts
# without polling. Replaced the old ``_judgePollTick`` 90-second
# global poll loop and its visibility-gate-bypass workaround.
assert "handleChildIntentVerdict" in body
assert "handleChildApprovalResolved" in body
assert "child_ws_intent_verdict" in body
assert "child_ws_approval_resolved" in body
# Reload parity for the coord-self approval gate: init() must
# consume the authoritative GET /workstreams snapshot's
# pending_approval_detail so a freshly opened tab can render
# Approve/Deny before SSE replay arrives.
assert "wsSnapshot.pending_approval_detail" in body
assert "appendToolBatch(pendingDetail.items" in body
# Tool-batch construct (PR #447) — the inline replacement for the
# pinned approval-dock pattern. These helpers carry the
# state-machine that pairs each tool call with its result and
# embeds the approval flow. Refactors that rename or drop them
# silently regress the entire coord-self approval surface — the
# most novel and risky behavior in the PR.
assert "function appendToolBatch" in body
assert "function _morphBatchResolved" in body
assert "function _resolveBatchAction" in body
assert "function _refreshBatchTier" in body
assert "function _refreshRowStatus" in body
# State modifiers driven by the upgrade-in-place path
# (--running orphan promoted to --pending or --auto when SSE
# arrives with the authoritative shape). Both class names must
# remain reachable from JS — dropping either breaks the reload
# state machine that PR #447's review pass surfaced.
assert "coord-tool-batch--running" in body
assert "coord-tool-batch--pending" in body
# History replay's outcome classifier — denied / errored tool
# turns must render with the correct batch state on reload, not
# the contradictory "✓ approved" pill that pre-fix showed for
# any prior denial. bug-1 / bug-3 from the second /review pass.
assert "Denied by user" in body
assert "callOutcomes" in body
# User-message attachment pills — both live send (coordSend) and
# history replay route through appendUserMessageWithAttachments.
# Renaming or dropping the helper would silently regress the
# attachment affordance to the pre-fix plain-text bubble, which
# would only surface in manual testing of an attached-file flow.
# The CSS class is the visual anchor (coordinator.css) — keeping
# both literals in the smoke layer covers JS↔CSS drift in either
# direction.
assert "function appendUserMessageWithAttachments" in body
assert "msg-user-attach" in body
# PR #487 — whitespace-only assistant content (Qwen3 with vLLM
# ``--reasoning-parser`` strips ``<think>…</think>`` and emits only
# ``"\n\n"`` as content before a tool call) must be skipped on
# history replay or the empty ``.msg.assistant`` card surfaces as
# a phantom row. The literal substring ``content.trim()`` is the
# single-line guard the rendering branch uses; a refactor that
# drops the trim() (e.g. simplifies to ``if (!content)``) silently
# regresses the phantom-card fix on the multi-node coord path.
# Mirrors ``test_app_js.py``'s same-shape pin on ``app.js``.
assert "content.trim()" in body
# PR #487 — coord history replay must render the assistant content
# card BEFORE the tool batch, not after, so DOM order matches the
# chronological order the model emitted (text → dispatch → results).
# Pre-fix the tool_calls branch sat at the role-agnostic top of the
# loop and rendered ahead of the assistant text that announced the
# batch, putting parallel fan-outs visually above their narrating
# message. The fix hoisted the synthesis into ``renderAssistantToolBatch``
# called from inside the assistant branch AFTER the content card —
# asserting the helper name lets a refactor that re-inlines or
# renames it surface here instead of via manual reload testing.
assert "function renderAssistantToolBatch" in body
assert "renderAssistantToolBatch(m)" in body
def test_coordinator_js_handle_child_state_no_longer_reads_sse_pending_approval_detail():
"""Stage 3 cleanup — ``pending_approval_detail`` is no longer
piggybacked on child_ws_state events. Approval items now arrive
via bulk fetch on the activity_state="approval" transition;
verdicts via the explicit ``child_ws_intent_verdict`` event class;
resolution via ``child_ws_approval_resolved``. A refactor that
re-introduces the piggyback would silently re-open the
duplicate-path race the dedicated event classes were added to
eliminate.
Structural assertions (regex against multi-line source) symbol-
presence alone wouldn't catch a guard that keeps the names but
inverts the comparison or drops the ``prev.live`` check. This
codebase has no JS test framework, so locking the guard's shape
here is the next-best thing to a behavioral test."""
import re
from pathlib import Path
coord_js = Path(__file__).resolve().parent.parent / (
"turnstone/console/static/coordinator/coordinator.js"
)
body = coord_js.read_text(encoding="utf-8")
# The piggyback read is gone from handleChildState. (The string
# may still appear elsewhere — e.g. handleChildIntentVerdict
# reading from cache, or comments — but never as ``ev.pending_approval_detail``.)
assert "ev.pending_approval_detail" not in body
# The pre-fix urgent-fetch on activity_state transitions is gone.
assert "enteredApproval" not in body
assert "leftApproval" not in body
# ``pendingApproval`` flag derivation must check BOTH state and
# activity_state. The worker thread can fire the state transition
# to "attention" before approve_tools updates activity_state, so
# checking only activity_state misses children that legitimately
# need approval. Pin the disjunction so the regression doesn't
# silently re-introduce.
assert re.search(
r'existing\.state\s*===\s*"attention"\s*\|\|\s*'
r'existing\.activity_state\s*===\s*"approval"',
body,
), (
"handleChildState must derive pendingApproval from "
"(state==='attention' || activity_state==='approval')"
)
# SSE-authoritative window constant is defined and used.
assert re.search(r"\bconst\s+SSE_AUTHORITATIVE_MS\s*=\s*\d+", body), (
"SSE_AUTHORITATIVE_MS constant must be defined as a numeric literal"
)
# SSE writers tag entries with sseUpdatedAt: Date.now() so the
# merge guard in flushLiveFetches preserves them against stale
# bulk-fetch responses. handleChildState only stamps when it
# AUTHORITATIVELY clears the detail (off-approval transition);
# writers that stamp unconditionally are intent_verdict (verdict
# stamp), approval_resolved (clear), and the optimistic-clear
# path in submitChildApproval. Pinning the literal Date.now()
# call keeps a refactor that drops the SSE-source tag entirely
# from sneaking in.
assert re.search(
r"sseUpdatedAt:\s*Date\.now\(\)",
body,
), "Critical SSE writers must stamp sseUpdatedAt: Date.now()"
# flushLiveFetches' merge guard structure: SSE-set pending_approval
# / _detail wins over a stale bulk-poll snapshot when (live) AND
# (prev exists) AND (prev.sseUpdatedAt set) AND (within window)
# AND (prev.live exists). Inverting the comparison or dropping
# any of these guards reopens the clobber bug.
merge_guard = re.search(
r"if\s*\(\s*live\s*&&\s*prev\s*&&\s*prev\.sseUpdatedAt\s*&&\s*"
r"now\s*-\s*prev\.sseUpdatedAt\s*<\s*SSE_AUTHORITATIVE_MS\s*&&\s*"
r"prev\.live\s*\)",
body,
)
assert merge_guard is not None, (
"flushLiveFetches merge guard must be the conjunction "
"(live && prev && prev.sseUpdatedAt && now - prev.sseUpdatedAt < "
"SSE_AUTHORITATIVE_MS && prev.live). An inverted comparison or "
"missing prev.live check would let a stale bulk-poll clobber a "
"fresh SSE-set approval."
)
# The merge body must preserve BOTH pending_approval and
# pending_approval_detail from prev — preserving only one would
# render a row with a phantom badge but no buttons (or vice versa).
merge_body = re.search(
r"mergedLive\s*=\s*Object\.assign\(\s*\{\}\s*,\s*live\s*,\s*\{"
r"[^}]*pending_approval:\s*prev\.live\.pending_approval[^}]*"
r"pending_approval_detail:\s*prev\.live\.pending_approval_detail",
body,
)
assert merge_body is not None, (
"Merge body must preserve both pending_approval AND "
"pending_approval_detail from prev.live — preserving only one "
"creates a half-rendered approval row."
)
# flushLiveFetches must forward sseUpdatedAt onto the new cache
# entry so the SSE-source tag survives the bulk-poll write back —
# without this, every bulk-poll resets the window and the next
# late-arriving poll silently clobbers.
assert re.search(
r"sseUpdatedAt:\s*prev\s*\?\s*prev\.sseUpdatedAt",
body,
), (
"flushLiveFetches must forward prev.sseUpdatedAt onto the new "
"cache entry (preserving the SSE-source window across bulk-poll "
"cycles) — without this, the second bulk-poll after an SSE "
"transition silently clobbers."
)
def test_coord_history_renders_user_interjection_advisory_after_tool_block():
"""Queued user messages spliced into the last tool-result envelope
of a batch (Seam 1) persist on the tool DB row as a wrapped
``<tool_output>`` envelope. ``decorate_history_messages`` extracts
the advisory back out and the wire layer projects it onto
``m.advisories``; the coord history loop must invoke the shared
``replayAdvisoriesAfterTool`` helper (defined in
``shared/utils.js``) so each ``user_interjection`` renders through
``appendUserMessageWithAttachments`` and the bubble looks identical
to a Seam 2/3 user row.
This test pins the call site so a refactor that drops the helper
invocation regresses the queued-during-batch replay shape silently.
Mirrors ``test_app_js.py``'s same-shape pin on interactive's
``replayHistory``."""
import re
from pathlib import Path
coord_js = Path(__file__).resolve().parent.parent / (
"turnstone/console/static/coordinator/coordinator.js"
)
body = coord_js.read_text(encoding="utf-8")
assert "replayAdvisoriesAfterTool(m.advisories" in body, (
"Coord history loop must invoke replayAdvisoriesAfterTool with "
"m.advisories so queued messages spliced into the tool envelope "
"render as user bubbles after the tool block."
)
# The renderer callback routes through appendUserMessageWithAttachments
# so the bubble matches a normal user-row replay.
assert re.search(
r"appendUserMessageWithAttachments\(\s*text",
body,
), (
"Coord history loop's renderer callback must route the extracted "
"advisory text through appendUserMessageWithAttachments so the "
"rendered bubble matches a normal user-row replay."
)
-395
View File
@@ -1,395 +0,0 @@
"""Tests for the coordinator ``/quota`` GET + POST endpoints.
Covers the admin partial-update surface for spawn-budget and
spawn-rate parallel to the /trust + /restrict shape in
``test_coordinator_governance.py``. Kept in its own file so PR B's
review surface stays tight.
"""
from __future__ import annotations
import json
from unittest.mock import MagicMock
import pytest
from starlette.applications import Starlette
from starlette.middleware import Middleware
from starlette.routing import Route
from starlette.testclient import TestClient
from tests._coord_test_helpers import (
_AuthMiddleware,
_build_mgr,
_fake_registry,
_FakeConfigStore,
)
from turnstone.console.server import (
coordinator_quota_get,
coordinator_quota_post,
)
from turnstone.core.spawn_quota import SpawnBudget, TokenBucket
from turnstone.core.storage._sqlite import SQLiteBackend
@pytest.fixture
def storage(tmp_path):
return SQLiteBackend(str(tmp_path / "coord.db"))
_COORD_HEADERS = {"X-Test-User": "user-1", "X-Test-Perms": "admin.coordinator"}
def _make_client(storage, *, coord_mgr, alias="my-model", registry=None) -> TestClient:
app = Starlette(
routes=[
Route(
"/v1/api/coordinator/{ws_id}/quota",
coordinator_quota_get,
methods=["GET"],
),
Route(
"/v1/api/coordinator/{ws_id}/quota",
coordinator_quota_post,
methods=["POST"],
),
],
middleware=[Middleware(_AuthMiddleware)],
)
app.state.coord_mgr = coord_mgr
app.state.config_store = _FakeConfigStore({"coordinator.model_alias": alias})
app.state.coord_registry = registry
app.state.coord_registry_error = "" if coord_mgr else "registry missing"
app.state.auth_storage = storage
app.state.jwt_secret = "x" * 64
return TestClient(app)
def _install_quota(coord) -> tuple[SpawnBudget, TokenBucket]:
"""Attach a real budget + bucket to the coord session under test."""
budget = SpawnBudget(20)
bucket = TokenBucket(5.0, 10)
session = MagicMock()
session._spawn_budget = budget
session._spawn_bucket = bucket
session._coord_client = MagicMock()
def _get_state():
return {
"spawn_budget": budget.budget,
"spawn_rate": {
"tokens_per_minute": bucket.tokens_per_minute,
"burst": bucket.burst,
"tokens_available": bucket.tokens,
},
}
def _set_budget(n):
budget.set_budget(int(n))
def _set_rate(tpm, brst):
bucket.set_rate(float(tpm), int(brst))
session.get_quota_state.side_effect = _get_state
session.set_spawn_budget.side_effect = _set_budget
session.set_spawn_rate.side_effect = _set_rate
coord.session = session
return budget, bucket
# ---------------------------------------------------------------------------
# GET
# ---------------------------------------------------------------------------
def test_quota_get_returns_live_snapshot(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.get(
f"/v1/api/coordinator/{coord.id}/quota",
headers=_COORD_HEADERS,
)
assert resp.status_code == 200
body = resp.json()
assert body["status"] == "ok"
assert body["spawn_budget"] == 20
assert body["spawn_rate"]["tokens_per_minute"] == 5.0
assert body["spawn_rate"]["burst"] == 10
assert 0 <= body["spawn_rate"]["tokens_available"] <= 10
def test_quota_get_404_when_session_not_loaded(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
coord.session = None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.get(
f"/v1/api/coordinator/{coord.id}/quota",
headers=_COORD_HEADERS,
)
assert resp.status_code == 404
# ---------------------------------------------------------------------------
# POST — happy path
# ---------------------------------------------------------------------------
def test_quota_post_updates_budget_only_and_audits(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
budget, bucket = _install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_budget": 42},
headers=_COORD_HEADERS,
)
assert resp.status_code == 200
body = resp.json()
assert body["spawn_budget"] == 42
# Rate left untouched — the partial update didn't widen it.
assert body["spawn_rate"]["tokens_per_minute"] == 5.0
assert body["spawn_rate"]["burst"] == 10
assert budget.budget == 42
events = [e for e in storage.list_audit_events() if e["action"] == "coordinator.quota.updated"]
assert len(events) == 1
detail = json.loads(events[0]["detail"])
assert detail["before"]["spawn_budget"] == 20
assert detail["after"]["spawn_budget"] == 42
def test_quota_post_accepts_nested_spawn_rate(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_budget, bucket = _install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_rate": {"tokens_per_minute": 30.0, "burst": 15}},
headers=_COORD_HEADERS,
)
assert resp.status_code == 200
body = resp.json()
assert body["spawn_rate"]["tokens_per_minute"] == 30.0
assert body["spawn_rate"]["burst"] == 15
assert bucket.burst == 15
def test_quota_post_accepts_flat_aliases(storage):
"""The admin UI may flatten the rate object — both shapes must work."""
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_budget, bucket = _install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"tokens_per_minute": 12.0, "burst": 4},
headers=_COORD_HEADERS,
)
assert resp.status_code == 200
assert bucket.tokens_per_minute == 12.0
assert bucket.burst == 4
def test_quota_post_burst_only_preserves_refill_rate(storage):
"""Changing only burst shouldn't zero the refill rate — a previous
bug-prone shape in partial-update handlers that overwrite missing
fields with defaults. Here the handler must read current state
for the missing dimension."""
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_budget, bucket = _install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"burst": 3},
headers=_COORD_HEADERS,
)
assert resp.status_code == 200
assert bucket.tokens_per_minute == 5.0 # unchanged
assert bucket.burst == 3
def test_quota_post_updates_all_three_knobs_at_once(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
budget, bucket = _install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_budget": 50, "tokens_per_minute": 0.0, "burst": 1},
headers=_COORD_HEADERS,
)
assert resp.status_code == 200
assert budget.budget == 50
assert bucket.tokens_per_minute == 0.0
assert bucket.burst == 1
# ---------------------------------------------------------------------------
# POST — validation failures
# ---------------------------------------------------------------------------
def test_quota_post_rejects_empty_body(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={},
headers=_COORD_HEADERS,
)
assert resp.status_code == 400
def test_quota_post_rejects_out_of_range_budget(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
for bad in (0, -5, 10_000):
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_budget": bad},
headers=_COORD_HEADERS,
)
assert resp.status_code == 400, f"expected 400 for {bad}"
def test_quota_post_rejects_non_numeric_rate(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"tokens_per_minute": "fast"},
headers=_COORD_HEADERS,
)
assert resp.status_code == 400
def test_quota_post_rejects_out_of_range_rate(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
for bad_tpm in (-1.0, 1_000.0):
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"tokens_per_minute": bad_tpm},
headers=_COORD_HEADERS,
)
assert resp.status_code == 400
def test_quota_post_rejects_out_of_range_burst(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
for bad in (0, -1, 10_000):
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"burst": bad},
headers=_COORD_HEADERS,
)
assert resp.status_code == 400
def test_quota_post_rejects_mixed_nested_and_flat_body(storage):
"""Schema description says 'don't mix' — the handler enforces it with 400.
Silently picking one side would make the admin UI's behaviour
unpredictable when it accidentally sends both shapes (e.g. during
a form-rewrite transition).
"""
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_rate": {"burst": 5}, "burst": 9},
headers=_COORD_HEADERS,
)
assert resp.status_code == 400
assert "conflicting" in resp.json()["error"]
def test_quota_post_rejects_non_object_spawn_rate(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_rate": "not-an-object"},
headers=_COORD_HEADERS,
)
assert resp.status_code == 400
def test_quota_post_rejects_bool_as_numeric_field(storage):
"""``True`` passes ``isinstance(x, int)`` in Python — explicit reject."""
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
for payload in (
{"spawn_budget": True},
{"burst": True},
{"tokens_per_minute": True},
):
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json=payload,
headers=_COORD_HEADERS,
)
assert resp.status_code == 400, f"expected 400 for {payload}"
def test_quota_post_404_when_session_not_loaded(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
coord.session = None
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_budget": 5},
headers=_COORD_HEADERS,
)
assert resp.status_code == 404
def test_quota_post_without_admin_coordinator_is_rejected(storage):
mgr = _build_mgr(storage)
coord = mgr.create(user_id="user-1", name="coord-a")
_install_quota(coord)
client = _make_client(storage, coord_mgr=mgr, registry=_fake_registry())
resp = client.post(
f"/v1/api/coordinator/{coord.id}/quota",
json={"spawn_budget": 5},
headers={"X-Test-User": "user-1", "X-Test-Perms": ""},
)
assert resp.status_code in (401, 403)
File diff suppressed because it is too large Load Diff
+10
View File
@@ -15,6 +15,16 @@ class TestIsSecret:
assert _is_secret("TURNSTONE_JWT_SECRET") is True
assert _is_secret("AWS_SECRET_ACCESS_KEY") is True
def test_tool_config_paths_scrubbed(self):
"""Tool-config env vars whose target files load executable
directives must be scrubbed even though they don't match a
secret-suffix pattern. Defence-in-depth alongside on-CLI
``--no-config`` for ripgrep and friends."""
assert _is_secret("RIPGREP_CONFIG_PATH") is True
assert _is_secret("GIT_CONFIG") is True
assert _is_secret("GIT_CONFIG_GLOBAL") is True
assert _is_secret("GIT_CONFIG_SYSTEM") is True
def test_suffix_matching(self):
assert _is_secret("MY_CUSTOM_API_KEY") is True
assert _is_secret("DB_PASSWORD") is True
+8
View File
@@ -8,6 +8,7 @@ from __future__ import annotations
from datetime import UTC, datetime
import pytest
import sqlalchemy as sa
# ---------------------------------------------------------------------------
@@ -317,6 +318,13 @@ class TestPromptTemplateCRUD:
def test_get_prompt_template_nonexistent(self, db):
assert db.get_prompt_template("missing") is None
def test_create_prompt_template_duplicate_id_raises_conflict(self, db):
from turnstone.core.storage._protocol import StorageConflictError
db.create_prompt_template("dup", "first", "general", "A")
with pytest.raises(StorageConflictError, match="prompt_template conflict"):
db.create_prompt_template("dup", "second", "general", "B")
def test_list_prompt_templates_ordered_by_name(self, db):
db.create_prompt_template("t2", "beta", "general", "B")
db.create_prompt_template("t1", "alpha", "general", "A")
+677
View File
@@ -0,0 +1,677 @@
"""Unit tests for ``turnstone.core.history_decoration``.
The decoration helpers are shared between two surfaces interactive's
SSE replay (``_build_history``) and the lifted ``/history`` REST
endpoint (``make_history_handler``, used by both interactive and
coord). Pinning the wire shape here lets a future schema/projection
change land in one file rather than spread across the two surfaces.
"""
from __future__ import annotations
from turnstone.core.history_decoration import (
build_output_assessment_payload,
build_verdict_payload,
decorate_history_messages,
decorate_tool_call,
)
class TestBuildVerdictPayload:
"""The wire-shape projection that's the single source of truth for
what intent_verdict fields ship to the client."""
def test_skips_unflagged_baseline(self) -> None:
"""``risk_level`` "none" is the unflagged-tool baseline; the
client filters those anyway, so projecting None at the wire
layer keeps the payload tight on long workstreams."""
row = {"risk_level": "none", "recommendation": "approve", "tier": "heuristic"}
assert build_verdict_payload(row) is None
def test_drops_call_id_and_func_name(self) -> None:
"""The client already has these on ``tc.id`` / ``tc.name``;
re-shipping them per-tool_call would balloon long replays."""
row = {
"call_id": "call_abc",
"func_name": "bash",
"risk_level": "medium",
"recommendation": "review",
"confidence": 0.8,
"intent_summary": "summary",
"tier": "heuristic",
}
out = build_verdict_payload(row)
assert out is not None
assert "call_id" not in out
assert "func_name" not in out
# Sanity — the kept fields are the ones renderVerdictBadge reads.
assert out["risk_level"] == "medium"
assert out["recommendation"] == "review"
assert out["confidence"] == 0.8
assert out["intent_summary"] == "summary"
assert out["tier"] == "heuristic"
def test_includes_reasoning_for_either_tier_when_present(self) -> None:
"""Heuristic verdicts in this project emit structured
rationales (one per matched pattern) e.g.
``policy.py`` writes a reasoning string per heuristic hit.
Ship the field for either tier when it has content; only
omit when the row didn't write one."""
for tier in ("heuristic", "llm"):
row = {
"risk_level": "high",
"tier": tier,
"reasoning": "The command exfiltrates ~/.ssh/id_rsa over an external connection.",
}
out = build_verdict_payload(row)
assert out is not None
assert "id_rsa" in out["reasoning"]
def test_omits_reasoning_when_empty(self) -> None:
"""An absent / empty reasoning string shouldn't ship as
``reasoning: ""`` the rationale ``<details>`` block on the
client renders an empty disclosure when the field is present
but empty."""
row = {"risk_level": "high", "tier": "heuristic", "reasoning": ""}
out = build_verdict_payload(row)
assert out is not None
assert "reasoning" not in out
def test_includes_judge_model_when_present(self) -> None:
"""``judge_model`` rides through so the batch tier badge can
render `` llm:claude-haiku-4`` on history-only replays
rather than the bare `` llm`` label."""
row = {"risk_level": "high", "tier": "llm", "judge_model": "claude-haiku-4"}
out = build_verdict_payload(row)
assert out is not None
assert out["judge_model"] == "claude-haiku-4"
def test_omits_judge_model_when_empty(self) -> None:
row = {"risk_level": "medium", "tier": "heuristic", "judge_model": ""}
out = build_verdict_payload(row)
assert out is not None
assert "judge_model" not in out
class TestBuildOutputAssessmentPayload:
"""Output-guard wire shape — flags decoded from JSON string at
this layer so the client never has to parse twice."""
def test_skips_unflagged_baseline(self) -> None:
row = {"risk_level": "none", "flags": "[]"}
assert build_output_assessment_payload(row) is None
def test_decodes_flags_from_json(self) -> None:
row = {"risk_level": "high", "flags": '["api_key","email"]', "redacted": 1}
out = build_output_assessment_payload(row)
assert out is not None
assert out["flags"] == ["api_key", "email"]
assert out["redacted"] is True
assert out["risk_level"] == "high"
def test_handles_malformed_flags_json(self) -> None:
"""Bad JSON in ``flags`` must not block the rest of the
assessment from rendering degrade to empty list."""
row = {"risk_level": "medium", "flags": "not-json", "redacted": 0}
out = build_output_assessment_payload(row)
assert out is not None
assert out["flags"] == []
assert out["redacted"] is False
class TestDecorateToolCall:
"""In-place mutation of either OpenAI-format or flattened tool_call
entries both shapes carry ``id`` at the top level."""
def test_attaches_verdict_when_present(self) -> None:
tc: dict[str, object] = {"id": "call_1", "function": {"name": "bash", "arguments": "{}"}}
verdicts = {
"call_1": {
"risk_level": "medium",
"recommendation": "review",
"confidence": 0.7,
"intent_summary": "summary",
"tier": "heuristic",
}
}
decorate_tool_call(tc, verdicts, {})
assert "verdict" in tc
assert tc["verdict"]["risk_level"] == "medium" # type: ignore[index]
def test_skips_when_no_call_id_match(self) -> None:
tc: dict[str, object] = {"id": "call_other", "name": "bash"}
verdicts = {
"call_1": {"risk_level": "medium", "tier": "heuristic"},
}
decorate_tool_call(tc, verdicts, {})
assert "verdict" not in tc
def test_skips_unflagged_verdict(self) -> None:
"""``build_verdict_payload`` returns None for unflagged rows;
decorate_tool_call must not stamp ``verdict`` in that case."""
tc: dict[str, object] = {"id": "call_1", "name": "bash"}
verdicts = {"call_1": {"risk_level": "none", "tier": "heuristic"}}
decorate_tool_call(tc, verdicts, {})
assert "verdict" not in tc
def test_handles_empty_id(self) -> None:
"""A tool_call with no id can't be paired against the lookup
table must not raise (or stamp the wrong row's verdict)."""
tc: dict[str, object] = {"id": "", "name": "bash"}
verdicts = {"call_1": {"risk_level": "high", "tier": "heuristic"}}
decorate_tool_call(tc, verdicts, {})
assert "verdict" not in tc
class TestDecorateHistoryMessages:
"""End-to-end mutation of a /history-shaped message list — covers
the full transform applied by ``make_history_handler``."""
def test_decorates_tool_calls_with_verdict_and_assessment(self) -> None:
verdicts = {
"call_a": {
"risk_level": "high",
"recommendation": "deny",
"confidence": 0.95,
"intent_summary": "exfil",
"tier": "llm",
"reasoning": "ssh key access",
}
}
assessments = {
"call_a": {"risk_level": "high", "flags": '["secret"]', "redacted": 1},
}
messages: list[dict[str, object]] = [
{"role": "user", "content": "hi"},
{
"role": "assistant",
"content": "running",
"tool_calls": [
{
"id": "call_a",
"function": {"name": "bash", "arguments": "{}"},
}
],
},
{"role": "tool", "tool_call_id": "call_a", "content": "long output"},
{"role": "tool", "tool_call_id": "call_b", "content": "short"},
]
decorate_history_messages(messages, verdicts, assessments)
# Assistant tool_calls got both decorations.
tc = messages[1]["tool_calls"][0] # type: ignore[index]
assert tc["verdict"]["risk_level"] == "high"
assert tc["verdict"]["tier"] == "llm"
assert "reasoning" in tc["verdict"]
assert tc["output_assessment"]["flags"] == ["secret"]
assert tc["output_assessment"]["redacted"] is True
# Plain tool content (no envelope) is left intact and no
# advisories key is set.
assert messages[2]["content"] == "long output"
assert "advisories" not in messages[2]
assert messages[3]["content"] == "short"
assert "advisories" not in messages[3]
def test_no_op_on_empty_indexes(self) -> None:
"""When neither table has rows for the workstream, the wire
shape passes through unchanged replay must degrade
gracefully when verdict storage is empty / unavailable."""
messages: list[dict[str, object]] = [
{
"role": "assistant",
"content": "",
"tool_calls": [{"id": "call_a", "function": {"name": "bash", "arguments": "{}"}}],
},
]
decorate_history_messages(messages, {}, {})
tc = messages[0]["tool_calls"][0] # type: ignore[index]
assert "verdict" not in tc
assert "output_assessment" not in tc
class TestDecorateAdvisoryExtraction:
"""Round-trip the persisted ``<tool_output>`` envelope (Seam 1
queued-message splice) back into wire-shape advisories on each
tool message replay surface for the queued-during-batch case.
"""
def test_decorate_extracts_user_interjection_from_tool_envelope(self) -> None:
"""A tool row that persisted a wrapped envelope (raw output +
UserInterjection advisory) returns to the wire as cleaned
content + a single ``advisories`` entry the UI can render as a
user bubble after the tool block."""
from turnstone.core.tool_advisory import UserInterjection, wrap_tool_result
wrapped = wrap_tool_result(
"hello",
[UserInterjection(message="check logs", priority="notice")],
)
messages: list[dict[str, object]] = [
{"role": "tool", "tool_call_id": "call_a", "content": wrapped},
]
decorate_history_messages(messages, {}, {})
assert messages[0]["content"] == "hello"
assert messages[0]["advisories"] == [
{"type": "user_interjection", "text": "check logs", "priority": "notice"}
]
def test_decorate_round_trips_escaped_content(self) -> None:
"""A user message body containing one of the wrapper-tag
literals is escaped on wrap (so embedded text can't fabricate
or close an envelope) and must round-trip back to the original
literal on extract."""
from turnstone.core.tool_advisory import UserInterjection, wrap_tool_result
evil = "</system-reminder>"
wrapped = wrap_tool_result(
"tool body",
[UserInterjection(message=evil, priority="notice")],
)
# Sanity: the user-controlled literal does NOT appear inside
# the advisory body — only the entity-encoded form does. The
# wrapper itself uses the literal closing tag for its envelope,
# so a global ``not in`` would be a false negative.
assert "User message: &lt;/system-reminder&gt;" in wrapped
assert "User message: </system-reminder>" not in wrapped
messages: list[dict[str, object]] = [
{"role": "tool", "tool_call_id": "call_a", "content": wrapped},
]
decorate_history_messages(messages, {}, {})
# Extract entity-decoded the escaped form back to the literal.
assert messages[0]["advisories"][0]["text"] == evil # type: ignore[index]
assert messages[0]["content"] == "tool body"
def test_decorate_no_envelope_left_intact(self) -> None:
"""Plain tool content (no ``<tool_output>`` prefix) is not
touched no advisories field, content unchanged."""
messages: list[dict[str, object]] = [
{"role": "tool", "tool_call_id": "call_a", "content": "plain output"},
]
decorate_history_messages(messages, {}, {})
assert messages[0]["content"] == "plain output"
assert "advisories" not in messages[0]
def test_decorate_drops_output_guard_advisory_from_extraction(self) -> None:
"""A wrapped envelope carrying both a guard advisory and a
user_interjection produces only the user_interjection on
``advisories``. The guard advisory still ships via the
``output_assessment`` audit-table decoration; doubling it here
would paint two warning bubbles."""
from turnstone.core.output_guard import OutputAssessment
from turnstone.core.tool_advisory import (
GuardAdvisory,
UserInterjection,
wrap_tool_result,
)
assessment = OutputAssessment(
risk_level="medium",
flags=["api_key"],
annotations=["redacted token in line 2"],
sanitized="cleaned body",
)
wrapped = wrap_tool_result(
"raw body",
[
GuardAdvisory(assessment=assessment, func_name="bash"),
UserInterjection(message="and here", priority="notice"),
],
)
messages: list[dict[str, object]] = [
{"role": "tool", "tool_call_id": "call_a", "content": wrapped},
]
decorate_history_messages(messages, {}, {})
adv = messages[0]["advisories"]
assert len(adv) == 1 # type: ignore[arg-type]
assert adv[0]["type"] == "user_interjection" # type: ignore[index]
def test_decorate_handles_important_priority(self) -> None:
"""The MUST-address preamble round-trips to ``priority=important``."""
from turnstone.core.tool_advisory import UserInterjection, wrap_tool_result
wrapped = wrap_tool_result(
"out",
[UserInterjection(message="urgent", priority="important")],
)
messages: list[dict[str, object]] = [
{"role": "tool", "tool_call_id": "call_a", "content": wrapped},
]
decorate_history_messages(messages, {}, {})
adv = messages[0]["advisories"][0] # type: ignore[index]
assert adv["priority"] == "important"
assert adv["text"] == "urgent"
def test_decorate_suppresses_empty_advisory_body(self) -> None:
"""``queue_message`` doesn't reject empty / whitespace-only
text, so an advisory with an empty body can round-trip through
``wrap_tool_result``. ``_classify_advisory`` must filter those
out so replay doesn't paint a featureless empty user bubble.
Removing the ``if not body.strip(): return None`` guard in
``_classify_advisory`` breaks this test."""
from turnstone.core.tool_advisory import UserInterjection, wrap_tool_result
wrapped = wrap_tool_result(
"tool body",
[UserInterjection(message="", priority="notice")],
)
messages: list[dict[str, object]] = [
{"role": "tool", "tool_call_id": "call_a", "content": wrapped},
]
decorate_history_messages(messages, {}, {})
# Envelope is still stripped from content (the cleaning side
# of decoration runs unconditionally), but no advisories
# surface — the empty body is filtered.
assert messages[0]["content"] == "tool body"
assert "advisories" not in messages[0]
def test_decorate_suppresses_whitespace_only_advisory_body(self) -> None:
"""Whitespace-only bodies are similarly suppressed — same
reasoning as the empty-body case."""
from turnstone.core.tool_advisory import UserInterjection, wrap_tool_result
wrapped = wrap_tool_result(
"tool body",
[UserInterjection(message=" \n\t ", priority="notice")],
)
messages: list[dict[str, object]] = [
{"role": "tool", "tool_call_id": "call_a", "content": wrapped},
]
decorate_history_messages(messages, {}, {})
assert messages[0]["content"] == "tool body"
assert "advisories" not in messages[0]
def test_wrap_extract_round_trips_preexisting_entities(self) -> None:
"""A user message body containing literal HTML-entity references
matching the wrapper-escape forms must round-trip identically
through ``wrap_tool_result + extract_advisories_from_tool_envelope``.
Without escaping ``&`` first in the encode step, encodedecode
would produce the bare wrapper tag, fabricating an envelope the
wrapper layer never produced.
"""
from turnstone.core.history_decoration import (
extract_advisories_from_tool_envelope,
)
from turnstone.core.tool_advisory import UserInterjection, wrap_tool_result
tricky = "I describe XML tags like &lt;tool_output&gt; in my docs."
wrapped = wrap_tool_result(
"tool body",
[UserInterjection(message=tricky, priority="notice")],
)
result = extract_advisories_from_tool_envelope(wrapped)
assert result is not None
cleaned, advisories = result
assert cleaned == "tool body"
assert len(advisories) == 1
# The original literal entity-reference text round-trips
# identically — the parser does not silently turn it into a
# bare wrapper tag.
assert advisories[0]["text"] == tricky
def test_save_load_decorate_round_trips_envelope(self, backend) -> None:
"""End-to-end round-trip pinning the persisted-envelope
contract. Persists a wrapped tool-output envelope via
``save_message``, loads via ``load_messages``, runs
``decorate_history_messages``, asserts the wire shape carries
the extracted advisory + cleaned content. Pins the contract
every component in the chain participates in (persistence
layer in-memory replay wire projection) so a schema drift,
an envelope-format change, or a parser regression surfaces
here rather than only in production.
"""
from turnstone.core.tool_advisory import UserInterjection, wrap_tool_result
wrapped = wrap_tool_result(
"command output",
[UserInterjection(message="check the logs", priority="notice")],
)
backend.register_workstream("ws_rt_1")
backend.save_message("ws_rt_1", "user", "go")
backend.save_message(
"ws_rt_1",
"assistant",
None,
tool_calls='[{"id":"call_a","type":"function","function":{"name":"bash","arguments":"{}"}}]',
)
backend.save_message(
"ws_rt_1",
"tool",
wrapped,
tool_call_id="call_a",
)
msgs = backend.load_messages("ws_rt_1")
# Persisted shape — content survives the storage layer
# untouched. Symmetry with in-memory ``self.messages[i]['content']``
# is what makes envelope extraction lossless on replay.
tool_msg = next(m for m in msgs if m["role"] == "tool")
assert tool_msg["content"] == wrapped
# Decorate (the /history shared transform) — extracts the
# advisory and strips the envelope.
decorate_history_messages(msgs, {}, {})
tool_msg = next(m for m in msgs if m["role"] == "tool")
assert tool_msg["content"] == "command output"
assert tool_msg["advisories"] == [
{"type": "user_interjection", "text": "check the logs", "priority": "notice"}
]
class TestExtractReasoningForHistory:
"""``extract_reasoning_for_history`` — Phase 1 surfaces stored
Anthropic thinking blocks on assistant messages and strips
``_provider_content`` from the wire payload.
Drives through the real ``AnthropicProvider.extract_reasoning_text``
(no mock-of-extractor) the helper test and the provider unit
test (``tests/test_provider_anthropic_reasoning.py``) together
catch a regression at either layer distinctly.
"""
def _anthropic_thinking_msg(self, text: str = "let me think") -> dict[str, object]:
return {
"role": "assistant",
"content": "Final answer.",
"_provider_content": [
{"type": "thinking", "thinking": text, "signature": "sig"},
{"type": "text", "text": "Final answer."},
],
}
def test_extract_thinking_surfaces_reasoning_field(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [self._anthropic_thinking_msg("let me think")]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert messages[0]["reasoning"] == "let me think"
def test_strips_provider_content_after_extraction(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [self._anthropic_thinking_msg("anything")]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert "_provider_content" not in messages[0]
def test_strips_provider_content_when_flag_false(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [self._anthropic_thinking_msg("anything")]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=False)
# Strip is unconditional; reasoning is the conditional bit.
assert "_provider_content" not in messages[0]
assert "reasoning" not in messages[0]
def test_first_block_thinking_dispatches_to_anthropic(self) -> None:
# Even when text and tool_use blocks follow, the first-block-type
# discriminator routes thinking-prefixed payloads correctly.
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [
{
"role": "assistant",
"content": "x",
"_provider_content": [
{"type": "thinking", "thinking": "first", "signature": "s"},
{"type": "text", "text": "spoken"},
{"type": "tool_use", "id": "t1", "name": "f", "input": {}},
],
}
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert messages[0]["reasoning"] == "first"
def test_first_block_reasoning_dispatches_to_openai_responses(self) -> None:
# Phase 3: dispatcher routes type=="reasoning" to the
# OpenAI Responses extractor, which now returns the
# summary[*].text concatenation. Pre-Phase-3 this asserted
# "" (the stub); the assertion was tightened once the wire
# path landed.
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [
{
"role": "assistant",
"content": "x",
"_provider_content": [
{"type": "reasoning", "summary": [{"type": "summary_text", "text": "s"}]}
],
}
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert messages[0]["reasoning"] == "s"
assert "_provider_content" not in messages[0]
def test_unknown_first_block_type_no_op(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [
{
"role": "assistant",
"content": "x",
"_provider_content": [{"type": "text", "text": "no reasoning here"}],
}
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert "reasoning" not in messages[0]
assert "_provider_content" not in messages[0]
def test_skips_messages_without_provider_content(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [{"role": "assistant", "content": "plain"}]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert "reasoning" not in messages[0]
assert messages[0]["content"] == "plain"
def test_user_and_tool_messages_untouched(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages: list[dict[str, object]] = [
{"role": "user", "content": "hi"},
{"role": "tool", "tool_call_id": "c1", "content": "out"},
self._anthropic_thinking_msg("only this one"),
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert "reasoning" not in messages[0]
assert "reasoning" not in messages[1]
assert messages[2]["reasoning"] == "only this one"
def test_empty_provider_content_no_extraction(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [{"role": "assistant", "content": "x", "_provider_content": []}]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert "reasoning" not in messages[0]
# Empty-list provider_content is still stripped from the wire.
assert "_provider_content" not in messages[0]
def test_first_block_not_a_dict_skipped(self) -> None:
from turnstone.core.history_decoration import extract_reasoning_for_history
messages: list[dict[str, object]] = [
{
"role": "assistant",
"content": "x",
"_provider_content": ["bogus"],
}
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert "reasoning" not in messages[0]
assert "_provider_content" not in messages[0]
def test_first_block_reasoning_text_dispatches_to_openai_chat(self) -> None:
# Phase 3 path 3: synthetic ``reasoning_text`` blocks (stamped
# by ChatSession._maybe_synth_reasoning_block for vLLM /
# llama.cpp / Gemini-compat conversations) dispatch to
# OpenAIChatCompletionsProvider.extract_reasoning_text.
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [
{
"role": "assistant",
"content": "answer",
"_provider_content": [
{"type": "reasoning_text", "text": "synth thought", "source": "vllm"},
],
}
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert messages[0]["reasoning"] == "synth thought"
assert "_provider_content" not in messages[0]
def test_dispatcher_scans_past_unrecognized_first_blocks(self) -> None:
# Regression for Copilot finding: dispatcher used to inspect
# only provider_content[0]['type']. OpenAI Responses captures
# EVERY output_item.done event into provider_blocks (not just
# reasoning), so a hypothetical [message, reasoning, ...]
# ordering would have silently dropped the reasoning. Now
# walks the list for the first recognised reasoning-bearing
# type and dispatches the whole list to that provider.
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [
{
"role": "assistant",
"content": "answer",
"_provider_content": [
# First block is a non-reasoning OpenAI Responses item.
{"type": "message", "role": "assistant", "content": "answer"},
# Reasoning sits later in the list.
{
"type": "reasoning",
"id": "r_1",
"summary": [{"type": "summary_text", "text": "deferred"}],
},
],
}
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert messages[0]["reasoning"] == "deferred"
assert "_provider_content" not in messages[0]
def test_first_block_redacted_thinking_dispatches_to_anthropic(self) -> None:
# Anthropic's extended-thinking API documents that
# ``redacted_thinking`` blocks (sealed by the safety system)
# can appear before, after, or interleaved with regular
# ``thinking`` blocks. When the redacted block lands first,
# the dispatcher must still route to AnthropicProvider so the
# surrounding real thinking text surfaces — without this the
# reasoning bubble silently disappears on history rehydration.
# Pinned by registering "redacted_thinking" as a second key
# in _BLOCK_TYPE_PROVIDER_FACTORY pointing at the Anthropic
# factory; Anthropic's extractor's type=="thinking" filter
# already correctly skips the redacted block.
from turnstone.core.history_decoration import extract_reasoning_for_history
messages = [
{
"role": "assistant",
"content": "answer",
"_provider_content": [
{"type": "redacted_thinking", "data": "sealed-blob"},
{"type": "thinking", "thinking": "real thought", "signature": "s"},
{"type": "text", "text": "answer"},
],
}
]
extract_reasoning_for_history(messages, surface_persisted_reasoning_flag=True)
assert messages[0]["reasoning"] == "real thought"
assert "_provider_content" not in messages[0]
+391
View File
@@ -0,0 +1,391 @@
"""Boundary-crossing integration test for the wake trigger pipeline.
Drives a *real* :class:`SessionManager` + a *real* :class:`ChatSession`
+ a *real* :class:`IdleNudgeWatcher` end-to-end. The only stub is the
LLM provider (patched ``_create_stream_with_retry``); every other layer
is production code:
* ``SessionManager.set_state`` snapshotting + iterating subscribers
* ``IdleNudgeWatcher._on_state`` peeking the queue
* ``session_worker.send`` atomic-spawn + daemon thread
* ``ChatSession.deliver_wake_nudge_from_queue`` opening / closing
``_wake_source_tag``
* ``ChatSession.send`` chat loop short-circuiting metacog detection
* ``_append_user_turn`` stamping ``_source = "system_nudge"``
* ``_attach_pending_user_reminders`` draining ``USER_DRAIN``
* ``_apply_reminders_for_provider`` splicing the rendered envelope
onto empty content
Per ``feedback_tests_through_boundaries.md``: direct injection tests
that bypass these boundaries silently mask wiring bugs. This test is
the structural integration gate.
"""
from __future__ import annotations
import time
from typing import Any
from unittest.mock import MagicMock, patch
import pytest
from tests.test_session_manager import FakeStorage
from turnstone.core.idle_nudge_watcher import IdleNudgeWatcher
from turnstone.core.session import ChatSession
from turnstone.core.session_manager import SessionManager
from turnstone.core.workstream import Workstream, WorkstreamKind, WorkstreamState
# ---------------------------------------------------------------------------
# Minimal fake adapter / UI for this integration test. Storage reuses
# the canonical FakeStorage from test_session_manager.py to avoid the
# drift risk of a parallel fake.
# ---------------------------------------------------------------------------
class _FakeUI:
"""Minimal UI surface for ChatSession + SessionManager.cleanup_ui."""
def __init__(self) -> None:
self.events: list[tuple[str, Any]] = []
def _unblock(self) -> None: # SessionManager.close calls this
pass
def broadcast_ws_closed(self) -> None:
pass
# ChatSession callbacks (no-op for this test)
def on_turn_start(self) -> None:
pass
def on_turn_committed(self) -> None:
pass
def on_thinking_start(self) -> None:
pass
def on_thinking_end(self) -> None:
pass
def on_state_change(self, state: str) -> None:
self.events.append(("state", state))
def on_user_reminder(self, reminders: Any, source: str | None = None) -> None:
self.events.append(("user_reminder", reminders, source))
def on_error(self, message: str) -> None:
pass
def on_rename(self, name: str) -> None:
pass
def on_output_warning(self, call_id: Any, assessment: Any) -> None:
pass
def __getattr__(self, name: str) -> Any:
# Catch-all for any UI hook not enumerated above so the chat
# loop's ``self.ui.<something>()`` call doesn't blow up.
return MagicMock()
class _BuildRealSessionAdapter:
"""Adapter that returns a real :class:`ChatSession` instead of a stub.
Tracks emit_* events the integration test asserts on. Mirrors the
``SessionKindAdapter`` + ``SessionEventEmitter`` Protocol surface
that production ``WebUI`` / coord adapters expose.
"""
def __init__(self, kind: WorkstreamKind = WorkstreamKind.INTERACTIVE) -> None:
self.kind = kind
self.events: list[str] = []
self.cleaned_up: list[str] = []
def emit_created(self, ws: Workstream) -> None:
self.events.append(f"created:{ws.id}")
def emit_rehydrated(self, ws: Workstream) -> None:
self.events.append(f"rehydrated:{ws.id}")
def emit_state(self, ws: Workstream, state: WorkstreamState) -> None:
self.events.append(f"state:{ws.id}:{state.value}")
def emit_closed(self, ws_id: str, *, reason: str = "closed", name: str = "") -> None:
self.events.append(f"closed:{ws_id}")
def cleanup_ui(self, ws: Workstream) -> None:
# Real production cleanup_ui calls ws.session.cancel() + close().
# We don't need that here — the test exits cleanly via pytest
# teardown without exercising the cleanup path. Just record
# the call for any test that wants to assert on it.
self.cleaned_up.append(ws.id)
def build_ui(self, ws: Workstream) -> Any:
return _FakeUI()
def build_session(
self,
ws: Workstream,
*,
skill: Any = None,
model: Any = None,
client_type: Any = None,
**extra: Any,
) -> Any:
# Mirror SessionManager.create's keyword set so config-threading
# bugs surface here rather than being silently swallowed by
# **kwargs. ``model`` flows to the real ChatSession; the rest
# are accepted but not used by this test.
client = MagicMock()
return ChatSession(
client=client,
model=str(model) if model else "test-model",
ui=ws.ui,
instructions=None,
temperature=0.5,
max_tokens=4096,
tool_timeout=30,
)
# ---------------------------------------------------------------------------
# Test
# ---------------------------------------------------------------------------
@pytest.fixture
def real_mgr() -> tuple[SessionManager, _BuildRealSessionAdapter]:
"""Real SessionManager wired to an adapter that builds real ChatSessions.
No StateWriter is wired so ``set_state`` writes directly to storage
on the calling thread (we want subscriber dispatch to fire in the
same thread the test invokes ``set_state`` on).
"""
adapter = _BuildRealSessionAdapter()
storage = FakeStorage()
mgr = SessionManager(
adapter,
storage=storage,
max_active=5,
event_emitter=adapter,
)
return mgr, adapter
def _wait_for_worker_done(ws: Workstream, timeout: float = 5.0) -> None:
"""Poll ``ws._worker_running`` until it clears or timeout elapses."""
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
with ws._lock:
if not ws._worker_running:
return
time.sleep(0.01)
raise AssertionError(f"worker thread for ws={ws.id[:8]} didn't exit within {timeout}s")
def test_idle_event_through_real_session_manager_drives_wake_send(real_mgr, tmp_db):
"""The full wake pipeline, no direct-injection shortcuts.
Boundary path under test:
enqueue mgr.set_state(IDLE)
SessionManager._state_subscribers iteration (real)
IdleNudgeWatcher._on_state (real)
session_worker.send (real)
real daemon thread
ChatSession.deliver_wake_nudge_from_queue (real)
ChatSession.send("") (real, with patched LLM stream)
_append_user_turn stamps ``_source``
_attach_pending_user_reminders drains ``{"user","any"}``
_apply_reminders_for_provider splices envelope onto empty content
"""
mgr, _adapter = real_mgr
watcher = IdleNudgeWatcher(mgr)
watcher.start()
try:
ws = mgr.create(user_id="u1", name="wake-int", skill=None)
assert ws.session is not None
# Patch the LLM-facing surface so send() runs the chat loop end-to-end
# without any real provider. We patch on the just-built ChatSession;
# the patches are reverted by the `with` block.
with (
patch.object(ws.session, "_create_stream_with_retry", return_value=iter([])),
patch.object(
ws.session,
"_stream_response",
return_value={"role": "assistant", "content": "ok"},
),
patch.object(ws.session, "_update_token_table"),
patch.object(ws.session, "_print_status_line"),
patch.object(ws.session, "_visible_memory_count", return_value=0),
patch("turnstone.core.session.save_message"),
):
# Suppress the auto-title side-thread; orthogonal to wake.
ws.session._title_generated = True
# Enqueue an any-channel nudge — the future ``idle_children`` shape.
ws.session._nudge_queue.enqueue("idle_children", "your kids", "any")
assert len(ws.session._nudge_queue) == 1
# Trigger IDLE. This runs subscriber dispatch synchronously on
# the calling thread → IdleNudgeWatcher._on_state → session_worker.send
# → spawn daemon thread → deliver_wake_nudge_from_queue.
mgr.set_state(ws.id, WorkstreamState.IDLE)
# Wait for the daemon thread to clear ``_worker_running`` so the
# post-conditions are stable.
_wait_for_worker_done(ws)
# Queue fully drained by the wake.
assert len(ws.session._nudge_queue) == 0
# The synthesized empty user message landed in history with the
# ``_source`` audit tag and the reminder side-channel populated.
user_msgs = [m for m in ws.session.messages if m.get("role") == "user"]
assert user_msgs, "expected a synthesized user message from the wake"
wake_msg = user_msgs[-1]
assert wake_msg["content"] == ""
assert wake_msg.get("_source") == "system_nudge"
assert wake_msg.get("_reminders") == [{"type": "idle_children", "text": "your kids"}]
# The wake-source tag is reset post-send so subsequent activity
# behaves normally.
assert ws.session._wake_source_tag == ""
finally:
watcher.shutdown()
def test_idle_event_with_empty_queue_does_not_dispatch_wake(real_mgr, tmp_db):
"""Non-empty queue is the gate. An IDLE event on a workstream with
nothing queued must NOT call ``session_worker.send``.
Patches the dispatch primitive directly rather than racing a
``time.sleep`` against an erroneous spawn the question is
whether the watcher's gate fired, which is a deterministic
decision the patch captures.
"""
mgr, _adapter = real_mgr
watcher = IdleNudgeWatcher(mgr)
watcher.start()
try:
ws = mgr.create(user_id="u1", name="empty-int", skill=None)
# No enqueue.
with patch("turnstone.core.session_worker.send") as mock_send:
mgr.set_state(ws.id, WorkstreamState.IDLE)
assert mock_send.call_count == 0, "wake must not dispatch for an empty queue"
finally:
watcher.shutdown()
@pytest.fixture
def coord_mgr() -> tuple[SessionManager, _BuildRealSessionAdapter, FakeStorage]:
"""Real coord-side SessionManager with the adapter's kind set to
COORDINATOR. Same shape as ``real_mgr`` but for the coord half of
the lifespan. No StateWriter wired so subscriber dispatch fires
synchronously on the test thread.
"""
adapter = _BuildRealSessionAdapter(kind=WorkstreamKind.COORDINATOR)
storage = FakeStorage()
mgr = SessionManager(
adapter,
storage=storage,
max_active=5,
event_emitter=adapter,
)
return mgr, adapter, storage
def test_coord_idle_with_active_children_emits_envelope_via_real_managers(coord_mgr, tmp_db):
"""Full coord-path integration test (matches design doc §7.4).
Drives the production install order ``CoordinatorIdleObserver``
registered FIRST, then ``IdleNudgeWatcher`` and asserts the
full chain: observer enqueues on IDLE watcher peeks wake
spawns a worker ``deliver_wake_nudge_from_queue`` drains and
runs the synthetic empty-user turn reminder envelope reaches
the synthesized user message via the side-channel.
The boundary-crossing path tested here mirrors what
``console/server.py``'s lifespan does at production startup; if
the install order is ever reversed, this test fails.
"""
from turnstone.console.coordinator_idle_observer import CoordinatorIdleObserver
from turnstone.core.workstream import WorkstreamKind as _Kind
mgr, adapter, storage = coord_mgr
# Observer FIRST, then watcher. Same order as
# ``console/server.py:4435-4443`` — production correctness depends
# on subscribers firing in registration order on the same IDLE.
observer = CoordinatorIdleObserver(mgr, storage)
observer.start()
watcher = IdleNudgeWatcher(mgr)
watcher.start()
try:
coord = mgr.create(user_id="u1", name="parent-coord", skill=None)
assert coord.session is not None
# Two interactive children of the coord, both running. Use
# the storage's register_workstream API so the rows match
# production shape (the observer queries via list_workstreams).
storage.register_workstream(
"child-a",
user_id="u1",
name="research-pricing",
kind=_Kind.INTERACTIVE,
parent_ws_id=coord.id,
state="running",
)
storage.register_workstream(
"child-b",
user_id="u1",
name="draft-rfc",
kind=_Kind.INTERACTIVE,
parent_ws_id=coord.id,
state="thinking",
)
# Pretend the coord has already had a real conversation so
# ``should_nudge``'s message_count > 1 gate passes.
coord.session.messages.append({"role": "user", "content": "spawn 2"})
coord.session.messages.append({"role": "assistant", "content": "ok"})
with (
patch.object(coord.session, "_create_stream_with_retry", return_value=iter([])),
patch.object(
coord.session,
"_stream_response",
return_value={"role": "assistant", "content": "ack"},
),
patch.object(coord.session, "_full_messages", return_value=[]),
patch.object(coord.session, "_update_token_table"),
patch.object(coord.session, "_print_status_line"),
patch.object(coord.session, "_visible_memory_count", return_value=0),
patch("turnstone.core.session.save_message"),
):
coord.session._title_generated = True
mgr.set_state(coord.id, WorkstreamState.IDLE)
_wait_for_worker_done(coord)
# Queue drained — the wake delivered the observer's enqueue.
assert len(coord.session._nudge_queue) == 0
# The synthetic empty-user turn landed with a reminder containing
# both children.
user_msgs = [m for m in coord.session.messages if m.get("role") == "user"]
# Two real msgs (user + assistant context above) plus the wake.
wake_msg = user_msgs[-1]
assert wake_msg["content"] == ""
assert wake_msg.get("_source") == "system_nudge"
reminders = wake_msg.get("_reminders") or []
assert len(reminders) == 1
assert reminders[0]["type"] == "idle_children"
text = reminders[0]["text"]
assert "research-pricing" in text
assert "draft-rfc" in text
assert "child-a" in text
assert "child-b" in text
assert "wait_for_workstream" in text
finally:
watcher.shutdown()
observer.shutdown()
+165
View File
@@ -0,0 +1,165 @@
"""Unit tests for :class:`IdleNudgeWatcher`.
Drives a fake :class:`SessionManager` that mimics the real one's
``subscribe_to_state`` / ``get`` contract. The watcher itself
dispatches via ``turnstone.core.session_worker.send``; we patch that
module-level function to capture calls without spawning real threads.
"""
from __future__ import annotations
import contextlib
import threading
from typing import Any
from unittest.mock import patch
import pytest
from turnstone.core.idle_nudge_watcher import IdleNudgeWatcher
from turnstone.core.nudge_queue import NudgeQueue
from turnstone.core.workstream import WorkstreamState
class _FakeSession:
def __init__(self) -> None:
self._nudge_queue = NudgeQueue()
self.deliver_wake_nudge_from_queue_called = 0
def deliver_wake_nudge_from_queue(self) -> None:
self.deliver_wake_nudge_from_queue_called += 1
class _FakeWorkstream:
def __init__(self, ws_id: str = "ws-test") -> None:
self.id = ws_id
self.session: _FakeSession | None = _FakeSession()
self._lock = threading.Lock()
self._worker_running = False
self._closed = False
self.worker_thread: Any = None
class _FakeManager:
"""Mimics SessionManager's subscribe-to-state surface without a DB."""
def __init__(self) -> None:
self._workstreams: dict[str, _FakeWorkstream] = {}
self._subscribers: list[Any] = []
self._subscribers_lock = threading.Lock()
def add_ws(self, ws: _FakeWorkstream) -> None:
self._workstreams[ws.id] = ws
def get(self, ws_id: str) -> _FakeWorkstream | None:
return self._workstreams.get(ws_id)
def subscribe_to_state(self, callback: Any) -> None:
with self._subscribers_lock:
self._subscribers.append(callback)
def unsubscribe_from_state(self, callback: Any) -> None:
with self._subscribers_lock, contextlib.suppress(ValueError):
self._subscribers.remove(callback)
def fire_state(self, ws_id: str, state: WorkstreamState) -> None:
"""Mirror SessionManager.set_state's subscriber-fan-out behaviour."""
with self._subscribers_lock:
subs = list(self._subscribers)
for cb in subs:
# Match contextlib.suppress(Exception) in real SessionManager.
with contextlib.suppress(Exception):
cb(ws_id, state)
@pytest.fixture
def fake_mgr_and_ws() -> tuple[_FakeManager, _FakeWorkstream]:
mgr = _FakeManager()
ws = _FakeWorkstream()
mgr.add_ws(ws)
return mgr, ws
class TestIdleNudgeWatcher:
def test_idle_event_with_empty_queue_no_op(self, fake_mgr_and_ws):
mgr, ws = fake_mgr_and_ws
watcher = IdleNudgeWatcher(mgr)
watcher.start()
with patch("turnstone.core.session_worker.send") as mock_send:
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert mock_send.call_count == 0
def test_idle_event_with_pending_nudge_dispatches(self, fake_mgr_and_ws):
mgr, ws = fake_mgr_and_ws
ws.session._nudge_queue.enqueue("idle_children", "your kids", "any")
watcher = IdleNudgeWatcher(mgr)
watcher.start()
with patch("turnstone.core.session_worker.send") as mock_send:
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert mock_send.call_count == 1
kwargs = mock_send.call_args.kwargs
# `enqueue=lambda: None` — verify by calling and checking no-op.
assert kwargs["enqueue"]() is None
# `run` should call deliver_wake_nudge_from_queue when invoked.
kwargs["run"]()
assert ws.session.deliver_wake_nudge_from_queue_called == 1
assert kwargs["thread_name"].startswith("wake-nudge-")
def test_non_idle_state_ignored(self, fake_mgr_and_ws):
mgr, ws = fake_mgr_and_ws
ws.session._nudge_queue.enqueue("foo", "bar", "any")
watcher = IdleNudgeWatcher(mgr)
watcher.start()
with patch("turnstone.core.session_worker.send") as mock_send:
for state in (
WorkstreamState.RUNNING,
WorkstreamState.THINKING,
WorkstreamState.ATTENTION,
WorkstreamState.ERROR,
):
mgr.fire_state(ws.id, state)
assert mock_send.call_count == 0
def test_unknown_ws_ignored(self, fake_mgr_and_ws):
mgr, _ws = fake_mgr_and_ws
watcher = IdleNudgeWatcher(mgr)
watcher.start()
with patch("turnstone.core.session_worker.send") as mock_send:
mgr.fire_state("ghost", WorkstreamState.IDLE)
assert mock_send.call_count == 0
def test_session_none_ignored(self, fake_mgr_and_ws):
mgr, ws = fake_mgr_and_ws
ws.session = None # workstream loaded but session not yet built
watcher = IdleNudgeWatcher(mgr)
watcher.start()
with patch("turnstone.core.session_worker.send") as mock_send:
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert mock_send.call_count == 0
def test_start_is_idempotent(self, fake_mgr_and_ws):
mgr, ws = fake_mgr_and_ws
watcher = IdleNudgeWatcher(mgr)
watcher.start()
watcher.start() # no-op
ws.session._nudge_queue.enqueue("foo", "bar", "any")
with patch("turnstone.core.session_worker.send") as mock_send:
mgr.fire_state(ws.id, WorkstreamState.IDLE)
# Only one subscriber was registered despite the double-start.
assert mock_send.call_count == 1
def test_shutdown_unsubscribes(self, fake_mgr_and_ws):
mgr, ws = fake_mgr_and_ws
ws.session._nudge_queue.enqueue("foo", "bar", "any")
watcher = IdleNudgeWatcher(mgr)
watcher.start()
watcher.shutdown()
with patch("turnstone.core.session_worker.send") as mock_send:
mgr.fire_state(ws.id, WorkstreamState.IDLE)
assert mock_send.call_count == 0
def test_shutdown_is_idempotent(self, fake_mgr_and_ws):
mgr, _ws = fake_mgr_and_ws
watcher = IdleNudgeWatcher(mgr)
watcher.start()
watcher.shutdown()
watcher.shutdown() # no error
+248
View File
@@ -0,0 +1,248 @@
"""Tests for InteractiveAdapter.
Focus: the ``emit_closed`` transport contract (sole path for
``ws_closed`` onto the process-wide queue) and ``cleanup_ui``
behavior (unblock pending events, broadcast ``ws_closed`` to per-UI
listeners, cancel + close session). The SessionManager-level tests
in ``test_session_manager.py`` cover the adapter-agnostic lifecycle.
The other three :class:`SessionEventEmitter` methods
(``emit_created`` / ``emit_state`` / ``emit_rehydrated``) are
documented no-op stubs ``ws_created`` is fired by the create HTTP
handler after attachment validation, and ``ws_state`` is fired by
``WebUI._broadcast_state`` with the full payload. No-op assertions
on those methods would be tautological given the class docstring,
so they're not retested here.
"""
from __future__ import annotations
import queue
import threading
from typing import Any
from unittest.mock import MagicMock
from turnstone.core.adapters.interactive_adapter import InteractiveAdapter
from turnstone.core.workstream import Workstream, WorkstreamKind
class _StubUI:
"""Stub matching the subset of WebUI the adapter touches."""
def __init__(self) -> None:
self._approval_event = threading.Event()
self._approval_result: tuple[bool, str | None] = (True, "initial")
self._plan_event = threading.Event()
self._plan_result: str = "accept"
self._fg_event = threading.Event()
self._listeners_lock = threading.Lock()
self._listeners: list[queue.Queue[dict[str, Any]]] = []
class _StubSession:
def __init__(self) -> None:
self.cancelled = False
self.closed = False
self.model = "gpt-5"
self.model_alias = "default"
def cancel(self) -> None:
self.cancelled = True
def close(self) -> None:
self.closed = True
def _make_adapter(
*,
ui_factory: Any = None,
session_factory: Any = None,
) -> tuple[InteractiveAdapter, queue.Queue[dict[str, Any]]]:
gq: queue.Queue[dict[str, Any]] = queue.Queue(maxsize=100)
adapter = InteractiveAdapter(
global_queue=gq,
ui_factory=ui_factory or (lambda ws: _StubUI()),
session_factory=session_factory or (lambda *a, **kw: _StubSession()),
)
return adapter, gq
def _make_ws(**overrides: Any) -> Workstream:
ws = Workstream(id="ws-1", name="hello")
ws.kind = WorkstreamKind.INTERACTIVE
ws.user_id = "u1"
ws.ui = _StubUI()
ws.session = _StubSession()
for k, v in overrides.items():
setattr(ws, k, v)
return ws
# ---------------------------------------------------------------------------
# Transport — emit_closed (the only emit_* with real behavior on interactive;
# emit_created / emit_state / emit_rehydrated are documented no-op stubs)
# ---------------------------------------------------------------------------
def test_emit_closed_defaults_to_closed_reason() -> None:
adapter, gq = _make_adapter()
adapter.emit_closed("ws-1", name="my-ws")
event = gq.get_nowait()
assert event == {
"type": "ws_closed",
"ws_id": "ws-1",
"reason": "closed",
"name": "my-ws",
}
def test_emit_closed_propagates_evicted_reason_and_name() -> None:
adapter, gq = _make_adapter()
adapter.emit_closed("ws-1", reason="evicted", name="my-ws")
event = gq.get_nowait()
assert event["reason"] == "evicted"
assert event["name"] == "my-ws"
def test_emit_closed_default_name_is_empty_string() -> None:
adapter, gq = _make_adapter()
adapter.emit_closed("ws-1")
assert gq.get_nowait()["name"] == ""
def test_emit_swallows_queue_full_without_raising() -> None:
gq: queue.Queue[dict[str, Any]] = queue.Queue(maxsize=1)
gq.put({"type": "filler"})
adapter = InteractiveAdapter(
global_queue=gq,
ui_factory=lambda ws: _StubUI(),
session_factory=lambda *a, **kw: _StubSession(),
)
adapter.emit_closed("ws-1") # must not raise even though queue is full
assert gq.qsize() == 1 # nothing added on a full queue
# ---------------------------------------------------------------------------
# cleanup_ui
# ---------------------------------------------------------------------------
def test_cleanup_ui_unblocks_pending_approval_plan_fg_events() -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
# Simulate pending events
ws.ui._approval_event.clear() # type: ignore[attr-defined]
ws.ui._plan_event.clear() # type: ignore[attr-defined]
ws.ui._fg_event.clear() # type: ignore[attr-defined]
adapter.cleanup_ui(ws)
assert ws.ui._approval_event.is_set() # type: ignore[attr-defined]
assert ws.ui._plan_event.is_set() # type: ignore[attr-defined]
assert ws.ui._fg_event.is_set() # type: ignore[attr-defined]
# Approval result flipped to "deny" so the waiter sees a sensible value.
assert ws.ui._approval_result == (False, None) # type: ignore[attr-defined]
assert ws.ui._plan_result == "reject" # type: ignore[attr-defined]
def test_cleanup_ui_broadcasts_ws_closed_to_listener_queues() -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
lq1: queue.Queue[dict[str, Any]] = queue.Queue(maxsize=10)
lq2: queue.Queue[dict[str, Any]] = queue.Queue(maxsize=10)
ws.ui._listeners.extend([lq1, lq2]) # type: ignore[attr-defined]
adapter.cleanup_ui(ws)
assert lq1.get_nowait() == {"type": "ws_closed"}
assert lq2.get_nowait() == {"type": "ws_closed"}
# Listeners cleared so subsequent events don't fan out to dead generators.
assert ws.ui._listeners == [] # type: ignore[attr-defined]
def test_cleanup_ui_broadcast_evicts_stale_head_when_listener_queue_full() -> None:
"""Per the old _cleanup_ui fallback: when a listener queue is full,
drop the oldest event and put ws_closed. Ensures an unresponsive
browser tab doesn't block close."""
adapter, _ = _make_adapter()
ws = _make_ws()
lq: queue.Queue[dict[str, Any]] = queue.Queue(maxsize=1)
lq.put_nowait({"type": "stale"})
ws.ui._listeners.append(lq) # type: ignore[attr-defined]
adapter.cleanup_ui(ws)
assert lq.get_nowait() == {"type": "ws_closed"}
assert lq.empty()
def test_cleanup_ui_cancels_and_closes_session() -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
adapter.cleanup_ui(ws)
assert ws.session.cancelled is True # type: ignore[attr-defined]
assert ws.session.closed is True # type: ignore[attr-defined]
def test_cleanup_ui_tolerates_missing_session_and_ui() -> None:
"""A placeholder workstream whose session build failed may arrive
at cleanup_ui with session=None or ui=None. Must not crash."""
adapter, _ = _make_adapter()
ws = _make_ws()
ws.session = None
ws.ui = None
adapter.cleanup_ui(ws) # no crash
def test_cleanup_ui_tolerates_stub_ui_without_events() -> None:
"""A stub UI missing _approval_event / etc. (test scaffolding
code) must not crash cleanup_ui the hasattr guards matter."""
adapter, _ = _make_adapter()
ws = _make_ws()
ws.ui = MagicMock(spec=[]) # empty spec — attribute accesses miss
adapter.cleanup_ui(ws) # no crash
# ---------------------------------------------------------------------------
# Construction passthrough
# ---------------------------------------------------------------------------
def test_build_ui_delegates_to_ui_factory() -> None:
captured_ws: list[Workstream] = []
def _ui_factory(ws: Workstream) -> Any:
captured_ws.append(ws)
return _StubUI()
adapter, _ = _make_adapter(ui_factory=_ui_factory)
ws = _make_ws()
result = adapter.build_ui(ws)
assert captured_ws == [ws]
assert isinstance(result, _StubUI)
def test_build_session_forwards_all_kwargs_to_session_factory() -> None:
captured: dict[str, Any] = {}
def _sf(ui: Any, model: str | None, ws_id: str, **kwargs: Any) -> Any:
captured["ui"] = ui
captured["model"] = model
captured["ws_id"] = ws_id
captured.update(kwargs)
return _StubSession()
adapter, _ = _make_adapter(session_factory=_sf)
ws = _make_ws()
adapter.build_session(
ws, skill="coder", model="gpt-5", client_type="web", judge_model="gpt-4.1"
)
assert captured["ui"] is ws.ui
assert captured["model"] == "gpt-5"
assert captured["ws_id"] == ws.id
assert captured["skill"] == "coder"
assert captured["client_type"] == "web"
assert captured["kind"] == WorkstreamKind.INTERACTIVE
assert captured["parent_ws_id"] is None
# Kind-specific passthrough — interactive session_factory accepts judge_model.
assert captured["judge_model"] == "gpt-4.1"
+52
View File
@@ -777,6 +777,58 @@ class TestModelAliasResolution:
assert judge._client_factory_args["api_key"] == "alias-key"
assert judge._client_factory_args["provider_name"] == "openai"
def test_unknown_alias_inherits_session_model(self):
"""``judge.model`` is alias-only. A value that doesn't resolve
through the registry inherits the session model (same path as
an empty config.model) rather than getting pinned onto the
session provider as a raw model id that legacy behavior
silently broke whenever the session provider didn't speak the
configured model id (Anthropic session, ``judge.model =
"gpt-5-mini"`` every verdict came back as ``llm_fallback``)."""
session_provider = _make_mock_provider()
session_provider.provider_name = "anthropic"
session_client = MagicMock()
session_client.base_url = "https://session.example/v1"
session_client.api_key = "session-key"
registry = MagicMock()
registry.has_alias.return_value = False # judge.model isn't an alias
config = JudgeConfig(enabled=True, model="gpt-5-mini")
judge = IntentJudge(
config=config,
session_provider=session_provider,
session_client=session_client,
session_model="session-default-model",
context_window=100_000,
model_registry=registry,
)
assert judge._provider is session_provider
assert judge._model == "session-default-model"
# Context window mirrors the session, not the (uncalled) caps lookup.
assert judge._judge_context_window == 100_000
def test_empty_model_inherits_session_model(self):
"""Empty ``config.model`` is the documented self-consistency path."""
session_provider = _make_mock_provider()
session_provider.provider_name = "openai"
session_client = MagicMock()
session_client.base_url = "https://session.example/v1"
session_client.api_key = "session-key"
config = JudgeConfig(enabled=True, model="")
judge = IntentJudge(
config=config,
session_provider=session_provider,
session_client=session_client,
session_model="session-default-model",
context_window=100_000,
)
assert judge._provider is session_provider
assert judge._model == "session-default-model"
def test_coordinator_tool_call_returns_llm_verdict_not_fallback(self):
"""Happy-path regression for coordinator tool calls: with a properly
resolved provider, the verdict tier must be ``llm`` the
+53
View File
@@ -114,6 +114,59 @@ class TestIntentVerdictCRUD:
assert ok is False
# ---------------------------------------------------------------------------
# Bulk insert
# ---------------------------------------------------------------------------
class TestIntentVerdictBulkInsert:
"""Coverage for ``create_intent_verdicts_bulk`` — backs the
``approve_tools`` per-turn heuristic-verdict persistence path so a
fan-out turn pays one commit instead of N.
"""
def test_bulk_insert_creates_all_rows(self, db):
db.create_intent_verdicts_bulk(
[
_make_verdict_kwargs(verdict_id="b1", call_id="c1"),
_make_verdict_kwargs(verdict_id="b2", call_id="c2"),
_make_verdict_kwargs(verdict_id="b3", call_id="c3"),
]
)
for vid in ("b1", "b2", "b3"):
v = db.get_intent_verdict(vid)
assert v is not None
assert v["verdict_id"] == vid
def test_bulk_insert_empty_list_is_noop(self, db):
# Must not raise and must not commit a phantom row.
db.create_intent_verdicts_bulk([])
assert db.list_intent_verdicts() == []
def test_bulk_insert_preserves_distinct_field_values(self, db):
db.create_intent_verdicts_bulk(
[
_make_verdict_kwargs(
verdict_id="b1",
risk_level="low",
tier="heuristic",
confidence=0.4,
),
_make_verdict_kwargs(
verdict_id="b2",
risk_level="high",
tier="llm",
confidence=0.95,
),
]
)
v1 = db.get_intent_verdict("b1")
v2 = db.get_intent_verdict("b2")
assert v1 is not None and v2 is not None
assert v1["risk_level"] == "low" and v1["tier"] == "heuristic"
assert v2["risk_level"] == "high" and v2["tier"] == "llm"
# ---------------------------------------------------------------------------
# List queries
# ---------------------------------------------------------------------------
+3 -1
View File
@@ -420,7 +420,9 @@ class TestSkillCatalogDisclosure:
session.system_messages = []
session._agent_system_messages = []
session.reasoning_effort = "medium"
session._pending_nudge = []
from turnstone.core.nudge_queue import NudgeQueue
session._nudge_queue = NudgeQueue()
session._tool_search = None
session._mcp_client = None
session._notify_on_complete = "{}"
File diff suppressed because it is too large Load Diff
+786 -148
View File
File diff suppressed because it is too large Load Diff
+117
View File
@@ -0,0 +1,117 @@
"""Structural gate against the Phase 7b sibling-bug pattern.
Phase 7b's bug-1 was a single ``f"MCP X error: {e}"`` site dropping a
structured-error JSON. Phase 8 introduces the ``consent_url`` field on
the same JSON envelope: every ``_structured_error(...)`` invocation
that emits ``mcp_consent_required`` or ``mcp_insufficient_scope`` MUST
also pass a ``consent_url=`` kwarg, otherwise the dashboard renderer
can't surface a re-consent button.
This test is purely structural it scans the source of
:mod:`turnstone.core.mcp_client` and asserts every consent-required /
insufficient-scope ``_structured_error`` call carries
``consent_url=``. It catches future regressions where a new exec path
adds a fourth call site and forgets the kwarg.
"""
from __future__ import annotations
import re
from pathlib import Path
import turnstone.core.mcp_client as _mcp_client_module
_USER_ACTIONABLE_CODES = ("mcp_consent_required", "mcp_insufficient_scope")
def _read_source() -> str:
path = Path(_mcp_client_module.__file__)
return path.read_text(encoding="utf-8")
def _find_structured_error_blocks(source: str) -> list[tuple[int, str]]:
"""Return ``(line_no, block)`` pairs for every ``_structured_error(...)``.
Each block is the call's argument list expanded across however many
lines the formatter chose. Uses a paren-counting walk so multi-line
kwargs and nested expressions are captured correctly.
"""
blocks: list[tuple[int, str]] = []
needle = "_structured_error("
idx = 0
while True:
loc = source.find(needle, idx)
if loc < 0:
break
# Skip the function definition itself.
if source[loc - 4 : loc] == "def ":
idx = loc + len(needle)
continue
line_no = source.count("\n", 0, loc) + 1
depth = 1
end = loc + len(needle)
while end < len(source) and depth > 0:
ch = source[end]
if ch == "(":
depth += 1
elif ch == ")":
depth -= 1
end += 1
blocks.append((line_no, source[loc:end]))
idx = end
return blocks
def test_every_user_actionable_structured_error_passes_consent_url() -> None:
source = _read_source()
blocks = _find_structured_error_blocks(source)
user_actionable_blocks = [
(ln, blk)
for ln, blk in blocks
if any(f'code="{code}"' in blk for code in _USER_ACTIONABLE_CODES)
]
# Sanity check: ensure we actually scanned the file the audit cares
# about (a stale path or import would otherwise silently pass with
# zero matches).
assert user_actionable_blocks, (
"No mcp_consent_required / mcp_insufficient_scope _structured_error "
"call sites found — has the audit been pointed at the wrong file?"
)
missing: list[tuple[int, str]] = []
for ln, blk in user_actionable_blocks:
if "consent_url=" not in blk:
# Strip whitespace and truncate so the failure message is
# readable in CI.
collapsed = re.sub(r"\s+", " ", blk).strip()
missing.append((ln, collapsed[:200]))
assert not missing, (
"Sibling-bug regression: the following consent-required / "
"insufficient-scope _structured_error sites are missing the "
"consent_url= kwarg.\n" + "\n".join(f" line {ln}: {snippet}" for ln, snippet in missing)
)
def test_audit_finds_all_known_user_actionable_sites() -> None:
"""Lock the count so accidental deletions are caught.
There are 13 user-actionable ``_structured_error`` call sites today
(4 each in the tool / resource / prompt token-classify branches +
3 in the post-retry-failed branches + 1 in ``_handle_auth_403``'s
insufficient-scope branch). If a new exec path is added the count
can rise; if a branch is removed the count can fall both are
fine, but require an intentional bump of this number to confirm
the change went through review.
"""
source = _read_source()
blocks = _find_structured_error_blocks(source)
user_actionable_count = sum(
1 for _, blk in blocks if any(f'code="{code}"' in blk for code in _USER_ACTIONABLE_CODES)
)
assert user_actionable_count == 13, (
f"Expected 13 user-actionable _structured_error sites, got "
f"{user_actionable_count}. If this is intentional, bump the "
f"expected count and document why in the commit message."
)
+235
View File
@@ -0,0 +1,235 @@
"""Tests for ``turnstone.core.mcp_crypto`` cipher + config loading.
Covers token-at-rest encryption for OAuth-MCP.
"""
from __future__ import annotations
import base64
import pytest
from cryptography.fernet import Fernet
from turnstone.core.mcp_crypto import (
MCPTokenCipher,
MCPTokenCipherConfig,
MCPTokenDecryptError,
MCPTokenKeyConfigError,
_key_fingerprint,
_validate_key,
load_mcp_token_cipher_config,
)
def _new_raw_key() -> bytes:
"""Return a fresh 32-byte Fernet key as raw bytes (post-base64-decode)."""
return base64.urlsafe_b64decode(Fernet.generate_key())
# ---------------------------------------------------------------------------
# Cipher round-trip
# ---------------------------------------------------------------------------
class TestCipherRoundTrip:
def test_round_trip_single_key(self) -> None:
cipher = MCPTokenCipher(MCPTokenCipherConfig(keys=(_new_raw_key(),)))
plaintext = b"access_token_12345"
ct = cipher.encrypt(plaintext)
assert ct != plaintext
assert cipher.decrypt(ct) == plaintext
def test_round_trip_unicode_token(self) -> None:
cipher = MCPTokenCipher(MCPTokenCipherConfig(keys=(_new_raw_key(),)))
# Tokens may legitimately carry UTF-8 bytes (e.g. JWT with
# non-ASCII claim values). Round-trip a multi-byte sequence.
plaintext = "tok_é中💯".encode()
ct = cipher.encrypt(plaintext)
assert cipher.decrypt(ct) == plaintext
def test_wrong_key_raises_decrypt_error(self) -> None:
cipher_a = MCPTokenCipher(MCPTokenCipherConfig(keys=(_new_raw_key(),)))
cipher_b = MCPTokenCipher(MCPTokenCipherConfig(keys=(_new_raw_key(),)))
ct = cipher_a.encrypt(b"secret")
with pytest.raises(MCPTokenDecryptError) as exc_info:
cipher_b.decrypt(ct)
# Audit-trail correlation: error must carry the fingerprints of
# the keys actually attempted, not a placeholder.
assert exc_info.value.key_fingerprints_attempted
assert exc_info.value.key_fingerprints_attempted == cipher_b.key_fingerprints
# ---------------------------------------------------------------------------
# Rotation (MultiFernet behavior)
# ---------------------------------------------------------------------------
class TestRotation:
def test_rotation_forward(self) -> None:
"""Encrypt with a new-only cipher, decrypt with a [v2, v1] cluster.
Mirrors the operational situation where a node already has the
rotated key list installed and a peer just wrote a row under v2.
"""
v1 = _new_raw_key()
v2 = _new_raw_key()
new_only = MCPTokenCipher(MCPTokenCipherConfig(keys=(v2,)))
cluster = MCPTokenCipher(MCPTokenCipherConfig(keys=(v2, v1)))
ct = new_only.encrypt(b"hello")
assert cluster.decrypt(ct) == b"hello"
def test_rotation_backward_keeps_old_decryptable(self) -> None:
"""A row written under the OLD key (v1) must still decrypt after
rotation places v2 first and keeps v1 as fallback."""
v1 = _new_raw_key()
v2 = _new_raw_key()
old_only = MCPTokenCipher(MCPTokenCipherConfig(keys=(v1,)))
rotated = MCPTokenCipher(MCPTokenCipherConfig(keys=(v2, v1)))
ct = old_only.encrypt(b"legacy")
assert rotated.decrypt(ct) == b"legacy"
# ---------------------------------------------------------------------------
# Config loader
# ---------------------------------------------------------------------------
def _patch_load_config(monkeypatch: pytest.MonkeyPatch, payload: dict) -> None:
"""Override ``turnstone.core.config.load_config`` to return ``payload``
when the ``"security"`` section is requested."""
def fake(section: str | None = None) -> dict:
if section == "security":
return payload
return {}
import turnstone.core.config as cfg_mod
monkeypatch.setattr(cfg_mod, "load_config", fake)
class TestLoadConfig:
def test_load_singular_key(self, monkeypatch: pytest.MonkeyPatch) -> None:
key = Fernet.generate_key().decode()
_patch_load_config(monkeypatch, {"mcp_token_encryption_key": key})
cfg = load_mcp_token_cipher_config()
assert cfg is not None
assert len(cfg.keys) == 1
def test_load_plural_overrides_singular(self, monkeypatch: pytest.MonkeyPatch) -> None:
plural = [Fernet.generate_key().decode(), Fernet.generate_key().decode()]
_patch_load_config(
monkeypatch,
{
"mcp_token_encryption_keys": plural,
"mcp_token_encryption_key": Fernet.generate_key().decode(),
},
)
cfg = load_mcp_token_cipher_config()
assert cfg is not None
assert len(cfg.keys) == 2 # plural wins, singular ignored
def test_load_returns_none_when_absent(self, monkeypatch: pytest.MonkeyPatch) -> None:
_patch_load_config(monkeypatch, {})
assert load_mcp_token_cipher_config() is None
def test_load_empty_plural_falls_through_to_singular(
self, monkeypatch: pytest.MonkeyPatch
) -> None:
"""Operator wrote ``mcp_token_encryption_keys = []`` AND set a
singular value: empty plural is treated as absent."""
key = Fernet.generate_key().decode()
_patch_load_config(
monkeypatch,
{"mcp_token_encryption_keys": [], "mcp_token_encryption_key": key},
)
cfg = load_mcp_token_cipher_config()
assert cfg is not None
assert len(cfg.keys) == 1
def test_load_invalid_base64_raises(self, monkeypatch: pytest.MonkeyPatch) -> None:
_patch_load_config(monkeypatch, {"mcp_token_encryption_key": "###not-base64###"})
with pytest.raises(MCPTokenKeyConfigError) as exc_info:
load_mcp_token_cipher_config()
# Operator-facing hint is part of every error message.
assert "regenerate with:" in str(exc_info.value)
def test_load_wrong_length_raises(self, monkeypatch: pytest.MonkeyPatch) -> None:
# 24 raw bytes → 32 base64 chars; not 32 raw bytes after decode.
short_key = base64.urlsafe_b64encode(b"\x00" * 24).decode()
_patch_load_config(monkeypatch, {"mcp_token_encryption_key": short_key})
with pytest.raises(MCPTokenKeyConfigError) as exc_info:
load_mcp_token_cipher_config()
assert "32 bytes" in str(exc_info.value)
def test_load_non_list_plural_raises(self, monkeypatch: pytest.MonkeyPatch) -> None:
_patch_load_config(monkeypatch, {"mcp_token_encryption_keys": "single-string-not-list"})
with pytest.raises(MCPTokenKeyConfigError) as exc_info:
load_mcp_token_cipher_config()
assert "list" in str(exc_info.value).lower()
def test_load_non_string_in_plural_raises(self, monkeypatch: pytest.MonkeyPatch) -> None:
_patch_load_config(monkeypatch, {"mcp_token_encryption_keys": [12345]})
with pytest.raises(MCPTokenKeyConfigError):
load_mcp_token_cipher_config()
# ---------------------------------------------------------------------------
# Fingerprint stability
# ---------------------------------------------------------------------------
class TestFingerprint:
def test_key_fingerprint_stable_and_short(self) -> None:
key = _new_raw_key()
fp1 = _key_fingerprint(key)
fp2 = _key_fingerprint(key)
assert fp1 == fp2
# 8 bytes -> 16 hex characters.
assert len(fp1) == 16
assert all(c in "0123456789abcdef" for c in fp1)
def test_different_keys_have_different_fingerprints(self) -> None:
fp1 = _key_fingerprint(_new_raw_key())
fp2 = _key_fingerprint(_new_raw_key())
assert fp1 != fp2
def test_cipher_fingerprints_match_keys(self) -> None:
v1 = _new_raw_key()
v2 = _new_raw_key()
cipher = MCPTokenCipher(MCPTokenCipherConfig(keys=(v1, v2)))
assert cipher.key_fingerprints == (
_key_fingerprint(v1),
_key_fingerprint(v2),
)
# ---------------------------------------------------------------------------
# Direct ``_validate_key`` — exercises edge cases not reachable via loader
# ---------------------------------------------------------------------------
class TestValidateKey:
def test_empty_string_rejected(self) -> None:
with pytest.raises(MCPTokenKeyConfigError):
_validate_key("", label="x")
def test_whitespace_only_rejected(self) -> None:
with pytest.raises(MCPTokenKeyConfigError):
_validate_key(" ", label="x")
def test_label_propagated_in_error(self) -> None:
with pytest.raises(MCPTokenKeyConfigError) as exc_info:
_validate_key("###", label="my_label_42")
assert "my_label_42" in str(exc_info.value)
# ---------------------------------------------------------------------------
# MCPTokenCipher constructor guard
# ---------------------------------------------------------------------------
class TestCipherConstructorGuard:
def test_empty_keys_rejected(self) -> None:
with pytest.raises(MCPTokenKeyConfigError):
MCPTokenCipher(MCPTokenCipherConfig(keys=()))
+28 -27
View File
@@ -4,6 +4,7 @@ from __future__ import annotations
from typing import Any
from tests.conftest import _seed_static_state
from turnstone.core.mcp_client import MCPClientManager
# ---------------------------------------------------------------------------
@@ -105,14 +106,18 @@ class TestRemoveServerSync:
"""remove_server_sync cleans up all per-server state dicts."""
mgr = MCPClientManager({"test": {"command": "echo"}})
# Simulate state as if the server was connected
mgr._per_server_tools["test"] = [_fake_openai_tool()]
mgr._per_server_resources["test"] = [_fake_resource_dict()]
mgr._per_server_prompts["test"] = [_fake_prompt_dict()]
mgr._supports_list_changed["test"] = True
mgr._supports_resources["test"] = True
mgr._supports_resource_list_changed["test"] = True
mgr._supports_prompts["test"] = True
mgr._supports_prompt_list_changed["test"] = True
_seed_static_state(
mgr,
"test",
tools=[_fake_openai_tool()],
resources=[_fake_resource_dict()],
prompts=[_fake_prompt_dict()],
supports_list_changed=True,
supports_resources=True,
supports_resource_list_changed=True,
supports_prompts=True,
supports_prompt_list_changed=True,
)
mgr._rebuild_tools()
mgr._rebuild_resources()
mgr._rebuild_prompts()
@@ -127,14 +132,7 @@ class TestRemoveServerSync:
assert len(mgr.get_tools()) == 0
assert mgr.resource_count == 0
assert mgr.prompt_count == 0
assert "test" not in mgr._per_server_tools
assert "test" not in mgr._per_server_resources
assert "test" not in mgr._per_server_prompts
assert "test" not in mgr._supports_list_changed
assert "test" not in mgr._supports_resources
assert "test" not in mgr._supports_resource_list_changed
assert "test" not in mgr._supports_prompts
assert "test" not in mgr._supports_prompt_list_changed
assert "test" not in mgr._static_servers
def test_removes_config_to_prevent_reconnect(self) -> None:
"""remove_server_sync removes from _server_configs to prevent reconnect."""
@@ -146,8 +144,8 @@ class TestRemoveServerSync:
def test_preserves_other_servers(self) -> None:
"""Removing one server does not affect another server's state."""
mgr = MCPClientManager({"srv_a": {}, "srv_b": {}})
mgr._per_server_tools["srv_a"] = [_fake_openai_tool("mcp__srv_a__foo")]
mgr._per_server_tools["srv_b"] = [_fake_openai_tool("mcp__srv_b__bar")]
_seed_static_state(mgr, "srv_a", tools=[_fake_openai_tool("mcp__srv_a__foo")])
_seed_static_state(mgr, "srv_b", tools=[_fake_openai_tool("mcp__srv_b__bar")])
mgr._rebuild_tools()
assert len(mgr.get_tools()) == 2
@@ -179,13 +177,17 @@ class TestGetServerStatus:
"""Status of a connected server reports correct tool/resource/prompt counts."""
mgr = MCPClientManager({"test": {}})
# Simulate connected state
mgr._sessions["test"] = object() # any truthy value
mgr._per_server_tools["test"] = [
_fake_openai_tool("mcp__test__a"),
_fake_openai_tool("mcp__test__b"),
]
mgr._per_server_resources["test"] = [_fake_resource_dict()]
mgr._per_server_prompts["test"] = [_fake_prompt_dict()]
_seed_static_state(
mgr,
"test",
session=object(), # any truthy value
tools=[
_fake_openai_tool("mcp__test__a"),
_fake_openai_tool("mcp__test__b"),
],
resources=[_fake_resource_dict()],
prompts=[_fake_prompt_dict()],
)
status = mgr.get_server_status("test")
assert status["connected"] is True
@@ -225,8 +227,7 @@ class TestGetAllServerStatus:
def test_mixed_connected_and_disconnected(self) -> None:
"""Status correctly reflects a mix of connected and disconnected servers."""
mgr = MCPClientManager({"up": {}, "down": {}})
mgr._sessions["up"] = object()
mgr._per_server_tools["up"] = [_fake_openai_tool("mcp__up__x")]
_seed_static_state(mgr, "up", session=object(), tools=[_fake_openai_tool("mcp__up__x")])
statuses = mgr.get_all_server_status()
assert statuses["up"]["connected"] is True
+242
View File
@@ -0,0 +1,242 @@
"""Unit tests for ``turnstone.core.mcp_http_parsers``.
The parser replaces the prior hand-rolled scanners that used
``header.lower().find("scope")`` to locate parameter names that approach
misparsed ``scope`` embedded inside other tokens (``xscope``) or inside
quoted-string values of preceding params. Each adversarial case below
asserts the new tokenizer respects RFC 7235 ``challenge auth-param``
boundaries; the docstrings document the equivalent input that broke the
naive parser. Negative-test verification: temporarily reverting
``parse_www_authenticate_scope`` to delegate to ``header.lower().find("scope")``
makes ``test_scope_inside_realm_value`` and ``test_scope_inside_xscope`` fail.
"""
from __future__ import annotations
import time
import pytest
from turnstone.core.mcp_http_parsers import (
parse_www_authenticate_bearer,
parse_www_authenticate_error,
parse_www_authenticate_scope,
)
class TestParseScope:
def test_basic_scope(self) -> None:
header = 'Bearer error="insufficient_scope", scope="files:read mail:send"'
assert parse_www_authenticate_scope(header) == ("files:read", "mail:send")
def test_no_scope_param(self) -> None:
assert parse_www_authenticate_scope('Bearer error="invalid_token"') == ()
def test_unterminated_quoted_string_returns_empty(self) -> None:
assert parse_www_authenticate_scope('Bearer scope="files:read') == ()
def test_escaped_chars_in_value_drops_invalid_scope_token(self) -> None:
# RFC 7230 §3.2.6 backslash escapes decode the literal scope to
# ``files:read "weird"``. RFC 6749 §3.3 ``scope-token`` forbids
# ``"``, so ``"weird"`` is dropped and only ``files:read``
# survives the post-split validation.
header = r'Bearer scope="files:read \"weird\""'
assert parse_www_authenticate_scope(header) == ("files:read",)
def test_empty_string(self) -> None:
assert parse_www_authenticate_scope("") == ()
def test_unquoted_scope_value(self) -> None:
# Unquoted single token.
assert parse_www_authenticate_scope("Bearer scope=files:read") == ("files:read",)
# --- the four headline misparse cases ---
def test_scope_inside_xscope(self) -> None:
"""``Bearer xscope="value"`` must NOT be read as ``scope``.
The naive ``find("scope")`` matched at position 7 inside
``xscope`` and returned ``("value",)``.
"""
assert parse_www_authenticate_scope('Bearer xscope="value"') == ()
def test_scope_inside_realm_value(self) -> None:
"""``Bearer realm="my scope=fake", scope="real"`` must return ``("real",)``.
The naive parser found ``scope=`` inside the quoted ``realm``
value first and returned ``("fake",)``.
"""
header = 'Bearer realm="my scope=fake", scope="real"'
assert parse_www_authenticate_scope(header) == ("real",)
def test_scope_inside_quoted_realm_with_escaped_quotes(self) -> None:
"""``Bearer realm="foo scope=\\"admin:write\\" bar"`` returns ``()``.
The inner ``scope=`` is wholly inside the quoted-string value of
``realm`` there is no top-level ``scope`` auth-param, so the
result is empty.
"""
header = r'Bearer realm="foo scope=\"admin:write\" bar"'
assert parse_www_authenticate_scope(header) == ()
def test_scope_token_validation_drops_control_bytes(self) -> None:
"""Tokens containing CR / LF / tab / DEL / quote are dropped.
RFC 6749 §3.3 restricts ``scope-token`` to visible ASCII
excluding ``"`` and ``\\``. The splitter applies that
validation so a malicious AS cannot smuggle CRLF (or the like)
through a future log / notification path that prints the scope
list verbatim. ``"a\\rb"`` and ``"\\nc"`` fail validation;
``"d"`` survives. The legitimate space separator splits ``d``
into its own token.
"""
# Build via concatenation so the assertion stays intelligible.
header = 'Bearer scope="a\rb \nc d"'
assert parse_www_authenticate_scope(header) == ("d",)
class TestParseError:
def test_basic_quoted_error(self) -> None:
assert (
parse_www_authenticate_error('Bearer error="insufficient_scope"')
== "insufficient_scope"
)
def test_other_quoted_error_tokens(self) -> None:
assert parse_www_authenticate_error('Bearer error="invalid_token"') == "invalid_token"
assert parse_www_authenticate_error('Bearer error="invalid_request"') == "invalid_request"
def test_no_error_param(self) -> None:
assert parse_www_authenticate_error("Bearer realm=foo") is None
def test_error_description_does_not_match_error(self) -> None:
"""``error_description`` is its own auth-param key, not ``error``.
The tokenizer reads ``_`` as part of the token (RFC 7230 ``tchar``),
so ``error_description`` becomes one key, ``error`` another.
"""
assert parse_www_authenticate_error('Bearer error_description="bad"') is None
def test_unquoted_error(self) -> None:
# Some ASes don't quote the error token.
assert (
parse_www_authenticate_error("Bearer error=insufficient_scope") == "insufficient_scope"
)
def test_empty_string(self) -> None:
assert parse_www_authenticate_error("") is None
def test_error_inside_realm_value(self) -> None:
"""``Bearer realm="my error=fake", error="real"`` must return ``"real"``.
Naive parser grabbed ``fake`` from inside the ``realm`` quoted
value.
"""
header = 'Bearer realm="my error=fake", error="real"'
assert parse_www_authenticate_error(header) == "real"
class TestBearerDict:
def test_returns_lowercased_keys(self) -> None:
header = 'Bearer Realm="x", Error="y", Scope="a b"'
params = parse_www_authenticate_bearer(header)
assert params == {"realm": "x", "error": "y", "scope": "a b"}
def test_non_bearer_scheme_returns_empty(self) -> None:
assert parse_www_authenticate_bearer('Basic realm="x"') == {}
def test_no_scheme(self) -> None:
assert parse_www_authenticate_bearer('realm="x"') == {}
def test_bearer_only_no_params(self) -> None:
assert parse_www_authenticate_bearer("Bearer ") == {}
def test_bearer_with_no_space_returns_empty(self) -> None:
# ``BearerToken`` is not a Bearer challenge (no separator).
assert parse_www_authenticate_bearer("BearerToken") == {}
def test_first_value_wins_on_duplicate(self) -> None:
# If a malformed AS sends two ``scope=`` params we keep the first.
# The earlier ``find()``-based scanner would have returned the
# last; either choice is legal for malformed input but we need
# to be consistent.
header = 'Bearer scope="first", scope="second"'
assert parse_www_authenticate_bearer(header) == {"scope": "first"}
def test_trailing_comma(self) -> None:
header = 'Bearer error="x",'
assert parse_www_authenticate_bearer(header) == {"error": "x"}
def test_multiple_commas(self) -> None:
header = 'Bearer ,, error="x",,, scope="y"'
assert parse_www_authenticate_bearer(header) == {"error": "x", "scope": "y"}
def test_embedded_escaped_quote(self) -> None:
header = r'Bearer realm="he said \"hi\""'
assert parse_www_authenticate_bearer(header) == {"realm": 'he said "hi"'}
def test_param_without_value_skipped(self) -> None:
header = 'Bearer realm, error="x"'
# ``realm`` without ``=`` is dropped; ``error`` survives.
assert parse_www_authenticate_bearer(header) == {"error": "x"}
@pytest.mark.parametrize(
"header,expected",
[
("", {}),
("Bearer", {}),
('Bearer realm=""', {"realm": ""}),
('Bearer realm="", scope=""', {"realm": "", "scope": ""}),
],
)
def test_edge_cases(self, header: str, expected: dict[str, str]) -> None:
assert parse_www_authenticate_bearer(header) == expected
class TestPathologicalInput:
def test_oversized_pathological_input_rejected_under_50ms(self) -> None:
"""Headers longer than the defensive cap return ``{}`` immediately.
The cap is set to 4096 bytes real ASes emit a few hundred bytes
at most. This guards both ``parse_www_authenticate_bearer``
callers against pathological input from a misbehaving server.
The previous ``header.lower().find("scope", i)`` loop was
O(N**2) a 100 KB header with no ``=`` took ~330 ms because
each ``find`` rescanned the entire suffix. The single-pass
tokenizer (capped at 4 KB) reduces this to a one-shot length
check that returns ``{}`` in microseconds, so the budget is
generous regardless of which side of the cap was hit.
"""
big = "Bearer scope=" + "a" * 10_000
start = time.perf_counter()
result = parse_www_authenticate_scope(big)
elapsed = time.perf_counter() - start
assert result == ()
assert elapsed < 0.05, f"oversized-header reject took {elapsed * 1000:.1f}ms"
def test_within_cap_long_header_under_50ms(self) -> None:
"""A 4 KB header with thousands of ``find`` candidates still parses fast.
Stays under the cap so the tokenizer actually runs end to end
the goal is to prove the inner loop is O(N), not just that the
cap rejects oversized input.
"""
# Pack the header right up to the cap with non-matching
# auth-params, then put the real ``scope`` at the end.
filler_parts = []
size = len("Bearer ")
i = 0
while size < 3900:
part = f'xscope{i}="ignore", '
if size + len(part) > 3900:
break
filler_parts.append(part)
size += len(part)
i += 1
header = "Bearer " + "".join(filler_parts) + 'scope="real"'
assert len(header) <= 4096
start = time.perf_counter()
result = parse_www_authenticate_scope(header)
elapsed = time.perf_counter() - start
assert result == ("real",)
assert elapsed < 0.05, f"4kb tokenize took {elapsed * 1000:.1f}ms"
+81 -66
View File
@@ -14,6 +14,7 @@ from unittest.mock import AsyncMock, MagicMock
import pytest
from tests.conftest import _seed_static_state
from turnstone.core.mcp_client import MCPClientManager
from turnstone.core.storage._sqlite import SQLiteBackend
@@ -109,13 +110,15 @@ class TestFullLifecycleResourcesPrompts:
def test_rebuild_resources_produces_merged_state(self, mgr: MCPClientManager) -> None:
"""_rebuild_resources merges per-server resources into a unified list."""
mgr._per_server_resources["alpha"] = [
_make_resource("file:///a.txt", "a", "alpha"),
_make_resource("file:///b.txt", "b", "alpha"),
]
mgr._per_server_resources["beta"] = [
_make_resource("file:///c.txt", "c", "beta"),
]
_seed_static_state(
mgr,
"alpha",
resources=[
_make_resource("file:///a.txt", "a", "alpha"),
_make_resource("file:///b.txt", "b", "alpha"),
],
)
_seed_static_state(mgr, "beta", resources=[_make_resource("file:///c.txt", "c", "beta")])
mgr._rebuild_resources()
@@ -130,13 +133,19 @@ class TestFullLifecycleResourcesPrompts:
def test_rebuild_prompts_produces_merged_state(self, mgr: MCPClientManager) -> None:
"""_rebuild_prompts merges per-server prompts into a unified list."""
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__greet", "greet", "alpha", "Say hello"),
]
mgr._per_server_prompts["beta"] = [
_make_prompt("mcp__beta__summarize", "summarize", "beta", "Summarize text"),
_make_prompt("mcp__beta__translate", "translate", "beta", "Translate text"),
]
_seed_static_state(
mgr,
"alpha",
prompts=[_make_prompt("mcp__alpha__greet", "greet", "alpha", "Say hello")],
)
_seed_static_state(
mgr,
"beta",
prompts=[
_make_prompt("mcp__beta__summarize", "summarize", "beta", "Summarize text"),
_make_prompt("mcp__beta__translate", "translate", "beta", "Translate text"),
],
)
mgr._rebuild_prompts()
@@ -164,10 +173,12 @@ class TestFullLifecycleResourcesPrompts:
try:
# Populate session and resource map
session = _make_mock_session()
mgr._sessions["alpha"] = session
mgr._per_server_resources["alpha"] = [
_make_resource("file:///readme.md", "readme", "alpha"),
]
_seed_static_state(
mgr,
"alpha",
session=session,
resources=[_make_resource("file:///readme.md", "readme", "alpha")],
)
mgr._rebuild_resources()
result = mgr.read_resource_sync("file:///readme.md", timeout=5)
@@ -194,18 +205,22 @@ class TestFullLifecycleResourcesPrompts:
try:
session = _make_mock_session()
mgr._sessions["alpha"] = session
# Register a template resource (no concrete resources)
mgr._per_server_resources["alpha"] = [
{
"uri": "db://tables/{table}/rows/{id}",
"name": "row",
"description": "Fetch a row",
"mimeType": "application/json",
"server": "alpha",
"template": True,
},
]
_seed_static_state(
mgr,
"alpha",
session=session,
resources=[
{
"uri": "db://tables/{table}/rows/{id}",
"name": "row",
"description": "Fetch a row",
"mimeType": "application/json",
"server": "alpha",
"template": True,
},
],
)
mgr._rebuild_resources()
# Template should not be in _resource_map
@@ -230,10 +245,12 @@ class TestFullLifecycleResourcesPrompts:
try:
session = _make_mock_session()
mgr._sessions["alpha"] = session
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__greet", "greet", "alpha", "Say hello"),
]
_seed_static_state(
mgr,
"alpha",
session=session,
prompts=[_make_prompt("mcp__alpha__greet", "greet", "alpha", "Say hello")],
)
mgr._rebuild_prompts()
messages = mgr.get_prompt_sync(
@@ -314,40 +331,42 @@ class TestFullLifecycleResourcesPrompts:
def test_shutdown_clears_all_state(self, mgr: MCPClientManager) -> None:
"""shutdown() clears sessions, tools, resources, prompts, and listeners."""
# Populate state
mgr._sessions["alpha"] = MagicMock()
mgr._per_server_tools["alpha"] = [
{
"type": "function",
"function": {
"name": "mcp__alpha__search",
"description": "Search",
"parameters": {},
_seed_static_state(
mgr,
"alpha",
session=MagicMock(),
tools=[
{
"type": "function",
"function": {
"name": "mcp__alpha__search",
"description": "Search",
"parameters": {},
},
}
],
resources=[
_make_resource("file:///a.txt", "a", "alpha"),
{
"uri": "db://tables/{table}",
"name": "table",
"description": "",
"mimeType": "",
"server": "alpha",
"template": True,
},
}
]
],
prompts=[_make_prompt("mcp__alpha__greet", "greet", "alpha")],
)
mgr._rebuild_tools()
mgr._per_server_resources["alpha"] = [
_make_resource("file:///a.txt", "a", "alpha"),
{
"uri": "db://tables/{table}",
"name": "table",
"description": "",
"mimeType": "",
"server": "alpha",
"template": True,
},
]
mgr._rebuild_resources()
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__greet", "greet", "alpha"),
]
mgr._rebuild_prompts()
mgr._listeners.append(lambda: None)
mgr._resource_listeners.append(lambda: None)
mgr._prompt_listeners.append(lambda: None)
# Verify populated
assert len(mgr._sessions) == 1
assert len(mgr._static_servers) == 1
assert len(mgr._tools) == 1
assert len(mgr._resources) == 2 # 1 concrete + 1 template
assert len(mgr._template_prefixes) == 1
@@ -355,7 +374,7 @@ class TestFullLifecycleResourcesPrompts:
mgr.shutdown()
assert len(mgr._sessions) == 0
assert len(mgr._static_servers) == 0
assert len(mgr._tools) == 0
assert len(mgr._tool_map) == 0
assert len(mgr._resources) == 0
@@ -376,19 +395,15 @@ class TestFullLifecycleResourcesPrompts:
mgr.add_resource_listener(lambda: resource_fired.append(1))
mgr.add_prompt_listener(lambda: prompt_fired.append(1))
mgr._per_server_tools["alpha"] = []
_seed_static_state(mgr, "alpha", tools=[])
mgr._rebuild_tools()
assert len(tool_fired) == 1
mgr._per_server_resources["alpha"] = [
_make_resource("file:///x.txt", "x", "alpha"),
]
_seed_static_state(mgr, "alpha", resources=[_make_resource("file:///x.txt", "x", "alpha")])
mgr._rebuild_resources()
assert len(resource_fired) == 1
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__p1", "p1", "alpha"),
]
_seed_static_state(mgr, "alpha", prompts=[_make_prompt("mcp__alpha__p1", "p1", "alpha")])
mgr._rebuild_prompts()
assert len(prompt_fired) == 1
+780
View File
@@ -0,0 +1,780 @@
"""Integration tests for the MCP OAuth ``/connections`` endpoints.
Covers the list and revoke handlers that surface user-owned MCP server
consents to the settings UI:
* ``GET /v1/api/mcp/oauth/connections`` non-secret projection only.
* ``DELETE /v1/api/mcp/oauth/connections/{server_name}`` best-effort
upstream revoke (RFC 7009) followed by the authoritative local
delete; cross-user attempts return 404 with the exact same body
shape as a never-existed row to avoid leaking tenant existence.
"""
from __future__ import annotations
import asyncio
from typing import TYPE_CHECKING, Any
from unittest.mock import AsyncMock, MagicMock, patch
import httpx
import pytest
from starlette.applications import Starlette
from starlette.middleware import Middleware
from starlette.middleware.base import BaseHTTPMiddleware
from starlette.routing import Mount, Route
from starlette.testclient import TestClient
from tests.conftest import make_mcp_token_cipher
from turnstone.core.auth import AuthResult
from turnstone.core.mcp_crypto import MCPTokenStore
from turnstone.core.mcp_oauth import (
handle_mcp_oauth_list_connections,
handle_mcp_oauth_revoke_connection,
)
from turnstone.core.oidc import OIDCConfig
from turnstone.core.storage._sqlite import SQLiteBackend
if TYPE_CHECKING:
from starlette.requests import Request
from starlette.responses import Response
# ---------------------------------------------------------------------------
# Fixtures + helpers (mirror tests/test_mcp_oauth_handlers.py)
# ---------------------------------------------------------------------------
class _InjectAuthMiddleware(BaseHTTPMiddleware):
"""Stamp a fixed authenticated user on every request."""
def __init__(self, app: Any, user_id: str = "user-1") -> None:
super().__init__(app)
self._user_id = user_id
async def dispatch(self, request: Request, call_next: Any) -> Response:
request.state.auth_result = AuthResult(
user_id=self._user_id,
scopes=frozenset({"write"}),
token_source="config",
permissions=frozenset({"read", "write"}),
)
return await call_next(request)
class _NoAuthMiddleware(BaseHTTPMiddleware):
"""Leave ``request.state.auth_result`` unset so handlers see anon."""
async def dispatch(self, request: Request, call_next: Any) -> Response:
return await call_next(request)
async def _list_handler(request: Request) -> Response:
return await handle_mcp_oauth_list_connections(request)
async def _revoke_handler(request: Request) -> Response:
return await handle_mcp_oauth_revoke_connection(request)
def _build_app(
*,
storage: SQLiteBackend,
http_client: httpx.AsyncClient | MagicMock,
token_store: MCPTokenStore | None,
user_id: str = "user-1",
mcp_client: Any = None,
authenticated: bool = True,
) -> Starlette:
middleware: list[Middleware]
if authenticated:
middleware = [Middleware(_InjectAuthMiddleware, user_id=user_id)]
else:
middleware = [Middleware(_NoAuthMiddleware)]
app = Starlette(
routes=[
Mount(
"/v1",
routes=[
Route("/api/mcp/oauth/connections", _list_handler),
Route(
"/api/mcp/oauth/connections/{server_name}",
_revoke_handler,
methods=["DELETE"],
),
],
),
],
middleware=middleware,
)
app.state.auth_storage = storage
app.state.mcp_token_store = token_store
app.state.mcp_oauth_http_client = http_client
app.state.mcp_oauth_refresh_locks = {}
app.state.mcp_oauth_dcr_locks = {}
app.state.mcp_oauth_metadata_cache = {}
app.state.mcp_oauth_last_cleanup_monotonic = 0.0
app.state.oidc_config = OIDCConfig(enabled=False, redirect_base="https://testserver")
if mcp_client is not None:
app.state.mcp_client = mcp_client
return app
def _make_token_store(backend: SQLiteBackend) -> MCPTokenStore:
return MCPTokenStore(backend, make_mcp_token_cipher(), node_id="test")
def _seed_oauth_user_server(
backend: SQLiteBackend,
*,
name: str = "srv-oauth",
server_id: str = "srv-id-1",
cached_issuer: str | None = "https://as.example.com",
) -> str:
backend.create_mcp_server(
server_id=server_id,
name=name,
transport="streamable-http",
url="https://mcp.example.com/sse",
auth_type="oauth_user",
oauth_client_id="client-abc",
oauth_scopes="openid profile",
oauth_audience="https://mcp.example.com",
oauth_authorization_server_url=None,
)
if cached_issuer is not None:
backend.update_mcp_server(server_id, oauth_as_issuer_cached=cached_issuer)
return server_id
def _seed_user_token(
token_store: MCPTokenStore,
*,
user_id: str = "user-1",
server_name: str = "srv-oauth",
refresh_token: str | None = "refresh-secret",
) -> None:
token_store.create_user_token(
user_id,
server_name,
access_token="access-secret",
refresh_token=refresh_token,
expires_at="2099-12-31T00:00:00",
scopes="openid profile",
as_issuer="https://as.example.com",
audience="https://mcp.example.com",
)
def _good_as_metadata_doc(
*, revocation_endpoint: str | None = "https://as.example.com/revoke"
) -> dict[str, Any]:
doc: dict[str, Any] = {
"issuer": "https://as.example.com",
"authorization_endpoint": "https://as.example.com/authorize",
"token_endpoint": "https://as.example.com/token",
"registration_endpoint": "https://as.example.com/register",
"jwks_uri": "https://as.example.com/jwks",
"code_challenge_methods_supported": ["S256"],
"token_endpoint_auth_methods_supported": ["none", "client_secret_basic"],
}
if revocation_endpoint is not None:
doc["revocation_endpoint"] = revocation_endpoint
return doc
def _mk_response(
status_code: int = 200,
json_body: Any = None,
headers: dict[str, str] | None = None,
) -> MagicMock:
import json as _json
resp = MagicMock(spec=httpx.Response)
resp.status_code = status_code
resp.headers = headers or {}
body_str = _json.dumps(json_body) if json_body is not None else ""
resp.content = body_str.encode("utf-8")
if json_body is not None:
resp.json.return_value = json_body
else:
resp.json.side_effect = ValueError("no body")
resp.text = body_str
return resp
def _public_addr_patch():
return patch("socket.getaddrinfo", return_value=[(2, 1, 6, "", ("93.184.216.34", 0))])
def _drain_revoke_upstream_tasks(client: TestClient, timeout: float = 2.0) -> None:
"""Block until all in-flight upstream-revoke tasks complete.
Phase 8 perf-1 made the RFC 7009 AS round-trip a fire-and-forget
task so the user-visible 204 isn't gated on the AS. The tasks were
scheduled on the TestClient's portal loop; we re-enter that loop
via :attr:`TestClient.portal` to await them. Tests that assert
against the upstream POST must call this helper before the
assertion.
"""
from turnstone.core.mcp_oauth import _revoke_upstream_tasks
portal = getattr(client, "portal", None)
if portal is None:
return
async def _drain() -> None:
pending = list(_revoke_upstream_tasks)
if pending:
async with asyncio.timeout(timeout):
await asyncio.gather(*pending, return_exceptions=True)
portal.call(_drain)
@pytest.fixture
def storage(tmp_path: Any) -> SQLiteBackend:
backend = SQLiteBackend(str(tmp_path / "test.db"))
backend.create_user("user-1", "user1", "User One", "hash")
backend.create_user("user-2", "user2", "User Two", "hash")
return backend
@pytest.fixture
def http_client_mock() -> MagicMock:
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock()
client.post = AsyncMock()
return client
# ---------------------------------------------------------------------------
# GET /connections
# ---------------------------------------------------------------------------
class TestListConnections:
def test_list_connections_unauthenticated_401(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
token_store = _make_token_store(storage)
app = _build_app(
storage=storage,
http_client=http_client_mock,
token_store=token_store,
authenticated=False,
)
client = TestClient(app, raise_server_exceptions=False)
resp = client.get("/v1/api/mcp/oauth/connections")
assert resp.status_code == 401
assert resp.json() == {"error": "Authentication required"}
def test_list_connections_no_token_store_503(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
app = _build_app(storage=storage, http_client=http_client_mock, token_store=None)
client = TestClient(app, raise_server_exceptions=False)
resp = client.get("/v1/api/mcp/oauth/connections")
assert resp.status_code == 503
def test_list_connections_empty_user_returns_empty_list(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
token_store = _make_token_store(storage)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
resp = client.get("/v1/api/mcp/oauth/connections")
assert resp.status_code == 200
assert resp.json() == {"connections": []}
def test_list_connections_returns_users_consents(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage, name="srv-a", server_id="srv-id-a")
_seed_oauth_user_server(storage, name="srv-b", server_id="srv-id-b")
token_store = _make_token_store(storage)
_seed_user_token(token_store, server_name="srv-a")
_seed_user_token(token_store, server_name="srv-b")
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
resp = client.get("/v1/api/mcp/oauth/connections")
assert resp.status_code == 200
body = resp.json()
assert "connections" in body
servers = sorted(row["server_name"] for row in body["connections"])
assert servers == ["srv-a", "srv-b"]
def test_list_connections_isolates_by_user(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, user_id="user-1", server_name="srv-oauth")
_seed_user_token(token_store, user_id="user-2", server_name="srv-oauth")
# User-1 sees only user-1's row.
app = _build_app(
storage=storage, http_client=http_client_mock, token_store=token_store, user_id="user-1"
)
client = TestClient(app, raise_server_exceptions=False)
resp = client.get("/v1/api/mcp/oauth/connections")
rows = resp.json()["connections"]
assert all(row["user_id"] == "user-1" for row in rows)
assert len(rows) == 1
# User-2 sees only user-2's row.
app2 = _build_app(
storage=storage, http_client=http_client_mock, token_store=token_store, user_id="user-2"
)
client2 = TestClient(app2, raise_server_exceptions=False)
resp2 = client2.get("/v1/api/mcp/oauth/connections")
rows2 = resp2.json()["connections"]
assert all(row["user_id"] == "user-2" for row in rows2)
assert len(rows2) == 1
def test_list_connections_does_not_leak_secret_fields(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
resp = client.get("/v1/api/mcp/oauth/connections")
rows = resp.json()["connections"]
assert rows
for row in rows:
for forbidden in (
"access_token",
"refresh_token",
"access_token_ct",
"refresh_token_ct",
):
assert forbidden not in row, f"secret field {forbidden!r} leaked in {row!r}"
# ---------------------------------------------------------------------------
# DELETE /connections/{server_name}
# ---------------------------------------------------------------------------
class TestRevokeConnection:
def test_revoke_connection_unauthenticated_401(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store)
app = _build_app(
storage=storage,
http_client=http_client_mock,
token_store=token_store,
authenticated=False,
)
client = TestClient(app, raise_server_exceptions=False)
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 401
def test_revoke_connection_missing_row_404(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
token_store = _make_token_store(storage)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
resp = client.delete("/v1/api/mcp/oauth/connections/srv-nonexistent")
assert resp.status_code == 404
assert resp.json() == {"error": "No such connection"}
def test_revoke_connection_local_delete_succeeds_204(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
# No refresh token → upstream revoke is skipped entirely.
_seed_user_token(token_store, refresh_token=None)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
# Local row is gone.
assert token_store.get_user_token("user-1", "srv-oauth") is None
# Upstream not contacted.
http_client_mock.post.assert_not_called()
def test_revoke_connection_with_revocation_endpoint_calls_upstream(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token="refresh-secret")
http_client_mock.get.return_value = _mk_response(200, _good_as_metadata_doc())
http_client_mock.post.return_value = _mk_response(200)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
# ``with TestClient(...)`` keeps a persistent portal so the
# fire-and-forget upstream-revoke task isn't cancelled when
# the request handler returns. See ``_drain_revoke_upstream_tasks``.
# The SSRF-validator's ``socket.getaddrinfo`` patch must wrap
# the drain too — the discovery call now runs on the background
# task and resolves the AS hostname after the request returns.
with (
TestClient(app, raise_server_exceptions=False) as client,
_public_addr_patch(),
):
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
# Local row is gone.
assert token_store.get_user_token("user-1", "srv-oauth") is None
# The upstream RFC 7009 POST is fire-and-forget post-Phase-8 perf-1
# so the test must drain the in-flight task set before asserting.
_drain_revoke_upstream_tasks(client)
# Upstream POSTed to revocation_endpoint with refresh-token grant.
assert http_client_mock.post.await_count == 1
call = http_client_mock.post.await_args
assert call.args[0] == "https://as.example.com/revoke"
data = call.kwargs.get("data") or {}
assert data.get("token") == "refresh-secret"
assert data.get("token_type_hint") == "refresh_token"
assert data.get("client_id") == "client-abc"
def test_revoke_connection_without_revocation_endpoint_skips_upstream(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token="refresh-secret")
http_client_mock.get.return_value = _mk_response(
200, _good_as_metadata_doc(revocation_endpoint=None)
)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
# ``with TestClient(...)`` keeps the portal alive for the
# background task drain.
with (
TestClient(app, raise_server_exceptions=False) as client,
_public_addr_patch(),
):
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
# Local row gone, upstream POST never made.
assert token_store.get_user_token("user-1", "srv-oauth") is None
# Drain the fire-and-forget discovery task before asserting on
# the AS POST — the task runs ``discover_authorization_server``
# but does NOT proceed to POST because revocation_endpoint is
# absent.
_drain_revoke_upstream_tasks(client)
http_client_mock.post.assert_not_called()
def test_revoke_connection_upstream_failure_still_204(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token="refresh-secret")
http_client_mock.get.return_value = _mk_response(200, _good_as_metadata_doc())
# AS returns 500 — local delete must still succeed.
http_client_mock.post.return_value = _mk_response(500)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
with _public_addr_patch():
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
assert token_store.get_user_token("user-1", "srv-oauth") is None
def test_revoke_connection_audit_event_emitted_with_user_revoked_reason(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token=None)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
# Audit row was written via the storage API (tests don't poke at
# the SQLite schema directly — the table name is an internal
# detail).
events = storage.list_audit_events(action="mcp_server.oauth.token_revoked")
assert len(events) == 1
ev = events[0]
assert ev["user_id"] == "user-1"
# resource_id is the immutable server_id PK, not the name.
assert ev["resource_id"] == "srv-id-1"
import json as _json
detail = _json.loads(ev["detail"]) if isinstance(ev["detail"], str) else ev["detail"]
assert detail["reason"] == "user_revoked"
assert detail["upstream_revoke_outcome"] == "no_refresh_token"
assert detail["server_name"] == "srv-oauth"
def test_revoke_connection_cross_user_attempt_404(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
# Owned by user-2, not user-1.
_seed_user_token(token_store, user_id="user-2", server_name="srv-oauth")
app = _build_app(
storage=storage, http_client=http_client_mock, token_store=token_store, user_id="user-1"
)
client = TestClient(app, raise_server_exceptions=False)
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
# Cross-user attempt MUST surface as a generic 404, byte-identical
# body to the never-existed case (no tenant existence leak).
assert resp.status_code == 404
assert resp.json() == {"error": "No such connection"}
# Drain pending tasks defensively, then confirm the upstream
# endpoint was NEVER contacted on the 404-cross-user path. A
# bug that scheduled the AS round-trip before the cross-user
# check would leak existence via the AS-side 200/4xx response.
_drain_revoke_upstream_tasks(client)
http_client_mock.post.assert_not_called()
# User-2's row is untouched.
assert token_store.get_user_token("user-2", "srv-oauth") is not None
def test_revoke_connection_evicts_pool_session(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token=None)
mcp_client_mock = MagicMock()
# ``evict_user_session`` is the public sync surface on
# MCPClientManager; mirror its signature here so the handler's
# ``hasattr`` gate triggers.
mcp_client_mock.evict_user_session = MagicMock(return_value=None)
app = _build_app(
storage=storage,
http_client=http_client_mock,
token_store=token_store,
mcp_client=mcp_client_mock,
)
client = TestClient(app, raise_server_exceptions=False)
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
mcp_client_mock.evict_user_session.assert_called_once_with("user-1", "srv-oauth")
def test_revoke_connection_pool_eviction_failure_does_not_block_204(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token=None)
mcp_client_mock = MagicMock()
mcp_client_mock.evict_user_session = MagicMock(side_effect=RuntimeError("loop closed"))
app = _build_app(
storage=storage,
http_client=http_client_mock,
token_store=token_store,
mcp_client=mcp_client_mock,
)
client = TestClient(app, raise_server_exceptions=False)
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
# Local delete still happened.
assert token_store.get_user_token("user-1", "srv-oauth") is None
def test_revoke_connection_204_not_gated_on_slow_upstream(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
"""The user-visible 204 must return promptly even when the
upstream AS round-trip is slow / hanging. Pre-perf-1 the
handler awaited ``revoke_token_at_as`` synchronously, so a
stuck AS could block the user's revoke confirmation. The
fire-and-forget refactor moves the call onto a background task
so the 204 returns in well under 1s regardless of AS latency.
Bound is conservative for CI runner jitter.
"""
import time
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token="refresh-secret")
http_client_mock.get.return_value = _mk_response(200, _good_as_metadata_doc())
async def _slow_post(*_args: Any, **_kwargs: Any) -> Any:
# Simulate a slow / unreachable AS — must NOT gate the
# user-visible 204 on this round-trip.
await asyncio.sleep(5.0)
return _mk_response(200)
http_client_mock.post = AsyncMock(side_effect=_slow_post)
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
client = TestClient(app, raise_server_exceptions=False)
with _public_addr_patch():
start = time.monotonic()
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
elapsed = time.monotonic() - start
assert resp.status_code == 204
# 1s ceiling — the 204 must return on the local-delete path
# without waiting on the AS POST (which sleeps 5s above). Bound
# is intentionally generous for CI runner jitter; the actual
# path is on the order of milliseconds.
assert elapsed < 1.0, (
f"204 returned in {elapsed:.3f}s — should be <1s; the "
"fire-and-forget upstream revoke isn't decoupled from the "
"response."
)
# The local row IS gone — the authoritative delete ran before
# the 204 returned, even though the AS round-trip is still
# in flight.
assert token_store.get_user_token("user-1", "srv-oauth") is None
# Cancel any in-flight tasks so the test client can exit cleanly.
from turnstone.core.mcp_oauth import _revoke_upstream_tasks
portal = getattr(client, "portal", None)
if portal is not None:
for task in list(_revoke_upstream_tasks):
portal.call(task.cancel)
def test_revoke_connection_sheds_upstream_when_task_set_full(
self, storage: SQLiteBackend, http_client_mock: MagicMock
) -> None:
"""Round-2 q-2 regression: the soft cap on ``_revoke_upstream_tasks``
is the only protection against unbounded background-task pile-up
under a coordinated mass-revoke. When the set is full, the local
delete still runs but no upstream task is scheduled; the audit
detail records ``upstream_revoke_outcome="shed_by_cap"`` and
the AS endpoint is never contacted.
"""
from turnstone.core.mcp_oauth import (
_REVOKE_UPSTREAM_TASKS_MAX,
_revoke_upstream_tasks,
)
_seed_oauth_user_server(storage)
token_store = _make_token_store(storage)
_seed_user_token(token_store, refresh_token="refresh-secret")
app = _build_app(storage=storage, http_client=http_client_mock, token_store=token_store)
sentinel_event_holder: dict[str, asyncio.Event] = {}
# Use ``with TestClient(...)`` so the portal stays alive — we
# need to schedule sentinel tasks on the portal's loop and the
# tasks must outlive the request to actually fill the set.
with (
TestClient(app, raise_server_exceptions=False) as client,
_public_addr_patch(),
):
portal = client.portal
assert portal is not None
async def _create_sentinel_event() -> asyncio.Event:
event = asyncio.Event()
sentinel_event_holder["event"] = event
return event
sentinel_event = portal.call(_create_sentinel_event)
async def _wait_on_event() -> None:
await sentinel_event.wait()
async def _fill_task_set() -> list[asyncio.Task[None]]:
tasks: list[asyncio.Task[None]] = []
for _ in range(_REVOKE_UPSTREAM_TASKS_MAX):
t = asyncio.create_task(_wait_on_event())
_revoke_upstream_tasks.add(t)
tasks.append(t)
return tasks
sentinels = portal.call(_fill_task_set)
assert len(_revoke_upstream_tasks) >= _REVOKE_UPSTREAM_TASKS_MAX
try:
resp = client.delete("/v1/api/mcp/oauth/connections/srv-oauth")
assert resp.status_code == 204
# Local row is still gone — authoritative delete ran.
assert token_store.get_user_token("user-1", "srv-oauth") is None
# AS endpoint MUST NOT have been contacted.
http_client_mock.post.assert_not_called()
# Audit detail records the categorical shed outcome.
events = storage.list_audit_events(action="mcp_server.oauth.token_revoked")
assert len(events) == 1
detail = events[0]["detail"]
if isinstance(detail, str):
import json as _json
detail = _json.loads(detail)
assert detail["upstream_revoke_outcome"] == "shed_by_cap"
finally:
# Release sentinels so the portal can shut down cleanly.
async def _release() -> None:
sentinel_event.set()
for t in sentinels:
t.cancel()
await asyncio.gather(*sentinels, return_exceptions=True)
portal.call(_release)
# ---------------------------------------------------------------------------
# evict_user_session helper sanity checks
# ---------------------------------------------------------------------------
class TestEvictUserSession:
def test_evict_user_session_no_loop_is_silent_noop(self) -> None:
from turnstone.core.mcp_client import MCPClientManager
mgr = MCPClientManager.__new__(MCPClientManager)
mgr._loop = None # type: ignore[attr-defined]
# Must not raise.
mgr.evict_user_session("user-1", "srv-oauth")
def test_evict_user_session_dispatches_to_loop(self) -> None:
from turnstone.core.mcp_client import MCPClientManager
mgr = MCPClientManager.__new__(MCPClientManager)
loop = asyncio.new_event_loop()
try:
mgr._loop = loop # type: ignore[attr-defined]
mgr._user_pool_entries = {} # type: ignore[attr-defined]
mgr._last_pool_notification_refresh = {} # type: ignore[attr-defined]
evicted: list[tuple[str, str]] = []
def _fake_evict(key: tuple[str, str]) -> None:
evicted.append(key)
mgr._evict_session = _fake_evict # type: ignore[method-assign]
# Run the dispatch on a separate thread so the loop can drain.
import threading
done = threading.Event()
def _run_loop() -> None:
loop.call_later(0.05, loop.stop)
loop.run_forever()
done.set()
t = threading.Thread(target=_run_loop, daemon=True)
t.start()
mgr.evict_user_session("user-1", "srv-oauth")
done.wait(timeout=1.0)
assert evicted == [("user-1", "srv-oauth")]
finally:
if not loop.is_closed():
loop.close()
+626
View File
@@ -0,0 +1,626 @@
"""Discovery tests for the per-(user, server) MCP OAuth flow.
Covers PRM (RFC 9728) and AS metadata (RFC 8414) discovery, including:
- override URL takes precedence
- PRM happy path: server URL -> .well-known/oauth-protected-resource
-> ``authorization_servers[0]``
- PRM 401 + ``WWW-Authenticate: Bearer resource_metadata="..."`` follows
the URL.
- AS metadata without S256 -> :class:`MCPOAuthDiscoveryError`.
- SSRF rejection on AS issuer URL.
- In-memory cache hit/miss + persistent cache write to
``mcp_servers.oauth_as_issuer_cached``.
"""
from __future__ import annotations
import asyncio
import time
from typing import Any
from unittest.mock import AsyncMock, MagicMock, patch
import httpx
import pytest
from turnstone.core.mcp_oauth import (
ASMetadata,
MCPOAuthDiscoveryError,
_parse_prm_url_from_www_authenticate,
discover_authorization_server,
)
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _mk_response(
status_code: int = 200,
json_body: Any = None,
headers: dict[str, str] | None = None,
) -> MagicMock:
"""Build a MagicMock that quacks like ``httpx.Response``."""
resp = MagicMock(spec=httpx.Response)
resp.status_code = status_code
resp.headers = headers or {}
resp.content = (str(json_body) if json_body is not None else "").encode("utf-8")
if json_body is not None:
resp.json.return_value = json_body
else:
resp.json.side_effect = ValueError("no body")
resp.text = str(json_body) if json_body is not None else ""
return resp
def _good_as_metadata_doc() -> dict[str, Any]:
return {
"issuer": "https://as.example.com",
"authorization_endpoint": "https://as.example.com/authorize",
"token_endpoint": "https://as.example.com/token",
"jwks_uri": "https://as.example.com/jwks",
"code_challenge_methods_supported": ["S256"],
"token_endpoint_auth_methods_supported": ["none", "client_secret_basic"],
"registration_endpoint": "https://as.example.com/register",
}
def _public_addr_patch():
return patch("socket.getaddrinfo", return_value=[(2, 1, 6, "", ("93.184.216.34", 0))])
def _mk_storage_mock(server_id: str = "srv-id") -> MagicMock:
storage = MagicMock()
storage.update_mcp_server.return_value = True
return storage
# ---------------------------------------------------------------------------
# PRM parser
# ---------------------------------------------------------------------------
class TestParsePRMUrl:
def test_extracts_resource_metadata_url(self) -> None:
header = (
'Bearer error="invalid_token", '
'resource_metadata="https://srv.example.com/.well-known/oauth-protected-resource"'
)
url = _parse_prm_url_from_www_authenticate(header)
assert url == "https://srv.example.com/.well-known/oauth-protected-resource"
def test_returns_none_when_absent(self) -> None:
assert _parse_prm_url_from_www_authenticate('Bearer realm="x"') is None
def test_handles_empty_header(self) -> None:
assert _parse_prm_url_from_www_authenticate("") is None
def test_handles_escaped_quote_in_value(self) -> None:
"""RFC 7230 quoted-string allows ``\\"`` — naive ``[^"]+`` truncates.
A malicious or buggy resource server could send an embedded
escaped quote; the parser must yield the unescaped value, not
the prefix up to the escaped quote.
"""
header = 'Bearer resource_metadata="https://srv.example.com/with\\"quote"'
url = _parse_prm_url_from_www_authenticate(header)
assert url == 'https://srv.example.com/with"quote'
def test_handles_escaped_backslash(self) -> None:
header = 'Bearer resource_metadata="https://srv.example.com/back\\\\slash"'
url = _parse_prm_url_from_www_authenticate(header)
assert url == "https://srv.example.com/back\\slash"
def test_unterminated_quoted_string_returns_none(self) -> None:
# Closing quote missing — naive regex would still match, but
# the proper parser should reject malformed input.
header = 'Bearer resource_metadata="https://srv.example.com/no-close'
assert _parse_prm_url_from_www_authenticate(header) is None
# ---------------------------------------------------------------------------
# discover_authorization_server happy paths
# ---------------------------------------------------------------------------
class TestDiscoveryOverride:
def test_override_url_skips_prm(self) -> None:
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, _good_as_metadata_doc()))
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
meta = asyncio.run(_run())
assert isinstance(meta, ASMetadata)
assert meta.token_endpoint == "https://as.example.com/token"
# Only the AS metadata URL was hit, not PRM.
called_urls = [c.args[0] for c in client.get.call_args_list]
assert all("oauth-authorization-server" in u for u in called_urls)
class TestDiscoveryPRM:
def test_prm_happy_path(self) -> None:
async def _get(url, *args, **kwargs):
if url.endswith("/oauth-protected-resource"):
return _mk_response(
200,
{
"resource": "https://mcp.example.com",
"authorization_servers": ["https://as.example.com"],
},
)
if url.endswith("/oauth-authorization-server"):
return _mk_response(200, _good_as_metadata_doc())
raise AssertionError(f"unexpected URL: {url}")
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(side_effect=_get)
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url=None,
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
meta = asyncio.run(_run())
assert meta.issuer == "https://as.example.com"
def test_prm_401_follows_www_authenticate(self) -> None:
async def _get(url, *args, **kwargs):
if url == "https://mcp.example.com/.well-known/oauth-protected-resource":
return _mk_response(
401,
headers={
"www-authenticate": (
'Bearer error="invalid_token", '
"resource_metadata="
'"https://meta.example.com/prm"'
)
},
json_body=None,
)
if url == "https://meta.example.com/prm":
return _mk_response(
200,
{
"authorization_servers": ["https://as.example.com"],
},
)
if url.endswith("/oauth-authorization-server"):
return _mk_response(200, _good_as_metadata_doc())
raise AssertionError(f"unexpected URL: {url}")
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(side_effect=_get)
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url=None,
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
meta = asyncio.run(_run())
assert meta.token_endpoint == "https://as.example.com/token"
def test_prm_401_without_resource_metadata_raises(self) -> None:
async def _get(url, *args, **kwargs):
return _mk_response(401, headers={"www-authenticate": "Basic realm=x"})
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(side_effect=_get)
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url=None,
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
with pytest.raises(MCPOAuthDiscoveryError, match="resource_metadata"):
asyncio.run(_run())
# ---------------------------------------------------------------------------
# AS metadata validation
# ---------------------------------------------------------------------------
class TestASMetadataValidation:
def test_no_s256_raises(self) -> None:
doc = _good_as_metadata_doc()
doc["code_challenge_methods_supported"] = ["plain"]
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, doc))
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
with pytest.raises(MCPOAuthDiscoveryError, match="S256"):
asyncio.run(_run())
def test_missing_endpoints_raises(self) -> None:
doc = _good_as_metadata_doc()
del doc["token_endpoint"]
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, doc))
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
with pytest.raises(MCPOAuthDiscoveryError, match="missing required"):
asyncio.run(_run())
def test_third_party_endpoint_rejected(self) -> None:
doc = _good_as_metadata_doc()
doc["token_endpoint"] = "https://attacker.example.com/token"
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, doc))
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
with pytest.raises(MCPOAuthDiscoveryError, match="token_endpoint"):
asyncio.run(_run())
def test_ssrf_on_override_rejected(self) -> None:
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock()
storage = _mk_storage_mock()
async def _run():
# Resolve to private 10.x — SSRF guard fires before any HTTP call.
with patch(
"socket.getaddrinfo",
return_value=[(2, 1, 6, "", ("10.0.0.1", 0))],
):
await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://internal.corp.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
with pytest.raises(MCPOAuthDiscoveryError):
asyncio.run(_run())
client.get.assert_not_called()
# ---------------------------------------------------------------------------
# Caching
# ---------------------------------------------------------------------------
class TestMetadataCache:
def test_cache_miss_then_hit(self) -> None:
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, _good_as_metadata_doc()))
storage = _mk_storage_mock()
cache: dict[str, tuple[ASMetadata, float]] = {}
async def _run():
with _public_addr_patch():
first = await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
metadata_cache=cache,
)
second = await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer="https://as.example.com",
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
metadata_cache=cache,
)
return first, second
first, second = asyncio.run(_run())
assert first.token_endpoint == second.token_endpoint
# First call hit AS metadata; second call hit the cache.
assert client.get.call_count == 1
def test_cache_expiry_refetches(self) -> None:
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, _good_as_metadata_doc()))
storage = _mk_storage_mock()
# Pre-populate cache with a very stale entry.
stale_meta = ASMetadata(
issuer="https://as.example.com",
authorization_endpoint="https://as.example.com/authorize",
token_endpoint="https://as.example.com/token",
registration_endpoint=None,
revocation_endpoint=None,
jwks_uri=None,
code_challenge_methods_supported=("S256",),
token_endpoint_auth_methods_supported=(),
)
cache = {"https://as.example.com": (stale_meta, time.monotonic() - 10**6)}
async def _run():
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
metadata_cache=cache,
)
meta = asyncio.run(_run())
# Stale entry was bypassed -> we hit the network.
assert client.get.call_count == 1
assert meta.token_endpoint == "https://as.example.com/token"
def test_persistent_cache_write_on_first_resolution(self) -> None:
async def _get(url, *args, **kwargs):
if url.endswith("/oauth-protected-resource"):
return _mk_response(200, {"authorization_servers": ["https://as.example.com"]})
return _mk_response(200, _good_as_metadata_doc())
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(side_effect=_get)
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url=None,
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
asyncio.run(_run())
# update_mcp_server was called once with the cached issuer.
storage.update_mcp_server.assert_called_once_with(
"srv-id", oauth_as_issuer_cached="https://as.example.com"
)
def test_persistent_cache_skip_when_already_cached(self) -> None:
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, _good_as_metadata_doc()))
storage = _mk_storage_mock()
async def _run():
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url=None,
cached_issuer="https://as.example.com",
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
asyncio.run(_run())
storage.update_mcp_server.assert_not_called()
# ---------------------------------------------------------------------------
# sec-3 — cached_issuer re-validated on read
# ---------------------------------------------------------------------------
class TestCachedIssuerSSRFRevalidation:
"""A cached issuer URL must still pass SSRF validation on every read.
Defense-in-depth: an admin who points ``oauth_as_issuer_cached`` at a
private address (or a hostname that has rebound to one) should not
bypass the guard just because the value was already in the row.
"""
def test_cached_issuer_rejected_clears_row_and_falls_through_to_prm(self) -> None:
async def _get(url: str, *args: Any, **kwargs: Any) -> MagicMock:
if url.endswith("/oauth-protected-resource"):
return _mk_response(200, {"authorization_servers": ["https://as.example.com"]})
if url.endswith("/oauth-authorization-server"):
return _mk_response(200, _good_as_metadata_doc())
raise AssertionError(f"unexpected URL: {url}")
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(side_effect=_get)
storage = _mk_storage_mock()
# cached_issuer points at a private host. SSRF guard fires on
# the cached value first, the row is cleared, and PRM
# discovery runs as a fallback.
async def _run() -> Any:
with patch(
"socket.getaddrinfo",
# Private resolution for "internal.corp", public for everything else.
side_effect=lambda host, *a, **kw: [
(2, 1, 6, "", ("10.0.0.1" if "internal" in host else "93.184.216.34", 0))
],
):
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url=None,
cached_issuer="https://internal.corp.example.com",
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
meta = asyncio.run(_run())
assert meta.token_endpoint == "https://as.example.com/token"
# The bad cached_issuer was cleared from the row.
clear_calls = [
c
for c in storage.update_mcp_server.call_args_list
if c.kwargs.get("oauth_as_issuer_cached") is None
]
assert clear_calls, "cached_issuer should have been cleared"
# ---------------------------------------------------------------------------
# revocation_endpoint parsing (RFC 8414)
# ---------------------------------------------------------------------------
class TestASMetadataRevocationEndpoint:
def test_as_metadata_parses_revocation_endpoint(self) -> None:
doc = _good_as_metadata_doc()
doc["revocation_endpoint"] = "https://as.example.com/revoke"
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, doc))
storage = _mk_storage_mock()
async def _run() -> ASMetadata:
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
meta = asyncio.run(_run())
assert meta.revocation_endpoint == "https://as.example.com/revoke"
def test_as_metadata_revocation_endpoint_absent(self) -> None:
doc = _good_as_metadata_doc()
doc.pop("revocation_endpoint", None)
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, doc))
storage = _mk_storage_mock()
async def _run() -> ASMetadata:
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
meta = asyncio.run(_run())
assert meta.revocation_endpoint is None
def test_as_metadata_revocation_endpoint_rejected_when_cross_origin(self) -> None:
doc = _good_as_metadata_doc()
doc["revocation_endpoint"] = "https://attacker.example.com/revoke"
client = MagicMock(spec=httpx.AsyncClient)
client.get = AsyncMock(return_value=_mk_response(200, doc))
storage = _mk_storage_mock()
async def _run() -> ASMetadata:
with _public_addr_patch():
return await discover_authorization_server(
server_name="srv-x",
server_url="https://mcp.example.com/sse",
override_url="https://as.example.com",
cached_issuer=None,
http_client=client,
storage=storage,
server_id="srv-id",
trusted_hosts=frozenset(),
)
with pytest.raises(MCPOAuthDiscoveryError, match="revocation_endpoint"):
asyncio.run(_run())
File diff suppressed because it is too large Load Diff
+133
View File
@@ -0,0 +1,133 @@
"""Smoke tests for the new OAuth-MCP storage tables.
Phase 2 only adds the schema token CRUD lands in Phase 3 and pending-
state CRUD in Phase 4. These tests verify the tables exist after
``init_storage`` and accept the documented row shape via raw SQL.
"""
from __future__ import annotations
import sqlalchemy as sa
from turnstone.core.storage._schema import mcp_oauth_pending, mcp_user_tokens
class TestMcpUserTokensTable:
def test_table_exists_and_accepts_row(self, backend) -> None:
with backend._engine.connect() as conn:
conn.execute(
sa.insert(mcp_user_tokens),
{
"user_id": "u1",
"server_name": "srv-a",
"access_token_ct": b"\x00ciphertext-a",
"refresh_token_ct": b"\x00ciphertext-r",
"expires_at": "2026-05-04T12:00:00",
"scopes": "openid profile",
"as_issuer": "https://auth.example.com",
"audience": "https://mcp.example.com",
"created": "2026-05-04T11:00:00",
"last_refreshed": None,
},
)
conn.commit()
row = conn.execute(
sa.select(mcp_user_tokens).where(
(mcp_user_tokens.c.user_id == "u1") & (mcp_user_tokens.c.server_name == "srv-a")
)
).one()
assert row.access_token_ct == b"\x00ciphertext-a"
assert row.refresh_token_ct == b"\x00ciphertext-r"
assert row.scopes == "openid profile"
assert row.audience == "https://mcp.example.com"
def test_composite_pk_distinguishes_user_server(self, backend) -> None:
"""Same user, different server => two rows; same (user, server) => conflict."""
with backend._engine.connect() as conn:
conn.execute(
sa.insert(mcp_user_tokens),
[
{
"user_id": "u1",
"server_name": "srv-a",
"access_token_ct": b"a",
"refresh_token_ct": None,
"expires_at": None,
"scopes": None,
"as_issuer": "https://auth.example.com",
"audience": "https://a.example.com",
"created": "2026-05-04T11:00:00",
"last_refreshed": None,
},
{
"user_id": "u1",
"server_name": "srv-b",
"access_token_ct": b"b",
"refresh_token_ct": None,
"expires_at": None,
"scopes": None,
"as_issuer": "https://auth.example.com",
"audience": "https://b.example.com",
"created": "2026-05-04T11:00:00",
"last_refreshed": None,
},
],
)
conn.commit()
count = conn.execute(sa.select(sa.func.count()).select_from(mcp_user_tokens)).scalar()
assert count == 2
class TestMcpOauthPendingTable:
def test_table_exists_and_accepts_row(self, backend) -> None:
with backend._engine.connect() as conn:
conn.execute(
sa.insert(mcp_oauth_pending),
{
"state": "rand-state-xyz",
"user_id": "u1",
"server_name": "srv-a",
"code_verifier": "verifier-blob",
"return_url": "/admin/mcp-servers",
"created_at": "2026-05-04T11:00:00",
},
)
conn.commit()
row = conn.execute(
sa.select(mcp_oauth_pending).where(mcp_oauth_pending.c.state == "rand-state-xyz")
).one()
assert row.user_id == "u1"
assert row.server_name == "srv-a"
assert row.return_url == "/admin/mcp-servers"
def test_state_pk_unique(self, backend) -> None:
"""A second insert with the same state value raises IntegrityError."""
with backend._engine.connect() as conn:
conn.execute(
sa.insert(mcp_oauth_pending),
{
"state": "dup-state",
"user_id": "u1",
"server_name": "srv-a",
"code_verifier": "v",
"return_url": "/x",
"created_at": "2026-05-04T11:00:00",
},
)
conn.commit()
import pytest
from sqlalchemy.exc import IntegrityError
with pytest.raises(IntegrityError), backend._engine.connect() as conn:
conn.execute(
sa.insert(mcp_oauth_pending),
{
"state": "dup-state",
"user_id": "u2",
"server_name": "srv-b",
"code_verifier": "v",
"return_url": "/y",
"created_at": "2026-05-04T11:01:00",
},
)
conn.commit()
+53
View File
@@ -0,0 +1,53 @@
"""PKCE pair-generation tests for the MCP OAuth flow.
Verifies the contract documented in RFC 7636 §4.1 and §4.2:
- ``code_verifier`` is a high-entropy 43..128 character urlsafe-base64 string.
- ``code_challenge`` is the BASE64URL-NO-PADDING encoding of
``SHA256(verifier)``.
"""
from __future__ import annotations
import base64
import hashlib
import string
from turnstone.core.mcp_oauth import generate_pkce_pair
_URLSAFE_CHARS = set(string.ascii_letters + string.digits + "-_")
class TestGeneratePkcePair:
def test_returns_tuple_of_strings(self) -> None:
verifier, challenge = generate_pkce_pair()
assert isinstance(verifier, str)
assert isinstance(challenge, str)
def test_verifier_length_in_rfc_range(self) -> None:
for _ in range(20):
verifier, _ = generate_pkce_pair()
assert 43 <= len(verifier) <= 128
def test_verifier_is_urlsafe(self) -> None:
for _ in range(20):
verifier, _ = generate_pkce_pair()
assert all(ch in _URLSAFE_CHARS for ch in verifier)
def test_challenge_matches_sha256_of_verifier(self) -> None:
for _ in range(20):
verifier, challenge = generate_pkce_pair()
digest = hashlib.sha256(verifier.encode("ascii")).digest()
expected = base64.urlsafe_b64encode(digest).rstrip(b"=").decode("ascii")
assert challenge == expected
def test_challenge_has_no_padding(self) -> None:
for _ in range(20):
_, challenge = generate_pkce_pair()
assert "=" not in challenge
def test_pairs_are_unique(self) -> None:
pairs = {generate_pkce_pair() for _ in range(50)}
# 50 random draws shouldn't collide; if they do we have a much
# bigger problem than this assertion.
assert len(pairs) == 50
File diff suppressed because it is too large Load Diff
+399
View File
@@ -0,0 +1,399 @@
"""Tests for :func:`turnstone.core.mcp_oauth.revoke_token_at_as`.
The helper is best-effort RFC 7009 token revocation. It must:
- skip cleanly when the AS metadata doesn't advertise a revocation endpoint
- POST the form body when one is present (with optional client_secret)
- never raise on non-2xx, network errors, or timeouts caller doesn't
want try/except in cleanup paths
- never use ``exc_info=True`` chained ``__context__`` may carry an
``httpx.Request`` whose ``Authorization`` header holds a bearer; the
bearer-leak invariant requires structured fields with type names only
"""
from __future__ import annotations
import asyncio
from typing import Any
from unittest.mock import AsyncMock, MagicMock, patch
import httpx
from turnstone.core.mcp_oauth import (
ASMetadata,
MCPOAuthDiscoveryError,
_attempt_upstream_revoke,
revoke_token_at_as,
)
def _make_as_metadata(
*,
revocation_endpoint: str | None = "https://as.example.com/revoke",
) -> ASMetadata:
return ASMetadata(
issuer="https://as.example.com",
authorization_endpoint="https://as.example.com/authorize",
token_endpoint="https://as.example.com/token",
registration_endpoint=None,
revocation_endpoint=revocation_endpoint,
jwks_uri=None,
code_challenge_methods_supported=("S256",),
token_endpoint_auth_methods_supported=("client_secret_basic",),
)
def _mk_response(status_code: int) -> MagicMock:
resp = MagicMock(spec=httpx.Response)
resp.status_code = status_code
return resp
class TestRevocationUnsupported:
def test_revoke_token_skipped_when_revocation_endpoint_none(self) -> None:
as_meta = _make_as_metadata(revocation_endpoint=None)
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock()
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
)
)
client.post.assert_not_called()
info_events = [c.args[0] for c in mock_log.info.call_args_list]
assert "mcp_server.oauth.revocation_unsupported" in info_events
class TestRevocationSuccess:
def test_revoke_token_succeeds_on_200(self) -> None:
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock(return_value=_mk_response(200))
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret="s-secret",
)
)
# POST shape — URL + form body keys.
client.post.assert_awaited_once()
call_args = client.post.call_args
assert call_args.args[0] == "https://as.example.com/revoke"
body = call_args.kwargs["data"]
assert body == {
"token": "r-secret",
"token_type_hint": "refresh_token",
"client_id": "client-1",
"client_secret": "s-secret",
}
info_events = [c.args[0] for c in mock_log.info.call_args_list]
assert "mcp_server.oauth.revocation_succeeded" in info_events
def test_revoke_token_omits_client_secret_when_none(self) -> None:
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock(return_value=_mk_response(200))
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
)
)
body = client.post.call_args.kwargs["data"]
assert "client_secret" not in body
assert body["token"] == "r-secret"
assert body["token_type_hint"] == "refresh_token"
assert body["client_id"] == "client-1"
def test_revoke_token_succeeds_on_204(self) -> None:
# RFC 7009 says the AS MAY return any 2xx; treat the whole range
# as success.
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock(return_value=_mk_response(204))
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
)
)
info_events = [c.args[0] for c in mock_log.info.call_args_list]
assert "mcp_server.oauth.revocation_succeeded" in info_events
class TestRevocationFailureLogged:
def _run_and_capture(self, status: int) -> list[Any]:
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock(return_value=_mk_response(status))
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
)
)
return mock_log.info.call_args_list
def test_revoke_token_logs_on_400_does_not_raise(self) -> None:
calls = self._run_and_capture(400)
events = [c.args[0] for c in calls]
assert "mcp_server.oauth.revocation_failed" in events
# Must include status field.
failed_call = next(c for c in calls if c.args[0] == "mcp_server.oauth.revocation_failed")
assert failed_call.kwargs.get("status") == 400
def test_revoke_token_logs_on_401_does_not_raise(self) -> None:
calls = self._run_and_capture(401)
events = [c.args[0] for c in calls]
assert "mcp_server.oauth.revocation_failed" in events
failed_call = next(c for c in calls if c.args[0] == "mcp_server.oauth.revocation_failed")
assert failed_call.kwargs.get("status") == 401
def test_revoke_token_logs_on_403_does_not_raise(self) -> None:
calls = self._run_and_capture(403)
events = [c.args[0] for c in calls]
assert "mcp_server.oauth.revocation_failed" in events
def test_revoke_token_logs_on_5xx_does_not_raise(self) -> None:
calls = self._run_and_capture(500)
events = [c.args[0] for c in calls]
assert "mcp_server.oauth.revocation_failed" in events
failed_call = next(c for c in calls if c.args[0] == "mcp_server.oauth.revocation_failed")
assert failed_call.kwargs.get("status") == 500
class TestRevocationExceptionPaths:
def test_revoke_token_handles_network_error(self) -> None:
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock(side_effect=httpx.ConnectError("boom"))
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
)
)
events = [c.args[0] for c in mock_log.info.call_args_list]
assert "mcp_server.oauth.revocation_failed" in events
failed_call = next(
c
for c in mock_log.info.call_args_list
if c.args[0] == "mcp_server.oauth.revocation_failed"
)
assert failed_call.kwargs.get("error") == "ConnectError"
def test_revoke_token_handles_httpx_timeout(self) -> None:
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock(side_effect=httpx.TimeoutException("slow"))
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
)
)
events = [c.args[0] for c in mock_log.info.call_args_list]
assert "mcp_server.oauth.revocation_failed" in events
failed_call = next(
c
for c in mock_log.info.call_args_list
if c.args[0] == "mcp_server.oauth.revocation_failed"
)
assert failed_call.kwargs.get("error") == "TimeoutException"
def test_revoke_token_handles_asyncio_timeout(self) -> None:
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
async def _slow(*_args: Any, **_kwargs: Any) -> Any:
await asyncio.sleep(10.0)
raise AssertionError("should have timed out")
client.post = AsyncMock(side_effect=_slow)
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
timeout_seconds=0.05,
)
)
events = [c.args[0] for c in mock_log.info.call_args_list]
assert "mcp_server.oauth.revocation_failed" in events
failed_call = next(
c
for c in mock_log.info.call_args_list
if c.args[0] == "mcp_server.oauth.revocation_failed"
)
# ``asyncio.timeout`` raises ``TimeoutError`` (Python's builtin)
# on cancellation.
assert failed_call.kwargs.get("error") == "TimeoutError"
def test_revoke_token_no_exc_info_in_logs(self) -> None:
"""Bearer-leak invariant: the revoke path must NEVER set
``exc_info=True``. Chained ``__context__`` may include an
``httpx.Request`` whose ``Authorization`` header holds a
bearer; the traceback formatter would render it.
"""
as_meta = _make_as_metadata()
client = MagicMock(spec=httpx.AsyncClient)
client.post = AsyncMock(side_effect=httpx.ConnectError("boom"))
with patch("turnstone.core.mcp_oauth.log") as mock_log:
asyncio.run(
revoke_token_at_as(
as_metadata=as_meta,
http_client=client,
refresh_token="r-secret",
client_id="client-1",
client_secret=None,
)
)
# No info call may carry exc_info.
for call in mock_log.info.call_args_list:
assert "exc_info" not in call.kwargs, (
f"mcp_server.oauth log info({call.args[0]!r}) used exc_info — "
"this violates the bearer-leak invariant"
)
# Defensively: also check warning + exception levels for the
# same call site.
for call in mock_log.warning.call_args_list:
assert "exc_info" not in call.kwargs
mock_log.exception.assert_not_called()
class TestAttemptUpstreamRevokeNeverRaises:
"""Round-2 q-3 regression: ``_attempt_upstream_revoke``'s docstring
claims ``Never raises``. Background-task semantics make this load-
bearing a propagated exception logs ``Task exception was never
retrieved`` because the ``set.discard`` done-callback doesn't read
``task.exception()``.
The wrapper's narrow inner ``except`` clauses (``MCPOAuthDiscoveryError``,
``MCPTokenDecryptError``) leave room for any other exception type
raised by ``discover_authorization_server`` /
``storage.get_mcp_oauth_client_secret_ct`` / ``token_store.cipher.decrypt``
to escape. The outer ``try/except Exception`` is what keeps the
contract honest. These tests pin that gate.
"""
def _build_args(self) -> dict[str, Any]:
token_store = MagicMock()
token_store.cipher = MagicMock()
token_store.cipher.decrypt.return_value = b"shh"
storage = MagicMock()
storage.get_mcp_oauth_client_secret_ct.return_value = None
return {
"http_client": MagicMock(spec=httpx.AsyncClient),
"metadata_cache": None,
"storage": storage,
"token_store": token_store,
"server_name": "srv-oauth",
"server_row": {
"url": "https://mcp.example.com",
"oauth_client_id": "client-1",
"oauth_authorization_server_url": None,
"oauth_as_issuer_cached": None,
},
"server_id_for_audit": "srv-id-1",
"refresh_token": "r-secret",
}
def test_attempt_upstream_revoke_swallows_unexpected_exception(self) -> None:
"""A generic exception from a path the inner handlers don't
cover MUST be caught at the outer boundary and logged with type
name only (no exc_info=True per the bearer-leak invariant).
"""
args = self._build_args()
async def _boom(*_a: Any, **_kw: Any) -> Any:
raise RuntimeError("network blew up")
with (
patch("turnstone.core.mcp_oauth.discover_authorization_server", side_effect=_boom),
patch("turnstone.core.mcp_oauth.log") as mock_log,
):
# MUST NOT raise.
asyncio.run(_attempt_upstream_revoke(**args))
events = [call.args[0] for call in mock_log.info.call_args_list]
assert "mcp_server.oauth.upstream_revoke_failed" in events, (
"outer try/except must log mcp_server.oauth.upstream_revoke_failed "
"with the exception type name when an unexpected exception escapes "
"the narrow inner handlers"
)
for call in mock_log.info.call_args_list:
assert "exc_info" not in call.kwargs, (
"outer-block log must not use exc_info=True — chained "
"__context__ may carry an httpx.Request bearer"
)
def test_attempt_upstream_revoke_logs_discovery_failure(self) -> None:
"""Round-2 bug-1: ``MCPOAuthDiscoveryError`` MUST emit
``upstream_revoke_discovery_failed`` so operators have visibility
into a silent-discovery-failure path that previously logged
nothing while the audit row recorded ``upstream_revoke_outcome=scheduled``.
"""
args = self._build_args()
async def _disc_fail(*_a: Any, **_kw: Any) -> Any:
raise MCPOAuthDiscoveryError("PRM fetch 503")
with (
patch(
"turnstone.core.mcp_oauth.discover_authorization_server",
side_effect=_disc_fail,
),
patch("turnstone.core.mcp_oauth.log") as mock_log,
):
asyncio.run(_attempt_upstream_revoke(**args))
events = [call.args[0] for call in mock_log.info.call_args_list]
assert "mcp_server.oauth.upstream_revoke_discovery_failed" in events
assert "mcp_server.oauth.upstream_revoke_failed" not in events

Some files were not shown because too many files have changed in this diff Show More