Compare commits

...

19 Commits

Author SHA1 Message Date
Patrick Buckley 35f462a46d chore: bump version to 1.5.6 2026-05-03 13:44:50 -07:00
Patrick Buckley ec74334e74 feat(providers): api_surface toggle + mistral medium reasoning fix (#469)
* feat(providers): api_surface toggle + mistral medium reasoning fix

Mistral medium open-weights served by vLLM expects reasoning_effort via
the Responses API (`reasoning.effort`), not as a `chat_template_kwargs`
entry on Chat Completions.  The session was unconditionally injecting
`{"reasoning_effort": ...}` into `chat_template_kwargs` for every
openai-compatible request, which corrupted the prompt rendering for any
backend whose chat template didn't consume that key (Mistral medium,
Mistral cloud, Groq, OpenRouter).

Changes:
- Add `api_surface` ("chat" | "responses") to `ModelConfig.server_compat`
  and thread it through `create_provider` / `model_registry.get_provider`.
  `openai-compatible` defaults to Chat Completions; operators can flip
  individual aliases to Responses for endpoints that support it.
- New `vllm-mistral-medium` profile that pre-fills api_surface=responses
  on Detect for known Mistral medium model ids.
- Drop the unconditional `reasoning_effort` injection into
  `chat_template_kwargs`.  Operators running gpt-oss-style local
  templates that consume `reasoning_effort` from the chat template now
  opt in via `server_compat.extra_body.chat_template_kwargs`.
- New "API Surface" select in the Models admin tab; allowlist-validated
  server-side at create/update time; pre-filled by Detect via the
  profile suggestion.
- Evict the cached provider singleton in `ModelRegistry.reload()` when
  api_surface changes (previously only cfg.provider triggered eviction).
- Fix `_run_agent` fallback path to inherit the session's primary alias
  for capability and server_compat resolution; previously the fallback
  passed `alias=None`, which silently dropped per-model caps on the
  agent path.

Tests: 5117 passed (-m "not live"); ruff + mypy clean.

* fix(providers): don't auto-suggest Responses for Mistral medium

vLLM's Responses API surface for Mistral medium open-weights doesn't
wire up the Mistral tool-call parser as of vLLM 0.x — tool calls leak
into the response as ``[TOOL_CALLS]<name>{...}`` text instead of
structured tool_calls.  Chat Completions on the same engine handles
tools cleanly via ``--tool-call-parser mistral``, and reasoning can be
turned on via the vLLM CLI ``--reasoning-parser`` flag.

Drop the auto-suggest mapping so Detect falls back to the generic
``vllm`` profile.  Keep the ``vllm-mistral-medium`` profile definition
in place so an operator who specifically wants per-request effort and
accepts the tool-calling limitation can still pick "Responses API"
manually in the admin UI.

* fix(providers): address Copilot review on PR #469

- providers/__init__.py: drop the redundant *_responses_provider /
  *_chat_provider names; have create_provider use _openai_provider and
  _openai_compat_provider directly so they're not flagged as unused
  globals.
- console/server.py: tighten _validate_api_surface to a strict equality
  match against the canonical {"chat", "responses"} set.  The previous
  strip().lower() membership check accepted ' Responses '/'CHAT' but
  stored the raw string verbatim, which then failed to round-trip
  through the admin <select>.
- console/static/admin.js: gate the entire server_compat block (server
  type, api_surface, extra_body) on provider == "openai-compatible" at
  save time so toggling provider away can't leave a stale hidden surface
  selection in the persisted capabilities JSON.
- tests/test_session.py: splat the bad kwarg via **dict so CodeQL no
  longer flags the call as a wrong-name keyword (the point of the test
  is the runtime contract, not the static type).
- tests/test_admin_model_registry_refresh.py: add endpoint-level tests
  for the api_surface validation on both create and update — covers the
  bogus-value rejection, non-canonical-string rejection, and the happy
  path persisting through to the refreshed registry.
2026-05-03 13:40:30 -07:00
Patrick Buckley 2fd0c29a92 fix(memory): query-aware candidate selection + OR-of-terms search (#468)
* fix(memory): query-aware candidate selection + OR-of-terms search

The system-message memory composition path used a recency-ordered
candidate set (`_list_visible_memories(limit=fetch_limit)`).  On
deployments with more than `fetch_limit` (default 50) visible
memories, BM25 only ever ranked the 50 most-recently-touched memories
— a relevant memory written months ago was silently invisible
regardless of how well it matched the recent context.  Multi-word
search at the SQL layer used AND-of-terms, killing recall on any
multi-word query without an exact field overlap.

## Functional changes

- `_init_system_messages` (`turnstone/core/session.py`): extract
  recent context first, then `_search_visible_memories(context)` to
  pull query-aware candidates.  Search hits below `fetch_limit` union
  with the recency list (deduped by memory_id) so the BM25 candidate
  pool is always a SUPERSET of the prior recency-only pool — even on
  noisy queries where the cap fills with stopwords, the recency-50
  the original bug surfaced still reaches BM25.  Empty context skips
  search entirely.  Candidate-selection logic extracted into
  `_select_memory_candidates`.

- `search_structured_memories` (PostgreSQL + SQLite): per-term
  clauses join with OR instead of AND.  A row matches if ANY term
  matches ANY of name/description/content.  Downstream BM25 narrows
  back down by relevance.

## Perf hardening

- Collapse the 1-3 fanned scope queries into a single SQL.  New
  backend methods `list_visible_structured_memories` /
  `search_visible_structured_memories` union the visibility scopes
  into one WHERE OR-group, so a composition rebuild now hits the DB
  at most twice (search + recency) instead of up to six times.

- Cap and normalize search terms.  Composition can hand a multi-KB
  pasted message to ILIKE-based search; without a cap, every distinct
  token would emit one unindexable predicate per scope-fanned query.
  `normalize_search_terms` (`storage/_utils.py`) de-dupes
  case-insensitively, drops <2-char tokens, and hard-caps at 16.

- Per-turn search cache.  `_init_system_messages` fires from many
  call sites within one turn (state transitions, MCP refresh, tool
  results) and the recent-context query is identical across them.
  Session-instance cache keyed by (query, mem_type, limit) absorbs
  the duplicates; invalidated in `_append_user_turn` and after
  memory save/delete tool actions.

- Stable secondary sort by `memory_id`.  `updated` is second-precision
  and `touch_structured_memories` can land a batch on identical
  timestamps; without a tie-breaker SQL returns rows in
  implementation-defined order, BM25 input shuffles, and the
  LLM-side prompt cache misses across calls.  All four backend ORDER
  BYs now break ties on `memory_id ASC`.

## Quality cleanups

- Coalesce `memory.search.term_count` + `memory.search.zero_results`
  into a single `memory.search` log carrying both `term_count` and
  `result_count`.
- New `memory.composition` log: source / candidates / injected.
- Promote a shared `make_chat_session` factory to `tests/_helpers.py`.
- Rename SQL builder local `extra` -> `scope_filters` for clarity.
- Add docstrings on `search_structured_memories` so the AND->OR flip
  survives future readers.

## Tests

Adds 20 tests across `tests/test_structured_memory.py`,
`tests/test_structured_memory_storage.py`, and
`tests/test_memory_relevance.py`: recency-ceiling regression,
empty-query fallback, sparse-match union, recency-preserved-when-
search-returns-noise (locks in the pool-superset invariant),
OR-of-terms on both backends, scope filtering preserved,
search-facade multi-word behavior, term-cap normalization, the new
visible-scope helpers (list + search + empty-scopes guard),
coord-scope composition isolation, end-to-end
`memory(action='search')` tool execution, per-turn cache hit +
invalidation, and stable ordering under tied `updated` timestamps.

Memory test sweep: 102/102.  Broader regression
(session, storage, coordinator, load_skill): 411/411.

* fix(memory): address Copilot review on PR #468

Three follow-ups from Copilot's inline review:

1. SUPERSET invariant violation (Copilot, session.py:5510).
   `(search_hits + extra)[:fetch_limit]` capped the union back down to
   fetch_limit, evicting the recency tail when search added distinct
   hits.  Recency tail is exactly where ancient-but-recently-touched
   memories live — the recall this PR is supposed to improve — so
   tail eviction recreated the bug for the narrow case where a query
   term fell off the 16-cap and the matching memory sat in
   recency[40-49].  Drop the cap; both halves are already SQL-capped
   at fetch_limit, so the union is at most 2 × fetch_limit (~100 with
   defaults).  BM25 over 100 candidates in pure Python is sub-ms;
   irrelevant recency fillers get score=0 and don't pollute ranking.
   Updates the docstring to actually be honest about the invariant.
   Adds `test_recency_tail_preserved_when_search_adds_distinct_hits`
   that locks the behavior in: 5 search hits + 10 recency = 15-item
   pool, every recency item present, source="union".

2. Unbounded `query.split()` in normalize_search_terms (Copilot,
   _utils.py:74).  `str.split()` allocates the full token list before
   the cap-after-16 break, so a 100KB pasted query did MB of throwaway
   work even though only 16 tokens entered SQL.  Switch to
   `re.finditer(r'\S+', query)` — streaming iterator, stops scanning
   at the first 16 normalized terms regardless of input size.

3. Misleading + unbounded log term_count (Copilot, session.py:8571).
   `len(item["query"].split())` had two problems: same unbounded
   split as #2, and the value reported the raw input token count
   rather than the normalized term count that actually hit the SQL
   WHERE clause — misleading metric for an operator trying to
   understand storage-side behavior.  Switch to
   `len(normalize_search_terms(item["query"]))` — accurate count, and
   bounded for free via #2.

Refuted: github-code-quality flagged `...` bodies in the new Protocol
methods as "statement has no effect."  False positive — `...` is the
canonical Protocol body convention, used 213 other times in the same
file.

Memory test sweep: 103/103.  Broader regression: 411/411.
2026-05-03 13:40:30 -07:00
Patrick Buckley 1207d27363 fix(tests): isolate metrics-singleton swaps so they don't leak across files
CI failure on main: test_publish_records_metric_outcome saw an empty
calls list — its monkeypatch was patching a different metrics
instance from the one `_publish_models_metadata` reads.

Two changes:

- test_close_reason_persistence.py: replace the bare
  `srv_mod._metrics = MetricsCollector()` assignment in `_make_app`
  with an autouse `monkeypatch.setattr(srv_mod, "_metrics", ...)`
  fixture so the test's metrics swap auto-restores. Other test
  files (test_auth.py, test_server_attachments_endpoints.py) carry
  the same anti-pattern; left for a follow-up since they're not on
  the critical path here.

- test_server_node_models_metadata.py: switch the publish-helper
  metric test to a string-form `monkeypatch.setattr("turnstone.
  server._metrics", FakeMetrics())` so it replaces whatever binding
  the live module currently holds, regardless of what other tests
  did to it. Robust against future leaks of the same shape.
2026-05-03 13:40:29 -07:00
Patrick Buckley 733c9818d4 feat(coord): expose healthy model aliases per node on list_nodes (#466)
* feat(coord): expose healthy model aliases per node on list_nodes

Surfaces a `model_aliases` field on each `list_nodes` row so a
coordinator can discover which model aliases each cluster node will
accept on `spawn_workstream(model=...)` without an HTTP fan-out.

Each server projects its registry into a `models` entry on
`node_metadata` (`{alias, provider, healthy}` per alias) at lifespan
startup, on every 30s heartbeat tick, and after `internal_model_reload`.
The publish helper short-circuits on a payload-equality cache so a
stable cluster doesn't pay UPSERT churn — exposed via the new
`turnstone_node_models_publish_total{outcome="written|skipped"}`
Prometheus counter so operators can graph cache hit-rate.

Coord client filters the per-alias rows to healthy aliases only and
drops the provider-side model identifier (`cfg.model`) — coords kept
reaching for it when they should pass the local alias.

* fix(coord): address Copilot+CodeQL feedback on list_nodes models work

- internal_model_reload: reuse a single get_storage() local across the
  registry load and the metadata publish (Copilot:3047)
- _collect_node_models_metadata: iterate sorted aliases so two
  structurally identical registries built in different insertion orders
  serialize to the same JSON — directly improves the publish-cache hit
  rate exposed via turnstone_node_models_publish_total (Copilot:3105)
- tests: drop mixed turnstone.server import style flagged by CodeQL —
  hoist _metrics into the from-import block, and use sys.modules in
  the shutdown-race regression test instead of `import as srv`
2026-05-03 13:40:29 -07:00
Patrick Buckley 28a2779c10 fix(core): scope rehydrate fallback to manager, fix resume orphan
Address Copilot feedback on PR #465:

1. The has_alias fallback in both session_factories silently rewrote
   any unknown caller-supplied alias to the default, including on the
   fresh-create path where the create handler maps the factory's
   ValueError to a 503 with operator-friendly text. A typo in
   body.model would now silently start a workstream on the default
   instead of telling the caller their requested model could not be
   resolved. Move the fallback out of the factories: each factory
   raises again on unknown aliases, and SessionManager filters stale
   aliases out of the rehydrate path via a new ``model_validator``
   constructor kwarg (production wiring passes ``registry.has_alias``
   on both interactive and coordinator).

2. ChatSession.resume()'s elif branch flipped self.model to the
   persisted model name even when the alias was unresolvable, leaving
   the session paired with the constructor's default provider/client
   but a removed model name — a broken state whose next API call
   fails. Drop the model copy: keep the constructor's coherent
   default (provider + model + capabilities) and just log the
   unreachable saved values so the missing alias is auditable.

Tests:
- Move stale-alias coverage from the factory level into
  SessionManager (tests/test_session_manager.py): validator drops
  stale aliases before reaching build_session; live aliases pass
  through unchanged.
- tests/test_sessions.py renamed test_resume_restores_model →
  test_resume_keeps_defaults_when_alias_unresolvable to match the new
  contract.
2026-05-03 13:40:29 -07:00
Patrick Buckley 56364b0b5b fix(core): preserve workstream model + config on rehydrate
SessionManager.open() was calling build_session(ws) without a model
arg on the rehydrate path. The session_factory then resolved the
*current* default alias, ChatSession.__init__'s _save_config() (INSERT
OR REPLACE per-key) clobbered the persisted workstream_config with
those defaults, and the subsequent resume() "restored" what was now
the default — silently resetting model_alias, model, temperature,
reasoning_effort, max_tokens, skill, creative_mode, instructions,
token_budget, and notify_on_complete on every reopen and every
service restart, for both interactive and coordinator workstreams.

Three layers:

1. SessionManager.open() now reads workstream_config via
   self._storage.load_workstream_config(ws_id) and threads the saved
   model_alias into build_session(ws, model=saved_alias).

2. ChatSession.__init__ now skips its initial _save_config() when a
   workstream_config row already exists for self._ws_id — protects
   every other persisted knob without having to plumb each one
   through the adapter signature, and catches any future construction
   path that forgets to thread model through build_session.

3. Both session_factories (server.py interactive, console
   session_factory.py coordinator) now treat an unknown caller-
   supplied alias the same as an unset alias: fall back to the
   runtime default rather than raising. Without this, a workstream
   pinned to an alias an operator has since removed from the registry
   would 500 on every reopen — defeating the "best effort restore,
   default if the original is gone" contract this fix is meant to
   deliver. Mirrors _effective_default_alias's existing has_alias
   guard against a stale ConfigStore default.
2026-05-03 13:40:29 -07:00
Patrick Buckley afb5804a7c fix(console): address Copilot feedback on Models → Roles sub-tab
Three changes from PR review:

- Permission gating: hide the Roles sub-tab button when the user
  lacks ``admin.settings``.  The sub-tab loads/saves through
  ``/v1/api/admin/settings``, so an admin with ``admin.models`` but
  no ``admin.settings`` would otherwise see a perpetual 403 loader.
  When Roles is the active sub-tab and the permission check fails,
  snap the panel back to Definitions so the user lands somewhere
  usable.

- Drop the redundant ``/v1/api/admin/model-definitions`` fetch from
  ``loadAdminModelRoles``.  Both entry points (initial Models-tab
  open + ``models_changed`` SSE refresh) flow through
  ``loadAdminModels`` first, which already populates ``_modelDefs``
  + ``_modelDefaultAlias``; ``_saveModelRole`` doesn't touch model
  definitions, so the cached snapshot stays accurate when the save
  chains back here.  Halves the per-render request count and
  removes a wasted round-trip on every cluster-wide model edit.

- Add ``test_models_changed_event.py`` covering the SSE fanout the
  prior commit introduced: each model-definition CRUD endpoint
  emits exactly one ``models_changed``, settings PUT/DELETE only
  emit for keys in ``_MODEL_AFFECTING_SETTING_KEYS`` (parametrised
  over all eight), and unrelated settings (e.g.
  ``session.retention_days``) don't trigger spurious refreshes.
  The expected key set is pinned in the test so a stray addition
  to the allowlist doesn't silently bypass coverage.
2026-05-03 13:40:29 -07:00
Patrick Buckley 4b508a1319 feat(console): add plan_agent + task_agent to Models → Roles
Same shape as the coordinator/judge rows already there: alias dropdown
+ reasoning_effort dropdown sourced from the existing
``model.plan_alias`` / ``model.plan_effort`` and
``model.task_alias`` / ``model.task_effort`` settings.  Adds the four
keys to the SSE ``models_changed`` allowlist so changes from the
Settings API also trigger a live dropdown refresh, and filters them
out of the Settings tab so they only render in one place.
2026-05-03 13:40:29 -07:00
Patrick Buckley 9c2cb185e1 feat(console): consolidate role-model settings + live-refresh dropdowns
Lifts judge and coordinator model assignments out of their respective
admin tabs and into a new Models → Roles sub-tab so role overrides live
next to the model definitions they reference. Forward-looking shape for
the upcoming perception.{audio,image,video} model settings — adding a
new role is one entry in the declarative MODEL_ROLES array.

Also drops the misleading "Coordinator subsystem not configured" home
banner. The session factory already falls back to the registry's
default model when coordinator.model_alias is unset, so the banner was
nagging on fresh installs where the system was actually working. The
related _probeCoordSubsystem / _homeCoordReady plumbing went with it.

Wires SSE-driven live refresh: the console now emits a models_changed
event when a model definition is created/updated/deleted/reloaded, or
when a model-affecting setting (model.default_alias, judge.model,
coordinator.model_alias, coordinator.reasoning_effort) changes.
Connected browsers refetch /v1/api/models on receipt so the home
composer's model dropdown and the Roles sub-tab stay accurate without
a manual reload — fixes the case where editing the underlying model
for an existing alias left the dropdown showing the old model id.

Companion cleanups:
- Renamed .judge-section-* CSS classes to .admin-subtab-* and shared
  them with the Models sub-tab switcher (same a11y attrs, arrow-key
  nav). Old names had no other callers.
- Filtered judge.model out of the Judge Settings sub-tab and
  coordinator.model_alias / coordinator.reasoning_effort out of the
  Settings tab — they live exclusively under Models → Roles now.
- Reworded the _require_coord_mgr 503 messages to point operators at
  the Models tab instead of suggesting they set coordinator.model_alias.
2026-05-03 13:40:29 -07:00
Patrick Buckley 0519b847bd docs(skills): add import-conversation-history SKILL.md
Source-agnostic guide that teaches an agent Turnstone's destination
contracts (workstream + conversations schema, ws_id routing, OpenAI
message shape, tool-call/result pairing, provider_data fidelity blob,
attachment lifecycle) so it can map any external chat export onto them.
Validated against turnstone.core.skill_parser.
2026-05-03 13:40:29 -07:00
Patrick Buckley bbc8b99a9f fix(console): home composer attachments + coord chat user-message pills (#462)
* fix(console): home composer attachments + coord chat user-message pills

Two parity gaps in the console's coordinator surface:

- The embedded creator on the home page accepted only text — the
  paperclip / paste / drop pipeline that the in-coord composer and the
  interactive new-ws modal both expose was missing, so a user couldn't
  attach files at create time. Stage Files in memory (no ws_id yet) and
  ship them multipart on Start; the coord create endpoint already accepts
  multipart via create_supports_attachments=True.

- User messages with attachments rendered as plain text on both live
  send and history replay — no chip cluster like the interactive pane.
  Added appendUserMessageWithAttachments and a structured userAttachments
  list built from _attachments_meta (preferred) or the multipart parts
  themselves, then rendered the same .msg-user-attach pill strip the
  interactive pane uses.

Polish from a designer pass:

- Pill background was --panel-2, equal to the .msg bubble background in
  both themes (border contrast ≈1.4:1, below WCAG 1.4.11). Switched to
  --panel so the pill sits on a different surface than the bubble.
- Capped chip filename width inside the home composer (max-width 200px +
  ellipsis) so a long filename doesn't push the strip past the textarea.
- aria-live="assertive" → "polite" on #home-coord-error; client-side
  validation isn't an interrupt-level event.
- Reserved min-height on .home-composer-error and dropped the
  display: none/block toggling so validation messages no longer reflow
  the active-coordinators list below.

* fix(console): address PR #462 review feedback

- Block home-composer submit when files are staged but the task field is
  empty.  Server's _coord_create_post_install short-circuits on an empty
  initial_message, so the multipart upload would create pending
  attachment rows that never reserve onto a turn — orphaned until the
  GC sweep.  Fail in the browser instead.
- Drop the redundant `part &&` guard in coordinator.js's history-replay
  multipart loop; the earlier `if (!part || ...) continue` already
  filtered.
- Rewrite the home-mount .composer-chip-name CSS comment.  shared/chat.css
  defines .composer-chip{,-size,-remove} but no .composer-chip-name rule
  — the span inherits the parent chip font with no width cap.
- Add smoke-guard string assertions in test_coordinator_page.py for
  appendUserMessageWithAttachments and msg-user-attach so a future
  rename can't silently regress the attachment affordance.
2026-05-03 13:40:29 -07:00
Patrick Buckley bd9f780b21 chore: bump version to 1.5.5 2026-05-01 14:09:38 -07:00
Patrick Buckley 5d14b5f675 fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration (#461)
* fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration

Loading a saved workstream silently dropped tool results and missed
verdict / output-guard / truncation signals on replay. Root cause was
in `Pane.prototype.replayHistory`: an assistant message carrying both
content and tool_calls cleared the `lastToolBlock` anchor before the
following tool-result iteration could attach. The fix reorders content
to render before the tool block (matching live SSE order) and
restructures the tool-result branch to anchor by `data-call-id` so
multi-tool batches render `[hdr A][out A][hdr B][out B]` rather than
bunching outputs at the bottom.

Beyond the bug, replay now reaches near-parity with the live UX:

- Persisted intent verdicts and output_assessments flow through both
  the SSE replay (`_build_history`) and the `/history` REST endpoint
  used by coord. Single shared helper module owns the wire shape.
- Memory/recall calls persist instead of being filtered at storage
  time — full audit trail; UI dims them by default with hover-reveal
  so heavy memory usage doesn't crowd the narrative.
- Truncation indicator surfaces as a sibling pill (consistent across
  interactive + coord) when a tool result hit the 2000-char cap.
- `replayHistory` wraps DOM work in `aria-busy` so screen readers
  don't get a chatty announce-flood on long replays.
- `_build_history`'s storage I/O moves off the event loop via a new
  `events_replay_prepare` async hook for the SSE path; other async
  callers wrap in `asyncio.to_thread`.

Coord parity:

- `/history` REST endpoint decorates tool_calls with verdict +
  output_assessment + truncation flag (was previously raw
  `load_messages` output).
- Coord JS stamps `judge_verdict` / `heuristic_verdict` from
  history-loaded `tc.verdict` so the existing batch render paints
  the persisted pill, seeds the verdict cache to dedupe later live
  SSE events, and emits an inline `.coord-tool-row-warning` chip
  per call instead of a generic chat line.
- Memory/recall dim rule mirrored on `.coord-tool-row[data-tool-name=...]`.

* fix(replay): address PR #461 review feedback + raise tool-result storage cap

Copilot review feedback:

- Sibling-chain dim rule (memory/recall) now adds :focus-within
  alongside :hover for .tool-output / .media-embed / .output-warning
  / .tool-output-truncated — keyboard users tabbing into a faded
  subtree now get full opacity.
- ``cfg.open_post_load`` is now invoked via ``await asyncio.to_thread``
  so its sync ``_build_history`` call (storage I/O for verdict
  indexes + message reconstruction) doesn't block the event loop on
  every workstream open. Mirrors the SSE replay path that's already
  protected via ``events_replay_prepare``.
- Replaced the hardcoded ``2000`` literal in server.py and session.py
  with ``TOOL_RESULT_STORAGE_CAP`` from the shared decoration module
  so the UI truncation-pill detection can't silently desync from the
  storage write side.

While here:

- Raised ``TOOL_RESULT_STORAGE_CAP`` from 2000 → 10000. A 2000-char
  clip routinely cut grep / file-read bodies mid-line, leaving the
  audit trail useless for retrospective debugging. FTS5 + row size
  grow proportionally; the per-tool upper bound is still bounded
  upstream by ``_truncate_output``'s context-budget clamp.
- Updated the user-visible truncation-pill tooltip on both
  interactive and coord to reflect the new cap.
- ``test_decorates_tool_calls_and_marks_truncated`` now references
  the constant instead of a literal so it stays correct on future
  cap changes.
2026-05-01 14:06:08 -07:00
Patrick Buckley 4693fa95f1 chore: bump version to 1.5.4 2026-04-30 23:51:06 -07:00
renovate[bot] c3423d6606 chore(deps): update ghcr.io/astral-sh/uv docker tag to v0.11.8 2026-04-30 23:48:53 -07:00
Patrick Buckley 53f1222c22 refactor(coord): remove priority queue + queue depth indicator + broken CSS
Speculative reliability machinery from the Stage 3 push that turned
out not to address any user-visible bug. The actual fixes (state /
activity disjunction in handleChildState, bulk-fetch race fix in
_fetch_live_block, push approve_request via cluster bus) are what
resolved the wedged-row issues. Manual testing showed the per-tab
SSE listener queue depth never climbed past single digits even when
rows were stuck — overflow was never the cause.

Removed
- ``_CRITICAL_EVENT_TYPES`` + ``_put_with_priority`` helper.
- Per-tab listener queue selective drop (back to plain
  ``contextlib.suppress(queue.Full)`` everywhere).
- ``ClusterCollector._fanout`` reverts to the same.
- WebUI ``_broadcast_intent_verdict`` / ``_broadcast_approval_resolved``
  / ``_broadcast_approve_request`` revert to plain ``put_nowait``.
- ``_queue_stats`` periodic SSE emit + frontend status-bar indicator
  + the supporting CSS rules.
- Broken ``.approval-block`` ``transition: max-height`` /
  ``max-height: 80vh`` / ``overflow: hidden`` rules — the transition
  never fired (nothing toggled max-height) and ``overflow: hidden``
  clipped long verdict reasoning. Layout-shift on auto-expand jumps
  again, which is preferable to clipped content (Copilot review).

Tidied
- ``_CollectorProtocol`` / ``_ManagerProtocol`` method bodies switch
  from ``...`` ellipsis to docstring-only bodies, silencing four
  CodeQL "statement has no effect" warnings without changing the
  Protocol contract.

5024 passed, ruff + mypy clean.
2026-04-30 23:48:53 -07:00
Patrick Buckley 802d87a57f feat(coord): Stage 3 SessionManager Children primitive lift + cluster bus push paths
Lift the Children primitive out of CoordinatorAdapter into universal
SessionManager core primitives, replace the fragile poll + state-event
piggyback paths with first-class cluster bus event types for inline
approval delivery, and clean up the resulting frontend reducer.

Architecture
- New `turnstone/core/children_registry.py` — universal parent → children
  + reverse-lookup primitive with atomic `add_child` (returns parent UI
  for race-free dispatch). Lifted from `CoordinatorAdapter`.
- New `turnstone/core/child_source.py` — `ChildSource` Protocol with
  `SameNodeChildSource` (in-process via SessionManager state observer)
  and `ClusterChildSource` (cross-node via ClusterCollector listener).
- `SessionManager._on_state_change` upgraded to multi-subscriber
  (`subscribe_to_state` / `unsubscribe_from_state`) under a dedicated
  lock; CLI consumer migrated.
- `CoordinatorAdapter` shrunk: 731 → ~640 LOC. Children data lives in
  the registry; fan-out lives in ClusterChildSource. Backward-compat
  property facades dropped; tests updated to use the registry surface.

Cluster bus event vocabulary
- New event types `intent_verdict`, `approval_resolved`,
  `approve_request` flow through both `ClusterCollector._apply_delta`
  (translation from node SSE) and `emit_console_ws_*` (synthesis on
  console pseudo-node).
- `CoordinatorAdapter._dispatch_child_event` re-emits as
  `child_ws_intent_verdict` / `child_ws_approval_resolved` /
  `child_ws_approve_request` on the parent coord's SSE stream.
- New `_broadcast_intent_verdict` / `_broadcast_approval_resolved` /
  `_broadcast_approve_request` no-op hooks on `SessionUIBase`. WebUI
  pushes to the global queue; ConsoleCoordinatorUI pushes to the
  collector. `approve_tools` calls `_broadcast_approve_request` right
  after setting `_pending_approval` so the items reach the coord tree
  immediately, eliminating the bulk-fetch race.

Cleanups
- `pending_approval_detail` piggyback on `ws_state` / `cluster_state`
  removed end-to-end. Bulk fetch + explicit verdict / approve-request
  push are the canonical carriers.
- Browser `_judgePollTick` 90-second poll loop deleted; push path is
  authoritative.
- `urgent` flag on `scheduleLiveFetch` deleted (only caller was 409
  retry; replaced with `invalidateLiveBadge` + standard schedule).
- Console `_fetch_live_block` derives `pending_approval` from a
  disjunction (`activity_state="approval"` OR `state="attention"`
  OR detail present) so the bulk fetch can't return false during the
  state-transition race window.
- Coord-side merge guard in `flushLiveFetches` no longer clobbered:
  `handleChildState` only stamps `sseUpdatedAt` when authoritatively
  clearing detail.
- `child_locality` capability flag removed (was inert dead code).

Reliability
- Selective drop on listener queue overflow: critical event types
  (verdicts, approvals, ws_closed, child_ws_*) evict one oldest item
  to make room rather than dropping themselves on a full queue.
  Best-effort events (state ticks, content tokens, status, activity)
  drop as before. Applied to `SessionUIBase._enqueue`,
  `ClusterCollector._fanout`, and the `WebUI._global_queue` puts in
  the new broadcast hooks.
- `_state_subscribers` snapshot under a dedicated lock so concurrent
  subscribe / unsubscribe during dispatch can't shift the iterator.

UX / a11y
- Loading placeholder in renderChildRow keeps row height stable while
  the bulk fetch is in-flight (sr-friendly aria-label).
- Focus preservation across `_renderChildrenNow` (capture +
  restore by row + marker) and across targeted `_updateChildRow` swaps.
- Layout-shift transition on the approval block max-height; respects
  `prefers-reduced-motion`.
- Sidebar pending count: `(N children · M pending)`.
- Risk pill `aria-label` spells out level + confidence for SR users.
- Per-coord SSE listener queue depth surfaced in the status bar
  (`queue N/500`) with color escalation (warn at >50%, danger at >80%).

Tests
- 305+ test changes across 8 files. New unit tests for
  `ChildrenRegistry`, `ChildSource` (both impls + multi-subscriber
  observer), the new collector emit + apply_delta cases, the dispatch
  cases for new event types, the broadcast hook overrides on both
  WebUI and ConsoleCoordinatorUI, and the focus / placeholder /
  pending-count frontend assertions in `test_coordinator_page.py`.

5024 passed, ruff + mypy clean.
2026-04-30 23:48:53 -07:00
Patrick Buckley 8349d9994d feat(console): multi-select delete UX for Saved Coordinators (#458)
* feat(console): multi-select delete UX for Saved Coordinators

Mirror the per-server "Saved Workstreams" multi-select delete onto the
console's "Saved Coordinators" section.  Coordinator deletes go through
the existing routing proxy at POST /v1/api/route/workstreams/delete
(body-keyed by ws_id, since coordinators live on the node that owns
them) — no backend change required.

Pagination caps the visible page (and therefore the Select-All fan-out)
at 24.  Without it, a Select-All on a busy cluster would pin the
console proxy pool with hundreds of parallel deletes through the
fan-out router.  While in delete mode the saved-coordinators list is
frozen against SSE re-renders so visible cards don't shuffle out from
under the user's selections (drained on cancel / post-delete close).

Refactor: shared logic now lives in turnstone/shared_static/cards.{css,js}.

  * .ws-delete-* CSS moved out of ui/static/style.css into the shared
    sheet alongside .dashboard-card; the existing ui/static modal
    markup picks up class hooks instead of id-scoped rules.
  * createSavedCardsController() owns mode state, checkbox decoration,
    toolbar wiring, focus trap, modal lifecycle, and batch fan-out.
    Both ui/static (Saved Workstreams) and console/static (Saved
    Coordinators) instantiate one controller; ui/static is now ~300
    LOC lighter as a result.
  * Internalises stale-selection prune across SSE re-renders, the
    wsId->item lookup map (was O(selected x N)), and the aria-hidden
    wrap on the toggle button's emoji glyph.

Designer review tightened the affordance:

  * Modal close restores focus to the toggle button (was landing on
    <body>) — WCAG 2.4.3.
  * Modal [role="alert"] gets a red-chip treatment when populated,
    stays invisible at rest via :not(:empty).
  * Pagination consolidated onto the existing .pagination control
    (terse "X / Y" label + arrow-glyph buttons) instead of a parallel
    .coord-pagination treatment.
  * Filled destructive buttons darkened to #dc2626 in dark theme so
    the white label clears WCAG AA contrast (was 3.0:1 on --red).
    Light theme keeps --red unchanged (5.9:1 already passes).
  * Toolbar wraps below 700px viewport — Delete Selected drops to its
    own full-width row underneath count + Cancel + Select All for
    thumb-target separation.
  * .ws-card-check:focus-visible outline + word-break on
    .ws-delete-item for narrow-modal long aliases.

* fix(cards): address Copilot review feedback on PR #458

* closeModal focus restore now falls back to the section toggle button
  (opts.buttonId) when prevFocus is hidden or detached.  The post-delete
  Close path runs cancel() before closeModal(), which puts the bar at
  display:none — so the captured prevFocus (the bar's "Delete Selected"
  button) is no longer focusable and focus would land on <body>,
  defeating the WCAG 2.4.3 fix.  Esc / Cancel paths still land on the
  original focus owner because the bar stays visible in those flows.

* Saved Coordinators onClose drains _savedCoordsRetry before reloading.
  Without it, SSE events that arrived during the delete-mode freeze
  leave the retry flag true, so loadSavedCoordinators's .finally()
  re-fires a second fetch immediately after the first resolves.  Mirrors
  the same idiom in cancelCoordDeleteMode.
2026-04-30 23:48:53 -07:00
69 changed files with 8817 additions and 1704 deletions
+1 -1
View File
@@ -8,7 +8,7 @@ FROM python:3.14-slim
LABEL org.opencontainers.image.title="turnstone" \
org.opencontainers.image.description="Multi-node AI orchestration platform"
COPY --from=ghcr.io/astral-sh/uv:0.11.7 /uv /usr/local/bin/uv
COPY --from=ghcr.io/astral-sh/uv:0.11.8 /uv /usr/local/bin/uv
# Remove the slim image's man page exclusion so man-db has actual content
RUN rm -f /etc/dpkg/dpkg.cfg.d/docker
@@ -0,0 +1,233 @@
---
name: import-conversation-history
description: Use this skill when the user wants to import or migrate conversation history from another LLM chat or coding tool (e.g. ChatGPT, Claude.ai, Cursor, Copilot Chat, Aider, Gemini, a custom JSON export) into Turnstone. The skill teaches Turnstone's destination contracts — workstream identity, the OpenAI-shaped message rows, tool-call/result pairing, provider-fidelity blobs, attachments, and archive-vs-resumable choice — so the agent can map any source format onto them. Trigger phrases: "import my chats", "migrate this transcript into Turnstone", "bring my Claude.ai history over", "load this export as a workstream".
version: 1.0.0
---
# Importing Conversation History into Turnstone
## Overview
Source formats vary; the destination does not. Your job is to translate whatever the user hands you (JSON dump, ZIP export, scraped HTML, screenshot OCR, raw transcript) into Turnstone's internal shape: **one workstream row** plus an ordered sequence of **conversation rows** in OpenAI message format. This skill documents the destination so you can write a correct mapper for any source.
Two questions to settle with the user before writing anything:
1. **Archive or resumable?** An archive ("saved" workstream — `state="closed"`) is read-only history. A resumable workstream (`state="idle"`) lets the user continue the conversation; this only works cleanly when the source LLM matches a Turnstone-supported provider/model and tool definitions still resolve.
2. **One workstream per source thread, or merge?** Default to one-to-one unless the user explicitly asks to merge.
Default to **archive** when in doubt — resuming a foreign transcript with mismatched tool schemas or stale provider signatures will fail at the next turn.
## Turnstone Data Model (the destination)
Two tables carry the conversation:
### `workstreams` (one row per imported thread)
| Column | Required | Notes |
|---|---|---|
| `ws_id` | yes | 32-char lowercase hex. Auto-generate with `secrets.token_hex(16)` if you don't already have one. **First 4 hex chars are the routing bucket** — see "Identity & Routing" below. |
| `name` | yes | Short title. Pull from source thread title; fall back to first ~60 chars of first user message. |
| `state` | yes | `"closed"` for archive, `"idle"` for resumable. Never set `"running"` on import. |
| `kind` | yes | `"interactive"` for normal threads. Do NOT use `"coordinator"` for imports — that's reserved for cluster-spawned coordinator workstreams. |
| `parent_ws_id` | no | Leave NULL. Only set if you're importing a coordinator-spawned subtree and re-parenting it; rare. |
| `user_id` | yes | Owner. Must exist in `users`; importer must know which Turnstone user owns the imported history. |
| `node_id` | yes (multi-node) | Denormalized cache of the node that owns this `ws_id`'s bucket. Single-node deployments can leave it NULL or set it to the only node. |
| `alias` | no | Human-typeable short name. Optional; must be unique cluster-wide if set. |
| `title` | no | Auto-titled later by the LLM; safe to leave NULL on import. |
| `skill_id`, `skill_version` | yes | Default `""` and `0` unless the source thread was scoped to a Turnstone skill. |
| `created`, `updated` | yes | ISO8601 strings. Use the source's first/last message timestamps when available. |
### `conversations` (many rows per thread, ordered by `id`/`timestamp`)
| Column | Notes |
|---|---|
| `ws_id` | The workstream this row belongs to. |
| `timestamp` | ISO8601 string. Preserve source timestamps; fall back to monotonically increasing values if unknown. **Order is canonical via `id` (autoincrement), not `timestamp`** — but always insert in conversational order so both agree. |
| `role` | One of `system`, `user`, `assistant`, `tool`, `developer`. See role mapping below. |
| `content` | Text. May be NULL for assistant rows that are *only* tool calls. |
| `tool_name` | Set on `role="tool"` rows (the tool whose result this is). NULL otherwise. |
| `tool_call_id` | Set on `role="tool"` rows (matches the assistant row's `tool_calls[].id`). NULL otherwise. |
| `tool_calls` | JSON-encoded list, on `role="assistant"` rows that issued tool calls. OpenAI shape — see "Tool Calls" below. |
| `provider_data` | JSON blob preserving provider-native content blocks (Anthropic `signature`, Gemini `thought_signature`, etc.). Optional; only matters for **resumable** imports against the same provider. Skip for archives. |
The internal format is **OpenAI-shaped**, even when the source was Anthropic or Gemini. Providers translate at their own API boundary; storage stays uniform.
## Identity & Routing (`ws_id`)
- `ws_id` is **32-char lowercase hex** (i.e. `secrets.token_hex(16)`).
- The **routing bucket** is `int(ws_id[:4], 16)` — the first 4 hex chars place this workstream on a specific node via the consistent hash ring.
- For multi-node imports: either insert through the console's routing proxy (which forwards to the owning node), or generate `ws_id`s and write directly to each node's database in batches grouped by bucket.
- For single-node imports: bucket math is irrelevant; any `ws_id` works.
- **Do not reuse the source platform's IDs as `ws_id`** unless they happen to be 32-char hex. Generate fresh; if you need the old ID for traceability, store it in `workstream_config` under a key like `import.source_id`.
## Recommended Import Path
Three options, in order of preference:
### 1. Storage protocol (recommended for full history)
Use `turnstone.core.storage.Storage.save_messages_bulk(rows)`. This is the canonical bulk-insert primitive and bypasses the LLM round-trip entirely.
```python
from turnstone.core.storage import get_storage # construct via the same path the server uses
storage = get_storage(...) # see turnstone.core.storage.__init__ for the project's wiring
storage.create_workstream( # or whatever the project's exposed creator is — check turnstone/core/storage/_protocol.py
ws_id=ws_id,
user_id=user_id,
name=name,
state="closed",
kind="interactive",
...
)
storage.save_messages_bulk([
{"ws_id": ws_id, "role": "user", "content": "Hello"},
{"ws_id": ws_id, "role": "assistant", "content": "Hi! What can I help with?"},
{"ws_id": ws_id, "role": "assistant", "content": None,
"tool_calls": json.dumps([{"id": "call_1", "type": "function",
"function": {"name": "search", "arguments": "{\"q\":\"x\"}"}}])},
{"ws_id": ws_id, "role": "tool", "tool_name": "search", "tool_call_id": "call_1",
"content": "result text"},
# ...
])
```
`save_messages_bulk` handles `timestamp` and the workstream's `updated` column internally, so you don't need to compute them per row. **Verify the exact creator signature** by reading `turnstone/core/storage/_protocol.py` — table layout has shifted across migrations and the Storage protocol is the source of truth.
### 2. SDK `create_workstream(resume_ws=...)` (when the source is already a Turnstone workstream)
Only useful for *Turnstone → Turnstone* re-parenting. Not relevant for foreign sources.
### 3. SDK `create_workstream(initial_message=...)` + `send()` per turn (last resort)
Only fits archives where the source had **no tool calls** and you don't care about preserving assistant turns verbatim. Each `send()` triggers a real LLM round-trip, which is expensive and rewrites assistant content. Don't use this for full history.
## Role Mapping
Common source-role conventions and how they map to Turnstone:
| Source role | Turnstone `role` | Notes |
|---|---|---|
| `user`, `human` | `user` | Direct map. |
| `assistant`, `ai`, `model`, `bot` | `assistant` | Direct map. |
| `system` | `system` | Preserve only if it's content the user wrote (custom instructions). Drop boilerplate provider preambles — Turnstone composes its own system message. |
| `developer` (OpenAI o-series) | `developer` | Preserve. |
| `tool`, `function`, `tool_result` | `tool` | Must carry `tool_name` and `tool_call_id` matching the prior assistant row's `tool_calls[].id`. |
| `tool_use` (Anthropic) | `assistant` with `tool_calls` | Anthropic emits tool calls *inside* an assistant message; flatten to OpenAI shape. |
| `human_feedback`, `revision` | `user` | Treat as a follow-up user turn. |
## Tool Calls (the most error-prone part)
Turnstone stores tool calls in OpenAI's nested-function shape on the assistant row, and matches them with `role="tool"` result rows by `tool_call_id`.
### Assistant row with tool calls
```json
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "search_web",
"arguments": "{\"query\":\"turnstone import\"}"
}
}
]
}
```
`tool_calls[].function.arguments` is **a JSON-encoded string**, not an object. Source formats commonly get this wrong — Anthropic stores arguments as a parsed object, Gemini as a struct. Always re-serialize to a string.
### Tool result row
```json
{
"role": "tool",
"tool_name": "search_web",
"tool_call_id": "call_abc123",
"content": "..."
}
```
Pairing rules:
- Every assistant `tool_calls[].id` MUST be followed by exactly one `role="tool"` row with the matching `tool_call_id`, before the next user/assistant turn.
- If the source dropped the tool result (cut-off transcript), insert a synthetic `role="tool"` row with `content="[tool result missing in source]"` to keep the chain valid. An assistant row with an unanswered `tool_calls[].id` will break replay and any LLM round-trip.
- Multi-tool assistant turns: one `role="tool"` row per call, in any order, all before the next non-tool row.
### Tool ID generation
If the source used opaque tool IDs that aren't unique within a thread (some platforms reuse them), regenerate with a stable scheme like `f"call_{i}"` where `i` is a per-thread counter. Update both the assistant and tool rows together.
## Provider Fidelity (`provider_data`)
Skip this entirely for **archive** imports.
For **resumable** imports against the same provider, populate `provider_data` to preserve provider-specific tool-call metadata that the next API round-trip will require:
- **Anthropic**: `signature` field on thinking blocks; required for round-tripping extended-thinking responses.
- **Gemini**: `thought_signature` on tool calls; required for fidelity.
- **OpenAI**: typically nothing to preserve.
The runtime-side dict key is `_provider_content` (a list of provider-native blocks); the persisted column is `provider_data` (the same list, JSON-encoded). If you don't have provider-native blocks from the source — and you usually won't, because a foreign export won't include them — leave `provider_data` NULL. The first new turn will succeed without it, but the previous assistant turn's reasoning won't replay back to the model.
## Attachments
If the source thread had image or file attachments:
- **Size limits**: images ≤ 4 MiB, text documents ≤ 512 KiB. Reject or downsample anything bigger.
- **Allowed types**: server validates magic bytes for images and UTF-8-decodes for text. Binary blobs that aren't images won't pass.
- **Lifecycle**: pending → reserved → consumed. For imports, the cleanest path is to upload as pending and immediately consume by attaching to the relevant `conversations.id`.
Two import paths:
1. **Bulk-insert + post-attach**: insert messages first, get back the assistant/user `conversations.id`, then write `workstream_attachments` rows linking the file to `message_id`.
2. **SDK multipart create**: `create_workstream(attachments=[...], initial_message=...)` for the *first* turn only — the server reserves and consumes them onto that turn. Doesn't help for mid-thread attachments.
For full-history imports with multiple attachments at different turns, path (1) is the only option.
## Validation Checklist
Before declaring success, verify:
- [ ] `ws_id` is 32-char lowercase hex.
- [ ] `workstreams` row exists with the right `user_id`, `state`, `kind`.
- [ ] Conversation rows are inserted **in order** (autoincrement `id` will reflect insert order).
- [ ] Every assistant `tool_calls[].id` has a matching `role="tool"` row with the same `tool_call_id`.
- [ ] `tool_calls[].function.arguments` is a JSON-encoded **string**, not a parsed object.
- [ ] First message is typically `role="user"` (not `system`) — Turnstone composes its own system prompt at runtime.
- [ ] No empty assistant rows (`content=NULL` AND `tool_calls=NULL` is invalid).
- [ ] If multi-node: the `ws_id`'s bucket maps to a node that exists; `workstreams.node_id` matches.
- [ ] Round-trip test: run `Storage.load_messages(ws_id)` and confirm the reconstructed list matches what you inserted (modulo timestamps).
## Anti-patterns
- **Don't import the source provider's system prompt verbatim.** Provider boilerplate ("You are Claude...", "You are ChatGPT...") will conflict with Turnstone's composed system message and confuse the model on resume. Drop it; preserve only user-authored custom instructions.
- **Don't preserve foreign tool definitions as Turnstone tools.** If the source had custom tools that don't exist in Turnstone, the assistant rows that called them are still valid history (archive), but the workstream is **not resumable** — mark `state="closed"`.
- **Don't fabricate `tool_call_id`s without re-pairing.** Mismatched ids silently break the replay chain on the next turn.
- **Don't skip the `tool_name` field on `role="tool"` rows.** Some load paths use it for display and audit; NULL there will render as "unknown tool".
- **Don't write through the LLM (`send()` per turn) for full history.** It's expensive, rewrites assistant turns, and rate-limits will bite long imports.
## Quick Reference
| Task | Path |
|---|---|
| Generate ws_id | `secrets.token_hex(16)` |
| Bulk insert messages | `Storage.save_messages_bulk(rows)` |
| Archive (read-only) | `state="closed"`, skip `provider_data` |
| Resumable | `state="idle"`, populate `provider_data` if same provider |
| Tool call id | OpenAI shape: `{"id": ..., "type": "function", "function": {"name": ..., "arguments": "<json string>"}}` |
| Tool result row | `role="tool"`, `tool_name`, `tool_call_id`, `content` |
| Source role → Turnstone role | See "Role Mapping" table |
| Per-thread metadata | Store source IDs in `workstream_config` under `import.*` keys |
## Files to read before writing the importer
- `turnstone/core/storage/_schema.py` — authoritative table definitions.
- `turnstone/core/storage/_protocol.py``save_message`, `save_messages_bulk`, `load_messages` signatures.
- `turnstone/core/session.py` (around the message-save section) — how the runtime constructs in-memory message dicts; mirror this shape on import to round-trip cleanly.
- `turnstone/api/server_schemas.py` — Pydantic shapes for the SDK paths if you go through HTTP.
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "turnstone"
version = "1.5.3"
version = "1.5.6"
description = "Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation."
readme = "README.md"
license = "BUSL-1.1"
+3 -4
View File
@@ -35,11 +35,10 @@ def _seed_children(
The production path populates the registry via the cluster-event
fan-out thread observing ``ws_created`` events. These tests just
need a known-children set for the endpoint handlers to iterate —
inject directly under ``_children_lock`` rather than spinning up
the collector + fan-out plumbing.
inject directly via the registry's bulk-merge surface rather than
spinning up the collector + fan-out plumbing.
"""
with adapter._children_lock:
adapter._merge_child_ids_locked(coord_ws_id, child_ws_ids)
adapter._registry.merge_children(coord_ws_id, child_ws_ids)
class _AuthMiddleware(BaseHTTPMiddleware):
+28
View File
@@ -0,0 +1,28 @@
"""Shared test helpers — kept out of conftest.py since these are factories,
not fixtures, and several test files want to import them directly."""
from __future__ import annotations
from typing import Any
from unittest.mock import MagicMock
def make_chat_session(**overrides: Any) -> Any:
"""Build a minimal ``ChatSession`` with sane test defaults.
Caller passes any constructor arg as a kwarg to override the default —
e.g. ``make_chat_session(memory_config=MemoryConfig(fetch_limit=5))``.
"""
from turnstone.core.session import ChatSession
defaults: dict[str, Any] = {
"client": MagicMock(),
"model": "test-model",
"ui": MagicMock(),
"instructions": None,
"temperature": 0.5,
"max_tokens": 4096,
"tool_timeout": 30,
}
defaults.update(overrides)
return ChatSession(**defaults)
@@ -352,6 +352,90 @@ def test_update_endpoint_skips_refresh_on_empty_body(
assert calls == [] # gate held: empty body did not trigger a refresh
def test_create_rejects_invalid_api_surface(storage: SQLiteBackend) -> None:
"""POST with a bogus server_compat.api_surface returns 400 rather than
persisting a value that would make get_provider() raise on every later
ChatSession init for the alias."""
_seed_model_def(storage, definition_id="m1", alias="local", model="m")
registry = _make_registry(alias="local", model="m")
client = _make_client(storage, registry)
resp = client.post(
"/v1/api/admin/model-definitions",
json={
"alias": "bad",
"model": "x",
"provider": "openai-compatible",
"base_url": "http://localhost:9000/v1",
"api_key": "sk-x",
"capabilities": {"server_compat": {"api_surface": "BOGUS"}},
},
)
assert resp.status_code == 400, resp.text
assert "api_surface" in resp.json()["error"]
# And the alias is not persisted
assert not registry.has_alias("bad")
def test_create_rejects_non_canonical_api_surface(storage: SQLiteBackend) -> None:
"""Strict validation: ' Responses ' / 'CHAT' don't round-trip through the
admin <select>, so they're rejected even though they'd survive a
case-insensitive membership check."""
_seed_model_def(storage, definition_id="m1", alias="local", model="m")
registry = _make_registry(alias="local", model="m")
client = _make_client(storage, registry)
for bad in (" responses ", "RESPONSES", "Chat"):
resp = client.post(
"/v1/api/admin/model-definitions",
json={
"alias": "noncanon",
"model": "x",
"provider": "openai-compatible",
"base_url": "http://localhost:9000/v1",
"api_key": "sk-x",
"capabilities": {"server_compat": {"api_surface": bad}},
},
)
assert resp.status_code == 400, f"{bad!r}: {resp.text}"
def test_create_accepts_valid_api_surface(storage: SQLiteBackend) -> None:
"""Canonical 'chat' / 'responses' / unset are all accepted and persisted."""
_seed_model_def(storage, definition_id="m1", alias="local", model="m")
registry = _make_registry(alias="local", model="m")
client = _make_client(storage, registry)
resp = client.post(
"/v1/api/admin/model-definitions",
json={
"alias": "responses-alias",
"model": "x",
"provider": "openai-compatible",
"base_url": "http://localhost:9000/v1",
"api_key": "sk-x",
"capabilities": {"server_compat": {"api_surface": "responses"}},
},
)
assert resp.status_code == 200, resp.text
assert registry.has_alias("responses-alias")
def test_update_rejects_invalid_api_surface(storage: SQLiteBackend) -> None:
"""PUT path also gates the validation, so an admin can't smuggle a bad
value into an existing alias."""
_seed_model_def(storage, definition_id="m1", alias="local", model="m")
registry = _make_registry(alias="local", model="m")
client = _make_client(storage, registry)
resp = client.put(
"/v1/api/admin/model-definitions/m1",
json={"capabilities": {"server_compat": {"api_surface": "junk"}}},
)
assert resp.status_code == 400, resp.text
assert "api_surface" in resp.json()["error"]
def test_delete_endpoint_refreshes_registry(storage: SQLiteBackend) -> None:
"""DELETE drops the alias from the in-process registry too — a
coord session that tried to resolve the deleted alias would
+68
View File
@@ -92,3 +92,71 @@ def test_tool_error_does_not_overwrite_approval_badge() -> None:
"badge instead so the approval verdict stays visible alongside "
"the error."
)
def test_replay_history_renders_content_before_tool_block() -> None:
"""In ``replayHistory``'s ``role === "assistant"`` branch, the
``msg.content`` render must precede the ``msg.tool_calls`` render.
Two reasons, both load-bearing:
1. **Structural** — the next loop iteration's ``role === "tool"``
message anchors to ``lastToolBlock``. The tool-block branch sets
that anchor; the content branch clears it. If content runs after
the tool block, the clear silently drops the upcoming tool
result. Pre-fix, every interactive tool result was missing from
saved-workstream replays whenever the assistant turn carried
both narration and tool calls (very common output shape).
2. **Visual** — the live SSE path renders content first
(``stream_text`` streams before ``tool_info`` /
``approve_request``), so replay should match.
The test pins the order via the offsets of the ``msg.content`` and
``msg.tool_calls`` branch headers inside the function body."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.replayHistory = function")
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
fn = body[start:end]
# Locate the assistant branch and bound the search to its body —
# the function also handles user / tool roles which would otherwise
# confuse the offset comparison.
asst_start = fn.index('msg.role === "assistant"')
asst_end = fn.index('msg.role === "tool"', asst_start)
asst = fn[asst_start:asst_end]
content_idx = asst.index("if (msg.content)")
tool_calls_idx = asst.index("if (msg.tool_calls && msg.tool_calls.length)")
assert content_idx < tool_calls_idx, (
"replayHistory must render msg.content BEFORE msg.tool_calls "
"inside the assistant branch — otherwise the lastToolBlock "
"anchor is clobbered before the next iteration's tool result "
"can attach to it (and the visual order also drifts from the "
"live SSE flow)."
)
def test_replay_history_renders_persisted_verdict_badge() -> None:
"""Saved-workstream replays must paint the persisted intent verdict
next to each tool div, using the same ``renderVerdictBadge`` helper
the live ``showInlineToolBlock`` path uses. Pre-fix the audit trail
was complete in storage (``intent_verdicts`` table) but never
surfaced on replay — operators reviewing a saved workstream
couldn't see what the heuristic / LLM judge thought of any tool
call. This test pins the call site so a refactor that drops the
decoration regresses the audit surface."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.replayHistory = function")
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
fn = body[start:end]
# Match a `renderVerdictBadge(<something>.verdict, ...)` call inside
# the replay loop. Loose on whitespace + identifier so a future
# rename of the iteration variable doesn't trip CI.
badge_call_re = re.compile(
r"renderVerdictBadge\(\s*\w+\.verdict\b",
)
assert badge_call_re.search(fn), (
"replayHistory must call renderVerdictBadge(tc.verdict, ...) "
"when a persisted verdict is attached to a tool_call entry — "
"otherwise the audit-trail data persisted to intent_verdicts "
"doesn't surface on saved-workstream replays."
)
+270
View File
@@ -0,0 +1,270 @@
"""Unit tests for :mod:`turnstone.core.child_source`.
Covers both strategies in isolation against fakes — no live collector,
no live SessionManager. Adapter-level integration coverage continues to
live in ``test_coordinator_adapter.py``.
"""
from __future__ import annotations
import contextlib
import time
from typing import TYPE_CHECKING, Any
from turnstone.core.child_source import ClusterChildSource, SameNodeChildSource
from turnstone.core.children_registry import ChildrenRegistry
from turnstone.core.workstream import WorkstreamState
if TYPE_CHECKING:
import queue
# ---------------------------------------------------------------------------
# SameNodeChildSource
# ---------------------------------------------------------------------------
class _FakeManager:
"""Minimal SessionManager stand-in implementing the subscribe API."""
def __init__(self) -> None:
self.subscribers: list[Any] = []
def subscribe_to_state(self, callback: Any) -> None:
self.subscribers.append(callback)
def unsubscribe_from_state(self, callback: Any) -> None:
with contextlib.suppress(ValueError):
self.subscribers.remove(callback)
def fire(self, ws_id: str, state: WorkstreamState) -> None:
for cb in self.subscribers:
cb(ws_id, state)
class TestSameNodeChildSource:
def test_start_subscribes_to_manager(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
sink_calls: list[dict[str, Any]] = []
src.start(sink=sink_calls.append)
assert len(mgr.subscribers) == 1
def test_state_change_for_known_child_pushes_to_sink(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
registry.install("p1", object())
registry.add_child("p1", "c1")
src = SameNodeChildSource(mgr, registry)
sink_calls: list[dict[str, Any]] = []
src.start(sink=sink_calls.append)
mgr.fire("c1", WorkstreamState.RUNNING)
assert len(sink_calls) == 1
ev = sink_calls[0]
assert ev["type"] == "cluster_state"
assert ev["ws_id"] == "c1"
assert ev["state"] == "running"
def test_state_change_for_unknown_workstream_is_dropped(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
sink_calls: list[dict[str, Any]] = []
src.start(sink=sink_calls.append)
# No registry entry — pre-filter drops the event without
# invoking the sink.
mgr.fire("ws-unknown", WorkstreamState.IDLE)
assert sink_calls == []
def test_shutdown_unsubscribes(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
src.start(sink=lambda ev: None)
assert len(mgr.subscribers) == 1
src.shutdown()
assert mgr.subscribers == []
def test_start_is_idempotent(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
src = SameNodeChildSource(mgr, registry)
src.start(sink=lambda ev: None)
src.start(sink=lambda ev: None)
# Second start is a no-op; only one subscription.
assert len(mgr.subscribers) == 1
def test_sink_exception_does_not_propagate(self) -> None:
mgr = _FakeManager()
registry = ChildrenRegistry()
registry.install("p1", object())
registry.add_child("p1", "c1")
src = SameNodeChildSource(mgr, registry)
def bad_sink(ev: dict[str, Any]) -> None:
raise RuntimeError("sink boom")
src.start(sink=bad_sink)
# Should not raise — the strategy catches sink failures and logs.
mgr.fire("c1", WorkstreamState.RUNNING)
# ---------------------------------------------------------------------------
# ClusterChildSource
# ---------------------------------------------------------------------------
class _FakeCollector:
"""Minimal ClusterCollector stand-in providing the listener API."""
def __init__(self, snapshot: dict[str, Any] | None = None) -> None:
self._snapshot = snapshot or {"nodes": []}
self.queues: list[queue.Queue[dict[str, Any]]] = []
self.unregistered: list[queue.Queue[dict[str, Any]]] = []
def get_snapshot_and_register(self, q: queue.Queue[dict[str, Any]]) -> dict[str, Any]:
self.queues.append(q)
return self._snapshot
def unregister_listener(self, q: queue.Queue[dict[str, Any]]) -> None:
self.unregistered.append(q)
def emit(self, event: dict[str, Any]) -> None:
"""Push an event to all registered listener queues."""
for q in self.queues:
q.put(event)
class TestClusterChildSource:
def test_start_subscribes_to_collector(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
try:
src.start(sink=lambda ev: None)
assert len(coll.queues) == 1
finally:
src.shutdown()
def test_start_primes_registry_from_snapshot(self) -> None:
snapshot = {
"nodes": [
{
"workstreams": [
{"id": "c1", "parent_ws_id": "p1"},
{"id": "c2", "parent_ws_id": "p1"},
# Unknown parent — dropped
{"id": "x", "parent_ws_id": "p-unknown"},
],
},
],
}
coll = _FakeCollector(snapshot)
registry = ChildrenRegistry()
registry.install("p1", object())
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=lambda: ["p1"],
)
try:
src.start(sink=lambda ev: None)
assert set(registry.children_of("p1")) == {"c1", "c2"}
assert registry.parent_for("x") is None
finally:
src.shutdown()
def test_event_dispatched_to_sink(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
sink_calls: list[dict[str, Any]] = []
try:
src.start(sink=sink_calls.append)
coll.emit({"type": "cluster_state", "ws_id": "c1", "state": "running"})
# Daemon thread loop has 1.0s queue timeout; poll briefly.
for _ in range(20):
if sink_calls:
break
time.sleep(0.05)
assert len(sink_calls) == 1
assert sink_calls[0]["ws_id"] == "c1"
finally:
src.shutdown()
def test_shutdown_unregisters_and_joins_thread(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
src.start(sink=lambda ev: None)
src.shutdown()
assert coll.unregistered == coll.queues
# Second shutdown is a no-op (idempotent).
src.shutdown()
def test_start_is_idempotent(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
try:
src.start(sink=lambda ev: None)
src.start(sink=lambda ev: None)
assert len(coll.queues) == 1
finally:
src.shutdown()
def test_sink_exception_does_not_kill_thread(self) -> None:
coll = _FakeCollector()
registry = ChildrenRegistry()
src = ClusterChildSource(
collector=coll,
registry=registry,
parents_provider=list,
)
survived_calls: list[dict[str, Any]] = []
call_count = [0]
def flaky_sink(ev: dict[str, Any]) -> None:
call_count[0] += 1
if call_count[0] == 1:
raise RuntimeError("first one boom")
survived_calls.append(ev)
try:
src.start(sink=flaky_sink)
coll.emit({"type": "cluster_state", "ws_id": "c1", "state": "x"})
coll.emit({"type": "cluster_state", "ws_id": "c2", "state": "y"})
for _ in range(40):
if survived_calls:
break
time.sleep(0.05)
assert len(survived_calls) == 1
assert survived_calls[0]["ws_id"] == "c2"
finally:
src.shutdown()
# Multi-subscriber observer tests for ``SessionManager.subscribe_to_state``
# / ``unsubscribe_from_state`` live in ``test_session_manager.py`` where
# the proper FakeAdapter / FakeStorage construction helpers already exist.
+239
View File
@@ -0,0 +1,239 @@
"""Unit tests for :class:`turnstone.core.children_registry.ChildrenRegistry`.
The registry was lifted from ``CoordinatorAdapter`` in Stage 3 Step 1.
Adapter-level coverage for the integrated behavior already lives in
``test_coordinator_adapter.py``; this file pins the data structure
invariants in isolation so the registry can be reused by future
``ChildSource`` strategies (Step 2) without re-deriving the behavior
from the adapter test surface.
"""
from __future__ import annotations
import threading
import pytest
from turnstone.core.children_registry import ChildrenRegistry
class _Sentinel:
"""Lightweight UI stand-in; identity-comparable, no behavior."""
@pytest.fixture
def registry() -> ChildrenRegistry:
return ChildrenRegistry()
# ---------------------------------------------------------------------------
# install / uninstall
# ---------------------------------------------------------------------------
class TestInstallUninstall:
def test_install_seeds_empty_child_set_and_presence(self, registry: ChildrenRegistry) -> None:
ui = _Sentinel()
registry.install("p1", ui)
assert registry.children_of("p1") == []
assert registry.ui_for("p1") is ui
assert registry.parents() == ["p1"]
def test_install_is_idempotent_repoints_ui_keeps_children(
self, registry: ChildrenRegistry
) -> None:
ui_a = _Sentinel()
ui_b = _Sentinel()
registry.install("p1", ui_a)
registry.merge_children("p1", ["c1", "c2"])
registry.install("p1", ui_b)
assert registry.ui_for("p1") is ui_b
assert set(registry.children_of("p1")) == {"c1", "c2"}
def test_uninstall_clears_forward_reverse_and_presence(
self, registry: ChildrenRegistry
) -> None:
ui = _Sentinel()
registry.install("p1", ui)
registry.merge_children("p1", ["c1", "c2"])
registry.uninstall("p1")
assert registry.children_of("p1") == []
assert registry.ui_for("p1") is None
assert registry.parents() == []
assert registry.parent_for("c1") is None
assert registry.parent_for("c2") is None
def test_uninstall_unknown_parent_is_noop(self, registry: ChildrenRegistry) -> None:
registry.uninstall("never-installed") # must not raise
def test_uninstall_does_not_clobber_other_parents(self, registry: ChildrenRegistry) -> None:
registry.install("p1", _Sentinel())
registry.install("p2", _Sentinel())
registry.merge_children("p1", ["c1"])
registry.merge_children("p2", ["c2"])
registry.uninstall("p1")
assert registry.parent_for("c1") is None
assert registry.parent_for("c2") == "p2"
assert registry.parents() == ["p2"]
# ---------------------------------------------------------------------------
# add_child — atomic check-and-route
# ---------------------------------------------------------------------------
class TestAddChild:
def test_add_child_returns_ui_on_success(self, registry: ChildrenRegistry) -> None:
ui = _Sentinel()
registry.install("p1", ui)
assert registry.add_child("p1", "c1") is ui
assert registry.parent_for("c1") == "p1"
assert registry.children_of("p1") == ["c1"]
def test_add_child_returns_none_when_parent_not_installed(
self, registry: ChildrenRegistry
) -> None:
assert registry.add_child("absent", "c1") is None
assert registry.parent_for("c1") is None
def test_add_child_returns_none_on_duplicate(self, registry: ChildrenRegistry) -> None:
ui = _Sentinel()
registry.install("p1", ui)
assert registry.add_child("p1", "c1") is ui
# second add for same child returns None — caller must not
# double-dispatch.
assert registry.add_child("p1", "c1") is None
assert registry.children_of("p1") == ["c1"]
# ---------------------------------------------------------------------------
# merge_children — bulk seeding
# ---------------------------------------------------------------------------
class TestMergeChildren:
def test_merge_seeds_forward_and_reverse(self, registry: ChildrenRegistry) -> None:
registry.merge_children("p1", ["c1", "c2", "c3"])
assert set(registry.children_of("p1")) == {"c1", "c2", "c3"}
for cid in ("c1", "c2", "c3"):
assert registry.parent_for(cid) == "p1"
def test_merge_is_idempotent(self, registry: ChildrenRegistry) -> None:
registry.merge_children("p1", ["c1"])
registry.merge_children("p1", ["c1"])
assert registry.children_of("p1") == ["c1"]
def test_merge_skips_empty_or_falsy_ids(self, registry: ChildrenRegistry) -> None:
registry.merge_children("p1", ["", "c1", "", "c2"])
assert set(registry.children_of("p1")) == {"c1", "c2"}
def test_merge_does_not_require_install(self, registry: ChildrenRegistry) -> None:
# Snapshot-priming may run before the parent's install fires —
# the merge still seeds the forward set so the install picks
# the children up. (Storage-seeded rebuild relies on this.)
registry.merge_children("p1", ["c1"])
assert registry.children_of("p1") == ["c1"]
# ui_for is still None because install hasn't run
assert registry.ui_for("p1") is None
# ---------------------------------------------------------------------------
# Lookups — return copies, not live refs
# ---------------------------------------------------------------------------
class TestLookups:
def test_children_of_returns_copy(self, registry: ChildrenRegistry) -> None:
registry.install("p1", _Sentinel())
registry.merge_children("p1", ["c1", "c2"])
snap = registry.children_of("p1")
snap.append("c3-injected")
assert "c3-injected" not in registry.children_of("p1")
def test_children_of_unknown_parent_returns_empty(self, registry: ChildrenRegistry) -> None:
assert registry.children_of("absent") == []
def test_parent_for_unknown_child_returns_none(self, registry: ChildrenRegistry) -> None:
assert registry.parent_for("absent") is None
def test_parents_returns_copy(self, registry: ChildrenRegistry) -> None:
registry.install("p1", _Sentinel())
snap = registry.parents()
snap.append("p2-injected")
assert "p2-injected" not in registry.parents()
# ---------------------------------------------------------------------------
# Concurrency — concurrent add_child must not exceed the unique-set
# invariant or leave a half-installed reverse-index entry.
# ---------------------------------------------------------------------------
class TestConcurrency:
def test_concurrent_add_child_returns_ui_exactly_once_per_unique(
self, registry: ChildrenRegistry
) -> None:
ui = _Sentinel()
registry.install("p1", ui)
results: list[object] = []
results_lock = threading.Lock()
def attempt_add(child_id: str) -> None:
r = registry.add_child("p1", child_id)
with results_lock:
results.append(r)
threads = [threading.Thread(target=attempt_add, args=("c1",)) for _ in range(20)]
for t in threads:
t.start()
for t in threads:
t.join()
# Exactly one thread sees the UI; the remaining 19 see None
# (duplicate). The forward + reverse indexes carry exactly one
# entry for c1.
successes = [r for r in results if r is ui]
nones = [r for r in results if r is None]
assert len(successes) == 1
assert len(nones) == 19
assert registry.children_of("p1") == ["c1"]
assert registry.parent_for("c1") == "p1"
def test_concurrent_install_and_add_child_no_resurrect(
self, registry: ChildrenRegistry
) -> None:
# add_child racing with uninstall: either lands first (registry
# populated) or the parent is gone (returns None). Must NOT
# leave a forward-set entry without presence — that would be
# the "resurrected after close" leak the locked dispatch path
# was guarding against.
ui = _Sentinel()
registry.install("p1", ui)
outcomes: list[object] = []
def adder() -> None:
outcomes.append(registry.add_child("p1", "c1"))
def uninstaller() -> None:
registry.uninstall("p1")
threads = [
threading.Thread(target=adder),
threading.Thread(target=uninstaller),
]
for t in threads:
t.start()
for t in threads:
t.join()
# If add_child landed first: c1 is in the forward set, then
# uninstall clears everything. End state: nothing.
# If uninstall landed first: add_child sees no presence,
# returns None, no entry added. End state: nothing.
# Either way, the leak invariant holds: child set is empty or
# parent is gone, never "child set populated but no presence".
children = registry.children_of("p1")
ui_present = registry.ui_for("p1") is not None
if children:
assert ui_present, "registry leaked: children set without presence"
+18 -2
View File
@@ -30,9 +30,25 @@ def _full_hdr() -> dict[str, str]:
}
@pytest.fixture(autouse=True)
def _isolate_metrics(monkeypatch):
"""Swap ``turnstone.server._metrics`` for a fresh collector
per-test, with auto-restore.
Bare ``srv_mod._metrics = MetricsCollector()`` (the prior
pattern) leaks into any test file that already bound the name
via ``from turnstone.server import _metrics`` at import time —
those tests' patches then operate on a different instance from
the one the live ``_publish_models_metadata`` reads, and the
monkeypatch silently no-ops. ``monkeypatch.setattr`` restores
after the test, so the leak is contained.
"""
fresh = MetricsCollector()
fresh.model = "test-model"
monkeypatch.setattr(srv_mod, "_metrics", fresh)
def _make_app(storage: Any) -> TestClient:
srv_mod._metrics = MetricsCollector()
srv_mod._metrics.model = "test-model"
mock_session = MagicMock()
mock_ws = MagicMock()
mock_ws.id = "ws-target"
+151 -24
View File
@@ -314,12 +314,14 @@ class TestCollectorSnapshot:
assert event["ws_id"] == "ws1"
assert event["state"] == "running"
def test_apply_snapshot_state_change_forwards_pending_approval_detail(self):
"""Reconnect-via-snapshot is the resync path after every console
restart or network blip. Without forwarding the field here,
a child sitting in approval-pending across the gap renders as
``activity_state=approval`` with no buttons until the next
state change — broken UX during the most common re-sync event."""
def test_apply_snapshot_state_change_does_not_carry_pending_approval_detail(self):
"""Stage 3 cleanup — the snapshot-resync cluster_state event no
longer piggybacks ``pending_approval_detail`` (the field is
gone from cluster_state entirely). On reconnect the browser's
bulk fetch — triggered by the ``activity_state="approval"``
transition in the reducer — pulls the items directly from
``ui.serialize_pending_approval_detail()`` via the dashboard
endpoint."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
@@ -329,10 +331,6 @@ class TestCollectorSnapshot:
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
detail = {
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": False,
}
c._apply_snapshot(
"node-a",
{
@@ -344,7 +342,6 @@ class TestCollectorSnapshot:
"name": "same",
"state": "running",
"activity_state": "approval",
"pending_approval_detail": detail,
}
],
"health": {},
@@ -354,7 +351,8 @@ class TestCollectorSnapshot:
event = q.get_nowait()
assert event["type"] == "cluster_state"
assert event["pending_approval_detail"] == detail
assert event["activity_state"] == "approval"
assert "pending_approval_detail" not in event
def test_apply_snapshot_skips_empty_id_workstream(self):
c = _make_collector()
@@ -401,12 +399,12 @@ class TestCollectorDelta:
# Verify in-memory state was updated
assert c._nodes["node-a"].workstreams["ws1"]["state"] == "running"
def test_apply_delta_ws_state_forwards_pending_approval_detail(self):
"""The rich approval payload now travels on the cluster bus so
coord tabs can render inline approve/deny buttons in lockstep
with the activity_state transition. Collector must forward
the field verbatim — the adapter does the child-routing on
top, but the bus carries the data."""
def test_apply_delta_ws_state_does_not_carry_pending_approval_detail(self):
"""Stage 3 cleanup — ``cluster_state`` no longer carries the
``pending_approval_detail`` piggyback. Approval items now arrive
via bulk fetch on activity_state transition; verdicts via the
explicit ``intent_verdict`` event class. Symmetric event flow,
no piggyback to dedupe against."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
@@ -416,10 +414,6 @@ class TestCollectorDelta:
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
detail = {
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": False,
}
c._apply_delta(
"node-a",
{
@@ -427,13 +421,13 @@ class TestCollectorDelta:
"ws_id": "ws1",
"state": "running",
"activity_state": "approval",
"pending_approval_detail": detail,
},
)
event = q.get_nowait()
assert event["type"] == "cluster_state"
assert event["pending_approval_detail"] == detail
assert event["activity_state"] == "approval"
assert "pending_approval_detail" not in event
def test_apply_delta_ws_created(self):
c = _make_collector()
@@ -480,6 +474,139 @@ class TestCollectorDelta:
assert event["name"] == "new-name"
assert c._nodes["node-a"].workstreams["ws1"]["name"] == "new-name"
def test_apply_delta_intent_verdict_forwards_verbatim(self):
"""Stage 3 Step 5 — node-emitted intent_verdict events flow
through _apply_delta to cluster fan-out so coord adapters can
re-emit as child_ws_intent_verdict on the parent's SSE."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
verdict = {
"call_id": "c1",
"risk_level": "low",
"confidence": 0.9,
"recommendation": "approve",
}
c._apply_delta(
"node-a",
{"type": "intent_verdict", "ws_id": "ws1", "verdict": verdict},
)
event = q.get_nowait()
assert event["type"] == "intent_verdict"
assert event["ws_id"] == "ws1"
assert event["node_id"] == "node-a"
assert event["verdict"] == verdict
def test_apply_delta_intent_verdict_drops_when_ws_id_missing(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta("node-a", {"type": "intent_verdict", "verdict": {}})
assert q.empty()
def test_apply_delta_approval_resolved_forwards_verbatim(self):
"""Stage 3 Step 5 — paired with intent_verdict; clears the
coord tree's pending-approval pill in lockstep with the
actual decision rather than waiting for the state-change
piggyback."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta(
"node-a",
{
"type": "approval_resolved",
"ws_id": "ws1",
"approved": True,
"feedback": "lgtm",
"always": False,
},
)
event = q.get_nowait()
assert event["type"] == "approval_resolved"
assert event["ws_id"] == "ws1"
assert event["node_id"] == "node-a"
assert event["approved"] is True
assert event["feedback"] == "lgtm"
assert event["always"] is False
def test_apply_delta_approve_request_forwards_detail(self):
"""Push path for the initial approval items — eliminates the
bulk-fetch race that left the coord row stuck on a loading
placeholder when the bulk fetch landed in the gap between
_emit_state(ATTENTION) and approve_tools setting _pending_approval."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
detail = {
"type": "approve_request",
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": True,
}
c._apply_delta(
"node-a",
{"type": "approve_request", "ws_id": "ws1", "detail": detail},
)
event = q.get_nowait()
assert event["type"] == "approve_request"
assert event["ws_id"] == "ws1"
assert event["node_id"] == "node-a"
assert event["detail"] == detail
def test_apply_delta_approve_request_drops_when_ws_id_missing(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta("node-a", {"type": "approve_request", "detail": {}})
assert q.empty()
def test_apply_delta_approval_resolved_coerces_missing_fields(self):
"""Defensive: ``approved`` / ``always`` / ``feedback`` may be
omitted by older nodes mid-rolling-upgrade; collector coerces
to safe defaults."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "test", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_delta("node-a", {"type": "approval_resolved", "ws_id": "ws1"})
event = q.get_nowait()
assert event["approved"] is False
assert event["feedback"] == ""
assert event["always"] is False
def test_apply_delta_health_changed(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
+145
View File
@@ -464,3 +464,148 @@ def test_coord_budget_override_survives_wildcard_allow_policy() -> None:
assert approved is True
types = [e.get("type") for e in captured_events]
assert "approve_request" in types, "Wildcard allow must not strip the budget-override prompt"
# ---------------------------------------------------------------------------
# Cluster-bus broadcast hooks — _broadcast_intent_verdict / _approval_resolved
# ---------------------------------------------------------------------------
class TestBroadcastIntentVerdict:
"""``ConsoleCoordinatorUI._broadcast_intent_verdict`` overrides the
no-op base hook to push the verdict onto the cluster bus via
``ClusterCollector.emit_console_ws_intent_verdict``. The far more
common path is the per-node ``WebUI`` override (covered in
test_webui_content.py); this lights up the rare coord-self path
(a coord that runs its own LLM judge).
"""
def test_calls_collector_emit_with_ws_id_and_verdict(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
verdict = {
"call_id": "c1",
"risk_level": "high",
"confidence": 0.91,
}
ui._broadcast_intent_verdict(verdict)
collector.emit_console_ws_intent_verdict.assert_called_once_with(
"coord-a",
verdict,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_no_op_when_collector_unset(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
ConsoleCoordinatorUI._collector = None
# Doesn't raise.
ui._broadcast_intent_verdict({"call_id": "c1"})
def test_collector_exception_swallowed(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
collector.emit_console_ws_intent_verdict.side_effect = RuntimeError("boom")
ConsoleCoordinatorUI._collector = collector
try:
# Doesn't raise — collector failures are observational only.
ui._broadcast_intent_verdict({"call_id": "c1"})
finally:
ConsoleCoordinatorUI._collector = None
class TestBroadcastApprovalResolved:
"""``ConsoleCoordinatorUI._broadcast_approval_resolved`` overrides
the base hook to push the resolution onto the cluster bus via
``ClusterCollector.emit_console_ws_approval_resolved``."""
def test_calls_collector_with_decision_fields(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
ui._broadcast_approval_resolved(True, "lgtm", always=True)
collector.emit_console_ws_approval_resolved.assert_called_once_with(
"coord-a",
approved=True,
feedback="lgtm",
always=True,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_normalises_none_feedback_to_empty_string(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
ui._broadcast_approval_resolved(False, None)
collector.emit_console_ws_approval_resolved.assert_called_once_with(
"coord-a",
approved=False,
feedback="",
always=False,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_no_op_when_collector_unset(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
ConsoleCoordinatorUI._collector = None
# Doesn't raise.
ui._broadcast_approval_resolved(True, None)
def test_collector_exception_swallowed(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
collector.emit_console_ws_approval_resolved.side_effect = RuntimeError("boom")
ConsoleCoordinatorUI._collector = collector
try:
# Doesn't raise.
ui._broadcast_approval_resolved(True, "ok")
finally:
ConsoleCoordinatorUI._collector = None
class TestBroadcastApproveRequest:
"""Coord-side override for the approve_request push. Same rationale
as the WebUI override — the coord-self path is rare today, but
parity keeps the override symmetric with the rest of the broadcast
family."""
def test_calls_collector_emit_with_ws_id_and_detail(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
ConsoleCoordinatorUI._collector = collector
try:
detail = {
"type": "approve_request",
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": True,
}
ui._broadcast_approve_request(detail)
collector.emit_console_ws_approve_request.assert_called_once_with(
"coord-a",
detail,
)
finally:
ConsoleCoordinatorUI._collector = None
def test_no_op_when_collector_unset(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
ConsoleCoordinatorUI._collector = None
# Doesn't raise.
ui._broadcast_approve_request({"items": []})
def test_collector_exception_swallowed(self) -> None:
ui = ConsoleCoordinatorUI(ws_id="coord-a", user_id="u1")
collector = MagicMock()
collector.emit_console_ws_approve_request.side_effect = RuntimeError("boom")
ConsoleCoordinatorUI._collector = collector
try:
# Doesn't raise.
ui._broadcast_approve_request({"items": []})
finally:
ConsoleCoordinatorUI._collector = None
+181 -74
View File
@@ -378,13 +378,21 @@ class TestCoordinatorAdapterWorkerDispatch:
class TestCoordinatorAdapterChildrenRegistry:
def test_emit_created_seeds_empty_children_set(self) -> None:
"""Adapter-level integration with :class:`ChildrenRegistry`.
Pure-registry invariants (forward/reverse consistency, idempotent
merge, locking) live in ``test_children_registry.py``. These
tests cover the adapter's wiring: that ``emit_*`` paths drive the
registry correctly and that the snapshot-priming bridge between
a collector snapshot and the registry preserves merge semantics.
"""
def test_emit_created_installs_parent(self) -> None:
adapter, _ = _make_adapter()
ws = _make_ws()
adapter.emit_created(ws)
assert ws.id in adapter._children
assert adapter._children[ws.id] == set()
assert adapter._active_coords[ws.id] is ws.ui
assert adapter._registry.children_of(ws.id) == []
assert adapter._registry.ui_for(ws.id) is ws.ui
def test_emit_rehydrated_calls_rebuild(self) -> None:
adapter, _ = _make_adapter()
@@ -398,42 +406,36 @@ class TestCoordinatorAdapterChildrenRegistry:
adapter.emit_rehydrated(ws)
assert calls == [ws.id]
def test_emit_closed_clears_forward_and_reverse_indexes(self) -> None:
def test_emit_closed_uninstalls_parent_and_clears_children(self) -> None:
adapter, _ = _make_adapter()
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-a1", "child-a2"])
adapter._merge_child_ids_locked("coord-b", ["child-b1"])
adapter._active_coords["coord-a"] = object()
adapter._active_coords["coord-b"] = object()
adapter._registry.install("coord-a", object())
adapter._registry.install("coord-b", object())
adapter._registry.merge_children("coord-a", ["child-a1", "child-a2"])
adapter._registry.merge_children("coord-b", ["child-b1"])
adapter.emit_closed("coord-a")
assert "coord-a" not in adapter._children
assert "coord-a" not in adapter._active_coords
assert "child-a1" not in adapter._child_to_coord
assert "child-a2" not in adapter._child_to_coord
assert adapter._registry.ui_for("coord-a") is None
assert adapter._registry.children_of("coord-a") == []
assert adapter._registry.parent_for("child-a1") is None
assert adapter._registry.parent_for("child-a2") is None
# coord-b untouched
assert adapter._child_to_coord["child-b1"] == "coord-b"
assert "coord-b" in adapter._children
def test_merge_child_ids_locked_is_idempotent(self) -> None:
adapter, _ = _make_adapter()
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-1"])
adapter._merge_child_ids_locked("coord-a", ["child-1"])
assert adapter._children["coord-a"] == {"child-1"}
assert adapter._child_to_coord == {"child-1": "coord-a"}
assert adapter._registry.parent_for("child-b1") == "coord-b"
assert adapter._registry.ui_for("coord-b") is not None
def test_prime_children_from_snapshot_merges_without_overwriting(self) -> None:
# Snapshot priming now lives on ClusterChildSource (production
# path). The adapter no longer carries its own duplicate copy.
from turnstone.core.child_source import ClusterChildSource
adapter, _ = _make_adapter()
# Seed one in-memory coord + one existing child
coord_ws = _make_ws()
coord_ws.id = "coord-a"
mgr = MagicMock()
mgr.list_all.return_value = [coord_ws]
adapter.attach(mgr)
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-a1"])
adapter._registry.merge_children("coord-a", ["child-a1"])
source = ClusterChildSource(
collector=MagicMock(),
registry=adapter._registry,
parents_provider=lambda: ["coord-a"],
)
snapshot = {
"nodes": [
@@ -448,10 +450,13 @@ class TestCoordinatorAdapterChildrenRegistry:
},
],
}
adapter._prime_children_from_snapshot(snapshot)
assert adapter._children["coord-a"] == {"child-a1", "child-a2"}
assert adapter._child_to_coord["child-a2"] == "coord-a"
assert "child-x" not in adapter._child_to_coord
source._prime_from_snapshot(snapshot)
assert set(adapter._registry.children_of("coord-a")) == {
"child-a1",
"child-a2",
}
assert adapter._registry.parent_for("child-a2") == "coord-a"
assert adapter._registry.parent_for("child-x") is None
# ---------------------------------------------------------------------------
@@ -478,9 +483,7 @@ class TestCoordinatorAdapterDispatchChildEvent:
coord_ws.id = coord_id
recorder = _UIRecorder()
coord_ws.ui = recorder # type: ignore[assignment]
with adapter._children_lock:
adapter._children.setdefault(coord_id, set())
adapter._active_coords[coord_id] = recorder
adapter._registry.install(coord_id, recorder)
adapter.attach(_StubManager(coord_ws)) # type: ignore[arg-type]
return adapter, recorder, coord_ws
@@ -510,12 +513,11 @@ class TestCoordinatorAdapterDispatchChildEvent:
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
# Reverse index updated for subsequent cluster_state events.
assert adapter._child_to_coord["child-a1"] == "coord-a"
assert adapter._registry.parent_for("child-a1") == "coord-a"
def test_dispatch_cluster_state_routes_via_reverse_index(self) -> None:
adapter, recorder, _ = self._setup()
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-a1"])
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "cluster_state",
@@ -533,8 +535,7 @@ class TestCoordinatorAdapterDispatchChildEvent:
def test_dispatch_ws_closed_routes_to_parent_coord(self) -> None:
adapter, recorder, _ = self._setup()
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-a1"])
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{"type": "ws_closed", "ws_id": "child-a1", "reason": "evicted"}
)
@@ -548,8 +549,7 @@ class TestCoordinatorAdapterDispatchChildEvent:
"""perf-6: _enqueue_on_ui mutates the payload dict in place with
the coord's ws_id so the browser can discriminate child events."""
adapter, recorder, _ = self._setup()
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-a1"])
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "cluster_state",
@@ -559,53 +559,160 @@ class TestCoordinatorAdapterDispatchChildEvent:
)
assert recorder.enqueued[0]["ws_id"] == "coord-a"
def test_dispatch_cluster_state_forwards_pending_approval_detail(self) -> None:
"""The rich approval payload now rides on child_ws_state directly so
the browser can mutate liveBadgeCache without a separate live-bulk
fetch. Drift here means the inline approve/deny buttons would
regress to chasing the dashboard cache (the load-storm pattern
Shape A is unwinding)."""
def test_dispatch_cluster_state_does_not_carry_pending_approval_detail(
self,
) -> None:
"""Stage 3 cleanup — the ``pending_approval_detail`` piggyback
on ``cluster_state`` is gone. Approval items now arrive via
bulk fetch (triggered by ``activity_state="approval"`` in the
browser); verdicts via ``child_ws_intent_verdict``; resolution
via ``child_ws_approval_resolved``. The state event carries
only state + activity_state — no detail field."""
adapter, recorder, _ = self._setup()
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-a1"])
detail = {
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": False,
}
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "cluster_state",
"ws_id": "child-a1",
"state": "running",
"activity_state": "approval",
"pending_approval_detail": detail,
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_state"
assert payload["activity_state"] == "approval"
assert payload["pending_approval_detail"] == detail
assert "pending_approval_detail" not in payload
def test_dispatch_cluster_state_pending_approval_detail_none_passes_through(
self,
) -> None:
"""Missing pending_approval_detail (no approval pending, or pre-fix
node mid-rolling-upgrade) must forward as None — not raise, not
omit — so the browser's handleChildState treats it as "no SSE-
supplied detail, fall back to cached value"."""
def test_dispatch_intent_verdict_emits_child_ws_intent_verdict(self) -> None:
"""Stage 3 Step 6 — explicit verdict events are re-emitted as
child_ws_intent_verdict on the parent's SSE so the tree UI
renders the risk pill without polling."""
adapter, recorder, _ = self._setup()
with adapter._children_lock:
adapter._merge_child_ids_locked("coord-a", ["child-a1"])
adapter._registry.merge_children("coord-a", ["child-a1"])
verdict = {
"call_id": "c1",
"risk_level": "low",
"confidence": 0.92,
"recommendation": "approve",
}
adapter._dispatch_child_event(
{
"type": "cluster_state",
"type": "intent_verdict",
"ws_id": "child-a1",
"state": "running",
"activity_state": "tool",
"node_id": "node-1",
"verdict": verdict,
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert "pending_approval_detail" in payload
assert payload["pending_approval_detail"] is None
assert payload["type"] == "child_ws_intent_verdict"
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
assert payload["node_id"] == "node-1"
assert payload["verdict"] == verdict
def test_dispatch_intent_verdict_unknown_child_drops(self) -> None:
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{
"type": "intent_verdict",
"ws_id": "ws-orphan",
"verdict": {"call_id": "c1"},
}
)
assert recorder.enqueued == []
def test_dispatch_approval_resolved_emits_child_ws_approval_resolved(
self,
) -> None:
"""Stage 3 Step 6 — paired with intent_verdict; clears the
pending-approval pill on the parent's tree UI in lockstep
with the actual decision."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event(
{
"type": "approval_resolved",
"ws_id": "child-a1",
"node_id": "node-1",
"approved": True,
"feedback": "lgtm",
"always": False,
}
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_approval_resolved"
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
assert payload["approved"] is True
assert payload["feedback"] == "lgtm"
assert payload["always"] is False
def test_dispatch_approval_resolved_coerces_missing_fields(self) -> None:
"""Older nodes mid-rolling-upgrade may omit approved / always /
feedback; dispatch coerces to safe defaults."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
adapter._dispatch_child_event({"type": "approval_resolved", "ws_id": "child-a1"})
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["approved"] is False
assert payload["feedback"] == ""
assert payload["always"] is False
def test_dispatch_approval_resolved_unknown_child_drops(self) -> None:
"""Symmetric to the intent_verdict drop test — events for
ws_ids the registry doesn't know about silently drop instead
of fanning out to a parent that has no business seeing them."""
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{
"type": "approval_resolved",
"ws_id": "ws-orphan",
"approved": True,
},
)
assert recorder.enqueued == []
def test_dispatch_approve_request_emits_child_ws_approve_request(
self,
) -> None:
"""Push path for the initial approval items — eliminates the
bulk-fetch race that left the coord row stuck on a loading
placeholder when the bulk fetch landed in the gap between
_emit_state(ATTENTION) and approve_tools setting _pending_approval."""
adapter, recorder, _ = self._setup()
adapter._registry.merge_children("coord-a", ["child-a1"])
detail = {
"type": "approve_request",
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": True,
}
adapter._dispatch_child_event(
{
"type": "approve_request",
"ws_id": "child-a1",
"node_id": "node-1",
"detail": detail,
},
)
assert len(recorder.enqueued) == 1
payload = recorder.enqueued[0]
assert payload["type"] == "child_ws_approve_request"
assert payload["child_ws_id"] == "child-a1"
assert payload["parent_ws_id"] == "coord-a"
assert payload["node_id"] == "node-1"
assert payload["detail"] == detail
def test_dispatch_approve_request_unknown_child_drops(self) -> None:
adapter, recorder, _ = self._setup()
adapter._dispatch_child_event(
{
"type": "approve_request",
"ws_id": "ws-orphan",
"detail": {"items": []},
},
)
assert recorder.enqueued == []
+131
View File
@@ -949,6 +949,137 @@ def test_list_nodes_empty_on_no_matching_filters(storage_with_nodes):
assert result["truncated"] is False
def test_list_nodes_surfaces_healthy_model_aliases(tmp_path):
"""The node's heartbeat loop projects its registry into a ``models``
metadata entry shaped like ``[{alias, provider, healthy}, ...]``.
``list_nodes`` flattens that to the healthy-alias list at the top
level (under ``model_aliases``) so a coordinator can pass aliases
straight to ``spawn_workstream(model=)`` without having to
introspect the metadata blob. The provider-side model identifier
(``cfg.model``) is intentionally NOT in the payload — coords kept
reaching for it when they should pass the local alias."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-x",
[
("arch", "x86_64", "auto"),
(
"models",
[
{"alias": "gpt5", "provider": "openai", "healthy": True},
{"alias": "claude-opus-47", "provider": "anthropic", "healthy": True},
{"alias": "broken", "provider": "openai", "healthy": False},
],
"auto",
),
],
)
_register_service(st, "node-x")
client = _make_read_client(st)
result = client.list_nodes()
node = result["nodes"][0]
assert node["model_aliases"] == ["gpt5", "claude-opus-47"]
# Full per-alias info still available under metadata for callers
# that want provider / healthy detail (e.g. surfacing degraded
# aliases in a UI).
full = node["metadata"]["models"]["value"]
assert {row["alias"] for row in full} == {"gpt5", "claude-opus-47", "broken"}
# ``model`` (the provider-side identifier) is intentionally absent
# — keep the payload to the three values a coord actually uses.
for row in full:
assert "model" not in row
def test_list_nodes_model_aliases_distinct_from_metadata_models(tmp_path):
"""Pin the naming distinction explicitly: the top-level shortlist
(``model_aliases``, list of strings) and the rich metadata blob
(``metadata.models.value``, list of dicts) live under different
keys so a caller that confuses them gets a clear KeyError rather
than a silent shape mismatch."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-x",
[
(
"models",
[{"alias": "a", "provider": "openai", "healthy": True}],
"auto",
),
],
)
_register_service(st, "node-x")
client = _make_read_client(st)
node = client.list_nodes()["nodes"][0]
# No top-level ``models`` field — only ``model_aliases``.
assert "models" not in node
assert node["model_aliases"] == ["a"]
# Rich shape stays under metadata.
assert isinstance(node["metadata"]["models"]["value"], list)
assert isinstance(node["metadata"]["models"]["value"][0], dict)
def test_list_nodes_model_aliases_empty_when_node_has_not_published(tmp_path):
"""Nodes from older builds — or a node mid-startup before its first
metadata write — won't have a ``models`` entry. The top-level
``model_aliases`` field defaults to ``[]`` rather than being
omitted so coordinators can rely on the key being present."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(st, "node-y", [("arch", "x86_64", "auto")])
_register_service(st, "node-y")
client = _make_read_client(st)
result = client.list_nodes()
assert result["nodes"][0]["model_aliases"] == []
def test_list_nodes_models_tolerates_malformed_entries(tmp_path):
"""If a node ever stores a malformed ``models`` entry (wrong outer
type, missing alias, non-bool healthy), the projection drops the
bad rows rather than raising — the rest of the response should
still be useful."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-z",
[
(
"models",
[
{"alias": "ok", "provider": "p", "healthy": True},
"not-a-dict",
{"provider": "p", "healthy": True}, # missing alias
{"alias": "", "healthy": True}, # empty alias
{"alias": "degraded", "healthy": False},
{"alias": 42, "healthy": True}, # non-string alias
],
"auto",
),
],
)
_register_service(st, "node-z")
client = _make_read_client(st)
result = client.list_nodes()
assert result["nodes"][0]["model_aliases"] == ["ok"]
def test_list_nodes_models_handles_non_list_payload(tmp_path):
"""A node with a corrupted models entry (dict, scalar, null) shouldn't
blow up the whole list_nodes call. ``model_aliases`` falls back to ``[]``."""
st = SQLiteBackend(str(tmp_path / "nodes.db"))
_set_meta(
st,
"node-w",
[
("models", {"oops": "not a list"}, "auto"),
],
)
_register_service(st, "node-w")
client = _make_read_client(st)
result = client.list_nodes()
assert result["nodes"][0]["model_aliases"] == []
# ---------------------------------------------------------------------------
# list_skills
# ---------------------------------------------------------------------------
+70 -34
View File
@@ -79,9 +79,10 @@ def test_coordinator_js_exposes_inline_approval_helpers():
assert "function submitChildApproval" in body or "submitChildApproval(" in body
# The shared approve POST helper (parameterized for child ws_ids)
assert "function approveWorkstream" in body or "approveWorkstream(" in body
# The urgent live-bulk fetch option that fires on activity_state
# transitions in/out of "approval"
assert "{ urgent: true }" in body or "urgent: true" in body
# The 409 stale-call_id retry path uses invalidateLiveBadge +
# scheduleLiveFetch (Stage 3 cleanup removed the urgent flag —
# cache invalidation makes the TTL gate fall through naturally).
assert "invalidateLiveBadge(targetWsId)" in body
# Server-side payload field — drift here means the JS reads stale keys
assert "pending_approval_detail" in body
# Reconnect parity (chunk 4): the SSE re-open handler must drop
@@ -90,10 +91,10 @@ def test_coordinator_js_exposes_inline_approval_helpers():
# can't render zombie approve/deny buttons on a row whose
# approval was resolved during the gap. The implementation
# iterates the cache and deletes only !permanent entries —
# asserting the literal Map iteration form keeps a refactor
# back to liveBadgeCache.clear() (which would re-pay 403s on
# every reconnect for denied ids) from sneaking in.
assert "liveBadgeCache.delete" in body
# asserting the literal helper call keeps a refactor back to
# _liveBadgeCacheClear() (which would re-pay 403s on every
# reconnect for denied ids) from sneaking in.
assert "_liveBadgeCacheDelete" in body
# Edge-case matrix sentinel labels — POLICY-BLOCKED renders when
# an item has error set + needs_approval=False (server-side
# tool policy already blocked the call); "(judge unavailable)"
@@ -113,15 +114,19 @@ def test_coordinator_js_exposes_inline_approval_helpers():
# coord-self ws_id (the coord lives on the console process).
# Children live on cluster nodes and 404 without the prefix.
assert "/v1/api/route/workstreams/" in body
# Late-judge polling — the LLM judge runs async on the child
# node and never pushes a signal that reaches the coord, so
# the row's pending_approval_detail with judge_pending=true
# would freeze on heuristic verdicts forever without this
# poll loop. The poller is GLOBAL (not per-row) so off-screen
# rows still refresh — a per-row poller's scheduleLiveFetch
# call short-circuits on non-visible rows, leaving them stuck.
assert "_maybeStartJudgePoll" in body
assert "_judgePollTick" in body
# Late-arriving LLM judge verdicts — Stage 3 Step 5 promoted
# ``intent_verdict`` and ``approval_resolved`` to first-class
# cluster-bus event types, so the coord adapter dispatches them
# as ``child_ws_intent_verdict`` / ``child_ws_approval_resolved``
# on the parent's SSE stream. The browser handlers write
# directly to liveBadgeCache (bypassing scheduleLiveFetch's
# visibility gate cleanly) so off-screen rows pick up verdicts
# without polling. Replaced the old ``_judgePollTick`` 90-second
# global poll loop and its visibility-gate-bypass workaround.
assert "handleChildIntentVerdict" in body
assert "handleChildApprovalResolved" in body
assert "child_ws_intent_verdict" in body
assert "child_ws_approval_resolved" in body
# Reload parity for the coord-self approval gate: init() must
# consume the authoritative GET /workstreams snapshot's
# pending_approval_detail so a freshly opened tab can render
@@ -152,20 +157,31 @@ def test_coordinator_js_exposes_inline_approval_helpers():
# any prior denial. bug-1 / bug-3 from the second /review pass.
assert "Denied by user" in body
assert "callOutcomes" in body
# User-message attachment pills — both live send (coordSend) and
# history replay route through appendUserMessageWithAttachments.
# Renaming or dropping the helper would silently regress the
# attachment affordance to the pre-fix plain-text bubble, which
# would only surface in manual testing of an attached-file flow.
# The CSS class is the visual anchor (coordinator.css) — keeping
# both literals in the smoke layer covers JS↔CSS drift in either
# direction.
assert "function appendUserMessageWithAttachments" in body
assert "msg-user-attach" in body
def test_coordinator_js_handle_child_state_reads_sse_pending_approval_detail():
"""Lock the Shape A behavior change: child_ws_state SSE events now
carry ``pending_approval_detail`` directly so the browser mutates
``liveBadgeCache`` without firing an urgent live-bulk fetch on
every activity_state transition into/out of approval. A refactor
that re-introduces the urgent-fetch path on routine transitions
(or drops the SSE-source merge guard in flushLiveFetches) would
re-open the load-storm pattern this PR is fixing.
def test_coordinator_js_handle_child_state_no_longer_reads_sse_pending_approval_detail():
"""Stage 3 cleanup — ``pending_approval_detail`` is no longer
piggybacked on child_ws_state events. Approval items now arrive
via bulk fetch on the activity_state="approval" transition;
verdicts via the explicit ``child_ws_intent_verdict`` event class;
resolution via ``child_ws_approval_resolved``. A refactor that
re-introduces the piggyback would silently re-open the
duplicate-path race the dedicated event classes were added to
eliminate.
Structural assertions (regex against multi-line source) — symbol-
presence alone wouldn't catch a guard that keeps the names but
inverts the comparison or drops the ``prev.live`` check. This
inverts the comparison or drops the ``prev.live`` check. This
codebase has no JS test framework, so locking the guard's shape
here is the next-best thing to a behavioral test."""
import re
@@ -176,26 +192,46 @@ def test_coordinator_js_handle_child_state_reads_sse_pending_approval_detail():
)
body = coord_js.read_text(encoding="utf-8")
# handleChildState now reads the SSE-supplied detail.
assert "ev.pending_approval_detail" in body
# The pre-fix urgent-fetch on activity_state transitions is
# gone (the 409 retry path keeps its own ``{ urgent: true }``
# for stale-call_id refresh — that's a different scenario).
# The piggyback read is gone from handleChildState. (The string
# may still appear elsewhere — e.g. handleChildIntentVerdict
# reading from cache, or comments — but never as ``ev.pending_approval_detail``.)
assert "ev.pending_approval_detail" not in body
# The pre-fix urgent-fetch on activity_state transitions is gone.
assert "enteredApproval" not in body
assert "leftApproval" not in body
# ``pendingApproval`` flag derivation must check BOTH state and
# activity_state. The worker thread can fire the state transition
# to "attention" before approve_tools updates activity_state, so
# checking only activity_state misses children that legitimately
# need approval. Pin the disjunction so the regression doesn't
# silently re-introduce.
assert re.search(
r'existing\.state\s*===\s*"attention"\s*\|\|\s*'
r'existing\.activity_state\s*===\s*"approval"',
body,
), (
"handleChildState must derive pendingApproval from "
"(state==='attention' || activity_state==='approval')"
)
# SSE-authoritative window constant is defined and used.
assert re.search(r"\bconst\s+SSE_AUTHORITATIVE_MS\s*=\s*\d+", body), (
"SSE_AUTHORITATIVE_MS constant must be defined as a numeric literal"
)
# handleChildState writes sseUpdatedAt = Date.now() into the cache
# entry it sets. This is the SSE-source tag; without it, the
# merge guard in flushLiveFetches has nothing to gate on.
# SSE writers tag entries with sseUpdatedAt: Date.now() so the
# merge guard in flushLiveFetches preserves them against stale
# bulk-fetch responses. handleChildState only stamps when it
# AUTHORITATIVELY clears the detail (off-approval transition);
# writers that stamp unconditionally are intent_verdict (verdict
# stamp), approval_resolved (clear), and the optimistic-clear
# path in submitChildApproval. Pinning the literal Date.now()
# call keeps a refactor that drops the SSE-source tag entirely
# from sneaking in.
assert re.search(
r"sseUpdatedAt:\s*Date\.now\(\)",
body,
), "handleChildState must write sseUpdatedAt: Date.now() onto liveBadgeCache entries"
), "Critical SSE writers must stamp sseUpdatedAt: Date.now()"
# flushLiveFetches' merge guard structure: SSE-set pending_approval
# / _detail wins over a stale bulk-poll snapshot when (live) AND
+232
View File
@@ -0,0 +1,232 @@
"""Unit tests for ``turnstone.core.history_decoration``.
The decoration helpers are shared between two surfaces — interactive's
SSE replay (``_build_history``) and the lifted ``/history`` REST
endpoint (``make_history_handler``, used by both interactive and
coord). Pinning the wire shape here lets a future schema/projection
change land in one file rather than spread across the two surfaces.
"""
from __future__ import annotations
from turnstone.core.history_decoration import (
build_output_assessment_payload,
build_verdict_payload,
decorate_history_messages,
decorate_tool_call,
)
class TestBuildVerdictPayload:
"""The wire-shape projection that's the single source of truth for
what intent_verdict fields ship to the client."""
def test_skips_unflagged_baseline(self) -> None:
"""``risk_level`` "none" is the unflagged-tool baseline; the
client filters those anyway, so projecting None at the wire
layer keeps the payload tight on long workstreams."""
row = {"risk_level": "none", "recommendation": "approve", "tier": "heuristic"}
assert build_verdict_payload(row) is None
def test_drops_call_id_and_func_name(self) -> None:
"""The client already has these on ``tc.id`` / ``tc.name``;
re-shipping them per-tool_call would balloon long replays."""
row = {
"call_id": "call_abc",
"func_name": "bash",
"risk_level": "medium",
"recommendation": "review",
"confidence": 0.8,
"intent_summary": "summary",
"tier": "heuristic",
}
out = build_verdict_payload(row)
assert out is not None
assert "call_id" not in out
assert "func_name" not in out
# Sanity — the kept fields are the ones renderVerdictBadge reads.
assert out["risk_level"] == "medium"
assert out["recommendation"] == "review"
assert out["confidence"] == 0.8
assert out["intent_summary"] == "summary"
assert out["tier"] == "heuristic"
def test_includes_reasoning_for_either_tier_when_present(self) -> None:
"""Heuristic verdicts in this project emit structured
rationales (one per matched pattern) — e.g.
``policy.py`` writes a reasoning string per heuristic hit.
Ship the field for either tier when it has content; only
omit when the row didn't write one."""
for tier in ("heuristic", "llm"):
row = {
"risk_level": "high",
"tier": tier,
"reasoning": "The command exfiltrates ~/.ssh/id_rsa over an external connection.",
}
out = build_verdict_payload(row)
assert out is not None
assert "id_rsa" in out["reasoning"]
def test_omits_reasoning_when_empty(self) -> None:
"""An absent / empty reasoning string shouldn't ship as
``reasoning: ""`` — the rationale ``<details>`` block on the
client renders an empty disclosure when the field is present
but empty."""
row = {"risk_level": "high", "tier": "heuristic", "reasoning": ""}
out = build_verdict_payload(row)
assert out is not None
assert "reasoning" not in out
def test_includes_judge_model_when_present(self) -> None:
"""``judge_model`` rides through so the batch tier badge can
render ``⚖ llm:claude-haiku-4`` on history-only replays
rather than the bare ``⚖ llm`` label."""
row = {"risk_level": "high", "tier": "llm", "judge_model": "claude-haiku-4"}
out = build_verdict_payload(row)
assert out is not None
assert out["judge_model"] == "claude-haiku-4"
def test_omits_judge_model_when_empty(self) -> None:
row = {"risk_level": "medium", "tier": "heuristic", "judge_model": ""}
out = build_verdict_payload(row)
assert out is not None
assert "judge_model" not in out
class TestBuildOutputAssessmentPayload:
"""Output-guard wire shape — flags decoded from JSON string at
this layer so the client never has to parse twice."""
def test_skips_unflagged_baseline(self) -> None:
row = {"risk_level": "none", "flags": "[]"}
assert build_output_assessment_payload(row) is None
def test_decodes_flags_from_json(self) -> None:
row = {"risk_level": "high", "flags": '["api_key","email"]', "redacted": 1}
out = build_output_assessment_payload(row)
assert out is not None
assert out["flags"] == ["api_key", "email"]
assert out["redacted"] is True
assert out["risk_level"] == "high"
def test_handles_malformed_flags_json(self) -> None:
"""Bad JSON in ``flags`` must not block the rest of the
assessment from rendering — degrade to empty list."""
row = {"risk_level": "medium", "flags": "not-json", "redacted": 0}
out = build_output_assessment_payload(row)
assert out is not None
assert out["flags"] == []
assert out["redacted"] is False
class TestDecorateToolCall:
"""In-place mutation of either OpenAI-format or flattened tool_call
entries — both shapes carry ``id`` at the top level."""
def test_attaches_verdict_when_present(self) -> None:
tc: dict[str, object] = {"id": "call_1", "function": {"name": "bash", "arguments": "{}"}}
verdicts = {
"call_1": {
"risk_level": "medium",
"recommendation": "review",
"confidence": 0.7,
"intent_summary": "summary",
"tier": "heuristic",
}
}
decorate_tool_call(tc, verdicts, {})
assert "verdict" in tc
assert tc["verdict"]["risk_level"] == "medium" # type: ignore[index]
def test_skips_when_no_call_id_match(self) -> None:
tc: dict[str, object] = {"id": "call_other", "name": "bash"}
verdicts = {
"call_1": {"risk_level": "medium", "tier": "heuristic"},
}
decorate_tool_call(tc, verdicts, {})
assert "verdict" not in tc
def test_skips_unflagged_verdict(self) -> None:
"""``build_verdict_payload`` returns None for unflagged rows;
decorate_tool_call must not stamp ``verdict`` in that case."""
tc: dict[str, object] = {"id": "call_1", "name": "bash"}
verdicts = {"call_1": {"risk_level": "none", "tier": "heuristic"}}
decorate_tool_call(tc, verdicts, {})
assert "verdict" not in tc
def test_handles_empty_id(self) -> None:
"""A tool_call with no id can't be paired against the lookup
table — must not raise (or stamp the wrong row's verdict)."""
tc: dict[str, object] = {"id": "", "name": "bash"}
verdicts = {"call_1": {"risk_level": "high", "tier": "heuristic"}}
decorate_tool_call(tc, verdicts, {})
assert "verdict" not in tc
class TestDecorateHistoryMessages:
"""End-to-end mutation of a /history-shaped message list — covers
the full transform applied by ``make_history_handler``."""
def test_decorates_tool_calls_and_marks_truncated(self) -> None:
verdicts = {
"call_a": {
"risk_level": "high",
"recommendation": "deny",
"confidence": 0.95,
"intent_summary": "exfil",
"tier": "llm",
"reasoning": "ssh key access",
}
}
assessments = {
"call_a": {"risk_level": "high", "flags": '["secret"]', "redacted": 1},
}
# Tool result content of exactly TOOL_RESULT_STORAGE_CAP chars
# hits the storage cap (longer is impossible — storage clamps
# at the cap). Reference the constant rather than a literal so
# this test stays correct if the cap moves again.
from turnstone.core.history_decoration import TOOL_RESULT_STORAGE_CAP
truncated_content = "x" * TOOL_RESULT_STORAGE_CAP
messages: list[dict[str, object]] = [
{"role": "user", "content": "hi"},
{
"role": "assistant",
"content": "running",
"tool_calls": [
{
"id": "call_a",
"function": {"name": "bash", "arguments": "{}"},
}
],
},
{"role": "tool", "tool_call_id": "call_a", "content": truncated_content},
{"role": "tool", "tool_call_id": "call_b", "content": "short"},
]
decorate_history_messages(messages, verdicts, assessments)
# Assistant tool_calls got both decorations.
tc = messages[1]["tool_calls"][0] # type: ignore[index]
assert tc["verdict"]["risk_level"] == "high"
assert tc["verdict"]["tier"] == "llm"
assert "reasoning" in tc["verdict"]
assert tc["output_assessment"]["flags"] == ["secret"]
assert tc["output_assessment"]["redacted"] is True
# Truncated tool message got the flag; the short one did not.
assert messages[2].get("truncated") is True
assert "truncated" not in messages[3]
def test_no_op_on_empty_indexes(self) -> None:
"""When neither table has rows for the workstream, the wire
shape passes through unchanged — replay must degrade
gracefully when verdict storage is empty / unavailable."""
messages: list[dict[str, object]] = [
{
"role": "assistant",
"content": "",
"tool_calls": [{"id": "call_a", "function": {"name": "bash", "arguments": "{}"}}],
},
]
decorate_history_messages(messages, {}, {})
tc = messages[0]["tool_calls"][0] # type: ignore[index]
assert "verdict" not in tc
assert "output_assessment" not in tc
+265
View File
@@ -1,6 +1,9 @@
"""Tests for turnstone.core.memory_relevance — scoring, formatting, context extraction."""
from unittest.mock import patch
from turnstone.core.memory_relevance import (
MemoryConfig,
build_memory_context,
extract_recent_context,
score_memories,
@@ -192,3 +195,265 @@ class TestExtractRecentContext:
def test_empty_messages(self):
assert extract_recent_context([]) == ""
# ---------------------------------------------------------------------------
# Composition candidate-selection (_init_system_messages)
# ---------------------------------------------------------------------------
def _make_mem(name: str, content: str = "", memory_id: str | None = None) -> dict[str, str]:
return {
"name": name,
"memory_id": memory_id or f"mid_{name}",
"type": "project",
"scope": "global",
"scope_id": "",
"description": "",
"content": content or name,
"updated": "2024-01-01T00:00:00",
}
def _make_session(fetch_limit: int = 5, relevance_k: int = 3, **kwargs: object):
"""Composition tests need a real ChatSession (constructor calls
``_init_system_messages`` once, unpatched, before the test gets a chance
to install patches). ``tmp_db`` initializes the storage singleton that
constructor needs; tests then patch the visibility helpers and call
``_init_system_messages`` a second time to exercise the new logic.
"""
from tests._helpers import make_chat_session
return make_chat_session(
memory_config=MemoryConfig(fetch_limit=fetch_limit, relevance_k=relevance_k),
**kwargs,
)
class TestCompositionCandidateSelection:
"""Verify the query-aware candidate set in _init_system_messages."""
def test_recency_ceiling_regression(self, tmp_db):
"""Old relevant memory not in recency top-N still injected via search path."""
session = _make_session(fetch_limit=5, relevance_k=3)
session.messages = [{"role": "user", "content": "postgres database configuration"}]
old_mem = _make_mem(
"ancient_db_config",
content="postgres database configuration connection host port",
memory_id="m_old",
)
# Recency top-5 do not include old_mem
recent = [_make_mem(f"recent_{i}", memory_id=f"mr{i}") for i in range(5)]
with (
patch.object(session, "_search_visible_memories", return_value=[old_mem]),
patch.object(session, "_list_visible_memories", return_value=recent),
):
session._init_system_messages()
joined = "\n".join(m["content"] for m in session.system_messages if m["role"] == "system")
# With the fix, old_mem enters the candidate pool via search and wins BM25
assert "ancient_db_config" in joined
def test_empty_query_falls_back_to_recency(self, tmp_db):
"""No user messages → empty context → recency path, search never called."""
session = _make_session()
session.messages = [] # extract_recent_context returns ""
recency = [_make_mem("note_alpha"), _make_mem("note_beta")]
with (
patch.object(session, "_list_visible_memories", return_value=recency),
patch.object(session, "_search_visible_memories") as search_mock,
):
session._init_system_messages()
search_mock.assert_not_called()
joined = "\n".join(m["content"] for m in session.system_messages if m["role"] == "system")
assert "note_alpha" in joined
def test_sparse_match_union_fills_candidate_pool(self, tmp_db):
"""Search returning < fetch_limit results unions with recency fillers."""
session = _make_session(fetch_limit=5, relevance_k=4)
session.messages = [{"role": "user", "content": "unique_term xyzzy"}]
hit_a = _make_mem("hit_alpha", content="unique_term xyzzy alpha", memory_id="m_ha")
hit_b = _make_mem("hit_beta", content="unique_term xyzzy beta", memory_id="m_hb")
search_hits = [hit_a, hit_b] # 2 < fetch_limit=5 → triggers union
# Recency overlaps on hit_a/hit_b and adds 3 fillers
filler = [_make_mem(f"filler_{i}", memory_id=f"mf{i}") for i in range(3)]
recency = [hit_a, hit_b] + filler
with (
patch.object(session, "_search_visible_memories", return_value=search_hits),
patch.object(session, "_list_visible_memories", return_value=recency),
):
session._init_system_messages()
joined = "\n".join(m["content"] for m in session.system_messages if m["role"] == "system")
# Both hits match "unique_term xyzzy" well → appear after BM25 ranking
assert "hit_alpha" in joined
assert "hit_beta" in joined
def test_recency_preserved_when_search_returns_noise_above_relevance_k(self, tmp_db):
"""Pool guarantee: recency-50 always reaches BM25, even when search
returns enough noise hits to clear ``relevance_k``.
Closes the narrow regression vs. the original bug — without the
``fetch_limit`` threshold, a stopword-dominated cap-search that
returned >= relevance_k irrelevant hits would short-circuit and
evict the recency-only memory the bug had been surfacing.
"""
session = _make_session(fetch_limit=10, relevance_k=3)
session.messages = [{"role": "user", "content": "configure host"}]
# Search returns relevance_k=3 noise hits — enough to skip recency
# under the OLD threshold, not enough to fill fetch_limit=10.
noise = [
_make_mem(f"noise_{i}", content="generic content", memory_id=f"mn{i}") for i in range(3)
]
# The memory the user actually wants — distinctive, in recency,
# but its content doesn't share any token with the noise hits.
wanted = _make_mem(
"host_config_v2",
content="host=localhost port=5432 db=production",
memory_id="m_wanted",
)
recency = [wanted] + [_make_mem(f"recent_{i}", memory_id=f"mr{i}") for i in range(5)]
with (
patch.object(session, "_search_visible_memories", return_value=noise),
patch.object(session, "_list_visible_memories", return_value=recency),
):
session._init_system_messages()
joined = "\n".join(m["content"] for m in session.system_messages if m["role"] == "system")
# ``wanted`` reached BM25 via the union and matched "host" → injected.
assert "host_config_v2" in joined
def test_recency_tail_preserved_when_search_adds_distinct_hits(self, tmp_db):
"""SUPERSET invariant: every recency item is in the candidate pool
when search adds hits, even if the resulting union exceeds
fetch_limit. Truncating the union at fetch_limit (the prior
behavior) evicted the recency tail — which is exactly where
ancient-but-recently-touched memories live, the recall this PR
sets out to improve.
"""
session = _make_session(fetch_limit=10, relevance_k=3)
session.messages = [{"role": "user", "content": "alpha"}]
# 5 search hits, none of which appear in recency.
search_hits = [
_make_mem(f"search_{i}", content="alpha", memory_id=f"ms{i}") for i in range(5)
]
# 10 recency items; without the union uncap, the 5 oldest of these
# would be displaced by the 5 search hits.
recency = [_make_mem(f"recency_{i}", memory_id=f"mr{i}") for i in range(10)]
with (
patch.object(session, "_search_visible_memories", return_value=search_hits),
patch.object(session, "_list_visible_memories", return_value=recency),
):
candidates, source = session._select_memory_candidates("alpha")
candidate_ids = {c["memory_id"] for c in candidates}
# Pool is search_hits recency — 15 items, no truncation.
assert len(candidates) == 15
assert source == "union"
# Every recency item present (no tail eviction).
for i in range(10):
assert f"mr{i}" in candidate_ids, f"recency item {i} evicted"
# And every search hit is also in the pool.
for i in range(5):
assert f"ms{i}" in candidate_ids, f"search hit {i} missing"
def test_coord_scope_isolated_visibility(self, tmp_db):
"""Coord composition queries the coord scope alone, never the
global/workstream/user union."""
from turnstone.core.workstream import WorkstreamKind
coord = _make_session(
fetch_limit=5,
relevance_k=3,
ws_id="coord-1",
user_id="user-1",
kind=WorkstreamKind.COORDINATOR,
)
scopes = coord._visible_scopes()
assert scopes == [("coordinator", "coord-1")]
# And: search uses those same scopes (no global/user fan-in)
coord.messages = [{"role": "user", "content": "anything"}]
with patch(
"turnstone.core.session.search_visible_structured_memories",
return_value=[],
) as search_mock:
coord._search_visible_memories("anything", limit=5)
search_mock.assert_called_once()
# Second positional arg is the scopes list
assert search_mock.call_args.args[1] == [("coordinator", "coord-1")]
class TestMemorySearchToolExecution:
"""End-to-end test of ``memory(action='search')`` through _exec_memory.
Drives the actual tool dispatch (not just the storage facade) so the
OR-of-terms fix and the coalesced ``memory.search`` log get exercised
together.
"""
def test_search_action_returns_or_of_terms_results(self, tmp_db):
"""Multi-word query returns rows where ANY term matches — not all."""
from turnstone.core.memory import save_structured_memory
save_structured_memory("postgres_notes", "host=localhost port=5432")
save_structured_memory("redis_notes", "host=redis port=6379")
save_structured_memory("unrelated", "completely different")
session = _make_session()
item = session._prepare_memory(
"call-1",
{"action": "search", "query": "postgres no_such_word_a no_such_word_b"},
)
# Sanity: prepare returned a search-ready dispatch (not an error item)
assert item.get("action") == "search"
call_id, msg = session._exec_memory(item)
assert call_id == "call-1"
assert "postgres_notes" in msg
# Other memories don't match any query term
assert "unrelated" not in msg
class TestPerTurnSearchCache:
"""The per-turn cache spares redundant SQL across mid-turn rebuilds."""
def test_repeated_search_in_same_turn_hits_cache(self, tmp_db):
from turnstone.core.memory import save_structured_memory
save_structured_memory("hello_mem", "alpha beta gamma")
session = _make_session()
with patch(
"turnstone.core.session.search_visible_structured_memories",
return_value=[],
) as backend_mock:
session._search_visible_memories("alpha beta", limit=5)
session._search_visible_memories("alpha beta", limit=5)
session._search_visible_memories("alpha beta", limit=5)
# 3 calls but only 1 backend hit — cache absorbed the rest
assert backend_mock.call_count == 1
def test_user_turn_invalidates_cache(self, tmp_db):
from turnstone.core.memory import save_structured_memory
save_structured_memory("hello_mem", "alpha")
session = _make_session()
with patch(
"turnstone.core.session.search_visible_structured_memories",
return_value=[],
) as backend_mock:
session._search_visible_memories("alpha", limit=5)
session._invalidate_memory_cache() # simulates new user turn
session._search_visible_memories("alpha", limit=5)
assert backend_mock.call_count == 2
+78 -6
View File
@@ -1126,24 +1126,57 @@ class TestSessionAgentModel:
def _captured_effort(captured: dict[str, Any]) -> str | None:
"""Pull reasoning_effort out of provider-specific shapes.
openai-compatible servers receive it via extra_body.chat_template_kwargs;
commercial providers receive it as a top-level kwarg.
Chat Completions delivers it as a top-level ``reasoning_effort`` kwarg
(when the model's caps permit it). Operators who route reasoning_effort
through ``chat_template_kwargs`` (gpt-oss-style local templates) get
it inside ``extra_body.chat_template_kwargs``.
"""
if "reasoning_effort" in captured:
return captured["reasoning_effort"]
eb = captured.get("extra_body") or {}
ctk = eb.get("chat_template_kwargs") or {}
return ctk.get("reasoning_effort") or captured.get("reasoning_effort")
return ctk.get("reasoning_effort")
@staticmethod
def _effort_caps() -> dict[str, Any]:
"""Capabilities that allow Chat-Completions reasoning_effort to flow."""
return {
"reasoning_effort_values": [
"minimal",
"low",
"medium",
"high",
"max",
],
}
def _three_model_registry(self, **kwargs: Any) -> ModelRegistry:
caps = self._effort_caps()
return ModelRegistry(
models={
"main": ModelConfig(
"main", "http://m/v1", "k", "main-model", provider="openai-compatible"
"main",
"http://m/v1",
"k",
"main-model",
provider="openai-compatible",
capabilities=dict(caps),
),
"smart": ModelConfig(
"smart", "http://s/v1", "k", "smart-model", provider="openai-compatible"
"smart",
"http://s/v1",
"k",
"smart-model",
provider="openai-compatible",
capabilities=dict(caps),
),
"fast": ModelConfig(
"fast", "http://f/v1", "k", "fast-model", provider="openai-compatible"
"fast",
"http://f/v1",
"k",
"fast-model",
provider="openai-compatible",
capabilities=dict(caps),
),
},
default="main",
@@ -1249,6 +1282,45 @@ class TestSessionAgentModel:
session._run_agent([{"role": "user", "content": "x"}], label="plan", agent_alias="fast")
assert captured["model"] == "fast-model"
def test_session_fallback_inherits_primary_alias_for_caps(self) -> None:
"""When _run_agent has no registry agent route, it must fall back to
the session's primary alias for capability and server_compat lookup —
otherwise per-model caps (reasoning_effort_values, server_compat) get
silently dropped on the agent path."""
reg = self._three_model_registry() # no agent_model / plan_model set
session = _make_session(registry=reg, model_alias="main")
# Probe what _run_agent passes to _provider_extra_params and
# _resolve_capabilities by recording the model_alias on each call.
captured_extra_alias: list[str | None] = []
captured_resolve_alias: list[str | None] = []
original_extra = session._provider_extra_params
original_resolve = session._resolve_capabilities
def spy_extra(*args: Any, **kwargs: Any) -> Any:
captured_extra_alias.append(kwargs.get("model_alias"))
return original_extra(*args, **kwargs)
def spy_resolve(*args: Any, **kwargs: Any) -> Any:
# _resolve_capabilities(provider, model, alias)
alias = args[2] if len(args) >= 3 else kwargs.get("alias")
captured_resolve_alias.append(alias)
return original_resolve(*args, **kwargs)
session._provider_extra_params = spy_extra # type: ignore[method-assign]
session._resolve_capabilities = spy_resolve # type: ignore[method-assign]
self._capture_on(session.client) # patch client.chat.completions.create
session._run_agent([{"role": "user", "content": "x"}], label="plan")
assert captured_extra_alias and captured_extra_alias[-1] == "main", (
f"agent fallback path did not inherit primary alias for extra_params: "
f"{captured_extra_alias!r}"
)
assert captured_resolve_alias and captured_resolve_alias[-1] == "main", (
f"agent fallback path did not inherit primary alias for caps: "
f"{captured_resolve_alias!r}"
)
def test_invalid_alias_raises_in_run_agent(self) -> None:
"""Defence-in-depth: _prepare_* validates first, but _run_agent
rejects unknown aliases too rather than silently falling back."""
+274
View File
@@ -0,0 +1,274 @@
"""``models_changed`` SSE fanout coverage.
The console pushes a ``models_changed`` cluster event whenever a model
definition is created / updated / deleted / reloaded, or whenever a
setting in :data:`turnstone.console.server._MODEL_AFFECTING_SETTING_KEYS`
is updated or reset. Connected browsers refetch ``/v1/api/models`` on
receipt so the home composer dropdown + admin Models Roles sub-tab
reflect alias edits without a manual reload.
These tests pin two contracts:
- every model-definition CRUD path emits exactly one ``models_changed``
fanout (so the browser stays in sync with the DB);
- settings PUT / DELETE only emit the fanout when the key is
model-affecting unrelated keys (e.g. ``session.retention_days``)
must not trigger spurious dropdown re-renders across the cluster.
"""
from __future__ import annotations
from typing import Any
from unittest.mock import MagicMock
import pytest
from starlette.applications import Starlette
from starlette.middleware import Middleware
from starlette.routing import Route
from starlette.testclient import TestClient
from tests._coord_test_helpers import _AuthMiddleware
from turnstone.console.server import (
_MODEL_AFFECTING_SETTING_KEYS,
admin_create_model_definition,
admin_delete_model_definition,
admin_delete_setting,
admin_model_reload,
admin_update_model_definition,
admin_update_setting,
)
from turnstone.core.storage._sqlite import SQLiteBackend
@pytest.fixture
def storage(tmp_path: Any) -> SQLiteBackend:
return SQLiteBackend(str(tmp_path / "models_changed.db"))
def _seed(storage: SQLiteBackend, *, definition_id: str, alias: str) -> None:
storage.create_model_definition(
definition_id=definition_id,
alias=alias,
model="model-x",
provider="openai-compatible",
base_url="http://localhost:8000/v1",
api_key="sk-test",
context_window=8192,
capabilities="{}",
enabled=True,
created_by="admin",
)
def _make_client(storage: SQLiteBackend) -> tuple[TestClient, MagicMock]:
"""Build a TestClient + return the stub collector for assertion.
Wires the four model-definition CRUD/reload routes plus the two
settings mutation routes. Collector is a MagicMock so each
``emit_models_changed`` call lands as a recorded call without
spinning up the full SSE listener queue.
"""
app = Starlette(
routes=[
Route(
"/v1/api/admin/model-definitions",
admin_create_model_definition,
methods=["POST"],
),
Route(
"/v1/api/admin/model-definitions/reload",
admin_model_reload,
methods=["POST"],
),
Route(
"/v1/api/admin/model-definitions/{definition_id}",
admin_update_model_definition,
methods=["PUT"],
),
Route(
"/v1/api/admin/model-definitions/{definition_id}",
admin_delete_model_definition,
methods=["DELETE"],
),
Route(
"/v1/api/admin/settings/{key:path}",
admin_update_setting,
methods=["PUT"],
),
Route(
"/v1/api/admin/settings/{key:path}",
admin_delete_setting,
methods=["DELETE"],
),
],
middleware=[Middleware(_AuthMiddleware)],
)
app.state.auth_storage = storage
app.state.coord_registry = None # CRUD endpoints handle this gracefully
collector = MagicMock()
collector.get_all_nodes.return_value = []
app.state.collector = collector
app.state.proxy_client = MagicMock()
app.state.config_store = MagicMock()
client = TestClient(app)
client.headers.update(
{
"X-Test-User": "admin",
"X-Test-Perms": "admin.models,admin.settings",
}
)
return client, collector
# ---------------------------------------------------------------------------
# Model-definition CRUD endpoints fan out ``models_changed``
# ---------------------------------------------------------------------------
def test_create_emits_models_changed(storage: SQLiteBackend) -> None:
client, collector = _make_client(storage)
resp = client.post(
"/v1/api/admin/model-definitions",
json={
"alias": "fast",
"model": "fast-model",
"provider": "openai-compatible",
"base_url": "http://localhost:9000/v1",
"api_key": "sk-x",
"context_window": 4096,
},
)
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 1
def test_update_emits_models_changed(storage: SQLiteBackend) -> None:
_seed(storage, definition_id="m1", alias="local")
client, collector = _make_client(storage)
resp = client.put(
"/v1/api/admin/model-definitions/m1",
json={"model": "swapped-model"},
)
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 1
def test_update_with_empty_body_does_not_emit(storage: SQLiteBackend) -> None:
"""Empty-body PUT writes no rows + skips the registry refresh — no
SSE fanout either, since nothing actually changed."""
_seed(storage, definition_id="m1", alias="local")
client, collector = _make_client(storage)
resp = client.put("/v1/api/admin/model-definitions/m1", json={})
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 0
def test_delete_emits_models_changed(storage: SQLiteBackend) -> None:
_seed(storage, definition_id="m1", alias="local")
client, collector = _make_client(storage)
resp = client.delete("/v1/api/admin/model-definitions/m1")
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 1
def test_reload_emits_models_changed(storage: SQLiteBackend) -> None:
_seed(storage, definition_id="m1", alias="local")
client, collector = _make_client(storage)
resp = client.post("/v1/api/admin/model-definitions/reload")
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 1
# ---------------------------------------------------------------------------
# Settings PUT / DELETE only emit for model-affecting keys
# ---------------------------------------------------------------------------
# Pinned snapshot of the role-related keys we expect the allowlist to
# cover today. The frozenset itself is asserted further down so a
# stray addition doesn't silently bypass coverage.
_EXPECTED_AFFECTING_KEYS = frozenset(
{
"model.default_alias",
"model.plan_alias",
"model.plan_effort",
"model.task_alias",
"model.task_effort",
"coordinator.model_alias",
"coordinator.reasoning_effort",
"judge.model",
}
)
def _value_for_key(key: str) -> str:
"""Return a registry-valid value for ``key``.
``reasoning_effort`` keys have a fixed choice list; alias-shaped
keys accept arbitrary strings. Avoids per-key custom payloads.
"""
if (
key.endswith("reasoning_effort")
or key.endswith("plan_effort")
or key.endswith("task_effort")
):
return "low"
return "anything"
def test_affecting_keys_set_matches_expected() -> None:
"""Lock in the allowlist so an unintentional removal is caught."""
assert _MODEL_AFFECTING_SETTING_KEYS == _EXPECTED_AFFECTING_KEYS
@pytest.mark.parametrize("key", sorted(_EXPECTED_AFFECTING_KEYS))
def test_settings_put_emits_for_model_affecting_key(storage: SQLiteBackend, key: str) -> None:
client, collector = _make_client(storage)
resp = client.put(
f"/v1/api/admin/settings/{key}",
json={"value": _value_for_key(key)},
)
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 1
@pytest.mark.parametrize("key", sorted(_EXPECTED_AFFECTING_KEYS))
def test_settings_delete_emits_for_model_affecting_key(storage: SQLiteBackend, key: str) -> None:
client, collector = _make_client(storage)
# Seed a row so DELETE has something to remove (otherwise 404).
client.put(
f"/v1/api/admin/settings/{key}",
json={"value": _value_for_key(key)},
)
collector.emit_models_changed.reset_mock()
resp = client.delete(f"/v1/api/admin/settings/{key}")
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 1
def test_settings_put_does_not_emit_for_unrelated_key(
storage: SQLiteBackend,
) -> None:
"""Updating a non-model setting (here: a session retention knob)
must not trigger a cluster-wide dropdown refresh."""
client, collector = _make_client(storage)
resp = client.put(
"/v1/api/admin/settings/session.retention_days",
json={"value": 30},
)
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 0
def test_settings_delete_does_not_emit_for_unrelated_key(
storage: SQLiteBackend,
) -> None:
client, collector = _make_client(storage)
client.put(
"/v1/api/admin/settings/session.retention_days",
json={"value": 30},
)
collector.emit_models_changed.reset_mock()
resp = client.delete("/v1/api/admin/settings/session.retention_days")
assert resp.status_code == 200, resp.text
assert collector.emit_models_changed.call_count == 0
+28
View File
@@ -1383,6 +1383,34 @@ class TestProviderFactory:
p2 = create_provider("openai")
assert p1 is p2
def test_create_provider_compat_responses_surface(self) -> None:
"""openai-compatible + api_surface=responses returns the Responses provider."""
from turnstone.core.providers import OpenAIResponsesProvider, create_provider
provider = create_provider("openai-compatible", api_surface="responses")
assert isinstance(provider, OpenAIResponsesProvider)
def test_create_provider_compat_chat_surface_default(self) -> None:
"""openai-compatible defaults to Chat Completions."""
from turnstone.core.providers import create_provider
for surface in (None, "", "chat"):
provider = create_provider("openai-compatible", api_surface=surface)
assert isinstance(provider, OpenAIChatCompletionsProvider)
def test_create_provider_invalid_api_surface(self) -> None:
from turnstone.core.providers import create_provider
with pytest.raises(ValueError, match="Unknown api_surface"):
create_provider("openai-compatible", api_surface="bogus")
def test_create_provider_openai_ignores_api_surface(self) -> None:
"""Cloud OpenAI is always Responses regardless of api_surface."""
from turnstone.core.providers import OpenAIResponsesProvider, create_provider
provider = create_provider("openai", api_surface="chat")
assert isinstance(provider, OpenAIResponsesProvider)
# -- Google provider -------------------------------------------------------
def test_create_provider_google(self) -> None:
+80 -35
View File
@@ -89,6 +89,27 @@ class TestSuggestProfile:
p = suggest_profile("vllm", "Google/GEMMA-4-31B-IT")
assert p["capabilities"]["thinking_mode"] == "manual"
def test_vllm_mistral_medium_not_auto_suggested(self) -> None:
"""Mistral medium falls back to the generic vLLM profile.
We don't auto-suggest the Responses surface for Mistral medium because
vLLM's Responses API tool-call parser isn't wired up for it yet
operators who want per-request reasoning effort must pick "Responses
API" manually in the admin UI and accept the tool-calling limitation.
"""
p = suggest_profile("vllm", "mistralai/Mistral-Medium-3-Instruct")
assert p["server_compat"]["server_type"] == "vllm"
assert "api_surface" not in p["server_compat"]
assert "capabilities" not in p
def test_vllm_mistral_medium_profile_still_available(self) -> None:
"""The vllm-mistral-medium profile remains in _PROFILES so an operator
who explicitly opts in via the admin UI gets the Responses surface."""
from turnstone.core.server_compat import _PROFILES
assert "vllm-mistral-medium" in _PROFILES
assert _PROFILES["vllm-mistral-medium"]["server_compat"]["api_surface"] == "responses"
def test_holo_requires_holo2(self) -> None:
"""Short 'holo' prefix shouldn't false-match; 'holo2' should match."""
p_short = suggest_profile("vllm", "some-org/hologram-7b")
@@ -110,32 +131,47 @@ class TestSuggestProfile:
class TestMergeServerCompat:
def test_empty_compat_returns_base_only(self) -> None:
def test_empty_base_and_compat_is_empty(self) -> None:
"""No base, no compat → no extra_body needed."""
assert merge_server_compat(None, {}) == {}
assert merge_server_compat({}, {}) == {}
def test_explicit_base_passes_through(self) -> None:
"""Explicit chat_template_kwargs base is forwarded as-is."""
base = {"reasoning_effort": "medium"}
result = merge_server_compat(base, {})
assert result == {"chat_template_kwargs": {"reasoning_effort": "medium"}}
def test_extra_body_merged_top_level(self) -> None:
base = {"reasoning_effort": "medium"}
compat = {"extra_body": {"skip_special_tokens": False}}
result = merge_server_compat(base, compat)
assert result["skip_special_tokens"] is False
assert "chat_template_kwargs" in result
def test_extra_body_merged_top_level_no_base(self) -> None:
"""Server-level overrides forward without a chat_template_kwargs wrapper."""
result = merge_server_compat(None, {"extra_body": {"skip_special_tokens": False}})
assert result == {"skip_special_tokens": False}
def test_full_vllm_gemma_compat(self) -> None:
base = {"reasoning_effort": "medium"}
def test_full_vllm_gemma_compat_no_base(self) -> None:
"""vLLM workaround forwards on its own."""
compat = {
"server_type": "vllm",
"extra_body": {"skip_special_tokens": False},
}
result = merge_server_compat(base, compat)
result = merge_server_compat(None, compat)
assert result == {"skip_special_tokens": False}
def test_operator_chat_template_kwargs_only(self) -> None:
"""Operator can set chat_template_kwargs explicitly without seeding the base."""
compat = {
"extra_body": {
"chat_template_kwargs": {"reasoning_effort": "high"},
"skip_special_tokens": False,
},
}
result = merge_server_compat(None, compat)
assert result == {
"chat_template_kwargs": {"reasoning_effort": "medium"},
"chat_template_kwargs": {"reasoning_effort": "high"},
"skip_special_tokens": False,
}
def test_extra_body_chat_template_kwargs_deep_merged(self) -> None:
"""chat_template_kwargs in extra_body is deep-merged, operator wins."""
def test_extra_body_chat_template_kwargs_deep_merged_with_base(self) -> None:
"""Operator chat_template_kwargs deep-merges over the seeded base."""
base = {"reasoning_effort": "medium"}
compat = {
"extra_body": {
@@ -144,17 +180,15 @@ class TestMergeServerCompat:
},
}
result = merge_server_compat(base, compat)
# Operator values win over base
assert result["chat_template_kwargs"]["custom_flag"] is True
# Operator value wins over seeded base
assert result["chat_template_kwargs"]["reasoning_effort"] == "high"
assert result["skip_special_tokens"] is False
def test_extra_body_chat_template_kwargs_non_dict_ignored(self) -> None:
"""Non-dict chat_template_kwargs in extra_body is safely ignored."""
base = {"reasoning_effort": "medium"}
compat = {"extra_body": {"chat_template_kwargs": "bad"}}
result = merge_server_compat(base, compat)
assert result["chat_template_kwargs"] == {"reasoning_effort": "medium"}
assert merge_server_compat(None, compat) == {}
def test_base_not_mutated(self) -> None:
base = {"reasoning_effort": "medium"}
@@ -164,9 +198,7 @@ class TestMergeServerCompat:
def test_non_dict_extra_body_ignored(self) -> None:
"""Gracefully handle malformed server_compat."""
base = {"reasoning_effort": "medium"}
result = merge_server_compat(base, {"extra_body": 42})
assert result == {"chat_template_kwargs": {"reasoning_effort": "medium"}}
assert merge_server_compat(None, {"extra_body": 42}) == {}
# ---------------------------------------------------------------------------
@@ -178,45 +210,58 @@ class TestEndToEndRequestShaping:
"""Compose both layers — session builds extra_params, provider applies thinking."""
def test_vllm_gemma_full_flow(self) -> None:
"""Session merges server workarounds, provider adds thinking param."""
"""Session forwards server workarounds, provider adds thinking param."""
caps = ModelCapabilities(thinking_mode="manual", thinking_param="enable_thinking")
base_ctk = {"reasoning_effort": "medium"}
server_compat = {
"server_type": "vllm",
"extra_body": {"skip_special_tokens": False},
}
# Step 1: session merges
extra_params = merge_server_compat(base_ctk, server_compat)
# Step 2: provider finalises
# Step 1: session forwards (no auto-injection of reasoning_effort).
extra_params = merge_server_compat(None, server_compat)
# Step 2: provider injects thinking param into chat_template_kwargs.
extra_body = dict(extra_params)
OpenAIChatCompletionsProvider._apply_thinking_mode(extra_body, caps)
assert extra_body == {
"chat_template_kwargs": {
"reasoning_effort": "medium",
"enable_thinking": True,
},
"chat_template_kwargs": {"enable_thinking": True},
"skip_special_tokens": False,
}
def test_granite_thinking_key(self) -> None:
"""Granite uses 'thinking' instead of 'enable_thinking'."""
caps = ModelCapabilities(thinking_mode="manual", thinking_param="thinking")
extra_params = merge_server_compat({"reasoning_effort": "low"}, {})
extra_params = merge_server_compat(None, {})
extra_body = dict(extra_params)
OpenAIChatCompletionsProvider._apply_thinking_mode(extra_body, caps)
assert extra_body["chat_template_kwargs"]["thinking"] is True
assert "enable_thinking" not in extra_body["chat_template_kwargs"]
assert extra_body == {"chat_template_kwargs": {"thinking": True}}
def test_non_thinking_model_no_injection(self) -> None:
"""Non-thinking model gets no thinking params."""
"""Non-thinking model gets no chat_template_kwargs at all."""
caps = ModelCapabilities() # thinking_mode="none"
extra_params = merge_server_compat({"reasoning_effort": "medium"}, {})
extra_params = merge_server_compat(None, {})
extra_body = dict(extra_params)
OpenAIChatCompletionsProvider._apply_thinking_mode(extra_body, caps)
assert extra_body == {"chat_template_kwargs": {"reasoning_effort": "medium"}}
assert extra_body == {}
def test_operator_reasoning_effort_passthrough(self) -> None:
"""Operator-supplied reasoning_effort under chat_template_kwargs is preserved."""
caps = ModelCapabilities(thinking_mode="manual", thinking_param="enable_thinking")
compat = {
"server_type": "vllm",
"extra_body": {"chat_template_kwargs": {"reasoning_effort": "high"}},
}
extra_params = merge_server_compat(None, compat)
extra_body = dict(extra_params)
OpenAIChatCompletionsProvider._apply_thinking_mode(extra_body, caps)
assert extra_body == {
"chat_template_kwargs": {
"reasoning_effort": "high",
"enable_thinking": True,
},
}
# ---------------------------------------------------------------------------
+346
View File
@@ -0,0 +1,346 @@
"""Tests for the per-node ``models`` metadata pipeline.
Two helpers in ``server.py`` carry the load:
- ``_collect_node_models_metadata`` projects the live ``ModelRegistry``
into the node_metadata row shape ``[{alias, provider, healthy}, ...]``.
- ``_publish_models_metadata`` short-circuits redundant writes via a
payload cache on ``app_state`` and is the helper called from both
the heartbeat loop and ``internal_model_reload``.
These tests pin the projection shape, the health-flag wiring, the
cache short-circuit, and the model-reload integration.
"""
from __future__ import annotations
import json
from types import SimpleNamespace
from unittest.mock import MagicMock
from turnstone.core.healthcheck import HealthTrackerRegistry
from turnstone.core.model_registry import ModelConfig, ModelRegistry
from turnstone.server import (
_collect_node_models_metadata,
_publish_models_metadata,
)
def _registry(*aliases_with_url: tuple[str, str]) -> ModelRegistry:
"""Build a registry from ``(alias, base_url)`` pairs.
Two aliases sharing a ``base_url`` deliberately share a tracker
that's the contract the cluster-level health surface needs to
preserve, and it's worth pinning in a test.
"""
models = {
alias: ModelConfig(alias=alias, base_url=url, api_key="k", model=alias, provider="openai")
for alias, url in aliases_with_url
}
default = aliases_with_url[0][0]
return ModelRegistry(models, default=default)
def test_returns_none_when_registry_missing():
state = SimpleNamespace()
assert _collect_node_models_metadata(state) is None
def test_projects_all_aliases_with_default_healthy_when_no_tracker():
"""Without a ``health_registry`` (or before any request has flowed
through a backend), every alias surfaces as ``healthy=True``
operators shouldn't get an empty ``models`` list on a freshly
started node just because the backends haven't been exercised."""
reg = _registry(("a", "http://x"), ("b", "http://y"))
state = SimpleNamespace(registry=reg)
entry = _collect_node_models_metadata(state)
assert entry is not None
key, value, source = entry
assert key == "models"
assert source == "auto"
rows = json.loads(value)
assert len(rows) == 2
aliases = {r["alias"] for r in rows}
assert aliases == {"a", "b"}
assert all(r["healthy"] is True for r in rows)
assert all(r["provider"] == "openai" for r in rows)
# Provider-side model identifier intentionally omitted — coords
# kept passing it as ``spawn_workstream(model=...)`` when they
# should have passed the local alias. Lock the projected keys
# so a future contributor doesn't reintroduce the footgun.
for row in rows:
assert set(row.keys()) == {"alias", "provider", "healthy"}
def test_health_flag_reflects_tracker_state():
reg = _registry(("a", "http://x"), ("b", "http://y"))
health_reg = HealthTrackerRegistry(failure_threshold=2)
# Seed the tracker for "a"'s backend and drive it into the degraded
# state — two consecutive failures cross the threshold.
bad_tracker = health_reg.get_tracker(provider="openai", base_url="http://x")
bad_tracker.record_failure()
bad_tracker.record_failure()
assert bad_tracker.is_degraded
# "b" gets a tracker that has only seen successes.
good_tracker = health_reg.get_tracker(provider="openai", base_url="http://y")
good_tracker.record_success()
state = SimpleNamespace(registry=reg, health_registry=health_reg)
rows = json.loads(_collect_node_models_metadata(state)[1])
by_alias = {r["alias"]: r for r in rows}
assert by_alias["a"]["healthy"] is False
assert by_alias["b"]["healthy"] is True
def test_two_aliases_sharing_a_backend_share_a_tracker():
"""Two aliases that point at the same ``(provider, base_url)``
share a single :class:`BackendHealthTracker` degrading one is
expected to surface as degraded on the other. The list_nodes
projection should respect that, otherwise a coord could see
``alias-a`` healthy and ``alias-b`` degraded for the same
backend."""
reg = _registry(("alpha", "http://shared"), ("beta", "http://shared"))
health_reg = HealthTrackerRegistry(failure_threshold=1)
tracker = health_reg.get_tracker(provider="openai", base_url="http://shared")
tracker.record_failure() # threshold=1 — degraded immediately
state = SimpleNamespace(registry=reg, health_registry=health_reg)
rows = json.loads(_collect_node_models_metadata(state)[1])
assert {r["alias"]: r["healthy"] for r in rows} == {"alpha": False, "beta": False}
def test_alias_with_no_tracker_yet_defaults_to_healthy():
"""An alias the registry knows about but whose backend hasn't been
invoked yet has no tracker. Default to healthy so a brand-new
alias is immediately visible to coordinators rather than waiting
for the first request to seed a tracker.
The collector calls ``health_reg.get_tracker(...)`` which mints a
fresh tracker on first lookup that's the path under test here.
The freshly minted tracker reports ``is_healthy=True`` (default
state), so the projection labels the alias healthy.
"""
reg = _registry(("a", "http://x"))
health_reg = HealthTrackerRegistry() # empty — no trackers seeded
state = SimpleNamespace(registry=reg, health_registry=health_reg)
rows = json.loads(_collect_node_models_metadata(state)[1])
assert rows[0]["healthy"] is True
# ---------------------------------------------------------------------------
# _publish_models_metadata — cache short-circuit + projection wiring
# ---------------------------------------------------------------------------
def _publish_state() -> SimpleNamespace:
"""Build an ``app_state`` with a minimal registry + health surface."""
reg = _registry(("a", "http://x"))
return SimpleNamespace(registry=reg, health_registry=HealthTrackerRegistry())
def test_publish_writes_when_payload_changes():
"""First publish has nothing in the cache — write happens; cache
fills. Second publish on the same unchanged registry skips the
write entirely."""
state = _publish_state()
storage = MagicMock()
_publish_models_metadata(state, storage, "node-a")
assert storage.set_node_metadata_bulk.call_count == 1
cached = state._last_models_payload
assert isinstance(cached, str) and "alias" in cached
# Second call, same registry, same health: cached payload matches
# — write must be skipped to avoid the per-30s UPSERT churn.
_publish_models_metadata(state, storage, "node-a")
assert storage.set_node_metadata_bulk.call_count == 1
def test_publish_records_metric_outcome(monkeypatch):
"""The publish helper feeds ``record_node_models_publish`` so
Prometheus can expose the hit-rate. Storage failures must NOT
record either outcome counters should reflect actual cache
decisions, not transient DB errors that will retry.
Replaces the module-level ``turnstone.server._metrics`` binding
via string-form monkeypatch (with auto-restore) rather than
patching an instance attribute on the imported singleton. Other
tests in the suite reassign ``srv_mod._metrics`` (some without
using monkeypatch), so an instance captured at import time can
diverge from the binding the live ``_publish_models_metadata``
reads on each call.
"""
state = _publish_state()
storage = MagicMock()
calls: list[bool] = []
class _FakeMetrics:
def record_node_models_publish(self, *, written: bool) -> None:
calls.append(written)
monkeypatch.setattr("turnstone.server._metrics", _FakeMetrics())
_publish_models_metadata(state, storage, "node-a") # first → write
_publish_models_metadata(state, storage, "node-a") # second → skip
assert calls == [True, False]
# Storage error: no metric recorded.
storage.set_node_metadata_bulk.side_effect = RuntimeError("db down")
state._last_models_payload = None # invalidate cache to force a write attempt
_publish_models_metadata(state, storage, "node-a")
assert calls == [True, False] # unchanged
def test_publish_rewrites_when_health_flips():
"""A health-tracker state change must invalidate the cache and
drive a fresh write otherwise the discovery surface would lag
a flip indefinitely."""
state = _publish_state()
storage = MagicMock()
_publish_models_metadata(state, storage, "node-a")
assert storage.set_node_metadata_bulk.call_count == 1
# Drive the only tracker to degraded.
tracker = state.health_registry.get_tracker(provider="openai", base_url="http://x")
for _ in range(10):
tracker.record_failure()
assert tracker.is_degraded
_publish_models_metadata(state, storage, "node-a")
assert storage.set_node_metadata_bulk.call_count == 2
def test_publish_swallows_storage_error_without_updating_cache():
"""A storage failure must NOT poison the cache — the next call
should retry the write rather than think it succeeded."""
state = _publish_state()
storage = MagicMock()
storage.set_node_metadata_bulk.side_effect = RuntimeError("db down")
_publish_models_metadata(state, storage, "node-a")
assert storage.set_node_metadata_bulk.call_count == 1
assert getattr(state, "_last_models_payload", None) is None
# Recover: a subsequent successful call writes again.
storage.set_node_metadata_bulk.side_effect = None
_publish_models_metadata(state, storage, "node-a")
assert storage.set_node_metadata_bulk.call_count == 2
assert state._last_models_payload is not None
def test_publish_skips_when_registry_missing():
"""Without a registry there's nothing to project; nothing should
be written and the cache must not be set."""
state = SimpleNamespace()
storage = MagicMock()
_publish_models_metadata(state, storage, "node-a")
assert storage.set_node_metadata_bulk.call_count == 0
assert getattr(state, "_last_models_payload", None) is None
# ---------------------------------------------------------------------------
# internal_model_reload — integration: registry change must rewrite the row
# ---------------------------------------------------------------------------
def test_model_reload_endpoint_rewrites_models_metadata(monkeypatch, tmp_path):
"""A successful ``internal_model_reload`` must refresh
``node_metadata.models`` so a coordinator sees the new alias on
its next ``list_nodes`` without waiting up to 30s for the
heartbeat tick.
The endpoint pulls a fresh registry from
``load_model_registry(...)`` and reloads in-place we stub the
loader to return a registry with a different alias set so the
publish-cache invalidation is exercised end-to-end.
"""
from turnstone.core.storage._sqlite import SQLiteBackend
from turnstone.server import internal_model_reload
storage = SQLiteBackend(str(tmp_path / "reload.db"))
# Old registry — single alias "a".
old_reg = _registry(("a", "http://x"))
# New registry that ``load_model_registry`` will return — adds "b".
new_reg = ModelRegistry(
{
"a": ModelConfig(
alias="a", base_url="http://x", api_key="k", model="a", provider="openai"
),
"b": ModelConfig(
alias="b", base_url="http://y", api_key="k", model="b", provider="openai"
),
},
default="a",
)
health_reg = HealthTrackerRegistry()
app_state = SimpleNamespace(
registry=old_reg,
health_registry=health_reg,
cli_model_args={
"base_url": "",
"api_key": "",
"model": "",
"context_window": 0,
"provider": "openai",
},
config_store=None,
node_id="node-a",
)
request = SimpleNamespace(app=SimpleNamespace(state=app_state))
# Patch the loader and storage accessors used inside the endpoint.
# ``internal_model_reload`` does ``from turnstone.core.storage._registry
# import get_storage`` inline, so patching the symbol on that module
# is what intercepts the call.
monkeypatch.setattr("turnstone.core.model_registry.load_model_registry", lambda **_kw: new_reg)
monkeypatch.setattr("turnstone.core.storage._registry.get_storage", lambda: storage)
# The endpoint also broadcasts schema refreshes to active sessions
# — stub this out, it's irrelevant to the metadata-write path.
monkeypatch.setattr("turnstone.server._broadcast_agent_tool_schema_refresh", lambda _s: None)
response = internal_model_reload(request) # type: ignore[arg-type]
assert response.status_code == 200
rows = storage.get_node_metadata("node-a")
by_key = {r["key"]: r for r in rows}
assert "models" in by_key
payload = json.loads(by_key["models"]["value"])
assert {r["alias"] for r in payload} == {"a", "b"}
# ---------------------------------------------------------------------------
# Shutdown race: heartbeat write must NOT resurrect post-shutdown delete
# ---------------------------------------------------------------------------
def test_heartbeat_write_awaits_before_shutdown_delete():
"""Pin the shutdown-race fix.
Before the fix, the lifespan shutdown sequence was:
1. ``_heartbeat_task.cancel()`` fire-and-forget
2. ``delete_node_metadata_by_source(node_id, "auto")``
A heartbeat tick already inside ``asyncio.to_thread(...)`` for
the ``set_node_metadata_bulk`` call would complete AFTER step 2,
resurrecting the deleted ``models`` row. The fix awaits the
cancelled task with ``contextlib.suppress(...)`` between (1) and
(2), so the in-flight write lands first.
We verify the fix by introspecting ``server.py`` source the
real lifespan is hard to test deterministically without a full
Starlette app, but the textual ordering between
``_heartbeat_task.cancel()`` and the delete is a stable contract
that catches the regression cheaply.
"""
import inspect
import sys
src = inspect.getsource(sys.modules[_collect_node_models_metadata.__module__])
cancel_idx = src.find("_heartbeat_task.cancel()")
delete_idx = src.find('delete_node_metadata_by_source, _svc_node_id, "auto"')
await_idx = src.find("await _heartbeat_task", cancel_idx)
assert cancel_idx != -1
assert delete_idx != -1
assert await_idx != -1
# The fix-line must sit BETWEEN the cancel and the delete.
assert cancel_idx < await_idx < delete_idx, (
"Shutdown race regression: "
"_heartbeat_task.cancel() must be followed by `await _heartbeat_task` "
"BEFORE delete_node_metadata_by_source(..., 'auto') so an in-flight "
"set_node_metadata_bulk lands before the delete."
)
+44 -66
View File
@@ -1182,7 +1182,7 @@ class TestAgentOutputGuard:
class TestProviderExtraParams:
"""Tests for _provider_extra_params — local-only chat_template_kwargs."""
"""Tests for _provider_extra_params — server_compat passthrough only."""
def _session_with_provider(self, provider_name: str, tmp_db) -> ChatSession:
from turnstone.core.providers import create_provider
@@ -1191,68 +1191,36 @@ class TestProviderExtraParams:
session._provider = create_provider(provider_name)
return session
def test_openai_compatible_returns_chat_template_kwargs(self, tmp_db):
def test_openai_compatible_no_compat_returns_none(self, tmp_db):
"""No server_compat → no extra_body needed (no auto-injection)."""
session = self._session_with_provider("openai-compatible", tmp_db)
result = session._provider_extra_params()
assert result is not None
assert "chat_template_kwargs" in result
assert result["chat_template_kwargs"]["reasoning_effort"] == "medium"
assert session._provider_extra_params() is None
def test_openai_commercial_returns_none(self, tmp_db):
def test_openai_commercial_no_compat_returns_none(self, tmp_db):
"""Cloud OpenAI without server_compat → None."""
session = self._session_with_provider("openai", tmp_db)
result = session._provider_extra_params()
assert result is None
assert session._provider_extra_params() is None
def test_anthropic_returns_none(self, tmp_db):
session = self._session_with_provider("anthropic", tmp_db)
result = session._provider_extra_params()
assert result is None
assert session._provider_extra_params() is None
def test_reasoning_effort_override(self, tmp_db):
def test_no_reasoning_effort_kwarg(self, tmp_db):
"""reasoning_effort is not part of the surface; passing it should TypeError.
Splatted via ``**kwargs`` so static analyzers (CodeQL "wrong-name
argument" / mypy) don't flag the call — the point of this test is the
runtime contract, not the static type.
"""
import pytest
bad_kwargs = {"reasoning_effort": "high"}
session = self._session_with_provider("openai-compatible", tmp_db)
result = session._provider_extra_params(reasoning_effort="high")
assert result is not None
assert result["chat_template_kwargs"]["reasoning_effort"] == "high"
with pytest.raises(TypeError):
session._provider_extra_params(**bad_kwargs)
def test_explicit_openai_provider_overrides_session(self, tmp_db):
"""Passing an explicit commercial OpenAI provider returns None even
when the session's own provider is openai-compatible."""
from turnstone.core.providers import create_provider
session = self._session_with_provider("openai-compatible", tmp_db)
openai_prov = create_provider("openai")
result = session._provider_extra_params(provider=openai_prov)
assert result is None
def test_server_compat_extra_body_merged(self, tmp_db):
"""server_compat.extra_body workarounds are merged into extra_params."""
from turnstone.core.model_registry import ModelConfig, ModelRegistry
session = self._session_with_provider("openai-compatible", tmp_db)
cfg = ModelConfig(
alias="test",
base_url="http://localhost:8000/v1",
api_key="none",
model="google/gemma-4-31B-it",
server_compat={
"extra_body": {"skip_special_tokens": False},
},
)
session._registry = ModelRegistry(models={"test": cfg}, default="test")
session._model_alias = "test"
result = session._provider_extra_params()
assert result is not None
assert result["chat_template_kwargs"]["reasoning_effort"] == "medium"
assert result["skip_special_tokens"] is False
def test_empty_server_compat_backwards_compatible(self, tmp_db):
"""Empty server_compat produces same output as before."""
session = self._session_with_provider("openai-compatible", tmp_db)
result = session._provider_extra_params()
assert result == {"chat_template_kwargs": {"reasoning_effort": "medium"}}
def test_server_compat_with_reasoning_effort_override(self, tmp_db):
"""reasoning_effort override works alongside server_compat."""
def test_server_compat_extra_body_passes_through(self, tmp_db):
"""server_compat.extra_body workarounds forward as extra_params."""
from turnstone.core.model_registry import ModelConfig, ModelRegistry
session = self._session_with_provider("openai-compatible", tmp_db)
@@ -1265,10 +1233,25 @@ class TestProviderExtraParams:
)
session._registry = ModelRegistry(models={"test": cfg}, default="test")
session._model_alias = "test"
result = session._provider_extra_params(reasoning_effort="high")
assert result is not None
assert result["chat_template_kwargs"]["reasoning_effort"] == "high"
assert result["skip_special_tokens"] is False
result = session._provider_extra_params()
assert result == {"skip_special_tokens": False}
def test_operator_chat_template_kwargs_pass_through(self, tmp_db):
"""Operator-set chat_template_kwargs (e.g. for gpt-oss) forwards verbatim."""
from turnstone.core.model_registry import ModelConfig, ModelRegistry
session = self._session_with_provider("openai-compatible", tmp_db)
cfg = ModelConfig(
alias="test",
base_url="http://localhost:8000/v1",
api_key="none",
model="openai/gpt-oss-120b",
server_compat={"extra_body": {"chat_template_kwargs": {"reasoning_effort": "high"}}},
)
session._registry = ModelRegistry(models={"test": cfg}, default="test")
session._model_alias = "test"
result = session._provider_extra_params()
assert result == {"chat_template_kwargs": {"reasoning_effort": "high"}}
def test_model_alias_resolves_target_compat(self, tmp_db):
"""model_alias parameter selects compat from the target, not the primary."""
@@ -1297,14 +1280,9 @@ class TestProviderExtraParams:
session._model_alias = "primary"
# Primary alias → gets Gemma workaround
result_primary = session._provider_extra_params()
assert result_primary is not None
assert result_primary["skip_special_tokens"] is False
# Fallback alias → no compat, just base kwargs
result_fallback = session._provider_extra_params(model_alias="fallback")
assert result_fallback == {"chat_template_kwargs": {"reasoning_effort": "medium"}}
assert "skip_special_tokens" not in result_fallback
assert session._provider_extra_params() == {"skip_special_tokens": False}
# Fallback alias → no compat at all
assert session._provider_extra_params(model_alias="fallback") is None
class TestSafePrepareTool:
+252 -2
View File
@@ -19,7 +19,10 @@ import threading
import time
from dataclasses import dataclass
from datetime import UTC, datetime
from typing import Any
from typing import TYPE_CHECKING, Any
if TYPE_CHECKING:
from collections.abc import Callable
from unittest.mock import MagicMock
import pytest
@@ -98,6 +101,7 @@ class FakeAdapter:
self.cleaned_up: list[str] = []
self.build_session_calls = 0
self.build_session_raises = build_session_raises
self.last_build_model: object | None = None
# Slow down session build so concurrent tests can race.
self.build_session_delay = 0.0
@@ -141,8 +145,13 @@ class FakeAdapter:
def build_ui(self, ws: Workstream) -> Any:
return FakeUI()
def build_session(self, ws: Workstream, **_: object) -> Any:
def build_session(self, ws: Workstream, **kwargs: object) -> Any:
self.build_session_calls += 1
# Record the ``model`` kwarg (None on fresh-create, the saved
# alias on rehydrate) so tests can assert SessionManager.open()
# threads the persisted alias through to construction instead
# of letting the adapter resolve the *current* default alias.
self.last_build_model = kwargs.get("model")
if self.build_session_delay:
time.sleep(self.build_session_delay)
if self.build_session_raises:
@@ -181,6 +190,13 @@ class FakeStorage:
# "no peers alive" (every row unprotected by liveness).
self.live_services: dict[str, list[str]] = {}
self.list_services_raises = False
# Per-ws config (model_alias, temperature, …). Populated by
# tests that exercise the rehydrate-preserves-config path; the
# SessionManager.open() rehydrate path reads this through
# ``self._storage.load_workstream_config`` so it can pass the
# saved alias into ``build_session`` and avoid clobbering the
# original on construction.
self.ws_config: dict[str, dict[str, str]] = {}
@staticmethod
def _now_iso() -> str:
@@ -291,6 +307,18 @@ class FakeStorage:
def count_skill_versions(self, template_id: str) -> int:
return 0
def load_workstream_config(self, ws_id: str) -> dict[str, str]:
with self.lock:
return dict(self.ws_config.get(ws_id, {}))
def save_workstream_config(self, ws_id: str, config: dict[str, str]) -> None:
# Mirrors the real backend's INSERT OR REPLACE per-key semantics
# — callers expect a partial save to overwrite only the keys
# they pass, not the whole row.
with self.lock:
row = self.ws_config.setdefault(ws_id, {})
row.update(config)
_EMITTER_DEFAULT = object()
@@ -302,6 +330,7 @@ def _make_manager(
storage: FakeStorage | None = None,
event_emitter: Any = _EMITTER_DEFAULT,
node_id: str | None = None,
model_validator: Callable[[str], bool] | None = None,
) -> tuple[SessionManager, FakeAdapter, FakeStorage]:
"""Build a SessionManager wired to a FakeAdapter for both Protocols.
@@ -321,6 +350,7 @@ def _make_manager(
max_active=max_active,
event_emitter=emitter,
node_id=node_id,
model_validator=model_validator,
)
return mgr, adapter, storage
@@ -654,6 +684,102 @@ def test_open_resurrects_closed_state() -> None:
assert ws_id in [e.ws_id for e in adapter.events_of("rehydrated")]
def test_open_threads_saved_model_alias_into_build_session() -> None:
"""Reopening a closed ws must build the session with the *original*
model alias, not the current registry default.
Without this, ``build_session(ws)`` is called with ``model=None``
the production session_factory resolves ``_effective_default_alias()``
ChatSession's ``__init__`` writes those defaults to
``workstream_config`` (INSERT OR REPLACE) the subsequent
``resume()`` restores what is now the default. Net effect: every
persisted knob (model, temperature, reasoning_effort, max_tokens,
skill, creative_mode, instructions, ) silently resets on every
reopen and on every service restart.
"""
mgr, adapter, storage = _make_manager()
ws = mgr.create(user_id="u1")
ws_id = ws.id
# Pretend the user set a non-default alias when the ws was created;
# the real path goes through ChatSession._save_config but the
# FakeSession in this suite doesn't model that, so seed directly.
storage.ws_config[ws_id] = {"model_alias": "gpt-5-pro"}
mgr.close(ws_id)
adapter.last_build_model = "<unset>" # sentinel — must be overwritten
reopened = mgr.open(ws_id)
assert reopened is not None
assert adapter.last_build_model == "gpt-5-pro"
def test_open_drops_saved_alias_when_validator_rejects() -> None:
"""When the persisted alias is no longer in the registry, the
manager must drop it before reaching ``build_session``. The
factory still raises on unknown aliases on the fresh-create path
(so a typo in body.model surfaces as 503), so the rehydrate path
has to filter the alias here rather than relying on factory-side
fallback. Without this filter, every reopen of a workstream pinned
to a since-removed alias 500s."""
mgr, adapter, storage = _make_manager(
# Validator says "alias is no longer in the registry".
model_validator=lambda alias: False,
)
ws = mgr.create(user_id="u1")
ws_id = ws.id
storage.ws_config[ws_id] = {"model_alias": "since-removed-alias"}
mgr.close(ws_id)
adapter.last_build_model = "<unset>"
reopened = mgr.open(ws_id)
assert reopened is not None
assert adapter.last_build_model is None # alias dropped before reaching build_session
def test_open_keeps_saved_alias_when_validator_accepts() -> None:
"""Sanity: an alias that still resolves must be passed through
unchanged. Filter only fires for stale aliases."""
accepted: list[str] = []
def validator(alias: str) -> bool:
accepted.append(alias)
return True
mgr, adapter, storage = _make_manager(model_validator=validator)
ws = mgr.create(user_id="u1")
ws_id = ws.id
storage.ws_config[ws_id] = {"model_alias": "still-live"}
mgr.close(ws_id)
adapter.last_build_model = "<unset>"
reopened = mgr.open(ws_id)
assert reopened is not None
assert accepted == ["still-live"]
assert adapter.last_build_model == "still-live"
def test_open_falls_back_to_none_when_no_saved_alias() -> None:
"""Reopening a ws with no saved alias must pass ``model=None`` to
``build_session`` so the adapter's session_factory can fall back to
the current default matching the user's intent: best effort
restore, default when the original is gone."""
mgr, adapter, storage = _make_manager()
ws = mgr.create(user_id="u1")
ws_id = ws.id
# No ws_config row — simulates "alias was never saved" or "saved
# alias was empty string".
assert ws_id not in storage.ws_config
mgr.close(ws_id)
adapter.last_build_model = "<unset>"
reopened = mgr.open(ws_id)
assert reopened is not None
assert adapter.last_build_model is None
def test_open_touches_workstream_on_rehydrate() -> None:
"""Rehydrating a workstream must bump its ``updated`` so a concurrent
close_idle pass-2 in this same process can't clobber the freshly-loaded
@@ -1368,3 +1494,127 @@ class TestSessionManagerWithStateWriter:
assert "running" not in ws_writes, (
f"set_state after close enqueued through buffer: {ws_writes}"
)
# ---------------------------------------------------------------------------
# Multi-subscriber observer — subscribe_to_state / unsubscribe_from_state
# ---------------------------------------------------------------------------
class TestStateSubscribers:
"""Multi-subscriber observer for ``set_state``.
Used by the CLI's background-attention notifier and by
``SameNodeChildSource``. Subscribe / unsubscribe must be safe under
concurrent dispatch, and dispatch must not skip / repeat callbacks
when subscribers register or unregister mid-iteration.
"""
def test_subscribe_fires_on_set_state(self) -> None:
mgr, _, _ = _make_manager()
ws = mgr.create(user_id="u1", name="ws", skill=None)
events: list[tuple[str, str]] = []
def cb(ws_id: str, state: WorkstreamState) -> None:
events.append((ws_id, state.value))
mgr.subscribe_to_state(cb)
mgr.set_state(ws.id, WorkstreamState.RUNNING)
assert events == [(ws.id, "running")]
def test_unsubscribe_stops_firing(self) -> None:
mgr, _, _ = _make_manager()
ws = mgr.create(user_id="u1", name="ws", skill=None)
events: list[str] = []
def cb(_ws_id: str, state: WorkstreamState) -> None:
events.append(state.value)
mgr.subscribe_to_state(cb)
mgr.unsubscribe_from_state(cb)
mgr.set_state(ws.id, WorkstreamState.RUNNING)
assert events == []
def test_unsubscribe_unknown_is_noop(self) -> None:
mgr, _, _ = _make_manager()
# Doesn't raise.
mgr.unsubscribe_from_state(lambda *_: None)
def test_multiple_subscribers_fire_in_registration_order(self) -> None:
mgr, _, _ = _make_manager()
ws = mgr.create(user_id="u1", name="ws", skill=None)
order: list[int] = []
def make(i: int) -> Callable[[str, WorkstreamState], None]:
def cb(_ws_id: str, _state: WorkstreamState) -> None:
order.append(i)
return cb
mgr.subscribe_to_state(make(1))
mgr.subscribe_to_state(make(2))
mgr.subscribe_to_state(make(3))
mgr.set_state(ws.id, WorkstreamState.IDLE)
assert order == [1, 2, 3]
def test_subscriber_exception_does_not_block_others(self) -> None:
mgr, _, _ = _make_manager()
ws = mgr.create(user_id="u1", name="ws", skill=None)
survived: list[str] = []
def boom(*_: Any) -> None:
raise RuntimeError("subscriber crash")
def good(_ws_id: str, state: WorkstreamState) -> None:
survived.append(state.value)
mgr.subscribe_to_state(boom)
mgr.subscribe_to_state(good)
mgr.set_state(ws.id, WorkstreamState.RUNNING)
assert survived == ["running"]
@pytest.mark.parametrize("n_threads", [10, 50])
def test_concurrent_subscribe(self, n_threads: int) -> None:
"""Subscribe from many threads; all callbacks land in the list.
Validates the lock around mutation without it the underlying
list.append could lose entries under contention.
"""
mgr, _, _ = _make_manager()
callbacks = [lambda *_, i=i: None for i in range(n_threads)]
threads = [threading.Thread(target=mgr.subscribe_to_state, args=(cb,)) for cb in callbacks]
for t in threads:
t.start()
for t in threads:
t.join()
# Snapshot under the lock to read the count safely.
with mgr._state_subscribers_lock:
assert len(mgr._state_subscribers) == n_threads
def test_subscribe_during_dispatch_does_not_corrupt_iteration(self) -> None:
"""A subscriber that calls subscribe_to_state during its own
callback must not affect the in-flight dispatch (snapshot
isolation). This is the bug-1 invariant: mutation during
iteration can't shift the iterator's index because dispatch
iterates a snapshot, not the live list.
"""
mgr, _, _ = _make_manager()
ws = mgr.create(user_id="u1", name="ws", skill=None)
fired: list[str] = []
def late(_ws_id: str, state: WorkstreamState) -> None:
fired.append("late:" + state.value)
def first(_ws_id: str, state: WorkstreamState) -> None:
fired.append("first:" + state.value)
mgr.subscribe_to_state(late) # mid-dispatch addition
mgr.subscribe_to_state(first)
mgr.set_state(ws.id, WorkstreamState.RUNNING)
# ``late`` was added during dispatch but the snapshot was
# already frozen — so it doesn't fire on this round.
assert fired == ["first:running"]
# Next round it does fire, in registration order after first.
fired.clear()
mgr.set_state(ws.id, WorkstreamState.IDLE)
assert fired == ["first:idle", "late:idle"]
+108 -5
View File
@@ -523,8 +523,15 @@ class TestWorkstreamConfig:
assert session.instructions == "be concise"
assert session.creative_mode is True
def test_resume_restores_model(self, tmp_db):
"""ChatSession.resume() should restore the model from workstream config."""
def test_resume_keeps_defaults_when_alias_unresolvable(self, tmp_db):
"""When the saved alias is empty or no longer in the registry,
``resume()`` must NOT copy ``saved_model`` onto the constructor's
default provider. Pairing a removed model name with a default
provider that doesn't know about it produces a broken session
whose next API call fails the exact regression Copilot flagged
on PR #465. The constructor already resolved a coherent default
(provider + model + capabilities); resume should leave it intact
and just log the unreachable saved values."""
client = MagicMock()
client.models.list.return_value.data = [MagicMock(id="test-model")]
ui = MagicMock()
@@ -533,13 +540,14 @@ class TestWorkstreamConfig:
ui.on_state_change = MagicMock()
ui.on_rename = MagicMock()
# Create a workstream that was using a specific model
register_workstream("model_ws")
save_message("model_ws", "user", "hello")
save_message("model_ws", "assistant", "hi")
# Empty alias + an orphan model name — same shape resume sees
# when an operator removes an alias from the registry that the
# workstream was originally pinned to.
save_workstream_config("model_ws", {"model": "gpt-5", "model_alias": ""})
# Resume into a session that was created with a different model
session = ChatSession(
client=client,
model="gpt-5-nano",
@@ -552,7 +560,102 @@ class TestWorkstreamConfig:
assert session.model == "gpt-5-nano"
result = session.resume("model_ws")
assert result is True
assert session.model == "gpt-5"
# Constructor's coherent default is preserved — saved orphan
# model name is NOT copied over.
assert session.model == "gpt-5-nano"
def test_init_does_not_clobber_existing_config(self, tmp_db):
"""ChatSession.__init__ must NOT overwrite existing
``workstream_config`` keys when constructing for an already-
persisted ws_id.
This is the fix for the rehydrate bug: ``SessionManager.open()``
builds a ChatSession with the persisted ws_id; the legacy
``__init__`` unconditionally called ``_save_config()`` which is
``INSERT OR REPLACE`` per-key silently resetting model_alias,
temperature, reasoning_effort, max_tokens, skill, creative_mode,
and instructions to the constructor defaults *before*
``resume()`` got a chance to read them back.
"""
client = MagicMock()
client.models.list.return_value.data = [MagicMock(id="test-model")]
ui = MagicMock()
ui.on_info = MagicMock()
ui.on_error = MagicMock()
ui.on_state_change = MagicMock()
ui.on_rename = MagicMock()
register_workstream("rehydrate_ws")
save_workstream_config(
"rehydrate_ws",
{
"model": "gpt-5-pro",
"model_alias": "gpt-5-pro",
"temperature": "0.2",
"reasoning_effort": "high",
"max_tokens": "8192",
"creative_mode": "True",
"instructions": "preserve me",
},
)
ChatSession(
client=client,
model="some-default-model",
ui=ui,
instructions=None,
temperature=0.7,
max_tokens=4096,
tool_timeout=30,
reasoning_effort="medium",
ws_id="rehydrate_ws",
)
loaded = load_workstream_config("rehydrate_ws")
assert loaded["model"] == "gpt-5-pro"
assert loaded["model_alias"] == "gpt-5-pro"
assert loaded["temperature"] == "0.2"
assert loaded["reasoning_effort"] == "high"
assert loaded["max_tokens"] == "8192"
assert loaded["creative_mode"] == "True"
assert loaded["instructions"] == "preserve me"
def test_init_writes_config_on_fresh_create(self, tmp_db):
"""The opposite half of the contract: when no config row exists
yet, ``__init__`` must still persist the constructor's values so
a later resume can find them. This is the path that previously
worked the fix must not break it."""
client = MagicMock()
client.models.list.return_value.data = [MagicMock(id="test-model")]
ui = MagicMock()
ui.on_info = MagicMock()
ui.on_error = MagicMock()
ui.on_state_change = MagicMock()
ui.on_rename = MagicMock()
# No save_workstream_config() before ChatSession() — this is
# the fresh-create path the SessionManager.create() flow takes.
register_workstream("fresh_ws")
assert load_workstream_config("fresh_ws") == {}
ChatSession(
client=client,
model="gpt-5-mini",
ui=ui,
instructions="be terse",
temperature=0.4,
max_tokens=2048,
tool_timeout=30,
reasoning_effort="low",
ws_id="fresh_ws",
)
loaded = load_workstream_config("fresh_ws")
assert loaded["model"] == "gpt-5-mini"
assert loaded["temperature"] == "0.4"
assert loaded["reasoning_effort"] == "low"
assert loaded["max_tokens"] == "2048"
assert loaded["instructions"] == "be terse"
# ── Prune workstreams ─────────────────────────────────────────────────
+36
View File
@@ -67,6 +67,42 @@ class TestSearchStructuredMemories:
assert len(results) >= 1
assert any(r["name"] == "db_host" for r in results)
def test_multiword_or_matches_partial(self, tmp_db):
"""OR-of-terms: memory matching only 1 of 3 query terms is returned."""
save_structured_memory("postgres_config", "host=localhost port=5432")
save_structured_memory("redis_config", "host=redis port=6379")
save_structured_memory("unrelated", "nothing relevant here")
# "postgres missing_word_a missing_word_b": only postgres_config matches "postgres"
results = search_structured_memories("postgres missing_word_a missing_word_b")
names = {r["name"] for r in results}
assert "postgres_config" in names
assert "unrelated" not in names
def test_multiword_or_multiple_partial_matches(self, tmp_db):
"""Multiple memories each matching different terms are all returned."""
save_structured_memory("key_alpha", "alpha content here")
save_structured_memory("key_beta", "beta content here")
save_structured_memory("key_other", "completely different")
results = search_structured_memories("alpha beta")
names = {r["name"] for r in results}
assert "key_alpha" in names
assert "key_beta" in names
assert "key_other" not in names
def test_search_scope_filtering_preserved(self, tmp_db):
"""Search with scope filter only returns memories in that scope."""
save_structured_memory("ws1_fact", "alpha info", scope="workstream", scope_id="ws1")
save_structured_memory("ws2_fact", "alpha info", scope="workstream", scope_id="ws2")
save_structured_memory("global_fact", "alpha info", scope="global")
results = search_structured_memories("alpha", scope="workstream", scope_id="ws1")
names = {r["name"] for r in results}
assert "ws1_fact" in names
assert "ws2_fact" not in names
assert "global_fact" not in names
class TestGetStructuredMemoryByName:
def test_get_existing(self, tmp_db):
+145
View File
@@ -126,3 +126,148 @@ class TestCount:
backend.create_structured_memory("m2", "b", "", "project", "workstream", "ws1", "2")
assert backend.count_structured_memories(scope="global") == 1
assert backend.count_structured_memories(scope="workstream") == 1
class TestSearchOrOfTerms:
"""Verify that multi-word search uses OR-of-terms (any term matches → row included)."""
def test_single_matching_term_in_multi_word_query(self, backend):
"""Memory with content 'apple' found when query is 'apple banana cherry'."""
backend.create_structured_memory("m1", "apple_mem", "", "project", "global", "", "apple")
backend.create_structured_memory("m2", "other_mem", "", "project", "global", "", "grape")
results = backend.search_structured_memories("apple banana cherry")
names = {r["name"] for r in results}
assert "apple_mem" in names # matches "apple" — OR-of-terms keeps it
assert "other_mem" not in names # "grape" matches nothing in the query
def test_partial_overlap_across_memories(self, backend):
"""Each memory matches one of three terms; all three are returned."""
backend.create_structured_memory("m1", "alpha_doc", "", "project", "global", "", "alpha")
backend.create_structured_memory("m2", "beta_doc", "", "project", "global", "", "beta")
backend.create_structured_memory("m3", "gamma_doc", "", "project", "global", "", "gamma")
backend.create_structured_memory("m4", "unrelated", "", "project", "global", "", "delta")
results = backend.search_structured_memories("alpha beta gamma")
names = {r["name"] for r in results}
assert "alpha_doc" in names
assert "beta_doc" in names
assert "gamma_doc" in names
assert "unrelated" not in names # "delta" doesn't appear in the query
def test_scope_filter_preserved(self, backend):
"""OR-of-terms search still respects scope / scope_id filters."""
backend.create_structured_memory(
"m1", "ws1_note", "", "project", "workstream", "ws1", "info"
)
backend.create_structured_memory(
"m2", "ws2_note", "", "project", "workstream", "ws2", "info"
)
backend.create_structured_memory("m3", "global_note", "", "project", "global", "", "info")
results = backend.search_structured_memories("info", scope="workstream", scope_id="ws1")
names = {r["name"] for r in results}
assert "ws1_note" in names
assert "ws2_note" not in names
assert "global_note" not in names
def test_term_cap_normalizes_unbounded_query(self, backend):
"""A multi-KB query collapses to <= MAX terms (de-dupe + length filter)."""
backend.create_structured_memory("m1", "alpha_doc", "", "project", "global", "", "alpha")
backend.create_structured_memory(
"m2", "other_doc", "", "project", "global", "", "irrelevant"
)
# Build a noisy query: same word repeated, plus 1-char tokens that
# the normalizer drops, plus the actual signal "alpha".
noisy = " ".join(["x"] * 100 + ["alpha"] * 50)
results = backend.search_structured_memories(noisy)
names = {r["name"] for r in results}
assert "alpha_doc" in names
class TestVisibleStructuredMemories:
"""Single-query union helpers used by the composition path."""
def test_list_visible_unions_global_workstream_user(self, backend):
backend.create_structured_memory("m1", "g_note", "", "project", "global", "", "g")
backend.create_structured_memory("m2", "ws_note", "", "project", "workstream", "ws1", "w")
backend.create_structured_memory("m3", "u_note", "", "project", "user", "u1", "u")
backend.create_structured_memory("m4", "other_ws", "", "project", "workstream", "ws2", "x")
scopes = [("global", ""), ("workstream", "ws1"), ("user", "u1")]
rows = backend.list_visible_structured_memories(scopes)
names = {r["name"] for r in rows}
assert names == {"g_note", "ws_note", "u_note"} # ws2 excluded
def test_search_visible_unions_scopes_and_terms(self, backend):
backend.create_structured_memory("m1", "g_alpha", "", "project", "global", "", "alpha")
backend.create_structured_memory(
"m2", "ws_beta", "", "project", "workstream", "ws1", "beta"
)
backend.create_structured_memory(
"m3", "ws_other", "", "project", "workstream", "ws2", "alpha"
)
scopes = [("global", ""), ("workstream", "ws1")]
rows = backend.search_visible_structured_memories("alpha beta", scopes)
names = {r["name"] for r in rows}
assert "g_alpha" in names # global, matches "alpha"
assert "ws_beta" in names # ws1, matches "beta"
assert "ws_other" not in names # ws2 -> outside visibility
def test_visible_helpers_handle_empty_scopes(self, backend):
backend.create_structured_memory("m1", "anything", "", "project", "global", "", "x")
assert backend.list_visible_structured_memories([]) == []
assert backend.search_visible_structured_memories("x", []) == []
class TestStableOrderingOnTimestampTies:
"""When two memories share an `updated` timestamp, secondary sort on
memory_id keeps the order deterministic across calls.
`updated` is second-precision, and touch_structured_memories() can bump
a batch to identical timestamps without a tie-breaker BM25 input
order shuffles run-to-run, busting the LLM-side prompt cache.
"""
def _seed_with_shared_timestamp(self, backend):
# Create three memories then force their `updated` columns equal —
# mirrors the real-world case where a touch_structured_memories
# batch lands them in the same second.
for mid in ("zebra_id", "apple_id", "mango_id"):
backend.create_structured_memory(
mid, f"name_{mid}", "", "project", "global", "", "shared content"
)
import sqlalchemy as sa
with backend._conn() as conn:
conn.execute(sa.text("UPDATE structured_memories SET updated = '2024-01-01T00:00:00'"))
conn.commit()
def test_list_stable_order_under_tied_updated(self, backend):
self._seed_with_shared_timestamp(backend)
first = [r["memory_id"] for r in backend.list_structured_memories()]
second = [r["memory_id"] for r in backend.list_structured_memories()]
# Deterministic across calls AND sorted by memory_id ASC for ties
assert first == second
assert first == ["apple_id", "mango_id", "zebra_id"]
def test_search_stable_order_under_tied_updated(self, backend):
self._seed_with_shared_timestamp(backend)
first = [r["memory_id"] for r in backend.search_structured_memories("shared")]
second = [r["memory_id"] for r in backend.search_structured_memories("shared")]
assert first == second
assert first == ["apple_id", "mango_id", "zebra_id"]
def test_visible_search_stable_order_under_tied_updated(self, backend):
self._seed_with_shared_timestamp(backend)
scopes = [("global", "")]
first = [
r["memory_id"] for r in backend.search_visible_structured_memories("shared", scopes)
]
second = [
r["memory_id"] for r in backend.search_visible_structured_memories("shared", scopes)
]
assert first == second
assert first == ["apple_id", "mango_id", "zebra_id"]
+129 -37
View File
@@ -157,20 +157,17 @@ class TestContentAccumulation:
assert len(idle_events[0]["content"]) <= _MAX_TURN_CONTENT_CHARS + 1024
class TestPendingApprovalDetailGate:
"""The Shape A SSE plumbing carries ``pending_approval_detail`` on the
``ws_state`` event so the coord tree UI can render inline approve/deny
buttons in lockstep with the activity_state transition. The gate
(``if self._pending_approval is not None``) keeps the per-broadcast
serializer cost off the common no-approval-pending path these tests
lock both branches down."""
class TestPendingApprovalDetailNotPiggybacked:
"""Stage 3 cleanup — ``pending_approval_detail`` is no longer
piggybacked on ``ws_state`` events. Approval items now arrive via
bulk fetch when the coord tree's reducer sees the
``activity_state="approval"`` transition; verdicts via the explicit
``intent_verdict`` event class; resolution via
``approval_resolved``. These tests lock the no-piggyback contract
down so a future regression doesn't silently re-introduce the
duplicated path."""
def test_state_broadcast_omits_field_when_no_approval_pending(self):
"""Common case: no approval pending → field absent from event so the
per-broadcast verdict-cache deepcopy in
``serialize_pending_approval_detail`` never runs. A regression
that drops the gate would silently 10x the cost of every state
broadcast in the steady state."""
ui = _make_ui()
assert ui._pending_approval is None
ui._broadcast_state("running")
@@ -180,17 +177,12 @@ class TestPendingApprovalDetailGate:
assert len(running_events) == 1
assert "pending_approval_detail" not in running_events[0]
def test_state_broadcast_includes_field_when_approval_pending(self):
"""When an approval is pending the broadcast must carry the rich
payload the coord tree UI reads it directly to render inline
approve/deny buttons. Without this, a coord browser would have
to chase a separate ``cluster/ws/live`` fetch on every
activity_state transition (the load-storm pattern Shape A is
unwinding)."""
def test_state_broadcast_omits_field_even_when_approval_pending(self):
"""The piggyback is gone: even when ``_pending_approval`` is set,
the state broadcast must NOT carry ``pending_approval_detail``.
The browser triggers a bulk fetch off the
``activity_state="approval"`` transition to get the items."""
ui = _make_ui()
# Mirror the shape ``pause_for_approval`` writes (session_ui_base
# lines 576-580) — items with call_id + header is the minimum
# the serializer needs to project.
ui._pending_approval = {
"type": "approve_request",
"items": [
@@ -209,21 +201,9 @@ class TestPendingApprovalDetailGate:
events = _drain_global()
attn = [e for e in events if e.get("state") == "attention"]
assert len(attn) == 1
# Field present and structurally sound — the serializer's
# full shape is covered by tests/test_session_ui_base.py;
# here we only need to confirm the gate fires and the
# serializer's output is what lands on the event.
assert "pending_approval_detail" in attn[0]
detail = attn[0]["pending_approval_detail"]
assert detail is not None
assert detail.get("items")
assert detail["items"][0]["call_id"] == "c1"
assert "pending_approval_detail" not in attn[0]
def test_field_cleared_after_approval_resolves(self):
"""Once ``_pending_approval`` is cleared, subsequent state
broadcasts must drop the field again without this, the
browser would render stale approve/deny buttons until the
next bulk-poll TTL window expired."""
def test_field_stays_absent_after_approval_resolves(self):
ui = _make_ui()
ui._pending_approval = {
"type": "approve_request",
@@ -231,7 +211,7 @@ class TestPendingApprovalDetailGate:
"judge_pending": False,
}
ui._broadcast_state("attention")
_drain_global() # discard the with-detail event
_drain_global()
ui._pending_approval = None
ui._broadcast_state("running")
@@ -239,3 +219,115 @@ class TestPendingApprovalDetailGate:
running = [e for e in events if e.get("state") == "running"]
assert len(running) == 1
assert "pending_approval_detail" not in running[0]
class TestBroadcastIntentVerdict:
"""Producer-side coverage for ``WebUI._broadcast_intent_verdict``.
The collector-side test (``test_apply_delta_intent_verdict_*`` in
test_console.py) covers consumption; this pins the event shape the
producer puts on the global queue. A field rename or missed key
here would slip past the consumer test because the consumer reads
via ``data.get(...)``.
"""
def test_pushes_intent_verdict_event_to_global_queue(self):
ui = _make_ui()
verdict = {
"call_id": "c1",
"risk_level": "low",
"confidence": 0.92,
"recommendation": "approve",
"reasoning": "tool reads only",
}
ui._broadcast_intent_verdict(verdict)
events = _drain_global()
assert len(events) == 1
ev = events[0]
assert ev["type"] == "intent_verdict"
assert ev["ws_id"] == "ws-test"
assert ev["verdict"] == verdict
def test_no_op_when_global_queue_unset(self):
WebUI._global_queue = None
ui = _make_ui()
# Doesn't raise.
ui._broadcast_intent_verdict({"call_id": "c1"})
def test_queue_full_swallowed(self):
# Force a tiny queue then fill it so the next put_nowait
# raises queue.Full — the broadcast must absorb it without
# propagating (matches _broadcast_state's queue.Full handling).
WebUI._global_queue = queue.Queue(maxsize=1)
WebUI._global_queue.put_nowait({"sentinel": True})
ui = _make_ui()
# Doesn't raise.
ui._broadcast_intent_verdict({"call_id": "c1"})
class TestBroadcastApprovalResolved:
"""Producer-side coverage for ``WebUI._broadcast_approval_resolved``."""
def test_pushes_approval_resolved_event_to_global_queue(self):
ui = _make_ui()
ui._broadcast_approval_resolved(True, "lgtm", always=False)
events = _drain_global()
assert len(events) == 1
ev = events[0]
assert ev["type"] == "approval_resolved"
assert ev["ws_id"] == "ws-test"
assert ev["approved"] is True
assert ev["feedback"] == "lgtm"
assert ev["always"] is False
def test_normalises_none_feedback_to_empty_string(self):
ui = _make_ui()
ui._broadcast_approval_resolved(False, None)
events = _drain_global()
assert events[0]["feedback"] == ""
assert events[0]["approved"] is False
assert events[0]["always"] is False
def test_always_kwarg_propagates(self):
ui = _make_ui()
ui._broadcast_approval_resolved(True, "ok", always=True)
events = _drain_global()
assert events[0]["always"] is True
def test_no_op_when_global_queue_unset(self):
WebUI._global_queue = None
ui = _make_ui()
# Doesn't raise.
ui._broadcast_approval_resolved(True, None)
class TestBroadcastApproveRequest:
"""Producer-side coverage for ``WebUI._broadcast_approve_request`` —
push path for the initial approval items so a coord parent's tree
UI can render the inline approve/deny block immediately without
waiting for a bulk-fetch round-trip."""
def test_pushes_approve_request_event_to_global_queue(self):
ui = _make_ui()
detail = {
"type": "approve_request",
"items": [{"call_id": "c1", "header": "tool x"}],
"judge_pending": True,
}
ui._broadcast_approve_request(detail)
events = _drain_global()
assert len(events) == 1
ev = events[0]
assert ev["type"] == "approve_request"
assert ev["ws_id"] == "ws-test"
assert ev["detail"] == detail
def test_no_op_when_global_queue_unset(self):
WebUI._global_queue = None
ui = _make_ui()
# Doesn't raise.
ui._broadcast_approve_request({"items": []})
+1 -1
View File
@@ -1,3 +1,3 @@
"""turnstone - Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation."""
__version__ = "1.5.3"
__version__ = "1.5.6"
+1 -1
View File
@@ -1242,7 +1242,7 @@ def main() -> None:
)
sys.stderr.flush()
manager._on_state_change = _bg_attention_notify
manager.subscribe_to_state(_bg_attention_notify)
# Print banner
print(f"\n{bold('Chat')} with {cyan(model)}")
+132 -14
View File
@@ -498,7 +498,6 @@ class ClusterCollector:
"kind": WorkstreamKind.from_raw(new_w.get("kind")),
"parent_ws_id": new_w.get("parent_ws_id"),
"activity_state": new_w.get("activity_state", ""),
"pending_approval_detail": new_w.get("pending_approval_detail"),
}
)
old_name = old_ws.get("title", "") or old_ws.get("name", "")
@@ -558,18 +557,6 @@ class ClusterCollector:
ws["kind"] = data["kind"]
if "parent_ws_id" in data:
ws["parent_ws_id"] = data["parent_ws_id"]
# ``pending_approval_detail`` overwrites (no
# ``ws.get`` fallback): the node's broadcast gate
# on ``_pending_approval is not None`` means the
# field is absent from ``data`` exactly when no
# approval is pending — falling back to the cached
# value would resurrect a stale detail after the
# approval resolved. Without this assignment the
# cached ``node.workstreams`` dict served by
# ``get_node_detail`` / ``get_snapshot`` between
# reconciliations would render stale approve/deny
# buttons on closed approvals.
ws["pending_approval_detail"] = data.get("pending_approval_detail")
pending_events.append(
{
"type": "cluster_state",
@@ -581,7 +568,6 @@ class ClusterCollector:
"kind": WorkstreamKind.from_raw(ws.get("kind")),
"parent_ws_id": ws.get("parent_ws_id"),
"activity_state": ws.get("activity_state", ""),
"pending_approval_detail": data.get("pending_approval_detail"),
}
)
@@ -648,6 +634,60 @@ class ClusterCollector:
ws["name"] = name
pending_events.append({"type": "ws_rename", "ws_id": ws_id, "name": name})
elif etype == "intent_verdict":
# Pure pass-through for LLM judge verdicts. The node's
# WebUI._broadcast_intent_verdict pushes onto the
# global queue (Stage 3 Step 5); the cluster fan-out
# forwards to subscribers (e.g. CoordinatorAdapter,
# which dispatches to the parent's coord SSE as
# ``child_ws_intent_verdict``).
ws_id = data.get("ws_id", "")
if ws_id:
pending_events.append(
{
"type": "intent_verdict",
"ws_id": ws_id,
"node_id": node_id,
"verdict": data.get("verdict") or {},
}
)
elif etype == "approval_resolved":
# Pure pass-through for approve/deny resolutions. Pairs
# with ``intent_verdict`` above so the parent
# coordinator's tree UI can clear the pending-approval
# pill the moment the user decides, rather than waiting
# for the subsequent state-change piggyback.
ws_id = data.get("ws_id", "")
if ws_id:
pending_events.append(
{
"type": "approval_resolved",
"ws_id": ws_id,
"node_id": node_id,
"approved": bool(data.get("approved", False)),
"feedback": data.get("feedback", "") or "",
"always": bool(data.get("always", False)),
}
)
elif etype == "approve_request":
# Push path for the initial approval items. Eliminates
# the bulk-fetch race that otherwise leaves the coord
# tree stuck on a loading placeholder when the bulk
# fetch lands in the gap between the state transition
# to ATTENTION and ``_pending_approval`` being set.
ws_id = data.get("ws_id", "")
if ws_id:
pending_events.append(
{
"type": "approve_request",
"ws_id": ws_id,
"node_id": node_id,
"detail": data.get("detail") or {},
}
)
elif etype == "health_changed":
# Update the health dict's backend status in-place
bstatus = data.get("backend_status", "")
@@ -1217,3 +1257,81 @@ class ClusterCollector:
return
entry["name"] = name
self._fanout({"type": "ws_rename", "ws_id": ws_id, "name": name})
def emit_models_changed(self) -> None:
"""Fan out a ``models_changed`` notice to all SSE listeners.
Browsers re-fetch :http:get:`/v1/api/models` on receipt so the
coordinator-composer model dropdown + the admin Models tab
reflect alias / underlying-model edits without a manual reload.
Body is intentionally empty listeners refetch authoritative
state rather than diffing the event payload.
"""
self._fanout({"type": "models_changed"})
def emit_console_ws_intent_verdict(self, ws_id: str, verdict: dict[str, Any]) -> None:
"""Fan an LLM intent-judge verdict for a console-pseudo-node ws.
Gives coord-spawned approval flows a first-class cluster-bus
event so the parent's tree UI can render the risk pill +
verdict result without polling.
``CoordinatorAdapter._dispatch_child_event`` re-emits these as
``child_ws_intent_verdict`` for the parent coordinator's SSE
stream. The dispatch path filters by registry membership;
downstream subscribers tolerate verdicts for ws_ids they don't
own (silently drop), so we skip the membership pre-check that
would otherwise add a lock acquisition per emit on a path
that fires once per tool-call during heuristic+LLM judging.
"""
self._fanout(
{
"type": "intent_verdict",
"ws_id": ws_id,
"node_id": self.CONSOLE_PSEUDO_NODE_ID,
"verdict": verdict,
}
)
def emit_console_ws_approval_resolved(
self,
ws_id: str,
*,
approved: bool,
feedback: str = "",
always: bool = False,
) -> None:
"""Fan an ``approval_resolved`` decision for a console-pseudo-node ws.
Paired with :meth:`emit_console_ws_intent_verdict` so the
coord tree UI clears the pending-approval pill in lockstep
with the actual decision. Same lock-skip rationale as the
intent-verdict emit above.
"""
self._fanout(
{
"type": "approval_resolved",
"ws_id": ws_id,
"node_id": self.CONSOLE_PSEUDO_NODE_ID,
"approved": approved,
"feedback": feedback,
"always": always,
}
)
def emit_console_ws_approve_request(self, ws_id: str, detail: dict[str, Any]) -> None:
"""Fan an ``approve_request`` payload for a console-pseudo-node ws.
Push path for the initial approval items so a coord parent's
tree UI can render the inline approve/deny block immediately
without a bulk-fetch round-trip. ``CoordinatorAdapter._dispatch_child_event``
re-emits as ``child_ws_approve_request`` for the parent
coordinator's SSE stream.
"""
self._fanout(
{
"type": "approve_request",
"ws_id": ws_id,
"node_id": self.CONSOLE_PSEUDO_NODE_ID,
"detail": detail,
}
)
+156 -247
View File
@@ -15,17 +15,17 @@ for the storage-seeded children rebuild.
from __future__ import annotations
import queue
import threading
from typing import TYPE_CHECKING, Any
from turnstone.core import session_worker
from turnstone.core.adapters._ui_cleanup import cleanup_session_ui
from turnstone.core.child_source import ClusterChildSource
from turnstone.core.children_registry import ChildrenRegistry
from turnstone.core.log import get_logger
from turnstone.core.workstream import Workstream, WorkstreamKind, WorkstreamState
if TYPE_CHECKING:
from collections.abc import Callable, Iterable
from collections.abc import Callable
from turnstone.console.collector import ClusterCollector
from turnstone.console.coordinator_ui import ConsoleCoordinatorUI
@@ -56,38 +56,25 @@ class CoordinatorAdapter:
# the manager to the adapter's ``__init__`` — break the cycle
# with a setter called from console startup.
self._manager: SessionManager | None = None
# Per-coordinator known-child ws_id set. Populated lazily on
# create/open from storage and updated live as the cluster fan-out
# thread sees ws_created events with matching parent_ws_id.
# Closed / deleted children stay in the registry so the tree UI
# can keep rendering them grayed out; the authoritative render
# path reads storage for state. Bounded by eventual coordinator
# close/eviction.
self._children: dict[str, set[str]] = {}
# Reverse index for O(1) child → coord lookup on every cluster
# event. Without this, every cluster event incurs a linear scan
# over every coordinator's child set while holding the fan-out
# lock — a hot-path tax that scales with both active coordinators
# and their retained-history depth.
self._child_to_coord: dict[str, str] = {}
self._children_lock = threading.Lock()
# Cluster-event fan-out: subscribes to the ClusterCollector's
# listener channel, filters by known child ws_ids, and re-emits
# child_ws_* events on the matching coordinator's UI. Configured
# lazily via ``start_child_event_fanout(collector)`` from the
# console lifespan once both the manager and collector exist.
self._collector_queue: queue.Queue[dict[str, Any]] | None = None
self._fanout_thread: threading.Thread | None = None
self._fanout_stop = threading.Event()
# Coord-ws-id → UI map for the fan-out dispatch path. Read
# and written under ``self._children_lock`` alongside the
# forward/reverse child maps. (Previously the value was
# ``(user_id, ui)`` with a copy-on-write dict swap so the
# dispatch could read it lock-free — but _dispatch_child_event
# already re-validates the parent under _children_lock anyway,
# so the lock-free snapshot was premature. The user_id half is
# also dead after a46dab1 dropped row-level ownership gates.)
self._active_coords: dict[str, Any] = {}
# Children registry — universal parent → children + reverse
# lookup primitive. Lifted from inline data on this adapter to
# ``turnstone.core.children_registry.ChildrenRegistry`` (Stage 3
# Step 1) so the same primitive can serve interactive
# workstreams when they gain spawn capability. The dispatch
# path (`_dispatch_child_event`) calls into the registry
# atomically; everything else delegates through the legacy
# method names which still exist as thin shims for the
# cluster-routing + cleanup callers.
self._registry = ChildrenRegistry()
# Cross-node child events arrive via ``ClusterChildSource``
# (Stage 3 Step 2): a strategy that subscribes to the
# collector's listener channel and runs a daemon thread that
# drains the queue, pushing each event to the sink. The sink
# is :meth:`_dispatch_child_event` so the existing translation
# logic (cluster_state → child_ws_state, etc.) stays in one
# place. Constructed lazily by ``start_child_event_fanout`` so
# the collector reference is available.
self._child_source: ClusterChildSource | None = None
def attach(self, manager: SessionManager) -> None:
"""Late-bind the owning :class:`SessionManager`.
@@ -109,7 +96,7 @@ class CoordinatorAdapter:
# needs the empty forward/presence entries so
# ``_dispatch_child_event`` recognises this coordinator when
# its first child is spawned.
self._install_coord_registry(ws)
self._registry.install(ws.id, ws.ui)
self._fanout_console_ws_created(ws)
def emit_rehydrated(self, ws: Workstream) -> None:
@@ -117,21 +104,10 @@ class CoordinatorAdapter:
# it from storage after the registry seed so a ``ws_created``
# for an already-spawned child that fires mid-rebuild merges
# cleanly.
self._install_coord_registry(ws)
self._registry.install(ws.id, ws.ui)
self._rebuild_children_registry(ws.id)
self._fanout_console_ws_created(ws)
def _install_coord_registry(self, ws: Workstream) -> None:
"""Seed the children registry + presence map for ``ws``.
Shared by ``emit_created`` and ``emit_rehydrated`` the
difference between the two is purely whether we then rebuild
from storage.
"""
with self._children_lock:
self._children.setdefault(ws.id, set())
self._active_coords[ws.id] = ws.ui
def _fanout_console_ws_created(self, ws: Workstream) -> None:
try:
self._collector.emit_console_ws_created(
@@ -197,14 +173,11 @@ class CoordinatorAdapter:
# toast only fire for real-node (interactive) ws_closed events.
del reason, name
# Drop the coordinator's children-registry entries AND its
# presence slot. Mirrors the eviction/close paths from the old
# CoordinatorManager (which did the same under _children_lock
# + _lock respectively). A plain _children.pop without clearing
# the reverse index would leak every evicted coordinator's
# child→parent pointers forever.
with self._children_lock:
self._pop_coord_registry_locked(ws_id)
self._active_coords.pop(ws_id, None)
# presence slot. A plain pop without clearing the reverse
# index would leak every evicted coordinator's child→parent
# pointers forever — :meth:`ChildrenRegistry.uninstall`
# handles forward set + reverse entries + presence atomically.
self._registry.uninstall(ws_id)
try:
self._collector.emit_console_ws_closed(ws_id)
except Exception:
@@ -356,25 +329,6 @@ class CoordinatorAdapter:
# Children registry
# ------------------------------------------------------------------
def _merge_child_ids_locked(self, coord_ws_id: str, child_ids: Iterable[str]) -> None:
"""Merge ``child_ids`` into ``coord_ws_id``'s forward + reverse maps.
Caller MUST hold ``self._children_lock``. Idempotent re-adding
an existing child is a no-op (the reverse-index pointer is
already correct). Empty / falsy entries in ``child_ids`` are
skipped.
Sole write-path for bulk registry updates so
``_rebuild_children_registry`` (storage-seeded) and
``_prime_children_from_snapshot`` (collector-seeded) agree on
ordering and reverse-index invariants.
"""
existing = self._children.setdefault(coord_ws_id, set())
for cid in child_ids:
if cid and cid not in existing:
existing.add(cid)
self._child_to_coord[cid] = coord_ws_id
def _rebuild_children_registry(self, coord_ws_id: str) -> None:
"""Populate ``self._children[coord_ws_id]`` from storage.
@@ -445,174 +399,88 @@ class CoordinatorAdapter:
if not child_id:
continue
child_ids.append(child_id)
with self._children_lock:
self._merge_child_ids_locked(coord_ws_id, child_ids)
def _coord_for_child(self, child_ws_id: str) -> str | None:
"""Reverse-lookup: which coordinator owns this child ws_id?
O(1) via the ``_child_to_coord`` reverse index. Cluster events
fire on every token tick across the cluster; a linear scan here
turned into a hot-path tax as the retained-history set grew.
"""
with self._children_lock:
return self._child_to_coord.get(child_ws_id)
self._registry.merge_children(coord_ws_id, child_ids)
def children_snapshot(self, coord_ws_id: str) -> list[str]:
"""Return a snapshot of the coordinator's direct child ws_ids.
Used by ``stop_cascade`` to iterate children without holding the
registry lock during the per-child HTTP dispatch. A mutation
racing with the snapshot (child spawned mid-cascade) either
lands before the snapshot and gets cancelled, or lands after
and is out of scope for this batch both outcomes are safe.
Returns an empty list for unknown coordinators.
Used by ``stop_cascade`` to iterate children without holding
the registry lock during the per-child HTTP dispatch. A
mutation racing with the snapshot (child spawned mid-cascade)
either lands before (cancelled) or after (out of scope for
this batch) both safe. Returns an empty list for unknown
coordinators.
"""
with self._children_lock:
child_set = self._children.get(coord_ws_id)
return list(child_set) if child_set else []
def _pop_coord_registry_locked(self, coord_ws_id: str) -> None:
"""Remove a coordinator's forward set + reverse-index entries.
Caller MUST hold ``self._children_lock``. Used by close /
eviction paths so stale coordinators don't leak registry
entries. No-op if the coordinator is unknown.
"""
child_set = self._children.pop(coord_ws_id, None)
if child_set is None:
return
for cid in child_set:
# Defensive: only clear the reverse entry if it still points
# at THIS coordinator. If a child has since been reassigned
# (unusual but possible on schema changes), we don't want to
# orphan the new owner's entry.
if self._child_to_coord.get(cid) == coord_ws_id:
self._child_to_coord.pop(cid, None)
return self._registry.children_of(coord_ws_id)
# ------------------------------------------------------------------
# Cluster-event fan-out thread
# ------------------------------------------------------------------
def start_child_event_fanout(self, collector: ClusterCollector) -> None:
"""Subscribe to cluster events and start the filter + re-emit thread.
"""Subscribe to cluster events via :class:`ClusterChildSource`.
Idempotent calling twice is a no-op (already-started fan-out
thread stays). Called once from the console lifespan after both
the collector and the session manager are constructed.
Idempotent already-started ChildSource stays. Called once
from the console lifespan after both the collector and the
session manager are constructed.
"""
if self._fanout_thread is not None and self._fanout_thread.is_alive():
if self._child_source is not None:
return
self._collector = collector
self._collector_queue = queue.Queue(maxsize=1000)
# Ensure the "console" pseudo-node exists in the snapshot map so
# emit_console_ws_* calls from create / close / open land on a
# real node entry the snapshot will surface.
# Coord-specific transport setup: the "console" pseudo-node
# must exist in the snapshot map BEFORE any
# ``emit_console_ws_*`` calls (and before the snapshot is
# taken inside ``ChildSource.start``) so those emits land on
# a real node entry the snapshot surfaces.
collector.ensure_console_pseudo_node()
# Register with the collector — use the existing listener channel
# the browser SSE fan-out uses; the collector treats our queue as
# just another subscriber.
snapshot = collector.get_snapshot_and_register(self._collector_queue)
# Prime the child registry from the snapshot so a coordinator
# that opens right after a console restart sees already-live
# children without waiting for the next ``ws_state`` tick to
# discover them via the fan-out path.
self._prime_children_from_snapshot(snapshot)
# Seed the pseudo-node with any coordinators already loaded in
# memory when the collector binds. Prevents a race where early
# creates happened before the collector was wired up and their
# rows never showed on the snapshot.
mgr = self._manager
if mgr is not None:
for ws in mgr.list_all():
try:
collector.emit_console_ws_created(
ws.id,
name=ws.name,
user_id=ws.user_id or "",
kind=WorkstreamKind.COORDINATOR.value,
state=ws.state.value,
parent_ws_id=None,
)
except Exception:
log.debug(
"coord_adapter.collector_seed_failed ws=%s",
ws.id[:8],
exc_info=True,
)
self._fanout_stop.clear()
t = threading.Thread(
target=self._fanout_loop,
name="coord-adapter-child-fanout",
daemon=True,
)
self._fanout_thread = t
t.start()
def _prime_children_from_snapshot(self, snapshot: dict[str, Any]) -> None:
"""Populate ``_children`` + ``_child_to_coord`` from a collector snapshot.
The snapshot's per-node workstreams carry ``parent_ws_id``. For
every workstream whose parent is an in-memory coordinator,
record the child so the fan-out filter sees it immediately.
"""
nodes = snapshot.get("nodes", []) if isinstance(snapshot, dict) else []
if not nodes:
return
mgr = self._manager
if mgr is None:
raise RuntimeError(
"CoordinatorAdapter: manager not attached — call attach(mgr) after construction"
)
by_parent: dict[str, list[str]] = {}
known = {ws.id for ws in mgr.list_all()}
for node in nodes:
for entry in node.get("workstreams", []) or []:
parent = entry.get("parent_ws_id") or ""
child_id = entry.get("id") or ""
if not parent or not child_id or parent not in known:
continue
by_parent.setdefault(parent, []).append(child_id)
if not by_parent:
return
with self._children_lock:
for parent, kids in by_parent.items():
self._merge_child_ids_locked(parent, kids)
# Build + start the strategy. The sink is
# :meth:`_dispatch_child_event` so cluster events flow through
# the same translation path that synthesises ``child_ws_*``
# payloads for the parent's UI.
source = ClusterChildSource(
collector=collector,
registry=self._registry,
parents_provider=lambda: [ws.id for ws in mgr.list_all()],
)
source.start(sink=self._dispatch_child_event)
self._child_source = source
# Seed the pseudo-node with any coordinators already loaded in
# memory when the collector binds. Prevents a race where early
# creates happened before the collector was wired up and their
# rows never showed on the snapshot. (Coord-specific — interactive
# has no analogous pseudo-node.)
for ws in mgr.list_all():
try:
collector.emit_console_ws_created(
ws.id,
name=ws.name,
user_id=ws.user_id or "",
kind=WorkstreamKind.COORDINATOR.value,
state=ws.state.value,
parent_ws_id=None,
)
except Exception:
log.debug(
"coord_adapter.collector_seed_failed ws=%s",
ws.id[:8],
exc_info=True,
)
def shutdown(self) -> None:
"""Stop the fan-out thread and unregister from the collector.
"""Stop the ChildSource and unregister from the collector.
Safe to call multiple times; idempotent. Invoked from the
console lifespan teardown so SSE listener queues don't leak.
"""
self._fanout_stop.set()
t = self._fanout_thread
q = self._collector_queue
coll = self._collector
self._fanout_thread = None
self._collector_queue = None
if coll is not None and q is not None:
try:
coll.unregister_listener(q)
except Exception:
log.debug("coord_adapter.unregister_listener_failed", exc_info=True)
if t is not None:
t.join(timeout=2.0)
def _fanout_loop(self) -> None:
"""Drain collector events, filter by known children, dispatch."""
q = self._collector_queue
if q is None:
return
while not self._fanout_stop.is_set():
try:
event = q.get(timeout=1.0)
except queue.Empty:
continue
try:
self._dispatch_child_event(event)
except Exception:
log.debug("coord_adapter.fanout.dispatch_failed", exc_info=True)
src = self._child_source
self._child_source = None
if src is not None:
src.shutdown()
def _dispatch_child_event(self, event: dict[str, Any]) -> None:
"""Match a cluster event to a coordinator and re-emit on its UI.
@@ -622,10 +490,12 @@ class CoordinatorAdapter:
- ``ws_created`` with ``parent_ws_id`` matching an in-memory
coordinator add to registry + re-emit as
``child_ws_created``.
- ``cluster_state`` / ``ws_closed`` / ``ws_rename`` whose
``ws_id`` is in any coordinator's known-children registry →
re-emit as ``child_ws_state`` / ``child_ws_closed`` /
``child_ws_rename``.
- ``cluster_state`` / ``ws_closed`` / ``ws_rename`` /
``intent_verdict`` / ``approval_resolved`` whose ``ws_id``
is in any coordinator's known-children registry → re-emit as
``child_ws_state`` / ``child_ws_closed`` /
``child_ws_rename`` / ``child_ws_intent_verdict`` /
``child_ws_approval_resolved``.
Events for ws_ids we don't own silently drop — the filter lives
on the server so each coordinator's SSE stream stays small.
@@ -639,23 +509,17 @@ class CoordinatorAdapter:
parent = event.get("parent_ws_id") or ""
if not parent:
return
# Presence check + registry mutation under the same lock:
# a concurrent close()/eviction can pop the entry between
# the check and the mutation, after which a bare setdefault
# would resurrect the entry — leaking the registry key and
# enqueuing onto the closed coordinator's UI. Trusted-team
# posture (#400 / a46dab1) means no per-event tenant gate
# here; scope-level auth at the SSE endpoint is the only
# boundary.
with self._children_lock:
coord_ui = self._active_coords.get(parent)
if coord_ui is None:
return
existing = self._children.setdefault(parent, set())
if ws_id in existing:
return
existing.add(ws_id)
self._child_to_coord[ws_id] = parent
# Atomic check-and-route under the registry's lock: a
# concurrent close()/eviction can pop the parent's entry
# between the presence check and the mutation, so the two
# must happen together. ``add_child`` returns the parent's
# UI on success or None on (a) parent not installed, or
# (b) duplicate child. Trusted-team posture (#400 / a46dab1)
# means no per-event tenant gate here; scope-level auth at
# the SSE endpoint is the only boundary.
coord_ui = self._registry.add_child(parent, ws_id)
if coord_ui is None:
return
payload = {
"type": "child_ws_created",
"ws_id": ws_id,
@@ -668,8 +532,15 @@ class CoordinatorAdapter:
_enqueue_on_ui(coord_ui, parent, payload)
return
if etype in ("cluster_state", "ws_closed", "ws_rename"):
coord_id = self._coord_for_child(ws_id)
if etype in (
"cluster_state",
"ws_closed",
"ws_rename",
"intent_verdict",
"approval_resolved",
"approve_request",
):
coord_id = self._registry.parent_for(ws_id)
if coord_id is None:
return
mgr = self._manager
@@ -684,14 +555,13 @@ class CoordinatorAdapter:
"state": event.get("state", ""),
"tokens": event.get("tokens", 0),
"node_id": event.get("node_id", ""),
# activity_state lets the JS detect approval-state
# transitions; pending_approval_detail rides on
# the same event so the browser can mutate
# liveBadgeCache directly and render inline
# approve/deny buttons in lockstep with the
# transition, no separate dashboard refetch.
# ``activity_state`` lets the JS detect approval
# transitions and trigger a bulk fetch for the
# initial detail. The detail itself no longer
# piggybacks here (Stage 3 cleanup) — verdicts
# arrive via the explicit ``intent_verdict`` event
# class and resolution via ``approval_resolved``.
"activity_state": event.get("activity_state", ""),
"pending_approval_detail": event.get("pending_approval_detail"),
}
elif etype == "ws_closed":
child_event = {
@@ -700,13 +570,52 @@ class CoordinatorAdapter:
"parent_ws_id": coord_id,
"reason": event.get("reason", ""),
}
else: # ws_rename
elif etype == "ws_rename":
child_event = {
"type": "child_ws_rename",
"child_ws_id": ws_id,
"parent_ws_id": coord_id,
"name": event.get("name", ""),
}
elif etype == "intent_verdict":
# Per-coord re-emit of an explicit verdict event so
# the tree UI can render the risk pill + verdict
# result without polling. The corresponding
# ``cluster_state`` event no longer carries the
# detail (the piggyback was removed end-to-end);
# the bulk fetch on initial approval entry plus this
# explicit event class are the only carriers.
child_event = {
"type": "child_ws_intent_verdict",
"child_ws_id": ws_id,
"parent_ws_id": coord_id,
"node_id": event.get("node_id", ""),
"verdict": event.get("verdict") or {},
}
elif etype == "approval_resolved":
child_event = {
"type": "child_ws_approval_resolved",
"child_ws_id": ws_id,
"parent_ws_id": coord_id,
"node_id": event.get("node_id", ""),
"approved": bool(event.get("approved", False)),
"feedback": event.get("feedback", "") or "",
"always": bool(event.get("always", False)),
}
else: # approve_request
# Push path for the initial approval items —
# eliminates the bulk-fetch race that previously left
# the coord row stuck on a loading placeholder when
# the bulk fetch landed in the gap between the state
# transition to ATTENTION and ``_pending_approval``
# being set inside ``approve_tools``.
child_event = {
"type": "child_ws_approve_request",
"child_ws_id": ws_id,
"parent_ws_id": coord_id,
"node_id": event.get("node_id", ""),
"detail": event.get("detail") or {},
}
_enqueue_on_ui(owning_ws.ui, coord_id, child_event)
+32 -1
View File
@@ -1081,7 +1081,38 @@ class CoordinatorClient:
"value": decoded,
"source": str(r.get("source", "")),
}
nodes.append({"node_id": nid, "metadata": meta})
# Project ``metadata.models`` (a list of
# ``{alias, provider, healthy}`` written by the node's
# heartbeat loop — see ``_collect_node_models_metadata``
# in ``turnstone/server.py``) down to the healthy-alias
# shortlist the coordinator passes back as ``model=`` to
# ``spawn_workstream`` / ``spawn_batch``. The top-level
# field is named ``model_aliases`` (not ``models``) so it
# doesn't collide with ``metadata.models`` — the two
# carry different shapes (list of strings vs list of
# dicts) and a coord that conflates them gets a runtime
# error. Empty list when the node hasn't published a
# models entry yet — older nodes without the heartbeat-
# side projection, or a brand new node mid-startup before
# the first metadata write.
models_entry = meta.get("models", {}).get("value")
healthy_aliases: list[str] = []
if isinstance(models_entry, list):
for row in models_entry:
if not isinstance(row, dict):
continue
if not row.get("healthy", False):
continue
alias = row.get("alias")
if isinstance(alias, str) and alias:
healthy_aliases.append(alias)
nodes.append(
{
"node_id": nid,
"metadata": meta,
"model_aliases": healthy_aliases,
}
)
return {"nodes": nodes, "truncated": truncated}
def list_skills(
+75
View File
@@ -181,6 +181,81 @@ class ConsoleCoordinatorUI(SessionUIBase):
with self._ws_lock:
self._last_broadcast_activity = current
def _broadcast_intent_verdict(self, verdict: dict[str, Any]) -> None:
"""Fan an LLM intent-judge verdict to the cluster collector.
Stage 3 Step 5 overrides the no-op base hook so a coord
workstream that produces its own verdict (rare coord agents
don't typically run the LLM judge) gets a first-class
cluster-bus event for the parent's tree UI. The far more
common path is a CHILD workstream firing its verdict on a
real node; that path goes through ``WebUI._broadcast_intent_verdict``
global SSE collector ``_apply_delta``.
"""
collector = ConsoleCoordinatorUI._collector
if collector is None:
return
try:
collector.emit_console_ws_intent_verdict(self.ws_id, verdict)
except Exception:
log.debug(
"coord_ui.intent_verdict_fanout_failed ws=%s",
self.ws_id,
exc_info=True,
)
def _broadcast_approval_resolved(
self,
approved: bool,
feedback: str | None = None,
*,
always: bool = False,
) -> None:
"""Fan an ``approval_resolved`` decision to the cluster collector.
Stage 3 Step 5 paired with :meth:`_broadcast_intent_verdict`.
Same rationale: coord-direct approvals are rare; the typical
path is a child workstream resolving on its node, with that
node's WebUI broadcasting through the global queue.
"""
collector = ConsoleCoordinatorUI._collector
if collector is None:
return
try:
collector.emit_console_ws_approval_resolved(
self.ws_id,
approved=approved,
feedback=feedback or "",
always=always,
)
except Exception:
log.debug(
"coord_ui.approval_resolved_fanout_failed ws=%s",
self.ws_id,
exc_info=True,
)
def _broadcast_approve_request(self, detail: dict[str, Any]) -> None:
"""Fan an ``approve_request`` payload to the cluster collector.
Push path for the initial approval items. Same rationale as
the other two broadcast hooks: coord-self approvals are rare
(the LLM judge isn't wired on the console coord today), but
the override exists for symmetry and lights up the same path
a future coord-self gate would use.
"""
collector = ConsoleCoordinatorUI._collector
if collector is None:
return
try:
collector.emit_console_ws_approve_request(self.ws_id, detail)
except Exception:
log.debug(
"coord_ui.approve_request_fanout_failed ws=%s",
self.ws_id,
exc_info=True,
)
def on_state_change(self, state: str) -> None:
# Flow state transitions through the unified SessionManager so
# the storage write + adapter emit_state fan-out stay in lockstep
+95 -5
View File
@@ -846,8 +846,20 @@ async def _fetch_live_block(
live = {k: entry.get(k) for k in _CLUSTER_WS_LIVE_KEYS if k in entry}
# Derived field — kept in lockstep with
# _coordinator_live_snapshot so both origins produce the
# same keys.
live["pending_approval"] = live.get("activity_state") == "approval"
# same keys. ``state="attention"`` is the canonical signal;
# ``activity_state="approval"`` is set inside approve_tools
# AFTER the state transition fires, so a bulk fetch that
# races with that window can see state=attention and
# activity_state="" simultaneously. A non-null
# ``pending_approval_detail`` is also a definitive signal
# (the serializer only emits non-None when ``_pending_approval``
# is set on the UI). Any of the three flips this true; the
# frontend reducer mirrors the same disjunction.
live["pending_approval"] = (
live.get("activity_state") == "approval"
or entry.get("state") == "attention"
or live.get("pending_approval_detail") is not None
)
return live
return None
@@ -2401,7 +2413,7 @@ def _require_coord_mgr(request: Request) -> tuple[Any, JSONResponse | None]:
if coord_mgr is None:
registry_err = getattr(request.app.state, "coord_registry_error", "") or ""
msg = "Coordinator subsystem not initialized. " + (
registry_err or "Check coordinator.model_alias and Models tab configuration."
registry_err or "Add a model definition in the admin Models tab."
)
return None, JSONResponse({"error": msg}, status_code=503)
if config_store is None:
@@ -2432,8 +2444,7 @@ def _require_coord_mgr(request: Request) -> tuple[Any, JSONResponse | None]:
{
"error": (
f"{hint} does not resolve: {exc}. "
"Configure a model in the admin Models tab, or set "
"``coordinator.model_alias`` in Settings to an existing alias."
"Add or enable a model in the admin Models tab."
)
},
status_code=503,
@@ -4014,6 +4025,11 @@ async def _lifespan(app: Starlette) -> AsyncGenerator[None, None]:
# the cluster collector's pseudo-node so the
# dashboard tree mirrors child state.
event_emitter=coord_adapter,
# Filter out persisted aliases that no longer resolve
# so a coordinator pinned to a since-removed alias
# still rehydrates (on the registry default) instead
# of 500-ing on every reopen.
model_validator=coord_registry.has_alias,
)
# Late-bind the manager onto the adapter so
# ``_rebuild_children_registry`` / ``send`` /
@@ -6964,6 +6980,35 @@ async def admin_delete_memory(request: Request) -> JSONResponse:
# ---------------------------------------------------------------------------
def _emit_models_changed(request: Request) -> None:
"""Fan a ``models_changed`` SSE notice to connected browsers, if any.
Best-effort: silently no-ops when the collector isn't attached
(e.g. test fixtures that bypass the cluster collector).
"""
collector = getattr(request.app.state, "collector", None)
if collector is not None:
collector.emit_models_changed()
# Settings whose change should refresh the model dropdown / Roles UI in
# every connected browser — covers the global default plus the per-role
# overrides surfaced in the admin Models → Roles sub-tab. Additions
# here are purely additive (e.g. future ``perception.*.model`` keys).
_MODEL_AFFECTING_SETTING_KEYS: frozenset[str] = frozenset(
{
"model.default_alias",
"model.plan_alias",
"model.plan_effort",
"model.task_alias",
"model.task_effort",
"coordinator.model_alias",
"coordinator.reasoning_effort",
"judge.model",
}
)
async def _publish_config_change(request: Request) -> None:
"""Fan out config-reload to all known server nodes (best-effort, async).
@@ -7166,6 +7211,8 @@ async def admin_update_setting(request: Request) -> JSONResponse:
)
await _publish_config_change(request)
if key in _MODEL_AFFECTING_SETTING_KEYS:
_emit_models_changed(request)
return JSONResponse(
{
@@ -7221,6 +7268,8 @@ async def admin_delete_setting(request: Request) -> JSONResponse:
)
await _publish_config_change(request)
if key in _MODEL_AFFECTING_SETTING_KEYS:
_emit_models_changed(request)
return JSONResponse({"status": "ok", "key": key, "default": defn.default})
@@ -8041,6 +8090,33 @@ _MODEL_PROVIDERS = frozenset({"openai", "anthropic", "openai-compatible", "googl
_REASONING_EFFORT_CHOICES = frozenset(
{"", "none", "minimal", "low", "medium", "high", "xhigh", "max"}
)
# Keep in sync with turnstone.core.providers._VALID_API_SURFACES.
_API_SURFACE_CHOICES = frozenset({"chat", "responses"})
def _validate_api_surface(caps: Any) -> str | None:
"""Return an error message if ``caps["server_compat"]["api_surface"]`` is invalid.
Strict equality match (no strip/lower normalisation): the persisted value
is bound directly to the admin ``<select>`` whose options are the canonical
``"chat"`` / ``"responses"`` strings, so anything else fails to round-trip
through edit/save. The provider factory raises ``ValueError`` at request
time for an unknown surface; validating here turns that into a 400 at
write time so an admin can't poison a model alias via direct API calls.
"""
if not isinstance(caps, dict):
return None
sc = caps.get("server_compat")
if not isinstance(sc, dict):
return None
raw = sc.get("api_surface")
if raw is None or raw == "":
return None
if not isinstance(raw, str) or raw not in _API_SURFACE_CHOICES:
return f"Invalid server_compat.api_surface: {raw!r}"
return None
# Keep in sync with turnstone.core.providers._google.GOOGLE_DEFAULT_BASE_URL
_PROVIDER_DEFAULT_URLS: dict[str, str] = {
"openai": "https://api.openai.com/v1",
@@ -8342,6 +8418,9 @@ async def admin_create_model_definition(request: Request) -> JSONResponse:
ctx_raw = body.get("context_window", 32768)
context_window = max(0, int(ctx_raw)) if isinstance(ctx_raw, (int, float)) else 0
caps = body.get("capabilities", {})
err_msg = _validate_api_surface(caps)
if err_msg:
return JSONResponse({"error": err_msg}, status_code=400)
capabilities = json.dumps(caps) if isinstance(caps, dict) else "{}"
enabled = bool(body.get("enabled", True))
@@ -8402,6 +8481,7 @@ async def admin_create_model_definition(request: Request) -> JSONResponse:
)
await asyncio.to_thread(_refresh_coord_registry, request.app.state, storage)
_emit_models_changed(request)
created = storage.get_model_definition(definition_id)
if created is None:
@@ -8495,6 +8575,9 @@ async def admin_update_model_definition(request: Request) -> JSONResponse:
updates["context_window"] = max(0, int(ctx_raw)) if isinstance(ctx_raw, (int, float)) else 0
if "capabilities" in body:
caps = body["capabilities"]
err_msg = _validate_api_surface(caps)
if err_msg:
return JSONResponse({"error": err_msg}, status_code=400)
updates["capabilities"] = json.dumps(caps) if isinstance(caps, dict) else "{}"
if "enabled" in body:
updates["enabled"] = bool(body["enabled"])
@@ -8562,6 +8645,7 @@ async def admin_update_model_definition(request: Request) -> JSONResponse:
if updates:
await asyncio.to_thread(_refresh_coord_registry, request.app.state, storage)
_emit_models_changed(request)
model_def = storage.get_model_definition(definition_id)
return JSONResponse(_mask_model_secrets(model_def or {}))
@@ -8599,6 +8683,7 @@ async def admin_delete_model_definition(request: Request) -> JSONResponse:
)
await asyncio.to_thread(_refresh_coord_registry, request.app.state, storage)
_emit_models_changed(request)
return JSONResponse({"status": "ok", "definition_id": definition_id})
@@ -8625,6 +8710,7 @@ async def admin_model_reload(request: Request) -> JSONResponse:
# otherwise the coord LLM keeps calling the prior model name even
# after a successful reload.
await asyncio.to_thread(_refresh_coord_registry, request.app.state, storage)
_emit_models_changed(request)
results = await _notify_nodes_model_reload(request)
return JSONResponse({"status": "ok", "results": results})
@@ -9115,6 +9201,8 @@ async def admin_update_judge_setting(request: Request) -> JSONResponse:
ip,
)
await _publish_config_change(request)
if key in _MODEL_AFFECTING_SETTING_KEYS:
_emit_models_changed(request)
effective = config_store.get(key, defn.default)
return JSONResponse(
@@ -9151,6 +9239,8 @@ async def admin_delete_judge_setting(request: Request) -> JSONResponse:
audit_uid, ip = _audit_context(request)
record_audit(storage, audit_uid, "setting.delete", "setting", key, {}, ip)
await _publish_config_change(request)
if key in _MODEL_AFFECTING_SETTING_KEYS:
_emit_models_changed(request)
return JSONResponse({"status": "ok", "key": key, "default": defn.default})
+330 -17
View File
@@ -2625,11 +2625,22 @@ function loadSettings() {
schemaMap[schemaArr[i].key] = schemaArr[i];
}
// Merge values + schema
// Merge values + schema. Skip role-assignment settings owned by
// the Models → Roles sub-tab (judge.* settings still live on the
// Judge tab; the four model-tab roles render only there).
var merged = {};
var roleKeys = {
"coordinator.model_alias": 1,
"coordinator.reasoning_effort": 1,
"model.plan_alias": 1,
"model.plan_effort": 1,
"model.task_alias": 1,
"model.task_effort": 1,
};
for (var j = 0; j < valuesArr.length; j++) {
var v = valuesArr[j];
if (v.key.startsWith("judge.")) continue;
if (roleKeys[v.key]) continue;
var s = schemaMap[v.key] || {};
merged[v.key] = {
key: v.key,
@@ -4440,7 +4451,69 @@ var _modelDefaultAlias = "";
var _modelCreateTrap = null;
var _modelCreateTrigger = null;
// Roles surfaced in the Models → Roles sub-tab. Each entry maps a
// settings-registry key onto a UX label. ``effortKey`` is optional —
// roles whose registry entry has a paired ``*.reasoning_effort``
// setting render a second selector inline. Adding a new role (e.g.
// ``perception.audio.model``) is purely additive: drop a row here once
// the SettingDef lands in turnstone/core/settings_registry.py.
var MODEL_ROLES = [
{
label: "Coordinator",
description:
"Console-hosted coordinator sessions that drive child workstreams.",
aliasKey: "coordinator.model_alias",
effortKey: "coordinator.reasoning_effort",
},
{
label: "Judge",
description:
"Intent-validation judge that scores tool calls before approval.",
aliasKey: "judge.model",
},
{
label: "Plan agent",
description:
"plan_agent sub-agent — produces high-level plans before task dispatch.",
aliasKey: "model.plan_alias",
effortKey: "model.plan_effort",
},
{
label: "Task agent",
description:
"task_agent sub-agent — runs autonomous subtasks dispatched by the parent.",
aliasKey: "model.task_alias",
effortKey: "model.task_effort",
},
];
// Roles sub-tab reads/writes via ``/v1/api/admin/settings`` which
// requires ``admin.settings`` — different from the ``admin.models``
// permission gating the Models tab itself. When the user has Models
// access but not Settings, hide the sub-tab button + force the
// Definitions panel visible so they don't see a perpetual 403 loader.
function _modelRolesAccessible() {
var perms = sessionStorage.getItem("turnstone_permissions") || "";
return perms.split(",").indexOf("admin.settings") !== -1;
}
function _applyModelRolesPermission() {
var btn = document.getElementById("models-tab-roles");
if (!btn) return;
if (_modelRolesAccessible()) {
btn.style.display = "";
return;
}
btn.style.display = "none";
// If Roles was the active sub-tab, snap back to Definitions so the
// user isn't staring at a hidden panel.
if (btn.classList.contains("active")) {
switchModelsSection("models-list");
}
}
function loadAdminModels() {
_applyModelRolesPermission();
authFetch("/v1/api/admin/model-definitions")
.then(function (r) {
if (!r.ok) throw new Error("Failed");
@@ -4450,6 +4523,12 @@ function loadAdminModels() {
_modelDefs = data.models || [];
_modelDefaultAlias = data.default_alias || "";
_renderModels(_modelDefs);
// Roles sub-tab piggybacks on the model list; skip it when the
// user has no settings permission since the underlying API will
// 403 anyway.
if (_modelRolesAccessible()) {
loadAdminModelRoles();
}
})
.catch(function () {
var el = document.getElementById("admin-models-table");
@@ -4461,6 +4540,223 @@ function loadAdminModels() {
});
}
function switchModelsSection(section) {
var sections = document.querySelectorAll("#admin-models .models-section");
for (var i = 0; i < sections.length; i++) sections[i].style.display = "none";
var switcher = document.querySelector("#admin-models .admin-subtab-switcher");
var btns = switcher ? switcher.querySelectorAll(".admin-subtab-btn") : [];
for (var k = 0; k < btns.length; k++) {
var isActive = btns[k].getAttribute("data-section") === section;
btns[k].classList.toggle("active", isActive);
btns[k].setAttribute("aria-selected", isActive ? "true" : "false");
btns[k].setAttribute("tabindex", isActive ? "0" : "-1");
}
var target = document.getElementById(section + "-section");
if (target) target.style.display = "";
}
// Arrow key navigation for Models sub-tabs (matches the Judge tab).
(function () {
var switcher = document.querySelector("#admin-models .admin-subtab-switcher");
if (!switcher) return;
switcher.addEventListener("keydown", function (e) {
if (e.key !== "ArrowLeft" && e.key !== "ArrowRight") return;
var btns = switcher.querySelectorAll(".admin-subtab-btn");
var secs = [];
for (var i = 0; i < btns.length; i++)
secs.push(btns[i].getAttribute("data-section"));
var current = switcher.querySelector(".admin-subtab-btn.active");
var idx = secs.indexOf(current ? current.getAttribute("data-section") : "");
if (e.key === "ArrowRight") idx = (idx + 1) % secs.length;
else idx = (idx - 1 + secs.length) % secs.length;
e.preventDefault();
switchModelsSection(secs[idx]);
btns[idx].focus();
});
})();
function _modelRolesError(container, msg) {
while (container.firstChild) container.removeChild(container.firstChild);
var d = document.createElement("div");
d.className = "dashboard-empty";
d.textContent = msg;
container.appendChild(d);
}
function loadAdminModelRoles() {
var c = document.getElementById("admin-models-roles-container");
if (!c) return;
// Reads ``_modelDefs`` / ``_modelDefaultAlias`` populated by the most
// recent ``loadAdminModels`` — both entry points into the Models tab
// (initial open + ``models_changed`` SSE refresh) go through
// ``loadAdminModels`` first, so the cached snapshot is fresh. Role
// saves don't change model definitions, so the snapshot stays
// accurate after ``_saveModelRole`` chains back here.
Promise.all([
authFetch("/v1/api/admin/settings").then(function (r) {
if (!r.ok) throw new Error("settings " + r.status);
return r.json();
}),
authFetch("/v1/api/admin/settings/schema").then(function (r) {
if (!r.ok) throw new Error("schema " + r.status);
return r.json();
}),
])
.then(function (results) {
var values = {};
var arr = results[0].settings || [];
for (var i = 0; i < arr.length; i++) values[arr[i].key] = arr[i];
var schema = {};
var sa = results[1].schema || [];
for (var j = 0; j < sa.length; j++) schema[sa[j].key] = sa[j];
_renderModelRoles(c, values, schema);
})
.catch(function () {
_modelRolesError(c, "Failed to load roles");
});
}
function _renderModelRoles(container, values, schema) {
var enabledAliases = [];
for (var i = 0; i < _modelDefs.length; i++) {
if (_modelDefs[i].enabled) enabledAliases.push(_modelDefs[i]);
}
container.textContent = "";
for (var r = 0; r < MODEL_ROLES.length; r++) {
var role = MODEL_ROLES[r];
var aliasInfo = values[role.aliasKey];
if (!aliasInfo) continue; // setting not registered (e.g. older server)
var row = document.createElement("div");
row.className = "model-role-row";
// The dropdown's selected-option text is the single source of
// truth for default vs override — when nothing is set it shows
// "(default — <alias>)", otherwise it shows the chosen alias. No
// separate badge: redundant with the select, and prone to
// confusing color contrasts on freshly-rendered rows.
var head = document.createElement("div");
head.className = "model-role-head";
var nameEl = document.createElement("span");
nameEl.className = "model-role-label";
nameEl.textContent = role.label;
head.appendChild(nameEl);
row.appendChild(head);
if (role.description) {
var desc = document.createElement("div");
desc.className = "model-role-desc";
desc.textContent = role.description;
row.appendChild(desc);
}
var controls = document.createElement("div");
controls.className = "model-role-controls";
// Alias dropdown
var aliasWrap = document.createElement("label");
aliasWrap.className = "model-role-control";
var aliasLabel = document.createElement("span");
aliasLabel.className = "model-role-control-label";
aliasLabel.textContent = "Model";
aliasWrap.appendChild(aliasLabel);
var aliasSel = document.createElement("select");
aliasSel.setAttribute("data-role-key", role.aliasKey);
aliasSel.setAttribute(
"aria-label",
role.label + " model (empty = default)",
);
var blank = document.createElement("option");
blank.value = "";
blank.textContent = _modelDefaultAlias
? "(default — " + _modelDefaultAlias + ")"
: "(default)";
aliasSel.appendChild(blank);
var currentAlias = aliasInfo.value || "";
var matched = false;
for (var m = 0; m < enabledAliases.length; m++) {
var md = enabledAliases[m];
var opt = document.createElement("option");
opt.value = md.alias;
opt.textContent =
md.alias === md.model ? md.alias : md.alias + " (" + md.model + ")";
if (currentAlias && currentAlias === md.alias) {
opt.selected = true;
matched = true;
}
aliasSel.appendChild(opt);
}
if (currentAlias && !matched) {
var manual = document.createElement("option");
manual.value = currentAlias;
manual.textContent = currentAlias + " (manual)";
manual.selected = true;
aliasSel.appendChild(manual);
}
aliasSel.addEventListener("change", function () {
_saveModelRole(this.getAttribute("data-role-key"), this.value);
});
aliasWrap.appendChild(aliasSel);
controls.appendChild(aliasWrap);
// Optional reasoning effort dropdown
if (role.effortKey && values[role.effortKey] && schema[role.effortKey]) {
var effortWrap = document.createElement("label");
effortWrap.className = "model-role-control";
var effortLabel = document.createElement("span");
effortLabel.className = "model-role-control-label";
effortLabel.textContent = "Reasoning effort";
effortWrap.appendChild(effortLabel);
var effortSel = document.createElement("select");
effortSel.setAttribute("data-role-key", role.effortKey);
effortSel.setAttribute("aria-label", role.label + " reasoning effort");
var choices = schema[role.effortKey].choices || [];
var currentEffort = values[role.effortKey].value;
for (var c2 = 0; c2 < choices.length; c2++) {
var eo = document.createElement("option");
eo.value = choices[c2];
eo.textContent = choices[c2] === "" ? "(inherit)" : choices[c2];
if (currentEffort === choices[c2]) eo.selected = true;
effortSel.appendChild(eo);
}
effortSel.addEventListener("change", function () {
_saveModelRole(this.getAttribute("data-role-key"), this.value);
});
effortWrap.appendChild(effortSel);
controls.appendChild(effortWrap);
}
row.appendChild(controls);
container.appendChild(row);
}
if (!container.children.length) {
_modelRolesError(container, "No model roles configured");
}
}
function _saveModelRole(key, value) {
authFetch("/v1/api/admin/settings/" + encodeURIComponent(key), {
method: "PUT",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ value: value }),
})
.then(function (r) {
if (!r.ok)
return r.json().then(function (d) {
throw new Error(d.error || "Failed");
});
return r.json();
})
.then(function () {
showToast("Saved");
loadAdminModelRoles();
})
.catch(function (e) {
showToast("Error: " + (e && e.message ? e.message : "save failed"));
});
}
function _renderModels(items) {
var el = document.getElementById("admin-models-table");
// Clear previous content
@@ -4705,6 +5001,7 @@ function showCreateModelModal() {
document.getElementById("model-max-tokens").value = "";
document.getElementById("model-reasoning-effort").value = "";
document.getElementById("model-server-type").value = "";
document.getElementById("model-api-surface").value = "";
document.getElementById("model-thinking-mode").value = "";
document.getElementById("model-thinking-param").value = "";
document.getElementById("model-thinking-param-row").style.display = "none";
@@ -4781,8 +5078,9 @@ function showEditModelModal(definitionId) {
document.getElementById("model-thinking-param").value = "";
}
_toggleThinkingParam();
// Server compat: server_type and extra_body workarounds
// Server compat: server_type, api_surface, and extra_body workarounds
document.getElementById("model-server-type").value = sc.server_type || "";
document.getElementById("model-api-surface").value = sc.api_surface || "";
var eb = sc.extra_body || {};
var ebText = JSON.stringify(eb, null, 2);
document.getElementById("model-extra-body").value =
@@ -4864,26 +5162,34 @@ function submitCreateModel() {
if (savedParam) caps.thinking_param = savedParam;
}
// Build server_compat from structured fields
// Build server_compat from structured fields. Only meaningful for
// openai-compatible aliases — for other providers the section is hidden
// but the form values can linger after a provider switch, so gate the
// whole block on the active provider to keep persisted state honest.
var serverCompat = {};
var serverType = document.getElementById("model-server-type").value;
if (serverType) serverCompat.server_type = serverType;
var providerVal = document.getElementById("model-provider").value;
var ebEl = document.getElementById("model-extra-body");
var ebText = ebEl.value.trim();
ebEl.removeAttribute("aria-invalid");
ebEl.style.borderColor = "";
if (ebText) {
try {
var ebParsed = JSON.parse(ebText);
if (!_isPlainObject(ebParsed)) {
throw new Error("not an object");
if (providerVal === "openai-compatible") {
var serverType = document.getElementById("model-server-type").value;
if (serverType) serverCompat.server_type = serverType;
var apiSurface = document.getElementById("model-api-surface").value;
if (apiSurface) serverCompat.api_surface = apiSurface;
var ebText = ebEl.value.trim();
if (ebText) {
try {
var ebParsed = JSON.parse(ebText);
if (!_isPlainObject(ebParsed)) {
throw new Error("not an object");
}
serverCompat.extra_body = ebParsed;
} catch (e) {
ebEl.setAttribute("aria-invalid", "true");
ebEl.style.borderColor = "var(--red)";
_showModelError("Extra body params must be a JSON object");
return;
}
serverCompat.extra_body = ebParsed;
} catch (e) {
ebEl.setAttribute("aria-invalid", "true");
ebEl.style.borderColor = "var(--red)";
_showModelError("Extra body params must be a JSON object");
return;
}
}
if (Object.keys(serverCompat).length > 0) {
@@ -5114,6 +5420,13 @@ function detectModel() {
stOpts2.indexOf(ssc.server_type) !== -1
)
stEl2.value = ssc.server_type;
// Restrict to the known set so a hostile detect response can't
// smuggle a non-listed value into the form.
var _SURFACE_SUGGESTABLE = { chat: 1, responses: 1 };
if (ssc.api_surface && _SURFACE_SUGGESTABLE[ssc.api_surface]) {
var asEl = document.getElementById("model-api-surface");
if (!asEl.value) asEl.value = ssc.api_surface;
}
if (ssc.extra_body) {
var ebEl2 = document.getElementById("model-extra-body");
if (!ebEl2.value.trim()) {
+389 -84
View File
@@ -5,17 +5,13 @@ window.onLoginSuccess = function () {
if (typeof _refreshHomeComposerVisibility === "function") {
_refreshHomeComposerVisibility();
}
// Re-populate the home-composer skill dropdown and re-probe the
// coordinator subsystem now that auth has landed. The initial
// page-load pass runs before login completes, so /v1/api/skills
// and /v1/api/workstreams both 401; without this re-run the
// dropdown stays empty and the 503 banner never flips correctly.
// Re-populate the home-composer skill dropdown now that auth has
// landed. The initial page-load pass runs before login completes,
// so /v1/api/skills 401s; without this re-run the dropdown stays
// empty.
if (typeof _populateHomeSkillDropdown === "function") {
_populateHomeSkillDropdown();
}
if (typeof _probeCoordSubsystem === "function") {
_probeCoordSubsystem();
}
// Active-coordinators list is SSE-driven via the console pseudo-node
// (#9) — no poller to restart after login. The home-view renderer
// reads from clusterState.nodes["console"].workstreams on every SSE
@@ -427,6 +423,22 @@ function handleClusterEvent(data) {
if (data.type === "ws_closed" && data.reason === "evicted") {
showToast("Evicted" + (data.name ? ": " + data.name : "") + " (capacity)");
}
if (data.type === "models_changed") {
// Server emits this when a model definition or a role-assignment
// setting (model.default_alias, judge.model, coordinator.model_alias,
// coordinator.reasoning_effort) changes. Refresh anything that
// renders model aliases so labels stay accurate without a reload.
if (typeof _populateHomeModelDropdowns === "function") {
_populateHomeModelDropdowns();
}
if (
typeof _adminTab !== "undefined" &&
_adminTab === "models" &&
typeof loadAdminModels === "function"
) {
loadAdminModels();
}
}
}
// --- Home View ---
@@ -1445,8 +1457,8 @@ function _hasCoordPermission() {
// POST /v1/api/workstreams/new. Accepts the three request fields
// directly + an errEl / setBusy callback so the caller owns the
// loading-state UX (button label swap, composer disabled flag, etc.).
// On success redirects to /coordinator/{ws_id}; on 503 invokes on503
// so the caller can surface the "subsystem not configured" banner.
// On success redirects to /coordinator/{ws_id}; on failure surfaces
// the server's error text inline through errEl.
function _createCoordinator(opts) {
var name = (opts.name || "").trim();
var skill = opts.skill || "";
@@ -1455,10 +1467,12 @@ function _createCoordinator(opts) {
var task = (opts.task || "").trim();
var errEl = opts.errEl;
var setBusy = opts.setBusy || function () {};
var on503 = opts.on503 || function () {};
var onSuccess = opts.onSuccess || function () {};
errEl.style.display = "none";
// Error region is always rendered with reserved min-height (see
// .home-composer-error in style.css) so toggling validation messages
// doesn't reflow the active-coordinators list below — clear the
// textContent only, no display toggle.
errEl.textContent = "";
setBusy(true);
@@ -1469,11 +1483,30 @@ function _createCoordinator(opts) {
if (judgeModel) body.judge_model = judgeModel;
if (task) body.initial_message = task;
authFetch("/v1/api/workstreams/new", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(body),
})
// Multipart when files are staged — the coord create endpoint
// accepts a `meta` JSON field plus zero-or-more `file` parts and
// reserves attachments for the very first turn (same flow the
// interactive UI's new-ws modal uses against the server). Plain
// JSON stays the default when no files are attached.
var files = Array.isArray(opts.files) ? opts.files : [];
var fetchOpts;
if (files.length > 0) {
var form = new FormData();
form.append("meta", JSON.stringify(body));
for (var i = 0; i < files.length; i++) {
form.append("file", files[i], files[i].name);
}
// Don't set Content-Type — the browser adds the correct boundary.
fetchOpts = { method: "POST", body: form };
} else {
fetchOpts = {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(body),
};
}
authFetch("/v1/api/workstreams/new", fetchOpts)
.then(function (r) {
return r.json().then(function (data) {
return { ok: r.ok, status: r.status, data: data };
@@ -1481,14 +1514,9 @@ function _createCoordinator(opts) {
})
.then(function (res) {
setBusy(false);
if (res.status === 503) {
on503(res);
return;
}
if (!res.ok || !res.data || !res.data.ws_id) {
errEl.textContent =
(res.data && res.data.error) || "HTTP " + res.status;
errEl.style.display = "block";
return;
}
onSuccess(res);
@@ -1498,7 +1526,6 @@ function _createCoordinator(opts) {
.catch(function () {
setBusy(false);
errEl.textContent = "Request failed";
errEl.style.display = "block";
});
}
@@ -1511,19 +1538,166 @@ function _createCoordinator(opts) {
// ---------------------------------------------------------------------------
var _homeComposerInit = false;
var _homeCoordReady = null; // tri-state: null = unknown, true = ready, false = 503
var _homeCoordComposer = null; // shared Composer instance
var _homeCoordBusy = false;
// Single owner for sendBtn.disabled: disabled if EITHER busy OR the
// subsystem probe flipped to 503. Every setter for _homeCoordBusy /
// _homeCoordReady ends with a call here so the two inputs can't drift
// out of sync (and a probe resolving mid-submit can't re-enable the
// button under an in-flight request).
// Attachment staging for the home coord composer. The coord ws_id
// doesn't exist until the create POST resolves, so we hold File
// objects in memory and ship them as multipart parts on submit (same
// pattern interactive uses for its new-ws modal + dashboard composer).
var _homeStagedFiles = [];
// Per-kind size caps + allowlist mirrored from turnstone/core/attachments.py
// so the browser can fail fast. Keep in sync with the interactive
// UI's _ATTACH_* constants in turnstone/ui/static/app.js.
var _HOME_IMAGE_CAP = 4 * 1024 * 1024;
var _HOME_TEXT_CAP = 512 * 1024;
var _HOME_MAX_FILES = 10;
var _HOME_IMAGE_MIMES = ["image/png", "image/jpeg", "image/gif", "image/webp"];
var _HOME_TEXT_APP_MIMES = [
"application/json",
"application/xml",
"application/x-yaml",
"application/yaml",
"application/toml",
];
var _HOME_TEXT_EXTENSIONS = [
".c",
".conf",
".cpp",
".css",
".go",
".h",
".hpp",
".html",
".ini",
".java",
".js",
".json",
".jsx",
".md",
".py",
".rs",
".sh",
".sql",
".toml",
".ts",
".tsx",
".txt",
".xml",
".yaml",
".yml",
];
function _homeFormatSize(n) {
if (n < 1024) return n + " B";
if (n < 1024 * 1024) return (n / 1024).toFixed(1) + " KB";
return (n / (1024 * 1024)).toFixed(1) + " MB";
}
function _homeIsAttachmentAllowed(file) {
var mime = (file.type || "").toLowerCase();
if (_HOME_IMAGE_MIMES.indexOf(mime) !== -1) return true;
if (mime.indexOf("text/") === 0) return true;
if (_HOME_TEXT_APP_MIMES.indexOf(mime) !== -1) return true;
var name = (file.name || "").toLowerCase();
var dot = name.lastIndexOf(".");
if (dot >= 0 && _HOME_TEXT_EXTENSIONS.indexOf(name.substr(dot)) !== -1) {
return true;
}
return false;
}
function _homeShowError(msg) {
var errEl = document.getElementById("home-coord-error");
if (!errEl) return;
// Element is always rendered (min-height reserves the row); just
// toggle the message text so layout doesn't shift on validation.
errEl.textContent = msg || "";
}
function _homeRenderChips() {
if (!_homeCoordComposer || !_homeCoordComposer.chipsEl) return;
var chipsEl = _homeCoordComposer.chipsEl;
chipsEl.textContent = "";
for (var i = 0; i < _homeStagedFiles.length; i++) {
(function (idx) {
var f = _homeStagedFiles[idx];
var isImage = (f.type || "").indexOf("image/") === 0;
var chip = document.createElement("span");
chip.className =
"composer-chip composer-chip-" + (isImage ? "image" : "text");
chip.setAttribute("role", "listitem");
var icon = document.createElement("span");
icon.className = "composer-chip-icon";
icon.setAttribute("aria-hidden", "true");
icon.textContent = isImage ? "🖼" : "📄";
chip.appendChild(icon);
var name = document.createElement("span");
name.className = "composer-chip-name";
name.textContent = f.name;
name.title = f.name + " (" + f.size + " bytes)";
chip.appendChild(name);
var size = document.createElement("span");
size.className = "composer-chip-size";
size.textContent = _homeFormatSize(f.size);
chip.appendChild(size);
var rm = document.createElement("button");
rm.type = "button";
rm.className = "composer-chip-remove";
rm.setAttribute("aria-label", "Remove " + f.name);
rm.title = "Remove";
rm.textContent = "×";
rm.onclick = function () {
_homeStagedFiles.splice(idx, 1);
_homeRenderChips();
};
chip.appendChild(rm);
chipsEl.appendChild(chip);
})(i);
}
}
function _homeStageFile(file) {
if (!file) return;
if (_homeStagedFiles.length >= _HOME_MAX_FILES) {
_homeShowError(
"At most " + _HOME_MAX_FILES + " attachments per coordinator",
);
return;
}
if (!_homeIsAttachmentAllowed(file)) {
_homeShowError(
"Unsupported file type: " +
file.name +
" (allowed: png/jpeg/gif/webp images, text)",
);
return;
}
var isImage = (file.type || "").indexOf("image/") === 0;
var cap = isImage ? _HOME_IMAGE_CAP : _HOME_TEXT_CAP;
if (file.size > cap) {
_homeShowError(file.name + " exceeds the " + _homeFormatSize(cap) + " cap");
return;
}
_homeShowError("");
_homeStagedFiles.push(file);
_homeRenderChips();
}
function _homeClearStagedFiles() {
_homeStagedFiles = [];
_homeRenderChips();
}
// Sole owner of sendBtn.disabled: disables while a submit is in flight.
function _refreshHomeCoordSubmitEnabled() {
if (!_homeCoordComposer) return;
_homeCoordComposer.sendBtn.disabled =
_homeCoordBusy || _homeCoordReady === false;
_homeCoordComposer.sendBtn.disabled = _homeCoordBusy;
}
function _ensureHomeComposerInit() {
@@ -1532,7 +1706,6 @@ function _ensureHomeComposerInit() {
_mountHomeCoordComposer();
_populateHomeSkillDropdown();
_populateHomeModelDropdowns();
_probeCoordSubsystem();
_refreshHomeComposerVisibility();
}
@@ -1596,6 +1769,12 @@ function _mountHomeCoordComposer() {
},
],
},
attachments: {
onAttach: function (file) {
_homeStageFile(file);
},
},
dragDrop: { targetEl: mount, dropClass: "home-coord-drop" },
onSend: function (text) {
submitHomeCoord(text);
},
@@ -1646,40 +1825,6 @@ function _populateHomeModelDropdowns() {
});
}
// Probe GET /v1/api/workstreams — 200 = subsystem ready; 503 = no model
// alias resolvable, show remediation banner. 4xx (auth / permission) is
// treated as "unknown, don't flip the banner" because the probe cannot
// actually tell us anything about subsystem readiness in that case —
// the caller is expected to re-invoke this after login lands so a real
// answer can arrive. Leaving the submit button enabled on unknown
// keeps first-paint usable; a subsequent 503 from the actual submit
// flips the banner via _createCoordinator's on503 hook.
//
// Skip the probe entirely for users without admin.coordinator — they
// can't see the composer anyway (see _refreshHomeComposerVisibility),
// and the endpoint returns 403 for them, producing a useless network
// round-trip on every login.
function _probeCoordSubsystem() {
if (!_hasCoordPermission()) return;
authFetch("/v1/api/workstreams")
.then(function (r) {
if (r.status === 503) {
_homeCoordReady = false;
} else if (r.ok) {
_homeCoordReady = true;
} else {
_homeCoordReady = null;
return;
}
var banner = document.getElementById("coord-composer-503");
if (banner) banner.style.display = _homeCoordReady ? "none" : "";
_refreshHomeCoordSubmitEnabled();
})
.catch(function () {
/* network error — leave banner hidden; submit will surface a retryable error */
});
}
function _refreshHomeComposerVisibility() {
var panel = document.getElementById("coord-composer-panel");
if (!panel) return;
@@ -1698,23 +1843,36 @@ function submitHomeCoord(textFromComposer) {
var task =
textFromComposer != null ? textFromComposer : _homeCoordComposer.value;
var opts = _homeCoordComposer.getOptionValues();
// Snapshot at submit time so a chip remove mid-request can't race
// the multipart payload (the actual reset only fires on the success
// branch, after the response lands).
var files = _homeStagedFiles.slice();
// Files-without-text would upload pending attachment rows but the
// server's _coord_create_post_install only reserves+dispatches when
// initial_message is non-empty — uploaded files would orphan as
// pending storage rows until the GC sweep. Require text whenever
// attachments are staged so the first turn always picks them up.
if (files.length > 0 && !(task || "").trim()) {
_homeShowError(
"Add a task message — attachments need an initial turn to dispatch on.",
);
return;
}
_createCoordinator({
name: opts.name || "",
skill: opts.skill || "",
model: opts.model || "",
judge_model: opts.judge_model || "",
task: task,
files: files,
errEl: document.getElementById("home-coord-error"),
setBusy: function (b) {
_homeCoordBusy = b;
if (_homeCoordComposer) _homeCoordComposer.setBusy(b);
_refreshHomeCoordSubmitEnabled();
},
on503: function () {
_homeCoordReady = false;
var banner = document.getElementById("coord-composer-503");
if (banner) banner.style.display = "";
_refreshHomeCoordSubmitEnabled();
onSuccess: function () {
_homeClearStagedFiles();
},
});
}
@@ -1730,9 +1888,6 @@ document.addEventListener("keydown", function (e) {
var mount = document.getElementById("home-coord-composer-mount");
if (!mount || !mount.contains(e.target)) return;
e.preventDefault();
// sendBtn.disabled is the single reconciler of busy + 503-ready —
// checking it here is enough to avoid double-submits or submits
// while the subsystem is down.
if (!_homeCoordComposer.sendBtn.disabled) submitHomeCoord();
});
@@ -1851,6 +2006,16 @@ var _savedCoordsRetry = false;
function loadSavedCoordinators() {
if (!_hasCoordPermission()) return;
// Freeze the list while the user is multi-selecting — re-rendering
// mid-mode would shuffle the visible page out from under them. The
// delete-mode wrapper drains the retry flag on cancel/onClose.
if (
typeof _coordDeleteController !== "undefined" &&
_coordDeleteController.inMode()
) {
_savedCoordsRetry = true;
return;
}
if (_savedCoordsInFlight) {
_savedCoordsRetry = true;
return;
@@ -1861,6 +2026,16 @@ function loadSavedCoordinators() {
return r.ok ? r.json() : { workstreams: [] };
})
.then(function (data) {
// Belt-and-braces: if the user entered delete mode while this
// fetch was already in flight, defer the render — re-rendering
// mid-selection would shuffle visible cards and reshape selections.
if (
typeof _coordDeleteController !== "undefined" &&
_coordDeleteController.inMode()
) {
_savedCoordsRetry = true;
return;
}
renderSavedCoordinators(data.workstreams || []);
})
.catch(function () {
@@ -1878,7 +2053,52 @@ function loadSavedCoordinators() {
});
}
// Saved Coordinators: paginated card list + multi-select delete.
// The shared controller (createSavedCardsController in /shared/cards.js)
// owns mode state, checkbox decoration, the toolbar, and the modal.
// Pagination caps Select-All fan-out at COORD_PAGE_SIZE — the controller
// only ever sees the visible page, so a confirm-all batch is bounded to
// COORD_PAGE_SIZE parallel POSTs against the routing proxy.
var COORD_PAGE_SIZE = 24;
var _coordPage = 0;
var _coordSavedItems = [];
var _coordDeleteController = createSavedCardsController({
idPrefix: "coord-delete",
buttonId: "coord-delete-btn",
noun: "coordinator",
activateLabel: function (s) {
return "Resume coordinator: " + (s.alias || s.title || s.name || s.ws_id);
},
// Coordinators live on whichever node owns the ws_id, so we can't fire
// a path-keyed delete the way ui/static does. The router proxy reads
// ws_id from the body, resolves the owning node via the consistent-
// hash ring, and forwards to that node's POST workstreams/{ws_id}/delete.
buildDeleteRequest: function (wsId) {
return {
url: "/v1/api/route/workstreams/delete",
options: {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ ws_id: wsId }),
},
};
},
render: function () {
renderSavedCoordinators(_coordSavedItems);
},
onClose: function () {
// Drain queued retries before the explicit reload — without this,
// _savedCoordsRetry is still true from SSE events that arrived
// during the freeze, so loadSavedCoordinators's .finally() would
// re-fire a second fetch immediately after the first resolves.
// Same idiom as cancelCoordDeleteMode below.
_savedCoordsRetry = false;
loadSavedCoordinators();
},
});
function renderSavedCoordinators(items) {
_coordSavedItems = items;
var section = document.getElementById("saved-coordinators");
var cards = document.getElementById("saved-coord-cards");
var countEl = document.getElementById("saved-coord-count");
@@ -1887,24 +2107,31 @@ function renderSavedCoordinators(items) {
section.style.display = "none";
cards.replaceChildren();
if (countEl) countEl.textContent = "";
_coordPage = 0;
if (_coordDeleteController.inMode()) _coordDeleteController.cancel();
_renderCoordPagination();
return;
}
// Clamp the page index after deletes (or upstream churn) shrink the list.
var pages = Math.max(1, Math.ceil(items.length / COORD_PAGE_SIZE));
if (_coordPage > pages - 1) _coordPage = pages - 1;
if (_coordPage < 0) _coordPage = 0;
var visible = items.slice(
_coordPage * COORD_PAGE_SIZE,
(_coordPage + 1) * COORD_PAGE_SIZE,
);
_coordDeleteController.setItems(visible);
section.style.display = "";
if (countEl) countEl.textContent = "(" + items.length + ")";
cards.replaceChildren();
items.forEach(function (sess) {
visible.forEach(function (sess) {
var card = renderSessionCard(sess, {
ariaLabel: function (s) {
return (
"Resume coordinator: " + (s.alias || s.title || s.name || s.ws_id)
);
},
ariaLabel: _coordDeleteController.ariaLabel,
onActivate: function (s, cardEl) {
if (_coordDeleteController.blockActivate()) return;
// POST /open BEFORE navigating so capacity issues surface as a
// toast instead of a broken-looking detail page. The /open
// endpoint calls the same lazy-rehydrate path the GET would,
// but we get the status code synchronously so the user learns
// "all slots in use" instead of staring at a 404.
// toast instead of a broken-looking detail page.
cardEl.classList.add("is-busy");
authFetch(
"/v1/api/workstreams/" + encodeURIComponent(s.ws_id) + "/open",
@@ -1936,8 +2163,86 @@ function renderSavedCoordinators(items) {
});
},
});
_coordDeleteController.decorateCard(card, sess);
cards.appendChild(card);
});
if (_coordDeleteController.inMode()) _coordDeleteController.refreshBar();
_renderCoordPagination();
}
function _renderCoordPagination() {
var pag = document.getElementById("coord-pagination");
if (!pag) return;
var total = _coordSavedItems.length;
var pages = Math.max(1, Math.ceil(total / COORD_PAGE_SIZE));
// Single-page lists and delete-mode hide the controls — page changes
// would invalidate the user's checkbox selections, so we lock them out.
if (pages <= 1 || _coordDeleteController.inMode()) {
pag.style.display = "none";
return;
}
pag.style.display = "";
var label = document.getElementById("coord-page-label");
if (label) {
/* Visible text uses the terse "X / Y" form to match the
filtered-pagination control elsewhere in the console; the long
form sits on the parent's aria-label so screen readers still get
a full sentence. */
label.textContent = _coordPage + 1 + " / " + pages;
pag.setAttribute(
"aria-label",
"Saved coordinators pagination — page " +
(_coordPage + 1) +
" of " +
pages,
);
}
var prev = document.getElementById("coord-page-prev");
if (prev) prev.disabled = _coordPage <= 0;
var next = document.getElementById("coord-page-next");
if (next) next.disabled = _coordPage >= pages - 1;
}
function coordPagePrev() {
if (_coordPage > 0) {
_coordPage--;
renderSavedCoordinators(_coordSavedItems);
}
}
function coordPageNext() {
var pages = Math.max(1, Math.ceil(_coordSavedItems.length / COORD_PAGE_SIZE));
if (_coordPage < pages - 1) {
_coordPage++;
renderSavedCoordinators(_coordSavedItems);
}
}
// HTML inline-onclick wrappers — keep the global names the markup binds
// to and forward to the shared controller.
function startCoordDeleteMode() {
_coordDeleteController.start();
}
function cancelCoordDeleteMode() {
_coordDeleteController.cancel();
// The freeze gate (see loadSavedCoordinators) may have queued retries
// while we were multi-selecting; drain them now that we're idle again.
if (_savedCoordsRetry) {
_savedCoordsRetry = false;
loadSavedCoordinators();
}
}
function toggleCoordSelectAll() {
_coordDeleteController.toggleAll();
}
function confirmCoordDeleteSelection() {
_coordDeleteController.confirmSelection();
}
function cancelCoordDelete() {
_coordDeleteController.closeModal();
}
function confirmCoordDelete() {
_coordDeleteController.confirm();
}
// --- Init ---
@@ -396,6 +396,87 @@
background: color-mix(in srgb, var(--err) 12%, var(--panel-2));
}
/* Output-guard finding rendered under its specific .coord-tool-row.
Stays anchored to the call that tripped the guard rather than
floating into the chat log as a generic "[output guard]" line
matches interactive's `.output-warning` placement convention.
Severity drives the hue (matches .verdict-badge.verdict-* palette
so an operator scanning a workstream reads risk consistently
across both surfaces). */
.coord-tool-row-warning {
display: inline-flex;
align-items: center;
gap: 6px;
margin-top: 4px;
padding: 2px 8px;
font-size: 11px;
font-family: var(--font-mono);
border: 1px solid var(--hair);
border-radius: 3px;
border-left-width: 3px;
background: var(--panel-2);
color: var(--ink-2);
max-width: max-content;
}
.coord-tool-row-warning--low {
color: color-mix(in srgb, var(--ok) 70%, var(--ink-2));
border-left-color: var(--ok);
background: color-mix(in srgb, var(--ok) 12%, var(--panel-2));
}
.coord-tool-row-warning--medium {
color: color-mix(in srgb, var(--warn) 70%, var(--ink-2));
border-left-color: var(--warn);
background: var(--warn-tint);
}
.coord-tool-row-warning--high,
.coord-tool-row-warning--critical {
color: color-mix(in srgb, var(--err) 70%, var(--ink-2));
border-left-color: var(--err);
background: color-mix(in srgb, var(--err) 12%, var(--panel-2));
}
.coord-tool-row-warning-redacted {
color: var(--ink-3);
font-size: 10px;
}
/* Storage-truncation indicator same convention as the interactive
UI's `.tool-output-truncated` pill (transparent bg, dim border,
small font) so the operator reads the affordance the same way on
both surfaces. Sibling node next to .coord-tool-row-result rather
than text-in-content so a future "best-effort JSON repair" pass
on the result body doesn't have to strip a marker string. */
.coord-tool-truncated {
display: inline-block;
margin-top: 4px;
margin-left: 6px;
padding: 1px 6px;
font-size: 10px;
font-family: var(--font-mono);
color: var(--ink-3);
background: transparent;
border: 1px solid var(--ink-3);
border-radius: 3px;
}
/* memory/recall calls are background metadata the audit trail is
useful but they crowd the tree on workstreams with heavy memory
usage. Dim the row by default; full opacity on hover so they
stay inspectable. Mirrors the interactive UI's metacog dim rule
(style.css `.ts-approval-tool[data-func-name="memory"]`). The
row stamps `data-tool-name` from item.func_name in coordinator.js
so this selector has something to match. */
.coord-tool-row[data-tool-name="memory"],
.coord-tool-row[data-tool-name="recall"] {
opacity: 0.55;
transition: opacity 120ms ease-out;
}
.coord-tool-row[data-tool-name="memory"]:hover,
.coord-tool-row[data-tool-name="memory"]:focus-within,
.coord-tool-row[data-tool-name="recall"]:hover,
.coord-tool-row[data-tool-name="recall"]:focus-within {
opacity: 1;
}
/* Tool result block paired under its row, mono pre-block. Capped
at 240px with internal scroll so a long tool output doesn't push
the rest of the chat off-screen. The interactive UI uses a
@@ -537,6 +618,43 @@
color: var(--ink-3);
}
/* User-message attachment pills rendered beneath the bubble for
live sends and on history replay. Mirrors the .msg-user-attach*
rules in turnstone/ui/static/style.css so coord and interactive
surfaces show the same affordance for attached files. */
.msg-user-attach {
display: flex;
flex-wrap: wrap;
gap: 4px;
margin-top: 6px;
}
.msg-user-attach-pill {
display: inline-flex;
align-items: center;
gap: 4px;
/* --panel (not --panel-2) the .msg bubble is already --panel-2,
so pulling the pill onto the alternate surface keeps it visible
against the bubble in both themes (WCAG 1.4.11 non-text contrast).
Mirrors the interactive UI's --bg-surface vs .msg --panel-2 split. */
background: var(--panel);
color: var(--ink-2);
border: 1px solid var(--hair);
border-radius: 4px;
padding: 2px 6px;
font-family: var(--font-mono);
font-size: 10px;
}
.msg-user-attach-icon {
font-size: 11px;
opacity: 0.7;
}
.msg-user-attach-name {
max-width: 200px;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
/* Mobile (<700px) — keep action targets ≥44px for WCAG 2.5.5. */
@media (max-width: 700px) {
.coord-tool-actions {
File diff suppressed because it is too large Load Diff
@@ -182,6 +182,36 @@
font-size: 11px;
line-height: 1.4;
}
/* Loading-state placeholder rendered while the bulk fetch is
in-flight. Same outer container as the real block so the row
height is stable when the real content swaps in. */
.ch-row .approval-block-loading {
color: var(--ink-3);
}
.ch-row .approval-loading-spin {
display: inline-block;
width: 8px;
height: 8px;
border-radius: 50%;
border: 1.5px solid var(--accent);
border-top-color: transparent;
margin-right: 5px;
animation: ts-spin 0.9s linear infinite;
vertical-align: middle;
}
/* Inline status note shown when the operator clicks Approve/Deny
but the call_id is stale (already resolved on another channel).
Quieter than a toast — the row is about to be replaced wholesale
by the refresh, so this only needs to bridge ~350ms. */
.ch-row .approval-stale-note {
font-size: 11px;
color: var(--ink-3);
font-style: italic;
margin-top: 4px;
}
@media (prefers-reduced-motion: reduce) {
.ch-row .approval-loading-spin { animation: none; }
}
/* Auto-approved pill — same 22px indent as .approval-block / .meta
so a child row showing both stacks cleanly. Lower visual weight
than the live approve/deny block (this is informational, not
+6 -4
View File
@@ -2791,7 +2791,8 @@ var _eogpTriggerEl = null;
function switchJudgeSection(section) {
var sections = document.querySelectorAll(".judge-section");
for (var i = 0; i < sections.length; i++) sections[i].style.display = "none";
var btns = document.querySelectorAll(".judge-section-btn");
var switcher = document.querySelector("#admin-judge .admin-subtab-switcher");
var btns = switcher ? switcher.querySelectorAll(".admin-subtab-btn") : [];
for (var i = 0; i < btns.length; i++) {
var isActive = btns[i].getAttribute("data-section") === section;
btns[i].classList.toggle("active", isActive);
@@ -2804,15 +2805,15 @@ function switchJudgeSection(section) {
// Arrow key navigation for judge sub-section tabs
(function () {
var switcher = document.querySelector(".judge-section-switcher");
var switcher = document.querySelector("#admin-judge .admin-subtab-switcher");
if (!switcher) return;
switcher.addEventListener("keydown", function (e) {
if (e.key !== "ArrowLeft" && e.key !== "ArrowRight") return;
var btns = switcher.querySelectorAll(".judge-section-btn");
var btns = switcher.querySelectorAll(".admin-subtab-btn");
var secs = [];
for (var i = 0; i < btns.length; i++)
secs.push(btns[i].getAttribute("data-section"));
var current = switcher.querySelector(".judge-section-btn.active");
var current = switcher.querySelector(".admin-subtab-btn.active");
var idx = secs.indexOf(current ? current.getAttribute("data-section") : "");
if (e.key === "ArrowRight") idx = (idx + 1) % secs.length;
else idx = (idx - 1 + secs.length) % secs.length;
@@ -2874,6 +2875,7 @@ function renderJudgeSettings() {
var html = "";
for (var i = 0; i < _judgeSettings.length; i++) {
var s = _judgeSettings[i];
if (s.key === "judge.model") continue;
var shortKey = s.key.replace("judge.", "");
var inputHtml = "";
var currentVal = s.value;
+215 -56
View File
@@ -87,33 +87,16 @@
<div id="view-home">
<!-- Persistent "start a new coordinator task" composer. Visibility is
gated on the admin.coordinator permission (same rule the existing
+coordinator header button + modal use). Shows a remediation
banner when the create endpoint would return 503 (no coordinator
model alias resolves). -->
+coordinator header button + modal use). Submission errors surface
inline via #home-coord-error; the create endpoint falls back to
the registry default model when ``coordinator.model_alias`` is
unset, so no proactive readiness probe is needed. -->
<section
id="coord-composer-panel"
class="home-panel"
style="display: none"
aria-label="Start a new orchestration task"
>
<div
id="coord-composer-503"
class="home-composer-banner"
role="status"
style="display: none"
>
Coordinator subsystem not configured —
<a
href="#"
onclick="
showAdmin();
switchAdminTab('models');
return false;
"
>open Admin → Models</a
>
to set <code>coordinator.model_alias</code> or a registry default.
</div>
<!-- Composer DOM is built by shared_static/composer.js into this
mount — stacked layout (textarea above, options toggle +
Start button below) with Name + Skill in an Options
@@ -122,9 +105,8 @@
<div
id="home-coord-error"
class="home-composer-error"
role="alert"
aria-live="assertive"
style="display: none"
role="status"
aria-live="polite"
></div>
</section>
@@ -165,6 +147,14 @@
<h2 class="home-section-title">
<span>Saved Coordinators</span>
<span id="saved-coord-count" class="home-section-count"></span>
<button
id="coord-delete-btn"
class="ws-delete-btn home-section-action"
onclick="startCoordDeleteMode()"
title="Delete coordinators"
>
<span aria-hidden="true">&#x1f5d1;</span> Delete
</button>
</h2>
<div
id="saved-coord-cards"
@@ -172,6 +162,61 @@
role="list"
aria-live="polite"
></div>
<div
id="coord-pagination"
class="pagination"
style="display: none"
role="navigation"
aria-label="Saved coordinators pagination"
>
<button
id="coord-page-prev"
type="button"
onclick="coordPagePrev()"
>
&#x25c4; Prev
</button>
<span id="coord-page-label" aria-live="polite" aria-atomic="true">
</span>
<button
id="coord-page-next"
type="button"
onclick="coordPageNext()"
>
Next &#x25ba;
</button>
</div>
<div id="coord-delete-bar" class="ws-delete-bar">
<span
class="ws-delete-count-label"
id="coord-delete-bar-count"
role="status"
aria-live="polite"
aria-atomic="true"
>0 selected</span
>
<button
class="ws-delete-cancel-btn"
onclick="cancelCoordDeleteMode()"
>
Cancel
</button>
<button
class="ws-delete-selectall-btn"
id="coord-delete-bar-select-all"
onclick="toggleCoordSelectAll()"
>
Select All
</button>
<button
class="ws-delete-bar-btn"
id="coord-delete-bar-delete"
onclick="confirmCoordDeleteSelection()"
disabled
>
Delete Selected
</button>
</div>
</section>
<!-- Cluster details — node list, always visible. The list is
@@ -820,13 +865,13 @@
<!-- Sub-panel switcher -->
<div
class="judge-section-switcher"
class="admin-subtab-switcher"
role="tablist"
aria-label="Judge sections"
>
<button
id="judge-tab-settings"
class="judge-section-btn active"
class="admin-subtab-btn active"
role="tab"
aria-selected="true"
aria-controls="judge-settings-section"
@@ -838,7 +883,7 @@
</button>
<button
id="judge-tab-heuristic"
class="judge-section-btn"
class="admin-subtab-btn"
role="tab"
aria-selected="false"
aria-controls="judge-heuristic-section"
@@ -850,7 +895,7 @@
</button>
<button
id="judge-tab-output-guard"
class="judge-section-btn"
class="admin-subtab-btn"
role="tab"
aria-selected="false"
aria-controls="judge-output-guard-section"
@@ -1733,37 +1778,106 @@
style="display: none"
>
<div class="admin-toolbar">
<span class="section-header">Models</span>
<button
id="model-sync-btn"
class="admin-action-btn admin-action-btn-ghost"
onclick="reloadModelNodes()"
title="Push model config to all cluster nodes"
>
Sync to Nodes
</button>
<button
class="admin-action-btn"
onclick="showCreateModelModal()"
>
+ Add Model
</button>
</div>
<div class="admin-colheaders models-grid" aria-hidden="true">
<span class="admin-col">ALIAS</span>
<span class="admin-col">MODEL</span>
<span class="admin-col">PROVIDER</span>
<span class="admin-col">CTX WINDOW</span>
<span class="admin-col">STATUS</span>
<span class="admin-col">ACTIONS</span>
<span class="section-header" style="margin: 0">MODELS</span>
</div>
<!-- Sub-panel switcher -->
<div
id="admin-models-table"
role="list"
aria-label="Model definitions"
aria-live="polite"
class="admin-subtab-switcher"
role="tablist"
aria-label="Models sections"
>
<div class="dashboard-empty">Loading...</div>
<button
id="models-tab-list"
class="admin-subtab-btn active"
role="tab"
aria-selected="true"
aria-controls="models-list-section"
tabindex="0"
data-section="models-list"
onclick="switchModelsSection('models-list')"
>
Definitions
</button>
<button
id="models-tab-roles"
class="admin-subtab-btn"
role="tab"
aria-selected="false"
aria-controls="models-roles-section"
tabindex="-1"
data-section="models-roles"
onclick="switchModelsSection('models-roles')"
>
Roles
</button>
</div>
<!-- Models list section -->
<div
id="models-list-section"
class="models-section"
role="tabpanel"
aria-labelledby="models-tab-list"
>
<div class="admin-toolbar" style="margin-bottom: 12px">
<span style="font-size: 13px; color: var(--fg-dim)"
>Model definitions used by sessions across the cluster</span
>
<button
id="model-sync-btn"
class="admin-action-btn admin-action-btn-ghost"
onclick="reloadModelNodes()"
title="Push model config to all cluster nodes"
>
Sync to Nodes
</button>
<button
class="admin-action-btn"
onclick="showCreateModelModal()"
>
+ Add Model
</button>
</div>
<div class="admin-colheaders models-grid" aria-hidden="true">
<span class="admin-col">ALIAS</span>
<span class="admin-col">MODEL</span>
<span class="admin-col">PROVIDER</span>
<span class="admin-col">CTX WINDOW</span>
<span class="admin-col">STATUS</span>
<span class="admin-col">ACTIONS</span>
</div>
<div
id="admin-models-table"
role="list"
aria-label="Model definitions"
aria-live="polite"
>
<div class="dashboard-empty">Loading...</div>
</div>
</div>
<!-- Roles section (Coordinator / Judge / future perception roles) -->
<div
id="models-roles-section"
class="models-section"
role="tabpanel"
aria-labelledby="models-tab-roles"
style="display: none"
>
<div style="margin-bottom: 12px">
<span style="font-size: 13px; color: var(--fg-dim)"
>Per-role model assignments. Empty = use the default
model.</span
>
</div>
<div
id="admin-models-roles-container"
aria-live="polite"
style="max-width: 720px"
>
<div class="dashboard-empty">Loading&hellip;</div>
</div>
</div>
</div>
@@ -3806,6 +3920,17 @@
<option value="llama.cpp">llama.cpp</option>
<option value="openai-compatible">Other OpenAI-compatible</option>
</select>
<label for="model-api-surface"
>API Surface
<span style="font-weight: 400; text-transform: none"
>(Chat Completions vs Responses)</span
></label
>
<select id="model-api-surface">
<option value="">Inherit (Chat Completions)</option>
<option value="chat">Chat Completions (pinned)</option>
<option value="responses">Responses API</option>
</select>
<label for="model-thinking-mode"
>Thinking Mode
<span style="font-weight: 400; text-transform: none"
@@ -3907,6 +4032,40 @@
</div>
</div>
<!-- Delete coordinators confirmation modal (batch) -->
<div
id="coord-delete-overlay"
class="ws-delete-modal-overlay"
style="display: none"
role="dialog"
aria-modal="true"
aria-labelledby="coord-delete-title"
>
<div id="coord-delete-box" class="ws-delete-modal-box">
<h3 id="coord-delete-title">Delete Coordinators</h3>
<div id="coord-delete-error" role="alert" aria-live="assertive"></div>
<p id="coord-delete-count"></p>
<div id="coord-delete-list" class="ws-delete-modal-list"></div>
<div id="coord-delete-buttons" class="ws-delete-modal-buttons">
<button
id="coord-delete-cancel-btn"
type="button"
onclick="cancelCoordDelete()"
>
Cancel
</button>
<button
id="coord-delete-confirm-btn"
class="ws-delete-confirm"
type="button"
onclick="confirmCoordDelete()"
>
Delete
</button>
</div>
</div>
</div>
<script src="/static/admin.js"></script>
<script src="/static/governance.js"></script>
<script src="/static/app.js"></script>
+109 -22
View File
@@ -88,27 +88,45 @@
text-transform: uppercase;
}
.home-composer-banner {
background: var(--bg-surface);
border: 1px solid var(--yellow);
border-left-width: 3px;
border-radius: var(--radius-sm);
padding: 8px 10px;
color: var(--fg-bright);
font-size: 12px;
}
.home-composer-banner a {
color: var(--accent);
text-decoration: underline;
text-decoration-thickness: 2px;
}
.home-composer-error {
/* Always rendered (no display toggle in JS) so toggling validation
messages doesn't reflow the active-coordinators list below. The
min-height holds a single 12px line + padding so an empty state
reserves the same space the rendered error will occupy. */
min-height: 20px;
color: var(--red);
font-size: 12px;
padding: 4px 2px 0;
}
/* Drag-over feedback for the home coord composer wired by Composer's
dragDrop option (dropClass: home-coord-drop) on the composer mount.
Mirrors the dashed-outline affordance on coord-main so users see the
same drop visual the in-coord composer uses. */
#home-coord-composer-mount {
position: relative;
}
#home-coord-composer-mount.home-coord-drop {
outline: 2px dashed var(--accent);
outline-offset: -6px;
border-radius: var(--radius);
}
/* Cap chip filename width inside the home composer so a long-name
attachment doesn't push the strip wider than the textarea or wrap
unpredictably across multiple rows. shared_static/chat.css defines
.composer-chip / .composer-chip-size / .composer-chip-remove but
leaves .composer-chip-name unstyled the span just inherits the
.composer-chip font with no width cap, fine for chat-pane width but
too loose for the narrower home column. Apply the same ellipsis
cap the .msg-user-attach-pill rule uses on the user bubble. */
#home-coord-composer-mount .composer-chip-name {
max-width: 200px;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.home-section {
display: flex;
flex-direction: column;
@@ -133,6 +151,17 @@
font-size: 11px;
font-variant-numeric: tabular-nums;
}
/* Right-align action buttons (e.g. Saved Coordinators "Delete") inside
.home-section-title without breaking the count's natural left position. */
.home-section-title .home-section-action {
margin-left: auto;
}
/* Saved Coordinators reuses the existing .pagination control (see the
"Pagination" block below) the visible page is capped at
COORD_PAGE_SIZE so Select-All fan-out is bounded. Pagination is
hidden in delete mode and when there's only one page (see
_renderCoordPagination). */
.home-coord-list {
border: 1px solid var(--border);
@@ -2306,15 +2335,15 @@ textarea.skill-content-area {
}
/* ==========================================================================
Judge sub-section tabs
Admin sub-section tabs (used by Judge + Models tabs)
========================================================================== */
.judge-section-switcher {
.admin-subtab-switcher {
display: flex;
gap: 8px;
margin: 12px 0 16px;
border-bottom: 1px solid var(--border-strong);
}
.judge-section-btn {
.admin-subtab-btn {
padding: 6px 14px;
background: none;
border: none;
@@ -2327,14 +2356,14 @@ textarea.skill-content-area {
color 0.15s,
border-color 0.15s;
}
.judge-section-btn:hover {
.admin-subtab-btn:hover {
color: var(--fg);
}
.judge-section-btn.active {
.admin-subtab-btn.active {
border-bottom-color: var(--accent);
color: var(--fg);
}
.judge-section-btn:focus-visible {
.admin-subtab-btn:focus-visible {
outline: 2px solid var(--accent);
outline-offset: -2px;
}
@@ -3797,6 +3826,64 @@ textarea.skill-content-area {
letter-spacing: 0.02em;
}
/* Models → Roles sub-tab rows */
.model-role-row {
padding: 12px 0;
border-bottom: 1px solid var(--border-strong);
}
.model-role-row:last-child {
border-bottom: none;
}
.model-role-head {
display: flex;
align-items: center;
gap: 8px;
margin-bottom: 4px;
}
.model-role-label {
font-size: 13px;
font-weight: 600;
color: var(--fg);
}
.model-role-desc {
font-size: 11px;
color: var(--fg-dim);
margin-bottom: 8px;
}
.model-role-controls {
display: grid;
grid-template-columns: minmax(240px, 1fr) minmax(180px, auto);
column-gap: 16px;
row-gap: 8px;
align-items: end;
}
.model-role-control {
display: flex;
flex-direction: column;
gap: 4px;
min-width: 0;
}
.model-role-control-label {
font-size: 10px;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--fg-dim);
}
.model-role-control select {
width: 100%;
padding: 4px 8px;
background: var(--bg);
border: 1px solid var(--border-strong);
color: var(--fg);
border-radius: 3px;
font-family: var(--font-ui);
font-size: 12px;
}
.model-role-control select:focus-visible {
outline: 2px solid var(--accent);
outline-offset: 1px;
}
/* Modal section divider for field groups */
.modal-section-divider {
font-family: var(--font-ui);
@@ -3844,7 +3931,7 @@ textarea.skill-content-area {
.admin-btn-danger,
.admin-btn-caution,
.admin-btn-action,
.judge-section-btn {
.admin-subtab-btn {
transition: none;
}
.settings-toggle-slider,
+268
View File
@@ -0,0 +1,268 @@
"""Strategy interface for delivering child workstream lifecycle events.
Two implementations bind to the unified :class:`ChildrenRegistry`:
- :class:`SameNodeChildSource` subscribes to a :class:`SessionManager`'s
state-change callbacks. In-process, no transport. For interactive
workstreams that spawn children locally (no cluster routing).
- :class:`ClusterChildSource` subscribes to a :class:`ClusterCollector`'s
listener channel and runs a daemon thread that drains the queue and
pushes events to the sink. For coordinator workstreams whose
children are routed across the cluster by hash bucket.
The strategy doesn't translate events into UI-shaped payloads; that's
the sink's job. This split keeps the strategy generic across kinds and
lets the consumer (e.g. ``CoordinatorAdapter._dispatch_child_event``)
own the per-kind translation.
Sink signature: ``Callable[[dict[str, Any]], None]``. The sink is
responsible for filtering by registry membership; strategies push raw
events without registry-side filtering so the sink can decide whether
to act based on its own state.
"""
from __future__ import annotations
import queue
import threading
from typing import TYPE_CHECKING, Any, Protocol
from turnstone.core.log import get_logger
if TYPE_CHECKING:
from collections.abc import Callable, Iterable
from turnstone.core.children_registry import ChildrenRegistry
from turnstone.core.workstream import WorkstreamState
log = get_logger(__name__)
class ChildSource(Protocol):
"""Subscription strategy for child workstream lifecycle events."""
def start(self, sink: Callable[[dict[str, Any]], None]) -> None:
"""Begin delivering events to ``sink``. Idempotent."""
def shutdown(self) -> None:
"""Stop the strategy. Idempotent; safe to call multiple times."""
class _CollectorProtocol(Protocol):
"""Subset of :class:`ClusterCollector` that :class:`ClusterChildSource` consumes.
Defined here (not imported) to keep ``turnstone/core/`` free of
``turnstone/console/`` imports AND to let test fakes satisfy the
type signature without subclassing the real collector.
"""
def get_snapshot_and_register(self, q: queue.Queue[dict[str, Any]]) -> dict[str, Any]:
"""Register ``q`` and return the current snapshot."""
def unregister_listener(self, q: queue.Queue[dict[str, Any]]) -> None:
"""Drop ``q`` from the listener set."""
class _ManagerProtocol(Protocol):
"""Subset of :class:`SessionManager` that :class:`SameNodeChildSource` consumes.
Lets test fakes participate in the strategy's typed surface
without forcing a full SessionManager construction.
"""
def subscribe_to_state(self, callback: Callable[[str, WorkstreamState], None]) -> None:
"""Register ``callback`` for state-change events."""
def unsubscribe_from_state(self, callback: Callable[[str, WorkstreamState], None]) -> None:
"""Remove a previously-registered ``callback``."""
class SameNodeChildSource:
"""In-process child events via :class:`SessionManager` state observer.
Subscribes to the manager's state-change callbacks (registered via
:meth:`SessionManager.subscribe_to_state`). For each transition on
a workstream that's a known child (per the
:class:`ChildrenRegistry` reverse index), synthesises a
cluster-state-shaped event and pushes it to the sink.
Used by interactive workstreams when they gain spawn capability
children live on the same node as the parent, so no cluster routing
is needed and event fan-out is in-process.
"""
def __init__(
self,
manager: _ManagerProtocol,
registry: ChildrenRegistry,
) -> None:
self._manager = manager
self._registry = registry
self._sink: Callable[[dict[str, Any]], None] | None = None
self._callback: Callable[[str, WorkstreamState], None] | None = None
def start(self, sink: Callable[[dict[str, Any]], None]) -> None:
if self._callback is not None:
return # idempotent — already started
self._sink = sink
def _on_state(ws_id: str, state: WorkstreamState) -> None:
sink_fn = self._sink
if sink_fn is None:
return
# Cheap pre-filter: skip dispatch for transitions on
# workstreams that aren't children of any in-memory parent.
# ``has_children`` is a lock-free dict-truthiness read; the
# ``parent_for`` call below would otherwise acquire the
# registry lock on every state change even when no
# children exist (the steady state for an interactive
# manager). The cluster strategy can't pre-filter because
# the collector queue carries all events; here we have the
# information to skip the synthesis entirely.
if not self._registry.has_children():
return
if self._registry.parent_for(ws_id) is None:
return
# ``pending_approval_detail`` deliberately omitted — the
# field was removed from cluster_state end-to-end in the
# Stage 3 cleanup pass. Approval items arrive via bulk
# fetch; verdicts via the explicit intent_verdict event;
# resolution via approval_resolved.
event = {
"type": "cluster_state",
"ws_id": ws_id,
"state": state.value,
"node_id": "",
"tokens": 0,
"activity_state": "",
}
try:
sink_fn(event)
except Exception:
log.debug("same_node_child_source.sink_failed", exc_info=True)
self._callback = _on_state
self._manager.subscribe_to_state(_on_state)
def shutdown(self) -> None:
cb = self._callback
if cb is None:
return
try:
self._manager.unsubscribe_from_state(cb)
except Exception:
log.debug("same_node_child_source.unsubscribe_failed", exc_info=True)
self._callback = None
self._sink = None
class ClusterChildSource:
"""Cross-node child events via :class:`ClusterCollector` subscription.
Refactor of the existing fan-out machinery from
``CoordinatorAdapter`` (was ``_collector_queue`` +
``_fanout_thread`` + ``_fanout_loop``). Subscribes as a listener on
the collector's broadcast channel and runs a daemon thread that
drains the queue, pushing each event to the sink.
On :meth:`start`, also primes the registry from the collector's
snapshot so a parent that re-installs after a console restart sees
its already-live children without waiting for the next state tick.
The ``parents_provider`` callback returns the set of in-memory
parent ws_ids for snapshot filtering only children whose parent
is currently installed get merged.
"""
def __init__(
self,
collector: _CollectorProtocol,
registry: ChildrenRegistry,
*,
parents_provider: Callable[[], Iterable[str]],
) -> None:
self._collector = collector
self._registry = registry
self._parents_provider = parents_provider
self._sink: Callable[[dict[str, Any]], None] | None = None
self._queue: queue.Queue[dict[str, Any]] | None = None
self._thread: threading.Thread | None = None
self._stop = threading.Event()
def start(self, sink: Callable[[dict[str, Any]], None]) -> None:
if self._thread is not None and self._thread.is_alive():
return # idempotent — already started
self._sink = sink
self._queue = queue.Queue(maxsize=1000)
snapshot = self._collector.get_snapshot_and_register(self._queue)
self._prime_from_snapshot(snapshot)
self._stop.clear()
t = threading.Thread(
target=self._loop,
name="cluster-child-source",
daemon=True,
)
self._thread = t
t.start()
def shutdown(self) -> None:
self._stop.set()
t = self._thread
q = self._queue
coll = self._collector
self._thread = None
self._queue = None
if coll is not None and q is not None:
try:
coll.unregister_listener(q)
except Exception:
log.debug(
"cluster_child_source.unregister_listener_failed",
exc_info=True,
)
if t is not None:
t.join(timeout=2.0)
self._sink = None
def _loop(self) -> None:
q = self._queue
if q is None:
return
while not self._stop.is_set():
try:
event = q.get(timeout=1.0)
except queue.Empty:
continue
sink = self._sink
if sink is None:
continue
try:
sink(event)
except Exception:
log.debug("cluster_child_source.dispatch_failed", exc_info=True)
def _prime_from_snapshot(self, snapshot: dict[str, Any]) -> None:
"""Populate the registry from a collector snapshot.
For every workstream in the snapshot whose ``parent_ws_id``
names a currently-installed parent (per ``parents_provider``),
merge it into the registry. Caller-installed parents that
appear in the snapshot are seeded; unknown parents are skipped
they'll be picked up by the live fan-out path once their
``ws_created`` event arrives.
"""
nodes = snapshot.get("nodes", []) if isinstance(snapshot, dict) else []
if not nodes:
return
known_parents = set(self._parents_provider())
if not known_parents:
return
by_parent: dict[str, list[str]] = {}
for node in nodes:
for entry in node.get("workstreams", []) or []:
parent = entry.get("parent_ws_id") or ""
child_id = entry.get("id") or ""
if not parent or not child_id or parent not in known_parents:
continue
by_parent.setdefault(parent, []).append(child_id)
for parent, kids in by_parent.items():
self._registry.merge_children(parent, kids)
+177
View File
@@ -0,0 +1,177 @@
"""Universal parent → children registry for SessionManager.
Pure data + lookups; no IO, no transport. Lifted from
:class:`turnstone.console.coordinator_adapter.CoordinatorAdapter` where
it lived bound to the coordinator kind. The lift is what lets the
``ChildSource`` strategies (Step 2) plug into a single shared primitive
regardless of whether children are local (interactive) or cluster-routed
(coordinator).
Storage rebuild and snapshot priming happen in the caller typically
the ``ChildSource`` implementation that owns the relevant transport.
The registry exposes :meth:`merge_children` for bulk seeding so callers
that compute child id lists from any source can feed them in without
the registry needing to know about storage shapes or collector
snapshots.
Threading: every public method is internally locked. Helpers suffixed
``_locked`` require the caller to already hold :attr:`_lock`.
"""
from __future__ import annotations
import threading
from typing import TYPE_CHECKING, Any
if TYPE_CHECKING:
from collections.abc import Iterable
class ChildrenRegistry:
"""Tracks parent → children + reverse lookup for in-memory parents.
The forward index (``_children``) is parent_ws_id set of child ws_ids.
The reverse index (``_child_to_parent``) is child_ws_id parent_ws_id.
The presence map (``_active``) is parent_ws_id UI ref, used by the
dispatch path to atomically check-and-route in one lock acquisition.
Closed / deleted children stay in the registry until their owning
parent is uninstalled the tree UI keeps rendering them grayed out;
state authority lives in storage, not here.
"""
def __init__(self) -> None:
self._children: dict[str, set[str]] = {}
self._child_to_parent: dict[str, str] = {}
self._lock = threading.Lock()
self._active: dict[str, Any] = {}
# ------------------------------------------------------------------
# Lifecycle — install / uninstall a parent
# ------------------------------------------------------------------
def install(self, parent_ws_id: str, ui: Any) -> None:
"""Seed the forward set + presence map for a new parent.
Idempotent re-installing re-points the UI but leaves the
existing child set intact. Mirrors the original
``_install_coord_registry`` semantics so a coordinator that
rehydrates after a crash doesn't lose its known-children.
"""
with self._lock:
self._children.setdefault(parent_ws_id, set())
self._active[parent_ws_id] = ui
def uninstall(self, parent_ws_id: str) -> None:
"""Drop a parent: forward set, reverse-index entries, presence.
No-op if the parent is unknown. Used by close / eviction paths.
"""
with self._lock:
self._uninstall_locked(parent_ws_id)
# ------------------------------------------------------------------
# Mutation — register children under a parent
# ------------------------------------------------------------------
def add_child(self, parent_ws_id: str, child_ws_id: str) -> Any | None:
"""Register a child under a parent. Returns parent's UI or None.
Returns the parent's UI on success (so the dispatch path can
atomically check-and-route in one lock acquisition). Returns
``None`` if the parent isn't installed (concurrent close /
eviction) or if the child is already registered (duplicate
ws_created from the cluster fan-out).
"""
with self._lock:
ui = self._active.get(parent_ws_id)
if ui is None:
return None
existing = self._children.setdefault(parent_ws_id, set())
if child_ws_id in existing:
return None
existing.add(child_ws_id)
self._child_to_parent[child_ws_id] = parent_ws_id
return ui
def merge_children(self, parent_ws_id: str, child_ws_ids: Iterable[str]) -> None:
"""Bulk-merge child_ids under a parent. Idempotent.
Sole bulk write-path used by both storage-seeded rebuilds and
snapshot-seeded priming so reverse-index ordering invariants
hold regardless of which seed source races first.
"""
with self._lock:
self._merge_locked(parent_ws_id, child_ws_ids)
# ------------------------------------------------------------------
# Lookups
# ------------------------------------------------------------------
def parent_for(self, child_ws_id: str) -> str | None:
"""Reverse lookup: which parent owns this child? O(1)."""
with self._lock:
return self._child_to_parent.get(child_ws_id)
def has_children(self) -> bool:
"""Lock-free fast path: any child registered under any parent?
Reads ``bool(self._child_to_parent)`` without taking the lock.
Dict-truthiness is a single GIL-atomic read, so callers on
the hot state-broadcast path can short-circuit without paying
the lock acquisition when the registry is empty (the steady
state for an interactive manager today). The answer is best-
effort if a child is added concurrently with the read the
caller may falsely return ``False``, but the next state event
will pick up the change correctly.
"""
return bool(self._child_to_parent)
def children_of(self, parent_ws_id: str) -> list[str]:
"""Snapshot copy of the parent's child ws_ids.
Returned list is a copy so callers can iterate without holding
the registry lock during per-child work. A mutation racing with
the snapshot either lands before (included) or after (excluded)
both outcomes are safe for cascade-style dispatch.
"""
with self._lock:
child_set = self._children.get(parent_ws_id)
return list(child_set) if child_set else []
def ui_for(self, parent_ws_id: str) -> Any | None:
"""Look up the UI registered for a parent."""
with self._lock:
return self._active.get(parent_ws_id)
def parents(self) -> list[str]:
"""Snapshot copy of installed parent ws_ids."""
with self._lock:
return list(self._active)
# ------------------------------------------------------------------
# Locked helpers — caller must hold ``self._lock``
# ------------------------------------------------------------------
def _merge_locked(self, parent_ws_id: str, child_ws_ids: Iterable[str]) -> None:
"""Idempotent merge under caller's lock. Empty/falsy ids skipped."""
existing = self._children.setdefault(parent_ws_id, set())
for cid in child_ws_ids:
if cid and cid not in existing:
existing.add(cid)
self._child_to_parent[cid] = parent_ws_id
def _uninstall_locked(self, parent_ws_id: str) -> None:
"""Pop forward set + presence + own reverse-index entries.
Defensive: only clears reverse entries that still point at
``parent_ws_id``. Schema-shaped reassignments (rare but
possible) shouldn't orphan the new owner's entry.
"""
child_set = self._children.pop(parent_ws_id, None)
self._active.pop(parent_ws_id, None)
if child_set is None:
return
for cid in child_set:
if self._child_to_parent.get(cid) == parent_ws_id:
self._child_to_parent.pop(cid, None)
+205
View File
@@ -0,0 +1,205 @@
"""Shared history-replay decoration helpers.
Both surfaces that build a history wire payload interactive's SSE
``_build_history`` and the lifted ``make_history_handler`` REST
endpoint need the same audit-trail data attached to each
``tool_calls`` entry: the persisted intent verdict (``intent_verdicts``
table) and the output-guard assessment (``output_assessments`` table).
Centralising the lookup + decoration here keeps the two surfaces from
drifting on which fields ship to the client and how they're shaped.
The shared helpers also let us project only the fields the UI actually
renders, dropping redundant ones (``call_id``/``func_name`` already
carried on ``tc.id``/``tc.name``) so the wire payload stays tight.
All functions are pure I/O or pure transforms safe to call from
either an async caller (via ``asyncio.to_thread``) or a sync hook.
"""
from __future__ import annotations
import json
from typing import Any
from turnstone.core.log import get_logger
log = get_logger(__name__)
# Tool results are clamped at this length per row at storage time
# (see ``session.py``'s ``store_text = raw_output[:TOOL_RESULT_STORAGE_CAP]``).
# Keeping the constant here lets the truncation flag detection in
# ``decorate_history_messages`` stay in sync without a magic number
# duplicated across server.py / session.py.
#
# Raised from 2000 → 10000 because a 2000-char clip routinely cut
# the body of a single grep / file read mid-line, leaving the
# historical record useless for retrospective debugging. FTS5
# index + row size grow proportionally; the per-tool upper bound is
# still bounded upstream by ``_truncate_output``'s context-budget
# clamp (so a single huge result can't blow past the live context
# window).
TOOL_RESULT_STORAGE_CAP = 10000
def load_verdict_indexes(
ws_id: str,
) -> tuple[dict[str, dict[str, Any]], dict[str, dict[str, Any]]]:
"""Bulk-load intent verdicts and output assessments for a workstream.
Returns ``(verdicts_by_call_id, assessments_by_call_id)``. Both
tables are indexed by ws_id so the queries are O(rows-for-ws); the
DESC ordering plus first-seen-wins dedupe leaves the newest
verdict per call_id (LLM upgrade beats heuristic when both exist).
Pure storage I/O safe to run in ``asyncio.to_thread`` from an
async caller. Returns empty dicts when storage is unavailable or
the lookup raises (best-effort: replay must never block on
audit-trail decoration).
"""
verdicts_by_call_id: dict[str, dict[str, Any]] = {}
assessments_by_call_id: dict[str, dict[str, Any]] = {}
if not ws_id:
return verdicts_by_call_id, assessments_by_call_id
try:
from turnstone.core.storage._registry import get_storage
storage = get_storage()
if storage is None:
return verdicts_by_call_id, assessments_by_call_id
for v in storage.list_intent_verdicts(ws_id=ws_id, limit=10000):
cid = v.get("call_id") or ""
if cid and cid not in verdicts_by_call_id:
verdicts_by_call_id[cid] = v
for a in storage.list_output_assessments(ws_id=ws_id, limit=10000):
cid = a.get("call_id") or ""
if cid and cid not in assessments_by_call_id:
assessments_by_call_id[cid] = a
except Exception:
# Missing storage / migration drift / driver error must not
# block replay — degrade to an unannotated history.
log.debug(
"verdict/assessment lookup failed; replay continues unannotated",
exc_info=True,
)
return verdicts_by_call_id, assessments_by_call_id
def build_verdict_payload(vrow: dict[str, Any]) -> dict[str, Any] | None:
"""Project a stored ``intent_verdicts`` row into the wire shape.
Returns ``None`` when the verdict is the unflagged baseline
(``risk_level == "none"``) the client's ``renderVerdictBadge``
helper would suppress those anyway, so skipping at the wire layer
keeps the payload tight on long workstreams.
Drops ``call_id`` and ``func_name`` from the wire payload they're
already carried on the parent ``tc.id`` / ``tc.name`` fields.
Ships ``reasoning`` for either tier when the row has non-empty
prose (heuristic rules in this project DO write meaningful
rationales e.g. ``policy.py`` emits structured reasoning per
matched pattern). ``judge_model`` rides through so the batch tier
badge can render `` llm:claude-haiku-4`` on history-only batches
rather than the bare `` llm`` label.
"""
if (vrow.get("risk_level") or "none") == "none":
return None
payload: dict[str, Any] = {
"risk_level": vrow.get("risk_level", "medium"),
"recommendation": vrow.get("recommendation", "review"),
"confidence": vrow.get("confidence", 0.0),
"intent_summary": vrow.get("intent_summary", ""),
"tier": vrow.get("tier", "heuristic"),
}
if vrow.get("reasoning"):
payload["reasoning"] = vrow.get("reasoning", "")
judge_model = vrow.get("judge_model") or ""
if judge_model:
payload["judge_model"] = judge_model
return payload
def build_output_assessment_payload(arow: dict[str, Any]) -> dict[str, Any] | None:
"""Project a stored ``output_assessments`` row into the wire shape.
Returns ``None`` when the assessment is the unflagged baseline
(``risk_level == "none"``) same skip-on-clean pattern as
:func:`build_verdict_payload`.
Decodes ``flags`` from its JSON string form here so the client
never has to parse twice. Falls back to an empty list on bad JSON
rather than raising the rest of the assessment is still useful.
"""
if (arow.get("risk_level") or "none") == "none":
return None
flags_raw = arow.get("flags") or "[]"
try:
flags = json.loads(flags_raw) if isinstance(flags_raw, str) else flags_raw
except (ValueError, TypeError):
flags = []
return {
"risk_level": arow.get("risk_level", "none"),
"flags": flags if isinstance(flags, list) else [],
"redacted": bool(arow.get("redacted", 0)),
}
def decorate_tool_call(
tc: dict[str, Any],
verdicts_by_call_id: dict[str, dict[str, Any]],
assessments_by_call_id: dict[str, dict[str, Any]],
) -> None:
"""Mutate ``tc`` in place, attaching ``verdict`` / ``output_assessment``.
Works on either tool_call shape:
- OpenAI format (``{id, function: {name, arguments}}``) used by
``/history`` REST.
- Flattened format (``{id, name, arguments}``) used by SSE replay.
Both carry ``id`` at the top level, which is the only field this
helper reads. No-ops cleanly when the call_id has no matching
row (unflagged tools stay clean).
"""
call_id = tc.get("id", "") or ""
if not call_id:
return
vrow = verdicts_by_call_id.get(call_id)
if vrow is not None:
verdict = build_verdict_payload(vrow)
if verdict is not None:
tc["verdict"] = verdict
arow = assessments_by_call_id.get(call_id)
if arow is not None:
assessment = build_output_assessment_payload(arow)
if assessment is not None:
tc["output_assessment"] = assessment
def decorate_history_messages(
messages: list[dict[str, Any]],
verdicts_by_call_id: dict[str, dict[str, Any]],
assessments_by_call_id: dict[str, dict[str, Any]],
) -> None:
"""Mutate a list of OpenAI-format messages, decorating tool_calls.
Used by the ``/history`` REST endpoint after ``load_messages``
returns. For each assistant message with ``tool_calls``, runs
:func:`decorate_tool_call` on every entry. For each tool message
whose content hits the storage cap, sets ``truncated: True`` so
the client can render the "… truncated in storage" pill.
Pure transform no I/O. Async callers should pre-load the
indexes via :func:`load_verdict_indexes` (in ``to_thread``) and
pass them in.
"""
for msg in messages:
role = msg.get("role")
if role == "assistant":
tcs = msg.get("tool_calls")
if isinstance(tcs, list):
for tc in tcs:
if isinstance(tc, dict):
decorate_tool_call(tc, verdicts_by_call_id, assessments_by_call_id)
elif role == "tool":
content = msg.get("content")
if isinstance(content, str) and len(content) >= TOOL_RESULT_STORAGE_CAP:
msg["truncated"] = True
+31
View File
@@ -757,6 +757,37 @@ def search_structured_memories(
return []
def list_visible_structured_memories(
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 100,
) -> list[dict[str, str]]:
"""Single-query union across visible (scope, scope_id) pairs."""
try:
return get_storage().list_visible_structured_memories(
scopes, mem_type=mem_type, limit=limit
)
except Exception:
log.warning("Failed to list visible structured memories", exc_info=True)
return []
def search_visible_structured_memories(
query: str,
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 20,
) -> list[dict[str, str]]:
"""OR-of-terms search joined with a single visibility OR-group."""
try:
return get_storage().search_visible_structured_memories(
query, scopes, mem_type=mem_type, limit=limit
)
except Exception:
log.warning("Failed to search visible structured memories", exc_info=True)
return []
def touch_structured_memories(keys: list[tuple[str, str, str]]) -> int:
"""Batch-touch memories (bump last_accessed, increment access_count).
+40
View File
@@ -32,6 +32,9 @@ class MetricsCollector:
# counters (continued)
self._ratelimit_rejects: int = 0 # counter: total 429 responses
self._evictions: int = 0 # counter: workstreams evicted
# node_models publish (heartbeat-loop refresh of node_metadata.models)
self._node_models_publish_written: int = 0
self._node_models_publish_skipped: int = 0
# judge metrics
self._judge_verdicts: dict[tuple[str, str], int] = defaultdict(int)
self._judge_latency: dict[str, Any] = {
@@ -104,6 +107,22 @@ class MetricsCollector:
with self._lock:
self._evictions += 1
def record_node_models_publish(self, *, written: bool) -> None:
"""Record one heartbeat-loop attempt to refresh ``node_metadata.models``.
``written=True`` means the projected payload differed from the
cached one and we ran an UPSERT. ``written=False`` means the
cache short-circuited the call. In a stable cluster the
skipped:written ratio runs ~100:1 a sustained drop in that
ratio is the signal an operator wants (backend health flapping
or a runaway model-reload loop).
"""
with self._lock:
if written:
self._node_models_publish_written += 1
else:
self._node_models_publish_skipped += 1
def set_judge_enabled(self, enabled: bool) -> None:
with self._lock:
self._judge_enabled = enabled
@@ -170,6 +189,8 @@ class MetricsCollector:
judge_verdicts = dict(self._judge_verdicts)
judge_latency = dict(self._judge_latency)
judge_enabled = self._judge_enabled
node_models_publish_written = self._node_models_publish_written
node_models_publish_skipped = self._node_models_publish_skipped
# turnstone_build_info
lines.append("# HELP turnstone_build_info Server version and model info")
@@ -277,6 +298,25 @@ class MetricsCollector:
evictions,
)
# turnstone_node_models_publish_total — split by outcome so an
# operator can compute hit-rate as
# ``rate(skipped) / (rate(skipped) + rate(written))``. In a
# stable cluster this ratio sits very close to 1.0; sustained
# dips signal backend health flapping or reload churn.
lines.append(
"# HELP turnstone_node_models_publish_total "
"node_metadata.models refresh attempts by outcome"
)
lines.append("# TYPE turnstone_node_models_publish_total counter")
lines.append(
f'turnstone_node_models_publish_total{{outcome="written"}} '
f"{node_models_publish_written}"
)
lines.append(
f'turnstone_node_models_publish_total{{outcome="skipped"}} '
f"{node_models_publish_skipped}"
)
# turnstone_judge_enabled
gauge(
"turnstone_judge_enabled",
+25 -5
View File
@@ -44,6 +44,19 @@ class ModelConfig:
server_compat: dict[str, Any] = field(default_factory=dict)
def _api_surface_of(cfg: ModelConfig) -> str | None:
"""Extract the operator-pinned api_surface from *cfg*, or ``None``.
Used both at provider-cache lookup time and at reload-eviction time so the
two sites stay in sync. Returns ``None`` when the field is absent, blank,
or not a string matching the "inherit provider default" semantics.
"""
raw = cfg.server_compat.get("api_surface") if isinstance(cfg.server_compat, dict) else None
if isinstance(raw, str) and raw.strip():
return raw
return None
# ---------------------------------------------------------------------------
# Registry
# ---------------------------------------------------------------------------
@@ -127,7 +140,9 @@ class ModelRegistry:
raise ValueError(f"Unknown model alias: {alias}")
if alias not in self._providers:
cfg = self._models[alias]
self._providers[alias] = create_provider(cfg.provider)
self._providers[alias] = create_provider(
cfg.provider, api_surface=_api_surface_of(cfg)
)
return self._providers[alias]
def get_config(self, alias: str) -> ModelConfig:
@@ -257,13 +272,18 @@ class ModelRegistry:
if hasattr(client, "close"):
client.close()
del self._clients[alias]
# Providers are keyed on alias but only depend on
# ``cfg.provider`` — drop only when the provider string
# changed or the alias was removed.
# Providers are keyed on alias and depend on (cfg.provider,
# cfg.server_compat["api_surface"]) — drop when either changes
# or the alias was removed.
for alias in list(self._providers.keys()):
old_cfg = old_models.get(alias)
new_cfg = self._models.get(alias)
if new_cfg is None or old_cfg is None or old_cfg.provider != new_cfg.provider:
if (
new_cfg is None
or old_cfg is None
or old_cfg.provider != new_cfg.provider
or _api_surface_of(old_cfg) != _api_surface_of(new_cfg)
):
del self._providers[alias]
def shutdown(self) -> None:
+38 -3
View File
@@ -33,7 +33,9 @@ __all__ = [
"lookup_model_capabilities",
]
# Singleton instances (stateless, safe to share)
# Singleton instances (stateless, safe to share). ``_openai_provider``
# is reused for both cloud OpenAI and ``openai-compatible`` with
# ``api_surface="responses"`` — see the ``create_provider`` docstring.
_provider_lock = threading.Lock()
_openai_provider = OpenAIResponsesProvider()
_openai_compat_provider = OpenAIChatCompletionsProvider()
@@ -41,12 +43,45 @@ _anthropic_provider: LLMProvider | None = None
_google_provider: LLMProvider | None = None
def create_provider(provider_name: str) -> LLMProvider:
"""Return a provider adapter for the given provider name. Thread-safe."""
_VALID_API_SURFACES = ("chat", "responses")
def create_provider(
provider_name: str,
*,
api_surface: str | None = None,
) -> LLMProvider:
"""Return a provider adapter for the given provider name. Thread-safe.
*api_surface* selects the OpenAI-compatible API surface for
``provider_name="openai-compatible"``:
- ``"chat"`` (default) Chat Completions (vLLM, llama.cpp, SGLang).
- ``"responses"`` Responses API (commercial OpenAI-compat
endpoints like Mistral cloud, or local servers that expose the
Responses surface).
Ignored for non-OpenAI providers. ``provider_name="openai"`` always
uses the Responses API regardless of *api_surface*.
Note: the ``OpenAIResponsesProvider`` singleton is reused for both
cloud OpenAI and ``openai-compatible`` + responses, so its
``provider_name`` reports ``"openai"`` even when serving an
openai-compatible config. Code that needs to distinguish the two
must read ``ModelConfig.provider`` and ``server_compat["api_surface"]``
rather than ``provider.provider_name``.
"""
global _anthropic_provider, _google_provider # noqa: PLW0603
if provider_name == "openai":
return _openai_provider
if provider_name == "openai-compatible":
normalised = (api_surface or "").strip().lower()
if normalised and normalised not in _VALID_API_SURFACES:
raise ValueError(
f"Unknown api_surface: {api_surface!r}. Supported: {', '.join(_VALID_API_SURFACES)}"
)
if normalised == "responses":
return _openai_provider
return _openai_compat_provider
if provider_name == "anthropic":
with _provider_lock:
+44 -10
View File
@@ -1,14 +1,20 @@
"""Server compatibility profiles for OpenAI-compatible backends.
Different local model servers (vLLM, llama.cpp, SGLang) need different
request shaping. This module separates two concerns:
request shaping. This module separates three concerns:
1. **Model capabilities** ``thinking_mode`` and ``thinking_param`` are
properties of the *model* (Gemma thinks, Llama doesn't). These go
into the ``capabilities`` dict and flow through ``ModelCapabilities``
so the provider can act on them (just like Anthropic's thinking mode).
2. **Server workarounds** ``extra_body`` overrides like
2. **API surface** ``api_surface`` selects which OpenAI-compatible
API surface the provider talks to: ``"chat"`` (Chat Completions,
the default) or ``"responses"`` (Responses API, native reasoning).
Stored under ``server_compat`` because it's an endpoint property,
not a model property.
3. **Server workarounds** ``extra_body`` overrides like
``skip_special_tokens=false`` are properties of the *server* (vLLM
bug workaround). These stay in ``server_compat`` and get merged
into the request's ``extra_body`` at call time.
@@ -82,6 +88,23 @@ _PROFILES: dict[str, dict[str, Any]] = {
"server_type": "vllm",
},
},
"vllm-mistral-medium": {
# Mistral medium open-weights served by vLLM can deliver reasoning
# via either surface, but the trade-off is asymmetric:
# * Chat Completions — tool calling works (``--tool-call-parser
# mistral``); reasoning is enabled via the vLLM CLI
# (``--reasoning-parser``) rather than per-request.
# * Responses API — reasoning effort is per-request and clean,
# but as of vLLM 0.x the tool-call parser is not wired up on
# this surface so tool calls leak as ``[TOOL_CALLS]`` text.
# We do **not** auto-suggest this profile from Detect; an operator
# who needs per-request effort and accepts the tool-calling
# limitation can pick "Responses API" manually in the admin UI.
"server_compat": {
"server_type": "vllm",
"api_surface": "responses",
},
},
"vllm": {
"server_compat": {
"server_type": "vllm",
@@ -125,6 +148,9 @@ _VLLM_MODEL_PROFILES: list[tuple[str, str]] = [
("granite3", "vllm-granite-thinking"),
("deepseek-r1", "vllm-deepseek-thinking"),
("holo2", "vllm-holo-thinking"),
# Mistral medium intentionally omitted — see ``vllm-mistral-medium``
# profile docstring for the Chat-vs-Responses trade-off; operator
# picks manually rather than letting Detect auto-suggest Responses.
]
# llama.cpp model-family → profile key mapping.
@@ -175,31 +201,39 @@ def suggest_profile(server_type: str, model_id: str) -> dict[str, Any]:
def merge_server_compat(
base_chat_template_kwargs: dict[str, Any],
base_chat_template_kwargs: dict[str, Any] | None,
server_compat: dict[str, Any],
) -> dict[str, Any]:
"""Build the ``extra_body`` dict by merging server compat into base kwargs.
*base_chat_template_kwargs* always contains at least ``reasoning_effort``.
*server_compat* comes from ``ModelConfig.server_compat``.
*base_chat_template_kwargs* is an explicit ``chat_template_kwargs`` dict
to seed the request with, or ``None``/empty to skip seeding. Operator-
supplied entries in ``server_compat["extra_body"]["chat_template_kwargs"]``
are deep-merged on top. Top-level ``extra_body`` keys (``skip_special_tokens``,
``reasoning_format``, etc.) are forwarded as-is.
Note: thinking-mode params (``enable_thinking``, ``thinking``) are **not**
merged here the provider handles those via ``ModelCapabilities``.
This function only merges server workarounds from ``extra_body``.
This function only merges what the operator stored in ``server_compat``.
Returns the complete dict to pass as ``extra_body`` to the OpenAI client.
May be empty when there is nothing to send.
"""
extra: dict[str, Any] = {"chat_template_kwargs": dict(base_chat_template_kwargs)}
extra: dict[str, Any] = {}
if base_chat_template_kwargs:
extra["chat_template_kwargs"] = dict(base_chat_template_kwargs)
# Merge top-level extra_body overrides (skip_special_tokens, etc.)
compat_eb = server_compat.get("extra_body")
if isinstance(compat_eb, dict):
for key, value in compat_eb.items():
if key == "chat_template_kwargs":
# Deep-merge: operator values in extra_body win over the
# base dict (which has reasoning_effort). This lets
# operators intentionally extend chat_template_kwargs.
# Deep-merge with operator values winning so an operator
# can intentionally extend chat_template_kwargs (e.g. set
# ``reasoning_effort`` for gpt-oss-style local templates).
if isinstance(value, dict):
if "chat_template_kwargs" not in extra:
extra["chat_template_kwargs"] = {}
extra["chat_template_kwargs"].update(value)
continue
extra[key] = value
+167 -107
View File
@@ -44,6 +44,7 @@ from turnstone.core.attachments import (
)
from turnstone.core.config import get_tavily_key
from turnstone.core.edit import find_occurrences, pick_nearest
from turnstone.core.history_decoration import TOOL_RESULT_STORAGE_CAP
from turnstone.core.log import get_logger
from turnstone.core.memory import (
count_structured_memories,
@@ -57,6 +58,7 @@ from turnstone.core.memory import (
list_default_skills,
list_skills_by_activation,
list_structured_memories,
list_visible_structured_memories,
list_workstreams_with_history,
load_messages,
load_workstream_config,
@@ -70,6 +72,7 @@ from turnstone.core.memory import (
search_history,
search_history_recent,
search_structured_memories,
search_visible_structured_memories,
set_workstream_alias,
unreserve_attachments,
update_workstream_title,
@@ -91,6 +94,7 @@ from turnstone.core.providers import create_provider
from turnstone.core.safety import is_command_blocked, sanitize_command
from turnstone.core.sandbox import execute_math_sandboxed
from turnstone.core.storage._registry import get_storage
from turnstone.core.storage._utils import normalize_search_terms
from turnstone.core.tool_advisory import escape_wrapper_tags, render_system_reminder
from turnstone.core.tool_search import ToolSearchManager
from turnstone.core.tools import (
@@ -427,6 +431,11 @@ class ChatSession:
except Exception:
log.debug("rule_registry.init_failed", exc_info=True)
self._memory_config = memory_config or MemoryConfig()
# Per-turn cache for _search_visible_memories — _init_system_messages
# fires many times within one turn (state transitions, MCP refresh,
# tool results) and the recent-context string is identical across
# them. Invalidated on user-turn append and on memory write/delete.
self._mem_search_cache: dict[tuple[str, str, int], list[dict[str, str]]] = {}
self._ws_id = ws_id or uuid.uuid4().hex
self._title_generated = False
self._read_files: set[str] = set()
@@ -578,7 +587,15 @@ class ChatSession:
self._skill_resources_dir: str | None = None
self._load_skills()
self._init_system_messages()
self._save_config()
# Skip on rehydrate — ``_save_config`` is ``INSERT OR
# REPLACE`` per-key, and the persisted row is what
# ``ChatSession.resume`` is about to read back. Pairs with
# ``SessionManager.open``'s saved-alias threading; together
# they keep reopened workstreams on their original model and
# settings instead of silently resetting to constructor
# defaults.
if not load_workstream_config(self._ws_id):
self._save_config()
@property
def ws_id(self) -> str:
@@ -1339,16 +1356,23 @@ class ChatSession:
model_name,
cfg.context_window,
)
elif saved_model and saved_model != self.model:
# No alias or alias no longer in registry — at least set the model name
self.model = saved_model
self._model_alias = None
self._cached_capabilities = None
elif saved_alias or saved_model:
# Saved alias is unset or no longer in the registry.
# Don't copy ``saved_model`` onto the constructor's
# default provider/client — pairing a removed model
# name with the default provider produces an API call
# the default provider can't service, which is exactly
# the broken state operators see today on the reopen
# path. The constructor already resolved a coherent
# default; keep it intact and warn so the missing
# alias is auditable.
log.warning(
"Resume: alias %r not in registry, keeping default provider=%s for model=%s",
"Resume: saved alias=%r model=%r unreachable; "
"keeping default provider=%s model=%s",
saved_alias,
type(self._provider).__name__,
saved_model,
type(self._provider).__name__,
self.model,
)
if "temperature" in config:
self.temperature = float(config["temperature"])
@@ -1434,11 +1458,6 @@ class ChatSession:
"""
new_system_messages: list[dict[str, Any]] = []
# -- Chat template kwargs --
self._chat_template_kwargs_base: dict[str, Any] = {
"reasoning_effort": self.reasoning_effort,
}
# -- Developer message --
if self.creative_mode:
dev_parts = [
@@ -1617,10 +1636,16 @@ class ChatSession:
if self.instructions:
dev_parts.append("")
dev_parts.append(self.instructions)
visible_mems = self._list_visible_memories(limit=self._mem_cfg.fetch_limit)
context = extract_recent_context(self.messages)
visible_mems, candidate_source = self._select_memory_candidates(context)
if visible_mems:
context = extract_recent_context(self.messages)
relevant = score_memories(visible_mems, context, k=self._mem_cfg.relevance_k)
log.info(
"memory.composition",
source=candidate_source,
candidates=len(visible_mems),
injected=len(relevant),
)
if relevant:
dev_parts.append("")
dev_parts.append(build_memory_context(relevant))
@@ -1814,48 +1839,48 @@ class ChatSession:
def _provider_extra_params(
self,
reasoning_effort: str | None = None,
provider: LLMProvider | None = None,
model_alias: str | None = None,
) -> dict[str, Any] | None:
"""Build provider-specific extra parameters.
``chat_template_kwargs`` is only meaningful for local model servers
(``openai-compatible``). Commercial OpenAI rejects it as an unknown
parameter, and handles ``reasoning_effort`` natively.
Forwards operator-supplied ``server_compat["extra_body"]`` overrides
(``skip_special_tokens``, ``reasoning_format``, or explicit
``chat_template_kwargs``) to the OpenAI SDK ``extra_body``. Operators
running gpt-oss-style local templates that consume ``reasoning_effort``
from ``chat_template_kwargs`` should set it explicitly under
``server_compat["extra_body"]["chat_template_kwargs"]``.
Merges server workarounds (``skip_special_tokens``, etc.) from
``ModelConfig.server_compat`` into the request's ``extra_body``.
Thinking-mode params (``enable_thinking``) are handled separately
by the provider based on ``ModelCapabilities.thinking_mode``.
Thinking-mode params (``enable_thinking``, ``thinking``) are added
separately by ``OpenAIChatCompletionsProvider._apply_thinking_mode``
based on ``ModelCapabilities.thinking_mode`` the Responses API
surface handles reasoning natively and ignores ``extra_body``.
*model_alias* controls which model config supplies server compat
*model_alias* selects which stored config supplies server compat
settings. When ``None``, defaults to the session's primary alias.
"""
from turnstone.core.server_compat import merge_server_compat
prov = provider or self._provider
if prov.provider_name == "openai-compatible":
ctk_base = dict(self._chat_template_kwargs_base)
if reasoning_effort:
ctk_base["reasoning_effort"] = reasoning_effort
return merge_server_compat(
ctk_base,
self._get_server_compat(model_alias),
)
return None
# Only OpenAI-shaped providers consume extra_body. Anthropic/Google
# have their own param paths handled inside their providers.
if prov.provider_name not in ("openai", "openai-compatible"):
return None
extra = merge_server_compat(None, self._get_server_compat(model_alias))
return extra or None
def _get_server_compat(self, model_alias: str | None = None) -> dict[str, Any]:
"""Get server compatibility settings from a model config.
*model_alias* selects the config to read. Falls back to the
session's primary alias when ``None``.
session's primary alias when ``None``. The returned dict is the
live ``ModelConfig.server_compat`` reference callers must not
mutate it. ``merge_server_compat`` reads only.
"""
alias = model_alias or self._model_alias
if self._registry and alias:
try:
cfg = self._registry.get_config(alias)
return dict(cfg.server_compat)
return self._registry.get_config(alias).server_compat
except (ValueError, KeyError):
pass
return {}
@@ -1884,7 +1909,7 @@ class ChatSession:
max_tokens=clamped,
temperature=temperature,
reasoning_effort=reasoning_effort,
extra_params=self._provider_extra_params(reasoning_effort=reasoning_effort),
extra_params=self._provider_extra_params(),
capabilities=caps,
)
@@ -2178,6 +2203,9 @@ class ChatSession:
consume step adds it to the WHERE clause so a stale send can't
steal rows reserved to a different one.
"""
# New user content invalidates the per-turn memory-search cache
# (composition will see a different recent-context string).
self._invalidate_memory_cache()
user_content: str | list[dict[str, Any]]
if attachments:
parts: list[dict[str, Any]] = [{"type": "text", "text": user_input}]
@@ -2414,16 +2442,7 @@ class ChatSession:
if assistant_msg.get("_provider_content"):
provider_data = json.dumps(assistant_msg["_provider_content"])
# Build tool_calls JSON (excluding memory tools)
tool_calls_json: str | None = None
if tc:
filtered_tc = [
call
for call in tc
if call.get("function", {}).get("name", "") not in ("memory", "recall")
]
if filtered_tc:
tool_calls_json = json.dumps(filtered_tc)
tool_calls_json: str | None = json.dumps(tc) if tc else None
# Save assistant message atomically (content + tool_calls in one row)
if content or provider_data is not None or tool_calls_json:
@@ -2571,27 +2590,27 @@ class ChatSession:
tok_est = max(1, int(len(output) / self._chars_per_token))
self._msg_tokens.append(tok_est)
# Log tool result (skip memory tools to avoid noise).
# Use raw_output (pre-advisory-wrap) so DB stores clean
# tool output without ephemeral advisory XML.
# Log tool result. Use raw_output (pre-advisory-wrap)
# so the DB stores clean tool output without ephemeral
# advisory XML. memory/recall persist alongside every
# other tool: replays show the full audit trail, and
# output already passes through _truncate_output above
# so size is bounded by the same budget every other
# tool uses.
_tname = _tc_names.get(tc_id, "")
if _tname not in (
"memory",
"recall",
):
if isinstance(raw_output, list):
store_text = " ".join(
p.get("text", "") for p in raw_output if p.get("type") == "text"
)[:2000]
else:
store_text = raw_output[:2000]
save_message(
self._ws_id,
"tool",
store_text,
_tname,
tool_call_id=tc_id,
)
if isinstance(raw_output, list):
store_text = " ".join(
p.get("text", "") for p in raw_output if p.get("type") == "text"
)[:TOOL_RESULT_STORAGE_CAP]
else:
store_text = raw_output[:TOOL_RESULT_STORAGE_CAP]
save_message(
self._ws_id,
"tool",
store_text,
_tname,
tool_call_id=tc_id,
)
# Inject user feedback from approval prompt (e.g. "y, use full path")
if user_feedback:
self.messages.append({"role": "user", "content": user_feedback})
@@ -5406,60 +5425,90 @@ class ChatSession:
n += count_structured_memories(scope="user", scope_id=self._user_id)
return n
def _visible_scopes(self) -> list[tuple[str, str]]:
"""Return the (scope, scope_id) pairs visible to this session.
Coord sessions see ONLY their coord-scope; interactive sessions see
global + their workstream + their user (when uid present). Drives
the single-query visibility helpers.
"""
if self._kind == WorkstreamKind.COORDINATOR:
return [("coordinator", self._ws_id)]
scopes: list[tuple[str, str]] = [("global", ""), ("workstream", self._ws_id)]
if self._user_id:
scopes.append(("user", self._user_id))
return scopes
def _list_visible_memories(self, mem_type: str = "", limit: int = 50) -> list[dict[str, str]]:
"""List memories visible to this session with optional type filter.
Single SQL round-trip collapses the prior per-scope fan-out.
See :meth:`_visible_memory_count` for the coord-isolation rule.
"""
if self._kind == WorkstreamKind.COORDINATOR:
return list_structured_memories(
mem_type=mem_type,
scope="coordinator",
scope_id=self._ws_id,
limit=limit,
)
global_mems = list_structured_memories(mem_type=mem_type, scope="global", limit=limit)
ws_mems = list_structured_memories(
mem_type=mem_type, scope="workstream", scope_id=self._ws_id, limit=limit
return list_visible_structured_memories(
self._visible_scopes(), mem_type=mem_type, limit=limit
)
user_mems: list[dict[str, str]] = []
if self._user_id:
user_mems = list_structured_memories(
mem_type=mem_type, scope="user", scope_id=self._user_id, limit=limit
)
combined = global_mems + ws_mems + user_mems
combined.sort(key=lambda m: m.get("updated", ""), reverse=True)
return combined[:limit]
def _search_visible_memories(
self, query: str, mem_type: str = "", limit: int = 20
) -> list[dict[str, str]]:
"""Search memories visible to this session (scope-filtered).
Single SQL round-trip with a per-turn cache: ``_init_system_messages``
is invoked many times within a turn (state transitions, MCP refresh,
tool results) and the recent-context query is identical across them.
Cache is cleared on each new user turn and after memory writes/deletes.
See :meth:`_visible_memory_count` for the coord-isolation rule.
"""
if self._kind == WorkstreamKind.COORDINATOR:
return search_structured_memories(
query,
mem_type=mem_type,
scope="coordinator",
scope_id=self._ws_id,
limit=limit,
)
global_mems = search_structured_memories(
query, mem_type=mem_type, scope="global", limit=limit
cache_key = (query, mem_type, limit)
cached = self._mem_search_cache.get(cache_key)
if cached is not None:
return cached
rows = search_visible_structured_memories(
query, self._visible_scopes(), mem_type=mem_type, limit=limit
)
ws_mems = search_structured_memories(
query, mem_type=mem_type, scope="workstream", scope_id=self._ws_id, limit=limit
)
user_mems: list[dict[str, str]] = []
if self._user_id:
user_mems = search_structured_memories(
query, mem_type=mem_type, scope="user", scope_id=self._user_id, limit=limit
)
combined = global_mems + ws_mems + user_mems
combined.sort(key=lambda m: m.get("updated", ""), reverse=True)
return combined[:limit]
self._mem_search_cache[cache_key] = rows
return rows
def _invalidate_memory_cache(self) -> None:
"""Drop the per-turn search cache; call on user-turn append + memory writes."""
self._mem_search_cache.clear()
def _select_memory_candidates(self, context: str) -> tuple[list[dict[str, str]], str]:
"""Pick the candidate set fed into BM25 ranking.
Returns ``(memories, source_label)`` where source is one of:
``recency`` (no context, or search returned nothing),
``search`` (search saturated the fetch_limit budget alone), or
``union`` (search hits recency, deduped by memory_id).
Invariant: the candidate pool is always a SUPERSET of the
recency-only pool the original bug used recency is fully
preserved (not truncated) whenever it gets unioned. Worst
case the union is 2 × fetch_limit candidates (~100 with
defaults), which BM25 ranks in pure Python in well under a
millisecond. BM25's score>0 cutoff in bm25.py drops anything
that doesn't match the query, so unranked recency tail items
cost nothing on irrelevant candidates while saving the
relevant ones.
Capping the union at fetch_limit (the prior behavior) would
evict the recency tail when search added distinct hits and
the recency tail is exactly where ancient-but-recently-touched
memories live, which is the recall the PR sets out to improve.
"""
fetch_limit = self._mem_cfg.fetch_limit
if not context:
return self._list_visible_memories(limit=fetch_limit), "recency"
search_hits = self._search_visible_memories(context, limit=fetch_limit)
if len(search_hits) >= fetch_limit:
return search_hits, "search"
recency = self._list_visible_memories(limit=fetch_limit)
seen = {m["memory_id"] for m in search_hits}
extra = [m for m in recency if m["memory_id"] not in seen]
if not search_hits:
return extra, "recency"
return search_hits + extra, ("union" if extra else "search")
def _check_metacognitive_nudge(self, user_message: str) -> tuple[str, str] | None:
"""Check if a metacognitive nudge should fire for *user_message*.
@@ -7983,6 +8032,10 @@ class ChatSession:
agent_client = self.client
agent_model = self.model
agent_provider = self._provider
# When falling through to the session's primary model, use the
# session's primary alias for capability and server_compat
# resolution so the agent sees the same caps as the main loop.
agent_alias = self._model_alias
# Per-kind reasoning effort. Explicit caller arg wins; otherwise
# delegate to the registry which knows the per-kind default (plan
@@ -8005,7 +8058,6 @@ class ChatSession:
# Build extra params for agent calls — resolve server compat from the
# agent's own model alias, not the session's primary model.
agent_extra = self._provider_extra_params(
reasoning_effort=reasoning_effort,
provider=agent_provider,
model_alias=agent_alias,
)
@@ -8444,6 +8496,7 @@ class ChatSession:
msg = f"Error: failed to save memory '{item['name']}'"
self._report_tool_result(call_id, "memory", msg, is_error=True)
return call_id, msg
self._invalidate_memory_cache()
self._init_system_messages()
if old is not None:
msg = f"Updated memory '{item['name']}' (type={item['mem_type']}, scope={item['scope']})"
@@ -8489,6 +8542,7 @@ class ChatSession:
msg = f"Error: memory '{item['name']}' not found (searched scopes: {tried})"
self._report_tool_result(call_id, "memory", msg, is_error=True)
else:
self._invalidate_memory_cache()
self._init_system_messages()
msg = f"Deleted memory '{item['name']}' (scope={deleted_scope})"
self._report_tool_result(call_id, "memory", msg)
@@ -8516,6 +8570,12 @@ class ChatSession:
mem_type=item.get("mem_type", ""),
limit=item["limit"],
)
log.info(
"memory.search",
term_count=len(normalize_search_terms(item["query"])),
result_count=len(rows),
query=item["query"][:120],
)
if rows:
lines = []
for m in rows:
+81 -10
View File
@@ -180,6 +180,7 @@ class SessionManager:
node_id: str | None = None,
state_writer: StateWriter | None = None,
event_emitter: SessionEventEmitter | None = None,
model_validator: Callable[[str], bool] | None = None,
) -> None:
if max_active < 1:
raise ValueError(f"max_active must be >= 1, got {max_active}")
@@ -199,6 +200,15 @@ class SessionManager:
# effects, and reserved for future kinds whose lifecycle
# transitions don't fan out anywhere.
self._event_emitter = event_emitter
# Optional registry-membership check applied to the persisted
# ``model_alias`` on the rehydrate path before threading it
# into ``build_session``. Production wiring passes
# ``registry.has_alias``; an alias that has been removed from
# the registry since the workstream was created is filtered
# out so the session_factory falls back to its default rather
# than raising. Restricted to the rehydrate path — fresh
# creates still want unknown aliases to surface as 503.
self._model_validator = model_validator
self._node_id = node_id
self._workstreams: dict[str, Workstream] = {}
self._order: list[str] = []
@@ -214,13 +224,21 @@ class SessionManager:
# manager never reads them.
self._active_id: str | None = None
self._eviction_count: int = 0
# Optional state-change observer. The CLI sets this to a
# callback that prints a background-attention notification
# when a non-focused workstream transitions to ATTENTION.
# Web/coord paths use the event_emitter's emit_state for their
# own fan-out; this is a second, manager-level hook for callers
# that don't consume SSE.
self._on_state_change: Callable[[str, WorkstreamState], None] | None = None
# State-change subscribers. Multi-subscriber to support the
# CLI's background-attention notification AND the in-process
# ``SameNodeChildSource`` strategy that delivers child
# workstream state changes to a parent's UI without going
# through the cluster bus. Each callback fires under
# exception-suppression so one failing subscriber doesn't
# block the others. Subscribers register via
# :meth:`subscribe_to_state`. ``_state_subscribers_lock``
# guards mutation + snapshot — set_state copies the list
# under the lock then iterates the snapshot unlocked so a
# slow subscriber doesn't block subscribe/unsubscribe (and
# so concurrent subscribe/unsubscribe during a state event
# can't shift the iterator's index — caught by /review bug-1).
self._state_subscribers: list[Callable[[str, WorkstreamState], None]] = []
self._state_subscribers_lock = threading.Lock()
# ------------------------------------------------------------------
# Properties
@@ -595,8 +613,34 @@ class SessionManager:
evicted.id, reason="evicted", name=evicted.name
)
# Thread the persisted ``model_alias`` into
# ``build_session`` so reopened workstreams keep the
# model they were created with. Pairs with the
# ``ChatSession.__init__`` skip-save guard: without
# both halves, ``_save_config`` clobbers persisted
# config with constructor defaults before
# ``ChatSession.resume`` reads them back. When
# ``model_validator`` is wired and the saved alias is
# no longer in the registry, drop it so the factory
# falls back to its default — the session_factory
# itself still raises on unknown aliases, since
# fresh-create paths want that to surface as a 503.
saved_cfg = self._storage.load_workstream_config(ws_id)
saved_alias = (saved_cfg.get("model_alias") or None) if saved_cfg else None
if (
saved_alias
and self._model_validator is not None
and not self._model_validator(saved_alias)
):
log.warning(
"session_mgr.stale_alias_dropped ws=%s alias=%s",
ws_id[:8],
saved_alias,
)
saved_alias = None
try:
ws.session = self._adapter.build_session(ws)
ws.session = self._adapter.build_session(ws, model=saved_alias)
except Exception:
# Clean up the UI the adapter built before re-raising
# so any listener/lock resources are released.
@@ -816,9 +860,36 @@ class SessionManager:
log.debug("session_mgr.state_update_failed ws=%s", ws_id[:8], exc_info=True)
if self._event_emitter is not None:
self._event_emitter.emit_state(ws, state)
if self._on_state_change is not None:
# Snapshot under the subscribers lock so concurrent
# subscribe / unsubscribe can't shift the iterator's index
# mid-dispatch (skipping or repeating callbacks). Iterate
# the snapshot WITHOUT the lock so a slow callback doesn't
# block subscribe / unsubscribe.
with self._state_subscribers_lock:
subscribers = list(self._state_subscribers)
for callback in subscribers:
with contextlib.suppress(Exception):
self._on_state_change(ws_id, state)
callback(ws_id, state)
# ------------------------------------------------------------------
# State-change subscription
# ------------------------------------------------------------------
def subscribe_to_state(self, callback: Callable[[str, WorkstreamState], None]) -> None:
"""Register ``callback`` to fire on every workstream state change.
Multiple subscribers are supported and fire in registration order.
Each callback is wrapped in exception-suppression so a failing
subscriber doesn't block the others. Use
:meth:`unsubscribe_from_state` to remove.
"""
with self._state_subscribers_lock:
self._state_subscribers.append(callback)
def unsubscribe_from_state(self, callback: Callable[[str, WorkstreamState], None]) -> None:
"""Remove a previously-registered state-change callback. No-op if absent."""
with self._state_subscribers_lock, contextlib.suppress(ValueError):
self._state_subscribers.remove(callback)
def cancel(self, ws_id: str) -> bool:
"""Cancel in-flight generation and unblock any pending approval / plan.
+61 -1
View File
@@ -394,6 +394,15 @@ class SessionEndpointConfig:
# separate ``/history`` endpoint and doesn't render the per-tab
# status bar). Kinds that don't need pre-replay wire ``None``.
events_replay: EventsReplay | None = None
# async (ws, ui, request) -> None. Kind-specific async pre-step
# the lifted ``events`` body awaits BEFORE iterating
# ``events_replay``. Lets a kind move blocking storage I/O off
# the event loop (via ``asyncio.to_thread``) and stash results
# on ``request.state`` for the sync replay generator to read.
# Interactive uses it to pre-load intent_verdicts +
# output_assessments so ``_build_history``'s decoration stays
# off the hot path. Coord wires ``None``.
events_replay_prepare: Callable[..., Any] | None = None
# (request) -> Executor for the SSE live-loop's blocking
# ``queue.get`` wait. Interactive returns the dedicated
# ``request.app.state.sse_executor`` (200-thread pool) so SSE
@@ -1203,6 +1212,8 @@ def make_open_handler(
"""
async def open_ws(request: Request) -> Response:
import asyncio
if cfg.permission_gate is not None:
err = cfg.permission_gate(request)
if err is not None:
@@ -1304,7 +1315,14 @@ def make_open_handler(
# emit_rehydrated path).
if cfg.open_post_load is not None:
try:
cfg.open_post_load(request, ws)
# Off-loop: interactive's post_load runs the sync
# ``_build_history`` (storage I/O for verdict
# indexes + message reconstruction) — without the
# to_thread wrap this blocks the event loop on every
# workstream open, mirroring the SSE replay path
# that's already protected via
# ``events_replay_prepare``.
await asyncio.to_thread(cfg.open_post_load, request, ws)
except Exception:
# Post-load is observational — never let a hook bug
# block the open. Log + continue.
@@ -1454,6 +1472,20 @@ def make_events_handler(cfg: SessionEndpointConfig) -> Handler:
# 500-slot cap on a chatty mid-generation workstream)
# while replay was being built.
if replay_cb is not None:
# Kind-specific async prep — runs before the sync
# replay generator iterates so blocking storage
# I/O lands in the executor pool rather than the
# event loop's hot path. Interactive uses this
# to pre-load verdict indexes; coord skips.
if cfg.events_replay_prepare is not None:
try:
await cfg.events_replay_prepare(ws, ui, request)
except Exception:
log.debug(
"ws.events.replay_prepare_failed ws=%s",
ws_id[:8],
exc_info=True,
)
try:
for ev in replay_cb(ws, ui, request):
yield {"data": json.dumps(ev)}
@@ -2240,6 +2272,34 @@ def make_history_handler(cfg: SessionEndpointConfig) -> Handler:
messages = await asyncio.to_thread(storage.load_messages, ws_id, limit=limit)
except Exception:
log.debug("ws.history.load_failed ws=%s", ws_id[:8], exc_info=True)
# Audit-trail decoration — attach persisted intent_verdict and
# output_assessment data to each assistant.tool_calls entry so
# the dashboard's history replay paints the same verdict pills
# / output-warning bubbles the live SSE path shows. Both
# storage queries are off-loop via ``to_thread``. Best-effort:
# any failure leaves messages undecorated — replay degrades to
# the pre-decoration shape rather than 500-ing.
if messages:
try:
from turnstone.core.history_decoration import (
decorate_history_messages,
load_verdict_indexes,
)
indexes = await asyncio.to_thread(load_verdict_indexes, ws_id)
decorate_history_messages(messages, indexes[0], indexes[1])
except Exception:
# Operationally interesting: a persistent decoration
# failure (missing migration, driver mismatch, schema
# drift) silently strips verdict pills + output
# warnings from every reload of every workstream.
# Log at warning so it surfaces in normal log review
# rather than only when DEBUG is on.
log.warning(
"ws.history.decoration_failed ws=%s",
ws_id[:8],
exc_info=True,
)
return JSONResponse({"ws_id": ws_id, "messages": messages})
return history
+62
View File
@@ -41,6 +41,7 @@ log = get_logger(__name__)
# from bloating memory.
_DEFAULT_LISTENER_QUEUE_MAX = 500
# Cap on the per-turn assistant content accumulator. The accumulator
# is piggybacked onto the ``ws_state:idle`` broadcast payload so the
# cluster collector / dashboard can render the freshly-emitted assistant
@@ -327,6 +328,11 @@ class SessionUIBase:
"always": bool(always),
}
)
# Kind-specific cross-stream broadcast — ConsoleCoordinatorUI
# overrides to push onto the cluster bus so a coord parent's
# tree UI clears the pending-approval pill in lockstep with
# the actual decision. Stage 3 Step 4.
self._broadcast_approval_resolved(approved, feedback, always=always)
self._approval_event.set()
@staticmethod
@@ -579,6 +585,20 @@ class SessionUIBase:
"judge_pending": judge_pending,
}
self._enqueue(self._pending_approval)
# Cross-stream broadcast — push the items via the cluster bus
# so a coord parent's tree UI can render the inline approve/deny
# block without waiting for a bulk fetch. Without this, the
# bulk fetch races with this assignment: the state transition
# to ATTENTION fires upstream BEFORE approve_tools runs (see
# session.py:_emit_state("attention") preceding ui.approve_tools),
# so a bulk fetch landing in the ~50-200ms window between
# _emit_state and this point sees ``_pending_approval=None``
# and returns ``pending_approval_detail: null``. The 5s TTL
# then locks the coord row on a "loading" placeholder until
# the next state event triggers a refresh — which never comes
# while parked on _approval_event.wait. The push path
# eliminates the race.
self._broadcast_approve_request(self._pending_approval)
if not self._approval_event.wait(timeout=self._APPROVAL_WAIT_TIMEOUT):
# Approval timed out (e.g., user disconnected). Deny via
# resolve_approval so verdicts and state are updated consistently.
@@ -632,6 +652,12 @@ class SessionUIBase:
del self._llm_verdicts[oldest_key]
self._llm_verdicts[call_id] = verdict
self._enqueue({"type": "intent_verdict", **verdict})
# Kind-specific cross-stream broadcast — ConsoleCoordinatorUI
# overrides to push onto the cluster bus so a coord parent's
# tree UI sees the verdict without polling. Default is no-op
# (the per-ws ``_enqueue`` above already covers WebUI's own
# SSE listeners). Stage 3 Step 4.
self._broadcast_intent_verdict(verdict)
self._persist_intent_verdict(verdict)
# Decision check + either queue or flag-for-persist happen
# under ONE lock acquisition so resolve_approval can't swap-
@@ -1351,6 +1377,42 @@ class SessionUIBase:
Default: no-op. Subclasses override.
"""
def _broadcast_intent_verdict(self, verdict: dict[str, Any]) -> None: # noqa: ARG002 — hook stub
"""Fan an LLM intent-judge verdict out to the kind's transport.
Default: no-op. ``ConsoleCoordinatorUI`` overrides to push a
``intent_verdict`` event onto the cluster bus so the parent
coordinator's tree UI can render the risk pill + verdict
result without polling. Stage 3 Step 4: hook only the
cluster-bus event class lands in Step 5.
"""
def _broadcast_approval_resolved(
self,
approved: bool, # noqa: ARG002 — hook stub
feedback: str | None = None, # noqa: ARG002 — hook stub
*,
always: bool = False, # noqa: ARG002 — hook stub
) -> None:
"""Fan an ``approval_resolved`` decision out to the kind's transport.
Default: no-op. ``ConsoleCoordinatorUI`` overrides to push to
the cluster bus so the parent coordinator's tree UI can clear
the pending-approval pill in sync with the actual decision.
"""
def _broadcast_approve_request(self, detail: dict[str, Any]) -> None: # noqa: ARG002 — hook stub
"""Fan an ``approve_request`` payload out to the kind's transport.
Default: no-op. ``WebUI`` and ``ConsoleCoordinatorUI`` override
to push the items list (the same dict that landed in
``_pending_approval``) onto their respective transports. The
push path eliminates the bulk-fetch race that otherwise
leaves coord rows stuck on a loading placeholder when the
bulk fetch lands in the gap between the state transition to
ATTENTION and ``_pending_approval`` being set.
"""
# ------------------------------------------------------------------
# State-broadcast snapshot helper
# ------------------------------------------------------------------
+110 -8
View File
@@ -87,6 +87,9 @@ from turnstone.core.storage._utils import (
from turnstone.core.storage._utils import (
VERDICT_MUTABLE as _VERDICT_MUTABLE,
)
from turnstone.core.storage._utils import (
normalize_search_terms as _normalize_search_terms,
)
from turnstone.core.storage._utils import (
reconstruct_messages as _reconstruct_messages,
)
@@ -3315,7 +3318,10 @@ class PostgreSQLBackend:
limit: int = 100,
) -> list[dict[str, str]]:
with self._conn() as conn:
q = sa.select(structured_memories).order_by(structured_memories.c.updated.desc())
q = sa.select(structured_memories).order_by(
structured_memories.c.updated.desc(),
structured_memories.c.memory_id.asc(),
)
if mem_type:
q = q.where(structured_memories.c.type == mem_type)
if scope:
@@ -3334,11 +3340,16 @@ class PostgreSQLBackend:
scope_id: str = "",
limit: int = 20,
) -> list[dict[str, str]]:
"""OR-of-terms ILIKE search; ranking is the caller's job (BM25 downstream)."""
if not query or not query.strip():
return self.list_structured_memories(
mem_type=mem_type, scope=scope, scope_id=scope_id, limit=limit
)
terms = query.split()
terms = _normalize_search_terms(query)
if not terms:
return self.list_structured_memories(
mem_type=mem_type, scope=scope, scope_id=scope_id, limit=limit
)
with self._conn() as conn:
clauses = []
params: dict[str, str] = {}
@@ -3352,25 +3363,116 @@ class PostgreSQLBackend:
params[f"n{i}"] = f"%{escaped}%"
params[f"d{i}"] = f"%{escaped}%"
params[f"c{i}"] = f"%{escaped}%"
where = " AND ".join(clauses)
term_clause = " OR ".join(clauses)
scope_filters = ""
if mem_type:
where += " AND type = :type_filter"
scope_filters += " AND type = :type_filter"
params["type_filter"] = mem_type
if scope:
where += " AND scope = :scope_filter"
scope_filters += " AND scope = :scope_filter"
params["scope_filter"] = scope
if scope_id and scope:
where += " AND scope_id = :scope_id_filter"
scope_filters += " AND scope_id = :scope_id_filter"
params["scope_id_filter"] = scope_id
rows = conn.execute(
sa.text(
f"SELECT * FROM structured_memories WHERE {where} "
f"ORDER BY updated DESC LIMIT :lim"
f"SELECT * FROM structured_memories WHERE ({term_clause}){scope_filters} "
f"ORDER BY updated DESC, memory_id ASC LIMIT :lim"
),
{**params, "lim": limit},
).fetchall()
return [dict(r._mapping) for r in rows]
def list_visible_structured_memories(
self,
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 100,
) -> list[dict[str, str]]:
"""Single-query union across visible (scope, scope_id) pairs.
Replaces the per-scope fan-out (one query per visible scope) so the
composition path issues 1 round-trip instead of 3.
"""
if not scopes:
return []
with self._conn() as conn:
scope_clauses, params = self._build_scope_or_clause(scopes)
extra = ""
if mem_type:
extra = " AND type = :type_filter"
params["type_filter"] = mem_type
rows = conn.execute(
sa.text(
f"SELECT * FROM structured_memories WHERE ({scope_clauses}){extra} "
f"ORDER BY updated DESC, memory_id ASC LIMIT :lim"
),
{**params, "lim": limit},
).fetchall()
return [dict(r._mapping) for r in rows]
def search_visible_structured_memories(
self,
query: str,
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 20,
) -> list[dict[str, str]]:
"""OR-of-terms search joined with a single visibility OR-group.
Replaces the per-scope search fan-out; ranking is the caller's job.
"""
if not scopes:
return []
if not query or not query.strip():
return self.list_visible_structured_memories(scopes, mem_type=mem_type, limit=limit)
terms = _normalize_search_terms(query)
if not terms:
return self.list_visible_structured_memories(scopes, mem_type=mem_type, limit=limit)
with self._conn() as conn:
scope_clauses, params = self._build_scope_or_clause(scopes)
term_clauses = []
for i, t in enumerate(terms):
escaped = _escape_ilike(t)
term_clauses.append(
f"(name ILIKE :n{i} ESCAPE '\\' "
f"OR description ILIKE :d{i} ESCAPE '\\' "
f"OR content ILIKE :c{i} ESCAPE '\\')"
)
params[f"n{i}"] = f"%{escaped}%"
params[f"d{i}"] = f"%{escaped}%"
params[f"c{i}"] = f"%{escaped}%"
term_clause = " OR ".join(term_clauses)
extra = ""
if mem_type:
extra = " AND type = :type_filter"
params["type_filter"] = mem_type
rows = conn.execute(
sa.text(
f"SELECT * FROM structured_memories "
f"WHERE ({scope_clauses}) AND ({term_clause}){extra} "
f"ORDER BY updated DESC, memory_id ASC LIMIT :lim"
),
{**params, "lim": limit},
).fetchall()
return [dict(r._mapping) for r in rows]
@staticmethod
def _build_scope_or_clause(
scopes: list[tuple[str, str]],
) -> tuple[str, dict[str, str]]:
"""Build a parameterized OR-group of (scope[, scope_id]) predicates."""
params: dict[str, str] = {}
clauses: list[str] = []
for i, (s, sid) in enumerate(scopes):
params[f"sc{i}"] = s
if sid:
params[f"sid{i}"] = sid
clauses.append(f"(scope = :sc{i} AND scope_id = :sid{i})")
else:
clauses.append(f"scope = :sc{i}")
return " OR ".join(clauses), params
def touch_structured_memories(self, keys: list[tuple[str, str, str]]) -> int:
"""Batch-touch multiple memories by (name, scope, scope_id)."""
if not keys:
+28
View File
@@ -386,6 +386,34 @@ class StorageBackend(Protocol):
"""Search structured memories by query. Returns matching memory dicts."""
...
def list_visible_structured_memories(
self,
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 100,
) -> list[dict[str, str]]:
"""List memories matching ANY of the (scope, scope_id) pairs in *scopes*.
A pair with an empty ``scope_id`` matches the scope alone (used for
``("global", "")``). Single SQL query replaces the per-scope fan-out
pattern that issued one query per visible scope.
"""
...
def search_visible_structured_memories(
self,
query: str,
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 20,
) -> list[dict[str, str]]:
"""OR-of-terms search across memories visible under *scopes*.
Single SQL query joining the scope OR-group with the term OR-group.
Ranking is the caller's job (BM25 downstream).
"""
...
def touch_structured_memories(self, keys: list[tuple[str, str, str]]) -> int:
"""Batch-touch multiple memories.
+103 -8
View File
@@ -87,6 +87,9 @@ from turnstone.core.storage._utils import (
from turnstone.core.storage._utils import (
VERDICT_MUTABLE as _VERDICT_MUTABLE,
)
from turnstone.core.storage._utils import (
normalize_search_terms as _normalize_search_terms,
)
from turnstone.core.storage._utils import (
reconstruct_messages as _reconstruct_messages,
)
@@ -3454,7 +3457,10 @@ class SQLiteBackend:
limit: int = 100,
) -> list[dict[str, str]]:
with self._conn() as conn:
q = sa.select(structured_memories).order_by(structured_memories.c.updated.desc())
q = sa.select(structured_memories).order_by(
structured_memories.c.updated.desc(),
structured_memories.c.memory_id.asc(),
)
if mem_type:
q = q.where(structured_memories.c.type == mem_type)
if scope:
@@ -3473,11 +3479,16 @@ class SQLiteBackend:
scope_id: str = "",
limit: int = 20,
) -> list[dict[str, str]]:
"""OR-of-terms LIKE search; ranking is the caller's job (BM25 downstream)."""
if not query or not query.strip():
return self.list_structured_memories(
mem_type=mem_type, scope=scope, scope_id=scope_id, limit=limit
)
terms = query.split()
terms = _normalize_search_terms(query)
if not terms:
return self.list_structured_memories(
mem_type=mem_type, scope=scope, scope_id=scope_id, limit=limit
)
with self._conn() as conn:
clauses = []
params: dict[str, str] = {}
@@ -3491,25 +3502,109 @@ class SQLiteBackend:
params[f"n{i}"] = f"%{escaped}%"
params[f"d{i}"] = f"%{escaped}%"
params[f"c{i}"] = f"%{escaped}%"
where = " AND ".join(clauses)
term_clause = " OR ".join(clauses)
scope_filters = ""
if mem_type:
where += " AND type = :type_filter"
scope_filters += " AND type = :type_filter"
params["type_filter"] = mem_type
if scope:
where += " AND scope = :scope_filter"
scope_filters += " AND scope = :scope_filter"
params["scope_filter"] = scope
if scope_id and scope:
where += " AND scope_id = :scope_id_filter"
scope_filters += " AND scope_id = :scope_id_filter"
params["scope_id_filter"] = scope_id
rows = conn.execute(
sa.text(
f"SELECT * FROM structured_memories WHERE {where} "
f"ORDER BY updated DESC LIMIT :lim"
f"SELECT * FROM structured_memories WHERE ({term_clause}){scope_filters} "
f"ORDER BY updated DESC, memory_id ASC LIMIT :lim"
),
{**params, "lim": limit},
).fetchall()
return [dict(r._mapping) for r in rows]
def list_visible_structured_memories(
self,
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 100,
) -> list[dict[str, str]]:
"""Single-query union across visible (scope, scope_id) pairs."""
if not scopes:
return []
with self._conn() as conn:
scope_clauses, params = self._build_scope_or_clause(scopes)
extra = ""
if mem_type:
extra = " AND type = :type_filter"
params["type_filter"] = mem_type
rows = conn.execute(
sa.text(
f"SELECT * FROM structured_memories WHERE ({scope_clauses}){extra} "
f"ORDER BY updated DESC, memory_id ASC LIMIT :lim"
),
{**params, "lim": limit},
).fetchall()
return [dict(r._mapping) for r in rows]
def search_visible_structured_memories(
self,
query: str,
scopes: list[tuple[str, str]],
mem_type: str = "",
limit: int = 20,
) -> list[dict[str, str]]:
"""OR-of-terms search joined with a single visibility OR-group."""
if not scopes:
return []
if not query or not query.strip():
return self.list_visible_structured_memories(scopes, mem_type=mem_type, limit=limit)
terms = _normalize_search_terms(query)
if not terms:
return self.list_visible_structured_memories(scopes, mem_type=mem_type, limit=limit)
with self._conn() as conn:
scope_clauses, params = self._build_scope_or_clause(scopes)
term_clauses = []
for i, t in enumerate(terms):
escaped = _escape_like(t)
term_clauses.append(
f"(name LIKE :n{i} ESCAPE '\\' "
f"OR description LIKE :d{i} ESCAPE '\\' "
f"OR content LIKE :c{i} ESCAPE '\\')"
)
params[f"n{i}"] = f"%{escaped}%"
params[f"d{i}"] = f"%{escaped}%"
params[f"c{i}"] = f"%{escaped}%"
term_clause = " OR ".join(term_clauses)
extra = ""
if mem_type:
extra = " AND type = :type_filter"
params["type_filter"] = mem_type
rows = conn.execute(
sa.text(
f"SELECT * FROM structured_memories "
f"WHERE ({scope_clauses}) AND ({term_clause}){extra} "
f"ORDER BY updated DESC, memory_id ASC LIMIT :lim"
),
{**params, "lim": limit},
).fetchall()
return [dict(r._mapping) for r in rows]
@staticmethod
def _build_scope_or_clause(
scopes: list[tuple[str, str]],
) -> tuple[str, dict[str, str]]:
"""Build a parameterized OR-group of (scope[, scope_id]) predicates."""
params: dict[str, str] = {}
clauses: list[str] = []
for i, (s, sid) in enumerate(scopes):
params[f"sc{i}"] = s
if sid:
params[f"sid{i}"] = sid
clauses.append(f"(scope = :sc{i} AND scope_id = :sid{i})")
else:
clauses.append(f"scope = :sc{i}")
return " OR ".join(clauses), params
def touch_structured_memories(self, keys: list[tuple[str, str, str]]) -> int:
"""Batch-touch multiple memories by (name, scope, scope_id)."""
if not keys:
+34
View File
@@ -5,6 +5,7 @@ from __future__ import annotations
import base64
import contextlib
import json
import re
from typing import Any
from turnstone.core.attachments import unreadable_placeholder
@@ -48,6 +49,39 @@ def _attachment_to_content_part(att: dict[str, Any]) -> dict[str, Any] | None:
return None
# ---------------------------------------------------------------------------
# Search-term normalization
# ---------------------------------------------------------------------------
# Composition can hand a multi-KB pasted user message to ILIKE-based search;
# without a cap, every distinct token would emit one unindexable predicate
# per scope-fanned query, producing hundreds of seq-scan clauses on a single
# rebuild. Cap + dedupe + length filter keeps the SQL bounded.
_MAX_SEARCH_TERMS = 16
_MIN_TERM_LEN = 2
# Streaming tokenizer — finditer doesn't allocate a full list up front,
# so a multi-KB pasted query stops being scanned the moment the cap is
# hit instead of after splitting every token.
_TOKEN_RE = re.compile(r"\S+")
def normalize_search_terms(query: str) -> list[str]:
"""De-dupe (case-insensitive), drop short tokens, and cap at MAX terms."""
seen: set[str] = set()
terms: list[str] = []
for match in _TOKEN_RE.finditer(query):
raw = match.group()
lowered = raw.lower()
if len(lowered) < _MIN_TERM_LEN or lowered in seen:
continue
seen.add(lowered)
terms.append(raw)
if len(terms) >= _MAX_SEARCH_TERMS:
break
return terms
# ---------------------------------------------------------------------------
# Text sanitization
# ---------------------------------------------------------------------------
+360 -26
View File
@@ -53,6 +53,15 @@ from turnstone.core.auth import (
_DenyFilter,
jwt_version_slot,
)
from turnstone.core.history_decoration import (
TOOL_RESULT_STORAGE_CAP,
)
from turnstone.core.history_decoration import (
decorate_tool_call as _decorate_tool_call,
)
from turnstone.core.history_decoration import (
load_verdict_indexes as _load_verdict_indexes,
)
from turnstone.core.log import get_logger
from turnstone.core.metrics import metrics as _metrics
from turnstone.core.ratelimit import resolve_client_ip
@@ -194,18 +203,14 @@ class WebUI(SessionUIBase):
}
if state == "idle":
event["content"] = payload["content"]
# Coord tree-UI renders inline approve/deny buttons off
# ``pending_approval_detail``; carrying it on the
# state-change broadcast lets the cluster bus update those
# buttons in lockstep with ``activity_state`` instead of
# forcing the browser to chase a separate dashboard fetch.
# Gated on existence so we don't pay the serializer's
# per-broadcast verdict-cache deepcopy on the common
# no-approval-pending path.
if self._pending_approval is not None:
detail = self.serialize_pending_approval_detail()
if detail is not None:
event["pending_approval_detail"] = detail
# ``pending_approval_detail`` is NO LONGER piggybacked on
# state-change events (Stage 3 cleanup). Symmetric event
# flow now: initial approval items arrive via bulk fetch
# triggered by the ``activity_state="approval"`` transition,
# individual verdicts via the explicit
# ``intent_verdict`` event class, and resolution via
# ``approval_resolved``. Reducer no longer has to dedupe
# the piggyback path against the explicit one.
try:
WebUI._global_queue.put_nowait(event)
except queue.Full:
@@ -230,6 +235,73 @@ class WebUI(SessionUIBase):
}
)
def _broadcast_intent_verdict(self, verdict: dict[str, Any]) -> None:
"""Send an LLM intent-judge verdict to the global SSE channel.
Stage 3 Step 5 the cluster collector's ``_apply_delta``
forwards this verbatim to the cluster bus, where coord
adapters dispatch it as ``child_ws_intent_verdict`` for the
owning parent's tree UI. Unlike the existing
``pending_approval_detail`` piggyback on ``ws_state``, this
fires WHENEVER a verdict lands including the common case
where the judge daemon writes during ``attention`` with no
state transition to ride along on.
"""
if WebUI._global_queue is not None:
with contextlib.suppress(queue.Full):
WebUI._global_queue.put_nowait(
{
"type": "intent_verdict",
"ws_id": self.ws_id,
"verdict": verdict,
}
)
def _broadcast_approval_resolved(
self,
approved: bool,
feedback: str | None = None,
*,
always: bool = False,
) -> None:
"""Send an ``approval_resolved`` decision to the global SSE channel.
Clears the parent's pending-approval pill in lockstep with
the actual decision rather than waiting for the next
state-change piggyback.
"""
if WebUI._global_queue is not None:
with contextlib.suppress(queue.Full):
WebUI._global_queue.put_nowait(
{
"type": "approval_resolved",
"ws_id": self.ws_id,
"approved": approved,
"feedback": feedback or "",
"always": bool(always),
}
)
def _broadcast_approve_request(self, detail: dict[str, Any]) -> None:
"""Send an ``approve_request`` payload to the global SSE channel.
Push path for the initial approval items so a coord parent's
tree UI can render the inline approve/deny block immediately
without waiting for a bulk-fetch round-trip. The bulk fetch
races with ``_pending_approval`` being set inside
``approve_tools`` (the state transition to ATTENTION fires
upstream first); the push path eliminates that race entirely.
"""
if WebUI._global_queue is not None:
with contextlib.suppress(queue.Full):
WebUI._global_queue.put_nowait(
{
"type": "approve_request",
"ws_id": self.ws_id,
"detail": detail,
}
)
# --- SessionUI protocol ---
#
# ``on_thinking_start`` / ``on_thinking_stop`` / ``on_reasoning_token``
@@ -345,8 +417,20 @@ class WebUI(SessionUIBase):
# ---------------------------------------------------------------------------
# Verdict + output-assessment decoration helpers (``_decorate_tool_call``,
# ``_load_verdict_indexes``) are imported at module top alongside the
# rest of ``turnstone.core.*``. Both this builder and
# :func:`make_history_handler` (the /history REST endpoint coord uses
# as its primary history loader) share them so the two surfaces don't
# drift on the wire shape they emit.
def _build_history(
session: ChatSession, has_pending_approval: bool = False
session: ChatSession,
has_pending_approval: bool = False,
*,
verdicts: dict[str, dict[str, Any]] | None = None,
assessments: dict[str, dict[str, Any]] | None = None,
) -> list[dict[str, Any]]:
"""Build a history replay list from ChatSession messages.
@@ -358,6 +442,12 @@ def _build_history(
``"denied": True``, and the corresponding assistant entry that
issued the tool calls is also marked ``"denied": True`` so the
client can render the correct badge.
``verdicts`` and ``assessments`` are optional pre-loaded
``{call_id row}`` dicts (see :func:`_load_verdict_indexes`).
Async callers should pre-load via ``asyncio.to_thread`` and pass
them in to avoid blocking the event loop on storage I/O. When
omitted, the storage call runs inline (sync call sites).
"""
# Metacognitive nudges live on the message dict's ``_reminders``
# side-channel — user messages carry user-channel nudges
@@ -369,6 +459,18 @@ def _build_history(
# ``content`` never carries the ``<system-reminder>`` envelope —
# that splice is transient, applied to a wire-bound copy in
# ``ChatSession._apply_reminders_for_provider``.
#
# Verdict + output-assessment lookup tables — populated either
# inline (sync call sites) or pre-loaded by an async caller via
# asyncio.to_thread (see _load_verdict_indexes). Pre-loading is
# what keeps _build_history off the event loop's hot path on the
# SSE replay generator path.
if verdicts is not None and assessments is not None:
verdicts_by_call_id = verdicts
assessments_by_call_id = assessments
else:
ws_id = getattr(session, "_ws_id", "") or ""
verdicts_by_call_id, assessments_by_call_id = _load_verdict_indexes(ws_id)
history = []
for msg in session.messages:
content = msg.get("content")
@@ -433,17 +535,45 @@ def _build_history(
if clean_reminders:
entry["reminders"] = clean_reminders
if msg.get("tool_calls"):
entry["tool_calls"] = [
{
"id": tc.get("id", ""),
tc_entries: list[dict[str, Any]] = []
for tc in msg["tool_calls"]:
tc_entry: dict[str, Any] = {
"id": tc.get("id", "") or "",
"name": tc["function"]["name"],
"arguments": tc["function"].get("arguments", ""),
}
for tc in msg["tool_calls"]
]
# Decorate with persisted verdict + output_assessment
# via the shared helper (also used by
# ``make_history_handler``). Skips unflagged
# ("risk_level == 'none'") rows so the wire stays
# tight; ships only the fields the UI renders.
_decorate_tool_call(
tc_entry,
verdicts_by_call_id,
assessments_by_call_id,
)
tc_entries.append(tc_entry)
entry["tool_calls"] = tc_entries
# Detect denied/blocked/errored tool results by their content prefix.
if msg.get("role") == "tool":
content = msg.get("content", "")
# Propagate tool_call_id so replayHistory can anchor the
# rendered output to the specific .ts-approval-tool element
# by data-call-id (mirrors the live appendToolOutput path).
# Without this, multi-tool batches render every result at
# the bottom of the block rather than under each header.
result_call_id = msg.get("tool_call_id")
if result_call_id:
entry["tool_call_id"] = str(result_call_id)
# Tool results are clamped to TOOL_RESULT_STORAGE_CAP
# chars per row at storage time (session.py). Surface
# that on replay so the user knows the visible output is
# a clipped view of what the live session saw, rather
# than the full result. Reference the shared constant
# rather than a literal so the UI pill logic can't
# silently desync if the cap ever changes.
if isinstance(content, str) and len(content) >= TOOL_RESULT_STORAGE_CAP:
entry["truncated"] = True
if isinstance(content, str):
if content.startswith("Denied by user") or content.startswith("Blocked"):
entry["denied"] = True
@@ -684,6 +814,31 @@ def _audit_close_workstream(
)
async def _interactive_events_replay_prepare(ws: Workstream, ui: Any, request: Request) -> None:
"""Async pre-step run before ``_interactive_events_replay`` iterates.
Loads ``intent_verdicts`` + ``output_assessments`` for the
workstream off the event loop (via ``asyncio.to_thread``) and
stashes the result on ``request.state.verdict_indexes``. The sync
replay generator reads from there and passes the dicts into
``_build_history`` so the storage I/O never blocks the event loop
on the SSE replay path.
Best-effort: if the workstream has no session or no ws_id, leaves
``request.state.verdict_indexes`` unset and ``_build_history``
falls back to the inline storage call (sync path).
"""
del ui # not needed; lookup is keyed on ws.session._ws_id
session = ws.session
if session is None:
return
ws_id = getattr(session, "_ws_id", "") or ""
if not ws_id:
return
indexes = await asyncio.to_thread(_load_verdict_indexes, ws_id)
request.state.verdict_indexes = indexes
def _interactive_events_replay(
ws: Workstream, ui: Any, request: Request
) -> Iterable[dict[str, Any]]:
@@ -701,7 +856,6 @@ def _interactive_events_replay(
Pure read never mutates ``ws`` / ``ui`` / ``session``.
"""
del request # not needed; replay reads ws/ui/session state
session = ws.session
if session is None:
# Defensive — the lifted body's UI presence check guarantees
@@ -716,9 +870,22 @@ def _interactive_events_replay(
# History replay — pending-approval flag rides on the last
# assistant entry's tool_calls so the client renders them as
# awaiting approval rather than already approved.
# awaiting approval rather than already approved. Verdict /
# assessment indexes were pre-loaded off the event loop by
# _interactive_events_replay_prepare; passing them in here keeps
# _build_history's storage I/O out of the sync generator path.
pending_approval = getattr(ui, "_pending_approval", None)
history = _build_history(session, has_pending_approval=pending_approval is not None)
cached_indexes = getattr(request.state, "verdict_indexes", None)
if isinstance(cached_indexes, tuple) and len(cached_indexes) == 2:
verdicts, assessments = cached_indexes
else:
verdicts, assessments = None, None
history = _build_history(
session,
has_pending_approval=pending_approval is not None,
verdicts=verdicts,
assessments=assessments,
)
if history:
yield {"type": "history", "messages": history}
@@ -1393,13 +1560,13 @@ async def command(request: Request) -> JSONResponse:
ui._enqueue({"type": "clear_ui"})
elif cmd_word == "/resume":
ui._enqueue({"type": "clear_ui"})
history = _build_history(ws.session)
history = await asyncio.to_thread(_build_history, ws.session)
if history:
ui._enqueue({"type": "history", "messages": history})
elif cmd_word in ("/rewind", "/retry"):
# Refresh frontend with truncated history
ui._enqueue({"type": "clear_ui"})
history = _build_history(ws.session)
history = await asyncio.to_thread(_build_history, ws.session)
if history:
ui._enqueue({"type": "history", "messages": history})
# Audit trail
@@ -1874,7 +2041,7 @@ async def _interactive_create_post_install(
ui = ws.ui
if isinstance(ui, WebUI):
ui._enqueue({"type": "clear_ui"})
history = _build_history(ws.session)
history = await asyncio.to_thread(_build_history, ws.session)
if history:
ui._enqueue({"type": "history", "messages": history})
with contextlib.suppress(queue.Full):
@@ -2800,13 +2967,14 @@ def internal_model_reload(request: Request) -> JSONResponse:
if registry is None or cli_args is None:
return JSONResponse({"status": "error", "reason": "no registry"}, status_code=503)
storage = get_storage()
new_registry = load_model_registry(
base_url=cli_args["base_url"],
api_key=cli_args["api_key"],
model=cli_args["model"],
context_window=cli_args["context_window"],
provider=cli_args["provider"],
storage=get_storage(),
storage=storage,
)
cs = getattr(request.app.state, "config_store", None)
if cs is not None:
@@ -2872,6 +3040,13 @@ def internal_model_reload(request: Request) -> JSONResponse:
# `model` parameter descriptions reflect the current registry.
_broadcast_agent_tool_schema_refresh(request.app.state)
# Refresh the per-node ``models`` metadata entry the coord reads on
# ``list_nodes``. Without this, the heartbeat loop's 30s tick would
# be the coord's first chance to see new aliases an admin just added.
node_id = getattr(request.app.state, "node_id", "")
if node_id:
_publish_models_metadata(request.app.state, storage, node_id)
return JSONResponse({"status": "ok", "aliases": registry.list_aliases()})
@@ -2897,6 +3072,109 @@ def internal_model_status(request: Request) -> JSONResponse:
return JSONResponse({"models": models})
def _collect_node_models_metadata(app_state: Any) -> tuple[str, str, str] | None:
"""Build the ``("models", json_value, "auto")`` node_metadata entry.
Each model alias on the live registry is projected to
``{alias, provider, healthy}`` the alias is what the coordinator
passes back as ``spawn_workstream(model=...)``, ``provider`` lets
coordinators classify or filter (e.g. "any anthropic node"), and
``healthy`` reflects the backend's :class:`BackendHealthTracker`
state at call time. The underlying model identifier (``cfg.model``)
is intentionally omitted coordinators kept reaching for the
provider-side string when they should have been passing the local
alias, and dropping it removes the footgun. Operators who need
the model string can hit ``/v1/api/_internal/model-status`` on the
node directly.
Trackers are eagerly seeded for every alias at server startup and
on every model-reload, so ``health_reg.get_tracker(...)`` returns
the existing tracker rather than minting a fresh one in steady
state. In the unlikely race where a tracker hasn't been seeded
yet, the freshly created tracker reports ``is_healthy=True``
(default state) which matches the prior "default to True when
no tracker" behavior, just routed through the tracker object.
Returns ``None`` when the registry has not yet been built (caller
should skip the write rather than zero out a previous snapshot).
"""
registry = getattr(app_state, "registry", None)
if registry is None:
return None
health_reg = getattr(app_state, "health_registry", None)
aliases_info: list[dict[str, Any]] = []
# Iterate aliases in a stable order — ``list_aliases`` returns dict
# insertion order, so two structurally identical registries built
# from different sources (config.toml vs. DB rows in different
# commit order) would otherwise serialize to different JSON and
# defeat the publish-cache hit-rate that the
# ``turnstone_node_models_publish_total`` metric tracks.
for alias in sorted(registry.list_aliases()):
try:
cfg = registry.get_config(alias)
except (ValueError, KeyError):
continue
healthy = True
if health_reg is not None:
# Direct keyed lookup — ``get_tracker_for_alias`` would
# do a second ``registry.get_config(alias)`` internally,
# but ``cfg`` is already in hand here.
tracker = health_reg.get_tracker(provider=cfg.provider, base_url=cfg.base_url)
healthy = tracker.is_healthy
aliases_info.append(
{
"alias": alias,
"provider": cfg.provider,
"healthy": healthy,
}
)
return ("models", json.dumps(aliases_info), "auto")
def _publish_models_metadata(app_state: Any, storage: Any, node_id: str) -> None:
"""Refresh the per-node ``models`` row when the projection changed.
Caches the last-written JSON on ``app_state._last_models_payload``
so back-to-back heartbeat ticks with no health flip don't churn
the row without this, the ``updated`` timestamp on every node's
``models`` row advances every 30s across the whole cluster.
Records the cache outcome on the metrics collector so
``turnstone_node_models_publish_total{outcome=...}`` exposes the
hit/miss ratio to Prometheus. Storage-error attempts don't
record either outcome the next call will retry and the
counters reflect actual cache decisions, not transient DB
failures.
Sync callers on the asyncio loop wrap with ``asyncio.to_thread``.
Concurrent callers (heartbeat tick vs. ``internal_model_reload``)
can race on the cache attribute; the worst case is a redundant
write, never a stale row, so we skip the lock.
"""
from turnstone.core.storage._registry import StorageUnavailableError
try:
entry = _collect_node_models_metadata(app_state)
except Exception:
log.warning("server.node_models_projection_failed", exc_info=True)
return
if entry is None:
return
payload = entry[1]
if payload == getattr(app_state, "_last_models_payload", None):
_metrics.record_node_models_publish(written=False)
return
try:
storage.set_node_metadata_bulk(node_id, [entry])
except StorageUnavailableError:
return # storage layer already logged
except Exception:
log.exception("server.node_models_publish_failed")
return
app_state._last_models_payload = payload
_metrics.record_node_models_publish(written=True)
# ---------------------------------------------------------------------------
# Global SSE fan-out
# ---------------------------------------------------------------------------
@@ -3173,6 +3451,22 @@ async def _lifespan(app: Starlette) -> AsyncGenerator[None, None]:
]
_cfg_meta = _load_meta_config("metadata")
_meta_entries.extend((k, json.dumps(v), "config") for k, v in _cfg_meta.items())
# Project the live model registry into a ``models`` entry so
# coord-side ``list_nodes`` can surface healthy aliases per
# node without a fan-out HTTP probe. Re-collected on each
# heartbeat tick so health flips converge within ~30s.
# Wrapped in its own try/except so a projection failure
# doesn't take out the auto+config metadata write — losing
# the discovery surface is recoverable on the next heartbeat
# tick, but losing ``arch`` / ``os`` / ``cpu_count`` blinds
# the cluster's capability filters until the next restart.
try:
_models_entry = _collect_node_models_metadata(app.state)
except Exception:
log.warning("server.node_models_projection_failed", exc_info=True)
_models_entry = None
if _models_entry is not None:
_meta_entries.append(_models_entry)
if _meta_entries:
# Clear stale auto/config rows from a prior run before upserting
_svc_storage.delete_node_metadata_by_source(_svc_node_id, "auto")
@@ -3183,11 +3477,26 @@ async def _lifespan(app: Starlette) -> AsyncGenerator[None, None]:
node_id=_svc_node_id,
count=len(_meta_entries),
)
# Seed the publish-cache so the first heartbeat tick
# doesn't redundant-write the same payload we just put
# in the bulk above.
if _models_entry is not None:
app.state._last_models_payload = _models_entry[1]
except Exception:
log.warning("server.node_metadata_failed", node_id=_svc_node_id, exc_info=True)
async def _heartbeat_loop() -> None:
"""Periodically update service heartbeat."""
"""Periodically update service heartbeat and refresh models metadata.
The ``models`` entry on ``node_metadata`` doubles as the
coord-side discovery surface for healthy model aliases per
node refreshed every 30s so health flips and registry
reloads converge promptly without a fan-out HTTP probe on
the coord's ``list_nodes`` path. The publish step short-
circuits when the projection is byte-identical to the
last write (cache lives on ``app.state``), so a stable
cluster doesn't pay UPSERT churn here.
"""
from turnstone.core.storage._registry import StorageUnavailableError
while True:
@@ -3198,6 +3507,12 @@ async def _lifespan(app: Starlette) -> AsyncGenerator[None, None]:
pass # already logged by storage layer
except Exception:
log.exception("server.heartbeat_failed")
# Both projection and write happen in the worker thread
# — keeps the registry-lock acquisition off the loop and
# bundles the round-trip into a single offload.
await asyncio.to_thread(
_publish_models_metadata, app.state, _svc_storage, _svc_node_id
)
_heartbeat_task = asyncio.create_task(_heartbeat_loop())
@@ -3205,6 +3520,13 @@ async def _lifespan(app: Starlette) -> AsyncGenerator[None, None]:
# Shutdown
if _heartbeat_task is not None:
_heartbeat_task.cancel()
# Wait for the cancel to land before we run the metadata
# delete below — a heartbeat tick mid-write would otherwise
# complete its ``set_node_metadata_bulk`` AFTER our
# ``delete_node_metadata_by_source(..., "auto")`` and
# resurrect the row we just cleared.
with contextlib.suppress(asyncio.CancelledError, Exception):
await _heartbeat_task
if _svc_node_id and _svc_url:
from turnstone.core.storage import get_storage as _get_svc_dereg
@@ -3370,6 +3692,7 @@ def create_app(
open_resolve_alias=_resolve_workstream_alias,
open_post_load=_interactive_open_post_load,
events_replay=_interactive_events_replay,
events_replay_prepare=_interactive_events_replay_prepare,
# Pre-lift ``events_sse`` used the dedicated 200-thread
# ``sse_executor`` so SSE polling stayed isolated from
# every other ``asyncio.to_thread`` caller in the process
@@ -3851,6 +4174,13 @@ def main() -> None:
assert ui is not None
# Resolve the effective alias once and use it consistently
# for both client resolution and ChatSession.model_alias.
# Unknown aliases here raise ValueError — the create handler
# maps that to a 503 with operator-friendly text so a typo or
# removed alias in body.model surfaces instead of silently
# starting on the default. SessionManager.open's rehydrate
# path is the one place where unknown aliases must NOT fail
# loud; the manager filters those out via its model_validator
# before the alias reaches this factory.
model_alias = model_alias or _effective_default_alias()
r_client, r_model, r_cfg = registry.resolve(model_alias)
# Read MCP client from shared ref — may have been replaced after startup
@@ -3980,6 +4310,10 @@ def main() -> None:
# emit_rehydrated are no-ops because those events fire from
# out-of-band paths (create handler + WebUI._broadcast_state).
event_emitter=interactive_adapter,
# Filter out persisted aliases that no longer resolve so a
# workstream pinned to a since-removed alias still rehydrates
# (on the registry default) instead of 500-ing on every reopen.
model_validator=registry.has_alias,
)
interactive_adapter.attach(manager)
WebUI._workstream_mgr = manager
+283 -3
View File
@@ -3,9 +3,9 @@
console/static (Saved Coordinators). Single source of truth so the two
surfaces don't drift on hover affordance, padding, or typography.
==========================================================================
Class names match the original ui/static rules they replaced; the
delete-mode subset stays in ui/static/style.css until coordinator gets
the same UX (then it can move here too).
Class names match the original ui/static rules they replaced. Delete-mode
selectors live here too so console (Saved Coordinators) and ui/static
(Saved Workstreams) share one card + delete affordance.
========================================================================== */
.dashboard-cards {
@@ -76,3 +76,283 @@
grid-template-columns: 1fr;
}
}
/* ==========================================================================
Delete UX section-level "Delete" toggle, per-card checkboxes, bottom
toolbar, and confirmation modal. Moved out of ui/static/style.css when
the console grew the same multi-select delete on Saved Coordinators.
========================================================================== */
.ws-delete-btn {
background: transparent;
border: 1px solid var(--border);
border-radius: var(--radius);
color: var(--fg-dim);
font-size: 12px;
padding: 4px 10px;
cursor: pointer;
transition:
color 0.15s,
border-color 0.15s;
}
.ws-delete-btn:hover {
color: var(--red);
border-color: var(--red);
}
/* Delete mode */
.dashboard-card.ws-delete-mode {
cursor: pointer;
}
.dashboard-card.ws-delete-mode:hover {
border-color: var(--red);
background: rgba(248, 113, 113, 0.04);
}
.dashboard-card.ws-delete-mode.ws-selected {
cursor: default;
}
.dashboard-card.ws-delete-mode.ws-selected:hover {
border-color: var(--red);
background: rgba(248, 113, 113, 0.08);
}
[data-theme="light"] .dashboard-card.ws-delete-mode:hover {
background: rgba(220, 38, 38, 0.04);
}
[data-theme="light"] .dashboard-card.ws-delete-mode.ws-selected:hover {
background: rgba(220, 38, 38, 0.08);
}
.ws-card-check {
position: absolute;
top: 8px;
right: 8px;
width: 18px;
height: 18px;
accent-color: var(--red);
cursor: pointer;
z-index: 1;
opacity: 0;
animation: ws-check-fadein 0.2s ease-out forwards;
}
.ws-card-check:focus-visible {
/* Card click-handler proxies space/enter to the checkbox, so focus
usually rests on the card; if a screen reader / power user tabs
directly onto the checkbox the UA outline can be killed by
adjacent rules this guarantees a visible affordance. */
outline: 2px solid var(--accent);
outline-offset: 2px;
}
@keyframes ws-check-fadein {
to {
opacity: 1;
}
}
.dashboard-card.ws-selected {
border-color: var(--red);
background: rgba(248, 113, 113, 0.08);
}
[data-theme="light"] .dashboard-card.ws-selected {
background: rgba(220, 38, 38, 0.08);
}
.ws-delete-bar {
display: none;
align-items: center;
gap: 12px;
margin-top: 12px;
padding: 8px 12px;
background: var(--bg-surface);
border: 1px solid var(--border);
border-radius: var(--radius);
}
.ws-delete-bar.visible {
display: flex;
animation: ws-bar-slide 0.2s ease-out;
}
@keyframes ws-bar-slide {
from {
opacity: 0;
transform: translateY(-8px);
}
to {
opacity: 1;
transform: translateY(0);
}
}
@media (prefers-reduced-motion: reduce) {
.ws-card-check {
animation: none;
opacity: 1;
}
.ws-delete-bar.visible {
animation: none;
}
}
/* Below ~700px the four-pill bar's preferred width (~480-520px) starts
eating into hit-targets and risking horizontal overflow. Wrap onto
two rows: count + Cancel + Select All on the first, the destructive
Delete Selected on its own full-width row underneath also a better
thumb-target separation than the desktop layout. */
@media (max-width: 700px) {
.ws-delete-bar {
flex-wrap: wrap;
gap: 8px;
}
.ws-delete-bar .ws-delete-bar-btn {
margin-left: 0;
flex: 1 1 100%;
order: 99;
}
}
.ws-delete-bar .ws-delete-count-label {
font-size: 12px;
color: var(--fg-dim);
}
.ws-delete-bar .ws-delete-bar-btn {
margin-left: auto;
/* Dark-theme --red (#f87171) on #fff is only 3.0:1 below WCAG AA
for normal text on a destructive button. Use the deeper red
(#dc2626 4.85:1) on the filled state so the button label clears
AA in the default theme. Light theme already uses --red (#b91c1c,
5.9:1) and stays put the override below pins it. */
background: #dc2626;
color: #fff;
border: none;
border-radius: var(--radius);
padding: 6px 16px;
font-size: 12px;
font-weight: 600;
cursor: pointer;
transition: filter 0.15s;
}
[data-theme="light"] .ws-delete-bar .ws-delete-bar-btn {
background: var(--red);
}
.ws-delete-bar .ws-delete-bar-btn:hover:not(:disabled) {
filter: brightness(1.1);
}
.ws-delete-bar .ws-delete-bar-btn:disabled {
opacity: 0.4;
cursor: not-allowed;
}
.ws-delete-bar .ws-delete-cancel-btn {
background: transparent;
color: var(--fg-dim);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 6px 12px;
font-size: 12px;
cursor: pointer;
}
.ws-delete-bar .ws-delete-cancel-btn:hover {
color: var(--fg-bright);
border-color: var(--border-strong);
background: var(--bg-highlight);
}
.ws-delete-bar .ws-delete-selectall-btn {
background: transparent;
color: var(--fg-bright);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 6px 12px;
font-size: 12px;
font-weight: 500;
cursor: pointer;
}
.ws-delete-bar .ws-delete-selectall-btn:hover {
border-color: var(--border-strong);
background: var(--bg-highlight);
}
/* Delete modal id-scoped so the surface that owns it controls visibility.
Both ui/static (#ws-delete-overlay) and console/static
(#coord-delete-overlay) share the same shape via the .ws-delete-modal
class hooks below. */
.ws-delete-modal-overlay {
position: fixed;
inset: 0;
background: rgba(0, 0, 0, 0.6);
display: flex;
align-items: center;
justify-content: center;
z-index: 10000;
}
.ws-delete-modal-box {
background: var(--bg-surface);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 24px;
max-width: 480px;
width: 90%;
}
.ws-delete-modal-box h3 {
margin: 0 0 12px;
font-size: 16px;
color: var(--fg-bright);
}
.ws-delete-modal-list {
max-height: 200px;
overflow-y: auto;
margin: 12px 0;
}
.ws-delete-modal-list .ws-delete-item {
padding: 6px 0;
font-size: 13px;
color: var(--fg-bright);
border-bottom: 1px solid var(--border);
/* Long aliases / raw ws_ids in the confirm + results list shouldn't
punch out of the modal at narrow viewports. */
word-break: break-word;
}
/* Modal alert region only painted when the controller writes a
message. Both close paths clear it, so :not(:empty) keeps the box
invisible at rest and avoids an empty-frame artefact. */
.ws-delete-modal-box [role="alert"]:not(:empty) {
background: rgba(248, 113, 113, 0.08);
border: 1px solid var(--red);
color: var(--red);
border-radius: var(--radius);
padding: 8px 10px;
font-size: 12px;
margin-bottom: 8px;
}
[data-theme="light"] .ws-delete-modal-box [role="alert"]:not(:empty) {
background: rgba(220, 38, 38, 0.06);
}
.ws-delete-modal-list .ws-delete-item:last-child {
border-bottom: none;
}
.ws-delete-modal-list .ws-delete-item.ws-delete-error {
color: var(--red);
}
.ws-delete-modal-buttons {
display: flex;
gap: 8px;
justify-content: flex-end;
margin-top: 16px;
}
.ws-delete-modal-buttons button {
padding: 8px 20px;
border-radius: var(--radius);
font-size: 13px;
font-weight: 500;
cursor: pointer;
border: 1px solid var(--border);
background: transparent;
color: var(--fg-bright);
}
.ws-delete-modal-buttons button.ws-delete-confirm {
/* Mirror .ws-delete-bar-btn's contrast bump same destructive
filled-button treatment, same dark-theme AA fix. */
background: #dc2626;
color: #fff;
border-color: #dc2626;
}
[data-theme="light"] .ws-delete-modal-buttons button.ws-delete-confirm {
background: var(--red);
border-color: var(--red);
}
/* `.ws-delete-close` is a state-marker, not a colour rule the
controller drops `.ws-delete-confirm` when the modal transitions to
the post-delete "Close" state, and the default
`.ws-delete-modal-buttons button` rule above already provides the
transparent / fg-bright / border styling. The class itself is useful
for DOM inspection and as a future hook. */
+430
View File
@@ -77,3 +77,433 @@ function renderSessionCard(sess, opts) {
card.appendChild(metaEl);
return card;
}
/* createSavedCardsController shared multi-select-delete behaviour for
the dashboard / home "saved cards" surfaces. ui/static (Saved
Workstreams) and console/static (Saved Coordinators) both instantiate
one of these; the controller owns:
- delete-mode state (active flag + selected ws_id set)
- card decoration (checkbox + key/click overrides)
- the bottom toolbar wiring (count, Select All, Delete Selected)
- the confirmation modal (focus trap, batch fan-out, results view)
It does NOT own how cards get fetched or rendered the caller's
render() is invoked when the controller needs the list redrawn (mode
transitions, Select-All toggles).
Required opts:
idPrefix DOM-id prefix shared by the toolbar + modal
(e.g. "ws-delete" / "coord-delete"). The DOM
must already contain `${idPrefix}-bar`,
`${idPrefix}-bar-count`, `${idPrefix}-bar-delete`,
`${idPrefix}-bar-select-all`, `${idPrefix}-overlay`,
`${idPrefix}-box`, `${idPrefix}-error`,
`${idPrefix}-count`, `${idPrefix}-list`,
`${idPrefix}-confirm-btn`, `${idPrefix}-cancel-btn`.
buttonId id of the section's start/cancel toggle button.
noun singular display word for the item kind, e.g.
"workstream" / "coordinator". Used in toast +
modal copy.
activateLabel sess => string; aria-label for the card when NOT
in delete mode (e.g. "Resume: foo").
buildDeleteRequest wsId => { url, options }; what authFetch should
send to delete one item.
render () => void; redraw the visible cards. Called by
the controller on mode start/cancel and Select-
All toggle. Caller is responsible for calling
setItems(items) + decorateCard() inside it.
onClose optional () => void; called once after the user
closes the post-delete results modal. Typical
use: re-fetch the saved list.
*/
function createSavedCardsController(opts) {
var state = { mode: false, selected: {}, items: [] };
var batchTrap = null;
/* Element that owned focus when the modal opened restored in
closeModal() so keyboard users land back on the toggle button (or
wherever they came from) instead of <body>. WCAG 2.4.3. */
var prevFocus = null;
function $(id) {
return document.getElementById(opts.idPrefix + "-" + id);
}
/* Replace the toggle button's content with a glyph + label, keeping
the glyph in an aria-hidden span so screen readers only read the
label. Built from DOM nodes (no innerHTML) same shape as the
section-header markup the JS replaces. */
function setIconButton(btn, glyph, label) {
btn.replaceChildren();
var span = document.createElement("span");
span.setAttribute("aria-hidden", "true");
span.textContent = glyph;
btn.appendChild(span);
btn.appendChild(document.createTextNode(" " + label));
}
function setItems(items) {
state.items = items;
/* Drop any selections whose ws_id is no longer on the visible page
SSE-driven re-renders or pagination jumps shouldn't leave ghost
entries inflating the count and 404-ing on confirm. */
if (state.mode) {
var byId = {};
items.forEach(function (s) {
byId[s.ws_id] = true;
});
Object.keys(state.selected).forEach(function (id) {
if (!byId[id]) delete state.selected[id];
});
}
}
function inMode() {
return state.mode;
}
function blockActivate() {
return state.mode;
}
function isSelected(wsId) {
return !!state.selected[wsId];
}
function ariaLabel(sess) {
var label = sess.alias || sess.title || sess.name || sess.ws_id;
if (state.mode) return "Select " + opts.noun + ": " + label;
return typeof opts.activateLabel === "function"
? opts.activateLabel(sess)
: "Activate: " + label;
}
/* Decorate an already-rendered .dashboard-card with the checkbox +
event overrides used in delete mode. Idempotent guard: only acts
when the controller is active. */
function decorateCard(card, sess) {
if (!state.mode) return;
card.classList.add("ws-delete-mode");
card.removeAttribute("role");
var chk = document.createElement("input");
chk.type = "checkbox";
chk.className = "ws-card-check";
chk.checked = !!state.selected[sess.ws_id];
var label = sess.alias || sess.title || sess.name || sess.ws_id;
chk.setAttribute("aria-label", "Select " + label + " for deletion");
chk.onclick = function (e) {
e.stopPropagation();
if (chk.checked) state.selected[sess.ws_id] = true;
else delete state.selected[sess.ws_id];
card.classList.toggle("ws-selected", chk.checked);
refreshBar();
};
card.insertBefore(chk, card.firstChild);
card.onclick = function (e) {
if (e.target === chk) return;
chk.checked = !chk.checked;
chk.onclick(e);
};
card.onkeydown = function (e) {
if (e.key === "Enter" || e.key === " ") {
e.preventDefault();
chk.checked = !chk.checked;
chk.onclick(e);
}
};
if (state.selected[sess.ws_id]) card.classList.add("ws-selected");
}
function refreshBar() {
var count = Object.keys(state.selected).length;
var label = $("bar-count");
if (label) label.textContent = count + " selected";
var delBtn = $("bar-delete");
if (delBtn) delBtn.disabled = count === 0;
var selBtn = $("bar-select-all");
if (selBtn) {
var allSelected = count === state.items.length && state.items.length > 0;
selBtn.textContent = allSelected ? "Deselect All" : "Select All";
}
}
function start() {
if (!state.items.length) {
if (typeof showToast === "function") {
showToast("No saved " + opts.noun + "s to delete");
}
return;
}
state.mode = true;
state.selected = {};
opts.render();
var btn = document.getElementById(opts.buttonId);
if (btn) {
setIconButton(btn, "✕", "Cancel");
btn.onclick = cancel;
}
var bar = $("bar");
if (bar) bar.classList.add("visible");
refreshBar();
}
function cancel() {
state.mode = false;
state.selected = {};
opts.render();
var btn = document.getElementById(opts.buttonId);
if (btn) {
setIconButton(btn, "\u{1f5d1}", "Delete");
btn.onclick = start;
}
var bar = $("bar");
if (bar) bar.classList.remove("visible");
}
function toggleAll() {
var allSelected =
Object.keys(state.selected).length === state.items.length &&
state.items.length > 0;
if (allSelected) {
state.selected = {};
} else {
state.items.forEach(function (s) {
state.selected[s.ws_id] = true;
});
}
opts.render();
refreshBar();
}
function _byId() {
/* Single-pass index over the visible items so the modal + fan-out
paths don't repeat O(N) `find` calls per selection. */
var map = {};
state.items.forEach(function (s) {
map[s.ws_id] = s;
});
return map;
}
function confirmSelection() {
var selected = Object.keys(state.selected);
if (!selected.length) {
if (typeof showToast === "function") {
showToast("No " + opts.noun + "s selected");
}
return;
}
var byId = _byId();
var overlay = $("overlay");
var countEl = $("count");
var listEl = $("list");
var errorEl = $("error");
if (errorEl) errorEl.textContent = "";
if (countEl) {
countEl.textContent =
selected.length + " " + opts.noun + "(s) will be permanently deleted:";
}
if (listEl) {
listEl.replaceChildren();
selected.forEach(function (wsId) {
var item = byId[wsId];
var name = item ? item.alias || item.title || item.name || wsId : wsId;
var div = document.createElement("div");
div.className = "ws-delete-item";
div.textContent = name;
listEl.appendChild(div);
});
}
var delBtn = $("confirm-btn");
if (delBtn) {
delBtn.textContent = "Delete";
delBtn.disabled = false;
delBtn.classList.remove("ws-delete-close");
delBtn.classList.add("ws-delete-confirm");
delBtn.onclick = confirm;
}
var cancelBtn = $("cancel-btn");
if (cancelBtn) cancelBtn.disabled = false;
if (overlay) overlay.style.display = "flex";
if (batchTrap) document.removeEventListener("keydown", batchTrap);
batchTrap = function (e) {
if (e.key === "Escape") {
e.preventDefault();
closeModal();
return;
}
if (e.key === "Tab") {
var box = $("box");
if (!box) return;
var focusable = box.querySelectorAll("button:not(:disabled)");
if (!focusable.length) return;
var first = focusable[0];
var last = focusable[focusable.length - 1];
if (e.shiftKey && document.activeElement === first) {
e.preventDefault();
last.focus();
} else if (!e.shiftKey && document.activeElement === last) {
e.preventDefault();
first.focus();
}
}
};
document.addEventListener("keydown", batchTrap);
/* Snapshot the pre-modal focus owner so closeModal() can return to
it. Captured before we move focus into the dialog so the
restore-target is the caller, not the dialog itself. */
prevFocus = document.activeElement;
if (cancelBtn) cancelBtn.focus();
}
function closeModal() {
var overlay = $("overlay");
if (overlay) overlay.style.display = "none";
if (batchTrap) {
document.removeEventListener("keydown", batchTrap);
batchTrap = null;
}
/* Pick the most useful focus target:
1. prevFocus (where the user came from), if it's still in the
DOM and visible. Esc / Cancel paths land here the bar is
still on screen, so focus returns to "Delete Selected".
2. The section toggle button always present, semantic exit
point for the flow. Used when prevFocus has been hidden by
cancel() (post-delete Close path: cancel() ran first and
put `.ws-delete-bar` at display:none, so the bar's button
is no longer focusable). */
var target = prevFocus;
if (!target || target.offsetParent === null) {
target = document.getElementById(opts.buttonId);
}
if (target && typeof target.focus === "function") {
try {
target.focus();
} catch (_) {
/* node detached between open and close — give up silently */
}
}
prevFocus = null;
}
function confirm() {
var selected = Object.keys(state.selected);
if (!selected.length) return;
var byId = _byId();
var errorEl = $("error");
var listEl = $("list");
var countEl = $("count");
var delBtn = $("confirm-btn");
var cancelBtn = $("cancel-btn");
if (errorEl) errorEl.textContent = "";
if (delBtn) {
delBtn.disabled = true;
delBtn.textContent = "Deleting...";
}
if (cancelBtn) cancelBtn.disabled = true;
var results = [];
var promises = selected.map(function (wsId) {
var shortId = wsId.substring(0, 8);
var item = byId[wsId];
var name = item ? item.alias || item.title || item.name || wsId : wsId;
var req = opts.buildDeleteRequest(wsId);
return authFetch(req.url, req.options)
.then(function (r) {
var status = r.status;
var contentType = r.headers.get("content-type") || "";
if (r.ok) {
results.push({ name: name, shortId: shortId, ok: true });
return;
}
return r.text().then(function (body) {
var errMsg = shortId + ": HTTP " + status;
if (contentType.includes("json")) {
try {
var j = JSON.parse(body);
if (j.error) errMsg = shortId + ": " + j.error;
} catch (_) {
/* fall through */
}
} else if (body) {
errMsg = shortId + ": " + body.substring(0, 200);
}
results.push({
name: name,
shortId: shortId,
ok: false,
error: errMsg,
});
});
})
.catch(function (err) {
results.push({
name: name,
shortId: shortId,
ok: false,
error: shortId + ": " + err.message,
});
});
});
Promise.all(promises).then(function () {
if (listEl) {
listEl.replaceChildren();
results.forEach(function (r) {
var div = document.createElement("div");
div.className = "ws-delete-item" + (r.ok ? "" : " ws-delete-error");
div.textContent =
(r.ok ? "✓ " : "✗ ") + r.name + (r.error ? " — " + r.error : "");
listEl.appendChild(div);
});
}
var okCount = results.filter(function (r) {
return r.ok;
}).length;
var failCount = results.filter(function (r) {
return !r.ok;
}).length;
if (countEl) {
countEl.textContent = okCount + " deleted, " + failCount + " failed";
}
if (delBtn) {
delBtn.disabled = false;
delBtn.textContent = "Close";
/* Swap modifier classes so styling is intent-driven instead of
cascade-positional: the Close button picks up the default
".ws-delete-modal-buttons button" rule once .ws-delete-confirm
is removed. */
delBtn.classList.remove("ws-delete-confirm");
delBtn.classList.add("ws-delete-close");
delBtn.onclick = function () {
/* Order matters: cancel() reshapes the toggle button via
setIconButton(), which preserves the element identity but
swaps its subtree. closeModal() then focuses prevFocus
which IS that toggle button landing on a freshly rebuilt
"Delete" affordance instead of <body>. */
cancel();
closeModal();
if (typeof opts.onClose === "function") opts.onClose();
};
}
if (cancelBtn) cancelBtn.disabled = false;
});
}
return {
setItems: setItems,
inMode: inMode,
blockActivate: blockActivate,
isSelected: isSelected,
ariaLabel: ariaLabel,
decorateCard: decorateCard,
refreshBar: refreshBar,
start: start,
cancel: cancel,
toggleAll: toggleAll,
confirmSelection: confirmSelection,
closeModal: closeModal,
confirm: confirm,
};
}
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "list_nodes",
"description": "List active cluster nodes with their metadata. By default only nodes with a fresh service-registry heartbeat (within 120s) are returned. Pass arbitrary `key=value` filters to narrow; all filters must match (AND). Pair with `target_node` on spawn_workstream to pin a child to a node that matches a capability. The 120s heartbeat is a sliding window, so a node returned here can drop out before a follow-up spawn lands — the race produces `\"No available node for routing\"`; omit `target_node` to let rendezvous pick from the still-healthy set, or retry after re-listing if a specific node is required. The `interfaces` key (container IPs, interface names) is stripped by default — routing should use capability/region tags, not IPs; pass `include_network_detail=true` only for debugging. Pass `include_inactive=true` to surface stale registrations (those nodes will reject `target_node` pinning).",
"description": "List active cluster nodes with their metadata. By default only nodes with a fresh service-registry heartbeat (within 120s) are returned. Each row carries `node_id`, `metadata`, and `model_aliases` — the latter being a list of healthy model aliases the node will accept on `spawn_workstream(model=...)` / `spawn_batch` (refreshed every 30s by the node's heartbeat). Pass arbitrary `key=value` filters to narrow; all filters must match (AND). Pair with `target_node` on spawn_workstream to pin a child to a node that matches a capability. The 120s heartbeat is a sliding window, so a node returned here can drop out before a follow-up spawn lands — the race produces `\"No available node for routing\"`; omit `target_node` to let rendezvous pick from the still-healthy set, or retry after re-listing if a specific node is required. The `interfaces` key (container IPs, interface names) is stripped by default — routing should use capability/region tags, not IPs; pass `include_network_detail=true` only for debugging. Pass `include_inactive=true` to surface stale registrations (those nodes will reject `target_node` pinning).",
"parameters": {
"type": "object",
"properties": {
+1 -1
View File
@@ -24,7 +24,7 @@
},
"model": {
"type": "string",
"description": "Optional model alias."
"description": "Optional model alias. Discover available aliases per node via `list_nodes.model_aliases`."
},
"target_node": {
"type": "string",
+1 -1
View File
@@ -18,7 +18,7 @@
},
"model": {
"type": "string",
"description": "Optional model alias from the registry. Omit to use the coordinator's default model (or the one the skill prescribes)."
"description": "Optional model alias from the registry. Discover available aliases per node via `list_nodes` (the `model_aliases` field on each row lists the healthy aliases that node will accept). Omit to use the coordinator's default model (or the one the skill prescribes)."
},
"target_node": {
"type": "string",
+241 -298
View File
@@ -1036,6 +1036,19 @@ Pane.prototype.replayHistory = function (messages) {
this.showEmptyState();
return;
}
// Suppress the polite live region while we batch-build the replay
// — messagesEl is aria-live="polite" so a fresh replay would otherwise
// queue an announcement for every approved/denied/verdict pill we
// insert. Restored after the loop so live SSE updates announce
// normally. WCAG 4.1.3 — historical content should not behave like
// real-time updates.
this.messagesEl.setAttribute("aria-busy", "true");
// pendingAssessments[call_id] = output_assessment dict. Populated
// from the assistant branch, consumed by the role==="tool" branch
// (or after the loop, for legacy rows missing tool_call_id).
// Replaces a JSON.stringify→dataset→JSON.parse round-trip with an
// in-memory map keyed by call_id.
var pendingAssessments = {};
var lastToolBlock = null;
for (var i = 0; i < messages.length; i++) {
var msg = messages[i];
@@ -1052,6 +1065,25 @@ Pane.prototype.replayHistory = function (messages) {
}
lastToolBlock = null;
} else if (msg.role === "assistant") {
// Render content BEFORE the tool block so the visual order
// matches the live SSE flow (stream_text streams content first,
// then tool_info / approve_request paints the tool block, then
// tool_result fills it in). Order also matters structurally:
// the tool-result message in the NEXT iteration anchors via
// lastToolBlock, which the tool-block branch sets last — so
// content must run first to avoid clobbering that anchor.
if (msg.content) {
var el = document.createElement("div");
el.className = "msg assistant";
var bodyEl = document.createElement("div");
bodyEl.className = "msg-body";
var rendered = renderMarkdown(msg.content);
bodyEl.innerHTML = rendered;
el.appendChild(bodyEl);
postRenderMarkdown(el);
self.messagesEl.appendChild(el);
lastToolBlock = null;
}
if (msg.tool_calls && msg.tool_calls.length) {
if (msg.pending) {
lastToolBlock = null;
@@ -1096,7 +1128,35 @@ Pane.prototype.replayHistory = function (messages) {
cmd.textContent = tc.arguments.substring(0, 100);
}
div.appendChild(cmd);
// Verdict badge — anchor to THIS tool's row (div) rather
// than the whole block, so a multi-tool batch with one
// flagged call doesn't drift the badge above unrelated
// calls. Same renderVerdictBadge helper as live; pass
// judgePending=false because any verdict on replay is
// final — no spinner.
if (tc.verdict) {
div.insertAdjacentHTML(
"beforeend",
renderVerdictBadge(tc.verdict, false),
);
}
block.appendChild(div);
// Output-guard finding — defer insertion until the tool
// result lands so the warning anchors under the output
// (mirrors live showOutputWarning placement). Stash in
// a function-local map keyed by call_id so the
// role==="tool" branch below can pick it up; legacy rows
// missing tool_call_id are flushed at end-of-replay.
if (
tc.output_assessment &&
tc.output_assessment.risk_level &&
tc.output_assessment.risk_level !== "none"
) {
pendingAssessments[tc.id || ""] = {
assessment: tc.output_assessment,
toolDiv: div,
};
}
});
var badge = document.createElement("div");
badge.setAttribute("role", "status");
@@ -1112,18 +1172,6 @@ Pane.prototype.replayHistory = function (messages) {
lastToolBlock = block;
}
}
if (msg.content) {
var el = document.createElement("div");
el.className = "msg assistant";
var bodyEl = document.createElement("div");
bodyEl.className = "msg-body";
var rendered = renderMarkdown(msg.content);
bodyEl.innerHTML = rendered;
el.appendChild(bodyEl);
postRenderMarkdown(el);
self.messagesEl.appendChild(el);
lastToolBlock = null;
}
} else if (msg.role === "tool") {
if (lastToolBlock) {
var stripped = stripAnsi(msg.content || "").trim();
@@ -1132,27 +1180,76 @@ Pane.prototype.replayHistory = function (messages) {
/^Denied by user/.test(stripped) ||
/^Blocked/.test(stripped);
var isToolError = !!msg.is_error;
// Anchor the rendered output to the specific .ts-approval-tool
// element matching this result's tool_call_id — mirrors the
// live appendToolOutput path so multi-tool batches show
// [hdr A][out A][hdr B][out B] rather than [A][B][out A][out B].
// Falls back to "before badge" when tool_call_id is absent
// (legacy rows pre-dating the wire-format addition).
var resultTarget = null;
if (msg.tool_call_id) {
resultTarget = lastToolBlock.querySelector(
'.ts-approval-tool[data-call-id="' +
CSS.escape(msg.tool_call_id) +
'"]',
);
}
// Cursor-style append: cursor advances after each insert so
// the next sibling lands AFTER the previous one. Fixes the
// bug where calling resultTarget.after(node) twice put the
// second node BETWEEN resultTarget and the first (the second
// .after call was always relative to the same anchor).
// Resulting order with all three present:
// [tool div][output][truncation pill][output-warning]
var insertCursor = resultTarget;
var insertChained = function (node) {
if (insertCursor) {
insertCursor.after(node);
insertCursor = node;
} else {
var bdg = lastToolBlock.querySelector(".ts-approval-badge");
if (bdg) lastToolBlock.insertBefore(node, bdg);
else lastToolBlock.appendChild(node);
}
};
if (stripped && !isDenied) {
var media = !isToolError ? tryParseMedia(stripped) : null;
if (media) {
var embed = buildMediaEmbed(media, stripped);
var bdg = lastToolBlock.querySelector(".ts-approval-badge");
if (bdg) lastToolBlock.insertBefore(embed, bdg);
else lastToolBlock.appendChild(embed);
insertChained(buildMediaEmbed(media, stripped));
} else {
var out = renderToolOutput(stripped, isToolError);
if (out.textContent.split("\n").length > 10) {
makeCollapsible(out);
}
var bdg = lastToolBlock.querySelector(".ts-approval-badge");
if (bdg) lastToolBlock.insertBefore(out, bdg);
else lastToolBlock.appendChild(out);
insertChained(out);
}
// Truncation pill — server marks this when the stored row
// hit the 2000-char cap. Live tool_result events carry full
// output so they don't need the indicator.
if (msg.truncated) {
var pill = document.createElement("span");
pill.className = "tool-output-truncated";
pill.textContent = "… truncated in storage";
pill.title =
"The full tool output was sent to the model live; only the first 10000 characters are persisted to the conversation row.";
insertChained(pill);
}
}
if (isToolError && !lastToolBlock.classList.contains("denied")) {
lastToolBlock.classList.add("error");
appendToolErrorBadge(lastToolBlock);
}
// Output-guard warning — pull the assessment out of the
// function-local pendingAssessments map (populated in the
// assistant branch). Skip when the tool result was denied —
// the ✗ denied badge already signals the deny path.
if (!isDenied && msg.tool_call_id) {
var pending = pendingAssessments[msg.tool_call_id];
if (pending) {
insertChained(_buildOutputWarningEl(pending.assessment));
delete pendingAssessments[msg.tool_call_id];
}
}
}
// Tool-channel metacog reminders (tool_error / repeat) attach
// to the LAST tool message in a batch; on replay we render the
@@ -1165,10 +1262,74 @@ Pane.prototype.replayHistory = function (messages) {
}
}
}
// Flush any output_assessments left in the map — these correspond
// to assistant tool_calls whose tool result row didn't carry a
// tool_call_id (legacy / migrated rows pre-dating the wire-format
// addition). Render the warning under the tool div itself rather
// than dropping the safety information silently.
var leftoverIds = Object.keys(pendingAssessments);
for (var p = 0; p < leftoverIds.length; p++) {
var leftover = pendingAssessments[leftoverIds[p]];
if (!leftover) continue;
leftover.toolDiv.insertAdjacentElement(
"afterend",
_buildOutputWarningEl(leftover.assessment),
);
}
this._attachRetryToLastAssistant();
this.scrollToBottom();
// Focus the input so keyboard users land on the next-action target
// after replay finishes — but only when this is the focused pane,
// there's no pending approval competing for focus, and an input
// element actually exists. Skipping when not the focused pane
// avoids stealing focus from another tab the user is interacting
// with while a background replay completes.
if (
this.id === focusedPaneId &&
!this.pendingApproval &&
this.inputEl &&
!this.busy
) {
try {
this.inputEl.focus({ preventScroll: true });
} catch (_) {
this.inputEl.focus();
}
}
// Restore live-region semantics now that the batch build is done.
this.messagesEl.removeAttribute("aria-busy");
};
// Shared output-warning DOM builder — used by both replayHistory
// (saved-workstream rendering) and the live appendToolOutput path
// via showOutputWarning. Single source of truth keeps the two
// surfaces from drifting on role / class / escape semantics.
function _buildOutputWarningEl(assessment) {
var risk = (assessment && assessment.risk_level) || "medium";
var flags = (assessment && assessment.flags) || [];
var warning = document.createElement("div");
warning.className = "output-warning output-warning-" + risk;
// role="status" (polite) rather than "alert" (assertive) — these
// are findings, not emergencies; the assertive announcement live
// would interrupt the user mid-typing on a high-risk match, which
// is more disruptive than informative.
warning.setAttribute("role", "status");
var labelEl = document.createElement("span");
labelEl.className = "output-warning-label";
labelEl.textContent = "⚠ " + String(risk).toUpperCase();
warning.appendChild(labelEl);
if (flags.length) {
warning.appendChild(document.createTextNode(" " + flags.join(", ")));
}
if (assessment && assessment.redacted) {
var redacted = document.createElement("span");
redacted.className = "output-warning-redacted";
redacted.textContent = " (credentials redacted)";
warning.appendChild(redacted);
}
return warning;
}
Pane.prototype._attachRetryToLastAssistant = function () {
// Remove any previous retry buttons
var old = this.messagesEl.querySelectorAll(".msg.assistant .msg-actions");
@@ -1176,6 +1337,21 @@ Pane.prototype._attachRetryToLastAssistant = function () {
// Find the last assistant message with content and add retry.
// Reasoning blocks emit as .msg.reasoning (distinct modifier) so the
// .msg.assistant selector already excludes them — no extra guard needed.
//
// Skip retry attachment when the most recent semantic turn is
// tool-only — last DOM child is a .ts-approval block. Walk back
// past .user-reminder bubbles (added via addToolReminder /
// addUserReminder AFTER the .ts-approval block they advise) so the
// guard fires correctly even when the tool turn carried a metacog
// reminder. Without this skip, retry lands on a stale prior
// assistant content bubble belonging to an earlier turn.
var lastChild = this.messagesEl.lastElementChild;
while (lastChild && lastChild.classList.contains("user-reminder")) {
lastChild = lastChild.previousElementSibling;
}
if (lastChild && lastChild.classList.contains("ts-approval")) {
return;
}
var assistants = this.messagesEl.querySelectorAll(".msg.assistant");
if (assistants.length) {
this._addRetryAction(assistants[assistants.length - 1]);
@@ -1512,20 +1688,14 @@ Pane.prototype.showOutputWarning = function (evt) {
'.ts-approval-tool[data-call-id="' + escapedId + '"]',
);
if (!toolDiv) return;
var risk = evt.risk_level || "medium";
var flags = evt.flags || [];
var warning = document.createElement("div");
warning.className = "output-warning output-warning-" + risk;
warning.setAttribute("role", "alert");
warning.innerHTML =
'<span class="output-warning-label">\u26a0 ' +
escapeHtml(risk.toUpperCase()) +
"</span> " +
flags.map(escapeHtml).join(", ");
if (evt.redacted) {
warning.innerHTML +=
' <span class="output-warning-redacted">(credentials redacted)</span>';
}
// Shared DOM-builder with replayHistory \u2014 single source of truth for
// role / class / escape semantics. Argument shape mirrors the
// server-side output_assessment dict (risk_level / flags / redacted).
var warning = _buildOutputWarningEl({
risk_level: evt.risk_level,
flags: evt.flags,
redacted: evt.redacted,
});
var nextEl = toolDiv.nextElementSibling;
if (nextEl && nextEl.classList.contains("tool-output")) {
nextEl.insertAdjacentElement("afterend", warning);
@@ -3668,12 +3838,35 @@ function updateDashFooter(agg) {
}
}
var _wsDeleteMode = false;
var _wsDeleteSelected = {};
// Saved Workstreams cache + multi-select delete controller. The
// controller (from /shared/cards.js) owns mode state, checkbox
// decoration, the toolbar wiring, and the confirmation modal — see
// createSavedCardsController for the shared bits.
var _wsSavedItems = [];
var _wsDeleteController = createSavedCardsController({
idPrefix: "ws-delete",
buttonId: "ws-delete-btn",
noun: "workstream",
activateLabel: function (s) {
return "Resume: " + (s.alias || s.title || s.ws_id);
},
buildDeleteRequest: function (wsId) {
return {
url: "/v1/api/workstreams/" + encodeURIComponent(wsId) + "/delete",
options: { method: "POST" },
};
},
render: function () {
renderSavedWorkstreams(_wsSavedItems);
},
onClose: function () {
loadDashboard();
},
});
function renderSavedWorkstreams(items) {
_wsSavedItems = items;
_wsDeleteController.setItems(items);
var c = document.getElementById("dashboard-saved-cards");
c.replaceChildren();
if (!items.length) {
@@ -3684,289 +3877,39 @@ function renderSavedWorkstreams(items) {
return;
}
items.forEach(function (sess) {
// Default card shape (title + meta + wsid + Resume click) comes from
// the shared /shared/cards.js helper so console (Saved Coordinators)
// and ui/static (Saved Workstreams) stay in lock-step. Delete mode
// is interactive-only; we layer the checkbox + selection wiring on
// top of the shared card after construction.
var card = renderSessionCard(sess, {
ariaLabel: function (s) {
var label = s.alias || s.title || s.ws_id;
return _wsDeleteMode ? "Select: " + label : "Resume: " + label;
},
ariaLabel: _wsDeleteController.ariaLabel,
onActivate: function (s) {
// Suppressed in delete mode \u2014 the layered checkbox handler below
// owns clicks while delete-mode is active.
if (_wsDeleteMode) return;
if (_wsDeleteController.blockActivate()) return;
dashboardResumeSession(s.ws_id);
},
});
if (_wsDeleteMode) {
card.classList.add("ws-delete-mode");
card.removeAttribute("role"); // becomes a checkbox host, not a button
var chk = document.createElement("input");
chk.type = "checkbox";
chk.className = "ws-card-check";
chk.checked = !!_wsDeleteSelected[sess.ws_id];
var label = sess.alias || sess.title || sess.ws_id;
chk.setAttribute("aria-label", "Select " + label + " for deletion");
chk.onclick = function (e) {
e.stopPropagation();
if (chk.checked) _wsDeleteSelected[sess.ws_id] = true;
else delete _wsDeleteSelected[sess.ws_id];
card.classList.toggle("ws-selected", chk.checked);
updateWsDeleteBar();
};
card.insertBefore(chk, card.firstChild);
// Override the shared helper's onclick/onkeydown \u2014 in delete mode
// a card click toggles the checkbox instead of activating Resume.
card.onclick = function (e) {
if (e.target === chk) return;
chk.checked = !chk.checked;
chk.onclick(e);
};
card.onkeydown = function (e) {
if (e.key === "Enter" || e.key === " ") {
e.preventDefault();
chk.checked = !chk.checked;
chk.onclick(e);
}
};
if (_wsDeleteSelected[sess.ws_id]) card.classList.add("ws-selected");
}
_wsDeleteController.decorateCard(card, sess);
c.appendChild(card);
});
if (_wsDeleteController.inMode()) _wsDeleteController.refreshBar();
}
// HTML inline-onclick wrappers — keep the global names the existing
// markup binds to (`onclick="startWsDeleteMode()"` etc.) and forward
// to the controller.
function startWsDeleteMode() {
_wsDeleteMode = true;
_wsDeleteSelected = {};
renderSavedWorkstreams(_wsSavedItems);
var btn = document.getElementById("ws-delete-btn");
if (btn) {
btn.textContent = "\u2715 Cancel";
btn.onclick = cancelWsDeleteMode;
}
var bar = document.getElementById("ws-delete-bar");
if (bar) bar.classList.add("visible");
_wsDeleteController.start();
}
function cancelWsDeleteMode() {
_wsDeleteMode = false;
_wsDeleteSelected = {};
renderSavedWorkstreams(_wsSavedItems);
var btn = document.getElementById("ws-delete-btn");
if (btn) {
btn.innerHTML = "&#x1f5d1; Delete";
btn.onclick = startWsDeleteMode;
}
var bar = document.getElementById("ws-delete-bar");
if (bar) bar.classList.remove("visible");
_wsDeleteController.cancel();
}
function updateWsDeleteBar() {
var count = Object.keys(_wsDeleteSelected).length;
var label = document.getElementById("ws-delete-bar-count");
if (label) label.textContent = count + " selected";
var delBtn = document.getElementById("ws-delete-bar-delete");
if (delBtn) delBtn.disabled = count === 0;
var selBtn = document.getElementById("ws-delete-bar-select-all");
if (selBtn) {
var allSelected =
count === _wsSavedItems.length && _wsSavedItems.length > 0;
selBtn.textContent = allSelected ? "Deselect All" : "Select All";
}
}
function toggleSelectAll() {
var allSelected =
Object.keys(_wsDeleteSelected).length === _wsSavedItems.length &&
_wsSavedItems.length > 0;
if (allSelected) {
_wsDeleteSelected = {};
} else {
_wsSavedItems.forEach(function (s) {
_wsDeleteSelected[s.ws_id] = true;
});
}
renderSavedWorkstreams(_wsSavedItems);
updateWsDeleteBar();
_wsDeleteController.toggleAll();
}
var _wsDeleteBatchTrap = null;
function confirmWsDeleteSelection() {
var selected = Object.keys(_wsDeleteSelected);
if (!selected.length) {
showToast("No workstreams selected", "warning");
return;
}
var overlay = document.getElementById("ws-delete-overlay");
var countEl = document.getElementById("ws-delete-count");
var listEl = document.getElementById("ws-delete-list");
var errorEl = document.getElementById("ws-delete-error");
errorEl.textContent = "";
countEl.textContent =
selected.length + " workstream(s) will be permanently deleted:";
listEl.innerHTML = "";
selected.forEach(function (wsId) {
var item = _wsSavedItems.find(function (s) {
return s.ws_id === wsId;
});
var name = item ? item.alias || item.title || wsId : wsId;
var div = document.createElement("div");
div.className = "ws-delete-item";
div.textContent = name;
listEl.appendChild(div);
});
// Reset confirm button handler (may have been overwritten to "Close" by previous run)
var delBtn = document.getElementById("ws-delete-confirm-btn");
if (delBtn) {
delBtn.textContent = "Delete";
delBtn.disabled = false;
delBtn.classList.remove("ws-delete-close");
delBtn.onclick = confirmWsDelete;
}
var cancelBtn = document.getElementById("ws-delete-cancel-btn");
if (cancelBtn) cancelBtn.disabled = false;
overlay.style.display = "flex";
// Focus trap + Escape
if (_wsDeleteBatchTrap)
document.removeEventListener("keydown", _wsDeleteBatchTrap);
_wsDeleteBatchTrap = function (e) {
if (e.key === "Escape") {
e.preventDefault();
cancelWsDelete();
return;
}
if (e.key === "Tab") {
var box = document.getElementById("ws-delete-box");
var focusable = box.querySelectorAll("button:not(:disabled)");
var first = focusable[0];
var last = focusable[focusable.length - 1];
if (e.shiftKey && document.activeElement === first) {
e.preventDefault();
last.focus();
} else if (!e.shiftKey && document.activeElement === last) {
e.preventDefault();
first.focus();
}
}
};
document.addEventListener("keydown", _wsDeleteBatchTrap);
if (cancelBtn) cancelBtn.focus();
_wsDeleteController.confirmSelection();
}
function cancelWsDelete() {
document.getElementById("ws-delete-overlay").style.display = "none";
if (_wsDeleteBatchTrap) {
document.removeEventListener("keydown", _wsDeleteBatchTrap);
_wsDeleteBatchTrap = null;
}
_wsDeleteController.closeModal();
}
function confirmWsDelete() {
var selected = Object.keys(_wsDeleteSelected);
if (!selected.length) return;
var overlay = document.getElementById("ws-delete-overlay");
var errorEl = document.getElementById("ws-delete-error");
var listEl = document.getElementById("ws-delete-list");
var countEl = document.getElementById("ws-delete-count");
var delBtn = document.getElementById("ws-delete-confirm-btn");
var cancelBtn = document.getElementById("ws-delete-cancel-btn");
errorEl.textContent = "";
// Disable buttons during deletion
if (delBtn) {
delBtn.disabled = true;
delBtn.textContent = "Deleting...";
}
if (cancelBtn) cancelBtn.disabled = true;
var results = [];
var promises = selected.map(function (wsId) {
var shortId = wsId.substring(0, 8);
var item = _wsSavedItems.find(function (s) {
return s.ws_id === wsId;
});
var name = item ? item.alias || item.title || wsId : wsId;
var url = "/v1/api/workstreams/" + encodeURIComponent(wsId) + "/delete";
return authFetch(url, { method: "POST" })
.then(function (r) {
var status = r.status;
var contentType = r.headers.get("content-type") || "";
if (r.ok) {
results.push({ name: name, shortId: shortId, ok: true });
return;
}
// Read body as text first to avoid JSON parse errors
return r.text().then(function (body) {
var errMsg = shortId + ": HTTP " + status;
if (contentType.includes("json")) {
try {
var j = JSON.parse(body);
if (j.error) errMsg = shortId + ": " + j.error;
} catch (_) {
/* fall through */
}
} else if (body) {
errMsg = shortId + ": " + body.substring(0, 200);
}
results.push({
name: name,
shortId: shortId,
ok: false,
error: errMsg,
});
});
})
.catch(function (err) {
results.push({
name: name,
shortId: shortId,
ok: false,
error: shortId + ": " + err.message,
});
});
});
Promise.all(promises).then(function () {
// Rebuild the list with results
listEl.innerHTML = "";
results.forEach(function (r) {
var div = document.createElement("div");
div.className = "ws-delete-item" + (r.ok ? "" : " ws-delete-error");
div.textContent =
(r.ok ? "\u2713 " : "\u2717 ") +
r.name +
(r.error ? " — " + r.error : "");
listEl.appendChild(div);
});
var okCount = results.filter(function (r) {
return r.ok;
}).length;
var failCount = results.filter(function (r) {
return !r.ok;
}).length;
countEl.textContent = okCount + " deleted, " + failCount + " failed";
if (delBtn) {
delBtn.disabled = false;
delBtn.textContent = "Close";
delBtn.classList.add("ws-delete-close");
delBtn.onclick = function () {
cancelWsDelete();
cancelWsDeleteMode();
loadDashboard();
};
}
if (cancelBtn) cancelBtn.disabled = false;
});
_wsDeleteController.confirm();
}
// --- Workstream title management ---
+5 -3
View File
@@ -363,17 +363,18 @@
<!-- Delete workstreams confirmation modal (batch) -->
<div
id="ws-delete-overlay"
class="ws-delete-modal-overlay"
style="display: none"
role="dialog"
aria-modal="true"
aria-labelledby="ws-delete-title"
>
<div id="ws-delete-box">
<div id="ws-delete-box" class="ws-delete-modal-box">
<h3 id="ws-delete-title">Delete Workstreams</h3>
<div id="ws-delete-error" role="alert" aria-live="assertive"></div>
<p id="ws-delete-count"></p>
<div id="ws-delete-list"></div>
<div id="ws-delete-buttons">
<div id="ws-delete-list" class="ws-delete-modal-list"></div>
<div id="ws-delete-buttons" class="ws-delete-modal-buttons">
<button
id="ws-delete-cancel-btn"
type="button"
@@ -383,6 +384,7 @@
</button>
<button
id="ws-delete-confirm-btn"
class="ws-delete-confirm"
type="button"
onclick="confirmWsDelete()"
>
+82 -250
View File
@@ -1631,6 +1631,77 @@ body {
.ts-approval-tool .tool-diff .diff-warn {
color: var(--yellow);
}
/* memory/recall calls are background metadata the audit trail is
valuable but they crowd the narrative when a workstream contains
dozens of them. Dim by default; full opacity on hover/focus so
they remain inspectable without permanently competing for
attention. General-sibling combinator (~) extends the fade past
any verdict-badge or output-warning sitting between the tool row
and its output, so the whole sub-tree fades together rather than
leaving a full-opacity badge stranded next to a dim row. */
.ts-approval-tool[data-func-name="memory"],
.ts-approval-tool[data-func-name="recall"] {
opacity: 0.55;
transition: opacity 120ms ease-out;
}
.ts-approval-tool[data-func-name="memory"]:hover,
.ts-approval-tool[data-func-name="memory"]:focus-within,
.ts-approval-tool[data-func-name="recall"]:hover,
.ts-approval-tool[data-func-name="recall"]:focus-within {
opacity: 1;
}
.ts-approval-tool[data-func-name="memory"] ~ .tool-output,
.ts-approval-tool[data-func-name="recall"] ~ .tool-output,
.ts-approval-tool[data-func-name="memory"] ~ .media-embed,
.ts-approval-tool[data-func-name="recall"] ~ .media-embed,
.ts-approval-tool[data-func-name="memory"] ~ .output-warning,
.ts-approval-tool[data-func-name="recall"] ~ .output-warning,
.ts-approval-tool[data-func-name="memory"] ~ .tool-output-truncated,
.ts-approval-tool[data-func-name="recall"] ~ .tool-output-truncated {
opacity: 0.55;
transition: opacity 120ms ease-out;
}
/* Reveal on hover OR focus-within across the entire dimmed
subtree. Without :focus-within on the siblings, a keyboard user
tabbing into a link or collapsible toggle inside .tool-output
sees the content remain dimmed a11y regression. Cover the
warning + truncation pills too so they fully reveal alongside
the result they decorate. */
.ts-approval-tool[data-func-name="memory"] ~ .tool-output:hover,
.ts-approval-tool[data-func-name="recall"] ~ .tool-output:hover,
.ts-approval-tool[data-func-name="memory"] ~ .tool-output:focus-within,
.ts-approval-tool[data-func-name="recall"] ~ .tool-output:focus-within,
.ts-approval-tool[data-func-name="memory"] ~ .media-embed:hover,
.ts-approval-tool[data-func-name="recall"] ~ .media-embed:hover,
.ts-approval-tool[data-func-name="memory"] ~ .media-embed:focus-within,
.ts-approval-tool[data-func-name="recall"] ~ .media-embed:focus-within,
.ts-approval-tool[data-func-name="memory"] ~ .output-warning:hover,
.ts-approval-tool[data-func-name="recall"] ~ .output-warning:hover,
.ts-approval-tool[data-func-name="memory"] ~ .output-warning:focus-within,
.ts-approval-tool[data-func-name="recall"] ~ .output-warning:focus-within,
.ts-approval-tool[data-func-name="memory"] ~ .tool-output-truncated:hover,
.ts-approval-tool[data-func-name="recall"] ~ .tool-output-truncated:hover,
.ts-approval-tool[data-func-name="memory"] ~ .tool-output-truncated:focus-within,
.ts-approval-tool[data-func-name="recall"] ~ .tool-output-truncated:focus-within {
opacity: 1;
}
/* Truncation indicator the persisted tool result is clamped at
2000 chars per row in storage; surface that on replay so users
know they're seeing a clipped view rather than the full output
the live session saw. Aligns with .output-warning's left gutter
(margin-left: 16px) and uses transparent background + dim border
so it reads as quiet metadata rather than a foreign element. */
.tool-output-truncated {
display: inline-block;
margin-top: 4px;
margin-left: 16px;
padding: 1px 6px;
font-size: 10px;
color: var(--fg-dim);
background: transparent;
border-radius: 3px;
border: 1px solid var(--fg-dim);
}
/* .ts-approval (chat.css) stacks its children with flex gap, so a
border-top on the body would float above a strip of container
background instead of sitting flush against the previous tool row.
@@ -2366,224 +2437,10 @@ audio.media-player {
color: var(--accent);
margin: 0;
}
.ws-delete-btn {
background: transparent;
border: 1px solid var(--border);
border-radius: var(--radius);
color: var(--fg-dim);
font-size: 12px;
padding: 4px 10px;
cursor: pointer;
transition:
color 0.15s,
border-color 0.15s;
}
.ws-delete-btn:hover {
color: var(--red);
border-color: var(--red);
}
/* .dashboard-cards / .dashboard-card / .card-title / .card-meta moved to
/shared/cards.css so console (Saved Coordinators) and ui/static (Saved
Workstreams) share one source of truth for the basic card primitive.
Delete-mode rules stay below they're ui/static-only until coordinator
gets the same UX. */
/* Delete mode */
.dashboard-card.ws-delete-mode {
cursor: pointer;
}
.dashboard-card.ws-delete-mode:hover {
border-color: var(--red);
background: rgba(248, 113, 113, 0.04);
}
.dashboard-card.ws-delete-mode.ws-selected {
cursor: default;
}
.dashboard-card.ws-delete-mode.ws-selected:hover {
border-color: var(--red);
background: rgba(248, 113, 113, 0.08);
}
[data-theme="light"] .dashboard-card.ws-delete-mode:hover {
background: rgba(220, 38, 38, 0.04);
}
[data-theme="light"] .dashboard-card.ws-delete-mode.ws-selected:hover {
background: rgba(220, 38, 38, 0.08);
}
.ws-card-check {
position: absolute;
top: 8px;
right: 8px;
width: 18px;
height: 18px;
accent-color: var(--red);
cursor: pointer;
z-index: 1;
opacity: 0;
animation: ws-check-fadein 0.2s ease-out forwards;
}
@keyframes ws-check-fadein {
to {
opacity: 1;
}
}
.dashboard-card.ws-selected {
border-color: var(--red);
background: rgba(248, 113, 113, 0.08);
}
[data-theme="light"] .dashboard-card.ws-selected {
background: rgba(220, 38, 38, 0.08);
}
.ws-delete-bar {
display: none;
align-items: center;
gap: 12px;
margin-top: 12px;
padding: 8px 12px;
background: var(--bg-surface);
border: 1px solid var(--border);
border-radius: var(--radius);
}
.ws-delete-bar.visible {
display: flex;
animation: ws-bar-slide 0.2s ease-out;
}
@keyframes ws-bar-slide {
from {
opacity: 0;
transform: translateY(-8px);
}
to {
opacity: 1;
transform: translateY(0);
}
}
@media (prefers-reduced-motion: reduce) {
.ws-card-check {
animation: none;
opacity: 1;
}
.ws-delete-bar.visible {
animation: none;
}
}
.ws-delete-bar .ws-delete-count-label {
font-size: 12px;
color: var(--fg-dim);
}
.ws-delete-bar .ws-delete-bar-btn {
margin-left: auto;
background: var(--red);
color: #fff;
border: none;
border-radius: var(--radius);
padding: 6px 16px;
font-size: 12px;
font-weight: 600;
cursor: pointer;
transition: filter 0.15s;
}
.ws-delete-bar .ws-delete-bar-btn:hover:not(:disabled) {
filter: brightness(1.1);
}
.ws-delete-bar .ws-delete-bar-btn:disabled {
opacity: 0.4;
cursor: not-allowed;
}
.ws-delete-bar .ws-delete-cancel-btn {
background: transparent;
color: var(--fg-dim);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 6px 12px;
font-size: 12px;
cursor: pointer;
}
.ws-delete-bar .ws-delete-cancel-btn:hover {
color: var(--fg-bright);
border-color: var(--border-strong);
background: var(--bg-highlight);
}
.ws-delete-bar .ws-delete-selectall-btn {
background: transparent;
color: var(--fg-bright);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 6px 12px;
font-size: 12px;
font-weight: 500;
cursor: pointer;
}
.ws-delete-bar .ws-delete-selectall-btn:hover {
border-color: var(--border-strong);
background: var(--bg-highlight);
}
/* Delete modal */
#ws-delete-overlay {
position: fixed;
inset: 0;
background: rgba(0, 0, 0, 0.6);
display: flex;
align-items: center;
justify-content: center;
z-index: 10000;
}
#ws-delete-box {
background: var(--bg-surface);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 24px;
max-width: 480px;
width: 90%;
}
#ws-delete-box h3 {
margin: 0 0 12px;
font-size: 16px;
color: var(--fg-bright);
}
#ws-delete-list {
max-height: 200px;
overflow-y: auto;
margin: 12px 0;
}
#ws-delete-list .ws-delete-item {
padding: 6px 0;
font-size: 13px;
color: var(--fg-bright);
border-bottom: 1px solid var(--border);
}
#ws-delete-list .ws-delete-item:last-child {
border-bottom: none;
}
#ws-delete-list .ws-delete-item.ws-delete-error {
color: var(--red);
}
#ws-delete-buttons {
display: flex;
gap: 8px;
justify-content: flex-end;
margin-top: 16px;
}
#ws-delete-buttons button {
padding: 8px 20px;
border-radius: var(--radius);
font-size: 13px;
font-weight: 500;
cursor: pointer;
border: 1px solid var(--border);
background: transparent;
color: var(--fg-bright);
}
#ws-delete-buttons button:last-child {
background: var(--red);
color: #fff;
border-color: var(--red);
}
#ws-delete-buttons button.ws-delete-close {
background: transparent;
color: var(--fg-bright);
border-color: var(--border);
}
/* .dashboard-cards / .dashboard-card / .card-title / .card-meta and the
delete-mode + modal rules are owned by /shared/cards.css so console
(Saved Coordinators) and ui/static (Saved Workstreams) share one source
of truth. */
/* Server dashboard row — clickable */
.dash-row {
@@ -2689,38 +2546,13 @@ audio.media-player {
}
}
/* ==========================================================================
Verdict badges (intent judge)
========================================================================== */
.verdict-badge {
padding: 4px 10px;
font-size: 11px;
font-weight: 600;
display: flex;
align-items: center;
gap: 8px;
margin-top: 2px;
/* No top separator .ts-verdict-badge (chat.css) shrinks to
max-content width, and a 1px border-top would extend only under
the badge text and read as a truncated line. */
}
.verdict-low {
color: var(--green);
border-left: 3px solid var(--green);
}
.verdict-medium {
color: var(--yellow);
border-left: 3px solid var(--yellow);
}
.verdict-high {
color: var(--red);
border-left: 3px solid var(--red);
}
.verdict-critical {
color: var(--red);
border-left: 3px solid var(--red);
background: rgba(255, 80, 80, 0.05);
}
/* Verdict-badge styling lives in the color-mix block further down
in this file (.verdict-badge.verdict-{low,medium,high,critical}).
The earlier flat-palette duplicate that lived here was removed
two competing .verdict-badge rule sets caused subtle cascade drift
(the color-mix block won for backgrounds, the flat one won for the
bare .verdict-low/medium/high/critical class names) which made
tweaks fragile. Single source of truth now. */
.verdict-detail {
padding: 6px 12px;
Generated
+1 -1
View File
@@ -2533,7 +2533,7 @@ wheels = [
[[package]]
name = "turnstone"
version = "1.5.3"
version = "1.5.6"
source = { editable = "." }
dependencies = [
{ name = "alembic" },