mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-24 21:04:48 -06:00
ef13f40cf5e2f0664ded1a5497ee6bb17fbdb921
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
38a0d9c3b6 |
feat(coord): Stage 3 SessionManager Children primitive lift + cluster bus push paths
Lift the Children primitive out of CoordinatorAdapter into universal SessionManager core primitives, replace the fragile poll + state-event piggyback paths with first-class cluster bus event types for inline approval delivery, and clean up the resulting frontend reducer. Architecture - New `turnstone/core/children_registry.py` — universal parent → children + reverse-lookup primitive with atomic `add_child` (returns parent UI for race-free dispatch). Lifted from `CoordinatorAdapter`. - New `turnstone/core/child_source.py` — `ChildSource` Protocol with `SameNodeChildSource` (in-process via SessionManager state observer) and `ClusterChildSource` (cross-node via ClusterCollector listener). - `SessionManager._on_state_change` upgraded to multi-subscriber (`subscribe_to_state` / `unsubscribe_from_state`) under a dedicated lock; CLI consumer migrated. - `CoordinatorAdapter` shrunk: 731 → ~640 LOC. Children data lives in the registry; fan-out lives in ClusterChildSource. Backward-compat property facades dropped; tests updated to use the registry surface. Cluster bus event vocabulary - New event types `intent_verdict`, `approval_resolved`, `approve_request` flow through both `ClusterCollector._apply_delta` (translation from node SSE) and `emit_console_ws_*` (synthesis on console pseudo-node). - `CoordinatorAdapter._dispatch_child_event` re-emits as `child_ws_intent_verdict` / `child_ws_approval_resolved` / `child_ws_approve_request` on the parent coord's SSE stream. - New `_broadcast_intent_verdict` / `_broadcast_approval_resolved` / `_broadcast_approve_request` no-op hooks on `SessionUIBase`. WebUI pushes to the global queue; ConsoleCoordinatorUI pushes to the collector. `approve_tools` calls `_broadcast_approve_request` right after setting `_pending_approval` so the items reach the coord tree immediately, eliminating the bulk-fetch race. Cleanups - `pending_approval_detail` piggyback on `ws_state` / `cluster_state` removed end-to-end. Bulk fetch + explicit verdict / approve-request push are the canonical carriers. - Browser `_judgePollTick` 90-second poll loop deleted; push path is authoritative. - `urgent` flag on `scheduleLiveFetch` deleted (only caller was 409 retry; replaced with `invalidateLiveBadge` + standard schedule). - Console `_fetch_live_block` derives `pending_approval` from a disjunction (`activity_state="approval"` OR `state="attention"` OR detail present) so the bulk fetch can't return false during the state-transition race window. - Coord-side merge guard in `flushLiveFetches` no longer clobbered: `handleChildState` only stamps `sseUpdatedAt` when authoritatively clearing detail. - `child_locality` capability flag removed (was inert dead code). Reliability - Selective drop on listener queue overflow: critical event types (verdicts, approvals, ws_closed, child_ws_*) evict one oldest item to make room rather than dropping themselves on a full queue. Best-effort events (state ticks, content tokens, status, activity) drop as before. Applied to `SessionUIBase._enqueue`, `ClusterCollector._fanout`, and the `WebUI._global_queue` puts in the new broadcast hooks. - `_state_subscribers` snapshot under a dedicated lock so concurrent subscribe / unsubscribe during dispatch can't shift the iterator. UX / a11y - Loading placeholder in renderChildRow keeps row height stable while the bulk fetch is in-flight (sr-friendly aria-label). - Focus preservation across `_renderChildrenNow` (capture + restore by row + marker) and across targeted `_updateChildRow` swaps. - Layout-shift transition on the approval block max-height; respects `prefers-reduced-motion`. - Sidebar pending count: `(N children · M pending)`. - Risk pill `aria-label` spells out level + confidence for SR users. - Per-coord SSE listener queue depth surfaced in the status bar (`queue N/500`) with color escalation (warn at >50%, danger at >80%). Tests - 305+ test changes across 8 files. New unit tests for `ChildrenRegistry`, `ChildSource` (both impls + multi-subscriber observer), the new collector emit + apply_delta cases, the dispatch cases for new event types, the broadcast hook overrides on both WebUI and ConsoleCoordinatorUI, and the focus / placeholder / pending-count frontend assertions in `test_coordinator_page.py`. 5024 passed, ruff + mypy clean. |
||
|
|
88facd260e |
feat(coord): pass pending_approval_detail on child_ws_state SSE events
Inline child approve/deny in the coord tree UI was rendering downstream
of the bulk-live cache (``GET /v1/api/cluster/ws/live``), not the SSE
stream. ``child_ws_state`` events were tiny notifications that fired
an urgent live-bulk fetch on every activity_state transition into/out
of "approval", just to pick up the rich ``pending_approval_detail``
payload. With multiple coord tabs and multi-child workstreams, that
urgent-fetch pattern compounded the SSE-executor pressure Shape A
is unwinding.
Thread the field through every layer so the SSE event itself carries
the rich payload — browser mutates ``liveBadgeCache`` directly,
no urgent fetch:
1. Node ``WebUI._broadcast_state`` emits ``pending_approval_detail``
on ``ws_state`` events. Gated on ``_pending_approval is not None``
so the per-broadcast verdict-cache deepcopy only runs when there
is actually an approval pending. ``_build_node_snapshot`` also
projects the field so the console's reconnect-via-snapshot
resync path delivers it (without this the new collector
forwarding would never see the field on a snapshot row).
2. Console ``ClusterCollector._apply_delta`` (live ``ws_state``
forwarding) and ``_reconcile_node`` (snapshot resync diff) both
forward the field on the emitted ``cluster_state`` event, AND
``_apply_delta`` persists it on the cached ``ws`` dict so the
``get_node_detail`` / ``get_snapshot`` endpoints between
reconciliations don't render stale approve/deny buttons.
3. ``CoordinatorAdapter._dispatch_child_event`` re-emits the field
on the ``child_ws_state`` event sent to coord listener queues.
4. Frontend ``handleChildState`` reads ``ev.pending_approval_detail``
and writes it directly into ``liveBadgeCache``, tagging the
entry with ``sseUpdatedAt``. ``flushLiveFetches`` honors that
tag for ``SSE_AUTHORITATIVE_MS`` (3s) — the upstream
``/dashboard`` cache has its own ~2s TTL, so a bulk-poll
landing right after a transition can otherwise clobber the
fresh SSE-set state with pre-transition data.
The pre-fix ``enteredApproval`` / ``leftApproval`` urgent-fetch
branch is removed. The 409 stale-call_id retry path keeps its own
urgent fetch — that's a different scenario.
Tests cover the forwarding contract at every layer, the broadcast
gate (event includes the field when an approval is pending,
omits it otherwise, and clears after resolution), and the
``flushLiveFetches`` merge-guard structural shape so a refactor
that keeps the symbols but inverts the comparison or drops the
``prev.live`` check can't pass silently.
|
||
|
|
3abd2c441b |
feat(console): coord rich ws_state payload + live activity broadcast (#420)
* feat(console): coord rich ws_state payload + live activity broadcast (Stage 2 follow-up)
Pre-lift coord's cluster broadcast was state-only — the dashboard's
coord rows showed the state column flipping but ``tokens`` /
``context_ratio`` / ``activity`` / ``content`` were all hardcoded
to zero / empty. The lift makes coord populate the same per-ws
metric fields interactive does and broadcasts them through the
cluster collector with the rich kwargs.
**Architecture changes:**
- Lift ``on_status`` / ``on_content_token`` / ``on_thinking_start`` /
``on_thinking_stop`` / ``on_stream_end`` / ``on_tool_result`` /
``on_reasoning_token`` / ``on_tool_output_chunk`` / ``on_info`` /
``on_error`` from ``WebUI`` to :class:`SessionUIBase` as base
implementations. Coord inherits the bodies; the per-ws metric
fields it had at the base but never populated now flow.
- ``WebUI`` keeps overrides for ``on_status`` / ``on_tool_result`` /
``on_error`` to layer Prometheus ``_metrics.record_*`` calls
on top of ``super()`` (node-only — the console isn't a node).
``WebUI._broadcast_state`` now uses the new
:meth:`SessionUIBase.snapshot_and_consume_state_payload` helper
for the rich-payload snapshot read.
- ``ConsoleCoordinatorUI`` adds a ``_broadcast_activity`` override
that calls the new
:meth:`ClusterCollector.update_console_ws_activity` (in-memory
pseudo-node row update; named ``update_*`` rather than ``emit_*``
to flag the no-fanout asymmetry vs. the rest of the
``emit_console_ws_*`` family).
- ``coord_adapter.emit_state`` reads ``ws.ui``'s snapshot under
``_ws_lock`` and passes the rich kwargs to the extended
:meth:`ClusterCollector.emit_console_ws_state`. Defensive when
``ws.ui is None`` mid-eviction (broadcasts state-only).
- ``coord_endpoint_config`` wires a new ``_coord_spawn_metrics``
hook so per-spawn ``_ws_messages`` / ``_ws_turn_tool_calls``
bookkeeping fires on coord too.
- ``_MAX_TURN_CONTENT_CHARS`` moved from ``turnstone.server`` to
``turnstone.core.session_ui_base`` so coord enforces the same
per-turn content cap.
**Three observable behaviour changes** (CHANGELOG-callout-worthy):
- Coord persists ``usage_event`` storage rows on every status
emission (governance dashboards / token-spend queries gain
coord visibility).
- Coord broadcasts live activity transitions to the cluster
collector (dashboard's coord rows show activity ticks between
state changes the same way interactive does), with last-emitted
dedup so a tool-heavy turn's repeated ``activity=""`` clears
don't hammer the collector lock.
- Cluster ``cluster_state`` events for coord rows now carry
non-zero ``tokens`` / ``content``. Frontend rendering that
conditionally hid these on coord can drop the branch.
**Tests:** 23 new tests in ``tests/test_coord_rich_ws_state_payload.py``
(per-ws metric writes, snapshot helper drain semantics +
single-lock-acquisition, adapter rich-payload pass-through +
None-UI defensive handling, activity broadcast wire + dedup +
failure swallow + no-op-when-collector-unset, spawn_metrics
hook, concurrent-writes-during-snapshot stress with reader
cycling through running/idle/error so drain branches actually
run, on_stream_end activity-clear pin). Plus WebUI override
regression tests confirming ``_metrics.record_*`` still fires
on top of the lifted bodies. Existing
``tests/test_webui_content.py`` updated to import
``_MAX_TURN_CONTENT_CHARS`` from its new home;
``tests/test_coordinator_adapter.py`` updated to expect the
rich-payload kwargs (default zeros) on
``emit_console_ws_state``. Total: ``4491 → 4514``.
``ruff check`` clean, ``mypy`` clean on touched files.
**/review pipeline** (4 finders → verify → dedupe) caught 14
findings → 12 unique (3 collapsed as duplicates of the lockless
``on_content_token`` writer):
- bug-1 Minor: ``on_status`` regressed coord's defensive
``usage.get(...)`` indexing → restored ``.get(..., 0)`` for
``prompt_tokens`` / ``completion_tokens`` on both base + WebUI
override.
- bug-2 Nit: concurrent-snapshot reader only used ``"running"`` →
cycled through ``("running", "idle", "error")`` so drain
branches run; also captures + re-raises thread exceptions
instead of silently passing.
- bug-3 + sec-2 + perf-3 Nit (merged): ``on_content_token``
mutated ``_ws_turn_content`` lockless while the snapshot drained
under lock → wrapped the cap-check + append + size-update in
``_ws_lock``.
- perf-2 Minor: collector lock contention from per-event activity
broadcasts → cached last-emitted ``(activity, activity_state)``
on the UI; subsequent identical ticks return early without
acquiring the collector lock.
- perf-4 Nit: join-under-lock in snapshot helper → swap-then-join
pattern (capture list reference under lock, reassign to empty,
join the captured list outside the lock). Halves the lock
hold and decouples the join walk from concurrent appenders.
- q-1 Minor: ``emit_console_ws_activity`` was misleading (no
``_fanout`` call, unlike the rest of the ``emit_console_ws_*``
family) → renamed to ``update_console_ws_activity`` + docstring
call-out for the asymmetry.
- q-2 + q-3 Minor/Nit: stale docstrings on
``coordinator_ui.py`` (still claimed "no per-node metrics —
Phase D") and ``_interactive_spawn_metrics`` (still claimed
"counters live on WebUI only") → both updated to reflect the
lifted base class + coord's new hook.
- q-4 Nit: broken Sphinx cross-ref
``:meth:\`_snapshot_and_consume_state_payload\``` → dropped
the leading underscore.
- q-5 Nit: missing ``test_coord_on_stream_end_clears_activity``
→ added.
**Two findings explicitly deferred** (out-of-scope follow-ups,
documented in CHANGELOG):
- perf-1: synchronous ``record_usage_event`` INSERT on coord
worker thread per status tick. Parity with WebUI is the lift's
goal; if throughput becomes a concern, batch usage_event writes
on a background flusher (would apply to both kinds).
- sec-1: coord assistant content now flows on the cluster SSE
stream, which has no per-user filter today. Pre-existing
exposure for interactive ``cluster_state`` events; the lift
extends to coord rows. Proper fix needs SSE auth gating
(``admin.cluster.inspect``) or per-listener user_id filtering
— separate security project, doesn't gate this lift.
* fix(console): apply review feedback on PR #420
Three review findings, all confirmed against source:
1. **Copilot — dedup-state-vs-failure race in `_broadcast_activity`**
(correctness bug): pre-fix ``self._last_broadcast_activity = current``
was assigned inside the ``_ws_lock`` block BEFORE the collector call.
If the collector raised mid-broadcast, the exception was swallowed
but the dedup state was already updated, so subsequent identical
activity ticks would be deduped and never retried — leaving the
dashboard's coord row stranded at the pre-failure activity until
the activity actually changed.
Fix: move the dedup-state update OUT of the lock and place it AFTER
a successful collector call. On failure, ``_last_broadcast_activity``
stays unchanged so the next identical tick retries. Two new
regression tests pin both the failure-recovery (``test_coord_ui_
broadcast_activity_failure_does_not_strand_dedup``) and the
happy-path dedup behavior (``test_coord_ui_broadcast_activity_
dedup_skips_identical_after_success``).
2. **Copilot — stale `emit_console_ws_activity` reference in
CHANGELOG**: the method was renamed to ``update_console_ws_activity``
per /review's q-1 finding before the original commit landed, but the
CHANGELOG entry was written ahead of the rename. Updated to match
the actual API + added the no-fanout asymmetry rationale inline so
readers don't have to chase the method name.
3. **code-quality bot ×2 — `except BaseException` in test workers**:
the concurrent-snapshot stress test caught thread-worker exceptions
with ``except BaseException`` (with a noqa to suppress BLE001).
``BaseException`` is overkill for a thread worker — ``SystemExit``
/ ``KeyboardInterrupt`` are main-thread signals and ``Exception``
is the right scope. Narrowed to ``except Exception`` on both
workers; ``writer_exc`` / ``reader_exc`` types narrowed from
``list[BaseException]`` to ``list[Exception]``.
Tests: ``4514 → 4516`` (+2 regression tests for the dedup race fix).
``ruff check`` clean, ``mypy`` clean. No code-path changes outside
the dedup-state placement; the rich-payload broadcast surface is
unchanged.
|
||
|
|
27349e1c13 |
refactor: move bridge content buffer to server-side single source of truth
Eliminate dual accumulation by piggybacking assistant response text on the server's ws_state:idle SSE event. The bridge no longer maintains its own _ws_content_buffer — it reads content directly from the idle event and passes it through to TurnCompleteEvent unchanged. Server-side: WebUI accumulates tokens in on_content_token(), joins and includes in the idle broadcast, then resets (with 256 KB cap). Downstream consumers (Discord bidi DM forwarding, catch-up) are unaffected — TurnCompleteEvent.content is still populated. |