mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-25 05:14:47 -06:00
eaabc79eb3
* fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration Loading a saved workstream silently dropped tool results and missed verdict / output-guard / truncation signals on replay. Root cause was in `Pane.prototype.replayHistory`: an assistant message carrying both content and tool_calls cleared the `lastToolBlock` anchor before the following tool-result iteration could attach. The fix reorders content to render before the tool block (matching live SSE order) and restructures the tool-result branch to anchor by `data-call-id` so multi-tool batches render `[hdr A][out A][hdr B][out B]` rather than bunching outputs at the bottom. Beyond the bug, replay now reaches near-parity with the live UX: - Persisted intent verdicts and output_assessments flow through both the SSE replay (`_build_history`) and the `/history` REST endpoint used by coord. Single shared helper module owns the wire shape. - Memory/recall calls persist instead of being filtered at storage time — full audit trail; UI dims them by default with hover-reveal so heavy memory usage doesn't crowd the narrative. - Truncation indicator surfaces as a sibling pill (consistent across interactive + coord) when a tool result hit the 2000-char cap. - `replayHistory` wraps DOM work in `aria-busy` so screen readers don't get a chatty announce-flood on long replays. - `_build_history`'s storage I/O moves off the event loop via a new `events_replay_prepare` async hook for the SSE path; other async callers wrap in `asyncio.to_thread`. Coord parity: - `/history` REST endpoint decorates tool_calls with verdict + output_assessment + truncation flag (was previously raw `load_messages` output). - Coord JS stamps `judge_verdict` / `heuristic_verdict` from history-loaded `tc.verdict` so the existing batch render paints the persisted pill, seeds the verdict cache to dedupe later live SSE events, and emits an inline `.coord-tool-row-warning` chip per call instead of a generic chat line. - Memory/recall dim rule mirrored on `.coord-tool-row[data-tool-name=...]`. * fix(replay): address PR #461 review feedback + raise tool-result storage cap Copilot review feedback: - Sibling-chain dim rule (memory/recall) now adds :focus-within alongside :hover for .tool-output / .media-embed / .output-warning / .tool-output-truncated — keyboard users tabbing into a faded subtree now get full opacity. - ``cfg.open_post_load`` is now invoked via ``await asyncio.to_thread`` so its sync ``_build_history`` call (storage I/O for verdict indexes + message reconstruction) doesn't block the event loop on every workstream open. Mirrors the SSE replay path that's already protected via ``events_replay_prepare``. - Replaced the hardcoded ``2000`` literal in server.py and session.py with ``TOOL_RESULT_STORAGE_CAP`` from the shared decoration module so the UI truncation-pill detection can't silently desync from the storage write side. While here: - Raised ``TOOL_RESULT_STORAGE_CAP`` from 2000 → 10000. A 2000-char clip routinely cut grep / file-read bodies mid-line, leaving the audit trail useless for retrospective debugging. FTS5 + row size grow proportionally; the per-tool upper bound is still bounded upstream by ``_truncate_output``'s context-budget clamp. - Updated the user-visible truncation-pill tooltip on both interactive and coord to reflect the new cap. - ``test_decorates_tool_calls_and_marks_truncated`` now references the constant instead of a literal so it stays correct on future cap changes.
163 lines
7.9 KiB
Python
163 lines
7.9 KiB
Python
"""Static smoke guards for ``turnstone/ui/static/app.js``.
|
|
|
|
The interactive WebUI's app.js has no JS test framework on the
|
|
project side. This file holds Python-side string-presence assertions
|
|
that catch regressions on critical paths — the kind of one-line
|
|
deletion or rename that breaks the UI silently and only surfaces in
|
|
manual testing.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import re
|
|
from pathlib import Path
|
|
|
|
_APP_JS = Path(__file__).resolve().parent.parent / "turnstone/ui/static/app.js"
|
|
|
|
|
|
def test_switch_tab_bootstraps_pane_when_none_exists() -> None:
|
|
"""``switchTab`` must create a pane when none exists. A fresh-
|
|
loaded interactive UI with no workstreams shows the dashboard
|
|
and creates no panes (per ``initWorkstreams``); the user's first
|
|
``create`` or ``open`` then calls ``switchTab(newWsId)``. Pre-fix,
|
|
the early ``if (!pane) return;`` left switchTab with nowhere to
|
|
attach — the chat UI never connected SSE for the freshly-created
|
|
workstream, and only a page refresh fixed it. This test guards
|
|
against accidentally re-introducing the early-return."""
|
|
body = _APP_JS.read_text(encoding="utf-8")
|
|
start = body.index("function switchTab(wsId) {")
|
|
# Bound the search to the function body — switchTab is short.
|
|
fn = body[start : start + 2000]
|
|
assert "if (!pane) return;" not in fn, (
|
|
"switchTab must not early-return when no pane exists — that's "
|
|
"the no-chat-after-first-create bug. Bootstrap a pane instead."
|
|
)
|
|
# Affirmatively check the bootstrap path exists.
|
|
assert "createPane(wsId)" in fn, (
|
|
"switchTab must call createPane(wsId) to bootstrap the first "
|
|
"pane when getFocusedPane returns null"
|
|
)
|
|
|
|
|
|
def test_tool_error_does_not_overwrite_approval_badge() -> None:
|
|
"""When an approved tool subsequently errors, the existing
|
|
``✓ approved`` (or ``✓ auto-approved``) pill must remain visible —
|
|
the error indicator is appended as a sibling pill, not by mutating
|
|
the approval pill in place. Pre-fix, both ``appendToolOutput``
|
|
(live) and ``replayHistory`` (history reconstruction) located the
|
|
existing approval badge via ``querySelector(".ts-approval-badge")``
|
|
and overwrote its className + textContent with the ``--error``
|
|
state, so the user lost the record that they had approved the
|
|
call. This test pins the new append-sibling behaviour."""
|
|
body = _APP_JS.read_text(encoding="utf-8")
|
|
# Affirmatively check that an idempotency guard exists somewhere:
|
|
# a ``querySelector(".ts-approval-badge--error")`` lookup is the
|
|
# structural marker of the fix. Pre-fix the modifier never appeared
|
|
# in app.js at all. Loose on quote style and surrounding form (the
|
|
# guard might be a negated ``if (!q) {build...}`` block at a call
|
|
# site, or a positive ``if (q) return;`` early-exit inside an
|
|
# extracted helper) so a later refactor doesn't trip CI on
|
|
# cosmetics.
|
|
error_guard_re = re.compile(
|
|
r"""querySelector\(\s*['"]\.ts-approval-badge--error['"]\s*\)""",
|
|
)
|
|
assert error_guard_re.search(body), (
|
|
"The error-badge code path must guard creation with a "
|
|
"querySelector for .ts-approval-badge--error so duplicate fires "
|
|
"(live + history re-render) do not stack badges."
|
|
)
|
|
# Forbid the mutate-existing-badge sequence: a generic
|
|
# ``.ts-approval-badge`` lookup followed within a handful of lines
|
|
# by mutating that same handle into the ``--error`` state. Two
|
|
# unrelated call sites (history rendering + live tool-output
|
|
# insertion) legitimately query ``.ts-approval-badge`` to position
|
|
# output above it, so the bare query alone is not the anti-pattern;
|
|
# the close pairing with an ``--error`` class mutation is. Accept
|
|
# either quote style and catch both ``className = "..."`` and
|
|
# ``classList.add("ts-approval-badge--error")`` forms.
|
|
overwrite_re = re.compile(
|
|
r"""(\w+)\s*=\s*\w+\.querySelector\(\s*(["'])\.ts-approval-badge\2\s*\)\s*;"""
|
|
r""".{0,200}?"""
|
|
r"""(?:"""
|
|
r"""\1\.className\s*=\s*(["'])[^"']*\bts-approval-badge--error\b[^"']*\3"""
|
|
r"""|"""
|
|
r"""\1\.classList\.add\([^)]*(["'])ts-approval-badge--error\4[^)]*\)"""
|
|
r""")""",
|
|
re.DOTALL,
|
|
)
|
|
assert not overwrite_re.search(body), (
|
|
"Found the badge-overwrite anti-pattern: a queried "
|
|
".ts-approval-badge handle is mutated into the --error variant "
|
|
"(via className overwrite or classList.add). Append a sibling "
|
|
"badge instead so the approval verdict stays visible alongside "
|
|
"the error."
|
|
)
|
|
|
|
|
|
def test_replay_history_renders_content_before_tool_block() -> None:
|
|
"""In ``replayHistory``'s ``role === "assistant"`` branch, the
|
|
``msg.content`` render must precede the ``msg.tool_calls`` render.
|
|
|
|
Two reasons, both load-bearing:
|
|
|
|
1. **Structural** — the next loop iteration's ``role === "tool"``
|
|
message anchors to ``lastToolBlock``. The tool-block branch sets
|
|
that anchor; the content branch clears it. If content runs after
|
|
the tool block, the clear silently drops the upcoming tool
|
|
result. Pre-fix, every interactive tool result was missing from
|
|
saved-workstream replays whenever the assistant turn carried
|
|
both narration and tool calls (very common output shape).
|
|
|
|
2. **Visual** — the live SSE path renders content first
|
|
(``stream_text`` streams before ``tool_info`` /
|
|
``approve_request``), so replay should match.
|
|
|
|
The test pins the order via the offsets of the ``msg.content`` and
|
|
``msg.tool_calls`` branch headers inside the function body."""
|
|
body = _APP_JS.read_text(encoding="utf-8")
|
|
start = body.index("Pane.prototype.replayHistory = function")
|
|
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
|
|
fn = body[start:end]
|
|
# Locate the assistant branch and bound the search to its body —
|
|
# the function also handles user / tool roles which would otherwise
|
|
# confuse the offset comparison.
|
|
asst_start = fn.index('msg.role === "assistant"')
|
|
asst_end = fn.index('msg.role === "tool"', asst_start)
|
|
asst = fn[asst_start:asst_end]
|
|
content_idx = asst.index("if (msg.content)")
|
|
tool_calls_idx = asst.index("if (msg.tool_calls && msg.tool_calls.length)")
|
|
assert content_idx < tool_calls_idx, (
|
|
"replayHistory must render msg.content BEFORE msg.tool_calls "
|
|
"inside the assistant branch — otherwise the lastToolBlock "
|
|
"anchor is clobbered before the next iteration's tool result "
|
|
"can attach to it (and the visual order also drifts from the "
|
|
"live SSE flow)."
|
|
)
|
|
|
|
|
|
def test_replay_history_renders_persisted_verdict_badge() -> None:
|
|
"""Saved-workstream replays must paint the persisted intent verdict
|
|
next to each tool div, using the same ``renderVerdictBadge`` helper
|
|
the live ``showInlineToolBlock`` path uses. Pre-fix the audit trail
|
|
was complete in storage (``intent_verdicts`` table) but never
|
|
surfaced on replay — operators reviewing a saved workstream
|
|
couldn't see what the heuristic / LLM judge thought of any tool
|
|
call. This test pins the call site so a refactor that drops the
|
|
decoration regresses the audit surface."""
|
|
body = _APP_JS.read_text(encoding="utf-8")
|
|
start = body.index("Pane.prototype.replayHistory = function")
|
|
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
|
|
fn = body[start:end]
|
|
# Match a `renderVerdictBadge(<something>.verdict, ...)` call inside
|
|
# the replay loop. Loose on whitespace + identifier so a future
|
|
# rename of the iteration variable doesn't trip CI.
|
|
badge_call_re = re.compile(
|
|
r"renderVerdictBadge\(\s*\w+\.verdict\b",
|
|
)
|
|
assert badge_call_re.search(fn), (
|
|
"replayHistory must call renderVerdictBadge(tc.verdict, ...) "
|
|
"when a persisted verdict is attached to a tool_call entry — "
|
|
"otherwise the audit-trail data persisted to intent_verdicts "
|
|
"doesn't surface on saved-workstream replays."
|
|
)
|