Files
turnstone/tests/test_app_js.py
T
Patrick Buckley eaabc79eb3 fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration (#461)
* fix(replay): repair saved-workstream tool result rendering + extend audit-trail decoration

Loading a saved workstream silently dropped tool results and missed
verdict / output-guard / truncation signals on replay. Root cause was
in `Pane.prototype.replayHistory`: an assistant message carrying both
content and tool_calls cleared the `lastToolBlock` anchor before the
following tool-result iteration could attach. The fix reorders content
to render before the tool block (matching live SSE order) and
restructures the tool-result branch to anchor by `data-call-id` so
multi-tool batches render `[hdr A][out A][hdr B][out B]` rather than
bunching outputs at the bottom.

Beyond the bug, replay now reaches near-parity with the live UX:

- Persisted intent verdicts and output_assessments flow through both
  the SSE replay (`_build_history`) and the `/history` REST endpoint
  used by coord. Single shared helper module owns the wire shape.
- Memory/recall calls persist instead of being filtered at storage
  time — full audit trail; UI dims them by default with hover-reveal
  so heavy memory usage doesn't crowd the narrative.
- Truncation indicator surfaces as a sibling pill (consistent across
  interactive + coord) when a tool result hit the 2000-char cap.
- `replayHistory` wraps DOM work in `aria-busy` so screen readers
  don't get a chatty announce-flood on long replays.
- `_build_history`'s storage I/O moves off the event loop via a new
  `events_replay_prepare` async hook for the SSE path; other async
  callers wrap in `asyncio.to_thread`.

Coord parity:

- `/history` REST endpoint decorates tool_calls with verdict +
  output_assessment + truncation flag (was previously raw
  `load_messages` output).
- Coord JS stamps `judge_verdict` / `heuristic_verdict` from
  history-loaded `tc.verdict` so the existing batch render paints
  the persisted pill, seeds the verdict cache to dedupe later live
  SSE events, and emits an inline `.coord-tool-row-warning` chip
  per call instead of a generic chat line.
- Memory/recall dim rule mirrored on `.coord-tool-row[data-tool-name=...]`.

* fix(replay): address PR #461 review feedback + raise tool-result storage cap

Copilot review feedback:

- Sibling-chain dim rule (memory/recall) now adds :focus-within
  alongside :hover for .tool-output / .media-embed / .output-warning
  / .tool-output-truncated — keyboard users tabbing into a faded
  subtree now get full opacity.
- ``cfg.open_post_load`` is now invoked via ``await asyncio.to_thread``
  so its sync ``_build_history`` call (storage I/O for verdict
  indexes + message reconstruction) doesn't block the event loop on
  every workstream open. Mirrors the SSE replay path that's already
  protected via ``events_replay_prepare``.
- Replaced the hardcoded ``2000`` literal in server.py and session.py
  with ``TOOL_RESULT_STORAGE_CAP`` from the shared decoration module
  so the UI truncation-pill detection can't silently desync from the
  storage write side.

While here:

- Raised ``TOOL_RESULT_STORAGE_CAP`` from 2000 → 10000. A 2000-char
  clip routinely cut grep / file-read bodies mid-line, leaving the
  audit trail useless for retrospective debugging. FTS5 + row size
  grow proportionally; the per-tool upper bound is still bounded
  upstream by ``_truncate_output``'s context-budget clamp.
- Updated the user-visible truncation-pill tooltip on both
  interactive and coord to reflect the new cap.
- ``test_decorates_tool_calls_and_marks_truncated`` now references
  the constant instead of a literal so it stays correct on future
  cap changes.
2026-05-01 14:05:07 -07:00

163 lines
7.9 KiB
Python

"""Static smoke guards for ``turnstone/ui/static/app.js``.
The interactive WebUI's app.js has no JS test framework on the
project side. This file holds Python-side string-presence assertions
that catch regressions on critical paths — the kind of one-line
deletion or rename that breaks the UI silently and only surfaces in
manual testing.
"""
from __future__ import annotations
import re
from pathlib import Path
_APP_JS = Path(__file__).resolve().parent.parent / "turnstone/ui/static/app.js"
def test_switch_tab_bootstraps_pane_when_none_exists() -> None:
"""``switchTab`` must create a pane when none exists. A fresh-
loaded interactive UI with no workstreams shows the dashboard
and creates no panes (per ``initWorkstreams``); the user's first
``create`` or ``open`` then calls ``switchTab(newWsId)``. Pre-fix,
the early ``if (!pane) return;`` left switchTab with nowhere to
attach — the chat UI never connected SSE for the freshly-created
workstream, and only a page refresh fixed it. This test guards
against accidentally re-introducing the early-return."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("function switchTab(wsId) {")
# Bound the search to the function body — switchTab is short.
fn = body[start : start + 2000]
assert "if (!pane) return;" not in fn, (
"switchTab must not early-return when no pane exists — that's "
"the no-chat-after-first-create bug. Bootstrap a pane instead."
)
# Affirmatively check the bootstrap path exists.
assert "createPane(wsId)" in fn, (
"switchTab must call createPane(wsId) to bootstrap the first "
"pane when getFocusedPane returns null"
)
def test_tool_error_does_not_overwrite_approval_badge() -> None:
"""When an approved tool subsequently errors, the existing
``✓ approved`` (or ``✓ auto-approved``) pill must remain visible —
the error indicator is appended as a sibling pill, not by mutating
the approval pill in place. Pre-fix, both ``appendToolOutput``
(live) and ``replayHistory`` (history reconstruction) located the
existing approval badge via ``querySelector(".ts-approval-badge")``
and overwrote its className + textContent with the ``--error``
state, so the user lost the record that they had approved the
call. This test pins the new append-sibling behaviour."""
body = _APP_JS.read_text(encoding="utf-8")
# Affirmatively check that an idempotency guard exists somewhere:
# a ``querySelector(".ts-approval-badge--error")`` lookup is the
# structural marker of the fix. Pre-fix the modifier never appeared
# in app.js at all. Loose on quote style and surrounding form (the
# guard might be a negated ``if (!q) {build...}`` block at a call
# site, or a positive ``if (q) return;`` early-exit inside an
# extracted helper) so a later refactor doesn't trip CI on
# cosmetics.
error_guard_re = re.compile(
r"""querySelector\(\s*['"]\.ts-approval-badge--error['"]\s*\)""",
)
assert error_guard_re.search(body), (
"The error-badge code path must guard creation with a "
"querySelector for .ts-approval-badge--error so duplicate fires "
"(live + history re-render) do not stack badges."
)
# Forbid the mutate-existing-badge sequence: a generic
# ``.ts-approval-badge`` lookup followed within a handful of lines
# by mutating that same handle into the ``--error`` state. Two
# unrelated call sites (history rendering + live tool-output
# insertion) legitimately query ``.ts-approval-badge`` to position
# output above it, so the bare query alone is not the anti-pattern;
# the close pairing with an ``--error`` class mutation is. Accept
# either quote style and catch both ``className = "..."`` and
# ``classList.add("ts-approval-badge--error")`` forms.
overwrite_re = re.compile(
r"""(\w+)\s*=\s*\w+\.querySelector\(\s*(["'])\.ts-approval-badge\2\s*\)\s*;"""
r""".{0,200}?"""
r"""(?:"""
r"""\1\.className\s*=\s*(["'])[^"']*\bts-approval-badge--error\b[^"']*\3"""
r"""|"""
r"""\1\.classList\.add\([^)]*(["'])ts-approval-badge--error\4[^)]*\)"""
r""")""",
re.DOTALL,
)
assert not overwrite_re.search(body), (
"Found the badge-overwrite anti-pattern: a queried "
".ts-approval-badge handle is mutated into the --error variant "
"(via className overwrite or classList.add). Append a sibling "
"badge instead so the approval verdict stays visible alongside "
"the error."
)
def test_replay_history_renders_content_before_tool_block() -> None:
"""In ``replayHistory``'s ``role === "assistant"`` branch, the
``msg.content`` render must precede the ``msg.tool_calls`` render.
Two reasons, both load-bearing:
1. **Structural** — the next loop iteration's ``role === "tool"``
message anchors to ``lastToolBlock``. The tool-block branch sets
that anchor; the content branch clears it. If content runs after
the tool block, the clear silently drops the upcoming tool
result. Pre-fix, every interactive tool result was missing from
saved-workstream replays whenever the assistant turn carried
both narration and tool calls (very common output shape).
2. **Visual** — the live SSE path renders content first
(``stream_text`` streams before ``tool_info`` /
``approve_request``), so replay should match.
The test pins the order via the offsets of the ``msg.content`` and
``msg.tool_calls`` branch headers inside the function body."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.replayHistory = function")
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
fn = body[start:end]
# Locate the assistant branch and bound the search to its body —
# the function also handles user / tool roles which would otherwise
# confuse the offset comparison.
asst_start = fn.index('msg.role === "assistant"')
asst_end = fn.index('msg.role === "tool"', asst_start)
asst = fn[asst_start:asst_end]
content_idx = asst.index("if (msg.content)")
tool_calls_idx = asst.index("if (msg.tool_calls && msg.tool_calls.length)")
assert content_idx < tool_calls_idx, (
"replayHistory must render msg.content BEFORE msg.tool_calls "
"inside the assistant branch — otherwise the lastToolBlock "
"anchor is clobbered before the next iteration's tool result "
"can attach to it (and the visual order also drifts from the "
"live SSE flow)."
)
def test_replay_history_renders_persisted_verdict_badge() -> None:
"""Saved-workstream replays must paint the persisted intent verdict
next to each tool div, using the same ``renderVerdictBadge`` helper
the live ``showInlineToolBlock`` path uses. Pre-fix the audit trail
was complete in storage (``intent_verdicts`` table) but never
surfaced on replay — operators reviewing a saved workstream
couldn't see what the heuristic / LLM judge thought of any tool
call. This test pins the call site so a refactor that drops the
decoration regresses the audit surface."""
body = _APP_JS.read_text(encoding="utf-8")
start = body.index("Pane.prototype.replayHistory = function")
end = body.index("Pane.prototype._attachRetryToLastAssistant", start)
fn = body[start:end]
# Match a `renderVerdictBadge(<something>.verdict, ...)` call inside
# the replay loop. Loose on whitespace + identifier so a future
# rename of the iteration variable doesn't trip CI.
badge_call_re = re.compile(
r"renderVerdictBadge\(\s*\w+\.verdict\b",
)
assert badge_call_re.search(fn), (
"replayHistory must call renderVerdictBadge(tc.verdict, ...) "
"when a persisted verdict is attached to a tool_call entry — "
"otherwise the audit-trail data persisted to intent_verdicts "
"doesn't surface on saved-workstream replays."
)