mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-24 21:04:48 -06:00
e60c19befd5e31376bb606cd380c3564ac4e27df
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3c7a3c1375 |
fix(webui): wedge-proof the live-session pipeline and de-O(N) hot paths
Long sessions (5000+ messages, several compactions) degraded steadily and could stop rendering entirely while the backend stayed healthy. Four hard failure mechanisms, each sufficient on its own: - Unguarded event pipeline: one throw escaping onmessage/handleEvent (e.g. renderMarkdown stack overflow on a few KB of nested "> ") stranded the streaming refs, so every later delta painted into the poisoned segment. stream_end now resets segment refs BEFORE the finalize render with a plain-text fallback (the coordinator pane's existing pattern); onmessage guards both parse and dispatch; renderMarkdown is depth-capped with throw-safe footnote-scope accounting; the streaming buffer is marked rendered only on success. - Rebuild-vs-live races: clear_ui/replay_truncated re-renders wiped events painted in the snapshot->replaceChildren window (never redelivered) and left deltas writing into detached nodes. Rebuilds now quiesce the event stream behind a token-owned queue flushed after the render; streaming refs reset on every rebuild path including refetch FAILURE; a mid-stream replay_truncated defers its re-sync to the idle edge instead of dropping the repair. - Ignored recovery floor: the global stream now handles node_snapshot and replay_truncated. Roster eviction (with a "Session ended" toast for open panes) happens only from the stream-ordered snapshot; the REST resync is merge-only and r.ok-gated so a mid-restart 503 body cannot read as an authoritative empty roster. - Unbounded growth: _agentCards released on rebuild — deliberately NOT on transport-only reconnects, which must preserve the maps or the next child event builds a duplicate card; orphan grace timers cancelled on full reload/destroy; toast queue capped with duplicate coalescing; diff previews capped at 400 rendered lines (the spread-append could throw RangeError before the approval gate painted) with the omission notice below the scroll box; raw results clamped at 64KiB. Per-event O(N) work removed from the hot paths: thinking-indicator instance ref; near-bottom cached from a passive scroll listener and re-checked at rAF pin time (a user scroll-up landing in the coalescing window wins; ResizeObserver re-engages follow after layout changes); rAF-coalesced outer and per-stream scroll pins; self-healing call_id->row/stream lookup caches; verdict lookup scoped to the row's batch; tracked retry holder; queue-controller Set replaces the whole-transcript idle sweep; rail renders rAF-coalesced; coordinator child_ws_state ticks routed to single-row updates (full render only on terminal-boundary crossings) with observer unobserve on replace. Also: the coordinator SSE-error 401 probe is un-deadened (raw fetch — authFetch never resolves a 401 — with the body inspected so a version_mismatch still takes auth.js's upgrade-reload path via the new noteVersionMismatch export); the console cluster-SSE reconnect timer is tracked across logout; the mermaid render chain is rejection-proof per link and paints errors on the containers the failing link had already claimed. Measured with scripts/livepass.py --perf (n=3000 history + 20-turn live storm): full replay 1060ms -> 238ms; re-render cycles 836-1071ms -> ~94ms flat; chunk path now flat vs transcript size; worst longtask 1080ms -> ~500ms; agent-card retention across rebuilds 4 -> 0. Known limit (needs a server-side event watermark on /history): a turn completing inside the refetch window can paint twice after the quiesce flush — rare, visible, and strictly better than the silent loss it replaces. |
||
|
|
221e804ef3 |
test(ui): de-modulize the renderer harness sources
test_renderer_js drives renderer.js behaviorally through node via vm.runInThisContext — script semantics, which choke on the import/ export syntax renderer.js and utils.js now carry (all 68 tests failed at harness setup). The harness now evaluates _demodulize()d source: imports drop (the shared vm context resolves cross-file bindings as globals, exactly like the pre-module classic scripts) and export keywords peel off. Deliberately NOT switched to dynamic import(): the mermaid harness pokes renderer-internal state (_mermaidState = 'ready') that script evaluation exposes but a real module would encapsulate. Module semantics are covered by test_shell_js's .mjs parse sweep; these tests pin renderer behavior. |
||
|
|
ad16de001d |
fix(renderer): drop single-$ inline math to stop false positives in prose
Single-$ inline math is too ambiguous in conversational text: currency
amounts ("$5 and $10 each"), shell variables ("$HOME and $PATH"), and
shell prompts all produced false-positive KaTeX spans because the
regex matched any non-$/non-newline span between two dollar signs.
Inline math now requires the unambiguous \(...\) form, which is what
GPT-5 / o-series / Claude with reasoning effort emit by default anyway.
Display math ($$...$$ and \[...\]) is unchanged — the doubled
delimiter has enough mass that ambiguity is not a practical problem.
Three former positive tests are inverted into regression guards so a
future regex change can't quietly resurrect the bug, and new tests
name the currency and env-var cases explicitly. The web env prompt is
updated to advertise \(...\) and to tell the model why $...$ is gone.
|
||
|
|
821108310f |
fix(ui): drop local escapeHtml in renderer (Copilot review on #553)
The local-escape posture from the prior commit double-encoded values that inlineMarkdown had already escaped: leading escapeHtml(text) turns `&` into `&`, the local escapeHtml(url) then turned that into `&amp;`, which breaks query-string URLs after browser parse + getAttribute + new URL round-trip. Switch to convention-rename: regex callback params renamed to safeAlt / safeUrl / safeLabel to signal the upstream-escape invariant. The attribute-context lint enforces all future attribute-context concat sites maintain the safe* convention or call escapeHtml explicitly — defense-in-depth preserved without the regression. Two added pin tests verify `&` survives with single (not double) entity encoding through image data-src and link href. Also addresses two test issues from the same review: - Docstring listed `safe[A-Z_]…` but code only checked isupper(). Drop the underscore option (JS uses camelCase anyway). - `_all_attr_names` only recorded attribute-bearing tags, so a bare `<script>` injection would have false-negatived the link- label pin test. Refactored to `_parse_renderer_html` returning both start tags and (tag, attr) pairs. |
||
|
|
849364c49e |
fix(ui): renderer local-escape + attribute-context CI lint (#553)
inlineMarkdown's image and link renderers now escapeHtml each interpolated value (url, alt, label, domain) at the call site instead of relying on the upstream escape pass. Defence-in-depth: a future refactor calling those renderers from outside inlineMarkdown would otherwise silently regress. New CI lint scans renderer.js for `attr="' + ident` patterns; ident must be escapeHtml(...), safe*, or in the reviewer-approved allowlist. Four pin tests use html.parser.HTMLParser to verify attacker URLs and labels don't materialize event-handler attributes on rendered DOM. |
||
|
|
f8f076cf20 |
fix(renderer): mermaid streaming parser errors + progressive hljs (#510)
* fix(renderer): mermaid streaming parser errors + progressive hljs
Live streaming was rendering mermaid diagrams with `Parse error,
got 'PS'` messages — bare `(`, `[`, `{` inside unquoted edge / node
labels re-entered Mermaid's shape parser. Two unrelated streaming-
specific issues in the renderer pile-up here; this commit addresses
both plus a follow-on UX improvement for code highlighting.
## Mermaid label autoquoter
`_normalizeMermaidSource` wraps two label forms that Mermaid rejects
when they contain bare shape-delimiter chars:
1. Edge labels: `|content|` → `|"content"|`
2. Rectangle node labels: `ID[content]` → `ID["content"]`
Shapes whose syntax already nests delimiters — cylinders `[(...)`,
subroutines `[[...]]`, trapezoids `[/.../]` `[\...\]`, circles
`((...))`, hexagons `{{...}}`, diamonds `{...}` — are intentionally
left alone (their inner delimiters are part of the shape syntax;
quoting would corrupt them). Labels already wrapped in `"..."` are
also left alone. The rewrite is idempotent and runs before the
mermaid SVG cache lookup so identical malformed input hits the
cache on re-render rather than re-quoting per tick.
## Markdown fence-pair regex
The old fence regex `/(```+)([^\s`]*)\n([\s\S]*?)\1/g` would, mid-
stream, pair an unclosed ```mermaid open with the OPENING backticks
of a later ```python fence as the "close", handing mermaid a
truncated source. New regex:
/(```+)([^\s`]*)\n((?:(?!\1)[\s\S])*?)\1[ \t]*(?=\n|$)/g
Two constraints close the gap:
- `(?!\1)` inside the content quantifier blocks the lazy matcher
from extending across another N-backtick run. Smaller inner
counts (e.g. 3-backtick inner inside a 4-backtick outer) still
pass since `\1` is the open's actual count.
- `[ \t]*(?=\n|$)` after `\1` forces the close to a line
boundary, so ```python (open with a language tag) can't
masquerade as a previous fence's close.
Together: an unclosed fence stays as plain markdown until its true
close arrives, so neither mermaid nor hljs ever sees a mid-stream
truncated source.
## Progressive hljs
Extracted `postRenderHljs` from `postRenderMarkdown` with a source-
keyed `_hljsCache` (FIFO, cap 64, keyed on `language:source`) and
wired it into `_streamingRenderApply`. Closed code fences are now
syntax-highlighted as they stream in, matching the progressive
mermaid pattern from #426. Per-tick cost stays cheap because the
cache returns the pre-tokenized HTML synchronously on hit; only
unique (language, source) pairs pay `hljs.highlightElement`.
## Internal cleanup from the review pipeline
- `_cacheFifoEntry(cache, key, value, max)` replaces the duplicated
`_cacheHljsEntry` and `_cacheMermaidEntry`. Single tested
implementation across four caches (hljs, mermaid svg, mermaid
error, mermaid normalize memo). The "don't evict on overwrite"
invariant is pinned per-cache in tests.
- `_mermaidNormalizeCache` memoizes raw textContent → normalized
output so the per-rAF-tick autoquoter split + regex doesn't
repeat for unchanged diagrams. Eviction shares
`_MERMAID_CACHE_MAX` with the SVG cache it feeds.
## Tests
The fake DOM in tests/test_renderer_js.py grew a few capabilities
to drive these paths:
- `classList` is now array-like (length + indexed access) so the
hljs language-extraction loop works.
- `textContent` setter mirrors the real-DOM side effect of
entity-escaping into innerHTML, so `escapeHtml()` round-trips
(otherwise every `renderMarkdown` returns empty `<p>` tags).
- `querySelectorAll` handles both `pre code.language-mermaid`
and `pre code[class*='language-']`.
Added: 6 fence-pairing regression cases, 9 hljs-progressive cases
(cache hit / distinct sources / language separation / NO_HIGHLIGHT
langs / terminal class / eviction / overwrite / postRenderMarkdown
wraps hljs / _streamingRenderApply invokes hljs), 11 autoquoter
cases including both diagram sources from the live screenshot
encoded verbatim as parametrized regressions, and 3 normalize-memo
cases (populates on first call, consulted before normalize via
sentinel pre-seed, distinct sources cache separately).
Total: 104 renderer tests pass (was 67).
* fix(renderer): apply Copilot review feedback on #510
Two doc / harness adjustments from the PR review — no behavior
change in production code.
- The `_mermaidNormalizeCache` comment claimed eviction "stays in
lockstep with the SVG cache". That was misleading: the two
caches key on different things (raw textContent vs normalized
source) and evict independently. Updated the comment to describe
what they actually share (the cap, for memory footprint) and
what they don't (positional coupling), and to note that the memo
deliberately survives `_initMermaid` since normalize output is
theme-independent.
- The fake DOM in tests/test_renderer_js.py had `innerHTML` setter
clear `children` but leave `_textContent` intact, so subsequent
`textContent` reads could return stale data after an innerHTML
mutation (real DOM invalidates textContent on innerHTML write).
No current test triggered this, but it would mask future bugs
that depend on innerHTML/textContent consistency. Setter now
clears `_textContent`; the children-derived fallback in the
getter returns `''` after the wholesale replace.
All 104 renderer tests still pass; ruff + mypy clean.
|
||
|
|
438e6f41ba |
feat(renderer): progressive mermaid rendering during streaming (#426)
* feat(renderer): progressive mermaid rendering during streaming
Mermaid diagrams used to materialize all-at-once at stream_end via
streamingRenderFinalize, which felt laggy on long responses with
multiple diagrams. Now closed mermaid fences render progressively
as each fence completes during streaming.
The blocker was streamingRender's wholesale `el.innerHTML = html`
on every rAF tick, which destroys any rendered SVG nodes — without
caching, calling postRenderMermaid per tick would re-trigger an
async mermaid.render every time, thrashing the renderer.
Added a source-keyed SVG cache (_mermaidSvgCache, FIFO-bounded at
64 entries):
- Cache hit on identical source: synchronous innerHTML swap, no
loading flash, no async work. Mermaid is deterministic for a
given init, so identical source ⇒ identical SVG, safe to reuse.
- Cache miss: queue async render, populate cache on success.
- Errored sources cached separately (_mermaidErrorCache) so a
syntactically-broken diagram doesn't re-thrash mermaid on every
tick. The user can fix the diagram and the new source string
misses the cache, triggering a fresh render.
_streamingRenderApply now calls postRenderMermaid after the
innerHTML replace. Per-stream cost: each unique mermaid source
pays mermaid.render once, then synchronous cache hits for every
subsequent rAF tick. hljs syntax highlighting stays deferred to
streamingRenderFinalize (it's a separate pass and benefits less
from progressive rendering — code blocks tend to be short and
already legible without color).
Tests: built a richer Node-driven harness with a fake DOM that
tracks attributes / classList / parent chain / replaceWith, plus
a stubbed mermaid.render with a call counter. 5 new tests cover:
cache-hit skips render, distinct sources render independently,
errors cache to avoid thrash, FIFO eviction at cap, and a static
guard that _streamingRenderApply actually calls postRenderMermaid.
* fix(renderer): apply Copilot feedback on PR #426
Six review items, all real:
1. _cacheMermaidEntry evicted on overwrite — overwriting an
existing source unnecessarily dropped the oldest entry.
Now: only evict when inserting a new key.
2. _initMermaid didn't clear caches — a theme change via
reRenderAllMermaid (which calls _initMermaid) would serve
stale SVG keyed by source-only, since rendered output
depends on themeVariables. Now clears both caches on
(re-)init.
3. bindFunctions never re-applied on cache hits — mermaid's
bindFunctions attaches link/click handlers to each rendered
SVG instance. Pre-fix, only the first render got bindings;
subsequent cache hits via raw innerHTML left the SVG inert.
Cache value is now {svg, bindFunctions}; cache hits go
through _applyMermaidSvg which re-applies bindings on each
new container instance.
4. Truthiness checks on cache lookups — empty-string SVG / error
would have masqueraded as a miss. Switched to cache.has()
(and then .get) so intent is explicit.
5. Concurrent mermaid.render — postRenderMermaid now fires on
every streaming rAF tick, so multiple ticks could overlap
while earlier render Promises pend. mermaid.render uses
module-level state internally — concurrent calls clobber it.
Two layers of serialization fix this:
- _mermaidPending: per-source. While a render is in flight
for source X, additional containers asking for X are queued
and the single render result fans out to all pending
containers when it lands.
- _mermaidRenderChain: across-source. Promises chain so
mermaid.render runs at most one at a time globally.
- Detached containers (no longer in the DOM by the time the
render completes) are skipped via isConnected guard —
wholesale innerHTML replace during streaming detaches them
and a later tick is already taking care of the live one.
6. Test brittleness — _streamingRenderApply guard used
body.index("\\n}\\n", start) which would stop at the first
inner-block closing brace inside the function. Switched to a
bounded-window string search (Copilot's suggestion).
Three new tests added: overwrite doesn't evict; _initMermaid
clears caches; cache hit re-applies bindFunctions. Existing tests
updated for the new {svg, bindFunctions} cache shape and the
async serialization (drain via setTimeout hops instead of bare
microtask resolves).
Test harness fix: fake DOM elements now have an isConnected
getter derived from the parent chain, so the new guard
exercises correctly under test.
|
||
|
|
33d16d19ce |
fix(renderer): handle LaTeX-style \(...\) and \[...\] math delimiters (#425)
* fix(renderer): handle LaTeX-style \(...\) and \[...\] math delimiters The browser renderer at turnstone/shared_static/renderer.js only recognized TeX-style $...$ / $$...$$ delimiters. Most modern LLMs (GPT-5 / o-series, Claude with reasoning effort) emit LaTeX-style \(...\) for inline math and \[...\] for display by default — those slipped through as raw text in the coord + interactive WebUIs, making KaTeX appear "broken when nested inside a markdown block" (actually broken everywhere, the surrounding markdown just made the failure noticeable). Added a second pass for each delimiter style alongside the existing $...$ / $$...$$ patterns. Both styles now feed the same mathBlocks / inlineMaths placeholder pipeline so all the existing nested-block handling (lists, blockquotes, tables, bold, headings, details, post-render KaTeX markup) Just Works. Edge cases verified by the new test_renderer_js.py harness: - \(...\) inside inline code stays literal - \(...\) inside fenced code blocks stays literal - Solo \[ with no closing \] doesn't trigger spurious math - Markdown links [text](url) untouched (regex uses \[ \], not [ ]) - Mixed TeX + LaTeX delimiters in one message both render The harness drives renderer.js through Node via vm.runInThisContext with stubbed document/katex globals — first JS-side regression guard for the renderer; previously it had no test coverage at all. * fix(renderer): apply Copilot feedback on PR #425 Three review items from Copilot: 1. Display-math sentinel could leak through inline-code spans. The original ordering ran $$...$$ / \[...\] extraction BEFORE inline code, so a backtick span around math (e.g. `$$x$$` or `\[x\]`) had its delimiters consumed by the math regex and replaced with \x00MB…\x00. Inline code then captured the sentinel; restore order put MB after IC, leaving the null-byte placeholder visible inside the rendered <code>. Reorder: inline code first, then display math, then inline math. Code spans now seal their content before any math regex sees it. The reverse edge case (math containing backticks, e.g. \verb|`x`|) is much rarer and KaTeX rejects \verb anyway. 2. Inline LaTeX-style \(...\) regex used [\s\S]+? which allowed newlines, so an unterminated \( on one line would eat the next paragraph until it found a closing \). Aligned with the existing $...$ behavior by switching to [^\n]+? — display math (\[...\] / $$...$$) stays multi-line by design. 3. tests/test_renderer_js.py was guarded with a node-availability skip, but CI's test + test-postgres jobs didn't explicitly install Node, so the suite would have silently no-op'd if the runner image dropped Node. Added actions/setup-node@v5 to both jobs. Four new regression tests cover the leak (both delimiter styles inside backticks must stay literal) and the cross-paragraph span (both \(...\) and $...$ must not eat newlines). |