mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
cf7cfe8932a1918e6a6bcce2f7ea6d1d10e71dbf
23 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
32c76499fa |
fix(mcp): harden obo mint path and admin lifecycle after review
Mint engine: guard the credential-rotation persist so a storage blip cannot escape the classified-result contract mid-mint (and cannot brick the user's other obo servers on strict-rotation IdPs); stop borrowing the login flow's httpx client across event loops — mints use a transient per-request client (obo_http_client remains as a test seam); retry OIDC discovery at runtime (cooldown-gated, single-flight) so a node that booted during an IdP outage can mint again without a restart; key the under-lock force-refresh reuse gate on created, which delete+create makes the mint time (obo rows never set last_refreshed, so the copied oauth_user gate never fired and serialized waiters each re-redeemed). Cross-node consent badges: the cleared-pairs set becomes a TTL map with bounded growth, so a badge written by another node after this node's last clear self-heals within one TTL window instead of surviving until a restart. Admin lifecycle: purge the mint cache when oauth_scopes changes on an obo row (an rfc8693 privilege reduction now applies immediately, like audience changes); normalize no-op scope/audience re-sends out of updates — the admin form re-submits pre-filled fields on every save, which both re-triggered purges and made entra-profile rows with legacy scopes un-editable; make flip-into-obo scope handling grant-profile aware (entra clears the carry-over, rfc8693 honors the request); clear obo-era audience/scopes when flipping back to oauth_user (the IdP-side app identifier is not a resource indicator); mirror the same column policy in the create handler. Revocation honesty: hide obo mint-cache rows from the user connections list and refuse the per-server disconnect with 409 — deleting the row returned 204, audited token_revoked, and then session-start priming silently re-minted from the surviving captured credential. Console form: keep the audience-from-URL autofill off for sign-in passthrough (the audience there is an IdP application identifier, and the prefilled URL passed every validation layer then failed every mint); clear the autofill artifact when switching modes; omit unchanged scopes from submissions. Dispatchers: route tool/resource/prompt through one shared lookup-error mapping and an auth-model-aware 401-exhausted detail (obo users are no longer pointed at a consent flow that does not exist). The consent-url audit count drops 13 → 7: the three per-dispatcher mapping copies collapsed into _pool_lookup_error. Priming: skip all obo servers for users with no captured credential via one existence SELECT (previously three reads per server per session). Also: USER_SCOPED_AUTH_TYPES now lives in storage._protocol so the backend SQL predicates share the application layer's set; docs describe the actual purge-on-transition behavior (the orphan-and-reactivate claims were wrong); the entra e2e setup script no longer aborts silently under set -e with suppressed stderr. |
||
|
|
44e9d46e40 |
fix(mcp): address pre-push review — obo scope/audience/priming defects
Frontend↔backend interaction bugs the backend-only rounds couldn't see: - flip oauth_user->oauth_obo: the admin form re-submits the pre-filled oauth_user scopes, so the flip-clear (gated on 'oauth_scopes' not in body) was skipped -> rfc8693 mints broke permanently. Clear now compares to the existing value, robust to the re-send. - entra edit-lockout: update validated the MERGED scopes, so a pre-existing scoped obo row under the entra profile became un-editable (every PUT 400'd). Reject only when the request actually SETS scopes. - flush-cache button never rendered: consented_users_count is now populated for oauth_obo rows too, not just oauth_user. Mint engine + priming: - audience guard: a cached token minted for a since-narrowed audience is no longer served (extracted _is_fresh_obo_cache_row, used pre/post-lock, checks refresh-less + audience-match + fresh). _persist_obo_cache_row now delete+creates so the row's audience column tracks the mint (a plain update kept the stale audience -> re-mint loop). - obo session priming passes revoke_ambiguous_escalation=False (new param threaded through get_obo_...), so an IdP wobble during a bulk prime can't escalate-revoke obo cache rows cluster-wide. Cross-node + lifecycle: - pending-consent success-clear now clears once-per-failure-cycle via a _pending_consent_cleared set (was gated on 'we wrote it' -> never fired cross-node/after-restart -> stale badge). Still no per-call SQL. - identity-unlink cache purge: per-server try/except so one failure doesn't leave other servers' bearers un-purged. - entra ignored-scopes: warn once per audience (was per-mint flood -> downgraded to debug -> no signal on a profile switch). - entra_setup.sh writes single-quoted .env values (secret may contain $). +6 regression tests. 1892 mcp/oidc/console tests green; mypy clean. Refs #551. |
||
|
|
3e88c54751 |
test(mcp): check in oauth_obo e2e harnesses under scripts/obo-e2e
Manual (non-CI) harnesses that exercise the real oauth_obo mint path against a live IdP, kept for future validation of the feature: - entra_e2e.py: real Entra tenant, one interactive sign-in, drives get_obo_access_token_classified -> _obo_mint_entra (E1-E7) - keycloak_e2e.py + .sh: ephemeral Keycloak, fully headless, drives the rfc8693 leg (refresh grant -> token exchange) - entra_spike.py: raw-OAuth wire probe (pre-implementation reference) - entra_setup.sh: creates the Entra spike app registrations - .env.example template; real creds stay in a gitignored .env Both legs pass E1-E7 (mint + aud, cache hit, single-credential->multi- audience, rotation write-back, force_refresh, unconsented->credential survives, flush->re-mint). Not wired into CI. Refs #551. |
||
|
|
41e7d5b7d7 |
feat(livepass): add long-session perf harness (--perf)
New /perf/livepass.html mounts the real InteractivePane at production scroll geometry and drives production-shaped SSE events through handleEvent/replayHistory in real time (no virtual-time budget, no forced reduced-motion — both corrupt the measurement), reporting: replayHistory wall time at N messages, per-turn live-storm cost on top of that transcript, tool_output_chunk throughput, busy/idle churn, heap + node + agent-card counts across repeated replay cycles (the detached-DOM leak probe), and longtask counts. The --perf runner builds, serves, and launches headless Chrome with --js-flags=--expose-gc and --enable-precise-memory-info so heap numbers are real floors; the page POSTs its JSON report to /perf/report. Reports carry a per-attempt run token the runner validates, so a straggler POST from a killed prior attempt cannot be misattributed to the next size, and the wait loop polls the Chrome process so a sandbox startup failure bails to the --no-sandbox fallback in seconds instead of burning the full timeout. |
||
|
|
8dd356b7e6 |
fix(task-agent): keep sub-tool steps nested + preserve denial reasons
Address the Copilot review on #732 plus a task-agent sub-tool nesting race surfaced alongside it. Nesting (web UI): - A sub-tool step whose task_agent row hasn't painted yet (the 4-wide tool pool's ordering window) buffers and nests when the row lands, instead of escaping to a top-level row that looks main-harness-issued. - A row that never paints (id-correlation mismatch / aborted agent) escapes its buffered steps back to a visible top-level paint after a grace window, so steps are never buffered invisibly or leaked. - The nested card survives the parent row's pending->resolved rebuild; a call_id reused across turns builds a fresh card rather than stealing the prior agent's steps. - tool_info routes through the same nesting path (no duplicate top-level row); a namespaced sub-tool result no longer grafts onto an unrelated top-level row. Denial reasons (backend): - Preserve the specific denial reason a gate already stamped (operator feedback, or the matched policy pattern; web and CLI contracts) instead of clobbering it with a flat "Denied by user" -- in both the sub-agent and the main tool loop. Verified with the livepass task_agent harness (race + orphan-escape scenarios, headless) and unit tests. |
||
|
|
77cb76c006 |
feat(task-agent): recall sub-trajectory + per-agent read isolation
Final chunk of the task_agent modernization: rebuild a finished task
agent's card from /history (reload / reopen while the workstream is in
memory) and isolate each sub-agent's file-read tracking.
Recall: _project_agent_steps projects a sub-agent's trajectory into step
items (FIFO-per-call_id pairing via _iter_agent_tool_results, shared with
_cancel_ledger; output/arguments/count capped); _stash_agent_trajectory
keeps them on the UI in an LRU-bounded store; make_history_handler
attaches them as agent_steps to each task_agent tool_call, and
replayHistory/_replayAgentCard rebuild the collapsed card. In-memory only
(durable persistence deferred); a cold/evicted entry renders the flat
parent row ("not retained"), never a fabricated 0-step card.
Read isolation: _read_files (the blind-overwrite guard's memory) is now
per-sub-agent via the _active_read_files contextvar -- _exec_task copies
the parent's set on spawn and merges the agent's reads back on
completion, so a sibling in the 4-wide pool can't suppress another
agent's guard.
Also: _exec_task now self-reports the task_agent tool_result on every
path (the parent loop only reports error/denied results centrally) --
without it the live card never completed and a failed task recorded
is_error=False in the canonical trajectory. is_error flows from
_tool_error_flags to the recalled step; on_info suppression is per-thread
so a parallel sibling tool's progress isn't dropped.
|
||
|
|
ca7958329a |
feat(task-agent): nest sub-tool steps in an expandable card
Route a task agent's sub-tool events (tool_pending / approve_request, tagged with parent_call_id) into a collapsible card under the task_agent row, replacing the blue on_info turn-legs. - conversation.js / interactive.js: buildAgentCardBody + _routeAgentItems / _ensureAgentCard nest steps by parent_call_id. Collapsed by default (a task agent can run 100+ steps and the parent fans out many in parallel); the label carries the live count + state. Auto-expand when a nested approval is pending so the blocking prompt can't hide behind the toggle. - session.py / session_ui_base.py: on_agent_step paints auto-tool step rows; namespace child call_ids by parent so the 4-wide task pool can't collide on local sequential ids (call_0); suppress sub-agent on_info on the web pane (no call_id to nest by — the card carries steps + result). - cli.py: on_agent_step prints a dim step leg (no card on the CLI, which keeps its on_info). - livepass.py: task-agent card harness driving the real InteractivePane. |
||
|
|
9ad447ca33 |
fix(attachments): design-review polish for preview chips/pills
Two-reviewer + sanity pass over the attachment previews: - composer audio chip is icon+name+size only; the native <audio> player renders on the sent message, not the staging chip (too heavy at chip scale) - cap sent-message pills (+ in-pill audio/snippet) so they no longer overflow the bubble at narrow widths; player and snippet drop to their own row - clamp the chip filename in shared chat.css so long names ellipsize instead of wrapping (console main + coordinator previously left it unclamped) - merge the duplicated .composer-chip rule; drop unused kind-modifier classes and inert vertical-align / inline-block declarations - fix undefined var(--bg-base) -> var(--bg-surface) thumbnail backing - label the <audio> control (aria-label) and drop the decorative snippet from the a11y tree scripts/livepass.py: add an attachments harness that drives the real createAttachmentController + Pane.addUserMessage so these surfaces render headlessly for review. |
||
|
|
44c11efb53 |
feat(ui): split-view follow-ups — per-pane ✕, child-opens-beside, close-on-ws_closed
Four refinements from first live use: - Per-pane ✕ chip, top-right of every visible pane. Split mode: hide that cell (closeCell — the tab stays, the sibling absorbs the space). Single-pane: close the pane outright (withheld from the unclosable Dashboard). The click decides at click time; the label tracks the mode. Manager-injected into the pane section — content untouched. - Coordinator child links open BESIDE the coordinator (openPaneBeside: split right of the focused cell, seeded with the child pane) instead of replacing it — the parent stays on screen. Degrades to the plain focused-cell swap when the split is denied (cap / narrow viewport). splitFocused() gained an optional explicit-fill parameter for this. - Tier-1 ws_closed now CLOSES the open interactive pane (tab gone, a split cell collapses) — the coordinator-closes-its-child flow, matching the standalone's pane-auto-close. The dead-banner lane stays for streams that die without a ws_closed (node crash/network), where the session may still be revivable. - Paint bug: the focused-cell ring was an inset box-shadow on the section, which paints in the element's own background layer — UNDER opaque children touching the edges, so the status bar / composer strip occluded it. The ring now rides a click-transparent ::after overlay above pane content; the 2px top bar sits above the ring line. The livepass shell surface's demo panes grew a .ws-status-bar footer so the occlusion bug class stays visible to future passes. |
||
|
|
f8f7152d63 |
feat(ui): split view returns to the L-shell — PaneManager layout tree
Revives the split-pane feature retired with ui/static (step 6), rebuilt on PaneManager: an optional binary layout tree (null = the one-pane-per- tab behaviour, unchanged) renders visible panes as %-inset cells — no reparenting, so live stream DOM, scroll state and media survive layout changes. Tabs stay global: the active tab is the focused cell, a backgrounded tab swaps into it, clicking inside a visible pane focuses its cell, .shown marks visible-unfocused tabs. Separators resize by pointer-capture drag and arrow keys (role=separator + aria-value*); the tree persists in the working-set blob and rehydrate prunes leaves whose pane did not restore. Limits: 6 cells, 200x150 cell minimums, denials toast the manager's reason. Affordance: Split right / Split down / Unsplit buttons in the tab-bar tail replace the redundant [+] (the permanent Dashboard tab is the launcher) — deliberately no contextmenu override this time. The dead TS_APP.focusLauncher seam goes with it. Measured chrome: the focused cell wears a 2px accent top bar (no thin tinted ring clears 3:1 in both themes) plus a 55%-mix inset ring; separators rest at --ink-4 with solid-accent hover/drag/focus; .shown tabs carry an accent underline; the tail cluster is fenced and lifted to --ink-3. scripts/livepass.py grows a third surface: shell/livepass.html boots the real shell.js + pane.js and drives ?split=right|down|three|none (+ &theme=light), stamping SPLIT-READY-<cells> / SPLIT-FAILED-<reason>. |
||
|
|
b3c3acc5d1 |
fix(scripts): livepass dialog-tier riders + loud open-failure + dock-displacement probe
Designer-review round on the scroll fix found the harness's dialog-tier gate silently green: confirm-dialog (and install/coord-delete) markup lives OUTSIDE #admin-layout, so the fragment extraction never embedded it — ?open=confirm threw at showConfirmModal and screenshot a normal, dialog-less page. build() now injects every hatch dialog the fragment does not already contain, and a driven ?open= that ends with no open dialog stamps OPEN-FAILED-<state> into the title instead of passing. Also upstreams the review's probe states: &focuslast=1 focuses the last shelf-body control (the displaced-dock regression class — only .sh-body may scroll; head/foot must stay pinned) and &scrolled=bottom shows the 24px scroll tail. |
||
|
|
6bb47cad2d |
fix(console): manage-pane scroll regressions — interior scroller, clip the hatch-host, anchor hidden inputs
The L-shell height-pins the admin chain and .hatch-host clipped it, so no box below the pane could scroll: tabs taller than the pane were cut dead, and the overflow:hidden host doubled as a hidden scroll container that focus-into-view silently scrolled — visually-hidden toggle/cap/radio inputs escape the .sh-body scroller (abspos under an unpositioned label), overhang the shelf, and a Tab keypress shoved the docked hatch off its head with no scrollbar to recover by. - .admin-content becomes the manage pane's interior scroller (the #main precedent); switchAdminTab resets it on real tab changes only - .hatch-host: overflow hidden -> clip — paint clipping without a scroll container, so focus can never displace the dock - position:relative anchors on the three hidden-input labels (toggle-switch, .sh-body .cap, segmented-option); .settings-toggle already carried one - livepass: the console harness wraps the fragment in the REAL L-shell chain (its bespoke height pin is exactly how this bug class stayed invisible to the screenshot gates) and gains a ?tall=1/&scrolled=1 scroll state |
||
|
|
b991dc2e83 |
fix: CI lint pin + copilot-thread hardening
The wiring-lint test used percent-formatted regex patterns — UP031 under the ruff 0.15.6 the CI pre-commit pins (the older venv binary let it through; checked repo-wide against the exact pin now). f-strings with doubled quantifier braces, plus one over-long fixture line in the livepass generator split. Copilot threads, both validated rather than blindly applied: - closeShelf's scrim-ownership scan now skips detached entries. The thread's throw scenario doesn't occur on the real removal path (a pane close detaches an ANCESTOR, so _hostOf still resolves inside the detached subtree) — but a detached shelf is genuinely not a scrim owner, so the guard is correct beyond being defensive. - toast.js drops the popover attribute via removeAttribute instead of the null assignment. The claim that null leaves popover="null" is refuted — the IDL is nullable and null removes the attribute (verified empirically in headless Chrome) — but removeAttribute reads correct without requiring that spec knowledge. |
||
|
|
9102f858a4 |
chore(scripts): commit the livepass harness generator
The livepass harness — the headless-render rig that verified every converted modal surface and click-drives submits (the dead-Save bug class) — lived as ad hoc files in /tmp and got wiped once already. The durable piece is the GENERATOR: the markup is extracted fresh from the index files at build time (a committed snapshot would drift) and the stylesheets/scripts are symlinked so edits are live on refresh. scripts/livepass.py builds both harnesses into /tmp/livepass/ (ui: all six dialog-tier surfaces incl. the real cards.js batch controller drive; console: the admin-pane fragment hosting the shelves, with schedule/model/policy/confirm/token fixtures and the model-save click drive that flips document.title to PUT-OK-<n>). --serve included; the chrome screenshot incantation and the ?open= registry are in the module docstring. Governance fixtures (roles/HR/OGP/memory/skill) are documented seams for when those surfaces need driving. |
||
|
|
80d201a67b |
chore(ui): L-shell step 6 (5e.2f) — retire dead split-pane CSS from ui/static
Removes the structurally-dead CSS the L-shell superseded — 794 lines: the tab bar (.ws-tab*, #tab-bar, #split-btn, .ws-tab-dropdown-*), the binary split-pane machinery (.split-*, #split-root, .pane-ctx-*), the old approval/verdict card (.ts-approval-*, .verdict-*, the judge spinner), the fixed .dashboard-overlay, and the retired appbar/settings-overlay bits — plus their [data-theme=light] overrides. The standalone now styles its conversation from the shared sheets (chat/conversation/interactive.css); the dashboard table + saved list were always shared (base/cards.css). Method: a conservative token-diff — a rule is dropped only when EVERY selector's class/id token is absent (word-boundary, comments stripped) from the standalone runtime (index.html + every JS it loads, incl. the vendored hljs/katex/mermaid so their runtime-built classes aren't mistaken for dead). Mixed/any-live rules are kept verbatim (no reformatting), so ~50 dead-but-harmless rules that share a generic token like `.active` survive — safe over-keep. The markdown / syntax / math / diagram theme lives ONLY in this style.css (the shared sheets don't carry it), so the hljs/katex/mermaid families are protected from removal. Also fixes four dead tab-DOM pokes in app.js (editWorkstreamTitle / confirmDeleteWorkstream read the title from the workstreams roster now, not the retired .ws-tab .tab-name; the cancel handlers drop the gone .tab-chevron focus restore). Verified: braces balanced (370/370), the headless harness still builds clean (errs:[]), git diff confirms zero live dashboard/render rules removed, and the css_specificity_audit (manifest synced to the standalone's new sheet set) shows the SAME 10 pre-existing findings before/after — zero new cascade flips (removing a rule for a non-existent selector can't change any live element's cascade). 121 JS guards green, ruff/mypy clean. |
||
|
|
215f7506ba |
feat(rerank): wire endpoint-backed reranking into BM25 retrieval surfaces
Reuse the shipped Cohere/Jina rerank client as an optional post-process on the BM25 surfaces (tool search, skill search, memory composition) via one seam: BM25Index gains an injected reranker + a two-stage search (BM25 recall top-50 -> rerank -> top-k). No new storage. Gated on a configured endpoint plus tools.rerank_bm25 (default on, matching rerank_web_search). tools.rerank_bm25_threshold (default 0.0 = off) is a relevance FLOOR for proactive memory surfacing: BM25 always returns something, so without a floor every-turn memory injection spends tokens on the top-k of whatever lexically matched; the reranker score is what makes a meaningful "inject nothing" gate possible. Two reranker modes (BM25Index rerank_filters): - REORDER (reactive tool/skill search): the reranker must never drop results -> fall back to BM25 order on empty, backfill omitted pool items, so a misbehaving endpoint can't silently lose tools. - FILTER (memory, rerank_filters = threshold > 0): a clean empty/short result is honoured (inject nothing) -- a deliberate divergence from web_search._rerank_results. Parse/endpoint failure is a discrete branch from the floor: an empty result for non-empty input means an unparseable response (a conforming reranker scores every doc), so the closure raises RerankError and BM25Index falls back to BM25 order in BOTH modes -- the floor only acts on valid scores. Also: cap the rerank client timeout at 15s (the per-turn memory path can't afford tools.timeout's 120s default); move the Reranker alias to rerank.py (shared, no import cycle); document the endpoint egress in the rerank_bm25 help, the admin Reranker-role description, and docs/tools.md; add scripts/bench_bm25_rerank.py (manual, needs a live endpoint) to measure precision@k/MRR lift and recommend a threshold default. Negative-tested: reorder fallback-on-empty and omitted-item backfill, filter-mode honor-empty, singleton-still-floored, the parse-fail RerankError raise, the >= floor boundary, and pool-position-to-doc-index mapping -- each guard reverted to confirm its test fails, then restored. |
||
|
|
a8eec0d740 |
fix(vendor): widen update-vendored-js sweep to catch shared_static/ + .py
The shared_static exclude in scripts/update-vendored-js.sh was meant to skip self-references inside vendored libraries, but it also hid shared_static/renderer.js — which loads the vendored libs and pinned mermaid-11.14.0 across every renovate bump since #426. Tests under tests/test_web_helpers.py were similarly invisible because the include list omitted *.py. Replace the broad shared_static exclude with the specific old-versioned vendor directory (about to be rm -rf'd next anyway), and add *.py to the include list. Bump renderer.js to mermaid-11.15.0 to repair the live 404, and refresh the test fixtures to current vendor versions so they stop drifting. |
||
|
|
4b5edce8c5 |
fix(css): two cascade-flip bugs found by specificity audit (#433)
* fix(css): two cascade-flip bugs found by specificity audit PR #431 stripped [data-design="v1"] from ~400 rules, dropping each by a specificity tier; two cascade flips (#header outranking .appbar, #header h1 outranking .appbar-title) were caught visually during that PR's review and fixed by renaming id="header" → id="ui-header" on the per-node UI page. This is the audit follow-up; it found two more: - textarea.skill-content-area (was .skill-content-area) — bumped to (0,1,1) so the rule ties with `.admin-modal textarea` (0,1,1) and wins on source order. Without the bump, min-height: 220px was clobbered to 40px by the modal default and the spec-content textarea rendered short. The three !important markers (font-family/size/line-height) are now redundant against the modal's font: inherit shorthand and are dropped. - h3.skill-spec-heading — removed `font-size: inherit;`. The author wrote it to "reset UA defaults" but it locked font-size to the parent's (~14-16px) at (0,1,1), silently overriding `.skill-spec-heading`'s 10px at (0,1,0). The bare class already beats UA `h3` on specificity (class > tag), so no font-size reset was needed; the `margin-block: 0` line stays because the bare class's `margin: 14px 0 6px` shorthand may not reset the UA's logical margin-block-start/end on every engine. Adds scripts/css_specificity_audit.py — the audit tool. It parses every CSS file referenced from the project's three HTML entry points, computes selector specificity (incl. :not/:is/:has math, attribute selectors, and !important), and flags every place an unscoped legacy rule could outrank a bare-class designed primitive. Honours per-page stylesheet manifests, state-pseudo subset gating (a `:hover` rule overriding a resting-state base rule is intentional, not a flip), and shorthand→longhand expansion for font/padding/margin/border/background. Triage of remaining findings (26 id-tier in default mode, 74 total at --all-tiers) confirmed all are intentional designer overrides — id-scoped buttons, BEM modifier classes, contextual ancestor selectors, last-child margin reset, [hidden] toggle. * fix(css-audit): correct two cascade-resolution bugs flagged by Copilot 1. _parse_declarations dict insertion order didn't update on overwrite, so a sequence like `font-size: 13px; font: inherit; font-size: 12px;` would iterate as (font-size=12px, font=inherit) and the shorthand expansion then clobbered font-size back to `inherit` — wrong. Delete-then-insert on overwrite so the last occurrence lands at the dict's tail and the shorthand expansion sees the real source order. 2. The cascade-winner tie-break used `rule.line_no` only, ignoring the stylesheet load order. A rule at line 1000 of `base.css` looked "later" than a rule at line 50 of `style.css`, even though the page loads `base.css` BEFORE `style.css`. Sort by `(file_index, line_no)` keyed off the element's per-page stylesheet manifest instead. |
||
|
|
7968f1b361 |
feat: auto-invalidate JWT and static assets on version upgrade (#307)
* feat: auto-invalidate JWT and static assets on version upgrade
Add a `ver` claim (major.minor) to user-facing JWTs so tokens from
previous versions are rejected after upgrade, triggering re-login.
Service tokens are excluded for rolling-deployment safety. Tokens
without a `ver` claim (pre-upgrade) are accepted for backward compat.
Inject `?v={__version__}` query strings into static asset URLs at
startup so browsers fetch fresh JS/CSS after any release. Vendored
libraries (KaTeX, Highlight.js, etc.) are skipped since they already
carry version numbers in directory paths. HTML responses now include
`Cache-Control: no-cache` to ensure browsers always revalidate.
Frontend detects upgrade-specific 401s and shows a contextual subtitle
("The server was updated — please sign in again"), then performs a full
page reload after re-auth to load the new versioned assets.
* refactor: address PR review — public API name, single decode, idempotent regex
Rename _version_slot() → jwt_version_slot() to make the cross-module
import explicit rather than relying on a private name.
Move version gating from validate_jwt() into check_request() via a new
AuthResult.token_version field. This eliminates the double JWT decode
that occurred on version-mismatch detection — the token is now decoded
once and the version compared afterward.
Guard version_html() regex against double-apply by excluding URLs that
already contain a query string ([^"?]+ instead of [^"]+).
* feat: structured version_mismatch code, ETag, cross-tab auth sync
Add structured "code": "version_mismatch" field to the 401 response
so the frontend detects upgrade-triggered re-auth without string
matching on the error message.
Add ETag headers to HTML index responses (server, console, and proxied
node UI). Combined with Cache-Control: no-cache, browsers send
conditional GETs and receive 304 between upgrades, saving bandwidth.
Add BroadcastChannel-based cross-tab auth sync so logging in on one
tab dismisses the login modal on all other tabs (and vice-versa for
logout).
Add a reminder to the vendored JS update script about the
version_html() regex lookahead.
* fix: remove unused import in test_web_helpers
|
||
|
|
db0baefeb2 |
feat: render rich media embeds for MCP tool results (#292)
* feat: render rich media embeds for MCP tool results Detect structured media JSON (stream_url, results, sessions) in MCP tool output and render interactive cards instead of plain text. Web UI: media cards with thumbnail, title, metadata, and click-to-play video/audio. HLS via lazy-loaded hls.js with direct-stream preference. Collapsed raw JSON (API keys redacted) for inspection. Discord: rich embeds with proxied thumbnail images (fetched by the bot since Discord CDN cannot reach private media servers). Search results as numbered lists, session state as "Now Playing" cards. Stream URLs never exposed in embeds — web_url used for safe clickable links. CI: vendor hls.js 1.6.15 with renovate tracking and update script. * fix: address PR #292 review — SSRF guards, streaming fetch, tests - URL validation: reject non-http(s) schemes and userinfo in thumbnail URLs. Private IPs intentionally allowed (media servers are on LAN). - Streaming fetch: use http.stream() with aiter_bytes() and a running byte count to enforce the 2MB cap without buffering the full response. Validate content-type is image/* before downloading. - Resilience: wrap try_build_media_embed in try/except in bot.py so a media embed failure falls through to the code-block path. - LICENSE: download hls.js LICENSE from npm on update instead of only copying from old dir. - Tests: add 19 new tests — try_parse_media (8 cases), _is_safe_image_url (7 cases), embed builders (4 cases including stream_url exclusion and string season/episode safety). * chore: add LICENSE file for vendored hls.js * fix: remove ANSI escape codes from tool preview fields Preview text (tool args, URLs, queries) was wrapped in DIM/RESET ANSI codes at the source in session.py, which leaked into SSE events and rendered as raw escape sequences in Discord and the web UI. Move ANSI styling to the CLI consumer (cli.py) where it belongs. Also escape markdown in Discord tool name titles to prevent __ from being interpreted as underline formatting. * fix: drop [MCP: server] prefix from tool descriptions The prefix made MCP tools look second-class compared to builtins, causing models to hesitate using them. The server name is already encoded in the tool name (mcp__server__tool). * feat: pretty-print JSON tool output, player error state, broader key redaction - JSON tool results are detected and pretty-printed with 2-space indent instead of rendering as a wall of text - API key redaction extended to cover api_key, apiKey, api-key, and token query params across all tool output (not just media embeds) - Video/audio player shows styled error message when stream fails to load instead of leaving a broken player element - Both appendToolOutput and replayHistory use shared renderToolOutput() * fix: designer review — player error retry, contrast, tool-cmd cap - Player error: role="alert" for screen readers, retry button that reuses existing play handler, includes media title in error message - Light theme: darken --red from #dc2626 to #b91c1c (5.7:1 contrast on --code-bg, was 4.3:1 failing WCAG AA at 12px) - Pretty-print collapsed raw JSON in media embeds (was missed earlier) - Cap .tool-cmd at 120px to prevent tools with many args from making approval blocks disproportionately tall in history replay - Dedicated .media-player-error class instead of reusing .tool-output * fix: Discord tool info name matching regression, suppress deprecation warning The escape_markdown call on tool names was stored for matching against ToolResultEvent.name, but event.name is raw/unescaped. The escaped name never matched, so the "Running → Done" transition silently failed and previews disappeared from the status embed. Fix: store raw name for matching, use escaped name only for display. Also suppress discord.py's re.sub count deprecation warning (Python 3.13+ issue, fixed upstream). * fix: update MCP tool description tests to match prefix removal * fix: address PR #292 review round 2 - Retry button: handle missing span children in click handler so retry buttons from player error state don't throw - Footer count: use len(lines) instead of min(len(results), 10) to reflect actual rendered count after char budget truncation - Null display: use "null" instead of "None" in JS tool arg preview - Broader redaction: also redact JSON "api_key": "..." patterns - SSRF hardening: block loopback and link-local IPs plus cloud metadata hostnames in thumbnail fetch (private LAN IPs still allowed) |
||
|
|
57080f4615 |
chore: release infrastructure for dual-track stable/experimental (#282)
* chore: release infrastructure for dual-track stable/experimental CI/CD changes for the 1.0 release: - Gate PyPI publish and Docker publish on CI success via workflow_run - Add docker-publish.yml: builds and pushes to GHCR with smart tagging (stable gets :X.Y.Z/:X.Y/:stable/:latest, pre-release gets :experimental) - Add stable/* and v* tags to CI and docker-scan triggers - Remove stale [mq] extra and types-redis from CI (Redis MQ deleted) - Remove stale redis from Renovate package rules Release tooling: - scripts/release.sh: bump version, uv lock, commit, tag (with --push) - docs/releasing.md: documents stable/experimental workflow Docker: - Add /workspace mount point (WORKSPACE_MOUNT env var, defaults to empty volume) - Update .env.example: remove stale Redis/auth-token refs, add workspace/model/discord README: - Remove beta warning, add hero image and release tracks table * fix: derive release tag from git instead of workflow_run.head_branch Use git tag --points-at HEAD after checkout to resolve the release tag instead of relying on workflow_run.head_branch, which may not reliably be the tag name for tag-triggered CI runs. Both publish and docker-publish workflows now skip cleanly when no v* tag exists at the checked-out commit. |
||
|
|
c5d5d0b7cd |
fix: update-vendored-js.sh detects old version from filesystem
The script detected the old version from pyproject.toml, which Renovate had already updated. This caused OLD_DIR == NEW_DIR, so the script downloaded files then immediately deleted them. Fix: detect old version from the actual directory on disk. Add a guard that errors if old == new version to prevent silent data loss. Also: run the fixed script to vendor katex 0.16.40 (fonts + css + js). |
||
|
|
22402e89de |
feat: add dependency management with Renovate, uv.lock, and security … (#83)
* feat: add dependency management with Renovate, uv.lock, and security scanning Adds automated dependency update detection and vulnerability scanning across all dependency layers (Python, vendored JS, TypeScript SDK, Docker, GitHub Actions). - Renovate config with 10 package groups and custom regex managers for vendored JS (KaTeX, Highlight.js, Mermaid) tracking via npm registry - uv.lock for reproducible builds (80 packages) - Dockerfile switched to uv sync --frozen with layer caching - CI: pip-audit (via lock file), npm audit, lock-check jobs - CI: lint job uses pre-commit for ruff version consistency - Docker security scan workflow (weekly Trivy, HIGH/CRITICAL) - Helper script for vendored JS library updates * fix: resolve CI failures and address review feedback - Update pre-commit hooks: ruff v0.9.10 -> v0.15.6 (fixes deprecated UP038 rule), mypy v1.14.1 -> v1.19.1 - Add per-file-ignore for N802 on sandbox.py (ast visitor convention) - Fix pip-audit: install into uv venv so uv run can find it - Pin uv-version in CI to match lock file generator (0.9.18) - Upgrade vitest ^2.0 -> ^4.1 to fix esbuild GHSA-67mh-4wv8-2f99 - Vendored JS script: use grep -rl for auto-discovery of version refs (catches docs/architecture.md), fix LICENSE comment, portable grep |