mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-16 17:01:40 -06:00
cf2811fcebec9919db0701df83155eb4ea75bc97
17 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d7db73305 |
fix(models): review feedback — separator vocabulary, constraints stub, import style
The scopes sanitize now shares the registry guard's separator vocabulary: tab/newline/CR read as spaces, and every other C0 byte — including the U+001C–U+001F block str.split() would silently promote to separators — strips like the control it is, so a control byte inside a token can never split it into two valid-looking scopes (pinned alongside the registry's refusal). The livepass auth-constraints stub serves the new app_identity_auth_modes field so the pass exercises the served-data path for the model list's auth badge, and the session-module import in the mint tests drops to the string-path monkeypatch spelling (single-style imports). |
||
|
|
8605c9783d |
feat(models): rfc8693_obo auth mode, per-alias exchange scopes, identity-keyed mint cache
Adds the dedicated `rfc8693_obo` model auth mode (#955): model definitions gain an `obo_scopes` column (migration 069), the mint threads the scopes to the token-exchange leg (RFC 8693), and every dynamic mode pins its grant leg — a mode is a dialect commitment, not a hint the deployment profile resolves. Exchange-capable IdPs refuse an audience whose scope was not requested; this closes the structurally unmintable model-OBO path on token-exchange deployments. The model mint-cache is identity-keyed on the owning definition's alias (`__model_obo__:<alias>` per user, `__model_app__:<alias>` under the shared app principal), matching the MCP discipline where rows key on the unique server name. The bearer's shape lives in the row's audience/scopes columns and the freshness gate compares it on every read, so a re-aimed alias refuses its old row and overwrites the same key in place. Admin lifecycle (rename, re-aim, scope change, delete) purges a definition's own rows through one shared helper — sound because one definition owns each key; a sibling's rows are untouchable by construction. Cooldown and backoff additionally key on the dispatch shape, so an operator's config repair is an instant clean slate. Cause records, cooldowns, locks and memoization are per-alias end to end, and the session heartbeat reads refusal causes under the same keys. Console: default-deny write gating for dynamic rows (value-diff over the full column ladder, admin.mcp escalation, a never-blockable pure-disable carve-out), a two-tier validator (audience allow-list on every write; deployment-posture checks when the pair is chosen), one shared scopes parser whose omit-unchanged arm keeps over-cap DB-direct residue rows disarmable without ungating real changes, and served constraints (dynamic/scopes/app-identity mode lists, mode-to-profile pairing) so the shelf tracks the registry by data. The admin shelf gains the mode option, a scopes input with residue affordances, pairing-aware option greying, and a derived auth badge. Registry load refuses control characters in alias, audience, and scopes — including the C0 separator block that str.split() would silently collapse — and the C0/DEL class has one exported spelling shared by every surface. Profile-mismatch visibility warns at reload and boot with the mode-correct cause, gated on OIDC being enabled. Breaking: a stored `entra_obo` alias on a deployment whose `[oidc] obo_grant_profile` is `rfc8693` (or the inverse pairing) no longer mints via the profile-driven overload — the mint refuses before any IdP traffic with cause `grant_profile_mismatch`, and the `model.auth_fail_closed` policy governs static fallback. Such rows never minted usefully on scope-gating IdPs; the shelf now surfaces the pairing and the per-turn heartbeat names the refusal cause. Live-verified end to end: scoped token exchange mints, the warm cache serves with zero IdP calls, and the mode/profile mismatch refuses with zero IdP traffic (scripts/obo-e2e/keycloak_e2e.sh); the refresh-redemption profile's E1-E7 hold via scripts/obo-e2e/entra_e2e.py. Closes #955. |
||
|
|
33ace975d2 |
feat(models): default-deny governance and admin UI for per-alias backend auth
Follow-up to the per-alias Entra OBO/app-identity backend auth: the console write path now applies default-deny field classification, the admin shelf gains full backend-auth support, and the session/registry rebind machinery is hardened for config changes landing under live sessions. Console write gate: - Default-deny classification: any non-neutral change to a row that is or becomes dynamic requires admin.mcp plus validation; the provably auth-neutral columns are enumerated (MODEL_AUTH_NEUTRAL_FIELDS) and a live-schema classification test forces every future column to be classified. The derivation is a pure function (_derive_auth_gate) with unit-pinned exclusivity invariants. - Two-tier validation mirroring the MCP oauth_obo validator: the row tier (audience allow-list) runs on every gated write; the posture tier (OIDC configured, token store present) runs on pair changes and on enable-arming. - Pure-disable carve-out: disabling a dynamic row is de-escalation and is never blocked — admin.models suffices and validation is skipped, including for rows with corrupt or skewed stored values. - Capabilities are compared canonically (key order, integral floats), the audience compare normalizes both sides, and staging an audience on a static row is refused on both write twins. - Calibrate writes the capabilities column under an enforced confinement invariant with a compare-and-swap persist. Admin shelf: - Backend-auth section with a per-open constraints fetch (GET /model-definitions/auth-constraints: audience allow-list, grant profile, dynamic modes), datalist audience suggestions, server-defined modes preserved on round-trip, and permission-aware visibility built on cache-skew-safe helpers shared through auth.js. - Refused live-registry swaps surface as an amber registry_warning on the write, delete, reload, and calibrate responses; audit rows carry auth_gated / auth_disarmed markers visible in the audit view. Registry and sessions: - The encryption-key requirement for dynamic auth is enforced inside ModelRegistry.reload() itself — nodes refuse with 503 and the console records coord_registry_error — and reload bumps the generation before the map swap so a racing reader can never pair a stale generation with new maps. - resolve()/resolve_binding() return the generation from inside the registry lock; sessions rebind per send on generation change with atomic client/provider/config commits, fallback-first handling of removed or unconstructable aliases, and judge/limiter resets only when the binding actually changed. - Mint refusals record per-user causes surfaced in the per-turn heartbeat logs; misconfiguration warnings are deduplicated with bounded state. Verification: 10417 tests (99 added on this branch), a 71-scenario browser harness over the real admin shelf, and a live rfc8693 token-exchange e2e run (MCP legs verified end to end; the model-leg scope gap is tracked as #955 under a narrow known-gap signature). Closes #950. |
||
|
|
729a02a833 |
feat(ui): copy-to-clipboard for messages and rendered blocks
Three idle-only affordances on every chat surface: a persistent copy button in each assistant bubble's actions bar, a pointer-only floating button over the hovered markdown block (fence, mermaid diagram, table), and Enter on a focused block for keyboard users, with the outcome flashed on the block itself. Copy resolves to SOURCE, not rendered text. The renderer stashes each table's raw markdown in data-md-source at render time — span sentinels restored in reverse mask order, footnote-definition bodies restored to raw before their recursive render — and whole-message copy reads the streaming pipeline's per-frame stash. The clipboard transport falls back to the legacy execCommand path for plain-HTTP LAN nodes, cloning and restoring the user's selection and focus. Outcomes surface button-local only: flash + title + one live-region announcement through the shared makeAnnouncer factory (also adopted by the interactive voice/tool announcers, whose lazily created regions swallowed their first announcement). Busy refusals answer with their own message. Coordinator retry and admin token-copy keep zero-module-dependency degrade paths. |
||
|
|
6f194d338f |
fix(console): restore back-to-console from a proxied node view
The console proxies a node's web UI at /node/{id}/ and injects a shim into
the page it fetches upstream. The only way back was a node-picker menu the
shim built into #ui-header — an element the L-shell renovation (
|
||
|
|
3a28dc2f16 |
fix(providers): review round 5 — same-id fragment merge, post-finish blip tolerance, chat finish shim
Correctness: - ToolCallSlotter's reannounce split is gated to ID-LESS deltas: id equality proves the same call, so compat servers that repeat the id+name header on every argument fragment merge back into one call with valid JSON (round 4's ungated heuristic split them into duplicate half-JSON calls — execution-confirmed by the review). The residual id-less repeat-name-per-fragment shape is documented as inherently ambiguous; ids are the only disambiguator. - drain_stream keeps a completed result when the transport blips AFTER the finish reason (trailing usage chunk / citation footer window): the generation is in hand, so forfeit the trailing metadata instead of discarding a fully-delivered verdict or re-paying a compaction. - The chat iterator shims finish_reason="stop" when a stream ends CLEANLY after delivering content or tool calls — the deleted non-streaming `or "stop"` default for lax finish-reason-less servers, now safe to restore because abrupt deaths surface as httpx.TransportError (round 4) rather than clean exhaustion. This supersedes the round-3 keep-the-gate ruling: the httpx catch changed the calculus, and the Anthropic/Responses lanes already got their marker-based shims. Empty/reasoning-only streams still fail the complete-or-error gate. Two streaming tests gained the shim chunk. Dispositions held: o1-era stream-rejecting models (third re-report) stay a release-note remediation per the earlier ruling. Cleanup: the two Responses terminal branches collapse into one path (status derived from the event type when the payload is missing — also fixes the end-of-stream debug log reporting finish_reason=None for completed lax streams); the annotations walk is one shared helper (the two copies had already diverged on None-content guarding); _raise_responses_failure is annotated NoReturn; scripts/livepass.py drops the phantom supports_streaming key; test_model_registry's capture helpers ride scripted_chat_client; _openai_stream_chunk points at its fake_chat_stream shape-twin for future consolidation. |
||
|
|
41e7d5b7d7 |
feat(livepass): add long-session perf harness (--perf)
New /perf/livepass.html mounts the real InteractivePane at production scroll geometry and drives production-shaped SSE events through handleEvent/replayHistory in real time (no virtual-time budget, no forced reduced-motion — both corrupt the measurement), reporting: replayHistory wall time at N messages, per-turn live-storm cost on top of that transcript, tool_output_chunk throughput, busy/idle churn, heap + node + agent-card counts across repeated replay cycles (the detached-DOM leak probe), and longtask counts. The --perf runner builds, serves, and launches headless Chrome with --js-flags=--expose-gc and --enable-precise-memory-info so heap numbers are real floors; the page POSTs its JSON report to /perf/report. Reports carry a per-attempt run token the runner validates, so a straggler POST from a killed prior attempt cannot be misattributed to the next size, and the wait loop polls the Chrome process so a sandbox startup failure bails to the --no-sandbox fallback in seconds instead of burning the full timeout. |
||
|
|
8dd356b7e6 |
fix(task-agent): keep sub-tool steps nested + preserve denial reasons
Address the Copilot review on #732 plus a task-agent sub-tool nesting race surfaced alongside it. Nesting (web UI): - A sub-tool step whose task_agent row hasn't painted yet (the 4-wide tool pool's ordering window) buffers and nests when the row lands, instead of escaping to a top-level row that looks main-harness-issued. - A row that never paints (id-correlation mismatch / aborted agent) escapes its buffered steps back to a visible top-level paint after a grace window, so steps are never buffered invisibly or leaked. - The nested card survives the parent row's pending->resolved rebuild; a call_id reused across turns builds a fresh card rather than stealing the prior agent's steps. - tool_info routes through the same nesting path (no duplicate top-level row); a namespaced sub-tool result no longer grafts onto an unrelated top-level row. Denial reasons (backend): - Preserve the specific denial reason a gate already stamped (operator feedback, or the matched policy pattern; web and CLI contracts) instead of clobbering it with a flat "Denied by user" -- in both the sub-agent and the main tool loop. Verified with the livepass task_agent harness (race + orphan-escape scenarios, headless) and unit tests. |
||
|
|
77cb76c006 |
feat(task-agent): recall sub-trajectory + per-agent read isolation
Final chunk of the task_agent modernization: rebuild a finished task
agent's card from /history (reload / reopen while the workstream is in
memory) and isolate each sub-agent's file-read tracking.
Recall: _project_agent_steps projects a sub-agent's trajectory into step
items (FIFO-per-call_id pairing via _iter_agent_tool_results, shared with
_cancel_ledger; output/arguments/count capped); _stash_agent_trajectory
keeps them on the UI in an LRU-bounded store; make_history_handler
attaches them as agent_steps to each task_agent tool_call, and
replayHistory/_replayAgentCard rebuild the collapsed card. In-memory only
(durable persistence deferred); a cold/evicted entry renders the flat
parent row ("not retained"), never a fabricated 0-step card.
Read isolation: _read_files (the blind-overwrite guard's memory) is now
per-sub-agent via the _active_read_files contextvar -- _exec_task copies
the parent's set on spawn and merges the agent's reads back on
completion, so a sibling in the 4-wide pool can't suppress another
agent's guard.
Also: _exec_task now self-reports the task_agent tool_result on every
path (the parent loop only reports error/denied results centrally) --
without it the live card never completed and a failed task recorded
is_error=False in the canonical trajectory. is_error flows from
_tool_error_flags to the recalled step; on_info suppression is per-thread
so a parallel sibling tool's progress isn't dropped.
|
||
|
|
ca7958329a |
feat(task-agent): nest sub-tool steps in an expandable card
Route a task agent's sub-tool events (tool_pending / approve_request, tagged with parent_call_id) into a collapsible card under the task_agent row, replacing the blue on_info turn-legs. - conversation.js / interactive.js: buildAgentCardBody + _routeAgentItems / _ensureAgentCard nest steps by parent_call_id. Collapsed by default (a task agent can run 100+ steps and the parent fans out many in parallel); the label carries the live count + state. Auto-expand when a nested approval is pending so the blocking prompt can't hide behind the toggle. - session.py / session_ui_base.py: on_agent_step paints auto-tool step rows; namespace child call_ids by parent so the 4-wide task pool can't collide on local sequential ids (call_0); suppress sub-agent on_info on the web pane (no call_id to nest by — the card carries steps + result). - cli.py: on_agent_step prints a dim step leg (no card on the CLI, which keeps its on_info). - livepass.py: task-agent card harness driving the real InteractivePane. |
||
|
|
9ad447ca33 |
fix(attachments): design-review polish for preview chips/pills
Two-reviewer + sanity pass over the attachment previews: - composer audio chip is icon+name+size only; the native <audio> player renders on the sent message, not the staging chip (too heavy at chip scale) - cap sent-message pills (+ in-pill audio/snippet) so they no longer overflow the bubble at narrow widths; player and snippet drop to their own row - clamp the chip filename in shared chat.css so long names ellipsize instead of wrapping (console main + coordinator previously left it unclamped) - merge the duplicated .composer-chip rule; drop unused kind-modifier classes and inert vertical-align / inline-block declarations - fix undefined var(--bg-base) -> var(--bg-surface) thumbnail backing - label the <audio> control (aria-label) and drop the decorative snippet from the a11y tree scripts/livepass.py: add an attachments harness that drives the real createAttachmentController + Pane.addUserMessage so these surfaces render headlessly for review. |
||
|
|
44c11efb53 |
feat(ui): split-view follow-ups — per-pane ✕, child-opens-beside, close-on-ws_closed
Four refinements from first live use: - Per-pane ✕ chip, top-right of every visible pane. Split mode: hide that cell (closeCell — the tab stays, the sibling absorbs the space). Single-pane: close the pane outright (withheld from the unclosable Dashboard). The click decides at click time; the label tracks the mode. Manager-injected into the pane section — content untouched. - Coordinator child links open BESIDE the coordinator (openPaneBeside: split right of the focused cell, seeded with the child pane) instead of replacing it — the parent stays on screen. Degrades to the plain focused-cell swap when the split is denied (cap / narrow viewport). splitFocused() gained an optional explicit-fill parameter for this. - Tier-1 ws_closed now CLOSES the open interactive pane (tab gone, a split cell collapses) — the coordinator-closes-its-child flow, matching the standalone's pane-auto-close. The dead-banner lane stays for streams that die without a ws_closed (node crash/network), where the session may still be revivable. - Paint bug: the focused-cell ring was an inset box-shadow on the section, which paints in the element's own background layer — UNDER opaque children touching the edges, so the status bar / composer strip occluded it. The ring now rides a click-transparent ::after overlay above pane content; the 2px top bar sits above the ring line. The livepass shell surface's demo panes grew a .ws-status-bar footer so the occlusion bug class stays visible to future passes. |
||
|
|
f8f7152d63 |
feat(ui): split view returns to the L-shell — PaneManager layout tree
Revives the split-pane feature retired with ui/static (step 6), rebuilt on PaneManager: an optional binary layout tree (null = the one-pane-per- tab behaviour, unchanged) renders visible panes as %-inset cells — no reparenting, so live stream DOM, scroll state and media survive layout changes. Tabs stay global: the active tab is the focused cell, a backgrounded tab swaps into it, clicking inside a visible pane focuses its cell, .shown marks visible-unfocused tabs. Separators resize by pointer-capture drag and arrow keys (role=separator + aria-value*); the tree persists in the working-set blob and rehydrate prunes leaves whose pane did not restore. Limits: 6 cells, 200x150 cell minimums, denials toast the manager's reason. Affordance: Split right / Split down / Unsplit buttons in the tab-bar tail replace the redundant [+] (the permanent Dashboard tab is the launcher) — deliberately no contextmenu override this time. The dead TS_APP.focusLauncher seam goes with it. Measured chrome: the focused cell wears a 2px accent top bar (no thin tinted ring clears 3:1 in both themes) plus a 55%-mix inset ring; separators rest at --ink-4 with solid-accent hover/drag/focus; .shown tabs carry an accent underline; the tail cluster is fenced and lifted to --ink-3. scripts/livepass.py grows a third surface: shell/livepass.html boots the real shell.js + pane.js and drives ?split=right|down|three|none (+ &theme=light), stamping SPLIT-READY-<cells> / SPLIT-FAILED-<reason>. |
||
|
|
b3c3acc5d1 |
fix(scripts): livepass dialog-tier riders + loud open-failure + dock-displacement probe
Designer-review round on the scroll fix found the harness's dialog-tier gate silently green: confirm-dialog (and install/coord-delete) markup lives OUTSIDE #admin-layout, so the fragment extraction never embedded it — ?open=confirm threw at showConfirmModal and screenshot a normal, dialog-less page. build() now injects every hatch dialog the fragment does not already contain, and a driven ?open= that ends with no open dialog stamps OPEN-FAILED-<state> into the title instead of passing. Also upstreams the review's probe states: &focuslast=1 focuses the last shelf-body control (the displaced-dock regression class — only .sh-body may scroll; head/foot must stay pinned) and &scrolled=bottom shows the 24px scroll tail. |
||
|
|
6bb47cad2d |
fix(console): manage-pane scroll regressions — interior scroller, clip the hatch-host, anchor hidden inputs
The L-shell height-pins the admin chain and .hatch-host clipped it, so no box below the pane could scroll: tabs taller than the pane were cut dead, and the overflow:hidden host doubled as a hidden scroll container that focus-into-view silently scrolled — visually-hidden toggle/cap/radio inputs escape the .sh-body scroller (abspos under an unpositioned label), overhang the shelf, and a Tab keypress shoved the docked hatch off its head with no scrollbar to recover by. - .admin-content becomes the manage pane's interior scroller (the #main precedent); switchAdminTab resets it on real tab changes only - .hatch-host: overflow hidden -> clip — paint clipping without a scroll container, so focus can never displace the dock - position:relative anchors on the three hidden-input labels (toggle-switch, .sh-body .cap, segmented-option); .settings-toggle already carried one - livepass: the console harness wraps the fragment in the REAL L-shell chain (its bespoke height pin is exactly how this bug class stayed invisible to the screenshot gates) and gains a ?tall=1/&scrolled=1 scroll state |
||
|
|
b991dc2e83 |
fix: CI lint pin + copilot-thread hardening
The wiring-lint test used percent-formatted regex patterns — UP031 under the ruff 0.15.6 the CI pre-commit pins (the older venv binary let it through; checked repo-wide against the exact pin now). f-strings with doubled quantifier braces, plus one over-long fixture line in the livepass generator split. Copilot threads, both validated rather than blindly applied: - closeShelf's scrim-ownership scan now skips detached entries. The thread's throw scenario doesn't occur on the real removal path (a pane close detaches an ANCESTOR, so _hostOf still resolves inside the detached subtree) — but a detached shelf is genuinely not a scrim owner, so the guard is correct beyond being defensive. - toast.js drops the popover attribute via removeAttribute instead of the null assignment. The claim that null leaves popover="null" is refuted — the IDL is nullable and null removes the attribute (verified empirically in headless Chrome) — but removeAttribute reads correct without requiring that spec knowledge. |
||
|
|
9102f858a4 |
chore(scripts): commit the livepass harness generator
The livepass harness — the headless-render rig that verified every converted modal surface and click-drives submits (the dead-Save bug class) — lived as ad hoc files in /tmp and got wiped once already. The durable piece is the GENERATOR: the markup is extracted fresh from the index files at build time (a committed snapshot would drift) and the stylesheets/scripts are symlinked so edits are live on refresh. scripts/livepass.py builds both harnesses into /tmp/livepass/ (ui: all six dialog-tier surfaces incl. the real cards.js batch controller drive; console: the admin-pane fragment hosting the shelves, with schedule/model/policy/confirm/token fixtures and the model-save click drive that flips document.title to PUT-OK-<n>). --serve included; the chrome screenshot incantation and the ?open= registry are in the module docstring. Governance fixtures (roles/HR/OGP/memory/skill) are documented seams for when those surfaces need driving. |