The scopes sanitize now shares the registry guard's separator
vocabulary: tab/newline/CR read as spaces, and every other C0 byte —
including the U+001C–U+001F block str.split() would silently promote to
separators — strips like the control it is, so a control byte inside a
token can never split it into two valid-looking scopes (pinned
alongside the registry's refusal).
The livepass auth-constraints stub serves the new
app_identity_auth_modes field so the pass exercises the served-data
path for the model list's auth badge, and the session-module import in
the mint tests drops to the string-path monkeypatch spelling
(single-style imports).
Adds the dedicated `rfc8693_obo` model auth mode (#955): model
definitions gain an `obo_scopes` column (migration 069), the mint
threads the scopes to the token-exchange leg (RFC 8693), and every
dynamic mode pins its grant leg — a mode is a dialect commitment, not a
hint the deployment profile resolves. Exchange-capable IdPs refuse an
audience whose scope was not requested; this closes the structurally
unmintable model-OBO path on token-exchange deployments.
The model mint-cache is identity-keyed on the owning definition's
alias (`__model_obo__:<alias>` per user, `__model_app__:<alias>` under
the shared app principal), matching the MCP discipline where rows key
on the unique server name. The bearer's shape lives in the row's
audience/scopes columns and the freshness gate compares it on every
read, so a re-aimed alias refuses its old row and overwrites the same
key in place. Admin lifecycle (rename, re-aim, scope change, delete)
purges a definition's own rows through one shared helper — sound
because one definition owns each key; a sibling's rows are untouchable
by construction. Cooldown and backoff additionally key on the dispatch
shape, so an operator's config repair is an instant clean slate. Cause
records, cooldowns, locks and memoization are per-alias end to end,
and the session heartbeat reads refusal causes under the same keys.
Console: default-deny write gating for dynamic rows (value-diff over
the full column ladder, admin.mcp escalation, a never-blockable
pure-disable carve-out), a two-tier validator (audience allow-list on
every write; deployment-posture checks when the pair is chosen), one
shared scopes parser whose omit-unchanged arm keeps over-cap DB-direct
residue rows disarmable without ungating real changes, and served
constraints (dynamic/scopes/app-identity mode lists, mode-to-profile
pairing) so the shelf tracks the registry by data. The admin shelf
gains the mode option, a scopes input with residue affordances,
pairing-aware option greying, and a derived auth badge.
Registry load refuses control characters in alias, audience, and
scopes — including the C0 separator block that str.split() would
silently collapse — and the C0/DEL class has one exported spelling
shared by every surface. Profile-mismatch visibility warns at reload
and boot with the mode-correct cause, gated on OIDC being enabled.
Breaking: a stored `entra_obo` alias on a deployment whose
`[oidc] obo_grant_profile` is `rfc8693` (or the inverse pairing) no
longer mints via the profile-driven overload — the mint refuses before
any IdP traffic with cause `grant_profile_mismatch`, and the
`model.auth_fail_closed` policy governs static fallback. Such rows
never minted usefully on scope-gating IdPs; the shelf now surfaces the
pairing and the per-turn heartbeat names the refusal cause.
Live-verified end to end: scoped token exchange mints, the warm cache
serves with zero IdP calls, and the mode/profile mismatch refuses with
zero IdP traffic (scripts/obo-e2e/keycloak_e2e.sh); the
refresh-redemption profile's E1-E7 hold via scripts/obo-e2e/entra_e2e.py.
Closes#955.
Follow-up to the per-alias Entra OBO/app-identity backend auth: the
console write path now applies default-deny field classification, the
admin shelf gains full backend-auth support, and the session/registry
rebind machinery is hardened for config changes landing under live
sessions.
Console write gate:
- Default-deny classification: any non-neutral change to a row that is
or becomes dynamic requires admin.mcp plus validation; the provably
auth-neutral columns are enumerated (MODEL_AUTH_NEUTRAL_FIELDS) and a
live-schema classification test forces every future column to be
classified. The derivation is a pure function (_derive_auth_gate)
with unit-pinned exclusivity invariants.
- Two-tier validation mirroring the MCP oauth_obo validator: the row
tier (audience allow-list) runs on every gated write; the posture
tier (OIDC configured, token store present) runs on pair changes and
on enable-arming.
- Pure-disable carve-out: disabling a dynamic row is de-escalation and
is never blocked — admin.models suffices and validation is skipped,
including for rows with corrupt or skewed stored values.
- Capabilities are compared canonically (key order, integral floats),
the audience compare normalizes both sides, and staging an audience
on a static row is refused on both write twins.
- Calibrate writes the capabilities column under an enforced
confinement invariant with a compare-and-swap persist.
Admin shelf:
- Backend-auth section with a per-open constraints fetch
(GET /model-definitions/auth-constraints: audience allow-list, grant
profile, dynamic modes), datalist audience suggestions,
server-defined modes preserved on round-trip, and permission-aware
visibility built on cache-skew-safe helpers shared through auth.js.
- Refused live-registry swaps surface as an amber registry_warning on
the write, delete, reload, and calibrate responses; audit rows carry
auth_gated / auth_disarmed markers visible in the audit view.
Registry and sessions:
- The encryption-key requirement for dynamic auth is enforced inside
ModelRegistry.reload() itself — nodes refuse with 503 and the
console records coord_registry_error — and reload bumps the
generation before the map swap so a racing reader can never pair a
stale generation with new maps.
- resolve()/resolve_binding() return the generation from inside the
registry lock; sessions rebind per send on generation change with
atomic client/provider/config commits, fallback-first handling of
removed or unconstructable aliases, and judge/limiter resets only
when the binding actually changed.
- Mint refusals record per-user causes surfaced in the per-turn
heartbeat logs; misconfiguration warnings are deduplicated with
bounded state.
Verification: 10417 tests (99 added on this branch), a 71-scenario
browser harness over the real admin shelf, and a live rfc8693
token-exchange e2e run (MCP legs verified end to end; the model-leg
scope gap is tracked as #955 under a narrow known-gap signature).
Closes#950.
Three idle-only affordances on every chat surface: a persistent copy
button in each assistant bubble's actions bar, a pointer-only floating
button over the hovered markdown block (fence, mermaid diagram, table),
and Enter on a focused block for keyboard users, with the outcome
flashed on the block itself.
Copy resolves to SOURCE, not rendered text. The renderer stashes each
table's raw markdown in data-md-source at render time — span sentinels
restored in reverse mask order, footnote-definition bodies restored to
raw before their recursive render — and whole-message copy reads the
streaming pipeline's per-frame stash. The clipboard transport falls
back to the legacy execCommand path for plain-HTTP LAN nodes, cloning
and restoring the user's selection and focus.
Outcomes surface button-local only: flash + title + one live-region
announcement through the shared makeAnnouncer factory (also adopted by
the interactive voice/tool announcers, whose lazily created regions
swallowed their first announcement). Busy refusals answer with their
own message. Coordinator retry and admin token-copy keep
zero-module-dependency degrade paths.
The console proxies a node's web UI at /node/{id}/ and injects a shim into
the page it fetches upstream. The only way back was a node-picker menu the
shim built into #ui-header — an element the L-shell renovation (cc508cf4,
shipped v1.6.0) removed from the server UI. buildPicker() has returned on
its first line ever since, so every supported node version has served a
proxied page with no in-UI way back; the browser back button or a
hand-edited URL were the only exits.
The shim now repoints the rail brand (.rail-brand .brand-home) at "/" and
relabels it, so the element users already read as "go home" goes home. It
captures at the document rather than on the button: shell.js binds a bubble
listener to that same element, and stopPropagation() keeps showHome() from
firing as well. The node's own dashboard stays reachable as the
non-closable first tab.
The dead picker goes with it — its JS, the CSS constant that styled only
its elements, the el() helper, and the NODE_ID_PLACEHOLDER substitution
whose only reader it was. It could only ever have run for nodes at or below
v1.5.x, which are not a supported configuration.
The failure mode here is silent by construction: the shim reaches across a
process boundary to select classes another file emits, and fails soft when
they stop matching. Nothing failed when #ui-header disappeared. So the
coupling is now pinned from both ends.
- tests/test_shell_js.py asserts both halves. The class names are derived
from the shim's own querySelector calls, so a newly selected class is
covered without editing the guard, and a vacuity floor keeps it from
going green if the selectors are removed entirely. The containment edges
are pinned separately, deriving shell.js's local variable names from
source: renaming a local stays green, re-parenting .brand-home out of
.rail-brand does not.
- tests/test_console.py executes the shim under node against a two-walk DOM
dispatcher — capture walk, then bubble walk, phase-filtered at every node
including the target. Modelling the real rule means the test accepts any
correct wiring rather than only the one that shipped.
- The injection test drives the real proxy_index against the real node
index. That also pins the bare <body> the literal replace() depends on;
an attribute there would silently drop the entire shim, prefix rewriting
included.
- scripts/livepass.py gains a proxybrand harness for manual verification:
an iframe host over the real shell.js and the real shim, reading the
frame's post-navigation location from the surviving top page. The shim is
read out of the source by text rather than imported, since scripts/ has
no sys.path guard and an import resolves to site-packages.
Correctness:
- ToolCallSlotter's reannounce split is gated to ID-LESS deltas: id
equality proves the same call, so compat servers that repeat the
id+name header on every argument fragment merge back into one call
with valid JSON (round 4's ungated heuristic split them into
duplicate half-JSON calls — execution-confirmed by the review). The
residual id-less repeat-name-per-fragment shape is documented as
inherently ambiguous; ids are the only disambiguator.
- drain_stream keeps a completed result when the transport blips AFTER
the finish reason (trailing usage chunk / citation footer window):
the generation is in hand, so forfeit the trailing metadata instead
of discarding a fully-delivered verdict or re-paying a compaction.
- The chat iterator shims finish_reason="stop" when a stream ends
CLEANLY after delivering content or tool calls — the deleted
non-streaming `or "stop"` default for lax finish-reason-less servers,
now safe to restore because abrupt deaths surface as
httpx.TransportError (round 4) rather than clean exhaustion. This
supersedes the round-3 keep-the-gate ruling: the httpx catch changed
the calculus, and the Anthropic/Responses lanes already got their
marker-based shims. Empty/reasoning-only streams still fail the
complete-or-error gate. Two streaming tests gained the shim chunk.
Dispositions held: o1-era stream-rejecting models (third re-report)
stay a release-note remediation per the earlier ruling.
Cleanup: the two Responses terminal branches collapse into one path
(status derived from the event type when the payload is missing —
also fixes the end-of-stream debug log reporting finish_reason=None
for completed lax streams); the annotations walk is one shared helper
(the two copies had already diverged on None-content guarding);
_raise_responses_failure is annotated NoReturn; scripts/livepass.py
drops the phantom supports_streaming key; test_model_registry's
capture helpers ride scripted_chat_client; _openai_stream_chunk points
at its fake_chat_stream shape-twin for future consolidation.
New /perf/livepass.html mounts the real InteractivePane at production
scroll geometry and drives production-shaped SSE events through
handleEvent/replayHistory in real time (no virtual-time budget, no
forced reduced-motion — both corrupt the measurement), reporting:
replayHistory wall time at N messages, per-turn live-storm cost on top
of that transcript, tool_output_chunk throughput, busy/idle churn,
heap + node + agent-card counts across repeated replay cycles (the
detached-DOM leak probe), and longtask counts.
The --perf runner builds, serves, and launches headless Chrome with
--js-flags=--expose-gc and --enable-precise-memory-info so heap
numbers are real floors; the page POSTs its JSON report to
/perf/report. Reports carry a per-attempt run token the runner
validates, so a straggler POST from a killed prior attempt cannot be
misattributed to the next size, and the wait loop polls the Chrome
process so a sandbox startup failure bails to the --no-sandbox
fallback in seconds instead of burning the full timeout.
Address the Copilot review on #732 plus a task-agent sub-tool nesting
race surfaced alongside it.
Nesting (web UI):
- A sub-tool step whose task_agent row hasn't painted yet (the 4-wide
tool pool's ordering window) buffers and nests when the row lands,
instead of escaping to a top-level row that looks main-harness-issued.
- A row that never paints (id-correlation mismatch / aborted agent)
escapes its buffered steps back to a visible top-level paint after a
grace window, so steps are never buffered invisibly or leaked.
- The nested card survives the parent row's pending->resolved rebuild; a
call_id reused across turns builds a fresh card rather than stealing the
prior agent's steps.
- tool_info routes through the same nesting path (no duplicate top-level
row); a namespaced sub-tool result no longer grafts onto an unrelated
top-level row.
Denial reasons (backend):
- Preserve the specific denial reason a gate already stamped (operator
feedback, or the matched policy pattern; web and CLI contracts) instead
of clobbering it with a flat "Denied by user" -- in both the sub-agent
and the main tool loop.
Verified with the livepass task_agent harness (race + orphan-escape
scenarios, headless) and unit tests.
Final chunk of the task_agent modernization: rebuild a finished task
agent's card from /history (reload / reopen while the workstream is in
memory) and isolate each sub-agent's file-read tracking.
Recall: _project_agent_steps projects a sub-agent's trajectory into step
items (FIFO-per-call_id pairing via _iter_agent_tool_results, shared with
_cancel_ledger; output/arguments/count capped); _stash_agent_trajectory
keeps them on the UI in an LRU-bounded store; make_history_handler
attaches them as agent_steps to each task_agent tool_call, and
replayHistory/_replayAgentCard rebuild the collapsed card. In-memory only
(durable persistence deferred); a cold/evicted entry renders the flat
parent row ("not retained"), never a fabricated 0-step card.
Read isolation: _read_files (the blind-overwrite guard's memory) is now
per-sub-agent via the _active_read_files contextvar -- _exec_task copies
the parent's set on spawn and merges the agent's reads back on
completion, so a sibling in the 4-wide pool can't suppress another
agent's guard.
Also: _exec_task now self-reports the task_agent tool_result on every
path (the parent loop only reports error/denied results centrally) --
without it the live card never completed and a failed task recorded
is_error=False in the canonical trajectory. is_error flows from
_tool_error_flags to the recalled step; on_info suppression is per-thread
so a parallel sibling tool's progress isn't dropped.
Route a task agent's sub-tool events (tool_pending / approve_request,
tagged with parent_call_id) into a collapsible card under the task_agent
row, replacing the blue on_info turn-legs.
- conversation.js / interactive.js: buildAgentCardBody +
_routeAgentItems / _ensureAgentCard nest steps by parent_call_id.
Collapsed by default (a task agent can run 100+ steps and the parent
fans out many in parallel); the label carries the live count + state.
Auto-expand when a nested approval is pending so the blocking prompt
can't hide behind the toggle.
- session.py / session_ui_base.py: on_agent_step paints auto-tool step
rows; namespace child call_ids by parent so the 4-wide task pool can't
collide on local sequential ids (call_0); suppress sub-agent on_info on
the web pane (no call_id to nest by — the card carries steps + result).
- cli.py: on_agent_step prints a dim step leg (no card on the CLI, which
keeps its on_info).
- livepass.py: task-agent card harness driving the real InteractivePane.
Two-reviewer + sanity pass over the attachment previews:
- composer audio chip is icon+name+size only; the native <audio> player
renders on the sent message, not the staging chip (too heavy at chip scale)
- cap sent-message pills (+ in-pill audio/snippet) so they no longer overflow
the bubble at narrow widths; player and snippet drop to their own row
- clamp the chip filename in shared chat.css so long names ellipsize instead
of wrapping (console main + coordinator previously left it unclamped)
- merge the duplicated .composer-chip rule; drop unused kind-modifier classes
and inert vertical-align / inline-block declarations
- fix undefined var(--bg-base) -> var(--bg-surface) thumbnail backing
- label the <audio> control (aria-label) and drop the decorative snippet from
the a11y tree
scripts/livepass.py: add an attachments harness that drives the real
createAttachmentController + Pane.addUserMessage so these surfaces render
headlessly for review.
Four refinements from first live use:
- Per-pane ✕ chip, top-right of every visible pane. Split mode: hide
that cell (closeCell — the tab stays, the sibling absorbs the space).
Single-pane: close the pane outright (withheld from the unclosable
Dashboard). The click decides at click time; the label tracks the
mode. Manager-injected into the pane section — content untouched.
- Coordinator child links open BESIDE the coordinator (openPaneBeside:
split right of the focused cell, seeded with the child pane) instead
of replacing it — the parent stays on screen. Degrades to the plain
focused-cell swap when the split is denied (cap / narrow viewport).
splitFocused() gained an optional explicit-fill parameter for this.
- Tier-1 ws_closed now CLOSES the open interactive pane (tab gone, a
split cell collapses) — the coordinator-closes-its-child flow,
matching the standalone's pane-auto-close. The dead-banner lane
stays for streams that die without a ws_closed (node crash/network),
where the session may still be revivable.
- Paint bug: the focused-cell ring was an inset box-shadow on the
section, which paints in the element's own background layer — UNDER
opaque children touching the edges, so the status bar / composer
strip occluded it. The ring now rides a click-transparent ::after
overlay above pane content; the 2px top bar sits above the ring line.
The livepass shell surface's demo panes grew a .ws-status-bar footer so
the occlusion bug class stays visible to future passes.
Revives the split-pane feature retired with ui/static (step 6), rebuilt
on PaneManager: an optional binary layout tree (null = the one-pane-per-
tab behaviour, unchanged) renders visible panes as %-inset cells — no
reparenting, so live stream DOM, scroll state and media survive layout
changes. Tabs stay global: the active tab is the focused cell, a
backgrounded tab swaps into it, clicking inside a visible pane focuses
its cell, .shown marks visible-unfocused tabs. Separators resize by
pointer-capture drag and arrow keys (role=separator + aria-value*); the
tree persists in the working-set blob and rehydrate prunes leaves whose
pane did not restore. Limits: 6 cells, 200x150 cell minimums, denials
toast the manager's reason.
Affordance: Split right / Split down / Unsplit buttons in the tab-bar
tail replace the redundant [+] (the permanent Dashboard tab is the
launcher) — deliberately no contextmenu override this time. The dead
TS_APP.focusLauncher seam goes with it.
Measured chrome: the focused cell wears a 2px accent top bar (no thin
tinted ring clears 3:1 in both themes) plus a 55%-mix inset ring;
separators rest at --ink-4 with solid-accent hover/drag/focus; .shown
tabs carry an accent underline; the tail cluster is fenced and lifted
to --ink-3.
scripts/livepass.py grows a third surface: shell/livepass.html boots
the real shell.js + pane.js and drives ?split=right|down|three|none
(+ &theme=light), stamping SPLIT-READY-<cells> / SPLIT-FAILED-<reason>.
Designer-review round on the scroll fix found the harness's dialog-tier
gate silently green: confirm-dialog (and install/coord-delete) markup
lives OUTSIDE #admin-layout, so the fragment extraction never embedded
it — ?open=confirm threw at showConfirmModal and screenshot a normal,
dialog-less page. build() now injects every hatch dialog the fragment
does not already contain, and a driven ?open= that ends with no open
dialog stamps OPEN-FAILED-<state> into the title instead of passing.
Also upstreams the review's probe states: &focuslast=1 focuses the last
shelf-body control (the displaced-dock regression class — only .sh-body
may scroll; head/foot must stay pinned) and &scrolled=bottom shows the
24px scroll tail.
The L-shell height-pins the admin chain and .hatch-host clipped it, so no
box below the pane could scroll: tabs taller than the pane were cut dead,
and the overflow:hidden host doubled as a hidden scroll container that
focus-into-view silently scrolled — visually-hidden toggle/cap/radio
inputs escape the .sh-body scroller (abspos under an unpositioned label),
overhang the shelf, and a Tab keypress shoved the docked hatch off its
head with no scrollbar to recover by.
- .admin-content becomes the manage pane's interior scroller (the #main
precedent); switchAdminTab resets it on real tab changes only
- .hatch-host: overflow hidden -> clip — paint clipping without a scroll
container, so focus can never displace the dock
- position:relative anchors on the three hidden-input labels
(toggle-switch, .sh-body .cap, segmented-option); .settings-toggle
already carried one
- livepass: the console harness wraps the fragment in the REAL L-shell
chain (its bespoke height pin is exactly how this bug class stayed
invisible to the screenshot gates) and gains a ?tall=1/&scrolled=1
scroll state
The wiring-lint test used percent-formatted regex patterns — UP031 under
the ruff 0.15.6 the CI pre-commit pins (the older venv binary let it
through; checked repo-wide against the exact pin now). f-strings with
doubled quantifier braces, plus one over-long fixture line in the
livepass generator split.
Copilot threads, both validated rather than blindly applied:
- closeShelf's scrim-ownership scan now skips detached entries. The
thread's throw scenario doesn't occur on the real removal path (a pane
close detaches an ANCESTOR, so _hostOf still resolves inside the
detached subtree) — but a detached shelf is genuinely not a scrim
owner, so the guard is correct beyond being defensive.
- toast.js drops the popover attribute via removeAttribute instead of
the null assignment. The claim that null leaves popover="null" is
refuted — the IDL is nullable and null removes the attribute (verified
empirically in headless Chrome) — but removeAttribute reads correct
without requiring that spec knowledge.
The livepass harness — the headless-render rig that verified every
converted modal surface and click-drives submits (the dead-Save bug
class) — lived as ad hoc files in /tmp and got wiped once already.
The durable piece is the GENERATOR: the markup is extracted fresh from
the index files at build time (a committed snapshot would drift) and
the stylesheets/scripts are symlinked so edits are live on refresh.
scripts/livepass.py builds both harnesses into /tmp/livepass/ (ui:
all six dialog-tier surfaces incl. the real cards.js batch controller
drive; console: the admin-pane fragment hosting the shelves, with
schedule/model/policy/confirm/token fixtures and the model-save click
drive that flips document.title to PUT-OK-<n>). --serve included;
the chrome screenshot incantation and the ?open= registry are in the
module docstring. Governance fixtures (roles/HR/OGP/memory/skill) are
documented seams for when those surfaces need driving.