Operator context moved to first-class system turns, leaving _reminders written
by nothing and read by nothing. Nulling it (the prior 060 step) left a writable
dead column — a foot-gun inviting accidental reuse. Drop it outright and remove
every reference in one shot so there is no half-alive state:
- migration 060: replace the wholesale null with batch_alter_table drop_column
(per migration 027); downgrade re-adds the empty column to match the 059
schema (the envelope un-wrap stays irreversible).
- _schema.py: remove the column.
- _sqlite / _postgresql: drop the reminders save param, the INSERT/bulk values,
and both SELECT columns.
- reconstruct_messages: the row tuple is now 8/9-tuple (event_id shifts from
index 9 to 8); _utils + the _row test helper updated.
- _protocol / memory save_message: drop the reminders param + docstrings.
- tests: replace the reminders-roundtrip tests with a _source-only file and a
060 drop-column assertion; remove the obsolete legacy-reminders wire test.
No production caller passed reminders=, and the SELECT no longer reads the
column, so an un-migrated DB simply ignores any residual values.
Phase-2 follow-ups to the mid-conversation-system consolidation:
- user_interjection framing (known #2): a queued message that drains mid-turn is
re-framed via render_user_interjection ("The user sent … User message: …") so
the user's words keep USER authority, not operator authority — the regression
mattered most on the native path, where the turn enters as a real role=system
message. Empty/whitespace interjections (e.g. a bare "!!!") are dropped (bug-2).
- empty-content user turns dropped at the wire boundary after the fold
(known #3): the wake pipeline's synthetic empty send("") leaves an empty user
turn on the native path (the nudge stays inline); an empty user message is
invalid on every provider. The drop runs after the fold so the fold-path wake
turn, which the nudge fills, survives.
- leading-system guard (_anthropic): a turn that converts to nothing no longer
lets a system message become messages[0] (the API requires messages[0]=user).
Newly reachable now that the empty-turn drop can expose it on a fresh-session
native wake.
- refresh stale .msg.watch-result comments (the card was removed) to describe
the current operator-bubble rendering.
Phase-1 review follow-ups:
- neutralize() now tolerates whitespace between '<' and the slash ('< /tag',
'< /tag'), matching output_guard's detection regex so a marker can no longer
be detected-but-not-defanged (a leaked-nonce break-out gap).
- Add direct tests for the sec-1 forge-in defence: _neutralize_host defangs a
forged <system-reminder_{nonce}> in both string- and list-content untrusted
hosts before the real fence is appended, and the host is defanged exactly once
so consecutive folds don't corrupt the first appended fence.
PlanReviewView was removed alongside the plan_agent built-in tool (110d44b0),
but its Discord owner-check tests were left behind importing a class that no
longer exists. They raise ImportError wherever discord.py is installed (green in
CI only because discord.py is absent there). Remove the dead test class and the
now-unused send_plan_feedback router mock from the shared bot double.
060 un-wrapped legacy <tool_output> envelopes with a loose guard (open + close)
that could irreversibly mis-rewrite a bare tool row resembling the open, and
entity-decoded the wrapper tags back to live form — re-activating injection the
old escape had neutralised (and downgrade cannot undo it).
- Require the full legacy signature (the exact </tool_output>\n\n<system-
reminder>\n join plus a trailing </system-reminder>), which wrap_tool_result
only ever emitted with advisories. A bare row with a matching close but no
advisory is left byte-for-byte untouched.
- Reverse only & -> & ; leave the wrapper-tag entities escaped so a
previously-defanged injection stays defanged.
Adds false-positive guard tests (open+close without advisory; missing tail).
Both the operator fold and the output-guard judge wrap spans in nonce-delimited
fences, but the two had drifted: the operator path minted a 32-bit nonce reused
per session with no body escaping, while the judge used a 64-bit per-call nonce
plus closing-tag escaping. Extract the shared mechanism (mint/neutralise/wrap)
into turnstone/core/fence.py and put both callers on it so they cannot diverge
again.
- Operator fold (sec-1): 64-bit nonce; fence.wrap neutralises the operator
body's close marker, and _fold_system_turns neutralises the untrusted host
turn's <system-reminder> markers once before the first fold, so a leaked or
guessed per-session nonce still cannot forge a trusted block. Per-session +
cached declaration kept (the declaration pins the exact value, so per-turn
rotation would bust the prompt cache). Marker is now <system-reminder_{nonce}>.
- Judge: refactored onto fence (behaviour-preserving; still per-call).
- Forgery detection: output_guard scans tool output for trust-fence markers —
an exact session-nonce match is HIGH (operator_marker_leak: the token has
leaked and is being replayed), any other marker LOW (operator_marker_forgery).
Removes mint_envelope_nonce / wrap_system_context (folded into fence.wrap).
Replace the two operator-context hacks (the <tool_output>/<system-reminder> content envelope and the transient _reminders side-channel) with one persistent {role: system, _source} trajectory turn. Adds supports_mid_conversation_system (claude-opus-4-8): native models take the turn inline; all others fold it into the preceding turn as a nonce-delimited <system-reminder> block declared in the system prompt as the sole trusted marker. Producers (advisories, metacog nudges, user interjections, idle/watch) emit system turns; the envelope/_reminders machinery, escaping round-trip, replay parser, and reminder SSE events are removed. Eager 060 migration un-wraps legacy envelopes. Net -1662 lines.
Known follow-ups from review (unfixed here): (1) the 060 un-wrap heuristic can irreversibly mis-rewrite bare tool rows that resemble the envelope, so do not run the migration until it is tightened; (2) user_interjection turns lost the user-framing/priority preamble (a regression, and a native-path authority-framing concern); (3) native-path wake nudge can emit empty user content.
- rerank_config.py: the runtime instruction fallback had a dead tail
(`get_rerank_instruction() or str(cs.get(...))` -- the cs.get term can only
return the registry default ""), via a stored_keys() branch that also diverged
from the calibrate CLI / endpoint. Collapse to the sibling idiom
(`cs.get(...) or get_rerank_instruction()`) so the instruction used at
calibration time matches the one used at runtime. Correct the module docstring:
ChatSession is the sole caller (the CLI/endpoint share only the instruction
precedence, not this function).
- session.py: the deferred first-turn memory recompose was gated on the flag
alone, so a synthetic wake send (empty user content -> flag stays False) re-ran
the full compose on every wake before the first real turn. Gate on a non-empty
query too, so wakes don't re-pay it and the real turn still fires exactly once
(+ test).
Accepted as-is: the __init__ compose (kept so system_messages/_agent_system_messages
are valid for early readers; one cheap extra compose per fresh session) and the
orphaned tools.rerank_* config rows (inert -- no read path, never listed or
redacted; a purge migration would collide with the 060 in flight on another branch).
Proactive memory selection scores candidates against the recent-user-message
query (extract_recent_context), but a fresh session composes the system prefix
once in __init__ while self.messages is still empty. That empty query takes the
no-context path: _select_memory_candidates returns recency order and
score_memories returns memories[:k] verbatim -- the 5 most recently UPDATED
memories, with BM25 and the reranker never invoked. send() never recomposes, so
those recency-only memories are what the model sees for the whole session (until
an unrelated event -- skill / MCP / model refresh / resume / memory write /
command -- happens to rebuild the prefix). Net effect: the injected memories are
unrelated to the actual question.
Fix: defer the memory-bearing compose to the first real user turn.
- Track _system_composed_with_context, set once extract_recent_context is
non-empty in _init_system_messages.
- send() recomposes once, right after _append_user_turn, while the flag is still
False -- so the opening turn's memory block is selected (and reranked) against
the real message. The flag then stays True, so the prefix is composed once and
stays cache-stable exactly as before (no per-turn prompt-cache churn).
This is the targeted fix; per-turn memory refresh (so later topic shifts also
re-rank) is the larger tail-injection redesign tracked on another branch. Two
adjacent gaps are left as-is for now: the reranker/BM25 only see content[:200],
and build_memory_context flat-truncates each memory to 500 chars (max_content is
the save cap, not an injection budget).
Tests: flag is False on a fresh session and after a whitespace-only wake turn,
flips True on a real query; send() runs the deferred recompose with the user
message in the query.
The reranker_alias -> model-definition path (added when reranking became a model
role) made the older global endpoint settings redundant. Resolve reranking
solely through the Reranker role and remove the parallel global config.
- Removed settings tools.rerank_url / rerank_model / rerank_api_key, their
config.py getters (+ $TURNSTONE_RERANK_URL / $TURNSTONE_RERANK_MODEL and the
module caches), and the fallback branch in resolve_rerank_client_from. The
resolver now returns a client only when a Reranker model (capability
supports_rerank, base_url = its /rerank endpoint) is selected, else None.
- Kept as global knobs: reranker_alias (the selector), rerank_web_search,
rerank_bm25, rerank_bm25_threshold, and rerank_instruction -- a task-level
query knob (Qwen3-style), not endpoint identity.
- The Settings tab is registry-driven, so the three fields disappear with their
SettingDefs. Updated the Reranker role help, example config, and docs/tools.md.
BREAKING: a reranker configured via [tools] rerank_url (config.toml / env /
Settings tab) no longer works -- add the reranker in the admin Models tab and
pick it under Models -> Roles -> Reranker. No migration: reranking is days old
and disabled by default, so any orphaned tools.rerank_* config rows are inert.
Tests: the resolver covers no-store / no-alias / non-rerank-alias -> None and the
model-definition happy path; the obsolete global-fallback tests are removed.
Adding a reranker model through the admin modal was impossible: Detect
(admin_detect_model) always ran the OpenAI /v1/models probe first and gated
calibration on its `reachable` result. A Cohere/Jina /rerank endpoint can't
answer /v1/models, so it failed both ways -- `$host/v1` passed the probe but
calibration POSTed to the wrong path, `$host/rerank` 404'd the probe outright.
- console/server.py: branch on `supports_rerank` BEFORE the probe and calibrate
the endpoint directly; that round-trip IS the reachability check (there is no
independent /v1/models signal for a rerank-only endpoint). Success
autopopulates the three calibration fields the way context_window does;
calibration failure -> reachable:False + error; empty base_url -> 400 with the
/rerank hint; reachable-but-no-clean-separation -> a note. Drops the now-dead
post-probe calibrate-on-detect block.
Reranker selection stays per-model (reranker_alias -> registry); recalibrate on
a saved reranker was already correct (it calibrates directly). Flagging a model
as a reranker still rides the capabilities JSON (supports_rerank).
Negative-tested: rerank detect skips the probe, autopopulates on success, notes
no-separation, reports unreachable on calibration failure, and 400s on an empty
base_url; the non-rerank detect path is unchanged.
Phase 3. Stores reranker calibration per-model on the model definition's
capabilities (rerank_threshold/rerank_scale/rerank_separated; a non-empty
rerank_scale is the "has been calibrated" marker) instead of a single global
threshold, populated automatically when a reranker endpoint is detected.
- ChatSession._bm25_rerank_threshold precedence: the active reranker model's
calibrated rerank_threshold (when separated) wins; calibrated-but-not-
separated -> 0 (no floor); else the global tools.rerank_bm25_threshold
fallback. Reads the raw caps dict, in-memory per turn.
- Detect (admin_detect_model) calibrates a supports_rerank endpoint and
autopopulates the three fields like context window; the create-model UI shows
a verdict chip (calibrated / no-clean-separation / not-calibrated).
- POST /api/admin/model-definitions/{id}/calibrate backs the Re-calibrate
button; calibration runs off the event loop (run_in_executor, bounded by
asyncio.timeout(90)), persists the fields + refreshes the registry, and is
graceful on a down/slow endpoint (never 500). Shared
merge_calibration_into_caps helper used by the endpoint and the CLI.
- turnstone-admin rerank-calibrate is now per-model: --model <alias> required;
--apply writes that model's caps (not the global setting). A no-separation
result records the marker (calibrated, no floor) consistently across CLI and
endpoint.
The serving lesson stays documented: Qwen3-Reranker needs vLLM --chat-template
or its scores are near-random (live-validated 0.6B + 4B: the calibrated floor
came out 0.95 vs 0.33 for the same task -- why per-model calibration exists).
Negative-tested: the floor-precedence branches (calibrated+separated -> per-
model, calibrated+!separated -> 0, uncalibrated -> global, no-alias -> global,
registry-must-not-be-consulted), endpoint persist/refresh + graceful failure,
detect-skips-non-rerankers, the CLI per-model write, and the caps-merge
preserving supports_rerank. Chip states verified via headless Chrome.
Phase 2 of BM25 reranking (follows #627). Makes the rerank_bm25_threshold floor
usable across reranker models and adds tooling to pick it.
- normalize_scores (rerank.py): map a rerank batch into a 0-1 relevance
probability -- sigmoid when any score falls outside [0,1] (logit endpoints
like bge/TEI), identity otherwise (Cohere/Jina/Qwen already 0-1). Applied in
the _bm25_reranker closure AND calibration so the threshold means the same on
every endpoint. Monotonic, so ranking order is unchanged.
- rerank_calibrate.py + `turnstone-admin rerank-calibrate [--apply]`: probe the
endpoint with labelled relevant/irrelevant groups, normalise, and recommend a
recall-biased floor -- or report "no clean separation" (a mis-served/weak
reranker). A warmup loop absorbs a cold endpoint's first-request compile so
calibration doesn't time out. Validated live against Qwen3-Reranker 0.6B and
4B: the calibrated floor differs sharply per model (~0.95 vs ~0.33 for the
same task) -- exactly why per-endpoint calibration exists.
- rerank_config.py: extract resolve_rerank_client_from(config_store, registry);
the alias/url precedence now lives in one place, shared by ChatSession (which
delegates) and the CLI.
- tools.rerank_instruction (config + setting + client): wrap the query as
<Instruct>:/<Query>: for instruction-aware rerankers (Qwen3) on endpoints that
don't apply the model's own chat template. Docs note the critical vLLM serving
detail: Qwen3-Reranker needs --chat-template or its scores are near-random and
reranking hurts retrieval.
Negative-tested: normalize sigmoid/identity branches, closure-normalises-before-
floor, calibration separation/recall-bias/warmup-absorbs-cold-start, the CLI
apply/no-apply/no-separation paths, and instruction query-wrapping through the
real httpx boundary.
Reuse the shipped Cohere/Jina rerank client as an optional post-process on
the BM25 surfaces (tool search, skill search, memory composition) via one
seam: BM25Index gains an injected reranker + a two-stage search (BM25 recall
top-50 -> rerank -> top-k). No new storage.
Gated on a configured endpoint plus tools.rerank_bm25 (default on, matching
rerank_web_search). tools.rerank_bm25_threshold (default 0.0 = off) is a
relevance FLOOR for proactive memory surfacing: BM25 always returns something,
so without a floor every-turn memory injection spends tokens on the top-k of
whatever lexically matched; the reranker score is what makes a meaningful
"inject nothing" gate possible.
Two reranker modes (BM25Index rerank_filters):
- REORDER (reactive tool/skill search): the reranker must never drop results
-> fall back to BM25 order on empty, backfill omitted pool items, so a
misbehaving endpoint can't silently lose tools.
- FILTER (memory, rerank_filters = threshold > 0): a clean empty/short result
is honoured (inject nothing) -- a deliberate divergence from
web_search._rerank_results.
Parse/endpoint failure is a discrete branch from the floor: an empty result
for non-empty input means an unparseable response (a conforming reranker
scores every doc), so the closure raises RerankError and BM25Index falls back
to BM25 order in BOTH modes -- the floor only acts on valid scores.
Also: cap the rerank client timeout at 15s (the per-turn memory path can't
afford tools.timeout's 120s default); move the Reranker alias to rerank.py
(shared, no import cycle); document the endpoint egress in the rerank_bm25
help, the admin Reranker-role description, and docs/tools.md; add
scripts/bench_bm25_rerank.py (manual, needs a live endpoint) to measure
precision@k/MRR lift and recommend a threshold default.
Negative-tested: reorder fallback-on-empty and omitted-item backfill,
filter-mode honor-empty, singleton-still-floored, the parse-fail RerankError
raise, the >= floor boundary, and pool-position-to-doc-index mapping -- each
guard reverted to confirm its test fails, then restored.
list()-materialize the reranker's output inside the guarded block so a
None / non-iterable / lazily-raising reranker falls back to native order
instead of raising out of web_search, and reject bool indices (an int
subclass) the same way _parse_hits already does.
Drop leftover web_fetch references from the rerank settings and Reranker
role help (reranking is wired into web_search only), and document both
endpoint paths (tools.reranker_alias and tools.rerank_url).
Reranking is delegated to an external Cohere/Jina-compatible /rerank endpoint
(self-hosted vLLM/TEI/llama.cpp, or hosted Cohere/Jina/Voyage); Turnstone runs
no reranker model itself. Disabled until an endpoint is configured.
- core/rerank.py: CohereJinaRerankClient (tolerant of results-wrapped and
bare-list responses) + resolver.
- web_search: rerank the SearxNG result pool by query relevance before top-k,
with a native-order fallback on error; answers/infoboxes untouched.
- Reranker as a model definition: add a model with the supports_rerank
capability and pick it under Models -> Roles -> Reranker
(tools.reranker_alias); takes precedence over the tools.rerank_url settings.
Settings: tools.rerank_url/model/api_key, tools.rerank_web_search,
tools.reranker_alias. Docs: docs/tools.md, turnstone.example.toml.
(web_fetch reranking was evaluated and dropped: for single-document chunk
selection it did not reliably beat head-truncation. Reranking is reserved for
multi-item ranking.)
`man` and `math` duplicated capabilities already reachable through
`bash`; `plan_agent` is better expressed as a `task_agent` running a
planning skill, and carried a large amount of special-case machinery
(plan-review gate, refinement loop, per-kind model routing). Removing
all three shrinks the tool surface and cuts per-call token cost.
Also removed, as dead-once-the-tools-are-gone:
- the `math` sandbox executor (`turnstone.core.sandbox`) and its
`[sandbox]` extra; the eval analyst now runs bash-only
- the read-only `AGENT_TOOLS` sub-agent tool set and the `agent`
tool-metadata key (`task_agent`/`TASK_AGENT_TOOLS` retained)
- the plan-review protocol end to end: the `on_plan_review` UI hook,
`resolve_plan`, `POST /v1/api/plan` + `POST /v1/api/route/plan`,
the `plan_review`/`plan_resolved` SSE events, and their Python SDK /
TypeScript SDK / OpenAPI / frontend / Discord+Slack bindings
- the `model.plan_alias` / `model.plan_effort` settings and the
registry `plan_model` / `plan_effort` routing fields
TOOLS 31->28, TASK_AGENT_TOOLS 13->11; COORDINATOR_TOOLS unchanged.
BREAKING CHANGE: removes the `man`, `math`, `plan_agent` tools, the
plan-review SSE/HTTP/SDK surface, and the plan_* model-routing settings
from the experimental 1.6 line.
Rename the web_search tool's `topic` parameter to `category` and expand the
enum to general/news/it/science, mapped to SearxNG `categories=`. The model can
now target the right corpus per query (e.g. `it` for code, `science` for
papers) — useful when generic engines rate-limit. The Tavily-era `finance`
topic (no SearxNG equivalent) is dropped. Threaded consistently through
_prepare_web_search / _exec_web_search / both search clients.
BREAKING: the web_search `topic` argument is now `category`.
The SearxNG change reworded _tool_write_compose's duplicate-skip return and
dropped the "identical content" phrase that test_identical_content_skipped
asserts on, failing CI (which then fail-fast-cancelled the parallel matrix
jobs). Restore the phrase (now covering all three bundled files) and assert
the searxng/settings.yml extraction in test_writes_compose_file.
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single
self-hosted SearxNG service bundled into the docker-compose stacks.
Core:
- New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to
(backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard
unchanged. _resolve_search_client follows storage -> toml -> env -> default
precedence (explicit "" disables, via ConfigStore.stored_keys()).
- Drop the Tavily-era topic=finance (no SearxNG category); topic is now
general/news.
Settings/config:
- Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key.
- Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines,
with get_searxng_url/get_searxng_engines.
Compose + bundled config:
- Internal-only searxng service (no published API port, :ro config, /healthz
healthcheck, persistent searxng-cache volume) in both stacks; bundle
turnstone/deploy/searxng/settings.yml (JSON output on, limiter off).
- Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in).
- bootstrap extractor + wheel packaging updated.
Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the
lxml/h2/brotli transitives).
Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG;
docs/docker.md carries the AGPL-3.0 §13 operator note.
BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg";
tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships
in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance.
Closes#545
run.sh autodetects Ubuntu/Debian, Fedora/RHEL, Arch, and WSL; ensures git and
Docker; clones, builds, picks free ports (Caddy prefers 443, Postgres 5432),
generates a .env with a JWT secret and Postgres password, asks how many nodes to
run, and starts the stack. The node count persists via an auto-loaded
compose.override.yaml, so a later plain `docker compose up -d` keeps it; a fresh
clone with no override still starts all 10. Ignores the generated override.
Lead with privacy, local-first, and no-telemetry; demote governance to an
optional team-controls line. Fix the dashboard URL (Caddy on 8443), add the
one-line installer, and link the Discord community.
The bundled production compose mounts ./Caddyfile, so write_compose has to write
it alongside compose.yaml or `docker compose up` fails to start Caddy. Update
the wizard's system prompt for the new model (no profiles, Caddy-fronted
dashboard, current ports).
Without a Discord/Slack token the channel gateway exited, which crash-loops
under `restart: unless-stopped`. Run the HTTP server and service heartbeat with
zero adapters instead — registered and idle — until a token is set. The Slack
token-pair mismatch stays a hard error.
A node refused to start with no model configured and no LLM reachable, so a
fresh cluster couldn't come up to be configured. load_model_registry now
accepts allow_empty and returns an empty registry; ModelRegistry permits the
empty state (default unset); the server passes allow_empty so a node registers
and shows in the console, then picks up models added in the admin UI live. The
CLI keeps failing fast — a REPL with no model is unusable.
Migrations run on every node at boot, and one migration rebuilds an index with
CREATE INDEX CONCURRENTLY, which can't run in a transaction and waits for all
concurrent transactions to drain. The advisory lock that serialises migrations
was held inside an open transaction, so the lock-holder's own idle-in-
transaction connection deadlocked the concurrent build when several nodes
started together. Acquire the lock on an AUTOCOMMIT connection and poll
pg_try_advisory_lock so no waiter pins a snapshot. Adds a Postgres concurrency
regression test (skipped on SQLite).
`docker compose up` from a clone builds one image and brings up the whole stack
— PostgreSQL, console, Caddy, channel, and 10 server nodes — sharing one
Postgres so the console discovers every node. The dashboard is reachable only
through Caddy (HTTP/2 avoids the browser's 6-connection cap on the dashboard's
SSE streams); the console's plain-HTTP port is no longer published. Postgres
binds 127.0.0.1 so a bare-metal turnstone-server can join the cluster — the
bare-metal overlay is folded in and removed. Insecure dev defaults keep it
zero-config; the bundled production stack mirrors the shape but pulls ghcr
images and requires real secrets.
Move the Caddyfile under turnstone/deploy so it ships in the wheel; update docs,
QUICKSTART, and .env.example to match.
The guard tests were appended via heredoc, bypassing the editor's
auto-format; ruff-format collapses one wrapped call onto a single line.
No behaviour change.
The intent-validation judge runs before the approval gate resolves, and
Smart Approvals (judge.smart_approvals) parks approve_tools on the async LLM
verdict for up to judge.timeout — so the tool-call card never reached the UI
until the judge had ruled. An operator could not see a committed call, let
alone Stop it, during that window.
approve_tools now emits a tool_pending event carrying the serialized batch at
the top of the gate, before the tool-policy lookup, the verdict wait, and the
human prompt. It is a UI paint only — no persistence, audit, or verdict
bookkeeping — so it cannot perturb the gate's accounting. The authoritative
tool_info / approve_request / tool_result events that follow upgrade the same
construct in place, keyed by call_id, and the Last-Event-ID replay slice
reconstructs it on reconnect. A ToolPendingEvent joins the SDK registry.
Coordinator: appendToolBatch was already idempotent on call_ids; the new
handler reuses the --running placeholder it already upgrades, with an
"Evaluating" kicker that swaps to "Running" on the auto-approve upgrade.
Interactive: showInlineToolBlock was create-only, so a second card would
duplicate. Added announceToolBlock + _takeAnnouncedBlock to reuse the
announced shell (matched on its call_id set) instead. The announced rail is
dashed amber and must out-specify the .msg.ts-approval--inline cyan-hold
(specificity 0,2,0) — at 0,1,0 it rendered cyan, indistinguishable from a
normal card — so the announced card is the one visually distinct surface in
the stream.
Screen-reader parity: the early paint announces politely through dedicated
off-screen aria-live regions on both surfaces (the messages log is
aria-live=off mid-stream, so the appended shell alone is inaudible), and the
announced shell carries aria-busy until the upgrade clears it. Polite, not
assertive — the human gate keeps its assertive announcement.
Tests cover the gate ordering (tool_pending precedes tool_info and the Smart
Approvals gate) plus string-guards on both UIs' wiring, the announced-rail
specificity, and the screen-reader regions.
risk_level is server-supplied and was interpolated straight into className
and data-risk at three sites — updateVerdictBadge, _buildOutputWarningEl,
renderVerdictBadge — as `risk_level || "medium"`. Whitespace, a stray case,
or a future relaxed-validation value would pass into the class string and
silently break the selectors that updateVerdictBadge, toggleVerdictDetail,
and the d-key handler rely on. It is not an XSS vector (className assignment
is text-typed), but a broken selector is a real failure.
Funnel all three through a normalizeRiskLevel() chokepoint backed by a
{low, medium, high, critical} allowlist; unknown or blank falls back to the
neutral "medium" default. Pre-existing; surfaced while rendering verdict
badges from the early-paint path.
The floor blocks only explicit heuristic deny/critical verdicts — it is
not a general "never lower the heuristic" rule. The heuristic default for
an unmatched tool is `review`, and letting a confident LLM `approve`
upgrade a `review` is the feature's purpose. Matches the implementation
and addresses PR review feedback.
Opt-in judge.smart_approvals (default off): when the intent-validation
LLM judge returns a high-confidence "approve" verdict, the tool batch is
approved automatically with no operator prompt. review/deny recommendations,
low confidence, judge errors (llm_fallback), and a deterministic heuristic
deny/critical finding all still require a human. Requires judge.enabled.
- Batch-atomic: a parallel tool batch auto-approves only if every call
qualifies; one non-qualifying call holds the whole batch for a human.
- Gate: tier==llm + recommendation==approve + confidence >=
judge.confidence_threshold (default raised 0.7 -> 0.95), with a floor
that never clears an explicit heuristic deny/critical verdict.
- approve_tools waits for the async LLM verdicts, finalises the audit
trail (AutoApproveReason.smart_approval), and re-emits verdicts after
the card so the live chip updates; the auto-approved row renders the
LLM verdict rather than the cautious heuristic carry-over.
- judge: always deliver exactly one verdict per call (fallback on error);
reject non-finite confidence so NaN can't clear the bar.
- Drop verdicts from a superseded judge generation so a reused call_id
from a prior turn's still-running daemon can't satisfy the gate's wait.
Config plumbed through the server/console/CLI builders and the live
_judge_cfg; admin Judge tab renders the toggle. Docs + example config
updated. ~35 tests covering the gate matrix, batch-atomicity, the
heuristic floor, audit stamping, the streaming re-emit, NaN/duplicate-id
defenses, and the cross-turn generation guard.
- Surface the degraded state in each node item's aria-label (only
unreachable was included), so screen readers announce it alongside
the visible DEGRADED word.
- Give the trigger an initial aria-label="Nodes" so it isn't an unnamed
control before the first snapshot render populates it; renderNodePicker
overwrites it with the live count/version once data arrives.
The trigger only showed a border on hover/open, so at rest it read as
plain text rather than a control. Add a persistent recessed box (subtle
fill + border) so it's visibly clickable, and brighten the border to
accent on hover and when open.
The always-visible NODES table dominated the coordinator-first landing
page for information most users glance at rarely. Replace it with a
compact node picker in the cluster status bar: the rightmost segment
shows "N nodes" + the cluster version (or a DRIFT chip on mixed
versions), and clicking it opens a dropdown of every compute node with
its live workstream count. Selecting a node navigates to /node/{id}/ —
the same destination the table rows linked to.
The picker reads the same /v1/api/cluster/snapshot + SSE data the table
did (via the retained buildNodeInfoFromSnapshot), so no backend change
was needed; the table was a pure client-side render. The node-grouping
/ prefix-collapsing JS and all the table CSS are removed.
Accessibility / design:
- Status is encoded by shape and colour (round = healthy, diamond =
degraded, square = unreachable), mirroring the .csb-state-dot
vocabulary, plus a spelled-out DEGRADED/DOWN word — colour alone
fails at 7px for color-blind users.
- DRIFT renders as a solid amber chip (dark text on fill) so it reads
as a real alert rather than yellow-on-yellow.
- role="menu"/menuitem (navigation, not selection), aria-haspopup,
aria-expanded; Escape / outside-click / Arrow / Home / End handling;
full node id surfaced via title when the name ellipsizes; menu height
clamped to the viewport so a long list never touches the top edge.
tests/test_console.py: assert the picker markup is served and the old
table markup (#view-overview / #node-table) stays gone.
Address PR review threads:
- _on_renewed now wraps the client-context hot-reload in try/except, so a
load_cert_chain failure can't abort the renewal callback before the
frontend-bundle update (matching the node-side reload hook).
- The startup gc_expired_certs() sweep is contained like the periodic one,
so a malformed legacy cert row can't abort the TLS block and skip the
proxy/collector mTLS client setup that follows it.
Enabling mTLS broke the cluster in three layered ways:
- Service certs were keyed on socket.gethostname() (the container ID) and
never carried the advertised service name as a SAN, so every collector and
routing-proxy handshake failed the hostname check. build_cert_hostnames()
now puts the advertised host first: it becomes the cert's primary domain
(hence a SAN) and a stable store key that survives container recreation.
- lacme's RenewalManager renews everything in the store; with the store shared
cluster-wide, every node renewed every other node's (and every dead
container's) cert — an N×M renewal storm. _SingleDomainStore scopes each
node's sweep to its own cert, and the console adds a periodic GC for the
certs of long-departed nodes.
- uvicorn loads its cert once at boot and never reloads, so renewed certs
never reached the listener and the served cert expired mid-process.
swap_context_cert() hot-swaps renewed material into the live SSL context
(server listener and console client context) via load_cert_chain.
Observability and browser access:
- The collector logged connection/TLS failures at DEBUG, so a persistent
mTLS-verify failure was invisible. It now logs the first failure per node
(reachable->unreachable) at WARNING and stays at DEBUG on retries.
- The console serves plain HTTP (it is the ACME bootstrap endpoint) and no
longer rewrites its advertised URL to https://. Browser->console TLS is
terminated by a reverse proxy: the cluster profile gains a caddy service
(browser h2/HTTPS -> caddy -> console h1.1/HTTP) plus browser-TLS docs.
Tests: tests/test_tls_san_renewal.py, tests/test_collector_reachability.py.
* feat(audio): voice I/O — speech-to-text + text-to-speech via model roles
Browser voice input/output over the OpenAI audio wire protocol, selected
through the existing model-roles system so the same code path serves OpenAI,
vLLM/vLLM-Omni, or any compatible backend — pure registry config, no new
in-process deps. Anthropic has no audio API, so it is capability-gated out of
the audio roles while remaining valid as the agent model.
Backend
- core/audio.py: role resolution + capability gating + transcribe()/synthesize()
over a registry-resolved client. Typed AudioUnavailableError (503) /
AudioBackendError (502 — body masked, SDK detail logged). Optional STT prompt.
- Endpoints POST /v1/api/workstreams/{ws_id}/speech-to-text and POST /v1/api/tts,
registered in v1_routes, write-scoped (direct + proxied), offloaded with
asyncio.to_thread. Silence -> 422; configured-but-failed backend -> masked 502.
- Model roles: audio.stt_model_alias / audio.tts_model_alias / audio.tts_voice /
audio.stt_prompt settings; Models -> Roles entries (capability-gated dropdowns,
"(disabled — voice off)" when unset). /v1/api/models exposes resolved
stt_default_alias / tts_default_alias + per-model capabilities.
- Capabilities: supports_transcription / supports_speech_synthesis on
ModelCapabilities; current OpenAI audio lineup (whisper-1, gpt-4o[-mini]-
transcribe, tts-1[-hd], gpt-4o-mini-tts) registered as known models, with a
name-inference backstop for local/openai-compatible aliases.
Frontend (interactive UI)
- Mic dictation (record -> transcribe -> fill composer for review) and
per-message playback, shown only when the role is configured.
- CSS-mask icon set, aria-pressed + live-region announcements, recording timer,
reduced-motion cue, error-typed toasts + persistent denial, mic disabled while
busy, code/math stripped before TTS.
Tests: new test_audio.py plus STT/TTS endpoint, settings, openapi, available-
models, and OpenAI-lineup capability coverage. ruff + mypy + node --check clean.
* fix(audio): use const for AUDIO_MODEL_HINTS (var-sweep invariant)