mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-25 13:24:46 -06:00
ecef600c0b86096b0d3d9aa0f532efea2d2fceca
74 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8f415f9c68 |
docs(judge): say llm_fallback, not 'heuristic fallback', for cancelled items
Review feedback: the cancel-path docstrings described undone items as degrading to 'heuristic fallback verdicts', but the emitted and persisted tier is llm_fallback (heuristic content relabeled). Aligned all eight occurrences — including the pre-existing _deliver_fallbacks docstring — so docs, logs, and audit rows use one vocabulary. |
||
|
|
fb282fab73 |
fix(judge): honor the cancel_on_approval=False run-to-completion contract
judge.cancel_on_approval=False (the default) promises the daemon evaluates every tool call to completion so all verdicts are available for later review. Two sites conspired to break that: the approval gate's finally set the cancel event unconditionally the moment a decision landed, and _evaluate_single's poll loop honors the event regardless of config — so every item the sequential judge hadn't reached degraded to a heuristic llm_fallback row. On a 22-call parallel batch, approving after the third verdict silently downgraded the other 19; the elaborate late-verdict machinery in on_intent_verdict was effectively dead code. Make the event a pure abort signal whose firing policy lives with the caller: the gate fires it only when cancel_on_approval is enabled, while generation supersede (next batch) and close() keep firing it unconditionally, bounding a stale daemon to one batch of real work. _run_judge drops its own config second-guessing — a fired event always fast-forwards the remainder to fallbacks (every call still gets exactly one verdict), and the fallback reason no longer claims 'user approval' for supersede/close aborts. |
||
|
|
effdb8f365 |
fix(judge): persist superseded late verdicts for the audit trail
ChatSession._on_verdict guards on judge-generation identity so a stale verdict can't ride a reused call_id into the Smart-Approvals cache — but it dropped those verdicts entirely, before persistence. Every ruling the sequential judge delivered after the next turn began left intent_verdicts claiming the judge never answered. Route superseded verdicts to a new persist-only hook (SessionUIBase.on_superseded_intent_verdict): the row lands with user_decision="superseded" while every live surface stays untouched (no SSE, no replay cache, no pending-decision park). The hook is duck-typed; display-only UIs (CLI/eval) don't define it and keep the plain drop. upsert_intent_verdict already excludes user_decision from its on-conflict SET, so a superseded fallback upgrading its heuristic row in place cannot clobber a decision already stamped there. |
||
|
|
4927efe942 |
refactor: pre-push review — content-addressed upload buffer, GC dedup, security hardening
- Buffer (attachment_buffer.py): content-address staged bytes once and track the
per-(ws_id,user_id) references to them, so identical bytes staged from two tabs
dedupe to one copy yet neither scope's send can drop the other's pending upload
(the prior hash-only key let one overwrite the other). Single lock; add a public
clear() that replaces test reaches into the private store.
- GC: lift the byte-identical _release_attachment_refs out of both backends into one
dialect-agnostic storage/_utils.release_attachment_refs with a portable searched-
CASE single-query decrement (was one UPDATE per id in a Python loop).
- _format_messages_for_summary: mark by-reference vision results
({type:image, attachment_id}) as [image], not just inline image_url.
- Security: escape_like() the attachment_referenced_in_ws LIKE needle on both
backends; secrets.compare_digest for the output-guard operator-fence leak check.
|
||
|
|
164f74dead |
feat(operator-context): deliver structured per-kind meta to the UI
Operator-context system turns (watch results, output-guard findings, idle children, user interjections) carried their kind (_source) and a flattened text content, but the structured per-kind fields were dropped at every persist/deliver boundary — so the UI rendered every kind as one generic operator bubble and the structured watch-result card was lost. Wire the structured meta through as the single source of truth: - Storage: new conversations.meta JSON column (migration 060); threaded through save_message/save_messages_bulk (facade + protocol + both backends) and rehydrated in reconstruct_turns onto Turn.meta.extra["source_meta"]. - Canonical: make_system_turn carries meta as one _source_meta dict; turn_from_dict/turn_to_dict bridge it to/from Turn.meta.extra. - Live + history: widen on_system_turn(content, source, meta) across all impls + the SSE payload; surface _source_meta -> meta in the /history projection. SDK HistoryEvent docs note the field. - Producers derive both the model-facing content text AND the card from one meta dict, so they cannot drift: render_output_guard_text, build_watch_ reminder carrying output, idle_children and user_interjection metadata. - Frontend: addSystemContext / renderSystemTurn dispatch by source to the watch-result, guard-finding, idle-children, and queued-message cards in both the interactive and coordinator panes; every untrusted field renders via textContent. The meta is a leading-underscore key, stripped before the wire (sanitize_ messages and the native mid-conversation path copy only role+content), so the per-provider wire payloads stay byte-identical. Additive column, no backfill: operator turns predating it reload as plain text bubbles. |
||
|
|
964a390e5e |
feat(attachments): resolve AttachmentRef at the translator; drop RawContentBlock
The by-reference content lane now materializes at the provider translator (the
C layer), not in the session. Each create_streaming / create_completion takes
a resolve_attachments callback and runs materialize_attachments() up front,
expanding {type:kind, attachment_id} placeholders to inline data-URI / document
parts by a content-addressed point-lookup the session hands down
(_resolve_attachments). _full_messages emits placeholders; the dict bridge
carries only placeholders.
RawContentBlock is removed — ContentBlock = TextBlock | AttachmentRef. A
resolved inline part is terminal (the wire payload / display output) and never
re-enters the canonical path, so turn_from_dict drops a stray inline image_url
rather than carrying bytes. resolve_attachment_parts / materialize_attachments
operate on the dict projection. Tool vision output rides by reference too
(_tool_content_by_reference): the turn carries placeholders, the bytes persist
content-addressed. The per-turn token estimate counts a by-ref image as one
fixed image budget; the document char budget lands at send (on resolution).
Wire harness byte-identical (the multipart fixture is a placeholder + a matching
resolver); full non-live suite green (7136).
|
||
|
|
dc88060b79 |
refactor(core): session.messages is the canonical Turn trajectory
ChatSession.messages flips from list[dict] to list[Turn] — the in-memory canonical trajectory. Reads migrate to typed fields (turn.role, turn.text, turn.tool_calls); appends and assignments go through turn_from_dict / turns_from_dicts; the fork bulk-save and retry's multipart check read via turn_to_dict. _full_messages lowers Turns→dicts at the wire boundary — the fold/repair and provider translators still consume dicts until the next slice. The token-accounting helpers accept a dict or a Turn. Non-session consumers migrate too: coordinator_idle_observer and eval to typed fields (mypy-enumerated), and server's last-assistant extractor via turn_to_dict (an Any-typed call site mypy could not flag). An all-text multipart content list (the unreadable-attachment placeholder path) now round-trips faithfully through the adapter (single text block → str, multiple → list). Tests that inspected session.messages as dicts read it through the dicts_from_turns / turn_to_dict bridge; those that built it pass dicts through turns_from_dicts / turn_from_dict. Byte-identical wire harness; full non-live suite green (7130). |
||
|
|
f3c96e6493 |
feat(skills): make skill hints first-class system turns; drop escape_wrapper_tags
_skill_hint spliced its guidance into the tool result as a bare <system-reminder>
block — but the operator declaration now tells the model to treat bare markers
as untrusted, silently demoting the hint. Make the hint first-class instead:
- _skill_hint returns the tool result verbatim and queues the guidance via
_queue_tool_advisory("skill_hint", ...); _collect_advisories drains it into a
{role:system, _source:"skill_hint"} turn after the clean result — folded in
the trusted nonce fence for non-native models, inline for native. (Queuing
no-ops mid-wake, like the other tool-channel advisories.)
- skill_hint added to SYSTEM_TURN_SOURCES (an advisory-producer source).
- escape_wrapper_tags removed outright: it was the last consumer, and its job
(defang a marker next to the bare block) is now covered at fold time by
_neutralize_host. The result message rides through verbatim. This also
collapses the two-escaping-mechanism confusion the review flagged.
Tests assert the clean result + the queued/drained hint, plus wake suppression.
|
||
|
|
8513016503 |
refactor(storage): drop the dead _reminders column instead of carrying it
Operator context moved to first-class system turns, leaving _reminders written by nothing and read by nothing. Nulling it (the prior 060 step) left a writable dead column — a foot-gun inviting accidental reuse. Drop it outright and remove every reference in one shot so there is no half-alive state: - migration 060: replace the wholesale null with batch_alter_table drop_column (per migration 027); downgrade re-adds the empty column to match the 059 schema (the envelope un-wrap stays irreversible). - _schema.py: remove the column. - _sqlite / _postgresql: drop the reminders save param, the INSERT/bulk values, and both SELECT columns. - reconstruct_messages: the row tuple is now 8/9-tuple (event_id shifts from index 9 to 8); _utils + the _row test helper updated. - _protocol / memory save_message: drop the reminders param + docstrings. - tests: replace the reminders-roundtrip tests with a _source-only file and a 060 drop-column assertion; remove the obsolete legacy-reminders wire test. No production caller passed reminders=, and the SELECT no longer reads the column, so an un-migrated DB simply ignores any residual values. |
||
|
|
99ba82e8ec |
fix(session): operator-turn wire correctness — framing, empty turns, leading system
Phase-2 follow-ups to the mid-conversation-system consolidation: - user_interjection framing (known #2): a queued message that drains mid-turn is re-framed via render_user_interjection ("The user sent … User message: …") so the user's words keep USER authority, not operator authority — the regression mattered most on the native path, where the turn enters as a real role=system message. Empty/whitespace interjections (e.g. a bare "!!!") are dropped (bug-2). - empty-content user turns dropped at the wire boundary after the fold (known #3): the wake pipeline's synthetic empty send("") leaves an empty user turn on the native path (the nudge stays inline); an empty user message is invalid on every provider. The drop runs after the fold so the fold-path wake turn, which the nudge fills, survives. - leading-system guard (_anthropic): a turn that converts to nothing no longer lets a system message become messages[0] (the API requires messages[0]=user). Newly reachable now that the empty-turn drop can expose it on a fresh-session native wake. - refresh stale .msg.watch-result comments (the card was removed) to describe the current operator-bubble rendering. |
||
|
|
c6b2288302 |
feat(session): consolidate operator-context into first-class system turns
Replace the two operator-context hacks (the <tool_output>/<system-reminder> content envelope and the transient _reminders side-channel) with one persistent {role: system, _source} trajectory turn. Adds supports_mid_conversation_system (claude-opus-4-8): native models take the turn inline; all others fold it into the preceding turn as a nonce-delimited <system-reminder> block declared in the system prompt as the sole trusted marker. Producers (advisories, metacog nudges, user interjections, idle/watch) emit system turns; the envelope/_reminders machinery, escaping round-trip, replay parser, and reminder SSE events are removed. Eager 060 migration un-wraps legacy envelopes. Net -1662 lines.
Known follow-ups from review (unfixed here): (1) the 060 un-wrap heuristic can irreversibly mis-rewrite bare tool rows that resemble the envelope, so do not run the migration until it is tightened; (2) user_interjection turns lost the user-framing/priority preamble (a regression, and a native-path authority-framing concern); (3) native-path wake nudge can emit empty user content.
|
||
|
|
24d75a690b |
fix(review): address review findings on the rerank/memory stack
- rerank_config.py: the runtime instruction fallback had a dead tail (`get_rerank_instruction() or str(cs.get(...))` -- the cs.get term can only return the registry default ""), via a stored_keys() branch that also diverged from the calibrate CLI / endpoint. Collapse to the sibling idiom (`cs.get(...) or get_rerank_instruction()`) so the instruction used at calibration time matches the one used at runtime. Correct the module docstring: ChatSession is the sole caller (the CLI/endpoint share only the instruction precedence, not this function). - session.py: the deferred first-turn memory recompose was gated on the flag alone, so a synthetic wake send (empty user content -> flag stays False) re-ran the full compose on every wake before the first real turn. Gate on a non-empty query too, so wakes don't re-pay it and the real turn still fires exactly once (+ test). Accepted as-is: the __init__ compose (kept so system_messages/_agent_system_messages are valid for early readers; one cheap extra compose per fresh session) and the orphaned tools.rerank_* config rows (inert -- no read path, never listed or redacted; a purge migration would collide with the 060 in flight on another branch). |
||
|
|
47d0b6c3c6 |
fix(memory): defer system-message compose to the first user turn so memory selection has a query
Proactive memory selection scores candidates against the recent-user-message query (extract_recent_context), but a fresh session composes the system prefix once in __init__ while self.messages is still empty. That empty query takes the no-context path: _select_memory_candidates returns recency order and score_memories returns memories[:k] verbatim -- the 5 most recently UPDATED memories, with BM25 and the reranker never invoked. send() never recomposes, so those recency-only memories are what the model sees for the whole session (until an unrelated event -- skill / MCP / model refresh / resume / memory write / command -- happens to rebuild the prefix). Net effect: the injected memories are unrelated to the actual question. Fix: defer the memory-bearing compose to the first real user turn. - Track _system_composed_with_context, set once extract_recent_context is non-empty in _init_system_messages. - send() recomposes once, right after _append_user_turn, while the flag is still False -- so the opening turn's memory block is selected (and reranked) against the real message. The flag then stays True, so the prefix is composed once and stays cache-stable exactly as before (no per-turn prompt-cache churn). This is the targeted fix; per-turn memory refresh (so later topic shifts also re-rank) is the larger tail-injection redesign tracked on another branch. Two adjacent gaps are left as-is for now: the reranker/BM25 only see content[:200], and build_memory_context flat-truncates each memory to 500 chars (max_content is the save cap, not an injection budget). Tests: flag is False on a fresh session and after a whitespace-only wake turn, flips True on a real query; send() runs the deferred recompose with the user message in the query. |
||
|
|
110d44b07e |
refactor(tools): remove man, math, and plan_agent built-in tools
`man` and `math` duplicated capabilities already reachable through `bash`; `plan_agent` is better expressed as a `task_agent` running a planning skill, and carried a large amount of special-case machinery (plan-review gate, refinement loop, per-kind model routing). Removing all three shrinks the tool surface and cuts per-call token cost. Also removed, as dead-once-the-tools-are-gone: - the `math` sandbox executor (`turnstone.core.sandbox`) and its `[sandbox]` extra; the eval analyst now runs bash-only - the read-only `AGENT_TOOLS` sub-agent tool set and the `agent` tool-metadata key (`task_agent`/`TASK_AGENT_TOOLS` retained) - the plan-review protocol end to end: the `on_plan_review` UI hook, `resolve_plan`, `POST /v1/api/plan` + `POST /v1/api/route/plan`, the `plan_review`/`plan_resolved` SSE events, and their Python SDK / TypeScript SDK / OpenAPI / frontend / Discord+Slack bindings - the `model.plan_alias` / `model.plan_effort` settings and the registry `plan_model` / `plan_effort` routing fields TOOLS 31->28, TASK_AGENT_TOOLS 13->11; COORDINATOR_TOOLS unchanged. BREAKING CHANGE: removes the `man`, `math`, `plan_agent` tools, the plan-review SSE/HTTP/SDK surface, and the plan_* model-routing settings from the experimental 1.6 line. |
||
|
|
948e413f66 |
feat(judge): add Smart Approvals (auto-approve trusted judge verdicts)
Opt-in judge.smart_approvals (default off): when the intent-validation LLM judge returns a high-confidence "approve" verdict, the tool batch is approved automatically with no operator prompt. review/deny recommendations, low confidence, judge errors (llm_fallback), and a deterministic heuristic deny/critical finding all still require a human. Requires judge.enabled. - Batch-atomic: a parallel tool batch auto-approves only if every call qualifies; one non-qualifying call holds the whole batch for a human. - Gate: tier==llm + recommendation==approve + confidence >= judge.confidence_threshold (default raised 0.7 -> 0.95), with a floor that never clears an explicit heuristic deny/critical verdict. - approve_tools waits for the async LLM verdicts, finalises the audit trail (AutoApproveReason.smart_approval), and re-emits verdicts after the card so the live chip updates; the auto-approved row renders the LLM verdict rather than the cautious heuristic carry-over. - judge: always deliver exactly one verdict per call (fallback on error); reject non-finite confidence so NaN can't clear the bar. - Drop verdicts from a superseded judge generation so a reused call_id from a prior turn's still-running daemon can't satisfy the gate's wait. Config plumbed through the server/console/CLI builders and the live _judge_cfg; admin Judge tab renders the toggle. Docs + example config updated. ~35 tests covering the gate matrix, batch-atomicity, the heuristic floor, audit stamping, the streaming re-emit, NaN/duplicate-id defenses, and the cross-turn generation guard. |
||
|
|
bdc1f35f94 |
fix(usage): correct dashboard totals + record auxiliary LLM token spend
The Usage dashboard summary cards read the oldest day bucket (`summary.breakdown[0]`) instead of the window SUM, so every headline (total/prompt/completion/tool-calls/cache) showed a single day's value — e.g. 30-day tool-calls reading lower than 7-day. Read `.summary[0]` and collapse the redundant two-request fetch into one (the response already carried both `summary` and `breakdown`). Only the main streaming loop (`on_status`) recorded `usage_events`. Auxiliary non-streaming calls — title generation, conversation compaction, web-fetch summarization, and plan/task sub-agents — bypassed that path and were never counted, undercounting real consumption by a large factor for agent-heavy workstreams. Add an `on_aux_usage` UI hook (storage row via a shared `_write_usage_row` helper with `on_status`; `WebUI` override feeds Prometheus) and route `_utility_completion` and sub-agent turns through it, attributed to the agent's own model. Judge token spend remains uncounted — deferred to a follow-up. |
||
|
|
30d670338e |
feat(judge): merge output-guard heuristic + LLM judge, annotate findings
Surface the output-guard LLM judge on the inline finding chip and merge it with the regex heuristic instead of one stage winning outright. Merge rule (issue #560, "show, annotated"): - risk_level = max(heuristic, llm); flags = union. The judge can escalate but never lower a heuristic positive — it evaluates adversarial tool output, so defeating it must not erase a deterministic regex finding. Credentials stay heuristic-only and are always redacted. - The judge's own verdict rides along as a dissent-aware annotation (judge_risk / confidence / reasoning / judge_model) on the chip in both the interactive and coordinator UIs, live and on reconnect. One shared merge_guard_display_payload drives both paths so they cannot drift. - The model is shown the merged risk + flags but never the judge's "benign" verdict — a fooled judge must not talk the model out of caution. Fixes a reconnect bug: a judge that ran but failed wrote a risk="none" row that won the replay dedup and hid the heuristic finding (it showed live but vanished on refresh). Failed judges now persist under tier="llm_error", excluded from the display merge; the max-merge also floors the displayed risk at the heuristic level so the chip never vanishes. Also adds a regression test confirming the LLM judge runs on every tool output, not just heuristic-flagged ones. Tests: merge unit tests, storage-backed replay regression, live/replay wire-shape parity, SDK-event drift guard. ruff + mypy clean. |
||
|
|
ee8dc7c1c3 |
refactor(history): project the /history wire shape server-side
Collapse the three hand-synced "raw storage -> render shape" projections into one server-side projection. The projection previously lived in a test-only `_build_history` (SSE-era reference impl), a client-side JS normaliser (`history_normalize.js`, the transitional bridge), and coord's inline `init()` handling -- drifting silently with no parity test. Add `project_history_messages` to `history_decoration.py` and run it as the final step of the `make_history_handler` pipeline (load_messages -> decorate -> extract_reasoning -> project), so `GET /history` emits the canonical render shape directly: flat tool_calls (with verdict / output_assessment), top-level source / reminders / attachments, collapsed multipart content, derived denied / is_error / pending, reasoning, and advisories. Interactive `replayHistory` now consumes the payload verbatim. Close two gaps the JS bridge deferred: - list-content <tool_output> advisory extraction (decorate handles only string content; the projection extracts list-carrier advisories, then joins remaining text parts to the string the renderers require); - orphan->pending marks ONLY the last orphan tool-call turn, so a mid-conversation cancelled tool still renders instead of vanishing. Delete `history_normalize.js` (+ its <script> tag and node test) and the test-only `_build_history` (+ orphaned imports); retarget its direct tests onto the projection helpers. Update the WorkstreamHistoryResponse description and the Web UI Resilience architecture note to the projected shape. Coord's `init()` still reads the raw side-channels; migrating it to the projected shape is the next commit, browser-verified separately. Refs #549. |
||
|
|
3233719856 |
feat(judge): output_guard LLM stage with capability gate (#560 mitigation #1)
Adds a second, LLM-driven stage to the output guard so domain-camouflaged prompt-injection payloads that the regex stage misses (arXiv:2605.22001 — Llama 3.1 8B evades the existing regex set on ~90% of camouflaged prompts) get caught before the tool output lands in the assistant's context. ## Surface * New `OutputGuardJudge` in `turnstone/core/output_guard_judge.py` — synchronous, single-shot LLM call. Inlines the alias-resolution + client-config + JSON-parsing helpers (copied verbatim from `IntentJudge` at `judge.py:917-969` / `1604-1659`) rather than going through a shared module — when `IntentJudge` lifts its own helpers, both copies move together. * JSON-in-content verdict with a 3-strategy parser (direct / markdown fence / balanced braces). `IntentJudge` ships a 4th regex-field fallback; OutputGuardJudge deliberately doesn't, because strategy-4 hits on broken LLM output can extract a "verdict" from the model's reasoning quote that lands in storage looking identical to a clean strategy-1 result. Failure of all three returns `error="unparseable_verdict"` and the heuristic stage stands. * `OutputJudgeVerdict` is a frozen dataclass with: `risk_level` (none/low/medium/high — normalises `critical`→`high` and `info[rmational]`→`low` for IntentJudge-echo safety), `flags: tuple[str, ...]`, `reasoning`, `confidence: float` (0.0-1.0, parsed + clamped from the LLM's self-report; pass-through to audit, no threshold gating), `judge_model`, `latency_ms`, `error`. * Real wall-clock timeout via `ThreadPoolExecutor.shutdown(wait=False, cancel_futures=True)` on the timeout/cancel path — `with ... as ex:` would block return until the worker drained. 1s `cancel_event` poll mirrors `IntentJudge._run_judge` at `judge.py:1117-1118`. * HTTP client lazy-init + reuse for the judge instance's lifetime. Session-side model swap drops the entire judge, dropping the client with it. * Untrusted tool output wrapped in per-call random-nonced `<tool_output_NONCE>...</tool_output_NONCE>` fence. Closing-tag substrings in the raw text are case-insensitively backslash-escaped first (`</tool_output` → `<\/tool_output`) so an attacker can't break out even if they guess the nonce. System prompt classifies the fenced region as UNTRUSTED DATA so directives inside are evaluated as content, not obeyed. * Judge user prompt carries the heuristic verdict (risk + flags + annotations), the tool description (looked up from the session's tools registry), and the tool args (truncated to 500 chars, also classified UNTRUSTED in the system prompt since they may be caller-supplied). Lets the judge defer to the regex on credential leaks and focus on injection signals the regex set misses; also enables output-vs-request plausibility reasoning. ## Session integration * `_evaluate_output(call_id, output, func_name, *, tool_args="")` — heuristic always runs; LLM stage runs when `judge.output_guard_llm` is enabled. When the LLM produces a usable verdict and the heuristic didn't detect credentials, the LLM verdict is acted on; otherwise the heuristic stands. * Credential redaction is a regex-only signal. When `heuristic. sanitized` is non-None, the heuristic owns the acted assessment regardless of what the LLM said — an LLM asked about prompt- injection can correctly label a credential-bearing output as "none" risk for injection, but the secret still needs redaction. * `_batch_evaluate_outputs` runs the per-tool guard concurrently (4-worker pool) when LLM is enabled and there are ≥2 string outputs — collapses N×LLM-latency to ⌈N/4⌉×latency on the common 5-20 tool-calls-per-turn turn. * Per-session `TokenBucket(rate=1.0, burst=60)` caps adversarial LLM-fan-out cost at 60 calls/min/session. * Pre-truncation: the per-tool loop truncates output before the judge sees it, so the judge evaluates exactly what enters the assistant's context (no wasted tokens on text that won't land). * Both heuristic and LLM tier rows persisted to `output_assessments` when the LLM ran (audit completeness); heuristic-only rows skip when matched-clean to keep the table focused. ## Storage Migration 057 extends `output_assessments` with five LLM-tier columns: `tier` (`heuristic` / `llm`, backfilled to `heuristic`), `reasoning`, `judge_model`, `latency_ms`, `confidence`. Tie-break on `(created DESC, tier='llm' first)` so downstream consumers see the acted verdict first when the two rows tie at second resolution. `StorageBackend.record_output_assessment` + sqlite/pg implementations + `SessionUIBase.record_output_assessment` + `SessionUI` protocol + the test stub overrides (cli, eval, 9 test files) all take the new LLM-tier kwargs. ## Config surface Three new judge.* settings in `settings_registry`: * `judge.output_guard_llm` (bool, default False) — capability gate. Default off; operators opt in once a small/fast model is pointed at `output_guard_model`. * `judge.output_guard_model` (str, default "") — alias for the LLM stage. Empty inherits the session model (same fallback shape as `judge.model`). * `judge.output_guard_llm_timeout` (float, default 30.0, min 1.0) — wall-clock budget per call. Both `server.py` and `console/session_factory.py` wire these into the `JudgeConfig` they hand to `ChatSession`. ## Notes * No backwards-compatibility shims — the LLM stage is purely additive. * No reasoning/threshold gating on confidence; it rides as an audit-only signal per maintainer direction. Surface it in the `on_output_warning` dict so live UI / cluster broadcast can sort flagged outputs by judge certainty. * Tests: 392 lines of judge-only coverage (`test_output_guard_judge. py`) + 629 lines of session-integration coverage in `test_session. py`, plus the storage and stub-shape updates. |
||
|
|
bf14fc8b45 |
fix(output_guard): harden against domain-camouflaged injection (#560) (#573)
* fix(output_guard): harden against domain-camouflaged injection (#560) Three layered mitigations against the camouflage attack class described in arXiv:2605.22001 (Pai, May 2026), which demonstrates 90.3% evasion on Llama 3.1 8B and 44.4% on Gemini 2.0 Flash against pattern-based detectors: - Sub-agent synthesis is now scanned by output_guard at the sub-agent boundary in _run_agent, in addition to the existing scan at the parent's tool-result loop. Covers all four return paths (clean exit, truncation, context-limit recovery, turn-limit forced synthesis), closing the cross-workstream summary laundering surface. - Adds pair-of-signals camouflage detection: imperative recommendation phrase combined with either an authority frame ("consistent with our risk framework") or a caps action verb (SELL/BUY/TRANSFER/...). New flag camouflaged_injection at medium risk; deliberately partial — the paper's augmented-detector approach recovers only ~10% on Llama-class models, so this is duct-tape pending a semantic-evaluator follow-up. - Bumps output_guard's wall-clock budget default from 5s to 30s and exposes it as judge.output_guard_budget_seconds in ConfigStore, so the expanded regex set has headroom on large tool outputs. * Fix test_budget_kwarg_is_honored to exercise deadline logic path The test previously passed an empty string which short-circuited evaluate_output() before budget_seconds was used. Now uses a non-empty input and monkeypatches time.monotonic() to deterministically verify the deadline path is exercised. |
||
|
|
471d48abd9 |
feat(skills): unify skill + list_skills into dual-kind action-multiplexed tool
Replaces the legacy `skill` (load + search) and `list_skills` tools with a single `skills(action=...)` tool serving both interactive and coordinator sessions. Stacks on the model.skills.write permission introduced in PR 1. Tool surface - `find`: filter by category/tag/risk_level/enabled_only/limit with optional BM25 query ranking; auto-approved on both kinds; kind-scoped at the storage filter (interactive sees interactive+any, coord sees coordinator+any). - `get`: fetch a single skill including content; cross-kind misses collapse to "not found" so a model can't enumerate the other surface by name-probing. - `load`: activate a skill in the current session (interactive-only; coord sessions get an explicit hint pointing at spawn_workstream). - `create`/`update`/`enable`/`disable`: require approval AND model.skills.write; permission re-checked at exec time to catch a revocation between approval and write. - No `delete` — hard-delete stays admin-UI exclusive; tool description documents the soft-delete-via-disable pattern. Defenses on the write surface - Approval cards surface projected risk_level (scanner re-run against the proposed final state) and warn explicitly when allowed_tools + auto_approve combine (auto-fire-on-load consequence is spelled out, not just shown as raw field values). - Toggle preview surfaces existing risk_level + allowed_tools count so re-enabling a critical-tier skill is never a one-click bypass. - Update path now re-fetches the row at exec to catch a readonly flip between approval and write, filters updates back to the runtime-only set if so, refuses if no fields survive. - Update path rejects empty content (hollow-out via emptying bypassed the soft-delete-via-disable invariant), non-list tags, and empty category — failures are loud rather than silent. - Permission denials audit `skill.write_denied` with actor_source=model so probing the permission state leaves a trail. Audit failures log at error (not warning) — a successful write without a row is the exact gap the trail exists to surface. - `_skill_hint` routes both message and system_reminder through escape_wrapper_tags so caller-controlled values can't close the <system-reminder> envelope and let the model fabricate directives in its own future context. Shared validation - `parse_skill_session_config` lifted from console/server.py to turnstone/core/skill_field_validation.py; both the HTTP admin path and the model-tool path consume it. Single source of truth so field rules can't drift between layers. - `SKILL_RUNTIME_CONFIG_FIELDS` lifted similarly (was duplicated as _SKILL_RUNTIME_CONFIG_FIELDS in server.py and _SKILLS_READONLY_FIELDS on ChatSession). - `notify_on_complete` validator now accepts list input from the JSON schema's `array` type — previously rejected because str() of a list yields Python repr that json.loads then refuses. Performance - Update prepare skips the projected-risk scan when neither content nor allowed_tools is changing (storage re-scans on write authoritatively). Metadata-only updates no longer pay the ~25 regex-pass scan cost. Cleanup - CoordinatorClient.list_skills deleted (-91 lines); model-tool path talks to storage directly via list_skills_filtered. - Roles admin UI gains a Model section exposing model.skills.write. - tests/test_load_skill.py renamed to tests/test_skills_tool.py and rewritten for the new tool — 48 tests covering registration, prepare dispatch, permission gating (including TOCTOU-revoked exec deny), audit actor_source on create + disable + permission-denied probe, BM25 ranking, invalid-kind branches, audit-failure swallow, and <system-reminder> envelope injection resistance. |
||
|
|
b4299f8888 |
fix(task_agent): address Copilot feedback on skill parameter
- Put ``skill`` back in the access-denial list in the tool
description with a clarification — TASK_AGENT_TOOLS does not
include the skill tool, so sub-agents cannot switch personas
mid-task. Removing the disclaimer entirely created an ambiguity
the LLM could misread.
- Minimize the skill_data carried on the approval item dict to
``name`` / ``content`` / ``risk_level`` only. ``get_skill_by_name``
returns the full ~30-column prompt_templates row including
``scan_report``, ``installed_by``, ``source_url`` — none of those
flow through ``_exec_task`` / ``_evaluate_intent``, and they
shouldn't ride along any future audit serializer that reads the
approval item shape.
- Regression test for ``skill=""``, whitespace-only, and ``\t\n``
values — pins the documented "empty value is acceptable" contract
at the ``(args.get("skill") or "").strip()`` chokepoint.
|
||
|
|
7d58df0d22 |
feat(task_agent): add optional skill parameter for per-call personas
The task_agent tool now accepts an optional ``skill=<name>`` argument that loads the named skill's content as the sub-agent's persona, substituting the hardcoded "# Task Agent" identity statement. The operating-guidance numbered list (one-shot, tool-use over narration, no follow-up questions) is layered on top of every persona and always applies — those are sub-agent semantics that a persona should ride on top of, not replace. Validation lives in ``_prepare_task`` so the approval surface tells the operator what they're consenting to: the validated skill dict (including content) rides on the item dict from prepare to exec to defeat TOCTOU between consent and execution. An unknown skill returns a clean error item with a hint pointing at ``skill(action='search')``; a disabled skill returns a distinct error so the LLM's recovery path can tell "not found" from "quarantined", mirroring the enabled gate that ``_exec_skill(action='load')`` and skill-search already apply. High and critical skills now surface their risk tier on the approval header (``, risk: critical``) and emit a ``task_agent.high_risk_skill`` warning — same signal ``_load_skills`` emits for session-level skills, so the operator sees the same flag whether the skill is loaded session-wide or per-call. ``_exec_task`` emits a ``task_agent.skill_invoked`` info log on the skill branch for forensic traceability — the approval row captures the choice at consent time, this log captures it at exec time so post-incident search doesn't have to cross-walk approval and exec tables. The ``_evaluate_intent`` func_args projection now includes the skill name — without it, heuristic ``arg_pattern`` rules targeting a risky persona name on ``task_agent`` silently no-op and the audit row loses the choice. Mirrors the long-standing ``spawn_workstream`` projection. |
||
|
|
ee163c0ae4 |
fix(session): prevent LLM bypass of per-role plan/task model overrides
The LLM was passing ``task_agent(model="default")`` (and the same for plan_agent) and routing to whichever backend the auto-created ``default`` alias was attached to at boot — flatspark in the verified case (ws_id 7dde674) — silently bypassing the operator-configured ``model.task_alias`` / ``model.plan_alias`` (gh200). Root fix: - ``load_model_registry`` only synthesises the back-compat ``default`` alias when neither DB nor ``[models.*]`` populate the registry. The shim was only ever meant for single-CLI-model setups; with a multi- model DB it became a phantom routing target aliasing ``LLM_BASE_URL``. - ``_render_agent_tool_descriptions`` filters ``default`` out of the LLM-visible alias list. The English reading of "default" trips the model into picking it explicitly even when the description tells it to omit ``model=`` for the per-role default. Defense-in-depth at the validator chokepoint (``_validate_agent_model_override``): explicit rejection of ``alias == "default"`` (post-strip) with corrective guidance; ``default`` filtered out of the unknown-alias retry list so an LLM probing with a bogus alias can't enumerate it back; the no-alternatives wording is distinguished from the no-registry-configured wording. The render path also always rewrites tool descriptions instead of returning early on filter-empty, so a reload that drops the registry to only ``default`` clears stale alias names left over from a prior render. |
||
|
|
6bdc6cf0bd |
feat(audit): emit memory tool save/update/delete events
Previously only the admin-console DELETE route emitted memory.delete audit rows, so a long-running session whose memory was deleted via the admin UI had no log trail showing what happened — masking out-of-band deletes as apparent tool bugs. The save branch now stamps memory.save (new row) or memory.update (upsert); the delete branch does a lookup-then-delete-by-id pair so the audit can record the resolved memory_id and type. All emissions are best-effort: failures log at debug and swallow so an audit hiccup never breaks the tool call itself. Reads (get/search/list) remain un-audited. |
||
|
|
29b850919f |
feat(sse): refresh-resume for mid-stream page reloads
Refreshing a coordinator or interactive workstream pane while the LLM is mid-stream now restores the partial assistant text + reasoning immediately and flips the composer back to stop-mode, instead of showing nothing until the response completes. Per-turn inflight buffers (`_ws_inflight_content`, `_ws_inflight_reasoning`, `_ws_inflight_seq`) on `SessionUIBase` are kept separate from the existing multi-turn `_ws_turn_content` buffer that drives the dashboard's IDLE-piggyback payload. New `on_turn_start` (top of send-loop, defensive) and `on_turn_committed` (right after `messages.append(assistant_msg)`, primary) lifecycle hooks reset inflight at turn boundaries. The seq counter is monotonic across turns so a long-lived subscriber's `snap_seq` cutoff stays valid for the lifetime of the connection — resetting per-turn would silently drop turn N+1's first M tokens (M = whatever was streamed pre-snapshot in turn N). `snapshot_and_consume_state_payload` also drains inflight at idle/error so cancel and exception paths don't leak stale text. New `register_listener_with_in_progress_snapshot` atomically registers a listener and snapshots the inflight buffers; `make_events_handler` emits a `state_change` event (so the JS busy machine flips to stop-mode) followed by a one-shot `in_progress_snapshot` after the kind-specific replay, then strips the internal `_seq` field from yielded live events while filtering against `snap_seq`. A per-listener shallow `dict` copy in the live drain prevents the multi-tab race where one listener's `del event["_seq"]` would corrupt another listener's filter view. `_synthesize_cancelled_results` now emits synthetic `on_tool_result` events for each cancelled tool so live coord tabs can drop the newly-additive `coord-tool-batch--running` indicator cleanly. The indicator now coexists with `--auto`/`--approved` (applied on `tool_info` and `approval_resolved` approved; removed when every row in the batch has a result), making live tool execution visually parallel to the replay-time orphan rendering. Frontend handlers in `app.js` (interactive) and `coordinator.js` (coord) absorb EventSource auto-reconnect re-replays via a length-based prefix check on the in-progress buffer. New `InProgressSnapshotEvent` + `StateChangeEvent` dataclasses in the Python and TypeScript SDKs with type guards. `_MAX_TURN_CONTENT_CHARS` lifted 256 KiB → 512 KiB (single constant for both buffers — headroom for current commercial models). Regression tests cover race-free composition under concurrent writers, seq-filter dedup invariants, the cross-turn seq monotonic invariant, idle/error inflight drain, synthesized `on_tool_result` on cancel (including UI-hook failure isolation), and the multi-listener shared-dict invariant. |
||
|
|
c2cb6a7ea5 |
fix(replay): apply PR #488 review findings
Four Copilot findings on
|
||
|
|
eca4bb79e4 |
fix(replay): seam 1 splice + storage symmetry for queued user messages
Reverses the seam-2-only design from the prior commits on this branch.
Queued user messages arriving DURING a tool batch (Seam 1) splice into
the last tool result's envelope as ``UserInterjection`` advisories via
``wrap_tool_result``. Messages arriving BETWEEN turns (Seam 2) drain
as a single trailing user row via ``_flush_queued_messages`` with
``user_feedback`` (operator text alongside an approval, e.g. "y, use
full path") folded in as a prefix. Cancel/exception drains (Seam 3)
keep the existing ``_flush_queued_messages()`` call unchanged.
Why all three seams:
* Strict-template providers (Mistral, Llama via vLLM with stock chat
templates) reject role-alternation violations. A literal ``user``
row mid-tool-batch breaks ``assistant(tool_calls) → tool → ... →
assistant``; back-to-back ``user → user`` rows on the wire also fail.
* The seam-2-only design produced back-to-back ``user`` whenever
``user_feedback`` and queued items both fired — bug-1 from the round-1
review. Folding ``user_feedback`` as a prefix to the queue-drain
collapses the two into one row.
* During-batch arrivals couldn't ride seam 2 — the splice was the only
way to deliver same-turn without violating role alternation.
Storage symmetry:
Tool DB rows now store the wrapped ``output`` (envelope + advisories)
unconditionally — ``self.messages[i]['content']`` and
``conversations.content`` match exactly. List-typed output (image /
structured MCP results) uses ``wrap_tool_result(raw_joined_text,
advisories)`` at save time so the persisted string is anchored on
``<tool_output>\n`` for the replay parser. ``TOOL_RESULT_STORAGE_CAP``
is removed entirely; tools are responsible for bounding their own
output, storage faithfully represents in-memory. Removing the cap
also simplifies the parser — no truncated-envelope edge case.
Replay extraction:
``decorate_history_messages`` (REST ``/history``) and ``_build_history``
(SSE replay, resume, rewind, retry, post-load, rename re-replay) both
call the public ``extract_advisories_from_tool_envelope`` helper to
pull the envelope back into structured ``advisories`` for JS replay.
Both string content and list-typed content (image+queued-message
combo) covered. JS renders extracted advisories as normal user
bubbles after the tool block via the shared ``replayAdvisoriesAfterTool``
helper in ``shared_static/utils.js``.
Wrapper-tag escape and provider splice:
``escape_wrapper_tags`` now encodes pre-existing ``&`` first using an
``&`` sentinel so tool output containing literal entity strings
(documentation viewers, code analyzers, web scrapers returning entity-
encoded markup) round-trips correctly. Both encode and decode helpers
short-circuit on absence of ``<`` / ``&``.
``_apply_reminders_for_provider`` detects already-wrapped content
(string body and list text-part) by ``startswith("<tool_output>\n")``
and skips re-escape so existing envelopes survive intact when a tool
message also carries ``_reminders`` (the queued-message + tool-error
co-occurrence case is now common).
``decorate_history_messages`` runs in ``asyncio.to_thread`` to keep
MB-scale string work off the event loop.
Other cleanup:
* ``_collect_advisories`` delegates the queue drain to a named helper
``_drain_queued_messages_to_advisories`` so the swap-and-clear pattern
lives next to ``_flush_queued_messages``'s identical pattern and the
side-effect is documented at the call site.
* Preamble strings + body marker for ``UserInterjection`` round-trip
detection moved to module-level constants in ``tool_advisory.py``;
imported by ``history_decoration.py`` so a producer-side rephrase
can't silently desync the parser.
* ``_send_with_mocks`` ctxmgr extracted in ``test_session.py`` — the
six new send-driven tests share an 8-deep ``patch.object`` block.
* ``replayAdvisoriesAfterTool`` shared helper in
``shared_static/utils.js``; ``app.js`` and ``coordinator.js`` both
invoke it.
* Dead truncation-pill CSS removed (``.tool-output-truncated`` and
``.coord-tool-truncated``); the JS that added these elements went
away with ``TOOL_RESULT_STORAGE_CAP``.
* Tautological tests (``TestBuildHistoryAdvisoryPropagation``)
replaced with production-realistic round-trip tests built from
``wrap_tool_result(...)`` envelopes — REST and SSE-replay surfaces
pinned to the same wire shape; full DB round-trip pinned end-to-end.
Negative-tested:
* Reverting the prefix-merge in ``_flush_queued_messages`` produces
back-to-back ``user`` rows, breaking
``test_user_feedback_and_queued_coexistence_single_row_with_prefix``.
* Reverting the ``extract_advisories_from_tool_envelope`` call in
``_build_history``'s tool branch leaves the envelope verbatim in
wire content, breaking the round-trip tests.
* Reverting the wrapper-detection in ``_apply_reminders_for_provider``
entity-encodes the existing envelope's literal tags, breaking both
the string-content and list-content envelope-preservation tests.
* Reverting the ``wrap_tool_result(raw_text, advisories)`` projection
at the DB save site produces a string starting with the original
raw text, breaking
``test_tool_db_row_round_trips_list_output_with_advisories``.
Tests: 5918 passed, 3 deselected. Lint + format + mypy clean on
touched files.
|
||
|
|
a032e71ff3 |
fix(replay): apply review findings q-2 through q-7
Round-1 ``/review`` apply-pass. Drops stale ``UserInterjection`` references from comments and docstrings that no longer describe the post-PR drain shape, asserts the two-stream invariant in the new queued-message persistence test, and pins the ``content.trim()`` + ``renderAssistantToolBatch`` invariants on coord-side so a future refactor can't silently regress the Qwen3 phantom-card fix or the chronological-order render fix. Deferred: * **bug-1** (back-to-back ``user`` row when ``user_feedback`` from the approval-prompt UI callback coexists with a queued-message drain). Reachable on strict OpenAI-compatible local templates (Anthropic and Anthropic-via-merge-consecutive collapse fine; vLLM-hosted Mistral / Llama enforcing role alternation can reject). The pre-PR splice guarded against this case by riding queued items inside the tool result envelope; that guard is what motivated the original UserInterjection design, so the fix lane needs a deliberate decision rather than a quick patch. Sleeping on it. * **q-1** (delete dead ``UserInterjection`` class + tests). Held for the bug-1 decision — if the chosen fix is to resume the splice for the ``user_feedback``+queue coexistence case, the advisory shape stays load-bearing. Class now carries a docstring note marking it retained-pending-decision so a passing reader doesn't grep for producers and assume it's actually dead. Apply-pass content: * ``q-2``: drop "queued user interjections" from the persistent- advisory parenthetical in ``send``'s tool-result loop comment; rewrite to point at ``_flush_queued_messages`` for the queue path. * ``q-3``: ``__init__`` channel-routing comment loses "and ``UserInterjection``" — only ``GuardAdvisory`` remains. * ``q-4``: ``_queue_tool_advisory`` docstring + the tool-error nudge comment lose the user-interjection mentions; the docstring also now describes the side-channel + ``_apply_reminders_for_provider`` splice path (the actual mechanism). * ``q-5``: ``AttachmentsNotQueueableError`` docstring rewritten to describe the post-PR ``_flush_queued_messages`` flow — the single-combined-turn ``\n\n``-join shape can't carry image / file blocks, and per-item separate user turns would expand the strict- template role-ordering surface that the post-batch drain already balances. * ``q-6``: the new ``test_queued_message_persists_as_user_row_after_tool_batch`` in ``test_session.py`` now asserts ``stream_idx == 2`` so a future regression where the post-batch flush runs but the send-loop short- circuits before the next iteration surfaces in CI rather than manual repro. * ``q-7``: ``test_coordinator_page.py`` gets two new string-grep pins mirroring the existing ``test_app_js.py`` shape — ``content.trim()`` on coord's assistant-replay branch and ``renderAssistantToolBatch`` for the hoisted helper that orders content card before tool batch. ## Test plan - [x] ``ruff check`` clean - [x] ``mypy turnstone/`` clean (189 source files) - [x] Affected test surface (``test_session.py`` + ``test_tool_advisory.py`` + ``test_app_js.py`` + ``test_coordinator_page.py``) — 240 passed |
||
|
|
c11692b327 |
fix(replay): coord render order + blank assistant cards + queued message persistence
Three independent rehydrate / replay regressions reported on long multi-turn conversations after the pull-model wake stack landed. **1. coord history replay rendered tool_calls above the assistant narration that announced them.** In ``coordinator.js``'s loadHistory loop, the ``role === "assistant"`` ``tool_calls`` branch sat above the role switch — every assistant turn with both narration AND tool dispatch produced ``[tool batch][content card]`` in the DOM, even though chronological order is content first. On a parallel fan-out (e.g. four ``close_workstream`` calls in one turn) operators saw the assistant text "Let me close them out and summarize" with NO tool batch between it and the next assistant message — the four-row batch had been rendered above the announcing text and was scrolled out of view. Hoisted the ``tool_calls`` synthesis into a local ``renderAssistantToolBatch(m)``, called from inside the assistant branch AFTER the content card. Live SSE order (text → dispatch → results) now matches replay order. **2. Whitespace-only assistant content rendered as a blank card on replay.** Models with vLLM's ``--reasoning-parser`` (Qwen3 in production) strip ``<think>…</think>`` and emit only the trailing ``"\n\n"`` as ``content`` before a tool call. ``content_parts = ["\n\n"]`` saves ``content = "\n\n"`` to the conversations row. Live the user only sees ``.msg.reasoning`` (the thinking content) — the empty ``.msg.assistant`` card lives next to it but reads as a thin divider. On rehydrate the reasoning bubble is gone (not persisted) and the empty assistant card is the only thing left, surfacing as "blank cards where the assistant message was." Both UIs now check ``content && content.trim()`` before rendering the body — whitespace-only content skips the card entirely instead of showing a phantom row. Live render unchanged. **3. Queued user messages disappeared on reconnect.** PR #474 routed queued user messages into the tool-result envelope via ``UserInterjection`` advisories — same-turn delivery, but no persisted user row. On page reload / cross-tab replay the optimistic ``.msg-queued`` bubble vanished: there was no DB row to rehydrate it. Dropped the ``UserInterjection`` splice in ``_collect_advisories``; the queue drains through ``_flush_queued_messages`` AFTER the tool batch completes instead. Sequence becomes ``assistant(tool_calls) → tool … tool → user(drained)``, which is valid for Mistral and Anthropic strict role validators (the only forbidden shape was user injected mid-batch BEFORE the tool result, which this still avoids). Persists a real user row → bubble survives reconnect, and stays in the session's wire-side context window on the next turn. ## Test plan - [x] ``ruff check`` clean - [x] ``mypy turnstone/`` clean (189 source files) - [x] ``pytest -m "not live"`` — 5798 passed, 3 deselected - [x] Updated ``test_collect_advisories_does_not_drain_queued_messages`` (was pinning the old UserInterjection shape) - [x] Added ``test_queued_message_persists_as_user_row_after_tool_batch`` (drives ``send`` end-to-end with a queued message arriving during the tool batch; asserts the user row lands in self.messages AND hits ``save_message``) - [x] Updated ``test_replay_history_renders_content_before_tool_block`` to tolerate the new ``msg.content && msg.content.trim()`` guard - [ ] Live browser pass on coord (close_workstream parallel fan-out rehydrates with the 4-row batch BETWEEN the announcing assistant text and the summary) and interactive (Qwen3 ``"\n\n"`` rows no longer paint blank cards on reload; queued bubble survives a tab refresh) |
||
|
|
b120ee2fd7 |
fix(session): trim tombstone refs + WHAT-narration in apply-pass comments
Closes round-2 review findings q-1 (minor), q-3 (nit), q-4 (nit), q-5 (nit). * **q-1:** Drop the ``post-migration 050`` clause from the fork-block comment — the apply-pass relocated rather than removed the tombstone-style temporal reference round-1 q-2 was supposed to fix. The bulk-row dict shape and ``_encode_reminders`` are self-explanatory; the WHY is pinned by ``test_fork_preserves_source_and_reminders``. * **q-3:** Replace ``DOES persist now`` framing on the wake-row save comment with a present-tense invariant. The ``now`` implies the reader knows the prior state, same family as the temporal tombstones. * **q-4:** Trim the 12-line WHAT-narration block above the resume-time ``_reminders_delivered = True`` loop to two lines stating the WHY only. The new regression test pins the contract. * **q-5:** Reframe ``test_fork_preserves_source_and_reminders`` docstring as a forward-looking invariant; drop the ``Dropping them was the original bug`` and ``post-migration 050`` fix-narration. Project convention: invariant statements, present tense; don't reference the current task / fix / migration number. |
||
|
|
7e35050b68 |
fix(metacog): cleanup batch — share watch-key constant, sanitize metadata, drop tombstones
Closes round-1 review findings q-2 (minor), q-5 (minor), q-6 (nit), q-7
(nit), sec-1 (nit), perf-4 (nit).
* **q-5:** Export ``_WATCH_REMINDER_OPTIONAL_KEYS`` from
``turnstone/core/watch.py`` and import in the dispatch closure
(session.py) and the replay filter (server.py:_build_history). The
three-place duplication of the literal tuple
``("watch_name", "command", "poll_count", "max_polls", "is_final")``
is gone; future field adds touch one constant.
* **sec-1:** Run ``sanitize_payload`` over string-typed metadata fields
(``watch_name`` / ``command``) before they enter the queue. Today's
consumers all use ``textContent``, but the asymmetry — sanitised
``text`` alongside unsanitised metadata — would survive forever in
DB rows and resurface if a future consumer used a non-textContent
sink (aria-label, copy-to-clipboard, markdown render).
* **q-7:** Drop the per-iteration ``isinstance(reminder, dict)`` from
the dispatch closure's metadata comprehension. By the time the
block runs, ``text = reminder.get("text", "") if isinstance(...)``
+ the ``if not sanitized: return`` guard above already established
``reminder`` is a non-empty dict.
* **q-2:** Strip tombstone-style references — "post-#482", "post-#484",
"Step 7 of the watch-card UX plan", "Post-Step-7 dispatch surface",
and the brittle line-anchor "session.py:2685-2686" — across
``session.py``, ``test_session.py``, ``test_watch.py``,
``test_watch_dispatch.py``, ``test_watch_integration.py``. Comment
intent preserved; historical anchors gone.
* **q-6:** Drop the ``del source`` line in ``cli.py``'s
``on_user_reminder``; the parallel ``on_tool_reminder`` ignores
``tool_call_id`` without ``del`` and the comment alone is enough.
* **perf-4:** Document the SQLite ``render_as_batch=True`` recreate
cost in migration 050's docstring — first deployment after upgrade
copies the conversations table twice (one per ``add_column``).
PostgreSQL is unaffected.
5734 non-live tests pass; ruff + mypy clean.
|
||
|
|
91e7f2daca |
fix(session): preserve _source/_reminders on fork + cap persisted reminder text
Closes round-1 review findings bug-2 (major), perf-1 (minor), perf-6 (nit).
* **bug-2:** ``ChatSession.resume(..., fork=True)``'s bulk-row builder
silently dropped the ``_source`` and ``_reminders`` side-channel
data the source workstream had persisted via ``_append_user_turn``.
Both backends' ``save_messages_bulk`` already accept these keys
(the columns exist post-migration 050) — the bulk builder just
didn't supply them. The fork's resumed transcript would then look
like the assistant turn answered out of nowhere: every wake marker
and every reminder bubble that survived to disk on the source got
dropped on the fork. New regression test
``test_fork_preserves_source_and_reminders`` pins the contract.
* **perf-6:** Extracts ``_encode_reminders(reminders) -> str | None``
near ``_apply_reminders_for_provider`` so the user-turn save path,
the tool-turn save path, and the new fork bulk builder share one
encoder. Eliminates the drift risk between three near-identical
``json.dumps(..., separators=(",", ":")) if X else None`` patterns.
* **perf-1:** The new helper clamps each entry's ``text`` field at
``REMINDER_TEXT_STORAGE_CAP = 8192`` characters before encoding so
a single rogue producer (a watch streaming unbounded shell output,
a corruption-class steering payload) can't blow the conversations
row width or the FTS5 index. The in-memory side-channel keeps the
full body — only the persisted JSON is clamped. Mirrors
``TOOL_RESULT_STORAGE_CAP`` on tool result rows.
5734 non-live tests pass; ruff + mypy clean.
|
||
|
|
f1466ca7e3 |
fix(session): flag persisted reminders delivered on resume
Persisted ``_reminders`` survive ``load_messages`` but the in-memory ``_reminders_delivered`` flag does not (it's session-scoped — set by ``_mark_reminders_delivered`` after each successful provider stream, never persisted alongside the JSON column). Without a re-splice guard at resume time, ``_apply_reminders_for_provider`` would walk every loaded message, see ``_reminders`` set + the flag falsy, and splice every historical ``<system-reminder>`` envelope onto the wire on the very next user turn — leaking each reminder a second time, the turn after it had already advised. Mirror the post-stream hook in ``resume()``: every loaded message that carries reminders has already been delivered (it survived to disk), so flag it accordingly so ``_apply_reminders_for_provider`` short-circuits on the pass-through path. Test pins the contract end-to-end — stage a workstream with a persisted reminder, resume into a fresh session, append a live user turn, run the wire transform, and assert the historical reminder body does NOT land in the rendered output. |
||
|
|
6ae6877acc |
feat(ui): structured watch-result card + system-nudge marker on replay
User-visible slice of the watch-card UX workstream — combines the
replay-path widening, both frontend renderers, the CSS, and the
cross-cutting Python tests.
server._build_history widens the reminder filter from {type, text} to
project on a known set of optional fields (watch_name, command,
poll_count, max_polls, is_final) and surfaces _source as
entry["source"] when set. The known-key filter narrows the blast
radius if a future producer accidentally stuffs sensitive fields
into the dict.
SessionUIBase.on_user_reminder takes a new source: str | None kwarg
that rides on the SSE event when set. _attach_pending_user_reminders
forwards user_msg["_source"] so non-originating tabs see the wake's
"system_nudge" tag and render the thin marker. Protocol + cli + eval
implementations widen accordingly.
Frontend (coordinator.js + app.js — touched in lockstep per project
memory's "logic that lands in BOTH UIs must touch both files"):
* Branch on r.type === "watch_triggered" for a structured
.msg.watch-result card with header / $ command / <pre> body /
poll N/M [· final] footer.
* New addSystemNudgeMarker (interactive) + appendSystemNudgeMarker
(coord) renders a thin .msg.user.system-nudge anchor for
wake-driven reminders, both live (source === "system_nudge" on the
SSE event) and replay (msg.source === "system_nudge").
* Default .msg.user-reminder rendering preserved for every other
metacog nudge type.
CSS (shared_static/chat.css):
* New .msg.watch-result rules — full-width treatment, cyan accent,
monospace body with word-break: break-word for mobile.
* New .msg.user.system-nudge rule — thin yellow marker.
* Bonus newline-collapse fix: .msg.user-reminder .msg-body now sets
white-space: pre-wrap so multi-line shell output / bulleted lists
stay readable inside the advisory bubble.
Plan reference: docs/design/watch-card-ux.md §4 Steps 9-12 + bonus
CSS §11 (Commit 4).
|
||
|
|
f64c3e7b10 |
feat(storage): persist _source + _reminders side-channels on conversations
Adds two TEXT-NULL columns to the conversations table so multi-tab / multi-device replay sees the same metacognitive bubble shape the originating tab saw live. Until now, reminders lived only on the in-memory ChatSession.messages dict, and the wake-driven empty user turn was not persisted at all (skip at session.py:2685-2686) — a second tab connecting via /history saw the assistant turn with no preceding wake context, and missed every other tab's reminder bubbles besides. Single Alembic revision 050 (head was 049) adds: * conversations._source — today only "system_nudge" for wake rows * conversations._reminders — JSON-encoded reminder list Both backends (sqlite + postgresql) thread the columns through save_message / save_messages_bulk / load_messages. reconstruct_messages unpacks the row tuple as 9 elements (was 7), JSON-decoding _reminders on the user AND tool branches with the same contextlib.suppress guard the existing provider_data / tool_calls decode uses. Tool-row reminders ride the same column so tool_error / repeat replay shape matches user-channel parity. session.py:2685-2686 wake-row persist skip is dropped; _append_user_turn JSON-encodes user_msg["_reminders"] and passes both source + reminders to save_message. The tool-message save site at session.py:3014-3020 mirrors with metacog_reminders. Plan reference: docs/design/watch-card-ux.md §4 Steps 1-5 (Commit 1). |
||
|
|
94ed79d488 |
feat(metacog): switchover — watches enqueue onto NudgeQueue not _watch_pending
Replaces the bespoke _make_watch_dispatch / _watch_pending /
_dispatch_pending_watch / _MAX_WATCH_CHAIN machinery with a single
NudgeQueue.enqueue("watch_triggered", ...) call inside
ChatSession.set_watch_runner. Watch results now drain at the same
<system-reminder> envelope seams as every other metacog nudge
(USER_DRAIN, TOOL_DRAIN, IdleNudgeWatcher IDLE wake) — no separate
worker-spawn, no recursive watch chain, no per-session queue.Queue.
The dispatch closure built inside set_watch_runner carries:
- producer-side sanitize_payload over the whole formatted message
before enqueue, so steering-vector / control-char shell output
can't tamper with the envelope at interpolation time
- a soft cap of 50 entries on per-session "watch_triggered" depth
via the new NudgeQueue.drop_oldest_by_type, replacing the prior
_watch_pending maxsize=20 + _MAX_WATCH_CHAIN=5 bounds; drop policy
is drop-oldest (latest output most useful), logged at WARNING
- a valid_until predicate that re-checks
storage.get_watch(watch_id)["active"] at drain time so a cancelled
watch's last splat doesn't ride out a future wake
Behavioural delta documented in the plan section 3.4: N back-to-back
watch fires now drain into ONE assistant turn responding to all N
(via the envelope splice) instead of N separate send turns. This is
intentional — fewer model invocations for noisy watches, and uniform
with the rest of the metacog pull-model surface introduced by #482.
Implements watch-switchover plan steps 5-8. Server-side simplifications
let the previously-load-bearing _make_watch_dispatch (47 lines), its
session_worker.send import, and the chat-loop _dispatch_pending_watch
seam at the no-tools IDLE branch all disappear. The obsolete
tests/test_watch_dispatch.py and the wake-tag test in test_session.py
(both pinning contracts that no longer exist) are removed; the
NudgeQueue-based replacement plus an integration test land in the
following commit.
|
||
|
|
3f106f98b2 |
fix(metacog): apply-pass fixes from pre-push full-stack review
Round-2 review caught 11 confirmed findings on the 3-commit metacog stack; this commit applies them. * **bug-1 (major)**: Wake source tag was leaking onto real user messages flushed during a wake send. ``_append_user_turn`` and ``send`` now take an explicit ``from_wake: bool`` parameter — only the wake's synthesized first turn passes True, so ``_flush_queued_messages``'s real user input no longer inherits the audit tag. Regression test pins the contract. * **perf-1 (major)**: ``CoordinatorIdleObserver._maybe_enqueue`` was issuing list_workstreams + visible_memory_count storage queries before the cheap cooldown gate could short-circuit. New ``_cooldown_allows`` read-only peek runs first; storage queries only fire when cooldown actually allows the nudge. * **q-1 (major)**: Added the missing coord-side integration test that exercises ``CoordinatorIdleObserver`` + ``IdleNudgeWatcher`` together in the production install order against a real ``SessionManager``, protecting the subscription-order contract from silent regression. * **perf-2/3 (minor)**: Cap check moved above ``_last_assistant_used_wait``; ``_fire_counts`` restructured as ``dict[str, dict[str, int]]`` keyed by ws_id so the leave-IDLE existence check is O(1). * **perf-4 (minor)**: ``NudgeQueue.drain`` fast-paths the all-match case (the common one for chat-loop drain seams) by swapping ``self._items`` directly instead of allocating a fresh ``kept`` deque + per-entry append. * **perf-5 (minor)**: Wake's synthesized empty user turn no longer writes a content-empty row to the conversations table — the ``_source`` audit tag isn't column-backed and the side-channel reminder is stripped before persist, so the row would carry nothing. * **q-3 (minor)**: Split ``IdleNudgeWatcher`` + ``install_*`` / ``shutdown_*`` helpers out of ``metacognition.py`` into the new ``turnstone/core/idle_nudge_watcher.py``; metacog stays a static-template module. * **sec-1 (nit)**: Widened ``_sanitize_child_name``'s control-char regex to cover Unicode bidi-overrides, zero-width chars, line/paragraph separators, BOM, and tag chars. * **q-4/q-5 (nits)**: Docstring referenced the wrong peek primitive (``has_pending`` → ``len()``); ``_last_assistant_used_wait``'s ``session`` parameter now typed ``ChatSession``. 5571 non-live tests pass; ruff + mypy clean. |
||
|
|
f0e7fea549 |
feat(metacog): wake trigger — IdleNudgeWatcher + ChatSession.deliver_wake_nudge_from_queue
Adds the third metacog channel: an out-of-band wake that converts a workstream's IDLE transition into a synthetic empty-user-turn ``send`` when the session has any-channel nudges queued. The ``IdleNudgeWatcher`` subscribes to ``SessionManager.subscribe_to_state``; on IDLE it dispatches via ``session_worker.send`` with a no-op ``enqueue`` callback so a busy-worker race silently drops without spawning a competing worker. Wake-source-tag plumbing on ``ChatSession`` short-circuits metacog detection on the synthetic empty input, suppresses queue producers during the wake's own tool dispatch, and stamps ``_source = "system_nudge"`` on the synthetic user-message for audit / replay distinction. The tag is saved / restored across ``_dispatch_pending_watch`` so watch chains recursing off the wake are processed as normal user turns rather than inheriting the wake's guards. Generic ``install_idle_nudge_watcher`` / ``shutdown_idle_nudge_watchers`` helpers wire the watcher into both the interactive and coord lifespans via a single ``app.state`` registry so both surfaces share the same teardown contract. Foundation for PR 3 (CoordinatorIdleObserver + idle_children formatter) and PR 4 (watch dispatcher switchover). |
||
|
|
94b3720916 |
refactor(metacog): unify advisory channels into pull-model NudgeQueue
Replaces the dual `_pending_user_advisories` / `_pending_tool_advisories`
list pair with a single channel-tagged `NudgeQueue` per session.
Producers tag entries with a channel ("user", "tool", or "any");
consumers drain by channel filter at their existing seams. Foundation
for the wake trigger (PR 2) and coordinator idle-children nudge (PR 3).
Existing nudges (start, correction, completion, denial, resume,
tool_error, repeat) keep their wire shape and drain timing — zero
behavior change. Cancel paths now `clear()` the unified queue.
|
||
|
|
39a6b7b447 |
fix(man): accept canonical name(section) page notation
Models often emit page references in the standard man-page form (``printf(3)``, ``open(2)``, ``perlfunc(3pm)``) rather than splitting them into ``page`` + ``section`` args. The page-name sanitizer was rejecting the parens as invalid input, killing the call. Parse the section out of the page string before sanitization (explicit ``section`` arg still wins) and widen the section validator to accept multi-letter suffixes like ``3pm`` / ``3perl`` that already appear on real systems. |
||
|
|
11f0813329 |
fix(session): properly inject queued user messages mid-loop (#474)
* fix(session): properly inject queued user messages mid-loop
Two queued-user-message bugs in ``ChatSession.send()``.
**Mid-tool-call: ``Unexpected role 'tool' after role 'user'`` on Mistral.**
The ``supports_tool_advisories`` capability flag (default False for
unknown openai-compatible models) routed cap-off providers down a
short-circuit branch in ``_collect_advisories`` that called
``_flush_queued_messages`` directly. That appended a ``user`` turn
between ``assistant(tool_calls)`` and ``tool``, which mistral-common's
``_validate_message_order`` rejects with a 400.
Drop the flag. All providers now run the unified path: queued user
messages become ``UserInterjection`` advisories that ride inside the
tool result envelope via ``wrap_tool_result``, splicing
``<system-reminder>`` text into the tool message's content. Role
sequence stays ``assistant → tool``. Live-confirmed on Mistral
medium and Qwen3 — both correctly distinguish system-reminder from
tool stdout in their reasoning.
**Mid-stream: queued message orphaned until next user send.**
After a no-tool assistant turn, ``_flush_queued_messages`` would
append the queued user message to history and the loop would
``break``, leaving the message at the tail of history with no
model response. Visible as "two sends to get one reply".
``_flush_queued_messages`` now returns ``bool``. The no-tool branch
``continue``s on drain instead of ``break``ing, so the model gets a
turn over the extended history.
Tests:
- ``test_collect_advisories_drains_text_queued_messages_to_persistent``
pins the unified-path drain (text-only queue → ``UserInterjection``,
no separate user turn appended to ``self.messages``).
- ``test_send_continues_when_messages_queued_during_streaming`` pins
the loop-continue behavior (fails with 1 stream call pre-fix,
passes with 2 post-fix).
* fix(session,ui): reject queued attachments + paperclip busy state
Copilot pointed out that the attachment-bearing branch in
``_collect_advisories`` had the same role-ordering bug as the
text-only path that
|
||
|
|
c339615e39 |
Bound search tool output against pathological inputs (#473)
* Bound search tool output against pathological inputs
Replaces the per-line truncation with a fully bounded pipeline so the
search tool can no longer overflow the LLM context — or OOM the parent —
on minified bundles, multi-GB JSONL records, or huge result sets.
Backend:
- Prefer ripgrep when on PATH; grep is the fallback. Detection is
cached via functools.cache.
- ripgrep flags do most of the bounding natively: --max-columns 1024
+ --max-columns-preview, --max-filesize 10M, --max-count 100,
--no-config, --no-messages, plus negative globs for the same
noisy directories grep has been excluding.
- ripgrep added to the Dockerfile.
Streaming subprocess (_search_capture):
- subprocess.Popen with a streaming, byte-capped stdout read (4 MB).
Defends against single-line files (training data, minified bundles)
that would have OOM'd the previous subprocess.run capture.
- threading.Timer watchdog enforces tool_timeout even when the
pipe read is blocked in the kernel — proc.wait(timeout=…) alone
was insufficient because the read sat ahead of it.
- Stderr drained in a daemon thread to avoid pipe-deadlock when the
child writes to stderr while we're still reading stdout. Cap on
captured stderr keeps a hostile child from growing the buffer.
Tier-based formatter (_format_search_results):
- Tier 1: full path:line:content output, stream-emitted with a
running-cost short-circuit so we never materialize past the budget.
- Tier 2: K samples per file with overflow notes; K is computed
analytically from budget / file_count / avg-line-length so we hit
the right ladder rung in a single pass.
- Tier 3: per-file counts only, also budget-bounded with a tail line
reporting the omitted files. Sorted by descending count.
- Total output budget (32 KB) is well under tool_truncation, so the
head+tail _truncate_output strategy never silently drops middle
files in a search result.
Argument injection fix:
- The ripgrep arg list was missing the `--` separator that the grep
branch already had. With auto_approve on the search tool, that was
exploitable: path='--pre=COMMAND' would have made ripgrep run the
script as a per-file preprocessor and surface its stdout. Added
`--` and a regression test.
State-machine cleanup in _exec_search:
- rc < 0 (signal-killed by something other than us) now surfaces a
dedicated 'killed by signal N' message instead of being parsed as
success.
- capped + zero parsed records (e.g. one multi-MB line with no \n)
now returns a dedicated byte-cap message instead of the malformed-
output message that previously masked the real cause.
- _report_tool_result descriptions now match the returned payload
(no more 'no matches' tag on a 'malformed' payload).
Defence-in-depth on env scrub:
- RIPGREP_CONFIG_PATH, GIT_CONFIG, GIT_CONFIG_GLOBAL, GIT_CONFIG_SYSTEM
added to _EXPLICIT_SCRUB. We pass --no-config on the rg CLI today,
but if a future caller forgets the flag, an attacker who can set
one of these env vars could plant a config containing --pre=… and
recreate the same RCE shape.
Tests:
- TestSearchLineTruncation rewritten to mock _search_capture instead
of subprocess.run (the previous tests passed ChatSession kwargs
that no longer satisfy the constructor).
- TestSearchBackendSelection covers rg/grep detection and arg
construction, including the --pre flag-injection regression.
- TestSearchOutputBudget exercises Tier 1/2/3 directly.
- TestSearchCaptureStreaming spawns real Python subprocess writers
to exercise the byte-cap trim, mega-line-no-newline edge case, the
watchdog timeout when the child writes nothing, and the stderr
drain under load.
- test_env_scrub picks up the new tool-config keys.
* Address Copilot review on #473
- Budget the Tier 2/3 header up front so the formatter's emission stays
strictly within _SEARCH_OUTPUT_BUDGET. Previously the fit checks only
counted body bytes, letting the final string overflow by ~120 chars
(header + separator) and triggering _truncate_output's head+tail
dropout — exactly the shape this code was trying to avoid.
- Restore the (5, 3, 1) ladder in Tier 2: the analytical K from perf-2
is kept as a starting estimate, but if that K's actual emission
doesn't fit (the estimate ignores the header and overweights shared-
path compression) we step down through the ladder before falling
through to Tier 3. The previous one-shot K could collapse to counts-
only when 3/file or 1/file would have fit.
- Only normalise rc to 0 in the capped-output path when rc < 0 (our
SIGKILL). There's a narrow race where the child can exit naturally
between our read and our kill; preserving a non-negative rc means
rg's rc=2 ('matches found but some files had errors') no longer
silently turns into a clean success when the byte cap also fires.
- Clarify _MAX_SEARCH_LINE_LENGTH doc: the cap applies to the content
portion (after path:lineno:), not the whole emitted line.
- Add explanatory comments on the two intentional `except Exception:
pass` blocks in _search_capture (stderr drain, pipe close in the
cleanup finally) so static analysis and future readers can see the
silence is deliberate.
- Tighten the budget tests: now assert strict `<= _SEARCH_OUTPUT_BUDGET`
instead of the +512-char slack that was masking the header overflow.
- New regression tests:
- Tier 2 ladder step-down (K=5 over budget, K=3 fits, no Tier 3 fall-through)
- capped + rc=2 surfaces stderr instead of being normalised to success
- capped + rc<0 (our SIGKILL) flows through as a partial-result success
* chore(search): post-review cleanup
Follow-up to the Copilot-review fixes in
|
||
|
|
32fd8f29c7 |
feat(providers): api_surface toggle + mistral medium reasoning fix (#469)
* feat(providers): api_surface toggle + mistral medium reasoning fix
Mistral medium open-weights served by vLLM expects reasoning_effort via
the Responses API (`reasoning.effort`), not as a `chat_template_kwargs`
entry on Chat Completions. The session was unconditionally injecting
`{"reasoning_effort": ...}` into `chat_template_kwargs` for every
openai-compatible request, which corrupted the prompt rendering for any
backend whose chat template didn't consume that key (Mistral medium,
Mistral cloud, Groq, OpenRouter).
Changes:
- Add `api_surface` ("chat" | "responses") to `ModelConfig.server_compat`
and thread it through `create_provider` / `model_registry.get_provider`.
`openai-compatible` defaults to Chat Completions; operators can flip
individual aliases to Responses for endpoints that support it.
- New `vllm-mistral-medium` profile that pre-fills api_surface=responses
on Detect for known Mistral medium model ids.
- Drop the unconditional `reasoning_effort` injection into
`chat_template_kwargs`. Operators running gpt-oss-style local
templates that consume `reasoning_effort` from the chat template now
opt in via `server_compat.extra_body.chat_template_kwargs`.
- New "API Surface" select in the Models admin tab; allowlist-validated
server-side at create/update time; pre-filled by Detect via the
profile suggestion.
- Evict the cached provider singleton in `ModelRegistry.reload()` when
api_surface changes (previously only cfg.provider triggered eviction).
- Fix `_run_agent` fallback path to inherit the session's primary alias
for capability and server_compat resolution; previously the fallback
passed `alias=None`, which silently dropped per-model caps on the
agent path.
Tests: 5117 passed (-m "not live"); ruff + mypy clean.
* fix(providers): don't auto-suggest Responses for Mistral medium
vLLM's Responses API surface for Mistral medium open-weights doesn't
wire up the Mistral tool-call parser as of vLLM 0.x — tool calls leak
into the response as ``[TOOL_CALLS]<name>{...}`` text instead of
structured tool_calls. Chat Completions on the same engine handles
tools cleanly via ``--tool-call-parser mistral``, and reasoning can be
turned on via the vLLM CLI ``--reasoning-parser`` flag.
Drop the auto-suggest mapping so Detect falls back to the generic
``vllm`` profile. Keep the ``vllm-mistral-medium`` profile definition
in place so an operator who specifically wants per-request effort and
accepts the tool-calling limitation can still pick "Responses API"
manually in the admin UI.
* fix(providers): address Copilot review on PR #469
- providers/__init__.py: drop the redundant *_responses_provider /
*_chat_provider names; have create_provider use _openai_provider and
_openai_compat_provider directly so they're not flagged as unused
globals.
- console/server.py: tighten _validate_api_surface to a strict equality
match against the canonical {"chat", "responses"} set. The previous
strip().lower() membership check accepted ' Responses '/'CHAT' but
stored the raw string verbatim, which then failed to round-trip
through the admin <select>.
- console/static/admin.js: gate the entire server_compat block (server
type, api_surface, extra_body) on provider == "openai-compatible" at
save time so toggling provider away can't leave a stale hidden surface
selection in the persisted capabilities JSON.
- tests/test_session.py: splat the bad kwarg via **dict so CodeQL no
longer flags the call as a wrong-name keyword (the point of the test
is the runtime contract, not the static type).
- tests/test_admin_model_registry_refresh.py: add endpoint-level tests
for the api_surface validation on both create and update — covers the
bogus-value rejection, non-canonical-string rejection, and the happy
path persisting through to the refreshed registry.
|
||
|
|
7a36ab95e4 |
fix(metacog): drop duplicate [repeat: tool()] info line
The themed ``tool_reminder`` bubble below the tool block already shows the metacog text, and the tool block immediately above it carries the tool name — so a separate gray ``[repeat: list_workstreams() called with same arguments]`` info line was just duplicate visual noise (operator-visible in the screenshot below the bubble). Drop the ``ui.on_info`` call inside ``_apply_post_execute_advisories`` that emitted the diagnostic line. Update the docstring to reflect that the bubble is the canonical signal. Rename ``test_emit_repeat_ui_line_on_streak_fire`` → ``test_no_legacy_repeat_info_line_on_streak_fire`` and invert the assertion. |
||
|
|
845dbab616 |
fix(metacog): drop write-success-clear so sequential same-call streaks fire
The repeat-detection block in ``_apply_post_execute_advisories`` had
a leftover "clear streak when a write tool succeeded" branch from
when ``RepeatDetector`` tracked cumulative counts. With the
consecutive-streak semantics introduced earlier in the branch the
branch became:
1. Redundant — any different (name, args) signature already resets
the streak via ``RepeatDetector.record``, so an intervening
read/write naturally breaks the streak.
2. Actively wrong — the clear runs ONCE at the top of each
``_apply_post_execute_advisories`` call, before the per-result
loop records sigs. In a single parallel batch
``[bash, bash, bash]`` the clear runs once and then three
``record`` calls accumulate to count=3 in the same call → fires.
But across three sequential turns, each turn calls
``_apply_post_execute_advisories`` fresh, the clear runs at the
top of each call, and only one ``record`` per call follows — so
the count never gets above 1 and the canonical
"small local model stuck on ``bash('echo test')``" pattern
never triggered the nudge.
The asymmetry only existed for successful calls — failures don't
satisfy the ``not _tool_error_flags.get(tc["id"])`` predicate, so
the clear didn't fire and sequential failures already worked. The
fix is to drop the clear entirely; ``RepeatDetector``'s
consecutive-streak semantics handle every case uniformly.
Tests:
- ``test_successful_write_clears_streak`` →
``test_intervening_different_call_resets_streak`` —
rewords the assertion to reflect the actual mechanism (any
different sig resets, write-or-otherwise) since "writes clear"
was the bug, not the contract.
- ``test_failed_write_does_not_clear_streak`` →
``test_sequential_bash_failures_fire_repeat`` — same shape, just
framing fixed.
- New ``test_sequential_bash_same_command_fires_repeat`` —
regression for the bug user hit (three sequential successful
``bash('echo test')`` calls now correctly fire the nudge).
|
||
|
|
c0fd951764 |
feat(metacog): themed reminder bubble unifies user + tool channels
The yellow themed reminder card introduced for user-channel nudges
(correction / denial / resume / start / completion) now also fronts
tool-channel nudges (tool_error / repeat). Pre-fix the tool channel
shipped its reminders inside the tool-result envelope via
``wrap_tool_result``, leaking the ``<system-reminder>`` block into
``self.messages`` content (same problem the user channel had before
the side-channel refactor) and surfacing the legacy gray
``[metacognition: nudge injected — …]`` info line as the only
operator-visible signal — duplicated alongside the new themed bubble
for user-channel nudges.
Tool-channel parity:
- ``_collect_advisories`` now returns
``(persistent_advisories, metacog_reminders)``. Persistent
advisories (``GuardAdvisory`` / ``UserInterjection``) keep
riding ``wrap_tool_result`` because they ARE conversation
history. Metacognitive reminders extract to the second tuple
element; the caller attaches them to the tool message dict's
``_reminders`` side-channel and emits ``on_tool_reminder``.
- ``_apply_reminders_for_provider`` already handles ``_reminders``
on any role, so the tool-channel splice into wire content is
free. ``_build_history`` also already propagates
``entry["reminders"]`` regardless of role, so reload renders the
bubble too.
- ``SessionUI`` Protocol gains ``on_tool_reminder(reminders,
tool_call_id)``; ``SessionUIBase`` enqueues a ``tool_reminder``
SSE event with the ``tool_call_id`` anchor.
- ``_emit_nudge_ping`` had no remaining callers and was removed —
the themed bubble (live SSE + ``/history`` reload) is the
canonical operator signal for both channels now.
UI polish (the four fixes the screenshot caught for the user
channel + their tool-channel mirror):
- Bubble renders BELOW the message it advises (semantically: a
hint to the model right before its turn). ``addUserReminder``
swaps ``insertBefore`` for ``insertAdjacentElement('afterend',
el)``; ``addToolReminder`` anchors below the ``.ts-approval``
block whose tool result triggered the batch's reminder.
- Label uses the full feature name ``metacognition`` (was the
``metacog`` shorthand).
- Card width / alignment inherits from the base ``.msg`` rule —
``align-self: flex-end`` and the explicit ``max-width`` are
gone, so the card matches the user / assistant column instead
of pinning right-aligned narrow.
- The legacy ``[metacognition: nudge injected — …]`` gray info
line is gone for both channels.
Frontend additions:
- ``Pane.prototype.addToolReminder(reminders, toolCallId)``
anchors below the ``.ts-approval`` block (live: by
``data-call-id``; replay: by "last block in messagesEl"
fallback, which is correct because messages render in order).
- SSE switch case ``"tool_reminder"`` calls ``addToolReminder``.
- ``replayHistory``'s tool-message branch now calls
``addToolReminder`` when ``msg.reminders`` is present.
- ``addUserReminder`` advances its anchor on each loop iteration
so multiple reminders stack in queued order rather than
reversed.
Coord console parity:
- ``coordinator.js`` gains ``appendReminderBubble`` /
``appendUserReminderLive`` / ``appendToolReminderLive`` mirroring
the interactive UI. The tool-channel anchor walks
``toolRows[callId].batch`` to attach below the
``.coord-tool-batch`` construct (one bubble per dispatch turn,
matching the "one nudge per batch even with many failing tools"
drain).
- SSE switch handles ``user_reminder`` and ``tool_reminder`` on
the coord conversation surface.
- ``/history`` replay propagates ``msg.reminders`` for user and
tool messages — same wire shape as the interactive pane.
- ``.msg.user-reminder`` styles moved to
``shared_static/chat.css`` so both surfaces inherit the same
yellow themed bubble from the shared base.
Defensive read on ``_apply_reminders_for_provider`` (per Copilot
review on the closed PR): a malformed ``_reminders`` entry (string,
None, etc. — corruption / partial state) used to abort ``send`` via
AttributeError on the ``.get("text", "")`` call. Filter to dicts
before building the block, mirroring the same filter
``_build_history`` already applies on the wire-out side; an
all-malformed list passes through as no-reminders.
Tests:
- ``test_collect_advisories_drains_tool_buffer_on_last_result``
rewritten to assert the ``(persistent, metacog)`` tuple shape
and that ``MetacognitiveAdvisory`` no longer appears in the
persistent list.
- ``test_collect_advisories_holds_*`` and ``_drops_*`` updated for
tuple return.
- ``test_attach_emits_visibility_ping`` /
``test_collect_advisories_emits_visibility_ping`` inverted to
assert the legacy gray line is gone on both channels.
- ``TestSessionUIBaseToolReminderHook`` covers the new SSE event
shape with the ``tool_call_id`` anchor.
- ``test_malformed_reminders_filtered_out`` and
``test_all_malformed_reminders_passes_through`` cover the
Copilot-flagged defensive filter.
|
||
|
|
3aa9f53fd8 |
fix(session): metacog reminders ride a side-channel, not user content
User-channel metacognitive nudges (correction, denial, resume, start,
completion) used to be spliced into ``user_msg["content"]`` permanently,
which leaked the ``<system-reminder>`` envelope into every consumer of
``self.messages`` — UI replay (mitigated by a regex strip in /history),
compaction, title generation, and any future channel adapter that
echoes conversation context. The /history strip was a band-aid;
compaction and title-gen still saw the raw spliced text.
Switch to a side-channel: ``_attach_pending_user_reminders`` writes the
rendered reminder list to ``user_msg["_reminders"]`` (sibling key,
leading-underscore convention shared with ``_attachments_meta`` /
``_provider_content``). At the provider boundary, a new
``_apply_reminders_for_provider`` builds a transient shallow-copy with
the reminder spliced into ``content``; the original message dict
stays clean. ``sanitize_messages`` drops the sibling key on the wire.
Once-per-session-not-per-turn semantics for the wire: after stream
success the loop calls ``_mark_reminders_delivered``, which flips a
``_reminders_delivered`` flag on every user message that carried
reminders into that call. ``_apply_reminders_for_provider`` skips
already-delivered messages so the model sees each reminder exactly
once (the turn it advised). ``_build_history`` ignores the delivered
flag entirely, so reconnecting tabs render the same nudge bubble the
originating tab saw via the live ``user_reminder`` SSE event.
UI surface:
- ``SessionUIBase.on_user_reminder`` enqueues a
``{type: "user_reminder", reminders: [...]}`` SSE event with the
same shape ``_build_history`` surfaces.
- ``app.js`` renders a ``.msg.user-reminder`` bubble (yellow accent,
pill-styled) anchored above the user message it advises, both
live and on history replay.
- ``replayHistory`` renders ``addUserMessage`` before
``addUserReminder`` so the anchor lookup finds the just-rendered
turn (not a prior one).
- Multi-tab caveat documented inline: non-originating tabs receive
no ``user_message`` SSE event today, so a reminder may anchor to
a stale prior bubble until ``/history`` reload corrects it.
Pre-existing bug surfaced by the audit: cancel handlers
(``GenerationCancelled`` / ``KeyboardInterrupt`` / generic
``Exception``) in ``ChatSession.send`` cleared
``_pending_tool_advisories`` but not the user-channel buffer. Both
now drain through a shared ``_drain_pending_advisories`` helper.
Removed the ``/history`` regex strip — the side-channel approach
makes it redundant. Hoisted ``escape_wrapper_tags`` +
``render_system_reminder`` imports to module top (called 2-3× per
turn).
Tests:
- ``TestApplyRemindersForProvider`` — pass-through-by-reference,
string + list content splice, escape on user-typed wrapper tags,
multi-reminder ordering, source-untouched invariant, delivered
flag skip path, fallback for unexpected content shape.
- ``TestMarkRemindersDelivered`` — flag idempotency, no-reminders
no-flag, only marks user messages with reminders.
- ``TestUpdateTokenTableMsgsParam`` — calibration uses pre-built
msgs when provided, falls back when not.
- ``TestUserAdvisoryCancelClear`` — all three cancel branches drain
the user buffer.
- ``TestReminderSidechannelIsolation`` — compaction's
``_format_messages_for_summary`` and the title-gen extraction
loop cannot see reminders by construction.
- ``TestSessionUIBaseUserReminderHook`` — ``on_user_reminder``
enqueues the right SSE shape.
- ``TestBuildHistoryReminderPropagation`` — ``entry["reminders"]``
propagation, absent / empty / multi / coexist-with-attachments
cases, malformed input filtering, all-malformed elision.
- ``test_sanitize_messages_strips_underscore_sibling_keys`` covers
``_reminders`` and ``_reminders_delivered``.
|
||
|
|
7c8cb8c595 |
fix(metacog): N>=3 streak detector + drop redundant error-prefix list
Cleanup pass on the metacognitive nudge stack — restores pre-split errored-counts-toward-repeat behaviour and tightens the is_error plumbing through the per-batch advisory hook. The per-batch hook in ``_run_loop`` was duplicating the is_error signal: ``self._tool_error_flags`` (set by ``_report_tool_result``) and a string-prefix tuple (``Error`` / ``JSON parse error`` / …). Two truth sources is what got us here — bash commands that exit non-zero with normal stdout matched the flag but not the prefix, the deny path matched the prefix but not the flag, and the result was that stuck-loop detection silently broke for the most common failure mode (the model bashing the same broken command). Single source of truth now: - ``_execute_tools.run_one`` deny branch routes through ``_report_tool_result(is_error=True)`` so denied calls populate ``_tool_error_flags`` like every other error path. - The error-prefix tuple is gone; the write-success-clear gate and the tool-error-nudge gate both read ``_tool_error_flags`` only. Repeat-detection state moves from a ``set[str]`` (fired on the second identical call, ignored errors entirely) to a ``RepeatDetector`` helper in ``metacognition.py`` with consecutive-streak semantics: - Threshold raised from 2 to 3 — two-in-a-row was noisy on legitimate transient retries; three is the cheapest stuck-loop signal. - Recording a different signature resets the count, so [A, A, B, A] is two short streaks of 2 and not a streak of 4. Bounded by O(1) state regardless of session length. - Errored calls now count toward the streak (the split into a separate metacog module unintentionally introduced a "skip errors" branch — restored). While there: - ``metacognition._COOLDOWN_SECS`` default aligned to 300s (matches ``MemoryConfig.nudge_cooldown`` and the ``memory.nudge_cooldown`` config-store default; was set to 30 by an earlier investigation). - The per-batch advisory block (~80 lines of mixed orchestration inside ``_run_loop``) is extracted to ``ChatSession._apply_post_execute_advisories`` so the wired behaviour is testable without driving ``_run_loop`` end-to-end. Producer extraction to a dedicated module is deferred to a follow-up; advisory producers all live on ``ChatSession`` for now per existing convention. - Frontend ``appendToolOutput`` (turnstone/ui/static/app.js) now skips rendering when the parent approval block is denied or the output starts with ``Denied by user`` / ``Blocked``, mirroring the history-replay guard at ``_build_history``. Previously the live SSE path didn't need this guard because the deny path never emitted a ``tool_result`` event; the is_error routing change above means it does now, so without this guard the badge from ``resolveApproval`` and the SSE output would both render. Tests: 8 unit tests for ``RepeatDetector`` covering streak, threshold, clear, and intervening-sig reset; 9 integration tests for ``_apply_post_execute_advisories`` covering the wired behaviour (3-identical fires warning + advisory + UI line, errored calls count toward streak as a regression guard, intervening sig resets streak, successful write clears, failed write does not, JSON outputs tracked but not inline-warned, tool_error nudge gates on memory_count, repeat UI line emitted on streak fire). |
||
|
|
64d5205dd6 |
perf(session): split metacognitive nudges out of the system message
The system-message developer block was rebuilt every turn with two unstable inputs: minute-precision current_datetime in the middle of the composed prefix, and _pending_nudge entries appended-then-cleared at the bottom. Both invalidated prompt-cache reuse on Anthropic / OpenAI for the entire prefix, every turn. current_datetime now rounds to the top of the hour. Nudges no longer ride on the system message at all — they drain through two channels: - tool_error and repeat ride the existing tool-result <system-reminder> envelope via a new MetacognitiveAdvisory ToolAdvisory subtype, drained in _collect_advisories alongside GuardAdvisory and UserInterjection. - correction, denial, resume, start, completion splice as <system-reminder> blocks at the trailing edge of the next user message via a new _splice_pending_user_advisories helper. User content passes through escape_wrapper_tags before concatenation so a user typing literal <system-reminder> tags cannot fabricate an envelope; the same escape now runs on advisory.render() output inside wrap_tool_result for defense-in-depth across all advisory types. Cancel handlers (GenerationCancelled / KeyboardInterrupt / bare Exception) now clear _pending_tool_advisories alongside the existing _flush_queued_messages so a queued nudge from an aborted batch cannot leak into the next generation. Visibility ping ([metacognition: nudge injected — ...]) preserved at both new attach points via a single _emit_nudge_ping helper. Also adds a "Session kind" line (interactive | coordinator) to the composed Session Context so the model can see which manager hosts its session. Tests: 4847 passing (+7 new in TestMetacognitiveBuffers and test_tool_advisory). ruff + mypy clean. |