Commit Graph

1045 Commits

Author SHA1 Message Date
Patrick Buckley caa589e39b feat(storage): tag the native lane with its producer ({producer, blocks})
Persist provider_data as a {producer, blocks} envelope (producer = the generating
provider's name) so the lowering layer can later replay the native lane verbatim only
to its producer and rebuild from neutral fields for any other.

The envelope is storage-only: prepare_provider_data_for_save runs the P1 mirror on the
bare block list and then wraps; reconstruct_messages dual-reads (new envelope OR legacy
bare list), unwraps to a bare _provider_content list (every consumer's contract), and
surfaces the producer on the stripped-before-wire _producer side channel. The producer
threads the same four save layers as is_error (facade -> protocol -> SQLite + Postgres);
the live assistant save tags it from self._provider.provider_name and the fork carries
_producer. Legacy rows need no migration to keep working (dual-read); the one-shot
backfill that tags them is a follow-up.

Sub-commit 2a of the canonical-trajectory storage cut.
2026-06-04 11:03:13 -07:00
Patrick Buckley bfa80f2159 feat(storage): persist tool-result is_error (migration 060)
Tool-result error state was an in-memory-only message key, lost on reload. Add an
is_error column to conversations (migration 060, backfilled False) and thread it through
the four save layers (memory facade → StorageBackend protocol → SQLite + PostgreSQL):
save_message/save_messages_bulk persist it, reconstruct_messages emits it on tool rows,
and the session tool-result + synthetic-cancel saves + the fork bulk-copy pass it. It
rides as the last conversations column so reconstruct's row-tuple positions stay stable
(legacy fixtures default False). history_decoration already prefers the persisted flag
over its text heuristic, so reload fidelity improves immediately.

First sub-commit of the canonical-trajectory storage cut (folds into rev 060).
2026-06-04 11:03:13 -07:00
Patrick Buckley 975fb4714c feat(core): add the canonical Turn trajectory model
The provider-neutral typed Turn (flat, role-discriminated; uniform tuple[ContentBlock]
content + .text; AttachmentRef by-reference content; ToolCall raw-arg str;
ProviderNative producer-tagged opaque lane; TurnMeta sidecar).  In-memory foundation
for the wire-shape narrow waist — no consumers yet; storage deserialization, the
lowering layer, and the provider translators wire onto it in subsequent steps.
2026-06-04 11:03:13 -07:00
Patrick Buckley 5323559e78 fix(storage): enforce the native/tool_calls mirror at the save boundary
A max-tokens truncation mid-tool_use can leave an orphan tool_use in the native lane
(provider_data / _provider_content) with no matching tool_calls; on a same-provider
resume that replays as a tool call with no result and the API rejects it.

normalize_native_for_save strips orphan client tool-call blocks (tool_use /
function_call / function) when tool_calls is empty, applied by save_message and
save_messages_bulk in both backends; strip_orphan_client_tool_blocks enforces the
same mirror in memory at message assembly (the truncation path). The mirror now holds
by construction, so the orphan-repair pass can read tool_calls alone and the Anthropic
pc_tool_ids fallback can be retired.
2026-06-04 11:03:13 -07:00
Patrick Buckley c9311e32e4 test(providers): freeze per-provider wire payloads as a refactor baseline
Captures the exact request kwargs each provider hands to its SDK seam (Anthropic
messages.stream, OpenAI chat/responses create, Google OpenAI-compat) for a
representative set of trajectories, asserted against committed goldens. This is the
behavior-equivalence net the canonical-trajectory wire-shape refactor is proven
against. Regenerate the baseline with UPDATE_WIRE_GOLDENS=1.
2026-06-04 11:03:13 -07:00
Patrick Buckley 6e260bddc5 fix(skills): don't echo model-controlled filter values into the trusted hint
Phase-2 review follow-ups:
- sec-1 (major): the skills find-zero hint interpolated the model-supplied
  filter values (query/category/tag/…) into the system_reminder, which now
  rides a TRUSTED operator system turn (fold fence / native system role). Under
  an indirect prompt injection the model could be steered to call
  skills(find, query='<directive>', category='nonexistent') so 0 rows match,
  laundering the attacker text into operator authority. Drop the filter echo —
  the count is harness-derived and the model already knows its own filters.
- q-1 (minor): refresh the stale :func:`escape_wrapper_tags` cross-reference in
  metacognition.sanitize_payload's docstring (the function was removed; fold-time
  fence.neutralize is the current marker defense).
2026-06-04 11:03:13 -07:00
Patrick Buckley f3c96e6493 feat(skills): make skill hints first-class system turns; drop escape_wrapper_tags
_skill_hint spliced its guidance into the tool result as a bare <system-reminder>
block — but the operator declaration now tells the model to treat bare markers
as untrusted, silently demoting the hint. Make the hint first-class instead:

- _skill_hint returns the tool result verbatim and queues the guidance via
  _queue_tool_advisory("skill_hint", ...); _collect_advisories drains it into a
  {role:system, _source:"skill_hint"} turn after the clean result — folded in
  the trusted nonce fence for non-native models, inline for native. (Queuing
  no-ops mid-wake, like the other tool-channel advisories.)
- skill_hint added to SYSTEM_TURN_SOURCES (an advisory-producer source).
- escape_wrapper_tags removed outright: it was the last consumer, and its job
  (defang a marker next to the bare block) is now covered at fold time by
  _neutralize_host. The result message rides through verbatim. This also
  collapses the two-escaping-mechanism confusion the review flagged.

Tests assert the clean result + the queued/drained hint, plus wake suppression.
2026-06-04 11:03:13 -07:00
Patrick Buckley 8513016503 refactor(storage): drop the dead _reminders column instead of carrying it
Operator context moved to first-class system turns, leaving _reminders written
by nothing and read by nothing. Nulling it (the prior 060 step) left a writable
dead column — a foot-gun inviting accidental reuse. Drop it outright and remove
every reference in one shot so there is no half-alive state:

- migration 060: replace the wholesale null with batch_alter_table drop_column
  (per migration 027); downgrade re-adds the empty column to match the 059
  schema (the envelope un-wrap stays irreversible).
- _schema.py: remove the column.
- _sqlite / _postgresql: drop the reminders save param, the INSERT/bulk values,
  and both SELECT columns.
- reconstruct_messages: the row tuple is now 8/9-tuple (event_id shifts from
  index 9 to 8); _utils + the _row test helper updated.
- _protocol / memory save_message: drop the reminders param + docstrings.
- tests: replace the reminders-roundtrip tests with a _source-only file and a
  060 drop-column assertion; remove the obsolete legacy-reminders wire test.

No production caller passed reminders=, and the SELECT no longer reads the
column, so an un-migrated DB simply ignores any residual values.
2026-06-04 11:03:13 -07:00
Patrick Buckley 99ba82e8ec fix(session): operator-turn wire correctness — framing, empty turns, leading system
Phase-2 follow-ups to the mid-conversation-system consolidation:

- user_interjection framing (known #2): a queued message that drains mid-turn is
  re-framed via render_user_interjection ("The user sent … User message: …") so
  the user's words keep USER authority, not operator authority — the regression
  mattered most on the native path, where the turn enters as a real role=system
  message. Empty/whitespace interjections (e.g. a bare "!!!") are dropped (bug-2).
- empty-content user turns dropped at the wire boundary after the fold
  (known #3): the wake pipeline's synthetic empty send("") leaves an empty user
  turn on the native path (the nudge stays inline); an empty user message is
  invalid on every provider. The drop runs after the fold so the fold-path wake
  turn, which the nudge fills, survives.
- leading-system guard (_anthropic): a turn that converts to nothing no longer
  lets a system message become messages[0] (the API requires messages[0]=user).
  Newly reachable now that the empty-turn drop can expose it on a fresh-session
  native wake.
- refresh stale .msg.watch-result comments (the card was removed) to describe
  the current operator-bubble rendering.
2026-06-04 11:03:13 -07:00
Patrick Buckley 2d64569cf8 fix(fence): close defang/detection whitespace gap; test host-escaping
Phase-1 review follow-ups:
- neutralize() now tolerates whitespace between '<' and the slash ('< /tag',
  '<  /tag'), matching output_guard's detection regex so a marker can no longer
  be detected-but-not-defanged (a leaked-nonce break-out gap).
- Add direct tests for the sec-1 forge-in defence: _neutralize_host defangs a
  forged <system-reminder_{nonce}> in both string- and list-content untrusted
  hosts before the real fence is appended, and the host is defanged exactly once
  so consecutive folds don't corrupt the first appended fence.
2026-06-04 11:03:13 -07:00
Patrick Buckley d9955805fd test(channels): drop orphaned PlanReviewView owner-check tests
PlanReviewView was removed alongside the plan_agent built-in tool (110d44b0),
but its Discord owner-check tests were left behind importing a class that no
longer exists. They raise ImportError wherever discord.py is installed (green in
CI only because discord.py is absent there). Remove the dead test class and the
now-unused send_plan_feedback router mock from the shared bot double.
2026-06-04 11:03:13 -07:00
Patrick Buckley 095f46a1fc fix(migration): tighten 060 envelope match; stop re-activating escaped tags
060 un-wrapped legacy <tool_output> envelopes with a loose guard (open + close)
that could irreversibly mis-rewrite a bare tool row resembling the open, and
entity-decoded the wrapper tags back to live form — re-activating injection the
old escape had neutralised (and downgrade cannot undo it).

- Require the full legacy signature (the exact </tool_output>\n\n<system-
  reminder>\n join plus a trailing </system-reminder>), which wrap_tool_result
  only ever emitted with advisories. A bare row with a matching close but no
  advisory is left byte-for-byte untouched.
- Reverse only &amp; -> & ; leave the wrapper-tag entities escaped so a
  previously-defanged injection stays defanged.

Adds false-positive guard tests (open+close without advisory; missing tail).
2026-06-04 11:03:13 -07:00
Patrick Buckley bb9e50c714 feat(fence): unify operator + judge trust fences on one primitive
Both the operator fold and the output-guard judge wrap spans in nonce-delimited
fences, but the two had drifted: the operator path minted a 32-bit nonce reused
per session with no body escaping, while the judge used a 64-bit per-call nonce
plus closing-tag escaping. Extract the shared mechanism (mint/neutralise/wrap)
into turnstone/core/fence.py and put both callers on it so they cannot diverge
again.

- Operator fold (sec-1): 64-bit nonce; fence.wrap neutralises the operator
  body's close marker, and _fold_system_turns neutralises the untrusted host
  turn's <system-reminder> markers once before the first fold, so a leaked or
  guessed per-session nonce still cannot forge a trusted block. Per-session +
  cached declaration kept (the declaration pins the exact value, so per-turn
  rotation would bust the prompt cache). Marker is now <system-reminder_{nonce}>.
- Judge: refactored onto fence (behaviour-preserving; still per-call).
- Forgery detection: output_guard scans tool output for trust-fence markers —
  an exact session-nonce match is HIGH (operator_marker_leak: the token has
  leaked and is being replayed), any other marker LOW (operator_marker_forgery).

Removes mint_envelope_nonce / wrap_system_context (folded into fence.wrap).
2026-06-04 11:03:13 -07:00
Patrick Buckley c6b2288302 feat(session): consolidate operator-context into first-class system turns
Replace the two operator-context hacks (the <tool_output>/<system-reminder> content envelope and the transient _reminders side-channel) with one persistent {role: system, _source} trajectory turn. Adds supports_mid_conversation_system (claude-opus-4-8): native models take the turn inline; all others fold it into the preceding turn as a nonce-delimited <system-reminder> block declared in the system prompt as the sole trusted marker. Producers (advisories, metacog nudges, user interjections, idle/watch) emit system turns; the envelope/_reminders machinery, escaping round-trip, replay parser, and reminder SSE events are removed. Eager 060 migration un-wraps legacy envelopes. Net -1662 lines.

Known follow-ups from review (unfixed here): (1) the 060 un-wrap heuristic can irreversibly mis-rewrite bare tool rows that resemble the envelope, so do not run the migration until it is tightened; (2) user_interjection turns lost the user-framing/priority preamble (a regression, and a native-path authority-framing concern); (3) native-path wake nudge can emit empty user content.
2026-06-04 11:03:13 -07:00
renovate[bot] ef2b18b62f chore(deps): lock file maintenance 2026-06-03 18:56:09 -07:00
renovate[bot] fb4c6eefe0 chore(deps): update helm release postgresql to ~18.7.0 2026-06-03 18:55:55 -07:00
renovate[bot] a80c4b025b chore(deps): update typescript sdk to v4.1.8 2026-06-03 18:55:42 -07:00
renovate[bot] 88683729d2 chore(deps): update docker images to v0.11.19 2026-06-03 18:55:29 -07:00
renovate[bot] 1696eeb9e5 chore(deps): update github actions 2026-06-03 18:55:16 -07:00
renovate[bot] e8211f12a0 chore(deps): update dependency aiohttp to v3.14.0 [security] 2026-06-03 18:41:51 -07:00
Patrick Buckley 87c51b43ca chore: bump version to 1.6.0a10 v1.6.0a10 2026-06-01 21:36:36 -07:00
Patrick Buckley 24d75a690b fix(review): address review findings on the rerank/memory stack
- rerank_config.py: the runtime instruction fallback had a dead tail
  (`get_rerank_instruction() or str(cs.get(...))` -- the cs.get term can only
  return the registry default ""), via a stored_keys() branch that also diverged
  from the calibrate CLI / endpoint. Collapse to the sibling idiom
  (`cs.get(...) or get_rerank_instruction()`) so the instruction used at
  calibration time matches the one used at runtime. Correct the module docstring:
  ChatSession is the sole caller (the CLI/endpoint share only the instruction
  precedence, not this function).

- session.py: the deferred first-turn memory recompose was gated on the flag
  alone, so a synthetic wake send (empty user content -> flag stays False) re-ran
  the full compose on every wake before the first real turn. Gate on a non-empty
  query too, so wakes don't re-pay it and the real turn still fires exactly once
  (+ test).

Accepted as-is: the __init__ compose (kept so system_messages/_agent_system_messages
are valid for early readers; one cheap extra compose per fresh session) and the
orphaned tools.rerank_* config rows (inert -- no read path, never listed or
redacted; a purge migration would collide with the 060 in flight on another branch).
2026-06-01 21:34:12 -07:00
Patrick Buckley 47d0b6c3c6 fix(memory): defer system-message compose to the first user turn so memory selection has a query
Proactive memory selection scores candidates against the recent-user-message
query (extract_recent_context), but a fresh session composes the system prefix
once in __init__ while self.messages is still empty. That empty query takes the
no-context path: _select_memory_candidates returns recency order and
score_memories returns memories[:k] verbatim -- the 5 most recently UPDATED
memories, with BM25 and the reranker never invoked. send() never recomposes, so
those recency-only memories are what the model sees for the whole session (until
an unrelated event -- skill / MCP / model refresh / resume / memory write /
command -- happens to rebuild the prefix). Net effect: the injected memories are
unrelated to the actual question.

Fix: defer the memory-bearing compose to the first real user turn.
- Track _system_composed_with_context, set once extract_recent_context is
  non-empty in _init_system_messages.
- send() recomposes once, right after _append_user_turn, while the flag is still
  False -- so the opening turn's memory block is selected (and reranked) against
  the real message. The flag then stays True, so the prefix is composed once and
  stays cache-stable exactly as before (no per-turn prompt-cache churn).

This is the targeted fix; per-turn memory refresh (so later topic shifts also
re-rank) is the larger tail-injection redesign tracked on another branch. Two
adjacent gaps are left as-is for now: the reranker/BM25 only see content[:200],
and build_memory_context flat-truncates each memory to 500 chars (max_content is
the save cap, not an injection budget).

Tests: flag is False on a fresh session and after a whitespace-only wake turn,
flips True on a real query; send() runs the deferred recompose with the user
message in the query.
2026-06-01 21:34:12 -07:00
Patrick Buckley 8b4b8b3fd5 refactor(rerank): reranker is a per-model definition only (drop global endpoint settings)
The reranker_alias -> model-definition path (added when reranking became a model
role) made the older global endpoint settings redundant. Resolve reranking
solely through the Reranker role and remove the parallel global config.

- Removed settings tools.rerank_url / rerank_model / rerank_api_key, their
  config.py getters (+ $TURNSTONE_RERANK_URL / $TURNSTONE_RERANK_MODEL and the
  module caches), and the fallback branch in resolve_rerank_client_from. The
  resolver now returns a client only when a Reranker model (capability
  supports_rerank, base_url = its /rerank endpoint) is selected, else None.
- Kept as global knobs: reranker_alias (the selector), rerank_web_search,
  rerank_bm25, rerank_bm25_threshold, and rerank_instruction -- a task-level
  query knob (Qwen3-style), not endpoint identity.
- The Settings tab is registry-driven, so the three fields disappear with their
  SettingDefs. Updated the Reranker role help, example config, and docs/tools.md.

BREAKING: a reranker configured via [tools] rerank_url (config.toml / env /
Settings tab) no longer works -- add the reranker in the admin Models tab and
pick it under Models -> Roles -> Reranker. No migration: reranking is days old
and disabled by default, so any orphaned tools.rerank_* config rows are inert.

Tests: the resolver covers no-store / no-alias / non-rerank-alias -> None and the
model-definition happy path; the obsolete global-fallback tests are removed.
2026-06-01 21:34:12 -07:00
Patrick Buckley a4a0965e85 fix(rerank): calibrate-as-probe for reranker detect (skip /v1/models)
Adding a reranker model through the admin modal was impossible: Detect
(admin_detect_model) always ran the OpenAI /v1/models probe first and gated
calibration on its `reachable` result. A Cohere/Jina /rerank endpoint can't
answer /v1/models, so it failed both ways -- `$host/v1` passed the probe but
calibration POSTed to the wrong path, `$host/rerank` 404'd the probe outright.

- console/server.py: branch on `supports_rerank` BEFORE the probe and calibrate
  the endpoint directly; that round-trip IS the reachability check (there is no
  independent /v1/models signal for a rerank-only endpoint). Success
  autopopulates the three calibration fields the way context_window does;
  calibration failure -> reachable:False + error; empty base_url -> 400 with the
  /rerank hint; reachable-but-no-clean-separation -> a note. Drops the now-dead
  post-probe calibrate-on-detect block.

Reranker selection stays per-model (reranker_alias -> registry); recalibrate on
a saved reranker was already correct (it calibrates directly). Flagging a model
as a reranker still rides the capabilities JSON (supports_rerank).

Negative-tested: rerank detect skips the probe, autopopulates on success, notes
no-separation, reports unreachable on calibration failure, and 400s on an empty
base_url; the non-rerank detect path is unchanged.
2026-06-01 21:34:12 -07:00
Patrick Buckley 68ed0b6c15 feat(rerank): per-model calibration on detect + calibrate endpoint
Phase 3. Stores reranker calibration per-model on the model definition's
capabilities (rerank_threshold/rerank_scale/rerank_separated; a non-empty
rerank_scale is the "has been calibrated" marker) instead of a single global
threshold, populated automatically when a reranker endpoint is detected.

- ChatSession._bm25_rerank_threshold precedence: the active reranker model's
  calibrated rerank_threshold (when separated) wins; calibrated-but-not-
  separated -> 0 (no floor); else the global tools.rerank_bm25_threshold
  fallback. Reads the raw caps dict, in-memory per turn.
- Detect (admin_detect_model) calibrates a supports_rerank endpoint and
  autopopulates the three fields like context window; the create-model UI shows
  a verdict chip (calibrated / no-clean-separation / not-calibrated).
- POST /api/admin/model-definitions/{id}/calibrate backs the Re-calibrate
  button; calibration runs off the event loop (run_in_executor, bounded by
  asyncio.timeout(90)), persists the fields + refreshes the registry, and is
  graceful on a down/slow endpoint (never 500). Shared
  merge_calibration_into_caps helper used by the endpoint and the CLI.
- turnstone-admin rerank-calibrate is now per-model: --model <alias> required;
  --apply writes that model's caps (not the global setting). A no-separation
  result records the marker (calibrated, no floor) consistently across CLI and
  endpoint.

The serving lesson stays documented: Qwen3-Reranker needs vLLM --chat-template
or its scores are near-random (live-validated 0.6B + 4B: the calibrated floor
came out 0.95 vs 0.33 for the same task -- why per-model calibration exists).

Negative-tested: the floor-precedence branches (calibrated+separated -> per-
model, calibrated+!separated -> 0, uncalibrated -> global, no-alias -> global,
registry-must-not-be-consulted), endpoint persist/refresh + graceful failure,
detect-skips-non-rerankers, the CLI per-model write, and the caps-merge
preserving supports_rerank. Chip states verified via headless Chrome.
2026-06-01 15:47:14 -07:00
Patrick Buckley f6bae70ea6 feat(rerank): calibration CLI, 0-1 normalization, instruction support
Phase 2 of BM25 reranking (follows #627). Makes the rerank_bm25_threshold floor
usable across reranker models and adds tooling to pick it.

- normalize_scores (rerank.py): map a rerank batch into a 0-1 relevance
  probability -- sigmoid when any score falls outside [0,1] (logit endpoints
  like bge/TEI), identity otherwise (Cohere/Jina/Qwen already 0-1). Applied in
  the _bm25_reranker closure AND calibration so the threshold means the same on
  every endpoint. Monotonic, so ranking order is unchanged.

- rerank_calibrate.py + `turnstone-admin rerank-calibrate [--apply]`: probe the
  endpoint with labelled relevant/irrelevant groups, normalise, and recommend a
  recall-biased floor -- or report "no clean separation" (a mis-served/weak
  reranker). A warmup loop absorbs a cold endpoint's first-request compile so
  calibration doesn't time out. Validated live against Qwen3-Reranker 0.6B and
  4B: the calibrated floor differs sharply per model (~0.95 vs ~0.33 for the
  same task) -- exactly why per-endpoint calibration exists.

- rerank_config.py: extract resolve_rerank_client_from(config_store, registry);
  the alias/url precedence now lives in one place, shared by ChatSession (which
  delegates) and the CLI.

- tools.rerank_instruction (config + setting + client): wrap the query as
  <Instruct>:/<Query>: for instruction-aware rerankers (Qwen3) on endpoints that
  don't apply the model's own chat template. Docs note the critical vLLM serving
  detail: Qwen3-Reranker needs --chat-template or its scores are near-random and
  reranking hurts retrieval.

Negative-tested: normalize sigmoid/identity branches, closure-normalises-before-
floor, calibration separation/recall-bias/warmup-absorbs-cold-start, the CLI
apply/no-apply/no-separation paths, and instruction query-wrapping through the
real httpx boundary.
2026-06-01 14:44:30 -07:00
Patrick Buckley 215f7506ba feat(rerank): wire endpoint-backed reranking into BM25 retrieval surfaces
Reuse the shipped Cohere/Jina rerank client as an optional post-process on
the BM25 surfaces (tool search, skill search, memory composition) via one
seam: BM25Index gains an injected reranker + a two-stage search (BM25 recall
top-50 -> rerank -> top-k). No new storage.

Gated on a configured endpoint plus tools.rerank_bm25 (default on, matching
rerank_web_search). tools.rerank_bm25_threshold (default 0.0 = off) is a
relevance FLOOR for proactive memory surfacing: BM25 always returns something,
so without a floor every-turn memory injection spends tokens on the top-k of
whatever lexically matched; the reranker score is what makes a meaningful
"inject nothing" gate possible.

Two reranker modes (BM25Index rerank_filters):
- REORDER (reactive tool/skill search): the reranker must never drop results
  -> fall back to BM25 order on empty, backfill omitted pool items, so a
  misbehaving endpoint can't silently lose tools.
- FILTER (memory, rerank_filters = threshold > 0): a clean empty/short result
  is honoured (inject nothing) -- a deliberate divergence from
  web_search._rerank_results.
Parse/endpoint failure is a discrete branch from the floor: an empty result
for non-empty input means an unparseable response (a conforming reranker
scores every doc), so the closure raises RerankError and BM25Index falls back
to BM25 order in BOTH modes -- the floor only acts on valid scores.

Also: cap the rerank client timeout at 15s (the per-turn memory path can't
afford tools.timeout's 120s default); move the Reranker alias to rerank.py
(shared, no import cycle); document the endpoint egress in the rerank_bm25
help, the admin Reranker-role description, and docs/tools.md; add
scripts/bench_bm25_rerank.py (manual, needs a live endpoint) to measure
precision@k/MRR lift and recommend a threshold default.

Negative-tested: reorder fallback-on-empty and omitted-item backfill,
filter-mode honor-empty, singleton-still-floored, the parse-fail RerankError
raise, the >= floor boundary, and pool-position-to-doc-index mapping -- each
guard reverted to confirm its test fails, then restored.
2026-06-01 12:56:45 -07:00
Patrick Buckley c6e0293185 fix(rerank): harden web_search rerank fallback and correct stale docs
list()-materialize the reranker's output inside the guarded block so a
None / non-iterable / lazily-raising reranker falls back to native order
instead of raising out of web_search, and reject bool indices (an int
subclass) the same way _parse_hits already does.

Drop leftover web_fetch references from the rerank settings and Reranker
role help (reranking is wired into web_search only), and document both
endpoint paths (tools.reranker_alias and tools.rerank_url).
2026-06-01 11:01:30 -07:00
Patrick Buckley 6a0bc852d9 feat(rerank): endpoint-backed reranking for web_search
Reranking is delegated to an external Cohere/Jina-compatible /rerank endpoint
(self-hosted vLLM/TEI/llama.cpp, or hosted Cohere/Jina/Voyage); Turnstone runs
no reranker model itself. Disabled until an endpoint is configured.

- core/rerank.py: CohereJinaRerankClient (tolerant of results-wrapped and
  bare-list responses) + resolver.
- web_search: rerank the SearxNG result pool by query relevance before top-k,
  with a native-order fallback on error; answers/infoboxes untouched.
- Reranker as a model definition: add a model with the supports_rerank
  capability and pick it under Models -> Roles -> Reranker
  (tools.reranker_alias); takes precedence over the tools.rerank_url settings.

Settings: tools.rerank_url/model/api_key, tools.rerank_web_search,
tools.reranker_alias. Docs: docs/tools.md, turnstone.example.toml.

(web_fetch reranking was evaluated and dropped: for single-document chunk
selection it did not reliably beat head-truncation. Reranking is reserved for
multi-item ranking.)
2026-06-01 11:01:30 -07:00
Patrick Buckley 40b49c1d6c chore: bump version to 1.6.0a9 v1.6.0a9 2026-05-31 20:45:06 -07:00
Patrick Buckley 110d44b07e refactor(tools): remove man, math, and plan_agent built-in tools
`man` and `math` duplicated capabilities already reachable through
`bash`; `plan_agent` is better expressed as a `task_agent` running a
planning skill, and carried a large amount of special-case machinery
(plan-review gate, refinement loop, per-kind model routing). Removing
all three shrinks the tool surface and cuts per-call token cost.

Also removed, as dead-once-the-tools-are-gone:
- the `math` sandbox executor (`turnstone.core.sandbox`) and its
  `[sandbox]` extra; the eval analyst now runs bash-only
- the read-only `AGENT_TOOLS` sub-agent tool set and the `agent`
  tool-metadata key (`task_agent`/`TASK_AGENT_TOOLS` retained)
- the plan-review protocol end to end: the `on_plan_review` UI hook,
  `resolve_plan`, `POST /v1/api/plan` + `POST /v1/api/route/plan`,
  the `plan_review`/`plan_resolved` SSE events, and their Python SDK /
  TypeScript SDK / OpenAPI / frontend / Discord+Slack bindings
- the `model.plan_alias` / `model.plan_effort` settings and the
  registry `plan_model` / `plan_effort` routing fields

TOOLS 31->28, TASK_AGENT_TOOLS 13->11; COORDINATOR_TOOLS unchanged.

BREAKING CHANGE: removes the `man`, `math`, `plan_agent` tools, the
plan-review SSE/HTTP/SDK surface, and the plan_* model-routing settings
from the experimental 1.6 line.
2026-05-31 19:54:43 -07:00
Patrick Buckley 15b3aad815 feat(web-search): let the model pick a SearxNG category
Rename the web_search tool's `topic` parameter to `category` and expand the
enum to general/news/it/science, mapped to SearxNG `categories=`. The model can
now target the right corpus per query (e.g. `it` for code, `science` for
papers) — useful when generic engines rate-limit. The Tavily-era `finance`
topic (no SearxNG equivalent) is dropped. Threaded consistently through
_prepare_web_search / _exec_web_search / both search clients.

BREAKING: the web_search `topic` argument is now `category`.
2026-05-31 19:03:16 -07:00
Patrick Buckley eea8bf9e96 fix(bootstrap): keep "identical content" in the write-skip message
The SearxNG change reworded _tool_write_compose's duplicate-skip return and
dropped the "identical content" phrase that test_identical_content_skipped
asserts on, failing CI (which then fail-fast-cancelled the parallel matrix
jobs). Restore the phrase (now covering all three bundled files) and assert
the searxng/settings.yml extraction in test_writes_compose_file.
2026-05-31 19:03:16 -07:00
Patrick Buckley 1728a4c0af feat(web-search): replace Tavily/DuckDuckGo backends with self-hosted SearxNG
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single
self-hosted SearxNG service bundled into the docker-compose stacks.

Core:
- New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to
  (backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard
  unchanged. _resolve_search_client follows storage -> toml -> env -> default
  precedence (explicit "" disables, via ConfigStore.stored_keys()).
- Drop the Tavily-era topic=finance (no SearxNG category); topic is now
  general/news.

Settings/config:
- Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key.
- Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines,
  with get_searxng_url/get_searxng_engines.

Compose + bundled config:
- Internal-only searxng service (no published API port, :ro config, /healthz
  healthcheck, persistent searxng-cache volume) in both stacks; bundle
  turnstone/deploy/searxng/settings.yml (JSON output on, limiter off).
- Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in).
- bootstrap extractor + wheel packaging updated.

Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the
lxml/h2/brotli transitives).

Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG;
docs/docker.md carries the AGPL-3.0 §13 operator note.

BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg";
tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships
in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance.

Closes #545
2026-05-31 19:03:16 -07:00
Patrick Buckley effc60297c feat(install): add a one-line curl|bash installer
run.sh autodetects Ubuntu/Debian, Fedora/RHEL, Arch, and WSL; ensures git and
Docker; clones, builds, picks free ports (Caddy prefers 443, Postgres 5432),
generates a .env with a JWT secret and Postgres password, asks how many nodes to
run, and starts the stack. The node count persists via an auto-loaded
compose.override.yaml, so a later plain `docker compose up -d` keeps it; a fresh
clone with no override still starts all 10. Ignores the generated override.
2026-05-31 16:05:21 -07:00
Patrick Buckley aac782ba78 docs(readme): reframe around local-first/no-telemetry and add Discord
Lead with privacy, local-first, and no-telemetry; demote governance to an
optional team-controls line. Fix the dashboard URL (Caddy on 8443), add the
one-line installer, and link the Discord community.
2026-05-31 16:05:21 -07:00
Patrick Buckley 7fde6f9122 fix(bootstrap): extract the Caddyfile and refresh the wizard
The bundled production compose mounts ./Caddyfile, so write_compose has to write
it alongside compose.yaml or `docker compose up` fails to start Caddy. Update
the wizard's system prompt for the new model (no profiles, Caddy-fronted
dashboard, current ports).
2026-05-31 16:05:21 -07:00
Patrick Buckley 8f02e688a7 fix(channels): stand by instead of exiting when no adapter token is set
Without a Discord/Slack token the channel gateway exited, which crash-loops
under `restart: unless-stopped`. Run the HTTP server and service heartbeat with
zero adapters instead — registered and idle — until a token is set. The Slack
token-pair mismatch stays a hard error.
2026-05-31 16:05:21 -07:00
Patrick Buckley cc4d9b1a8d fix(server): boot nodes with no models in a degraded state
A node refused to start with no model configured and no LLM reachable, so a
fresh cluster couldn't come up to be configured. load_model_registry now
accepts allow_empty and returns an empty registry; ModelRegistry permits the
empty state (default unset); the server passes allow_empty so a node registers
and shows in the console, then picks up models added in the admin UI live. The
CLI keeps failing fast — a REPL with no model is unusable.
2026-05-31 16:05:21 -07:00
Patrick Buckley 842826501d fix(storage): prevent concurrent-migration deadlock on the advisory lock
Migrations run on every node at boot, and one migration rebuilds an index with
CREATE INDEX CONCURRENTLY, which can't run in a transaction and waits for all
concurrent transactions to drain. The advisory lock that serialises migrations
was held inside an open transaction, so the lock-holder's own idle-in-
transaction connection deadlocked the concurrent build when several nodes
started together. Acquire the lock on an AUTOCOMMIT connection and poll
pg_try_advisory_lock so no waiter pins a snapshot. Adds a Postgres concurrency
regression test (skipped on SQLite).
2026-05-31 16:05:21 -07:00
Patrick Buckley 631f1b0021 feat(compose): cluster-by-default Caddy-fronted stack with bare-metal join
`docker compose up` from a clone builds one image and brings up the whole stack
— PostgreSQL, console, Caddy, channel, and 10 server nodes — sharing one
Postgres so the console discovers every node. The dashboard is reachable only
through Caddy (HTTP/2 avoids the browser's 6-connection cap on the dashboard's
SSE streams); the console's plain-HTTP port is no longer published. Postgres
binds 127.0.0.1 so a bare-metal turnstone-server can join the cluster — the
bare-metal overlay is folded in and removed. Insecure dev defaults keep it
zero-config; the bundled production stack mirrors the shape but pulls ghcr
images and requires real secrets.

Move the Caddyfile under turnstone/deploy so it ships in the wheel; update docs,
QUICKSTART, and .env.example to match.
2026-05-31 16:05:21 -07:00
Patrick Buckley c58df26a30 style(tests): ruff-format provider/registry empty-string tests 2026-05-31 13:06:24 -07:00
chrismuzyn 8737081373 add tests for anthropic api 2026-05-31 13:06:24 -07:00
chrismuzyn 85db6895e3 make sure empty strings don't get passed to the openapi sdk which don't
allow env var fallback
2026-05-31 13:06:24 -07:00
Patrick Buckley e322b69c8b chore: bump version to 1.6.0a8 v1.6.0a8 2026-05-31 00:58:44 -07:00
Patrick Buckley 1580e8e64b style(tests): ruff-format the appended early-paint guards
The guard tests were appended via heredoc, bypassing the editor's
auto-format; ruff-format collapses one wrapped call onto a single line.
No behaviour change.
2026-05-31 00:57:13 -07:00
Patrick Buckley 44f19f1401 feat(judge,console,ui): paint pending tool calls before the intent verdict
The intent-validation judge runs before the approval gate resolves, and
Smart Approvals (judge.smart_approvals) parks approve_tools on the async LLM
verdict for up to judge.timeout — so the tool-call card never reached the UI
until the judge had ruled. An operator could not see a committed call, let
alone Stop it, during that window.

approve_tools now emits a tool_pending event carrying the serialized batch at
the top of the gate, before the tool-policy lookup, the verdict wait, and the
human prompt. It is a UI paint only — no persistence, audit, or verdict
bookkeeping — so it cannot perturb the gate's accounting. The authoritative
tool_info / approve_request / tool_result events that follow upgrade the same
construct in place, keyed by call_id, and the Last-Event-ID replay slice
reconstructs it on reconnect. A ToolPendingEvent joins the SDK registry.

Coordinator: appendToolBatch was already idempotent on call_ids; the new
handler reuses the --running placeholder it already upgrades, with an
"Evaluating" kicker that swaps to "Running" on the auto-approve upgrade.

Interactive: showInlineToolBlock was create-only, so a second card would
duplicate. Added announceToolBlock + _takeAnnouncedBlock to reuse the
announced shell (matched on its call_id set) instead. The announced rail is
dashed amber and must out-specify the .msg.ts-approval--inline cyan-hold
(specificity 0,2,0) — at 0,1,0 it rendered cyan, indistinguishable from a
normal card — so the announced card is the one visually distinct surface in
the stream.

Screen-reader parity: the early paint announces politely through dedicated
off-screen aria-live regions on both surfaces (the messages log is
aria-live=off mid-stream, so the appended shell alone is inaudible), and the
announced shell carries aria-busy until the upgrade clears it. Polite, not
assertive — the human gate keeps its assertive announcement.

Tests cover the gate ordering (tool_pending precedes tool_info and the Smart
Approvals gate) plus string-guards on both UIs' wiring, the announced-rail
specificity, and the screen-reader regions.
2026-05-31 00:57:13 -07:00
Patrick Buckley 83ab25e611 fix(ui): normalize risk_level before className/data-risk interpolation (#562)
risk_level is server-supplied and was interpolated straight into className
and data-risk at three sites — updateVerdictBadge, _buildOutputWarningEl,
renderVerdictBadge — as `risk_level || "medium"`. Whitespace, a stray case,
or a future relaxed-validation value would pass into the class string and
silently break the selectors that updateVerdictBadge, toggleVerdictDetail,
and the d-key handler rely on. It is not an XSS vector (className assignment
is text-typed), but a broken selector is a real failure.

Funnel all three through a normalizeRiskLevel() chokepoint backed by a
{low, medium, high, critical} allowlist; unknown or blank falls back to the
neutral "medium" default. Pre-existing; surfaced while rendering verdict
badges from the early-paint path.
2026-05-31 00:57:13 -07:00
Patrick Buckley da07554693 docs(judge): correct Smart Approvals heuristic-floor wording
The floor blocks only explicit heuristic deny/critical verdicts — it is
not a general "never lower the heuristic" rule. The heuristic default for
an unmatched tool is `review`, and letting a confident LLM `approve`
upgrade a `review` is the feature's purpose. Matches the implementation
and addresses PR review feedback.
2026-05-30 19:48:17 -07:00