mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
eval/skill-adherence
175 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7053439e84 |
refactor(eval): split measurement core from prompt optimizer
turnstone-eval was misnamed: it was a prompt optimizer, not a measurement harness. Split the 3252-line turnstone/eval.py into a strictly one-way dependency (optimizer -> eval-core; core never imports the optimizer): - turnstone/eval/core.py measurement substrate — everything up to and including _run_iteration: provider detection, NullUI, HeadlessSession, the test runner, score_run, aggregation, and neutral reporting. - turnstone/eval/cli.py new measure-only `turnstone-eval` — the old --no-optimize path promoted to the whole job (one _run_iteration call, then print the summary table). - turnstone/optimizer.py the UCB self-modify loop and its multi-agent pipeline (analyst/optimizer/observer/diversifier/tool optimizer), now `turnstone-optimizer`; imports from eval.core only. - turnstone/eval/__init__.py re-exports the core public API for back-compat (score_run, _match_action, _run_iteration, HeadlessSession). _apply_tool_overrides lives in core (HeadlessSession needs it) rather than alongside the other tree helpers, so the dependency stays one-way. Breaking change: `turnstone-eval` now measures; use `turnstone-optimizer` to optimize. Both code paths are behaviour-preserving — the moved function bodies are byte-identical. |
||
|
|
dbf389783e |
refactor(personas): file-backed built-in prompts, explicit source column
Built-in persona base prompts move from inline DB text / base.md into prompts/personas/<slug>.md — code-owned, PR-reviewable, drift-proof. base.md / base_coordinator.md become personas/engineer.md / orchestrator.md. Prompt source is now explicit in storage instead of inferred in app logic: a new base_prompt_file column plus CHECK (base_prompt IS NOT NULL OR base_prompt_file IS NOT NULL) — two nullable columns, never both empty. Resolution is a coalesce (base_prompt else load(base_prompt_file)), frozen into the workstream stamp at creation. base_prompt_file marks a persona as built-in (code-only, un-archivable); an operator override on a built-in is allowed and wins over the file. "Inherit the kind default" is a workstream-creation act (is_default), not a persona-row state. Migration 063: - seeds reference their file (base_prompt NULL); no runtime file reads — the backfill's frozen prompt text is inlined as a point-in-time snapshot so migration history stays self-contained and reproducible. - every existing workstream is stamped by kind (creative -> writer, else the kind default), set-based (INSERT..SELECT via temp tables) with the persona column added after the bulk writes to shorten its lock window. Storage guards (both backends): operators must supply base_prompt; built-ins can't be archived or have base_prompt_file set via the API; clearing an operator persona's only source is rejected. Follow-ups reviewed alongside (#756): soft-set visibility docstring scoped to per-process; _apply_persona_snapshot / _current_persona_snapshot own the stamp round-trip; spawn approval-header args (skill/name/target_node) flattened+capped like persona; server-side tool injection generalized to replace-only (client-def gated, incl. the xAI include forwarding). Seed copy revised (researcher soft; de-costumed prose; engineer de-biased). New test_schema_parity asserts create_all matches the alembic head. Closes #683 groundwork; ruff + strict mypy clean, full suite green. |
||
|
|
b65e5cae0e |
docs(personas): accuracy sweep — spec models, protocol contracts, page corrections
Spec models now describe what the endpoints do: ListPersonasResponse declares the tool_inventory the shelf depends on, both console create models declare persona, CreatePersonaRequest declares org_id, and UpdatePersonaRequest documents the null-vs-absent split (null clears base_prompt/tool_allowlist, null on flags/kinds is ignored). Console OpenAPI regenerated. Protocol contracts match the implementations: update_persona's return covers the no-op case, create_persona's raises-list is complete, and both extended row-shape docstrings gain their tail columns plus the append-only rule. The workstreams.persona comments say slug, not display name. Page corrections from the docs review: personas.md documents the creative_mode-to-writer migration conversion, the mid-session /resume MCP-lever behavior, visibility-based nudge gating, the soft-set prompt-cache cost, and the executive tool list — and drops internal jargon. The changelog entry moves under [Unreleased] with the house breaking-marker style and the auto-conversion note. coordinator-skills and the API tour stop using persona to mean framing; governance, api-reference, sdk, console, tools, and memory pick up the new permission family, endpoints, kwargs, picker, and lever caveats. |
||
|
|
0d6d7ebae1 |
docs(personas): concept doc, CHANGELOG 1.7 entry with /creative breaking note
docs/personas.md covers the four levers, the resolve-once/stamp-forever snapshot semantics, the seed matrix, per-surface selection, authoring rules, and RBAC; architecture.md's config-persistence paragraph swaps the removed creative_mode for the persona stamp. |
||
|
|
c0be383f99 |
refactor(doctor): replace turnstone-bootstrap with turnstone-doctor (#718)
* refactor(doctor): replace turnstone-bootstrap with turnstone-doctor turnstone-bootstrap was an LLM setup wizard for Day-0; run.sh now owns install. Repurpose its LLM/conversation plumbing into turnstone-doctor — a diagnose-only tool for a running cluster. - Preflight detects the install kind (docker-compose/systemd/pip/source) from config.toml + TURNSTONE_* env, with secret redaction. - Self-configuring brain resolves the cluster's own model from config/env/storage read-only (no migrations, no create_all), falling back to interactive selection; the attempt itself is the LLM-backend health check. - Deterministic version check: installed version, cluster drift via the console's authoritative /health, and latest upstream stable/experimental (offline-safe). - Read-only diagnostic tools (read_file, compose/systemd/journal, http_health, check_llm_backend, node_health, finish) behind one secret-scrubbing chokepoint; no generic shell, so read-only is structural. - node_health reaches a node the right way for the detected install kind (exec-into-container for compose, direct HTTP otherwise), overridable per node for mixed clusters. - mTLS-aware: forwards [database] SSL params and reports node-mesh mTLS instead of mislabelling healthy nodes "unreachable". init_storage gains a backward-compatible create_tables override for read-only opens. Entry point turnstone-bootstrap -> turnstone-doctor; README/QUICKSTART/ architecture/docker docs, the bundled compose header, run.sh, and the CI smoke updated. CHANGELOG deferred. * fix(doctor): address Copilot + CodeQL review findings on #718 Validated all seven review findings (none false positives) and fixed: - check_llm_backend now applies the same scheme / metadata-host guard as http_health (extracted to _assert_safe_http_url), so a model-supplied base_url can't be steered at the cloud metadata endpoint or a file:// URL. - node_health no longer double-appends the default port when the operator passes host:port (regression: 10.0.0.5:8081 -> http://10.0.0.5:8081:8080). - node_health install_type enum uses "git-source" to match the label the rest of the module and the prompt/report show the model (a schema-strict provider would otherwise reject the value the model is told to use). - _read_api_creds takes base_url + api_key as a unit from the first config source that defines either field, then env-fills, instead of splicing the two across different config files into a pair that exists in no real config. - _mask_secrets masks assignment-shaped content inside comment lines, so a commented-out real secret can't leak through read_file / the report; prose comments (no KEY=value shape) still pass through untouched. - drop the mixed import styles CodeQL flagged in doctor.py and test_doctor.py. Adds 5 tests; ruff + mypy clean; full doctor suite passes (129). |
||
|
|
776430d860 |
feat(cancel): honest cancellation dispositions + coordinator subtree propagation
A cancelled agent previously discarded its own ledger and reported a bare "(task interrupted by user)" — fabricating the *outcome* (read downstream as "nothing happened"), which invites a double-send as readily as a dropped record causes an orphan. Make the fold-back honest, and propagate an owner's cancel down the coordinator subtree. - task_agent (single + parallel): on cancel, fold back a deterministic disposition built from the agent's in-memory ledger — actions completed, the in-flight action flagged outcome-UNKNOWN, and not-started calls — instead of the opaque interrupted string. - coordinator cancel now auto-propagates to its direct children via a post_cancel hook on the shared cancel handler (cooperative fan-out; no blocking drain). - synthesized cancelled tool results now read outcome-UNKNOWN rather than implying the call never ran. - remove the now-redundant stop_cascade operator endpoint (handler, route, OpenAPI spec + schema, tests, docs); a coordinator cancel supersedes it. |
||
|
|
0b4f77db33 |
fix(judge): daemon-thread call deadlines; raise local-model timeouts
The judges and the regex ReDoS probe ran a blocking call on a ThreadPoolExecutor and abandoned the worker with shutdown(wait=False) on timeout or cancel. concurrent.futures joins every executor worker from an atexit hook regardless of wait=False, so a wedged call could pin interpreter exit — and hang the test suite at shutdown. Add turnstone/core/deadline.py::run_with_deadline: run a blocking callable on a daemon thread bounded by a wall-clock timeout and an optional cancel event. A daemon worker is never joined at exit, so abandoning one is safe. Migrate three sites onto it: - OutputGuardJudge.evaluate() - IntentJudge._evaluate_single / _run_judge — this also removes _ExecutorPoisonedError and the executor-restart dance: per-call daemon threads can't poison a shared single-slot pool, so a timeout now returns None and the caller delivers one fallback verdict. - console/server.py _validate_regex_pattern (regex ReDoS probe) Also: - Double the default judge LLM timeouts for slower local models: judge.timeout 60->120s and judge.output_guard_llm_timeout 30->60s (settings registry, JudgeConfig dataclass, --judge-timeout CLI default, class docstring, docs). Correct a stale doc that described the per-turn timeout as a total budget across turns. - Raise the regex probe bound 0.5->3.0s so a legitimately complex pattern isn't false-flagged as catastrophic backtracking. - CI: run pytest with -v instead of -q so a hang names the offending test instead of riding the job timeout. - Tests: cover deadline.py and the regex validator; move test_judge.py off fixed sleeps onto the existing _wait_for helper. |
||
|
|
04b3a3abe4 |
feat(deploy): systemd units for a bare-metal turnstone-server node
Hardened service + slice + node-identity drop-in template + a README for running a turnstone-server outside Docker that joins the compose cluster — the production-shaped counterpart to the one-liner in docs/docker.md. Secrets stay in config.toml; per-host identity + cluster URLs go in the drop-in. The README notes the cross-host mTLS caveat (turnstonelabs/lacme#22). |
||
|
|
1f61350545 |
feat(compose): let bare-metal turnstone-servers join the cluster (incl. mTLS)
A turnstone-server running outside the compose network ("bare-metal", e.g. a
local-GPU box) couldn't fully join: it can't resolve the in-cluster console
(console:8090) to enroll its mTLS cert, and SearxNG was unreachable for
web_search. Only Postgres was published.
Publish the console's plain-HTTP ACME endpoint (:8090) and SearxNG (:8081)
alongside Postgres, all bound via one knob TURNSTONE_HOST_IP (default 127.0.0.1
-- nothing new on the LAN; set it to the host's LAN IP for a node on another
machine). Postgres keeps honoring the legacy POSTGRES_BIND as a fallback, so
existing .env files don't break.
The node's TLS client now honors TURNSTONE_CONSOLE_URL so a bare-metal node can
point at the published ACME endpoint instead of the unreachable in-cluster name
(empty = in-cluster service discovery, unchanged).
Docs (docker.md, tls.md), the run.sh-generated .env, and the bootstrap wizard
updated to match. The advertised host is the cert's primary SAN and the console
collector dials it back, so mTLS hostname verification holds both ways.
|
||
|
|
108714a48d |
fix(auth): isolate server/console session cookies by name
The server (:8080) and console (:8090) both set a cookie named `turnstone_auth`. Cookies ignore port (RFC 6265), so on a shared host (localhost dev, the Electron build, single-box installs) logging into one surface overwrote the other's cookie and 401'd the first session. Give each surface its own cookie name -- `turnstone_auth_server` / `turnstone_auth_console` -- threaded as a required `cookie_name` argument through the cookie builders, `check_request`, `AuthMiddleware`, and the six shared auth handlers (login/logout/setup/whoami/refresh/oidc_callback). Each app passes its own constant; the parameter is required (no default) so a forgotten caller fails loudly instead of silently reverting to the legacy name. Names key on role, not node: the cluster shares one JWT identity and the console->node proxy re-mints a bearer token (dropping Set-Cookie), so per-instance names would break identity portability and aren't used. Hard cutover: the legacy `turnstone_auth` cookie is no longer read and self-expires within its 24h TTL (one forced re-login). JWT audience was already enforced, so the shared cookie was a session clobber, not an auth bypass. |
||
|
|
30b590fb25 |
feat(memory): durable per-user coordinator scope + anonymous-coordinator guard
The coordinator memory scope was keyed by the session's ws_id, so every new coordinator session started with an empty namespace and its rows were orphaned on close — coordinator memory never actually persisted. Re-key the scope to the coordinator's creator user_id: one durable orchestration namespace per user, shared by all of that user's coordinator sessions (concurrent ones included; upsert-by-name is the collision rule). The child-containment threat model is unchanged: the gate is session KIND — children are always interactive and share the parent's user_id, so _validate_scope rejects them before scope resolution, and the REST memories API still rejects the coordinator scope outright. The implicit visibility lane now also fails closed on an empty scope_id to match the explicit search/list lanes (the storage helpers treat a falsy scope_id as 'no scope_id filter', which would have read every user's rows). Anonymous coordinators are no longer constructible: ChatSession refuses kind=COORDINATOR with an empty user_id at the constructor — the single choke point covering create, rehydration of legacy rows (surfaced by the open handler as a 503 with remediation text), and any future host — and the console no longer masks an empty uid as a phantom 'system' principal when minting coordinator JWTs, per CoordinatorTokenManager's documented 'sub = the real creator user_id' contract. Migration 061 carries existing coordinator rows across: rows whose owning workstream is gone or ownerless are deleted (unreachable under user keying), same-name collisions within a user keep the newest updated row (memory_id tiebreak), and survivors re-key to the owner's user_id. |
||
|
|
7ef04e576a |
fix(providers): require base_url for anthropic-compatible
Copilot review on #661: empty base_url let the SDK fall back to https://api.anthropic.com, sending compat-shaped requests to the commercial API. The lane is local-only by definition, and the /v1-strip edge case already established fail-loudly-over-silent-prod-retarget; apply the same principle to the empty case. create_client raises an actionable ValueError; the admin Detect path surfaces it as a clean error string via probe_model_endpoint's existing handler. |
||
|
|
12bd848c68 |
feat(providers): anthropic-compatible lane for local /v1/messages servers
Add provider id "anthropic-compatible": the existing AnthropicProvider pointed at Anthropic-compatible local servers (vLLM /v1/messages), mirroring the openai/openai-compatible split. Registry-only — configured via the admin Models tab or [models.*] toml, not exposed on the bare --provider flag, so the CLI/server prod-URL defaults are unreachable for the lane and real-Anthropic behavior is untouched. Lane behavior (live-verified against vLLM 0.22.1rc1 + DeepSeek-V4-Flash): - Capability defaults replace the Claude static table: token_param max_tokens, thinking_mode none, web_search/tool_search/vision off, reasoning replay on. vLLM rejects Anthropic server-side tool types (tools require input_schema) and ignores the thinking request param, so neither is sent; thinking blocks still stream back and round-trip through the native lane verbatim. - Reasoning toggles via server_compat extra_body chat_template_kwargs (first-class vLLM request field; request-level keys beat server defaults). _build_thinking_and_kwargs forwards non-internal extra_params as SDK extra_body; thinking_budget_tokens stays internal. - No temperature force: thinking_mode none skips the Claude-only temperature=1.0 requirement. Admin UI: provider option + URL placeholder (base_url without /v1 — the SDK appends /v1/messages); the server-compat section shows only the extra-body field for the lane. thinking_mode round-trips through the form dropdown for every provider except anthropic-compatible, where it stays in the raw capabilities JSON — the edit-load lift and save restore use the same predicate so stored overrides are never silently dropped. Docs: architecture.md gains the lane subsection incl. verified quirks (thinking param dropped by vLLM; stop_sequences cut inside thinking and report end_turn; usage has no cache fields; images need a multimodal model; mid-conversation system turns are per-model opt-in). Negative-tested: removing the _INTERNAL_EXTRA_PARAMS exclusion fails test_internal_keys_not_leaked; the live test drives a streamed turn with the chat_template_kwargs toggle and asserts no reasoning deltas. |
||
|
|
f44886a55f |
docs: 1.6.0 changelog + release-track policy (#653)
* docs: 1.6.0 changelog — roll up the 1.5→1.6 line for stable Replaces [Unreleased] with the 1.6.0 section: 320 main-only commits since the stable/1.5 divergence grouped into theme bullets (license, trajectory/migration-060, web search, rerank/memory, approvals/judge, L-shell, shelf, SSE, providers, cluster ops, security). Breaking changes aggregated up top; migration-060 backup callout reshaped from discussion #631 for the stable audience. * docs: add the stable/1.6 track to the changelog preamble * docs: retire the stable/1.4 track — current + one prior policy Changelog preamble down to three tracks with the policy stated; 1.4 retirement noted in the 1.6.0 Removed section (final release v1.4.0; tags/artifacts remain, BUSL-1.1 as shipped). releasing.md track table, policy bullet, and examples brought up to the 1.6.0 promote cycle — the doc was still describing the 1.4-stable era. |
||
|
|
d0e9aa3dbe |
fix(console): coordinator ws-ref validation, did-you-mean recovery, wait fail-fast
Field incident: the coordinator LLM hand-copied a child ws_id and collapsed its aaa run to a, producing a 30-char id. inspect said "not found", wait called it "denied", neither offered recovery, and the model concluded the child was dead and dropped the lane — silent report degradation while the child kept working. - validate model-supplied ws_id args at the tool boundary (send/close/cancel/delete/inspect/wait): full 32-hex ids pass through at unchanged storage cost; a child's exact legacy id still resolves; anything else fails fast with a did-you-mean (capped Levenshtein <=3 over the coord's own children) plus a child roster. Near-misses never auto-resolve; display names are not addresses (mutable, non-unique) — a name ref errors with a pointer at the right id - wait_for_workstream: rename per-entry state "denied" -> "not_found" with an honest sentinel; malformed refs error before any waiting (invalid_ws_ids); a well-formed id that is foreign, missing, or hard-deleted mid-wait aborts the wait on the tick that observes it instead of burning the timeout (mode=all was unsatisfiable) or riding along to complete=True (silent lane loss); mode=all completes only when every id is real-terminal; entries carry the child display name - one not-found payload across all verbs: foreign and nonexistent stay byte-identical (no existence oracle), hints reference only the coord's own children, echoed refs clipped in error strings; invalid_ws_ids and not_found share one per-ref shape with the roster hoisted top-level - inspect ownership now requires user_id parity via _row_in_own_subtree, matching the wait/mutating gates (#506) — closes the forged-parent cross-tenant read - session exec serializes the structured recovery payload (results + did_you_mean + children) on unresolvable-id wait errors instead of collapsing to the bare error string - tool JSON descriptions + coordinator docs updated to the new contract; incident regression test pins the captured aaa-collapse ids |
||
|
|
81a3eaecce |
chore: relicense BUSL-1.1 → Apache 2.0 for 1.6.0 (#651)
* chore: relicense BUSL-1.1 -> Apache 2.0 for 1.6.0 Flips every license artifact in the tree; 1.5.x and earlier remain BUSL-1.1 per their release-time LICENSE files. Contributor consent record: #548 (rationale: #546). - LICENSE: canonical Apache 2.0 text - NOTICE: new; copyright line + pointer to THIRD-PARTY-NOTICES - pyproject.toml: SPDX expression + explicit license-files trio - Dockerfile: COPY the license trio (hatchling needs them at build) - THIRD-PARTY-NOTICES: BUSL line reworded; bundled-version drift fixed (KaTeX 0.17.0, Mermaid 11.15.0, hls.js 1.6.16) - README badge + License section, CONTRIBUTING inbound-license line, TS SDK package(+lock), example pyproject - docs/pgbouncer.md: drop stray ':' introduced in #353 * docs: add CONTRIBUTORS.md * chore: drop LICENSE leading blank line The apache.org LICENSE-2.0.txt begins with a newline; the SPDX canonical text and GitHub license templates do not. Use the conventional form — detection is whitespace-normalized either way. |
||
|
|
b24b029c4d |
refactor(storage): single-source the bulk-insert race rationale
Review follow-up: the ON CONFLICT rationale lived verbatim in three places (protocol docstring + both backend comments). Keep the prose in the protocol — the contract's home — and point the backends at it. Also recommend cancel_on_approval=true in docs for deployments where the judge shares one local inference backend with the session model. |
||
|
|
c75afd704b | docs(judge): document verdict lifecycle — run-to-completion, superseded rows, replay parity | ||
|
|
a3ff07a86d |
fix(tls): mTLS-aware container healthcheck + boot-time init retry
A whole-stack restart races every node against the console for the CA fetch (compose re-enforces depends_on ordering only on `up`): losers logged one warning and served plain HTTP for their lifetime, while winners served mTLS that the plain-HTTP container healthcheck could never probe — leaving "healthy" plaintext nodes and "unhealthy" working ones. - TLSClient.init() grows attempts/base_delay retry (server passes 6 attempts, ~31 s backoff) absorbing the boot race; per-attempt CA-fetch failures log warning + debug traceback instead of error tracebacks. - healthcheck.py falls back to HTTPS when the plain probe fails, presenting the node's own cert as the client cert with the cluster CA pinned; dials localhost because the internal CA issues DNS SANs only. Default plain-HTTP deployments are unchanged. - The server writes boot PEMs under a fixed root (TURNSTONE_TLS_PEM_DIR, default <tmpdir>/turnstone-tls) so the probe can find them; boot clears stale dirs and refuses a symlinked/foreign-owned root; renewal rewrites the PEM dir so the probe's client cert never outlives the served cert. - /health reports tls: "active"|"fallback" (absent when TLS is disabled) so a silently downgraded node is observable. |
||
|
|
b3c1b9c9e0 |
build: promote anthropic, postgres, console, tls to core dependencies
The Anthropic SDK provider was the lone first-class provider gated behind an optional extra, while OpenAI ships in core and Google rides the OpenAI-compatible path. Fold anthropic, psycopg (postgres), croniter (console), and lacme (tls) into the base dependency set so a default `pip install turnstone` yields a complete single- or multi-node deployment; only the Discord/Slack channel gateways stay optional. - pyproject: four extras → base deps; `all` is now discord+slack; drop the redundant croniter from the `test` extra; regenerate uv.lock. - ci: the postgres test job installs `.[test]` (psycopg is base now). - providers: `_ensure_anthropic` becomes a thin SDK accessor for `create_client`; drop the now-redundant eager import-guard calls from the streaming/completion hot path (anthropic is always present). - bootstrap: import anthropic directly. - tests/docs: drop the anthropic importorskips and stale extra-install hints. |
||
|
|
8b4b8b3fd5 |
refactor(rerank): reranker is a per-model definition only (drop global endpoint settings)
The reranker_alias -> model-definition path (added when reranking became a model role) made the older global endpoint settings redundant. Resolve reranking solely through the Reranker role and remove the parallel global config. - Removed settings tools.rerank_url / rerank_model / rerank_api_key, their config.py getters (+ $TURNSTONE_RERANK_URL / $TURNSTONE_RERANK_MODEL and the module caches), and the fallback branch in resolve_rerank_client_from. The resolver now returns a client only when a Reranker model (capability supports_rerank, base_url = its /rerank endpoint) is selected, else None. - Kept as global knobs: reranker_alias (the selector), rerank_web_search, rerank_bm25, rerank_bm25_threshold, and rerank_instruction -- a task-level query knob (Qwen3-style), not endpoint identity. - The Settings tab is registry-driven, so the three fields disappear with their SettingDefs. Updated the Reranker role help, example config, and docs/tools.md. BREAKING: a reranker configured via [tools] rerank_url (config.toml / env / Settings tab) no longer works -- add the reranker in the admin Models tab and pick it under Models -> Roles -> Reranker. No migration: reranking is days old and disabled by default, so any orphaned tools.rerank_* config rows are inert. Tests: the resolver covers no-store / no-alias / non-rerank-alias -> None and the model-definition happy path; the obsolete global-fallback tests are removed. |
||
|
|
f6bae70ea6 |
feat(rerank): calibration CLI, 0-1 normalization, instruction support
Phase 2 of BM25 reranking (follows #627). Makes the rerank_bm25_threshold floor usable across reranker models and adds tooling to pick it. - normalize_scores (rerank.py): map a rerank batch into a 0-1 relevance probability -- sigmoid when any score falls outside [0,1] (logit endpoints like bge/TEI), identity otherwise (Cohere/Jina/Qwen already 0-1). Applied in the _bm25_reranker closure AND calibration so the threshold means the same on every endpoint. Monotonic, so ranking order is unchanged. - rerank_calibrate.py + `turnstone-admin rerank-calibrate [--apply]`: probe the endpoint with labelled relevant/irrelevant groups, normalise, and recommend a recall-biased floor -- or report "no clean separation" (a mis-served/weak reranker). A warmup loop absorbs a cold endpoint's first-request compile so calibration doesn't time out. Validated live against Qwen3-Reranker 0.6B and 4B: the calibrated floor differs sharply per model (~0.95 vs ~0.33 for the same task) -- exactly why per-endpoint calibration exists. - rerank_config.py: extract resolve_rerank_client_from(config_store, registry); the alias/url precedence now lives in one place, shared by ChatSession (which delegates) and the CLI. - tools.rerank_instruction (config + setting + client): wrap the query as <Instruct>:/<Query>: for instruction-aware rerankers (Qwen3) on endpoints that don't apply the model's own chat template. Docs note the critical vLLM serving detail: Qwen3-Reranker needs --chat-template or its scores are near-random and reranking hurts retrieval. Negative-tested: normalize sigmoid/identity branches, closure-normalises-before- floor, calibration separation/recall-bias/warmup-absorbs-cold-start, the CLI apply/no-apply/no-separation paths, and instruction query-wrapping through the real httpx boundary. |
||
|
|
215f7506ba |
feat(rerank): wire endpoint-backed reranking into BM25 retrieval surfaces
Reuse the shipped Cohere/Jina rerank client as an optional post-process on the BM25 surfaces (tool search, skill search, memory composition) via one seam: BM25Index gains an injected reranker + a two-stage search (BM25 recall top-50 -> rerank -> top-k). No new storage. Gated on a configured endpoint plus tools.rerank_bm25 (default on, matching rerank_web_search). tools.rerank_bm25_threshold (default 0.0 = off) is a relevance FLOOR for proactive memory surfacing: BM25 always returns something, so without a floor every-turn memory injection spends tokens on the top-k of whatever lexically matched; the reranker score is what makes a meaningful "inject nothing" gate possible. Two reranker modes (BM25Index rerank_filters): - REORDER (reactive tool/skill search): the reranker must never drop results -> fall back to BM25 order on empty, backfill omitted pool items, so a misbehaving endpoint can't silently lose tools. - FILTER (memory, rerank_filters = threshold > 0): a clean empty/short result is honoured (inject nothing) -- a deliberate divergence from web_search._rerank_results. Parse/endpoint failure is a discrete branch from the floor: an empty result for non-empty input means an unparseable response (a conforming reranker scores every doc), so the closure raises RerankError and BM25Index falls back to BM25 order in BOTH modes -- the floor only acts on valid scores. Also: cap the rerank client timeout at 15s (the per-turn memory path can't afford tools.timeout's 120s default); move the Reranker alias to rerank.py (shared, no import cycle); document the endpoint egress in the rerank_bm25 help, the admin Reranker-role description, and docs/tools.md; add scripts/bench_bm25_rerank.py (manual, needs a live endpoint) to measure precision@k/MRR lift and recommend a threshold default. Negative-tested: reorder fallback-on-empty and omitted-item backfill, filter-mode honor-empty, singleton-still-floored, the parse-fail RerankError raise, the >= floor boundary, and pool-position-to-doc-index mapping -- each guard reverted to confirm its test fails, then restored. |
||
|
|
6a0bc852d9 |
feat(rerank): endpoint-backed reranking for web_search
Reranking is delegated to an external Cohere/Jina-compatible /rerank endpoint (self-hosted vLLM/TEI/llama.cpp, or hosted Cohere/Jina/Voyage); Turnstone runs no reranker model itself. Disabled until an endpoint is configured. - core/rerank.py: CohereJinaRerankClient (tolerant of results-wrapped and bare-list responses) + resolver. - web_search: rerank the SearxNG result pool by query relevance before top-k, with a native-order fallback on error; answers/infoboxes untouched. - Reranker as a model definition: add a model with the supports_rerank capability and pick it under Models -> Roles -> Reranker (tools.reranker_alias); takes precedence over the tools.rerank_url settings. Settings: tools.rerank_url/model/api_key, tools.rerank_web_search, tools.reranker_alias. Docs: docs/tools.md, turnstone.example.toml. (web_fetch reranking was evaluated and dropped: for single-document chunk selection it did not reliably beat head-truncation. Reranking is reserved for multi-item ranking.) |
||
|
|
110d44b07e |
refactor(tools): remove man, math, and plan_agent built-in tools
`man` and `math` duplicated capabilities already reachable through `bash`; `plan_agent` is better expressed as a `task_agent` running a planning skill, and carried a large amount of special-case machinery (plan-review gate, refinement loop, per-kind model routing). Removing all three shrinks the tool surface and cuts per-call token cost. Also removed, as dead-once-the-tools-are-gone: - the `math` sandbox executor (`turnstone.core.sandbox`) and its `[sandbox]` extra; the eval analyst now runs bash-only - the read-only `AGENT_TOOLS` sub-agent tool set and the `agent` tool-metadata key (`task_agent`/`TASK_AGENT_TOOLS` retained) - the plan-review protocol end to end: the `on_plan_review` UI hook, `resolve_plan`, `POST /v1/api/plan` + `POST /v1/api/route/plan`, the `plan_review`/`plan_resolved` SSE events, and their Python SDK / TypeScript SDK / OpenAPI / frontend / Discord+Slack bindings - the `model.plan_alias` / `model.plan_effort` settings and the registry `plan_model` / `plan_effort` routing fields TOOLS 31->28, TASK_AGENT_TOOLS 13->11; COORDINATOR_TOOLS unchanged. BREAKING CHANGE: removes the `man`, `math`, `plan_agent` tools, the plan-review SSE/HTTP/SDK surface, and the plan_* model-routing settings from the experimental 1.6 line. |
||
|
|
15b3aad815 |
feat(web-search): let the model pick a SearxNG category
Rename the web_search tool's `topic` parameter to `category` and expand the enum to general/news/it/science, mapped to SearxNG `categories=`. The model can now target the right corpus per query (e.g. `it` for code, `science` for papers) — useful when generic engines rate-limit. The Tavily-era `finance` topic (no SearxNG equivalent) is dropped. Threaded consistently through _prepare_web_search / _exec_web_search / both search clients. BREAKING: the web_search `topic` argument is now `category`. |
||
|
|
1728a4c0af |
feat(web-search): replace Tavily/DuckDuckGo backends with self-hosted SearxNG
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single self-hosted SearxNG service bundled into the docker-compose stacks. Core: - New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to (backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard unchanged. _resolve_search_client follows storage -> toml -> env -> default precedence (explicit "" disables, via ConfigStore.stored_keys()). - Drop the Tavily-era topic=finance (no SearxNG category); topic is now general/news. Settings/config: - Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key. - Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines, with get_searxng_url/get_searxng_engines. Compose + bundled config: - Internal-only searxng service (no published API port, :ro config, /healthz healthcheck, persistent searxng-cache volume) in both stacks; bundle turnstone/deploy/searxng/settings.yml (JSON output on, limiter off). - Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in). - bootstrap extractor + wheel packaging updated. Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the lxml/h2/brotli transitives). Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG; docs/docker.md carries the AGPL-3.0 §13 operator note. BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg"; tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance. Closes #545 |
||
|
|
631f1b0021 |
feat(compose): cluster-by-default Caddy-fronted stack with bare-metal join
`docker compose up` from a clone builds one image and brings up the whole stack — PostgreSQL, console, Caddy, channel, and 10 server nodes — sharing one Postgres so the console discovers every node. The dashboard is reachable only through Caddy (HTTP/2 avoids the browser's 6-connection cap on the dashboard's SSE streams); the console's plain-HTTP port is no longer published. Postgres binds 127.0.0.1 so a bare-metal turnstone-server can join the cluster — the bare-metal overlay is folded in and removed. Insecure dev defaults keep it zero-config; the bundled production stack mirrors the shape but pulls ghcr images and requires real secrets. Move the Caddyfile under turnstone/deploy so it ships in the wheel; update docs, QUICKSTART, and .env.example to match. |
||
|
|
da07554693 |
docs(judge): correct Smart Approvals heuristic-floor wording
The floor blocks only explicit heuristic deny/critical verdicts — it is not a general "never lower the heuristic" rule. The heuristic default for an unmatched tool is `review`, and letting a confident LLM `approve` upgrade a `review` is the feature's purpose. Matches the implementation and addresses PR review feedback. |
||
|
|
948e413f66 |
feat(judge): add Smart Approvals (auto-approve trusted judge verdicts)
Opt-in judge.smart_approvals (default off): when the intent-validation LLM judge returns a high-confidence "approve" verdict, the tool batch is approved automatically with no operator prompt. review/deny recommendations, low confidence, judge errors (llm_fallback), and a deterministic heuristic deny/critical finding all still require a human. Requires judge.enabled. - Batch-atomic: a parallel tool batch auto-approves only if every call qualifies; one non-qualifying call holds the whole batch for a human. - Gate: tier==llm + recommendation==approve + confidence >= judge.confidence_threshold (default raised 0.7 -> 0.95), with a floor that never clears an explicit heuristic deny/critical verdict. - approve_tools waits for the async LLM verdicts, finalises the audit trail (AutoApproveReason.smart_approval), and re-emits verdicts after the card so the live chip updates; the auto-approved row renders the LLM verdict rather than the cautious heuristic carry-over. - judge: always deliver exactly one verdict per call (fallback on error); reject non-finite confidence so NaN can't clear the bar. - Drop verdicts from a superseded judge generation so a reused call_id from a prior turn's still-running daemon can't satisfy the gate's wait. Config plumbed through the server/console/CLI builders and the live _judge_cfg; admin Judge tab renders the toggle. Docs + example config updated. ~35 tests covering the gate matrix, batch-atomicity, the heuristic floor, audit stamping, the streaming re-emit, NaN/duplicate-id defenses, and the cross-turn generation guard. |
||
|
|
d820168f3f |
fix(tls): repair cluster mTLS — cert identity, renewal scoping, hot-reload
Enabling mTLS broke the cluster in three layered ways: - Service certs were keyed on socket.gethostname() (the container ID) and never carried the advertised service name as a SAN, so every collector and routing-proxy handshake failed the hostname check. build_cert_hostnames() now puts the advertised host first: it becomes the cert's primary domain (hence a SAN) and a stable store key that survives container recreation. - lacme's RenewalManager renews everything in the store; with the store shared cluster-wide, every node renewed every other node's (and every dead container's) cert — an N×M renewal storm. _SingleDomainStore scopes each node's sweep to its own cert, and the console adds a periodic GC for the certs of long-departed nodes. - uvicorn loads its cert once at boot and never reloads, so renewed certs never reached the listener and the served cert expired mid-process. swap_context_cert() hot-swaps renewed material into the live SSL context (server listener and console client context) via load_cert_chain. Observability and browser access: - The collector logged connection/TLS failures at DEBUG, so a persistent mTLS-verify failure was invisible. It now logs the first failure per node (reachable->unreachable) at WARNING and stays at DEBUG on retries. - The console serves plain HTTP (it is the ACME bootstrap endpoint) and no longer rewrites its advertised URL to https://. Browser->console TLS is terminated by a reverse proxy: the cluster profile gains a caddy service (browser h2/HTTPS -> caddy -> console h1.1/HTTP) plus browser-TLS docs. Tests: tests/test_tls_san_renewal.py, tests/test_collector_reachability.py. |
||
|
|
30d670338e |
feat(judge): merge output-guard heuristic + LLM judge, annotate findings
Surface the output-guard LLM judge on the inline finding chip and merge it with the regex heuristic instead of one stage winning outright. Merge rule (issue #560, "show, annotated"): - risk_level = max(heuristic, llm); flags = union. The judge can escalate but never lower a heuristic positive — it evaluates adversarial tool output, so defeating it must not erase a deterministic regex finding. Credentials stay heuristic-only and are always redacted. - The judge's own verdict rides along as a dissent-aware annotation (judge_risk / confidence / reasoning / judge_model) on the chip in both the interactive and coordinator UIs, live and on reconnect. One shared merge_guard_display_payload drives both paths so they cannot drift. - The model is shown the merged risk + flags but never the judge's "benign" verdict — a fooled judge must not talk the model out of caution. Fixes a reconnect bug: a judge that ran but failed wrote a risk="none" row that won the replay dedup and hid the heuristic finding (it showed live but vanished on refresh). Failed judges now persist under tier="llm_error", excluded from the display merge; the max-merge also floors the displayed risk at the heuristic level so the chip never vanishes. Also adds a regression test confirming the LLM judge runs on every tool output, not just heuristic-flagged ones. Tests: merge unit tests, storage-backed replay regression, live/replay wire-shape parity, SDK-event drift guard. ruff + mypy clean. |
||
|
|
ee8dc7c1c3 |
refactor(history): project the /history wire shape server-side
Collapse the three hand-synced "raw storage -> render shape" projections into one server-side projection. The projection previously lived in a test-only `_build_history` (SSE-era reference impl), a client-side JS normaliser (`history_normalize.js`, the transitional bridge), and coord's inline `init()` handling -- drifting silently with no parity test. Add `project_history_messages` to `history_decoration.py` and run it as the final step of the `make_history_handler` pipeline (load_messages -> decorate -> extract_reasoning -> project), so `GET /history` emits the canonical render shape directly: flat tool_calls (with verdict / output_assessment), top-level source / reminders / attachments, collapsed multipart content, derived denied / is_error / pending, reasoning, and advisories. Interactive `replayHistory` now consumes the payload verbatim. Close two gaps the JS bridge deferred: - list-content <tool_output> advisory extraction (decorate handles only string content; the projection extracts list-carrier advisories, then joins remaining text parts to the string the renderers require); - orphan->pending marks ONLY the last orphan tool-call turn, so a mid-conversation cancelled tool still renders instead of vanishing. Delete `history_normalize.js` (+ its <script> tag and node test) and the test-only `_build_history` (+ orphaned imports); retarget its direct tests onto the projection helpers. Update the WorkstreamHistoryResponse description and the Web UI Resilience architecture note to the projected shape. Coord's `init()` still reads the raw side-channels; migrating it to the projected shape is the next commit, browser-verified separately. Refs #549. |
||
|
|
b02e4312ac |
feat(tools): make notify dual-kind, expose to coordinator sessions (#559)
notify was interactive-only — a coord with a natural "fan-out complete" or "batch failed" beat could only post by spawning a child for the single message, which is a lot of ceremony. Routing is session-kind- agnostic in _prepare_notify / _exec_notify; this is a metadata flip that adds the coord flag (plus the explicit interactive flag the loader needs once coordinator is set) and updates the dual-kind whitelists, coord tool-set assertions, and skill-author docs accordingly. Adds two coord-session tests pinning the prepare dispatch contract (needs_approval=False matches notify.json auto_approve) and the exec → channel-gateway path. |
||
|
|
7404ae46db | chore: download vendored JS files | ||
|
|
83a602fef7 |
fix(skills): address /review on flatten — kind validation + stale text
Four /review findings collapsed to one code chokepoint + two
documentation fixes:
1. find's `kind` arg now validated against ``SkillKind`` (matching
create / update's existing pattern at session.py:8298 / :8512).
Closes two failure modes that shared the same root:
- typos (`kind="interactivee"`) silently produced
`kinds=["interactivee", "any"]` filtering to literal-`any` rows
only and masquerading as a narrowed catalog — now returns an
explicit "kind must be one of: ..." error;
- the documented enum value `kind="any"` degenerated to
`kinds=["any", "any"]` which narrowed to literal-`any` rows
instead of returning "every kind" — now collapses to ``None``
so the documented semantic holds.
2. docs/coordinator-skills.md "two-surface model" section rewritten
to reflect the post-flatten reality: kind is metadata, not an
enforcement boundary. The line-67 tools-table row updated from
the long-dead `list_skills` to `skills (action=find)` with the
opt-in kind-filter framing.
3. Three stale "interactive-only" comments in session.py
(:5514, :7857, :8210) that directly contradicted the
`_prepare_skills_load` docstring ("Both kinds can load") — drop
the qualifier so future grep-and-encode hazards don't reintroduce
the rejection.
Tests:
- test_find_kind_invalid_errors — typo case (replaces the silent
degenerate to literal-any-only)
- test_find_kind_any_means_no_filter — documented enum value matches
documented semantic (collapses to None at prepare)
- test_find_kind_narrow_passes_through — valid narrowing values
reach exec as expected
Deferred to release notes (no code change, intentional policy shift):
- skills(action='get') / load can now read full content + scan_report +
allowed_tools on cross-kind rows from any session. Operators with
pre-existing kind=coordinator skills authored under the prior
implicit visibility contract should audit those bodies for
sensitive content (allowed_tools allowlists, embedded credentials,
internal hostnames in examples) before upgrade.
|
||
|
|
c4495c0c48 |
fix(coord): address PR review threads on spawn_workstream rename
Three Copilot threads from PR #526: 1. ``_exec_spawn_workstream`` success path emitted ``{"child_ws_id": null}`` when the upstream response unexpectedly omitted ``ws_id`` (200-shape with no error field, no id field). Adds the missing guard — mirrors ``_exec_spawn_batch`` which already surfaces ``"spawn returned no ws_id"`` as a denied row. The LLM now sees a tool error and can retry instead of chasing a null id through follow-up tools. 2. ``docs/coordinator-skills.md`` UI render note said "keep the ws_id as the click-through key" in a paragraph that had just introduced ``child_ws_id`` — readable as "the ws_id value" but confusable as a field-name claim. Clarifies that the value class is the same regardless of which key carried it. 3. ``docs/bulk-endpoints.md`` ``spawn_batch`` example shows ``child_ws_id`` (coord-tool output shape). The doc title and the "model tool" column label already disambiguate it from HTTP API responses, but a reader landing at the example section directly could miss the framing. Adds one explicit sentence. |
||
|
|
6948ea21cb |
fix(coord): rename ws_id->child_ws_id in spawn return JSON
Coordinator LLMs on large fan-outs recency-bias on seeing `ws_id` in a `spawn_workstream` / `spawn_batch` return -- calling `spawn_workstream(ws_id=...)` again instead of progressing to `wait_for_workstream(ws_ids=[...])`. On 10+ child fan-outs this cascades into self-inflicted re-spawn loops. Rename to `child_ws_id` (already an existing project term -- see `tasks` tool, `child_event_bus.py`) defuses the recency bias. Scope is the LLM-facing JSON only -- the server HTTP API at the spawn endpoint still returns `ws_id`, and the internal reads of that HTTP response are unchanged. Also updates the two tool descriptions, the operator-facing skill doc, and the bulk-endpoints example so docs don't undo the rename. |
||
|
|
6998b442a9 | chore: download vendored JS files | ||
|
|
f28a3533a2 | chore: download vendored JS files | ||
|
|
adeb10bc2c |
feat(mcp): admin status, deferred-consent persistence, operator docs (Phase 9) (#516)
* feat(mcp): admin status, deferred-consent persistence, operator docs (Phase 9)
Completes the OAuth-MCP build-out (Phases 0-8 shipped) by closing the
operator + deferred-consent gaps:
1. **Per-(user, server) deferred-consent persistence** — when a
non-interactive run (scheduled / channel) hits ``mcp_consent_required``
or ``mcp_insufficient_scope``, the sync pool dispatchers now upsert a
row into a new ``mcp_pending_consent`` table. The dashboard hydrates
the gear-icon badge from this table on load, so users who weren't
online to see the in-flight SSE prompt still surface the deferred
work on next login. Cleared automatically by the OAuth callback
handler on consent completion; manual user dismiss via new DELETE
endpoints. Composite PK ``(user_id, server_name)`` collapses repeat
occurrences for the same server — no NULLs-not-distinct trap.
2. **Admin status pill + bulk-revoke** — the MCP Servers admin row now
shows ``consented_users_count`` for ``auth_type=oauth_user`` rows
when ≥1, with a two-step-confirm ``bulk-revoke`` button that drops
every user's token for the server via the existing
``delete_mcp_oauth_rows_by_server_name`` primitive. Upstream RFC
7009 revoke is intentionally NOT attempted in bulk (avoids N
upstream HTTP calls per admin click); audit detail records
``upstream_revoke_outcome=bulk_admin_no_upstream``. A "last
refresh" pill (age + outcome) renders on each row, sourced from a
new ``_last_refresh`` dict populated by ``_refresh_server`` on every
call (both manual ``refresh_sync`` and the ``_cb_auto_reconnect``
follow-up).
3. **ClientType.SCHEDULED** added to the prompts module + scheduler
passes it through to ``create_workstream``. ``ChatSession`` now
computes ``_is_interactive_for_consent`` at construction (WEB / CLI
are interactive; CHAT / SCHEDULED are not) and plumbs the flag
through ``call_tool_sync`` / ``read_resource_sync`` /
``get_prompt_sync`` to the three sync dispatchers. The wrap at the
``_is_structured_error`` gate routes consent codes to the new
``_record_pending_consent_best_effort`` helper for non-interactive
callers only; interactive sessions stay on the in-flight SSE path
Phase 8 ships unchanged.
4. **Operator docs** — ``docs/mcp-oauth.md`` (operator guide, parallel
to ``docs/oidc.md``: ``auth_type`` choice, OAuth client setup,
encryption-key rotation, troubleshooting matrix) and
``docs/operations/mcp-oauth-headless.md`` (one-paragraph runbook
per ``feedback_runbook_trust_llm.md``: pre-consent recipe for
scheduled / channel-driven runs).
Schema
- Migration 054_mcp_pending_consent.py — composite PK
``(user_id, server_name)``, ``occurrence_count`` + ``first_seen_at`` /
``last_seen_at`` for recency metadata, ``idx_mcp_pending_consent_user``
for the badge-load query. No FKs (matches the rest of the
oauth_user schema).
- Migration 055_mcp_user_tokens_server_index.py — adds
``idx_mcp_user_tokens_server`` on ``(server_name, expires_at)`` so
the admin pill's ``count_mcp_consented_users_*`` queries don't
full-scan against the leading-``user_id`` composite PK.
- Cross-backend: works on SQLite + PostgreSQL via dialect-specific
``on_conflict_do_update`` (PG ``postgresql.insert`` / SQLite
``sqlalchemy.dialects.sqlite.insert``). No ``NULLS NOT DISTINCT``
needed — the simplified PK eliminates the cross-version trap.
Endpoints
- ``GET /v1/api/mcp/oauth/pending`` — list deferred-consent records for
the authenticated user. Install-level gate via cached
``any_oauth_user_mcp_servers`` short-circuits to ``{pending: 0}`` on
installs with no oauth_user MCP servers — local-auth deployments
exercise zero new storage queries on this path. The gate result is
cached on ``app.state`` with a 60s TTL to spare repeat dashboard
loads.
- ``DELETE /v1/api/mcp/oauth/pending/{server_name}`` — single dismiss.
Returns 204 in both existed-and-deleted and never-existed cases
(no cross-tenant existence leak); audits
``mcp_server.oauth.pending_consent_dismissed`` with
``mode=single`` + ``cleared=0|1`` so a session-hijack attacker
scrubbing breadcrumbs leaves an audit trail.
- ``DELETE /v1/api/mcp/oauth/pending`` — bulk dismiss; audits
``mode=bulk`` + ``cleared=N``.
- ``POST /v1/api/admin/mcp-servers/{name}/bulk-revoke`` — admin
bulk-revoke for the named server's per-user tokens. Requires
``admin.mcp`` permission + 400s when the row isn't ``oauth_user``.
All four registered on both ``turnstone-server`` and
``turnstone-console`` (mirrors the Phase 8 ``/connections`` endpoint
shape).
Performance
- Admin list handler now uses a single ``GROUP BY`` bulk-count query
(``count_mcp_consented_users_grouped_by_server``) wrapped in
``asyncio.to_thread`` rather than N per-row sync DB round-trips
inside the async handler. Skipped entirely when no row is
oauth_user.
Frontend
- ``ui/static/app.js``: ``loadPendingConsents()`` hydrates the
existing ``_pendingConsentServers`` set on dashboard init + after
the user opens the settings modal. Endpoint failures stay silent
— the badge will be re-driven by the next in-flight tool error.
- ``console/static/admin.js``: ``consented_users_count`` pill +
``bulk-revoke`` button on each MCP row (only when ≥1 consented),
two-step confirm matching the existing delete pattern. ``last-
refresh`` age + outcome pill in the per-row status cell, sourced
from the freshest per-node entry in ``status[*].last_refresh_at`` /
``last_refresh_outcome``. CSS for the pills in ``style.css``.
Tests
- ``test_mcp_pending_consent_storage`` — 13 tests covering upsert
idempotency, list ordering, per-user isolation, single/bulk delete,
count-by-server + grouped variant, install-level gate.
- ``test_mcp_pending_consent_dispatch`` — 9 tests, including the
boundary-cross gate per ``feedback_tests_through_boundaries.md``:
drives the real ``call_tool_sync`` → ``_dispatch_pool_sync`` →
``_is_structured_error`` → ``_record_pending_consent_best_effort``
with a mocked classified-lookup so the structural plumb-through is
verified end-to-end. Includes a storage-failure test that pins
the docstring's "envelope unchanged on storage failure" promise.
- ``test_mcp_pending_consent_endpoints`` — 11 tests: install gate,
list-for-self, no-cross-user-leak, single/bulk delete, idempotent
not-found, audit emission on single + bulk + cross-tenant dismiss.
- ``test_chat_session_interactivity_flag`` — 7 tests pinning the
``ClientType`` → ``_is_interactive_for_consent`` mapping against
the module-level ``INTERACTIVE_CONSENT_CLIENT_TYPES`` frozenset.
- ``test_mcp_admin_bulk_revoke`` — 7 tests covering admin.mcp
permission gate, 404 on missing, 400 on non-oauth_user, 200 with
``rows_deleted`` + ``consented_users_before``, audit row with
``upstream_revoke_outcome=bulk_admin_no_upstream``, cross-server
isolation.
- ``test_mcp_oauth_handlers`` — 2 new callback tests pin the post-
callback ``delete_mcp_pending_consent`` invocation: success-clears
+ storage-failure-still-redirects.
- 636 tests pass on the impacted surface (47 new + Phase 0-8 OAuth-MCP
+ session + prompts + storage admin). ruff + mypy clean.
Hard invariants honored
- Static path byte-identical for ``auth_type ∈ {none, static}`` — the
flag flows only through the pool dispatchers, which only fire when
the row resolves to ``oauth_user``.
- ``asyncio.timeout`` (not ``asyncio.wait_for``) preserved on every
AS / SDK / pool-loop await — no new awaits added to the hot path.
- Install-level gate on the badge endpoint: cached
``any_oauth_user_mcp_servers`` returns False on a row-less
deployment → endpoint short-circuits without touching the pending-
consent table; 60s TTL bounds the staleness window after admin
flips ``auth_type``.
- Operator-actionable codes (key-unknown, url-insecure, *_forbidden)
explicitly filtered out of persistence — they're outside the
user-facing consent badge scope.
- Best-effort write: the structured-error envelope returned to the
agent is identical whether the persistence write succeeds or fails
(storage exception is logged with type name only — no chained
context that could carry an ``httpx.Request`` bearer header).
- No ``exc_info=True`` on any new path that can chain a bearer-bearing
``httpx.Request``.
- Defensive parsing: ``_parse_pending_consent_envelope`` mirrors
``_is_structured_error``'s ``isinstance(decoded, dict)`` guard plus
filters scope tokens through ``is_valid_scope_token`` capped at
``MAX_INSUFFICIENT_SCOPE_REPORTED`` — defense-in-depth even though
production callers already validate upstream.
- Audit events on every dismiss endpoint so a session-control attacker
scrubbing dashboard breadcrumbs still leaves a trail.
Cross-backend
- Tested on SQLite via the conftest backend fixture.
- PostgreSQL path uses ``postgresql.insert(...).on_conflict_do_update``
parallel to the existing ``mcp_user_tokens`` upsert in Phase 3.
Deferred (not Phase 9 blockers)
- Multi-node pool eviction on bulk-revoke: only local-node sessions
would be evicted if we built it, and there's no bulk-by-server
primitive on MCPClientManager today; remote nodes will surface as
a 401 on next dispatch which refreshes through the (now empty)
token row.
- RFC 8693 / Azure OBO ``auth_type=oauth_token_exchange`` — captured
in the design doc as a future architectural direction (~600 LOC +
IdP-side admin work); requires OIDC token capture and per-MCP-server
resource-trust configuration that v1 does not ship.
* docs(mcp): address Copilot review feedback on Phase 9
- Fix misleading admin.js comment that claimed the refresh pill rendered
"<short-relative> <outcome>" — the pill actually renders only the short
age, with outcome reflected via CSS class and tooltip.
- Replace broken feedback_secrets_not_in_env.md repo-root link in
mcp-oauth.md with the inlined rationale (env-borne secrets reachable
via shell tools / os.environ; TOML secrets are not).
|
||
|
|
752fea0fdd |
feat(console): reactive node discovery via PG LISTEN/NOTIFY dispatcher (#505)
* feat(console): reactive node discovery via PG LISTEN/NOTIFY dispatcher Add a console-side `NotifyDispatcher` that holds a dedicated PostgreSQL `LISTEN` connection and fans wake-ups out to per-channel handlers on a separate dispatch thread. Cluster collector subscribes to a new `services` channel and runs node discovery reactively — new-node / graceful-deregister visibility drops from up-to-60 s to ~500 ms on Postgres, with the 60 s discovery loop retained as the backstop for crash-shaped node loss (NOTIFY only fires on real writes). Storage layer gains a uniform `notify` / `listen` API: - PostgreSQL: real `pg_notify` / `LISTEN` on a dedicated session-mode connection that bypasses pgbouncer (mandatory: pgbouncer is required in transaction-pool mode per docs, which is incompatible with LISTEN). - SQLite: in-process fan-out + synthetic-sweep fallback so consumer code is identical across backends. `TURNSTONE_DB_LISTEN_URL` (or `[database] listen_url` in config.toml) points the dispatcher's connection direct-to-Postgres. Defaults to the main DB URL when unset. Migration 053 installs the `services_notify` trigger; it filters heartbeat-only UPDATEs in-trigger so the 30 s × N-nodes heartbeat tick stays quiet, while INSERT, DELETE, and url/metadata-changing UPDATE still fire. Dispatcher detail: - Two threads: listener (drains stream → bounded queue) and dispatch (invokes handlers under exception suppression). Same-channel notifies coalesce per dispatch batch so an N-node deploy burst is one `_discover_nodes` per channel. - Reconnect uses exponential backoff (1 s → 30 s cap). After any successful reopen — whether the prior failure was a stream-poll error or a connect / initial-LISTEN error — one synthetic Notify with payload="reconcile" is enqueued per channel so handlers re-read on the same code path they use for real events. Future consumers (ConfigStore live reload, scheduler immediate dispatch, audit live-tail) plug in by adding their channel to the dispatcher's construction list. Tests: 22 dispatcher tests (incl. reconnect + coalescing under stub storage), 7 SQLite notify-stream tests, 4 PG-gated trigger-filter tests, 4 collector wire-in tests. All pass; ruff + mypy clean. * fix(notify): address Copilot review on #505 - _sqlite.py: SQLiteBackend.listen() now de-dupes channel names via dict.fromkeys before constructing the stream — duplicates would otherwise register the queue twice and double-deliver each notify. - _sqlite.py: SQLiteBackend.listen() gains a keyword-only sweep_interval parameter (defaults to _SQLITE_NOTIFY_SWEEP_INTERVAL) — matches what the comment at the constant already promised, and lets future consumers without their own polling timer pick a tighter cadence without reaching into private stream attributes. - _sqlite.py: documented the `except queue.Empty: pass` end-of-drain termination so it's not mistaken for swallowing an unexpected error. - _postgresql.py: docstring referenced :func:`_pg_listen_url` which was renamed to _resolve_pg_listen_url during PR development. - notify_dispatcher.py: module docstring referenced a non-existent _bootstrap_console_subsystem; wire-in is at console/server.py::main. Refuted (no change, false positives from github-code-quality bot): - 4× "Statement has no effect" on Protocol-method `...` ellipsis bodies (idiomatic Python Protocol declaration, not dead code). - 2× "Mixed import style" in tests — `import ... as nd_mod` is intentional to allow attribute assignment for monkey-patching the module's `_RECONNECT_BACKOFF_INITIAL` constant inside try/finally. |
||
|
|
6f8574eef3 |
fix(reasoning): synthesize reasoning_text alongside non-reasoning provider_blocks
GoogleProvider attaches raw tool_call dicts as ``provider_blocks`` on the finish chunk for ``thought_signature`` round-trip (``_google.py:_iter_stream``). When the same turn streamed Gemini's ``reasoning_content`` as ``reasoning_delta`` chunks, the prior synthesizer bailed out the moment ``provider_blocks`` was non-empty — so the captured reasoning was visible live but lost on page reload. Replace the early-return-if-non-empty check with a reasoning-bearing type test (``thinking`` / ``redacted_thinking`` / ``reasoning`` / ``reasoning_text``). When none of those types appear, append the synthetic ``reasoning_text`` block to the existing list rather than replacing it — preserving Google's tool-call fidelity blocks. Also addresses two doc-accuracy review findings: - ``LLMProvider.extract_reasoning_text`` docstring no longer claims OpenAI Chat / Responses are unwired (Phase 3+4 shipped extractors). - Add the method to the Protocol methods table in ``docs/architecture.md`` (was missing alongside the class diagram). |
||
|
|
20e1e7b110 |
fix(reasoning): apply Copilot review feedback + docs sync
PR #498 round-robin review surfaced 5 findings. 4 applied; 1 rejected with rationale. Applied * **Copilot finding 5** (history_decoration.py:341): dispatcher inspected only ``provider_content[0]['type']``. OpenAI Responses captures EVERY ``output_item.done`` event into ``provider_blocks`` (not just reasoning) — in practice the order is ``[reasoning, message, ...]`` but the API doesn't guarantee that; a hypothetical ``[message, reasoning]`` ordering would silently drop the reasoning under an index-only check. Now walks the list for the first block whose type is in ``_BLOCK_TYPE_PROVIDER_FACTORY``, then dispatches the WHOLE list to that provider's extractor. Each provider's extractor already filters internally by its own block type, so passing the full list is correct. Regression test added (``test_dispatcher_scans_past_unrecognized_first_blocks``). * **Copilot finding 3** (migration 052 docstring): the previous review-fix wave used sed to rename ``persist_reasoning`` → ``surface_persisted_reasoning`` everywhere, which mangled a historical reference in the migration docstring ("The earlier name ``surface_persisted_reasoning`` was renamed..."). Restored to point at the actual pre-rename name (``persist_reasoning``). * **Copilot finding 4** (sdk/typescript/src/events.ts:26): ``HistoryEvent`` JSDoc still referenced ``persist_reasoning`` — the sed rename only walked ``turnstone/`` and ``tests/``, missing the TypeScript SDK. Updated to ``surface_persisted_reasoning``. Also widened the comment to cover all three reasoning-bearing block types (Anthropic ``thinking``, OpenAI Responses ``reasoning``, synthetic ``reasoning_text``) instead of mentioning only Anthropic. * **github-code-quality finding** (session.py:1120): ``_resolve_server_type`` had a bare ``except Exception: pass``. Replaced with a ``log.debug(..., exc_info=True)`` + explanatory comment. Behaviour unchanged (still returns ``""`` on any lookup failure); failures are now observable under DEBUG triage. Rejected (with rationale) * **github-code-quality finding** (_protocol.py:265): ``extract_reasoning_text``'s body is ``...`` per ``LLMProvider`` Protocol convention. Every method in the file uses ``...`` (PEP 544 idiomatic Protocol style). Changing only this one to ``raise NotImplementedError`` would be inconsistent with the rest of the file. CodeQL's "statement has no effect" warning is technically correct for ``...`` as a standalone expression but ignores the documented Python Protocol convention. No fix. Docs sync * docs/api-reference.md: ``history`` SSE event message-shape table gains the optional ``reasoning`` field. * docs/architecture.md: ``ModelCapabilities`` row in the type table gains ``supports_reasoning_replay``; ``StreamChunk`` and ``CompletionResult`` rows gain the existing ``provider_blocks`` field (was missing pre-PR). New "Per-model reasoning persistence" subsection under the Models config section, documenting the two flags + capability gate + three reasoning paths + cross-provider shape filter. * docs/settings.md: new "Reasoning persistence (per-model)" subsection with the two-flag table and capability-gate note. * docs/diagrams/03-core-engine-classes.puml: ``LLMProvider`` interface adds ``extract_reasoning_text`` + the new ``replay_reasoning_to_model`` kwarg; ``ModelCapabilities`` class adds ``supports_reasoning_replay``. PNG regenerated. Lint + test gate * ruff check + ruff format clean. * mypy clean (191 source files). * pytest -m 'not live' — 6116 passed (3 deselected), +1 net new test (``test_dispatcher_scans_past_unrecognized_first_blocks``). |
||
|
|
57cb09c871 |
docs(sse): document state_change + in_progress_snapshot events
Updates the docs that describe the per-workstream SSE event stream and the SessionUI lifecycle to match the refresh-resume changes: - api-reference.md: documented the `state_change` event (previously undocumented despite already being a live event) and the new `in_progress_snapshot` event; rewrote the multi-consumer fan-out paragraph to mention the kind-specific replay tail (state_change + optional in_progress_snapshot) so the "no catch-up needed" claim is no longer misleading. - architecture.md: bumped the SessionUI Protocol stub to 16 methods (added `on_turn_start` / `on_turn_committed`) and pointed at the in_progress_snapshot section in the API reference. - sdk.md: added rows for `state_change`, `in_progress_snapshot`, and `approval_resolved` (preexisting gap) to the per-workstream event table. - coordinator-api-tour.md: added an `in_progress_snapshot` row to the event table and rewrote the reconnection-contract paragraph to cover mid-stream content/reasoning restoration. - diagrams/04-conversation-turn.puml: added `on_turn_start()` before the thinking-start emit and `on_turn_committed()` immediately after `messages.append(assistant_msg)`, with notes explaining the inflight- buffer reset semantics. PNG regenerated. |
||
|
|
eb2a119da9 |
refactor(mcp): remove periodic refresh, add manual refresh/reconnect controls
Deletes the _periodic_refresh task and its supporting state
(_refresh_task, _refresh_failures, _refresh_backoff_until,
_REFRESH_BACKOFF_BASE/MAX, _DEFAULT_REFRESH_INTERVAL, refresh_interval
kwarg) from MCPClientManager. Push notifications and operator-driven
manual refresh now cover all catalog-update needs; the long-running
4-hour timer was dead complexity that obscured the per-user pool
work to come.
Catalog freshness on auto-reconnect is preserved by scheduling an
unblocking _refresh_server task on the mcp-loop after _connect_one
succeeds; the calling thread returns immediately so half-open
recovery latency does not double. Adds MCPClientManager.reconnect_sync
(clears the circuit, closes any existing session, calls _connect_one,
clears stale catalog on failure).
Wires a new pair of operator endpoints —
POST /v1/api/admin/mcp-servers/{name}/refresh and
/v1/api/admin/mcp-servers/{name}/reconnect — that fan out to all
nodes through the existing _internal route family, with per-row
"Refresh" and "Reconnect" buttons in the MCP Servers admin tab.
The new node-internal paths /api/_internal/mcp-{refresh,reconnect}/
are gated to the approve scope to prevent direct unprivileged
reconnects bypassing the console's admin.mcp gate. Internal
endpoints return generic error messages and a filtered status
payload (no command/url) to keep transport details admin-gated.
Drops the [mcp] refresh_interval setting, the
--mcp-refresh-interval CLI flag, and the matching config-mapping
entry; updates docs/architecture.md, docs/tools.md,
docs/settings.md, and the three PlantUML diagrams that referenced
the periodic loop.
Tradeoffs (intentional):
- Idle nodes will not auto-rejoin a recovered MCP server until
traffic arrives or an operator clicks Reconnect. The previous
background reconnection loop is gone by design — push
notifications + operator controls replace it.
- Console fan-out blocks on the slowest node (existing pattern);
not changed here.
This is Phase 1 of the OAuth-MCP series — feature subtraction
ahead of per-user state.
|
||
|
|
3cf87628d2 |
docs(oidc): document TRUSTED_ENDPOINT_HOSTS + fix three-vs-four required drift (cumulative q-1, q-2)
The 8-commit OIDC stack added TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS (operator allow-list for cross-host IdP discovery endpoints) and promoted TURNSTONE_OIDC_REDIRECT_BASE to required, but the docs drifted in two places: q-1 — Troubleshooting > "OIDC not configured" still listed three required env vars. An operator hitting the missing-redirect-base startup error landed on a debugging entry that didn't mention the variable they were missing. Fixed; added a separate troubleshooting entry naming the exact log message produced by initialize_oidc_state when redirect_base is unset. q-2 — TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS was undocumented entirely. Added a row to the env-var table and a new "Cross-host endpoints" section explaining when the knob is needed (Google is the canonical multi-origin IdP, but it's auto-handled; the env var is for any other IdP whose discovery doc legitimately references hosts beyond the issuer's origin). Added a troubleshooting entry pointing at the new section. |
||
|
|
bae4adca12 |
refactor(oidc): quality cleanup (bug-3, q-1/3/4/6/7/9/10/11/12/13)
Eleven small maintenance fixes; no behavior change beyond bug-3.
bug-3: pending.get('audience', audience) couldn't fall back because
pop_oidc_pending_state always returns a dict with the audience key
set verbatim from a non-null TEXT column. Replaced with
pending.get('audience') or audience to cover the empty-string case
defensively. Comment explains the security rationale.
q-1: extract _env_or_cfg_str / _env_or_cfg_bool helpers in oidc.py;
load_oidc_config's six near-identical env-or-config blocks collapse
to one-liners. role_map / trusted_endpoint_hosts / redirect_base
retain bespoke parsing.
q-3: discover_oidc narrows except (httpx.HTTPError, ValueError, KeyError)
with exc_info=True.
q-4: OIDC_STATE_TTL_SECONDS = 300 constant in oidc.py; auth.py imports
and passes it explicitly. Storage signatures keep the literal default
(storage layer doesn't know OIDC TTL semantics).
q-6: hoist runtime imports (OIDCError, OIDCKeyNotFoundError, exchange_code,
fetch_jwks, provision_oidc_user, validate_id_token, build_authorize_url,
generate_pkce_verifier) to module scope in auth.py. The genuine cycle
is only oidc._derive_username -> auth.is_valid_username, kept
function-scoped. test_oidc_handlers.py mock targets repointed to
turnstone.core.auth.X to match the new binding.
q-7: comment + docs explain the 'oidc' vs 'oidc-default' assigned_by
marker distinction.
q-9: OIDCIdentity / OIDCPendingState TypedDicts in storage protocol.
Implementations construct via TypedDict syntax so mypy structurally
verifies all required fields.
q-10: fetch_jwks narrows except (httpx.HTTPError, ValueError); docstring
matches.
q-11: rename generate_pkce_pair -> generate_pkce_verifier; return only
the verifier (build_authorize_url already recomputes the challenge).
q-12: extract _buildOidcRow helper in admin.js so future field additions
go in one place.
q-13: OIDCConfig docstring lists startup-config vs discovery-derived
field groups.
|
||
|
|
52aba17740 |
fix(oidc): require TURNSTONE_OIDC_REDIRECT_BASE; drop Host-header fallback (sec-2)
_build_oidc_redirect_uri previously fell back to the request Host
header when redirect_base was unset. With a permissive reverse proxy
or direct backend access, a spoofed Host minted an authorize URL
pointing to attacker-controlled host — combined with a permissive
IdP redirect_uri allowlist this enables auth-code interception.
There is no production scenario where a Host-derived redirect_uri is
correct, so this fails closed:
- initialize_oidc_state checks redirect_base after discovery succeeds
and disables OIDC (with an explicit error log naming the env var)
if it's empty. Runs before fetch_jwks so a misconfigured deploy
doesn't make a wasted JWKS call.
- _build_oidc_redirect_uri simplifies to f"{redirect_base}/v1/api/auth/oidc/callback".
request parameter dropped; both call sites (handle_oidc_authorize,
handle_oidc_callback) updated.
- docs/oidc.md promotes TURNSTONE_OIDC_REDIRECT_BASE from "Recommended"
to "Required" with the security rationale.
|
||
|
|
6c28ac828f |
docs(skills): add import-conversation-history SKILL.md
Source-agnostic guide that teaches an agent Turnstone's destination contracts (workstream + conversations schema, ws_id routing, OpenAI message shape, tool-call/result pairing, provider_data fidelity blob, attachment lifecycle) so it can map any external chat export onto them. Validated against turnstone.core.skill_parser. |