mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
0df9f6e2d4c6e5ca754513820ab84173a0306c24
17 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
460308241d |
fix(tools): lower working directory and workspace into fs tool descriptions
The process cwd was nowhere in the model's context: shells start in the inherited process cwd (spawn_group_leader passes no cwd), relative file paths resolve against it, but nothing told the model where it was standing — in stock Docker every shell ran in /data while user files sat in the /workspace mount, and the model's only recourse was to probe with pwd (#857, #833). Lower both facts into the tool schemas, where they gate intrinsically on tool availability (a persona without fs tools carries no note, and coordinator envelopes are untouched): - tools/*.json: cwd_note/workspace_note metadata templates on bash, read_file, write_file, edit_file, search, diff_file; bash also states the fresh-shell-per-call semantics (cd does not persist) and drops a stale reference to the removed man tool. - tools.apply_cwd_context(): renders the notes into descriptions; deep-copies noted tools (the fs dicts are shared across TOOLS/INTERACTIVE_TOOLS/TASK_AGENT_TOOLS and aliased through merge_mcp_tools), passes note-less tools through by reference. - ChatSession._apply_cwd_notes(): wraps every fresh interactive build of _tools AND _task_tools (construction, MCP catalog change, MCP disconnect) — assignment-time, so the wire tools block stays byte-stable for provider prompt caches. os.getcwd() is OSError-guarded (MCP rebuilds run on a background thread; eval tears down its workdir); the workspace hint drops when the dir is missing or equals the cwd. Task-agent sub-agents carry their own notes via _task_tools, independent of parent persona visibility. - config.get_workspace_dir(): [tools] workspace_dir with TURNSTONE_WORKSPACE env fallback (searxng pattern), informational only — no chdir, no path confinement (per-workstream working-dir grants are a separate planned feature). - Dockerfile: ENV TURNSTONE_WORKSPACE=/workspace so stock deployments surface the mount with zero operator config. - docs/docker.md: document the /data working directory, the working_dir: /workspace compose override as the operator-level fix, and the SQLite-fallback-DB-in-cwd caveat. Closes #857 |
||
|
|
8b4b8b3fd5 |
refactor(rerank): reranker is a per-model definition only (drop global endpoint settings)
The reranker_alias -> model-definition path (added when reranking became a model role) made the older global endpoint settings redundant. Resolve reranking solely through the Reranker role and remove the parallel global config. - Removed settings tools.rerank_url / rerank_model / rerank_api_key, their config.py getters (+ $TURNSTONE_RERANK_URL / $TURNSTONE_RERANK_MODEL and the module caches), and the fallback branch in resolve_rerank_client_from. The resolver now returns a client only when a Reranker model (capability supports_rerank, base_url = its /rerank endpoint) is selected, else None. - Kept as global knobs: reranker_alias (the selector), rerank_web_search, rerank_bm25, rerank_bm25_threshold, and rerank_instruction -- a task-level query knob (Qwen3-style), not endpoint identity. - The Settings tab is registry-driven, so the three fields disappear with their SettingDefs. Updated the Reranker role help, example config, and docs/tools.md. BREAKING: a reranker configured via [tools] rerank_url (config.toml / env / Settings tab) no longer works -- add the reranker in the admin Models tab and pick it under Models -> Roles -> Reranker. No migration: reranking is days old and disabled by default, so any orphaned tools.rerank_* config rows are inert. Tests: the resolver covers no-store / no-alias / non-rerank-alias -> None and the model-definition happy path; the obsolete global-fallback tests are removed. |
||
|
|
f6bae70ea6 |
feat(rerank): calibration CLI, 0-1 normalization, instruction support
Phase 2 of BM25 reranking (follows #627). Makes the rerank_bm25_threshold floor usable across reranker models and adds tooling to pick it. - normalize_scores (rerank.py): map a rerank batch into a 0-1 relevance probability -- sigmoid when any score falls outside [0,1] (logit endpoints like bge/TEI), identity otherwise (Cohere/Jina/Qwen already 0-1). Applied in the _bm25_reranker closure AND calibration so the threshold means the same on every endpoint. Monotonic, so ranking order is unchanged. - rerank_calibrate.py + `turnstone-admin rerank-calibrate [--apply]`: probe the endpoint with labelled relevant/irrelevant groups, normalise, and recommend a recall-biased floor -- or report "no clean separation" (a mis-served/weak reranker). A warmup loop absorbs a cold endpoint's first-request compile so calibration doesn't time out. Validated live against Qwen3-Reranker 0.6B and 4B: the calibrated floor differs sharply per model (~0.95 vs ~0.33 for the same task) -- exactly why per-endpoint calibration exists. - rerank_config.py: extract resolve_rerank_client_from(config_store, registry); the alias/url precedence now lives in one place, shared by ChatSession (which delegates) and the CLI. - tools.rerank_instruction (config + setting + client): wrap the query as <Instruct>:/<Query>: for instruction-aware rerankers (Qwen3) on endpoints that don't apply the model's own chat template. Docs note the critical vLLM serving detail: Qwen3-Reranker needs --chat-template or its scores are near-random and reranking hurts retrieval. Negative-tested: normalize sigmoid/identity branches, closure-normalises-before- floor, calibration separation/recall-bias/warmup-absorbs-cold-start, the CLI apply/no-apply/no-separation paths, and instruction query-wrapping through the real httpx boundary. |
||
|
|
6a0bc852d9 |
feat(rerank): endpoint-backed reranking for web_search
Reranking is delegated to an external Cohere/Jina-compatible /rerank endpoint (self-hosted vLLM/TEI/llama.cpp, or hosted Cohere/Jina/Voyage); Turnstone runs no reranker model itself. Disabled until an endpoint is configured. - core/rerank.py: CohereJinaRerankClient (tolerant of results-wrapped and bare-list responses) + resolver. - web_search: rerank the SearxNG result pool by query relevance before top-k, with a native-order fallback on error; answers/infoboxes untouched. - Reranker as a model definition: add a model with the supports_rerank capability and pick it under Models -> Roles -> Reranker (tools.reranker_alias); takes precedence over the tools.rerank_url settings. Settings: tools.rerank_url/model/api_key, tools.rerank_web_search, tools.reranker_alias. Docs: docs/tools.md, turnstone.example.toml. (web_fetch reranking was evaluated and dropped: for single-document chunk selection it did not reliably beat head-truncation. Reranking is reserved for multi-item ranking.) |
||
|
|
110d44b07e |
refactor(tools): remove man, math, and plan_agent built-in tools
`man` and `math` duplicated capabilities already reachable through `bash`; `plan_agent` is better expressed as a `task_agent` running a planning skill, and carried a large amount of special-case machinery (plan-review gate, refinement loop, per-kind model routing). Removing all three shrinks the tool surface and cuts per-call token cost. Also removed, as dead-once-the-tools-are-gone: - the `math` sandbox executor (`turnstone.core.sandbox`) and its `[sandbox]` extra; the eval analyst now runs bash-only - the read-only `AGENT_TOOLS` sub-agent tool set and the `agent` tool-metadata key (`task_agent`/`TASK_AGENT_TOOLS` retained) - the plan-review protocol end to end: the `on_plan_review` UI hook, `resolve_plan`, `POST /v1/api/plan` + `POST /v1/api/route/plan`, the `plan_review`/`plan_resolved` SSE events, and their Python SDK / TypeScript SDK / OpenAPI / frontend / Discord+Slack bindings - the `model.plan_alias` / `model.plan_effort` settings and the registry `plan_model` / `plan_effort` routing fields TOOLS 31->28, TASK_AGENT_TOOLS 13->11; COORDINATOR_TOOLS unchanged. BREAKING CHANGE: removes the `man`, `math`, `plan_agent` tools, the plan-review SSE/HTTP/SDK surface, and the plan_* model-routing settings from the experimental 1.6 line. |
||
|
|
1728a4c0af |
feat(web-search): replace Tavily/DuckDuckGo backends with self-hosted SearxNG
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single self-hosted SearxNG service bundled into the docker-compose stacks. Core: - New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to (backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard unchanged. _resolve_search_client follows storage -> toml -> env -> default precedence (explicit "" disables, via ConfigStore.stored_keys()). - Drop the Tavily-era topic=finance (no SearxNG category); topic is now general/news. Settings/config: - Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key. - Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines, with get_searxng_url/get_searxng_engines. Compose + bundled config: - Internal-only searxng service (no published API port, :ro config, /healthz healthcheck, persistent searxng-cache volume) in both stacks; bundle turnstone/deploy/searxng/settings.yml (JSON output on, limiter off). - Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in). - bootstrap extractor + wheel packaging updated. Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the lxml/h2/brotli transitives). Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG; docs/docker.md carries the AGPL-3.0 §13 operator note. BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg"; tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance. Closes #545 |
||
|
|
948e413f66 |
feat(judge): add Smart Approvals (auto-approve trusted judge verdicts)
Opt-in judge.smart_approvals (default off): when the intent-validation LLM judge returns a high-confidence "approve" verdict, the tool batch is approved automatically with no operator prompt. review/deny recommendations, low confidence, judge errors (llm_fallback), and a deterministic heuristic deny/critical finding all still require a human. Requires judge.enabled. - Batch-atomic: a parallel tool batch auto-approves only if every call qualifies; one non-qualifying call holds the whole batch for a human. - Gate: tier==llm + recommendation==approve + confidence >= judge.confidence_threshold (default raised 0.7 -> 0.95), with a floor that never clears an explicit heuristic deny/critical verdict. - approve_tools waits for the async LLM verdicts, finalises the audit trail (AutoApproveReason.smart_approval), and re-emits verdicts after the card so the live chip updates; the auto-approved row renders the LLM verdict rather than the cautious heuristic carry-over. - judge: always deliver exactly one verdict per call (fallback on error); reject non-finite confidence so NaN can't clear the bar. - Drop verdicts from a superseded judge generation so a reused call_id from a prior turn's still-running daemon can't satisfy the gate's wait. Config plumbed through the server/console/CLI builders and the live _judge_cfg; admin Judge tab renders the toggle. Docs + example config updated. ~35 tests covering the gate matrix, batch-atomicity, the heuristic floor, audit stamping, the streaming re-emit, NaN/duplicate-id defenses, and the cross-turn generation guard. |
||
|
|
9d50e90f67 |
feat(providers): add Claude Opus 4.8 model support
Register `claude-opus-4-8` in the Anthropic capability table. Opus 4.8 shares Opus 4.7's request/response surface exactly — adaptive-thinking only (`budget_tokens` rejected), sampling params removed, the low/medium/high/xhigh/max effort levels, `thinking.display` defaulting to omitted, 1M context, and 128K output — so the entry is a verbatim copy of the 4.7 row. `_lookup_capabilities` longest-prefix matching then resolves date-suffixed ids (e.g. `claude-opus-4-8-20260601`) without colliding with the 4.7 key. No provider code paths change: the existing 4.7 handling already covers all of 4.8's behavior. Models are selected via config.toml / the admin ConfigStore UI, so there is no catalog or dropdown to update. - _anthropic.py: new claude-opus-4-8 capability entry + effort comment - tests/test_providers.py: opus 4.8 bare + dated capability tests - turnstone.example.toml: bump the showcased model example to 4.8 |
||
|
|
752fea0fdd |
feat(console): reactive node discovery via PG LISTEN/NOTIFY dispatcher (#505)
* feat(console): reactive node discovery via PG LISTEN/NOTIFY dispatcher Add a console-side `NotifyDispatcher` that holds a dedicated PostgreSQL `LISTEN` connection and fans wake-ups out to per-channel handlers on a separate dispatch thread. Cluster collector subscribes to a new `services` channel and runs node discovery reactively — new-node / graceful-deregister visibility drops from up-to-60 s to ~500 ms on Postgres, with the 60 s discovery loop retained as the backstop for crash-shaped node loss (NOTIFY only fires on real writes). Storage layer gains a uniform `notify` / `listen` API: - PostgreSQL: real `pg_notify` / `LISTEN` on a dedicated session-mode connection that bypasses pgbouncer (mandatory: pgbouncer is required in transaction-pool mode per docs, which is incompatible with LISTEN). - SQLite: in-process fan-out + synthetic-sweep fallback so consumer code is identical across backends. `TURNSTONE_DB_LISTEN_URL` (or `[database] listen_url` in config.toml) points the dispatcher's connection direct-to-Postgres. Defaults to the main DB URL when unset. Migration 053 installs the `services_notify` trigger; it filters heartbeat-only UPDATEs in-trigger so the 30 s × N-nodes heartbeat tick stays quiet, while INSERT, DELETE, and url/metadata-changing UPDATE still fire. Dispatcher detail: - Two threads: listener (drains stream → bounded queue) and dispatch (invokes handlers under exception suppression). Same-channel notifies coalesce per dispatch batch so an N-node deploy burst is one `_discover_nodes` per channel. - Reconnect uses exponential backoff (1 s → 30 s cap). After any successful reopen — whether the prior failure was a stream-poll error or a connect / initial-LISTEN error — one synthetic Notify with payload="reconcile" is enqueued per channel so handlers re-read on the same code path they use for real events. Future consumers (ConfigStore live reload, scheduler immediate dispatch, audit live-tail) plug in by adding their channel to the dispatcher's construction list. Tests: 22 dispatcher tests (incl. reconnect + coalescing under stub storage), 7 SQLite notify-stream tests, 4 PG-gated trigger-filter tests, 4 collector wire-in tests. All pass; ruff + mypy clean. * fix(notify): address Copilot review on #505 - _sqlite.py: SQLiteBackend.listen() now de-dupes channel names via dict.fromkeys before constructing the stream — duplicates would otherwise register the queue twice and double-deliver each notify. - _sqlite.py: SQLiteBackend.listen() gains a keyword-only sweep_interval parameter (defaults to _SQLITE_NOTIFY_SWEEP_INTERVAL) — matches what the comment at the constant already promised, and lets future consumers without their own polling timer pick a tighter cadence without reaching into private stream attributes. - _sqlite.py: documented the `except queue.Empty: pass` end-of-drain termination so it's not mistaken for swallowing an unexpected error. - _postgresql.py: docstring referenced :func:`_pg_listen_url` which was renamed to _resolve_pg_listen_url during PR development. - notify_dispatcher.py: module docstring referenced a non-existent _bootstrap_console_subsystem; wire-in is at console/server.py::main. Refuted (no change, false positives from github-code-quality bot): - 4× "Statement has no effect" on Protocol-method `...` ellipsis bodies (idiomatic Python Protocol declaration, not dead code). - 2× "Mixed import style" in tests — `import ... as nd_mod` is intentional to allow attribute assignment for monkey-patching the module's `_RECONNECT_BACKOFF_INITIAL` constant inside try/finally. |
||
|
|
eb2a119da9 |
refactor(mcp): remove periodic refresh, add manual refresh/reconnect controls
Deletes the _periodic_refresh task and its supporting state
(_refresh_task, _refresh_failures, _refresh_backoff_until,
_REFRESH_BACKOFF_BASE/MAX, _DEFAULT_REFRESH_INTERVAL, refresh_interval
kwarg) from MCPClientManager. Push notifications and operator-driven
manual refresh now cover all catalog-update needs; the long-running
4-hour timer was dead complexity that obscured the per-user pool
work to come.
Catalog freshness on auto-reconnect is preserved by scheduling an
unblocking _refresh_server task on the mcp-loop after _connect_one
succeeds; the calling thread returns immediately so half-open
recovery latency does not double. Adds MCPClientManager.reconnect_sync
(clears the circuit, closes any existing session, calls _connect_one,
clears stale catalog on failure).
Wires a new pair of operator endpoints —
POST /v1/api/admin/mcp-servers/{name}/refresh and
/v1/api/admin/mcp-servers/{name}/reconnect — that fan out to all
nodes through the existing _internal route family, with per-row
"Refresh" and "Reconnect" buttons in the MCP Servers admin tab.
The new node-internal paths /api/_internal/mcp-{refresh,reconnect}/
are gated to the approve scope to prevent direct unprivileged
reconnects bypassing the console's admin.mcp gate. Internal
endpoints return generic error messages and a filtered status
payload (no command/url) to keep transport details admin-gated.
Drops the [mcp] refresh_interval setting, the
--mcp-refresh-interval CLI flag, and the matching config-mapping
entry; updates docs/architecture.md, docs/tools.md,
docs/settings.md, and the three PlantUML diagrams that referenced
the periodic loop.
Tradeoffs (intentional):
- Idle nodes will not auto-rejoin a recovered MCP server until
traffic arrives or an operator clicks Reconnect. The previous
background reconnection loop is gone by design — push
notifications + operator controls replace it.
- Console fan-out blocks on the slowest node (existing pattern);
not changed here.
This is Phase 1 of the OAuth-MCP series — feature subtraction
ahead of per-user state.
|
||
|
|
551fc43c15 |
feat: per-call model selection on plan_agent / task_agent (#361)
* feat: per-call model selection on plan_agent / task_agent The calling LLM can now pass `model="<alias>"` to plan_agent or task_agent to override the operator-configured per-kind model for that one invocation. Useful when subtask difficulty varies within a session: the model can downgrade to a cheap alias for trivial work and reach for a stronger one when the problem is hard. Tool descriptions list the live registered aliases (refreshed when the operator hits "sync to nodes" / internal_model_reload), so the calling LLM always sees the current options. Bad aliases return a corrective error dict with the available choices so the LLM retries cleanly rather than failing silently. No whitelist — any alias the registry knows is acceptable; cost control is intentionally ceded to the model. No per-call effort override (out of scope; effort stays operator-configured). Resolution precedence in _run_agent: explicit per-call agent_alias override > registry per-kind (plan_model/task_model) > legacy agent_model > session model. The plan retry path (when _validate_plan fails) reuses the same alias so coaching reflects real model behaviour rather than a different model masking the signal. Implementation: - plan_agent.json / task_agent.json: optional `model` parameter. - ChatSession._validate_agent_model_override extracts and validates the arg; mirrors the existing empty-prompt error pattern. - _prepare_plan / _prepare_task stash the override in item["model_override"]; _exec_* pass it through. - _run_agent gains agent_alias kwarg with defence-in-depth ValueError on unknown alias. - _render_agent_tool_descriptions deep-copies plan/task entries before mutating description so the module-level TOOLS constant stays untouched across sessions; rebuilds the BM25 tool-search index when active so its text matches what the LLM sees. - server._broadcast_agent_tool_schema_refresh walks active workstreams on internal_model_reload so descriptions update without restart. * fix: clarify no-registry placeholder + avoid double BM25 rebuild Addresses Copilot feedback on PR #361. 1. plan_agent.json / task_agent.json placeholder said the parameter falls back to the "operator-configured plan/task model". That text is what no-registry sessions see (registry-bearing sessions get the templated description with the live alias list); for those single-model sessions, omitting the param falls back to the current session model, not an operator-configured one. Reword so the no-registry user gets accurate guidance. 2. _on_mcp_tools_changed already calls _rebuild_tool_search after merging MCP tools. _render_agent_tool_descriptions also rebuilt the BM25 index when active, so the MCP refresh path was rebuilding twice per refresh. Move the BM25 rebuild out of the private render helper into the public refresh_agent_tool_schemas wrapper — _on_mcp_tools_changed keeps calling the render helper directly (no double rebuild), and registry-reload callers go through the wrapper which still keeps the index in sync. |
||
|
|
6c026710ff |
feat: ConfigStore + admin UI for plan/task agent model and effort (#360)
* feat: ConfigStore + admin UI for plan/task agent model and effort Per-kind sub-agent routing was added in #359 but only via config.toml. Operators can now switch the plan_agent / task_agent model and reasoning effort at runtime from the admin Model tab without restarting. Adds four ConfigStore-backed settings: model.plan_alias — alias for plan_agent model.task_alias — alias for task_agent model.plan_effort — reasoning effort for plan_agent model.task_effort — reasoning effort for task_agent Server startup and internal_model_reload both apply these as overrides on top of the registry's config.toml-loaded values; the new logic computes "effective" values for all five model-routing fields and only calls registry.reload() when at least one differs. Admin UI: extracts ALIAS_SETTING_KEYS to a const used by both the dynamic-alias-choice injection and the empty-option label rendering. Adds INHERIT_EMPTY_LABEL_KEYS so plan_effort / task_effort show "(inherit)" for empty — distinct from the literal "none" choice (which actually disables reasoning, very different from leaving unset). Also fixes Copilot review feedback from #359: - _validate_effort treats empty / whitespace as unset rather than warning on benign explicit-empty configs (with .strip().lower() normalisation; "HIGH" and " low " now parse correctly) - turnstone.example.toml's reasoning_effort comment lists the full set of accepted values (none, minimal, low, medium, high, xhigh, max) * fix: apply routing overrides on config-reload + skip no-op model-reload Addresses Copilot feedback on PR #360. 1. Admin settings updates fan out via /_internal/config-reload, which only reloaded the ConfigStore — plan/task routing changes weren't visible until a model-reload or restart, defeating the runtime configurability this PR is meant to add. 2. /_internal/model-reload always called registry.reload(), churning cached clients even when nothing changed. Risky when fanned out across nodes (could close in-flight clients). Extracts two helpers in server.py: - _effective_routing(cs, ...) pure function: overlay CS values on base - _apply_routing_overrides(reg, cs) reload only when something differs Used by the startup path, config_reload (new), and model_reload (now short-circuits with a noop response when models + routing are unchanged). |
||
|
|
54dd557476 |
feat: split plan_model and task_model, configurable agent reasoning effort
plan_agent and task_agent previously shared a single agent_model knob and
plan_agent hardcoded reasoning_effort="high" in three call sites. They
have different cost/latency profiles — plan is rare and benefits from a
stronger model, task is frequent and benefits from a cheaper one — so
sharing the knob undertunes both.
ModelRegistry gains plan_model, task_model, plan_effort, task_effort.
Per-kind overrides win over the legacy agent_model, which still works
as the single-knob fallback for both. resolve_agent_alias(kind) and
resolve_agent_effort(kind) centralise the resolution; PLAN_DEFAULT_EFFORT
captures the back-compat "high" default in one place rather than at
every call site.
session._run_agent delegates resolution by label ("plan" vs "task").
The three hardcoded reasoning_effort="high" arguments are removed —
behaviour is identical when no plan_effort is configured.
Loader validates effort against {none,minimal,low,medium,high,xhigh,max}
and warns + drops typos rather than passing them to the provider.
ConfigStore parity and admin UI for the new knobs are deferred to a
follow-up — config.toml-only is enough for the backend split.
|
||
|
|
30c89f46c6 |
feat: add Claude Opus 4.7 support (#357)
- Add claude-opus-4-7 capability entry (1M ctx, 128K output, adaptive thinking, supports_temperature=False, thinking_display=summarized) - Suppress temperature param for Opus 4.7 (API returns 400) - Add thinking display opt-in via new ModelCapabilities.thinking_display field - Opus 4.7 omits thinking by default, always send summarized - Add xhigh effort level to mapping and Opus 4.7 effort_levels - Add xhigh/max options to skill template dropdowns in admin console - Align reasoning effort label capitalization across all console dropdowns - Update example config to reference claude-opus-4-7 - 10 new tests with regression guards for Opus 4.6 backward compat Verified against live API: streaming and completion calls succeed. |
||
|
|
62d2a0fe6a |
fix: remove non-auth support from bootstrap wizard (#274)
* fix: remove non-auth support from bootstrap wizard Auth is now mandatory for all deployments. Remove the TURNSTONE_AUTH_ENABLED toggle and make JWT_SECRET and AUTH_TOKEN required in the wizard's system prompt. * fix: remove auth disable support from runtime and infra Remove AuthConfig.enabled field — auth is always on. Drop TURNSTONE_AUTH_ENABLED env var, config toggle, and the check_request bypass. Update compose.yaml, Helm chart, Terraform, docs, and tests to match. * feat: deprecate config tokens, require JWT secret, prefer JWT auth Phase 1 of config-token removal: - load_jwt_secret() now exits with error if no secret is configured (was: silently auto-generated ephemeral secret) - _authenticate_token() logs deprecation warning on config token use - CLI /cluster commands use ServiceTokenManager when JWT secret is set - turnstone-admin tls-list uses ServiceTokenManager when JWT secret is set - Update bootstrap wizard, docker.md, security.md to mark TURNSTONE_AUTH_TOKEN as deprecated and JWT_SECRET as required - Console test fixtures use auth token + headers (auth always enforced) * feat: add service scope for inter-service JWT auth Add "service" to VALID_SCOPES and SCOPE_HIERARCHY. Service tokens bypass require_permission() RBAC checks, replacing the old empty-user-id bypass that config tokens relied on. All ServiceTokenManager instances that need admin access now include "service" in their scopes (console proxy, channel gateway, CLI, admin CLI). Read-only services (collector, notification) unchanged. * feat: phase 2 config token deprecation - SDK doc examples now show API tokens (ts_) instead of config tokens - Remove _get_config_token() from admin CLI (dead code) - Block config token exchange in handle_auth_login — only password and API token login allowed - Update login tests to use password-based auth instead of config token exchange * feat: phase 3 — remove config tokens entirely Complete removal of config-file token authentication: - Delete AuthConfig.tokens, check(), _ROLE_TO_SCOPES, hmac dispatch branch, and config token loading from load_auth_config() - Remove auth_config parameter from _authenticate_token() and check_request() — callers updated throughout - Remove TURNSTONE_AUTH_TOKEN from compose.yaml, Helm charts, Terraform, turnstone.example.toml - Remove --auth-token CLI flags from turnstone, turnstone-admin, and turnstone-console - Simplify console main() — always use ServiceTokenManager (no fallback to static tokens) - Delete config-token-specific tests, rewrite check_request and integration tests to use JWT auth with proper audience claims - Remove all config token references from docs (security.md, docker.md, sdk.md, console.md, architecture.md, bootstrap prompt) * fix: address code review findings - Fix 33 broken tests: add JWT auth to test_api_versioning, test_console_routing_proxy, test_tls_admin, test_tls_manager, test_server_live (jwt_secret + audience-scoped auth headers) - Add TestRequirePermissionServiceScope: 4 tests covering the service scope RBAC bypass path - Remove stale comments referencing config tokens in auth.py and console/server.py - Remove dead proxy_auth_token parameter from console create_app() and static token fallback in _proxy_auth_headers() - Remove TURNSTONE_AUTH_TOKEN from env.py scrub list * fix: address Copilot review — JWT audience, compose require secret - CLI /cluster: add audience=JWT_AUD_CONSOLE to ServiceTokenManager (console validates audience, JWTs without it were rejected) - Admin CLI tls-list: same audience fix - compose.yaml: TURNSTONE_JWT_SECRET now uses :? to fail fast if unset - SDK console: fix default port from 8081 to 8090 * test: add auth enforcement tests for TLS admin endpoints 5 new tests: unauthenticated requests return 401 (list, renew, delete), read-only-scoped requests return 403 (renew, delete). Closes the TLS auth enforcement test gap noted in PROGRESS.md. * fix: address remaining Copilot review feedback - Fix token_source="config" → "test" in TLS test fixtures - Fix AuthResult.token_source docstring to include service origins - Require TURNSTONE_JWT_SECRET in cluster compose profile (:?) - Helm: add auth.jwtSecret + auth.existingSecret values, wire TURNSTONE_JWT_SECRET into secret.yaml and both deployments - Terraform: replace auth_token with jwt_secret variable + secret, remove orphaned auth_token resources and IAM reference - Remove [[auth.tokens]] from security.md config example * fix: address full code review — 10 findings Critical: - Terraform: replace concat(common_env, auth_env) with common_env (auth_env local was removed but still referenced) - Channel gateway: remove hmac static token auth from _check_auth(), use JWT-only validation. Remove --auth-token CLI arg from channel - Rebalancer: add token_manager support so migration requests carry JWT auth (was sending unauthenticated POST to /internal/migrate) Major: - Guard _permissions_to_scopes() against "service" privilege escalation from DB role permissions - Remove dead AuthConfig class, load_auth_config(), and all auth_config parameters from create_app() signatures - Helm: inject JWT secret for both inline and existingSecret paths Minor: - Remove dead auth_token param from ClusterCollector - Remove empty TestLoadAuthConfig class - Short JWT secret now exits instead of warning - Compose: add generation command comment above JWT_SECRET - Clean stale config token references from 6 doc files - Clean stale AUTH_TOKEN reference from bootstrap wizard prompt * fix: remove remaining stale config token references from docs - channels.md: remove --auth-token from options table - oidc.md: remove "config-file tokens still work" claim - security.md: remove config token section, fix JWT secret docs (now required/exits, no ephemeral fallback), remove hmac from ASCII diagram, remove --auth-token reference |
||
|
|
a7d9461735 |
refactor: channel router + scheduler use SDK clients
ChannelRouter: replace raw httpx with AsyncTurnstoneServer (single-node) and AsyncTurnstoneConsole route methods (multi-node). Remove _post() helper, _route_path(), and manual JSON construction. Scheduler: replace raw httpx.Client with TurnstoneServer (sync). Lazy per-node client cache with token rotation and stale client pruning. Clean remaining Redis/MQ references from tests, docs, and config: - test_tls_admin: redis.internal -> app.internal - test_config: [redis] test data -> [database] - docs/channels.md, console.md: rewrite for HTTP architecture - docs/api-reference.md, openshell.md: remove stale diagram/Redis refs - turnstone.example.toml: remove [redis] section - .pre-commit-config.yaml: remove types-redis dependency - QUICKSTART.md: remove bridge/Redis from deployment descriptions |
||
|
|
274c97135e |
feat: TLS storage backend + config for lacme integration (#178)
Storage layer for mTLS certificate management via lacme: - Migration 026: tls_account_keys, tls_ca, tls_certificates tables - 8 protocol methods on StorageBackend (save/load account keys, CA, certs; list/delete certs) - SQLite and PostgreSQL implementations with dialect-specific upserts - StorageStore adapter bridging lacme's Store protocol to turnstone storage (bytes↔str PEM conversion, CertBundle↔dict mapping) - Settings registry: tls.enabled (bool), tls.acme_directory (string) - Config.toml: [redis] TLS and [database] SSL passthrough params - lacme>=1.0.1 as optional [tls] dependency - 20 unit tests covering storage CRUD + adapter + crypto roundtrip |