Compare commits

...

37 Commits

Author SHA1 Message Date
Patrick Buckley f1f448277f chore: bump version to 0.5.6
Prompt template runtime wiring, security hardening, scheduler/channel/MQ
template support, migration 010.
2026-03-12 16:58:20 -07:00
Patrick Buckley 2f7f70825b feat: wire prompt templates into session startup with full creation-p… (#47)
* feat: wire prompt templates into session startup with full creation-path support

Prompt templates (prompt_templates table) now have runtime effect:

- is_default=true templates auto-apply as system message content,
  concatenated in name order before user instructions
- Per-workstream template selection via --template CLI flag, template
  field on POST /v1/api/workstreams/new, console creation modal dropdown,
  scheduled task config, and channel adapter config
- {{model}}, {{ws_id}}, {{node_id}} variable substitution via single-pass
  regex (prevents cross-variable injection)
- /template slash command for runtime switching, persisted across resume
- set_template() public API on ChatSession

Security hardening:
- MCP sync resets is_default=False on content update (prevents compromised
  server from injecting defaults)
- 32KB content cap on template create/update + defensive truncation
- Template existence validation returns 400 before workstream creation
- Single-pass regex eliminates cross-variable expansion

Template field plumbed through all creation paths: CLI, server API, MQ
protocol/bridge, console backend, scheduler dispatch, channel router,
MQ client. Migration 010 adds template column to scheduled_tasks.

Frontend: console workstream modal template dropdown, scheduler
create/edit template field, governance template UI variables auto-detected
from content (read-only display replaces editable input). Focus trap and
Enter-key accessibility fixes in workstream modal.

Docs: governance.md template runtime section, api-reference.md template
field, governance + MCP architecture diagrams updated.

Python + TypeScript SDKs, Pydantic schemas all updated. 29 new tests.

* fix: address PR #47 review feedback

- Defer template validation until after resume_ws — a bad template name
  no longer 400s when resume would have ignored it anyway
- Add template field to OpenAPI JSON specs (openapi-server.json,
  openapi-console.json) for SDK/docs consistency
- Validate template existence in schedule create and update endpoints —
  reject unknown template names with 400 instead of allowing schedules
  that would silently fail at dispatch time
2026-03-12 16:57:31 -07:00
Patrick Buckley 4866c9873c feat: ddgCluster compose profile with DuckDuckGo Search MCP sidecar (#48)
Add ddgCluster profile extending the 10-node cluster with a DuckDuckGo
Search MCP sidecar. All cluster nodes connect via streamable-http and
gain duckduckgo_web_search + duckduckgo_fetch_content tools. No API
key required.

Key implementation details learned during testing:
- MCP SDK DNS rebinding protection must be disabled for Docker
  internal networking (Host header uses container names)
- FastMCP server binds to 127.0.0.1 by default; must set
  mcp.settings.host='0.0.0.0' for cross-container access
- DDG CLI lacks --host/--port flags; settings configured via Python
  entry point that patches FastMCP.settings directly
- Safe search disabled by default

Also adds MCP_CONFIG env var support to all server commands (shell
conditional, no-op when empty) and moves default server/bridge to
production profile for cleaner profile separation.
2026-03-12 16:51:53 -07:00
Patrick Buckley 8b2e2130fc fix: MCP resource template URI expansion via prefix matching (#46)
* fix: MCP resource template URI expansion via prefix matching

Resource templates (RFC 6570 URI patterns like `db://tables/{table}/rows/{id}`)
were discovered from MCP servers but non-functional — `read_resource_sync()`
only accepted exact URIs from `_resource_map`, which excludes templates.

Add prefix-based fallback: extract the static prefix from each template
(everything before the first `{`), store a prefix→server mapping, and
fall back to longest-prefix matching when exact URI lookup fails. MCP
servers handle URI routing internally so we just need to route the
expanded URI to the correct server.

Also surface templates in the system message catalog and `/mcp` command
so the model knows they exist and can construct expanded URIs.

* fix: address PR #46 review feedback

- Template prefix collision now keeps more specific (longer) template
  URI instead of blindly overriding
- Fix _match_template docstring to accurately describe startswith
  matching on static prefixes (not full template matching)
- Add missing loop.close() in integration test finally block
- Rewrite test_template_longest_prefix_wins with genuinely different
  prefix lengths to avoid brittle collision-order dependency
2026-03-12 15:46:00 -07:00
Patrick Buckley f81c06761d chore: remove dead code, add MCP integration + collector tests (#45)
* chore: remove dead code, add MCP integration + collector tests

Remove unused delete_prompt_templates_by_server from protocol and
both storage backends (sync uses per-template deletion).

Add 10 MCP integration tests exercising full lifecycle: rebuild
resources/prompts, read_resource_sync/get_prompt_sync with real
asyncio loop, governance sync to real SQLite, shutdown cleanup,
listener notification isolation.

Add 3 console collector MCP aggregation tests: multi-node sums,
absent when zero, mixed nodes with/without MCP.

* fix: close event loops and SQLite backend in MCP integration tests
2026-03-12 15:16:51 -07:00
Patrick Buckley be165c1971 feat: MCP resource and prompt discovery with read_resource tool (#44)
* feat: MCP resource and prompt discovery with read_resource tool

Extends MCPClientManager with resource and prompt discovery alongside
existing tool support. Resources and prompts are discovered on connect,
cached per-server with copy-on-write rebuilds, and refreshed via push
notifications, periodic polling, or manual /mcp refresh.

New read_resource built-in tool reads MCP resources by URI. Requires
user approval (same as MCP tool calls) since resources are served by
external MCP servers. Resource catalog injected into system message
with XML delimiters. Error messages sanitized to prevent leaking
server internals to the model.

Prompt discovery stores prefixed names (mcp__server__prompt) and
exposes get_prompt_sync() for future use_prompt tool (Chunk D).

/mcp command now shows tools, resources, and prompts. Docs and
diagrams updated.

* feat: MCP prompt governance sync with origin tracking and readonly guards

Migration 009 adds origin, mcp_server, and readonly columns to
prompt_templates. MCP prompts discovered by MCPClientManager are
automatically synced into the governance table as read-only templates
with origin="mcp".

Sync engine handles: create on connect, update on prompt refresh,
delete when prompts are removed from server. Manual templates take
precedence on name collision (MCP prompt skipped with warning).

Admin API returns 403 on update/delete of readonly templates. Console
UI shows MCP origin badge and disables edit/delete buttons. Storage
backends gain get_prompt_template_by_name, list_prompt_templates_by_origin,
and delete_prompt_templates_by_server methods.

Also addresses PR #44 review feedback: concurrent.futures.TimeoutError
handling in sync dispatch, XML-escape resource catalog descriptions,
resource template entries excluded from _resource_map, URI collision
warnings, needs_periodic capability-aware computation, malformed JSON
primary key fallback for read_resource.

* feat: use_prompt tool, prompt catalog, and PR review hardening

New use_prompt built-in tool invokes MCP prompt templates by name,
expanding them into messages. Requires user approval (external MCP
servers). Prompt catalog injected into system message with XML
delimiters (up to 30 prompts, HTML-escaped).

Prompt listener registered in session for catalog rebuild on changes.

Addresses PR #44 review feedback:
- _init_system_messages() now uses copy-on-write (build locally,
  assign atomically) so background thread callbacks never see
  partial system messages
- sync_prompts_to_storage() serialized behind _sync_lock to prevent
  races between set_storage() (main thread) and MCP background thread
- shutdown() clears listener lists to release callback references

Docs and diagrams updated for 18 built-in tools.

* feat: granular tool policies for MCP resources, prompts, and tools

Policy evaluation now uses approval_label (falling back to func_name)
for fnmatch pattern matching, enabling fine-grained per-URI and
per-server policies:
- read_resource: mcp_resource__{normalized_uri}
- use_prompt: mcp__{server}__{prompt} (prefixed name)
- MCP tools: mcp__{server}__{tool} (was static "mcp_tool")

URI normalization resolves .. path segments to prevent traversal
bypasses in policy matching. Resource templates filtered from system
message catalog (not directly readable). use_prompt arguments
validated as dict with string coercion.

TypeScript SDK PromptTemplateInfo gains origin, mcp_server, readonly
fields. Governance docs updated with MCP policy patterns.

* feat: MCP visibility in server and console UIs

Server health endpoint includes mcp.servers, mcp.resources, mcp.prompts
counts. Server UI status bar shows magenta MCP indicator with tooltip.
Console cluster status bar shows MCP metrics with magenta LED dot.
Console node detail view shows per-node MCP summary. Console collector
aggregates MCP counts across nodes in overview.

Uses var(--magenta) design token with new --magenta-glow for theme
adaptation. ARIA roles on MCP status elements. Tooltips on console
MCP metric labels. Node MCP summary hidden on mobile (< 700px).

New diagram: 20-mcp-architecture.puml covering full MCP lifecycle
(connection, discovery, refresh, governance sync, policy, UI).

* fix: McpStatus in health schema, count properties, catalog name fidelity

Adds McpStatus model to HealthResponse (Python + TypeScript SDKs) so
typed clients see the mcp field from /health.

Addresses Copilot review feedback:
- resource_count/prompt_count properties avoid list allocation on
  /health and /metrics polls
- get_tools/resources/prompts return shallow-copied dicts to prevent
  callers from mutating internal cache
- Prompt names and arg names in system message catalog are NOT
  HTML-escaped (model must use exact strings in use_prompt calls);
  only descriptions are escaped

* fix: OpenAPI spec McpStatus + diagram approval column accuracy

Adds McpStatus schema and optional mcp field to HealthResponse in
openapi-server.json, matching the Python schema and TypeScript types.

Fixes tool pipeline diagram: math, web_fetch, web_search correctly
shown as auto-approve (not "Yes" for approval).
2026-03-12 14:49:58 -07:00
Patrick Buckley 3264fdefca fix: channel bidirectional routing — emit TurnCompleteEvent on all id… (#43)
* fix: channel bidirectional routing — emit TurnCompleteEvent on all idle transitions

Bridge previously only emitted TurnCompleteEvent for MQ-initiated turns
(those with a correlation_id in _active_sends). Server-UI-initiated turns
went idle without emitting TurnCompleteEvent, so the Discord bot's
StreamingMessage never finalized — content accumulated in the buffer and
collided with the next Discord-triggered response.

Now TurnCompleteEvent is emitted unconditionally on every idle transition.
correlation_id is empty for non-MQ turns; SDK client filters by
correlation_id so existing consumers are unaffected.

* fix: remove unused variable flagged by ruff
2026-03-12 11:57:15 -07:00
Patrick Buckley 28cb3a5c51 fix: approval timeout UI state and content flush before tool calls (#42)
* fix: approval timeout UI state and content flush before tool calls

Two bug fixes:

1. Approval timeout now shows denied state in UI — resolve_approval()
   emits an approval_resolved SSE event so the browser transitions
   from pending to denied (red border + badge). Also fixes the cancel-
   during-approval path. Frontend resolveInlineApproval() gains a
   skipPost parameter to avoid redundant POST when server-initiated.
   ApprovalResolvedEvent added to Python and TypeScript SDKs.

2. Content streaming flushes pending buffer before tool call deltas —
   _stream_response() held up to 13 trailing chars in the pending
   buffer (for <think> tag detection) when transitioning to tool calls.
   Now flushed eagerly when tool_call_deltas arrive, before clearing
   in_think so reasoning text is correctly categorized.

* fix: address Copilot review feedback on PR #42

Patch _execute_tools in stream flush test to prevent real bash execution,
simplify confusing nested comprehension, and update resolve_approval()
docstring to reflect cancel/timeout call paths.
2026-03-12 11:43:06 -07:00
Patrick Buckley 8b11e0a6f9 fix: bridge retries node_id fetch indefinitely with capped backoff 2026-03-12 11:12:03 -07:00
Patrick Buckley 648ba477e1 refactor: list_user_roles uses _row_to_dict instead of positional row mapping 2026-03-11 21:14:58 -07:00
Patrick Buckley 7960784786 fix: usage events now record per-request tool_calls delta, not cumulative total 2026-03-11 21:12:03 -07:00
Patrick Buckley e06554d1ec feat: add channel admin endpoints to console OpenAPI spec 2026-03-11 21:08:28 -07:00
Patrick Buckley 8eb8722346 Bump version to 0.5.5 2026-03-11 20:24:12 -07:00
Patrick Buckley a2e2ffacd8 feat: robust plan quality gate, iterative refinement, and amend UX (#41)
* feat: robust plan quality gate, iterative refinement, and amend UX

Plan agent output from weak models often produced garbage (11-char plans
that echo the prompt). Two fixes:

1. Quality validation (_validate_plan) checks length, section structure,
   echo detection, and refusal patterns. Fails trigger one automatic
   retry with a coaching message injected into the agent's existing
   conversation, preserving all prior exploration context.

2. Iterative feedback loop — user feedback at plan review re-runs the
   plan agent via _refine_plan() instead of appending text to the tool
   result. Up to 5 refinement rounds. The plan file path is always
   included in the tool result so the outer model knows where it lives.

UI improvements:
- Web: Reject button dynamically becomes "Amend" (amber) when feedback
  is typed. Key hint badges (Esc/Enter) on plan buttons. Main input
  disabled during review. Light-theme contrast fix via --on-color var.
- CLI: Prompt shows all three actions (approve/amend/reject).
- Bridge: Race condition fix — clear pending entry before HTTP POST so
  sequential plan reviews from the refinement loop aren't skipped.

15 new tests covering validation, retry, and refinement.

* fix: address PR 41 review feedback

- Escape key in plan dialog now mirrors the Amend button: if feedback is
  typed, Esc sends the feedback (amend); if empty, Esc rejects. Previously
  Esc always hard-coded "reject", discarding typed feedback.

- Coaching message for plan retry now says "should include at least two of"
  instead of "MUST include these", matching the actual validation rule
  (_MIN_PLAN_SECTIONS = 2).

* feat: render plan inline in chat after approval

After the plan review dialog closes, the plan content is now rendered
as a collapsible inline block in the chat stream — styled with a
status header (approved/rejected/amending), markdown-rendered body,
and feedback note when amending. Uses the same makeCollapsible pattern
as tool output blocks.

* fix: prevent plan approval hang when inline render fails

The authFetch call that unblocks the server must fire before the
cosmetic inline plan rendering. Previously _addInlinePlan ran first
and any JS error (e.g. from renderMarkdown) prevented the API call,
leaving the session thread blocked forever.

- Move authFetch before _addInlinePlan
- Wrap _addInlinePlan in try-catch
- Guard against empty content
- Only auto-collapse plans longer than 12 lines

* fix: address PR 41 review feedback (round 2)

- Max refinement rounds no longer implicitly approve: the loop now
  shows the final plan for explicit approve/reject before proceeding.
  Previously exhausting 5 rounds silently accepted the last revision.

- Plan inline block: correct aria-label from "Tool output" to
  "Plan content" when makeCollapsible is applied.

- XSS concern (not applicable): renderMarkdown is used for all
  assistant messages — plan content follows the same trust model.

- Test loop concern (acknowledged): refinement tests verify component
  logic; full _execute_tools integration would require extensive
  mocking for marginal coverage gain.

* feat: thinking spinner + inline plan hardening

* fix lint
2026-03-11 20:22:41 -07:00
Patrick Buckley c6ba8d59b0 feat: bootstrap wizard — LLM-guided interactive setup for deployments
Add `turnstone-bootstrap`, a new entry point that uses any LLM (OpenAI,
Anthropic, or local/vLLM) to conversationally walk users through
configuring a Turnstone deployment. Generates .env files, setup.sh
scripts, and optional docker-compose overrides.

- Fully interactive startup (zero CLI args) with provider/model selection
- Auto-detects available models on local OpenAI-compatible endpoints
- 7 tools: read_file, write_file, generate_secret, check_port,
  validate_api_key, check_docker, finish
- Path traversal protection on file read/write
- Duplicate write detection (skips identical content)
- Bounded retry loop (3 attempts) on LLM errors
- Anthropic message conversion with consecutive-role merging
2026-03-11 01:58:39 -07:00
Patrick Buckley 087f5b49f6 Bump version to 0.5.4 2026-03-10 20:47:40 -07:00
Patrick Buckley fd507c6a3c feat: generation cancellation — stop button, cancel API, cooperative … (#40)
* feat: generation cancellation — stop button, cancel API, cooperative cancel

Add cooperative cancellation via threading.Event on ChatSession. The cancel
signal is set from outside the worker thread (HTTP handler, MQ bridge, or
Escape key) and checked at defined checkpoints: per streaming chunk, before
tool execution, inside bash commands, and at each sub-agent turn.

Core: GenerationCancelled(BaseException) exception, cancel()/_check_cancelled()
methods, partial content preservation in _stream_response, clean rollback in
send() with idle state emission (no re-raise).

Server: POST /v1/api/cancel endpoint, CancelledEvent SSE emission, worker
thread safety net.

Frontend: Stop button (■ Stop) with send/stop swap via setBusy(), Escape key
shortcut, cancelled event handler. Accessible: aria-label, focus-visible
override, light theme contrast, non-color differentiation.

MQ: CancelMessage inbound type, bridge _handle_cancel routed handler.

SDK: cancel() on Python async+sync clients, CancelledEvent in Python+TypeScript
event registries, isCancelledEvent type guard.

OpenAPI: CancelRequest schema + endpoint spec.

Docs: API reference, architecture, SDK docs updated. Diagrams: conversation
turn, tool pipeline, MQ protocol, workstream states, SDK architecture.

* fix: address PR #40 review feedback

- setBusy() now resets stopBtn.disabled so stop button is re-enabled on
  next generation after a successful cancel
- Gate cancel side effects (resolve_approval, resolve_plan, cancelled SSE
  event) on worker_thread.is_alive() to avoid spurious events when idle
- Add /v1/api/cancel endpoint and CancelRequest schema to TypeScript
  openapi-server.json to keep it in sync with Python-generated spec
2026-03-10 20:43:52 -07:00
Patrick Buckley 562c3c8ab7 docs: add governance section and missing diagrams to README 2026-03-10 19:50:07 -07:00
Patrick Buckley 4773535bb8 docs: add governance architecture diagram PNG 2026-03-10 19:41:35 -07:00
Patrick Buckley 7492816ab2 feat: governance — RBAC, tool policies, prompt templates, usage track… (#39)
* feat: governance — RBAC, tool policies, prompt templates, usage tracking, audit logging

Add comprehensive governance layer for the admin console:

- RBAC with 15 granular permissions, 3 builtin roles (admin, operator, viewer),
  custom role CRUD, user-role assignment with privilege escalation prevention
- Tool policies with glob pattern matching, priority-ordered evaluation
  (allow/deny/ask), enforced before auto-approve in WebUI.approve_tools()
- Prompt templates with variable substitution, categories, default flag
- Usage tracking: per-LLM-request token/tool metrics, aggregated queries
  (group by day/model/user), automatic 90-day pruning via scheduler
- Audit logging: append-only event trail for all admin mutations,
  filterable/paginated queries, automatic 365-day pruning, X-Forwarded-For
  aware IP extraction
- require_permission() enforced on all 35+ admin endpoints (users, tokens,
  channels, schedules, watches, roles, orgs, policies, templates, usage, audit)
- Field allowlists on storage update methods prevent mass-assignment bugs
- Self-deletion guard on admin_delete_user, delete_user cascades user_roles
- _row_to_dict helper eliminates ~400 lines of fragile positional row mapping
- _audit_context helper deduplicates 18 instances of audit boilerplate
- Migration 008: 7 new tables, 3 builtin roles, org_id on users
- Console admin panel: 5 new tabs (Roles, Policies, Templates, Usage, Audit)
  with permission-gated visibility, 7 modal dialogs, full keyboard accessibility
- Python + TypeScript SDK methods for all governance endpoints
- 120+ new tests (1554 total)

* fix: address PR #39 review feedback

- Rebuild serialized items after policy evaluation so denied/allowed
  verdicts are reflected in tool_info/approve_request SSE payloads
- Make `since` query param optional in usage OpenAPI spec (handler
  already defaults to last 7 days)
- Add response_model=StatusResponse to DELETE role/policy/template
  and POST/DELETE role assignment endpoints in OpenAPI spec
- Add missing org_id/created/updated fields to UserRoleInfo schema
- Add missing created field to AuditEventInfo schema
- Show "no permissions" empty state instead of loading inaccessible
  tab when all admin tabs are permission-gated
- Fix "13 permissions" → "15 permissions" in architecture.md and
  security.md
- Fix import sorting in test_audit.py and test_tool_policy.py

* fix: address PR #39 round 2 review feedback

- Clear stale permissions from sessionStorage on config-token login
  (auth.js _storePermissions)
- Only trust X-Forwarded-For when behind a proxy that sets
  X-Forwarded-Proto (conditional on is_secure_request trust model)
- Thread user_id from auth into WebUI.on_status for usage events
- Add created field to TS AuditEventInfo type
- Return typed Pydantic models from all SDK governance methods instead
  of dict[str, Any] — both async and sync clients
- Validate group_by param against allowed enum in admin_usage handler
- Add deterministic secondary sort (event_id DESC) to
  list_audit_events in both SQLite and PostgreSQL backends
2026-03-10 19:32:37 -07:00
Patrick Buckley d6ba1d5e25 fix: mypy no-any-return in agent context overflow handler 2026-03-10 14:11:09 -07:00
Patrick Buckley 41d1b27d34 Bump version to 0.5.3 2026-03-10 13:46:50 -07:00
Patrick Buckley 8bc284c60e fix: agent context overflow — truncate tool output, catch context errors
Agent tool outputs are now truncated to 16k chars to prevent search
results (14M+ chars observed) from blowing past the model's context
limit. On context-exceeded API errors, the agent returns its last
content instead of crashing.
2026-03-10 13:43:01 -07:00
Patrick Buckley a322d6b1d1 fix: sub-agent context — clean plan, merged task, no Qwen template error
Plan agent: own identity only (no base system prompt needed).
Task agent: base system prompt merged with task identity into a single
system message (needs tool patterns for tool execution).
Neither agent receives conversation history.

Fixes Jinja template error on Qwen models that reject system messages
appearing after non-system messages.
2026-03-10 13:36:23 -07:00
Patrick Buckley 70d495aa5b fix: per-workstream SSE fan-out — multiple consumers no longer steal … (#38)
* fix: per-workstream SSE fan-out — multiple consumers no longer steal each other's tokens

After 4d665a5 removed the single-consumer SSE lock, the shared
_event_queue let concurrent consumers (browser, bridge, console proxy)
race on Queue.get(), each receiving ~1/N of content tokens and producing
garbled streaming text.

Replace the single queue with per-client fan-out: each SSE connection
registers its own bounded queue (maxsize=500) on WebUI._listeners, and
_enqueue() copies every event to all registered queues. On eviction or
close, a ws_closed sentinel is injected so SSE generators exit promptly.

* fix: address CI failures and Copilot review feedback

- Handle ws_closed sentinel in events_sse generator (break on close)
- Guarantee sentinel delivery by evicting one item when queue is full
- Clear listeners list after injecting sentinels on cleanup
- Fix test_slow_consumer to fill only slow queue directly
- Fix ruff SIM117 (nested with), unused import, mypy unused-ignore
2026-03-10 13:25:16 -07:00
Patrick Buckley de64535221 Feat/eval improvements (#37)
* ci: add GitHub Release creation on tag push

* refactor: rename plan tool to create_plan

Rename plan → create_plan to resolve cross-provider tool selection
failures. Models consistently treated "plan" as a reasoning concept
rather than a callable tool. The new name is an unambiguous verb+noun
action. Also rename the parameter from prompt → goal for clarity,
add web_search to the default system prompt tool patterns

* feat: eval harness improvements inspired by autoresearch patterns

Major enhancements to turnstone-eval:

- Per-test timeout (--test-timeout, default 300s) and suite timeout
  (--suite-timeout) prevent stuck runs from blocking the suite
- Fast-fail skips remaining runs after ceil(n/2) consecutive zeros
- Summary table with colored PASS/WEAK/FAIL and append-only TSV output
- Progress reporting with running pass rate, token count, and ETA
- Parallel test execution via ProcessPoolExecutor (--parallel N)
- Per-role model assignment: test/optimizer/observer can use different
  models and providers (--optimizer-model, --observer-model, etc.)
  with auto-detection from base URL
- Improved optimizer and observer system prompts with structured
  failure-mode diagnosis, keep/discard rules, and trend analysis
- Fixed token counting (prompt tokens use last-turn value, not sum)
- Added math-calculation and web-search-query test cases
- Fixed multi-file-edit test (both files now contain the target string)

* fix: address Copilot review feedback

- Revert prompt token counting to sum (reflects billed usage)
- Add tool_args to fast-fail skipped run dicts for schema consistency
- Align approval_label with func_name ("create_plan")
- Add timeout to future.result() in parallel path (test_timeout + 30s)
- Document thread-leak trade-off on serial timeout path
2026-03-10 13:06:39 -07:00
Patrick Buckley 187d004033 feat: watch tool — periodic command polling within workstreams (#36)
* feat: watch tool — periodic command polling within workstreams

Add a new `watch` tool that lets the model (or user) set up periodic
polling of a shell command. Results inject as synthetic user messages
that trigger LLM turns, enabling reactive workflows like PR monitoring,
CI/CD status tracking, and deployment health checks.

Key design:
- Single tool with create/list/cancel actions
- Python expression DSL for stop conditions (restricted eval)
- Server-owned WatchRunner daemon (DB-persisted, survives eviction + restart)
- Three dispatch paths: idle, busy, and evicted workstream restore
- REST API for console visibility (GET /v1/api/watches, POST cancel)
- Migration 007, 8 storage CRUD methods, 75 new tests (1383 total)

* fix: address Copilot review — condition errors, restore deadlock, docs

- Condition eval errors now deactivate the watch immediately instead
  of silently looping until max_polls
- Restored (evicted) workstreams set auto_approve=True to prevent
  approval deadlocks with no connected user
- Tool description clarifies first-poll baseline behavior for change
  detection mode
- Diagram updated: DELETE → POST /v1/api/watches/{id}/cancel
2026-03-10 08:18:28 -07:00
Patrick Buckley 7ea150fa71 Bump version to 0.5.2 2026-03-09 13:42:28 -07:00
Patrick Buckley 4d665a5f62 fix: SSE reconnect loop — remove _sse_generation single-consumer lock
The _sse_generation mechanism assumed one SSE consumer per workstream,
but the bridge also maintains an SSE connection to each workstream.
When a new client connected (browser, proxy, or test), it incremented
the generation counter, killing the bridge's connection. The bridge
reconnected, killing the new client's connection — creating a
mutual-kill cascade that closed every SSE connection after one ping
cycle (5s).

Fix: remove _sse_generation entirely. sse-starlette handles disconnect
detection via its own ASGI task. Also remove the redundant
request.is_disconnected() check which raced with sse-starlette's
disconnect listener in Starlette 0.52.

Root cause confirmed via raw socket test: the server was sending
a zero-length chunked terminator (0\r\n\r\n) at exactly 5s,
cleanly ending the HTTP response body.
2026-03-09 13:40:36 -07:00
Patrick Buckley 3bc3250869 fix: recovered workstreams invisible in console UI (#35)
* fix: recovered workstreams invisible in console UI

Bridge startup recovery (_recover_workstreams) re-registered workstream
ownership but never published WorkstreamCreatedEvent to the cluster
channel. The collector's poll loop would pick up the workstream in its
internal state, but _apply_poll never fanned out SSE events to connected
browsers. Combined, this made channel-resumed workstreams invisible in
the console while remaining accessible through the proxied node UI.

- Bridge: emit WorkstreamCreatedEvent for each recovered workstream
- Collector: diff poll results and fan out synthetic ws_created/ws_closed
  events for workstream additions and removals
- Skip workstreams with empty IDs in poll processing
- Add 4 tests for poll-diff fanout behavior
- Update console data-flow diagram and architecture docs

* fix: address PR review — filter empty ws IDs, stable event ordering

- Filter empty-string keys from old_ids to avoid phantom ws_closed
  events if a previous poll inserted a workstream under key "".
- Sort set diffs before iterating so ws_created/ws_closed fanout
  order is deterministic across poll cycles.
2026-03-09 13:39:47 -07:00
Patrick Buckley db937486cf Bump version to 0.5.1 2026-03-09 01:48:18 -07:00
Patrick Buckley 554257ac4d fix: SSE proxy Firefox reconnect — Connection: keep-alive header 2026-03-09 01:46:35 -07:00
Patrick Buckley 5f0004dc91 feat: add ClusterSnapshot for instant console UI state rebuild (#34)
* feat: add ClusterSnapshot for instant console UI state rebuild

The console web UI was SSE-driven with no initial state — reloads and
navigation caused blank/loading gaps while waiting for API re-fetches.

Server-side: GET /v1/api/cluster/snapshot returns the full cluster state
(all nodes with workstreams + overview aggregates) built under a single
lock. The SSE stream now emits this snapshot as the first event on
connect (snapshot taken before listener registration to avoid race).

Frontend: local clusterState object mirrors the snapshot, patched
incrementally by SSE events. View navigation renders from local state
with no API round-trips. Fixes popstate/pushState history corruption
on Back/Forward navigation (pre-existing bug). Stable node sorting
with node_id tie-breaker on both server and client.

SDK: snapshot() method on Python (sync + async) and TypeScript console
clients. ClusterSnapshotEvent in event registries.

* fix: address review feedback and SSE proxy reconnect bug

Copilot review fixes:
- Atomic snapshot+register: new get_snapshot_and_register() acquires
  both state and listener locks, eliminating the event gap between
  snapshot read and listener registration.
- Debounce patch renders: patchClusterState uses requestAnimationFrame
  to batch rapid SSE events into a single recompute+render cycle.
- Fix health type: dict[str, str] → dict[str, Any] on all three
  console schema models (ClusterNodeInfo, NodeDetailResponse,
  ClusterSnapshotNode) since /health payloads contain nested objects.
- TypeScript ClusterSnapshotEvent: use concrete ClusterSnapshotNode[]
  and ClusterOverviewResponse types instead of Record<string, unknown>.

SSE proxy reconnect fix:
- _proxy_sse raw_stream now emits `: proxy-ping` comments every 3s
  when no upstream data arrives, preventing the browser EventSource
  from dropping idle connections. The raw byte passthrough refactor
  (4d11078) removed the proxy's independent keepalive — this restores
  it without reverting to EventSourceResponse.
2026-03-09 01:13:35 -07:00
Patrick Buckley 6cc1b3a5bd feat: add vision/image support to read_file tool (#33)
* feat: add vision/image support to read_file tool

read_file now detects image files (PNG, JPEG, GIF, WebP, BMP, TIFF, ICO)
and returns base64-encoded content parts for vision-capable models.
Non-vision models receive a text description instead. A new
supports_vision flag on ModelCapabilities gates the feature, with
config.toml [models.*.capabilities] overrides for local models
(vLLM, llama.cpp, NIM).

* fix: address PR review feedback

- Discard _read_files on no-vision OSError path, include exception detail
- Discard _read_files on oversized image error (not a successful read)
- Validate capabilities type from config.toml (reject non-dict)
- Clarify tool description re: vision behavior and offset/limit scope
- Remove unused os import in tests, fix import sort order
- Handle list content (image tool results) in eval.py tool result loop
2026-03-08 23:43:42 -07:00
Patrick Buckley cc9afe94cd get title in collector for console 2026-03-08 22:32:52 -07:00
Patrick Buckley 136b75fdef Bump version to 0.5.0 2026-03-08 04:47:10 -07:00
Patrick Buckley 4d1107839b refactor: use raw streaming for SSE proxy to preserve event framing (#32)
* refactor: use raw streaming for SSE proxy to preserve event framing

- Replace httpx_sse aconnect_sse with raw httpx.stream for SSE proxy
- Stream bytes verbatim to preserve server-side ping comments and event framing
- Add StreamingResponse with proper headers (Cache-Control, X-Accel-Buffering)
- Update compose.yaml to add 'cluster' profile to the service

* Refactor SSE proxy to raw byte passthrough

- turnstone/console/server.py: Replace aconnect_sse + EventSourceResponse with
  httpx.stream() + StreamingResponse for raw byte passthrough. Server pings,
  events, and comments now flow through verbatim. Added per-request timeout
  override (read=None, pool=None) for long-lived SSE streams.

- tests/test_console.py: Add 3 new tests for SSE proxy:
  - Ping and event preservation
  - Upstream error status handling
  - Client disconnect handling

- docs/console.md: Update SSE Proxy section to reflect raw byte passthrough
  approach.
2026-03-08 04:46:34 -07:00
126 changed files with 20694 additions and 872 deletions
+8
View File
@@ -5,6 +5,7 @@ on:
tags: ["v*"]
permissions:
contents: write
id-token: write
jobs:
@@ -19,3 +20,10 @@ jobs:
- run: pip install build
- run: python -m build
- uses: pypa/gh-action-pypi-publish@release/v1
- name: Create GitHub Release
uses: softprops/action-gh-release@v2
with:
generate_release_notes: true
draft: false
prerelease: ${{ contains(github.ref, '-') }}
+92
View File
@@ -0,0 +1,92 @@
# Bootstrap Wizard
Interactive, AI-guided setup for Turnstone deployments. Instead of manually
editing `.env` files and reading deployment docs, the wizard walks you through
every decision conversationally and generates all the config files for you.
## Quick Start
```bash
turnstone-bootstrap
```
That's it — no flags, no arguments. The wizard prompts for everything.
## How It Works
1. **Pick a model** — Choose OpenAI, Anthropic, or a local/vLLM endpoint to
power the wizard. Local endpoints auto-detect available models.
2. **Answer questions** — The AI walks you through deployment mode, LLM
provider, database, authentication, ports, and optional features.
3. **Review generated files** — Each file is previewed before writing. You
confirm or reject every write.
4. **Start the stack** — The wizard prints the exact `docker compose` command
and a `setup.sh` script to create your first admin user, roles, and policies.
## What Gets Generated
| File | Purpose |
|------|---------|
| `.env` | All environment variables for `compose.yaml` |
| `setup.sh` | Post-start script: creates admin user, roles, tool policies, prompt templates via the API |
| `docker-compose.override.yaml` | Only if customizations beyond env vars are needed |
## Requirements
- **Python 3.11+** with turnstone installed (`pip install turnstone`)
- **An LLM API key** — for the wizard itself (OpenAI, Anthropic, or a local
model). This can differ from the LLM your deployment will use.
- **Docker & Docker Compose** — needed to run the stack. The wizard detects
whether Docker is installed and gives platform-specific install instructions
if it's missing. You can still generate config files without Docker.
## Deployment Modes
The wizard supports two deployment modes:
- **Single-node production** (`docker compose --profile production up`) —
1 server + bridge + console + PostgreSQL + Redis. Good for most use cases.
- **Multi-node cluster** (`docker compose --profile cluster up`) —
10-node server/bridge fleet + PostgreSQL + Redis. For high-throughput or
HA deployments.
## Example Session
```
$ turnstone-bootstrap
Turnstone Bootstrap Wizard v0.5.4
────────────────────────────────────────────────
Which provider for this wizard?
[1] OpenAI
[2] Anthropic
[3] OpenAI-compatible (local/vLLM)
> 3
Base URL [http://localhost:8000/v1]:
API key (press Enter for 'none'):
Querying http://localhost:8000/v1 for available models...
Found model: Qwen/Qwen3-32B
Connected to Qwen/Qwen3-32B. Handing off to AI assistant...
> (AI walks you through the rest interactively)
```
## Tips
- **Re-run safely** — running the wizard again detects your existing `.env`
and offers to update it rather than overwriting.
- **Duplicate writes are skipped** — if the LLM tries to write the same file
twice with identical content, it's silently ignored.
- **Type `quit` to exit** at any time during the conversation.
- **Ctrl+C** is handled gracefully — press once to interrupt, twice to exit.
## See Also
- [Docker Deployment](docker.md) — manual compose setup and profiles
- [Security](security.md) — auth architecture and token types
- [Governance](governance.md) — roles, policies, and templates
+21 -2
View File
@@ -17,6 +17,7 @@ Turnstone gives LLMs tools — shell, files, search, web, planning — and orche
- **Queue-driven agents** — trigger workstreams via message queue, stream progress, approve or auto-approve tool use
- **Multi-node clusters** — generic work load-balances across nodes, directed work routes to a specific server
- **Cluster dashboard** — real-time view of all nodes and workstreams, workstream creation with node targeting, reverse proxy for server UIs (only the console port needs network access)
- **Governance & compliance** — role-based access control, tool policies, usage tracking, and append-only audit logs
- **Cluster simulator** — test the stack at scale (up to 1000 nodes) without an LLM backend
<p align="center">
@@ -127,6 +128,23 @@ Detailed UML diagrams are available in [`docs/diagrams/`](docs/diagrams/):
| [Deployment](docs/diagrams/png/12-deployment.png) | Docker Compose service topology |
| [SDK Architecture](docs/diagrams/png/13-sdk-architecture.png) | Python + TypeScript client libraries |
| [Storage Architecture](docs/diagrams/png/14-storage-architecture.png) | Pluggable database backends (SQLite + PostgreSQL) |
| [Auth Architecture](docs/diagrams/png/15-auth-architecture.png) | JWT, scopes, token types, login flows |
| [Channel Architecture](docs/diagrams/png/16-channel-architecture.png) | Discord/Slack adapter protocol and routing |
| [Notify Flow](docs/diagrams/png/17-notify-flow.png) | Channel notification dispatch |
| [Watch Architecture](docs/diagrams/png/18-watch-architecture.png) | Periodic command polling daemon |
| [Governance Architecture](docs/diagrams/png/19-governance-architecture.png) | RBAC, policies, audit, usage enforcement flow |
### Governance
Turnstone includes a built-in governance layer for enterprise deployments — manage who can do what, which tools run unattended, and where every token goes.
- **RBAC** — 15 granular permissions, 3 built-in roles (admin / operator / viewer), custom roles, privilege escalation prevention
- **Tool policies** — glob-pattern rules (`allow` / `deny` / `ask`) with priority ordering; automate approvals or lock down dangerous tools
- **Prompt templates** — reusable system messages with `{{variable}}` substitution and categories
- **Usage tracking** — per-request token and tool metrics, aggregation by day / model / user, automatic 90-day pruning
- **Audit logging** — append-only event trail for all admin mutations, IP-aware, 365-day retention
All governance features are managed through the console admin panel (10 tabs) and the full REST API. See [docs/governance.md](docs/governance.md) for setup and configuration.
## Multi-node routing
@@ -151,12 +169,12 @@ Bridges BLPOP from their per-node queue (priority) then the shared queue. Direct
## Tools
15 built-in tools, 2 agent tools, plus external tools via MCP:
16 built-in tools, 2 agent tools, plus external tools via MCP:
| Tool | Description | Auto-approved |
|------|-------------|:---:|
| `bash` | Execute shell commands | |
| `read_file` | Read file contents | yes |
| `read_file` | Read file contents (text or images with vision models) | yes |
| `write_file` | Write/create files | |
| `edit_file` | Fuzzy-match file editing | |
| `search` | Search files by name/content | yes |
@@ -168,6 +186,7 @@ Bridges BLPOP from their per-node queue (priority) then the shared queue. Direct
| `recall` | Search memories and history | yes |
| `forget` | Remove a memory | yes |
| `notify` | Send notifications to linked channels | yes |
| `watch` | Periodic command polling with conditions | |
| `task` | Spawn autonomous sub-agent | |
| `plan` | Explore codebase, write .plan.md | |
| `mcp__*` | External tools from MCP servers | |
+57 -5
View File
@@ -2,10 +2,11 @@
# Turnstone Docker Compose Stack
#
# Usage:
# Default (SQLite): docker compose up
# Infra only: docker compose up
# Single node: docker compose --profile production up
# Production (PG): DB_BACKEND=postgresql docker compose --profile production up
# (or set DB_BACKEND=postgresql in .env)
# 10-node cluster: docker compose --profile cluster up
# Cluster + DDG: docker compose --profile ddgCluster up
# With simulator: docker compose --profile sim up
# =============================================================================
@@ -29,6 +30,7 @@ services:
profiles:
- production
- cluster
- ddgCluster
environment:
POSTGRES_DB: turnstone
POSTGRES_USER: ${POSTGRES_USER:-turnstone}
@@ -88,6 +90,8 @@ services:
build:
context: .
dockerfile: Dockerfile
profiles:
- production
command:
- sh
- -c
@@ -99,10 +103,12 @@ services:
--api-key "$${OPENAI_API_KEY}"
$${MODEL:+--model $$MODEL}
$${SKIP_PERMISSIONS:+--skip-permissions}
$${MCP_CONFIG:+--mcp-config $$MCP_CONFIG}
ports:
- "${SERVER_PORT:-8080}:8080"
volumes:
- turnstone-data:/data
- ./docker/mcp-ddg.json:/etc/turnstone/mcp-ddg.json:ro
environment:
- LLM_BASE_URL=${LLM_BASE_URL:-http://host.docker.internal:8000/v1}
- OPENAI_API_KEY=${OPENAI_API_KEY:-dummy}
@@ -112,6 +118,7 @@ services:
- TURNSTONE_AUTH_TOKEN=${TURNSTONE_AUTH_TOKEN:-}
- TURNSTONE_JWT_SECRET=${TURNSTONE_JWT_SECRET:-}
- MODEL=${MODEL:-}
- MCP_CONFIG=${MCP_CONFIG:-}
- TURNSTONE_DB_BACKEND=${DB_BACKEND:-sqlite}
- TURNSTONE_DB_URL=${DATABASE_URL:-}
- TURNSTONE_NODE_ID=${TURNSTONE_NODE_ID:-}
@@ -125,6 +132,9 @@ services:
postgres:
condition: service_healthy
required: false
ddg-search:
condition: service_healthy
required: false
healthcheck:
test: ["CMD", "python", "/usr/local/bin/healthcheck.py", "http://127.0.0.1:8080/health"]
interval: 10s
@@ -141,6 +151,8 @@ services:
build:
context: .
dockerfile: Dockerfile
profiles:
- production
command:
- turnstone-bridge
- --server-url=http://server:8080
@@ -207,6 +219,8 @@ services:
dockerfile: Dockerfile
profiles:
- production
- cluster
- ddgCluster
command:
- sh
- -c
@@ -234,6 +248,39 @@ services:
required: false
restart: unless-stopped
# -------------------------------------------------------------------
# ddg-search — DuckDuckGo Search MCP server (HTTP transport)
# Provides web search + content fetch tools to turnstone via MCP.
# No API key required.
#
# Start with: MCP_CONFIG=/etc/turnstone/mcp-ddg.json \
# docker compose --profile ddgCluster up
# -------------------------------------------------------------------
ddg-search:
image: python:3.13-slim
profiles:
- ddgCluster
command:
- sh
- -c
- >-
pip install --no-cache-dir duckduckgo-mcp-server &&
python -c "from mcp.server.transport_security import TransportSecuritySettings; import duckduckgo_mcp_server.server as s; s.safe_search=s.SafeSearchMode.OFF; s.mcp.settings.host='0.0.0.0'; s.mcp.settings.port=3000; s.mcp.settings.transport_security=TransportSecuritySettings(enable_dns_rebinding_protection=False); s.mcp.run(transport='streamable-http')"
networks:
- turnstone-net
healthcheck:
test: ["CMD-SHELL", "python -c \"import socket; s=socket.create_connection(('0.0.0.0',3000),2); s.close()\""]
interval: 10s
timeout: 5s
retries: 3
start_period: 30s
deploy:
resources:
limits:
memory: 256M
cpus: '0.25'
restart: unless-stopped
# -------------------------------------------------------------------
# turnstone-sim — Multi-node cluster simulator (no LLM needed)
# Start with: docker compose --profile sim up
@@ -287,7 +334,7 @@ services:
server-1: &cluster-server
build: { context: ., dockerfile: Dockerfile }
profiles: [cluster]
profiles: [cluster, ddgCluster]
command: &cluster-server-cmd
- sh
- -c
@@ -299,7 +346,10 @@ services:
--api-key "$${OPENAI_API_KEY}"
$${MODEL:+--model $$MODEL}
$${SKIP_PERMISSIONS:+--skip-permissions}
volumes: [turnstone-data:/data]
$${MCP_CONFIG:+--mcp-config $$MCP_CONFIG}
volumes:
- turnstone-data:/data
- ./docker/mcp-ddg.json:/etc/turnstone/mcp-ddg.json:ro
environment: &cluster-server-env
LLM_BASE_URL: ${LLM_BASE_URL:-http://host.docker.internal:8000/v1}
OPENAI_API_KEY: ${OPENAI_API_KEY:-dummy}
@@ -309,6 +359,7 @@ services:
TURNSTONE_AUTH_TOKEN: ${TURNSTONE_AUTH_TOKEN:-}
TURNSTONE_JWT_SECRET: ${TURNSTONE_JWT_SECRET:-}
MODEL: ${MODEL:-}
MCP_CONFIG: ${MCP_CONFIG:-}
TURNSTONE_DB_BACKEND: ${DB_BACKEND:-postgresql}
TURNSTONE_DB_URL: ${DATABASE_URL:-postgresql://${POSTGRES_USER:-turnstone}:${POSTGRES_PASSWORD:?}@postgres:5432/turnstone}
TURNSTONE_NODE_ID: node-1
@@ -317,6 +368,7 @@ services:
depends_on:
redis: { condition: service_healthy }
postgres: { condition: service_healthy }
ddg-search: { condition: service_healthy, required: false }
healthcheck:
test: ["CMD", "python", "/usr/local/bin/healthcheck.py", "http://127.0.0.1:8080/health"]
interval: 10s
@@ -360,7 +412,7 @@ services:
bridge-1: &cluster-bridge
build: { context: ., dockerfile: Dockerfile }
profiles: [cluster]
profiles: [cluster, ddgCluster]
command:
- turnstone-bridge
- --server-url=http://server-1:8080
+7
View File
@@ -0,0 +1,7 @@
{
"mcpServers": {
"ddg": {
"url": "http://ddg-search:3000/mcp"
}
}
}
+125 -6
View File
@@ -448,6 +448,14 @@ after `/clear` or `/new` commands).
{"type": "clear_ui"}
```
**`cancelled`** -- the generation was cancelled by the user (via the Stop
button or `POST /v1/api/cancel`). The client should finalize any in-progress
assistant message with whatever partial content was streamed.
```json
{"type": "cancelled"}
```
#### Keepalive
The server sends an SSE comment every 5 seconds when no events are pending:
@@ -460,13 +468,13 @@ The server sends an SSE comment every 5 seconds when no events are pending:
This prevents proxies and browsers from closing the connection due to
inactivity.
#### Generation mechanism
#### Multi-consumer fan-out
Each new SSE connection to a workstream increments an internal
`_sse_generation` counter. The previous SSE handler detects the generation
mismatch and exits its event loop, ensuring only one active SSE connection per
workstream at a time. The event queue is drained of stale events before the new
connection begins streaming.
Each SSE connection to a workstream receives its own delivery queue. Events
produced by the worker thread are fanned out to all registered listener queues,
so multiple consumers (browser, bridge, console proxy, SDK) can connect
simultaneously and each receives every event. On reconnect the client receives
a full history replay, so no catch-up mechanism is needed.
---
@@ -701,6 +709,43 @@ containing the resumed session's messages.
---
### `POST /v1/api/cancel`
Cancels the active generation in a workstream. Sets a cooperative cancellation
flag that is checked at multiple points in the generation loop (per streaming
chunk, before tool execution, inside bash commands). The session transitions to
`idle` state and preserves any partial content already streamed.
If the workstream is waiting for tool approval or plan review, the pending
prompt is automatically denied/rejected to unblock the worker thread.
Calling this endpoint when the workstream is already idle is a harmless no-op.
**Request body:**
```json
{"ws_id": "abc123"}
```
| Field | Type | Required | Description |
|--------|--------|----------|----------------------|
| `ws_id`| string | yes | Target workstream ID |
**Response:**
```json
{"status": "ok"}
```
**Error responses:**
| Status | Body | Condition |
|--------|------------------------------------|------------------------|
| 400 | `{"error": "No session"}` | Session not initialized|
| 404 | `{"error": "Unknown workstream"}` | `ws_id` not found |
---
### `POST /v1/api/workstreams/new`
Creates a new workstream. The server supports up to 10 concurrent workstreams.
@@ -719,6 +764,7 @@ All fields are optional. The body can be empty or an empty JSON object.
| `model` | string | default | Model alias from the registry (`[models.*]`) |
| `auto_approve` | bool | false | Auto-approve all tool calls for this workstream |
| `resume_ws` | string | "" | Workstream ID to resume atomically during creation (empty = fresh)|
| `template` | string | "" | Prompt template name (replaces default templates; 400 if not found)|
**Response (success):**
@@ -774,6 +820,79 @@ Status code: `400`
---
### `GET /v1/api/watches`
List active watches on this server node. Optionally filter by workstream.
Requires `write` scope.
**Query parameters:**
| Parameter | Type | Required | Description |
|-----------|--------|----------|------------------------------------|
| `ws_id` | string | no | Filter to watches for this workstream. If omitted, returns all watches on the node. |
**Response:**
```json
{
"watches": [
{
"watch_id": "abc123def456...",
"ws_id": "ws-1",
"node_id": "host_a1b2",
"name": "pr-review",
"command": "gh pr view --json state",
"interval_secs": 300.0,
"stop_on": "data[\"state\"] == \"MERGED\"",
"max_polls": 100,
"poll_count": 5,
"last_output": "{\"state\": \"OPEN\"}",
"last_poll": "2026-03-09T12:00:00",
"next_poll": "2026-03-09T12:05:00",
"active": 1,
"created": "2026-03-09T11:30:00"
}
]
}
```
---
### `POST /v1/api/watches/{watch_id}/cancel`
Cancel an active watch. Sets `active=0` and clears `next_poll`.
Requires `write` scope. Verifies node ownership in multi-node deployments.
**Path parameters:**
| Parameter | Type | Description |
|------------|--------|-----------------|
| `watch_id` | string | Watch ID to cancel |
**Response (success):**
```json
{"status": "ok", "watch_id": "abc123def456..."}
```
**Error (not found):**
```json
{"error": "Watch not found"}
```
Status code: `404`
**Error (wrong node):**
```json
{"error": "Watch belongs to another node"}
```
Status code: `403`
---
### `OPTIONS` (any path)
Handles CORS preflight requests.
+66 -8
View File
@@ -3,7 +3,7 @@
Turnstone is an AI orchestration platform with tool use, parallel workstreams, and persistent
memory. It connects to any OpenAI-compatible API (local vLLM, OpenAI, etc.) or
Anthropic's native Messages API via pluggable provider adapters, and gives the
model 14 built-in tools plus external tools via MCP (Model Context Protocol) for
model 18 built-in tools plus external tools via MCP (Model Context Protocol) for
reading, writing, searching, planning, and executing code.
The core design principle is a **UI-agnostic engine with pluggable frontends**.
@@ -44,6 +44,7 @@ turnstone/
tools.py Tool schema loader (JSON -> OpenAI function-calling format)
mcp_client.py MCPClientManager — MCP server connections, tool discovery, dynamic refresh, async-sync bridge
tool_search.py Dynamic tool search — BM25 index, session-scoped tool visibility
watch.py WatchRunner daemon — periodic command polling, condition DSL, result dispatch
model_registry.py ModelRegistry — named model configs, lazy client creation, fallback routing
memory.py Persistence facade (delegates to storage backend)
storage/ Pluggable storage: StorageBackend protocol, SQLite + PostgreSQL
@@ -128,6 +129,7 @@ A user message flows through the system as follows:
| on_reasoning_token() / on_content_token()
| accumulate tool_calls from deltas
| track finish_reason
| _check_cancelled() per chunk (cooperative cancel)
v
finish_reason check:
+--- "length" --> warn, discard partial tool_calls
@@ -173,11 +175,13 @@ Phase 2: APPROVE (serial, blocking)
_emit_state("running")
Phase 3: EXECUTE (parallel)
_check_cancelled() <-- cancellation checkpoint before execution starts
if len(items) == 1:
run_one(items[0])
else:
ThreadPoolExecutor(max_workers=4).map(run_one, items)
Bash tool streams stdout line-by-line via ui.on_tool_output_chunk(call_id, line)
(cancel_event also checked per line — kills process group on cancel)
Final output (stdout + stderr) delivered via ui.on_tool_result(call_id, name, output)
call_id links tool_info items → streaming chunks → final result
For plan tool: post-execution gate via ui.on_plan_review()
@@ -208,6 +212,11 @@ The engine emits state changes via `_emit_state()` which calls
"idle" ---> no more tool calls, turn complete
|
(or "error" ---> exception or KeyboardInterrupt)
cancel() may be called from any state. It sets a cooperative flag
checked at each streaming chunk, before tool execution, and inside
bash commands. The session transitions to "idle" with partial
content preserved, emitting on_info("[Generation cancelled]").
```
---
@@ -560,21 +569,23 @@ LLMProvider (protocol)
|------|--------|
| `StreamChunk` | `content_delta`, `reasoning_delta`, `tool_call_deltas`, `info_delta`, `usage`, `finish_reason` |
| `CompletionResult` | `content`, `tool_calls`, `finish_reason`, `usage` |
| `ModelCapabilities` | `context_window`, `max_output_tokens`, `supports_temperature`, `token_param`, `thinking_mode`, `supports_effort`, `supports_web_search` |
| `ModelCapabilities` | `context_window`, `max_output_tokens`, `supports_temperature`, `token_param`, `thinking_mode`, `supports_effort`, `supports_web_search`, `supports_tool_search`, `supports_vision` |
| `UsageInfo` | `prompt_tokens`, `completion_tokens`, `total_tokens` |
**OpenAIProvider** (`_openai.py`): passes messages through unchanged (they are
already in OpenAI format). Model capability lookup table covers
GPT-5/5.1/5.2, O-series, and search models (`gpt-5-search-api`).
already in OpenAI format), including multi-part content blocks (text + images)
in tool results. Model capability lookup table covers GPT-5/5.1/5.2/5.3/5.4,
O-series, and search models (`gpt-5-search-api`) — all with `supports_vision`.
For search models, injects `web_search_options` and removes the `web_search`
function tool (the model always searches). Citations from `url_citation`
annotations are formatted as footnotes. Unknown models (local servers) get
permissive defaults and use Tavily for web search.
permissive defaults with `supports_vision=False` and use Tavily for web search.
**AnthropicProvider** (`_anthropic.py`): converts OpenAI-format messages to
Anthropic content blocks, maps `system`/`developer` roles to the `system`
parameter, groups consecutive `tool` result messages into user-role content
blocks, and translates tool schemas from OpenAI function-calling format to
blocks (converting `image_url` parts to Anthropic's `image` source format),
and translates tool schemas from OpenAI function-calling format to
Anthropic's `input_schema` format. Supports both manual and adaptive thinking
modes, with effort parameter support for models like Claude Opus 4.6 and
Sonnet 4.6. Replaces the `web_search` function tool with Anthropic's native
@@ -620,6 +631,18 @@ agent_model = "claude"
Each `[models.*]` entry produces a `ModelConfig` with a `provider` field
(default: `"openai"`). Supported values: `"openai"` and `"anthropic"`.
An optional `[models.*.capabilities]` sub-table overrides per-model
`ModelCapabilities` flags (useful for local models whose capabilities
cannot be detected programmatically):
```toml
[models.qwen-vl]
base_url = "http://localhost:8000/v1"
model = "qwen-3.5-vl"
[models.qwen-vl.capabilities]
supports_vision = true
```
**Lifecycle:**
1. `load_model_registry()` reads `[models.*]` sections from config.toml and
@@ -1098,7 +1121,7 @@ context manager handles startup/shutdown (health monitor, MCP client,
registry).
Each workstream's `WebUI` has:
- `_event_queue` (per-workstream SSE events, `queue.Queue`)
- `_listeners` (per-client SSE queues, fan-out on `_enqueue()`)
- `_approval_event` / `_plan_event` (`threading.Event` for blocking)
- `_global_queue` (class variable, shared, for state broadcasts)
@@ -1158,6 +1181,10 @@ bridge auto-approves via `POST /v1/api/approve`. Otherwise, it publishes an
`BLPOP` of a Redis response queue (`turnstone:resp:{request_id}`) until the client pushes
a response or the approval timeout (default 3600s / 1 hour) expires.
**Cancellation:** The `CancelMessage` (type `"cancel"`) is a routed inbound message.
The bridge dispatches it to `POST /v1/api/cancel` on the server owning the workstream,
which sets the cooperative cancel flag and unblocks any pending approval/plan waits.
**Completion detection:** The bridge tracks which `correlation_id` maps to which
`ws_id` for active sends. When the global SSE reports `ws_state → idle` for a tracked
workstream, the bridge emits a synthetic `TurnCompleteEvent` with the correlation ID.
@@ -1172,6 +1199,9 @@ for existing workstreams are auto-routed via `turnstone:ws:{ws_id}` ownership ke
If a bridge picks up a shared-queue message for a workstream owned by another node, it
re-routes to that node's queue (1 extra hop). Bridges publish heartbeats to
`turnstone:node:{node_id}` with configurable TTL for node discovery.
On startup, `_recover_workstreams` re-registers ownership of existing
workstreams and publishes `WorkstreamCreatedEvent` to the cluster channel
so the console collector picks them up immediately.
### Cluster Console
@@ -1197,7 +1227,10 @@ The console HTTP layer is a Starlette/ASGI app served by uvicorn. The SSE
endpoint uses `EventSourceResponse` with the same listener queue pattern as
the main server. `ClusterCollector`'s background threads (event subscriber,
node discovery, poll loop) use sync Redis clients and `ThreadPoolExecutor`
for parallel HTTP polling.
for parallel HTTP polling. The poll loop diffs workstream IDs between poll
cycles and fans out synthetic `ws_created`/`ws_closed` SSE events for any
changes, ensuring browser clients stay in sync even when real-time cluster
events are missed (e.g. bridge startup recovery).
The console has two write-path capabilities:
@@ -1319,3 +1352,28 @@ gateway validates the JWT, resolves the target (username lookup via
the appropriate `ChannelAdapter.send()`. Delivery retries up to 3 times
with backoff, re-querying the service registry on each attempt. See
[Notification Flow diagram](diagrams/png/17-notify-flow.png).
---
## Governance
> See also: [Governance documentation](governance.md) | [Governance Architecture diagram](diagrams/19-governance-architecture.puml)
Turnstone governance extends the Phase 1 auth system with role-based access
control (RBAC), tool execution policies, prompt templates, usage tracking,
and audit logging. The permission model has two layers: legacy scopes
(`read`, `write`, `approve`) checked by `AuthMiddleware`, and 15 granular
permissions checked per-endpoint by `require_permission()`. Three built-in
roles (admin, operator, viewer) are seeded by migration 008; custom roles
can be created with any permission subset. JWTs carry both `scopes` and
`permissions` claims for backward compatibility.
Tool policies use glob pattern matching (`fnmatch`) with priority-ordered
first-match-wins evaluation to control tool execution (allow/deny/ask).
Prompt templates provide reusable system messages with `{{variable}}`
substitution. Usage events are recorded per-LLM-request for token
accounting. An append-only audit log captures all admin mutations.
The console admin panel adds 5 governance tabs (Roles, Policies, Templates,
Usage, Audit) for a total of 10 tabs, all permission-gated. Both Python
and TypeScript SDKs expose governance methods on the console client.
+38 -2
View File
@@ -61,6 +61,8 @@ The collector (`turnstone/console/collector.py`) maintains an in-memory snapshot
3. **Poll loop** — fetches `GET /v1/api/dashboard` and `GET /health` from each known node every 10 seconds. Uses `ThreadPoolExecutor(max_workers=50)` for parallelism. Each poll replaces the node's workstream list with the authoritative server data.
A `get_snapshot()` method builds the full cluster state under a single lock acquisition — overview aggregates and per-node workstream lists in one atomic read. This is served both as a REST endpoint and as the initial SSE event on client connect.
### Thread Safety
All reads and writes to the node/workstream map are protected by a single `threading.Lock`. Query methods acquire the lock, copy data, and release before returning.
@@ -146,6 +148,38 @@ Single node detail with all its workstreams.
}
```
### `GET /v1/api/cluster/snapshot`
Full cluster state in a single response — all nodes with their workstreams plus overview aggregates. Built under a single lock for internal consistency. Used by the browser on initial load and SSE reconnect.
```json
{
"nodes": [
{
"node_id": "db-west-04",
"server_url": "http://10.0.3.4:8080",
"max_ws": 10,
"reachable": true,
"version": "0.3.0",
"health": {"status": "ok", "version": "0.3.0"},
"aggregate": {"total_tokens": 48200, "total_tool_calls": 156},
"workstreams": [
{"id": "a1b2c3d4", "name": "perf-db-west", "state": "running", ...}
]
}
],
"overview": {
"nodes": 847,
"workstreams": 4219,
"states": {"running": 1847, "thinking": 312, "attention": 89, "idle": 1940, "error": 31},
"aggregate": {"total_tokens": 12400000, "total_tool_calls": 34200},
"version_drift": false,
"versions": ["0.3.0"]
},
"timestamp": 1709294400.0
}
```
### `POST /v1/api/cluster/workstreams/new`
Create a new workstream on a target node. Dispatches a `CreateWorkstreamMessage` through the Redis MQ pipeline — the bridge on the target node picks it up and creates the workstream on the server. Requires `write` scope.
@@ -182,7 +216,7 @@ Creation is asynchronous — the response confirms the MQ message was dispatched
### `GET /v1/api/cluster/events`
Server-Sent Events stream for real-time cluster updates.
Server-Sent Events stream for real-time cluster updates. The first event is always a `snapshot` containing the full cluster state (same shape as `GET /v1/api/cluster/snapshot` with an added `type: "snapshot"` field), followed by incremental events:
```
data: {"type":"cluster_state","ws_id":"a1b2","node_id":"db-west-04","state":"running"}
@@ -326,7 +360,7 @@ The server UI uses root-relative URLs (`/v1/api/send`, `/static/app.js`, `/share
### SSE Proxy
SSE streams (`/v1/api/events`, `/v1/api/events/global`) are proxied by creating a per-connection `httpx.AsyncClient(timeout=None)`, streaming the upstream response via `aiter_text()`, parsing SSE framing (`\n\n` delimiters), and re-emitting events through `EventSourceResponse`. Each proxied SSE stream requires its own httpx client since the shared client's 30-second timeout would kill long-lived connections.
SSE streams (`/v1/api/events`, `/v1/api/events/global`) are proxied as raw byte passthrough — the console opens an `httpx.AsyncClient.stream()` to the upstream server (with `read=None` and `pool=None` timeouts since SSE connections are long-lived) and relays every byte via `StreamingResponse`. This preserves server-side ping comments, event framing, and keepalives verbatim without parsing or re-encoding.
### Authentication
@@ -368,6 +402,8 @@ On submit, `POST /v1/api/cluster/workstreams/new` dispatches the creation reques
All five views receive live updates via SSE — state cards update counts, node rows update metrics, workstream rows update state indicators.
The browser maintains a local `clusterState` object that mirrors the cluster snapshot. It is initialized from the SSE `snapshot` event on connect (or via `GET /v1/api/cluster/snapshot` on initial page load) and updated incrementally by SSE events. View navigation reads from local state — no API round-trips needed after the initial snapshot.
### 5. Admin Panel
Accessed via the "admin" button in the header (visible when authenticated
+1 -1
View File
@@ -96,7 +96,7 @@ package "turnstone/sdk/" <<Rectangle>> {
' Tool schemas
package "turnstone/tools/" <<Rectangle>> {
component [*.json\n15 tool schemas] as schemas <<artifact>>
component [*.json\n18 tool schemas] as schemas <<artifact>>
}
' Entry point dependencies
+12 -1
View File
@@ -41,7 +41,7 @@ class "WorkstreamTerminalUI" as WsTermUI {
}
class "WebUI" as WebUI {
- _event_queue: Queue
- _listeners: list[Queue]
- _approval_event: Event
- _plan_event: Event
- _ws_prompt_tokens: int
@@ -109,6 +109,7 @@ class "ModelCapabilities" as ModelCaps <<frozen>> {
+ supports_effort: bool
+ supports_web_search: bool
+ supports_tool_search: bool
+ supports_vision: bool
}
' ChatSession
@@ -210,15 +211,23 @@ enum "WorkstreamState" as WsState {
class "MCPClientManager" as MCPMgr {
- _sessions: dict[str, ClientSession]
- _per_server_tools: dict[str, list[dict]]
- _per_server_resources: dict[str, list[dict]]
- _per_server_prompts: dict[str, list[dict]]
- _tools: list[dict]
- _tool_map: dict[str, tuple]
- _resource_map: dict[str, tuple]
- _prompt_map: dict[str, tuple]
- _supports_list_changed: dict[str, bool]
- _listeners: list[Callable]
--
+ start()
+ get_tools() → list[dict]
+ get_resources() → list[dict]
+ get_prompts() → list[dict]
+ is_mcp_tool(name) → bool
+ call_tool_sync(name, args) → str
+ read_resource_sync(uri) → str
+ get_prompt_sync(name, args?) → list[dict]
+ refresh_sync(server?) → dict
+ add_listener(callback)
+ remove_listener(callback)
@@ -229,6 +238,8 @@ class "MCPClientManager" as MCPMgr {
bridges async MCP SDK to
sync ChatSession dispatch.
Push + periodic + manual refresh.
Resources + prompts discovered
alongside tools at startup.
--
core/mcp_client.py
}
+15 -1
View File
@@ -57,6 +57,14 @@ group loop [while tool_calls present]
end
end
note right of CS
**Cancellation checkpoint:**
_check_cancelled() runs per chunk.
If cancel_event is set, raises
GenerationCancelled — preserves
partial content, emits idle state.
end note
LLM --> CS : stream complete (usage stats)
deactivate LLM
@@ -112,7 +120,7 @@ group loop [while tool_calls present]
note right of TP
Parallel execution:
bash → Popen + line-by-line streaming
read_file → open().read()
read_file → open().read() or base64 image
search → grep subprocess
edit_file → string replace
task/plan → _run_agent() sub-loop
@@ -144,6 +152,12 @@ group loop [while tool_calls present]
end
note right of CS : Loop back for next LLM call
else GenerationCancelled
CS -> CS : Preserve partial content\nor roll back incomplete tools
CS -> UI : on_info("[Generation cancelled]")
CS -> UI : on_state_change("idle")
CS --> User : return (no re-raise)
end
end
+30 -24
View File
@@ -24,29 +24,31 @@ partition "Phase 1: Prepare" #E8F5E9 {
:Dispatch to _prepare_{func_name}();
note right
**Dispatch table (16 tools):**
┌──────────────┬──────────────────┐
│ Tool │ Needs Approval? │
├──────────────┼──────────────────┤
│ bash │ ✓ Yes │
│ read_file │ ✗ Auto-approve │
│ write_file │ ✓ Yes │
│ edit_file │ ✓ Yes │
│ search │ ✗ Auto-approve │
│ math │ ✓ Yes
│ man │ ✗ Auto-approve │
│ web_fetch │ ✓ Yes
│ web_search │ ✓ Yes
│ tool_search │ ✗ Auto-approve │
│ task │ ✓ Yes │
│ plan │ ✓ Yes │
│ remember │ ✗ Auto-approve │
│ recall │ ✗ Auto-approve │
│ forget │ ✗ Auto-approve │
│ notify │ ✗ Auto-approve │
├──────────────┼──────────────────┤
mcp__* │ ✓ Yes (external)
────────────────────────────────
**Dispatch table (18 tools):**
┌──────────────┬──────────────────┐
│ Tool │ Needs Approval? │
├──────────────┼──────────────────┤
│ bash │ ✓ Yes │
│ read_file │ ✗ Auto-approve │
│ write_file │ ✓ Yes │
│ edit_file │ ✓ Yes │
│ search │ ✗ Auto-approve │
│ math ✗ Auto-approve
│ man │ ✗ Auto-approve │
│ web_fetch ✗ Auto-approve
│ web_search ✗ Auto-approve
│ tool_search │ ✗ Auto-approve │
│ task │ ✓ Yes │
│ plan │ ✓ Yes │
│ remember │ ✗ Auto-approve │
│ recall │ ✗ Auto-approve │
│ forget │ ✗ Auto-approve │
│ notify │ ✗ Auto-approve │
│ read_resource │ ✓ Yes │
use_prompt │ ✓ Yes
├─────────────────────────────────
│ mcp__* │ ✓ Yes (external) │
└───────────────┴──────────────────┘
end note
:Build item dict:
@@ -88,6 +90,8 @@ partition "Phase 2: Approve" #FFF3E0 {
}
partition "Phase 3: Execute" #E3F2FD {
:_check_cancelled();
note right: Cancellation checkpoint:\nraises GenerationCancelled if\ncancel event is set
if (single tool call?) then (yes)
:Execute sequentially:\nrun_one(items[0]);
else (multiple)
@@ -100,7 +104,7 @@ partition "Phase 3: Execute" #E3F2FD {
if item.denied → return denial message
else → item["execute"](item)
├─ _exec_bash: subprocess.run(["bash", script.sh])
├─ _exec_read_file: open().readlines()
├─ _exec_read_file: open().readlines() or _exec_read_image (base64)
├─ _exec_write_file: makedirs + write
├─ _exec_edit_file: find_occurrences + replace
├─ _exec_search: grep subprocess
@@ -115,6 +119,8 @@ partition "Phase 3: Execute" #E3F2FD {
├─ _exec_remember: SQLite INSERT OR REPLACE
├─ _exec_recall: SQLite FTS5/LIKE search
├─ _exec_forget: SQLite DELETE
├─ _exec_read_resource: MCPClientManager.read_resource_sync()
├─ _exec_use_prompt: MCPClientManager.get_prompt_sync()
└─ _exec_mcp_tool: MCPClientManager.call_tool_sync()
end note
+7
View File
@@ -79,6 +79,12 @@ package "Inbound Messages (Client → Bridge)" #FFF3E0 {
type = "list_nodes"
}
class CancelMessage {
type = "cancel"
--
+ ws_id: str
}
IM <|-- SendMessage
IM <|-- ApproveMessage
IM <|-- PlanFeedbackMessage
@@ -88,6 +94,7 @@ package "Inbound Messages (Client → Bridge)" #FFF3E0 {
IM <|-- ListWorkstreamsMessage
IM <|-- HealthMessage
IM <|-- ListNodesMessage
IM <|-- CancelMessage
}
package "Outbound Events (Bridge → Client)" #E3F2FD {
+6
View File
@@ -40,6 +40,12 @@ running --> error : Exception during\ntool execution
error --> thinking : New send() call\n_emit_state("thinking")
thinking --> idle : cancel() called\n_emit_state("idle")
running --> idle : cancel() called\n_emit_state("idle")
attention --> idle : cancel() unblocks\napproval/plan wait\n_emit_state("idle")
note right of thinking
**Emitted via:**
session._emit_state(state)
+25 -1
View File
@@ -78,7 +78,19 @@ activate NodeA
NodeA --> CC : {status:"ok", version:"0.3.0",\nmodel:"...", workstreams:{...}}
deactivate NodeA
CC -> CC : Diff old vs new workstream IDs
CC -> CC : Replace NodeSnapshot["nodeA"]\n.workstreams, .health, .aggregate
CC -> CC : _fanout(ws_created) for\nnewly appeared workstreams
CC -> CC : _fanout(ws_closed) for\nremoved workstreams
note right of CC
Poll-diff fanout ensures
browser SSE clients learn
about workstreams that
appeared without a real-time
cluster event (e.g. bridge
startup recovery).
end note
CC -x NodeB : (SKIPPED: sim:// URL)
@@ -89,10 +101,15 @@ deactivate CC
Browser -> Server : GET /v1/api/cluster/events
activate Server
Server -> CC : get_snapshot()
CC --> Server : ClusterSnapshot\n(full current state)
Server -> CC : register_listener(queue)
note right : Per-client queue.Queue(maxsize=500)\nSSE via EventSourceResponse + run_in_executor()
loop continuous
Server -> Browser : data: {"type":"snapshot",...}\n(full state as first SSE event)
loop continuous (incremental updates)
CC -> Server : event via listener queue\n(from any of the 3 threads)
Server -> Browser : data: {"type":"cluster_state",...}\n\n
end
@@ -105,6 +122,13 @@ Browser -> Server : connection closed
Server -> CC : unregister_listener(queue)
deactivate Server
== Browser REST: Snapshot ==
Browser -> Server : GET /v1/api/cluster/snapshot
Server -> CC : get_snapshot()
CC --> Server : ClusterSnapshot\n(full current state)
Server --> Browser : JSON response
== Browser REST Requests ==
Browser -> Server : GET /v1/api/cluster/overview
+3
View File
@@ -32,6 +32,7 @@ package "turnstone/sdk/ (Python)" {
+ approve()
+ plan_feedback()
+ command()
+ cancel(ws_id)
+ stream_events(ws_id)
+ stream_global_events()
+ send_and_wait()
@@ -45,6 +46,7 @@ package "turnstone/sdk/ (Python)" {
+ nodes()
+ workstreams()
+ node_detail()
+ snapshot()
+ create_workstream()
+ stream_cluster_events()
+ login() / logout()
@@ -129,6 +131,7 @@ package "sdk/typescript/ (TypeScript)" {
class "TurnstoneConsole" as TSConsole <<ts>> {
+ overview()
+ nodes()
+ snapshot()
+ clusterEvents()
...
}
+166
View File
@@ -0,0 +1,166 @@
@startuml
!theme plain
title Turnstone — Watch Tool Architecture
skinparam participant {
BackgroundColor<<server>> #FFE0B2
BackgroundColor<<storage>> #B3E5FC
BackgroundColor<<session>> #C8E6C9
BackgroundColor<<ui>> #E8EAF6
}
participant "ChatSession\n(session.py)" as Session <<session>>
participant "WatchRunner\n(watch.py)" as Runner <<server>>
participant "StorageBackend\n(SQLite)" as Storage <<storage>>
participant "WebUI / SSE\n(server.py)" as UI <<ui>>
== Create Phase ==
Session -> Session : _prepare_watch(action="create")
note right
Validates:
- command via is_command_blocked()
- poll_every → parse_duration()
- stop_on → validate_condition()
- max watches limit (5)
- duplicate name check
needs_approval = True
end note
Session -> Storage : create_watch(watch_id, ws_id,\nnode_id, command, interval,\nstop_on, max_polls, next_poll)
Session --> UI : tool_result:\n"Watch 'pr-review' created"
== Poll Phase (WatchRunner daemon, every 15s) ==
Runner -> Storage : list_due_watches(now)
Storage --> Runner : due_watches[]
note right
Filters:
active=1 AND
next_poll <= now AND
node_id matches
end note
loop for each due watch
Runner -> Runner : is_command_blocked()?
alt blocked
Runner -> Storage : update_watch(active=False)
else safe
Runner -> Runner : subprocess.run(command)
note right
timeout = tool_timeout
start_new_session = True
output truncated at 64KB
end note
Runner -> Runner : evaluate_condition(\nstop_on, output,\nexit_code, prev_output)
note right
**Variables:**
output, data, exit_code,
prev_output, changed
**Safe builtins only:**
len, str, int, sorted, ...
No import/open/exec/eval
**stop_on=None:**
fires on change (skip 1st poll)
end note
alt condition fired OR max_polls reached
Runner -> Storage : update_watch(\npoll_count++,\nlast_output, active=False)
Runner -> Runner : format_watch_message()
Runner -> Runner : _dispatch_result(ws_id, msg)
else not fired
Runner -> Storage : update_watch(\npoll_count++,\nlast_output, next_poll)
end
end
end
== Dispatch Phase ==
note over Runner, Session
**Three dispatch paths:**
end note
alt Path A: workstream active + idle
Runner -> Session : dispatch_fn(message)\n→ _watch_pending.put()
Session -> Session : _dispatch_pending_watch()\n→ self.send(message)
Session -> UI : SSE: thinking, content,\ntool calls...
note right
Watch result appears as
synthetic user message.
Model sees it and responds.
Depth guard: max 5 chains.
end note
else Path B: workstream active + busy
Runner -> Session : dispatch_fn(message)\n→ _watch_pending.put()
note right
Queued. Dispatched when
current send() reaches IDLE.
end note
else Path C: workstream evicted
Runner -> Runner : restore_fn(ws_id)
note right
1. mgr.create() — may evict
another idle workstream
2. session.resume(ws_id)
3. set_watch_runner()
4. register new dispatch_fn
end note
Runner -> Session : restored dispatch_fn(message)
end
== Cancel / List ==
Session -> Storage : list_watches_for_ws(ws_id)
note right : action="list" (auto-approve)
Session -> Storage : update_watch(active=False)
note right : action="cancel" (auto-approve)
== Server Lifecycle ==
note over Runner, Storage
**Startup:**
1. WatchRunner created in main() with storage + node_id
2. restore_fn closure captures WorkstreamManager
3. Initial workstream: session.set_watch_runner(runner)
4. _lifespan(): runner.start() — daemon thread begins
**New workstream:**
session.set_watch_runner(runner) in create_workstream()
→ registers dispatch_fn for ws_id
**Eviction / close:**
session.close() → runner.remove_dispatch_fn(ws_id)
Watches remain active in DB — WatchRunner uses restore_fn
**Restart recovery:**
Overdue watches fire ONE immediate poll
next_poll updated to now + interval
Normal cadence resumes
**Shutdown:**
_lifespan(): runner.stop() — joins thread
end note
== REST API ==
note over UI, Storage
**GET /v1/api/watches[?ws_id=X]**
List active watches (for node or workstream)
**POST /v1/api/watches/{watch_id}/cancel**
Cancel a watch (sets active=False)
Both require write scope
end note
@enduml
@@ -0,0 +1,81 @@
@startuml
!theme plain
skinparam backgroundColor #FFFFFF
skinparam defaultFontName "IBM Plex Mono"
skinparam componentStyle rectangle
title Turnstone Governance Architecture
package "Auth Flow" {
[Login/Token Auth] as auth
[_load_user_permissions()] as perms
[_permissions_to_scopes()] as scopes
[create_jwt()] as jwt
}
package "Middleware" {
[AuthMiddleware\n(scope check)] as mw
[require_permission()\n(granular check)] as rp
}
package "Governance Storage" {
database "roles" as roles_db
database "user_roles" as ur_db
database "orgs" as orgs_db
database "tool_policies" as tp_db
database "prompt_templates" as pt_db
database "usage_events" as ue_db
database "audit_events" as ae_db
}
package "Runtime Enforcement" {
[evaluate_tool_policies_batch()] as eval
[WebUI.approve_tools()] as approve
[record_usage_event()] as usage
[record_audit()] as audit
}
package "Template Runtime" {
[_load_templates()] as tload
[_render_template()\n{{model}}, {{ws_id}}, {{node_id}}] as trender
[_init_system_messages()] as tsys
[set_template() / /template] as tset
}
package "Console UI" {
[Admin Panel\n10 tabs] as ui
[governance.js] as govjs
[sessionStorage\npermissions] as ss
}
auth --> perms : user_id
perms --> roles_db : JOIN user_roles + roles
perms --> scopes : permission set
scopes --> jwt : scopes + permissions
jwt --> mw : JWT in cookie/header
mw --> rp : scope OK → check permission
rp --> ui : 403 or allow
eval --> tp_db : list_tool_policies()
approve --> eval : tool names
approve --> ae_db : (via audit)
usage --> ue_db : on_status()
audit --> ae_db : admin handlers
govjs --> roles_db : /v1/api/admin/roles
govjs --> tp_db : /v1/api/admin/policies
govjs --> pt_db : /v1/api/admin/templates
govjs --> ue_db : /v1/api/admin/usage
govjs --> ae_db : /v1/api/admin/audit
tload --> pt_db : list_default_templates()\nor get_by_name()
tload --> trender : template content
trender --> tsys : rendered content
tset --> tload : name or None
auth -[hidden]-> mw
mw -[hidden]-> approve
@enduml
+158
View File
@@ -0,0 +1,158 @@
@startuml
!theme plain
title Turnstone — MCP Architecture (Resources, Prompts, Tools)
skinparam participant {
BackgroundColor<<mcp>> #E1BEE7
BackgroundColor<<session>> #C8E6C9
BackgroundColor<<storage>> #B3E5FC
BackgroundColor<<server>> #FFE0B2
BackgroundColor<<ui>> #E8EAF6
}
participant "MCP Server\n(external)" as MCPSrv <<mcp>>
participant "MCPClientManager\n(mcp_client.py)" as MCPMgr <<mcp>>
participant "ChatSession\n(session.py)" as Session <<session>>
participant "StorageBackend\n(governance)" as Storage <<storage>>
participant "Server / Console\n(health + UI)" as UI <<server>>
== Startup: Connection & Discovery ==
MCPMgr -> MCPSrv : initialize (stdio or HTTP)
MCPSrv --> MCPMgr : capabilities\n(tools, resources, prompts)
MCPMgr -> MCPSrv : tools/list
MCPSrv --> MCPMgr : Tool[]
opt resources capability
MCPMgr -> MCPSrv : resources/list
MCPSrv --> MCPMgr : Resource[]
MCPMgr -> MCPSrv : resources/templates/list
MCPSrv --> MCPMgr : ResourceTemplate[]
end
opt prompts capability
MCPMgr -> MCPSrv : prompts/list
MCPSrv --> MCPMgr : Prompt[]
end
note over MCPMgr
Per-server storage:
_per_server_tools, _per_server_resources, _per_server_prompts
Copy-on-write rebuild into _tools, _resources, _prompts
Prefix: mcp__{server}__{name}
end note
MCPMgr -> Session : notify tool listeners
MCPMgr -> Session : notify resource listeners
== Governance Sync (on connect & refresh) ==
MCPMgr -> Storage : sync_prompts_to_storage()
note right
For each MCP prompt:
- Manual template exists? → skip
- MCP template exists? → update
(reset is_default=False)
- New? → create (origin="mcp",
readonly=True, is_default=False)
Removed prompts → delete
Protected by _sync_lock
end note
== set_storage() from entry point ==
UI -> MCPMgr : set_storage(backend)
note right
If servers already connected,
triggers immediate sync
end note
== Runtime: Tool Execution ==
Session -> Session : _prepare_mcp_tool(func_name, args)
note right
approval_label = func_name
(e.g. mcp__github__search)
needs_approval = True
end note
Session -> MCPMgr : call_tool_sync(name, args)
MCPMgr -> MCPSrv : tools/call
MCPSrv --> MCPMgr : ToolResult
MCPMgr --> Session : output (text)
== Runtime: Resource Read ==
Session -> Session : _prepare_read_resource(uri)
note right
approval_label = mcp_resource__{normalized_uri}
URI normalized (.. resolved)
needs_approval = True
end note
Session -> MCPMgr : read_resource_sync(uri)
MCPMgr -> MCPSrv : resources/read
MCPSrv --> MCPMgr : ReadResourceResult
MCPMgr --> Session : content (text/blob)
== Runtime: Prompt Invocation ==
Session -> Session : _prepare_use_prompt(name, arguments)
note right
approval_label = mcp__srv__prompt
Validated via is_mcp_prompt()
needs_approval = True
end note
Session -> MCPMgr : get_prompt_sync(name, args)
MCPMgr -> MCPSrv : prompts/get
MCPSrv --> MCPMgr : GetPromptResult
MCPMgr --> Session : messages [{role, content}]
== Three-Tier Refresh ==
group Push Notifications
MCPSrv -> MCPMgr : ToolListChangedNotification
MCPMgr -> MCPMgr : _refresh_server_tools()
MCPSrv -> MCPMgr : ResourceListChangedNotification
MCPMgr -> MCPMgr : _refresh_server_resources()
MCPSrv -> MCPMgr : PromptListChangedNotification
MCPMgr -> MCPMgr : _refresh_server_prompts()
MCPMgr -> Storage : sync_prompts_to_storage()
end
group Periodic Polling (default 4h)
MCPMgr -> MCPMgr : _periodic_refresh()
note right
Only polls capabilities
without push support.
Staggered per-server.
end note
end
group Manual Refresh
Session -> MCPMgr : refresh_sync()
note right: /mcp refresh [server]
end
== Policy Evaluation ==
note over Session
Tool policies use fnmatch on approval_label:
- mcp__github__* → allow (all GitHub tools/prompts)
- mcp_resource__file:///docs/* → allow
- mcp_resource__* → deny (block all resource reads)
- mcp__untrusted__* → ask
end note
== UI Visibility ==
UI -> MCPMgr : server_count, get_resources(), get_prompts()
note over UI
/health → mcp.servers, mcp.resources, mcp.prompts
Server UI: magenta status badge
Console: cluster status bar + node detail
System message: <mcp-resources> + <mcp-prompts> catalogs
end note
@enduml
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0ee0a9391bd19d92e9271bf6bd531e9c2e18baf8c5a11ead49b3c10db4d8939b
size 329625
oid sha256:c9daca81971ba7a8ed6736d23d5373c69435158fa6240b9880d14fc4759ab580
size 329673
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4512be48a51f7cd1136ea8e3344489a8d225c611cb1b961643b869351de76812
size 549668
oid sha256:01fbb3338df6426cefc2811541a865f268673b4febf32f524c264d120bc068fa
size 589546
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:24bdc6a83259e4db6aaa24f581bed59d351c7a83e1c282aac123c52b32d9f80d
size 288250
oid sha256:da9d32000e3d92d92ce621661ced60f276f9b5be652f5ed6123b400505415f4a
size 319702
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:027ad99469d69f1d6b2e73ee802b50a75617cd286392375aa69133c13d3683dc
size 256347
oid sha256:6dd3c923d1e1c49b5f91d8d342fb4b0d49a46d432460379ad146a9e3b075a05a
size 277234
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32a0665cceffcc0517265bde12cfb227688aa8585284b5e946ab23bcc52daee6
size 187650
oid sha256:d17f3feacf7bc9f64dfea19464143bc9b6ef0da5d55e6d57c0bc5a73d5724eba
size 184466
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e0a3f48cca1b8408862dc4ba04fd340703346f44d84048c99e9900f48e9c7e22
size 158867
oid sha256:7896c6e041b6dbb89d034468fa980c8fe645df5eb969d45ef966ccc6399edac2
size 200083
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:435a58aa09d0e6615e78c0be62e5fd9aa6d7329b1e96619744355c42ade649c9
size 196502
oid sha256:e7c3e40c10425d721f833390ae3531c09af501157fd3142531ba4eba86ff719d
size 197112
@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:96176a09e65e90dadc32d5e9ed778423842be89204d2cf382225f53a90cfaf01
size 258547
@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a889d4bb84c4afa3c822c3acb7021395a9463462f7aeae5382e583b783412814
size 144960
@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e6dc5142c7908314ce01229b3c4f13bf9450adcbb62a178838bd4cf81d9f4da
size 250417
+184
View File
@@ -0,0 +1,184 @@
# Governance
Turnstone governance provides role-based access control (RBAC), tool execution
policies, prompt templates, usage tracking, and audit logging for the admin
console.
## Architecture
See [diagram: 19-governance-architecture.puml](diagrams/19-governance-architecture.puml).
### RBAC (Roles & Permissions)
The permission model has two layers:
1. **Scopes** (legacy) — `read`, `write`, `approve`. Checked by `AuthMiddleware`
on every request based on URL path classification.
2. **Permissions** (granular) — 15 permission strings checked per-endpoint by
`require_permission()`.
**Built-in roles** (seeded by migration 008):
| Role | Permissions |
|------|-------------|
| admin | read, write, approve, admin.users, admin.roles, admin.orgs, admin.policies, admin.templates, admin.audit, admin.usage, admin.schedules, admin.watches, tools.approve, workstreams.create, workstreams.close |
| operator | read, write, workstreams.create, workstreams.close |
| viewer | read |
Custom roles can be created with any subset of the 15 valid permissions.
**Auth flow:**
1. User logs in (password or API token) → `_load_user_permissions()` aggregates
permissions from all assigned roles
2. `_permissions_to_scopes()` derives legacy scopes (any `admin.*``approve`)
3. JWT created with both `scopes` and `permissions` claims
4. Middleware checks scope → handler checks permission via `require_permission()`
### Tool Policies
Admin-defined rules that control tool execution:
- **Pattern matching**: Glob syntax via `fnmatch` (e.g., `bash*`, `file_write`, `*`)
- **Actions**: `allow` (auto-approve), `deny` (block), `ask` (normal approval flow)
- **Priority**: Higher priority evaluated first, first match wins
- **Enforcement**: `evaluate_tool_policies_batch()` called in `WebUI.approve_tools()`
before the `auto_approve` check
- **MCP granular policies**: MCP resources and prompts are evaluated using their
`approval_label` for fine-grained control:
- Resource reads: `mcp_resource__{uri}` (e.g., `mcp_resource__file:///docs/*` to allow,
`mcp_resource__*` to deny all)
- Prompt invocations: `mcp__{server}__{prompt}` (e.g., `mcp__trusted__*` to allow,
`mcp__*` to require approval for all)
- Built-in tools continue to use `func_name` for backward compatibility
### Prompt Templates
Admin-curated system message templates injected at workstream startup:
- **Runtime behavior**: Templates are loaded once at session creation and injected
into the system message *before* user `instructions`. Templates set the baseline;
instructions customize per-workstream behavior.
- **Default templates**: All `is_default=true` templates auto-apply to new
workstreams, concatenated in alphabetical order by name. Use name prefixes
(e.g. `01-safety`, `02-style`) to control ordering.
- **Explicit selection**: `--template <name>` CLI flag, `template` field on
`POST /v1/api/workstreams/new`, console creation modal dropdown, scheduled task
config, and channel adapter config. An explicit template *replaces* defaults.
- **Variables**: Three built-in placeholders resolved at load time:
`{{model}}` (active model name), `{{ws_id}}` (workstream ID),
`{{node_id}}` (server node ID). Unrecognized placeholders are kept as-is.
- **Runtime switching**: `/template <name>` to switch, `/template clear` to revert
to defaults, `/template` to show current. Persisted across resume.
- **Categories**: general, engineering, support, custom, mcp
- **Content limit**: 32 KB per template (enforced on create/update)
- **Storage**: `prompt_templates` table with JSON `variables` array. Migration 010
adds `template` column to `scheduled_tasks`.
- **MCP sync**: MCP server prompts auto-sync into prompt_templates with
`origin="mcp"`, `mcp_server` set, and `readonly=True`. Manual templates take
precedence on name collision. MCP-synced content updates reset `is_default` to
prevent compromised servers from injecting defaults. Admin UI shows origin badge
and disables edit/delete for MCP-sourced templates.
### Usage Tracking
Per-LLM-request token and tool call metrics:
- **Recording**: `on_status()` in `WebUI` records a `usage_event` after each
LLM response with prompt/completion tokens, tool call count, model, ws_id
- **Querying**: `GET /v1/api/admin/usage` with `group_by` (day/hour/model/user)
and time range filtering
- **Pruning**: `prune_usage_events(retention_days=90)` and
`prune_audit_events(retention_days=365)` run automatically via the
console scheduler's periodic cleanup cycle
### Audit Logging
Append-only trail of admin actions:
- **Recording**: `record_audit()` helper called from all admin mutation handlers
- **Events captured**: user.create, user.delete, token.create, token.revoke,
channel.link, channel.unlink, role.create, role.update, role.delete,
role.assign, role.unassign, policy.create, policy.update, policy.delete,
template.create, template.update, template.delete, org.update
- **Querying**: `GET /v1/api/admin/audit` with action/user/time filters + pagination
## Database Schema
Migration 008 adds 7 tables:
| Table | Purpose |
|-------|---------|
| `orgs` | Organizations (single default org for now) |
| `roles` | Named permission bundles (3 builtin + custom) |
| `user_roles` | User-to-role assignments (composite PK) |
| `tool_policies` | Per-tool approve/deny/ask rules |
| `prompt_templates` | Reusable system message templates |
| `usage_events` | Per-request token/tool metrics |
| `audit_events` | Admin action log |
Also adds `org_id` column to `users` table.
## API Endpoints
All under `/v1/api/admin/` (requires `approve` scope + granular permission).
| Group | Endpoints | Permission |
|-------|-----------|------------|
| Users / Tokens / Channels | 9 (CRUD) | `admin.users` |
| Roles | 7 (CRUD + assignment) | `admin.roles` / `admin.users` |
| Orgs | 3 (list, get, update) | `admin.orgs` |
| Tool Policies | 4 (CRUD) | `admin.policies` |
| Prompt Templates | 4 (CRUD) | `admin.templates` |
| Schedules | 6 (CRUD + runs) | `admin.schedules` |
| Watches | 3 (list, create, cancel) | `admin.watches` |
| Usage | 1 (aggregated query) | `admin.usage` |
| Audit | 1 (paginated, filtered) | `admin.audit` |
Full OpenAPI spec at `/openapi.json` and Swagger UI at `/docs`.
## Admin Console UI
5 new tabs added to the admin panel (10 total):
- **Roles** — CRUD roles, permission checkbox grid, user role assignment modal
- **Policies** — CRUD tool policies with colored action badges (green/red/amber)
- **Templates** — CRUD prompt templates with wide modal, textarea editor
- **Usage** — Summary readouts + CSS bar chart, time range + group-by selectors
- **Audit** — Filterable log with relative timestamps, load-more pagination
Tabs are permission-gated: hidden if the user lacks the required permission.
## SDK
Both Python and TypeScript console SDKs expose governance methods:
**Python** (`TurnstoneConsole` / `AsyncTurnstoneConsole`):
- `list_roles()`, `create_role()`, `update_role()`, `delete_role()`
- `list_user_roles()`, `assign_role()`, `unassign_role()`
- `list_orgs()`, `get_org()`, `update_org()`
- `list_policies()`, `create_policy()`, `update_policy()`, `delete_policy()`
- `list_templates()`, `create_template()`, `update_template()`, `delete_template()`
- `get_usage(since, group_by=...)`, `get_audit(action=..., limit=...)`
**TypeScript** (`TurnstoneConsole`):
- Same methods with camelCase naming and typed interfaces
## Security Considerations
- **Privilege escalation prevented**: `admin_assign_role` blocks self-assignment
and requires caller to hold a superset of the target role's permissions
- **Permission validation**: Role create/update validates permissions against
a 15-item allowlist (`_VALID_PERMISSIONS`)
- **Self-deletion blocked**: `admin_delete_user` rejects attempts to delete
your own account (matching the self-assignment guard on role endpoints)
- **Field allowlists**: Storage `update_*` methods filter fields against
allowlists (`_ROLE_MUTABLE`, `_POLICY_MUTABLE`, etc.) — handler bugs
cannot overwrite `role_id`, `builtin`, `created`, or other protected columns
- **Bootstrap safety**: `handle_auth_setup` fails and rolls back if admin role
assignment fails, preventing locked-out first user
- **API token RBAC**: `_authenticate_api_token` loads permissions from user's
roles, ensuring API tokens are subject to RBAC enforcement
- **Policy evaluation is fail-open**: If storage is unavailable, tool policies
degrade to the existing approval flow (not auto-approve)
- **Audit IP resolution**: `_audit_context()` prefers `X-Forwarded-For` for
client IP when behind a reverse proxy, falling back to `request.client.host`
+6
View File
@@ -75,6 +75,7 @@ Both `TurnstoneServer` (sync) and `AsyncTurnstoneServer` (async) expose:
| | `approve(*, ws_id, approved, feedback, always)` | `StatusResponse` |
| | `plan_feedback(*, ws_id, feedback)` | `StatusResponse` |
| | `command(*, ws_id, command)` | `StatusResponse` |
| | `cancel(ws_id)` | `StatusResponse` |
| **Streaming** | `stream_events(ws_id)` | `Iterator[ServerEvent]` |
| | `stream_global_events()` | `Iterator[ServerEvent]` |
| **High-level** | `send_and_wait(message, ws_id, *, timeout, on_event)` | `TurnResult` |
@@ -95,6 +96,7 @@ Both `TurnstoneConsole` (sync) and `AsyncTurnstoneConsole` (async) expose:
| | `nodes(*, sort, limit, offset)` | `ClusterNodesResponse` |
| | `workstreams(*, state, node, search, sort, page, per_page)` | `ClusterWorkstreamsResponse` |
| | `node_detail(node_id)` | `NodeDetailResponse` |
| | `snapshot()` | `ClusterSnapshotResponse` |
| | `create_workstream(*, node_id, name, model, initial_message)` | `ConsoleCreateWsResponse` |
| **Schedules** | `list_schedules()` | `ListSchedulesResponse` |
| | `create_schedule(*, name, schedule_type, initial_message, ...)` | `ScheduleInfo` |
@@ -128,6 +130,7 @@ SSE events are deserialized into typed dataclasses. Use `event.type` to discrimi
| `error` | `ErrorEvent` | `message` |
| `info` | `InfoEvent` | `message` |
| `stream_end` | `StreamEndEvent` | — |
| `cancelled` | `CancelledEvent` | — |
**Global events** (from `stream_global_events()`):
@@ -146,6 +149,9 @@ SSE events are deserialized into typed dataclasses. Use `event.type` to discrimi
| `node_lost` | `NodeLostEvent` | `node_id` |
| `cluster_state` | `ClusterStateEvent` | `ws_id`, `node_id`, `state`, `tokens` |
| `ws_created` | `ClusterWsCreatedEvent` | `ws_id`, `node_id`, `name` |
| `ws_closed` | `ClusterWsClosedEvent` | `ws_id` |
| `ws_rename` | `ClusterWsRenameEvent` | `ws_id`, `name` |
| `snapshot` | `ClusterSnapshotEvent` | `nodes`, `overview`, `timestamp` |
### TurnResult
+28
View File
@@ -92,6 +92,34 @@ Public paths bypass authentication entirely: `/`, `/health`, `/metrics`,
`/static/*`, `/shared/*`, `/docs`, `/openapi.json`, `/api/auth/login`,
`/api/auth/logout`, `/api/auth/status`, `/api/auth/setup`.
### RBAC (Granular Permissions)
> See also: [Governance documentation](governance.md)
Scopes provide coarse endpoint-level access control. For finer-grained
enforcement, the governance layer adds 15 named permissions checked
per-endpoint by `require_permission()`. Permissions are bundled into
roles; users are assigned roles via the `user_roles` join table.
At login, `_load_user_permissions()` aggregates all permissions from
the user's assigned roles. `_permissions_to_scopes()` derives legacy
scopes for backward compatibility (e.g., any `admin.*` permission
implies the `approve` scope). The JWT carries both `scopes` and
`permissions` claims.
Three built-in roles are seeded by migration 008:
| Role | Permissions |
|------|-------------|
| admin | All 15 permissions |
| operator | read, write, workstreams.create, workstreams.close |
| viewer | read |
Custom roles can be created with any subset of the valid permissions.
Role creation and update validate permissions against a static allowlist.
Self-assignment is blocked, and assigning a role requires the caller to
hold a superset of the target role's permissions.
---
## Login Flows
+231 -10
View File
@@ -1,6 +1,6 @@
# Tools Reference
turnstone exposes 15 built-in tools plus any number of external MCP tools to the
turnstone exposes 18 built-in tools plus any number of external MCP tools to the
LLM via the OpenAI function-calling interface. Built-in tools are defined as JSON
files under `turnstone/tools/` and loaded at startup by `turnstone/core/tools.py`.
MCP tools are discovered from configured MCP servers at startup by
@@ -46,12 +46,12 @@ schema plus turnstone-specific metadata keys:
| Name | Description |
|---------------------|-------------|
| `TOOLS` | All 15 tool definitions (sent to the model). |
| `TOOLS` | All 18 tool definitions (sent to the model). |
| `AGENT_TOOLS` | Tools with `agent: true` -- available to plan sub-agents. Read-only tools. |
| `TASK_AGENT_TOOLS` | Tools with `task_agent: true` -- available to task sub-agents. Includes write operations. |
| `AGENT_AUTO_TOOLS` | Set of tool names with `auto_approve: true` -- no user confirmation needed. |
| `TASK_AUTO_TOOLS` | Same as `AGENT_AUTO_TOOLS` (identical filter). |
| `BUILTIN_TOOL_NAMES`| Frozenset of all 15 built-in tool names. Used by tool search to distinguish always-on tools from deferrable MCP tools. |
| `BUILTIN_TOOL_NAMES`| Frozenset of all 18 built-in tool names. Used by tool search to distinguish always-on tools from deferrable MCP tools. |
| `PRIMARY_KEY_MAP` | Dict mapping tool name to its `primary_key` parameter name. |
---
@@ -69,7 +69,7 @@ Tool execution follows a three-phase pipeline inside `ChatSession._execute_tools
- Parses the JSON arguments (with fallback for malformed JSON).
- If JSON parsing fails entirely, uses `PRIMARY_KEY_MAP` to map a bare string
to the correct parameter.
- Dispatches to the matching `_prepare_{func_name}()` handler. There are 15
- Dispatches to the matching `_prepare_{func_name}()` handler. There are 18
built-in tools plus `tool_search` (synthetic, client-side BM25 fallback) and
the generic `_prepare_mcp_tool()` handler for MCP tools.
- Validates arguments and builds a preview dict containing:
@@ -168,6 +168,8 @@ Every tool defines a `primary_key`. The mapping is:
| `recall` | `query` |
| `forget` | `key` |
| `notify` | `message` |
| `read_resource` | `uri` |
| `use_prompt` | `name` |
---
@@ -189,15 +191,17 @@ Execute a bash command and return stdout + stderr.
### read_file
Read the contents of a file, returning numbered lines.
Read the contents of a file, returning numbered lines for text files or
base64-encoded image data for supported image formats.
| Parameter | Type | Required | Description |
|-----------|---------|----------|-------------|
| `path` | string | yes | Absolute or relative file path. |
| `offset` | integer | no | Line number to start from (1-based, default: 1). |
| `limit` | integer | no | Maximum number of lines to read. Omit for full file. |
| `offset` | integer | no | Line number to start from (1-based, default: 1). Text files only. |
| `limit` | integer | no | Maximum number of lines to read. Omit for full file. Text files only. |
- **What it does**: Reads the file and returns content with line numbers. Must be called before `edit_file` on the same path (the session tracks which files have been read).
- **What it does**: For text files, reads and returns content with line numbers. For image files (PNG, JPEG, GIF, WebP, BMP, TIFF, ICO), returns image data as multi-part content when the model supports vision, or a text description when it does not. SVG files are read as text. Images larger than 4 MB are rejected. Must be called before `edit_file` on the same path (the session tracks which files have been read).
- **Vision support**: Controlled by `ModelCapabilities.supports_vision`. All commercial OpenAI and Anthropic models have vision enabled. Local models (vLLM, llama.cpp, NIM) default to off — enable via `[models.*.capabilities] supports_vision = true` in config.toml.
- **Auto-approve**: Yes.
- **Agent availability**: `agent` and `task_agent`.
@@ -420,6 +424,81 @@ Provide either `username` for user-based targeting or `channel_type` +
---
### watch
Set up periodic polling of a shell command within the current workstream.
Results are injected back into the conversation as synthetic user messages,
triggering the model to respond and act. Use for monitoring CI/CD pipelines,
PR reviews, deployments, file changes, etc.
| Parameter | Type | Required | Description |
|-------------|---------|----------|-------------|
| `action` | string | yes | `create`, `list`, or `cancel`. |
| `command` | string | create | Shell command to poll periodically. |
| `poll_every`| string | no | Poll interval as duration (`30s`, `5m`, `1h`). Default: `5m`. |
| `stop_on` | string | no | Python expression for stop condition (see below). Omit for change detection. |
| `name` | string | create | Human-readable watch name (e.g. `pr-review`). Used as identifier for cancel. |
| `max_polls` | integer | no | Max poll cycles before auto-cancel. Default: 100. |
**Actions:**
- `create` — Start a new watch. Requires approval (same as bash — runs shell
commands). Persists to the `watches` table; the server-level `WatchRunner`
daemon polls every 15 seconds for due watches.
- `list` — Show all active watches in this workstream. Auto-approved.
- `cancel` — Stop a watch by name or ID prefix. Auto-approved.
**Stop condition DSL** — The `stop_on` parameter accepts a Python expression
evaluated after each poll. Available variables:
| Variable | Type | Description |
|---------------|------------|-------------|
| `output` | `str` | stdout (+stderr) of the command. |
| `data` | `Any` | `json.loads(output)`, or `None` if not valid JSON. |
| `exit_code` | `int` | Process exit code. |
| `prev_output` | `str|None` | Previous poll's stdout (`None` on first poll). |
| `changed` | `bool` | `True` if output differs from previous poll. |
Safe builtins: `len`, `str`, `int`, `float`, `bool`, `abs`, `min`, `max`,
`any`, `all`, `isinstance`, `sorted`. No `import`, `open`, `exec`, or
`eval`. Security model: equivalent to `bash` — the model already has shell
access.
**Examples:**
```
data["state"] == "MERGED"
"error" in output
exit_code != 0
changed and "ready" in output.lower()
data.get("mergedAt") is not None
```
**Lifecycle:**
1. Model calls `watch(action="create", ...)` — persisted to SQLite.
2. `WatchRunner` daemon polls for due watches every 15s.
3. Each poll runs the command, evaluates the condition.
4. When the condition fires (or max polls reached), the result is injected
as a synthetic user message and the watch auto-cancels.
5. If the workstream was evicted, it is restored before injection.
6. Watches survive server restart (overdue watches fire once on recovery).
**Constraints:**
- Max 5 active watches per workstream.
- Poll interval: 10s24h.
- Output truncated at 64 KB.
- Max 5 consecutive watch dispatches per worker thread (depth guard).
- Duplicate names rejected within the same workstream.
- **Auto-approve**: `create` requires approval; `list` and `cancel` are auto-approved.
- **Agent availability**: Main session only — not available to plan/task sub-agents.
> See [Watch Architecture](diagrams/png/18-watch-architecture.png) for the
> full poll → evaluate → dispatch flow.
---
## Summary Table
| Tool | Category | Auto-approve | agent | task_agent | primary_key |
@@ -439,6 +518,9 @@ Provide either `username` for user-based targeting or `channel_type` +
| `recall` | Memory | Yes | No | No | `query` |
| `forget` | Memory | Yes | No | No | `key` |
| `notify` | Notify | Yes | Yes | Yes | `message` |
| `watch` | Monitor | No (create) | No | No | `command` |
| `read_resource`| MCP | No | Yes | Yes | `uri` |
| `use_prompt` | MCP | No | Yes | Yes | `name` |
| `tool_search`| Search | Yes | No | No | `query` |
---
@@ -491,7 +573,7 @@ CLI flags override the config file:
search stays off and all tools are sent to the model directly.
2. **Partitioning**: When active, tools are split into two sets:
- **Always-on** -- the 15 built-in tools (members of `BUILTIN_TOOL_NAMES`).
- **Always-on** -- the 18 built-in tools (members of `BUILTIN_TOOL_NAMES`).
These are always visible to the model.
- **Deferred** -- all MCP tools. These are not sent in the tool list unless
the model searches for them.
@@ -515,6 +597,8 @@ where the model can interactively search for tools it needs.
## MCP Tools (External)
> See also: [MCP Architecture diagram](diagrams/png/20-mcp-architecture.png)
Turnstone supports the [Model Context Protocol](https://modelcontextprotocol.io/)
(MCP) for connecting external tool servers — GitHub, databases, filesystems, or any
MCP-compatible service.
@@ -532,7 +616,7 @@ MCP-compatible service.
3. **Schema conversion**: Each MCP tool's `inputSchema` is converted to OpenAI
function-calling format. The tool name is prefixed: `mcp__{server}__{tool}`.
4. **Merging**: MCP tools are appended after the 15 built-in tools via
4. **Merging**: MCP tools are appended after the 18 built-in tools via
`merge_mcp_tools()`. Built-in tools appear first, giving them natural LLM priority.
When dynamic tool search is active, MCP tools are deferred rather than directly
visible -- the model discovers them via search as needed (see
@@ -651,3 +735,140 @@ MCP refresh complete:
MCP refresh complete:
github: no changes
```
---
## MCP Resources
MCP servers can expose **resources** -- named data items (files, database rows,
API responses) addressable by URI. turnstone discovers resources at startup and
makes them available to the model via the `read_resource` built-in tool.
### Discovery
During the MCP `initialize` handshake, `MCPClientManager` checks each server's
capabilities for the `resources` capability. For servers that declare it:
1. `list_resources` fetches static resources (fixed URIs).
2. `list_resource_templates` fetches URI templates (parameterized patterns like
`db://tables/{table}/rows/{id}`).
Both are stored as `{uri, name, description, mimeType, server}` dicts and
merged into a unified catalog.
### Resource catalog in system message
The first 50 resources are injected into the system message as an XML-delimited
block so the model knows what URIs are available:
```xml
<mcp-resources>
file:///project/README.md Project readme
db://users/schema User table schema
</mcp-resources>
Use read_resource(uri='...') to access the resources listed above.
```
### read_resource tool
| Parameter | Type | Required | Description |
|-----------|--------|----------|-------------|
| `uri` | string | yes | The resource URI to read. |
- **What it does**: Reads the resource from its MCP server via `MCPClientManager.read_resource_sync()`. Returns text content for text resources or base64-encoded data for binary resources. Output is truncated by the standard tool output limiter.
- **Auto-approve**: No -- requires user confirmation (reads external data).
- **Agent availability**: `agent` and `task_agent`.
### Capability guards
The `read_resource` tool schema is always loaded (it is a built-in JSON schema),
but resource discovery only runs for servers that declare the `resources`
capability. Servers without the capability contribute zero resources to the
catalog.
### Refresh
Resource lists stay current through the same three-tier mechanism as tool lists:
1. **Push** -- Servers declaring `resources.listChanged: true` send
`notifications/resources/list_changed`, triggering an immediate refresh.
2. **Periodic** -- Servers without push are polled on the configured refresh
interval (default 4 hours, same timer as tools).
3. **Manual** -- `/mcp refresh` re-fetches resources alongside tools.
---
## MCP Prompts
MCP servers can also expose **prompts** -- reusable message templates with
optional arguments. turnstone discovers prompts at startup for servers that
declare the `prompts` capability.
### Discovery
Prompt discovery mirrors resource discovery: `list_prompts` is called during
the `initialize` handshake. Each prompt is stored with its prefixed name
(`mcp__{server}__{prompt}`), description, and argument schema.
### use_prompt tool
| Parameter | Type | Required | Description |
|-------------|--------|----------|-------------|
| `name` | string | yes | The prompt name (e.g. `mcp__server__prompt_name`). |
| `arguments` | object | no | Key-value argument pairs for the prompt. Values must be strings. |
- **What it does**: Invokes an MCP prompt template by name via `MCPClientManager.get_prompt_sync()`, expanding it into messages. Returns the expanded prompt content formatted as `[role]: content` blocks joined with blank lines. The prompt catalog is listed in the system message so the model knows which prompts are available. Output is truncated by the standard tool output limiter.
- **Auto-approve**: No -- requires user confirmation (invokes external prompt servers).
- **Agent availability**: `agent` and `task_agent`.
### Invocation
`MCPClientManager.get_prompt_sync()` calls the server's `get_prompt` method
with the provided arguments and returns the expanded messages. The `use_prompt`
built-in tool exposes this to the model as a function call.
### Governance Sync
Discovered MCP prompts are automatically synced into the `prompt_templates`
governance table as first-class governed templates:
- **Origin tracking**: MCP-sourced templates have `origin="mcp"` and
`mcp_server` set to the server name. Manual templates have
`origin="manual"`.
- **Read-only**: MCP-sourced templates are `readonly=True`. The admin API
returns 403 on update/delete attempts. The admin UI disables edit/delete
buttons and shows an origin badge.
- **Precedence**: If a manual template and MCP prompt share the same name,
the manual template wins and the MCP prompt is skipped (with a log
warning).
- **Lifecycle**: Templates are created on connect, updated on prompt list
refresh, and removed when the MCP server no longer exposes the prompt.
The sync runs automatically on connect, on `PromptListChangedNotification`,
and on manual `/mcp refresh`.
- **Schema**: Migration 009 adds `origin`, `mcp_server`, and `readonly`
columns to the `prompt_templates` table.
The `use_prompt` tool allows the model to invoke any discovered MCP prompt at
runtime. A catalog of up to 30 prompts is injected into the system message
inside `<mcp-prompts>` XML tags so the model can discover available prompts.
---
## MCP UI Visibility
MCP server, resource, and prompt counts are surfaced across the UI:
- **Server `/health` endpoint**: Returns `mcp.servers`, `mcp.resources`,
`mcp.prompts` when MCP is configured
- **Server UI**: Magenta status badge in the header showing server count,
with resource/prompt counts in tooltip
- **Console cluster status bar**: MCP metrics (servers/resources/prompts)
with magenta LED dot indicator, shown after a divider from workstream
metrics
- **Console node detail**: Per-node MCP summary showing server, resource,
and prompt counts
- **Console collector**: Aggregates MCP counts across all nodes in the
cluster overview
MCP indicators use the `--magenta` design token for consistent theming
across light and dark modes.
+2 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "turnstone"
version = "0.4.6"
version = "0.5.6"
description = "Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation."
readme = "README.md"
license = "BUSL-1.1"
@@ -62,6 +62,7 @@ turnstone-console = "turnstone.console.server:main"
turnstone-sim = "turnstone.sim.cli:main"
turnstone-admin = "turnstone.admin:main"
turnstone-channel = "turnstone.channels.cli:main"
turnstone-bootstrap = "turnstone.bootstrap:main"
[tool.hatch.build.targets.wheel]
include = [
+29 -78
View File
@@ -10,9 +10,7 @@
"get": {
"summary": "Cluster state summary",
"operationId": "v1_api_cluster_overview_get",
"tags": [
"Cluster"
],
"tags": ["Cluster"],
"responses": {
"200": {
"description": "Success",
@@ -31,9 +29,7 @@
"get": {
"summary": "Paginated node list",
"operationId": "v1_api_cluster_nodes_get",
"tags": [
"Cluster"
],
"tags": ["Cluster"],
"parameters": [
{
"name": "sort",
@@ -42,11 +38,7 @@
"schema": {
"type": "string",
"default": "activity",
"enum": [
"activity",
"tokens",
"name"
]
"enum": ["activity", "tokens", "name"]
},
"description": "Sort field"
},
@@ -89,9 +81,7 @@
"get": {
"summary": "Filtered workstream list",
"operationId": "v1_api_cluster_workstreams_get",
"tags": [
"Cluster"
],
"tags": ["Cluster"],
"parameters": [
{
"name": "state",
@@ -99,13 +89,7 @@
"required": false,
"schema": {
"type": "string",
"enum": [
"running",
"thinking",
"attention",
"idle",
"error"
]
"enum": ["running", "thinking", "attention", "idle", "error"]
},
"description": "Filter by state"
},
@@ -134,11 +118,7 @@
"schema": {
"type": "string",
"default": "state",
"enum": [
"state",
"tokens",
"name"
]
"enum": ["state", "tokens", "name"]
},
"description": "Sort field"
},
@@ -181,9 +161,7 @@
"get": {
"summary": "Single node detail",
"operationId": "v1_api_cluster_node_{node_id}_get",
"tags": [
"Cluster"
],
"tags": ["Cluster"],
"parameters": [
{
"name": "node_id",
@@ -222,9 +200,7 @@
"post": {
"summary": "Create workstream via MQ dispatch",
"operationId": "v1_api_cluster_workstreams_new_post",
"tags": [
"Cluster"
],
"tags": ["Cluster"],
"requestBody": {
"required": true,
"content": {
@@ -283,9 +259,7 @@
"get": {
"summary": "Cluster SSE event stream",
"operationId": "v1_api_cluster_events_get",
"tags": [
"Streaming"
],
"tags": ["Streaming"],
"description": "Server-Sent Events stream for real-time cluster updates. Returns text/event-stream with node_joined, node_lost, cluster_state, ws_created, ws_closed, ws_rename events.",
"responses": {
"200": {
@@ -298,9 +272,7 @@
"post": {
"summary": "Authenticate with a token",
"operationId": "v1_api_auth_login_post",
"tags": [
"Auth"
],
"tags": ["Auth"],
"requestBody": {
"required": true,
"content": {
@@ -339,9 +311,7 @@
"post": {
"summary": "Clear auth cookie",
"operationId": "v1_api_auth_logout_post",
"tags": [
"Auth"
],
"tags": ["Auth"],
"responses": {
"200": {
"description": "Success",
@@ -360,9 +330,7 @@
"get": {
"summary": "Console health check",
"operationId": "health_get",
"tags": [
"Observability"
],
"tags": ["Observability"],
"responses": {
"200": {
"description": "Success",
@@ -389,9 +357,7 @@
"type": "string"
}
},
"required": [
"error"
],
"required": ["error"],
"title": "ErrorResponse",
"type": "object"
},
@@ -400,9 +366,7 @@
"properties": {
"status": {
"default": "ok",
"examples": [
"ok"
],
"examples": ["ok"],
"title": "Status",
"type": "string"
}
@@ -419,9 +383,7 @@
"type": "string"
}
},
"required": [
"token"
],
"required": ["token"],
"title": "AuthLoginRequest",
"type": "object"
},
@@ -435,17 +397,12 @@
},
"role": {
"description": "Assigned role",
"examples": [
"full",
"read"
],
"examples": ["full", "read"],
"title": "Role",
"type": "string"
}
},
"required": [
"role"
],
"required": ["role"],
"title": "AuthLoginResponse",
"type": "object"
},
@@ -557,9 +514,7 @@
"type": "integer"
}
},
"required": [
"nodes"
],
"required": ["nodes"],
"title": "ClusterNodesResponse",
"type": "object"
},
@@ -632,9 +587,7 @@
"type": "string"
}
},
"required": [
"node_id"
],
"required": ["node_id"],
"title": "ClusterNodeInfo",
"type": "object"
},
@@ -668,9 +621,7 @@
"type": "integer"
}
},
"required": [
"workstreams"
],
"required": ["workstreams"],
"title": "ClusterWorkstreamsResponse",
"type": "object"
},
@@ -726,9 +677,7 @@
"type": "integer"
}
},
"required": [
"id"
],
"required": ["id"],
"title": "ClusterWorkstreamInfo",
"type": "object"
},
@@ -771,9 +720,7 @@
"type": "boolean"
}
},
"required": [
"node_id"
],
"required": ["node_id"],
"title": "NodeDetailResponse",
"type": "object"
},
@@ -802,6 +749,12 @@
"description": "Optional first message sent after creation",
"title": "Initial Message",
"type": "string"
},
"template": {
"default": "",
"description": "Prompt template name (replaces default templates)",
"title": "Template",
"type": "string"
}
},
"title": "ConsoleCreateWsRequest",
@@ -832,9 +785,7 @@
"properties": {
"status": {
"default": "ok",
"examples": [
"ok"
],
"examples": ["ok"],
"title": "Status",
"type": "string"
},
+142 -154
View File
@@ -10,9 +10,7 @@
"get": {
"summary": "List active workstreams",
"operationId": "v1_api_workstreams_get",
"tags": [
"Workstreams"
],
"tags": ["Workstreams"],
"responses": {
"200": {
"description": "Success",
@@ -31,9 +29,7 @@
"get": {
"summary": "Dashboard with workstream details and aggregates",
"operationId": "v1_api_dashboard_get",
"tags": [
"Workstreams"
],
"tags": ["Workstreams"],
"responses": {
"200": {
"description": "Success",
@@ -52,9 +48,7 @@
"post": {
"summary": "Create a new workstream",
"operationId": "v1_api_workstreams_new_post",
"tags": [
"Workstreams"
],
"tags": ["Workstreams"],
"requestBody": {
"required": true,
"content": {
@@ -93,9 +87,7 @@
"post": {
"summary": "Close a workstream",
"operationId": "v1_api_workstreams_close_post",
"tags": [
"Workstreams"
],
"tags": ["Workstreams"],
"requestBody": {
"required": true,
"content": {
@@ -134,9 +126,7 @@
"post": {
"summary": "Send a user message",
"operationId": "v1_api_send_post",
"tags": [
"Chat"
],
"tags": ["Chat"],
"requestBody": {
"required": true,
"content": {
@@ -185,9 +175,7 @@
"post": {
"summary": "Approve or deny a tool call",
"operationId": "v1_api_approve_post",
"tags": [
"Chat"
],
"tags": ["Chat"],
"requestBody": {
"required": true,
"content": {
@@ -226,9 +214,7 @@
"post": {
"summary": "Respond to a plan review",
"operationId": "v1_api_plan_post",
"tags": [
"Chat"
],
"tags": ["Chat"],
"requestBody": {
"required": true,
"content": {
@@ -267,9 +253,7 @@
"post": {
"summary": "Execute a slash command",
"operationId": "v1_api_command_post",
"tags": [
"Chat"
],
"tags": ["Chat"],
"requestBody": {
"required": true,
"content": {
@@ -314,13 +298,60 @@
}
}
},
"/v1/api/cancel": {
"post": {
"summary": "Cancel the active generation in a workstream",
"operationId": "v1_api_cancel_post",
"tags": ["Chat"],
"requestBody": {
"required": true,
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/CancelRequest"
}
}
}
},
"responses": {
"200": {
"description": "Success",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/StatusResponse"
}
}
}
},
"400": {
"description": "Error 400",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
},
"404": {
"description": "Error 404",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ErrorResponse"
}
}
}
}
}
}
},
"/v1/api/events": {
"get": {
"summary": "Per-workstream SSE event stream",
"operationId": "v1_api_events_get",
"tags": [
"Streaming"
],
"tags": ["Streaming"],
"description": "Opens a Server-Sent Events stream scoped to a single workstream. Returns text/event-stream. See API reference for event types.",
"parameters": [
{
@@ -354,9 +385,7 @@
"get": {
"summary": "Global SSE event stream",
"operationId": "v1_api_events_global_get",
"tags": [
"Streaming"
],
"tags": ["Streaming"],
"description": "Global Server-Sent Events stream for state-change broadcasts across all workstreams. Returns text/event-stream.",
"responses": {
"200": {
@@ -369,9 +398,7 @@
"get": {
"summary": "List saved workstreams",
"operationId": "v1_api_workstreams_saved_get",
"tags": [
"Workstreams"
],
"tags": ["Workstreams"],
"responses": {
"200": {
"description": "Success",
@@ -390,9 +417,7 @@
"post": {
"summary": "Authenticate with a token",
"operationId": "v1_api_auth_login_post",
"tags": [
"Auth"
],
"tags": ["Auth"],
"requestBody": {
"required": true,
"content": {
@@ -431,9 +456,7 @@
"post": {
"summary": "Create first admin user",
"operationId": "v1_api_auth_setup_post",
"tags": [
"Auth"
],
"tags": ["Auth"],
"requestBody": {
"required": true,
"content": {
@@ -492,9 +515,7 @@
"get": {
"summary": "Return auth state",
"operationId": "v1_api_auth_status_get",
"tags": [
"Auth"
],
"tags": ["Auth"],
"responses": {
"200": {
"description": "Success",
@@ -513,9 +534,7 @@
"post": {
"summary": "Clear auth cookie",
"operationId": "v1_api_auth_logout_post",
"tags": [
"Auth"
],
"tags": ["Auth"],
"responses": {
"200": {
"description": "Success",
@@ -534,9 +553,7 @@
"get": {
"summary": "Server health check",
"operationId": "health_get",
"tags": [
"Observability"
],
"tags": ["Observability"],
"responses": {
"200": {
"description": "Success",
@@ -563,9 +580,7 @@
"type": "string"
}
},
"required": [
"error"
],
"required": ["error"],
"title": "ErrorResponse",
"type": "object"
},
@@ -574,9 +589,7 @@
"properties": {
"status": {
"default": "ok",
"examples": [
"ok"
],
"examples": ["ok"],
"title": "Status",
"type": "string"
}
@@ -625,19 +638,14 @@
},
"role": {
"description": "Legacy role",
"examples": [
"full",
"read"
],
"examples": ["full", "read"],
"title": "Role",
"type": "string"
},
"scopes": {
"default": "",
"description": "Comma-separated scopes",
"examples": [
"read,write,approve"
],
"examples": ["read,write,approve"],
"title": "Scopes",
"type": "string"
},
@@ -648,9 +656,7 @@
"type": "string"
}
},
"required": [
"role"
],
"required": ["role"],
"title": "AuthLoginResponse",
"type": "object"
},
@@ -673,11 +679,7 @@
"type": "string"
}
},
"required": [
"username",
"display_name",
"password"
],
"required": ["username", "display_name", "password"],
"title": "AuthSetupRequest",
"type": "object"
},
@@ -714,10 +716,7 @@
"type": "string"
}
},
"required": [
"user_id",
"username"
],
"required": ["user_id", "username"],
"title": "AuthSetupResponse",
"type": "object"
},
@@ -737,11 +736,7 @@
"type": "boolean"
}
},
"required": [
"auth_enabled",
"has_users",
"setup_required"
],
"required": ["auth_enabled", "has_users", "setup_required"],
"title": "AuthStatusResponse",
"type": "object"
},
@@ -758,10 +753,7 @@
"type": "string"
}
},
"required": [
"message",
"ws_id"
],
"required": ["message", "ws_id"],
"title": "SendRequest",
"type": "object"
},
@@ -769,17 +761,12 @@
"properties": {
"status": {
"description": "'ok' or 'busy'",
"examples": [
"ok",
"busy"
],
"examples": ["ok", "busy"],
"title": "Status",
"type": "string"
}
},
"required": [
"status"
],
"required": ["status"],
"title": "SendResponse",
"type": "object"
},
@@ -815,10 +802,7 @@
"type": "string"
}
},
"required": [
"approved",
"ws_id"
],
"required": ["approved", "ws_id"],
"title": "ApproveRequest",
"type": "object"
},
@@ -835,10 +819,7 @@
"type": "string"
}
},
"required": [
"feedback",
"ws_id"
],
"required": ["feedback", "ws_id"],
"title": "PlanFeedbackRequest",
"type": "object"
},
@@ -855,13 +836,22 @@
"type": "string"
}
},
"required": [
"command",
"ws_id"
],
"required": ["command", "ws_id"],
"title": "CommandRequest",
"type": "object"
},
"CancelRequest": {
"properties": {
"ws_id": {
"description": "Target workstream ID",
"title": "Ws Id",
"type": "string"
}
},
"required": ["ws_id"],
"title": "CancelRequest",
"type": "object"
},
"CreateWorkstreamRequest": {
"properties": {
"name": {
@@ -887,6 +877,12 @@
"description": "Workstream ID to resume atomically during creation (empty = fresh start)",
"title": "Resume Ws",
"type": "string"
},
"template": {
"default": "",
"description": "Prompt template name (replaces default templates)",
"title": "Template",
"type": "string"
}
},
"title": "CreateWorkstreamRequest",
@@ -917,10 +913,7 @@
"type": "integer"
}
},
"required": [
"ws_id",
"name"
],
"required": ["ws_id", "name"],
"title": "CreateWorkstreamResponse",
"type": "object"
},
@@ -932,9 +925,7 @@
"type": "string"
}
},
"required": [
"ws_id"
],
"required": ["ws_id"],
"title": "CloseWorkstreamRequest",
"type": "object"
},
@@ -948,9 +939,7 @@
"type": "array"
}
},
"required": [
"workstreams"
],
"required": ["workstreams"],
"title": "ListWorkstreamsResponse",
"type": "object"
},
@@ -969,11 +958,7 @@
"type": "string"
}
},
"required": [
"id",
"name",
"state"
],
"required": ["id", "name", "state"],
"title": "WorkstreamInfo",
"type": "object"
},
@@ -990,10 +975,7 @@
"$ref": "#/components/schemas/DashboardAggregate"
}
},
"required": [
"workstreams",
"aggregate"
],
"required": ["workstreams", "aggregate"],
"title": "DashboardResponse",
"type": "object"
},
@@ -1093,11 +1075,7 @@
"type": "string"
}
},
"required": [
"id",
"name",
"state"
],
"required": ["id", "name", "state"],
"title": "DashboardWorkstream",
"type": "object"
},
@@ -1111,9 +1089,7 @@
"type": "array"
}
},
"required": [
"workstreams"
],
"required": ["workstreams"],
"title": "ListSavedWorkstreamsResponse",
"type": "object"
},
@@ -1160,22 +1136,14 @@
"type": "integer"
}
},
"required": [
"ws_id",
"created",
"updated",
"message_count"
],
"required": ["ws_id", "created", "updated", "message_count"],
"title": "SavedWorkstreamInfo",
"type": "object"
},
"HealthResponse": {
"properties": {
"status": {
"examples": [
"ok",
"degraded"
],
"examples": ["ok", "degraded"],
"title": "Status",
"type": "string"
},
@@ -1215,38 +1183,58 @@
}
],
"default": null
},
"mcp": {
"anyOf": [
{
"$ref": "#/components/schemas/McpStatus"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"status"
],
"required": ["status"],
"title": "HealthResponse",
"type": "object"
},
"McpStatus": {
"properties": {
"servers": {
"default": 0,
"title": "Servers",
"type": "integer"
},
"resources": {
"default": 0,
"title": "Resources",
"type": "integer"
},
"prompts": {
"default": 0,
"title": "Prompts",
"type": "integer"
}
},
"title": "McpStatus",
"type": "object"
},
"BackendStatus": {
"properties": {
"status": {
"examples": [
"up",
"down"
],
"examples": ["up", "down"],
"title": "Status",
"type": "string"
},
"circuit_state": {
"examples": [
"closed",
"open",
"half_open"
],
"examples": ["closed", "open", "half_open"],
"title": "Circuit State",
"type": "string"
}
},
"required": [
"status",
"circuit_state"
],
"required": ["status", "circuit_state"],
"title": "BackendStatus",
"type": "object"
},
+142
View File
@@ -1,23 +1,40 @@
import { BaseClient, type ClientOptions } from "./base.js";
import type { ClusterEvent } from "./events.js";
import type {
AuditQueryOptions,
AuditResponse,
AuthLoginResponse,
AuthSetupResponse,
AuthStatusResponse,
ClusterNodesResponse,
ClusterOverviewResponse,
ClusterSnapshotResponse,
ClusterWorkstreamsResponse,
ConsoleCreateWsRequest,
ConsoleCreateWsResponse,
ConsoleHealthResponse,
CreatePolicyOptions,
CreateRoleOptions,
CreateScheduleRequest,
CreateTemplateOptions,
ListScheduleRunsResponse,
ListSchedulesResponse,
NodeDetailResponse,
NodesOptions,
OrgInfo,
PromptTemplateInfo,
RoleInfo,
ScheduleInfo,
StatusResponse,
ToolPolicyInfo,
UpdateOrgOptions,
UpdatePolicyOptions,
UpdateRoleOptions,
UpdateScheduleRequest,
UpdateTemplateOptions,
UsageQueryOptions,
UsageResponse,
UserRoleInfo,
WorkstreamsOptions,
} from "./types.js";
@@ -33,6 +50,10 @@ export class TurnstoneConsole extends BaseClient {
return this.request("GET", "/v1/api/cluster/overview");
}
async snapshot(): Promise<ClusterSnapshotResponse> {
return this.request("GET", "/v1/api/cluster/snapshot");
}
async nodes(opts?: NodesOptions): Promise<ClusterNodesResponse> {
return this.request("GET", "/v1/api/cluster/nodes", {
params: {
@@ -152,4 +173,125 @@ export class TurnstoneConsole extends BaseClient {
params: { limit: opts?.limit ?? 50 },
});
}
// -- Governance: Roles ------------------------------------------------------
async listRoles(): Promise<{ roles: RoleInfo[] }> {
return this.request("GET", "/v1/api/admin/roles");
}
async createRole(opts: CreateRoleOptions): Promise<RoleInfo> {
return this.request("POST", "/v1/api/admin/roles", { json: opts });
}
async updateRole(roleId: string, opts: UpdateRoleOptions): Promise<RoleInfo> {
return this.request("PUT", `/v1/api/admin/roles/${roleId}`, {
json: opts,
});
}
async deleteRole(roleId: string): Promise<StatusResponse> {
return this.request("DELETE", `/v1/api/admin/roles/${roleId}`);
}
async listUserRoles(userId: string): Promise<{ roles: UserRoleInfo[] }> {
return this.request("GET", `/v1/api/admin/users/${userId}/roles`);
}
async assignRole(userId: string, roleId: string): Promise<StatusResponse> {
return this.request("POST", `/v1/api/admin/users/${userId}/roles`, {
json: { role_id: roleId },
});
}
async unassignRole(userId: string, roleId: string): Promise<StatusResponse> {
return this.request(
"DELETE",
`/v1/api/admin/users/${userId}/roles/${roleId}`,
);
}
// -- Governance: Organizations ----------------------------------------------
async listOrgs(): Promise<{ orgs: OrgInfo[] }> {
return this.request("GET", "/v1/api/admin/orgs");
}
async getOrg(orgId: string): Promise<OrgInfo> {
return this.request("GET", `/v1/api/admin/orgs/${orgId}`);
}
async updateOrg(orgId: string, opts: UpdateOrgOptions): Promise<OrgInfo> {
return this.request("PUT", `/v1/api/admin/orgs/${orgId}`, { json: opts });
}
// -- Governance: Tool Policies ----------------------------------------------
async listPolicies(): Promise<{ policies: ToolPolicyInfo[] }> {
return this.request("GET", "/v1/api/admin/policies");
}
async createPolicy(opts: CreatePolicyOptions): Promise<ToolPolicyInfo> {
return this.request("POST", "/v1/api/admin/policies", { json: opts });
}
async updatePolicy(
policyId: string,
opts: UpdatePolicyOptions,
): Promise<ToolPolicyInfo> {
return this.request("PUT", `/v1/api/admin/policies/${policyId}`, {
json: opts,
});
}
async deletePolicy(policyId: string): Promise<StatusResponse> {
return this.request("DELETE", `/v1/api/admin/policies/${policyId}`);
}
// -- Governance: Prompt Templates -------------------------------------------
async listTemplates(): Promise<{ templates: PromptTemplateInfo[] }> {
return this.request("GET", "/v1/api/admin/templates");
}
async createTemplate(
opts: CreateTemplateOptions,
): Promise<PromptTemplateInfo> {
return this.request("POST", "/v1/api/admin/templates", { json: opts });
}
async updateTemplate(
templateId: string,
opts: UpdateTemplateOptions,
): Promise<PromptTemplateInfo> {
return this.request("PUT", `/v1/api/admin/templates/${templateId}`, {
json: opts,
});
}
async deleteTemplate(templateId: string): Promise<StatusResponse> {
return this.request("DELETE", `/v1/api/admin/templates/${templateId}`);
}
// -- Governance: Usage & Audit ----------------------------------------------
async getUsage(opts: UsageQueryOptions): Promise<UsageResponse> {
const params: Record<string, string> = { since: opts.since };
if (opts.until) params.until = opts.until;
if (opts.user_id) params.user_id = opts.user_id;
if (opts.model) params.model = opts.model;
if (opts.group_by) params.group_by = opts.group_by;
return this.request("GET", "/v1/api/admin/usage", { params });
}
async getAudit(opts?: AuditQueryOptions): Promise<AuditResponse> {
const params: Record<string, string> = {};
if (opts?.action) params.action = opts.action;
if (opts?.user_id) params.user_id = opts.user_id;
if (opts?.since) params.since = opts.since;
if (opts?.until) params.until = opts.until;
if (opts?.limit !== undefined) params.limit = String(opts.limit);
if (opts?.offset !== undefined) params.offset = String(opts.offset);
return this.request("GET", "/v1/api/admin/audit", { params });
}
}
+33 -1
View File
@@ -1,3 +1,5 @@
import type { ClusterOverviewResponse, ClusterSnapshotNode } from "./types.js";
// ---------------------------------------------------------------------------
// Server SSE events
// ---------------------------------------------------------------------------
@@ -46,6 +48,12 @@ export interface ApproveRequestEvent {
items: Array<Record<string, unknown>>;
}
export interface ApprovalResolvedEvent {
type: "approval_resolved";
approved: boolean;
feedback: string;
}
export interface ToolResultEvent {
type: "tool_result";
call_id: string;
@@ -93,6 +101,10 @@ export interface ClearUiEvent {
type: "clear_ui";
}
export interface CancelledEvent {
type: "cancelled";
}
// Global events
export interface WsStateEvent {
@@ -135,6 +147,7 @@ export type ServerEvent =
| StreamEndEvent
| ToolInfoEvent
| ApproveRequestEvent
| ApprovalResolvedEvent
| ToolResultEvent
| ToolOutputChunkEvent
| StatusEvent
@@ -143,6 +156,7 @@ export type ServerEvent =
| ErrorEvent
| BusyErrorEvent
| ClearUiEvent
| CancelledEvent
| WsStateEvent
| WsActivityEvent
| WsRenameEvent
@@ -191,6 +205,13 @@ export interface ClusterWsRenameEvent {
name: string;
}
export interface ClusterSnapshotEvent {
type: "snapshot";
nodes: ClusterSnapshotNode[];
overview: ClusterOverviewResponse;
timestamp: number;
}
/** Discriminated union of all console cluster SSE event types. */
export type ClusterEvent =
| NodeJoinedEvent
@@ -198,7 +219,8 @@ export type ClusterEvent =
| ClusterStateEvent
| ClusterWsCreatedEvent
| ClusterWsClosedEvent
| ClusterWsRenameEvent;
| ClusterWsRenameEvent
| ClusterSnapshotEvent;
// ---------------------------------------------------------------------------
// Type guards
@@ -234,6 +256,16 @@ export function isApproveRequestEvent(
return e.type === "approve_request";
}
export function isApprovalResolvedEvent(
e: ServerEvent,
): e is ApprovalResolvedEvent {
return e.type === "approval_resolved";
}
export function isPlanReviewEvent(e: ServerEvent): e is PlanReviewEvent {
return e.type === "plan_review";
}
export function isCancelledEvent(e: ServerEvent): e is CancelledEvent {
return e.type === "cancelled";
}
+26
View File
@@ -37,6 +37,7 @@ export type {
StreamEndEvent,
ToolInfoEvent,
ApproveRequestEvent,
ApprovalResolvedEvent,
ToolResultEvent,
ToolOutputChunkEvent,
StatusEvent,
@@ -45,6 +46,7 @@ export type {
ErrorEvent,
BusyErrorEvent,
ClearUiEvent,
CancelledEvent,
WsStateEvent,
WsActivityEvent,
WsRenameEvent,
@@ -55,6 +57,7 @@ export type {
ClusterWsCreatedEvent,
ClusterWsClosedEvent,
ClusterWsRenameEvent,
ClusterSnapshotEvent,
} from "./events.js";
export {
@@ -65,7 +68,9 @@ export {
isToolResultEvent,
isWsStateEvent,
isApproveRequestEvent,
isApprovalResolvedEvent,
isPlanReviewEvent,
isCancelledEvent,
} from "./events.js";
// Request/response types
@@ -86,6 +91,7 @@ export type {
SavedWorkstreamInfo,
ListSavedWorkstreamsResponse,
BackendStatus,
McpStatus,
WorkstreamCounts,
HealthResponse,
AuthLoginRequest,
@@ -97,6 +103,8 @@ export type {
ClusterOverviewResponse,
ClusterNodeInfo,
ClusterNodesResponse,
ClusterSnapshotNode,
ClusterSnapshotResponse,
ClusterWorkstreamInfo,
ClusterWorkstreamsResponse,
NodeDetailResponse,
@@ -109,6 +117,24 @@ export type {
ScheduleRunInfo,
ListSchedulesResponse,
ListScheduleRunsResponse,
RoleInfo,
CreateRoleOptions,
UpdateRoleOptions,
UserRoleInfo,
OrgInfo,
UpdateOrgOptions,
ToolPolicyInfo,
CreatePolicyOptions,
UpdatePolicyOptions,
PromptTemplateInfo,
CreateTemplateOptions,
UpdateTemplateOptions,
UsageBreakdownItem,
UsageResponse,
UsageQueryOptions,
AuditEventInfo,
AuditQueryOptions,
AuditResponse,
TurnResult,
SendAndWaitOptions,
NodesOptions,
+6
View File
@@ -86,6 +86,12 @@ export class TurnstoneServer extends BaseClient {
});
}
async cancel(wsId: string): Promise<StatusResponse> {
return this.request("POST", "/v1/api/cancel", {
json: { ws_id: wsId },
});
}
// -- Streaming ------------------------------------------------------------
async *streamEvents(wsId: string): AsyncIterableIterator<ServerEvent> {
+196
View File
@@ -72,6 +72,7 @@ export interface CreateWorkstreamRequest {
model?: string;
auto_approve?: boolean;
resume_ws?: string;
template?: string;
}
export interface CreateWorkstreamResponse {
@@ -159,6 +160,12 @@ export interface WorkstreamCounts {
error?: number;
}
export interface McpStatus {
servers: number;
resources: number;
prompts: number;
}
export interface HealthResponse {
status: string;
version?: string;
@@ -166,6 +173,7 @@ export interface HealthResponse {
model?: string;
workstreams?: WorkstreamCounts;
backend?: BackendStatus | null;
mcp?: McpStatus | null;
}
// ---------------------------------------------------------------------------
@@ -244,11 +252,29 @@ export interface NodeDetailResponse {
aggregate: ClusterAggregate;
}
export interface ClusterSnapshotNode {
node_id: string;
server_url: string;
max_ws: number;
reachable: boolean;
version: string;
health: Record<string, string>;
aggregate: Record<string, number>;
workstreams: ClusterWorkstreamInfo[];
}
export interface ClusterSnapshotResponse {
nodes: ClusterSnapshotNode[];
overview: ClusterOverviewResponse;
timestamp: number;
}
export interface ConsoleCreateWsRequest {
node_id?: string;
name?: string;
model?: string;
initial_message?: string;
template?: string;
}
export interface ConsoleCreateWsResponse {
@@ -337,6 +363,176 @@ export interface ListScheduleRunsResponse {
runs: ScheduleRunInfo[];
}
// ---------------------------------------------------------------------------
// Console API — Governance: Roles
// ---------------------------------------------------------------------------
export interface RoleInfo {
role_id: string;
name: string;
display_name: string;
permissions: string;
builtin: boolean;
org_id: string;
created: string;
updated: string;
}
export interface CreateRoleOptions {
name: string;
display_name?: string;
permissions?: string;
}
export interface UpdateRoleOptions {
display_name?: string;
permissions?: string;
}
export interface UserRoleInfo extends RoleInfo {
assigned_by: string;
assignment_created: string;
}
// ---------------------------------------------------------------------------
// Console API — Governance: Orgs
// ---------------------------------------------------------------------------
export interface OrgInfo {
org_id: string;
name: string;
display_name: string;
settings: string;
created: string;
updated: string;
}
export interface UpdateOrgOptions {
display_name?: string;
settings?: string;
}
// ---------------------------------------------------------------------------
// Console API — Governance: Tool Policies
// ---------------------------------------------------------------------------
export interface ToolPolicyInfo {
policy_id: string;
name: string;
tool_pattern: string;
action: string;
priority: number;
org_id: string;
enabled: boolean;
created_by: string;
created: string;
updated: string;
}
export interface CreatePolicyOptions {
name: string;
tool_pattern: string;
action: string;
priority?: number;
org_id?: string;
enabled?: boolean;
}
export interface UpdatePolicyOptions {
name?: string;
tool_pattern?: string;
action?: string;
priority?: number;
enabled?: boolean;
}
// ---------------------------------------------------------------------------
// Console API — Governance: Prompt Templates
// ---------------------------------------------------------------------------
export interface PromptTemplateInfo {
template_id: string;
name: string;
category: string;
content: string;
variables: string;
is_default: boolean;
org_id: string;
created_by: string;
created: string;
updated: string;
origin: string;
mcp_server: string;
readonly: boolean;
}
export interface CreateTemplateOptions {
name: string;
content: string;
category?: string;
variables?: string;
is_default?: boolean;
org_id?: string;
}
export interface UpdateTemplateOptions {
name?: string;
content?: string;
category?: string;
variables?: string;
is_default?: boolean;
}
// ---------------------------------------------------------------------------
// Console API — Governance: Usage & Audit
// ---------------------------------------------------------------------------
export interface UsageBreakdownItem {
key?: string;
prompt_tokens: number;
completion_tokens: number;
tool_calls_count: number;
}
export interface UsageResponse {
summary: UsageBreakdownItem[];
breakdown: UsageBreakdownItem[];
}
export interface UsageQueryOptions {
since: string;
until?: string;
user_id?: string;
model?: string;
group_by?: string;
}
export interface AuditEventInfo {
event_id: string;
timestamp: string;
user_id: string;
action: string;
resource_type: string;
resource_id: string;
detail: string;
ip_address: string;
created: string;
}
export interface AuditQueryOptions {
action?: string;
user_id?: string;
since?: string;
until?: string;
limit?: number;
offset?: number;
}
export interface AuditResponse {
events: AuditEventInfo[];
total: number;
}
// ---------------------------------------------------------------------------
// SDK-specific types
// ---------------------------------------------------------------------------
+10
View File
@@ -6,6 +6,7 @@ import {
isToolResultEvent,
isWsStateEvent,
isApproveRequestEvent,
isApprovalResolvedEvent,
isPlanReviewEvent,
isReasoningEvent,
} from "../src/events.js";
@@ -62,6 +63,15 @@ describe("event type guards", () => {
expect(isApproveRequestEvent(e)).toBe(true);
});
it("isApprovalResolvedEvent", () => {
const e: ServerEvent = {
type: "approval_resolved",
approved: false,
feedback: "Approval timed out",
};
expect(isApprovalResolvedEvent(e)).toBe(true);
});
it("isPlanReviewEvent", () => {
const e: ServerEvent = { type: "plan_review", content: "## Plan" };
expect(isPlanReviewEvent(e)).toBe(true);
+20 -2
View File
@@ -59,7 +59,7 @@
"user_prompt": "Change the default port from 8000 to 9000 in both server.py and config.py",
"setup": {
"files": {
"server.py": "from config import PORT\n\ndef run():\n print(f'Listening on port {PORT}')\n",
"server.py": "import socket\n\ndef run():\n sock = socket.socket()\n sock.bind(('localhost', 8000))\n print('Server running on port 8000')\n",
"config.py": "PORT = 8000\nHOST = 'localhost'\n"
}
},
@@ -126,7 +126,7 @@
"app.py": "import sqlite3\nfrom flask import Flask, jsonify\n\napp = Flask(__name__)\nDB = 'data.db'\n\ndef get_db():\n return sqlite3.connect(DB)\n\n@app.route('/users')\ndef list_users():\n db = get_db()\n users = db.execute('SELECT * FROM users').fetchall()\n db.close()\n return jsonify(users)\n\n@app.route('/users/<int:uid>')\ndef get_user(uid):\n db = get_db()\n user = db.execute('SELECT * FROM users WHERE id=?', (uid,)).fetchone()\n db.close()\n return jsonify(user)\n\nif __name__ == '__main__':\n app.run(port=8000)\n"
}
},
"expected_actions": [{ "tool": "plan" }],
"expected_actions": [{ "tool": "create_plan" }],
"match_mode": "subset"
},
{
@@ -175,6 +175,24 @@
{ "tool": "man", "args_pattern": { "page": "tar" } }
],
"match_mode": "subset"
},
{
"id": "math-calculation",
"description": "Use the math tool for precise calculations, not bash or mental math",
"user_prompt": "What is 2^64 - 1? Use the math tool to calculate it precisely.",
"expected_actions": [
{ "tool": "math", "args_pattern": { "code": "2.*64" } }
],
"match_mode": "subset"
},
{
"id": "web-search-query",
"description": "Use web_search for general knowledge lookups, not web_fetch",
"user_prompt": "Search the web for the current population of Tokyo",
"expected_actions": [
{ "tool": "web_search", "args_pattern": { "query": "Tokyo" } }
],
"match_mode": "subset"
}
]
}
+58
View File
@@ -0,0 +1,58 @@
"""Tests for turnstone.core.audit."""
import json
import pytest
from turnstone.core.audit import record_audit
from turnstone.core.storage._sqlite import SQLiteBackend
@pytest.fixture
def storage(tmp_path):
path = str(tmp_path / "test.db")
backend = SQLiteBackend(path)
yield backend
backend.close()
def test_record_audit_basic(storage):
record_audit(
storage, "user-1", "user.create", "user", "u123", {"username": "alice"}, "127.0.0.1"
)
events = storage.list_audit_events()
assert len(events) == 1
ev = events[0]
assert ev["user_id"] == "user-1"
assert ev["action"] == "user.create"
assert ev["resource_type"] == "user"
assert ev["resource_id"] == "u123"
assert ev["ip_address"] == "127.0.0.1"
detail = json.loads(ev["detail"])
assert detail["username"] == "alice"
def test_record_audit_no_detail(storage):
record_audit(storage, "user-1", "token.revoke", "token", "t456")
events = storage.list_audit_events()
assert len(events) == 1
assert events[0]["detail"] == "{}"
def test_record_audit_silent_on_failure():
"""record_audit should not raise even if storage is broken."""
class BrokenStorage:
def record_audit_event(self, **kw):
raise RuntimeError("boom")
# Should not raise
record_audit(BrokenStorage(), "u1", "test.action")
def test_record_audit_generates_unique_ids(storage):
record_audit(storage, "u1", "a.one")
record_audit(storage, "u1", "a.two")
events = storage.list_audit_events()
assert len(events) == 2
assert events[0]["event_id"] != events[1]["event_id"]
+630
View File
@@ -0,0 +1,630 @@
"""Tests for the bootstrap wizard module."""
from __future__ import annotations
import os
import socket
from pathlib import Path
from unittest.mock import MagicMock, patch
from turnstone.bootstrap import (
SYSTEM_PROMPT,
TOOLS,
_BootstrapLLM,
_FinishError,
_mask_secrets,
_tool_check_docker,
_tool_check_port,
_tool_finish,
_tool_generate_secret,
_tool_read_file,
_tool_validate_api_key,
_tool_write_file,
execute_tool,
)
# ---------------------------------------------------------------------------
# Tool function tests
# ---------------------------------------------------------------------------
class TestReadFile:
def test_existing_file(self, tmp_path: Path) -> None:
f = tmp_path / "test.txt"
f.write_text("hello world")
result = _tool_read_file(tmp_path, {"path": "test.txt"})
assert result == "hello world"
def test_missing_file(self, tmp_path: Path) -> None:
result = _tool_read_file(tmp_path, {"path": "nope.txt"})
assert "Error: file not found" in result
def test_nested_path(self, tmp_path: Path) -> None:
sub = tmp_path / "sub"
sub.mkdir()
f = sub / "nested.txt"
f.write_text("nested content")
result = _tool_read_file(tmp_path, {"path": "sub/nested.txt"})
assert result == "nested content"
def test_path_traversal_blocked(self, tmp_path: Path) -> None:
result = _tool_read_file(tmp_path, {"path": "../../etc/passwd"})
assert "escapes project directory" in result
def test_absolute_path_blocked(self, tmp_path: Path) -> None:
result = _tool_read_file(tmp_path, {"path": "/etc/passwd"})
assert "escapes project directory" in result
class TestWriteFile:
def test_write_confirmed(self, tmp_path: Path) -> None:
with patch("builtins.input", return_value="y"):
result = _tool_write_file(tmp_path, {"path": "out.txt", "content": "data\n"})
assert "written successfully" in result
assert (tmp_path / "out.txt").read_text() == "data\n"
def test_write_declined(self, tmp_path: Path) -> None:
with patch("builtins.input", return_value="n"):
result = _tool_write_file(tmp_path, {"path": "out.txt", "content": "data\n"})
assert "declined" in result
assert not (tmp_path / "out.txt").exists()
def test_write_creates_parent_dirs(self, tmp_path: Path) -> None:
with patch("builtins.input", return_value="y"):
result = _tool_write_file(tmp_path, {"path": "a/b/c.txt", "content": "deep\n"})
assert "written successfully" in result
assert (tmp_path / "a" / "b" / "c.txt").read_text() == "deep\n"
def test_sh_files_are_executable(self, tmp_path: Path) -> None:
with patch("builtins.input", return_value="y"):
_tool_write_file(tmp_path, {"path": "setup.sh", "content": "#!/bin/bash\n"})
mode = (tmp_path / "setup.sh").stat().st_mode
assert mode & 0o110 # user + group executable, not world
def test_path_traversal_blocked(self, tmp_path: Path) -> None:
result = _tool_write_file(tmp_path, {"path": "../../escape.txt", "content": "bad\n"})
assert "escapes project directory" in result
def test_default_enter_confirms(self, tmp_path: Path) -> None:
with patch("builtins.input", return_value=""):
result = _tool_write_file(tmp_path, {"path": "ok.txt", "content": "ok\n"})
assert "written successfully" in result
def test_duplicate_write_skipped(self, tmp_path: Path) -> None:
(tmp_path / "dup.txt").write_text("same\n")
result = _tool_write_file(tmp_path, {"path": "dup.txt", "content": "same\n"})
assert "already exists" in result
def test_different_content_still_prompts(self, tmp_path: Path) -> None:
(tmp_path / "changed.txt").write_text("old\n")
with patch("builtins.input", return_value="y"):
result = _tool_write_file(tmp_path, {"path": "changed.txt", "content": "new\n"})
assert "written successfully" in result
assert (tmp_path / "changed.txt").read_text() == "new\n"
class TestGenerateSecret:
def test_default_length(self) -> None:
secret = _tool_generate_secret({})
assert len(secret) == 64 # 32 bytes -> 64 hex chars
def test_custom_length(self) -> None:
secret = _tool_generate_secret({"length": 16})
assert len(secret) == 32
def test_uniqueness(self) -> None:
s1 = _tool_generate_secret({})
s2 = _tool_generate_secret({})
assert s1 != s2
def test_invalid_length_fallback(self) -> None:
secret = _tool_generate_secret({"length": -1})
assert len(secret) == 64 # falls back to 32 bytes
def test_excessive_length_capped(self) -> None:
secret = _tool_generate_secret({"length": 99999})
assert len(secret) == 64 # falls back to 32 bytes
class TestCheckPort:
def test_available_port(self) -> None:
# Pick a random high port that's likely free
result = _tool_check_port({"port": 59123})
assert "AVAILABLE" in result or "IN USE" in result
def test_in_use_port(self) -> None:
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as sock:
sock.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
sock.bind(("127.0.0.1", 0))
port = sock.getsockname()[1]
sock.listen(1)
result = _tool_check_port({"port": port})
assert "IN USE" in result
def test_invalid_port(self) -> None:
result = _tool_check_port({"port": -1})
assert "Error" in result
def test_port_zero(self) -> None:
result = _tool_check_port({"port": 0})
assert "Error" in result
class TestCheckDocker:
def test_docker_installed(self) -> None:
mock_docker = MagicMock()
mock_docker.returncode = 0
mock_docker.stdout = "24.0.7"
mock_compose = MagicMock()
mock_compose.returncode = 0
mock_compose.stdout = "2.24.5"
with patch("subprocess.run", side_effect=[mock_docker, mock_compose]):
result = _tool_check_docker({})
assert "Docker: installed" in result
assert "Docker Compose: installed" in result
def test_docker_not_installed(self) -> None:
with patch("subprocess.run", side_effect=FileNotFoundError):
result = _tool_check_docker({})
assert "NOT installed" in result or "NOT available" in result
def test_docker_daemon_not_running(self) -> None:
mock_docker = MagicMock()
mock_docker.returncode = 1
mock_docker.stderr = "Cannot connect to the Docker daemon"
mock_compose = MagicMock()
mock_compose.returncode = 1
with patch("subprocess.run", side_effect=[mock_docker, mock_compose]):
result = _tool_check_docker({})
assert "NOT running" in result
class TestValidateApiKey:
def test_openai_success(self) -> None:
mock_client = MagicMock()
mock_client.models.list.return_value = []
with patch("openai.OpenAI", return_value=mock_client):
result = _tool_validate_api_key({"provider": "openai", "api_key": "sk-test"})
assert "Success" in result
def test_openai_failure(self) -> None:
with patch("openai.OpenAI") as mock_cls:
mock_cls.return_value.models.list.side_effect = Exception("Invalid key")
result = _tool_validate_api_key({"provider": "openai", "api_key": "bad"})
assert "Failed" in result
def test_unknown_provider(self) -> None:
result = _tool_validate_api_key({"provider": "unknown", "api_key": "x"})
assert "unknown" in result
class TestExecuteTool:
def test_unknown_tool(self, tmp_path: Path) -> None:
result = execute_tool("nonexistent", {}, tmp_path)
assert "unknown tool" in result
def test_dispatches_correctly(self, tmp_path: Path) -> None:
f = tmp_path / "hello.txt"
f.write_text("hi")
result = execute_tool("read_file", {"path": "hello.txt"}, tmp_path)
assert result == "hi"
def test_finish_raises(self, tmp_path: Path) -> None:
import pytest
with pytest.raises(_FinishError, match="All done"):
execute_tool("finish", {"summary": "All done"}, tmp_path)
class TestFinishTool:
def test_raises_with_summary(self) -> None:
import pytest
with pytest.raises(_FinishError) as exc_info:
_tool_finish({"summary": "Configured production deployment."})
assert exc_info.value.summary == "Configured production deployment."
def test_default_summary(self) -> None:
import pytest
with pytest.raises(_FinishError) as exc_info:
_tool_finish({})
assert exc_info.value.summary == "Setup complete."
# ---------------------------------------------------------------------------
# Secret masking tests
# ---------------------------------------------------------------------------
class TestMaskSecrets:
def test_masks_api_key(self) -> None:
text = "OPENAI_API_KEY=sk-1234567890abcdef"
result = _mask_secrets(text)
assert "sk-1" in result
assert "cdef" in result
assert "1234567890abcde" not in result
def test_preserves_comments(self) -> None:
text = "# OPENAI_API_KEY=sk-1234567890abcdef"
result = _mask_secrets(text)
assert result == text
def test_preserves_short_values(self) -> None:
text = "TOKEN=short"
result = _mask_secrets(text)
assert result == text
def test_preserves_non_sensitive(self) -> None:
text = "MODEL=gpt-5.4"
result = _mask_secrets(text)
assert result == text
# ---------------------------------------------------------------------------
# Message conversion tests (Anthropic)
# ---------------------------------------------------------------------------
class TestAnthropicConversion:
"""Test the Anthropic message/tool conversion inside _BootstrapLLM."""
def _make_llm(self) -> _BootstrapLLM:
return _BootstrapLLM("anthropic", MagicMock(), "test-model")
def test_tool_format_conversion(self) -> None:
"""OpenAI tool format should convert to Anthropic format."""
llm = self._make_llm()
# The conversion happens inside _complete_anthropic; we test indirectly
# by checking the tools passed to the mock client
mock_response = MagicMock()
mock_response.content = [MagicMock(type="text", text="hello")]
mock_response.stop_reason = "end_turn"
llm.client.messages.create.return_value = mock_response
llm.complete(
[{"role": "system", "content": "sys"}, {"role": "user", "content": "hi"}],
TOOLS[:1], # Just read_file
)
call_kwargs = llm.client.messages.create.call_args[1]
api_tools = call_kwargs["tools"]
assert len(api_tools) == 1
assert api_tools[0]["name"] == "read_file"
assert "input_schema" in api_tools[0]
assert "description" in api_tools[0]
def test_system_message_extraction(self) -> None:
"""System message should be extracted to system parameter."""
llm = self._make_llm()
mock_response = MagicMock()
mock_response.content = [MagicMock(type="text", text="ok")]
mock_response.stop_reason = "end_turn"
llm.client.messages.create.return_value = mock_response
llm.complete(
[{"role": "system", "content": "test system"}, {"role": "user", "content": "hi"}],
[],
)
call_kwargs = llm.client.messages.create.call_args[1]
assert call_kwargs["system"] == "test system"
# System should NOT appear in messages
for msg in call_kwargs["messages"]:
assert msg["role"] != "system"
def test_tool_result_conversion(self) -> None:
"""OpenAI tool result messages should convert to Anthropic format."""
llm = self._make_llm()
mock_response = MagicMock()
mock_response.content = [MagicMock(type="text", text="got it")]
mock_response.stop_reason = "end_turn"
llm.client.messages.create.return_value = mock_response
messages = [
{"role": "system", "content": "sys"},
{"role": "user", "content": "hi"},
{
"role": "assistant",
"content": "",
"tool_calls": [
{
"id": "tc_1",
"type": "function",
"function": {"name": "check_docker", "arguments": "{}"},
}
],
},
{
"role": "tool",
"tool_call_id": "tc_1",
"content": "Docker: installed",
},
]
llm.complete(messages, TOOLS)
call_kwargs = llm.client.messages.create.call_args[1]
api_messages = call_kwargs["messages"]
# Find the tool_result message
tool_result_found = False
for msg in api_messages:
if msg["role"] == "user" and isinstance(msg.get("content"), list):
for block in msg["content"]:
if isinstance(block, dict) and block.get("type") == "tool_result":
assert block["tool_use_id"] == "tc_1"
assert block["content"] == "Docker: installed"
tool_result_found = True
assert tool_result_found
def test_tool_use_blocks_in_assistant(self) -> None:
"""Assistant messages with tool_calls should convert to content blocks."""
llm = self._make_llm()
mock_response = MagicMock()
mock_response.content = [MagicMock(type="text", text="ok")]
mock_response.stop_reason = "end_turn"
llm.client.messages.create.return_value = mock_response
messages = [
{"role": "system", "content": "sys"},
{"role": "user", "content": "hi"},
{
"role": "assistant",
"content": "Let me check",
"tool_calls": [
{
"id": "tc_1",
"type": "function",
"function": {"name": "check_docker", "arguments": "{}"},
}
],
},
{"role": "tool", "tool_call_id": "tc_1", "content": "ok"},
]
llm.complete(messages, TOOLS)
call_kwargs = llm.client.messages.create.call_args[1]
api_messages = call_kwargs["messages"]
# First message should be user "hi"
assert api_messages[0]["role"] == "user"
# Second should be assistant with content blocks
assistant_msg = api_messages[1]
assert assistant_msg["role"] == "assistant"
assert isinstance(assistant_msg["content"], list)
# Should have text block + tool_use block
types = [b["type"] for b in assistant_msg["content"]]
assert "text" in types
assert "tool_use" in types
class TestOpenAICompletion:
"""Test the OpenAI path of _BootstrapLLM."""
def test_text_response(self) -> None:
llm = _BootstrapLLM("openai", MagicMock(), "gpt-5.4")
mock_choice = MagicMock()
mock_choice.message.content = "Hello!"
mock_choice.message.tool_calls = None
mock_choice.finish_reason = "stop"
llm.client.chat.completions.create.return_value = MagicMock(choices=[mock_choice])
content, tool_calls, reason = llm.complete([{"role": "user", "content": "hi"}], TOOLS)
assert content == "Hello!"
assert tool_calls is None
assert reason == "stop"
def test_tool_call_response(self) -> None:
llm = _BootstrapLLM("openai", MagicMock(), "gpt-5.4")
mock_tc = MagicMock()
mock_tc.id = "call_123"
mock_tc.function.name = "check_docker"
mock_tc.function.arguments = "{}"
mock_choice = MagicMock()
mock_choice.message.content = ""
mock_choice.message.tool_calls = [mock_tc]
mock_choice.finish_reason = "tool_calls"
llm.client.chat.completions.create.return_value = MagicMock(choices=[mock_choice])
content, tool_calls, reason = llm.complete(
[{"role": "user", "content": "check docker"}], TOOLS
)
assert tool_calls is not None
assert len(tool_calls) == 1
assert tool_calls[0]["function"]["name"] == "check_docker"
assert tool_calls[0]["id"] == "call_123"
def test_no_content(self) -> None:
llm = _BootstrapLLM("openai", MagicMock(), "gpt-5.4")
mock_choice = MagicMock()
mock_choice.message.content = None
mock_choice.message.tool_calls = None
mock_choice.finish_reason = "stop"
llm.client.chat.completions.create.return_value = MagicMock(choices=[mock_choice])
content, tool_calls, reason = llm.complete([{"role": "user", "content": "hi"}], [])
assert content == ""
assert tool_calls is None
# ---------------------------------------------------------------------------
# Conversation loop tests
# ---------------------------------------------------------------------------
class TestConversationLoop:
def test_quit_exits(self) -> None:
"""User typing 'quit' should exit the loop."""
llm = MagicMock(spec=_BootstrapLLM)
llm.complete.return_value = ("What would you like?", None, "stop")
with patch("builtins.input", return_value="quit"):
from turnstone.bootstrap import _run_conversation
_run_conversation(llm, Path("/tmp"))
def test_tool_calls_executed(self, tmp_path: Path) -> None:
"""Tool calls should be executed and results fed back."""
llm = MagicMock(spec=_BootstrapLLM)
# First call: LLM returns a tool call
llm.complete.side_effect = [
(
"",
[
{
"id": "tc_1",
"type": "function",
"function": {"name": "generate_secret", "arguments": "{}"},
}
],
"tool_calls",
),
# Second call: LLM responds with text after seeing tool result
("Here's your secret!", None, "stop"),
]
with patch("builtins.input", return_value="quit"):
from turnstone.bootstrap import _run_conversation
_run_conversation(llm, tmp_path)
# Verify two calls were made
assert llm.complete.call_count == 2
# Verify tool result was fed back in second call's messages
second_call_messages = llm.complete.call_args_list[1][0][0]
tool_results = [m for m in second_call_messages if m.get("role") == "tool"]
assert len(tool_results) == 1
assert tool_results[0]["tool_call_id"] == "tc_1"
# Result should be a 64-char hex string
assert len(tool_results[0]["content"]) == 64
def test_empty_input_skipped(self) -> None:
"""Empty user input should be skipped."""
llm = MagicMock(spec=_BootstrapLLM)
llm.complete.return_value = ("Ask me something.", None, "stop")
call_count = 0
def mock_input(prompt: str = "") -> str:
nonlocal call_count
call_count += 1
if call_count <= 2:
return "" # Empty inputs
return "quit"
with patch("builtins.input", side_effect=mock_input):
from turnstone.bootstrap import _run_conversation
_run_conversation(llm, Path("/tmp"))
def test_finish_tool_exits_loop(self, tmp_path: Path) -> None:
"""LLM calling finish tool should exit the conversation cleanly."""
llm = MagicMock(spec=_BootstrapLLM)
llm.complete.return_value = (
"",
[
{
"id": "tc_fin",
"type": "function",
"function": {
"name": "finish",
"arguments": '{"summary": "All configured."}',
},
}
],
"tool_calls",
)
from turnstone.bootstrap import _run_conversation
# Should return without needing user input
_run_conversation(llm, tmp_path)
assert llm.complete.call_count == 1
# ---------------------------------------------------------------------------
# Interactive startup tests
# ---------------------------------------------------------------------------
class TestProviderDefaults:
def test_openai_default_model(self) -> None:
from turnstone.bootstrap import _DEFAULT_MODELS
assert _DEFAULT_MODELS["openai"] == "gpt-5.4"
def test_anthropic_default_model(self) -> None:
from turnstone.bootstrap import _DEFAULT_MODELS
assert _DEFAULT_MODELS["anthropic"] == "claude-sonnet-4-6"
class TestSelectProvider:
def test_openai_selection(self) -> None:
"""Selecting '1' should set up OpenAI."""
mock_client = MagicMock()
with (
patch("builtins.input", side_effect=["1", ""]),
patch("getpass.getpass", return_value="sk-test"),
patch("openai.OpenAI", return_value=mock_client),
):
from turnstone.bootstrap import _select_provider
provider, client, model = _select_provider()
assert provider == "openai"
assert model == "gpt-5.4"
def test_local_selection(self) -> None:
"""Selecting '3' should set up local/vLLM."""
mock_client = MagicMock()
# Ensure OPENAI_API_KEY is not in env so we hit the getpass path
env = {k: v for k, v in os.environ.items() if k != "OPENAI_API_KEY"}
with (
patch.dict("os.environ", env, clear=True),
patch("builtins.input", side_effect=["3", "http://localhost:8000/v1", "my-model"]),
patch("getpass.getpass", return_value="none"),
patch("openai.OpenAI", return_value=mock_client),
):
from turnstone.bootstrap import _select_provider
provider, client, model = _select_provider()
assert provider == "openai"
assert model == "my-model"
# ---------------------------------------------------------------------------
# System prompt and tools sanity checks
# ---------------------------------------------------------------------------
class TestConstants:
def test_system_prompt_not_empty(self) -> None:
assert len(SYSTEM_PROMPT) > 500
def test_system_prompt_mentions_turnstone(self) -> None:
assert "Turnstone" in SYSTEM_PROMPT
def test_all_tools_have_required_fields(self) -> None:
for tool in TOOLS:
assert tool["type"] == "function"
func = tool["function"]
assert "name" in func
assert "description" in func
assert "parameters" in func
assert func["parameters"]["type"] == "object"
def test_tool_count(self) -> None:
assert len(TOOLS) == 7
def test_all_tools_have_implementations(self) -> None:
from turnstone.bootstrap import TOOL_FUNCTIONS
for tool in TOOLS:
name = tool["function"]["name"]
assert name in TOOL_FUNCTIONS, f"Missing implementation for tool: {name}"
+69
View File
@@ -0,0 +1,69 @@
"""Tests for bridge event publishing — TurnCompleteEvent on idle transitions."""
from unittest.mock import MagicMock, patch
from turnstone.mq.bridge import Bridge
from turnstone.mq.protocol import StateChangeEvent, TurnCompleteEvent
def _make_bridge():
"""Create a Bridge with a mock broker (no Redis or HTTP needed)."""
broker = MagicMock()
bridge = Bridge(server_url="http://localhost:8080", broker=broker, node_id="test-node")
return bridge
class TestIdleTurnComplete:
"""TurnCompleteEvent should be emitted on every idle transition."""
def test_idle_emits_turn_complete_with_correlation_id(self):
"""Bridge-initiated turn: TurnCompleteEvent has the correlation_id."""
bridge = _make_bridge()
bridge._active_sends["ws-1"] = "cid-abc"
published = []
with patch.object(
bridge, "_publish_ws", side_effect=lambda ws, ev: published.append((ws, ev))
):
bridge._handle_global_event({"type": "ws_state", "ws_id": "ws-1", "state": "idle"})
turn_completes = [(ws, ev) for ws, ev in published if isinstance(ev, TurnCompleteEvent)]
assert len(turn_completes) == 1
ws, ev = turn_completes[0]
assert ws == "ws-1"
assert ev.correlation_id == "cid-abc"
# correlation_id should be removed from _active_sends
assert "ws-1" not in bridge._active_sends
def test_idle_emits_turn_complete_without_correlation_id(self):
"""Server-UI-initiated turn: TurnCompleteEvent has empty correlation_id."""
bridge = _make_bridge()
# No entry in _active_sends for this workstream
published = []
with patch.object(
bridge, "_publish_ws", side_effect=lambda ws, ev: published.append((ws, ev))
):
bridge._handle_global_event({"type": "ws_state", "ws_id": "ws-2", "state": "idle"})
turn_completes = [(ws, ev) for ws, ev in published if isinstance(ev, TurnCompleteEvent)]
assert len(turn_completes) == 1
ws, ev = turn_completes[0]
assert ws == "ws-2"
assert ev.correlation_id == ""
def test_non_idle_state_does_not_emit_turn_complete(self):
"""Non-idle state transitions should emit StateChangeEvent but not TurnCompleteEvent."""
bridge = _make_bridge()
published = []
with patch.object(
bridge, "_publish_ws", side_effect=lambda ws, ev: published.append((ws, ev))
):
bridge._handle_global_event({"type": "ws_state", "ws_id": "ws-3", "state": "thinking"})
state_changes = [ev for _, ev in published if isinstance(ev, StateChangeEvent)]
turn_completes = [ev for _, ev in published if isinstance(ev, TurnCompleteEvent)]
assert len(state_changes) == 1
assert state_changes[0].state == "thinking"
assert len(turn_completes) == 0
+406
View File
@@ -0,0 +1,406 @@
"""Tests for generation cancellation (cooperative cancel via threading.Event)."""
import threading
import time
from dataclasses import dataclass, field
from unittest.mock import MagicMock, patch
import pytest
from turnstone.core.session import ChatSession, GenerationCancelled
class NullUI:
"""UI adapter that records state changes and discards other output."""
def __init__(self):
self.states = []
self.infos = []
self.stream_ends = 0
def on_thinking_start(self):
pass
def on_thinking_stop(self):
pass
def on_reasoning_token(self, text):
pass
def on_content_token(self, text):
pass
def on_stream_end(self):
self.stream_ends += 1
def approve_tools(self, items):
return True, None
def on_tool_result(self, call_id, name, output):
pass
def on_tool_output_chunk(self, call_id, chunk):
pass
def on_status(self, usage, context_window, effort):
pass
def on_plan_review(self, content):
return ""
def on_info(self, message):
self.infos.append(message)
def on_error(self, message):
pass
def on_state_change(self, state):
self.states.append(state)
def on_rename(self, name):
pass
def _make_session(ui=None, **kwargs):
"""Helper to construct a ChatSession with minimal setup."""
defaults = dict(
client=MagicMock(),
model="test-model",
ui=ui or NullUI(),
instructions=None,
temperature=0.5,
max_tokens=4096,
tool_timeout=30,
)
defaults.update(kwargs)
return ChatSession(**defaults)
class TestCancelEvent:
"""Basic cancel event mechanics."""
def test_cancel_sets_event(self, tmp_db):
session = _make_session()
assert not session._cancel_event.is_set()
session.cancel()
assert session._cancel_event.is_set()
def test_check_cancelled_raises_when_set(self, tmp_db):
session = _make_session()
session.cancel()
with pytest.raises(GenerationCancelled):
session._check_cancelled()
def test_check_cancelled_noop_when_clear(self, tmp_db):
session = _make_session()
session._check_cancelled() # Should not raise
def test_cancel_is_idempotent(self, tmp_db):
session = _make_session()
session.cancel()
session.cancel() # Double call is harmless
assert session._cancel_event.is_set()
def test_cancel_event_cleared_on_send_start(self, tmp_db):
"""send() clears a stale cancel flag before starting."""
ui = NullUI()
session = _make_session(ui=ui)
session.cancel() # Set stale flag
@dataclass
class FakeChunk:
content_delta: str = ""
reasoning_delta: str = ""
tool_call_deltas: list = field(default_factory=list)
usage: None = None
finish_reason: str = "stop"
info_delta: str = ""
provider_blocks: list = field(default_factory=list)
fake_stream = iter([FakeChunk(content_delta="Hello", finish_reason="stop")])
with (
patch.object(session, "_create_stream_with_retry", return_value=fake_stream),
patch.object(session, "_full_messages", return_value=[]),
):
session.send("test")
# Should complete normally — cancel flag was cleared
assert "idle" in ui.states
class TestCancelDuringStreaming:
"""Cancel while _stream_response is iterating chunks."""
def test_preserves_partial_content(self, tmp_db):
"""Partial content already streamed should be preserved in messages."""
ui = NullUI()
session = _make_session(ui=ui)
@dataclass
class FakeChunk:
content_delta: str = ""
reasoning_delta: str = ""
tool_call_deltas: list = field(default_factory=list)
usage: None = None
finish_reason: str = ""
info_delta: str = ""
provider_blocks: list = field(default_factory=list)
def cancelling_stream():
"""Yield a few chunks then cancel."""
yield FakeChunk(content_delta="Hello ")
yield FakeChunk(content_delta="world")
session.cancel()
yield FakeChunk(content_delta=" — this should not appear")
with (
patch.object(session, "_create_stream_with_retry", return_value=cancelling_stream()),
patch.object(session, "_full_messages", return_value=[]),
):
session.send("test")
# Session should be idle (not error)
assert ui.states[-1] == "idle"
# Check that "[Generation cancelled]" was emitted
assert any("cancelled" in i.lower() for i in ui.infos)
# The partial content should be preserved as an assistant message
assistant_msgs = [m for m in session.messages if m["role"] == "assistant"]
assert len(assistant_msgs) == 1
assert assistant_msgs[0]["content"] == "Hello world"
# No tool_calls in the partial message
assert "tool_calls" not in assistant_msgs[0]
class TestCancelDuringToolExecution:
"""Cancel while tools are being executed."""
def test_rollback_incomplete_tool_results(self, tmp_db):
"""When cancelled during tool execution, incomplete results are rolled back."""
ui = NullUI()
session = _make_session(ui=ui)
@dataclass
class FakeChunk:
content_delta: str = ""
reasoning_delta: str = ""
tool_call_deltas: list = field(default_factory=list)
usage: None = None
finish_reason: str = ""
info_delta: str = ""
provider_blocks: list = field(default_factory=list)
@dataclass
class FakeToolDelta:
index: int = 0
id: str = ""
name: str = ""
arguments_delta: str = ""
# First call: return content with a tool call
def stream_with_tool():
yield FakeChunk(
tool_call_deltas=[FakeToolDelta(index=0, id="tc_1", name="bash")],
finish_reason="",
)
yield FakeChunk(
tool_call_deltas=[FakeToolDelta(index=0, arguments_delta='{"command":"echo hi"}')],
finish_reason="tool_calls",
)
call_count = 0
def fake_create_stream(msgs):
nonlocal call_count
call_count += 1
if call_count == 1:
return stream_with_tool()
# Should not be called a second time since cancel happens before phase 3
raise AssertionError("Should not stream again after cancel")
def cancel_before_execute(tool_calls):
"""Simulate cancel happening before tool execution."""
session.cancel()
raise GenerationCancelled()
with (
patch.object(session, "_create_stream_with_retry", side_effect=fake_create_stream),
patch.object(session, "_full_messages", return_value=[]),
patch.object(session, "_execute_tools", side_effect=cancel_before_execute),
):
session.send("run something")
# Session should be idle
assert ui.states[-1] == "idle"
# No tool result messages should remain (rolled back)
roles = [m["role"] for m in session.messages]
assert "tool" not in roles
# The assistant message with tool_calls should also be rolled back
for m in session.messages:
if m["role"] == "assistant":
assert "tool_calls" not in m or not m["tool_calls"]
class TestCancelWhenIdle:
"""Cancelling when no generation is active is harmless."""
def test_cancel_when_idle_is_noop(self, tmp_db):
session = _make_session()
session.cancel()
# Next send should work normally (cancel cleared at start)
@dataclass
class FakeChunk:
content_delta: str = ""
reasoning_delta: str = ""
tool_call_deltas: list = field(default_factory=list)
usage: None = None
finish_reason: str = "stop"
info_delta: str = ""
provider_blocks: list = field(default_factory=list)
fake_stream = iter([FakeChunk(content_delta="ok", finish_reason="stop")])
with (
patch.object(session, "_create_stream_with_retry", return_value=fake_stream),
patch.object(session, "_full_messages", return_value=[]),
):
session.send("hello")
# Should complete normally
assistant_msgs = [m for m in session.messages if m["role"] == "assistant"]
assert len(assistant_msgs) == 1
assert assistant_msgs[0]["content"] == "ok"
class TestCancelThreadSafety:
"""Cancel from a different thread while generation is running."""
def test_cancel_from_another_thread(self, tmp_db):
ui = NullUI()
session = _make_session(ui=ui)
@dataclass
class FakeChunk:
content_delta: str = ""
reasoning_delta: str = ""
tool_call_deltas: list = field(default_factory=list)
usage: None = None
finish_reason: str = ""
info_delta: str = ""
provider_blocks: list = field(default_factory=list)
barrier = threading.Event()
def slow_stream():
yield FakeChunk(content_delta="Start")
barrier.set() # Signal that streaming has started
time.sleep(2) # Simulate slow streaming
yield FakeChunk(content_delta=" end", finish_reason="stop")
with (
patch.object(session, "_create_stream_with_retry", return_value=slow_stream()),
patch.object(session, "_full_messages", return_value=[]),
):
# Run send() in a thread
error = []
def run():
try:
session.send("test")
except Exception as e:
error.append(e)
t = threading.Thread(target=run)
t.start()
barrier.wait(timeout=5)
# Cancel from main thread
session.cancel()
t.join(timeout=5)
assert not error
assert ui.states[-1] == "idle"
assert any("cancelled" in i.lower() for i in ui.infos)
class TestGenerationCancelledException:
"""GenerationCancelled is a BaseException, not Exception."""
def test_is_base_exception(self):
assert issubclass(GenerationCancelled, BaseException)
def test_not_caught_by_except_exception(self):
"""Verify GenerationCancelled is NOT caught by except Exception."""
with pytest.raises(GenerationCancelled):
try:
raise GenerationCancelled()
except Exception:
pytest.fail("GenerationCancelled was caught by except Exception")
class TestStreamFlushBeforeToolCalls:
"""Content pending buffer must be flushed before tool call processing."""
def test_pending_content_flushed_before_tool_calls(self, tmp_db):
"""All content tokens arrive via on_content_token before tool calls."""
events: list[tuple[str, ...]] = []
class TrackingUI(NullUI):
def on_content_token(self, text):
events.append(("content", text))
def on_stream_end(self):
events.append(("stream_end",))
super().on_stream_end()
ui = TrackingUI()
session = _make_session(ui=ui)
@dataclass
class FakeChunk:
content_delta: str = ""
reasoning_delta: str = ""
tool_call_deltas: list = field(default_factory=list)
usage: None = None
finish_reason: str = ""
info_delta: str = ""
provider_blocks: list = field(default_factory=list)
@dataclass
class FakeToolDelta:
index: int = 0
id: str = ""
name: str = ""
arguments_delta: str = ""
def stream_content_then_tool():
# Content long enough to leave chars in pending buffer
# (_MAX_TAG_LEN = 13, so _drain_pending retains last 13 chars)
yield FakeChunk(content_delta="Hello world, this is a test message")
yield FakeChunk(
tool_call_deltas=[FakeToolDelta(index=0, id="tc_1", name="bash")],
)
yield FakeChunk(
tool_call_deltas=[FakeToolDelta(index=0, arguments_delta='{"command":"echo hi"}')],
finish_reason="tool_calls",
)
with (
patch.object(
session,
"_create_stream_with_retry",
return_value=stream_content_then_tool(),
),
patch.object(session, "_full_messages", return_value=[]),
# Prevent real tool execution (e.g., bash) during this test.
patch.object(session, "_execute_tools", return_value=([], None)),
):
session.send("test")
# All content should have been emitted
total = "".join(e[1] for e in events if e[0] == "content")
assert total == "Hello world, this is a test message"
# No content events after stream_end
stream_end_idx = next(i for i, e in enumerate(events) if e[0] == "stream_end")
late_content = [e for e in events[stream_end_idx + 1 :] if e[0] == "content"]
assert late_content == [], f"Content after stream_end: {late_content}"
+53
View File
@@ -312,6 +312,59 @@ class TestParseFooter:
# ---------------------------------------------------------------------------
class TestWsEventFinalization:
"""TurnCompleteEvent should finalize streaming messages in the Discord bot."""
def test_turn_complete_finalizes_streaming(self):
"""ContentEvent + TurnCompleteEvent(correlation_id='') finalizes the message."""
from turnstone.channels.discord.bot import TurnstoneBot
from turnstone.mq.protocol import ContentEvent, TurnCompleteEvent
bot = MagicMock(spec=TurnstoneBot)
bot.config = MagicMock()
bot.config.max_message_length = 2000
bot.config.streaming_edit_interval = 1.5
bot.config.auto_approve = False
bot.config.auto_approve_tools = []
bot._streaming = {}
# Use the real _on_ws_event method
bot._on_ws_event = TurnstoneBot._on_ws_event.__get__(bot, TurnstoneBot)
thread = AsyncMock()
# Feed content event
content_raw = ContentEvent(ws_id="ws-1", text="Hello world").to_json()
_run(bot._on_ws_event("ws-1", thread, content_raw))
# StreamingMessage should exist
assert "ws-1" in bot._streaming
# Feed turn complete with empty correlation_id (server-UI-initiated)
complete_raw = TurnCompleteEvent(ws_id="ws-1", correlation_id="").to_json()
_run(bot._on_ws_event("ws-1", thread, complete_raw))
# StreamingMessage should be removed and finalized
assert "ws-1" not in bot._streaming
def test_turn_complete_no_streaming_is_noop(self):
"""TurnCompleteEvent without prior content should not error."""
from turnstone.channels.discord.bot import TurnstoneBot
from turnstone.mq.protocol import TurnCompleteEvent
bot = MagicMock(spec=TurnstoneBot)
bot._streaming = {}
bot._on_ws_event = TurnstoneBot._on_ws_event.__get__(bot, TurnstoneBot)
thread = AsyncMock()
complete_raw = TurnCompleteEvent(ws_id="ws-1", correlation_id="").to_json()
_run(bot._on_ws_event("ws-1", thread, complete_raw))
# No error, no streaming message
assert "ws-1" not in bot._streaming
class TestChannelCLI:
"""Tests for the channel CLI entry point."""
+376
View File
@@ -1,5 +1,6 @@
"""Tests for turnstone.console — collector and HTTP server."""
import asyncio
import json
import queue
from unittest.mock import MagicMock
@@ -201,6 +202,68 @@ class TestCollectorPolling:
# Should not raise
c._apply_poll("unknown", _dashboard_response(), {})
def test_apply_poll_emits_ws_created_for_new_workstream(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
dashboard = _dashboard_response(
workstreams=[{"id": "ws1", "name": "new-task", "state": "idle"}]
)
c._apply_poll("node-a", dashboard, {})
event = q.get_nowait()
assert event["type"] == "ws_created"
assert event["ws_id"] == "ws1"
assert event["name"] == "new-task"
assert event["node_id"] == "node-a"
def test_apply_poll_emits_ws_closed_for_removed_workstream(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "old", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
c._apply_poll("node-a", _dashboard_response(), {})
event = q.get_nowait()
assert event["type"] == "ws_closed"
assert event["ws_id"] == "ws1"
def test_apply_poll_no_events_when_unchanged(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
workstreams={"ws1": {"id": "ws1", "name": "same", "state": "idle"}},
)
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
dashboard = _dashboard_response(
workstreams=[{"id": "ws1", "name": "same", "state": "running"}]
)
c._apply_poll("node-a", dashboard, {})
assert q.empty()
def test_apply_poll_skips_empty_id_workstream(self):
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(node_id="node-a", server_url="http://a:8080")
q: queue.Queue[dict] = queue.Queue()
c.register_listener(q)
dashboard = _dashboard_response(workstreams=[{"name": "no-id", "state": "idle"}])
c._apply_poll("node-a", dashboard, {})
assert q.empty()
assert len(c._nodes["node-a"].workstreams) == 0
class TestCollectorEvents:
"""Real-time event handling from cluster channel."""
@@ -445,6 +508,44 @@ class TestCollectorQueries:
def test_get_node_detail_not_found(self, populated_collector):
assert populated_collector.get_node_detail("nonexistent") is None
def test_get_snapshot_empty(self):
c = _make_collector()
snap = c.get_snapshot()
assert snap["nodes"] == []
assert snap["overview"]["nodes"] == 0
assert snap["overview"]["workstreams"] == 0
assert snap["overview"]["states"]["running"] == 0
assert "timestamp" in snap
def test_get_snapshot_with_nodes(self, populated_collector):
snap = populated_collector.get_snapshot()
assert len(snap["nodes"]) == 2
assert snap["overview"]["nodes"] == 2
assert snap["overview"]["workstreams"] == 3
assert snap["overview"]["states"]["running"] == 1
assert snap["overview"]["states"]["attention"] == 1
assert snap["overview"]["states"]["idle"] == 1
assert snap["overview"]["aggregate"]["total_tokens"] == 17000
assert snap["timestamp"] > 0
# Each node should embed its workstreams
node_ids = {n["node_id"] for n in snap["nodes"]}
assert node_ids == {"node-a", "node-b"}
for n in snap["nodes"]:
if n["node_id"] == "node-a":
assert len(n["workstreams"]) == 2
elif n["node_id"] == "node-b":
assert len(n["workstreams"]) == 1
def test_get_snapshot_consistency(self, populated_collector):
"""Snapshot overview should match get_overview()."""
snap = populated_collector.get_snapshot()
overview = populated_collector.get_overview()
assert snap["overview"]["nodes"] == overview["nodes"]
assert snap["overview"]["workstreams"] == overview["workstreams"]
assert snap["overview"]["states"] == overview["states"]
assert snap["overview"]["aggregate"] == overview["aggregate"]
assert snap["overview"]["version_drift"] == overview["version_drift"]
# ---------------------------------------------------------------------------
# ClusterStateEvent protocol tests
@@ -535,6 +636,31 @@ class TestConsoleHTTPEndpoints:
"workstreams": [],
"aggregate": {},
}
collector.get_snapshot.return_value = {
"nodes": [
{
"node_id": "node-a",
"server_url": "http://a:8080",
"max_ws": 10,
"reachable": True,
"version": "0.5.0",
"health": {},
"aggregate": {"total_tokens": 50000, "total_tool_calls": 200},
"workstreams": [
{"id": "ws1", "name": "test", "state": "running", "node": "node-a"},
],
},
],
"overview": {
"nodes": 3,
"workstreams": 15,
"states": {"running": 5, "thinking": 2, "attention": 1, "idle": 6, "error": 1},
"aggregate": {"total_tokens": 50000, "total_tool_calls": 200},
"version_drift": False,
"versions": ["0.5.0"],
},
"timestamp": 1234567890.0,
}
return collector
@pytest.fixture()
@@ -614,6 +740,16 @@ class TestConsoleHTTPEndpoints:
assert status == 404
assert "error" in data
def test_get_snapshot(self, client, mock_collector):
status, data = self._get(client, "/v1/api/cluster/snapshot")
assert status == 200
assert len(data["nodes"]) == 1
assert data["nodes"][0]["node_id"] == "node-a"
assert data["overview"]["nodes"] == 3
assert data["overview"]["workstreams"] == 15
assert data["timestamp"] == 1234567890.0
mock_collector.get_snapshot.assert_called_once()
def test_health_endpoint(self, client, mock_collector):
status, data = self._get(client, "/health")
assert status == 200
@@ -1299,3 +1435,243 @@ class TestProxySharedStatic:
resp = client.get("/node/unknown/shared/base.css")
assert resp.status_code == 404
client.close()
# ---------------------------------------------------------------------------
# SSE proxy — raw byte passthrough
# ---------------------------------------------------------------------------
class TestSSEProxy:
"""Verify _proxy_sse forwards raw bytes including ping comments."""
def test_proxy_sse_preserves_pings_and_events(self):
"""SSE proxy should forward ping comments and events verbatim."""
from turnstone.console.server import _proxy_sse
# Simulate an upstream SSE response with a ping comment and a real event
sse_payload = b': ping - 2026-03-08T12:00:00Z\n\nevent: message\ndata: {"type": "test"}\n\n'
class FakeResponse:
status_code = 200
headers = {"content-type": "text/event-stream"}
async def aiter_bytes(self):
yield sse_payload
async def aclose(self):
pass
async def __aenter__(self):
return self
async def __aexit__(self, *args):
pass
class FakeClient:
def stream(self, method, url, **kwargs):
return FakeResponse()
class FakeRequest:
class url: # noqa: N801
query = "ws_id=test123"
class app: # noqa: N801
class state: # noqa: N801
proxy_sse_client = FakeClient()
proxy_auth_token = ""
headers = {}
async def is_disconnected(self):
return False
async def _run():
response = await _proxy_sse(
FakeRequest(), "http://fake:8080", "events", api_prefix="v1/api"
)
assert response.media_type == "text/event-stream"
# Collect the streamed bytes
chunks: list[bytes] = []
async for chunk in response.body_iterator:
chunks.append(chunk if isinstance(chunk, bytes) else chunk.encode())
body = b"".join(chunks)
# Ping comment must be preserved (not filtered)
assert b": ping" in body
# Real event must be preserved
assert b"event: message" in body
assert b'"type": "test"' in body
asyncio.run(_run())
def test_proxy_sse_upstream_error_status(self):
"""Non-200 upstream status should yield an error event."""
from turnstone.console.server import _proxy_sse
class FakeResponse:
status_code = 502
async def aiter_bytes(self):
return
yield # make it an async generator
async def aclose(self):
pass
async def __aenter__(self):
return self
async def __aexit__(self, *args):
pass
class FakeClient:
def stream(self, method, url, **kwargs):
return FakeResponse()
class FakeRequest:
class url: # noqa: N801
query = ""
class app: # noqa: N801
class state: # noqa: N801
proxy_sse_client = FakeClient()
proxy_auth_token = ""
headers = {}
async def is_disconnected(self):
return False
async def _run():
response = await _proxy_sse(FakeRequest(), "http://fake:8080", "events")
chunks: list[bytes] = []
async for chunk in response.body_iterator:
chunks.append(chunk if isinstance(chunk, bytes) else chunk.encode())
body = b"".join(chunks)
assert b"event: error" in body
assert b"502" in body
asyncio.run(_run())
def test_proxy_sse_disconnect_handling(self):
"""Proxy should stop when browser disconnects."""
from turnstone.console.server import _proxy_sse
class FakeResponse:
status_code = 200
async def aiter_bytes(self):
yield b"data: chunk1\n\n"
yield b"data: chunk2\n\n" # should not be reached
yield b"data: chunk3\n\n"
async def aclose(self):
pass
async def __aenter__(self):
return self
async def __aexit__(self, *args):
pass
class FakeClient:
def stream(self, method, url, **kwargs):
return FakeResponse()
call_count = 0
class FakeRequest:
class url: # noqa: N801
query = ""
class app: # noqa: N801
class state: # noqa: N801
proxy_sse_client = FakeClient()
proxy_auth_token = ""
headers = {}
async def is_disconnected(self):
nonlocal call_count
call_count += 1
return call_count > 1 # disconnect after first chunk
async def _run():
response = await _proxy_sse(FakeRequest(), "http://fake:8080", "events")
chunks: list[bytes] = []
async for chunk in response.body_iterator:
chunks.append(chunk if isinstance(chunk, bytes) else chunk.encode())
body = b"".join(chunks)
assert b"chunk1" in body
# Should have stopped before chunk3
assert b"chunk3" not in body
asyncio.run(_run())
# ---------------------------------------------------------------------------
# Collector — MCP aggregation in get_overview()
# ---------------------------------------------------------------------------
class TestCollectorMCPAggregation:
"""Verify MCP server/resource/prompt aggregation in overview and snapshot."""
def test_overview_mcp_aggregation(self):
"""Two nodes with MCP data produce correct sums in the overview."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
health={"mcp": {"servers": 2, "resources": 5, "prompts": 3}},
)
c._nodes["node-b"] = NodeSnapshot(
node_id="node-b",
server_url="http://b:8080",
health={"mcp": {"servers": 1, "resources": 4, "prompts": 2}},
)
overview = c.get_overview()
assert overview["mcp_servers"] == 3
assert overview["mcp_resources"] == 9
assert overview["mcp_prompts"] == 5
def test_overview_mcp_absent_when_zero(self):
"""Nodes without MCP data produce no mcp_servers key in the overview."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
health={"status": "ok"},
)
c._nodes["node-b"] = NodeSnapshot(
node_id="node-b",
server_url="http://b:8080",
health={},
)
overview = c.get_overview()
assert "mcp_servers" not in overview
assert "mcp_resources" not in overview
assert "mcp_prompts" not in overview
def test_overview_mcp_mixed_nodes(self):
"""One node with MCP, one without — only the MCP node contributes."""
c = _make_collector()
c._nodes["node-a"] = NodeSnapshot(
node_id="node-a",
server_url="http://a:8080",
health={"mcp": {"servers": 3, "resources": 10, "prompts": 7}},
)
c._nodes["node-b"] = NodeSnapshot(
node_id="node-b",
server_url="http://b:8080",
health={"status": "ok"},
)
overview = c.get_overview()
assert overview["mcp_servers"] == 3
assert overview["mcp_resources"] == 10
assert overview["mcp_prompts"] == 7
+761
View File
@@ -0,0 +1,761 @@
"""Tests for governance admin API endpoints (roles, orgs, policies, templates, usage, audit)."""
from __future__ import annotations
from typing import TYPE_CHECKING, Any
import pytest
from starlette.applications import Starlette
from starlette.middleware import Middleware
from starlette.middleware.base import BaseHTTPMiddleware
from starlette.routing import Mount, Route
from starlette.testclient import TestClient
if TYPE_CHECKING:
from starlette.requests import Request
from starlette.responses import Response
from turnstone.console.server import (
admin_assign_role,
admin_audit,
admin_create_policy,
admin_create_role,
admin_create_template,
admin_delete_policy,
admin_delete_role,
admin_delete_template,
admin_delete_user,
admin_get_org,
admin_list_orgs,
admin_list_policies,
admin_list_roles,
admin_list_templates,
admin_list_user_roles,
admin_unassign_role,
admin_update_org,
admin_update_policy,
admin_update_role,
admin_update_template,
admin_usage,
)
from turnstone.core.auth import AuthResult
from turnstone.core.storage._sqlite import SQLiteBackend
# ---------------------------------------------------------------------------
# Auth bypass middleware — injects a full-access AuthResult on every request.
# ---------------------------------------------------------------------------
class _InjectAuthMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request: Request, call_next: Any) -> Response:
request.state.auth_result = AuthResult(
user_id="test-admin",
scopes=frozenset({"approve"}),
token_source="config",
permissions=frozenset(
{
"read",
"write",
"approve",
"admin.roles",
"admin.users",
"admin.orgs",
"admin.policies",
"admin.templates",
"admin.usage",
"admin.audit",
"admin.schedules",
"admin.watches",
"tools.approve",
"workstreams.create",
"workstreams.close",
}
),
)
return await call_next(request)
# ---------------------------------------------------------------------------
# Fixtures
# ---------------------------------------------------------------------------
@pytest.fixture
def storage(tmp_path):
"""Fresh SQLite backend for each test, seeded with test users."""
backend = SQLiteBackend(str(tmp_path / "test.db"))
# Seed users required by role assignment tests
backend.create_user("test-admin", "testadmin", "Test Admin", "hash")
backend.create_user("user-1", "user1", "User One", "hash")
return backend
@pytest.fixture
def client(storage):
"""TestClient with storage and auth bypassed."""
app = Starlette(
routes=[
Mount(
"/v1",
routes=[
# Roles
Route("/api/admin/roles", admin_list_roles),
Route("/api/admin/roles", admin_create_role, methods=["POST"]),
Route("/api/admin/roles/{role_id}", admin_update_role, methods=["PUT"]),
Route("/api/admin/roles/{role_id}", admin_delete_role, methods=["DELETE"]),
# Users
Route(
"/api/admin/users/{user_id}",
admin_delete_user,
methods=["DELETE"],
),
# User-role assignments
Route("/api/admin/users/{user_id}/roles", admin_list_user_roles),
Route(
"/api/admin/users/{user_id}/roles",
admin_assign_role,
methods=["POST"],
),
Route(
"/api/admin/users/{user_id}/roles/{role_id}",
admin_unassign_role,
methods=["DELETE"],
),
# Orgs
Route("/api/admin/orgs", admin_list_orgs),
Route("/api/admin/orgs/{org_id}", admin_get_org),
Route("/api/admin/orgs/{org_id}", admin_update_org, methods=["PUT"]),
# Policies
Route("/api/admin/policies", admin_list_policies),
Route("/api/admin/policies", admin_create_policy, methods=["POST"]),
Route(
"/api/admin/policies/{policy_id}",
admin_update_policy,
methods=["PUT"],
),
Route(
"/api/admin/policies/{policy_id}",
admin_delete_policy,
methods=["DELETE"],
),
# Templates
Route("/api/admin/templates", admin_list_templates),
Route("/api/admin/templates", admin_create_template, methods=["POST"]),
Route(
"/api/admin/templates/{template_id}",
admin_update_template,
methods=["PUT"],
),
Route(
"/api/admin/templates/{template_id}",
admin_delete_template,
methods=["DELETE"],
),
# Usage & Audit
Route("/api/admin/usage", admin_usage),
Route("/api/admin/audit", admin_audit),
],
),
],
middleware=[Middleware(_InjectAuthMiddleware)],
)
app.state.auth_storage = storage
return TestClient(app)
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _role_payload(**overrides: Any) -> dict[str, Any]:
defaults: dict[str, Any] = {
"name": "analyst",
"display_name": "Data Analyst",
"permissions": "read,write",
}
defaults.update(overrides)
return defaults
def _policy_payload(**overrides: Any) -> dict[str, Any]:
defaults: dict[str, Any] = {
"name": "Allow bash",
"tool_pattern": "bash_*",
"action": "allow",
"priority": 10,
}
defaults.update(overrides)
return defaults
def _template_payload(**overrides: Any) -> dict[str, Any]:
defaults: dict[str, Any] = {
"name": "Greeting",
"content": "Hello {{user}}, how can I help?",
"category": "system",
}
defaults.update(overrides)
return defaults
# ---------------------------------------------------------------------------
# Tests — Roles
# ---------------------------------------------------------------------------
class TestRoles:
def test_list_empty(self, client):
resp = client.get("/v1/api/admin/roles")
assert resp.status_code == 200
assert resp.json()["roles"] == []
def test_create_role(self, client):
resp = client.post("/v1/api/admin/roles", json=_role_payload())
assert resp.status_code == 200
role = resp.json()
assert role["name"] == "analyst"
assert role["display_name"] == "Data Analyst"
assert role["permissions"] == "read,write"
assert role["builtin"] is False
assert "role_id" in role
assert "created" in role
def test_create_role_missing_name(self, client):
resp = client.post("/v1/api/admin/roles", json=_role_payload(name=""))
assert resp.status_code == 400
assert "name" in resp.json()["error"].lower()
def test_create_role_invalid_name(self, client):
resp = client.post("/v1/api/admin/roles", json=_role_payload(name="bad name!@#"))
assert resp.status_code == 400
assert "name" in resp.json()["error"].lower()
def test_create_role_default_display_name(self, client):
resp = client.post(
"/v1/api/admin/roles",
json={"name": "ops", "permissions": ""},
)
assert resp.status_code == 200
role = resp.json()
# display_name defaults to name when not provided
assert role["display_name"] == "ops"
def test_list_after_create(self, client):
client.post("/v1/api/admin/roles", json=_role_payload())
resp = client.get("/v1/api/admin/roles")
assert resp.status_code == 200
roles = resp.json()["roles"]
assert len(roles) == 1
assert roles[0]["name"] == "analyst"
def test_update_role(self, client):
create_resp = client.post("/v1/api/admin/roles", json=_role_payload())
role_id = create_resp.json()["role_id"]
resp = client.put(
f"/v1/api/admin/roles/{role_id}",
json={"display_name": "Senior Analyst", "permissions": "read,write,approve"},
)
assert resp.status_code == 200
role = resp.json()
assert role["display_name"] == "Senior Analyst"
assert role["permissions"] == "read,write,approve"
def test_update_nonexistent_role(self, client):
resp = client.put(
"/v1/api/admin/roles/nonexistent",
json={"display_name": "Nope"},
)
assert resp.status_code == 404
def test_update_builtin_role_rejected(self, client, storage):
# Seed a builtin role directly via storage
storage.create_role(
role_id="builtin-admin",
name="admin",
display_name="Administrator",
permissions="*",
builtin=True,
)
resp = client.put(
"/v1/api/admin/roles/builtin-admin",
json={"display_name": "Hacked"},
)
assert resp.status_code == 400
assert "builtin" in resp.json()["error"].lower()
def test_delete_role(self, client):
create_resp = client.post("/v1/api/admin/roles", json=_role_payload())
role_id = create_resp.json()["role_id"]
resp = client.delete(f"/v1/api/admin/roles/{role_id}")
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
# Verify gone from listing
list_resp = client.get("/v1/api/admin/roles")
assert list_resp.json()["roles"] == []
def test_delete_nonexistent_role(self, client):
resp = client.delete("/v1/api/admin/roles/nonexistent")
assert resp.status_code == 404
def test_delete_builtin_role_rejected(self, client, storage):
storage.create_role(
role_id="builtin-viewer",
name="viewer",
display_name="Viewer",
permissions="read",
builtin=True,
)
resp = client.delete("/v1/api/admin/roles/builtin-viewer")
assert resp.status_code == 400
assert "builtin" in resp.json()["error"].lower()
# ---------------------------------------------------------------------------
# Tests — Role assignments
# ---------------------------------------------------------------------------
class TestRoleAssignments:
def test_list_user_roles_empty(self, client):
resp = client.get("/v1/api/admin/users/user-1/roles")
assert resp.status_code == 200
assert resp.json()["roles"] == []
def test_assign_role(self, client):
create_resp = client.post("/v1/api/admin/roles", json=_role_payload())
role_id = create_resp.json()["role_id"]
resp = client.post(
"/v1/api/admin/users/user-1/roles",
json={"role_id": role_id},
)
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
# Verify listed
list_resp = client.get("/v1/api/admin/users/user-1/roles")
roles = list_resp.json()["roles"]
assert len(roles) >= 1
def test_assign_role_missing_role_id(self, client):
resp = client.post(
"/v1/api/admin/users/user-1/roles",
json={},
)
assert resp.status_code == 400
assert "role_id" in resp.json()["error"].lower()
def test_unassign_role(self, client):
create_resp = client.post("/v1/api/admin/roles", json=_role_payload())
role_id = create_resp.json()["role_id"]
# Assign first
client.post(
"/v1/api/admin/users/user-1/roles",
json={"role_id": role_id},
)
# Now unassign
resp = client.delete(f"/v1/api/admin/users/user-1/roles/{role_id}")
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
# Verify removed
list_resp = client.get("/v1/api/admin/users/user-1/roles")
assert list_resp.json()["roles"] == []
def test_unassign_nonexistent(self, client):
resp = client.delete("/v1/api/admin/users/user-1/roles/nonexistent")
assert resp.status_code == 404
# ---------------------------------------------------------------------------
# Tests — Orgs
# ---------------------------------------------------------------------------
class TestOrgs:
def test_list_empty(self, client):
resp = client.get("/v1/api/admin/orgs")
assert resp.status_code == 200
assert resp.json()["orgs"] == []
def test_get_org(self, client, storage):
storage.create_org(
org_id="org-1",
name="acme",
display_name="Acme Corp",
settings='{"theme": "dark"}',
)
resp = client.get("/v1/api/admin/orgs/org-1")
assert resp.status_code == 200
org = resp.json()
assert org["org_id"] == "org-1"
assert org["name"] == "acme"
assert org["display_name"] == "Acme Corp"
def test_get_org_not_found(self, client):
resp = client.get("/v1/api/admin/orgs/nonexistent")
assert resp.status_code == 404
def test_update_org(self, client, storage):
storage.create_org(org_id="org-1", name="acme", display_name="Acme Corp")
resp = client.put(
"/v1/api/admin/orgs/org-1",
json={"display_name": "Acme Inc.", "settings": '{"theme": "light"}'},
)
assert resp.status_code == 200
org = resp.json()
assert org["display_name"] == "Acme Inc."
assert org["settings"] == '{"theme": "light"}'
def test_update_org_not_found(self, client):
resp = client.put(
"/v1/api/admin/orgs/nonexistent",
json={"display_name": "Nope"},
)
assert resp.status_code == 404
# ---------------------------------------------------------------------------
# Tests — Tool policies
# ---------------------------------------------------------------------------
class TestPolicies:
def test_list_empty(self, client):
resp = client.get("/v1/api/admin/policies")
assert resp.status_code == 200
assert resp.json()["policies"] == []
def test_create_policy(self, client):
resp = client.post("/v1/api/admin/policies", json=_policy_payload())
assert resp.status_code == 200
policy = resp.json()
assert policy["name"] == "Allow bash"
assert policy["tool_pattern"] == "bash_*"
assert policy["action"] == "allow"
assert policy["priority"] == 10
assert "policy_id" in policy
assert "created" in policy
def test_create_policy_missing_name(self, client):
resp = client.post("/v1/api/admin/policies", json=_policy_payload(name=""))
assert resp.status_code == 400
assert "name" in resp.json()["error"].lower()
def test_create_policy_missing_tool_pattern(self, client):
resp = client.post(
"/v1/api/admin/policies",
json=_policy_payload(tool_pattern=""),
)
assert resp.status_code == 400
assert "tool_pattern" in resp.json()["error"].lower()
def test_create_policy_invalid_action(self, client):
resp = client.post(
"/v1/api/admin/policies",
json=_policy_payload(action="yolo"),
)
assert resp.status_code == 400
assert "action" in resp.json()["error"].lower()
def test_list_after_create(self, client):
client.post("/v1/api/admin/policies", json=_policy_payload())
resp = client.get("/v1/api/admin/policies")
assert resp.status_code == 200
policies = resp.json()["policies"]
assert len(policies) == 1
assert policies[0]["name"] == "Allow bash"
def test_update_policy(self, client):
create_resp = client.post("/v1/api/admin/policies", json=_policy_payload())
policy_id = create_resp.json()["policy_id"]
resp = client.put(
f"/v1/api/admin/policies/{policy_id}",
json={"name": "Deny bash", "action": "deny", "priority": 20},
)
assert resp.status_code == 200
policy = resp.json()
assert policy["name"] == "Deny bash"
assert policy["action"] == "deny"
assert policy["priority"] == 20
def test_update_policy_invalid_action(self, client):
create_resp = client.post("/v1/api/admin/policies", json=_policy_payload())
policy_id = create_resp.json()["policy_id"]
resp = client.put(
f"/v1/api/admin/policies/{policy_id}",
json={"action": "nope"},
)
assert resp.status_code == 400
assert "action" in resp.json()["error"].lower()
def test_update_policy_not_found(self, client):
resp = client.put(
"/v1/api/admin/policies/nonexistent",
json={"name": "Nope"},
)
assert resp.status_code == 404
def test_delete_policy(self, client):
create_resp = client.post("/v1/api/admin/policies", json=_policy_payload())
policy_id = create_resp.json()["policy_id"]
resp = client.delete(f"/v1/api/admin/policies/{policy_id}")
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
# Verify gone
list_resp = client.get("/v1/api/admin/policies")
assert list_resp.json()["policies"] == []
def test_delete_policy_not_found(self, client):
resp = client.delete("/v1/api/admin/policies/nonexistent")
assert resp.status_code == 404
# ---------------------------------------------------------------------------
# Tests — Prompt templates
# ---------------------------------------------------------------------------
class TestTemplates:
def test_list_empty(self, client):
resp = client.get("/v1/api/admin/templates")
assert resp.status_code == 200
assert resp.json()["templates"] == []
def test_create_template(self, client):
resp = client.post("/v1/api/admin/templates", json=_template_payload())
assert resp.status_code == 200
tmpl = resp.json()
assert tmpl["name"] == "Greeting"
assert "{{user}}" in tmpl["content"]
assert tmpl["category"] == "system"
assert "template_id" in tmpl
assert "created" in tmpl
def test_create_template_missing_name(self, client):
resp = client.post(
"/v1/api/admin/templates",
json=_template_payload(name=""),
)
assert resp.status_code == 400
assert "name" in resp.json()["error"].lower()
def test_create_template_missing_content(self, client):
resp = client.post(
"/v1/api/admin/templates",
json=_template_payload(content=""),
)
assert resp.status_code == 400
assert "content" in resp.json()["error"].lower()
def test_list_after_create(self, client):
client.post("/v1/api/admin/templates", json=_template_payload())
resp = client.get("/v1/api/admin/templates")
assert resp.status_code == 200
templates = resp.json()["templates"]
assert len(templates) == 1
assert templates[0]["name"] == "Greeting"
def test_update_template(self, client):
create_resp = client.post("/v1/api/admin/templates", json=_template_payload())
template_id = create_resp.json()["template_id"]
resp = client.put(
f"/v1/api/admin/templates/{template_id}",
json={"name": "Welcome", "content": "Welcome, {{user}}!", "is_default": True},
)
assert resp.status_code == 200
tmpl = resp.json()
assert tmpl["name"] == "Welcome"
assert tmpl["content"] == "Welcome, {{user}}!"
assert tmpl["is_default"] is True
def test_update_template_not_found(self, client):
resp = client.put(
"/v1/api/admin/templates/nonexistent",
json={"name": "Nope"},
)
assert resp.status_code == 404
def test_delete_template(self, client):
create_resp = client.post("/v1/api/admin/templates", json=_template_payload())
template_id = create_resp.json()["template_id"]
resp = client.delete(f"/v1/api/admin/templates/{template_id}")
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
# Verify gone
list_resp = client.get("/v1/api/admin/templates")
assert list_resp.json()["templates"] == []
def test_delete_template_not_found(self, client):
resp = client.delete("/v1/api/admin/templates/nonexistent")
assert resp.status_code == 404
# ---------------------------------------------------------------------------
# Tests — Usage
# ---------------------------------------------------------------------------
class TestUsage:
def test_usage_defaults(self, client):
"""Query usage with no params — should return summary and breakdown."""
resp = client.get("/v1/api/admin/usage")
assert resp.status_code == 200
data = resp.json()
assert "summary" in data
assert "breakdown" in data
# Summary is a list with at least one row
assert isinstance(data["summary"], list)
assert len(data["summary"]) >= 1
# All-zeros when no data
assert data["summary"][0]["prompt_tokens"] == 0
def test_usage_with_data(self, client, storage):
"""Seed usage events and verify they appear in the query."""
storage.record_usage_event(
event_id="evt-1",
user_id="user-1",
model="gpt-5",
prompt_tokens=100,
completion_tokens=50,
tool_calls_count=2,
)
storage.record_usage_event(
event_id="evt-2",
user_id="user-1",
model="gpt-5",
prompt_tokens=200,
completion_tokens=75,
tool_calls_count=1,
)
resp = client.get("/v1/api/admin/usage")
assert resp.status_code == 200
summary = resp.json()["summary"]
assert summary[0]["prompt_tokens"] == 300
assert summary[0]["completion_tokens"] == 125
assert summary[0]["tool_calls_count"] == 3
def test_usage_with_filters(self, client, storage):
storage.record_usage_event(
event_id="evt-f1",
user_id="user-a",
model="gpt-5",
prompt_tokens=100,
completion_tokens=10,
)
storage.record_usage_event(
event_id="evt-f2",
user_id="user-b",
model="claude-4",
prompt_tokens=200,
completion_tokens=20,
)
resp = client.get("/v1/api/admin/usage?user_id=user-a")
assert resp.status_code == 200
summary = resp.json()["summary"]
assert summary[0]["prompt_tokens"] == 100
resp2 = client.get("/v1/api/admin/usage?model=claude-4")
assert resp2.status_code == 200
summary2 = resp2.json()["summary"]
assert summary2[0]["prompt_tokens"] == 200
# ---------------------------------------------------------------------------
# Tests — Audit
# ---------------------------------------------------------------------------
class TestAudit:
def test_audit_empty(self, client):
resp = client.get("/v1/api/admin/audit")
assert resp.status_code == 200
data = resp.json()
assert data["events"] == []
assert data["total"] == 0
def test_audit_populated_by_mutations(self, client):
"""Creating a role should produce an audit event."""
client.post("/v1/api/admin/roles", json=_role_payload())
resp = client.get("/v1/api/admin/audit")
assert resp.status_code == 200
data = resp.json()
assert data["total"] >= 1
actions = [e["action"] for e in data["events"]]
assert "role.create" in actions
def test_audit_filter_by_action(self, client):
# Create a role and a policy to produce different audit actions
client.post("/v1/api/admin/roles", json=_role_payload())
client.post("/v1/api/admin/policies", json=_policy_payload())
resp = client.get("/v1/api/admin/audit?action=policy.create")
assert resp.status_code == 200
data = resp.json()
assert data["total"] >= 1
assert all(e["action"] == "policy.create" for e in data["events"])
def test_audit_filter_by_user_id(self, client):
client.post("/v1/api/admin/roles", json=_role_payload())
resp = client.get("/v1/api/admin/audit?user_id=test-admin")
assert resp.status_code == 200
data = resp.json()
assert data["total"] >= 1
assert all(e["user_id"] == "test-admin" for e in data["events"])
def test_audit_pagination(self, client):
# Create several resources to produce multiple audit events
for i in range(5):
client.post(
"/v1/api/admin/roles",
json=_role_payload(name=f"role-{i}"),
)
resp = client.get("/v1/api/admin/audit?limit=2&offset=0")
assert resp.status_code == 200
data = resp.json()
assert len(data["events"]) == 2
assert data["total"] >= 5
resp2 = client.get("/v1/api/admin/audit?limit=2&offset=2")
assert resp2.status_code == 200
data2 = resp2.json()
assert len(data2["events"]) == 2
# The two pages should not overlap
ids_page1 = {e["event_id"] for e in data["events"]}
ids_page2 = {e["event_id"] for e in data2["events"]}
assert ids_page1.isdisjoint(ids_page2)
# ---------------------------------------------------------------------------
# Tests — User self-deletion guard
# ---------------------------------------------------------------------------
class TestUserSelfDeletion:
def test_cannot_delete_self(self, client):
"""Admin should not be able to delete their own account."""
resp = client.delete("/v1/api/admin/users/test-admin")
assert resp.status_code == 400
assert "own account" in resp.json()["error"].lower()
def test_can_delete_other_user(self, client):
resp = client.delete("/v1/api/admin/users/user-1")
assert resp.status_code == 200
assert resp.json()["status"] == "ok"
+808
View File
@@ -0,0 +1,808 @@
"""Tests for governance storage operations (SQLite backend).
Covers RBAC roles, organizations, tool policies, prompt templates,
usage events, and audit events.
"""
from __future__ import annotations
from datetime import UTC, datetime
import pytest
import sqlalchemy as sa
from turnstone.core.storage._sqlite import SQLiteBackend
@pytest.fixture()
def db(tmp_path):
"""Create a fresh SQLite backend for each test."""
return SQLiteBackend(str(tmp_path / "test.db"))
# ---------------------------------------------------------------------------
# Roles
# ---------------------------------------------------------------------------
class TestRoleCRUD:
def test_create_role(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
role = db.get_role("r1")
assert role is not None
assert role["role_id"] == "r1"
assert role["name"] == "editor"
assert role["display_name"] == "Editor"
assert role["permissions"] == "read,write"
assert role["builtin"] is False
assert role["org_id"] == ""
assert "created" in role
assert "updated" in role
def test_create_role_idempotent(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
# Second insert with same role_id should be silently ignored.
db.create_role("r1", "editor2", "Editor 2", "read", builtin=True, org_id="org1")
role = db.get_role("r1")
assert role is not None
# Original values preserved.
assert role["name"] == "editor"
assert role["display_name"] == "Editor"
def test_get_role_by_name(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
role = db.get_role_by_name("editor")
assert role is not None
assert role["role_id"] == "r1"
def test_get_role_by_name_nonexistent(self, db):
assert db.get_role_by_name("nope") is None
def test_list_roles(self, db):
db.create_role("r2", "beta", "Beta Role", "read", builtin=False, org_id="")
db.create_role("r1", "alpha", "Alpha Role", "write", builtin=False, org_id="")
roles = db.list_roles()
assert len(roles) == 2
# Ordered by name ascending.
assert roles[0]["name"] == "alpha"
assert roles[1]["name"] == "beta"
def test_list_roles_filter_org(self, db):
db.create_role("r1", "role_a", "A", "read", builtin=False, org_id="org1")
db.create_role("r2", "role_b", "B", "read", builtin=False, org_id="org2")
db.create_role("r3", "role_c", "C", "read", builtin=False, org_id="org1")
result = db.list_roles(org_id="org1")
assert len(result) == 2
assert {r["role_id"] for r in result} == {"r1", "r3"}
def test_update_role(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
ok = db.update_role("r1", permissions="read,write,approve", display_name="Senior Editor")
assert ok is True
role = db.get_role("r1")
assert role is not None
assert role["permissions"] == "read,write,approve"
assert role["display_name"] == "Senior Editor"
def test_update_role_nonexistent(self, db):
assert db.update_role("missing", permissions="read") is False
def test_delete_role(self, db):
db.create_role("r1", "editor", "Editor", "read", builtin=False, org_id="")
db.create_user("u1", "alice", "Alice", "$2b$hash")
db.assign_role("u1", "r1")
# Verify assignment exists.
assert len(db.list_user_roles("u1")) == 1
ok = db.delete_role("r1")
assert ok is True
assert db.get_role("r1") is None
# Cascade: user_roles for this role should be gone.
assert len(db.list_user_roles("u1")) == 0
def test_delete_role_nonexistent(self, db):
assert db.delete_role("missing") is False
def test_assign_role(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
db.create_user("u1", "alice", "Alice", "$2b$hash")
db.assign_role("u1", "r1", assigned_by="admin")
roles = db.list_user_roles("u1")
assert len(roles) == 1
assert roles[0]["role_id"] == "r1"
assert roles[0]["assigned_by"] == "admin"
def test_assign_role_idempotent(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
db.create_user("u1", "alice", "Alice", "$2b$hash")
db.assign_role("u1", "r1")
# Second assign should not raise.
db.assign_role("u1", "r1")
roles = db.list_user_roles("u1")
assert len(roles) == 1
def test_unassign_role(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
db.create_user("u1", "alice", "Alice", "$2b$hash")
db.assign_role("u1", "r1")
ok = db.unassign_role("u1", "r1")
assert ok is True
assert len(db.list_user_roles("u1")) == 0
def test_unassign_role_nonexistent(self, db):
assert db.unassign_role("u1", "r1") is False
def test_list_user_roles(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
db.create_role("r2", "viewer", "Viewer", "read", builtin=True, org_id="")
db.create_user("u1", "alice", "Alice", "$2b$hash")
db.assign_role("u1", "r1", assigned_by="admin")
db.assign_role("u1", "r2", assigned_by="system")
roles = db.list_user_roles("u1")
assert len(roles) == 2
# Each entry should have joined role fields plus assignment metadata.
for r in roles:
assert "role_id" in r
assert "name" in r
assert "permissions" in r
assert "assigned_by" in r
assert "assignment_created" in r
def test_get_user_permissions(self, db):
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
db.create_role("r2", "approver", "Approver", "approve,read", builtin=False, org_id="")
db.create_user("u1", "alice", "Alice", "$2b$hash")
db.assign_role("u1", "r1")
db.assign_role("u1", "r2")
perms = db.get_user_permissions("u1")
assert perms == {"read", "write", "approve"}
def test_get_user_permissions_no_roles(self, db):
db.create_user("u1", "alice", "Alice", "$2b$hash")
assert db.get_user_permissions("u1") == set()
# ---------------------------------------------------------------------------
# Organizations
# ---------------------------------------------------------------------------
class TestOrgCRUD:
def test_create_org(self, db):
db.create_org("org1", "acme", "Acme Corp", '{"plan":"pro"}')
org = db.get_org("org1")
assert org is not None
assert org["org_id"] == "org1"
assert org["name"] == "acme"
assert org["display_name"] == "Acme Corp"
assert org["settings"] == '{"plan":"pro"}'
assert "created" in org
assert "updated" in org
def test_get_org_nonexistent(self, db):
assert db.get_org("nope") is None
def test_create_org_idempotent(self, db):
db.create_org("org1", "acme", "Acme Corp")
db.create_org("org1", "acme2", "Acme 2")
org = db.get_org("org1")
assert org is not None
assert org["name"] == "acme"
def test_list_orgs(self, db):
db.create_org("o2", "beta", "Beta Inc")
db.create_org("o1", "alpha", "Alpha LLC")
orgs = db.list_orgs()
assert len(orgs) == 2
# Ordered by name ascending.
assert orgs[0]["name"] == "alpha"
assert orgs[1]["name"] == "beta"
def test_update_org(self, db):
db.create_org("org1", "acme", "Acme Corp")
ok = db.update_org(
"org1", display_name="Acme Corp Global", settings='{"plan":"enterprise"}'
)
assert ok is True
org = db.get_org("org1")
assert org is not None
assert org["display_name"] == "Acme Corp Global"
assert org["settings"] == '{"plan":"enterprise"}'
def test_update_org_nonexistent(self, db):
assert db.update_org("missing", display_name="X") is False
# ---------------------------------------------------------------------------
# Tool Policies
# ---------------------------------------------------------------------------
class TestToolPolicyCRUD:
def test_create_tool_policy(self, db):
db.create_tool_policy(
"p1",
"deny-bash",
"bash*",
"deny",
priority=100,
org_id="org1",
enabled=True,
created_by="admin",
)
pol = db.get_tool_policy("p1")
assert pol is not None
assert pol["policy_id"] == "p1"
assert pol["name"] == "deny-bash"
assert pol["tool_pattern"] == "bash*"
assert pol["action"] == "deny"
assert pol["priority"] == 100
assert pol["org_id"] == "org1"
assert pol["enabled"] is True
assert pol["created_by"] == "admin"
def test_get_tool_policy_nonexistent(self, db):
assert db.get_tool_policy("missing") is None
def test_list_tool_policies_ordered_by_priority(self, db):
db.create_tool_policy("p1", "low", "*", "allow", priority=10)
db.create_tool_policy("p2", "high", "*", "deny", priority=100)
db.create_tool_policy("p3", "mid", "*", "ask", priority=50)
policies = db.list_tool_policies()
assert len(policies) == 3
# DESC priority order.
assert policies[0]["priority"] == 100
assert policies[1]["priority"] == 50
assert policies[2]["priority"] == 10
def test_update_tool_policy(self, db):
db.create_tool_policy("p1", "deny-bash", "bash*", "deny", priority=100)
ok = db.update_tool_policy("p1", action="allow", priority=50)
assert ok is True
pol = db.get_tool_policy("p1")
assert pol is not None
assert pol["action"] == "allow"
assert pol["priority"] == 50
def test_update_tool_policy_nonexistent(self, db):
assert db.update_tool_policy("missing", action="deny") is False
def test_delete_tool_policy(self, db):
db.create_tool_policy("p1", "deny-bash", "bash*", "deny", priority=100)
ok = db.delete_tool_policy("p1")
assert ok is True
assert db.get_tool_policy("p1") is None
def test_delete_tool_policy_nonexistent(self, db):
assert db.delete_tool_policy("missing") is False
def test_enabled_as_bool(self, db):
db.create_tool_policy("p1", "on", "*", "allow", priority=0, enabled=True)
db.create_tool_policy("p2", "off", "*", "deny", priority=0, enabled=False)
p1 = db.get_tool_policy("p1")
p2 = db.get_tool_policy("p2")
assert p1 is not None
assert p2 is not None
assert p1["enabled"] is True
assert isinstance(p1["enabled"], bool)
assert p2["enabled"] is False
assert isinstance(p2["enabled"], bool)
def test_list_policies_filter_org(self, db):
db.create_tool_policy("p1", "a", "*", "allow", priority=0, org_id="org1")
db.create_tool_policy("p2", "b", "*", "deny", priority=0, org_id="org2")
db.create_tool_policy("p3", "c", "*", "ask", priority=0, org_id="org1")
result = db.list_tool_policies(org_id="org1")
assert len(result) == 2
assert {r["policy_id"] for r in result} == {"p1", "p3"}
# ---------------------------------------------------------------------------
# Prompt Templates
# ---------------------------------------------------------------------------
class TestPromptTemplateCRUD:
def test_create_prompt_template(self, db):
db.create_prompt_template(
"t1",
"greeting",
"general",
"Hello {{name}}!",
variables='["name"]',
is_default=True,
org_id="org1",
created_by="admin",
)
tpl = db.get_prompt_template("t1")
assert tpl is not None
assert tpl["template_id"] == "t1"
assert tpl["name"] == "greeting"
assert tpl["category"] == "general"
assert tpl["content"] == "Hello {{name}}!"
assert tpl["variables"] == '["name"]'
assert tpl["is_default"] is True
assert tpl["org_id"] == "org1"
assert tpl["created_by"] == "admin"
def test_get_prompt_template_nonexistent(self, db):
assert db.get_prompt_template("missing") is None
def test_list_prompt_templates_ordered_by_name(self, db):
db.create_prompt_template("t2", "beta", "general", "B")
db.create_prompt_template("t1", "alpha", "general", "A")
templates = db.list_prompt_templates()
assert len(templates) == 2
assert templates[0]["name"] == "alpha"
assert templates[1]["name"] == "beta"
def test_list_prompt_templates_filter_org(self, db):
db.create_prompt_template("t1", "a", "general", "A", org_id="org1")
db.create_prompt_template("t2", "b", "general", "B", org_id="org2")
result = db.list_prompt_templates(org_id="org1")
assert len(result) == 1
assert result[0]["template_id"] == "t1"
def test_update_prompt_template(self, db):
db.create_prompt_template("t1", "greeting", "general", "Hello!")
ok = db.update_prompt_template("t1", content="Hi there!", category="custom")
assert ok is True
tpl = db.get_prompt_template("t1")
assert tpl is not None
assert tpl["content"] == "Hi there!"
assert tpl["category"] == "custom"
def test_update_prompt_template_nonexistent(self, db):
assert db.update_prompt_template("missing", content="x") is False
def test_delete_prompt_template(self, db):
db.create_prompt_template("t1", "greeting", "general", "Hello!")
ok = db.delete_prompt_template("t1")
assert ok is True
assert db.get_prompt_template("t1") is None
def test_delete_prompt_template_nonexistent(self, db):
assert db.delete_prompt_template("missing") is False
def test_is_default_as_bool(self, db):
db.create_prompt_template("t1", "default_one", "general", "D", is_default=True)
db.create_prompt_template("t2", "not_default", "general", "N", is_default=False)
t1 = db.get_prompt_template("t1")
t2 = db.get_prompt_template("t2")
assert t1 is not None
assert t2 is not None
assert t1["is_default"] is True
assert isinstance(t1["is_default"], bool)
assert t2["is_default"] is False
assert isinstance(t2["is_default"], bool)
def test_create_with_mcp_origin(self, db):
db.create_prompt_template(
"t1",
"mcp__srv__prompt",
"mcp",
"content",
variables="[]",
is_default=False,
org_id="",
created_by="",
origin="mcp",
mcp_server="srv",
readonly=True,
)
tpl = db.get_prompt_template("t1")
assert tpl is not None
assert tpl["origin"] == "mcp"
assert tpl["mcp_server"] == "srv"
assert tpl["readonly"] is True
assert isinstance(tpl["readonly"], bool)
def test_default_origin_values(self, db):
db.create_prompt_template("t1", "basic", "general", "Hello")
tpl = db.get_prompt_template("t1")
assert tpl is not None
assert tpl["origin"] == "manual"
assert tpl["mcp_server"] == ""
assert tpl["readonly"] is False
def test_get_prompt_template_by_name(self, db):
db.create_prompt_template("t1", "greeting", "general", "Hello!")
tpl = db.get_prompt_template_by_name("greeting")
assert tpl is not None
assert tpl["template_id"] == "t1"
assert tpl["name"] == "greeting"
def test_get_prompt_template_by_name_nonexistent(self, db):
assert db.get_prompt_template_by_name("nope") is None
def test_list_default_templates(self, db):
db.create_prompt_template("t1", "alpha", "general", "A", is_default=True)
db.create_prompt_template("t2", "beta", "general", "B", is_default=False)
db.create_prompt_template("t3", "gamma", "general", "C", is_default=True)
result = db.list_default_templates()
assert len(result) == 2
assert result[0]["name"] == "alpha"
assert result[1]["name"] == "gamma"
def test_list_default_templates_empty(self, db):
db.create_prompt_template("t1", "alpha", "general", "A", is_default=False)
assert db.list_default_templates() == []
def test_list_prompt_templates_by_origin(self, db):
db.create_prompt_template("t1", "manual_one", "general", "A", origin="manual")
db.create_prompt_template("t2", "mcp_one", "mcp", "B", origin="mcp", mcp_server="srv1")
db.create_prompt_template("t3", "mcp_two", "mcp", "C", origin="mcp", mcp_server="srv2")
result = db.list_prompt_templates_by_origin("mcp")
assert len(result) == 2
names = [r["name"] for r in result]
assert "mcp_one" in names
assert "mcp_two" in names
# ---------------------------------------------------------------------------
# Usage Events
# ---------------------------------------------------------------------------
class TestUsageEvents:
def test_record_usage_event(self, db):
db.record_usage_event(
"ev1",
user_id="u1",
ws_id="ws1",
node_id="n1",
model="gpt-5",
prompt_tokens=100,
completion_tokens=50,
tool_calls_count=2,
)
# Verify via query_usage (no group_by returns summary).
result = db.query_usage(since="2000-01-01T00:00:00")
assert len(result) == 1
assert result[0]["prompt_tokens"] == 100
assert result[0]["completion_tokens"] == 50
assert result[0]["tool_calls_count"] == 2
def test_query_usage_summary(self, db):
db.record_usage_event("ev1", model="gpt-5", prompt_tokens=100, completion_tokens=50)
db.record_usage_event("ev2", model="gpt-5", prompt_tokens=200, completion_tokens=75)
result = db.query_usage(since="2000-01-01T00:00:00")
assert len(result) == 1
assert result[0]["prompt_tokens"] == 300
assert result[0]["completion_tokens"] == 125
def test_query_usage_by_day(self, db):
# Insert events with known timestamps by directly inserting rows.
from turnstone.core.storage._schema import usage_events
with db._engine.connect() as conn:
conn.execute(
sa.insert(usage_events),
[
{
"event_id": "e1",
"timestamp": "2026-03-01T10:00:00",
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "gpt-5",
"prompt_tokens": 100,
"completion_tokens": 50,
"tool_calls_count": 0,
"created": "2026-03-01T10:00:00",
},
{
"event_id": "e2",
"timestamp": "2026-03-01T14:00:00",
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "gpt-5",
"prompt_tokens": 50,
"completion_tokens": 25,
"tool_calls_count": 0,
"created": "2026-03-01T14:00:00",
},
{
"event_id": "e3",
"timestamp": "2026-03-02T08:00:00",
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "gpt-5",
"prompt_tokens": 200,
"completion_tokens": 100,
"tool_calls_count": 0,
"created": "2026-03-02T08:00:00",
},
],
)
conn.commit()
result = db.query_usage(since="2026-03-01T00:00:00", group_by="day")
assert len(result) == 2
assert result[0]["key"] == "2026-03-01"
assert result[0]["prompt_tokens"] == 150
assert result[1]["key"] == "2026-03-02"
assert result[1]["prompt_tokens"] == 200
def test_query_usage_by_model(self, db):
from turnstone.core.storage._schema import usage_events
with db._engine.connect() as conn:
conn.execute(
sa.insert(usage_events),
[
{
"event_id": "e1",
"timestamp": "2026-03-01T10:00:00",
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "gpt-5",
"prompt_tokens": 100,
"completion_tokens": 50,
"tool_calls_count": 0,
"created": "2026-03-01T10:00:00",
},
{
"event_id": "e2",
"timestamp": "2026-03-01T10:00:00",
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "claude-4",
"prompt_tokens": 200,
"completion_tokens": 100,
"tool_calls_count": 1,
"created": "2026-03-01T10:00:00",
},
],
)
conn.commit()
result = db.query_usage(since="2026-03-01T00:00:00", group_by="model")
assert len(result) == 2
keys = [r["key"] for r in result]
assert "gpt-5" in keys
assert "claude-4" in keys
def test_query_usage_by_user(self, db):
from turnstone.core.storage._schema import usage_events
with db._engine.connect() as conn:
conn.execute(
sa.insert(usage_events),
[
{
"event_id": "e1",
"timestamp": "2026-03-01T10:00:00",
"user_id": "u1",
"ws_id": "",
"node_id": "",
"model": "",
"prompt_tokens": 100,
"completion_tokens": 50,
"tool_calls_count": 0,
"created": "2026-03-01T10:00:00",
},
{
"event_id": "e2",
"timestamp": "2026-03-01T10:00:00",
"user_id": "u2",
"ws_id": "",
"node_id": "",
"model": "",
"prompt_tokens": 300,
"completion_tokens": 150,
"tool_calls_count": 2,
"created": "2026-03-01T10:00:00",
},
],
)
conn.commit()
result = db.query_usage(since="2026-03-01T00:00:00", group_by="user")
assert len(result) == 2
by_key = {r["key"]: r for r in result}
assert by_key["u1"]["prompt_tokens"] == 100
assert by_key["u2"]["prompt_tokens"] == 300
def test_query_usage_filter_model(self, db):
from turnstone.core.storage._schema import usage_events
with db._engine.connect() as conn:
conn.execute(
sa.insert(usage_events),
[
{
"event_id": "e1",
"timestamp": "2026-03-01T10:00:00",
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "gpt-5",
"prompt_tokens": 100,
"completion_tokens": 50,
"tool_calls_count": 0,
"created": "2026-03-01T10:00:00",
},
{
"event_id": "e2",
"timestamp": "2026-03-01T10:00:00",
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "claude-4",
"prompt_tokens": 200,
"completion_tokens": 100,
"tool_calls_count": 0,
"created": "2026-03-01T10:00:00",
},
],
)
conn.commit()
result = db.query_usage(since="2026-03-01T00:00:00", model="gpt-5")
assert len(result) == 1
assert result[0]["prompt_tokens"] == 100
def test_prune_usage_events(self, db):
from turnstone.core.storage._schema import usage_events
old_ts = "2020-01-01T00:00:00"
now_ts = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with db._engine.connect() as conn:
conn.execute(
sa.insert(usage_events),
[
{
"event_id": "old",
"timestamp": old_ts,
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "",
"prompt_tokens": 10,
"completion_tokens": 5,
"tool_calls_count": 0,
"created": old_ts,
},
{
"event_id": "new",
"timestamp": now_ts,
"user_id": "",
"ws_id": "",
"node_id": "",
"model": "",
"prompt_tokens": 20,
"completion_tokens": 10,
"tool_calls_count": 0,
"created": now_ts,
},
],
)
conn.commit()
pruned = db.prune_usage_events(retention_days=30)
assert pruned == 1
# Only the recent event should remain.
result = db.query_usage(since="2000-01-01T00:00:00")
assert result[0]["prompt_tokens"] == 20
# ---------------------------------------------------------------------------
# Audit Events
# ---------------------------------------------------------------------------
class TestAuditEvents:
def test_record_audit_event(self, db):
db.record_audit_event(
"a1",
user_id="u1",
action="role.create",
resource_type="role",
resource_id="r1",
detail='{"name":"editor"}',
ip_address="127.0.0.1",
)
events = db.list_audit_events()
assert len(events) == 1
ev = events[0]
assert ev["event_id"] == "a1"
assert ev["user_id"] == "u1"
assert ev["action"] == "role.create"
assert ev["resource_type"] == "role"
assert ev["resource_id"] == "r1"
assert ev["detail"] == '{"name":"editor"}'
assert ev["ip_address"] == "127.0.0.1"
def test_list_audit_events(self, db):
db.record_audit_event("a1", action="login")
db.record_audit_event("a2", action="logout")
events = db.list_audit_events()
assert len(events) == 2
# Ordered by timestamp DESC — most recent first.
# Both created in quick succession with same-second granularity,
# but the order should still be deterministic (DESC).
assert {e["event_id"] for e in events} == {"a1", "a2"}
def test_list_audit_events_filter_action(self, db):
db.record_audit_event("a1", action="login")
db.record_audit_event("a2", action="logout")
db.record_audit_event("a3", action="login")
events = db.list_audit_events(action="login")
assert len(events) == 2
assert all(e["action"] == "login" for e in events)
def test_list_audit_events_filter_user(self, db):
db.record_audit_event("a1", user_id="u1", action="login")
db.record_audit_event("a2", user_id="u2", action="login")
events = db.list_audit_events(user_id="u1")
assert len(events) == 1
assert events[0]["user_id"] == "u1"
def test_list_audit_events_pagination(self, db):
for i in range(5):
db.record_audit_event(f"a{i}", action="test")
page1 = db.list_audit_events(limit=2, offset=0)
page2 = db.list_audit_events(limit=2, offset=2)
page3 = db.list_audit_events(limit=2, offset=4)
assert len(page1) == 2
assert len(page2) == 2
assert len(page3) == 1
# No overlap.
ids = [e["event_id"] for e in page1 + page2 + page3]
assert len(set(ids)) == 5
def test_count_audit_events(self, db):
db.record_audit_event("a1", action="login")
db.record_audit_event("a2", action="logout")
db.record_audit_event("a3", action="login")
assert db.count_audit_events() == 3
assert db.count_audit_events(action="login") == 2
assert db.count_audit_events(action="logout") == 1
def test_count_audit_events_filter_user(self, db):
db.record_audit_event("a1", user_id="u1", action="login")
db.record_audit_event("a2", user_id="u2", action="login")
assert db.count_audit_events(user_id="u1") == 1
def test_prune_audit_events(self, db):
from turnstone.core.storage._schema import audit_events
old_ts = "2020-01-01T00:00:00"
now_ts = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with db._engine.connect() as conn:
conn.execute(
sa.insert(audit_events),
[
{
"event_id": "old",
"timestamp": old_ts,
"user_id": "",
"action": "test",
"resource_type": "",
"resource_id": "",
"detail": "{}",
"ip_address": "",
"created": old_ts,
},
{
"event_id": "new",
"timestamp": now_ts,
"user_id": "",
"action": "test",
"resource_type": "",
"resource_id": "",
"detail": "{}",
"ip_address": "",
"created": now_ts,
},
],
)
conn.commit()
pruned = db.prune_audit_events(retention_days=30)
assert pruned == 1
assert db.count_audit_events() == 1
+660 -1
View File
@@ -51,6 +51,83 @@ def _fake_openai_tool(name: str = "mcp__test__search") -> dict[str, Any]:
}
def _fake_mcp_resource(
uri: str = "file:///README.md",
name: str = "readme",
description: str = "Project readme",
mime_type: str = "text/plain",
) -> MagicMock:
"""Create a mock MCP Resource object matching the SDK's Resource type."""
res = MagicMock()
res.uri = uri
res.name = name
res.description = description
res.mimeType = mime_type
return res
def _fake_resource_dict(
uri: str = "file:///README.md",
name: str = "readme",
description: str = "Project readme",
mime_type: str = "text/plain",
server: str = "test",
) -> dict[str, Any]:
"""Create a fake resource dict as stored in per-server state."""
return {
"uri": uri,
"name": name,
"description": description,
"mimeType": mime_type,
"server": server,
}
def _fake_mcp_prompt(
name: str = "code_review",
description: str = "Generate a code review",
arguments: list[dict[str, Any]] | None = None,
) -> MagicMock:
"""Create a mock MCP Prompt object matching the SDK's Prompt type."""
prompt = MagicMock()
prompt.name = name
prompt.description = description
if arguments is None:
arg = MagicMock()
arg.name = "language"
arg.description = "Programming language"
arg.required = True
prompt.arguments = [arg]
else:
mock_args = []
for a in arguments:
arg = MagicMock()
arg.name = a["name"]
arg.description = a.get("description", "")
arg.required = a.get("required", False)
mock_args.append(arg)
prompt.arguments = mock_args
return prompt
def _fake_prompt_dict(
name: str = "mcp__test__code_review",
original_name: str = "code_review",
server: str = "test",
description: str = "Generate a code review",
) -> dict[str, Any]:
"""Create a fake prompt dict as stored in per-server state."""
return {
"name": name,
"original_name": original_name,
"server": server,
"description": description,
"arguments": [
{"name": "language", "description": "Programming language", "required": True}
],
}
# ---------------------------------------------------------------------------
# Schema conversion
# ---------------------------------------------------------------------------
@@ -453,6 +530,23 @@ class TestRebuildTools:
class TestRefreshServer:
@staticmethod
def _add_empty_resource_prompt_mocks(
mgr: MCPClientManager, server_name: str, mock_session: MagicMock
) -> None:
"""Add empty list_resources/list_prompts mocks so _refresh_server works."""
mgr._supports_resources[server_name] = True
mgr._supports_prompts[server_name] = True
empty_res = MagicMock()
empty_res.resources = []
mock_session.list_resources = AsyncMock(return_value=empty_res)
empty_tmpl = MagicMock()
empty_tmpl.resourceTemplates = []
mock_session.list_resource_templates = AsyncMock(return_value=empty_tmpl)
empty_prompts = MagicMock()
empty_prompts.prompts = []
mock_session.list_prompts = AsyncMock(return_value=empty_prompts)
def test_refresh_detects_added_tools(self):
async def _run() -> None:
mgr = MCPClientManager({})
@@ -463,6 +557,7 @@ class TestRefreshServer:
_fake_mcp_tool("create"), # new tool
]
mock_session.list_tools = AsyncMock(return_value=mock_result)
self._add_empty_resource_prompt_mocks(mgr, "github", mock_session)
mgr._sessions["github"] = mock_session
mgr._per_server_tools["github"] = [_fake_openai_tool("mcp__github__search")]
mgr._rebuild_tools()
@@ -481,6 +576,7 @@ class TestRefreshServer:
mock_result = MagicMock()
mock_result.tools = [] # all tools removed
mock_session.list_tools = AsyncMock(return_value=mock_result)
self._add_empty_resource_prompt_mocks(mgr, "github", mock_session)
mgr._sessions["github"] = mock_session
mgr._per_server_tools["github"] = [_fake_openai_tool("mcp__github__search")]
mgr._rebuild_tools()
@@ -499,6 +595,7 @@ class TestRefreshServer:
mock_result = MagicMock()
mock_result.tools = [_fake_mcp_tool("search")]
mock_session.list_tools = AsyncMock(return_value=mock_result)
self._add_empty_resource_prompt_mocks(mgr, "github", mock_session)
mgr._sessions["github"] = mock_session
mgr._per_server_tools["github"] = [_fake_openai_tool("mcp__github__search")]
mgr._rebuild_tools()
@@ -513,7 +610,7 @@ class TestRefreshServer:
async def _run() -> None:
mgr = MCPClientManager({})
with pytest.raises(RuntimeError, match="not connected"):
await mgr._refresh_server("ghost")
await mgr._refresh_server_tools("ghost")
asyncio.run(_run())
@@ -709,3 +806,565 @@ class TestSessionRefresh:
session.handle_command("/mcp refresh")
session.ui.on_error.assert_called_once()
assert "MCP refresh failed" in session.ui.on_error.call_args[0][0]
# ---------------------------------------------------------------------------
# MCP Resources
# ---------------------------------------------------------------------------
class TestMCPResources:
def test_resource_discovery(self):
"""Mock list_resources() returning 2 resources, verify get_resources()."""
mgr = MCPClientManager({})
mgr._per_server_resources = {
"fs": [
_fake_resource_dict("file:///a.txt", "a", "File A", "text/plain", "fs"),
_fake_resource_dict("file:///b.txt", "b", "File B", "text/plain", "fs"),
],
}
mgr._rebuild_resources()
resources = mgr.get_resources()
assert len(resources) == 2
uris = {r["uri"] for r in resources}
assert uris == {"file:///a.txt", "file:///b.txt"}
assert all(r["server"] == "fs" for r in resources)
def test_rebuild_resources_copy_on_write(self):
"""Verify mutation safety — get_resources() returns independent copy."""
mgr = MCPClientManager({})
mgr._per_server_resources = {
"a": [_fake_resource_dict("file:///x", "x", "", "", "a")],
}
mgr._rebuild_resources()
old_resources = mgr._resources
old_map = mgr._resource_map
mgr._per_server_resources["b"] = [_fake_resource_dict("file:///y", "y", "", "", "b")]
mgr._rebuild_resources()
assert mgr._resources is not old_resources
assert mgr._resource_map is not old_map
def test_get_resources_returns_copy(self):
mgr = MCPClientManager({})
mgr._per_server_resources = {
"a": [_fake_resource_dict("file:///x", "x", "", "", "a")],
}
mgr._rebuild_resources()
resources = mgr.get_resources()
assert len(resources) == 1
resources.clear()
assert len(mgr.get_resources()) == 1
def test_read_resource_sync(self):
"""Mock session.read_resource(), verify text extraction."""
mgr = MCPClientManager({})
mgr._resource_map = {"file:///readme": ("fs", "file:///readme")}
mock_session = MagicMock()
mgr._sessions["fs"] = mock_session
mgr._loop = asyncio.new_event_loop()
# Mock the read_resource result
text_content = MagicMock(spec=["text"])
text_content.text = "Hello, world!"
mock_result = MagicMock()
mock_result.contents = [text_content]
mock_session.read_resource = AsyncMock(return_value=mock_result)
thread = None
try:
thread = __import__("threading").Thread(target=mgr._loop.run_forever, daemon=True)
thread.start()
output = mgr.read_resource_sync("file:///readme", timeout=5)
assert output == "Hello, world!"
mock_session.read_resource.assert_awaited_once_with("file:///readme")
finally:
mgr._loop.call_soon_threadsafe(mgr._loop.stop)
if thread:
thread.join(timeout=5)
mgr._loop.close()
def test_read_resource_sync_blob(self):
"""Verify base64 blob extraction."""
mgr = MCPClientManager({})
mgr._resource_map = {"file:///img.png": ("fs", "file:///img.png")}
mock_session = MagicMock()
mgr._sessions["fs"] = mock_session
mgr._loop = asyncio.new_event_loop()
blob_content = MagicMock(spec=["blob"])
blob_content.blob = "aGVsbG8="
mock_result = MagicMock()
mock_result.contents = [blob_content]
mock_session.read_resource = AsyncMock(return_value=mock_result)
thread = None
try:
thread = __import__("threading").Thread(target=mgr._loop.run_forever, daemon=True)
thread.start()
output = mgr.read_resource_sync("file:///img.png", timeout=5)
assert output == "aGVsbG8="
finally:
mgr._loop.call_soon_threadsafe(mgr._loop.stop)
if thread:
thread.join(timeout=5)
mgr._loop.close()
def test_read_resource_sync_unknown_uri(self):
mgr = MCPClientManager({})
with pytest.raises(ValueError, match="Unknown MCP resource"):
mgr.read_resource_sync("file:///nonexistent")
def test_read_resource_sync_disconnected(self):
mgr = MCPClientManager({})
mgr._resource_map = {"file:///x": ("dead", "file:///x")}
with pytest.raises(RuntimeError, match="not connected"):
mgr.read_resource_sync("file:///x")
def test_read_resource_sync_timeout(self):
"""Verify timeout handling."""
mgr = MCPClientManager({})
mgr._resource_map = {"file:///x": ("fs", "file:///x")}
mock_session = MagicMock()
mgr._sessions["fs"] = mock_session
mgr._loop = asyncio.new_event_loop()
async def _slow_read(_uri: str) -> None:
await asyncio.sleep(10)
mock_session.read_resource = _slow_read
thread = None
try:
thread = __import__("threading").Thread(target=mgr._loop.run_forever, daemon=True)
thread.start()
with pytest.raises(TimeoutError):
mgr.read_resource_sync("file:///x", timeout=1)
finally:
mgr._loop.call_soon_threadsafe(mgr._loop.stop)
if thread:
thread.join(timeout=5)
mgr._loop.close()
def test_resource_listener_notification(self):
"""Verify callback fires on rebuild."""
mgr = MCPClientManager({})
calls: list[int] = []
mgr.add_resource_listener(lambda: calls.append(1))
mgr._per_server_resources = {"a": [_fake_resource_dict()]}
mgr._rebuild_resources()
assert len(calls) == 1
def test_resource_listener_remove(self):
mgr = MCPClientManager({})
calls: list[int] = []
cb = lambda: calls.append(1) # noqa: E731
mgr.add_resource_listener(cb)
mgr.remove_resource_listener(cb)
mgr._rebuild_resources()
assert calls == []
def test_resource_listener_error_does_not_propagate(self):
mgr = MCPClientManager({})
mgr.add_resource_listener(lambda: 1 / 0)
mgr._rebuild_resources() # should not raise
def test_resource_refresh_on_notification(self):
"""Mock notification, verify re-fetch of resources."""
async def _run() -> None:
mgr = MCPClientManager({})
mock_session = MagicMock()
mgr._sessions["fs"] = mock_session
mgr._supports_resources["fs"] = True
# Initial state
mgr._per_server_resources["fs"] = [
_fake_resource_dict("file:///old", server="fs"),
]
mgr._rebuild_resources()
assert len(mgr.get_resources()) == 1
# Mock the re-fetch returning a new resource
new_res = _fake_mcp_resource("file:///new", "new")
mock_res_result = MagicMock()
mock_res_result.resources = [new_res]
mock_session.list_resources = AsyncMock(return_value=mock_res_result)
mock_tmpl_result = MagicMock()
mock_tmpl_result.resourceTemplates = []
mock_session.list_resource_templates = AsyncMock(return_value=mock_tmpl_result)
await mgr._refresh_server_resources("fs")
resources = mgr.get_resources()
assert len(resources) == 1
assert resources[0]["uri"] == "file:///new"
asyncio.run(_run())
def test_rebuild_resources_empty(self):
mgr = MCPClientManager({})
mgr._per_server_resources = {}
mgr._rebuild_resources()
assert mgr._resources == []
assert mgr._resource_map == {}
def test_rebuild_resources_multi_server(self):
mgr = MCPClientManager({})
mgr._per_server_resources = {
"fs": [_fake_resource_dict("file:///a", server="fs")],
"db": [_fake_resource_dict("db://table", name="table", server="db")],
}
mgr._rebuild_resources()
assert len(mgr._resources) == 2
assert mgr._resource_map["file:///a"] == ("fs", "file:///a")
assert mgr._resource_map["db://table"] == ("db", "db://table")
def test_template_prefix_matching(self):
"""Expanded URI matches template by prefix."""
mgr = MCPClientManager({})
mgr._per_server_resources = {
"db": [
{
"uri": "db://tables/{table}/rows/{id}",
"name": "row",
"description": "A row",
"mimeType": "application/json",
"server": "db",
"template": True,
},
],
}
mgr._rebuild_resources()
# Template should not be in resource_map
assert "db://tables/{table}/rows/{id}" not in mgr._resource_map
# But prefix matching should find it
result = mgr._match_template("db://tables/users/rows/1")
assert result is not None
server, template_uri = result
assert server == "db"
assert template_uri == "db://tables/{table}/rows/{id}"
def test_template_longest_prefix_wins(self):
"""When two templates have overlapping prefixes, the longer one wins."""
mgr = MCPClientManager({})
# Use templates with genuinely different prefix lengths:
# "db://data/" (6 chars after scheme) vs "db://data/tables/" (13 chars after scheme)
mgr._per_server_resources = {
"short": [
{
"uri": "db://data/{collection}",
"name": "collection",
"description": "",
"mimeType": "",
"server": "short",
"template": True,
},
],
"long": [
{
"uri": "db://data/tables/{table}",
"name": "table",
"description": "",
"mimeType": "",
"server": "long",
"template": True,
},
],
}
mgr._rebuild_resources()
# "db://data/tables/users" matches both prefixes ("db://data/" and
# "db://data/tables/") — the longer one should win
result = mgr._match_template("db://data/tables/users")
assert result is not None
server, template_uri = result
assert server == "long"
assert template_uri == "db://data/tables/{table}"
# URI that only matches the short prefix
result2 = mgr._match_template("db://data/views/active")
assert result2 is not None
assert result2[0] == "short"
def test_template_no_match_raises(self):
"""Completely unrelated URI still raises ValueError."""
mgr = MCPClientManager({})
mgr._per_server_resources = {
"db": [
{
"uri": "db://tables/{table}",
"name": "table",
"description": "",
"mimeType": "",
"server": "db",
"template": True,
},
],
}
mgr._rebuild_resources()
assert mgr._match_template("file:///something") is None
with pytest.raises(ValueError, match="Unknown MCP resource"):
mgr.read_resource_sync("file:///something")
def test_read_resource_sync_with_template_uri(self):
"""End-to-end: template discovered, expanded URI dispatched to correct server."""
mgr = MCPClientManager({})
mgr._per_server_resources = {
"db": [
{
"uri": "db://tables/{table}/rows/{id}",
"name": "row",
"description": "A row",
"mimeType": "application/json",
"server": "db",
"template": True,
},
],
}
mgr._rebuild_resources()
mock_session = MagicMock()
mgr._sessions["db"] = mock_session
mgr._loop = asyncio.new_event_loop()
text_content = MagicMock(spec=["text"])
text_content.text = '{"name": "Alice"}'
mock_result = MagicMock()
mock_result.contents = [text_content]
mock_session.read_resource = AsyncMock(return_value=mock_result)
thread = None
try:
thread = __import__("threading").Thread(target=mgr._loop.run_forever, daemon=True)
thread.start()
output = mgr.read_resource_sync("db://tables/users/rows/1", timeout=5)
assert output == '{"name": "Alice"}'
mock_session.read_resource.assert_awaited_once_with("db://tables/users/rows/1")
finally:
mgr._loop.call_soon_threadsafe(mgr._loop.stop)
if thread:
thread.join(timeout=5)
mgr._loop.close()
# ---------------------------------------------------------------------------
# MCP Prompts
# ---------------------------------------------------------------------------
class TestMCPPrompts:
def test_prompt_discovery(self):
"""Mock list_prompts(), verify get_prompts() with correct prefixed names."""
mgr = MCPClientManager({})
mgr._per_server_prompts = {
"tmpl": [
_fake_prompt_dict("mcp__tmpl__code_review", "code_review", "tmpl"),
_fake_prompt_dict("mcp__tmpl__summarize", "summarize", "tmpl"),
],
}
mgr._rebuild_prompts()
prompts = mgr.get_prompts()
assert len(prompts) == 2
names = {p["name"] for p in prompts}
assert names == {"mcp__tmpl__code_review", "mcp__tmpl__summarize"}
# Verify map entries
assert mgr._prompt_map["mcp__tmpl__code_review"] == ("tmpl", "code_review")
assert mgr._prompt_map["mcp__tmpl__summarize"] == ("tmpl", "summarize")
def test_rebuild_prompts_copy_on_write(self):
"""Verify mutation safety."""
mgr = MCPClientManager({})
mgr._per_server_prompts = {
"a": [_fake_prompt_dict("mcp__a__p1", "p1", "a")],
}
mgr._rebuild_prompts()
old_prompts = mgr._prompts
old_map = mgr._prompt_map
mgr._per_server_prompts["b"] = [_fake_prompt_dict("mcp__b__p2", "p2", "b")]
mgr._rebuild_prompts()
assert mgr._prompts is not old_prompts
assert mgr._prompt_map is not old_map
def test_get_prompts_returns_copy(self):
mgr = MCPClientManager({})
mgr._per_server_prompts = {
"a": [_fake_prompt_dict("mcp__a__p1", "p1", "a")],
}
mgr._rebuild_prompts()
prompts = mgr.get_prompts()
assert len(prompts) == 1
prompts.clear()
assert len(mgr.get_prompts()) == 1
def test_get_prompt_sync(self):
"""Mock session.get_prompt(), verify message conversion."""
mgr = MCPClientManager({})
mgr._prompt_map = {"mcp__tmpl__review": ("tmpl", "review")}
mock_session = MagicMock()
mgr._sessions["tmpl"] = mock_session
mgr._loop = asyncio.new_event_loop()
# Build mock PromptMessage
msg1 = MagicMock()
msg1.role = "user"
msg1.content = MagicMock()
msg1.content.text = "Review this code"
msg2 = MagicMock()
msg2.role = "assistant"
msg2.content = MagicMock()
msg2.content.text = "Looks good!"
mock_result = MagicMock()
mock_result.messages = [msg1, msg2]
mock_session.get_prompt = AsyncMock(return_value=mock_result)
thread = None
try:
thread = __import__("threading").Thread(target=mgr._loop.run_forever, daemon=True)
thread.start()
messages = mgr.get_prompt_sync(
"mcp__tmpl__review", arguments={"language": "python"}, timeout=5
)
assert len(messages) == 2
assert messages[0] == {"role": "user", "content": "Review this code"}
assert messages[1] == {"role": "assistant", "content": "Looks good!"}
mock_session.get_prompt.assert_awaited_once_with(
"review", arguments={"language": "python"}
)
finally:
mgr._loop.call_soon_threadsafe(mgr._loop.stop)
if thread:
thread.join(timeout=5)
mgr._loop.close()
def test_get_prompt_sync_unknown(self):
mgr = MCPClientManager({})
with pytest.raises(ValueError, match="Unknown MCP prompt"):
mgr.get_prompt_sync("mcp__no__such")
def test_get_prompt_sync_disconnected(self):
mgr = MCPClientManager({})
mgr._prompt_map = {"mcp__dead__p": ("dead", "p")}
with pytest.raises(RuntimeError, match="not connected"):
mgr.get_prompt_sync("mcp__dead__p")
def test_get_prompt_sync_timeout(self):
"""Verify timeout handling."""
mgr = MCPClientManager({})
mgr._prompt_map = {"mcp__tmpl__slow": ("tmpl", "slow")}
mock_session = MagicMock()
mgr._sessions["tmpl"] = mock_session
mgr._loop = asyncio.new_event_loop()
async def _slow_prompt(_name: str, *, arguments: dict[str, str] | None = None) -> None:
await asyncio.sleep(10)
mock_session.get_prompt = _slow_prompt
thread = None
try:
thread = __import__("threading").Thread(target=mgr._loop.run_forever, daemon=True)
thread.start()
with pytest.raises(TimeoutError):
mgr.get_prompt_sync("mcp__tmpl__slow", timeout=1)
finally:
mgr._loop.call_soon_threadsafe(mgr._loop.stop)
if thread:
thread.join(timeout=5)
mgr._loop.close()
def test_prompt_listener_notification(self):
"""Verify callback fires on rebuild."""
mgr = MCPClientManager({})
calls: list[int] = []
mgr.add_prompt_listener(lambda: calls.append(1))
mgr._per_server_prompts = {"a": [_fake_prompt_dict()]}
mgr._rebuild_prompts()
assert len(calls) == 1
def test_prompt_listener_remove(self):
mgr = MCPClientManager({})
calls: list[int] = []
cb = lambda: calls.append(1) # noqa: E731
mgr.add_prompt_listener(cb)
mgr.remove_prompt_listener(cb)
mgr._rebuild_prompts()
assert calls == []
def test_prompt_listener_error_does_not_propagate(self):
mgr = MCPClientManager({})
mgr.add_prompt_listener(lambda: 1 / 0)
mgr._rebuild_prompts() # should not raise
def test_is_mcp_prompt(self):
"""Verify name lookup."""
mgr = MCPClientManager({})
mgr._prompt_map["mcp__tmpl__review"] = ("tmpl", "review")
assert mgr.is_mcp_prompt("mcp__tmpl__review") is True
assert mgr.is_mcp_prompt("nonexistent") is False
def test_prompt_refresh_on_notification(self):
"""Mock notification, verify re-fetch of prompts."""
async def _run() -> None:
mgr = MCPClientManager({})
mock_session = MagicMock()
mgr._sessions["tmpl"] = mock_session
mgr._supports_prompts["tmpl"] = True
# Initial state
mgr._per_server_prompts["tmpl"] = [
_fake_prompt_dict("mcp__tmpl__old", "old", "tmpl"),
]
mgr._rebuild_prompts()
assert len(mgr.get_prompts()) == 1
# Mock re-fetch returning a new prompt
new_prompt = _fake_mcp_prompt("new_prompt", "A new prompt")
mock_prompt_result = MagicMock()
mock_prompt_result.prompts = [new_prompt]
mock_session.list_prompts = AsyncMock(return_value=mock_prompt_result)
await mgr._refresh_server_prompts("tmpl")
prompts = mgr.get_prompts()
assert len(prompts) == 1
assert prompts[0]["name"] == "mcp__tmpl__new_prompt"
assert prompts[0]["original_name"] == "new_prompt"
asyncio.run(_run())
def test_rebuild_prompts_empty(self):
mgr = MCPClientManager({})
mgr._per_server_prompts = {}
mgr._rebuild_prompts()
assert mgr._prompts == []
assert mgr._prompt_map == {}
def test_rebuild_prompts_multi_server(self):
mgr = MCPClientManager({})
mgr._per_server_prompts = {
"a": [_fake_prompt_dict("mcp__a__p1", "p1", "a")],
"b": [_fake_prompt_dict("mcp__b__p2", "p2", "b")],
}
mgr._rebuild_prompts()
assert len(mgr._prompts) == 2
assert mgr._prompt_map["mcp__a__p1"] == ("a", "p1")
assert mgr._prompt_map["mcp__b__p2"] == ("b", "p2")
# ---------------------------------------------------------------------------
# Shutdown cleans up new state
# ---------------------------------------------------------------------------
class TestShutdownCleanup:
def test_shutdown_clears_resources_and_prompts(self):
mgr = MCPClientManager({})
mgr._per_server_resources = {"a": [_fake_resource_dict()]}
mgr._rebuild_resources()
mgr._per_server_prompts = {"a": [_fake_prompt_dict()]}
mgr._rebuild_prompts()
assert mgr.get_resources() != []
assert mgr.get_prompts() != []
mgr.shutdown()
assert mgr.get_resources() == []
assert mgr.get_prompts() == []
assert mgr._resource_map == {}
assert mgr._prompt_map == {}
+397
View File
@@ -0,0 +1,397 @@
"""Integration tests for MCPClientManager data flow.
Uses real storage (SQLite) and real MCPClientManager state manipulation,
but mock MCP sessions instead of wire-protocol connections. This validates
the full data pipeline: per-server data -> rebuild -> merged state ->
query methods -> storage sync -> shutdown cleanup.
"""
from __future__ import annotations
import asyncio
from typing import Any
from unittest.mock import AsyncMock, MagicMock
import pytest
from turnstone.core.mcp_client import MCPClientManager
from turnstone.core.storage._sqlite import SQLiteBackend
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _make_resource(
uri: str, name: str, server: str, description: str = "", mime: str = "text/plain"
) -> dict[str, Any]:
return {
"uri": uri,
"name": name,
"description": description,
"mimeType": mime,
"server": server,
}
def _make_prompt(
prefixed_name: str,
original_name: str,
server: str,
description: str = "",
arguments: list[dict[str, Any]] | None = None,
) -> dict[str, Any]:
return {
"name": prefixed_name,
"original_name": original_name,
"server": server,
"description": description,
"arguments": arguments or [],
}
def _make_mock_session(
read_resource_result: Any = None,
get_prompt_result: Any = None,
) -> AsyncMock:
"""Build a mock ClientSession with configurable async return values."""
session = AsyncMock()
if read_resource_result is not None:
session.read_resource.return_value = read_resource_result
else:
# Default: single text content
content_item = MagicMock()
content_item.text = "resource content"
result = MagicMock()
result.contents = [content_item]
session.read_resource.return_value = result
if get_prompt_result is not None:
session.get_prompt.return_value = get_prompt_result
else:
msg = MagicMock()
msg.role = "user"
msg.content = MagicMock()
msg.content.text = "Hello, World!"
result = MagicMock()
result.messages = [msg]
session.get_prompt.return_value = result
return session
# ---------------------------------------------------------------------------
# Integration test class
# ---------------------------------------------------------------------------
class TestFullLifecycleResourcesPrompts:
"""Integration test exercising real code paths with real SQLite storage
but mock MCP sessions.
Validates the complete data flow: per-server data population, rebuild
merging, query methods, resource/prompt dispatch through asyncio, storage
sync, and shutdown cleanup.
"""
@pytest.fixture()
def mgr(self) -> MCPClientManager:
"""Create an MCPClientManager with no server configs (no start())."""
return MCPClientManager({})
@pytest.fixture()
def db(self, tmp_path) -> SQLiteBackend:
"""Create a fresh SQLite backend for each test."""
backend = SQLiteBackend(str(tmp_path / "test.db"))
yield backend
backend.close()
def test_rebuild_resources_produces_merged_state(self, mgr: MCPClientManager) -> None:
"""_rebuild_resources merges per-server resources into a unified list."""
mgr._per_server_resources["alpha"] = [
_make_resource("file:///a.txt", "a", "alpha"),
_make_resource("file:///b.txt", "b", "alpha"),
]
mgr._per_server_resources["beta"] = [
_make_resource("file:///c.txt", "c", "beta"),
]
mgr._rebuild_resources()
resources = mgr.get_resources()
assert len(resources) == 3
uris = {r["uri"] for r in resources}
assert uris == {"file:///a.txt", "file:///b.txt", "file:///c.txt"}
# resource_map should have entries for all non-template resources
assert "file:///a.txt" in mgr._resource_map
assert "file:///c.txt" in mgr._resource_map
assert mgr.resource_count == 3
def test_rebuild_prompts_produces_merged_state(self, mgr: MCPClientManager) -> None:
"""_rebuild_prompts merges per-server prompts into a unified list."""
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__greet", "greet", "alpha", "Say hello"),
]
mgr._per_server_prompts["beta"] = [
_make_prompt("mcp__beta__summarize", "summarize", "beta", "Summarize text"),
_make_prompt("mcp__beta__translate", "translate", "beta", "Translate text"),
]
mgr._rebuild_prompts()
prompts = mgr.get_prompts()
assert len(prompts) == 3
names = {p["name"] for p in prompts}
assert names == {"mcp__alpha__greet", "mcp__beta__summarize", "mcp__beta__translate"}
# prompt_map should map prefixed -> (server, original)
assert mgr._prompt_map["mcp__alpha__greet"] == ("alpha", "greet")
assert mgr._prompt_map["mcp__beta__summarize"] == ("beta", "summarize")
assert mgr.prompt_count == 3
assert mgr.is_mcp_prompt("mcp__alpha__greet") is True
assert mgr.is_mcp_prompt("nonexistent") is False
def test_read_resource_sync_dispatches_correctly(self, mgr: MCPClientManager) -> None:
"""read_resource_sync dispatches to the correct session via a real asyncio loop."""
# Set up a real event loop in a thread (simulating start())
loop = asyncio.new_event_loop()
import threading
thread = threading.Thread(target=loop.run_forever, daemon=True)
thread.start()
mgr._loop = loop
try:
# Populate session and resource map
session = _make_mock_session()
mgr._sessions["alpha"] = session
mgr._per_server_resources["alpha"] = [
_make_resource("file:///readme.md", "readme", "alpha"),
]
mgr._rebuild_resources()
result = mgr.read_resource_sync("file:///readme.md", timeout=5)
assert result == "resource content"
session.read_resource.assert_awaited_once_with("file:///readme.md")
finally:
loop.call_soon_threadsafe(loop.stop)
thread.join(timeout=5)
loop.close()
def test_read_resource_sync_unknown_uri_raises(self, mgr: MCPClientManager) -> None:
"""read_resource_sync raises ValueError for an unknown URI."""
with pytest.raises(ValueError, match="Unknown MCP resource"):
mgr.read_resource_sync("file:///nonexistent")
def test_read_resource_via_template(self, mgr: MCPClientManager) -> None:
"""Expanded template URI dispatched to correct server via real asyncio loop."""
loop = asyncio.new_event_loop()
import threading
thread = threading.Thread(target=loop.run_forever, daemon=True)
thread.start()
mgr._loop = loop
try:
session = _make_mock_session()
mgr._sessions["alpha"] = session
# Register a template resource (no concrete resources)
mgr._per_server_resources["alpha"] = [
{
"uri": "db://tables/{table}/rows/{id}",
"name": "row",
"description": "Fetch a row",
"mimeType": "application/json",
"server": "alpha",
"template": True,
},
]
mgr._rebuild_resources()
# Template should not be in _resource_map
assert "db://tables/{table}/rows/{id}" not in mgr._resource_map
# But expanded URI should resolve via prefix matching
result = mgr.read_resource_sync("db://tables/users/rows/42", timeout=5)
assert result == "resource content"
session.read_resource.assert_awaited_once_with("db://tables/users/rows/42")
finally:
loop.call_soon_threadsafe(loop.stop)
thread.join(timeout=5)
loop.close()
def test_get_prompt_sync_dispatches_correctly(self, mgr: MCPClientManager) -> None:
"""get_prompt_sync dispatches to the correct session via a real asyncio loop."""
loop = asyncio.new_event_loop()
import threading
thread = threading.Thread(target=loop.run_forever, daemon=True)
thread.start()
mgr._loop = loop
try:
session = _make_mock_session()
mgr._sessions["alpha"] = session
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__greet", "greet", "alpha", "Say hello"),
]
mgr._rebuild_prompts()
messages = mgr.get_prompt_sync(
"mcp__alpha__greet", arguments={"name": "World"}, timeout=5
)
assert len(messages) == 1
assert messages[0]["role"] == "user"
assert messages[0]["content"] == "Hello, World!"
session.get_prompt.assert_awaited_once_with("greet", arguments={"name": "World"})
finally:
loop.call_soon_threadsafe(loop.stop)
thread.join(timeout=5)
loop.close()
def test_get_prompt_sync_unknown_name_raises(self, mgr: MCPClientManager) -> None:
"""get_prompt_sync raises ValueError for an unknown prompt name."""
with pytest.raises(ValueError, match="Unknown MCP prompt"):
mgr.get_prompt_sync("mcp__nosrv__nope")
def test_sync_prompts_to_storage_creates_templates(
self, mgr: MCPClientManager, db: SQLiteBackend
) -> None:
"""sync_prompts_to_storage creates governance templates in real SQLite."""
mgr.set_storage(db)
mgr._prompts = [
_make_prompt(
"mcp__alpha__greet",
"greet",
"alpha",
"Say hello",
[{"name": "user", "description": "Who to greet", "required": True}],
),
_make_prompt(
"mcp__beta__summarize",
"summarize",
"beta",
"Summarize text",
),
]
# Mark connected so set_storage triggers sync
mgr._connected.set()
# Re-set storage to trigger auto-sync
mgr.set_storage(db)
templates = db.list_prompt_templates()
assert len(templates) == 2
names = {t["name"] for t in templates}
assert names == {"mcp__alpha__greet", "mcp__beta__summarize"}
# Verify details on first template
tpl = db.get_prompt_template_by_name("mcp__alpha__greet")
assert tpl is not None
assert tpl["origin"] == "mcp"
assert tpl["mcp_server"] == "alpha"
assert tpl["readonly"] is True
assert tpl["category"] == "mcp"
assert "user" in tpl["variables"]
def test_sync_prompts_removes_stale_templates(
self, mgr: MCPClientManager, db: SQLiteBackend
) -> None:
"""sync_prompts_to_storage removes templates whose MCP prompts are gone."""
mgr.set_storage(db)
# Create an initial template via sync
mgr._prompts = [
_make_prompt("mcp__alpha__old", "old", "alpha", "Old prompt"),
]
mgr.sync_prompts_to_storage()
assert len(db.list_prompt_templates()) == 1
# Now the prompt is gone
mgr._prompts = []
result = mgr.sync_prompts_to_storage()
assert result["removed"] == ["mcp__alpha__old"]
assert len(db.list_prompt_templates()) == 0
def test_shutdown_clears_all_state(self, mgr: MCPClientManager) -> None:
"""shutdown() clears sessions, tools, resources, prompts, and listeners."""
# Populate state
mgr._sessions["alpha"] = MagicMock()
mgr._per_server_tools["alpha"] = [
{
"type": "function",
"function": {
"name": "mcp__alpha__search",
"description": "Search",
"parameters": {},
},
}
]
mgr._rebuild_tools()
mgr._per_server_resources["alpha"] = [
_make_resource("file:///a.txt", "a", "alpha"),
{
"uri": "db://tables/{table}",
"name": "table",
"description": "",
"mimeType": "",
"server": "alpha",
"template": True,
},
]
mgr._rebuild_resources()
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__greet", "greet", "alpha"),
]
mgr._rebuild_prompts()
mgr._listeners.append(lambda: None)
mgr._resource_listeners.append(lambda: None)
mgr._prompt_listeners.append(lambda: None)
# Verify populated
assert len(mgr._sessions) == 1
assert len(mgr._tools) == 1
assert len(mgr._resources) == 2 # 1 concrete + 1 template
assert len(mgr._template_prefixes) == 1
assert len(mgr._prompts) == 1
mgr.shutdown()
assert len(mgr._sessions) == 0
assert len(mgr._tools) == 0
assert len(mgr._tool_map) == 0
assert len(mgr._resources) == 0
assert len(mgr._resource_map) == 0
assert len(mgr._template_prefixes) == 0
assert len(mgr._prompts) == 0
assert len(mgr._prompt_map) == 0
assert len(mgr._listeners) == 0
assert len(mgr._resource_listeners) == 0
assert len(mgr._prompt_listeners) == 0
def test_listener_notifications_fire_on_rebuild(self, mgr: MCPClientManager) -> None:
"""Rebuild methods fire the appropriate listener callbacks."""
tool_fired = []
resource_fired = []
prompt_fired = []
mgr.add_listener(lambda: tool_fired.append(1))
mgr.add_resource_listener(lambda: resource_fired.append(1))
mgr.add_prompt_listener(lambda: prompt_fired.append(1))
mgr._per_server_tools["alpha"] = []
mgr._rebuild_tools()
assert len(tool_fired) == 1
mgr._per_server_resources["alpha"] = [
_make_resource("file:///x.txt", "x", "alpha"),
]
mgr._rebuild_resources()
assert len(resource_fired) == 1
mgr._per_server_prompts["alpha"] = [
_make_prompt("mcp__alpha__p1", "p1", "alpha"),
]
mgr._rebuild_prompts()
assert len(prompt_fired) == 1
# Tool and resource listeners should not have been fired again
assert len(tool_fired) == 1
assert len(resource_fired) == 1
+272
View File
@@ -0,0 +1,272 @@
"""Tests for MCP prompt → governance template sync and readonly API guards."""
from __future__ import annotations
from unittest.mock import MagicMock
import pytest
from turnstone.core.mcp_client import MCPClientManager
@pytest.fixture()
def mgr() -> MCPClientManager:
"""Create an MCPClientManager with no real servers (no start())."""
return MCPClientManager({})
def _make_storage() -> MagicMock:
"""Create a mock storage backend with prompt template methods."""
storage = MagicMock()
storage.get_prompt_template_by_name.return_value = None
storage.list_prompt_templates_by_origin.return_value = []
storage.create_prompt_template.return_value = None
storage.update_prompt_template.return_value = True
storage.delete_prompt_template.return_value = True
return storage
class TestSyncPromptsToStorage:
def test_sync_no_storage(self, mgr: MCPClientManager) -> None:
"""Without storage set, sync returns empty stats."""
result = mgr.sync_prompts_to_storage()
assert result == {"added": [], "removed": [], "skipped": []}
def test_sync_creates_mcp_templates(self, mgr: MCPClientManager) -> None:
"""New MCP prompts are created as templates."""
storage = _make_storage()
mgr.set_storage(storage)
# Populate internal prompts list directly
mgr._prompts = [
{
"name": "mcp__test__greeting",
"original_name": "greeting",
"server": "test",
"description": "Say hello",
"arguments": [
{"name": "name", "description": "Who to greet", "required": True},
],
},
]
result = mgr.sync_prompts_to_storage()
assert result["added"] == ["mcp__test__greeting"]
assert result["removed"] == []
assert result["skipped"] == []
storage.create_prompt_template.assert_called_once()
call_kwargs = storage.create_prompt_template.call_args
assert call_kwargs[1]["name"] == "mcp__test__greeting"
assert call_kwargs[1]["origin"] == "mcp"
assert call_kwargs[1]["mcp_server"] == "test"
assert call_kwargs[1]["readonly"] is True
assert call_kwargs[1]["category"] == "mcp"
assert '"name"' in call_kwargs[1]["variables"]
def test_sync_skips_manual_overrides(self, mgr: MCPClientManager) -> None:
"""A manual template with the same name is not overwritten."""
storage = _make_storage()
storage.get_prompt_template_by_name.return_value = {
"template_id": "existing-id",
"name": "mcp__test__greeting",
"origin": "manual",
"readonly": False,
}
mgr.set_storage(storage)
mgr._prompts = [
{
"name": "mcp__test__greeting",
"original_name": "greeting",
"server": "test",
"description": "Say hello",
"arguments": [],
},
]
result = mgr.sync_prompts_to_storage()
assert result["skipped"] == ["mcp__test__greeting"]
assert result["added"] == []
storage.create_prompt_template.assert_not_called()
storage.update_prompt_template.assert_not_called()
def test_sync_updates_existing_mcp_template(self, mgr: MCPClientManager) -> None:
"""An existing MCP template gets its content/variables updated."""
storage = _make_storage()
storage.get_prompt_template_by_name.return_value = {
"template_id": "existing-id",
"name": "mcp__test__greeting",
"origin": "mcp",
"mcp_server": "test",
"readonly": True,
}
mgr.set_storage(storage)
mgr._prompts = [
{
"name": "mcp__test__greeting",
"original_name": "greeting",
"server": "test",
"description": "Updated description",
"arguments": [
{"name": "user", "description": "The user", "required": False},
],
},
]
result = mgr.sync_prompts_to_storage()
assert result["added"] == []
assert result["skipped"] == []
storage.create_prompt_template.assert_not_called()
storage.update_prompt_template.assert_called_once()
call_args = storage.update_prompt_template.call_args
assert call_args[0][0] == "existing-id"
assert "Updated description" in call_args[1]["content"]
assert "user" in call_args[1]["variables"]
# Security: is_default must be reset to prevent compromised MCP server
# from injecting content into a previously admin-promoted default
assert call_args[1]["is_default"] is False
def test_sync_resets_is_default_on_promoted_template(self, mgr: MCPClientManager) -> None:
"""An MCP template promoted to default by admin gets is_default reset on sync."""
storage = _make_storage()
storage.get_prompt_template_by_name.return_value = {
"template_id": "promoted-id",
"name": "mcp__test__greeting",
"origin": "mcp",
"mcp_server": "test",
"readonly": True,
"is_default": True, # admin toggled this
}
mgr.set_storage(storage)
mgr._prompts = [
{
"name": "mcp__test__greeting",
"original_name": "greeting",
"server": "test",
"description": "Potentially compromised content",
"arguments": [],
},
]
mgr.sync_prompts_to_storage()
call_args = storage.update_prompt_template.call_args
assert call_args[1]["is_default"] is False
def test_sync_removes_deleted_prompts(self, mgr: MCPClientManager) -> None:
"""MCP templates in storage with no matching prompt are deleted."""
storage = _make_storage()
storage.list_prompt_templates_by_origin.return_value = [
{
"template_id": "old-id",
"name": "mcp__test__old_prompt",
"origin": "mcp",
"mcp_server": "test",
},
]
mgr.set_storage(storage)
mgr._prompts = [] # No prompts at all
result = mgr.sync_prompts_to_storage()
assert result["removed"] == ["mcp__test__old_prompt"]
storage.delete_prompt_template.assert_called_once_with("old-id")
class TestSetStorageAutoSync:
"""set_storage() triggers an immediate sync when servers are already connected."""
def test_set_storage_syncs_when_connected(self, mgr) -> None:
storage = _make_storage()
mgr._prompts = [
{
"name": "mcp__srv__p1",
"original_name": "p1",
"server": "srv",
"description": "A prompt",
"arguments": [],
}
]
mgr._connected.set()
mgr.set_storage(storage)
# Should have called create_prompt_template for the discovered prompt
storage.create_prompt_template.assert_called_once()
call_kwargs = storage.create_prompt_template.call_args
assert call_kwargs[1]["name"] == "mcp__srv__p1"
assert call_kwargs[1]["origin"] == "mcp"
def test_set_storage_no_sync_when_not_connected(self, mgr) -> None:
storage = _make_storage()
mgr._prompts = [
{
"name": "mcp__srv__p1",
"original_name": "p1",
"server": "srv",
"description": "A prompt",
"arguments": [],
}
]
# _connected is NOT set
mgr.set_storage(storage)
# Should not have synced
storage.create_prompt_template.assert_not_called()
class TestReadonlyAPIGuards:
"""Test that the console server API guards reject edits to readonly templates."""
@pytest.fixture()
def db(self, tmp_path):
"""Create a fresh SQLite backend for each test."""
from turnstone.core.storage._sqlite import SQLiteBackend
return SQLiteBackend(str(tmp_path / "test.db"))
def test_readonly_guard_update(self, db) -> None:
"""Readonly templates cannot be updated via storage guard logic."""
db.create_prompt_template(
"t1",
"mcp__srv__prompt",
"mcp",
"content",
variables="[]",
is_default=False,
org_id="",
created_by="",
origin="mcp",
mcp_server="srv",
readonly=True,
)
tpl = db.get_prompt_template("t1")
assert tpl is not None
assert tpl["readonly"] is True
# Simulate API guard check
assert tpl.get("readonly") is True
def test_readonly_guard_delete(self, db) -> None:
"""Readonly templates are flagged for API-level rejection."""
db.create_prompt_template(
"t1",
"mcp__srv__prompt",
"mcp",
"content",
variables="[]",
is_default=False,
org_id="",
created_by="",
origin="mcp",
mcp_server="srv",
readonly=True,
)
existing = db.get_prompt_template("t1")
assert existing is not None
assert existing.get("readonly") is True
+393
View File
@@ -0,0 +1,393 @@
"""Tests for prompt template runtime wiring into ChatSession."""
from __future__ import annotations
from unittest.mock import MagicMock
from turnstone.core.session import ChatSession, _render_template
class NullUI:
"""UI adapter that discards all output."""
def on_thinking_start(self):
pass
def on_thinking_stop(self):
pass
def on_reasoning_token(self, text):
pass
def on_content_token(self, text):
pass
def on_stream_end(self):
pass
def approve_tools(self, items):
return True, None
def on_tool_result(self, call_id, name, output):
pass
def on_tool_output_chunk(self, call_id, chunk):
pass
def on_status(self, usage, context_window, effort):
pass
def on_plan_review(self, content):
return ""
def on_info(self, message):
pass
def on_error(self, message):
pass
def on_state_change(self, state):
pass
def on_rename(self, name):
pass
def _make_session(**kwargs):
defaults = dict(
client=MagicMock(),
model="test-model",
ui=NullUI(),
instructions=None,
temperature=0.5,
max_tokens=4096,
tool_timeout=30,
)
defaults.update(kwargs)
return ChatSession(**defaults)
def _sys_content(session: ChatSession) -> str:
"""Extract the system message content."""
msgs = [m for m in session.system_messages if m["role"] == "system"]
assert msgs
return msgs[0]["content"]
def _create_template(db, template_id, name, content, is_default=False, **kwargs):
"""Helper to create a prompt template in storage."""
db.create_prompt_template(
template_id=template_id,
name=name,
category=kwargs.get("category", "general"),
content=content,
variables=kwargs.get("variables", "[]"),
is_default=is_default,
org_id=kwargs.get("org_id", ""),
created_by=kwargs.get("created_by", "test"),
origin=kwargs.get("origin", "manual"),
mcp_server=kwargs.get("mcp_server", ""),
readonly=kwargs.get("readonly", False),
)
# ---------------------------------------------------------------------------
# _render_template unit tests
# ---------------------------------------------------------------------------
class TestRenderTemplate:
def test_basic_substitution(self):
result = _render_template("Hello {{name}}", {"name": "world"})
assert result == "Hello world"
def test_multiple_variables(self):
result = _render_template(
"Model: {{model}}, WS: {{ws_id}}", {"model": "gpt-5", "ws_id": "abc123"}
)
assert result == "Model: gpt-5, WS: abc123"
def test_unresolvable_variable_kept(self):
result = _render_template("Hello {{unknown}}", {"model": "gpt-5"})
assert result == "Hello {{unknown}}"
def test_empty_context(self):
result = _render_template("No vars here", {})
assert result == "No vars here"
def test_duplicate_placeholder(self):
result = _render_template("{{x}} and {{x}}", {"x": "val"})
assert result == "val and val"
def test_no_cross_variable_injection(self):
# If model contains {{ws_id}}, it must NOT be expanded
result = _render_template("Model: {{model}}", {"model": "{{ws_id}}", "ws_id": "secret"})
assert result == "Model: {{ws_id}}"
assert "secret" not in result
# ---------------------------------------------------------------------------
# Default templates in system message
# ---------------------------------------------------------------------------
class TestDefaultTemplates:
def test_default_templates_in_system_message(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "alpha", "You are a helpful assistant.", is_default=True)
_create_template(db, "t2", "beta", "Always be concise.", is_default=True)
session = _make_session()
content = _sys_content(session)
assert "You are a helpful assistant." in content
assert "Always be concise." in content
def test_default_templates_ordered_by_name(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t2", "b-template", "SECOND", is_default=True)
_create_template(db, "t1", "a-template", "FIRST", is_default=True)
session = _make_session()
content = _sys_content(session)
first_pos = content.index("FIRST")
second_pos = content.index("SECOND")
assert first_pos < second_pos
def test_no_default_templates(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "alpha", "Not default.", is_default=False)
session = _make_session()
content = _sys_content(session)
assert "Not default." not in content
def test_templates_before_instructions(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "tpl", "TEMPLATE_CONTENT", is_default=True)
session = _make_session(instructions="USER_INSTRUCTIONS")
content = _sys_content(session)
tpl_pos = content.index("TEMPLATE_CONTENT")
instr_pos = content.index("USER_INSTRUCTIONS")
assert tpl_pos < instr_pos
# ---------------------------------------------------------------------------
# Explicit template selection
# ---------------------------------------------------------------------------
class TestExplicitTemplate:
def test_explicit_template_replaces_defaults(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "default-tpl", "DEFAULT_CONTENT", is_default=True)
_create_template(db, "t2", "specific-tpl", "SPECIFIC_CONTENT", is_default=False)
session = _make_session(template="specific-tpl")
content = _sys_content(session)
assert "SPECIFIC_CONTENT" in content
assert "DEFAULT_CONTENT" not in content
def test_explicit_template_not_found(self, tmp_db):
session = _make_session(template="nonexistent")
content = _sys_content(session)
# Graceful degradation — no template content injected
assert "nonexistent" not in content
# ---------------------------------------------------------------------------
# Variable substitution in templates
# ---------------------------------------------------------------------------
class TestTemplateVariables:
def test_model_and_ws_id_substituted(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "vars-tpl", "Model: {{model}}, WS: {{ws_id}}", is_default=True)
session = _make_session()
content = _sys_content(session)
assert "Model: test-model" in content
assert f"WS: {session.ws_id}" in content
def test_node_id_substituted(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "node-tpl", "Node: {{node_id}}", is_default=True)
session = _make_session(node_id="node-42")
content = _sys_content(session)
assert "Node: node-42" in content
def test_unknown_variable_preserved(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "unknown-tpl", "Val: {{unknown_var}}", is_default=True)
session = _make_session()
content = _sys_content(session)
assert "Val: {{unknown_var}}" in content
# ---------------------------------------------------------------------------
# Template persistence and resume
# ---------------------------------------------------------------------------
class TestTemplatePersistence:
def test_template_persisted_in_config(self, tmp_db):
from turnstone.core.memory import load_workstream_config
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "my-tpl", "TPL_CONTENT", is_default=False)
session = _make_session(template="my-tpl")
config = load_workstream_config(session.ws_id)
assert config["template"] == "my-tpl"
def test_template_restored_on_resume(self, tmp_db):
from turnstone.core.memory import save_message
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "my-tpl", "PERSISTED_TEMPLATE", is_default=False)
# Create session with template, save a message so resume has history
session1 = _make_session(template="my-tpl")
ws_id = session1.ws_id
save_message(ws_id, "user", "hello")
# New session without template, then resume
session2 = _make_session()
assert session2._template_name is None
resumed = session2.resume(ws_id)
assert resumed
assert session2._template_name == "my-tpl"
content = _sys_content(session2)
assert "PERSISTED_TEMPLATE" in content
def test_empty_template_config_means_defaults(self, tmp_db):
from turnstone.core.memory import load_workstream_config
session = _make_session()
config = load_workstream_config(session.ws_id)
assert config["template"] == ""
# ---------------------------------------------------------------------------
# /template slash command
# ---------------------------------------------------------------------------
class TestTemplateSlashCommand:
def test_template_set(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "my-tpl", "SLASH_TEMPLATE", is_default=False)
session = _make_session()
content_before = _sys_content(session)
assert "SLASH_TEMPLATE" not in content_before
session.handle_command("/template my-tpl")
assert session._template_name == "my-tpl"
content_after = _sys_content(session)
assert "SLASH_TEMPLATE" in content_after
def test_template_clear(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "my-tpl", "EXPLICIT_TEMPLATE", is_default=False)
_create_template(db, "t2", "default-tpl", "DEFAULT_TEMPLATE", is_default=True)
session = _make_session(template="my-tpl")
assert "EXPLICIT_TEMPLATE" in _sys_content(session)
assert "DEFAULT_TEMPLATE" not in _sys_content(session)
session.handle_command("/template clear")
assert session._template_name is None
assert "DEFAULT_TEMPLATE" in _sys_content(session)
assert "EXPLICIT_TEMPLATE" not in _sys_content(session)
def test_template_not_found(self, tmp_db):
ui = NullUI()
ui.on_error = MagicMock()
session = _make_session(ui=ui)
session.handle_command("/template nonexistent")
ui.on_error.assert_called_once()
assert "not found" in ui.on_error.call_args[0][0].lower()
def test_template_show_current(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(db, "t1", "my-tpl", "content", is_default=False)
ui = NullUI()
ui.on_info = MagicMock()
session = _make_session(ui=ui, template="my-tpl")
session.handle_command("/template")
ui.on_info.assert_called_once()
assert "my-tpl" in ui.on_info.call_args[0][0]
# ---------------------------------------------------------------------------
# MCP-origin templates
# ---------------------------------------------------------------------------
class TestMCPTemplates:
def test_mcp_readonly_template_as_default(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(
db,
"t1",
"mcp__server__prompt",
"MCP_CONTENT",
is_default=True,
origin="mcp",
mcp_server="server",
readonly=True,
)
session = _make_session()
content = _sys_content(session)
assert "MCP_CONTENT" in content
def test_mcp_template_selectable_explicitly(self, tmp_db):
from turnstone.core.storage import get_storage
db = get_storage()
_create_template(
db,
"t1",
"mcp__server__code",
"MCP_EXPLICIT",
is_default=False,
origin="mcp",
mcp_server="server",
readonly=True,
)
session = _make_session(template="mcp__server__code")
content = _sys_content(session)
assert "MCP_EXPLICIT" in content
+16
View File
@@ -8,6 +8,7 @@ from turnstone.mq.protocol import (
AckEvent,
ApprovalRequestEvent,
ApproveMessage,
CancelMessage,
CloseWorkstreamMessage,
CommandMessage,
ContentEvent,
@@ -68,6 +69,7 @@ INBOUND_TYPES = [
(ListWorkstreamsMessage, {}),
(HealthMessage, {}),
(ListNodesMessage, {}),
(CancelMessage, {"ws_id": "abc"}),
]
@@ -207,6 +209,20 @@ def test_create_workstream_target_node():
assert restored.name == "debug-ws"
def test_create_workstream_template_field():
msg = CreateWorkstreamMessage(name="ws", template="code-review")
assert msg.template == "code-review"
raw = msg.to_json()
restored = InboundMessage.from_json(raw)
assert isinstance(restored, CreateWorkstreamMessage)
assert restored.template == "code-review"
def test_create_workstream_template_default_empty():
msg = CreateWorkstreamMessage(name="ws")
assert msg.template == ""
def test_list_nodes_round_trip():
msg = ListNodesMessage()
raw = msg.to_json()
+135
View File
@@ -2049,3 +2049,138 @@ class TestModelCapabilitiesToolSearch:
caps = ModelCapabilities()
assert caps.supports_tool_search is False
# ---------------------------------------------------------------------------
# Vision support
# ---------------------------------------------------------------------------
class TestVisionCapabilities:
"""Test supports_vision flag across providers."""
def test_default_is_false(self) -> None:
from turnstone.core.providers._protocol import ModelCapabilities
caps = ModelCapabilities()
assert caps.supports_vision is False
def test_openai_commercial_supports_vision(self) -> None:
provider = OpenAIProvider()
for model in ("gpt-5", "gpt-5-mini", "gpt-5.4", "o3", "o4-mini"):
caps = provider.get_capabilities(model)
assert caps.supports_vision is True, f"{model} should support vision"
def test_openai_default_no_vision(self) -> None:
"""Unknown models (local servers) default to no vision."""
provider = OpenAIProvider()
caps = provider.get_capabilities("some-local-model")
assert caps.supports_vision is False
def test_anthropic_supports_vision(self) -> None:
from turnstone.core.providers._anthropic import AnthropicProvider
provider = AnthropicProvider()
for model in ("claude-opus-4-6", "claude-sonnet-4-6", "claude-haiku-4-5"):
caps = provider.get_capabilities(model)
assert caps.supports_vision is True, f"{model} should support vision"
def test_anthropic_default_supports_vision(self) -> None:
"""Anthropic default (unknown Claude model) supports vision."""
from turnstone.core.providers._anthropic import AnthropicProvider
provider = AnthropicProvider()
caps = provider.get_capabilities("claude-unknown-9")
assert caps.supports_vision is True
class TestAnthropicVisionConversion:
"""Test image content conversion in _convert_messages."""
def setup_method(self) -> None:
from turnstone.core.providers._anthropic import AnthropicProvider
self.provider = AnthropicProvider()
def test_tool_result_with_image_content(self) -> None:
"""Tool result with list content converts image_url to Anthropic image."""
messages = [
{"role": "user", "content": "Read this image"},
{
"role": "assistant",
"content": "",
"tool_calls": [
{
"id": "call_1",
"function": {"name": "read_file", "arguments": '{"path": "img.png"}'},
}
],
},
{
"role": "tool",
"tool_call_id": "call_1",
"content": [
{"type": "text", "text": "Image file: img.png (1024 bytes)"},
{
"type": "image_url",
"image_url": {"url": "data:image/png;base64,iVBORw0KGgo="},
},
],
},
]
_, converted = self.provider._convert_messages(messages)
# Tool result should be in a user message
tool_user_msg = converted[2]
assert tool_user_msg["role"] == "user"
tool_result = tool_user_msg["content"][0]
assert tool_result["type"] == "tool_result"
assert tool_result["tool_use_id"] == "call_1"
# Content should be a list with converted image block
content = tool_result["content"]
assert isinstance(content, list)
assert content[0] == {"type": "text", "text": "Image file: img.png (1024 bytes)"}
assert content[1]["type"] == "image"
assert content[1]["source"]["type"] == "base64"
assert content[1]["source"]["media_type"] == "image/png"
assert content[1]["source"]["data"] == "iVBORw0KGgo="
def test_tool_result_with_string_content_unchanged(self) -> None:
"""Tool result with plain string content is unchanged."""
messages = [
{"role": "user", "content": "Read file"},
{
"role": "assistant",
"content": "",
"tool_calls": [
{
"id": "call_2",
"function": {"name": "read_file", "arguments": '{"path": "f.py"}'},
}
],
},
{
"role": "tool",
"tool_call_id": "call_2",
"content": " 1\tprint('hello')",
},
]
_, converted = self.provider._convert_messages(messages)
tool_result = converted[2]["content"][0]
assert tool_result["content"] == " 1\tprint('hello')"
def test_convert_content_parts_static_method(self) -> None:
"""_convert_content_parts handles both image_url and text."""
from turnstone.core.providers._anthropic import AnthropicProvider
parts = [
{"type": "text", "text": "description"},
{
"type": "image_url",
"image_url": {"url": "data:image/jpeg;base64,/9j/4AAQ"},
},
]
result = AnthropicProvider._convert_content_parts(parts)
assert result[0] == {"type": "text", "text": "description"}
assert result[1]["type"] == "image"
assert result[1]["source"]["media_type"] == "image/jpeg"
assert result[1]["source"]["data"] == "/9j/4AAQ"
+21
View File
@@ -2,11 +2,19 @@
from __future__ import annotations
from typing import TYPE_CHECKING, Any
import pytest
from starlette.applications import Starlette
from starlette.middleware import Middleware
from starlette.middleware.base import BaseHTTPMiddleware
from starlette.routing import Mount, Route
from starlette.testclient import TestClient
if TYPE_CHECKING:
from starlette.requests import Request
from starlette.responses import Response
from turnstone.console.server import (
admin_create_schedule,
admin_delete_schedule,
@@ -15,9 +23,21 @@ from turnstone.console.server import (
admin_list_schedules,
admin_update_schedule,
)
from turnstone.core.auth import AuthResult
from turnstone.core.storage._sqlite import SQLiteBackend
class _InjectAuthMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request: Request, call_next: Any) -> Response:
request.state.auth_result = AuthResult(
user_id="test-admin",
scopes=frozenset({"approve"}),
token_source="config",
permissions=frozenset({"admin.schedules"}),
)
return await call_next(request)
@pytest.fixture
def storage(tmp_path):
"""Fresh SQLite backend for each test."""
@@ -52,6 +72,7 @@ def client(storage):
],
),
],
middleware=[Middleware(_InjectAuthMiddleware)],
)
app.state.auth_storage = storage
return TestClient(app)
+10
View File
@@ -1,6 +1,7 @@
"""Tests for turnstone.sdk.events — SSE event deserialization."""
from turnstone.sdk.events import (
ApprovalResolvedEvent,
ApproveRequestEvent,
BusyErrorEvent,
ClearUiEvent,
@@ -92,6 +93,15 @@ def test_approve_request_event():
assert len(e.items) == 1
def test_approval_resolved_event():
e = ServerEvent.from_dict(
{"type": "approval_resolved", "approved": False, "feedback": "Approval timed out"}
)
assert isinstance(e, ApprovalResolvedEvent)
assert e.approved is False
assert e.feedback == "Approval timed out"
def test_tool_result_event():
e = ServerEvent.from_dict(
{"type": "tool_result", "call_id": "c1", "name": "search", "output": "found it"}
+453 -11
View File
@@ -1,9 +1,10 @@
"""Tests for turnstone.core.session — ChatSession construction."""
import base64
import json
from unittest.mock import MagicMock, patch
from turnstone.core.session import ChatSession
from turnstone.core.session import _IMAGE_EXTENSIONS, _IMAGE_SIZE_CAP, ChatSession
class NullUI:
@@ -143,12 +144,21 @@ class TestChatSessionConstruction:
class TestPlanExec:
"""Tests for _exec_plan: unique session-scoped plan file and existing-plan injection."""
def _run_plan(self, session, prompt, agent_return="# Plan\n\nDo the thing."):
_VALID_PLAN = (
"## Goal\n\nDo the thing.\n\n"
"## Current State\n\nFile foo.py has bar().\n\n"
"## Plan\n\n1. Edit foo.py line 10.\n\n"
"## Risks\n\nNone."
)
def _run_plan(self, session, prompt, agent_return=None):
"""Invoke _exec_plan with _run_agent patched to avoid LLM calls.
Returns (call_id_returned, content_returned, captured_messages) where
captured_messages is the agent_messages list passed to _run_agent.
"""
if agent_return is None:
agent_return = self._VALID_PLAN
captured = {}
def fake_run_agent(messages, **kwargs):
@@ -174,10 +184,9 @@ class TestPlanExec:
"""Written plan file contains the agent's output verbatim."""
monkeypatch.chdir(tmp_path)
session = _make_session()
plan_content = "## Goal\n\nAdd a new endpoint."
self._run_plan(session, "add endpoint", agent_return=plan_content)
self._run_plan(session, "add endpoint")
plan_file = tmp_path / f".plan-{session._ws_id}.md"
assert plan_file.read_text() == plan_content
assert plan_file.read_text() == self._VALID_PLAN
def test_two_sessions_produce_different_files(self, tmp_db, tmp_path, monkeypatch):
"""Two ChatSession instances never collide on the same plan file."""
@@ -202,8 +211,8 @@ class TestPlanExec:
"id": tc_id,
"type": "function",
"function": {
"name": "plan",
"arguments": json.dumps({"prompt": prior_prompt}),
"name": "create_plan",
"arguments": json.dumps({"goal": prior_prompt}),
},
}
],
@@ -238,7 +247,7 @@ class TestPlanExec:
m for m in messages if m["role"] == "assistant" and m.get("tool_calls")
]
assert len(assistant_with_tc) == 1
assert assistant_with_tc[0]["tool_calls"][0]["function"]["name"] == "plan"
assert assistant_with_tc[0]["tool_calls"][0]["function"]["name"] == "create_plan"
# The real tool result is forwarded with its original content
tool_msgs = [m for m in messages if m["role"] == "tool"]
@@ -261,7 +270,440 @@ class TestPlanExec:
"""_exec_plan returns (call_id, agent_output)."""
monkeypatch.chdir(tmp_path)
session = _make_session()
agent_output = "## Goal\n\nBuild it."
call_id, content, _ = self._run_plan(session, "do stuff", agent_return=agent_output)
call_id, content, _ = self._run_plan(session, "do stuff")
assert call_id == "test-call-1"
assert content == agent_output
assert content == self._VALID_PLAN
def test_exec_plan_retries_on_garbage(self, tmp_db, tmp_path, monkeypatch):
"""When _run_agent returns garbage, _exec_plan retries once."""
monkeypatch.chdir(tmp_path)
session = _make_session()
good_plan = (
"## Goal\n\nAdd feature X.\n\n"
"## Current State\n\nFile foo.py has bar().\n\n"
"## Plan\n\n1. Edit foo.py:bar()\n\n"
"## Risks\n\nNone."
)
call_count = 0
def fake_run_agent(messages, **kwargs):
nonlocal call_count
call_count += 1
if call_count == 1:
return "Sure, do the thing."
return good_plan
item = {"call_id": "c1", "prompt": "add feature X"}
with patch.object(session, "_run_agent", side_effect=fake_run_agent):
_, content = session._exec_plan(item)
assert call_count == 2
assert "## Goal" in content
def test_exec_plan_warning_on_double_failure(self, tmp_db, tmp_path, monkeypatch):
"""When both attempts produce garbage, content gets a warning prefix."""
monkeypatch.chdir(tmp_path)
session = _make_session()
def fake_run_agent(messages, **kwargs):
return "nope"
item = {"call_id": "c1", "prompt": "add feature X"}
with patch.object(session, "_run_agent", side_effect=fake_run_agent):
_, content = session._exec_plan(item)
assert content.startswith("[Warning:")
def test_retry_continues_agent_conversation(self, tmp_db, tmp_path, monkeypatch):
"""Retry appends coaching to the same agent_messages list."""
monkeypatch.chdir(tmp_path)
session = _make_session()
captured_messages: list[list] = []
def fake_run_agent(messages, **kwargs):
captured_messages.append(list(messages))
if len(captured_messages) == 1:
return "garbage"
return (
"## Goal\n\nDone.\n\n## Current State\n\nx\n\n## Plan\n\n1. x\n\n## Risks\n\nNone."
)
item = {"call_id": "c1", "prompt": "add feature X"}
with patch.object(session, "_run_agent", side_effect=fake_run_agent):
session._exec_plan(item)
assert len(captured_messages) == 2
# Second call should have more messages (coaching appended)
assert len(captured_messages[1]) > len(captured_messages[0])
# Last user message in second call is the coaching message
assert "did not follow" in captured_messages[1][-1]["content"]
# ---------------------------------------------------------------------------
# Plan validation
# ---------------------------------------------------------------------------
class TestPlanValidation:
"""Tests for ChatSession._validate_plan quality gate."""
GOOD_PLAN = (
"## Goal\n\nAdd authentication to the API.\n\n"
"## Current State\n\nFile server.py:45 has no auth middleware.\n\n"
"## Plan\n\n1. Add AuthMiddleware to server.py.\n"
"2. Create auth.py with JWT verification.\n\n"
"## Risks\n\nToken expiry handling may need tuning."
)
def test_valid_plan_passes(self):
valid, issues = ChatSession._validate_plan(self.GOOD_PLAN, "add auth")
assert valid
assert issues == []
def test_too_short_fails(self):
valid, issues = ChatSession._validate_plan("Do the thing.", "do stuff")
assert not valid
assert any("too short" in i for i in issues)
def test_no_sections_fails(self):
content = "A" * 150 # long enough but no sections
valid, issues = ChatSession._validate_plan(content, "build it")
assert not valid
assert any("missing plan sections" in i for i in issues)
def test_echo_detection(self):
goal = "deliver a simpsons quote from a specific episode"
content = "Deliver a Simpsons quote from a specific episode"
valid, issues = ChatSession._validate_plan(content, goal)
assert not valid
assert any("echo" in i for i in issues)
def test_refusal_detection(self):
content = "I cannot create a plan for this task because " + "x" * 100
valid, issues = ChatSession._validate_plan(content, "do stuff")
assert not valid
assert any("refusal" in i for i in issues)
def test_partial_sections_passes(self):
"""2 out of 4 sections is enough to pass."""
content = (
"## Goal\n\nFix the bug in parsing.\n\n"
"## Plan\n\n1. Edit parser.py line 42.\n"
"2. Add boundary check.\n"
"This is enough detail to proceed with confidence."
)
valid, issues = ChatSession._validate_plan(content, "fix bug")
assert valid
def test_one_section_fails(self):
"""Only 1 out of 4 sections is not enough."""
content = (
"## Goal\n\nFix the bug.\n\n"
"We should probably edit parser.py and add some checks "
"to the boundary handling code path for safety."
)
valid, issues = ChatSession._validate_plan(content, "fix bug")
assert not valid
assert any("missing plan sections" in i for i in issues)
# ---------------------------------------------------------------------------
# Plan refinement loop
# ---------------------------------------------------------------------------
class TestPlanRefinement:
"""Tests for the iterative plan refinement loop in _execute_tools."""
GOOD_PLAN = TestPlanValidation.GOOD_PLAN
def test_feedback_triggers_refinement(self, tmp_db, tmp_path, monkeypatch):
"""User feedback causes _refine_plan to run, then approval exits."""
monkeypatch.chdir(tmp_path)
session = _make_session()
refine_called = []
review_responses = iter(["add error handling", ""])
session.ui = MagicMock(spec_set=NullUI)
session.ui.on_plan_review.side_effect = lambda c: next(review_responses)
session.ui.on_info = MagicMock()
session.ui.on_state_change = MagicMock()
revised = self.GOOD_PLAN + "\n\n3. Add error handling."
def fake_refine(content, goal, feedback):
refine_called.append(feedback)
return revised
with patch.object(session, "_refine_plan", side_effect=fake_refine):
items = [
{
"func_name": "create_plan",
"call_id": "c1",
"prompt": "add auth",
}
]
results = [("c1", self.GOOD_PLAN)]
# Manually invoke the post-plan gate portion of _execute_tools.
# We test the loop by calling the gate code directly.
session.auto_approve = False
original_goal = items[0].get("prompt", "")
output = results[0][1]
refinement_round = 0
while refinement_round < session._MAX_PLAN_REFINEMENTS:
resp = session.ui.on_plan_review(output)
if resp.lower() in ("n", "no", "reject"):
break
elif resp:
output = session._refine_plan(output, original_goal, resp)
refinement_round += 1
else:
break
assert len(refine_called) == 1
assert refine_called[0] == "add error handling"
assert "error handling" in output
def test_reject_skips_refinement(self, tmp_db, tmp_path, monkeypatch):
"""Rejection exits immediately without calling _refine_plan."""
monkeypatch.chdir(tmp_path)
session = _make_session()
session.ui = MagicMock(spec_set=NullUI)
session.ui.on_plan_review.return_value = "reject"
with patch.object(session, "_refine_plan") as mock_refine:
output = self.GOOD_PLAN
resp = session.ui.on_plan_review(output)
if resp.lower() in ("n", "no", "reject"):
output += "\n\n---\nUser REJECTED"
elif resp:
output = session._refine_plan(output, "g", resp)
mock_refine.assert_not_called()
assert "REJECTED" in output
def test_approve_skips_refinement(self, tmp_db, tmp_path, monkeypatch):
"""Empty response (enter) approves without refinement."""
monkeypatch.chdir(tmp_path)
session = _make_session()
session.ui = MagicMock(spec_set=NullUI)
session.ui.on_plan_review.return_value = ""
with patch.object(session, "_refine_plan") as mock_refine:
output = self.GOOD_PLAN
resp = session.ui.on_plan_review(output)
if resp.lower() in ("n", "no", "reject"):
output += "\n\n---\nUser REJECTED"
elif resp:
output = session._refine_plan(output, "g", resp)
mock_refine.assert_not_called()
assert "REJECTED" not in output
def test_max_refinement_rounds(self, tmp_db, tmp_path, monkeypatch):
"""Loop stops after _MAX_PLAN_REFINEMENTS rounds with a final review."""
monkeypatch.chdir(tmp_path)
session = _make_session()
session.ui = MagicMock(spec_set=NullUI)
session.ui.on_plan_review.return_value = "more detail please"
session.ui.on_info = MagicMock()
refine_count = 0
def fake_refine(content, goal, feedback):
nonlocal refine_count
refine_count += 1
return content + f"\n(revision {refine_count})"
with patch.object(session, "_refine_plan", side_effect=fake_refine):
output = self.GOOD_PLAN
original_goal = "add auth"
refinement_round = 0
while True:
resp = session.ui.on_plan_review(output)
if (
resp.lower() in ("n", "no", "reject")
or not resp
or refinement_round >= session._MAX_PLAN_REFINEMENTS
):
break
output = session._refine_plan(output, original_goal, resp)
refinement_round += 1
assert refine_count == session._MAX_PLAN_REFINEMENTS
# User gets one extra review call after max rounds (the final prompt)
assert session.ui.on_plan_review.call_count == session._MAX_PLAN_REFINEMENTS + 1
def test_refine_plan_message_structure(self, tmp_db, tmp_path, monkeypatch):
"""_refine_plan passes system + prior plan + feedback to _run_agent."""
monkeypatch.chdir(tmp_path)
session = _make_session()
captured = {}
def fake_run_agent(messages, **kwargs):
captured["messages"] = list(messages)
return self.GOOD_PLAN
with patch.object(session, "_run_agent", side_effect=fake_run_agent):
session._refine_plan(self.GOOD_PLAN, "add auth", "add tests too")
msgs = captured["messages"]
assert msgs[0]["role"] == "system"
assert msgs[1]["role"] == "assistant"
assert msgs[1]["tool_calls"][0]["function"]["name"] == "create_plan"
assert msgs[2]["role"] == "tool"
assert msgs[2]["content"] == self.GOOD_PLAN
assert msgs[3]["role"] == "user"
assert "add tests too" in msgs[3]["content"]
# ---------------------------------------------------------------------------
# Vision / image support
# ---------------------------------------------------------------------------
class TestImageExtensions:
"""Test _IMAGE_EXTENSIONS constant and detection logic."""
def test_common_image_extensions(self):
for ext in (".png", ".jpg", ".jpeg", ".gif", ".webp", ".bmp", ".tiff", ".tif", ".ico"):
assert ext in _IMAGE_EXTENSIONS, f"{ext} should be in _IMAGE_EXTENSIONS"
def test_svg_excluded(self):
assert ".svg" not in _IMAGE_EXTENSIONS
def test_text_extensions_excluded(self):
for ext in (".py", ".txt", ".json", ".md", ".rs", ".go"):
assert ext not in _IMAGE_EXTENSIONS
class TestExecReadImage:
"""Test _exec_read_image method."""
def _make_png(self, path: str, size: int = 100) -> None:
"""Write a minimal valid-ish PNG header to a file."""
# 8-byte PNG signature + enough bytes to reach target size
header = b"\x89PNG\r\n\x1a\n"
with open(path, "wb") as f:
f.write(header + b"\x00" * max(0, size - len(header)))
def test_image_returns_content_parts(self, tmp_db, tmp_path):
"""read_file on a PNG with vision support returns content parts."""
img = tmp_path / "test.png"
self._make_png(str(img))
session = _make_session()
# Mock provider to report vision support
mock_caps = MagicMock()
mock_caps.supports_vision = True
session._provider.get_capabilities = MagicMock(return_value=mock_caps)
item = {"call_id": "c1", "path": str(img), "offset": None, "limit": None}
call_id, output = session._exec_read_file(item)
assert call_id == "c1"
assert isinstance(output, list)
assert len(output) == 2
assert output[0]["type"] == "text"
assert "test.png" in output[0]["text"]
assert output[1]["type"] == "image_url"
url = output[1]["image_url"]["url"]
assert url.startswith("data:image/png;base64,")
# Verify base64 round-trip
b64part = url.split(",", 1)[1]
decoded = base64.b64decode(b64part)
assert decoded == img.read_bytes()
def test_no_vision_returns_text(self, tmp_db, tmp_path):
"""read_file on image with non-vision model returns text description."""
img = tmp_path / "photo.jpg"
self._make_png(str(img), size=2048)
session = _make_session()
mock_caps = MagicMock()
mock_caps.supports_vision = False
session._provider.get_capabilities = MagicMock(return_value=mock_caps)
item = {"call_id": "c2", "path": str(img), "offset": None, "limit": None}
call_id, output = session._exec_read_file(item)
assert call_id == "c2"
assert isinstance(output, str)
assert "does not support vision" in output
assert "photo.jpg" in output
def test_oversized_image_returns_error(self, tmp_db, tmp_path):
"""Images exceeding _IMAGE_SIZE_CAP return an error string."""
img = tmp_path / "huge.png"
# Write slightly over the cap
with open(img, "wb") as f:
f.write(b"\x89PNG\r\n\x1a\n" + b"\x00" * _IMAGE_SIZE_CAP)
session = _make_session()
mock_caps = MagicMock()
mock_caps.supports_vision = True
session._provider.get_capabilities = MagicMock(return_value=mock_caps)
item = {"call_id": "c3", "path": str(img), "offset": None, "limit": None}
call_id, output = session._exec_read_file(item)
assert call_id == "c3"
assert isinstance(output, str)
assert "exceeds" in output
def test_missing_image_returns_error(self, tmp_db, tmp_path):
"""read_file on non-existent image returns error."""
session = _make_session()
mock_caps = MagicMock()
mock_caps.supports_vision = True
session._provider.get_capabilities = MagicMock(return_value=mock_caps)
item = {"call_id": "c4", "path": str(tmp_path / "nope.png"), "offset": None, "limit": None}
call_id, output = session._exec_read_file(item)
assert isinstance(output, str)
assert "not found" in output
def test_svg_read_as_text(self, tmp_db, tmp_path):
"""SVG files are read as text, not as images."""
svg = tmp_path / "icon.svg"
svg.write_text('<svg xmlns="http://www.w3.org/2000/svg"><circle r="10"/></svg>')
session = _make_session()
item = {"call_id": "c5", "path": str(svg), "offset": None, "limit": None}
call_id, output = session._exec_read_file(item)
assert isinstance(output, str)
assert "<svg" in output # Read as text
class TestGetCapabilitiesOverride:
"""Test _get_capabilities with config.toml overrides."""
def test_config_override_applies(self, tmp_db):
"""capabilities dict from ModelConfig is merged onto provider caps."""
from turnstone.core.model_registry import ModelConfig, ModelRegistry
from turnstone.core.providers._protocol import ModelCapabilities
cfg = ModelConfig(
alias="qwen-vl",
base_url="http://localhost:8000/v1",
api_key="dummy",
model="qwen-3.5-vl",
capabilities={"supports_vision": True},
)
registry = ModelRegistry(
models={"qwen-vl": cfg},
default="qwen-vl",
)
session = _make_session(registry=registry, model_alias="qwen-vl")
# Ensure provider returns a real ModelCapabilities (not MagicMock)
session._provider.get_capabilities = MagicMock(return_value=ModelCapabilities())
caps = session._get_capabilities()
assert caps.supports_vision is True
def test_no_override_uses_provider_default(self, tmp_db):
"""Without config override, provider defaults are used."""
session = _make_session()
caps = session._get_capabilities()
# Default OpenAI provider for unknown model → no vision
assert caps.supports_vision is False
+167
View File
@@ -0,0 +1,167 @@
"""Tests for turnstone.core.policy."""
import pytest
from turnstone.core.policy import evaluate_tool_policies_batch, evaluate_tool_policy
from turnstone.core.storage._sqlite import SQLiteBackend
@pytest.fixture
def storage(tmp_path):
path = str(tmp_path / "test.db")
backend = SQLiteBackend(path)
yield backend
backend.close()
def test_no_policies_returns_none(storage):
result = evaluate_tool_policy(storage, "bash")
assert result is None
def test_exact_match_allow(storage):
storage.create_tool_policy("p1", "allow-read", "read_file", "allow", 0)
assert evaluate_tool_policy(storage, "read_file") == "allow"
assert evaluate_tool_policy(storage, "write_file") is None
def test_glob_match_deny(storage):
storage.create_tool_policy("p1", "block-bash", "bash*", "deny", 0)
assert evaluate_tool_policy(storage, "bash") == "deny"
assert evaluate_tool_policy(storage, "bash_exec") == "deny"
assert evaluate_tool_policy(storage, "read_file") is None
def test_wildcard_match(storage):
storage.create_tool_policy("p1", "ask-all", "*", "ask", 0)
assert evaluate_tool_policy(storage, "anything") == "ask"
def test_priority_ordering(storage):
# Higher priority wins
storage.create_tool_policy("p1", "allow-all", "*", "allow", 0)
storage.create_tool_policy("p2", "deny-bash", "bash*", "deny", 100)
assert evaluate_tool_policy(storage, "bash") == "deny" # p2 matches first (higher priority)
assert evaluate_tool_policy(storage, "read_file") == "allow" # p1 matches
def test_disabled_policy_skipped(storage):
storage.create_tool_policy("p1", "block-bash", "bash*", "deny", 100, enabled=False)
storage.create_tool_policy("p2", "allow-all", "*", "allow", 0)
assert evaluate_tool_policy(storage, "bash") == "allow" # p1 disabled, falls through to p2
def test_batch_evaluation(storage):
storage.create_tool_policy("p1", "block-bash", "bash*", "deny", 100)
storage.create_tool_policy("p2", "allow-read", "read_*", "allow", 50)
results = evaluate_tool_policies_batch(storage, ["bash", "read_file", "write_file"])
assert results["bash"] == "deny"
assert results["read_file"] == "allow"
assert results["write_file"] is None
def test_storage_failure_returns_none():
"""Graceful degradation on storage failure."""
class BrokenStorage:
def list_tool_policies(self, org_id=""):
raise RuntimeError("boom")
assert evaluate_tool_policy(BrokenStorage(), "bash") is None
def test_batch_storage_failure():
class BrokenStorage:
def list_tool_policies(self, org_id=""):
raise RuntimeError("boom")
results = evaluate_tool_policies_batch(BrokenStorage(), ["a", "b"])
assert results == {"a": None, "b": None}
def test_first_match_wins(storage):
# Two policies match, first by priority wins
storage.create_tool_policy("p1", "deny-bash", "bash*", "deny", 100)
storage.create_tool_policy("p2", "allow-bash", "bash*", "allow", 50)
assert evaluate_tool_policy(storage, "bash_exec") == "deny"
# ---------------------------------------------------------------------------
# MCP resource and prompt policy patterns
# ---------------------------------------------------------------------------
def test_mcp_resource_wildcard_deny(storage):
"""Deny all MCP resource reads via glob pattern."""
storage.create_tool_policy("p1", "block-resources", "mcp_resource__*", "deny", 100)
assert evaluate_tool_policy(storage, "mcp_resource__file:///secret.txt") == "deny"
assert evaluate_tool_policy(storage, "mcp_resource__db://users") == "deny"
assert evaluate_tool_policy(storage, "read_file") is None # unrelated tool
def test_mcp_resource_per_server_pattern(storage):
"""Allow resources from a specific server, deny others."""
storage.create_tool_policy("p1", "block-all-resources", "mcp_resource__*", "deny", 50)
storage.create_tool_policy("p2", "allow-docs", "mcp_resource__file:///docs/*", "allow", 100)
assert evaluate_tool_policy(storage, "mcp_resource__file:///docs/readme.md") == "allow"
assert evaluate_tool_policy(storage, "mcp_resource__file:///etc/passwd") == "deny"
def test_mcp_prompt_wildcard_ask(storage):
"""Require approval for all MCP prompt invocations."""
storage.create_tool_policy("p1", "ask-prompts", "mcp__*", "ask", 100)
assert evaluate_tool_policy(storage, "mcp__github__code_review") == "ask"
assert evaluate_tool_policy(storage, "mcp__templates__greeting") == "ask"
assert evaluate_tool_policy(storage, "bash") is None
def test_mcp_prompt_per_server_allow(storage):
"""Auto-approve prompts from a trusted server."""
storage.create_tool_policy("p1", "ask-all-mcp", "mcp__*", "ask", 50)
storage.create_tool_policy("p2", "allow-trusted", "mcp__trusted__*", "allow", 100)
assert evaluate_tool_policy(storage, "mcp__trusted__greeting") == "allow"
assert evaluate_tool_policy(storage, "mcp__untrusted__evil") == "ask"
def test_mcp_batch_mixed(storage):
"""Batch evaluation with mixed MCP and built-in tools."""
storage.create_tool_policy("p1", "block-resources", "mcp_resource__*", "deny", 100)
storage.create_tool_policy("p2", "allow-prompts", "mcp__trusted__*", "allow", 100)
results = evaluate_tool_policies_batch(
storage,
["mcp_resource__file:///x", "mcp__trusted__greeting", "bash", "mcp__other__y"],
)
assert results["mcp_resource__file:///x"] == "deny"
assert results["mcp__trusted__greeting"] == "allow"
assert results["bash"] is None
assert results["mcp__other__y"] is None
def test_normalize_resource_uri_prevents_traversal():
"""URI normalization resolves .. segments to prevent policy traversal bypass."""
from turnstone.core.session import ChatSession
# Normal URI unchanged
assert ChatSession._normalize_resource_uri("file:///docs/readme.md") == "file:///docs/readme.md"
# Traversal resolved
assert ChatSession._normalize_resource_uri("file:///docs/../etc/passwd") == "file:///etc/passwd"
# Double traversal
assert ChatSession._normalize_resource_uri("file:///a/b/../../c") == "file:///c"
# Non-file scheme (netloc preserved, path normalized)
assert ChatSession._normalize_resource_uri("db://host/tables/../secrets") == "db://host/secrets"
# Percent-encoded traversal decoded before normalization
assert (
ChatSession._normalize_resource_uri("file:///docs/%2e%2e/etc/passwd")
== "file:///etc/passwd"
)
# Mixed percent-encoded and literal traversal
assert ChatSession._normalize_resource_uri("file:///a/%2e%2e/b/../c") == "file:///c"
def test_mcp_tool_granular_policy(storage):
"""MCP tool calls use their prefixed func_name for granular policy matching."""
storage.create_tool_policy("p1", "ask-all-mcp", "mcp__*", "ask", 50)
storage.create_tool_policy("p2", "allow-github", "mcp__github__*", "allow", 100)
# MCP tools now use func_name as approval_label
assert evaluate_tool_policy(storage, "mcp__github__search") == "allow"
assert evaluate_tool_policy(storage, "mcp__untrusted__exec") == "ask"
+16 -5
View File
@@ -72,16 +72,24 @@ class TestToolsMetadata:
"""Validate the metadata extracted from JSON files."""
def test_tool_count(self):
assert len(TOOLS) == 15
assert len(TOOLS) == 18
def test_agent_tools_count(self):
assert len(AGENT_TOOLS) == 7
assert len(AGENT_TOOLS) == 9
def test_task_agent_tools_count(self):
assert len(TASK_AGENT_TOOLS) == 10
assert len(TASK_AGENT_TOOLS) == 12
def test_auto_approve_sets_match(self):
expected = {"read_file", "search", "math", "man", "web_fetch", "web_search", "notify"}
expected = {
"read_file",
"search",
"math",
"man",
"web_fetch",
"web_search",
"notify",
}
assert expected == AGENT_AUTO_TOOLS
assert expected == TASK_AUTO_TOOLS
@@ -97,11 +105,14 @@ class TestToolsMetadata:
"web_fetch": "url",
"web_search": "query",
"task": "prompt",
"plan": "prompt",
"create_plan": "goal",
"remember": "key",
"recall": "query",
"forget": "key",
"notify": "message",
"watch": "command",
"read_resource": "uri",
"use_prompt": "name",
}
assert expected == PRIMARY_KEY_MAP
+8
View File
@@ -65,6 +65,14 @@ class TestUserCRUD:
db.delete_user("u1")
assert len(db.list_api_tokens("u1")) == 0
def test_delete_cascades_user_roles(self, db):
db.create_user("u1", "admin", "Admin", "$2b$hash")
db.create_role("r1", "editor", "Editor", "read,write", builtin=False, org_id="")
db.assign_role("u1", "r1")
assert len(db.list_user_roles("u1")) == 1
db.delete_user("u1")
assert len(db.list_user_roles("u1")) == 0
class TestApiTokenCRUD:
def test_create_and_lookup_by_hash(self, db):
+487
View File
@@ -0,0 +1,487 @@
"""Tests for the watch module — duration parsing, condition evaluation, WatchRunner."""
from __future__ import annotations
from datetime import UTC, datetime
from unittest.mock import MagicMock
import pytest
from turnstone.core.watch import (
WatchRunner,
evaluate_condition,
format_interval,
format_watch_message,
parse_duration,
validate_condition,
)
# ---------------------------------------------------------------------------
# parse_duration
# ---------------------------------------------------------------------------
class TestParseDuration:
def test_seconds(self):
assert parse_duration("30s") == 30.0
def test_minutes(self):
assert parse_duration("5m") == 300.0
def test_hours(self):
assert parse_duration("1h") == 3600.0
def test_compound(self):
assert parse_duration("2h30m") == 9000.0
def test_bare_number(self):
assert parse_duration("90") == 90.0
def test_bare_float(self):
assert parse_duration("10.5") == 10.5
def test_whitespace(self):
assert parse_duration(" 5m ") == 300.0
def test_case_insensitive(self):
assert parse_duration("1H30M") == 5400.0
def test_empty_raises(self):
with pytest.raises(ValueError, match="empty"):
parse_duration("")
def test_invalid_raises(self):
with pytest.raises(ValueError, match="invalid duration"):
parse_duration("abc")
def test_negative_raises(self):
with pytest.raises(ValueError, match="positive"):
parse_duration("-5")
def test_zero_raises(self):
with pytest.raises(ValueError, match="positive"):
parse_duration("0")
def test_zero_duration_raises(self):
with pytest.raises(ValueError, match="positive"):
parse_duration("0s")
# ---------------------------------------------------------------------------
# validate_condition
# ---------------------------------------------------------------------------
class TestValidateCondition:
def test_valid_expression(self):
assert validate_condition('data["state"] == "MERGED"') is None
def test_valid_simple(self):
assert validate_condition('"error" in output') is None
def test_valid_compound(self):
assert validate_condition('changed and "ready" in output.lower()') is None
def test_syntax_error(self):
result = validate_condition("if True:")
assert result is not None
assert "syntax" in result.lower()
def test_incomplete_expression(self):
result = validate_condition("==")
assert result is not None
# ---------------------------------------------------------------------------
# evaluate_condition
# ---------------------------------------------------------------------------
class TestEvaluateCondition:
def test_none_first_poll_no_fire(self):
"""With stop_on=None, first poll (prev_output=None) should not fire."""
fired, reason = evaluate_condition(None, "hello", 0, None)
assert not fired
def test_none_change_detected(self):
fired, reason = evaluate_condition(None, "world", 0, "hello")
assert fired
assert "changed" in reason
def test_none_no_change(self):
fired, reason = evaluate_condition(None, "same", 0, "same")
assert not fired
def test_string_match(self):
fired, reason = evaluate_condition('"error" in output', "has error here", 0, None)
assert fired
def test_string_no_match(self):
fired, reason = evaluate_condition('"error" in output', "all good", 0, None)
assert not fired
def test_exit_code(self):
fired, reason = evaluate_condition("exit_code != 0", "fail", 1, None)
assert fired
def test_exit_code_zero(self):
fired, reason = evaluate_condition("exit_code != 0", "ok", 0, None)
assert not fired
def test_json_data(self):
output = '{"state": "MERGED"}'
fired, reason = evaluate_condition('data["state"] == "MERGED"', output, 0, None)
assert fired
def test_json_data_no_match(self):
output = '{"state": "OPEN"}'
fired, reason = evaluate_condition('data["state"] == "MERGED"', output, 0, None)
assert not fired
def test_json_data_none_for_non_json(self):
"""Non-JSON output should have data=None."""
fired, reason = evaluate_condition("data is None", "plain text", 0, None)
assert fired
def test_changed_variable(self):
fired, reason = evaluate_condition("changed", "new", 0, "old")
assert fired
def test_changed_false(self):
fired, reason = evaluate_condition("changed", "same", 0, "same")
assert not fired
def test_compound_condition(self):
fired, reason = evaluate_condition(
'changed and "ready" in output.lower()',
"System Ready",
0,
"System Starting",
)
assert fired
def test_invalid_expression_no_crash(self):
fired, reason = evaluate_condition("1/0", "hello", 0, None)
assert not fired
assert "error" in reason.lower()
def test_no_import_builtin(self):
"""__import__ should not be accessible."""
fired, reason = evaluate_condition("__import__('os')", "hello", 0, None)
assert not fired
assert "error" in reason.lower()
def test_no_open_builtin(self):
fired, reason = evaluate_condition("open('/etc/passwd')", "hello", 0, None)
assert not fired
assert "error" in reason.lower()
def test_no_exec_builtin(self):
fired, reason = evaluate_condition("exec('print(1)')", "hello", 0, None)
assert not fired
assert "error" in reason.lower()
def test_no_eval_builtin(self):
fired, reason = evaluate_condition("eval('1+1')", "hello", 0, None)
assert not fired
assert "error" in reason.lower()
def test_no_compile_builtin(self):
fired, reason = evaluate_condition("compile('1','','eval')", "hello", 0, None)
assert not fired
assert "error" in reason.lower()
def test_safe_len(self):
fired, reason = evaluate_condition("len(output) > 0", "hello", 0, None)
assert fired
def test_safe_sorted(self):
fired, reason = evaluate_condition("sorted([3,1,2]) == [1,2,3]", "x", 0, None)
assert fired
def test_data_get_method(self):
output = '{"mergedAt": "2024-01-15"}'
fired, reason = evaluate_condition('data.get("mergedAt") is not None', output, 0, None)
assert fired
def test_prev_output_available(self):
fired, reason = evaluate_condition(
"prev_output is not None and output != prev_output",
"new",
0,
"old",
)
assert fired
# ---------------------------------------------------------------------------
# format_interval
# ---------------------------------------------------------------------------
class TestFormatInterval:
def test_seconds(self):
assert format_interval(30) == "30s"
def test_exactly_60(self):
assert format_interval(60) == "1m"
def test_minutes(self):
assert format_interval(300) == "5m"
def test_exactly_3600(self):
assert format_interval(3600) == "1h"
def test_hours_and_minutes(self):
assert format_interval(5400) == "1h30m"
def test_hours_only(self):
assert format_interval(7200) == "2h"
def test_large_value(self):
assert format_interval(86400) == "24h"
# ---------------------------------------------------------------------------
# format_watch_message
# ---------------------------------------------------------------------------
class TestFormatWatchMessage:
def test_basic(self):
msg = format_watch_message(
name="pr-review",
command="gh pr view --json state",
output='{"state": "MERGED"}',
poll_count=5,
max_polls=100,
elapsed_secs=1500,
stop_on='data["state"] == "MERGED"',
is_final=True,
reason='condition met: data["state"] == "MERGED"',
)
assert "pr-review" in msg
assert "poll #5/100" in msg
assert "25m" in msg
assert "gh pr view --json state" in msg
assert "MERGED" in msg
assert "auto-cancelled" in msg.lower()
# Model should see the condition it was waiting for
assert "condition:" in msg.lower()
def test_non_final(self):
msg = format_watch_message(
name="deploy",
command="curl -s http://localhost/health",
output="ok",
poll_count=3,
max_polls=50,
elapsed_secs=90,
stop_on=None,
is_final=False,
reason="",
)
assert "deploy" in msg
assert "auto-cancelled" not in msg.lower()
# Change-detection mode should be indicated
assert "output change" in msg.lower()
def test_max_polls_final(self):
msg = format_watch_message(
name="test",
command="echo hello",
output="hello",
poll_count=100,
max_polls=100,
elapsed_secs=6000,
stop_on=None,
is_final=True,
reason="",
)
assert "max polls" in msg.lower()
# ---------------------------------------------------------------------------
# WatchRunner
# ---------------------------------------------------------------------------
class TestWatchRunner:
def _make_runner(self, storage=None, **kwargs):
if storage is None:
storage = MagicMock()
storage.list_due_watches.return_value = []
return WatchRunner(
storage=storage,
node_id="test-node",
check_interval=0.1,
tool_timeout=5,
**kwargs,
)
def test_start_stop(self):
runner = self._make_runner()
runner.start()
assert runner._thread is not None
assert runner._thread.is_alive()
runner.stop()
assert runner._thread is None
def test_tick_calls_list_due(self):
storage = MagicMock()
storage.list_due_watches.return_value = []
runner = self._make_runner(storage=storage)
runner._tick()
storage.list_due_watches.assert_called_once()
def test_poll_watch_runs_command(self):
storage = MagicMock()
storage.update_watch.return_value = True
runner = self._make_runner(storage=storage)
dispatch_fn = MagicMock()
runner.set_dispatch_fn("ws-1", dispatch_fn)
watch_row = {
"watch_id": "abc123",
"ws_id": "ws-1",
"name": "test-watch",
"command": "echo hello",
"stop_on": '"hello" in output',
"max_polls": 100,
"poll_count": 0,
"last_output": None,
"interval_secs": 60,
"created": datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S"),
}
runner._poll_watch(watch_row)
# Should update the watch in storage
storage.update_watch.assert_called_once()
call_kwargs = storage.update_watch.call_args
assert call_kwargs[0][0] == "abc123" # watch_id
assert call_kwargs[1]["poll_count"] == 1
# Condition should fire (output contains "hello")
assert call_kwargs[1]["active"] is False # deactivated
# Should dispatch result
dispatch_fn.assert_called_once()
def test_poll_watch_no_fire_on_first_change_detection(self):
storage = MagicMock()
storage.update_watch.return_value = True
runner = self._make_runner(storage=storage)
dispatch_fn = MagicMock()
runner.set_dispatch_fn("ws-1", dispatch_fn)
watch_row = {
"watch_id": "abc123",
"ws_id": "ws-1",
"name": "test-watch",
"command": "echo hello",
"stop_on": None, # change detection
"max_polls": 100,
"poll_count": 0,
"last_output": None, # first poll
"interval_secs": 60,
"created": datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S"),
}
runner._poll_watch(watch_row)
# First poll with change detection should not fire
dispatch_fn.assert_not_called()
call_kwargs = storage.update_watch.call_args
# Watch should remain active
assert "active" not in call_kwargs[1] or call_kwargs[1].get("active") is not False
def test_max_polls_deactivates(self):
storage = MagicMock()
storage.update_watch.return_value = True
runner = self._make_runner(storage=storage)
dispatch_fn = MagicMock()
runner.set_dispatch_fn("ws-1", dispatch_fn)
watch_row = {
"watch_id": "abc123",
"ws_id": "ws-1",
"name": "test-watch",
"command": "echo hello",
"stop_on": '"never" in output', # won't fire
"max_polls": 5,
"poll_count": 4, # next is #5 = max
"last_output": "hello\n",
"interval_secs": 60,
"created": datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S"),
}
runner._poll_watch(watch_row)
call_kwargs = storage.update_watch.call_args
assert call_kwargs[1]["active"] is False
assert call_kwargs[1]["poll_count"] == 5
dispatch_fn.assert_called_once()
def test_blocked_command_deactivates(self):
storage = MagicMock()
storage.update_watch.return_value = True
runner = self._make_runner(storage=storage)
watch_row = {
"watch_id": "abc123",
"ws_id": "ws-1",
"name": "test-watch",
"command": "rm -rf /",
"stop_on": None,
"max_polls": 100,
"poll_count": 0,
"last_output": None,
"interval_secs": 60,
"created": datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S"),
}
runner._poll_watch(watch_row)
storage.update_watch.assert_called_once()
call_kwargs = storage.update_watch.call_args
assert call_kwargs[0][0] == "abc123"
assert call_kwargs[1]["active"] is False
def test_dispatch_fn_registry(self):
runner = self._make_runner()
fn1 = MagicMock()
fn2 = MagicMock()
runner.set_dispatch_fn("ws-1", fn1)
runner.set_dispatch_fn("ws-2", fn2)
runner._dispatch_result("ws-1", "msg1")
fn1.assert_called_once_with("msg1")
fn2.assert_not_called()
runner.remove_dispatch_fn("ws-1")
# After removal, dispatch should try restore_fn
runner._dispatch_result("ws-1", "msg2")
fn1.assert_called_once() # still just the one call
def test_restore_fn_called_for_evicted(self):
restored_fn = MagicMock()
restore_fn = MagicMock(return_value=restored_fn)
runner = self._make_runner(restore_fn=restore_fn)
runner._dispatch_result("ws-evicted", "hello")
restore_fn.assert_called_once_with("ws-evicted")
restored_fn.assert_called_once_with("hello")
def test_run_command_success(self):
runner = self._make_runner()
output, code = runner._run_command("echo hello")
assert "hello" in output
assert code == 0
def test_run_command_failure(self):
runner = self._make_runner()
output, code = runner._run_command("exit 42")
assert code == 42
def test_run_command_timeout(self):
runner = self._make_runner()
runner._tool_timeout = 1
output, code = runner._run_command("sleep 30")
assert "timed out" in output.lower()
assert code == -1
+130
View File
@@ -0,0 +1,130 @@
"""Tests for watches storage CRUD."""
from __future__ import annotations
import pytest
from turnstone.core.storage._sqlite import SQLiteBackend
@pytest.fixture
def db(tmp_path):
"""Fresh SQLite backend for each test."""
return SQLiteBackend(str(tmp_path / "test.db"))
def _make_watch_kwargs(**overrides):
"""Build default kwargs for create_watch."""
defaults = {
"watch_id": "watch_001",
"ws_id": "ws-abc",
"node_id": "node-1",
"name": "pr-review",
"command": "gh pr view --json state",
"interval_secs": 300.0,
"stop_on": 'data["state"] == "MERGED"',
"max_polls": 100,
"created_by": "model",
"next_poll": "2099-01-01T00:05:00",
}
defaults.update(overrides)
return defaults
class TestWatchCRUD:
def test_create_and_get(self, db):
db.create_watch(**_make_watch_kwargs())
w = db.get_watch("watch_001")
assert w is not None
assert w["name"] == "pr-review"
assert w["command"] == "gh pr view --json state"
assert w["interval_secs"] == 300.0
assert w["active"] == 1
assert w["poll_count"] == 0
def test_get_nonexistent(self, db):
assert db.get_watch("nope") is None
def test_create_idempotent(self, db):
db.create_watch(**_make_watch_kwargs())
db.create_watch(**_make_watch_kwargs()) # OR IGNORE
assert db.get_watch("watch_001") is not None
def test_update(self, db):
db.create_watch(**_make_watch_kwargs())
updated = db.update_watch(
"watch_001",
poll_count=5,
last_output="hello",
last_exit_code=0,
)
assert updated is True
w = db.get_watch("watch_001")
assert w["poll_count"] == 5
assert w["last_output"] == "hello"
assert w["last_exit_code"] == 0
def test_update_nonexistent(self, db):
assert db.update_watch("nope", poll_count=1) is False
def test_update_active_flag(self, db):
db.create_watch(**_make_watch_kwargs())
db.update_watch("watch_001", active=False)
w = db.get_watch("watch_001")
assert w["active"] == 0
def test_delete(self, db):
db.create_watch(**_make_watch_kwargs())
assert db.delete_watch("watch_001") is True
assert db.get_watch("watch_001") is None
def test_delete_nonexistent(self, db):
assert db.delete_watch("nope") is False
class TestWatchListQueries:
def test_list_for_ws(self, db):
db.create_watch(**_make_watch_kwargs(watch_id="w1", ws_id="ws-1", name="a"))
db.create_watch(**_make_watch_kwargs(watch_id="w2", ws_id="ws-1", name="b"))
db.create_watch(**_make_watch_kwargs(watch_id="w3", ws_id="ws-2", name="c"))
ws1 = db.list_watches_for_ws("ws-1")
assert len(ws1) == 2
assert {w["name"] for w in ws1} == {"a", "b"}
def test_list_for_ws_excludes_inactive(self, db):
db.create_watch(**_make_watch_kwargs(watch_id="w1", ws_id="ws-1"))
db.update_watch("w1", active=False)
assert db.list_watches_for_ws("ws-1") == []
def test_list_for_node(self, db):
db.create_watch(**_make_watch_kwargs(watch_id="w1", node_id="n1"))
db.create_watch(**_make_watch_kwargs(watch_id="w2", node_id="n1"))
db.create_watch(**_make_watch_kwargs(watch_id="w3", node_id="n2"))
n1 = db.list_watches_for_node("n1")
assert len(n1) == 2
def test_list_due(self, db):
# Due
db.create_watch(**_make_watch_kwargs(watch_id="w1", next_poll="2020-01-01T00:00:00"))
# Not due (far future)
db.create_watch(**_make_watch_kwargs(watch_id="w2", next_poll="2099-01-01T00:00:00"))
# Due but inactive
db.create_watch(**_make_watch_kwargs(watch_id="w3", next_poll="2020-01-01T00:00:00"))
db.update_watch("w3", active=False)
due = db.list_due_watches("2025-01-01T00:00:00")
assert len(due) == 1
assert due[0]["watch_id"] == "w1"
def test_delete_for_ws(self, db):
db.create_watch(**_make_watch_kwargs(watch_id="w1", ws_id="ws-1"))
db.create_watch(**_make_watch_kwargs(watch_id="w2", ws_id="ws-1"))
db.create_watch(**_make_watch_kwargs(watch_id="w3", ws_id="ws-2"))
count = db.delete_watches_for_ws("ws-1")
assert count == 2
assert db.get_watch("w1") is None
assert db.get_watch("w2") is None
assert db.get_watch("w3") is not None
+136
View File
@@ -665,6 +665,31 @@ class TestWebUI:
assert ui._approval_result == (True, "looks good")
t.join()
def test_resolve_approval_emits_event(self):
"""resolve_approval should enqueue an approval_resolved SSE event."""
from turnstone.server import WebUI
ui = WebUI(ws_id="test-emit")
listener = ui._register_listener()
# Drain any init events
while not listener.empty():
listener.get_nowait()
ui.resolve_approval(False, "Approval timed out")
# Collect events from the listener
events = []
while not listener.empty():
events.append(listener.get_nowait())
ui._unregister_listener(listener)
resolved = [e for e in events if e.get("type") == "approval_resolved"]
assert len(resolved) == 1
assert resolved[0]["approved"] is False
assert resolved[0]["feedback"] == "Approval timed out"
def test_resolve_plan(self):
from turnstone.server import WebUI
@@ -683,6 +708,117 @@ class TestWebUI:
t.join()
# ---------------------------------------------------------------------------
# WebUI SSE fan-out
# ---------------------------------------------------------------------------
class TestWebUIFanOut:
"""Verify per-client SSE fan-out on WebUI._enqueue / _register_listener."""
def test_enqueue_no_listeners(self):
"""Events silently dropped when no listeners are registered."""
from turnstone.server import WebUI
ui = WebUI(ws_id="test")
ui._enqueue({"type": "content", "text": "hello"}) # should not raise
def test_enqueue_single_listener(self):
"""Single listener receives the event."""
from turnstone.server import WebUI
ui = WebUI(ws_id="test")
q = ui._register_listener()
ui._enqueue({"type": "content", "text": "hello"})
assert q.get_nowait() == {"type": "content", "text": "hello"}
def test_enqueue_multiple_listeners(self):
"""All registered listeners receive an identical copy."""
from turnstone.server import WebUI
ui = WebUI(ws_id="test")
q1 = ui._register_listener()
q2 = ui._register_listener()
q3 = ui._register_listener()
event = {"type": "content", "text": "world"}
ui._enqueue(event)
assert q1.get_nowait() == event
assert q2.get_nowait() == event
assert q3.get_nowait() == event
def test_unregister_stops_delivery(self):
"""After unregister, the queue receives no further events."""
import queue as queue_mod
from turnstone.server import WebUI
ui = WebUI(ws_id="test")
q = ui._register_listener()
ui._unregister_listener(q)
ui._enqueue({"type": "content", "text": "gone"})
with pytest.raises(queue_mod.Empty):
q.get_nowait()
def test_slow_consumer_does_not_block(self):
"""A full queue doesn't block the producer or starve other listeners."""
from turnstone.server import WebUI
ui = WebUI(ws_id="test")
slow = ui._register_listener()
fast = ui._register_listener()
# Fill only the slow consumer's queue directly to capacity
for i in range(500):
slow.put_nowait({"type": "content", "text": f"fill-{i}"})
assert slow.qsize() == 500
assert fast.qsize() == 0
# Enqueue via fan-out — slow drops (full), fast receives
event = {"type": "content", "text": "overflow"}
ui._enqueue(event)
assert slow.qsize() == 500 # still full, overflow dropped
assert fast.qsize() == 1
assert fast.get_nowait() == event
def test_unregister_idempotent(self):
"""Double unregister does not raise."""
from turnstone.server import WebUI
ui = WebUI(ws_id="test")
q = ui._register_listener()
ui._unregister_listener(q)
ui._unregister_listener(q) # should not raise
def test_concurrent_enqueue_and_register(self):
"""Concurrent register/unregister and enqueue should not crash."""
from turnstone.server import WebUI
ui = WebUI(ws_id="test")
stop = threading.Event()
def register_loop():
while not stop.is_set():
q = ui._register_listener()
ui._unregister_listener(q)
def enqueue_loop():
for i in range(500):
ui._enqueue({"type": "content", "text": f"tok-{i}"})
t1 = threading.Thread(target=register_loop)
t2 = threading.Thread(target=enqueue_loop)
t1.start()
t2.start()
t2.join()
stop.set()
t1.join()
# ---------------------------------------------------------------------------
# Integration: WorkstreamManager + session state transitions
# ---------------------------------------------------------------------------
+1 -1
View File
@@ -1,3 +1,3 @@
"""turnstone - Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation."""
__version__ = "0.4.6"
__version__ = "0.5.5"
+245 -2
View File
@@ -2,6 +2,8 @@
from __future__ import annotations
from typing import Any
from pydantic import BaseModel, Field
# ---------------------------------------------------------------------------
@@ -48,7 +50,7 @@ class ClusterNodeInfo(BaseModel):
total_tokens: int = 0
started: float = 0.0
reachable: bool = True
health: dict[str, str] = Field(default_factory=dict)
health: dict[str, Any] = Field(default_factory=dict)
version: str = ""
@@ -91,12 +93,34 @@ class ClusterWorkstreamsResponse(BaseModel):
class NodeDetailResponse(BaseModel):
node_id: str
server_url: str = ""
health: dict[str, str] = Field(default_factory=dict)
health: dict[str, Any] = Field(default_factory=dict)
workstreams: list[ClusterWorkstreamInfo] = []
aggregate: dict[str, int] = Field(default_factory=dict)
reachable: bool = True
# ---------------------------------------------------------------------------
# Cluster snapshot
# ---------------------------------------------------------------------------
class ClusterSnapshotNode(BaseModel):
node_id: str
server_url: str = ""
max_ws: int = 10
reachable: bool = True
version: str = ""
health: dict[str, Any] = Field(default_factory=dict)
aggregate: dict[str, int] = Field(default_factory=dict)
workstreams: list[ClusterWorkstreamInfo] = []
class ClusterSnapshotResponse(BaseModel):
nodes: list[ClusterSnapshotNode]
overview: ClusterOverviewResponse
timestamp: float = 0.0
# ---------------------------------------------------------------------------
# Workstream creation
# ---------------------------------------------------------------------------
@@ -112,6 +136,9 @@ class ConsoleCreateWsRequest(BaseModel):
initial_message: str = Field(
default="", description="Optional first message sent after creation"
)
template: str = Field(
default="", description="Prompt template name (replaces default templates)"
)
class ConsoleCreateWsResponse(BaseModel):
@@ -132,3 +159,219 @@ class ConsoleHealthResponse(BaseModel):
workstreams: int = 0
version_drift: bool = False
versions: list[str] = []
# ---------------------------------------------------------------------------
# Governance: Roles
# ---------------------------------------------------------------------------
class RoleInfo(BaseModel):
role_id: str
name: str
display_name: str
permissions: str
builtin: bool
org_id: str
created: str
updated: str
class CreateRoleRequest(BaseModel):
name: str
display_name: str = ""
permissions: str = "read"
class UpdateRoleRequest(BaseModel):
display_name: str | None = None
permissions: str | None = None
class ListRolesResponse(BaseModel):
roles: list[RoleInfo]
class AssignRoleRequest(BaseModel):
role_id: str
class UserRoleInfo(BaseModel):
role_id: str
name: str
display_name: str
permissions: str
builtin: bool
org_id: str
created: str
updated: str
assigned_by: str
assignment_created: str
class ListUserRolesResponse(BaseModel):
roles: list[UserRoleInfo]
# ---------------------------------------------------------------------------
# Governance: Orgs
# ---------------------------------------------------------------------------
class OrgInfo(BaseModel):
org_id: str
name: str
display_name: str
settings: str
created: str
updated: str
class UpdateOrgRequest(BaseModel):
display_name: str | None = None
settings: str | None = None
class ListOrgsResponse(BaseModel):
orgs: list[OrgInfo]
# ---------------------------------------------------------------------------
# Governance: Tool Policies
# ---------------------------------------------------------------------------
class ToolPolicyInfo(BaseModel):
policy_id: str
name: str
tool_pattern: str
action: str
priority: int
org_id: str
enabled: bool
created_by: str
created: str
updated: str
class CreateToolPolicyRequest(BaseModel):
name: str
tool_pattern: str
action: str # allow, deny, ask
priority: int = 0
org_id: str = ""
enabled: bool = True
class UpdateToolPolicyRequest(BaseModel):
name: str | None = None
tool_pattern: str | None = None
action: str | None = None
priority: int | None = None
enabled: bool | None = None
class ListToolPoliciesResponse(BaseModel):
policies: list[ToolPolicyInfo]
# ---------------------------------------------------------------------------
# Governance: Prompt Templates
# ---------------------------------------------------------------------------
class PromptTemplateInfo(BaseModel):
template_id: str
name: str
category: str
content: str
variables: str
is_default: bool
org_id: str
created_by: str
origin: str = "manual"
mcp_server: str = ""
readonly: bool = False
created: str
updated: str
class CreatePromptTemplateRequest(BaseModel):
name: str
content: str
category: str = "general"
variables: str = "[]"
is_default: bool = False
org_id: str = ""
class UpdatePromptTemplateRequest(BaseModel):
name: str | None = None
content: str | None = None
category: str | None = None
variables: str | None = None
is_default: bool | None = None
class ListPromptTemplatesResponse(BaseModel):
templates: list[PromptTemplateInfo]
# ---------------------------------------------------------------------------
# Governance: Usage
# ---------------------------------------------------------------------------
class UsageBreakdownItem(BaseModel):
key: str = ""
prompt_tokens: int = 0
completion_tokens: int = 0
tool_calls_count: int = 0
class UsageResponse(BaseModel):
summary: list[UsageBreakdownItem]
breakdown: list[UsageBreakdownItem]
# ---------------------------------------------------------------------------
# Governance: Audit
# ---------------------------------------------------------------------------
class AuditEventInfo(BaseModel):
event_id: str
timestamp: str
user_id: str
action: str
resource_type: str
resource_id: str
detail: str
ip_address: str
created: str
class ListAuditEventsResponse(BaseModel):
events: list[AuditEventInfo]
# ---------------------------------------------------------------------------
# Channels
# ---------------------------------------------------------------------------
class ChannelUserInfo(BaseModel):
channel_type: str
channel_user_id: str
user_id: str
created: str
class ListChannelUsersResponse(BaseModel):
channels: list[ChannelUserInfo]
class CreateChannelUserRequest(BaseModel):
channel_type: str = Field(..., description="Channel type (e.g. discord, slack)")
channel_user_id: str = Field(..., description="External channel user identifier")
total: int
+273 -2
View File
@@ -8,13 +8,39 @@ if TYPE_CHECKING:
from pydantic import BaseModel
from turnstone.api.console_schemas import (
AssignRoleRequest,
AuditEventInfo,
ChannelUserInfo,
ClusterNodesResponse,
ClusterOverviewResponse,
ClusterSnapshotResponse,
ClusterWorkstreamsResponse,
ConsoleCreateWsRequest,
ConsoleCreateWsResponse,
ConsoleHealthResponse,
CreateChannelUserRequest,
CreatePromptTemplateRequest,
CreateRoleRequest,
CreateToolPolicyRequest,
ListAuditEventsResponse,
ListChannelUsersResponse,
ListOrgsResponse,
ListPromptTemplatesResponse,
ListRolesResponse,
ListToolPoliciesResponse,
ListUserRolesResponse,
NodeDetailResponse,
OrgInfo,
PromptTemplateInfo,
RoleInfo,
ToolPolicyInfo,
UpdateOrgRequest,
UpdatePromptTemplateRequest,
UpdateRoleRequest,
UpdateToolPolicyRequest,
UsageBreakdownItem,
UsageResponse,
UserRoleInfo,
)
from turnstone.api.openapi import EndpointSpec, QueryParam, build_openapi
from turnstone.api.schemas import (
@@ -97,14 +123,23 @@ CONSOLE_ENDPOINTS: list[EndpointSpec] = [
error_codes=[400, 404, 503],
tags=["Cluster"],
),
EndpointSpec(
"/v1/api/cluster/snapshot",
"GET",
"Full cluster state snapshot",
description="Returns the complete cluster state: all nodes with their workstreams "
"and overview aggregates. Used for initial load and reconnection.",
response_model=ClusterSnapshotResponse,
tags=["Cluster"],
),
# --- Streaming ---
EndpointSpec(
"/v1/api/cluster/events",
"GET",
"Cluster SSE event stream",
description="Server-Sent Events stream for real-time cluster updates. "
"Returns text/event-stream with node_joined, node_lost, cluster_state, "
"ws_created, ws_closed, ws_rename events.",
"First event is a 'snapshot' with full cluster state, followed by "
"node_joined, node_lost, cluster_state, ws_created, ws_closed, ws_rename events.",
tags=["Streaming"],
),
# --- Auth ---
@@ -187,6 +222,31 @@ CONSOLE_ENDPOINTS: list[EndpointSpec] = [
error_codes=[404],
tags=["Admin"],
),
# --- Channels ---
EndpointSpec(
"/v1/api/admin/users/{user_id}/channels",
"GET",
"List channel links for a user",
response_model=ListChannelUsersResponse,
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/users/{user_id}/channels",
"POST",
"Link a channel account to a user",
request_model=CreateChannelUserRequest,
response_model=ChannelUserInfo,
error_codes=[400, 404, 409],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/channels/{channel_type}/{channel_user_id}",
"DELETE",
"Unlink a channel account",
response_model=StatusResponse,
error_codes=[404],
tags=["Admin"],
),
# --- Schedules ---
EndpointSpec(
"/v1/api/admin/schedules",
@@ -242,6 +302,191 @@ CONSOLE_ENDPOINTS: list[EndpointSpec] = [
error_codes=[404],
tags=["Schedules"],
),
# --- Governance: Roles ---
EndpointSpec(
"/v1/api/admin/roles",
"GET",
"List all roles",
response_model=ListRolesResponse,
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/roles",
"POST",
"Create a custom role",
request_model=CreateRoleRequest,
response_model=RoleInfo,
error_codes=[400],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/roles/{role_id}",
"PUT",
"Update a role",
request_model=UpdateRoleRequest,
response_model=RoleInfo,
error_codes=[400, 404],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/roles/{role_id}",
"DELETE",
"Delete a custom role",
response_model=StatusResponse,
error_codes=[400, 404],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/users/{user_id}/roles",
"GET",
"List roles assigned to a user",
response_model=ListUserRolesResponse,
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/users/{user_id}/roles",
"POST",
"Assign a role to a user",
request_model=AssignRoleRequest,
response_model=StatusResponse,
error_codes=[400, 404],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/users/{user_id}/roles/{role_id}",
"DELETE",
"Unassign a role from a user",
response_model=StatusResponse,
error_codes=[404],
tags=["Admin"],
),
# --- Governance: Orgs ---
EndpointSpec(
"/v1/api/admin/orgs",
"GET",
"List organizations",
response_model=ListOrgsResponse,
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/orgs/{org_id}",
"GET",
"Get organization details",
response_model=OrgInfo,
error_codes=[404],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/orgs/{org_id}",
"PUT",
"Update organization settings",
request_model=UpdateOrgRequest,
response_model=OrgInfo,
error_codes=[404],
tags=["Admin"],
),
# --- Governance: Tool Policies ---
EndpointSpec(
"/v1/api/admin/policies",
"GET",
"List tool policies",
response_model=ListToolPoliciesResponse,
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/policies",
"POST",
"Create a tool policy",
request_model=CreateToolPolicyRequest,
response_model=ToolPolicyInfo,
error_codes=[400],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/policies/{policy_id}",
"PUT",
"Update a tool policy",
request_model=UpdateToolPolicyRequest,
response_model=ToolPolicyInfo,
error_codes=[404],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/policies/{policy_id}",
"DELETE",
"Delete a tool policy",
response_model=StatusResponse,
error_codes=[404],
tags=["Admin"],
),
# --- Governance: Prompt Templates ---
EndpointSpec(
"/v1/api/admin/templates",
"GET",
"List prompt templates",
response_model=ListPromptTemplatesResponse,
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/templates",
"POST",
"Create a prompt template",
request_model=CreatePromptTemplateRequest,
response_model=PromptTemplateInfo,
error_codes=[400],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/templates/{template_id}",
"PUT",
"Update a prompt template",
request_model=UpdatePromptTemplateRequest,
response_model=PromptTemplateInfo,
error_codes=[404],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/templates/{template_id}",
"DELETE",
"Delete a prompt template",
response_model=StatusResponse,
error_codes=[404],
tags=["Admin"],
),
# --- Governance: Usage & Audit ---
EndpointSpec(
"/v1/api/admin/usage",
"GET",
"Aggregated usage data",
response_model=UsageResponse,
query_params=[
QueryParam("since", "Start timestamp (ISO8601, defaults to last 7 days)"),
QueryParam("until", "End timestamp (ISO8601)"),
QueryParam("user_id", "Filter by user"),
QueryParam("model", "Filter by model"),
QueryParam(
"group_by",
"Group results",
enum=["day", "hour", "model", "user"],
),
],
tags=["Admin"],
),
EndpointSpec(
"/v1/api/admin/audit",
"GET",
"Paginated audit events",
response_model=ListAuditEventsResponse,
query_params=[
QueryParam("action", "Filter by action type"),
QueryParam("user_id", "Filter by user"),
QueryParam("since", "Start timestamp (ISO8601)"),
QueryParam("until", "End timestamp (ISO8601)"),
QueryParam("limit", "Page size", schema_type="integer", default=50),
QueryParam("offset", "Pagination offset", schema_type="integer", default=0),
],
tags=["Admin"],
),
# --- Observability ---
EndpointSpec(
"/health",
@@ -266,10 +511,14 @@ _ALL_MODELS: list[type[BaseModel]] = [
CreateTokenRequest,
CreateTokenResponse,
ListTokensResponse,
ChannelUserInfo,
CreateChannelUserRequest,
ListChannelUsersResponse,
ClusterOverviewResponse,
ClusterNodesResponse,
ClusterWorkstreamsResponse,
NodeDetailResponse,
ClusterSnapshotResponse,
ConsoleCreateWsRequest,
ConsoleCreateWsResponse,
ConsoleHealthResponse,
@@ -278,6 +527,28 @@ _ALL_MODELS: list[type[BaseModel]] = [
ScheduleInfo,
ListSchedulesResponse,
ListScheduleRunsResponse,
RoleInfo,
CreateRoleRequest,
UpdateRoleRequest,
ListRolesResponse,
AssignRoleRequest,
UserRoleInfo,
ListUserRolesResponse,
OrgInfo,
UpdateOrgRequest,
ListOrgsResponse,
ToolPolicyInfo,
CreateToolPolicyRequest,
UpdateToolPolicyRequest,
ListToolPoliciesResponse,
PromptTemplateInfo,
CreatePromptTemplateRequest,
UpdatePromptTemplateRequest,
ListPromptTemplatesResponse,
UsageBreakdownItem,
UsageResponse,
AuditEventInfo,
ListAuditEventsResponse,
]
+3
View File
@@ -175,6 +175,7 @@ class CreateScheduleRequest(BaseModel):
initial_message: str = Field(description="Message sent to the new workstream")
auto_approve: bool = Field(default=False)
auto_approve_tools: list[str] = Field(default_factory=list)
template: str = Field(default="", description="Prompt template name")
enabled: bool = Field(default=True)
@@ -191,6 +192,7 @@ class UpdateScheduleRequest(BaseModel):
initial_message: str | None = None
auto_approve: bool | None = None
auto_approve_tools: list[str] | None = None
template: str | None = None
enabled: bool | None = None
@@ -208,6 +210,7 @@ class ScheduleInfo(BaseModel):
initial_message: str
auto_approve: bool = False
auto_approve_tools: list[str] = Field(default_factory=list)
template: str = ""
enabled: bool = True
created_by: str = ""
last_run: str | None = None
+14
View File
@@ -35,6 +35,10 @@ class CommandRequest(BaseModel):
ws_id: str = Field(description="Target workstream ID")
class CancelRequest(BaseModel):
ws_id: str = Field(description="Target workstream ID")
class CreateWorkstreamRequest(BaseModel):
name: str = Field(default="", description="Workstream display name (auto-generated if empty)")
model: str = Field(default="", description="Model alias from registry")
@@ -43,6 +47,9 @@ class CreateWorkstreamRequest(BaseModel):
default="",
description="Workstream ID to resume atomically during creation (empty = fresh start)",
)
template: str = Field(
default="", description="Prompt template name (replaces default templates)"
)
class CreateWorkstreamResponse(BaseModel):
@@ -139,6 +146,12 @@ class WorkstreamCounts(BaseModel):
error: int = 0
class McpStatus(BaseModel):
servers: int = 0
resources: int = 0
prompts: int = 0
class HealthResponse(BaseModel):
status: str = Field(examples=["ok", "degraded"])
version: str = ""
@@ -146,3 +159,4 @@ class HealthResponse(BaseModel):
model: str = ""
workstreams: WorkstreamCounts = WorkstreamCounts()
backend: BackendStatus | None = None
mcp: McpStatus | None = None
+11
View File
@@ -19,6 +19,7 @@ from turnstone.api.schemas import (
)
from turnstone.api.server_schemas import (
ApproveRequest,
CancelRequest,
CloseWorkstreamRequest,
CommandRequest,
CreateWorkstreamRequest,
@@ -103,6 +104,15 @@ SERVER_ENDPOINTS: list[EndpointSpec] = [
error_codes=[400, 404],
tags=["Chat"],
),
EndpointSpec(
"/v1/api/cancel",
"POST",
"Cancel the active generation in a workstream",
request_model=CancelRequest,
response_model=StatusResponse,
error_codes=[400, 404],
tags=["Chat"],
),
# --- Streaming ---
EndpointSpec(
"/v1/api/events",
@@ -186,6 +196,7 @@ _ALL_MODELS: list[type[BaseModel]] = [
ApproveRequest,
PlanFeedbackRequest,
CommandRequest,
CancelRequest,
CreateWorkstreamRequest,
CreateWorkstreamResponse,
CloseWorkstreamRequest,
File diff suppressed because it is too large Load Diff
+1
View File
@@ -21,3 +21,4 @@ class ChannelConfig:
model: str = ""
auto_approve: bool = False
auto_approve_tools: list[str] = field(default_factory=list)
template: str = ""
+3
View File
@@ -49,11 +49,13 @@ class ChannelRouter:
*,
auto_approve: bool = False,
auto_approve_tools: list[str] | None = None,
template: str = "",
) -> None:
self._broker = broker
self._storage = storage
self._auto_approve = auto_approve
self._auto_approve_tools: list[str] = auto_approve_tools or []
self._template = template
self._pending: dict[str, asyncio.Event] = {}
self._pending_results: dict[str, str] = {}
self._global_task: asyncio.Task[None] | None = None
@@ -172,6 +174,7 @@ class ChannelRouter:
resume_ws=resume_ws,
auto_approve=self._auto_approve,
auto_approve_tools=list(self._auto_approve_tools),
template=self._template,
)
cid = msg.correlation_id
waiter = asyncio.Event()
+1
View File
@@ -141,6 +141,7 @@ class TurnstoneBot:
storage,
auto_approve=config.auto_approve,
auto_approve_tools=list(config.auto_approve_tools),
template=config.template,
)
self._subscribed_ws: set[str] = set()
+11 -1
View File
@@ -204,7 +204,8 @@ class TerminalUI(SessionUI):
try:
prompt_text = (
f" \001{BOLD}\002Plan ready.\001{RESET}\002 "
f"\001{DIM}\002[enter to approve, or give feedback]\001{RESET}\002 "
f"\001{DIM}\002[enter to approve, feedback to amend, "
f"ctrl-c to reject]\001{RESET}\002 "
)
resp = input(prompt_text).strip()
except EOFError:
@@ -724,6 +725,11 @@ def main() -> None:
default=None,
help="Developer instructions injected as developer message",
)
parser.add_argument(
"--template",
default=None,
help="Prompt template name (replaces default templates)",
)
parser.add_argument(
"--temperature",
type=float,
@@ -951,6 +957,7 @@ def main() -> None:
tool_search=args.tool_search,
tool_search_threshold=args.tool_search_threshold,
tool_search_max_results=args.tool_search_max_results,
template=args.template,
)
# Create workstream manager and initial workstream
@@ -1001,6 +1008,9 @@ def main() -> None:
mcp_tools = mcp_client.get_tools()
if mcp_tools:
print(f"MCP tools: {len(mcp_tools)} from {mcp_client.server_count} server(s)")
from turnstone.core.storage import get_storage as _cli_get_storage
mcp_client.set_storage(_cli_get_storage())
print("Type /help for commands, /ws for workstreams, /exit or Ctrl+D to quit.\n")
# Prompt string -- use a short display name
+140 -8
View File
@@ -150,7 +150,7 @@ class ClusterCollector:
"state": "idle",
"node": node_id,
"server_url": node.server_url,
"title": "",
"title": data.get("title", ""),
"tokens": 0,
"context_ratio": 0.0,
"activity": "",
@@ -273,6 +273,7 @@ class ClusterCollector:
"""Apply polled data to the in-memory node snapshot."""
ws_list = dashboard.get("workstreams", [])
aggregate = dashboard.get("aggregate", {})
pending_events: list[dict[str, Any]] = []
with self._lock:
node = self._nodes.get(node_id)
if not node:
@@ -281,12 +282,35 @@ class ClusterCollector:
node.reachable = True
node.health = health
node.aggregate = aggregate
# Replace workstreams entirely from the authoritative poll
node.workstreams = {}
# Build new workstream map
old_ids = {k for k in node.workstreams if k}
new_ws: dict[str, dict[str, Any]] = {}
for ws in ws_list:
ws_id = ws.get("id", "")
if not ws_id:
continue
ws["node"] = node_id
ws["server_url"] = node.server_url
node.workstreams[ws.get("id", "")] = ws
new_ws[ws_id] = ws
new_ids = set(new_ws.keys())
# Detect additions not yet known to SSE clients
for ws_id in sorted(new_ids - old_ids):
ws = new_ws[ws_id]
pending_events.append(
{
"type": "ws_created",
"ws_id": ws_id,
"name": ws.get("name", ""),
"node_id": node_id,
}
)
# Detect removals
for ws_id in sorted(old_ids - new_ids):
pending_events.append({"type": "ws_closed", "ws_id": ws_id})
node.workstreams = new_ws
# Fan out diffs to SSE listeners outside the lock
for event in pending_events:
self._fanout(event)
# -- query methods (thread-safe) -----------------------------------------
@@ -296,6 +320,9 @@ class ClusterCollector:
total_tokens = 0
total_tool_calls = 0
total_ws = 0
mcp_servers = 0
mcp_resources = 0
mcp_prompts = 0
versions: set[str] = set()
with self._lock:
for node in self._nodes.values():
@@ -308,8 +335,12 @@ class ClusterCollector:
ver = node.health.get("version", "")
if ver:
versions.add(ver)
mcp = node.health.get("mcp", {})
mcp_servers += mcp.get("servers", 0)
mcp_resources += mcp.get("resources", 0)
mcp_prompts += mcp.get("prompts", 0)
node_count = len(self._nodes)
return {
result: dict[str, Any] = {
"nodes": node_count,
"workstreams": total_ws,
"states": states,
@@ -320,6 +351,11 @@ class ClusterCollector:
"version_drift": len(versions) > 1,
"versions": sorted(versions),
}
if mcp_servers:
result["mcp_servers"] = mcp_servers
result["mcp_resources"] = mcp_resources
result["mcp_prompts"] = mcp_prompts
return result
def get_version_info(self) -> dict[str, Any]:
"""Return per-node version map and drift flag."""
@@ -379,11 +415,11 @@ class ClusterCollector:
)
total = len(items)
# Sort
# Sort (secondary key: node_id for stable ordering)
if sort_by == "activity":
items.sort(key=lambda n: n["ws_running"] + n["ws_attention"], reverse=True)
items.sort(key=lambda n: (-(n["ws_running"] + n["ws_attention"]), n["node_id"]))
elif sort_by == "tokens":
items.sort(key=lambda n: n["total_tokens"], reverse=True)
items.sort(key=lambda n: (-n["total_tokens"], n["node_id"]))
elif sort_by == "name":
items.sort(key=lambda n: n["node_id"])
@@ -455,6 +491,102 @@ class ClusterCollector:
"reachable": node.reachable,
}
def get_snapshot(self) -> dict[str, Any]:
"""Build a complete cluster snapshot under a single lock.
Returns everything the UI needs to render the full dashboard:
all nodes with their workstreams plus pre-computed overview aggregates.
"""
with self._lock:
return self._build_snapshot_locked()
def get_snapshot_and_register(self, q: queue.Queue[dict[str, Any]]) -> dict[str, Any]:
"""Build snapshot and register listener atomically.
Acquiring both locks ensures no event can be published between
the snapshot read and the listener registration the client
receives the snapshot followed by every subsequent event with
no gap.
"""
with self._lock:
snap = self._build_snapshot_locked()
with self._listeners_lock:
self._listeners.append(q)
return snap
def _build_snapshot_locked(self) -> dict[str, Any]:
"""Build snapshot data — caller must hold ``_lock``."""
nodes_out = []
states: dict[str, int] = {
"running": 0,
"thinking": 0,
"attention": 0,
"idle": 0,
"error": 0,
}
total_tokens = 0
total_tool_calls = 0
total_ws = 0
mcp_servers = 0
mcp_resources = 0
mcp_prompts = 0
versions: set[str] = set()
for node in self._nodes.values():
ws_list = []
for ws in node.workstreams.values():
ws_list.append(dict(ws))
s = ws.get("state", "idle")
states[s] = states.get(s, 0) + 1
total_ws += 1
total_tokens += node.aggregate.get("total_tokens", 0)
total_tool_calls += node.aggregate.get("total_tool_calls", 0)
ver = node.health.get("version", "")
if ver:
versions.add(ver)
mcp = node.health.get("mcp", {})
mcp_servers += mcp.get("servers", 0)
mcp_resources += mcp.get("resources", 0)
mcp_prompts += mcp.get("prompts", 0)
nodes_out.append(
{
"node_id": node.node_id,
"server_url": node.server_url,
"max_ws": node.max_ws,
"reachable": node.reachable,
"version": ver,
"health": dict(node.health),
"aggregate": dict(node.aggregate),
"workstreams": ws_list,
}
)
node_count = len(self._nodes)
overview: dict[str, Any] = {
"nodes": node_count,
"workstreams": total_ws,
"states": states,
"aggregate": {
"total_tokens": total_tokens,
"total_tool_calls": total_tool_calls,
},
"version_drift": len(versions) > 1,
"versions": sorted(versions),
}
if mcp_servers:
overview["mcp_servers"] = mcp_servers
overview["mcp_resources"] = mcp_resources
overview["mcp_prompts"] = mcp_prompts
return {
"nodes": nodes_out,
"overview": overview,
"timestamp": time.time(),
}
# -- SSE listener management ---------------------------------------------
def register_listener(self, q: queue.Queue[dict[str, Any]]) -> None:
+14
View File
@@ -112,6 +112,18 @@ class TaskScheduler:
pruned = self._storage.prune_task_runs(retention_days=90)
if pruned:
log.info("scheduler.pruned_runs", count=pruned)
try:
usage_pruned = self._storage.prune_usage_events(retention_days=90)
if usage_pruned:
log.info("scheduler.pruned_usage", count=usage_pruned)
except Exception:
log.warning("scheduler.prune_usage_error", exc_info=True)
try:
audit_pruned = self._storage.prune_audit_events(retention_days=365)
if audit_pruned:
log.info("scheduler.pruned_audit", count=audit_pruned)
except Exception:
log.warning("scheduler.prune_audit_error", exc_info=True)
finally:
# Only release our own lock (safe even if TTL expired and another took it)
self._broker._redis.eval( # type: ignore[no-untyped-call]
@@ -196,6 +208,7 @@ class TaskScheduler:
auto_approve=bool(task.get("auto_approve", 0)),
auto_approve_tools=self._parse_tools(task),
user_id=task.get("created_by", ""),
template=task.get("template", ""),
)
self._broker.push_inbound(msg.to_json(), node_id=node_id)
@@ -221,6 +234,7 @@ class TaskScheduler:
auto_approve=bool(task.get("auto_approve", 0)),
auto_approve_tools=self._parse_tools(task),
user_id=task.get("created_by", ""),
template=task.get("template", ""),
)
self._broker.push_inbound(msg.to_json())
File diff suppressed because it is too large Load Diff
+288 -10
View File
@@ -10,6 +10,7 @@ var _ctTrapHandler = null;
var _tcTrapHandler = null;
var _ccTrapHandler = null;
var _cfTrapHandler = null;
var _adminWatches = [];
var _confirmCallbackFn = null;
var _confirmTriggerEl = null;
@@ -28,11 +29,62 @@ function showAdmin() {
document.getElementById("breadcrumb-label").textContent = "Admin";
document.getElementById("main").scrollTop = 0;
history.pushState({ view: "admin" }, "");
loadAdminUsers();
// Permission gating: hide tabs the user cannot access
var perms = sessionStorage.getItem("turnstone_permissions") || "";
var tabPerms = {
users: "admin.users",
tokens: "admin.users",
channels: "admin.users",
schedules: "admin.schedules",
watches: "admin.watches",
roles: "admin.roles",
policies: "admin.policies",
templates: "admin.templates",
usage: "admin.usage",
audit: "admin.audit",
};
if (perms) {
var permSet = perms.split(",");
var tabs = document.querySelectorAll(".admin-tab");
for (var i = 0; i < tabs.length; i++) {
var tabName = tabs[i].getAttribute("data-tab");
var needed = tabPerms[tabName];
if (needed && permSet.indexOf(needed) < 0) {
tabs[i].style.display = "none";
} else {
tabs[i].style.display = "";
}
}
}
// Switch to the first visible tab
var visibleTabs = document.querySelectorAll(
'.admin-tab:not([style*="display: none"])',
);
if (visibleTabs.length > 0) {
switchAdminTab(visibleTabs[0].getAttribute("data-tab"));
} else {
// No tabs visible — show empty state instead of loading an inaccessible tab
var panels = document.querySelectorAll(".admin-panel");
for (var j = 0; j < panels.length; j++) panels[j].style.display = "none";
var empty = document.getElementById("admin-no-permissions");
if (!empty) {
empty = document.createElement("div");
empty.id = "admin-no-permissions";
empty.className = "dashboard-empty";
empty.textContent = "You do not have permissions to view any admin tabs.";
document.getElementById("view-admin").appendChild(empty);
}
empty.style.display = "";
}
}
function switchAdminTab(tab) {
_adminTab = tab;
// Hide no-permissions empty state if it was showing
var noPerms = document.getElementById("admin-no-permissions");
if (noPerms) noPerms.style.display = "none";
var tabs = document.querySelectorAll(".admin-tab");
for (var i = 0; i < tabs.length; i++) {
var isActive = tabs[i].getAttribute("data-tab") === tab;
@@ -40,19 +92,36 @@ function switchAdminTab(tab) {
tabs[i].setAttribute("aria-selected", isActive ? "true" : "false");
tabs[i].setAttribute("tabindex", isActive ? "0" : "-1");
}
document.getElementById("admin-users").style.display =
tab === "users" ? "" : "none";
document.getElementById("admin-tokens").style.display =
tab === "tokens" ? "" : "none";
document.getElementById("admin-channels").style.display =
tab === "channels" ? "" : "none";
document.getElementById("admin-schedules").style.display =
tab === "schedules" ? "" : "none";
var panels = [
"users",
"tokens",
"channels",
"schedules",
"watches",
"roles",
"policies",
"templates",
"usage",
"audit",
];
for (var p = 0; p < panels.length; p++) {
var el = document.getElementById("admin-" + panels[p]);
if (el) el.style.display = panels[p] === tab ? "" : "none";
}
if (tab === "users") loadAdminUsers();
if (tab === "tokens") _populateTokenUserSelect();
if (tab === "channels") _populateChannelUserSelect();
if (tab === "schedules") loadAdminSchedules();
if (tab === "watches") loadAdminWatches();
if (tab === "roles") loadGovRoles();
if (tab === "policies") loadGovPolicies();
if (tab === "templates") loadGovTemplates();
if (tab === "usage") loadGovUsage();
if (tab === "audit") {
_populateAuditUserFilter();
loadGovAudit();
}
}
// ---------------------------------------------------------------------------
@@ -98,6 +167,9 @@ function _renderUsers(users) {
escapeHtml(u.created || "").slice(0, 10) +
"</span>" +
'<span class="admin-col admin-col-actions">' +
'<button class="admin-btn-action" data-user-roles="' +
escapeHtml(u.user_id) +
'" title="Manage roles">roles</button>' +
'<button class="admin-btn-danger" data-delete-user="' +
escapeHtml(u.user_id) +
'" data-username="' +
@@ -107,6 +179,13 @@ function _renderUsers(users) {
"</div>";
}
container.innerHTML = html;
// Bind roles buttons
var roleBtns = container.querySelectorAll("[data-user-roles]");
for (var rj = 0; rj < roleBtns.length; rj++) {
roleBtns[rj].addEventListener("click", function () {
showUserRolesModal(this.getAttribute("data-user-roles"));
});
}
// Bind delete buttons via delegation (avoids inline JS injection)
var btns = container.querySelectorAll("[data-delete-user]");
for (var j = 0; j < btns.length; j++) {
@@ -582,6 +661,7 @@ function showCreateScheduleModal() {
document.getElementById("cs-target").value = "auto";
document.getElementById("cs-node").value = "";
document.getElementById("cs-model").value = "";
document.getElementById("cs-template").value = "";
document.getElementById("cs-message").value = "";
document.getElementById("cs-autoapprove").checked = false;
toggleScheduleTypeFields();
@@ -614,6 +694,7 @@ function submitCreateSchedule() {
var nodeId = (document.getElementById("cs-node").value || "").trim();
var model = (document.getElementById("cs-model").value || "").trim();
var message = (document.getElementById("cs-message").value || "").trim();
var template = (document.getElementById("cs-template").value || "").trim();
var autoApprove = document.getElementById("cs-autoapprove").checked;
var errEl = document.getElementById("create-schedule-error");
@@ -650,6 +731,7 @@ function submitCreateSchedule() {
model: model,
initial_message: message,
auto_approve: autoApprove,
template: template,
}),
})
.then(function (r) {
@@ -716,6 +798,7 @@ function showEditScheduleModal(taskId) {
? s.target_mode
: "";
document.getElementById("es-model").value = s.model || "";
document.getElementById("es-template").value = s.template || "";
document.getElementById("es-message").value = s.initial_message || "";
document.getElementById("es-autoapprove").checked = !!s.auto_approve;
document.getElementById("es-enabled").checked = !!s.enabled;
@@ -788,6 +871,7 @@ function submitEditSchedule() {
at_time: atTime,
target_mode: targetMode,
model: (document.getElementById("es-model").value || "").trim(),
template: (document.getElementById("es-template").value || "").trim(),
initial_message: (
document.getElementById("es-message").value || ""
).trim(),
@@ -887,6 +971,167 @@ function hideScheduleRunsModal() {
_runsScheduleTriggerEl = null;
}
// ---------------------------------------------------------------------------
// Watches
// ---------------------------------------------------------------------------
function _populateWatchNodeSelect() {
var sel = document.getElementById("admin-watch-node");
var current = sel.value;
var seen = {};
sel.innerHTML = '<option value="">All nodes</option>';
for (var i = 0; i < _adminWatches.length; i++) {
var nid = _adminWatches[i].node_id || "";
if (nid && !seen[nid]) {
seen[nid] = true;
var opt = document.createElement("option");
opt.value = nid;
opt.textContent = nid;
sel.appendChild(opt);
}
}
if (current) sel.value = current;
}
function loadAdminWatches() {
authFetch("/v1/api/admin/watches")
.then(function (r) {
if (!r.ok) throw new Error("Failed to load watches");
return r.json();
})
.then(function (data) {
_adminWatches = data.watches || [];
_populateWatchNodeSelect();
var nodeFilter = document.getElementById("admin-watch-node").value;
var filtered = _adminWatches;
if (nodeFilter) {
filtered = _adminWatches.filter(function (w) {
return w.node_id === nodeFilter;
});
}
_renderWatches(filtered);
})
.catch(function () {
document.getElementById("admin-watches-table").innerHTML =
'<div class="dashboard-empty">Failed to load watches</div>';
});
}
function _formatInterval(secs) {
if (!secs || secs <= 0) return "\u2014";
if (secs >= 3600) return Math.round(secs / 3600) + "h";
if (secs >= 60) return Math.round(secs / 60) + "m";
return secs + "s";
}
function _renderWatches(watches) {
var container = document.getElementById("admin-watches-table");
if (!watches.length) {
container.innerHTML =
'<div class="dashboard-empty">No active watches. Watches are created when workstreams use the watch tool.</div>';
return;
}
var html = "";
for (var i = 0; i < watches.length; i++) {
var w = watches[i];
var name = w.name || w.watch_id || "\u2014";
var nodeShort = (w.node_id || "").slice(0, 8);
var cmd = w.command || "";
var cmdTrunc = cmd.length > 40 ? cmd.slice(0, 40) + "\u2026" : cmd;
var interval = _formatInterval(w.interval_secs);
var pollMax = w.max_polls ? w.max_polls : "\u221e";
var pollLabel = (w.poll_count || 0) + "/" + pollMax;
var cond = w.stop_on || "on change";
var condTrunc = cond.length > 30 ? cond.slice(0, 30) + "\u2026" : cond;
var active = w.active;
var statusCls = active ? "watch-active" : "watch-completed";
var statusLabel = active ? "active" : "done";
var statusDot = active ? "\u25cf " : "\u25cb ";
var cancelBtn = active
? '<button class="admin-btn-danger" data-cancel-watch="' +
escapeHtml(w.watch_id) +
'" data-watch-node="' +
escapeHtml(w.node_id || "") +
'" data-watch-name="' +
escapeHtml(name) +
'" title="Cancel watch">cancel</button>'
: "";
html +=
'<div class="admin-row" role="listitem">' +
'<span class="admin-col admin-col-wname">' +
escapeHtml(name) +
"</span>" +
'<span class="admin-col admin-col-wnode" title="' +
escapeHtml(w.node_id || "") +
'"><code>' +
escapeHtml(nodeShort) +
"</code></span>" +
'<span class="admin-col admin-col-wcmd" title="' +
escapeHtml(cmd) +
'"><code>' +
escapeHtml(cmdTrunc) +
"</code></span>" +
'<span class="admin-col admin-col-winterval">' +
escapeHtml(interval) +
"</span>" +
'<span class="admin-col admin-col-wpoll"><code>' +
escapeHtml(pollLabel) +
"</code></span>" +
'<span class="admin-col admin-col-wcond" title="' +
escapeHtml(cond) +
'">' +
escapeHtml(condTrunc) +
"</span>" +
'<span class="admin-col admin-col-wstatus"><span class="' +
statusCls +
'">' +
statusDot +
statusLabel +
"</span></span>" +
'<span class="admin-col admin-col-actions">' +
cancelBtn +
"</span></div>";
}
container.innerHTML = html;
// Bind cancel buttons
var btns = container.querySelectorAll("[data-cancel-watch]");
for (var j = 0; j < btns.length; j++) {
btns[j].addEventListener("click", function () {
_cancelWatch(
this.getAttribute("data-cancel-watch"),
this.getAttribute("data-watch-node"),
this.getAttribute("data-watch-name"),
);
});
}
}
function _cancelWatch(watchId, nodeId, name) {
showConfirmModal(
"Cancel Watch",
"Cancel watch \u2018" + name + "\u2019? This will stop future polling.",
"Cancel watch",
function () {
authFetch(
"/v1/api/admin/watches/" + encodeURIComponent(watchId) + "/cancel",
{
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ node_id: nodeId }),
},
)
.then(function (r) {
if (!r.ok) throw new Error("Cancel failed");
showToast("Watch '" + name + "' cancelled");
loadAdminWatches();
})
.catch(function () {
showToast("Failed to cancel watch");
});
},
);
}
// ---------------------------------------------------------------------------
// Create Channel Link Modal
// ---------------------------------------------------------------------------
@@ -1197,6 +1442,14 @@ function _installTrap(overlayId, boxId, trapRef) {
else if (overlayId === "edit-schedule-overlay") hideEditScheduleModal();
else if (overlayId === "schedule-runs-overlay") hideScheduleRunsModal();
else if (overlayId === "confirm-overlay") hideConfirmModal();
else if (overlayId === "create-role-overlay") hideCreateRoleModal();
else if (overlayId === "edit-role-overlay") hideEditRoleModal();
else if (overlayId === "user-roles-overlay") hideUserRolesModal();
else if (overlayId === "create-policy-overlay") hideCreatePolicyModal();
else if (overlayId === "edit-policy-overlay") hideEditPolicyModal();
else if (overlayId === "create-template-overlay")
hideCreateTemplateModal();
else if (overlayId === "edit-template-overlay") hideEditTemplateModal();
}
};
}
@@ -1263,6 +1516,24 @@ document.addEventListener("keydown", function (e) {
hideConfirmModal();
return;
}
// Governance modals
var govOverlays = [
["create-role-overlay", hideCreateRoleModal],
["edit-role-overlay", hideEditRoleModal],
["user-roles-overlay", hideUserRolesModal],
["create-policy-overlay", hideCreatePolicyModal],
["edit-policy-overlay", hideEditPolicyModal],
["create-template-overlay", hideCreateTemplateModal],
["edit-template-overlay", hideEditTemplateModal],
];
for (var gi = 0; gi < govOverlays.length; gi++) {
var govEl = document.getElementById(govOverlays[gi][0]);
if (govEl && govEl.style.display !== "none") {
e.preventDefault();
govOverlays[gi][1]();
return;
}
}
});
// Tab arrow key navigation
@@ -1271,7 +1542,14 @@ document.addEventListener("keydown", function (e) {
if (!tablist) return;
tablist.addEventListener("keydown", function (e) {
if (e.key !== "ArrowLeft" && e.key !== "ArrowRight") return;
var tabOrder = ["users", "tokens", "channels", "schedules"];
var allTabs = document.querySelectorAll(
'.admin-tab:not([style*="display: none"])',
);
var tabOrder = [];
for (var ti = 0; ti < allTabs.length; ti++) {
tabOrder.push(allTabs[ti].getAttribute("data-tab"));
}
if (tabOrder.length === 0) return;
var idx = tabOrder.indexOf(_adminTab);
if (e.key === "ArrowRight") idx = (idx + 1) % tabOrder.length;
else idx = (idx - 1 + tabOrder.length) % tabOrder.length;
+403 -118
View File
@@ -1,9 +1,6 @@
// --- Shared hooks ---
window.onLoginSuccess = function () {
connectSSE();
if (currentView === "overview") loadOverview();
else if (currentView === "node") drillDownToNode(currentNodeId);
else if (currentView === "filtered") loadFilteredWorkstreams();
};
window.onLogout = function () {
if (evtSource) {
@@ -33,6 +30,8 @@ var _lastOverviewJson = "";
var _lastNodesJson = "";
var evtSource = null;
var retryDelay = 1000;
var clusterState = null;
var _navigatingFromPopstate = false;
// --- Constants ---
var STATE_DISPLAY = {
@@ -44,6 +43,265 @@ var STATE_DISPLAY = {
};
var STATE_ORDER = ["running", "thinking", "attention", "error", "idle"];
// --- Cluster State Model ---
function applySnapshot(data) {
clusterState = {
nodes: {},
overview: data.overview || {},
timestamp: data.timestamp || 0,
};
(data.nodes || []).forEach(function (n) {
clusterState.nodes[n.node_id] = n;
});
renderFromState();
}
function patchClusterState(data) {
if (!clusterState) return;
var t = data.type;
if (t === "cluster_state") {
var node = clusterState.nodes[data.node_id];
if (node) {
(node.workstreams || []).forEach(function (ws) {
if (ws.id === data.ws_id) {
if ("state" in data) ws.state = data.state;
if ("tokens" in data) ws.tokens = data.tokens;
if ("context_ratio" in data) ws.context_ratio = data.context_ratio;
if ("activity" in data) ws.activity = data.activity;
if ("activity_state" in data) ws.activity_state = data.activity_state;
}
});
}
} else if (t === "ws_created") {
var targetNode = clusterState.nodes[data.node_id];
if (targetNode) {
targetNode.workstreams = targetNode.workstreams || [];
targetNode.workstreams.push({
id: data.ws_id,
name: data.name || "",
state: "idle",
node: data.node_id,
server_url: targetNode.server_url || "",
title: data.title || "",
tokens: 0,
context_ratio: 0.0,
activity: "",
activity_state: "",
tool_calls: 0,
});
}
} else if (t === "ws_closed") {
Object.keys(clusterState.nodes).forEach(function (nid) {
var n = clusterState.nodes[nid];
n.workstreams = (n.workstreams || []).filter(function (ws) {
return ws.id !== data.ws_id;
});
});
} else if (t === "ws_rename") {
Object.keys(clusterState.nodes).forEach(function (nid) {
(clusterState.nodes[nid].workstreams || []).forEach(function (ws) {
if (ws.id === data.ws_id) ws.name = data.name || "";
});
});
} else if (t === "node_joined") {
if (!clusterState.nodes[data.node_id]) {
clusterState.nodes[data.node_id] = {
node_id: data.node_id,
server_url: "",
max_ws: 10,
reachable: true,
version: "",
health: {},
aggregate: {},
workstreams: [],
};
}
} else if (t === "node_lost") {
delete clusterState.nodes[data.node_id];
} else {
return;
}
scheduleRender();
}
var _renderTimer = null;
function scheduleRender() {
if (_renderTimer) return;
_renderTimer = requestAnimationFrame(function () {
_renderTimer = null;
recomputeOverview();
renderFromState();
});
}
function recomputeOverview() {
if (!clusterState) return;
var states = { running: 0, thinking: 0, attention: 0, idle: 0, error: 0 };
var totalTokens = 0,
totalToolCalls = 0,
totalWs = 0;
var mcpServers = 0,
mcpResources = 0,
mcpPrompts = 0;
var versions = {};
Object.keys(clusterState.nodes).forEach(function (nid) {
var node = clusterState.nodes[nid];
var nodeWsTokens = 0;
(node.workstreams || []).forEach(function (ws) {
var s = ws.state || "idle";
states[s] = (states[s] || 0) + 1;
totalWs++;
nodeWsTokens += ws.tokens || 0;
});
var aggTokens = (node.aggregate || {}).total_tokens || 0;
totalTokens += aggTokens || nodeWsTokens;
totalToolCalls += (node.aggregate || {}).total_tool_calls || 0;
if (node.version) versions[node.version] = true;
var mcp = (node.health || {}).mcp || {};
mcpServers += mcp.servers || 0;
mcpResources += mcp.resources || 0;
mcpPrompts += mcp.prompts || 0;
});
var versionList = Object.keys(versions).sort();
clusterState.overview = {
nodes: Object.keys(clusterState.nodes).length,
workstreams: totalWs,
states: states,
aggregate: {
total_tokens: totalTokens,
total_tool_calls: totalToolCalls,
},
version_drift: versionList.length > 1,
versions: versionList,
};
if (mcpServers > 0) {
clusterState.overview.mcp_servers = mcpServers;
clusterState.overview.mcp_resources = mcpResources;
clusterState.overview.mcp_prompts = mcpPrompts;
}
}
function buildNodeInfoFromSnapshot(node) {
var states = { running: 0, thinking: 0, attention: 0, idle: 0, error: 0 };
var ws = node.workstreams || [];
ws.forEach(function (w) {
var s = w.state || "idle";
states[s] = (states[s] || 0) + 1;
});
var aggTokens = (node.aggregate || {}).total_tokens || 0;
if (!aggTokens) {
ws.forEach(function (w) {
aggTokens += w.tokens || 0;
});
}
return {
node_id: node.node_id,
server_url: node.server_url || "",
ws_total: ws.length,
ws_running: states.running,
ws_thinking: states.thinking,
ws_attention: states.attention,
ws_idle: states.idle,
ws_error: states.error,
total_tokens: aggTokens,
ws_tokens: aggTokens,
max_ws: node.max_ws || 10,
started: node.started || 0,
reachable: node.reachable !== false,
health: node.health || {},
version: node.version || "",
};
}
function renderFromState() {
if (!clusterState) return;
renderStatusBar(clusterState.overview);
if (currentView === "overview") {
var nodesList = Object.keys(clusterState.nodes).map(function (nid) {
return buildNodeInfoFromSnapshot(clusterState.nodes[nid]);
});
nodesList.sort(function (a, b) {
var d = b.ws_running + b.ws_attention - (a.ws_running + a.ws_attention);
return d !== 0 ? d : a.node_id.localeCompare(b.node_id);
});
renderNodeGroups(nodesList, nodesList.length);
document.getElementById("cluster-summary").textContent =
clusterState.overview.nodes +
" nodes \u00b7 " +
formatCount(clusterState.overview.workstreams) +
" workstreams";
} else if (currentView === "node" && currentNodeId) {
var snapNode = clusterState.nodes[currentNodeId];
if (snapNode) {
var wsList = snapNode.workstreams || [];
var active = wsList.filter(function (w) {
return w.state !== "idle";
}).length;
document.getElementById("node-ws-summary").textContent =
active + " active \u00b7 " + wsList.length + " total";
var mcpSumEl = document.getElementById("node-mcp-summary");
if (mcpSumEl) {
var mcpInfo = snapNode.health && snapNode.health.mcp;
if (mcpInfo && mcpInfo.servers > 0) {
mcpSumEl.textContent =
mcpInfo.servers +
" MCP server" +
(mcpInfo.servers !== 1 ? "s" : "") +
" \u00b7 " +
mcpInfo.resources +
" resources \u00b7 " +
mcpInfo.prompts +
" prompts";
} else {
mcpSumEl.textContent = "";
}
}
renderWsTable(document.getElementById("node-ws-table"), wsList);
}
} else if (currentView === "filtered") {
var allWs = [];
Object.keys(clusterState.nodes).forEach(function (nid) {
(clusterState.nodes[nid].workstreams || []).forEach(function (ws) {
allWs.push(ws);
});
});
if (currentFilter.state) {
allWs = allWs.filter(function (ws) {
return ws.state === currentFilter.state;
});
}
if (currentFilter.node) {
allWs = allWs.filter(function (ws) {
return ws.node === currentFilter.node;
});
}
var stateOrder = {
running: 0,
thinking: 1,
attention: 2,
error: 3,
idle: 4,
};
allWs.sort(function (a, b) {
return (stateOrder[a.state] || 9) - (stateOrder[b.state] || 9);
});
var total = allWs.length;
var perPage = currentFilter.per_page || 50;
var pages = Math.max(1, Math.ceil(total / perPage));
var page = Math.min(currentFilter.page || 1, pages);
var start = (page - 1) * perPage;
var pageWs = allWs.slice(start, start + perPage);
document.getElementById("filtered-summary").textContent =
"Page " + page + " of " + pages + " (" + total + " total)";
renderWsTable(document.getElementById("filtered-ws-table"), pageWs);
renderPagination(
document.getElementById("filtered-pagination"),
page,
pages,
);
}
}
// --- SSE Connection ---
function connectSSE() {
if (evtSource) {
@@ -94,19 +352,11 @@ function connectSSE() {
};
}
var _refreshTimer = null;
function scheduleRefresh() {
if (_refreshTimer) return;
_refreshTimer = setTimeout(function () {
_refreshTimer = null;
if (currentView === "overview") loadOverview();
else if (currentView === "node" && currentNodeId)
loadNodeDetail(currentNodeId);
else if (currentView === "filtered") loadFilteredWorkstreams();
}, 250);
}
function handleClusterEvent(data) {
if (data.type === "snapshot") {
applySnapshot(data);
return;
}
if (
data.type === "cluster_state" ||
data.type === "ws_created" ||
@@ -115,7 +365,7 @@ function handleClusterEvent(data) {
data.type === "node_joined" ||
data.type === "node_lost"
) {
scheduleRefresh();
patchClusterState(data);
}
if (data.type === "ws_closed" && data.reason === "evicted") {
showToast("Evicted" + (data.name ? ": " + data.name : "") + " (capacity)");
@@ -135,28 +385,18 @@ function showOverview() {
if (adminView) adminView.style.display = "none";
document.getElementById("breadcrumb").style.display = "none";
document.getElementById("main").scrollTop = 0;
loadOverview();
history.pushState({ view: "overview" }, "");
if (clusterState) renderFromState();
else loadOverview();
if (!_navigatingFromPopstate) history.pushState({ view: "overview" }, "");
}
function loadOverview() {
var overviewP = authFetch("/v1/api/cluster/overview").then(function (r) {
return r.json();
});
var nodesP = authFetch("/v1/api/cluster/nodes?sort=activity&limit=1000").then(
function (r) {
authFetch("/v1/api/cluster/snapshot")
.then(function (r) {
return r.json();
},
);
Promise.all([overviewP, nodesP])
.then(function (res) {
renderStatusBar(res[0]);
renderNodeGroups(res[1].nodes, res[1].total);
document.getElementById("cluster-summary").textContent =
res[0].nodes +
" nodes \u00b7 " +
formatCount(res[0].workstreams) +
" workstreams";
})
.then(function (data) {
applySnapshot(data);
})
.catch(function () {
document.getElementById("node-table").innerHTML =
@@ -259,6 +499,43 @@ function renderStatusBar(overview) {
verEl.appendChild(verLbl);
metricsContainer.appendChild(verEl);
}
// MCP aggregate metrics
if (overview.mcp_servers && overview.mcp_servers > 0) {
var mcpDivider = document.createElement("span");
mcpDivider.className = "csb-divider";
mcpDivider.setAttribute("aria-hidden", "true");
metricsContainer.appendChild(mcpDivider);
var mcpTitles = {
mcp: "MCP servers",
rsrc: "MCP resources",
pmpt: "MCP prompts",
};
var mcpMetrics = [
{ value: overview.mcp_servers, label: "mcp" },
{ value: overview.mcp_resources, label: "rsrc" },
{ value: overview.mcp_prompts, label: "pmpt" },
];
mcpMetrics.forEach(function (m) {
var el = document.createElement("span");
el.className = "csb-metric";
el.title = mcpTitles[m.label] || "";
if (m.label === "mcp") {
var dot = document.createElement("span");
dot.className = "csb-mcp-dot";
dot.setAttribute("aria-hidden", "true");
el.appendChild(dot);
}
var valSpan = document.createElement("span");
valSpan.className = "csb-metric-value";
valSpan.textContent = formatCount(m.value);
var labelSpan = document.createElement("span");
labelSpan.className = "csb-metric-label";
labelSpan.textContent = m.label;
el.appendChild(valSpan);
el.appendChild(labelSpan);
metricsContainer.appendChild(el);
});
}
}
// --- Node Grouping ---
@@ -310,7 +587,8 @@ function groupNodes(nodes) {
});
groupOrder.forEach(function (prefix) {
groupMap[prefix].nodes.sort(function (a, b) {
return b.ws_running + b.ws_attention - (a.ws_running + a.ws_attention);
var d = b.ws_running + b.ws_attention - (a.ws_running + a.ws_attention);
return d !== 0 ? d : a.node_id.localeCompare(b.node_id);
});
});
var groups = groupOrder.map(function (p) {
@@ -653,38 +931,37 @@ function drillDownToNode(nodeId, serverUrl) {
link.href = "/node/" + encodeURIComponent(nodeId) + "/";
link.style.display = "";
document.getElementById("main").scrollTop = 0;
document.getElementById("node-ws-table").innerHTML =
'<div class="dashboard-empty">Loading workstreams...</div>';
loadNodeDetail(nodeId);
if (clusterState && clusterState.nodes[nodeId]) {
renderFromState();
} else {
document.getElementById("node-ws-table").innerHTML =
'<div class="dashboard-empty">Loading workstreams...</div>';
loadNodeDetail(nodeId);
}
document.getElementById("breadcrumb-home").focus();
history.pushState({ view: "node", nodeId: nodeId, serverUrl: serverUrl }, "");
if (!_navigatingFromPopstate)
history.pushState(
{ view: "node", nodeId: nodeId, serverUrl: serverUrl },
"",
);
}
function loadNodeDetail(nodeId) {
var detailP = authFetch(
"/v1/api/cluster/node/" + encodeURIComponent(nodeId),
).then(function (r) {
return r.json();
});
var overviewP = authFetch("/v1/api/cluster/overview").then(function (r) {
return r.json();
});
Promise.all([detailP, overviewP]).then(function (res) {
var data = res[0];
renderStatusBar(res[1]);
if (data.error) {
authFetch("/v1/api/cluster/snapshot")
.then(function (r) {
return r.json();
})
.then(function (data) {
applySnapshot(data);
if (!clusterState || !clusterState.nodes[nodeId]) {
document.getElementById("node-ws-table").innerHTML =
'<div class="dashboard-empty">Node not found</div>';
}
})
.catch(function () {
document.getElementById("node-ws-table").innerHTML =
'<div class="dashboard-empty">' + escapeHtml(data.error) + "</div>";
return;
}
var ws = data.workstreams || [];
var active = ws.filter(function (w) {
return w.state !== "idle";
}).length;
document.getElementById("node-ws-summary").textContent =
active + " active \u00b7 " + ws.length + " total";
renderWsTable(document.getElementById("node-ws-table"), ws);
});
'<div class="dashboard-empty">Failed to load</div>';
});
}
// --- Drill-down: Filtered ---
@@ -703,9 +980,11 @@ function drillDownByState(state) {
document.getElementById("filtered-title").textContent =
"WORKSTREAMS — " + sd.label.toUpperCase();
document.getElementById("main").scrollTop = 0;
loadFilteredWorkstreams();
if (clusterState) renderFromState();
else loadFilteredWorkstreams();
document.getElementById("breadcrumb-home").focus();
history.pushState({ view: "filtered", filter: currentFilter }, "");
if (!_navigatingFromPopstate)
history.pushState({ view: "filtered", filter: currentFilter }, "");
}
function drillDownByNode(nodeId) {
@@ -721,48 +1000,20 @@ function drillDownByNode(nodeId) {
document.getElementById("filtered-title").textContent =
"WORKSTREAMS — " + nodeId;
document.getElementById("main").scrollTop = 0;
loadFilteredWorkstreams();
if (clusterState) renderFromState();
else loadFilteredWorkstreams();
document.getElementById("breadcrumb-home").focus();
history.pushState({ view: "filtered", filter: currentFilter }, "");
if (!_navigatingFromPopstate)
history.pushState({ view: "filtered", filter: currentFilter }, "");
}
function loadFilteredWorkstreams() {
var params =
"page=" + currentFilter.page + "&per_page=" + currentFilter.per_page;
if (currentFilter.state)
params += "&state=" + encodeURIComponent(currentFilter.state);
if (currentFilter.node)
params += "&node=" + encodeURIComponent(currentFilter.node);
var wsP = authFetch("/v1/api/cluster/workstreams?" + params).then(
function (r) {
authFetch("/v1/api/cluster/snapshot")
.then(function (r) {
return r.json();
},
);
var overviewP = authFetch("/v1/api/cluster/overview").then(function (r) {
return r.json();
});
Promise.all([wsP, overviewP])
.then(function (res) {
var data = res[0];
renderStatusBar(res[1]);
document.getElementById("main").scrollTop = 0;
document.getElementById("filtered-summary").textContent =
"Page " +
data.page +
" of " +
data.pages +
" (" +
data.total +
" total)";
renderWsTable(
document.getElementById("filtered-ws-table"),
data.workstreams,
);
renderPagination(
document.getElementById("filtered-pagination"),
data.page,
data.pages,
);
})
.then(function (data) {
applySnapshot(data);
})
.catch(function () {
document.getElementById("filtered-ws-table").innerHTML =
@@ -778,7 +1029,8 @@ function renderPagination(container, page, pages) {
prev.disabled = page <= 1;
prev.onclick = function () {
currentFilter.page--;
loadFilteredWorkstreams();
if (clusterState) renderFromState();
else loadFilteredWorkstreams();
};
container.appendChild(prev);
var info = document.createElement("span");
@@ -789,7 +1041,8 @@ function renderPagination(container, page, pages) {
next.disabled = page >= pages;
next.onclick = function () {
currentFilter.page++;
loadFilteredWorkstreams();
if (clusterState) renderFromState();
else loadFilteredWorkstreams();
};
container.appendChild(next);
}
@@ -923,19 +1176,24 @@ function renderWsTable(container, wsList) {
window.addEventListener("popstate", function (e) {
var overlay = document.getElementById("login-overlay");
if (overlay && overlay.style.display !== "none") return;
if (!e.state) {
showOverview();
return;
}
if (e.state.view === "overview") showOverview();
else if (e.state.view === "admin" && typeof showAdmin === "function")
showAdmin();
else if (e.state.view === "node" && e.state.nodeId)
drillDownToNode(e.state.nodeId, e.state.serverUrl);
else if (e.state.view === "filtered" && e.state.filter) {
currentFilter = e.state.filter;
if (currentFilter.state) drillDownByState(currentFilter.state);
else if (currentFilter.node) drillDownByNode(currentFilter.node);
_navigatingFromPopstate = true;
try {
if (!e.state) {
showOverview();
return;
}
if (e.state.view === "overview") showOverview();
else if (e.state.view === "admin" && typeof showAdmin === "function")
showAdmin();
else if (e.state.view === "node" && e.state.nodeId)
drillDownToNode(e.state.nodeId, e.state.serverUrl);
else if (e.state.view === "filtered" && e.state.filter) {
currentFilter = e.state.filter;
if (currentFilter.state) drillDownByState(currentFilter.state);
else if (currentFilter.node) drillDownByNode(currentFilter.node);
}
} finally {
_navigatingFromPopstate = false;
}
});
@@ -982,6 +1240,27 @@ function showNewWsModal() {
.catch(function () {
/* ignore — auto is always available */
});
// Populate template dropdown
var tplSelect = document.getElementById("new-ws-template");
tplSelect.innerHTML = '<option value="">Use defaults</option>';
authFetch("/v1/api/admin/templates")
.then(function (r) {
return r.json();
})
.then(function (data) {
(data.templates || []).forEach(function (t) {
var opt = document.createElement("option");
opt.value = t.name;
var label = t.name;
if (t.is_default) label += " (default)";
if (t.origin === "mcp") label += " [MCP]";
opt.textContent = label;
tplSelect.appendChild(opt);
});
})
.catch(function () {
/* ignore — defaults still work */
});
document.getElementById("new-ws-name").value = "";
document.getElementById("new-ws-model").value = "";
document.getElementById("new-ws-task").value = "";
@@ -998,7 +1277,7 @@ function showNewWsModal() {
_newWsTrapHandler = function (e) {
if (e.key === "Tab") {
var box = document.getElementById("new-ws-box");
var focusable = box.querySelectorAll("select, input, button");
var focusable = box.querySelectorAll("select, input, textarea, button");
var first = focusable[0];
var last = focusable[focusable.length - 1];
if (e.shiftKey) {
@@ -1036,6 +1315,7 @@ function submitNewWs() {
var nodeId = document.getElementById("new-ws-node").value;
var name = document.getElementById("new-ws-name").value.trim();
var model = document.getElementById("new-ws-model").value.trim();
var template = document.getElementById("new-ws-template").value;
var task = document.getElementById("new-ws-task").value.trim();
var errEl = document.getElementById("new-ws-error");
var btn = document.getElementById("new-ws-submit");
@@ -1049,6 +1329,7 @@ function submitNewWs() {
if (name) body.name = name;
if (model) body.model = model;
if (task) body.initial_message = task;
if (template) body.template = template;
authFetch("/v1/api/cluster/workstreams/new", {
method: "POST",
@@ -1089,7 +1370,11 @@ document.addEventListener("keydown", function (e) {
e.preventDefault();
hideNewWsModal();
}
if (e.key === "Enter" && e.target.tagName !== "SELECT") {
if (
e.key === "Enter" &&
e.target.tagName !== "SELECT" &&
e.target.tagName !== "TEXTAREA"
) {
e.preventDefault();
var btn = document.getElementById("new-ws-submit");
if (btn && !btn.disabled) submitNewWs();
File diff suppressed because it is too large Load Diff
+306
View File
@@ -41,6 +41,7 @@
<div class="dash-header">
<span class="dash-header-title">WORKSTREAMS</span>
<span class="dash-header-summary" id="node-ws-summary"></span>
<span id="node-mcp-summary" aria-label="MCP status"></span>
</div>
<div class="dash-colheaders" aria-hidden="true">
<span class="dash-col dash-col-state">STATE</span>
@@ -81,6 +82,12 @@
<button id="tab-tokens" class="admin-tab" data-tab="tokens" role="tab" aria-selected="false" aria-controls="admin-tokens" tabindex="-1" onclick="switchAdminTab('tokens')">Tokens</button>
<button id="tab-channels" class="admin-tab" data-tab="channels" role="tab" aria-selected="false" aria-controls="admin-channels" tabindex="-1" onclick="switchAdminTab('channels')">Channels</button>
<button id="tab-schedules" class="admin-tab" data-tab="schedules" role="tab" aria-selected="false" aria-controls="admin-schedules" tabindex="-1" onclick="switchAdminTab('schedules')">Schedules</button>
<button id="tab-watches" class="admin-tab" data-tab="watches" role="tab" aria-selected="false" aria-controls="admin-watches" tabindex="-1" onclick="switchAdminTab('watches')">Watches</button>
<button id="tab-roles" class="admin-tab" data-tab="roles" role="tab" aria-selected="false" aria-controls="admin-roles" tabindex="-1" onclick="switchAdminTab('roles')">Roles</button>
<button id="tab-policies" class="admin-tab" data-tab="policies" role="tab" aria-selected="false" aria-controls="admin-policies" tabindex="-1" onclick="switchAdminTab('policies')">Policies</button>
<button id="tab-templates" class="admin-tab" data-tab="templates" role="tab" aria-selected="false" aria-controls="admin-templates" tabindex="-1" onclick="switchAdminTab('templates')">Templates</button>
<button id="tab-usage" class="admin-tab" data-tab="usage" role="tab" aria-selected="false" aria-controls="admin-usage" tabindex="-1" onclick="switchAdminTab('usage')">Usage</button>
<button id="tab-audit" class="admin-tab" data-tab="audit" role="tab" aria-selected="false" aria-controls="admin-audit" tabindex="-1" onclick="switchAdminTab('audit')">Audit</button>
</div>
<!-- Users Tab -->
@@ -163,6 +170,142 @@
<div class="dashboard-empty">Loading schedules...</div>
</div>
</div>
<!-- Watches Tab -->
<div id="admin-watches" class="admin-panel" role="tabpanel" aria-labelledby="tab-watches" style="display:none">
<div class="admin-toolbar">
<span class="section-header" style="margin:0">WATCHES</span>
<label for="admin-watch-node" class="sr-only">Filter watches by node</label>
<select id="admin-watch-node" onchange="loadAdminWatches()">
<option value="">All nodes</option>
</select>
</div>
<div class="admin-colheaders" aria-hidden="true">
<span class="admin-col admin-col-wname">NAME</span>
<span class="admin-col admin-col-wnode">NODE</span>
<span class="admin-col admin-col-wcmd">COMMAND</span>
<span class="admin-col admin-col-winterval">INTERVAL</span>
<span class="admin-col admin-col-wpoll">POLL</span>
<span class="admin-col admin-col-wcond">CONDITION</span>
<span class="admin-col admin-col-wstatus">STATUS</span>
<span class="admin-col admin-col-actions">ACTIONS</span>
</div>
<div id="admin-watches-table" role="list" aria-label="Watches" aria-live="polite">
<div class="dashboard-empty">Loading watches...</div>
</div>
</div>
<!-- Roles Tab -->
<div id="admin-roles" class="admin-panel" role="tabpanel" aria-labelledby="tab-roles" style="display:none">
<div class="admin-toolbar">
<span class="section-header" style="margin:0">ROLES</span>
<button class="admin-action-btn" onclick="showCreateRoleModal()">+ Create role</button>
</div>
<div class="admin-colheaders" aria-hidden="true">
<span class="admin-col admin-col-rname">NAME</span>
<span class="admin-col admin-col-rperms">PERMISSIONS</span>
<span class="admin-col admin-col-actions">ACTIONS</span>
</div>
<div id="admin-roles-table" role="list" aria-label="Roles" aria-live="polite">
<div class="dashboard-empty">Loading roles...</div>
</div>
</div>
<!-- Policies Tab -->
<div id="admin-policies" class="admin-panel" role="tabpanel" aria-labelledby="tab-policies" style="display:none">
<div class="admin-toolbar">
<span class="section-header" style="margin:0">TOOL POLICIES</span>
<button class="admin-action-btn" onclick="showCreatePolicyModal()">+ Create policy</button>
</div>
<div class="admin-colheaders" aria-hidden="true">
<span class="admin-col admin-col-pname">NAME</span>
<span class="admin-col admin-col-ppattern">PATTERN</span>
<span class="admin-col admin-col-paction">ACTION</span>
<span class="admin-col admin-col-ppriority">PRI</span>
<span class="admin-col admin-col-pstatus">STATUS</span>
<span class="admin-col admin-col-actions">ACTIONS</span>
</div>
<div id="admin-policies-table" role="list" aria-label="Tool policies" aria-live="polite">
<div class="dashboard-empty">Loading policies...</div>
</div>
</div>
<!-- Templates Tab -->
<div id="admin-templates" class="admin-panel" role="tabpanel" aria-labelledby="tab-templates" style="display:none">
<div class="admin-toolbar">
<span class="section-header" style="margin:0">PROMPT TEMPLATES</span>
<button class="admin-action-btn" onclick="showCreateTemplateModal()">+ Create template</button>
</div>
<div class="admin-colheaders" aria-hidden="true">
<span class="admin-col admin-col-tmname">NAME</span>
<span class="admin-col admin-col-tmcat">CATEGORY</span>
<span class="admin-col admin-col-tmvars">VARIABLES</span>
<span class="admin-col admin-col-actions">ACTIONS</span>
</div>
<div id="admin-templates-table" role="list" aria-label="Prompt templates" aria-live="polite">
<div class="dashboard-empty">Loading templates...</div>
</div>
</div>
<!-- Usage Tab -->
<div id="admin-usage" class="admin-panel" role="tabpanel" aria-labelledby="tab-usage" style="display:none">
<div class="admin-toolbar">
<span class="section-header" style="margin:0">USAGE</span>
<div class="usage-range-group" role="group" aria-label="Time range">
<button class="usage-range-btn" data-range="24h" aria-pressed="false" onclick="setUsageRange('24h')">24h</button>
<button class="usage-range-btn active" data-range="7d" aria-pressed="true" onclick="setUsageRange('7d')">7d</button>
<button class="usage-range-btn" data-range="30d" aria-pressed="false" onclick="setUsageRange('30d')">30d</button>
</div>
<div class="usage-range-group" role="group" aria-label="Group by">
<button class="usage-group-btn active" data-group="day" aria-pressed="true" onclick="setUsageGroupBy('day')">day</button>
<button class="usage-group-btn" data-group="model" aria-pressed="false" onclick="setUsageGroupBy('model')">model</button>
<button class="usage-group-btn" data-group="user" aria-pressed="false" onclick="setUsageGroupBy('user')">user</button>
</div>
</div>
<div id="admin-usage-content">
<div class="dashboard-empty">Loading usage data...</div>
</div>
</div>
<!-- Audit Tab -->
<div id="admin-audit" class="admin-panel" role="tabpanel" aria-labelledby="tab-audit" style="display:none">
<div class="admin-toolbar">
<span class="section-header" style="margin:0">AUDIT LOG</span>
<label for="audit-action-filter" class="sr-only">Filter by action</label>
<select id="audit-action-filter" onchange="loadGovAudit()">
<option value="">All actions</option>
<option value="user.create">user.create</option>
<option value="user.delete">user.delete</option>
<option value="token.create">token.create</option>
<option value="token.revoke">token.revoke</option>
<option value="role.create">role.create</option>
<option value="role.update">role.update</option>
<option value="role.delete">role.delete</option>
<option value="role.assign">role.assign</option>
<option value="role.unassign">role.unassign</option>
<option value="policy.create">policy.create</option>
<option value="policy.update">policy.update</option>
<option value="policy.delete">policy.delete</option>
<option value="template.create">template.create</option>
<option value="template.update">template.update</option>
<option value="template.delete">template.delete</option>
</select>
<label for="audit-user-filter" class="sr-only">Filter by user</label>
<select id="audit-user-filter" onchange="loadGovAudit()">
<option value="">All users</option>
</select>
</div>
<div class="admin-colheaders" aria-hidden="true">
<span class="admin-col admin-col-atime">TIME</span>
<span class="admin-col admin-col-auser">USER</span>
<span class="admin-col admin-col-aaction">ACTION</span>
<span class="admin-col admin-col-aresource">RESOURCE</span>
<span class="admin-col admin-col-adetail">DETAIL</span>
</div>
<div id="admin-audit-table" role="list" aria-label="Audit events" aria-live="polite">
<div class="dashboard-empty">Loading audit log...</div>
</div>
</div>
</div>
</div>
@@ -206,6 +349,10 @@ window.TURNSTONE_KB_SHORTCUTS = [
<input id="new-ws-name" type="text" placeholder="Auto-generated if empty" autocomplete="off">
<label for="new-ws-model">Model <span class="label-hint">optional</span></label>
<input id="new-ws-model" type="text" placeholder="Default model" autocomplete="off">
<label for="new-ws-template">Template <span class="label-hint">optional</span></label>
<select id="new-ws-template">
<option value="">Use defaults</option>
</select>
<label for="new-ws-task">Task <span class="label-hint">optional &mdash; sent as first message</span></label>
<textarea id="new-ws-task" rows="3" placeholder="What should this workstream work on?"></textarea>
<div id="new-ws-buttons">
@@ -341,6 +488,8 @@ window.TURNSTONE_KB_SHORTCUTS = [
</div>
<label for="cs-model">Model <span class="label-hint">optional</span></label>
<input id="cs-model" type="text" placeholder="Default model" autocomplete="off">
<label for="cs-template">Template <span class="label-hint">optional</span></label>
<input id="cs-template" type="text" placeholder="Prompt template name" autocomplete="off">
<label for="cs-message">Initial message</label>
<textarea id="cs-message" rows="3" placeholder="What should the workstream do?"></textarea>
<label class="admin-checkbox"><input id="cs-autoapprove" type="checkbox"> Auto-approve tool calls</label>
@@ -387,6 +536,8 @@ window.TURNSTONE_KB_SHORTCUTS = [
</div>
<label for="es-model">Model</label>
<input id="es-model" type="text" autocomplete="off">
<label for="es-template">Template <span class="label-hint">optional</span></label>
<input id="es-template" type="text" autocomplete="off">
<label for="es-message">Initial message</label>
<textarea id="es-message" rows="3"></textarea>
<label class="admin-checkbox"><input id="es-autoapprove" type="checkbox"> Auto-approve tool calls</label>
@@ -409,7 +560,162 @@ window.TURNSTONE_KB_SHORTCUTS = [
</div>
</div>
<!-- Create Role Modal -->
<div id="create-role-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="create-role-title">
<div id="create-role-box" class="admin-modal admin-modal-wide">
<h2 id="create-role-title">Create Role</h2>
<div id="create-role-error" role="alert" aria-live="assertive"></div>
<label for="cr-name">Name</label>
<input id="cr-name" type="text" placeholder="e.g. security-reviewer" autocomplete="off" spellcheck="false">
<label for="cr-displayname">Display name</label>
<input id="cr-displayname" type="text" placeholder="Security Reviewer" autocomplete="off">
<fieldset class="perm-fieldset"><legend>Permissions</legend>
<div id="cr-perms-container" role="group" aria-label="Permissions"></div>
</fieldset>
<div class="modal-buttons">
<button class="modal-cancel" onclick="hideCreateRoleModal()">Cancel</button>
<button id="cr-submit" class="modal-submit" onclick="submitCreateRole()">Create</button>
</div>
</div>
</div>
<!-- Edit Role Modal -->
<div id="edit-role-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="edit-role-title">
<div id="edit-role-box" class="admin-modal admin-modal-wide">
<h2 id="edit-role-title">Edit Role</h2>
<div id="edit-role-error" role="alert" aria-live="assertive"></div>
<input id="er-id" type="hidden">
<label for="er-name">Display name</label>
<input id="er-name" type="text" autocomplete="off">
<fieldset class="perm-fieldset"><legend>Permissions</legend>
<div id="er-perms-container" role="group" aria-label="Permissions"></div>
</fieldset>
<div class="modal-buttons">
<button class="modal-cancel" onclick="hideEditRoleModal()">Cancel</button>
<button id="er-submit" class="modal-submit" onclick="submitEditRole()">Save</button>
</div>
</div>
</div>
<!-- User Roles Modal -->
<div id="user-roles-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="user-roles-title">
<div id="user-roles-box" class="admin-modal">
<h2 id="user-roles-title">Assign Roles</h2>
<div id="user-roles-error" role="alert" aria-live="assertive"></div>
<input id="ur-user-id" type="hidden">
<div id="ur-roles-container"></div>
<div class="modal-buttons">
<button class="modal-cancel" onclick="hideUserRolesModal()">Cancel</button>
<button class="modal-submit" onclick="submitUserRoles()">Save</button>
</div>
</div>
</div>
<!-- Create Policy Modal -->
<div id="create-policy-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="create-policy-title">
<div id="create-policy-box" class="admin-modal">
<h2 id="create-policy-title">Create Tool Policy</h2>
<div id="create-policy-error" role="alert" aria-live="assertive"></div>
<label for="cp-name">Name</label>
<input id="cp-name" type="text" placeholder="e.g. Block shell access" autocomplete="off">
<label for="cp-pattern">Tool pattern <span class="label-hint">glob syntax: bash*, file_write, *</span></label>
<input id="cp-pattern" type="text" placeholder="bash*" autocomplete="off" spellcheck="false">
<label for="cp-action">Action</label>
<select id="cp-action">
<option value="ask">Ask (require approval)</option>
<option value="allow">Allow (auto-approve)</option>
<option value="deny">Deny (block)</option>
</select>
<label for="cp-priority">Priority <span class="label-hint">higher = evaluated first</span></label>
<input id="cp-priority" type="number" value="0" min="0" max="9999">
<div class="modal-buttons">
<button class="modal-cancel" onclick="hideCreatePolicyModal()">Cancel</button>
<button id="cp-submit" class="modal-submit" onclick="submitCreatePolicy()">Create</button>
</div>
</div>
</div>
<!-- Edit Policy Modal -->
<div id="edit-policy-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="edit-policy-title">
<div id="edit-policy-box" class="admin-modal">
<h2 id="edit-policy-title">Edit Tool Policy</h2>
<div id="edit-policy-error" role="alert" aria-live="assertive"></div>
<input id="ep-id" type="hidden">
<label for="ep-name">Name</label>
<input id="ep-name" type="text" autocomplete="off">
<label for="ep-pattern">Tool pattern</label>
<input id="ep-pattern" type="text" autocomplete="off" spellcheck="false">
<label for="ep-action">Action</label>
<select id="ep-action">
<option value="ask">Ask (require approval)</option>
<option value="allow">Allow (auto-approve)</option>
<option value="deny">Deny (block)</option>
</select>
<label for="ep-priority">Priority</label>
<input id="ep-priority" type="number" value="0" min="0" max="9999">
<label class="admin-checkbox"><input id="ep-enabled" type="checkbox" checked> Enabled</label>
<div class="modal-buttons">
<button class="modal-cancel" onclick="hideEditPolicyModal()">Cancel</button>
<button id="ep-submit" class="modal-submit" onclick="submitEditPolicy()">Save</button>
</div>
</div>
</div>
<!-- Create Template Modal -->
<div id="create-template-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="create-template-title">
<div id="create-template-box" class="admin-modal admin-modal-wide">
<h2 id="create-template-title">Create Prompt Template</h2>
<div id="create-template-error" role="alert" aria-live="assertive"></div>
<label for="ctm-name">Name</label>
<input id="ctm-name" type="text" placeholder="e.g. Code Review Agent" autocomplete="off">
<label for="ctm-category">Category</label>
<select id="ctm-category">
<option value="general">General</option>
<option value="engineering">Engineering</option>
<option value="support">Support</option>
<option value="custom">Custom</option>
</select>
<label for="ctm-content">Content <span class="label-hint">system message text, use {{model}}, {{ws_id}}, {{node_id}} for placeholders</span></label>
<textarea id="ctm-content" rows="6" placeholder="You are a code reviewer using {{model}}..."></textarea>
<label>Variables <span class="label-hint">auto-detected from content &mdash; available: model, ws_id, node_id</span></label>
<div id="ctm-variables" class="label-hint" style="padding:4px 0;min-height:1.2em"></div>
<label class="admin-checkbox"><input id="ctm-default" type="checkbox"> Set as default for new workstreams</label>
<div class="modal-buttons">
<button class="modal-cancel" onclick="hideCreateTemplateModal()">Cancel</button>
<button id="ctm-submit" class="modal-submit" onclick="submitCreateTemplate()">Create</button>
</div>
</div>
</div>
<!-- Edit Template Modal -->
<div id="edit-template-overlay" style="display:none" role="dialog" aria-modal="true" aria-labelledby="edit-template-title">
<div id="edit-template-box" class="admin-modal admin-modal-wide">
<h2 id="edit-template-title">Edit Prompt Template</h2>
<div id="edit-template-error" role="alert" aria-live="assertive"></div>
<input id="etm-id" type="hidden">
<label for="etm-name">Name</label>
<input id="etm-name" type="text" autocomplete="off">
<label for="etm-category">Category</label>
<select id="etm-category">
<option value="general">General</option>
<option value="engineering">Engineering</option>
<option value="support">Support</option>
<option value="custom">Custom</option>
</select>
<label for="etm-content">Content</label>
<textarea id="etm-content" rows="6"></textarea>
<label>Variables <span class="label-hint">auto-detected from content &mdash; available: model, ws_id, node_id</span></label>
<div id="etm-variables" class="label-hint" style="padding:4px 0;min-height:1.2em"></div>
<label class="admin-checkbox"><input id="etm-default" type="checkbox"> Set as default</label>
<div class="modal-buttons">
<button class="modal-cancel" onclick="hideEditTemplateModal()">Cancel</button>
<button id="etm-submit" class="modal-submit" onclick="submitEditTemplate()">Save</button>
</div>
</div>
</div>
<script src="/static/admin.js"></script>
<script src="/static/governance.js"></script>
<script src="/static/app.js"></script>
</body>
</html>
+302 -1
View File
@@ -166,6 +166,17 @@
opacity: 0.7;
}
/* MCP indicator dot — LED effect with magenta glow */
.csb-mcp-dot {
width: 7px;
height: 7px;
border-radius: 50%;
background: var(--magenta);
box-shadow: 0 0 4px var(--magenta-glow);
display: inline-block;
flex-shrink: 0;
}
.csb-loading { color: var(--fg-dim); font-size: 11px; font-style: italic; opacity: 0.8; }
#cluster-status-bar.stale { border-top-color: var(--yellow); }
@@ -454,6 +465,16 @@
}
.dash-cell-node:hover { text-decoration: underline; color: var(--fg-bright); }
/* ==========================================================================
MCP summary in node detail
========================================================================== */
#node-mcp-summary {
color: var(--magenta);
font-size: 11px;
font-family: var(--font-mono);
margin-left: 12px;
}
/* ==========================================================================
Node link
========================================================================== */
@@ -665,6 +686,7 @@
.node-group-header .node-group-cell:last-child { display: none; }
.ncol-version, .node-cell-version { display: none; }
.ncol-health, .node-cell-health { display: none; }
#node-mcp-summary { display: none; }
#main { padding: 16px; padding-bottom: 60px; }
}
@media (max-width: 480px) {
@@ -709,6 +731,11 @@
color: var(--accent);
border-bottom-color: var(--accent);
}
.admin-tab:focus-visible {
outline: 2px solid var(--accent);
outline-offset: -2px;
border-radius: var(--radius-sm);
}
.admin-toolbar {
display: flex;
@@ -863,6 +890,16 @@
.sched-disabled { color: var(--fg-dim); }
.sched-expired { color: var(--accent); }
/* Watches grid: NAME | NODE | COMMAND | INTERVAL | POLL | CONDITION | STATUS | ACTIONS */
#admin-watches .admin-colheaders,
#admin-watches .admin-row {
grid-template-columns: 1.2fr 80px 1.5fr 60px 70px 1fr 70px 70px;
}
/* Watch status indicators */
.watch-active { color: var(--green); font-weight: 500; }
.watch-completed { color: var(--accent); }
/* Wide modal variant for schedule forms */
.admin-modal-wide { width: 480px; }
@@ -979,7 +1016,10 @@
.modal-submit:disabled { opacity: 0.4; cursor: not-allowed; filter: none; }
#create-user-overlay, #create-token-overlay, #token-created-overlay, #create-channel-overlay, #confirm-overlay,
#create-schedule-overlay, #edit-schedule-overlay, #schedule-runs-overlay {
#create-schedule-overlay, #edit-schedule-overlay, #schedule-runs-overlay,
#create-role-overlay, #edit-role-overlay, #user-roles-overlay,
#create-policy-overlay, #edit-policy-overlay,
#create-template-overlay, #edit-template-overlay {
position: fixed;
inset: 0;
background: rgba(0, 0, 0, 0.7);
@@ -1027,6 +1067,267 @@
grid-template-columns: 1fr 60px 80px 130px;
}
.admin-col-sschedule, .admin-col-starget, .admin-col-snext { display: none; }
#admin-watches .admin-colheaders, #admin-watches .admin-row {
grid-template-columns: 1.2fr 80px 70px 70px 70px;
}
.admin-col-wcmd, .admin-col-wcond, .admin-col-winterval { display: none; }
}
/* ==========================================================================
Admin tabs horizontal scroll for 10+ tabs
========================================================================== */
.admin-tabs {
overflow-x: auto;
-webkit-overflow-scrolling: touch;
flex-wrap: nowrap;
scrollbar-width: thin;
}
/* ==========================================================================
Governance: Roles grid
========================================================================== */
#admin-roles .admin-colheaders,
#admin-roles .admin-row {
grid-template-columns: 160px 1fr 110px;
}
/* ==========================================================================
Governance: Tool Policies grid
========================================================================== */
#admin-policies .admin-colheaders,
#admin-policies .admin-row {
grid-template-columns: 1.2fr 1fr 70px 50px 80px 140px;
}
/* Policy action badges */
.policy-badge {
display: inline-block;
font-family: var(--font-display);
font-size: 9px;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
padding: 2px 8px;
border-radius: 2px;
}
.policy-allow {
color: var(--green);
background: var(--green-glow);
border: 1px solid var(--green-glow);
}
.policy-deny {
color: var(--red);
background: var(--red-glow);
border: 1px solid var(--red-glow);
}
.policy-ask {
color: var(--yellow);
background: var(--yellow-glow);
border: 1px solid var(--yellow-glow);
}
/* ==========================================================================
Governance: Prompt Templates grid
========================================================================== */
#admin-templates .admin-colheaders,
#admin-templates .admin-row {
grid-template-columns: 1.5fr 100px 1fr 140px;
}
/* ==========================================================================
Governance: Audit grid
========================================================================== */
#admin-audit .admin-colheaders,
#admin-audit .admin-row {
grid-template-columns: 80px 80px 1fr 120px 1.5fr;
}
/* Audit action badges */
.audit-badge {
display: inline-block;
font-family: var(--font-display);
font-size: 9px;
font-weight: 600;
text-transform: none;
letter-spacing: 0.02em;
padding: 1px 6px;
border-radius: 2px;
background: var(--bg-highlight);
color: var(--fg-dim);
border: 1px solid var(--border);
}
.audit-danger { color: var(--red); border-color: var(--red-glow); }
.audit-success { color: var(--green); border-color: var(--green-glow); }
/* ==========================================================================
Governance: Usage dashboard
========================================================================== */
.usage-summary {
display: flex;
gap: 24px;
padding: 16px 0 20px;
flex-wrap: wrap;
}
.usage-readout {
display: flex;
flex-direction: column;
gap: 2px;
}
.usage-readout-value {
font-size: 22px;
font-weight: 600;
color: var(--fg-bright);
font-variant-numeric: tabular-nums;
font-family: var(--font-mono);
letter-spacing: -0.02em;
}
.usage-readout-label {
font-size: 10px;
font-family: var(--font-display);
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.08em;
color: var(--fg-dim);
}
/* Usage bar chart */
.usage-chart { padding-top: 4px; }
.usage-bar-row {
display: grid;
grid-template-columns: 90px 1fr 60px;
align-items: center;
gap: 10px;
padding: 4px 0;
}
.usage-bar-label {
font-size: 11px;
color: var(--fg-dim);
text-align: right;
font-variant-numeric: tabular-nums;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.usage-bar-track {
height: 16px;
background: var(--bg-highlight);
border-radius: 2px;
overflow: hidden;
}
.usage-bar-fill {
height: 100%;
background: var(--accent);
border-radius: 2px;
min-width: 2px;
transition: width 0.3s ease;
box-shadow: 0 0 6px var(--accent-glow);
}
.usage-bar-value {
font-size: 11px;
color: var(--fg-dim);
font-variant-numeric: tabular-nums;
text-align: right;
}
/* Usage range/group buttons */
.usage-range-group {
display: flex;
gap: 2px;
border: 1px solid var(--border-strong);
border-radius: var(--radius-sm);
overflow: hidden;
}
.usage-range-btn, .usage-group-btn {
background: var(--bg);
color: var(--fg-dim);
border: none;
font-family: var(--font-display);
font-size: 10px;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
padding: 5px 10px;
cursor: pointer;
transition: background 0.15s, color 0.15s;
}
.usage-range-btn:hover, .usage-group-btn:hover {
background: var(--bg-highlight);
color: var(--fg);
}
.usage-range-btn.active, .usage-group-btn.active {
background: var(--accent-dim);
color: var(--accent);
}
.usage-range-btn:focus-visible, .usage-group-btn:focus-visible {
outline: 2px solid var(--accent);
outline-offset: -2px;
}
/* ==========================================================================
Governance: Permission grid (modal checkboxes)
========================================================================== */
.perm-fieldset {
border: none;
padding: 0;
margin: 12px 0 0;
}
.perm-fieldset legend {
font-family: var(--font-display);
font-size: 10px;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.08em;
color: var(--fg-dim);
padding: 0;
margin-bottom: 4px;
}
.perm-grid {
display: grid;
grid-template-columns: 1fr 1fr;
gap: 4px 16px;
padding: 8px 0;
}
.perm-checkbox {
display: flex;
align-items: center;
gap: 6px;
font-size: 11px;
font-family: var(--font-mono);
color: var(--fg);
padding: 3px 0;
cursor: pointer;
text-transform: none;
letter-spacing: normal;
}
.perm-checkbox input[type="checkbox"] {
width: auto;
margin: 0;
accent-color: var(--accent);
}
/* ==========================================================================
Governance: Responsive
========================================================================== */
@media (max-width: 700px) {
#admin-roles .admin-colheaders, #admin-roles .admin-row {
grid-template-columns: 1fr 100px;
}
.admin-col-rperms { display: none; }
#admin-policies .admin-colheaders, #admin-policies .admin-row {
grid-template-columns: 1fr 70px 50px 100px;
}
.admin-col-pstatus, .admin-col-ppriority { display: none; }
#admin-templates .admin-colheaders, #admin-templates .admin-row {
grid-template-columns: 1fr 100px;
}
.admin-col-tmcat, .admin-col-tmvars { display: none; }
#admin-audit .admin-colheaders, #admin-audit .admin-row {
grid-template-columns: 60px 1fr 100px;
}
.admin-col-auser, .admin-col-adetail { display: none; }
.usage-readout-value { font-size: 18px; }
.usage-bar-row { grid-template-columns: 70px 1fr 50px; }
.perm-grid { grid-template-columns: 1fr; }
}
/* ==========================================================================
+41
View File
@@ -0,0 +1,41 @@
"""Audit event recording helper.
Provides a fire-and-forget ``record_audit`` function that admin handlers
call after mutations to create a persistent audit trail.
"""
from __future__ import annotations
import json
import logging
import uuid
from typing import TYPE_CHECKING, Any
if TYPE_CHECKING:
from turnstone.core.storage._protocol import StorageBackend
log = logging.getLogger(__name__)
def record_audit(
storage: StorageBackend,
user_id: str,
action: str,
resource_type: str = "",
resource_id: str = "",
detail: dict[str, Any] | None = None,
ip_address: str = "",
) -> None:
"""Record an audit event. Silently logs on failure (never raises)."""
try:
storage.record_audit_event(
event_id=uuid.uuid4().hex,
user_id=user_id,
action=action,
resource_type=resource_type,
resource_id=resource_id,
detail=json.dumps(detail) if detail else "{}",
ip_address=ip_address,
)
except Exception:
log.warning("Failed to record audit event: %s %s", action, resource_id, exc_info=True)
+117 -4
View File
@@ -18,6 +18,7 @@ always accessible without authentication.
from __future__ import annotations
import contextlib
import hashlib
import hmac
import json
@@ -33,7 +34,7 @@ from typing import TYPE_CHECKING, Any
if TYPE_CHECKING:
from starlette.requests import Request
from starlette.responses import Response
from starlette.responses import JSONResponse, Response
from starlette.types import ASGIApp, Receive, Scope, Send
log = logging.getLogger(__name__)
@@ -80,6 +81,60 @@ _ROLE_TO_SCOPES: dict[str, frozenset[str]] = {
"full": frozenset({"read", "write", "approve"}),
}
# ---------------------------------------------------------------------------
# RBAC helpers
# ---------------------------------------------------------------------------
def _load_user_permissions(storage: Any, user_id: str) -> set[str]:
"""Load the union of all permissions from a user's assigned roles."""
try:
result: set[str] = storage.get_user_permissions(user_id)
return result
except Exception:
log.warning("Failed to load permissions for user %s", user_id)
return set()
def _permissions_to_scopes(permissions: set[str]) -> frozenset[str]:
"""Derive legacy scopes from a granular permission set."""
scopes: set[str] = set()
if not permissions:
scopes.add("read")
return frozenset(scopes)
for perm in permissions:
if perm in VALID_SCOPES:
scopes.update(SCOPE_HIERARCHY.get(perm, {perm}))
# Any admin.* permission requires access to admin endpoints → approve scope
if any(p.startswith("admin.") for p in permissions):
scopes.update(SCOPE_HIERARCHY["approve"])
if not scopes:
scopes.add("read")
return frozenset(scopes)
def require_permission(request: Request, permission: str) -> JSONResponse | None:
"""Return a 403 JSONResponse if the user lacks *permission*, else None.
Call from admin handlers after the middleware scope check passes.
Config-file tokens (no user_id) are treated as full-access.
"""
from starlette.responses import JSONResponse
auth_result: AuthResult | None = getattr(getattr(request, "state", None), "auth_result", None)
if auth_result is None:
return JSONResponse({"error": "Unauthorized"}, status_code=401)
# Config-file tokens (no user_id) are treated as full-access
if not auth_result.user_id:
return None
if auth_result.has_permission(permission):
return None
return JSONResponse(
{"error": f"Forbidden: missing '{permission}' permission"},
status_code=403,
)
# ---------------------------------------------------------------------------
# Path classification
# ---------------------------------------------------------------------------
@@ -104,6 +159,7 @@ WRITE_PATHS: frozenset[str] = frozenset(
"/api/send",
"/api/plan",
"/api/command",
"/api/cancel",
"/api/workstreams/new",
"/api/workstreams/close",
"/api/cluster/workstreams/new",
@@ -133,11 +189,16 @@ class AuthResult:
user_id: str # empty string for config-file tokens
scopes: frozenset[str]
token_source: str # "config", "jwt", "database"
permissions: frozenset[str] = frozenset()
def has_scope(self, scope: str) -> bool:
"""Return True if this result includes *scope*."""
return scope in self.scopes
def has_permission(self, permission: str) -> bool:
"""Return True if this result includes *permission*."""
return permission in self.permissions
# ---------------------------------------------------------------------------
# AuthConfig (unchanged from before — static config-file tokens)
@@ -249,8 +310,9 @@ def create_jwt(
secret: str,
expiry_hours: int = 24,
audience: str = "",
permissions: frozenset[str] = frozenset(),
) -> str:
"""Create a signed JWT with user identity and scopes."""
"""Create a signed JWT with user identity, scopes, and permissions."""
import jwt
now = int(time.time())
@@ -264,6 +326,8 @@ def create_jwt(
}
if audience:
payload["aud"] = audience
if permissions:
payload["permissions"] = ",".join(sorted(permissions))
return jwt.encode(payload, secret, algorithm="HS256")
@@ -293,11 +357,15 @@ def validate_jwt(token: str, secret: str, audience: str = "") -> AuthResult | No
user_id = payload.get("sub", "")
scopes_str = payload.get("scopes", "")
source = payload.get("src", "jwt")
perms_str = payload.get("permissions", "")
perms = frozenset(p for p in perms_str.split(",") if p) if perms_str else frozenset()
return AuthResult(
user_id=user_id,
scopes=parse_scopes(scopes_str),
token_source=source,
permissions=perms,
)
@@ -390,6 +458,13 @@ def required_scope(method: str, path: str) -> str:
# Write endpoints
if method == "POST" and normalized in WRITE_PATHS:
return "write"
# Watch cancel has a path parameter: /api/watches/{id}/cancel
if (
method == "POST"
and normalized.startswith("/api/watches/")
and normalized.endswith("/cancel")
):
return "write"
# Console proxy routes: /node/{node_id}/api/{tail} or /node/{node_id}/v1/api/{tail}
if method == "POST" and normalized.startswith("/node/"):
@@ -524,10 +599,12 @@ def _authenticate_api_token(token: str, storage: Any) -> AuthResult | None:
if exp_dt < now:
return None
perms = _load_user_permissions(storage, row["user_id"]) if storage else set()
return AuthResult(
user_id=row["user_id"],
scopes=parse_scopes(row["scopes"]),
token_source="database",
permissions=frozenset(perms),
)
@@ -828,10 +905,14 @@ async def handle_auth_login(request: Request, audience: str) -> Response:
if username and password and storage is not None:
user = storage.get_user_by_username(username)
if user and verify_password(password, user["password_hash"]):
# Derive scopes and permissions from assigned roles
perms = _load_user_permissions(storage, user["user_id"])
scopes = _permissions_to_scopes(perms)
result = AuthResult(
user_id=user["user_id"],
scopes=frozenset({"read", "write", "approve"}),
scopes=scopes,
token_source="password",
permissions=frozenset(perms),
)
elif body.get("token"):
result = _authenticate_token(
@@ -858,11 +939,14 @@ async def handle_auth_login(request: Request, audience: str) -> Response:
source=result.token_source,
secret=jwt_secret,
audience=audience,
permissions=result.permissions,
)
role = "full" if result.has_scope("write") else "read"
scopes_str = ",".join(sorted(result.scopes))
resp_body: dict[str, str] = {"status": "ok", "role": role, "scopes": scopes_str}
if result.permissions:
resp_body["permissions"] = ",".join(sorted(result.permissions))
if jwt_token:
resp_body["jwt"] = jwt_token
if result.user_id:
@@ -952,7 +1036,33 @@ async def handle_auth_setup(request: Request, audience: str) -> Response:
if not created:
return JSONResponse({"error": "Setup already completed"}, status_code=409)
scopes = frozenset({"read", "write", "approve"})
# Assign admin role to the first user — fail setup if this breaks,
# otherwise the admin is created with read-only access and locked out.
try:
storage.assign_role(user_id, "builtin-admin", "")
except Exception:
log.error("Failed to assign admin role to first user %s — aborting setup", user_id)
# Roll back the user creation so setup can be retried
with contextlib.suppress(Exception):
storage.delete_user(user_id)
return JSONResponse(
{"error": "Failed to assign admin role. Ensure migrations have run."},
status_code=503,
)
# Derive permissions from roles
perms = _load_user_permissions(storage, user_id)
if not perms:
log.error(
"First user %s has no permissions after role assignment — aborting setup", user_id
)
with contextlib.suppress(Exception):
storage.delete_user(user_id)
return JSONResponse(
{"error": "Failed to load permissions. Ensure migrations have run."},
status_code=503,
)
scopes = _permissions_to_scopes(perms)
jwt_token = ""
if jwt_secret:
jwt_token = create_jwt(
@@ -961,6 +1071,7 @@ async def handle_auth_setup(request: Request, audience: str) -> Response:
source="password",
secret=jwt_secret,
audience=audience,
permissions=frozenset(perms),
)
resp_body: dict[str, str] = {
@@ -970,6 +1081,8 @@ async def handle_auth_setup(request: Request, audience: str) -> Response:
"role": "full",
"scopes": ",".join(sorted(scopes)),
}
if perms:
resp_body["permissions"] = ",".join(sorted(perms))
if jwt_token:
resp_body["jwt"] = jwt_token
+611 -28
View File
@@ -1,16 +1,16 @@
"""MCP (Model Context Protocol) client manager.
Connects to external MCP tool servers and exposes their tools alongside
turnstone's built-in tools.
Connects to external MCP tool servers and exposes their tools, resources,
and prompts alongside turnstone's built-in capabilities.
Architecture: the MCP SDK is fully async, but turnstone's ChatSession is
synchronous. We bridge the two by running a dedicated asyncio event loop
in a daemon thread. ``call_tool_sync`` dispatches coroutines onto that loop
via ``asyncio.run_coroutine_threadsafe``.
Tool refresh: three mechanisms keep tool lists up-to-date without restart:
1. Push notifications servers declaring ``tools.listChanged`` trigger
immediate refresh via ``ToolListChangedNotification``.
Refresh: three mechanisms keep tool/resource/prompt lists up-to-date:
1. Push notifications servers declaring ``listChanged`` on the
respective capability trigger immediate refresh.
2. Periodic timer servers *without* push support are polled on a
staggered interval (configurable, default 4 h, seeded at launch).
3. Manual ``/mcp refresh [server]`` triggers ``refresh_sync()``.
@@ -19,6 +19,7 @@ Tool refresh: three mechanisms keep tool lists up-to-date without restart:
from __future__ import annotations
import asyncio
import concurrent.futures
import contextlib
import json
import logging
@@ -26,6 +27,7 @@ import os
import random
import threading
import time
import uuid
from contextlib import AsyncExitStack
from pathlib import Path
from typing import TYPE_CHECKING, Any
@@ -112,6 +114,31 @@ class MCPClientManager:
self._listeners: list[Callable[[], None]] = []
self._listeners_lock = threading.Lock()
# Resources — parallel to tools
self._per_server_resources: dict[str, list[dict[str, Any]]] = {}
self._resources: list[dict[str, Any]] = []
self._resource_map: dict[str, tuple[str, str]] = {} # uri → (server, uri)
self._supports_resources: dict[str, bool] = {} # server has resources capability
self._supports_resource_list_changed: dict[str, bool] = {}
self._resource_listeners: list[Callable[[], None]] = []
self._resource_listeners_lock = threading.Lock()
# Prompts — parallel to tools
self._per_server_prompts: dict[str, list[dict[str, Any]]] = {}
self._prompts: list[dict[str, Any]] = []
self._prompt_map: dict[str, tuple[str, str]] = {} # prefixed → (server, original)
self._supports_prompts: dict[str, bool] = {} # server has prompts capability
self._supports_prompt_list_changed: dict[str, bool] = {}
self._prompt_listeners: list[Callable[[], None]] = []
self._prompt_listeners_lock = threading.Lock()
# Template prefix → (server_name, full_template_uri) for URI expansion
self._template_prefixes: dict[str, tuple[str, str]] = {}
# Governance storage (optional — set via set_storage())
self._storage: Any = None
self._sync_lock = threading.Lock()
# Periodic refresh for servers without push notifications
self._refresh_interval = refresh_interval
self._refresh_task: asyncio.Task[None] | None = None
@@ -146,7 +173,16 @@ class MCPClientManager:
# Start periodic refresh for servers without push notifications
needs_periodic = any(
not self._supports_list_changed.get(name, False) for name in self._sessions
not self._supports_list_changed.get(name, False)
or (
self._supports_resources.get(name, False)
and not self._supports_resource_list_changed.get(name, False)
)
or (
self._supports_prompts.get(name, False)
and not self._supports_prompt_list_changed.get(name, False)
)
for name in self._sessions
)
if needs_periodic and self._refresh_interval > 0:
self._refresh_task = asyncio.get_running_loop().create_task(self._periodic_refresh())
@@ -178,20 +214,26 @@ class MCPClientManager:
)
read, write = await self._exit_stack.enter_async_context(stdio_client(params))
# Register notification handler — lightweight; only acts on
# ToolListChangedNotification, which is a no-op if the server
# never sends it.
# Register notification handler — dispatches tool, resource, and
# prompt list-change notifications to the appropriate refresh method.
async def _on_notification(
msg: Any, # RequestResponder | ServerNotification | Exception
) -> None:
if isinstance(msg, mcp_types.ServerNotification) and isinstance(
msg.root, mcp_types.ToolListChangedNotification
):
log.info("Received tools/list_changed from '%s'", name)
try:
await self._refresh_server(name)
except Exception:
log.warning("Refresh after notification failed for '%s'", name, exc_info=True)
if not isinstance(msg, mcp_types.ServerNotification):
return
root = msg.root
try:
if isinstance(root, mcp_types.ToolListChangedNotification):
log.info("Received tools/list_changed from '%s'", name)
await self._refresh_server_tools(name)
elif isinstance(root, mcp_types.ResourceListChangedNotification):
log.info("Received resources/list_changed from '%s'", name)
await self._refresh_server_resources(name)
elif isinstance(root, mcp_types.PromptListChangedNotification):
log.info("Received prompts/list_changed from '%s'", name)
await self._refresh_server_prompts(name)
except Exception:
log.warning("Refresh after notification failed for '%s'", name, exc_info=True)
session = await self._exit_stack.enter_async_context(
ClientSession(read, write, message_handler=_on_notification) # type: ignore[arg-type]
@@ -199,11 +241,22 @@ class MCPClientManager:
await session.initialize()
self._sessions[name] = session
# Check push notification support
# Check push notification support for each capability
caps = session.get_server_capabilities()
tools_cap = getattr(caps, "tools", None) if caps else None
self._supports_list_changed[name] = bool(getattr(tools_cap, "listChanged", False))
resources_cap = getattr(caps, "resources", None) if caps else None
self._supports_resources[name] = resources_cap is not None
self._supports_resource_list_changed[name] = bool(
getattr(resources_cap, "listChanged", False)
)
prompts_cap = getattr(caps, "prompts", None) if caps else None
self._supports_prompts[name] = prompts_cap is not None
self._supports_prompt_list_changed[name] = bool(getattr(prompts_cap, "listChanged", False))
# Discover tools
result = await session.list_tools()
server_tools: list[dict[str, Any]] = []
@@ -213,14 +266,88 @@ class MCPClientManager:
self._per_server_tools[name] = server_tools
self._rebuild_tools()
push_status = " (push)" if self._supports_list_changed[name] else ""
# Discover resources
resource_count = 0
if resources_cap is not None:
server_resources: list[dict[str, Any]] = []
res_result = await session.list_resources()
for r in res_result.resources:
server_resources.append(
{
"uri": str(r.uri),
"name": r.name or "",
"description": r.description or "",
"mimeType": r.mimeType or "",
"server": name,
}
)
# Also include resource templates (catalog-only — not directly
# readable via read_resource since they contain URI placeholders)
tmpl_result = await session.list_resource_templates()
for t in tmpl_result.resourceTemplates:
server_resources.append(
{
"uri": str(t.uriTemplate),
"name": t.name or "",
"description": t.description or "",
"mimeType": t.mimeType or "",
"server": name,
"template": True,
}
)
resource_count = len(server_resources)
self._per_server_resources[name] = server_resources
self._rebuild_resources()
# Discover prompts
prompt_count = 0
if prompts_cap is not None:
server_prompts: list[dict[str, Any]] = []
prompt_result = await session.list_prompts()
for p in prompt_result.prompts:
server_prompts.append(
{
"name": f"mcp__{name}__{p.name}",
"original_name": p.name,
"server": name,
"description": p.description or "",
"arguments": [
{
"name": a.name,
"description": a.description or "",
"required": a.required or False,
}
for a in (p.arguments or [])
],
}
)
prompt_count = len(server_prompts)
self._per_server_prompts[name] = server_prompts
self._rebuild_prompts()
push_parts: list[str] = []
if self._supports_list_changed[name]:
push_parts.append("tools")
if self._supports_resource_list_changed[name]:
push_parts.append("resources")
if self._supports_prompt_list_changed[name]:
push_parts.append("prompts")
push_status = f" (push: {','.join(push_parts)})" if push_parts else ""
log.info(
"Connected MCP server '%s'%d tool(s)%s",
"Connected MCP server '%s'%d tool(s), %d resource(s), %d prompt(s)%s",
name,
len(result.tools),
resource_count,
prompt_count,
push_status,
)
# Sync discovered prompts into governance storage
try:
self.sync_prompts_to_storage()
except Exception:
log.warning("Prompt sync after connect failed for '%s'", name, exc_info=True)
# -- tool refresh --------------------------------------------------------
def _rebuild_tools(self) -> None:
@@ -242,7 +369,7 @@ class MCPClientManager:
self._tool_map = new_map
self._notify_listeners()
async def _refresh_server(self, name: str) -> tuple[list[str], list[str]]:
async def _refresh_server_tools(self, name: str) -> tuple[list[str], list[str]]:
"""Re-fetch tools for one server. Returns ``(added, removed)`` names."""
session = self._sessions.get(name)
if session is None:
@@ -268,10 +395,21 @@ class MCPClientManager:
)
return added, removed
async def _refresh_server(self, name: str) -> tuple[list[str], list[str]]:
"""Re-fetch tools, resources, and prompts for one server.
Returns ``(added_tools, removed_tools)`` names (tool diff only,
for backward compatibility with ``/mcp refresh`` output).
"""
added, removed = await self._refresh_server_tools(name)
await self._refresh_server_resources(name)
await self._refresh_server_prompts(name)
return added, removed
async def _refresh_all(
self, server_name: str | None = None
) -> dict[str, tuple[list[str], list[str]]]:
"""Refresh tools for one or all servers.
"""Refresh tools, resources, and prompts for one or all servers.
For disconnected servers (in config but not connected), attempts
reconnect. Returns ``{server: (added, removed)}`` per server.
@@ -297,6 +435,13 @@ class MCPClientManager:
except Exception:
log.warning("Refresh failed for MCP server '%s'", name, exc_info=True)
results[name] = ([], [])
# Final sync to clean up templates from servers that are no longer connected
try:
self.sync_prompts_to_storage()
except Exception:
log.warning("Prompt sync after refresh_all failed", exc_info=True)
return results
def refresh_sync(
@@ -319,16 +464,169 @@ class MCPClientManager:
await asyncio.sleep(initial_delay)
while True:
for name in list(self._server_configs):
if self._supports_list_changed.get(name, False):
continue # has push — skip
if name not in self._sessions:
continue # not connected — skip (reconnect on manual refresh)
try:
await self._refresh_server(name)
if not self._supports_list_changed.get(name, False):
await self._refresh_server_tools(name)
if not self._supports_resource_list_changed.get(name, False):
await self._refresh_server_resources(name)
if not self._supports_prompt_list_changed.get(name, False):
await self._refresh_server_prompts(name)
except Exception:
log.warning("Periodic refresh failed for '%s'", name, exc_info=True)
await asyncio.sleep(self._refresh_interval)
# -- resource refresh ----------------------------------------------------
def _rebuild_resources(self) -> None:
"""Rebuild merged ``_resources`` and ``_resource_map`` from per-server state.
Uses copy-on-write: builds new objects, then assigns atomically.
"""
new_resources: list[dict[str, Any]] = []
new_map: dict[str, tuple[str, str]] = {}
for srv_name, srv_resources in self._per_server_resources.items():
for res in srv_resources:
uri: str = res["uri"]
new_resources.append(res)
if res.get("template"):
continue # templates are catalog-only, not directly readable
if uri in new_map:
log.warning(
"Resource URI collision: '%s' from '%s' overrides '%s'",
uri,
srv_name,
new_map[uri][0],
)
new_map[uri] = (srv_name, uri)
# Build template prefix map for URI expansion fallback
new_prefixes: dict[str, tuple[str, str]] = {}
for srv_name, srv_resources in self._per_server_resources.items():
for res in srv_resources:
if res.get("template"):
tmpl_uri = res["uri"]
brace = tmpl_uri.find("{")
prefix = tmpl_uri[:brace] if brace >= 0 else tmpl_uri
if prefix:
if prefix in new_prefixes:
existing_srv, existing_tmpl = new_prefixes[prefix]
if len(tmpl_uri) > len(existing_tmpl):
log.warning(
"Template prefix collision: '%s' from '%s' overrides '%s'"
" (keeping more specific template)",
prefix,
srv_name,
existing_srv,
)
new_prefixes[prefix] = (srv_name, tmpl_uri)
else:
log.warning(
"Template prefix collision: '%s' from '%s' ignored in"
" favor of '%s' (keeping more specific template)",
prefix,
srv_name,
existing_srv,
)
else:
new_prefixes[prefix] = (srv_name, tmpl_uri)
self._resources = new_resources
self._resource_map = new_map
self._template_prefixes = new_prefixes
self._notify_resource_listeners()
async def _refresh_server_resources(self, name: str) -> None:
"""Re-fetch resources for one server."""
if not self._supports_resources.get(name, False):
return
session = self._sessions.get(name)
if session is None:
return
server_resources: list[dict[str, Any]] = []
res_result = await session.list_resources()
for r in res_result.resources:
server_resources.append(
{
"uri": str(r.uri),
"name": r.name or "",
"description": r.description or "",
"mimeType": r.mimeType or "",
"server": name,
}
)
tmpl_result = await session.list_resource_templates()
for t in tmpl_result.resourceTemplates:
server_resources.append(
{
"uri": str(t.uriTemplate),
"name": t.name or "",
"description": t.description or "",
"mimeType": t.mimeType or "",
"server": name,
"template": True,
}
)
self._per_server_resources[name] = server_resources
self._rebuild_resources()
# -- prompt refresh ------------------------------------------------------
def _rebuild_prompts(self) -> None:
"""Rebuild merged ``_prompts`` and ``_prompt_map`` from per-server state.
Uses copy-on-write: builds new objects, then assigns atomically.
"""
new_prompts: list[dict[str, Any]] = []
new_map: dict[str, tuple[str, str]] = {}
for srv_name, srv_prompts in self._per_server_prompts.items():
for prompt in srv_prompts:
prefixed: str = prompt["name"]
new_prompts.append(prompt)
new_map[prefixed] = (srv_name, prompt["original_name"])
self._prompts = new_prompts
self._prompt_map = new_map
self._notify_prompt_listeners()
async def _refresh_server_prompts(self, name: str) -> None:
"""Re-fetch prompts for one server."""
if not self._supports_prompts.get(name, False):
return
session = self._sessions.get(name)
if session is None:
return
server_prompts: list[dict[str, Any]] = []
prompt_result = await session.list_prompts()
for p in prompt_result.prompts:
server_prompts.append(
{
"name": f"mcp__{name}__{p.name}",
"original_name": p.name,
"server": name,
"description": p.description or "",
"arguments": [
{
"name": a.name,
"description": a.description or "",
"required": a.required or False,
}
for a in (p.arguments or [])
],
}
)
self._per_server_prompts[name] = server_prompts
self._rebuild_prompts()
# Sync discovered prompts into governance storage
try:
self.sync_prompts_to_storage()
except Exception:
log.warning("Prompt sync after refresh failed for '%s'", name, exc_info=True)
# -- listener infrastructure ---------------------------------------------
def add_listener(self, callback: Callable[[], None]) -> None:
@@ -342,7 +640,7 @@ class MCPClientManager:
self._listeners.remove(callback)
def _notify_listeners(self) -> None:
"""Invoke all registered listeners (runs on MCP background thread)."""
"""Invoke all registered tool-change listeners."""
with self._listeners_lock:
listeners = list(self._listeners)
for cb in listeners:
@@ -351,6 +649,156 @@ class MCPClientManager:
except Exception:
log.warning("Tool-change listener raised", exc_info=True)
def add_resource_listener(self, callback: Callable[[], None]) -> None:
"""Register a callback invoked when the resource list changes."""
with self._resource_listeners_lock:
self._resource_listeners.append(callback)
def remove_resource_listener(self, callback: Callable[[], None]) -> None:
"""Unregister a resource-change callback."""
with self._resource_listeners_lock, contextlib.suppress(ValueError):
self._resource_listeners.remove(callback)
def _notify_resource_listeners(self) -> None:
"""Invoke all registered resource-change listeners."""
with self._resource_listeners_lock:
listeners = list(self._resource_listeners)
for cb in listeners:
try:
cb()
except Exception:
log.warning("Resource-change listener raised", exc_info=True)
def add_prompt_listener(self, callback: Callable[[], None]) -> None:
"""Register a callback invoked when the prompt list changes."""
with self._prompt_listeners_lock:
self._prompt_listeners.append(callback)
def remove_prompt_listener(self, callback: Callable[[], None]) -> None:
"""Unregister a prompt-change callback."""
with self._prompt_listeners_lock, contextlib.suppress(ValueError):
self._prompt_listeners.remove(callback)
def _notify_prompt_listeners(self) -> None:
"""Invoke all registered prompt-change listeners."""
with self._prompt_listeners_lock:
listeners = list(self._prompt_listeners)
for cb in listeners:
try:
cb()
except Exception:
log.warning("Prompt-change listener raised", exc_info=True)
# -- governance storage sync ---------------------------------------------
def set_storage(self, storage: Any) -> None:
"""Inject governance storage backend for prompt template sync.
If MCP servers are already connected, triggers an immediate sync
so prompts discovered during startup appear in governance storage
(``start()`` completes before ``set_storage()`` is called).
"""
self._storage = storage
if self._connected.is_set():
try:
self.sync_prompts_to_storage()
except Exception:
log.warning("Prompt sync after set_storage failed", exc_info=True)
def sync_prompts_to_storage(self) -> dict[str, Any]:
"""Sync discovered MCP prompts into the prompt_templates governance table.
Returns ``{"added": [...], "removed": [...], "skipped": [...]}``.
Thread-safe: serialized via ``_sync_lock`` to prevent races
between ``set_storage()`` (main thread) and MCP background thread.
"""
if self._storage is None:
return {"added": [], "removed": [], "skipped": []}
with self._sync_lock:
return self._sync_prompts_locked()
def _sync_prompts_locked(self) -> dict[str, Any]:
"""Inner sync logic — must be called under ``_sync_lock``."""
storage = self._storage
added: list[str] = []
removed: list[str] = []
skipped: list[str] = []
# Current MCP prompt names (the prefixed names used as template names)
current_names: set[str] = set()
for prompt in list(self._prompts):
name: str = prompt["name"][:256]
server: str = prompt["server"][:128]
current_names.add(name)
# Build content from description + argument schema
desc = prompt.get("description", "")[:4096]
args_list = prompt.get("arguments", [])
content_parts = [desc] if desc else []
if args_list:
content_parts.append("\nArguments:")
for arg in args_list:
req = " (required)" if arg.get("required") else ""
arg_desc = arg.get("description", "")[:512]
content_parts.append(f" - {arg['name'][:128]}{req}: {arg_desc}")
content = "\n".join(content_parts) if content_parts else name
# Variables = JSON list of argument names
variables = json.dumps([a["name"] for a in args_list])
existing = storage.get_prompt_template_by_name(name)
if existing is not None:
if existing.get("origin") == "manual":
log.info(
"Skipping MCP prompt '%s' — manual template with same name exists", name
)
skipped.append(name)
continue
# Existing MCP template — update content/variables.
# Reset is_default to prevent a compromised MCP server from
# injecting content into a previously admin-promoted default.
storage.update_prompt_template(
existing["template_id"],
content=content,
variables=variables,
is_default=False,
)
else:
# Create new MCP-sourced template
template_id = str(uuid.uuid4())
storage.create_prompt_template(
template_id=template_id,
name=name,
category="mcp",
content=content,
variables=variables,
is_default=False,
org_id="",
created_by="",
origin="mcp",
mcp_server=server,
readonly=True,
)
added.append(name)
# Remove MCP templates whose prompts no longer exist
existing_mcp = storage.list_prompt_templates_by_origin("mcp")
for tpl in existing_mcp:
if tpl["name"] not in current_names:
storage.delete_prompt_template(tpl["template_id"])
removed.append(tpl["name"])
if added or removed:
log.info(
"MCP prompt sync: +%d added, -%d removed, %d skipped",
len(added),
len(removed),
len(skipped),
)
return {"added": added, "removed": removed, "skipped": skipped}
# -- lifecycle (shutdown) ------------------------------------------------
def shutdown(self) -> None:
@@ -371,18 +819,62 @@ class MCPClientManager:
if self._thread:
self._thread.join(timeout=5)
# Clear all state
self._sessions.clear()
self._tools = []
self._tool_map = {}
self._per_server_tools.clear()
self._supports_list_changed.clear()
self._resources = []
self._resource_map = {}
self._template_prefixes = {}
self._per_server_resources.clear()
self._supports_resources.clear()
self._supports_resource_list_changed.clear()
self._prompts = []
self._prompt_map = {}
self._per_server_prompts.clear()
self._supports_prompts.clear()
self._supports_prompt_list_changed.clear()
# Clear listener lists to release callback references
self._listeners.clear()
self._resource_listeners.clear()
self._prompt_listeners.clear()
log.info("MCP client shut down")
# -- query methods -------------------------------------------------------
def get_tools(self) -> list[dict[str, Any]]:
"""Return MCP tools in OpenAI function-calling format."""
return list(self._tools)
return [dict(t) for t in self._tools]
def get_resources(self) -> list[dict[str, Any]]:
"""Return discovered MCP resources (shallow-copied dicts)."""
return [dict(r) for r in self._resources]
def get_prompts(self) -> list[dict[str, Any]]:
"""Return discovered MCP prompts (shallow-copied dicts)."""
return [dict(p) for p in self._prompts]
@property
def resource_count(self) -> int:
"""Number of discovered resources (no allocation)."""
return len(self._resources)
@property
def prompt_count(self) -> int:
"""Number of discovered prompts (no allocation)."""
return len(self._prompts)
def is_mcp_tool(self, func_name: str) -> bool:
"""Check whether *func_name* belongs to an MCP server."""
return func_name in self._tool_map
def is_mcp_prompt(self, name: str) -> bool:
"""Check whether *name* is a known MCP prompt."""
return name in self._prompt_map
@property
def server_count(self) -> int:
return len(self._sessions)
@@ -417,7 +909,10 @@ class MCPClientManager:
future = asyncio.run_coroutine_threadsafe(
session.call_tool(original_name, arguments), self._loop
)
result = future.result(timeout=timeout)
try:
result = future.result(timeout=timeout)
except concurrent.futures.TimeoutError:
raise TimeoutError(f"MCP tool call timed out after {timeout}s") from None
# Extract text from the content array
texts: list[str] = []
@@ -435,6 +930,94 @@ class MCPClientManager:
output = f"Error: {output}"
return output
# -- resource read -------------------------------------------------------
def _match_template(self, uri: str) -> tuple[str, str] | None:
"""Find the longest matching template prefix for an expanded URI.
Returns ``(server_name, template_uri)`` or *None* if no match.
The match uses the longest static prefix stored in
``_template_prefixes`` (the portion of each template URI before
the first ``{``), with simple ``startswith`` matching.
"""
best: tuple[str, str] | None = None
best_len = 0
for prefix, mapping in self._template_prefixes.items():
if uri.startswith(prefix) and len(prefix) > best_len:
best = mapping
best_len = len(prefix)
return best
def read_resource_sync(self, uri: str, timeout: int = 120) -> str:
"""Read a resource by URI synchronously (blocks the calling thread).
Returns text content for ``TextResourceContents``, or base64 data
for ``BlobResourceContents``.
"""
mapping = self._resource_map.get(uri)
if mapping is None:
# Fall back to template prefix matching for expanded URIs
mapping = self._match_template(uri)
if mapping is None:
raise ValueError(f"Unknown MCP resource: {uri}")
server_name, _ = mapping
session = self._sessions.get(server_name)
if session is None:
raise RuntimeError(f"MCP server '{server_name}' is not connected")
assert self._loop is not None
future = asyncio.run_coroutine_threadsafe(session.read_resource(uri), self._loop)
try:
result = future.result(timeout=timeout)
except concurrent.futures.TimeoutError:
raise TimeoutError(f"MCP resource read timed out after {timeout}s") from None
parts: list[str] = []
for item in result.contents:
if hasattr(item, "text"):
parts.append(item.text)
elif hasattr(item, "blob"):
parts.append(item.blob)
else:
parts.append(str(item))
return "\n".join(parts) if parts else "(empty resource)"
# -- prompt invocation ---------------------------------------------------
def get_prompt_sync(
self,
prefixed_name: str,
arguments: dict[str, str] | None = None,
timeout: int = 30,
) -> list[dict[str, Any]]:
"""Invoke an MCP prompt synchronously and return expanded messages.
Returns a list of ``{role: str, content: str}`` dicts.
"""
mapping = self._prompt_map.get(prefixed_name)
if mapping is None:
raise ValueError(f"Unknown MCP prompt: {prefixed_name}")
server_name, original_name = mapping
session = self._sessions.get(server_name)
if session is None:
raise RuntimeError(f"MCP server '{server_name}' is not connected")
assert self._loop is not None
future = asyncio.run_coroutine_threadsafe(
session.get_prompt(original_name, arguments=arguments), self._loop
)
try:
result = future.result(timeout=timeout)
except concurrent.futures.TimeoutError:
raise TimeoutError(f"MCP prompt retrieval timed out after {timeout}s") from None
messages: list[dict[str, Any]] = []
for msg in result.messages:
content = msg.content
text = content.text if hasattr(content, "text") else str(content)
messages.append({"role": msg.role, "content": text})
return messages
# ---------------------------------------------------------------------------
# Config loading
+19
View File
@@ -143,6 +143,25 @@ def load_workstream_config(ws_id: str) -> dict[str, str]:
return {}
# -- Prompt templates ---------------------------------------------------------
def list_default_templates(org_id: str = "") -> list[dict[str, Any]]:
"""Return all templates where is_default=True, ordered by name."""
try:
return get_storage().list_default_templates(org_id)
except Exception:
return []
def get_prompt_template_by_name(name: str) -> dict[str, Any] | None:
"""Lookup prompt template by name."""
try:
return get_storage().get_prompt_template_by_name(name)
except Exception:
return None
# -- Workstream metadata ------------------------------------------------------
+19
View File
@@ -102,6 +102,7 @@ class MetricsCollector:
workstream_states: dict[str, int],
total_workstreams: int,
workstream_metrics: list[dict[str, Any]] | None = None,
mcp_info: dict[str, int] | None = None,
) -> str:
"""Return Prometheus text exposition format (v0.0.4)."""
lines: list[str] = []
@@ -322,6 +323,24 @@ class MetricsCollector:
f"turnstone_workstream_context_ratio{lstr} {_fmt_value(wm['context_ratio'])}"
)
# MCP gauges (optional)
if mcp_info:
gauge(
"turnstone_mcp_servers",
"Number of connected MCP servers",
mcp_info.get("servers", 0),
)
gauge(
"turnstone_mcp_resources",
"Number of MCP resources available",
mcp_info.get("resources", 0),
)
gauge(
"turnstone_mcp_prompts",
"Number of MCP prompts available",
mcp_info.get("prompts", 0),
)
lines.append("") # trailing newline
return "\n".join(lines)
+4
View File
@@ -33,6 +33,7 @@ class ModelConfig:
model: str
context_window: int = 131072
provider: str = "openai"
capabilities: dict[str, Any] = field(default_factory=dict)
# ---------------------------------------------------------------------------
@@ -185,6 +186,9 @@ def load_model_registry(
model=model_name,
context_window=entry.get("context_window", context_window),
provider=entry.get("provider", "openai"),
capabilities=entry.get("capabilities", {})
if isinstance(entry.get("capabilities"), dict)
else {},
)
# Ensure a "default" entry from CLI args
+80
View File
@@ -0,0 +1,80 @@
"""Tool policy evaluation engine.
Evaluates tool calls against admin-defined policies to determine whether
a tool should be auto-allowed, denied, or require human approval.
"""
from __future__ import annotations
import fnmatch
import logging
from typing import TYPE_CHECKING
if TYPE_CHECKING:
from turnstone.core.storage._protocol import StorageBackend
log = logging.getLogger(__name__)
def evaluate_tool_policy(
storage: StorageBackend,
tool_name: str,
org_id: str = "",
) -> str | None:
"""Check tool policies for *tool_name*.
Policies are evaluated in priority order (highest first). The first
matching policy wins.
Returns ``"allow"``, ``"deny"``, or ``"ask"`` if a policy matches,
or ``None`` if no policy matches (caller should fall through to the
default approval behaviour).
"""
try:
policies = storage.list_tool_policies(org_id=org_id)
except Exception:
log.warning("Failed to load tool policies", exc_info=True)
return None
for policy in policies:
if not policy.get("enabled", True):
continue
pattern = policy.get("tool_pattern", "")
if fnmatch.fnmatch(tool_name, pattern):
action: str = policy.get("action", "ask")
if action in ("allow", "deny", "ask"):
return action
log.warning("Unknown policy action %r for policy %s", action, policy.get("policy_id"))
return "ask"
return None
def evaluate_tool_policies_batch(
storage: StorageBackend,
tool_names: list[str],
org_id: str = "",
) -> dict[str, str | None]:
"""Evaluate policies for multiple tools at once (single DB query).
Returns a dict mapping each tool name to its policy result.
"""
try:
policies = storage.list_tool_policies(org_id=org_id)
except Exception:
log.warning("Failed to load tool policies", exc_info=True)
return {name: None for name in tool_names}
results: dict[str, str | None] = {}
for name in tool_names:
result = None
for policy in policies:
if not policy.get("enabled", True):
continue
pattern = policy.get("tool_pattern", "")
if fnmatch.fnmatch(name, pattern):
action = policy.get("action", "ask")
result = action if action in ("allow", "deny", "ask") else "ask"
break
results[name] = result
return results
+50 -1
View File
@@ -75,6 +75,7 @@ _ANTHROPIC_DEFAULT = ModelCapabilities(
token_param="max_tokens",
thinking_mode="manual",
supports_web_search=True,
supports_vision=True,
)
_ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
@@ -87,6 +88,7 @@ _ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
effort_levels=("low", "medium", "high", "max"),
supports_web_search=True,
supports_tool_search=True,
supports_vision=True,
),
"claude-sonnet-4-6": ModelCapabilities(
context_window=200000,
@@ -97,6 +99,7 @@ _ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
effort_levels=("low", "medium", "high"),
supports_web_search=True,
supports_tool_search=True,
supports_vision=True,
),
"claude-haiku-4-5": ModelCapabilities(
context_window=200000,
@@ -104,6 +107,7 @@ _ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
token_param="max_tokens",
thinking_mode="manual",
supports_web_search=True,
supports_vision=True,
),
"claude-sonnet-4-5": ModelCapabilities(
context_window=200000,
@@ -111,6 +115,7 @@ _ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
token_param="max_tokens",
thinking_mode="manual",
supports_web_search=True,
supports_vision=True,
),
"claude-opus-4-5": ModelCapabilities(
context_window=200000,
@@ -120,6 +125,7 @@ _ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
supports_effort=True,
effort_levels=("low", "medium", "high"),
supports_web_search=True,
supports_vision=True,
),
"claude-opus-4": ModelCapabilities(
context_window=200000,
@@ -128,6 +134,7 @@ _ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
thinking_mode="manual",
supports_web_search=True,
supports_tool_search=True,
supports_vision=True,
),
"claude-sonnet-4": ModelCapabilities(
context_window=200000,
@@ -136,6 +143,7 @@ _ANTHROPIC_CAPABILITIES: dict[str, ModelCapabilities] = {
thinking_mode="manual",
supports_web_search=True,
supports_tool_search=True,
supports_vision=True,
),
}
@@ -324,11 +332,15 @@ class AnthropicProvider:
tool_results: list[dict[str, Any]] = []
while i < len(messages) and messages[i]["role"] == "tool":
tool_msg = messages[i]
content = tool_msg.get("content", "")
# Convert image_url parts to Anthropic image format
if isinstance(content, list):
content = self._convert_content_parts(content)
tool_results.append(
{
"type": "tool_result",
"tool_use_id": tool_msg.get("tool_call_id", ""),
"content": tool_msg.get("content", ""),
"content": content,
}
)
i += 1
@@ -346,6 +358,43 @@ class AnthropicProvider:
return "\n\n".join(system_parts), _merge_consecutive(converted)
@staticmethod
def _convert_content_parts(parts: list[dict[str, Any]]) -> list[dict[str, Any]]:
"""Convert OpenAI-format content parts to Anthropic format.
Transforms ``image_url`` parts (with ``data:`` URIs) to Anthropic's
``image`` source blocks. Text parts pass through unchanged.
"""
converted: list[dict[str, Any]] = []
for part in parts:
if part.get("type") == "image_url":
url = part.get("image_url", {}).get("url", "")
if url.startswith("data:") and "," in url:
# Parse "data:image/png;base64,<data>"
header, _, b64data = url.partition(",")
media_type = header.split(":", 1)[1].split(";", 1)[0]
converted.append(
{
"type": "image",
"source": {
"type": "base64",
"media_type": media_type,
"data": b64data,
},
}
)
else:
# URL-based image — pass as Anthropic URL source
converted.append(
{
"type": "image",
"source": {"type": "url", "url": url},
}
)
else:
converted.append(part)
return converted
# -- tool conversion -----------------------------------------------------
def convert_tools(
+17
View File
@@ -30,6 +30,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
supports_temperature=False,
reasoning_effort_values=("minimal", "low", "medium", "high"),
default_reasoning_effort="medium",
supports_vision=True,
),
"gpt-5-mini": ModelCapabilities(
context_window=400000,
@@ -37,6 +38,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
supports_temperature=False,
reasoning_effort_values=("minimal", "low", "medium", "high"),
default_reasoning_effort="medium",
supports_vision=True,
),
"gpt-5-nano": ModelCapabilities(
context_window=400000,
@@ -44,6 +46,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
supports_temperature=False,
reasoning_effort_values=("minimal", "low", "medium", "high"),
default_reasoning_effort="medium",
supports_vision=True,
),
# GPT-5 pro — high reasoning only, extended output
"gpt-5-pro": ModelCapabilities(
@@ -52,6 +55,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
supports_temperature=False,
reasoning_effort_values=("high",),
default_reasoning_effort="high",
supports_vision=True,
),
# GPT-5.1 — temperature OK when reasoning_effort=none (default)
"gpt-5.1": ModelCapabilities(
@@ -59,6 +63,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
max_output_tokens=128000,
reasoning_effort_values=("none", "low", "medium", "high"),
default_reasoning_effort="none",
supports_vision=True,
),
# GPT-5.2 — adds xhigh
"gpt-5.2": ModelCapabilities(
@@ -66,6 +71,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
max_output_tokens=128000,
reasoning_effort_values=("none", "low", "medium", "high", "xhigh"),
default_reasoning_effort="none",
supports_vision=True,
),
# GPT-5.2 pro — always-reasoning variant
"gpt-5.2-pro": ModelCapabilities(
@@ -74,6 +80,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
supports_temperature=False,
reasoning_effort_values=("medium", "high", "xhigh"),
default_reasoning_effort="medium",
supports_vision=True,
),
# GPT-5.3 — same capabilities as 5.2 (matches gpt-5.3-chat-latest, codex)
"gpt-5.3": ModelCapabilities(
@@ -81,6 +88,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
max_output_tokens=128000,
reasoning_effort_values=("none", "low", "medium", "high", "xhigh"),
default_reasoning_effort="none",
supports_vision=True,
),
# GPT-5.4 — 1M context window, native tool search
"gpt-5.4": ModelCapabilities(
@@ -89,6 +97,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
reasoning_effort_values=("none", "low", "medium", "high", "xhigh"),
default_reasoning_effort="none",
supports_tool_search=True,
supports_vision=True,
),
# GPT-5.4 pro — always-reasoning, 1M context, native tool search
"gpt-5.4-pro": ModelCapabilities(
@@ -98,6 +107,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
reasoning_effort_values=("medium", "high", "xhigh"),
default_reasoning_effort="medium",
supports_tool_search=True,
supports_vision=True,
),
# O-series reasoning models
"o1": ModelCapabilities(
@@ -105,33 +115,39 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
max_output_tokens=100000,
supports_temperature=False,
supports_streaming=False,
supports_vision=True,
),
"o1-mini": ModelCapabilities(
context_window=128000,
max_output_tokens=65536,
supports_temperature=False,
supports_streaming=False,
supports_vision=True,
),
"o3": ModelCapabilities(
context_window=200000,
max_output_tokens=100000,
supports_temperature=False,
supports_vision=True,
),
"o3-mini": ModelCapabilities(
context_window=200000,
max_output_tokens=100000,
supports_temperature=False,
supports_vision=True,
),
"o3-pro": ModelCapabilities(
context_window=200000,
max_output_tokens=100000,
supports_temperature=False,
supports_streaming=False,
supports_vision=True,
),
"o4-mini": ModelCapabilities(
context_window=200000,
max_output_tokens=100000,
supports_temperature=False,
supports_vision=True,
),
# Search models — always search on every request, no reasoning_effort
"gpt-5-search-api": ModelCapabilities(
@@ -140,6 +156,7 @@ _OPENAI_CAPABILITIES: dict[str, ModelCapabilities] = {
supports_temperature=False,
supports_web_search=True,
reasoning_effort_values=(),
supports_vision=True,
),
}
+1
View File
@@ -77,6 +77,7 @@ class ModelCapabilities:
default_reasoning_effort: str = "medium"
supports_web_search: bool = False
supports_tool_search: bool = False
supports_vision: bool = False
def _lookup_capabilities(
+1266 -149
View File
File diff suppressed because it is too large Load Diff
+736
View File
@@ -10,9 +10,16 @@ import sqlalchemy as sa
from turnstone.core.storage._schema import (
api_tokens,
audit_events,
conversations,
memories,
metadata,
orgs,
prompt_templates,
roles,
tool_policies,
usage_events,
user_roles,
users,
workstream_config,
workstreams,
@@ -22,6 +29,23 @@ from turnstone.core.storage._sqlite import _reconstruct_messages
log = logging.getLogger(__name__)
def _row_to_dict(row: Any, *bool_fields: str) -> dict[str, Any]:
"""Convert a SQLAlchemy row to a dict, casting named fields to bool."""
d = dict(row._mapping)
for key in bool_fields:
if key in d:
d[key] = bool(d[key])
return d
# -- Field allowlists for governance update methods ---------------------------
_ROLE_MUTABLE = frozenset({"display_name", "permissions"})
_ORG_MUTABLE = frozenset({"display_name", "settings"})
_POLICY_MUTABLE = frozenset({"name", "tool_pattern", "action", "priority", "enabled"})
_TEMPLATE_MUTABLE = frozenset({"name", "content", "category", "variables", "is_default"})
class PostgreSQLBackend:
"""PostgreSQL implementation of the StorageBackend protocol."""
@@ -531,6 +555,7 @@ class PostgreSQLBackend:
from turnstone.core.storage._schema import channel_users
with self._engine.connect() as conn:
conn.execute(sa.delete(user_roles).where(user_roles.c.user_id == user_id))
conn.execute(sa.delete(channel_users).where(channel_users.c.user_id == user_id))
conn.execute(sa.delete(api_tokens).where(api_tokens.c.user_id == user_id))
result = conn.execute(sa.delete(users).where(users.c.user_id == user_id))
@@ -838,6 +863,7 @@ class PostgreSQLBackend:
auto_approve_tools: list[str],
created_by: str,
next_run: str,
template: str = "",
) -> None:
from sqlalchemy.dialects import postgresql
@@ -859,6 +885,7 @@ class PostgreSQLBackend:
initial_message=initial_message,
auto_approve=1 if auto_approve else 0,
auto_approve_tools=",".join(auto_approve_tools),
template=template,
enabled=1,
created_by=created_by,
next_run=next_run,
@@ -901,6 +928,7 @@ class PostgreSQLBackend:
"initial_message",
"auto_approve",
"auto_approve_tools",
"template",
"enabled",
"last_run",
"next_run",
@@ -1011,6 +1039,139 @@ class PostgreSQLBackend:
conn.commit()
return result.rowcount
# -- Watches ---------------------------------------------------------------
def create_watch(
self,
watch_id: str,
ws_id: str,
node_id: str,
name: str,
command: str,
interval_secs: float,
stop_on: str | None,
max_polls: int,
created_by: str,
next_poll: str,
) -> None:
from sqlalchemy.dialects import postgresql
from turnstone.core.storage._schema import watches
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
postgresql.insert(watches)
.values(
watch_id=watch_id,
ws_id=ws_id,
node_id=node_id,
name=name,
command=command,
interval_secs=interval_secs,
stop_on=stop_on,
max_polls=max_polls,
poll_count=0,
active=1,
created_by=created_by,
next_poll=next_poll,
created=now,
updated=now,
)
.on_conflict_do_nothing()
)
conn.commit()
def get_watch(self, watch_id: str) -> dict[str, Any] | None:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
row = conn.execute(sa.select(watches).where(watches.c.watch_id == watch_id)).fetchone()
if row is None:
return None
return dict(row._mapping)
def list_watches_for_ws(self, ws_id: str) -> list[dict[str, Any]]:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(watches)
.where((watches.c.ws_id == ws_id) & (watches.c.active == 1))
.order_by(watches.c.created.desc())
).fetchall()
return [dict(r._mapping) for r in rows]
def list_watches_for_node(self, node_id: str) -> list[dict[str, Any]]:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(watches)
.where((watches.c.node_id == node_id) & (watches.c.active == 1))
.order_by(watches.c.created.desc())
).fetchall()
return [dict(r._mapping) for r in rows]
def list_due_watches(self, now: str) -> list[dict[str, Any]]:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(watches)
.where(
(watches.c.active == 1)
& (watches.c.next_poll <= now)
& (watches.c.next_poll != "")
)
.order_by(watches.c.next_poll)
.limit(100)
).fetchall()
return [dict(r._mapping) for r in rows]
_UPDATABLE_WATCH_FIELDS = frozenset(
{
"name",
"poll_count",
"last_output",
"last_exit_code",
"last_poll",
"next_poll",
"active",
"updated",
}
)
def update_watch(self, watch_id: str, **fields: Any) -> bool:
from turnstone.core.storage._schema import watches
fields = {k: v for k, v in fields.items() if k in self._UPDATABLE_WATCH_FIELDS}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
if "active" in fields:
fields["active"] = 1 if fields["active"] else 0
with self._engine.connect() as conn:
result = conn.execute(
sa.update(watches).where(watches.c.watch_id == watch_id).values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_watch(self, watch_id: str) -> bool:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
result = conn.execute(sa.delete(watches).where(watches.c.watch_id == watch_id))
conn.commit()
return result.rowcount > 0
def delete_watches_for_ws(self, ws_id: str) -> int:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
result = conn.execute(sa.delete(watches).where(watches.c.ws_id == ws_id))
conn.commit()
return result.rowcount
# -- Service registry ------------------------------------------------------
def register_service(
@@ -1083,6 +1244,581 @@ class PostgreSQLBackend:
conn.commit()
return result.rowcount > 0
# -- Roles -----------------------------------------------------------------
def create_role(
self,
role_id: str,
name: str,
display_name: str,
permissions: str,
builtin: bool,
org_id: str = "",
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
existing = conn.execute(
sa.select(roles.c.role_id).where(roles.c.role_id == role_id)
).fetchone()
if not existing:
conn.execute(
sa.insert(roles),
{
"role_id": role_id,
"name": name,
"display_name": display_name,
"permissions": permissions,
"builtin": 1 if builtin else 0,
"org_id": org_id,
"created": now,
"updated": now,
},
)
conn.commit()
def get_role(self, role_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(sa.select(roles).where(roles.c.role_id == role_id)).fetchone()
if row:
return _row_to_dict(row, "builtin")
return None
def get_role_by_name(self, name: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(sa.select(roles).where(roles.c.name == name)).fetchone()
if row:
return _row_to_dict(row, "builtin")
return None
def list_roles(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(roles).order_by(roles.c.name.asc())
if org_id:
q = q.where(roles.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "builtin") for r in rows]
def update_role(self, role_id: str, **fields: Any) -> bool:
dropped = set(fields) - _ROLE_MUTABLE
if dropped:
log.warning("update_role: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _ROLE_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(
sa.update(roles).where(roles.c.role_id == role_id).values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_role(self, role_id: str) -> bool:
with self._engine.connect() as conn:
conn.execute(sa.delete(user_roles).where(user_roles.c.role_id == role_id))
result = conn.execute(sa.delete(roles).where(roles.c.role_id == role_id))
conn.commit()
return result.rowcount > 0
def assign_role(self, user_id: str, role_id: str, assigned_by: str = "") -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
existing = conn.execute(
sa.select(user_roles.c.user_id).where(
(user_roles.c.user_id == user_id) & (user_roles.c.role_id == role_id)
)
).fetchone()
if not existing:
conn.execute(
sa.insert(user_roles),
{
"user_id": user_id,
"role_id": role_id,
"assigned_by": assigned_by,
"created": now,
},
)
conn.commit()
def unassign_role(self, user_id: str, role_id: str) -> bool:
with self._engine.connect() as conn:
result = conn.execute(
sa.delete(user_roles).where(
(user_roles.c.user_id == user_id) & (user_roles.c.role_id == role_id)
)
)
conn.commit()
return result.rowcount > 0
def list_user_roles(self, user_id: str) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(
roles.c.role_id,
roles.c.name,
roles.c.display_name,
roles.c.permissions,
roles.c.builtin,
roles.c.org_id,
roles.c.created,
roles.c.updated,
user_roles.c.assigned_by,
user_roles.c.created.label("assignment_created"),
)
.select_from(user_roles.join(roles, user_roles.c.role_id == roles.c.role_id))
.where(user_roles.c.user_id == user_id)
).fetchall()
return [_row_to_dict(r, "builtin") for r in rows]
def get_user_permissions(self, user_id: str) -> set[str]:
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(roles.c.permissions)
.select_from(user_roles.join(roles, user_roles.c.role_id == roles.c.role_id))
.where(user_roles.c.user_id == user_id)
).fetchall()
perms: set[str] = set()
for r in rows:
if r[0]:
for p in r[0].split(","):
p = p.strip()
if p:
perms.add(p)
return perms
# -- Organizations ---------------------------------------------------------
def create_org(self, org_id: str, name: str, display_name: str, settings: str = "{}") -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
existing = conn.execute(
sa.select(orgs.c.org_id).where(orgs.c.org_id == org_id)
).fetchone()
if not existing:
conn.execute(
sa.insert(orgs),
{
"org_id": org_id,
"name": name,
"display_name": display_name,
"settings": settings,
"created": now,
"updated": now,
},
)
conn.commit()
def get_org(self, org_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(sa.select(orgs).where(orgs.c.org_id == org_id)).fetchone()
if row:
return _row_to_dict(row)
return None
def list_orgs(self) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(sa.select(orgs).order_by(orgs.c.name)).fetchall()
return [_row_to_dict(r) for r in rows]
def update_org(self, org_id: str, **fields: Any) -> bool:
dropped = set(fields) - _ORG_MUTABLE
if dropped:
log.warning("update_org: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _ORG_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(sa.update(orgs).where(orgs.c.org_id == org_id).values(**fields))
conn.commit()
return result.rowcount > 0
# -- Tool policies ---------------------------------------------------------
def create_tool_policy(
self,
policy_id: str,
name: str,
tool_pattern: str,
action: str,
priority: int,
org_id: str = "",
enabled: bool = True,
created_by: str = "",
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(tool_policies),
{
"policy_id": policy_id,
"name": name,
"tool_pattern": tool_pattern,
"action": action,
"priority": priority,
"org_id": org_id,
"enabled": 1 if enabled else 0,
"created_by": created_by,
"created": now,
"updated": now,
},
)
conn.commit()
def get_tool_policy(self, policy_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(
sa.select(tool_policies).where(tool_policies.c.policy_id == policy_id)
).fetchone()
if row:
return _row_to_dict(row, "enabled")
return None
def list_tool_policies(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(tool_policies).order_by(tool_policies.c.priority.desc())
if org_id:
q = q.where(tool_policies.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "enabled") for r in rows]
def update_tool_policy(self, policy_id: str, **fields: Any) -> bool:
dropped = set(fields) - _POLICY_MUTABLE
if dropped:
log.warning("update_tool_policy: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _POLICY_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
if "enabled" in fields:
fields["enabled"] = int(fields["enabled"])
with self._engine.connect() as conn:
result = conn.execute(
sa.update(tool_policies)
.where(tool_policies.c.policy_id == policy_id)
.values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_tool_policy(self, policy_id: str) -> bool:
with self._engine.connect() as conn:
result = conn.execute(
sa.delete(tool_policies).where(tool_policies.c.policy_id == policy_id)
)
conn.commit()
return result.rowcount > 0
# -- Prompt templates ------------------------------------------------------
def create_prompt_template(
self,
template_id: str,
name: str,
category: str,
content: str,
variables: str = "[]",
is_default: bool = False,
org_id: str = "",
created_by: str = "",
origin: str = "manual",
mcp_server: str = "",
readonly: bool = False,
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(prompt_templates),
{
"template_id": template_id,
"name": name,
"category": category,
"content": content,
"variables": variables,
"is_default": 1 if is_default else 0,
"org_id": org_id,
"created_by": created_by,
"origin": origin,
"mcp_server": mcp_server,
"readonly": 1 if readonly else 0,
"created": now,
"updated": now,
},
)
conn.commit()
def get_prompt_template(self, template_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(
sa.select(prompt_templates).where(prompt_templates.c.template_id == template_id)
).fetchone()
if row:
return _row_to_dict(row, "is_default", "readonly")
return None
def get_prompt_template_by_name(self, name: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(
sa.select(prompt_templates).where(prompt_templates.c.name == name)
).fetchone()
if row:
return _row_to_dict(row, "is_default", "readonly")
return None
def list_prompt_templates(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(prompt_templates).order_by(prompt_templates.c.name)
if org_id:
q = q.where(prompt_templates.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "is_default", "readonly") for r in rows]
def list_default_templates(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = (
sa.select(prompt_templates)
.where(prompt_templates.c.is_default == 1)
.order_by(prompt_templates.c.name)
)
if org_id:
q = q.where(prompt_templates.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "is_default", "readonly") for r in rows]
def list_prompt_templates_by_origin(self, origin: str) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(prompt_templates)
.where(prompt_templates.c.origin == origin)
.order_by(prompt_templates.c.name)
).fetchall()
return [_row_to_dict(r, "is_default", "readonly") for r in rows]
def update_prompt_template(self, template_id: str, **fields: Any) -> bool:
dropped = set(fields) - _TEMPLATE_MUTABLE
if dropped:
log.warning("update_prompt_template: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _TEMPLATE_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
if "is_default" in fields:
fields["is_default"] = int(fields["is_default"])
with self._engine.connect() as conn:
result = conn.execute(
sa.update(prompt_templates)
.where(prompt_templates.c.template_id == template_id)
.values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_prompt_template(self, template_id: str) -> bool:
with self._engine.connect() as conn:
result = conn.execute(
sa.delete(prompt_templates).where(prompt_templates.c.template_id == template_id)
)
conn.commit()
return result.rowcount > 0
# -- Usage events ----------------------------------------------------------
def record_usage_event(
self,
event_id: str,
user_id: str = "",
ws_id: str = "",
node_id: str = "",
model: str = "",
prompt_tokens: int = 0,
completion_tokens: int = 0,
tool_calls_count: int = 0,
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(usage_events),
{
"event_id": event_id,
"timestamp": now,
"user_id": user_id,
"ws_id": ws_id,
"node_id": node_id,
"model": model,
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"tool_calls_count": tool_calls_count,
"created": now,
},
)
conn.commit()
def query_usage(
self,
since: str,
until: str = "",
user_id: str = "",
model: str = "",
group_by: str = "",
) -> list[dict[str, Any]]:
clauses = ["timestamp >= :since"]
params: dict[str, Any] = {"since": since}
if until:
clauses.append("timestamp <= :until")
params["until"] = until
if user_id:
clauses.append("user_id = :user_id")
params["user_id"] = user_id
if model:
clauses.append("model = :model")
params["model"] = model
where = " AND ".join(clauses)
if group_by == "day":
key_expr = "substring(timestamp from 1 for 10)"
elif group_by == "hour":
key_expr = "substring(timestamp from 1 for 13)"
elif group_by == "model":
key_expr = "model"
elif group_by == "user":
key_expr = "user_id"
else:
# No grouping — single summary row
sql = (
f"SELECT SUM(prompt_tokens), SUM(completion_tokens), "
f"SUM(tool_calls_count) FROM usage_events WHERE {where}"
)
with self._engine.connect() as conn:
row = conn.execute(sa.text(sql), params).fetchone()
if row:
return [
{
"prompt_tokens": row[0] or 0,
"completion_tokens": row[1] or 0,
"tool_calls_count": row[2] or 0,
}
]
return [{"prompt_tokens": 0, "completion_tokens": 0, "tool_calls_count": 0}]
sql = (
f"SELECT {key_expr} AS key, SUM(prompt_tokens), SUM(completion_tokens), "
f"SUM(tool_calls_count) FROM usage_events WHERE {where} "
f"GROUP BY {key_expr} ORDER BY key ASC"
)
with self._engine.connect() as conn:
rows = conn.execute(sa.text(sql), params).fetchall()
return [
{
"key": r[0],
"prompt_tokens": r[1] or 0,
"completion_tokens": r[2] or 0,
"tool_calls_count": r[3] or 0,
}
for r in rows
]
def prune_usage_events(self, retention_days: int = 90) -> int:
cutoff = (datetime.now(UTC) - timedelta(days=retention_days)).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(sa.delete(usage_events).where(usage_events.c.timestamp < cutoff))
conn.commit()
return result.rowcount
# -- Audit events ----------------------------------------------------------
def record_audit_event(
self,
event_id: str,
user_id: str = "",
action: str = "",
resource_type: str = "",
resource_id: str = "",
detail: str = "{}",
ip_address: str = "",
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(audit_events),
{
"event_id": event_id,
"timestamp": now,
"user_id": user_id,
"action": action,
"resource_type": resource_type,
"resource_id": resource_id,
"detail": detail,
"ip_address": ip_address,
"created": now,
},
)
conn.commit()
def list_audit_events(
self,
action: str = "",
user_id: str = "",
since: str = "",
until: str = "",
limit: int = 100,
offset: int = 0,
) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(
audit_events.c.event_id,
audit_events.c.timestamp,
audit_events.c.user_id,
audit_events.c.action,
audit_events.c.resource_type,
audit_events.c.resource_id,
audit_events.c.detail,
audit_events.c.ip_address,
audit_events.c.created,
).order_by(audit_events.c.timestamp.desc(), audit_events.c.event_id.desc())
if action:
q = q.where(audit_events.c.action == action)
if user_id:
q = q.where(audit_events.c.user_id == user_id)
if since:
q = q.where(audit_events.c.timestamp >= since)
if until:
q = q.where(audit_events.c.timestamp <= until)
q = q.limit(limit).offset(offset)
rows = conn.execute(q).fetchall()
return [
{
"event_id": r[0],
"timestamp": r[1],
"user_id": r[2],
"action": r[3],
"resource_type": r[4],
"resource_id": r[5],
"detail": r[6],
"ip_address": r[7],
"created": r[8],
}
for r in rows
]
def count_audit_events(
self,
action: str = "",
user_id: str = "",
since: str = "",
until: str = "",
) -> int:
with self._engine.connect() as conn:
q = sa.select(sa.func.count()).select_from(audit_events)
if action:
q = q.where(audit_events.c.action == action)
if user_id:
q = q.where(audit_events.c.user_id == user_id)
if since:
q = q.where(audit_events.c.timestamp >= since)
if until:
q = q.where(audit_events.c.timestamp <= until)
row = conn.execute(q).fetchone()
return row[0] if row else 0
def prune_audit_events(self, retention_days: int = 365) -> int:
cutoff = (datetime.now(UTC) - timedelta(days=retention_days)).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(sa.delete(audit_events).where(audit_events.c.timestamp < cutoff))
conn.commit()
return result.rowcount
# -- Lifecycle -------------------------------------------------------------
def close(self) -> None:
+266
View File
@@ -247,6 +247,7 @@ class StorageBackend(Protocol):
auto_approve_tools: list[str],
created_by: str,
next_run: str,
template: str = "",
) -> None:
"""Create a scheduled task. No-op if task_id already exists."""
...
@@ -293,6 +294,52 @@ class StorageBackend(Protocol):
"""Delete task runs older than retention_days. Returns count deleted."""
...
# -- Watches ---------------------------------------------------------------
def create_watch(
self,
watch_id: str,
ws_id: str,
node_id: str,
name: str,
command: str,
interval_secs: float,
stop_on: str | None,
max_polls: int,
created_by: str,
next_poll: str,
) -> None:
"""Create a watch. No-op if watch_id already exists."""
...
def get_watch(self, watch_id: str) -> dict[str, Any] | None:
"""Return watch dict or None."""
...
def list_watches_for_ws(self, ws_id: str) -> list[dict[str, Any]]:
"""Return active watches for a workstream, ordered by created DESC."""
...
def list_watches_for_node(self, node_id: str) -> list[dict[str, Any]]:
"""Return all active watches on a node, ordered by created DESC."""
...
def list_due_watches(self, now: str) -> list[dict[str, Any]]:
"""Return active watches whose next_poll <= now, ordered by next_poll."""
...
def update_watch(self, watch_id: str, **fields: Any) -> bool:
"""Update specified fields on a watch. Returns True if found."""
...
def delete_watch(self, watch_id: str) -> bool:
"""Delete a watch. Returns True if found."""
...
def delete_watches_for_ws(self, ws_id: str) -> int:
"""Delete all watches for a workstream. Returns count deleted."""
...
# -- Service registry ------------------------------------------------------
def register_service(
@@ -313,6 +360,225 @@ class StorageBackend(Protocol):
"""Remove a service registration. Returns True if existed."""
...
# -- Roles (RBAC) ----------------------------------------------------------
def create_role(
self,
role_id: str,
name: str,
display_name: str,
permissions: str,
builtin: bool,
org_id: str,
) -> None:
"""Create a role. No-op if role_id already exists."""
...
def get_role(self, role_id: str) -> dict[str, Any] | None:
"""Return role dict or None."""
...
def get_role_by_name(self, name: str) -> dict[str, Any] | None:
"""Lookup role by name. Returns same dict as get_role or None."""
...
def list_roles(self, org_id: str = "") -> list[dict[str, Any]]:
"""Return all roles, optionally filtered by org_id. Ordered by name."""
...
def update_role(self, role_id: str, **fields: Any) -> bool:
"""Update specified fields on a role. Returns True if found."""
...
def delete_role(self, role_id: str) -> bool:
"""Delete a custom role. Returns True if found."""
...
def assign_role(self, user_id: str, role_id: str, assigned_by: str) -> None:
"""Assign a role to a user. No-op if already assigned."""
...
def unassign_role(self, user_id: str, role_id: str) -> bool:
"""Unassign a role from a user. Returns True if existed."""
...
def list_user_roles(self, user_id: str) -> list[dict[str, Any]]:
"""List roles assigned to a user (joins user_roles with roles)."""
...
def get_user_permissions(self, user_id: str) -> set[str]:
"""Return the union of all permissions from the user's assigned roles."""
...
# -- Organizations ---------------------------------------------------------
def create_org(self, org_id: str, name: str, display_name: str, settings: str = "{}") -> None:
"""Create an organization. No-op if org_id already exists."""
...
def get_org(self, org_id: str) -> dict[str, Any] | None:
"""Return org dict or None."""
...
def list_orgs(self) -> list[dict[str, Any]]:
"""Return all organizations ordered by name."""
...
def update_org(self, org_id: str, **fields: Any) -> bool:
"""Update specified fields on an org. Returns True if found."""
...
# -- Tool policies ---------------------------------------------------------
def create_tool_policy(
self,
policy_id: str,
name: str,
tool_pattern: str,
action: str,
priority: int,
org_id: str,
enabled: bool,
created_by: str,
) -> None:
"""Create a tool policy."""
...
def get_tool_policy(self, policy_id: str) -> dict[str, Any] | None:
"""Return tool policy dict or None."""
...
def list_tool_policies(self, org_id: str = "") -> list[dict[str, Any]]:
"""Return all tool policies ordered by priority DESC."""
...
def update_tool_policy(self, policy_id: str, **fields: Any) -> bool:
"""Update specified fields on a tool policy. Returns True if found."""
...
def delete_tool_policy(self, policy_id: str) -> bool:
"""Delete a tool policy. Returns True if found."""
...
# -- Prompt templates ------------------------------------------------------
def create_prompt_template(
self,
template_id: str,
name: str,
category: str,
content: str,
variables: str,
is_default: bool,
org_id: str,
created_by: str,
origin: str = "manual",
mcp_server: str = "",
readonly: bool = False,
) -> None:
"""Create a prompt template."""
...
def get_prompt_template(self, template_id: str) -> dict[str, Any] | None:
"""Return prompt template dict or None."""
...
def get_prompt_template_by_name(self, name: str) -> dict[str, Any] | None:
"""Lookup prompt template by name. Returns same dict as get_prompt_template or None."""
...
def list_prompt_templates(self, org_id: str = "") -> list[dict[str, Any]]:
"""Return all prompt templates ordered by name."""
...
def list_default_templates(self, org_id: str = "") -> list[dict[str, Any]]:
"""Return all templates where is_default=True, ordered by name."""
...
def list_prompt_templates_by_origin(self, origin: str) -> list[dict[str, Any]]:
"""Return all prompt templates with the given origin, ordered by name."""
...
def update_prompt_template(self, template_id: str, **fields: Any) -> bool:
"""Update specified fields on a prompt template. Returns True if found."""
...
def delete_prompt_template(self, template_id: str) -> bool:
"""Delete a prompt template. Returns True if found."""
...
# -- Usage events ----------------------------------------------------------
def record_usage_event(
self,
event_id: str,
user_id: str,
ws_id: str,
node_id: str,
model: str,
prompt_tokens: int,
completion_tokens: int,
tool_calls_count: int,
) -> None:
"""Record a usage event (token counts, tool calls for one LLM request)."""
...
def query_usage(
self,
since: str,
until: str = "",
user_id: str = "",
model: str = "",
group_by: str = "",
) -> list[dict[str, Any]]:
"""Query aggregated usage data. group_by: 'day', 'hour', 'model', 'user'."""
...
def prune_usage_events(self, retention_days: int = 90) -> int:
"""Delete usage events older than retention_days. Returns count deleted."""
...
# -- Audit events ----------------------------------------------------------
def record_audit_event(
self,
event_id: str,
user_id: str,
action: str,
resource_type: str,
resource_id: str,
detail: str,
ip_address: str,
) -> None:
"""Record an audit event."""
...
def list_audit_events(
self,
action: str = "",
user_id: str = "",
since: str = "",
until: str = "",
limit: int = 100,
offset: int = 0,
) -> list[dict[str, Any]]:
"""List audit events with optional filters, ordered by timestamp DESC."""
...
def count_audit_events(
self,
action: str = "",
user_id: str = "",
since: str = "",
until: str = "",
) -> int:
"""Count audit events matching the filters."""
...
def prune_audit_events(self, retention_days: int = 365) -> int:
"""Delete audit events older than retention_days. Returns count deleted."""
...
# -- Lifecycle -------------------------------------------------------------
def close(self) -> None:
+150
View File
@@ -71,6 +71,7 @@ users = sa.Table(
sa.Column("username", sa.Text, nullable=False, unique=True),
sa.Column("display_name", sa.Text, nullable=False),
sa.Column("password_hash", sa.Text, nullable=False),
sa.Column("org_id", sa.Text, nullable=False, server_default=""),
sa.Column("created", sa.Text, nullable=False),
)
@@ -139,6 +140,7 @@ scheduled_tasks = sa.Table(
sa.Column("initial_message", sa.Text, nullable=False),
sa.Column("auto_approve", sa.Integer, nullable=False, server_default="0"),
sa.Column("auto_approve_tools", sa.Text, nullable=False, server_default=""),
sa.Column("template", sa.Text, nullable=False, server_default=""),
sa.Column("enabled", sa.Integer, nullable=False, server_default="1"),
sa.Column("created_by", sa.Text, nullable=False, server_default=""),
sa.Column("last_run", sa.Text),
@@ -170,6 +172,40 @@ sa.Index("idx_scheduled_task_runs_started", scheduled_task_runs.c.started)
# Service registry
# ---------------------------------------------------------------------------
# ---------------------------------------------------------------------------
# Watches — in-session periodic command polling
# ---------------------------------------------------------------------------
watches = sa.Table(
"watches",
metadata,
sa.Column("watch_id", sa.Text, primary_key=True),
sa.Column("ws_id", sa.Text, nullable=False),
sa.Column("node_id", sa.Text, nullable=False, server_default=""),
sa.Column("name", sa.Text, nullable=False),
sa.Column("command", sa.Text, nullable=False),
sa.Column("interval_secs", sa.Float, nullable=False),
sa.Column("stop_on", sa.Text), # Python expression, NULL = change detection
sa.Column("max_polls", sa.Integer, nullable=False, server_default="100"),
sa.Column("poll_count", sa.Integer, nullable=False, server_default="0"),
sa.Column("last_output", sa.Text),
sa.Column("last_exit_code", sa.Integer),
sa.Column("last_poll", sa.Text), # ISO8601
sa.Column("next_poll", sa.Text), # ISO8601
sa.Column("active", sa.Integer, nullable=False, server_default="1"),
sa.Column("created_by", sa.Text, nullable=False, server_default=""),
sa.Column("created", sa.Text, nullable=False),
sa.Column("updated", sa.Text, nullable=False),
)
sa.Index("idx_watches_active_next", watches.c.active, watches.c.next_poll)
sa.Index("idx_watches_ws_id", watches.c.ws_id)
sa.Index("idx_watches_node_id", watches.c.node_id)
# ---------------------------------------------------------------------------
# Service registry
# ---------------------------------------------------------------------------
services = sa.Table(
"services",
metadata,
@@ -183,3 +219,117 @@ services = sa.Table(
)
sa.Index("idx_services_type_heartbeat", services.c.service_type, services.c.last_heartbeat)
# ---------------------------------------------------------------------------
# Governance tables — RBAC, orgs, policies, templates, usage, audit
# ---------------------------------------------------------------------------
orgs = sa.Table(
"orgs",
metadata,
sa.Column("org_id", sa.Text, primary_key=True),
sa.Column("name", sa.Text, nullable=False, unique=True),
sa.Column("display_name", sa.Text, nullable=False),
sa.Column("settings", sa.Text, nullable=False, server_default="{}"),
sa.Column("created", sa.Text, nullable=False),
sa.Column("updated", sa.Text, nullable=False),
)
roles = sa.Table(
"roles",
metadata,
sa.Column("role_id", sa.Text, primary_key=True),
sa.Column("name", sa.Text, nullable=False, unique=True),
sa.Column("display_name", sa.Text, nullable=False),
sa.Column("permissions", sa.Text, nullable=False), # comma-separated
sa.Column("builtin", sa.Integer, nullable=False, server_default="0"),
sa.Column("org_id", sa.Text, nullable=False, server_default=""),
sa.Column("created", sa.Text, nullable=False),
sa.Column("updated", sa.Text, nullable=False),
)
user_roles = sa.Table(
"user_roles",
metadata,
sa.Column("user_id", sa.Text, nullable=False),
sa.Column("role_id", sa.Text, nullable=False),
sa.Column("assigned_by", sa.Text, nullable=False, server_default=""),
sa.Column("created", sa.Text, nullable=False),
sa.PrimaryKeyConstraint("user_id", "role_id"),
)
sa.Index("idx_user_roles_role_id", user_roles.c.role_id)
tool_policies = sa.Table(
"tool_policies",
metadata,
sa.Column("policy_id", sa.Text, primary_key=True),
sa.Column("name", sa.Text, nullable=False),
sa.Column("tool_pattern", sa.Text, nullable=False),
sa.Column("action", sa.Text, nullable=False), # allow / deny / ask
sa.Column("priority", sa.Integer, nullable=False, server_default="0"),
sa.Column("org_id", sa.Text, nullable=False, server_default=""),
sa.Column("enabled", sa.Integer, nullable=False, server_default="1"),
sa.Column("created_by", sa.Text, nullable=False, server_default=""),
sa.Column("created", sa.Text, nullable=False),
sa.Column("updated", sa.Text, nullable=False),
)
sa.Index("idx_tool_policies_priority", tool_policies.c.priority.desc())
sa.Index("idx_tool_policies_org", tool_policies.c.org_id)
prompt_templates = sa.Table(
"prompt_templates",
metadata,
sa.Column("template_id", sa.Text, primary_key=True),
sa.Column("name", sa.Text, nullable=False, unique=True),
sa.Column("category", sa.Text, nullable=False, server_default="general"),
sa.Column("content", sa.Text, nullable=False),
sa.Column("variables", sa.Text, nullable=False, server_default="[]"), # JSON array
sa.Column("is_default", sa.Integer, nullable=False, server_default="0"),
sa.Column("org_id", sa.Text, nullable=False, server_default=""),
sa.Column("created_by", sa.Text, nullable=False, server_default=""),
sa.Column("origin", sa.Text, nullable=False, server_default="manual"),
sa.Column("mcp_server", sa.Text, nullable=False, server_default=""),
sa.Column("readonly", sa.Integer, nullable=False, server_default="0"),
sa.Column("created", sa.Text, nullable=False),
sa.Column("updated", sa.Text, nullable=False),
)
usage_events = sa.Table(
"usage_events",
metadata,
sa.Column("event_id", sa.Text, primary_key=True),
sa.Column("timestamp", sa.Text, nullable=False),
sa.Column("user_id", sa.Text, nullable=False, server_default=""),
sa.Column("ws_id", sa.Text, nullable=False, server_default=""),
sa.Column("node_id", sa.Text, nullable=False, server_default=""),
sa.Column("model", sa.Text, nullable=False, server_default=""),
sa.Column("prompt_tokens", sa.Integer, nullable=False, server_default="0"),
sa.Column("completion_tokens", sa.Integer, nullable=False, server_default="0"),
sa.Column("tool_calls_count", sa.Integer, nullable=False, server_default="0"),
sa.Column("created", sa.Text, nullable=False),
)
sa.Index("idx_usage_events_timestamp", usage_events.c.timestamp)
sa.Index("idx_usage_events_user", usage_events.c.user_id, usage_events.c.timestamp)
sa.Index("idx_usage_events_model", usage_events.c.model, usage_events.c.timestamp)
sa.Index("idx_usage_events_ws", usage_events.c.ws_id)
audit_events = sa.Table(
"audit_events",
metadata,
sa.Column("event_id", sa.Text, primary_key=True),
sa.Column("timestamp", sa.Text, nullable=False),
sa.Column("user_id", sa.Text, nullable=False, server_default=""),
sa.Column("action", sa.Text, nullable=False),
sa.Column("resource_type", sa.Text, nullable=False, server_default=""),
sa.Column("resource_id", sa.Text, nullable=False, server_default=""),
sa.Column("detail", sa.Text, nullable=False, server_default="{}"),
sa.Column("ip_address", sa.Text, nullable=False, server_default=""),
sa.Column("created", sa.Text, nullable=False),
)
sa.Index("idx_audit_timestamp", audit_events.c.timestamp)
sa.Index("idx_audit_action", audit_events.c.action)
sa.Index("idx_audit_user", audit_events.c.user_id)
+719
View File
@@ -12,9 +12,16 @@ import sqlalchemy as sa
from turnstone.core.storage._schema import (
api_tokens,
audit_events,
conversations,
memories,
metadata,
orgs,
prompt_templates,
roles,
tool_policies,
usage_events,
user_roles,
users,
workstream_config,
workstreams,
@@ -38,6 +45,23 @@ def _fts5_query(query: str) -> str:
return " ".join(safe)
def _row_to_dict(row: Any, *bool_fields: str) -> dict[str, Any]:
"""Convert a SQLAlchemy row to a dict, casting named fields to bool."""
d = dict(row._mapping)
for key in bool_fields:
if key in d:
d[key] = bool(d[key])
return d
# -- Field allowlists for governance update methods ---------------------------
_ROLE_MUTABLE = frozenset({"display_name", "permissions"})
_ORG_MUTABLE = frozenset({"display_name", "settings"})
_POLICY_MUTABLE = frozenset({"name", "tool_pattern", "action", "priority", "enabled"})
_TEMPLATE_MUTABLE = frozenset({"name", "content", "category", "variables", "is_default"})
class SQLiteBackend:
"""SQLite implementation of the StorageBackend protocol."""
@@ -590,6 +614,7 @@ class SQLiteBackend:
from turnstone.core.storage._schema import channel_users
with self._engine.connect() as conn:
conn.execute(sa.delete(user_roles).where(user_roles.c.user_id == user_id))
conn.execute(sa.delete(channel_users).where(channel_users.c.user_id == user_id))
conn.execute(sa.delete(api_tokens).where(api_tokens.c.user_id == user_id))
result = conn.execute(sa.delete(users).where(users.c.user_id == user_id))
@@ -891,6 +916,7 @@ class SQLiteBackend:
auto_approve_tools: list[str],
created_by: str,
next_run: str,
template: str = "",
) -> None:
from turnstone.core.storage._schema import scheduled_tasks
@@ -910,6 +936,7 @@ class SQLiteBackend:
"initial_message": initial_message,
"auto_approve": 1 if auto_approve else 0,
"auto_approve_tools": ",".join(auto_approve_tools),
"template": template,
"enabled": 1,
"created_by": created_by,
"next_run": next_run,
@@ -951,6 +978,7 @@ class SQLiteBackend:
"initial_message",
"auto_approve",
"auto_approve_tools",
"template",
"enabled",
"last_run",
"next_run",
@@ -1062,6 +1090,136 @@ class SQLiteBackend:
conn.commit()
return result.rowcount
# -- Watches ---------------------------------------------------------------
def create_watch(
self,
watch_id: str,
ws_id: str,
node_id: str,
name: str,
command: str,
interval_secs: float,
stop_on: str | None,
max_polls: int,
created_by: str,
next_poll: str,
) -> None:
from turnstone.core.storage._schema import watches
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(watches).prefix_with("OR IGNORE"),
{
"watch_id": watch_id,
"ws_id": ws_id,
"node_id": node_id,
"name": name,
"command": command,
"interval_secs": interval_secs,
"stop_on": stop_on,
"max_polls": max_polls,
"poll_count": 0,
"active": 1,
"created_by": created_by,
"next_poll": next_poll,
"created": now,
"updated": now,
},
)
conn.commit()
def get_watch(self, watch_id: str) -> dict[str, Any] | None:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
row = conn.execute(sa.select(watches).where(watches.c.watch_id == watch_id)).fetchone()
if row is None:
return None
return dict(row._mapping)
def list_watches_for_ws(self, ws_id: str) -> list[dict[str, Any]]:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(watches)
.where((watches.c.ws_id == ws_id) & (watches.c.active == 1))
.order_by(watches.c.created.desc())
).fetchall()
return [dict(r._mapping) for r in rows]
def list_watches_for_node(self, node_id: str) -> list[dict[str, Any]]:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(watches)
.where((watches.c.node_id == node_id) & (watches.c.active == 1))
.order_by(watches.c.created.desc())
).fetchall()
return [dict(r._mapping) for r in rows]
def list_due_watches(self, now: str) -> list[dict[str, Any]]:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(watches)
.where(
(watches.c.active == 1)
& (watches.c.next_poll <= now)
& (watches.c.next_poll != "")
)
.order_by(watches.c.next_poll)
.limit(100)
).fetchall()
return [dict(r._mapping) for r in rows]
_UPDATABLE_WATCH_FIELDS = frozenset(
{
"name",
"poll_count",
"last_output",
"last_exit_code",
"last_poll",
"next_poll",
"active",
"updated",
}
)
def update_watch(self, watch_id: str, **fields: Any) -> bool:
from turnstone.core.storage._schema import watches
fields = {k: v for k, v in fields.items() if k in self._UPDATABLE_WATCH_FIELDS}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
if "active" in fields:
fields["active"] = 1 if fields["active"] else 0
with self._engine.connect() as conn:
result = conn.execute(
sa.update(watches).where(watches.c.watch_id == watch_id).values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_watch(self, watch_id: str) -> bool:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
result = conn.execute(sa.delete(watches).where(watches.c.watch_id == watch_id))
conn.commit()
return result.rowcount > 0
def delete_watches_for_ws(self, ws_id: str) -> int:
from turnstone.core.storage._schema import watches
with self._engine.connect() as conn:
result = conn.execute(sa.delete(watches).where(watches.c.ws_id == ws_id))
conn.commit()
return result.rowcount
# -- Service registry ------------------------------------------------------
def register_service(
@@ -1134,6 +1292,567 @@ class SQLiteBackend:
conn.commit()
return result.rowcount > 0
# -- Roles -----------------------------------------------------------------
def create_role(
self,
role_id: str,
name: str,
display_name: str,
permissions: str,
builtin: bool,
org_id: str = "",
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(roles).prefix_with("OR IGNORE"),
{
"role_id": role_id,
"name": name,
"display_name": display_name,
"permissions": permissions,
"builtin": 1 if builtin else 0,
"org_id": org_id,
"created": now,
"updated": now,
},
)
conn.commit()
def get_role(self, role_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(sa.select(roles).where(roles.c.role_id == role_id)).fetchone()
if row:
return _row_to_dict(row, "builtin")
return None
def get_role_by_name(self, name: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(sa.select(roles).where(roles.c.name == name)).fetchone()
if row:
return _row_to_dict(row, "builtin")
return None
def list_roles(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(roles).order_by(roles.c.name.asc())
if org_id:
q = q.where(roles.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "builtin") for r in rows]
def update_role(self, role_id: str, **fields: Any) -> bool:
dropped = set(fields) - _ROLE_MUTABLE
if dropped:
log.warning("update_role: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _ROLE_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(
sa.update(roles).where(roles.c.role_id == role_id).values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_role(self, role_id: str) -> bool:
with self._engine.connect() as conn:
conn.execute(sa.delete(user_roles).where(user_roles.c.role_id == role_id))
result = conn.execute(sa.delete(roles).where(roles.c.role_id == role_id))
conn.commit()
return result.rowcount > 0
def assign_role(self, user_id: str, role_id: str, assigned_by: str = "") -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(user_roles).prefix_with("OR IGNORE"),
{
"user_id": user_id,
"role_id": role_id,
"assigned_by": assigned_by,
"created": now,
},
)
conn.commit()
def unassign_role(self, user_id: str, role_id: str) -> bool:
with self._engine.connect() as conn:
result = conn.execute(
sa.delete(user_roles).where(
(user_roles.c.user_id == user_id) & (user_roles.c.role_id == role_id)
)
)
conn.commit()
return result.rowcount > 0
def list_user_roles(self, user_id: str) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(
roles.c.role_id,
roles.c.name,
roles.c.display_name,
roles.c.permissions,
roles.c.builtin,
roles.c.org_id,
roles.c.created,
roles.c.updated,
user_roles.c.assigned_by,
user_roles.c.created.label("assignment_created"),
)
.select_from(user_roles.join(roles, user_roles.c.role_id == roles.c.role_id))
.where(user_roles.c.user_id == user_id)
).fetchall()
return [_row_to_dict(r, "builtin") for r in rows]
def get_user_permissions(self, user_id: str) -> set[str]:
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(roles.c.permissions)
.select_from(user_roles.join(roles, user_roles.c.role_id == roles.c.role_id))
.where(user_roles.c.user_id == user_id)
).fetchall()
perms: set[str] = set()
for r in rows:
if r[0]:
for p in r[0].split(","):
p = p.strip()
if p:
perms.add(p)
return perms
# -- Organizations ---------------------------------------------------------
def create_org(self, org_id: str, name: str, display_name: str, settings: str = "{}") -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(orgs).prefix_with("OR IGNORE"),
{
"org_id": org_id,
"name": name,
"display_name": display_name,
"settings": settings,
"created": now,
"updated": now,
},
)
conn.commit()
def get_org(self, org_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(sa.select(orgs).where(orgs.c.org_id == org_id)).fetchone()
if row:
return _row_to_dict(row)
return None
def list_orgs(self) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(sa.select(orgs).order_by(orgs.c.name)).fetchall()
return [_row_to_dict(r) for r in rows]
def update_org(self, org_id: str, **fields: Any) -> bool:
dropped = set(fields) - _ORG_MUTABLE
if dropped:
log.warning("update_org: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _ORG_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(sa.update(orgs).where(orgs.c.org_id == org_id).values(**fields))
conn.commit()
return result.rowcount > 0
# -- Tool policies ---------------------------------------------------------
def create_tool_policy(
self,
policy_id: str,
name: str,
tool_pattern: str,
action: str,
priority: int,
org_id: str = "",
enabled: bool = True,
created_by: str = "",
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(tool_policies),
{
"policy_id": policy_id,
"name": name,
"tool_pattern": tool_pattern,
"action": action,
"priority": priority,
"org_id": org_id,
"enabled": 1 if enabled else 0,
"created_by": created_by,
"created": now,
"updated": now,
},
)
conn.commit()
def get_tool_policy(self, policy_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(
sa.select(tool_policies).where(tool_policies.c.policy_id == policy_id)
).fetchone()
if row:
return _row_to_dict(row, "enabled")
return None
def list_tool_policies(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(tool_policies).order_by(tool_policies.c.priority.desc())
if org_id:
q = q.where(tool_policies.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "enabled") for r in rows]
def update_tool_policy(self, policy_id: str, **fields: Any) -> bool:
dropped = set(fields) - _POLICY_MUTABLE
if dropped:
log.warning("update_tool_policy: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _POLICY_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
if "enabled" in fields:
fields["enabled"] = int(fields["enabled"])
with self._engine.connect() as conn:
result = conn.execute(
sa.update(tool_policies)
.where(tool_policies.c.policy_id == policy_id)
.values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_tool_policy(self, policy_id: str) -> bool:
with self._engine.connect() as conn:
result = conn.execute(
sa.delete(tool_policies).where(tool_policies.c.policy_id == policy_id)
)
conn.commit()
return result.rowcount > 0
# -- Prompt templates ------------------------------------------------------
def create_prompt_template(
self,
template_id: str,
name: str,
category: str,
content: str,
variables: str = "[]",
is_default: bool = False,
org_id: str = "",
created_by: str = "",
origin: str = "manual",
mcp_server: str = "",
readonly: bool = False,
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(prompt_templates),
{
"template_id": template_id,
"name": name,
"category": category,
"content": content,
"variables": variables,
"is_default": 1 if is_default else 0,
"org_id": org_id,
"created_by": created_by,
"origin": origin,
"mcp_server": mcp_server,
"readonly": 1 if readonly else 0,
"created": now,
"updated": now,
},
)
conn.commit()
def get_prompt_template(self, template_id: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(
sa.select(prompt_templates).where(prompt_templates.c.template_id == template_id)
).fetchone()
if row:
return _row_to_dict(row, "is_default", "readonly")
return None
def get_prompt_template_by_name(self, name: str) -> dict[str, Any] | None:
with self._engine.connect() as conn:
row = conn.execute(
sa.select(prompt_templates).where(prompt_templates.c.name == name)
).fetchone()
if row:
return _row_to_dict(row, "is_default", "readonly")
return None
def list_prompt_templates(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(prompt_templates).order_by(prompt_templates.c.name)
if org_id:
q = q.where(prompt_templates.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "is_default", "readonly") for r in rows]
def list_default_templates(self, org_id: str = "") -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = (
sa.select(prompt_templates)
.where(prompt_templates.c.is_default == 1)
.order_by(prompt_templates.c.name)
)
if org_id:
q = q.where(prompt_templates.c.org_id == org_id)
rows = conn.execute(q).fetchall()
return [_row_to_dict(r, "is_default", "readonly") for r in rows]
def list_prompt_templates_by_origin(self, origin: str) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
rows = conn.execute(
sa.select(prompt_templates)
.where(prompt_templates.c.origin == origin)
.order_by(prompt_templates.c.name)
).fetchall()
return [_row_to_dict(r, "is_default", "readonly") for r in rows]
def update_prompt_template(self, template_id: str, **fields: Any) -> bool:
dropped = set(fields) - _TEMPLATE_MUTABLE
if dropped:
log.warning("update_prompt_template: ignoring unknown fields: %s", dropped)
fields = {k: v for k, v in fields.items() if k in _TEMPLATE_MUTABLE}
fields["updated"] = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
if "is_default" in fields:
fields["is_default"] = int(fields["is_default"])
with self._engine.connect() as conn:
result = conn.execute(
sa.update(prompt_templates)
.where(prompt_templates.c.template_id == template_id)
.values(**fields)
)
conn.commit()
return result.rowcount > 0
def delete_prompt_template(self, template_id: str) -> bool:
with self._engine.connect() as conn:
result = conn.execute(
sa.delete(prompt_templates).where(prompt_templates.c.template_id == template_id)
)
conn.commit()
return result.rowcount > 0
# -- Usage events ----------------------------------------------------------
def record_usage_event(
self,
event_id: str,
user_id: str = "",
ws_id: str = "",
node_id: str = "",
model: str = "",
prompt_tokens: int = 0,
completion_tokens: int = 0,
tool_calls_count: int = 0,
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(usage_events),
{
"event_id": event_id,
"timestamp": now,
"user_id": user_id,
"ws_id": ws_id,
"node_id": node_id,
"model": model,
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"tool_calls_count": tool_calls_count,
"created": now,
},
)
conn.commit()
def query_usage(
self,
since: str,
until: str = "",
user_id: str = "",
model: str = "",
group_by: str = "",
) -> list[dict[str, Any]]:
clauses = ["timestamp >= :since"]
params: dict[str, Any] = {"since": since}
if until:
clauses.append("timestamp <= :until")
params["until"] = until
if user_id:
clauses.append("user_id = :user_id")
params["user_id"] = user_id
if model:
clauses.append("model = :model")
params["model"] = model
where = " AND ".join(clauses)
if group_by == "day":
key_expr = "substr(timestamp, 1, 10)"
elif group_by == "hour":
key_expr = "substr(timestamp, 1, 13)"
elif group_by == "model":
key_expr = "model"
elif group_by == "user":
key_expr = "user_id"
else:
# No grouping — single summary row
sql = (
f"SELECT SUM(prompt_tokens), SUM(completion_tokens), "
f"SUM(tool_calls_count) FROM usage_events WHERE {where}"
)
with self._engine.connect() as conn:
row = conn.execute(sa.text(sql), params).fetchone()
if row:
return [
{
"prompt_tokens": row[0] or 0,
"completion_tokens": row[1] or 0,
"tool_calls_count": row[2] or 0,
}
]
return [{"prompt_tokens": 0, "completion_tokens": 0, "tool_calls_count": 0}]
sql = (
f"SELECT {key_expr} AS key, SUM(prompt_tokens), SUM(completion_tokens), "
f"SUM(tool_calls_count) FROM usage_events WHERE {where} "
f"GROUP BY {key_expr} ORDER BY key ASC"
)
with self._engine.connect() as conn:
rows = conn.execute(sa.text(sql), params).fetchall()
return [
{
"key": r[0],
"prompt_tokens": r[1] or 0,
"completion_tokens": r[2] or 0,
"tool_calls_count": r[3] or 0,
}
for r in rows
]
def prune_usage_events(self, retention_days: int = 90) -> int:
cutoff = (datetime.now(UTC) - timedelta(days=retention_days)).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(sa.delete(usage_events).where(usage_events.c.timestamp < cutoff))
conn.commit()
return result.rowcount
# -- Audit events ----------------------------------------------------------
def record_audit_event(
self,
event_id: str,
user_id: str = "",
action: str = "",
resource_type: str = "",
resource_id: str = "",
detail: str = "{}",
ip_address: str = "",
) -> None:
now = datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
conn.execute(
sa.insert(audit_events),
{
"event_id": event_id,
"timestamp": now,
"user_id": user_id,
"action": action,
"resource_type": resource_type,
"resource_id": resource_id,
"detail": detail,
"ip_address": ip_address,
"created": now,
},
)
conn.commit()
def list_audit_events(
self,
action: str = "",
user_id: str = "",
since: str = "",
until: str = "",
limit: int = 100,
offset: int = 0,
) -> list[dict[str, Any]]:
with self._engine.connect() as conn:
q = sa.select(
audit_events.c.event_id,
audit_events.c.timestamp,
audit_events.c.user_id,
audit_events.c.action,
audit_events.c.resource_type,
audit_events.c.resource_id,
audit_events.c.detail,
audit_events.c.ip_address,
audit_events.c.created,
).order_by(audit_events.c.timestamp.desc(), audit_events.c.event_id.desc())
if action:
q = q.where(audit_events.c.action == action)
if user_id:
q = q.where(audit_events.c.user_id == user_id)
if since:
q = q.where(audit_events.c.timestamp >= since)
if until:
q = q.where(audit_events.c.timestamp <= until)
q = q.limit(limit).offset(offset)
rows = conn.execute(q).fetchall()
return [
{
"event_id": r[0],
"timestamp": r[1],
"user_id": r[2],
"action": r[3],
"resource_type": r[4],
"resource_id": r[5],
"detail": r[6],
"ip_address": r[7],
"created": r[8],
}
for r in rows
]
def count_audit_events(
self,
action: str = "",
user_id: str = "",
since: str = "",
until: str = "",
) -> int:
with self._engine.connect() as conn:
q = sa.select(sa.func.count()).select_from(audit_events)
if action:
q = q.where(audit_events.c.action == action)
if user_id:
q = q.where(audit_events.c.user_id == user_id)
if since:
q = q.where(audit_events.c.timestamp >= since)
if until:
q = q.where(audit_events.c.timestamp <= until)
row = conn.execute(q).fetchone()
return row[0] if row else 0
def prune_audit_events(self, retention_days: int = 365) -> int:
cutoff = (datetime.now(UTC) - timedelta(days=retention_days)).strftime("%Y-%m-%dT%H:%M:%S")
with self._engine.connect() as conn:
result = conn.execute(sa.delete(audit_events).where(audit_events.c.timestamp < cutoff))
conn.commit()
return result.rowcount
# -- Lifecycle -------------------------------------------------------------
def close(self) -> None:

Some files were not shown because too many files have changed in this diff Show More