* feat(coordinator): phase 7 — governance + skill metadata + cross-cutting invariants
Combines three stacked sub-PRs into a single coordinator phase-7
shipment against the phase-7 plan doc. The sub-PR structure (0 / A /
B) preserved on individual branches for reviewer drill-down; this
branch is the one reviewers should merge.
## Sub-PR 0 — service-auth boundary invariants
Shared helpers and contracts that lock the console ↔ node service-auth
boundary so later authz surfaces use them by construction.
- ``_effective_user_filter(request)`` in both ``turnstone.console.server``
and ``turnstone.server`` with a shared ``DENY_EMPTY_SUB`` sentinel
on ``turnstone.core.auth``. Three-way return — admin/service
bypass, scoped caller uid, or fail-closed sentinel on blank sub.
Four callsite migrations (``_coordinator_rows``,
``coordinator_children``, ``coordinator_metrics``,
``cluster_ws_live_bulk``).
- ``StorageBackend`` class docstring codifies the tenancy contract
(every list/count/aggregate method must accept ``user_id: str |
None = None`` and push ``WHERE user_id = :user_id`` into SQL) and
the ``_mapping`` row-access contract. New
``turnstone.testing.row_contract`` ships ``assert_row_like()``.
- ``_verify_collector_service_scope`` probes an upstream node at boot
with ``expected_node_id=_scope-probe_``; a 409 proves the scope
gate was passed, a 403/401 sets ``collector_scope_error`` and
causes ``cluster_snapshot`` / ``cluster_events_sse`` to return 503
with a remediation hint. Probe URL allowlist rejects non-http(s)
schemes and 169.254.0.0/16 hosts.
- 4xx log-level floor on ``_NodeDashboardCache.get``,
``_fetch_live_block``, and ``_proxy_sse`` — dotted-hierarchy
prefixes with bounded body previews. ``_bounded_body_preview`` and
``_bounded_stream_preview`` strip control chars.
## Sub-PR A — coordinator governance core
Mid-session governance surface for coordinator workstreams.
- **Trusted-session mode.** New ``coordinator.trust.send``
permission (migration 042). ``ChatSession.set_trust_send`` /
``revoke_tools`` methods with a ``_governance_lock``. ``POST
/v1/api/coordinator/{ws_id}/trust {send: bool}`` double-gated on
``admin.coordinator`` AND ``coordinator.trust.send`` with
``allow_service_bypass=False`` so service tokens can't escalate.
``_prepare_send_to_workstream`` auto-approves sends whose target is
in the coordinator's own subtree; foreign ws_ids still require
approval. ``_is_own_subtree`` checks both ``parent_ws_id`` AND
``user_id`` to defend against cross-tenant row corruption.
- **Audit-layer credential redaction.** ``record_audit`` walks
``detail`` (dicts, lists, tuples, sets, frozensets; keys too)
and routes every string through ``redact_credentials`` + a C0
control-char scrub. New kw-only ``raw_detail=True`` opt-out.
``_has_any_string`` fast-path. Audit action registry extended
with the four new governance sub-prefixes.
- **Mid-session revocation + cascading stop.** ``POST
/v1/api/coordinator/{ws_id}/restrict {revoke: [...]}`` caps 256
entries / 128 chars; ``_prepare_tool`` short-circuits with a
tool-error. ``POST /v1/api/coordinator/{ws_id}/stop_cascade``
cancels the coord's in-flight generation then dispatches
``cancel_workstream`` for every direct child in parallel via
``asyncio.gather`` bounded by ``Semaphore(16)``. Per-child
outcomes split into ``cancelled`` / ``failed`` / ``skipped``
(404 = already-gone rather than dispatch-broken). Both endpoints
apply ``allow_service_bypass=False`` on the admin gate.
- **Shared plumbing.** ``_resolve_coord_session`` helper collapses
the handler prelude three endpoints shared. ``_emit_coord_audit``
wraps ``record_audit`` in a dedicated ``ThreadPoolExecutor``
(``app.state.audit_executor``) so audit bursts don't starve cancel
dispatches. ``_require_json_object`` guards body parsing so non-
object JSON returns 400 instead of 500.
## Sub-PR B — skill metadata governance
- **Description validator (migration 043).** ``prompt_templates``
rows now require a non-empty ``description``. Existing empty rows
get backfilled with a ``"Skill: <name>"`` placeholder on upgrade.
The installer (``admin_skill_discover``) and MCP prompt sync both
synthesise a placeholder when the upstream description is blank
so non-admin write paths satisfy the invariant.
- **Skill kind classifier (migration 044).** New
``prompt_templates.kind`` column (``interactive`` / ``coordinator``
/ ``any``; defaults to ``any``). New
``turnstone.core.skill_kind.SkillKind`` StrEnum is the single
source of truth; Pydantic schemas type ``kind`` as ``SkillKind``
(OpenAPI advertises the enum) and the handler validator catches
the ValueError. ``list_skills_filtered`` gains a
``kinds: list[str] | None = None`` SQL filter.
``CoordinatorClient.list_skills`` defaults to
``kinds=["coordinator", "any"]`` so interactive-only skills are
hidden from the orchestrator.
- **``scan_status`` → ``risk_level`` rename (migration 045).**
Lossless column rename to align with ``IntentVerdict.risk_level``
terminology. Swept storage (both backends + schema + protocol),
handlers, API schemas, tool JSON, generated OpenAPI specs,
TypeScript SDK types, frontend (``governance.js``), tests, and
English prose in ``docs/judge.md`` + ``docs/tools.md``. The
user-facing on-load warning now reads ``has risk level:
{risk_tier}``. Tool JSON's ``risk_level`` enum corrected to the
scanner's actual taxonomy (``safe / low / medium / high /
critical``; was the never-shipped ``clean / flagged / unscanned /
pending``). Historical migration 021 left untouched.
## Migrations
042 (``coordinator.trust.send`` perm — PR A)
043 (description backfill — PR B)
044 (``kind`` column add — PR B)
045 (``scan_status`` → ``risk_level`` rename — PR B)
All four use position-anchored permission strings / host-side
parse-filter-rejoin on downgrade where SQL ``REPLACE`` could
corrupt prefix-overlapping values.
## Verification
- ``ruff check turnstone tests`` clean.
- ``mypy turnstone`` clean on 165 source files.
- ``pytest -m "not live"``: 4431 passed (+85 over the phase-6
baseline). Includes +32 tests in ``tests/test_service_auth_boundary.py``
and +38 in ``tests/test_coordinator_governance.py``; shared fixtures
extracted to ``tests/_coord_test_helpers.py``.
- Generated OpenAPI JSON (``sdk/typescript/openapi-{console,server}.json``)
regenerated via ``sdk/typescript/scripts/generate-types.py``; zero
``scan_status`` occurrences remaining outside the historical
migration 021 and the rename migration 045.
## Security reviews
Both reviews flagged by the phase-7 plan (items 1 + 5, plus 0a's
refuse-to-serve gate) ran through the multi-stage ``/review``
pipeline twice per sub-PR; all confirmed findings landed in-branch.
* fixup(phase-7): CI lint + PR #383 review fixups
Addresses the lint CI failure (ruff format) plus 12 findings from the
two automated PR reviewers.
Copilot:
- ``_sqlite.list_installed_skill_urls`` / ``_postgresql.list_installed_skill_urls``
used positional row indexing (``r[0]``/``r[1]``/``r[2]``) while this
same PR's ``StorageBackend`` class docstring forbids it. Switched
both to ``r._mapping["..."]`` access.
- ``list_skills.json`` previously advertised ``risk_level=""`` as a
filter for unscanned skills, but the implementation treats empty
strings as "no filter". Clarified the tool description to say
omit the filter entirely to include unscanned rows, and added an
explicit ``enum`` on the parameter restricting it to the scanner
tiers. ``_prepare_list_skills`` keeps the ``strip() or None``
normalisation — unscanned filtering now has an unambiguous contract.
- ``test_storage_skills_filtered.test_risk_level_filter`` used the
legacy ``clean`` / ``flagged`` values from the pre-rename column.
Rewritten with the scanner's actual taxonomy (``safe`` / ``high``).
github-code-quality (CodeQL):
- ``test_deny_sentinel_is_singleton`` previously asserted
``cs.DENY_EMPTY_SUB is cs.DENY_EMPTY_SUB`` — an identical-expression
comparison. Rewritten as two separate ``from ... import ... as`` aliases
(``FIRST_READ`` / ``SECOND_READ``) so the identity check is between
distinct bindings.
- ``test_restrict_empty_revoke_is_noop_but_audits`` unpacked ``state``
without using it. Renamed to ``_state``.
- Mixed import styles in ``test_service_auth_boundary.py`` — the
file previously used both ``import turnstone.console.server as cs``
and ``from turnstone.console.server import ...`` for the same
module (same story for ``turnstone.core.auth`` and
``turnstone.server``). Consolidated to the ``from X import Y`` style
used elsewhere in the file; the ``_fetch_live_block`` test now
patches via pytest's ``monkeypatch`` fixture instead of a manual
rebind through a module alias.
CI:
- ``ruff format`` reformatted one line in
``tests/test_coordinator_endpoints.py``.
Verification: ruff check + mypy clean (166 files); 4459 non-live
pytest pass.
* fix(tests): swap asyncio marker for anyio in service-auth boundary tests
PR #383 CI caught that the 13 ``@pytest.mark.asyncio`` decorators I
added in ``test_service_auth_boundary.py`` are an off-convention
choice — the rest of the repo uses ``@pytest.mark.anyio`` (148 sites
vs my 13). The CI environment pulls in ``anyio`` but not
``pytest-asyncio``, so every async test in this one file was failing
with "async def functions are not natively supported". It passed
locally by accident — my dev venv happens to have pytest-asyncio
installed ambiently.
Swapped all 13 marker sites to ``@pytest.mark.anyio``. No functional
change; the tests run under the same default asyncio backend anyio
provides.
Verification: ruff + mypy clean (166 files); 4459 non-live pytest
pass.
12 KiB
Governance
Turnstone governance provides role-based access control (RBAC), tool execution policies, skills, usage tracking, and audit logging for the admin console.
Architecture
See diagram: 19-governance-architecture.puml.
RBAC (Roles & Permissions)
The permission model has two layers:
- Scopes (legacy) —
read,write,approve. Checked byAuthMiddlewareon every request based on URL path classification. - Permissions (granular) — 15 permission strings checked per-endpoint by
require_permission().
Built-in roles (seeded by migration 008):
| Role | Permissions |
|---|---|
| admin | read, write, approve, admin.users, admin.roles, admin.orgs, admin.policies, admin.skills, admin.audit, admin.usage, admin.schedules, admin.watches, tools.approve, workstreams.create, workstreams.close |
| operator | read, write, workstreams.create, workstreams.close |
| viewer | read |
Custom roles can be created with any subset of the 15 valid permissions.
Auth flow:
- User logs in (password or API token) →
_load_user_permissions()aggregates permissions from all assigned roles _permissions_to_scopes()derives legacy scopes (anyadmin.*→approve)- JWT created with both
scopesandpermissionsclaims - Middleware checks scope → handler checks permission via
require_permission()
Tool Policies
Admin-defined rules that control tool execution:
- Pattern matching: Glob syntax via
fnmatch(e.g.,bash*,file_write,*) - Actions:
allow(auto-approve),deny(block),ask(normal approval flow) - Priority: Higher priority evaluated first, first match wins
- Enforcement:
evaluate_tool_policies_batch()called inWebUI.approve_tools()before theauto_approvecheck - MCP granular policies: MCP resources and prompts are evaluated using their
approval_labelfor fine-grained control:- Resource reads:
mcp_resource__{uri}(e.g.,mcp_resource__file:///docs/*to allow,mcp_resource__*to deny all) - Prompt invocations:
mcp__{server}__{prompt}(e.g.,mcp__trusted__*to allow,mcp__*to require approval for all) - Built-in tools continue to use
func_namefor backward compatibility
- Resource reads:
Skills
Admin-curated system message skills injected at workstream startup. Skills also include session configuration (model, temperature, auto-approve, token budget, etc.) since workstream templates were merged into the skills system in v0.8.0.
- Runtime behavior: Skills are loaded once at session creation and injected
into the system message before user
instructions. Skills set the baseline; instructions customize per-workstream behavior. - Default skills: All
is_default=trueskills auto-apply to new workstreams, concatenated in alphabetical order by name. Use name prefixes (e.g.01-safety,02-style) to control ordering. - Explicit selection:
--skill <name>CLI flag,skillfield onPOST /v1/api/workstreams/new, console creation modal dropdown, scheduled task config, and channel adapter config. An explicit skill replaces defaults. - Variables: Three built-in placeholders resolved at load time:
{{model}}(active model name),{{ws_id}}(workstream ID),{{node_id}}(server node ID). Unrecognized placeholders are kept as-is. - Runtime switching:
/skill <name>to switch,/skill clearto revert to defaults,/skillto show current. Persisted across resume. - Model-driven loading: The
skillbuilt-in tool lets the model discover and activate skills mid-conversation.searchaction finds skills by query (auto-approved);loadaction activates by name (requires user approval since it changes session behavior). Main session only. - Categories: general, engineering, support, custom, mcp
- Content limit: 32 KB per skill (enforced on create/update)
- Storage:
prompt_templatestable (stores skills) with JSONvariablesarray. Migration 010 addstemplatecolumn toscheduled_tasks. - MCP sync: MCP server prompts auto-sync into the
prompt_templatestable withorigin="mcp",mcp_serverset, andreadonly=True. Manual skills take precedence on name collision. MCP-synced content updates resetis_defaultto prevent compromised servers from injecting defaults. Admin UI shows origin badge and disables edit/delete for MCP-sourced skills. - Spec fields: Skills support the full Agent Skills standard frontmatter:
name,description,license,compatibility,metadata(author, version),allowed-tools. Thelicenseandcompatibilityfields are preserved on import and editable in the admin UI. See https://agentskills.io/specification. - Security scanning: Skills are automatically scanned at creation and update
time. The scanner evaluates four risk axes: content risk (command execution,
data exfiltration), supply chain risk (pipe-to-shell, transitive installs),
vulnerability risk (prompt injection, insecure credentials), and declared
capability risk (from
allowed-toolsin SKILL.md). Results populate therisk_level(safe/low/medium/high/critical) andscan_report(JSON breakdown) columns. These fields are system-managed and cannot be overwritten via the admin API. - Discovery: External skills can be discovered and installed from registries:
GET /v1/api/admin/skills/discover?q=...— search the skills.sh registry (or a custom registry viaskills.discovery_urlsetting)POST /v1/api/admin/skills/install— install from skills.sh or GitHub. Fetches theSKILL.mdfile, parses YAML frontmatter, creates a skill withorigin="source"andreadonly=True, stores bundled resources.- Admin UI: Skills tab has "Installed" / "Discover" pill toggle. Discovery view has search bar, result cards, and "Import from GitHub" modal.
- SDK:
discover_skills(q)andinstall_skill(source, skill_id=..., url=...)on both Python and TypeScript console clients.
- Runtime config on installed skills: Installed (readonly) skills can have
their runtime configuration edited — model, temperature, reasoning effort,
token budget, max tokens, agent max turns, auto-approve, allowed tools,
and enabled flag. The server restricts updates to these fields only via
_SKILL_RUNTIME_CONFIG_FIELDSfiltering; spec/content fields (name, description, tags, license, compatibility, content, activation) remain immutable. The admin UI shows "Save Config" instead of "Save" for these skills. Audit action:skill.update.config. - Admin UI: Create/Edit skill modals use a two-column spec manifest layout (left: Identity / Manifest / Deployment; right: Skill Content editor with monospace font). Runtime Config is a collapsible 3-column grid below. License uses an SPDX identifier dropdown (MIT, Apache-2.0, GPL-3.0, etc.). Installed skills show a cyan origin badge with source URL, spec fields are disabled, and all collapsible sections auto-expand in view mode.
Usage Tracking
Per-LLM-request token and tool call metrics:
- Recording:
on_status()inWebUIrecords ausage_eventafter each LLM response with prompt/completion tokens, cache tokens, tool call count, model, ws_id - Prompt caching: Anthropic automatic caching (
cache_control: ephemeral) and OpenAI extended retention (prompt_cache_retention: 24hfor GPT-5.x) are enabled by default.cache_creation_tokensandcache_read_tokensare tracked per request inusage_eventsand surfaced in the Usage admin tab - Querying:
GET /v1/api/admin/usagewithgroup_by(day/hour/model/user) and time range filtering — includes cache token aggregates - Prometheus:
turnstone_tokens_total{type="cache_creation|cache_read"}counters on/metrics - Pruning:
prune_usage_events(retention_days=90)andprune_audit_events(retention_days=365)run automatically via the console scheduler's periodic cleanup cycle
Audit Logging
Append-only trail of admin actions:
- Recording:
record_audit()helper called from all admin mutation handlers - Events captured: user.create, user.delete, token.create, token.revoke, channel.link, channel.unlink, role.create, role.update, role.delete, role.assign, role.unassign, policy.create, policy.update, policy.delete, template.create, template.update, template.delete, skill.create, skill.update, skill.delete, org.update
- Querying:
GET /v1/api/admin/auditwith action/user/time filters + pagination
Database Schema
Migration 008 adds 7 tables:
| Table | Purpose |
|---|---|
orgs |
Organizations (single default org for now) |
roles |
Named permission bundles (3 builtin + custom) |
user_roles |
User-to-role assignments (composite PK) |
tool_policies |
Per-tool approve/deny/ask rules |
prompt_templates |
Reusable system message skills |
usage_events |
Per-request token/tool/cache metrics |
audit_events |
Admin action log |
Also adds org_id column to users table.
API Endpoints
All under /v1/api/admin/ (requires approve scope + granular permission).
| Group | Endpoints | Permission |
|---|---|---|
| Users / Tokens / Channels | 9 (CRUD) | admin.users |
| Roles | 7 (CRUD + assignment) | admin.roles / admin.users |
| Orgs | 3 (list, get, update) | admin.orgs |
| Tool Policies | 4 (CRUD) | admin.policies |
| Skills | 4 (CRUD) | admin.skills |
| Schedules | 6 (CRUD + runs) | admin.schedules |
| Watches | 3 (list, create, cancel) | admin.watches |
| Usage | 1 (aggregated query) | admin.usage |
| Audit | 1 (paginated, filtered) | admin.audit |
Full OpenAPI spec at /openapi.json and Swagger UI at /docs.
Admin Console UI
Governance-related tabs within the 18-tab admin panel:
- Roles — CRUD roles, permission checkbox grid, user role assignment modal
- Policies — CRUD tool policies with colored action badges (green/red/amber)
- Prompts — Prompt-policy editor (heuristics for admin guardrails)
- Skills — CRUD skills with wide modal, textarea editor; Discover pill for installing from skills.sh / GitHub; per-row scan badges (safe/low/med/high/critical)
- Judge — Intent validation configuration and verdict history
- Usage — Summary readouts + CSS bar chart, time range + group-by selectors
- Audit — Filterable log with relative timestamps, load-more pagination
Tabs are permission-gated: hidden if the user lacks the required permission. See docs/console.md for the full tab list and docs/settings.md for the Settings tab that edits live ConfigStore values.
SDK
Both Python and TypeScript console SDKs expose governance methods:
Python (TurnstoneConsole / AsyncTurnstoneConsole):
list_roles(),create_role(),update_role(),delete_role()list_user_roles(),assign_role(),unassign_role()list_orgs(),get_org(),update_org()list_policies(),create_policy(),update_policy(),delete_policy()list_templates(),create_template(),update_template(),delete_template()get_usage(since, group_by=...),get_audit(action=..., limit=...)
TypeScript (TurnstoneConsole):
- Same methods with camelCase naming and typed interfaces
Security Considerations
- Privilege escalation prevented:
admin_assign_roleblocks self-assignment and requires caller to hold a superset of the target role's permissions - Permission validation: Role create/update validates permissions against
a 15-item allowlist (
_VALID_PERMISSIONS) - Self-deletion blocked:
admin_delete_userrejects attempts to delete your own account (matching the self-assignment guard on role endpoints) - Field allowlists: Storage
update_*methods filter fields against allowlists (_ROLE_MUTABLE,_POLICY_MUTABLE, etc.) — handler bugs cannot overwriterole_id,builtin,created, or other protected columns - Bootstrap safety:
handle_auth_setupfails and rolls back if admin role assignment fails, preventing locked-out first user - API token RBAC:
_authenticate_api_tokenloads permissions from user's roles, ensuring API tokens are subject to RBAC enforcement - Policy evaluation is fail-open: If storage is unavailable, tool policies degrade to the existing approval flow (not auto-approve)
- Audit IP resolution:
_audit_context()prefersX-Forwarded-Forfor client IP when behind a reverse proxy, falling back torequest.client.host