mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
9826ea15c5
* feat(coordinator): phase 7 — governance + skill metadata + cross-cutting invariants
Combines three stacked sub-PRs into a single coordinator phase-7
shipment against the phase-7 plan doc. The sub-PR structure (0 / A /
B) preserved on individual branches for reviewer drill-down; this
branch is the one reviewers should merge.
## Sub-PR 0 — service-auth boundary invariants
Shared helpers and contracts that lock the console ↔ node service-auth
boundary so later authz surfaces use them by construction.
- ``_effective_user_filter(request)`` in both ``turnstone.console.server``
and ``turnstone.server`` with a shared ``DENY_EMPTY_SUB`` sentinel
on ``turnstone.core.auth``. Three-way return — admin/service
bypass, scoped caller uid, or fail-closed sentinel on blank sub.
Four callsite migrations (``_coordinator_rows``,
``coordinator_children``, ``coordinator_metrics``,
``cluster_ws_live_bulk``).
- ``StorageBackend`` class docstring codifies the tenancy contract
(every list/count/aggregate method must accept ``user_id: str |
None = None`` and push ``WHERE user_id = :user_id`` into SQL) and
the ``_mapping`` row-access contract. New
``turnstone.testing.row_contract`` ships ``assert_row_like()``.
- ``_verify_collector_service_scope`` probes an upstream node at boot
with ``expected_node_id=_scope-probe_``; a 409 proves the scope
gate was passed, a 403/401 sets ``collector_scope_error`` and
causes ``cluster_snapshot`` / ``cluster_events_sse`` to return 503
with a remediation hint. Probe URL allowlist rejects non-http(s)
schemes and 169.254.0.0/16 hosts.
- 4xx log-level floor on ``_NodeDashboardCache.get``,
``_fetch_live_block``, and ``_proxy_sse`` — dotted-hierarchy
prefixes with bounded body previews. ``_bounded_body_preview`` and
``_bounded_stream_preview`` strip control chars.
## Sub-PR A — coordinator governance core
Mid-session governance surface for coordinator workstreams.
- **Trusted-session mode.** New ``coordinator.trust.send``
permission (migration 042). ``ChatSession.set_trust_send`` /
``revoke_tools`` methods with a ``_governance_lock``. ``POST
/v1/api/coordinator/{ws_id}/trust {send: bool}`` double-gated on
``admin.coordinator`` AND ``coordinator.trust.send`` with
``allow_service_bypass=False`` so service tokens can't escalate.
``_prepare_send_to_workstream`` auto-approves sends whose target is
in the coordinator's own subtree; foreign ws_ids still require
approval. ``_is_own_subtree`` checks both ``parent_ws_id`` AND
``user_id`` to defend against cross-tenant row corruption.
- **Audit-layer credential redaction.** ``record_audit`` walks
``detail`` (dicts, lists, tuples, sets, frozensets; keys too)
and routes every string through ``redact_credentials`` + a C0
control-char scrub. New kw-only ``raw_detail=True`` opt-out.
``_has_any_string`` fast-path. Audit action registry extended
with the four new governance sub-prefixes.
- **Mid-session revocation + cascading stop.** ``POST
/v1/api/coordinator/{ws_id}/restrict {revoke: [...]}`` caps 256
entries / 128 chars; ``_prepare_tool`` short-circuits with a
tool-error. ``POST /v1/api/coordinator/{ws_id}/stop_cascade``
cancels the coord's in-flight generation then dispatches
``cancel_workstream`` for every direct child in parallel via
``asyncio.gather`` bounded by ``Semaphore(16)``. Per-child
outcomes split into ``cancelled`` / ``failed`` / ``skipped``
(404 = already-gone rather than dispatch-broken). Both endpoints
apply ``allow_service_bypass=False`` on the admin gate.
- **Shared plumbing.** ``_resolve_coord_session`` helper collapses
the handler prelude three endpoints shared. ``_emit_coord_audit``
wraps ``record_audit`` in a dedicated ``ThreadPoolExecutor``
(``app.state.audit_executor``) so audit bursts don't starve cancel
dispatches. ``_require_json_object`` guards body parsing so non-
object JSON returns 400 instead of 500.
## Sub-PR B — skill metadata governance
- **Description validator (migration 043).** ``prompt_templates``
rows now require a non-empty ``description``. Existing empty rows
get backfilled with a ``"Skill: <name>"`` placeholder on upgrade.
The installer (``admin_skill_discover``) and MCP prompt sync both
synthesise a placeholder when the upstream description is blank
so non-admin write paths satisfy the invariant.
- **Skill kind classifier (migration 044).** New
``prompt_templates.kind`` column (``interactive`` / ``coordinator``
/ ``any``; defaults to ``any``). New
``turnstone.core.skill_kind.SkillKind`` StrEnum is the single
source of truth; Pydantic schemas type ``kind`` as ``SkillKind``
(OpenAPI advertises the enum) and the handler validator catches
the ValueError. ``list_skills_filtered`` gains a
``kinds: list[str] | None = None`` SQL filter.
``CoordinatorClient.list_skills`` defaults to
``kinds=["coordinator", "any"]`` so interactive-only skills are
hidden from the orchestrator.
- **``scan_status`` → ``risk_level`` rename (migration 045).**
Lossless column rename to align with ``IntentVerdict.risk_level``
terminology. Swept storage (both backends + schema + protocol),
handlers, API schemas, tool JSON, generated OpenAPI specs,
TypeScript SDK types, frontend (``governance.js``), tests, and
English prose in ``docs/judge.md`` + ``docs/tools.md``. The
user-facing on-load warning now reads ``has risk level:
{risk_tier}``. Tool JSON's ``risk_level`` enum corrected to the
scanner's actual taxonomy (``safe / low / medium / high /
critical``; was the never-shipped ``clean / flagged / unscanned /
pending``). Historical migration 021 left untouched.
## Migrations
042 (``coordinator.trust.send`` perm — PR A)
043 (description backfill — PR B)
044 (``kind`` column add — PR B)
045 (``scan_status`` → ``risk_level`` rename — PR B)
All four use position-anchored permission strings / host-side
parse-filter-rejoin on downgrade where SQL ``REPLACE`` could
corrupt prefix-overlapping values.
## Verification
- ``ruff check turnstone tests`` clean.
- ``mypy turnstone`` clean on 165 source files.
- ``pytest -m "not live"``: 4431 passed (+85 over the phase-6
baseline). Includes +32 tests in ``tests/test_service_auth_boundary.py``
and +38 in ``tests/test_coordinator_governance.py``; shared fixtures
extracted to ``tests/_coord_test_helpers.py``.
- Generated OpenAPI JSON (``sdk/typescript/openapi-{console,server}.json``)
regenerated via ``sdk/typescript/scripts/generate-types.py``; zero
``scan_status`` occurrences remaining outside the historical
migration 021 and the rename migration 045.
## Security reviews
Both reviews flagged by the phase-7 plan (items 1 + 5, plus 0a's
refuse-to-serve gate) ran through the multi-stage ``/review``
pipeline twice per sub-PR; all confirmed findings landed in-branch.
* fixup(phase-7): CI lint + PR #383 review fixups
Addresses the lint CI failure (ruff format) plus 12 findings from the
two automated PR reviewers.
Copilot:
- ``_sqlite.list_installed_skill_urls`` / ``_postgresql.list_installed_skill_urls``
used positional row indexing (``r[0]``/``r[1]``/``r[2]``) while this
same PR's ``StorageBackend`` class docstring forbids it. Switched
both to ``r._mapping["..."]`` access.
- ``list_skills.json`` previously advertised ``risk_level=""`` as a
filter for unscanned skills, but the implementation treats empty
strings as "no filter". Clarified the tool description to say
omit the filter entirely to include unscanned rows, and added an
explicit ``enum`` on the parameter restricting it to the scanner
tiers. ``_prepare_list_skills`` keeps the ``strip() or None``
normalisation — unscanned filtering now has an unambiguous contract.
- ``test_storage_skills_filtered.test_risk_level_filter`` used the
legacy ``clean`` / ``flagged`` values from the pre-rename column.
Rewritten with the scanner's actual taxonomy (``safe`` / ``high``).
github-code-quality (CodeQL):
- ``test_deny_sentinel_is_singleton`` previously asserted
``cs.DENY_EMPTY_SUB is cs.DENY_EMPTY_SUB`` — an identical-expression
comparison. Rewritten as two separate ``from ... import ... as`` aliases
(``FIRST_READ`` / ``SECOND_READ``) so the identity check is between
distinct bindings.
- ``test_restrict_empty_revoke_is_noop_but_audits`` unpacked ``state``
without using it. Renamed to ``_state``.
- Mixed import styles in ``test_service_auth_boundary.py`` — the
file previously used both ``import turnstone.console.server as cs``
and ``from turnstone.console.server import ...`` for the same
module (same story for ``turnstone.core.auth`` and
``turnstone.server``). Consolidated to the ``from X import Y`` style
used elsewhere in the file; the ``_fetch_live_block`` test now
patches via pytest's ``monkeypatch`` fixture instead of a manual
rebind through a module alias.
CI:
- ``ruff format`` reformatted one line in
``tests/test_coordinator_endpoints.py``.
Verification: ruff check + mypy clean (166 files); 4459 non-live
pytest pass.
* fix(tests): swap asyncio marker for anyio in service-auth boundary tests
PR #383 CI caught that the 13 ``@pytest.mark.asyncio`` decorators I
added in ``test_service_auth_boundary.py`` are an off-convention
choice — the rest of the repo uses ``@pytest.mark.anyio`` (148 sites
vs my 13). The CI environment pulls in ``anyio`` but not
``pytest-asyncio``, so every async test in this one file was failing
with "async def functions are not natively supported". It passed
locally by accident — my dev venv happens to have pytest-asyncio
installed ambiently.
Swapped all 13 marker sites to ``@pytest.mark.anyio``. No functional
change; the tests run under the same default asyncio backend anyio
provides.
Verification: ruff + mypy clean (166 files); 4459 non-live pytest
pass.
239 lines
12 KiB
Markdown
239 lines
12 KiB
Markdown
# Governance
|
|
|
|
Turnstone governance provides role-based access control (RBAC), tool execution
|
|
policies, skills, usage tracking, and audit logging for the admin console.
|
|
|
|
## Architecture
|
|
|
|
See [diagram: 19-governance-architecture.puml](diagrams/19-governance-architecture.puml).
|
|
|
|
### RBAC (Roles & Permissions)
|
|
|
|
The permission model has two layers:
|
|
|
|
1. **Scopes** (legacy) — `read`, `write`, `approve`. Checked by `AuthMiddleware`
|
|
on every request based on URL path classification.
|
|
2. **Permissions** (granular) — 15 permission strings checked per-endpoint by
|
|
`require_permission()`.
|
|
|
|
**Built-in roles** (seeded by migration 008):
|
|
|
|
| Role | Permissions |
|
|
|------|-------------|
|
|
| admin | read, write, approve, admin.users, admin.roles, admin.orgs, admin.policies, admin.skills, admin.audit, admin.usage, admin.schedules, admin.watches, tools.approve, workstreams.create, workstreams.close |
|
|
| operator | read, write, workstreams.create, workstreams.close |
|
|
| viewer | read |
|
|
|
|
Custom roles can be created with any subset of the 15 valid permissions.
|
|
|
|
**Auth flow:**
|
|
1. User logs in (password or API token) → `_load_user_permissions()` aggregates
|
|
permissions from all assigned roles
|
|
2. `_permissions_to_scopes()` derives legacy scopes (any `admin.*` → `approve`)
|
|
3. JWT created with both `scopes` and `permissions` claims
|
|
4. Middleware checks scope → handler checks permission via `require_permission()`
|
|
|
|
### Tool Policies
|
|
|
|
Admin-defined rules that control tool execution:
|
|
|
|
- **Pattern matching**: Glob syntax via `fnmatch` (e.g., `bash*`, `file_write`, `*`)
|
|
- **Actions**: `allow` (auto-approve), `deny` (block), `ask` (normal approval flow)
|
|
- **Priority**: Higher priority evaluated first, first match wins
|
|
- **Enforcement**: `evaluate_tool_policies_batch()` called in `WebUI.approve_tools()`
|
|
before the `auto_approve` check
|
|
- **MCP granular policies**: MCP resources and prompts are evaluated using their
|
|
`approval_label` for fine-grained control:
|
|
- Resource reads: `mcp_resource__{uri}` (e.g., `mcp_resource__file:///docs/*` to allow,
|
|
`mcp_resource__*` to deny all)
|
|
- Prompt invocations: `mcp__{server}__{prompt}` (e.g., `mcp__trusted__*` to allow,
|
|
`mcp__*` to require approval for all)
|
|
- Built-in tools continue to use `func_name` for backward compatibility
|
|
|
|
### Skills
|
|
|
|
Admin-curated system message skills injected at workstream startup. Skills also
|
|
include session configuration (model, temperature, auto-approve, token budget,
|
|
etc.) since workstream templates were merged into the skills system in v0.8.0.
|
|
|
|
- **Runtime behavior**: Skills are loaded once at session creation and injected
|
|
into the system message *before* user `instructions`. Skills set the baseline;
|
|
instructions customize per-workstream behavior.
|
|
- **Default skills**: All `is_default=true` skills auto-apply to new
|
|
workstreams, concatenated in alphabetical order by name. Use name prefixes
|
|
(e.g. `01-safety`, `02-style`) to control ordering.
|
|
- **Explicit selection**: `--skill <name>` CLI flag, `skill` field on
|
|
`POST /v1/api/workstreams/new`, console creation modal dropdown, scheduled task
|
|
config, and channel adapter config. An explicit skill *replaces* defaults.
|
|
- **Variables**: Three built-in placeholders resolved at load time:
|
|
`{{model}}` (active model name), `{{ws_id}}` (workstream ID),
|
|
`{{node_id}}` (server node ID). Unrecognized placeholders are kept as-is.
|
|
- **Runtime switching**: `/skill <name>` to switch, `/skill clear` to revert
|
|
to defaults, `/skill` to show current. Persisted across resume.
|
|
- **Model-driven loading**: The `skill` built-in tool lets the model
|
|
discover and activate skills mid-conversation. `search` action finds skills
|
|
by query (auto-approved); `load` action activates by name (requires user
|
|
approval since it changes session behavior). Main session only.
|
|
- **Categories**: general, engineering, support, custom, mcp
|
|
- **Content limit**: 32 KB per skill (enforced on create/update)
|
|
- **Storage**: `prompt_templates` table (stores skills) with JSON `variables`
|
|
array. Migration 010 adds `template` column to `scheduled_tasks`.
|
|
- **MCP sync**: MCP server prompts auto-sync into the `prompt_templates` table
|
|
with `origin="mcp"`, `mcp_server` set, and `readonly=True`. Manual skills take
|
|
precedence on name collision. MCP-synced content updates reset `is_default` to
|
|
prevent compromised servers from injecting defaults. Admin UI shows origin badge
|
|
and disables edit/delete for MCP-sourced skills.
|
|
- **Spec fields**: Skills support the full Agent Skills standard frontmatter:
|
|
`name`, `description`, `license`, `compatibility`, `metadata` (author, version),
|
|
`allowed-tools`. The `license` and `compatibility` fields are preserved on import
|
|
and editable in the admin UI. See https://agentskills.io/specification.
|
|
- **Security scanning**: Skills are automatically scanned at creation and update
|
|
time. The scanner evaluates four risk axes: content risk (command execution,
|
|
data exfiltration), supply chain risk (pipe-to-shell, transitive installs),
|
|
vulnerability risk (prompt injection, insecure credentials), and declared
|
|
capability risk (from `allowed-tools` in SKILL.md). Results populate the `risk_level`
|
|
(safe/low/medium/high/critical) and `scan_report` (JSON breakdown) columns.
|
|
These fields are system-managed and cannot be overwritten via the admin API.
|
|
- **Discovery**: External skills can be discovered and installed from registries:
|
|
- `GET /v1/api/admin/skills/discover?q=...` — search the skills.sh registry
|
|
(or a custom registry via `skills.discovery_url` setting)
|
|
- `POST /v1/api/admin/skills/install` — install from skills.sh or GitHub.
|
|
Fetches the `SKILL.md` file, parses YAML frontmatter, creates a skill with
|
|
`origin="source"` and `readonly=True`, stores bundled resources.
|
|
- Admin UI: Skills tab has "Installed" / "Discover" pill toggle.
|
|
Discovery view has search bar, result cards, and "Import from GitHub" modal.
|
|
- SDK: `discover_skills(q)` and `install_skill(source, skill_id=..., url=...)`
|
|
on both Python and TypeScript console clients.
|
|
- **Runtime config on installed skills**: Installed (readonly) skills can have
|
|
their runtime configuration edited — model, temperature, reasoning effort,
|
|
token budget, max tokens, agent max turns, auto-approve, allowed tools,
|
|
and enabled flag. The server restricts updates to these fields only via
|
|
`_SKILL_RUNTIME_CONFIG_FIELDS` filtering; spec/content fields (name,
|
|
description, tags, license, compatibility, content, activation) remain
|
|
immutable. The admin UI shows "Save Config" instead of "Save" for these
|
|
skills. Audit action: `skill.update.config`.
|
|
- **Admin UI**: Create/Edit skill modals use a two-column spec manifest layout
|
|
(left: Identity / Manifest / Deployment; right: Skill Content editor with
|
|
monospace font). Runtime Config is a collapsible 3-column grid below.
|
|
License uses an SPDX identifier dropdown (MIT, Apache-2.0, GPL-3.0, etc.).
|
|
Installed skills show a cyan origin badge with source URL, spec fields are
|
|
disabled, and all collapsible sections auto-expand in view mode.
|
|
|
|
### Usage Tracking
|
|
|
|
Per-LLM-request token and tool call metrics:
|
|
|
|
- **Recording**: `on_status()` in `WebUI` records a `usage_event` after each
|
|
LLM response with prompt/completion tokens, cache tokens, tool call count,
|
|
model, ws_id
|
|
- **Prompt caching**: Anthropic automatic caching (`cache_control: ephemeral`)
|
|
and OpenAI extended retention (`prompt_cache_retention: 24h` for GPT-5.x)
|
|
are enabled by default. `cache_creation_tokens` and `cache_read_tokens` are
|
|
tracked per request in `usage_events` and surfaced in the Usage admin tab
|
|
- **Querying**: `GET /v1/api/admin/usage` with `group_by` (day/hour/model/user)
|
|
and time range filtering — includes cache token aggregates
|
|
- **Prometheus**: `turnstone_tokens_total{type="cache_creation|cache_read"}`
|
|
counters on `/metrics`
|
|
- **Pruning**: `prune_usage_events(retention_days=90)` and
|
|
`prune_audit_events(retention_days=365)` run automatically via the
|
|
console scheduler's periodic cleanup cycle
|
|
|
|
### Audit Logging
|
|
|
|
Append-only trail of admin actions:
|
|
|
|
- **Recording**: `record_audit()` helper called from all admin mutation handlers
|
|
- **Events captured**: user.create, user.delete, token.create, token.revoke,
|
|
channel.link, channel.unlink, role.create, role.update, role.delete,
|
|
role.assign, role.unassign, policy.create, policy.update, policy.delete,
|
|
template.create, template.update, template.delete,
|
|
skill.create, skill.update, skill.delete, org.update
|
|
- **Querying**: `GET /v1/api/admin/audit` with action/user/time filters + pagination
|
|
|
|
## Database Schema
|
|
|
|
Migration 008 adds 7 tables:
|
|
|
|
| Table | Purpose |
|
|
|-------|---------|
|
|
| `orgs` | Organizations (single default org for now) |
|
|
| `roles` | Named permission bundles (3 builtin + custom) |
|
|
| `user_roles` | User-to-role assignments (composite PK) |
|
|
| `tool_policies` | Per-tool approve/deny/ask rules |
|
|
| `prompt_templates` | Reusable system message skills |
|
|
| `usage_events` | Per-request token/tool/cache metrics |
|
|
| `audit_events` | Admin action log |
|
|
|
|
Also adds `org_id` column to `users` table.
|
|
|
|
## API Endpoints
|
|
|
|
All under `/v1/api/admin/` (requires `approve` scope + granular permission).
|
|
|
|
| Group | Endpoints | Permission |
|
|
|-------|-----------|------------|
|
|
| Users / Tokens / Channels | 9 (CRUD) | `admin.users` |
|
|
| Roles | 7 (CRUD + assignment) | `admin.roles` / `admin.users` |
|
|
| Orgs | 3 (list, get, update) | `admin.orgs` |
|
|
| Tool Policies | 4 (CRUD) | `admin.policies` |
|
|
| Skills | 4 (CRUD) | `admin.skills` |
|
|
| Schedules | 6 (CRUD + runs) | `admin.schedules` |
|
|
| Watches | 3 (list, create, cancel) | `admin.watches` |
|
|
| Usage | 1 (aggregated query) | `admin.usage` |
|
|
| Audit | 1 (paginated, filtered) | `admin.audit` |
|
|
|
|
Full OpenAPI spec at `/openapi.json` and Swagger UI at `/docs`.
|
|
|
|
## Admin Console UI
|
|
|
|
Governance-related tabs within the 18-tab admin panel:
|
|
|
|
- **Roles** — CRUD roles, permission checkbox grid, user role assignment modal
|
|
- **Policies** — CRUD tool policies with colored action badges (green/red/amber)
|
|
- **Prompts** — Prompt-policy editor (heuristics for admin guardrails)
|
|
- **Skills** — CRUD skills with wide modal, textarea editor; Discover pill for
|
|
installing from skills.sh / GitHub; per-row scan badges (safe/low/med/high/critical)
|
|
- **Judge** — Intent validation configuration and verdict history
|
|
- **Usage** — Summary readouts + CSS bar chart, time range + group-by selectors
|
|
- **Audit** — Filterable log with relative timestamps, load-more pagination
|
|
|
|
Tabs are permission-gated: hidden if the user lacks the required permission.
|
|
See [docs/console.md](console.md) for the full tab list and
|
|
[docs/settings.md](settings.md) for the Settings tab that edits live
|
|
ConfigStore values.
|
|
|
|
## SDK
|
|
|
|
Both Python and TypeScript console SDKs expose governance methods:
|
|
|
|
**Python** (`TurnstoneConsole` / `AsyncTurnstoneConsole`):
|
|
- `list_roles()`, `create_role()`, `update_role()`, `delete_role()`
|
|
- `list_user_roles()`, `assign_role()`, `unassign_role()`
|
|
- `list_orgs()`, `get_org()`, `update_org()`
|
|
- `list_policies()`, `create_policy()`, `update_policy()`, `delete_policy()`
|
|
- `list_templates()`, `create_template()`, `update_template()`, `delete_template()`
|
|
- `get_usage(since, group_by=...)`, `get_audit(action=..., limit=...)`
|
|
|
|
**TypeScript** (`TurnstoneConsole`):
|
|
- Same methods with camelCase naming and typed interfaces
|
|
|
|
## Security Considerations
|
|
|
|
- **Privilege escalation prevented**: `admin_assign_role` blocks self-assignment
|
|
and requires caller to hold a superset of the target role's permissions
|
|
- **Permission validation**: Role create/update validates permissions against
|
|
a 15-item allowlist (`_VALID_PERMISSIONS`)
|
|
- **Self-deletion blocked**: `admin_delete_user` rejects attempts to delete
|
|
your own account (matching the self-assignment guard on role endpoints)
|
|
- **Field allowlists**: Storage `update_*` methods filter fields against
|
|
allowlists (`_ROLE_MUTABLE`, `_POLICY_MUTABLE`, etc.) — handler bugs
|
|
cannot overwrite `role_id`, `builtin`, `created`, or other protected columns
|
|
- **Bootstrap safety**: `handle_auth_setup` fails and rolls back if admin role
|
|
assignment fails, preventing locked-out first user
|
|
- **API token RBAC**: `_authenticate_api_token` loads permissions from user's
|
|
roles, ensuring API tokens are subject to RBAC enforcement
|
|
- **Policy evaluation is fail-open**: If storage is unavailable, tool policies
|
|
degrade to the existing approval flow (not auto-approve)
|
|
- **Audit IP resolution**: `_audit_context()` prefers `X-Forwarded-For` for
|
|
client IP when behind a reverse proxy, falling back to `request.client.host`
|