mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
d068366a61
* feat(rbac): editable builtin role permissions via overlay layer
Adds a ``role_permission_overrides`` table that stores per-(role_id,
permission) grant/revoke deltas, applied on top of the immutable
``roles.permissions`` baseline at permission-load time. Builtin roles
(``builtin-admin/operator/viewer``) become customizable through the
admin Roles UI without losing the "reset to default" guarantee — every
override is auditable and reversible.
Motivating case: ``model.skills.write`` is deliberately default-ungranted
on every role so operators must consciously opt in before a coordinator
session can mutate the skill catalog. Until now there was no UX path to
do that opt-in — the only options were dropping into SQL or running a
fresh migration. The overrides editor closes that gap.
Backend
- Migration 057 + storage methods on both sqlite + postgresql backends
- ``get_user_permissions`` merges baseline ∪ grants − revokes for builtin
rows; custom rows pass through unchanged
- ``GET /v1/api/admin/roles/{id}/effective`` for inspect
- ``PUT /v1/api/admin/roles/{id}/overrides`` for write — admin.roles gated,
audited, validates against ``_VALID_PERMISSIONS``, refuses non-builtin
targets, strips no-op grants/revokes before persisting
- Lockout guard: cannot revoke ``admin.roles`` if doing so would leave
zero users with the permission (returns 409)
- ``coordinator.trust.send`` added to ``_VALID_PERMISSIONS`` — was
seeded into builtin-admin by migration 042 but never registered with
the validator, so the very first round-trip through the editor 400'd
on it. Drift-detection test guards future migrations from recreating
the same gap
Frontend
- Roles tab redesign: chevron + permission-count chip replace the
"..." truncation; expand-on-click drawer groups perms by namespace
with baseline / grant (green +) / revoke (red −) chip variants
- Edit modal opens for builtin rows ("Customize Built-in Role" title);
toggles show baseline-default vs override state; submit diffs against
the rendered toggle universe (not raw baseline) so future taxonomy
drift can't silently strip unknown perms
- "Modified +N/-N" pill on rows with active overrides; "Reset to default"
drawer action clears the override set
- ``_PERMISSION_SECTIONS`` brought up to date with all currently-seeded
perms (admin.coordinator, admin.cluster.inspect, admin.models,
admin.nodes, admin.prompt_policies, conversation.modify,
coordinator.trust.send were missing)
Tests
- 7 storage tests covering set/list/clear/effective + overlay merge into
``get_user_permissions`` for both builtin and custom roles
- 11 endpoint tests covering effective/overrides happy paths, validation,
lockout guard, builtin-only restriction, no-op normalization, list
enrichment
* feat(rbac): enforce workstreams.{create,close} + tools.approve gates
These three permissions were declared in ``_VALID_PERMISSIONS``, seeded
into ``builtin-operator``'s baseline by migration 008/017, surfaced in
the admin Roles UI as toggles, and documented in ``bootstrap.py`` as
the operator role's capabilities — and never enforced anywhere. The
audit that ran out of the overlay PR found zero ``require_permission``
sites for any of them; any authenticated user could create workstreams,
close any workstream, or approve any pending tool regardless of role.
Behaviour change for callers without the perms:
- ``POST /v1/api/workstreams/new`` (node + console proxy variants)
now 403 without ``workstreams.create``
- ``POST /v1/api/workstreams/{ws_id}/close`` (and ``/route/`` proxy)
now 403 without ``workstreams.close``
- ``POST /v1/api/workstreams/{ws_id}/approve`` (and ``/route/`` proxy)
now 403 without ``tools.approve``
The OR-fallback to ``admin.coordinator`` keeps coord sessions spawning
interactive children unblocked without needing operator-style perms.
Service-scoped inter-cluster calls bypass via the existing
``allow_service_bypass`` path on the new ``require_any_permission``
helper. Builtin admin and operator both already carry these perms;
viewer correctly loses workstream create/close/approve (it already
couldn't do those in spirit).
Implementation
- ``require_any_permission`` (core/auth.py) — OR-semantics variant of
``require_permission`` with per-conditional comments documenting the
security policy at the choke point. 403 body names every accepted
perm so operators get an actionable remediation
- ``make_{create,close,approve}_handler`` (core/session_routes.py)
accept ``fallback_permissions: tuple[str, ...]`` — checked only when
``cfg.permission_gate is None`` (interactive case). Coord's
``permission_gate=_require_admin_coordinator`` continues to take
precedence on the coord-config side
- Console-side ``create_workstream`` and ``route_create`` inline the
same OR check before proxying — fail fast on a forbidden request
without burning a cluster round-trip
- ``route_proxy`` adds a verb-scoped gate on ``approve`` and ``close``
only; ``send``/``cancel``/``dequeue``/``command``/``plan`` remain
authenticated-only (pre-existing, out of scope for this audit)
Tests
- New ``TestPermissionGatesOnLifecycle`` (4 tests) in test_server_authz
pinning 403-without-perm + non-403-with-perm at the node lift sites
- New ``TestRouteProxyPermissionGates`` (5 tests) in
test_console_routing_proxy covering 403 paths, OR fallback via
``admin.coordinator``, and that ``send`` remains ungated
- ``_make_jwt`` helpers in test_server_authz, test_close_reason_
persistence, test_server_attachments_on_create updated to embed
operator-shaped perms by default so existing tests continue to
exercise the post-gate logic rather than 403'ing on the new check
Docs
- ``bootstrap.py`` operator role line corrected to list every perm
it actually carries (was missing ``tools.approve`` and
``conversation.modify``)
* fix(rbac): close lockout + escalation gaps in role-overrides editor
Three issues surfaced by /review of the overlay layer and gate uplift —
all in the RBAC/auth surface, treated as zero-days.
**F-1: lockout guard misses the grant-removal path.** PUT-replace
semantics on ``set_role_overrides`` mean an existing grant of
``admin.roles`` (added via override to e.g. builtin-operator) is
silently dropped when the new payload omits it. The previous guard
short-circuited on ``"admin.roles" not in revokes`` and never noticed.
Concrete cluster-bricking scenario: grant admin.roles to operator via
override, unassign builtin-admin, click "Reset to default" on operator
→ all users lose admin.roles, recoverable only via SQL.
The rewritten guard simulates the post-PUT effective set on the target
role directly: if ``(baseline | new_grants) - new_revokes`` lacks
admin.roles AND nobody holds it via another role, refuse the change.
The "via another role" question is answered by one bulk query rather
than the prior O(users × roles) round-trip loop.
**F-3: lockout check blocked the event loop on moderate deployments.**
The prior check called ``storage.list_user_roles`` per user and
``storage.effective_role_permissions`` per (user, role) pair —
synchronous SQL inside an async handler. 200 users × 5 roles = 1000
connection cycles long enough to trip reverse-proxy timeouts on a
permission revoke.
Replaced with ``storage.users_with_permission(perm, *,
exclude_role_id)`` — one join over ``user_roles ⋈ roles`` plus one IN
fetch on overrides for the builtin role ids in the result, folded
in-process. Two queries total, independent of cluster size. The whole
check now runs under ``asyncio.to_thread`` so even the bulk read
doesn't stall the loop.
**F-2 reframed: admin_assign_role's subset check ignored the overlay.**
The check at lines 6321-6328 reads ``target_role.get("permissions",
"")`` (baseline column) when computing the perms it requires the
caller to hold. After this branch, an admin.roles holder can grant
e.g. ``model.skills.write`` to builtin-operator via override; an
admin.users holder (who happens to NOT hold that perm) could then
assign operator to a new user, silently escalating the assignee. The
existing two-person-rule by perm split (admin.roles for catalog edits,
admin.users for assignments) only holds if the assignment-time check
considers the overlay. Switched ``target_perms`` to
``storage.effective_role_permissions(role_id)["effective"]``.
Note: this PR retains the existing model where admin.roles is the
catalog-edit superuser (admin_create_role, admin_update_role, and now
admin_role_overrides all skip the caller-holds-grants check). The
two-person rule against escalation lives at the assignment gate, which
this fix reinforces.
**F-7: delete_role left orphaned override rows.** No FK on
``role_permission_overrides.role_id`` (migration 057 omitted FKs to
match the rest of the governance schema). Added explicit cleanup in
both sqlite + postgresql ``delete_role`` implementations so a
re-seeded role_id (deterministic for builtins on schema reseed) can't
silently inherit stale overrides from the prior occupant.
Tests
- storage: ``test_users_with_permission_bulk`` exercises the new bulk
helper including ``exclude_role_id`` and overlay folding
- storage: ``test_delete_role_cleans_up_overrides`` pins the F-7 fix
- endpoint: ``test_overrides_lockout_guard_blocks_grant_removal`` is
the F-1 reproduction — operator-overlay grants admin.roles, builtin-
admin has it removed, attempting to reset operator's overrides 409s
- endpoint: ``test_assign_role_blocks_escalation_via_overlay_grant``
pins the F-2 reframed fix — overlay-poisoned operator can't be
assigned by a caller missing the overlay perms
* refactor(rbac): cleanup batch from /review (#584)
Five non-security findings folded into one commit so the security
batch stays focused. All consistent with the existing intent of
``feat/builtin-role-overrides``.
**F-4: presence check on ``_effectivePerms``.** ``governance.js`` was
guarding on ``Array.isArray(role.effective) && role.effective.length > 0``,
falling through to splitting ``role.permissions`` (the baseline) when
the array was empty. For a builtin role whose overrides legitimately
revoke every baseline perm, that path silently rendered the baseline
chips with no override indicators — the inspector lied about what the
role can do. ``_enrich_role`` always sets ``effective: []``, so
presence is the right sentinel.
**F-5: JS-side drift detector.** Commit 1 added a Python-side test
asserting ``_VALID_PERMISSIONS`` covers every baseline perm; the
mirror invariant on the frontend went uncaught. A new perm added to
``_VALID_PERMISSIONS`` without a matching entry in
``_PERMISSION_SECTIONS`` becomes silently un-customizable through the
admin UI (the only documented grant/revoke path). Test parses the
JS const out via regex and asserts set-equality both directions —
detects "missing in UI" and "extra in UI" so the toggle catalog and
validator can't fork.
**F-6: bulk enrich for ``admin_list_roles``.** Was ``1 +
2*builtin_count + 1*custom_count`` SELECTs per admin-tab open;
collapsed to one ``IN``-filtered query via new
``storage.effective_role_permissions_bulk(role_ids)``. Implemented
on both sqlite + postgresql backends following the existing
``effective_role_permissions`` shape.
**F-8: rename ``fallback_permissions`` → ``accepted_permissions``.**
The lift body uses ``if cfg.permission_gate / elif accepted_permissions``
— mutually exclusive — so when ``permission_gate`` is None this IS
the primary gate, not a fallback to anything. The "fallback" name
suggested a tier-2-after-tier-1 semantic that didn't exist. Renamed
across ``make_{approve,close,create}_handler`` factories, the three
call sites in ``turnstone/server.py``, and the docstrings.
**F-9: positive lift-level tests for ``admin.coordinator``-only.**
``TestPermissionGatesOnLifecycle`` previously had a single positive
test for ``workstreams.create`` alone, plus negative-403 tests for
each verb without perms. The OR-fallback to ``admin.coordinator``
(which keeps coord sessions spawning interactive children unblocked)
had no positive coverage at the lift code path — only at the proxy,
which exercises a different verb-dict gate. Added three tests
(create / close / approve) that pass ``admin.coordinator`` alone and
assert non-403, so a future tightening of the accepted_permissions
tuple can't silently regress coord-driven child workstreams.
Out of scope: nit perf-4 (event-delegation refactor on
``_renderGovRoles``). ``setSafeHtml`` rebuild is the existing
pattern across every admin tab; rewriting one tab's render path on
this branch would be drive-by inconsistent with the surrounding
codebase. Filed as a separate concern if the Roles tab grows past
the scale where it bites.
* fix(rbac-ui): aria-expanded + row-click on Roles drawer (#585)
Two Copilot review findings on governance.js:
- Expand button was missing aria-expanded — screen readers couldn't
announce drawer state. Now reflects the row's expanded flag.
- Comment said "row + chevron both work" but only the chevron was
wired. Added data-expand-role to the row element too so the
existing handler loop (querySelectorAll on the attribute) picks up
both — clicking anywhere in the role row toggles the drawer.
Edit/Delete handlers already stopPropagation so they aren't
triggered by the row-level click.
* fix(migrations): rebase role_permission_overrides to 058
PR #560 mitigation #1 landed 057_output_assessments_llm_judge.py on
main in parallel; my migration claimed the same number, forking
alembic's head and breaking postgres. Renumbered to 058 and
re-pointed down_revision at 057 so the chain stays linear.
No behaviour change — same DDL. Full sweep clean (6730 passed).
* fix(migrations): update 058 revision strings to match filename
Previous commit (ea86aefc) renamed 057_role_permission_overrides.py to
058_* but the in-file revision = "057" / down_revision = "056"
strings stayed — leftover from when the file shipped as 057. Tests
pass because alembic walks the chain by revision string, and the
strings now correctly read revision = "058" / down_revision = "057"
to make the chain linear with main's 057_output_assessments_llm_judge.
Caught locally before re-running CI; my prior `git mv` + content edit
landed as a staged rename + unstaged modification on the previous
push.
204 lines
7.1 KiB
Python
204 lines
7.1 KiB
Python
"""Server-side tests for the close_workstream handler's close_reason
|
|
persistence — guards the seam that lets coordinator inspect surface
|
|
why a workstream was retired without scraping the audit log.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import queue
|
|
import threading
|
|
from typing import Any
|
|
from unittest.mock import MagicMock
|
|
|
|
import pytest
|
|
from starlette.testclient import TestClient
|
|
|
|
import turnstone.server as srv_mod
|
|
from turnstone.core.auth import JWT_AUD_SERVER, create_jwt
|
|
from turnstone.core.metrics import MetricsCollector
|
|
from turnstone.core.storage._sqlite import SQLiteBackend
|
|
from turnstone.core.workstream import WorkstreamState
|
|
|
|
_JWT_SECRET = "test-jwt-secret-minimum-32-chars!"
|
|
|
|
|
|
def _full_hdr() -> dict[str, str]:
|
|
# ``workstreams.close`` is now a real gate on the close handler
|
|
# (was a vestigial perm, see PR adding 057_role_permission_overrides);
|
|
# tests that drive close need it embedded in the JWT.
|
|
return {
|
|
"Authorization": (
|
|
f"Bearer {create_jwt('u1', frozenset({'read', 'write', 'approve'}), 'test', _JWT_SECRET, audience=JWT_AUD_SERVER, permissions=frozenset({'workstreams.close'}))}"
|
|
)
|
|
}
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _isolate_metrics(monkeypatch):
|
|
"""Swap ``turnstone.server._metrics`` for a fresh collector
|
|
per-test, with auto-restore.
|
|
|
|
Bare ``srv_mod._metrics = MetricsCollector()`` (the prior
|
|
pattern) leaks into any test file that already bound the name
|
|
via ``from turnstone.server import _metrics`` at import time —
|
|
those tests' patches then operate on a different instance from
|
|
the one the live ``_publish_models_metadata`` reads, and the
|
|
monkeypatch silently no-ops. ``monkeypatch.setattr`` restores
|
|
after the test, so the leak is contained.
|
|
"""
|
|
fresh = MetricsCollector()
|
|
fresh.model = "test-model"
|
|
monkeypatch.setattr(srv_mod, "_metrics", fresh)
|
|
|
|
|
|
def _make_app(storage: Any) -> TestClient:
|
|
mock_session = MagicMock()
|
|
mock_ws = MagicMock()
|
|
mock_ws.id = "ws-target"
|
|
mock_ws.name = "test"
|
|
mock_ws.state = WorkstreamState.IDLE
|
|
mock_ws.session = mock_session
|
|
# Tenant gate (#375) checks ws.user_id == JWT subject; explicit set
|
|
# so MagicMock's auto-generated truthy attribute doesn't reject the
|
|
# request before the persistence path runs. kind / parent_ws_id
|
|
# land in the audit_detail dict alongside ``reason``.
|
|
mock_ws.user_id = "u1"
|
|
mock_ws.kind = "interactive"
|
|
mock_ws.parent_ws_id = None
|
|
mock_mgr = MagicMock()
|
|
mock_mgr.get.return_value = mock_ws
|
|
mock_mgr.close.return_value = True
|
|
mock_mgr.list_all.return_value = [mock_ws]
|
|
mock_mgr.max_active = 10
|
|
|
|
app = srv_mod.create_app(
|
|
workstreams=mock_mgr,
|
|
global_queue=queue.Queue(),
|
|
global_listeners=[],
|
|
global_listeners_lock=threading.Lock(),
|
|
skip_permissions=False,
|
|
jwt_secret=_JWT_SECRET,
|
|
auth_storage=storage,
|
|
cors_origins=["*"],
|
|
)
|
|
return TestClient(app, raise_server_exceptions=False)
|
|
|
|
|
|
@pytest.fixture
|
|
def storage(tmp_path):
|
|
return SQLiteBackend(str(tmp_path / "close.db"))
|
|
|
|
|
|
def test_close_with_reason_persists_to_workstream_config(storage):
|
|
client = _make_app(storage)
|
|
resp = client.post(
|
|
"/v1/api/workstreams/ws-target/close",
|
|
json={"reason": "task complete"},
|
|
headers=_full_hdr(),
|
|
)
|
|
assert resp.status_code == 200
|
|
cfg = storage.load_workstream_config("ws-target")
|
|
assert cfg.get("close_reason") == "task complete"
|
|
|
|
|
|
def test_close_without_reason_does_not_touch_config(storage):
|
|
client = _make_app(storage)
|
|
resp = client.post(
|
|
"/v1/api/workstreams/ws-target/close",
|
|
json={},
|
|
headers=_full_hdr(),
|
|
)
|
|
assert resp.status_code == 200
|
|
cfg = storage.load_workstream_config("ws-target")
|
|
assert "close_reason" not in cfg
|
|
|
|
|
|
def test_close_reason_capped_at_512_bytes(storage):
|
|
"""A model that dumps a multi-KB blob (or a captured secret) into the
|
|
close reason must not be able to grow the workstream_config row
|
|
without bound — the handler enforces a 512-byte ceiling. Tested
|
|
with ASCII (1B/char) so the byte cap and char count coincide."""
|
|
huge = "x" * 5000
|
|
client = _make_app(storage)
|
|
resp = client.post(
|
|
"/v1/api/workstreams/ws-target/close",
|
|
json={"reason": huge},
|
|
headers=_full_hdr(),
|
|
)
|
|
assert resp.status_code == 200
|
|
cfg = storage.load_workstream_config("ws-target")
|
|
stored = cfg.get("close_reason")
|
|
assert stored is not None
|
|
assert len(stored.encode("utf-8")) <= 512
|
|
|
|
|
|
def test_close_reason_byte_cap_holds_for_multibyte_utf8(storage):
|
|
"""Repro for the char-cap-vs-byte-cap mismatch: a CJK-only payload
|
|
of 600 chars would have leaked through a code-point slice at
|
|
600*3=1800 bytes. The byte-aware cap holds it at <=512 bytes."""
|
|
huge = "\u6f22" * 600 # 3 bytes/char in UTF-8
|
|
client = _make_app(storage)
|
|
resp = client.post(
|
|
"/v1/api/workstreams/ws-target/close",
|
|
json={"reason": huge},
|
|
headers=_full_hdr(),
|
|
)
|
|
assert resp.status_code == 200
|
|
cfg = storage.load_workstream_config("ws-target")
|
|
stored = cfg.get("close_reason")
|
|
assert stored is not None
|
|
assert len(stored.encode("utf-8")) <= 512
|
|
|
|
|
|
def test_close_with_non_string_reason_drops_silently(storage):
|
|
"""A malformed body (reason=dict / list / int) should not crash the
|
|
handler — non-string reasons are coerced to empty and the close
|
|
proceeds without writing to workstream_config."""
|
|
client = _make_app(storage)
|
|
resp = client.post(
|
|
"/v1/api/workstreams/ws-target/close",
|
|
json={"reason": {"unexpected": "shape"}},
|
|
headers=_full_hdr(),
|
|
)
|
|
assert resp.status_code == 200
|
|
cfg = storage.load_workstream_config("ws-target")
|
|
assert "close_reason" not in cfg
|
|
|
|
|
|
def test_close_reason_redacts_credentials(storage):
|
|
"""A model under prompt injection that captures a secret and stuffs
|
|
it into ``reason`` must not get to plant the plaintext secret in
|
|
audit logs / workstream_config. The output guard's credential-
|
|
redaction pass runs at the close handler boundary."""
|
|
client = _make_app(storage)
|
|
secret = "AKIAIOSFODNN7EXAMPLE" # AWS access key — output guard catches.
|
|
resp = client.post(
|
|
"/v1/api/workstreams/ws-target/close",
|
|
json={"reason": f"task done; key={secret}"},
|
|
headers=_full_hdr(),
|
|
)
|
|
assert resp.status_code == 200
|
|
cfg = storage.load_workstream_config("ws-target")
|
|
stored = cfg.get("close_reason")
|
|
assert stored is not None
|
|
assert secret not in stored
|
|
assert "[REDACTED:" in stored
|
|
|
|
|
|
def test_close_reason_persistence_failure_does_not_block_close(storage):
|
|
"""If the storage save raises, the close still succeeds — persistence
|
|
is best-effort; a transient storage error must not block the user
|
|
from closing a workstream."""
|
|
client = _make_app(storage)
|
|
|
|
def _boom(*args, **kwargs):
|
|
raise RuntimeError("storage down")
|
|
|
|
storage.save_workstream_config = _boom # type: ignore[method-assign]
|
|
resp = client.post(
|
|
"/v1/api/workstreams/ws-target/close",
|
|
json={"reason": "task complete"},
|
|
headers=_full_hdr(),
|
|
)
|
|
assert resp.status_code == 200
|