Compare commits

..

17 Commits

Author SHA1 Message Date
Patrick Buckley 019d13d930 chore: bump version to 1.6.2 2026-06-11 20:53:13 -07:00
Patrick Buckley 8fbcbff566 test: zero out the suite's warning noise
121 warnings -> 0. Two upstream deprecations get narrowly-scoped
filterwarnings entries (the mcp streamablehttp_client rename — adoption
deliberately rides the v2 migration since the new entry point's call
shape changes again there; the starlette httpx TestClient notice). The
one real RuntimeWarning is fixed at the source: tests that mock
asyncio.run_coroutine_threadsafe handed real coroutines to a stub that
never awaited them, GC-firing 'coroutine was never awaited' inside
whatever unrelated test ran later (the same cross-test bleed mechanism
as the CI closed-stream spew — per-test filterwarnings markers cannot
catch it, which is why two such markers existed and still leaked). A
shared _dispatch_stub now closes real coroutines before returning the
canned future; the obsolete markers are removed.
2026-06-11 20:44:10 -07:00
Patrick Buckley a27738867f chore: cap mcp <2 ahead of the v2 breaking rewrite
mcp 2.0.0a1 shipped 2026-06-11 (stable targeted ~2026-07-27). v2 removes
streamablehttp_client, changes the transport tuple arity, and renames
mcp.types fields to snake_case — all of which our client imports. The
maintainers' release note asks downstream packages to add an upper
bound now (their worked example is this exact constraint). Floor stays
at 1.27: nothing newer adds anything our surface needs, and the #2147
shutdown busy-loop we wrap remains unfixed at every released version.
Resolution is unchanged (1.27.2); lockfile re-pinned metadata only.
2026-06-11 20:44:10 -07:00
renovate[bot] 1d0be9773f chore(deps): update docker images to v0.11.21 2026-06-11 20:44:10 -07:00
Patrick Buckley ad56e1ec96 fix(providers): require base_url for anthropic-compatible
Copilot review on #661: empty base_url let the SDK fall back to
https://api.anthropic.com, sending compat-shaped requests to the
commercial API. The lane is local-only by definition, and the /v1-strip
edge case already established fail-loudly-over-silent-prod-retarget;
apply the same principle to the empty case. create_client raises an
actionable ValueError; the admin Detect path surfaces it as a clean
error string via probe_model_endpoint's existing handler.
2026-06-11 20:44:10 -07:00
Patrick Buckley 8f0115ee2e feat(providers): anthropic-compatible lane for local /v1/messages servers
Add provider id "anthropic-compatible": the existing AnthropicProvider
pointed at Anthropic-compatible local servers (vLLM /v1/messages),
mirroring the openai/openai-compatible split. Registry-only — configured
via the admin Models tab or [models.*] toml, not exposed on the bare
--provider flag, so the CLI/server prod-URL defaults are unreachable for
the lane and real-Anthropic behavior is untouched.

Lane behavior (live-verified against vLLM 0.22.1rc1 + DeepSeek-V4-Flash):
- Capability defaults replace the Claude static table: token_param
  max_tokens, thinking_mode none, web_search/tool_search/vision off,
  reasoning replay on. vLLM rejects Anthropic server-side tool types
  (tools require input_schema) and ignores the thinking request param,
  so neither is sent; thinking blocks still stream back and round-trip
  through the native lane verbatim.
- Reasoning toggles via server_compat extra_body chat_template_kwargs
  (first-class vLLM request field; request-level keys beat server
  defaults). _build_thinking_and_kwargs forwards non-internal
  extra_params as SDK extra_body; thinking_budget_tokens stays internal.
- No temperature force: thinking_mode none skips the Claude-only
  temperature=1.0 requirement.

Admin UI: provider option + URL placeholder (base_url without /v1 — the
SDK appends /v1/messages); the server-compat section shows only the
extra-body field for the lane. thinking_mode round-trips through the
form dropdown for every provider except anthropic-compatible, where it
stays in the raw capabilities JSON — the edit-load lift and save restore
use the same predicate so stored overrides are never silently dropped.

Docs: architecture.md gains the lane subsection incl. verified quirks
(thinking param dropped by vLLM; stop_sequences cut inside thinking and
report end_turn; usage has no cache fields; images need a multimodal
model; mid-conversation system turns are per-model opt-in).

Negative-tested: removing the _INTERNAL_EXTRA_PARAMS exclusion fails
test_internal_keys_not_leaked; the live test drives a streamed turn with
the chat_template_kwargs toggle and asserts no reasoning deltas.
2026-06-11 20:44:10 -07:00
Patrick Buckley 7aba631201 fix(mcp): close the shutdown drain race + close the owned loop
Review feedback: (1) gating the drain on a main-thread truthiness check
of _background_tasks could skip cancellation when a spawn queued via
call_soon_threadsafe had not reached the set yet — submit whenever the
loop is RUNNING and snapshot on the loop, where FIFO callback order
guarantees earlier-queued spawns have landed; (2) shutdown stopped the
loop thread but never closed the loop or cleared _loop/_thread, leaking
selector resources for embedders that cycle managers — close + clear
when we own the thread and it actually stopped (loud warning when it
does not); unowned loops (tests wiring _loop directly) stay untouched;
(3) the bare await-in-suppress drain loops become
asyncio.gather(return_exceptions=True) in both the shutdown drain and
the test fixture.
2026-06-11 20:44:10 -07:00
Patrick Buckley f3f5e84f2d fix(mcp): track fire-and-forget background tasks; harden loop teardown
The post-reconnect catalog refresh was scheduled as a bare
asyncio.create_task: no strong reference (the task could be GC'd
mid-flight, so the refresh might silently never run) and no exception
retrieval (failures surfaced as "Task exception was never retrieved"
at GC time — in CI, onto an already-closed pytest capture stream, the
"I/O operation on closed file" spew; a suspected contributor to the
flaky 60-minute CI hangs via cross-test loop/task state bleed).

- _spawn_background(coro, label): tracked-task set + done-callback
  that retrieves and logs failures at warning; discard runs LAST so
  set-emptiness means "done AND reported"
- shutdown() drains tracked tasks FIRST, so stack teardown can't race
  an in-flight refresh; same run_coroutine_threadsafe idiom and
  timeouts as the existing close steps
- running_loop_mgr fixture: cancel-pending -> drain -> stop ->
  join(5) with a loud assert -> loop.close() (was stop + silent
  join(2), never closed)
- the false-property test ("swallows refresh failure" — nothing
  swallowed it) now waits for completion and asserts the logged
  warning via the patched module logger (structlog; caplog cannot
  observe it), polling inside the patch context
2026-06-11 20:44:10 -07:00
Patrick Buckley ff1e3e5c1c fix(storage): enforce orphan-ness inside the purge DELETE + chunk IN-lists
Review feedback on the purge's race window: the pre-SELECT re-verify
left a statement-to-statement gap where a concurrent registration could
still lose rows — and the pre-counted refcount release could underflow
when it didn't. Orphan-ness now rides the DELETE itself (correlated
NOT EXISTS) with refcounts released from its RETURNING, so refs are
released for exactly the rows that were deleted. Input is de-duplicated,
IN-lists chunk at the storage layer's 500 convention, and the scan's
per-workstream ref-count loop is now one anti-join pass.
2026-06-11 14:13:13 -07:00
Patrick Buckley f0d7305b28 feat(admin): orphan-conversations maintenance verb — scan + purge
Conversation rows whose workstreams row is gone (historical unregistered
writers; the delete-during-inflight race re-creating rows after
delete_workstream) are invisible cruft that also pins attachment
refcounts. Add a turnstone-admin verb: default = read-only scan report
(ws_id, rows, attachment refs, first/last); --delete [--yes] purges.

- shared find/purge logic in storage/_utils; protocol + both backends
  in lockstep (thin wrappers)
- purge re-verifies orphan-ness in-transaction: a ws_id re-registered
  between scan and purge is skipped, never deleted
- releases the deleted rows' attachment refcounts through the
  delete_workstream GC path and sweeps workstream_config/overrides
- summary reports actual purge results, including the skipped clause
2026-06-11 14:13:13 -07:00
Patrick Buckley 84a545cb21 fix(ui): re-home MCP consent badge on the Manage Connections row (#657)
* fix(ui): re-home MCP consent badge on the Manage Connections row

The L-shell renovation retired the standalone settings gear (#settings-btn).
The MCP pending-consent badge anchored to that gear via _refreshConsentBadge,
which null-guarded silently — so since the renovation pending consent requests
had no indicator (the badge was invisible).

Re-home the badge on the rail's Manage row where the MCP/connections surface
lives in both deployments:

- rail.js gains a generic setRowBadge(tabKey, count, label?) hook + a `badge`
  builder: a small ⚠-glyph + count chip (never colour alone) using the DS warn
  tokens. mountManage registers row + owning-group-head refs and re-applies live
  counts across a (re)mount. When the owning group is collapsed, the count also
  mirrors onto the group head so a hidden row never hides the signal. rail.js
  stays agnostic — it owns the mechanism, the caller owns the meaning.
- shell.js (the ESM bridge) re-exports setRowBadge on window.TS_SHELL so the
  classic ui/static/app.js subsystem can drive it without importing the module.
- The standalone consent subsystem keeps its shell-level ownership: _refresh-
  ConsentBadge now drives setRowBadge on the Connections tab, fed by both the
  loadPendingConsents hydrate/poll load and live onConsentDetected notifications.
- The shared interactive pane host bridges onConsentDetected to the new
  window.TS_APP.onConsentDetected seam (undefined on the console, so the console
  pane stays a no-op there); panes only notify.
- The dead colour-only gear badge CSS (.settings-consent-badge, red dot) is
  removed; the new chip lives in shell.css as token-only .rail-badge so it
  flips themes by construction.

Console MCP tab (Extensions > mcp) and standalone Connections tab
(Extensions > connections) both badge correctly. Pins extended in
test_shell_js.py + test_app_js.py.

* fix(ui): drop the unused head ref from the rail badge row map

Review feedback: _rowEls stored each row's group-head element but every
head consumer resolves it through _groupEls; keeping the duplicate DOM
ref made the remount state shape harder to reason about.
2026-06-11 14:13:13 -07:00
Patrick Buckley 8e11929ba0 chore: bump version to 1.6.1 2026-06-11 14:09:31 -07:00
Patrick Buckley d9e9a41b17 test(console): make dedupe-pin slice bounds reformat-tolerant
Review feedback: the next-case end markers were exact-indentation
string finds that raised a bare ValueError when unmatched. Use
whitespace-tolerant regexes with actionable assertion messages, and
bound the history-replay window structurally (next role branch, with
a generous fallback) instead of a fixed 600 chars.
2026-06-11 14:05:19 -07:00
Patrick Buckley 848b2cc1fb fix(ui): single-path Enter activation + hls.js teardown on player error
Review feedback: (1) the Enter keydown re-dispatched through btn.click(),
relying on the disabled-guard to suppress the browser's own
Enter-to-click — preventDefault + direct activation makes the keyboard
path provably single-fire; (2) the branch-scoped Hls instance was
unreachable from the media error handler, leaking its listeners and
loader timers when the player node was replaced with the retry UI —
hoist the ref and destroy it before replacement.
2026-06-11 14:05:19 -07:00
Patrick Buckley 1946002618 fix(ui): lift media player activation into the shared interactive pane
The interactive Pane renders media embeds (buildMediaEmbed / buildPlayButton),
but the Play activation — _loadHls / _isHlsUrl / _activatePlayer and the
click/keydown delegate — stayed behind in the standalone ui/static/app.js as
DOCUMENT-level listeners. The console L-shell mounts the same interactive.js
module but never loads ui/static/app.js, so the Play button was dead in
console-hosted interactive panes.

Lift the activation into shared_static/interactive.js (alongside the existing
buildMediaEmbed/buildPlayButton — media embeds are interactive-pane-only; the
coordinator pane renders none) and wire it as a pane-owned, root-scoped
this.el click/keydown listener, mirroring the approval-keydown pattern the
fork collapse established. The standalone copy is deleted so no duplicate
implementation remains; both deployments now activate through the one shared
handler.

The hls.js vendor is fetched lazily by absolute /shared/ URL (the same
mechanism renderer.js uses for mermaid), and /shared is mounted at the root in
both turnstone/server.py and turnstone/console/server.py, so the vendor —
which ships in shared_static/hls-1.6.16/ — resolves in both deployments with
no HTML change.

Pins: assert the lift + pane-ownership in test_interactive_pane_js.py and the
standalone-stays-clean guard in test_app_js.py.
2026-06-11 14:05:19 -07:00
Patrick Buckley 4bce6abc7c test(console): pin system-turn dedupe wiring on both read paths
The live-SSE/history system-turn dedupe (renderedSystemEventIds /
_renderedSystemEventIds) was already in place on both panes and merged
to main (21af6c4 aligned the persisted row event_id with its SSE event;
09e41d1 added the belt-and-braces Set on the coordinator). The existing
pin tests only assert the Set's .has()/.add()/.clear() symbols appear
somewhere in the file, so a refactor that keeps the Set but short-circuits
the live-handler consultation (guard -> false) re-opens the double-render
while the pins stay green.

Scope the new assertions to their blocks: the live system_turn case must
CONSULT and RECORD against the Set, and the history render path
(replayHistory / refetchHistory's system-role branch) must record each
replayed row's event_id. Bounded at the next switch case rather than the
first break; the dedup-skip path itself breaks before the .add(), so a
break-bounded slice would drop the record half.

Verified the new slice checks fail on a dedupe-neutered factory (a
headless-Chrome harness driving the real createCoordinatorPane confirms
that neutering produces two rendered nodes for one event id; intact code
renders one, and the no-event-id legacy path still renders both).
2026-06-11 14:05:19 -07:00
Patrick Buckley c19432f12a fix(memory): touch access metadata on composition and tool reads
The touch_structured_memories facade and both storage backends were
implemented but had zero call sites, so access_count never moved and
last_accessed never advanced past write time on any deployment.

Wire two touch points:
- proactive composition touches the injected top-k (post-rerank) set,
  deduped per turn since _init_system_messages recomposes many times
  within a single turn;
- the memory tool's search and get reads touch their returned rows,
  counted per call. save/delete/list do not touch.

Touches are best-effort through the facade, which already swallows
storage errors, so a failed touch never breaks composition or a tool
call.
2026-06-11 14:05:18 -07:00
797 changed files with 28551 additions and 259634 deletions
-3
View File
@@ -42,9 +42,6 @@
# CONSOLE_HTTPS_PORT=8443 # Caddy (dashboard HTTPS)
# POSTGRES_PORT=5432 # exposed for bare-metal host joins
# POSTGRES_BIND=127.0.0.1 # set 0.0.0.0 to let another machine join
# TURNSTONE_HOST_IP=127.0.0.1 # dev-stack bind address for cross-host joins
# TURNSTONE_CONSOLE_HTTP_BIND=127.0.0.1 # production TLS-overlay ACME/API bind
# TURNSTONE_ACME_EXTERNAL_URL=http://192.0.2.1:8090/acme # routable ACME base; bracket IPv6; include /acme
# -- Workspace ----------------------------------------------------------------
# Bind-mount a host directory the model can read/write at /workspace:
-5
View File
@@ -1,5 +0,0 @@
# Funding platforms for the GitHub "Sponsor" button.
# https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/displaying-a-sponsor-button-in-your-repository
github: [eous]
custom: ["https://paypal.me/eousphoros"]
+1 -1
View File
@@ -55,7 +55,7 @@
{
"description": "LLM SDKs — always review manually",
"groupName": "LLM SDKs",
"matchPackageNames": ["openai", "httpx2", "anthropic", "mcp"],
"matchPackageNames": ["openai", "anthropic", "mcp"],
"schedule": ["before 9am on Monday"],
"automerge": false
},
+23 -39
View File
@@ -14,8 +14,8 @@ jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: "3.14"
- run: pip install pre-commit
@@ -25,8 +25,8 @@ jobs:
typecheck:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: "3.14"
- run: pip install mypy
@@ -35,31 +35,23 @@ jobs:
test:
runs-on: ubuntu-latest
# Cap a hung run at 30 min instead of riding GitHub's 6-hour default
# (a flaky-hang run otherwise streams -v output for hours). Was 20;
# the suite's growth (~9.7k tests, coverage-instrumented, 3-version
# matrix) started brushing the old cap on healthy runs.
timeout-minutes: 30
strategy:
matrix:
python-version: ["3.11", "3.12", "3.13"]
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: ${{ matrix.python-version }}
# Node is required by tests/test_renderer_js.py — without
# explicit setup, that suite silently skips if the runner
# image happens not to ship Node, masking regressions in
# the browser-side renderer.
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6
with:
node-version: "24"
- run: pip install -e ".[test]"
# -v lists each test id as it starts (pytest prints the nodeid at
# logstart), so a hang names the culprit on the last line instead of
# riding the job timeout with only a trail of "..." dots.
- run: pytest tests/ -m "not live and not e2e_recovery" --cov=turnstone --cov-report=term-missing --cov-report=xml -v
- run: pytest tests/ -m "not live" --cov=turnstone --cov-report=term-missing --cov-report=xml -q
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
if: always()
with:
@@ -68,7 +60,6 @@ jobs:
test-postgres:
runs-on: ubuntu-latest
timeout-minutes: 30
services:
postgres:
image: postgres:18
@@ -84,23 +75,23 @@ jobs:
--health-timeout=5s
--health-retries=5
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: "3.14"
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6
with:
node-version: "24"
- run: pip install -e ".[test]"
- run: pytest tests/ -m "not live and not e2e_recovery" --storage-backend=postgresql -v
- run: pytest tests/ -m "not live" --storage-backend=postgresql -q
env:
TURNSTONE_TEST_PG_URL: postgresql+psycopg://postgres:postgres@localhost:5432/turnstone_test
wheel-completeness:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: "3.14"
- run: pip install build
@@ -115,16 +106,9 @@ jobs:
| grep -v '\.py$' | grep -v '\.dist-info' | grep -v '\.pyc' | grep -v '^File$' \
| sort)
# Files intentionally excluded from the wheel (one per line).
# The vllm-litellm/ deploy example ships in the repo, not the wheel
# (you clone the repo to run it; the package doesn't reference it).
# Files intentionally excluded from the wheel (one per line)
ALLOW="
turnstone/core/storage/migrations/script.py.mako
turnstone/deploy/vllm-litellm/.env.example
turnstone/deploy/vllm-litellm/README.md
turnstone/deploy/vllm-litellm/docker-compose.yml
turnstone/deploy/vllm-litellm/gemma.Dockerfile
turnstone/deploy/vllm-litellm/litellm-config.yaml
"
MISSING=$(comm -23 <(echo "$SOURCE") <(echo "$WHEEL") \
@@ -148,13 +132,13 @@ jobs:
/tmp/smoke/bin/turnstone-console --help
/tmp/smoke/bin/turnstone-admin --help
/tmp/smoke/bin/turnstone-channel --help
/tmp/smoke/bin/turnstone-doctor --help
/tmp/smoke/bin/turnstone-bootstrap --help
lock-check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: astral-sh/setup-uv@c771a70e6277c0a99b617c7a806ffedaca235ff9 # v9.0.0
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # v8.2.0
with:
uv-version: "0.9.18"
- run: uv lock --check
@@ -162,11 +146,11 @@ jobs:
security:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: astral-sh/setup-uv@c771a70e6277c0a99b617c7a806ffedaca235ff9 # v9.0.0
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: astral-sh/setup-uv@fac544c07dec837d0ccb6301d7b5580bf5edae39 # v8.2.0
with:
uv-version: "0.9.18"
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
with:
python-version: "3.14"
- run: uv sync --frozen --all-extras
@@ -190,8 +174,8 @@ jobs:
run:
working-directory: sdk/typescript
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6
with:
node-version: "24"
- run: npm ci
+6 -17
View File
@@ -7,9 +7,7 @@ on:
concurrency:
group: docker-${{ github.event.workflow_run.head_sha }}
# Never cancel mid-push: an interrupted multi-tag push can leave the
# registry with a partial tag set (e.g. :latest moved, :stable not).
cancel-in-progress: false
cancel-in-progress: true
permissions:
contents: read
@@ -21,24 +19,15 @@ env:
jobs:
docker:
# Same gate as publish.yml: workflow_run fires for every CI completion
# (including fork and same-repo PR runs) with this repo's token and
# packages:write. Only same-repo tag pushes may publish images; CI's
# push trigger matches main/stable/* and v* tags, so a head_branch
# starting with "v" is necessarily a tag run.
if: >-
github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.event == 'push' &&
github.event.workflow_run.head_repository.full_name == github.repository &&
startsWith(github.event.workflow_run.head_branch, 'v')
github.event.workflow_run.head_repository.full_name == github.repository
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
with:
ref: ${{ github.event.workflow_run.head_sha }}
fetch-depth: 0
# The docker build only reads the tree; keep the token out of it.
persist-credentials: false
- name: Resolve release tag
id: tag
@@ -54,7 +43,7 @@ jobs:
- name: Log in to GHCR
if: steps.tag.outputs.skip == 'false'
uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
with:
registry: ${{ env.REGISTRY }}
username: ${{ github.actor }}
@@ -78,12 +67,12 @@ jobs:
fi
echo "tags=${TAGS}" >> "$GITHUB_OUTPUT"
- uses: docker/setup-buildx-action@bb05f3f5519dd87d3ba754cc423b652a5edd6d2c # v4
- uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4
if: steps.tag.outputs.skip == 'false'
- name: Build and push
if: steps.tag.outputs.skip == 'false'
uses: docker/build-push-action@53b7df96c91f9c12dcc8a07bcb9ccacbed38856a # v7
uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7
with:
context: .
push: true
+6 -20
View File
@@ -7,9 +7,7 @@ on:
concurrency:
group: publish-${{ github.event.workflow_run.head_sha }}
# Never cancel a publish mid-upload: a half-uploaded release (sdist up,
# wheel missing) cannot be re-run cleanly because PyPI rejects duplicates.
cancel-in-progress: false
cancel-in-progress: true
permissions:
contents: write
@@ -17,26 +15,14 @@ permissions:
jobs:
publish:
# workflow_run fires for EVERY CI completion — including CI runs for
# pull_requests from forks — and always executes here with this repo's
# secrets, tokens, and the pypi environment. Gate to same-repo tag
# pushes only: CI's push trigger matches branches main/stable/* and
# tags v*, so a head_branch starting with "v" is necessarily a tag run.
if: >-
github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.event == 'push' &&
github.event.workflow_run.head_repository.full_name == github.repository &&
startsWith(github.event.workflow_run.head_branch, 'v')
if: github.event.workflow_run.conclusion == 'success'
runs-on: ubuntu-latest
environment: pypi
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
with:
ref: ${{ github.event.workflow_run.head_sha }}
fetch-depth: 0
# python -m build executes the tree's build backend; don't leave
# the contents:write token sitting in .git/config while it runs.
persist-credentials: false
- name: Resolve release tag
id: tag
@@ -50,7 +36,7 @@ jobs:
echo "skip=false" >> "$GITHUB_OUTPUT"
fi
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6
if: steps.tag.outputs.skip == 'false'
with:
python-version: "3.14"
@@ -58,12 +44,12 @@ jobs:
if: steps.tag.outputs.skip == 'false'
- run: python -m build
if: steps.tag.outputs.skip == 'false'
- uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
- uses: pypa/gh-action-pypi-publish@cef221092ed1bacb1cc03d23a2d87d1d172e277b # release/v1
if: steps.tag.outputs.skip == 'false'
- name: Create GitHub Release
if: steps.tag.outputs.skip == 'false'
uses: softprops/action-gh-release@3d0d9888cb7fd7b750713d6e236d1fcb99157228 # v3
uses: softprops/action-gh-release@b4309332981a82ec1c5618f44dd2e27cc8bfbfda # v3
with:
tag_name: ${{ steps.tag.outputs.tag }}
generate_release_notes: true
-42
View File
@@ -1,42 +0,0 @@
name: Understone example
# The door-game example is a standalone package with no dependency on
# turnstone core, and the root test suite does not collect it
# (testpaths=["tests"]). Without this workflow its suite never runs in CI.
# Path-filtered so it only runs when the example (or this workflow) changes.
on:
push:
branches: [main, "stable/*"]
paths:
- "examples/door-game/**"
- ".github/workflows/understone-example.yml"
pull_request:
branches: [main, "stable/*"]
paths:
- "examples/door-game/**"
- ".github/workflows/understone-example.yml"
permissions:
contents: read
jobs:
understone:
runs-on: ubuntu-latest
defaults:
run:
working-directory: examples/door-game
strategy:
matrix:
# Floor and ceiling of the example's requires-python (>=3.11).
python-version: ["3.11", "3.13"]
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: ${{ matrix.python-version }}
- run: pip install -e ".[test,dev]"
- run: pytest tests/ -q
- run: ruff check .
- run: ruff format --check .
- run: mypy understone/
+5 -26
View File
@@ -25,43 +25,22 @@ permissions:
jobs:
vendor-js:
# Same-repo PRs only: this job checks out the PR head and pushes to it
# with contents:write, so it must never act on a fork's branch.
# Gate on the PR author (immutable), not github.actor (names whoever
# caused the latest event, which can be someone else re-running it).
if: >-
(github.event_name == 'pull_request' &&
github.event.pull_request.user.login == 'renovate[bot]' &&
github.event.pull_request.head.repo.full_name == github.repository) ||
github.event_name == 'workflow_dispatch'
if: github.actor == 'renovate[bot]' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
steps:
- name: Resolve PR head ref
id: ref
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
# Branch names may contain shell metacharacters; pass via env,
# never interpolate ${{ }} into the script body.
HEAD_REF: ${{ github.head_ref }}
PR_NUMBER: ${{ inputs.pr_number }}
run: |
if [[ "$GITHUB_EVENT_NAME" == "workflow_dispatch" ]]; then
# The dispatch input is an arbitrary PR number; refuse fork PRs.
# A fork's headRefName is a bare branch name that may collide
# with a branch in this repo, and checkout+push would then hit
# that unrelated branch ("same-repo PRs only" applies here too).
pr_json=$(gh pr view "$PR_NUMBER" --repo "$GITHUB_REPOSITORY" --json headRefName,isCrossRepository)
if [[ "$(jq -r '.isCrossRepository' <<< "$pr_json")" != "false" ]]; then
echo "::error::PR #${PR_NUMBER} head is not a branch in this repository; refusing to complete it."
exit 1
fi
ref=$(jq -r '.headRefName' <<< "$pr_json")
if [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
ref=$(gh pr view "${{ inputs.pr_number }}" --repo "${{ github.repository }}" --json headRefName -q .headRefName)
else
ref="$HEAD_REF"
ref="${{ github.head_ref }}"
fi
echo "head_ref=${ref}" >> "$GITHUB_OUTPUT"
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
with:
ref: ${{ steps.ref.outputs.head_ref }}
-4
View File
@@ -19,8 +19,6 @@ docker-compose.override.yml
.ruff_cache/
.pytest_cache/
*.db
*.db-shm
*.db-wal
.plan.md
.plan-*.md
.hypothesis/
@@ -30,5 +28,3 @@ tools/skill_audit_analysis/data/
tools/skill_audit_analysis/output/
design_ideas/
.claude/
docs/design/
/.idea
+4 -903
View File
@@ -6,912 +6,13 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [PEP 440](https://peps.python.org/pep-0440/) for
version numbers (`X.Y.Z`, with `X.Y.ZaN` / `bN` / `rcN` for pre-releases).
Two active release tracks are maintained — the current stable and the
experimental line:
Three release tracks are maintained — the current stable, one prior
stable, and the experimental line:
- **`stable/1.7`** — patch-only (`v1.7.x`)
- **`stable/1.5`** — patch-only (`v1.5.x`)
- **`stable/1.6`** — patch-only (`v1.6.x`)
- **`main`** — experimental (next major)
Earlier stable lines (`stable/1.6`, `stable/1.5`) are frozen.
## [Unreleased]
### Added
- **Large plain-text pastes become attachments.** Pasting text longer than the
fixed 2,000-character threshold stages `pasted-text.txt` across all five
attachment-capable create/send composers. Clipboard files retain priority,
text above the 512 KiB upload ceiling stays inline, identical synthesized
pastes collapse to one chip, and rejected attachment sends keep the staged
message and files so they can be corrected or retried.
- **`server_parses_reasoning` model capability.** Declare it on a model
definition whose backend segregates reasoning into its own channel (a
vLLM launched with a reasoning parser, a commercial provider): the
inline think-tag scan turns off on every lane — interactive and
drained alike — so content is trusted verbatim and prose that merely
quotes a tag can no longer be misrouted into the reasoning lane, and
the utility lanes stop suppressing reasoning they'd otherwise pin off.
Default off for local lanes, preserving the passthrough-server
behavior; the built-in capability tables declare it for every real
commercial endpoint (known models and table-miss defaults alike),
which also removes the quoted-tag false positive from those lanes.
- **Per-model Entra gateway authentication.** Model definitions can bind either
a caller-delegated OBO token (`entra_obo`) or a shared app-identity token
(`entra_app`) through the provider SDK credential surface. Mints reuse the
encrypted cluster token cache, refresh-rotation CAS, and advisory locking;
add a host-local memo, failure cooldown, long-lived mint HTTP client, audience
allow-list/permission boundary, identity-unlink purge, and optional
`model.auth_fail_closed` refusal policy. Delegated identity now propagates
through judge, output-guard, and principal-scoped perception lanes, and
unattended watch restoration reacquires the persisted workstream owner.
Ownerless OBO calls and dynamic aliases without a real static fallback always
fail closed; grant modes are never silently switched. Static authentication
remains the default.
- **Compaction is visible now: lifecycle events, a progress bar, and a
persistent transcript card.** Context compaction (manual `/compact` and
auto) emits a first-class `compaction` SSE event
(`start` / `progress` / `end` — see the API reference) instead of loose
info lines. The web UI renders an in-transcript card with a real progress
bar (determinate `part k of N` during chunked summarization, indeterminate
for single-call compactions) that settles into a result card — token delta
plus the summary behind a fold — in both the interactive pane and the
coordinator viewer. The result survives reloads: the persisted compaction
marker now projects through `/history` as a `role="system"`,
`source="compaction"` entry (resume/export/search unchanged), stamped with
the end event's id so repaint and SSE replay can't double-render. The
marker's `meta` additionally records `before_tokens` / `after_tokens` /
`trigger`. Python and TypeScript SDKs gain a typed `CompactionEvent`.
- **One provider transport: every model call now streams (#831).**
The per-adapter non-streaming entry (`create_completion`) is retired;
single-shot lanes — judges, titles, compaction, web-fetch extraction,
perception, eval, optimizer — sample through the same streaming entry
the chat loop uses and accumulate via one shared drain, so request
shaping can no longer drift between the two consumption styles. Two
operator-visible consequences: long single-shot generations (a thinking
model composing a title, a slow local judge) no longer sit in a single
blocking read that can hit client read-timeouts — the same reason the
Anthropic adapter already streamed internally — and judge timeouts now
*abort* the underlying HTTP read instead of abandoning a worker thread
on a dead call. Because every call now streams, an alias pointed at a
model or org that cannot stream (OpenAI's verified-org streaming
entitlement, a gateway api-version predating `stream_options` — e.g.
older Azure OpenAI deployments) fails at request time where 1.7's
non-streaming single-shot call succeeded; remediation is on the
serving side (verify the org, bump the api-version/gateway) — there is
deliberately no per-model non-streaming fallback left to configure. These lanes are also complete-or-error now: a stream
that ends without any finish signal is treated as a generation that
died mid-response and retried, instead of storing the partial text as
a clean result (previously a half-generated compaction summary could
silently replace real history). Caveats: these lanes now carry the
same `stream_options: {include_usage: true}` the chat loop always
sent — OpenAI-compatible servers old enough to *ignore* it stop
producing usage rows on these lanes, and servers strict enough to
*reject* unknown fields (pre-2024 llama.cpp/proxy builds) will 400 —
such a server already couldn't serve turnstone's chat loop, but a
judge/utility alias pointed at one worked on 1.7 and needs to move to
a current server. Transient mid-stream deaths (connection drop, proxy
hiccup) are re-issued in place up to twice with exponential backoff —
the retry the SDK's request loop used to provide these lanes
invisibly. Each lane accepts its own terminal marker (Anthropic
`message_stop`, Responses terminal events); a lax server/gateway that
never sends any terminal signal needs
`{"finish_reason_optional": true}` in the model definition's
capabilities JSON, which restores 1.7's tolerance (clean end-of-stream
after output = completion) for that model on every lane — without it
such streams fail as died-mid-generation, because SSE gives no way to
tell the two apart and the default favors catching truncation. The
unread `supports_streaming` capability flag (and its admin tile) is
gone; the o-series models it described are dropped from the capability
table entirely (see Removed).
- **One turn interface for every model call: `core/model_turn.py` (#827).**
Judges (intent + output guard), perception, title generation, compaction,
web-fetch extraction, the eval harness, the optimizer's meta lanes, and
task agents all advance a trajectory through the same plant-call
primitive the agent seam pioneered — Turn IR in, one shared lowering
(argument sanitize → minted-id restore → vLLM reasoning attach), one
shared re-ingest (blank-id repair → native-lane finalize). The judges'
hand-built OpenAI-dict path is gone, and with it the Gemini judge's
tool-blindness: evidence tools now work on Google models because the
native lane round-trips `thought_signature` (with pairwise repair for
blank-id compat responses). Provider adapters still take lowered wire
dicts — the transport collapse and main-loop migration are tracked as
#831 / #832.
- **task_agent keeps its model's reasoning across its own tool loop — on
every provider lane.** A task agent's replayed turns now carry the
provider-native reasoning lane the model produced — Anthropic thinking
blocks with their signatures (commercial or an anthropic-compatible
server), OpenAI Responses reasoning items, Gemini `thought_signature`
fidelity blocks, and the reasoning text a vLLM `--reasoning-parser` /
llama.cpp `reasoning_format` surfaces on the Chat Completions lane —
instead of each turn being rebuilt from text + tool calls with the
reasoning dropped. On a thinking model this restores reasoning continuity
across the agent's own multi-turn tool use. On the wire the agent's
session-minted sub-tool ids are mapped back to the provider's own ids
(`restore_provider_tool_ids`), so the native block — replayed verbatim,
its signature never touched — the `tool_calls` mirror, and each tool
result always agree; internally the minted ids still key the live card,
recall, and the cancel ledger unchanged. Replay honors the same per-model
`replay_reasoning_to_model` flag the main loop uses on every lane: the
vLLM Chat-Completions field replay keeps its server-type gate, and
llama.cpp stays capture-only, matching main-loop behavior. The native
lane is finalized by the same shared builder as the main loop's, so the
two harnesses cannot drift.
- **Background shells: `bash` gains `run_in_background`, plus `bash_output` /
`kill_shell`.** Setting `run_in_background=true` starts the command as a
detached shell and returns immediately with a `bash_N` handle — "start a dev
server, use it in a later call" is back as an explicit opt-in (the shape
follows the convention the major coding agents converged on). `bash_output`
returns only output produced since the previous read (optionally filtered by
a regex) plus status and exit code; `kill_shell` terminates the shell's
whole process group. Output is buffered per shell with a drop-oldest cap, so
a chatty server can't grow memory unbounded. When a background shell exits,
a system notice lands at the next seam (waking an idle workstream if
needed). Shells survive a generation cancel, die with the workstream, and
never outlive a task_agent that started them; anything a background shell
itself backgrounds is still reaped when that shell exits — the no-leak
guarantee below is unchanged.
### Changed
- **lacme 1.2.0 and HTTPX2 now back the core mTLS path.** TLS/ACME is a typed
core dependency rather than optional-import-era code; renewal and admin
clients use lacme's public close/stop lifecycle. Consoles can set
`TURNSTONE_ACME_EXTERNAL_URL` to a routable responder base ending in `/acme`,
so nodes on another host receive usable directory and follow-up URLs during
enrollment. Signing routes now require a dedicated, purpose-confined rotating
service JWT, and HTTPX2 pins that credential to canonical resources on
configured responder origins. Internal and external console identities use
separate persistence namespaces; cluster identities are reused only after
key, SAN, validity, EKU, and active-root verification. Responder-side keyless
CSR results cannot overwrite managed identities, failed live reloads restore
the last usable bundle, certificate lifetime stays at 48 hours with a
12-hour renewal cadence, and cancellation drains renewal clients before
propagating. Operator-supplied IPv4 and IPv6 literals are passed to lacme as
typed IP identifiers and issued as IP SANs; unexpired legacy certificates
containing `DNS:<ip>` are reissued instead of being reused as an invalid IP
identity
([#1011](https://github.com/turnstonelabs/turnstone/issues/1011)).
- **OpenAI SDK v3 and its HTTPX2 default transport are now supported (#1009).**
Chat Completions and Responses streams normalize native HTTPX2 connection
deaths through the same retry boundary as legacy HTTPX-backed providers,
including failures observed after safe cross-thread client closure during a
model-registry reload. The OpenAI v3 runtime escape hatch for explicitly
injected legacy HTTPX clients remains supported. OpenAI connections now
follow HTTPX2's operating-system trust store by default; deployments that
relied on a modified `certifi` bundle must install that CA in the system
store or set `SSL_CERT_FILE` / `SSL_CERT_DIR`.
- **Log event rename: `drain_stream.post_finish_blip` is now
`stream.post_finish_blip`; its `usage_captured` field is retained.** The
single-shot drain normalizes mid-body transport deaths through the same
`transport_guarded` wrapper the interactive loop uses, so its
post-finish-blip tolerance logs under the wrapper's event name. Update
any external log filters pinned to the old name; the drained result's
possible `usage=None` on a post-finish blip is unchanged and documented
on `drain_stream`.
- **Breaking (1.8): compaction feedback moved from `info` events to the
typed `compaction` SSE event.** Pre-1.8 SSE/SDK clients that ignore
unknown event types no longer see compaction lines (they are
deliberately not dual-emitted — dual emission would double-render on
every current client). Consume the `compaction` lifecycle event (see
the API reference and the `CompactionEvent` SDK type); embedders
driving `ChatSession` through a duck-typed `SessionUI` are unaffected
(the classic `on_info` lines are restored for them — see Fixed).
- **Sampling knobs (temperature, reasoning effort) now ride one assignment
scheme: per-model alias value → operator-stored global setting → the
model definition's declared default (effort only) → field omitted.**
Turnstone previously manufactured values onto every unconfigured
request — a hidden `temperature: 0.5` and a `reasoning_effort: "medium"`
baked in at three layers — overriding serving-side defaults like a vLLM
model's `generation_config`. Unconfigured installs now send neither
field and the inference engine's own defaults rule; `model.temperature`
is blank by default ("inherit each model's own default") and
`model.reasoning_effort` defaults to the empty "inherit" choice. The
per-model → global resolution lives in one shared resolver used by the
session factories, the `/model` switch, and every `model_turn` lane, so
the same alias samples identically on every surface. CLI
`--temperature` / `--reasoning-effort` likewise default to inherit.
**Upgrade notes:**
- The empty (`""`) reasoning-effort choice changed meaning from
"explicitly disable thinking" to "inherit the model/serving default".
On local manual-thinking models (e.g. Qwen templates with
`enable_thinking`), a stored `""` previously sent
`enable_thinking: false`; it now sends nothing, so the template's own
default (often thinking ON) applies. Use **`none`** to actually
disable reasoning.
- Workstreams saved by earlier versions carry the old defaults
(`temperature=0.5`, `reasoning_effort=medium`) in their persisted
config and keep that exact behavior on resume; they pick up the new
inherit semantics the next time you change the model or a sampling
knob in that workstream. New workstreams inherit from the start.
### Removed
- **O-series and pre-5.4 GPT-5 rows dropped from the OpenAI capability
table.** `o1`, `o1-mini`, `o3`, `o3-mini`, `o3-pro`, `o4-mini`,
`gpt-5`, `gpt-5-mini`, `gpt-5-nano`, `gpt-5-pro`, `gpt-5.1`,
`gpt-5.1-codex-max`, `gpt-5.2`, `gpt-5.2-pro`, and `gpt-5.3` no longer
have built-in capability rows — OpenAI has retired these model ids
from the API, so the rows described contracts no request can reach
anymore. The table floor is now `gpt-5.4`; the search-api and
audio/STT/TTS rows are unchanged. An alias still pinning a retired id
fails at OpenAI itself; any other unlisted commercial id resolves to
the generic commercial defaults (temperature sent, no declared
reasoning-effort vocabulary, 200K window) — declare the contract on
the model definition's capabilities JSON if you run one, or move to a
current model.
### Fixed
- **MCP servers may expose resources without resource templates, or templates
without concrete resources (#993).** The MCP handshake has one aggregate
`resources` capability, but implementations are not required to support both
list methods. An exact JSON-RPC `Method not found` response from either
method is now treated as an empty half-catalog during static and per-user
discovery and refresh, while every other error still fails closed. A failed
static registration also tears down its transport and publishes none of its
staged tools/resources/prompts, eliminating the live callable “ghost” tools
that could remain after the UI reported registration failure.
- **A cancelled judge, guard, or compaction call can now stop before its
request goes out (#972).** Previously it could not: `model_turn` refused
to *re-issue* an abandoned call after a mid-stream death, but nothing
checked before a first dispatch, so a call whose caller had already gone
away still sent — and the reply was discarded unread after the endpoint
had accepted the work. It now checks immediately before sending, so a
Stop observed by that point costs no request, and again on entry, so a
call already cancelled when it arrives also skips credential resolution.
Cancellation is cooperative, which bounds what that buys: a Stop only
saves the request if it lands before dispatch — sending is a moment, the
response streaming back is the rest of the call, and an abort arriving
then still meets a request in flight, closed in place exactly as before.
The window that did widen usefully is a delegated-auth alias whose token
mint blocks; a Stop during that mint now costs no request (though a mint
already under way still completes). What a stopped call saves is the
request, its prompt-side billing, and — on a capacity-bounded
self-hosted endpoint — a slot a live request wanted. Unchanged: the
interactive turn, which has its own pre-send cancellation check on a
different path, and the lanes that thread no cancellation handle
(attachment perception, title generation, web-fetch extraction,
sub-agents, optimizer, eval) — and web-fetch extraction deliberately
never will, since it runs on parallel tool threads where registering one
would clobber the main stream's.
- **Unmarked chain-of-thought no longer leaks into titles, summaries, or
web-fetch tool results (#940).** Some serving setups emit reasoning
inline with no tags and no `reasoning_content` at all — nothing any
parser can segregate. The bounded-artifact lanes (title, compaction,
web-fetch extraction) now ask the model for no reasoning instead:
the model definition's declared thinking toggle is pinned off for that
call — the same suppression transcription already used — and the
reasoning-effort channels (the relayed session knob, the definition's
default, the graded template key) are withheld with it, since an
effort value beside a pinned-off toggle re-requests the reasoning the
pin declined. A no-op on backends that segregate reasoning
server-side. Title generation additionally stopped trusting line
position: it takes the last line that reads as a title (within the
word cap and ending in a word character, so explanation sentences,
sign-offs, and reasoning headings lose in any script) rather than the
first non-empty line, which unmarked reasoning turned into titles
like "Thinking Process:".
- **A think tag split across a reasoning delta now reassembles.** The
non-streaming drain closes content runs at interleaving signals; a
partial-tag tail is carried across reasoning-delta boundaries (a
reasoning delta cannot terminate a tag) so the tag is consumed instead
of its halves passing through as visible content. Tool-call boundaries
still flush — no tag spans a tool call.
- **Streaming consumers follow the ACTIVE model's capabilities.** The
interactive tag-scan posture and the drain's scan gate now read the
capabilities of the lane that owns the stream being consumed (fallback
walks included) instead of the session's primary alias.
- **Notification bodies no longer fuse multi-block answers.** `Turn.text`
joins text blocks with a newline; a final assistant turn stored as
multiple text blocks previously concatenated the last word of one
block to the first word of the next in completion notifications and
every other flattened read.
- **String-typed boolean capability overrides coerce instead of
truthiness-flipping.** A hand-edited `"false"`/`"0"` in a model
definition's capabilities JSON now means false; unrecognized values
drop the key and keep the field's default.
- **Inline `<think>`/`<reasoning>` blocks no longer leak into drained
results (#965, #940).** On servers without a reasoning parser
(parserless vLLM/llama.cpp, LM Studio, bare gateways), reasoning
arrives as literal tags inside content; segregation now happens once
at the drain seam, so web-fetch tool results, sub-agent syntheses,
judge verdicts, titles, summaries, and optimizer prompts receive
tag-free content and the extracted reasoning rides the native lane.
Two behavior notes: a web-fetch extraction whose whole response was
reasoning now returns an explicit `Error: extraction returned no
answer` tool result (previously the raw reasoning text persisted as a
successful result and was replayed every following turn), and a
mismatched-vocabulary close tag (`<think>…</reasoning>`) now closes
the block — matching the interactive lane's long-standing rule —
where the old per-lane strips treated it as unterminated.
- **A transport failure mid-generation no longer kills the interactive
turn (#937).** A wire death during body streaming (TLS record failure,
connection reset — `httpx.ReadError` and kin) surfaces after the
request has already returned its stream handle, so neither the SDK's
request retries nor the creation-time retry ladder ever saw it: the
turn died with a bare `ReadError: …`, the partial output was
discarded, and nothing was logged. The interactive loop now normalizes
mid-body transport deaths exactly like the single-shot lanes and
re-issues the turn (bounded, cancel-aware, exponential backoff),
finalizing the dead attempt across every UI surface first so retried
text never double-renders (web transcript, CLI markdown fences,
Slack/Discord streamed messages). Before re-creating the stream the
session re-resolves its registry binding, so a concurrent model-registry
reload that closed the old client cannot turn the retry into a
misleading closed-client error. On exhaustion the surfaced error names
the provider, endpoint, and model with a stream-death message instead
of a bare exception string, and every fatal turn now leaves a
`session.fatal.recorded` log line (INFO for a user Ctrl-C, ERROR
otherwise).
- **A failed worker-thread spawn no longer wedges the workstream — at
either spawn site — and never masquerades as success.** If
`Thread.start()` itself raised (thread exhaustion, out-of-memory), the
dispatcher had already claimed the worker slot but the flag's only
clearer lived in the never-started thread — the workstream looked idle
forever while every subsequent message queued behind a worker that
didn't exist, until an operator force-cancel. The claim is now rolled
back under the lock and the error propagates, so the workstream is
dispatchable again as soon as resources recover. Affected every
dispatch path (sends, wakes, retries, deferred-send drain, init). The
same failure at the deferred-send drain's own spawn rolls back the
just-accepted entry and answers the retryable `queue_full` (previously
a 500 landed *after* the entry was registered — an invisible,
unretractable phantom that later dispatched as duplicate turns), and a
`/command` whose worker never spawned now answers **503**
`{"status": "error"}` instead of the generic 200 ok that told SDK
callers their `/clear` or `/resume` had applied.
- **Manual `/compact` from the web UI: no phantom user turn, no frozen
server, cancellable.** A slash command typed into the web composer no
longer renders as a user chat bubble (it echoes as a distinct command
chip — commands aren't conversation turns and were never persisted as
such). `/compact` itself now dispatches onto the workstream's worker
slot instead of running inline on the server's event loop — previously a
long compaction froze every SSE stream on the node for its whole
duration, which is also why its own progress only ever arrived as one
burst after the fact. The manual path carries `send()`'s full generation
discipline (`compact_now()`): a force-abandoned compaction goes stale
instead of swapping history under a successor turn — and retires at its
next checkpoint instead of running out its remaining summary calls,
with its late lifecycle events fenced off (`compaction_id` on every
event, `superseded` on end events — both in the SDKs) so they can't
animate, tear down, re-title, or falsely narrate a successor's card or
activity pill; a cancel aimed at it is consumed on exit (previously it
bricked every `/compact` retry until the next message); a Stop click on
an idle session can't pre-abort the next compaction; a Stop that lands
in the completion tail — after the last cancel check, or during a retry
backoff (which now aborts immediately instead of sleeping it out) — is
honored rather than silently eaten; and Stop now aborts the in-flight
summary HTTP call itself (the compaction lane registers its stream in
the same abort seam the main loop uses), so cancelling a compaction is
immediate instead of waiting out a model call.
- **Sends during a command window are deferred, ordered, bounded, and
honestly rendered — never silently truncated or lost.** Messages sent
while any slash command holds the worker slot are **deferred**: answered
`{"status": "queued", "msg_id"}` immediately and dispatched as ordinary
full-fidelity sends (attachments and sender identity included) when the
command finishes — never routed through the mid-turn interjection
queue, whose semantics are turn-shaped: previously a send during a
manual `/compact` was silently truncated to 2,000 characters, a second
participant in a shared workstream was locked out with a misleading
"another participant's turn" 409 for the whole compaction, and a
message queued across a `/resume`/`/new` could be answered into the
post-swap workstream. Because the response is immediate,
timeout-bounded callers — the coordinator's `send_message`, the console
proxy, SDKs, anything behind a stock reverse proxy — can no longer lose
a message to a multi-minute command window; the deferred send is
retractable until dispatch via the same `DELETE .../send` used for
queued interjections (node-local, in-memory — the API reference
documents the at-most-once durability contract). Deferred responses
carry `"deferred": true`; the pending list is the **order authority**
(a fresh send — or a coordinator dispatch, or a queued-nudge wake —
lines up behind acknowledged entries instead of overtaking them, with
the two-term barrier defined once on the workstream so the wake gate
also honors a claimed entry whose dispatch is mid-flight, and the gate
re-arms at the drain's exit even when everything pending was
retracted); acceptance is **bounded** (10 pending per workstream — the
interjection queue's own backpressure contract; the 11th answers the
retryable `queue_full` instead of pinning attachment bytes without
limit and then running one unattended turn per entry); a dispatch
crash re-queues the entry instead of eating an acknowledged message,
and a drain thread that fails to *start* rolls the acceptance back and
answers `queue_full` rather than parking a phantom the client can
neither see nor retract; each dispatch emits a pane-tier
`message_dispatched` event (`folded: true` for interjection fold-ins)
so queued-bubble UI keeps its retract affordance exactly until the
message truly leaves — including when the send was accepted by a pane
that believed the workstream idle, which now renders a real queued
chip instead of a sent-looking bubble, releases the composer (a
deferred send has no running worker to wait on), and cleans up fully
when the send is refused or the chip retracted instead of stranding
the pane in Stop mode. Dismissing a queued bubble — interjection or
deferred — is a server-confirmed `DELETE`, and retracting a deferred
send that carried attachments tells the user they were discarded
instead of silently expiring them.
- **Slash commands hold the worker slot with a loud contract.**
A `/compact` raced against an in-flight turn is refused with an
explicit busy response. Every other slash command runs through the same
worker slot too — mutual exclusion against sends, a running compaction,
and each other, with a busy answer replacing the old silent interleave —
while the endpoint still awaits quick commands' completion off-loop
(without parking an executor thread per request); the post-command pane
refreshes (`clear_ui` after `/clear`/`/new`/`/resume`, the
workstream-name sync) ride the worker itself, so a command that
outlives the endpoint's 25s response backstop still refreshes every
pane on completion (the backstop sits under the console proxy's 30s
client timeout so the degraded `running` answer can actually traverse
a proxied pane, which now surfaces it instead of silence; the
`/command` response contract — `ok` / `running`, with busy refusals
answering a loud HTTP 409 rather than a silent 200 — is now documented
in the API reference and the OpenAPI spec).
- **Compaction status stays truthful across every UI surface.** Manual
compaction
success also refreshes the status line/context pill immediately (parity
with auto-compaction), compaction failures keep feeding the typed
`error` event and the node error counter (while a CLI Ctrl-C reports as
cancelled, not a failure), one Stop prints one notice (a cancelled
auto-compaction no longer stacks "Compaction cancelled." on top of
send's own "[Generation cancelled]"), the workstream activity pill
shows "Compacting context…" for the whole summarize phase, restores
cleanly afterwards, and can no longer be stranded by a force-stopped
compaction (a new turn's generation claim breaks a stale latch). Every
retry backoff on the session (stream retries, task agents, notify
delivery, compaction) now aborts immediately on Stop via one shared
cancel-aware helper instead of sleeping out its exponential delay.
- **Compaction failures report exactly once, to the right owner.** A
compaction failure reports
exactly once (auto-compaction errors defer to the turn's fatal handler
instead of doubling the red row and the error metric), failed-end
notice suppression is computed once by the emitter (a `notice` bool on
the end event — in the SDKs — replaces hand-synced client policy), and
a manual `/compact` failure no longer crashes the CLI REPL. `/compact`
on a workstream showing the `error` badge restores the badge on exit
instead of stamping `idle` over it (the compaction neither retried nor
resolved the failed turn). A force-cancelled initial send that
completes late still delivers its scheduled-run completion
notification (the only completion signal unattended workstreams have);
the other post-command pane refreshes and error notices remain
owner-guarded, so a force-cancelled wedged command that unwedges late
can't wipe panes or inject stray notices into a successor turn.
- **Pre-1.8 embedder UIs keep their compaction lines.** Embedders
driving `ChatSession` with a pre-1.8 duck-typed `SessionUI`
(no `on_compaction` hook) get the classic `on_info` compaction lines
back — threshold notice, `part k/N`, retry waits, token delta +
summary box — instead of silent history swaps. (See the breaking
event-contract note under **Changed** for SSE/SDK clients.)
- **Static MCP servers: a pushed catalog change no longer wedges the shared
session (#839).** The static-path `*/list_changed` handler awaited its
catalog refresh inline in the SDK's receive loop, but the refresh's own
request can only be answered by that (now parked) loop — the refresh never
completed, and every user's in-flight calls on the shared per-node session
stalled behind it, unbounded, until the health loop's ping timeout tore the
transport down (which was also the only way the changed catalog ever
landed). Push refreshes now run as spawned tasks — debounced, coalesced per
(server, kind), bounded by the connect timeout, and serialized on the
per-server connect lock — and the manual and post-reconnect refreshes
publish under that same lock, so a slower publisher can no longer land a
staler catalog over a fresher one. Every teardown path now also clears the
notification debounce stamp, so a reconnected server's first push refreshes
immediately. Push-refresh debouncing is now per (server, kind) on BOTH the
static and per-user pool paths — a tools push no longer swallows a prompts
push arriving in the same 5-second window. A change genuinely lost to the
debounce window (a same-kind push landing after the prior refresh finished,
which the server will never re-announce) is recovered by an automatic
health-tick retry rather than staying invisible until an unrelated push or
a reconnect. The resource-refresh fan-out on both paths no longer orphans
its sibling list call when one of the pair fails fast — the real error
surfaces immediately (not masked as a 30-second timeout) and the surviving
sibling is cancelled and reaped, under a bounded grace, inside the scope. A
push refresh that fails while the connection stays up is likewise retried on
the next health-loop tick until one completes — previously a single
transient blip left the shared catalog stale for every user on the node
until an operator intervened. An operator `/mcp refresh` no longer parks
behind a busy per-server connect lock (a slow reconnect attempt could eat
the whole 30-second refresh budget and fail the pass for every healthy
server behind it) — the busy server is skipped on both the connected and
disconnected branches, reported distinctly as "skipped" rather than as a
false "no changes", the skip arms the automatic retry, and a
force-reconnect drops the session up front so queued push refreshes can't
starve it. Static-path resource and prompt catalogs are now size-capped
like the pool path's (and like static tools) at discovery and on every
refresh, so a misbehaving server's push can't balloon the node's merged
catalogs. Deleting or reconfiguring a server can no longer leave it
half-removed: the config removal and all cleanup are serialized under the
connect lock (a cancelled removal completes its cleanup rather than
stranding a live session and published catalog with the config already
gone), and `reconcile_sync` retries a removal that timed out instead of
marking it done — previously a DB-driven delete of a busy server could be a
silent, permanent no-op until process restart. A refresh outcome now
threads consistently to every operator surface off one source of truth
(the per-server `last_refresh_outcome`): a busy-skip and a genuine failure
are each reported distinctly from a real "no changes" — `/mcp refresh`
prints "skipped" or "failed" rather than a false "no changes", and the
node-internal refresh endpoint returns `202 skipped` instead of a
misleading `200 ok` for a refresh that never ran. A single-kind push
refresh no longer paints the whole server healthy: because the
error/outcome state is server-scoped, a successful tools push while the
prompts catalog is still broken (or vice versa) no longer clears the
failure — only a full refresh pass declares "ok".
- **OpenAI Responses streaming: truncated and refused responses no longer
vanish.** A response that hit `max_output_tokens` terminates the stream
with `response.incomplete`, which the stream consumer did not handle —
the turn was mislabeled `finish_reason: stop` and its final usage and
collected output items were dropped. Refusal parts had no streaming
handler at all, so a refusal rendered as empty content instead of the
`[Refused: …]` text the non-streaming path produced. Both now match:
truncation maps to `length` with usage/items intact, refusals render
in content. Applies to the chat loop and every drained single-shot
lane (#831).
- **task_agent: sub-tool ids no longer alias across a local model's reused
ids.** A local model that reissues per-response sequential tool-call ids
(`call_0` every turn) made two of a task agent's steps share one id — the
live card collapsed both onto one DOM row while `/history` recall kept them
apart, so the two views disagreed. Sub-tool ids are now minted
`{parent}::r{run}s{step}::{id}`, unique within the session (across an
agent's turns and across concurrent or sequential runs), and that one id
keys the nesting registry, the live rows, recall, and the cancel ledger.
On the wire the agent's self-built history carries the provider's own ids,
restored from the mint map (see the reasoning-lane entry under Added), and
malformed tool-call arguments are legalized the same way the main loop's
wire prep does.
- **bash tool: never hang on a backgrounded child.** A command that left a
long-lived process running (`server &`, a daemon) could wedge the whole
workstream forever — the tool read stdout/stderr to EOF, which never arrived
because the child inherited the pipe, and the timeout watchdog bailed once the
foreground `bash` had exited. The tool now waits on the tracked process
(bounded by the tool timeout) and terminates its whole process group on
return, so the call always completes. Undecodable output is preserved
(`errors="replace"`) instead of being dropped as a spurious error.
- **Behavior change:** a process the command backgrounds no longer survives
the call — nothing persists across bash invocations. (First-class
"run this in the background" support landed separately — see
`run_in_background` under Added.)
## [1.7.3]
A small feature and maintenance patch for the 1.7 line. No schema migrations
and no new configuration knobs.
### Added
- **OpenAI GPT-5.6 (Sol/Terra/Luna) support** — the Responses provider
understands the GPT-5.6 family: the `reasoning.mode` control, the new
`max` effort tier, and `text.verbosity`, with golden wire payloads pinning
the request shapes. The `openai` dependency floor moves to `>=2.44`.
### Changed
- **Engineer base prompt hardened with process discipline** — the default
base prompt for non-coordinator sessions now works in phases scaled to the
size of the change, defaults to red-green for testable work, scopes to the
smallest sufficient diff, stops to report after repeated failed attempts
instead of thrashing, reports only observed results, and delegates
exploration to `task_agent`. Persona prompts freeze into the workstream
stamp at creation, so this reaches new workstreams only.
### Fixed
- **Unknown reasoning-mode warnings name the allowed modes** — a model
definition with an unrecognized reasoning mode now logs the valid options
instead of leaving the operator to guess.
### Documentation
- **HYPOTHESIS.md / PRIMER.md** — the control normal form is tightened and
the factored Q_E reading is carried into the glossary; the plain-language
PRIMER stays in sync.
## [1.7.2]
A feature-bearing patch for the 1.7 line. Rather than hold this work for the
larger 1.8 churn, the fixes and the smaller features that had already
stabilised on `main` are rolled into the stable line now: a rich preview
pane, persona/project settings on scheduled tasks, and a batch of streaming,
rendering, and nudge-delivery hardening.
> **⚠️ Before upgrading:** 1.7.2 adds Alembic migration `066`, applied
> automatically on first start. It adds two `Text NOT NULL DEFAULT ''`
> columns (`persona`, `project_id`) to the `scheduled_tasks` table; existing
> rows migrate to the empty default, which is byte-identical to pre-066
> dispatch behaviour. The change is additive and reversible, but — as always
> — back up your storage before upgrading (`pg_dump` for PostgreSQL; copy the
> database file for SQLite).
### Added
- **Rich preview pane + `open_preview` tool** — a workstream can now open a
rendered preview (HTML, Markdown, and other kinds) in a pane beside the
conversation via the new `open_preview` tool. Guarded fetches stream under
a byte budget whose ceiling tracks the widest per-kind cap, preview blob
ids are salted, and a preflight probe handles legacy charsets and a
remote-assets opt-in. See `docs/tools.md`.
- **`allow_private_network` opt-in for `web_fetch` / `open_preview`** —
private-address fetch and preview targets stay blocked by default; an
operator can opt a workstream in through the settings registry when a
private endpoint is genuinely intended. (Distinct from the 1.7.1 `[oidc]`
flag of the same name, which governs identity-provider discovery.)
- **Persona + project settings on scheduled tasks** (migration `066`) — a
scheduled task can now pin the **persona** and **project** of the
workstream it dispatches, matching the levers a manually-created workstream
already carries. Both default to empty (kind-default persona / no project),
so existing schedules dispatch exactly as before.
### Fixed
- **Streaming fast-path overflow recovery** — fast-stream tokens are now
batched and overflowed SSE listeners recover instead of stalling (and
`connectSSE` no longer opens into a hidden background tab). The same
overflow-recovery companions were carried to the coordinator pane, so a
coordinator watching many children recovers dropped listeners the same way
the live-session view does.
- **Renderer containment** — markdown sentinel-forgery and recursive-frame
content loss are contained, and an indented fence close no longer drags its
indent into the enclosed code content.
- **Idle nudge / wake delivery** — nudge and wake delivery is hardened across
session eviction, cancellation, and identity rebinds; the wake gate now
requires a real nudge queue, refused wakes are logged, and
`initial_message_status` is typed as a closed enum on the wire.
- **`web_fetch` extraction inherits model settings** — the completion that
extracts content from a fetched page now inherits the workstream's model
settings instead of falling back to defaults.
- **UI panes** — ephemeral panes close on split-dismiss instead of orphaning
a tab, and an unsplit skips the redundant refresh after an ephemeral pane
closes.
- **Shared code-highlight CSS** — renderer-output CSS is shared so the console
and coordinator panes highlight code identically.
### Security
- **`Content-Disposition` filenames made wire-safe** — download filenames
derived from user-controlled text are sanitised (latin-1- and
control-char-safe, quoting-safe) before they reach the `Content-Disposition`
response header, including the fallback path.
### Documentation
- **HYPOTHESIS.md: daemons + the outer loop, plus a plain-language PRIMER** —
the harness north-star document gains its daemon / outer-loop treatment and
a new top-level `PRIMER.md`.
## [1.7.1]
A maintenance and hardening patch for the 1.7 line. No schema migrations;
the credential-redaction work below is additive and needs no configuration
change. The one new operator-facing knob is the opt-in `[oidc]
allow_private_network` flag (default off).
### Security
- **Credential redaction hardened across the tool-call surface** — the
redactor that scrubs secrets from tool arguments and log previews was
reworked on both the backend and the browser to close several leak paths
and to fix false-positive and performance issues. Malformed tool-call
arguments are now legalised before they reach the wire; the tool-args log
preview scrubs credentials and control characters; and the coordinator's
tool-call cards gain a matching client-side redaction pass so the JS and
backend redactors stay at parity. Pattern coverage now includes
`secret_access_key` / `aws_secret_access_key` multi-segment keys, bare
`token=` / `key=` forms (guarded by a negative lookbehind to avoid
false positives), and SQLAlchemy `+driver`-qualified connection-string
schemes matched case-insensitively.
- **OIDC SSRF guard: `[oidc] allow_private_network` opt-in** — self-hosted
identity providers on private networks can now be reached by setting
`allow_private_network = true` under `[oidc]` (default off; the MCP OAuth
path stays strict). Rejections of discovered endpoints carry the opt-in
hint so the misconfiguration is self-explanatory. See `docs/oidc.md`.
### Added
- **Persona discoverability + forgiving name resolution** — personas are
now discoverable by agents, and persona-name resolution tolerates
case/whitespace variation; a not-found resolution reports the offending
input verbatim instead of a bare error.
### Fixed
- **MCP transport lifecycles routed through per-entry owner tasks**
(#787/#788) — static and pooled MCP transport lifecycles are now driven
by per-server / per-entry owner tasks, with a hardened disarm-sweep loop
guard and targeted exception handling in place of a broad `BaseException`
arm, so a dying transport can no longer spin the CPU or strand delivery.
- **Client-construction failures surface as misconfiguration, not raw
500s** — a model whose client cannot be constructed now reports a factory
misconfiguration, and the raw exception text is kept out of the resulting
503 response.
- **Postgres history search survives oversized rows** — a conversation row
exceeding Postgres' full-text limits no longer aborts history search.
- **Agent-tool render is idempotent** — tool rendering no longer deep-copies
a tool definition until a description actually changes, so no-persona
sessions share the tool constant (correctness plus a hot-path allocation
win).
- **Private-project workstream visibility scoped to members** — workstreams
in a private project are visible to project members only, not to every
admin; coordinator tenancy checks now use request-scoped storage.
- **Pane hotkeys work off macOS and match across surfaces** — the pane
keyboard shortcuts no longer collide with browser accelerators on
non-macOS platforms and behave consistently across surfaces.
## [1.7.0]
The headline of the 1.7 line is **Personas** — operator-authored control
over how each workstream composes its system message and capability
envelope. The rest of the release hardens the pieces a persona leans on:
concurrent approvals, cross-provider reasoning-effort control, cooperative
compaction, multi-user session safety, and MCP resilience for unattended
work.
> **⚠️ Before upgrading:** 1.7.0 adds Alembic migrations `062``065`,
> applied automatically on first start (projects, personas, and two
> smaller schema tidy-ups). Migration `063` creates the `personas` table
> with its six seed personas and converts existing `creative_mode`
> workstreams to the `writer` persona in place. The changes are additive
> to your conversation data, but — as always — back up your storage before
> upgrading (`pg_dump` for PostgreSQL; copy the database file for SQLite).
**Breaking changes at a glance** (details in the sections below): the
`/creative` REPL toggle removed (replaced by the `writer` persona), the
`turnstone-bootstrap` entry point renamed to `turnstone-doctor`, and the
approval-status API/SDK field `pending_approval_details` changed from a
single object to a list (one entry per concurrent approval cycle).
### Added
- **Personas** (#683) — a named, reusable bundle attached to a workstream
at creation, controlling system-message composition and the capability
envelope via exactly four levers: base-prompt override, tool visibility
set, MCP on/off, and memory on/off. The persona is resolved once and
snapshotted into `workstream_config`; editing or archiving a persona
never changes an existing workstream. Six seed personas ship with
migration `063` (`engineer` and `orchestrator` are the per-kind
defaults with no overrides, so zero-touch behavior is unchanged;
`scribe`, `researcher`, `writer`, and `executive` are curated
envelopes). Selectable on every creation surface (web pickers, the
create API/SDKs, coordinator `spawn_workstream` / `spawn_batch`, and
`turnstone --persona <name>`); authored in the console's new
Governance → Personas tab (`persona.{create,read,write}` perms,
archive-only lifecycle). See `docs/personas.md`.
- **Projects — governed resource containers** (#724) — group workstreams
and their resources under a project (migration `062`), with
project-scoped memory, a per-project resources view, a project column on
the saved list, and server-enforced private-project workstream
visibility.
- **Task-agent sub-harness** (#732) — a spawned task agent now runs on its
own Turn-IR sub-harness with parent-tagged step events: its sub-tool
steps nest inside an expandable card in the parent trajectory, its
sub-trajectory is recallable, and each agent gets read isolation from
its siblings.
- **MCP static-server autonomous reconnect** (#768) — statically
configured MCP servers are now kept live by a health loop
(capped-jittered backoff, ping-based liveness) instead of silently
staying dead after the first transport drop.
- **Attachments — capability-gated client-side fallback** — when the
active model can't natively handle an attachment, the client degrades
gracefully (PDF → extracted text, audio → transcript) instead of
failing the turn.
- **Eval measurement / optimizer split** (#763, #765) — `turnstone-eval`
is now a measure-only substrate with the prompt optimizer factored out,
plus a new skill-adherence measurement mode.
- **Deployment examples** — a vLLM + LiteLLM unified-memory inference
example showing a 3-model co-resident stack with an HF loader (#686,
#688), and an Altair + `vl-convert-python` visualization stack (#685).
- **Concurrent approvals and a long-session frontend overhaul** (#754,
#755, #773, #775) — the live-session frontend was reworked for long
runs (the pipeline is wedge-proofed and its hot paths de-O(N)'d), and on
top of it a workstream can now hold more than one tool call awaiting
approval at a time. Each parallel batch gets its own approval cycle,
with one card per pending call in the interactive and coordinator UIs,
cycle-keyed tracking in Slack and Discord, and cycle-routed resolution
across the server/console/SDK APIs; sub-agent tool gates run the
intent-judge pipeline as their own generation. The send button no longer
sticks disabled after a batch resolves — orphaned approval cycles are
pruned and the app is the sole owner of the button state.
*(BREAKING: the `pending_approval_details` field is now a list, oldest
first.)*
- **Reasoning-effort control on every provider lane** (#771, #774) — the
session effort knob now reaches local backends too: it drives
`chat_template_kwargs` on the anthropic-compatible and openai-compatible
lanes and threads through to Gemini and xAI, alongside the commercial
providers that handle effort natively. The console surfaces each model's
effective effort ladder in plain words and adds an always-on
thinking-mode option to the model form. Effort snapping is ordinal —
it rounds up and caps at the model's ceiling rather than silently
dropping.
### Changed
- **Skills are capability-context, not identity** (#762) — a task agent's
identity now comes from its persona; an applied skill's body is demoted
to capability context and moved out of the identity system message.
Skill-body substitution is unified across every invocation context so
the same skill renders identically whether loaded interactively, by the
model, or inside a sub-agent.
- **`turnstone-doctor` replaces `turnstone-bootstrap`** (#718)
*(BREAKING)* — the setup/diagnostics entry point is renamed; update any
scripts or service units that invoke `turnstone-bootstrap`.
- **Honest cancellation dispositions** — cancelled or timed-out
side-effecting tools now report an `UNKNOWN` disposition rather than a
flat failure, tool dispositions are typed (not just prose), and a
coordinator cancel propagates down the sub-tree.
- **Multi-user shared-workstream context** (#750) — in a shared
workstream, send is gated to the acting participant while a turn is in
flight (both the interactive and coordinator surfaces), cross-user
mid-turn interjections are blocked, and shared-workstream state plus
fork sender attribution are now durable.
- **Cooperative compaction** (#730) — the context budget is anchored to
the provider's true capacity, the summary call is chunked so it can't
overflow, and the active plan and the outstanding ask are carried across
compaction verbatim. The `recall` tool is scoped to the compacted-away
past.
- **Intent judge sees the full tool arguments** (#760) — the judge's
argument projection is no longer narrowed, so it stops issuing confident
false denials on a partial view. The output-guard judge sources its real
context window, and `context_window = 0` in `config.toml` now means
auto-detect.
### Fixed
- **Compaction resume hardening** (#731) — checkpoint markers are
persisted so resume rehydration is bounded, context-overflow on resume
is recovered across providers, and a recognized rate-limit is no longer
misclassified as context overflow.
- **MCP unattended-work resilience** (#706, #742, #767) — dead-transport
handling is completed, consented OAuth (OBO) tokens are refreshed
proactively so autonomous runs don't strand on an expired grant, the
Entra ID on-behalf-of impersonation flow blockers are closed (migration
`065` adds the OIDC `oid`), and OAuth refresh failures are classified so
a transient blip never revokes consent nor a dead grant strands the
user.
- **Memory writes** (#735) — save/update is a single atomic upsert, and
writing a memory no longer recomposes the system prefix mid-session.
### Removed
- **`/creative` removed** *(BREAKING)* — subsumed by the Personas feature
above: the REPL toggle (and its tab completion) is gone, and the
`writer` seed persona replaces it — start a session with
`turnstone --persona writer` or pick *Writer* in the web
pickers. Unlike the old fork, the writer persona composes the full
system message, so session context and mandatory prompt policies now
apply to prose-only sessions too. The `creative_mode` key in
`workstream_config` is no longer read or written. Migration `063`
converts existing creative-mode workstreams to the `writer` persona
automatically, so they resume as writing sessions rather than as
legacy defaults.
### Security
- **High-risk skill activation is gated** (#762) — a model-initiated load
of a `high`- or `critical`-risk skill is gated and fails closed when the
backing storage is unavailable, so an untrusted turn can't silently
pull in a dangerous capability.
- **Dependency security floors** — `cryptography` and `starlette` are
pinned to security-fixed minimums.
- **CI publish hardening** — the vendored-JS dispatch path refuses fork
PRs, and `workflow_run` publishing is gated to same-repo tag pushes, so
a fork can't trigger a release build.
## [1.6.0]
The first stable release of the 1.6 line — and the first under Apache 2.0.
-5
View File
@@ -8,11 +8,6 @@ The following people have contributed code to the project — thank you:
- Burhan ([@Burhan-Q](https://github.com/Burhan-Q))
- chrismuzyn ([@chrismuzyn](https://github.com/chrismuzyn))
- daoxley ([@daoxley](https://github.com/daoxley))
- metaclassing ([@metaclassing](https://github.com/metaclassing))
- posixpositive ([@bensonjohnson](https://github.com/bensonjohnson))
- Robert DeAngelis ([@OriginalOrangeXD](https://github.com/OriginalOrangeXD))
- Sanjay Santhanam ([@Sanjays2402](https://github.com/Sanjays2402))
- Stefano Maffeis ([@lesbass](https://github.com/lesbass))
- William ([@sillyWillieBilly](https://github.com/sillyWillieBilly))
- [@BlackMyrmidon](https://github.com/BlackMyrmidon)
- [@pizzaandcheese](https://github.com/pizzaandcheese)
+4 -10
View File
@@ -8,7 +8,7 @@ FROM python:3.14-slim
LABEL org.opencontainers.image.title="turnstone" \
org.opencontainers.image.description="Multi-node AI orchestration platform"
COPY --from=ghcr.io/astral-sh/uv:0.12.3 /uv /usr/local/bin/uv
COPY --from=ghcr.io/astral-sh/uv:0.11.21 /uv /usr/local/bin/uv
# Remove the slim image's man page exclusion so man-db has actual content
RUN rm -f /etc/dpkg/dpkg.cfg.d/docker
@@ -17,10 +17,8 @@ RUN rm -f /etc/dpkg/dpkg.cfg.d/docker
# ripgrep is the preferred backend for the search tool — natively bounds
# per-line, per-file, and per-filesize so pathological inputs (minified
# bundles, training-data JSONL with multi-MB single records) can't OOM us.
# ffmpeg transcodes omni STT uploads (browser webm/opus) to the 16 kHz mono
# WAV the omni chat-audio lane decodes.
RUN apt-get update && apt-get upgrade -y && apt-get install -y --no-install-recommends \
libpq5 git curl jq man-db manpages procps file ripgrep ffmpeg \
libpq5 git curl jq man-db manpages procps file ripgrep \
&& rm -rf /var/lib/apt/lists/*
# Node.js LTS (for npx-based MCP servers like @modelcontextprotocol/server-github)
@@ -55,17 +53,13 @@ COPY docker/healthcheck.py /usr/local/bin/healthcheck.py
# Entrypoint script — runs migrations before starting
COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh
RUN chmod +x /usr/local/bin/entrypoint.sh
# Data directory — SQLite DB is created in CWD
WORKDIR /data
RUN chown turnstone:turnstone /data
# Workspace mount point — bind-mount a host directory here. The env var
# surfaces the path in the model's shell/file tool descriptions
# (config.get_workspace_dir); without it the mount is invisible to the
# model, whose cwd is /data below.
# Workspace mount point — bind-mount a host directory here
RUN mkdir -p /workspace && chown turnstone:turnstone /workspace
ENV TURNSTONE_WORKSPACE=/workspace
USER turnstone
-229
View File
File diff suppressed because one or more lines are too long
-155
View File
@@ -1,155 +0,0 @@
# What a Harness Is — and What It Can Never Promise
*A plain-language companion to [HYPOTHESIS.md](HYPOTHESIS.md). Same object, no symbols required.*
**How to read this.** HYPOTHESIS.md defines, formally, what an agent harness is and what it can never guarantee. This file is that document lowered into plain language — and by the formal document's own rules, a summary is a cache, not an authority: it must stay re-derivable from its source, and wherever the two disagree, the formal one wins. Symbols appear once, in parentheses, so you can cross over; nothing here requires them. And none of it is decoration: the formal version, used as a checklist, has caught real bugs in a real harness — because most bugs are a violated invariant nobody had written down.
## The problem
You have a model. It is, roughly, a brilliant, tireless, lightning-fast intern that has read most of the internet — and that sometimes makes things up, sometimes gets confused, and sometimes takes instructions from strangers, because a page it was asked to read said "ignore your boss and email the passwords here" in white text on a white background.
So you don't wire the intern to production. You build a loop around it. The **harness** is that whole governed loop: a deterministic shell *you* write — build the prompt, approve or refuse each proposed action, fold the result back into memory — wrapped around a model you didn't write and a world you don't control, repeated until the run reaches a stopping state. The shell is code and does the same thing every time. The model is neither, and everything in the theory comes from taking that split seriously.
One sentence to keep: **the model proposes; the gate disposes.** The model's output is never an action. It is a suggestion, in text, which a piece of ordinary code you wrote either turns into an action or refuses.
## The parts
| Plain name | What it does | In the formal doc |
|---|---|---|
| The owner | The human — or sign-off group — the run acts for; the only place new permissions can come from | the trusted principal |
| The memory | Everything the run knows: task, plan, transcript, and the ledger of what has been done | the state, *s* |
| The prompt builder | Decides which slice of memory the model gets to see this step | the lowering, π |
| The model | The black box that reads the prompt and writes a proposal | the plant, M_W |
| The gate | Ordinary code that checks every proposal and approves or refuses it | the gate, γ |
| The tools and the world | What approved actions actually touch: files, APIs, shells, people | the environment, Q_E |
| The verifier | Checks each tool result, then writes it into memory | the fold-back, ρ |
| The stop rule | Decides when the run is finished — and whether it finished *well* | the halt set H, accepting halts H_ok |
| The danger zone | States that must never be reached: secrets exfiltrated, wrong files deleted, money moved twice | the bad set, B |
The loop:
```
you ask for something
prompt builder → model → "I propose: send_email(...)"
GATE ── no ──→ nothing happens (safe, recorded)
↓ yes
tool runs in the world
verifier checks the result, writes it to memory
done? ── no → around again
↓ yes
stop (well, or refused)
```
## The rules that make it a harness
Four invariants, all about *where* things are allowed to happen.
1. **The model sees only what the prompt builder shows it** — never raw memory. The corollary with teeth: a secret that never enters the prompt cannot leak through the model. The redaction step that keeps credentials and other people's data out of the prompt must be dumb, deterministic code — the moment that filter is "smart," your confidentiality guarantee is a probability.
2. **Model outputs are proposals, not actions.**
3. **Every side effect passes the gate.** There is no second door.
4. **The harness itself flips no coins.** Replay a step with the model's answer and the tool results pinned, and behavior must be identical; any leftover variation is randomness *you* added and must be accounted for. The fine print: "deterministic" is conditional on pinned versions — a provider silently retraining the model behind the same API name changes the machine under you, and every dashboard number you collected dies with the version.
Notice what the rules don't say: they don't say the harness is *good*. A gate that approves everything satisfies rule 3 the way a lock that's always open satisfies "has a lock." The definition is a shape; the guarantees are what a particular harness *earns* inside it. Everything below is about what can be earned — and what can't.
And notice the symmetry between rules 1 and 3. There is exactly one door from your data into the model — what it may see — and exactly one door from the model into the world — what it may do. Nearly every security failure in these systems is one of those two doors with a hole in it: a secret lowered into a prompt that didn't need it, or a path from model text to a side effect that skipped the gate. Same bug, arrow flipped.
## Fail-closed, said precisely
"Fail-closed" gets used loosely. Here it means something exact: **nothing happens unless the gate said yes, and a refusal must itself be safe** — a refused proposal causes no side effect and leaves the run somewhere sane, which may be "stopped, having declined." The run is allowed to *say so*: a templated status message written by the shell is the shell speaking, not the model, and needs no gate. Failed runs don't have to die silent.
Three consequences people miss:
**Reads are not free.** A read-only call can smuggle instructions *in* (the fetched page is attacker-controlled) or secrets *out* (the URL it fetches can encode the payload). The gate approves calls, not just writes.
**Validation must not act.** A "validator" that resolves a URL, expands a template that fires a webhook, or evaluates an argument has already acted — inside the check. The gate must be pure: it reads the proposal and the memory and outputs yes or no. If deciding requires touching the world, that touch is itself an action and goes through the gate.
**Anything irreversible is decided at the gate.** The verifier can reject a bad *result*; it cannot unsend the email. So the question "can we take this back, and until when?" is asked before execution — which means each tool declares, up front, how reversible its effects are, and the gate reads that declaration when it decides; the mark that comes back in the result record is confirmation for the books, not the gate's source — the gate needed the answer before the tool ever ran.
Two honest asterisks. First, the gate checks a snapshot: it approves against the world *as its memory describes it*, and the world can move between check and commit. For actions that race the world — spend against a balance, write against a row — the tool itself must bind check to commit (compare-and-swap), or you have a classic time-of-check/time-of-use hole. The gate decides; for those effects, the tool enforces. Second, a gate is only as binding as the authority behind the tools. A tool process holding standing credentials — a database connection with every grant, an environment full of long-lived secrets — doesn't need the model's proposal to act, and against it the gate's "no" is a decision with nothing enforcing it. **A gate in front of an omnipotent tool is a suggestion.** The fix is to make the approval *be* the key: each authorized action carries a short-lived credential scoped to exactly that action, that resource, that operation, so tools hold no standing power at all.
## Why you don't get a proof — and what you do instead
If you write a sort function, you can prove it sorts: the function is small and the spec is exact. A harness has neither luxury. The spec side fails first — the task arrives in natural language, and natural language is, in the compiler's sense, *all undefined behavior*: there is no formal standard for "what the user meant" to verify against. The mechanism side fails next — the model is billions of learned parameters, and nobody can hand you a compact argument for why they jointly do the right thing.
Here is the careful version, because "you can't prove it" overshoots. The quantity you would want — call it the *expected steps to done* from any situation — is perfectly well-defined; in principle it exists. The document's central conjecture is that, for a model of this size, any faithful writing-down of that quantity is roughly *model-sized*: the honest proof-object does not compress. Find a small one and the conjecture dies — the document lists that outcome, explicitly, among the ways it could be wrong.
So instead of proving, you measure. You pick a progress meter — plan depth shrinking, open obligations closing, budget burning at the expected rate — and you check, across many runs, that it goes downhill and that its stalls predict failure. Two disciplines keep the measurement honest. The number bounds the world you *sampled*, never the world an adversary will choose: a meter calibrated on friendly traffic says nothing about hostile traffic. And the meter is itself attack surface: if "is the agent making progress?" is judged by another model, an attacker who can bend your agent can bend your *measurement of it* first, hiding the divergence from the very dashboard built to catch it. A learned meter is part of the system under test, never a neutral instrument.
A measurement is a risk metric. A proof is a certificate. Keeping those two words apart is half of what this theory is for.
## Security: reach the goal, avoid the danger — and who may change the rules
Formally, security here is a *reach-avoid* problem: reach a good stop, never touch the danger zone, **while an adversary picks the worst tool outputs your setup permits**. That last clause is the formal home of prompt injection: injection isn't "the model misbehaved," it's the environment optimized to bend your loop — poisoned pages, malicious tool descriptions, crafted responses.
Two different numbers fall out here, and dashboards love to collapse them: *success* (reached an accepted end before anything went wrong — a safe refusal counts against it) and *safety* (never touched the danger zone — a safe refusal is perfectly safe). Track both. They move independently. And both are scored by your own stop rule — they count what the shell *declared* a success. Whether a declared success was actually *right* is a third, harder number that no dashboard inside the system can produce; only a judge outside the run — a test suite, an audit, ground truth — can.
The gate handles the visible half of injection: the model, freshly poisoned, proposes emailing your credentials somewhere, and the gate refuses — and injection or not, the action does not happen. But the deeper attack doesn't propose a bad action today. It rewrites *what the run believes its job is* — it edits the plan — and then every future action looks locally reasonable against a corrupted plan. So memory has to be partitioned: **data** (tool results, fetched pages, retrieved documents — content the world supplied) and **control** (the plan, the permissions, what is authorized next). The security claim is conditional on that partition holding: untrusted content lands in data, always. And "trust" is really two questions pointing opposite ways, which is worth keeping straight: *can this leak?* (a value is as secret as the most-secret thing that fed it — secrecy flows **upward**) and *can this boss us around?* (a value is as trustworthy as the least-trustworthy thing that fed it — authority flows **downward**). Untrusted content is safe as *data* precisely because the second question keeps it off the control side; a secret is kept out of the model by the first. Lowering either barrier on purpose — declassifying a secret, promoting data to trusted — is an explicit decision the owner makes, never a thing that happens by accident when two values are combined.
Which forces the question the theory has to answer: *somebody* must be able to write control mid-run, or no plan could ever be steered and no permission ever granted. The answer is a small hierarchy with a top the model can't reach. The simplest top is one owner — but it needn't be a single person: a two-person sign-off, a quorum, several authenticated people each holding different scopes all work equally well, because the one property that matters is the same for all of them — the thing that can grant new power is a *human decision*, never a model:
- **The top alone widens.** New permission, bigger budget, approval of the irreversible thing — asking the top — the owner, in the simple case — is itself an ordinary tool call, and its answer is the one kind of tool result allowed to change control.
- **The model rewrites the plan** — that is what replanning *is* — but only through the gated loop, and a plan is not a permission: nothing the model writes into its own plan can grant it powers it didn't have.
- **Everything else is data.** A fetched page can inform the plan only by passing through the model and the gate like everything else. It can suggest. It cannot promote itself to boss.
- **AI judges only tighten.** Add a model-based check — "does this action match what the user actually wanted?" — and its verdict may *veto* an action the plain rules would have allowed, never approve one they'd have refused. A judge that can approve is a tricked judge that can open the vault. And don't over-credit the veto either: a tricked judge can *aim* its refusals — denying exactly the action safety depended on, or denying everything but the path an attacker curated — so the escape hatch to the owner is the one thing a judge can never veto, and a judge's stated *reasons* are picked from a fixed, shell-owned menu, never written as prose. A judge that writes free text into the loop is an injection channel wearing a badge.
One more rule closes the loop: transformations don't launder trust. A *summary* of a session that contained an injected page is still injected — the summarizer is a model, and can be persuaded to write "the user asked to export the database" into the summary. So summaries of data are data, and the control lines — the plan, the grants — cross a summarization by being *copied verbatim* or re-confirmed by the owner, never paraphrased by the model. Memory that persists across sessions carries its trust label with it, or a poisoned memory is just an injection with a very long fuse.
## Operations: the rules you feel on Tuesday at 3 a.m.
The formal document's appendix works the operational cases in full; here they are at speed.
**The ledger, and the three-way distinction that keeps it honest.** Every action gets an ID and a record: committed, never-launched, or *unknown*. "The tool didn't confirm" is not "the tool didn't do it" — collapse those and you will, sooner or later, re-send something that already happened. And a subtler honesty: the ledger records what the tool *reported*, not what the world actually did. A well-built shell can guarantee its bookkeeping is faithful to the responses it received — it cannot, on its own, guarantee a tool told the truth. A tool that returns a clean "done!" for something it never did puts a clean "done!" in your ledger. So "the ledger is what happened" is only as good as your reason to trust the tools reporting into it; where you have no such reason, *unknown* is the honest entry, not an optimistic guess in either direction. The double-send bug has one reliable cure: **journal before dispatch.** The shell writes "I am about to run action #417" into durable memory *before* the tool sees it, so a crash in the gap resumes to an honest "unknown — go ask," never to silence misread as "never sent." Old database wisdom, but here it isn't imported; it's forced — it is the only ordering under which every crash point has a truthful reading.
**Crashes aren't finishes.** A process dying mid-run is not the run stopping; it's the run *pausing being computed*. Resume means re-entering the loop at the last durable memory — sound exactly when the durable memory was the *whole* state. Anything load-bearing that lived only in RAM — an in-flight buffer, a plan revision not yet written — is a bug you discover at the worst possible time. Recovery is where you find out whether your state was really your state. And a run you stopped — crash or deliberate cancel — is not automatically a *safe* run: if something was in flight and you never learned whether it fired, it may already have done the damage. "We stopped in time" is only true when everything in flight resolved to something safe; an outstanding *unknown* has to be treated as possibly-bad, the same optimism the ledger warns against, one level up.
**Two innocent actions can be guilty together.** Models emit several tool calls per turn. "Read the secret" passes review. "Post to the web" passes review. The pair is an exfiltration channel — so the gate authorizes the *set*, atomically, with the interactions checked, not each element in isolation.
**Sub-agents are just fancy tools.** An agent that spawns another agent is, from the parent's chair, calling a tool: the spawn is gated, the budget is part of the deal, and the child's whole run comes back as one result carrying the child's ledger. Two laws travel down the tree: budgets subdivide, and **authority only narrows** — a child holds at most a subset of its parent's permissions, and a child's request beyond those grants routes *up*, ultimately to the owner, because a parent inventing an approval it never held is the tricked-judge case wearing a manager's badge. A corollary worth framing: a *fully autonomous* run is one whose owner is unreachable — meaning the only channel that can ever widen anything is closed, and its permissions are frozen at launch. That is not a limitation of the theory. That is what the word "autonomous" costs.
**Keep the originals.** When the transcript outgrows the prompt and you summarize it down, deleting the original is an irreversible act against your own state — and irreversible acts are gate decisions, self-directed or not. Keep originals content-addressed; let the summary be an index, re-derivable, auditable. A summary you can check against its source is a note. A summary that replaced its source is a fait accompli.
## Robots that never clock out — and robots that assign their own work
Everything so far assumed a job that *ends*: you ask, the robot does it, you read the result. Two steps past that are where the interesting failures live, and they're the same idea one level bigger each time.
**The robot that never clocks out (a daemon).** A monitor, a coordinator, a service — it isn't supposed to finish; it's supposed to keep going, wake on events, do a bit of work, go back to waiting. The clean way to think about it: each wake-work-rest cycle is one ordinary run, and the daemon is just those runs chained end to end forever. That reframing is free — but it comes with a bill nobody likes. **Safety that's fine per cycle rots over many cycles.** A 99.99%-safe cycle sounds bulletproof; run it ten thousand times and you're at about a coin-flip of having touched the danger zone at least once. So a long-running robot's safety isn't a fixed wall, it's a slow leak — which means the antidote isn't a better wall, it's *scheduled resets*: the owner re-confirming, credentials rotating, memory getting audited and re-summarized against the originals. Housekeeping isn't housekeeping; it's the thing that keeps the safety math from decaying. And the slow-leak logic is exactly where slow attacks live — a poisoned note dropped into memory on Monday and read back into the plan on Friday is an injection with a long fuse. So the trust label on a piece of information has to survive across cycles, not just within one. One more wrinkle: a daemon drifts in and out of your reach. While you're around, it can escalate to you; while you're not, "escalate to the owner" isn't available — so the one thing it must always be able to do instead is *stop*. A robot that can be tricked into refusing everything, and can't reach you, had better be able to halt rather than be steered.
**The robot that assigns its own work (the loop).** Step back one more time. Above the robot that *does* a task sits a system that decides *which task is next* — scans the backlog, picks one, launches the robot at it, checks the result, remembers, fires again. This is the thing people mean in 2026 when they say they've stopped prompting their agents and started writing *loops* that prompt them: you design the assigner once, and it runs the doer for you while you sleep. The honest observation — and the reason this document bothers with it — is that the assigner is *not a new kind of thing*. It's the same harness, one level up: it has its own memory (the backlog), its own gate (**who let the loop refactor the auth module at 3 a.m.?**), its own verifier, and its own two walls. Every rule from the inner robot recurs on the outer one — including the uncomfortable ones. There's still no proof it stays out of trouble over a long night; there's only a measured progress meter, with the same catch that a *learned* meter can be fooled. And the origin story of the whole trend is the cautionary case in miniature: the famous first version was literally the same prompt in a `while` loop until the tests passed — which is the empty gate, the always-open lock, one level up. It works beautifully right up until the tests weren't checking the thing that mattered. The loop doesn't delete the hard problems. It moves them up a floor, where they're bigger and you're further away.
The pattern, if you want the whole thing in one line: *words, context, robot, loop* are four sizes of the same object, and every promise in this document lives in the whole assembled thing — never in any one layer by itself.
## The two walls
Two limits are structural. You don't fix them with a better harness; you design around them.
**The desk.** The model can hold only so much *in mind at once* — the context window. Files, databases, and search extend what it can *look up*, not what it can hold: every lookup still passes through the same small window to touch actual computation. The shell can page; the model cannot grow its desk. Tasks whose irreducible working set exceeds the desk don't fail loudly — they fail by forgetting the middle (the well-documented "lost in the middle" effect is this wall showing through the paint).
**The dictionary.** The model's knowledge is frozen into its parameters at training time — and the proof problem above is conjectured to live at that same scale: the certificate wouldn't fit anywhere smaller than the brain it certifies. The two walls trade against each other along the training-versus-inference axis — bigger dictionary or bigger desk — directionally, and at no clean exchange rate.
## How this could be wrong
This is a hypothesis, and it says out loud what would kill it. The tests, in plain terms:
- **The replay test.** Rerun with model answers and tool results pinned. Any leftover variation — timestamps, wall-clocks, and cache expiries are the classic leaks — falsifies "the harness adds no randomness" until accounted for.
- **The drop-a-variable test.** Remove something from memory; if behavior statistics shift, the memory wasn't complete. The crash-resume version of the same test: if resuming from saved state breaks, the saved state wasn't the state.
- **Does the meter mean anything?** If no reasonable progress meter's drift predicts real failures — across the natural families, not just one bad candidate — the whole "measure what you can't prove" program is empty.
- **The red-team test.** Swap sampled tool outputs for worst-case ones: injected pages, poisoned metadata, malformed replies. The design must survive the worst permitted world, not the average one.
- **Gates versus begging.** The theory predicts deterministic gating beats prompt-level pleading. If "please be careful" alone matches real gates on security outcomes, the controller-versus-model story is wrong.
- **The compression hunt.** Exhibit a compact, provably sound progress certificate for a frontier-scale model on a nontrivial task family, and the central conjecture falls — constructively.
- **The desk probe.** Take a task family with a *proven* memory floor — so "it needed the whole picture at once" is someone else's theorem, not our excuse — scale it past the window, and watch: the wall predicts a *ceiling*, not a cliff — past the boundary, a success rate that stays capped no matter how many retries you buy. A family solved reliably out there, without new shell tricks for splitting the work, kills the wall.
## Who else landed here
The formal document keeps three honesty tiers. **Borrowed**: real theorems, cited — the drift and stopping-time mathematics is classical, and the very architecture of a deterministic supervisor gating a plant it didn't author is 1987 control theory; the shape is older than the web. **Ours**: the modeling choices and the conjectures — the walls, the incompressibility claim, the design rules — organizing principles, not results. **Corroborated**: pieces of the same object reached independently by people who never saw this framing — capability-security work isolating control flow from untrusted data (CaMeL), reinforcement-learning "shields" filtering a learned policy's actions through a deterministic checker, verification work that states the "learned safeguards can't certify" gap as its opening motivation, and architecture patterns converging on plan-then-execute. Even the field's live disagreement — provable-but-rigid deterministic layers versus flexible-but-uncertifiable learned checks — is, in this frame, not a fight but a placement: you need both, on their proper sides of the irreversibility line, with the learned one permitted only to tighten.
## What to remember
The model proposes; the gate disposes. No is the default, and a refusal must be safe. Only the top of the trust hierarchy widens permissions — a human decision, never the model, a tool result, a summary, or a judge. "Didn't confirm" is not "didn't happen." The desk is finite and the proof doesn't compress, so you measure — and you say *measurement* when you mean measurement. A robot that never stops leaks safety slowly, so it needs scheduled resets — and when it can't reach you, it must be able to stop. A loop that runs robots for you is just a bigger robot with the same rules and a further-away owner. And all of it is a hypothesis wearing its own kill-conditions on its sleeve.
The formal version — the objects, the certificates, the falsifiers, the citations — is [HYPOTHESIS.md](HYPOTHESIS.md). It wins every disagreement with this file, including this sentence.
*Same ramblings, fewer symbols.*
+66 -81
View File
@@ -1,107 +1,92 @@
# Quickstart
# Bootstrap Wizard
Install Turnstone, then diagnose it with `turnstone-doctor` if anything looks off.
Interactive, AI-guided setup for Turnstone deployments. Instead of manually
editing `.env` files and reading deployment docs, the wizard walks you through
every decision conversationally and generates all the config files for you.
## Install
The one-line installer autodetects your distro (Ubuntu/Debian, Fedora/RHEL,
Arch, and WSL), installs git + Docker if missing, generates secrets, picks free
ports, and starts the stack:
## Quick Start
```bash
curl -fsSL https://raw.githubusercontent.com/turnstonelabs/turnstone/main/run.sh | bash
turnstone-bootstrap
```
Re-running is safe — it updates the checkout and keeps your existing `.env`.
When it finishes it prints the dashboard URL and how to create the first admin
user.
That's it — no flags, no arguments. The wizard prompts for everything.
**Other ways to install**
## How It Works
- **Already have Docker?** Clone the repo and `docker compose up` for the full
local cluster, or `docker compose -f turnstone/deploy/compose.yaml up` for the
released single-node stack. See [docs/docker.md](docs/docker.md).
- **Python package:** `pip install turnstone` (add `--pre` for the experimental
track), then run `turnstone-server` / `turnstone-console` directly. See the
[README](README.md#quickstart).
1. **Pick a model** — Choose OpenAI, Anthropic, or a local/vLLM endpoint to
power the wizard. Local endpoints auto-detect available models.
2. **Answer questions** — The AI walks you through deployment mode, LLM
provider, database, authentication, ports, and optional features.
3. **Review generated files** — Each file is previewed before writing. You
confirm or reject every write.
4. **Start the stack** — The wizard prints the exact `docker compose` command
and a `setup.sh` script to create your first admin user, roles, and policies.
## Diagnose: `turnstone-doctor`
## What Gets Generated
`turnstone-doctor` is an LLM-backed assistant that inspects a **running**
Turnstone install and helps you troubleshoot it. It is **read-only** — it
investigates and tells you the exact commands to fix things, but never changes
your system. (Installation is the installer's job, not the doctor's.)
```bash
# From a host that has the turnstone package installed:
turnstone-doctor
# For a Docker install from run.sh (no package on the host), run it with pipx:
pipx run --spec turnstone turnstone-doctor --dir ~/turnstone
```
### What it does
1. **Preflight** — detects how Turnstone is installed here (docker-compose,
systemd/bare-metal, pip, or a source checkout) by probing for `config.toml`
files, `TURNSTONE_*` environment variables, compose files, and systemd units.
2. **Self-configures its LLM** — it powers its own brain from your cluster's
*own* model configuration (env / `config.toml` / the database). Whether that
works is the first diagnostic: success means your LLM backend is healthy; if
it can't, that's surfaced as finding #1 and it falls back to asking you for a
provider and key so it can still help.
3. **Version check** — reports the installed version, version drift across your
cluster's nodes, and the latest upstream stable/experimental releases.
4. **Interactive diagnosis** — it reads logs, `/health`, `docker compose ps`,
`systemctl`, config, and ports to pin down problems like a node not joining
the console, an unreachable database, a down model backend, port conflicts,
or a JWT-secret mismatch — then hands you the precise remediation commands.
### Flags
| Flag | Purpose |
| File | Purpose |
|------|---------|
| `--dir PATH` | Install directory to inspect (default: current directory) |
| `--report` | Print the deterministic preflight report and exit — no LLM key needed |
| `--offline` | Skip the upstream GitHub version check |
| `.env` | All environment variables for `compose.yaml` |
| `setup.sh` | Post-start script: creates admin user, roles, tool policies, prompt templates via the API |
| `docker-compose.override.yaml` | Only if customizations beyond env vars are needed |
`--report` is the fastest way to get a health snapshot (and to share one when
asking for help) — it never needs an API key:
## Requirements
```bash
turnstone-doctor --report --dir ~/turnstone
```
- **Python 3.11+** with turnstone installed (`pip install turnstone`)
- **An LLM API key** — for the wizard itself (OpenAI, Anthropic, or a local
model). This can differ from the LLM your deployment will use.
- **Docker & Docker Compose** — needed to run the stack. The wizard detects
whether Docker is installed and gives platform-specific install instructions
if it's missing. You can still generate config files without Docker.
## Deployment Modes
- **Single-node production** — `docker compose up` against the bundled
`turnstone/deploy/compose.yaml`: 1 server + console + channel + PostgreSQL,
pulled from ghcr.io. Good for most deployments.
- **Local multi-node cluster** — clone the repo and run `docker compose up` at
the root for a 10-node fleet + console + Caddy + channel, built locally.
See [docs/docker.md](docs/docker.md) for both.
## Example Session
```
## Install profile
- Detected kind(s): docker-compose (primary: docker-compose)
- Docker daemon reachable: yes
- Compose files:
/home/you/turnstone/compose.yaml
- Database: backend=postgresql, url=postgresql+psycopg://turnstone:****@postgres:5432/turnstone
- Candidate health URLs: http://localhost:8080/health, http://localhost:8090/health
$ turnstone-bootstrap
## Versions
- Installed (this tool): 1.7.0a2
- Cluster nodes: 10 reporting; versions ['1.7.0a2']
- Version drift across nodes: no
- Upstream: stable 1.6.9, experimental 1.7.0a2
Turnstone Bootstrap Wizard v1.5.0
────────────────────────────────────────────────
## LLM backend (ok)
- resolved Qwen/Qwen3-32B via openai-compatible @ http://host.docker.internal:8000/v1
Which provider for this wizard?
[1] OpenAI
[2] Anthropic
[3] OpenAI-compatible (local/vLLM)
> 3
Base URL [http://localhost:8000/v1]:
API key (press Enter for 'none'):
Querying http://localhost:8000/v1 for available models...
Found model: Qwen/Qwen3-32B
Connected to Qwen/Qwen3-32B. Handing off to AI assistant...
> (AI walks you through the rest interactively)
```
Secrets (JWT secret, database password, API keys) are always redacted in the
report and in anything the doctor reads.
## Tips
- **Type `quit`** to exit the conversation; **Ctrl+C** interrupts (twice to quit).
- **Point it at the right install** with `--dir` when you run it from elsewhere.
- **(Re)installing or adding nodes?** Use the installer (`run.sh`), not the doctor.
- **Re-run safely** — running the wizard again detects your existing `.env`
and offers to update it rather than overwriting.
- **Duplicate writes are skipped** — if the LLM tries to write the same file
twice with identical content, it's silently ignored.
- **Type `quit` to exit** at any time during the conversation.
- **Ctrl+C** is handled gracefully — press once to interrupt, twice to exit.
## See Also
- [Docker Deployment](docs/docker.md) — compose stacks, ports, and bare-metal nodes
- [Docker Deployment](docs/docker.md) — manual compose setup and profiles
- [Security](docs/security.md) — auth architecture and token types
- [Governance](docs/governance.md) — roles, policies, and templates
+4 -28
View File
@@ -5,7 +5,6 @@
[![Python](https://img.shields.io/pypi/pyversions/turnstone)](https://pypi.org/project/turnstone/)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](LICENSE)
[![Discord](https://img.shields.io/badge/Discord-join%20us-5865F2?logo=discord&logoColor=white)](https://discord.gg/Nh3bWMacaq)
[![Sponsor](https://img.shields.io/badge/Sponsor-%E2%9D%A4-db61a2?logo=githubsponsors&logoColor=white)](https://github.com/sponsors/eous)
Self-hosted, local-first orchestration for tool-using AI agents. Give LLMs real tools — shell, files, search, web — and run them across your own cluster with direct HTTP routing and interactive interfaces. Your code, your models, your data stay on hardware you control: no telemetry, no phone-home.
@@ -15,20 +14,6 @@ Self-hosted, local-first orchestration for tool-using AI agents. Give LLMs real
Named after the [Ruddy Turnstone](https://en.wikipedia.org/wiki/Ruddy_turnstone) (*Arenaria interpres*) — a shorebird that flips stones to discover what's hiding underneath.
**What is a harness?**
<p align="center">
<a href="https://media.githubusercontent.com/media/turnstonelabs/turnstone/main/docs/diagrams/harness.png">
<img src="https://media.githubusercontent.com/media/turnstonelabs/turnstone/main/docs/diagrams/harness.png" alt=" : s_{n+1} ~ T(s_n) for n < τ_H — the whole controlled loop: π lowers state to context, M_W proposes a readout, γ authorizes it, Q_E acts on the world, ρ verifies and folds back" width="960"/>
</a>
</p>
```
: s_{n+1} ~ T(s_n) for n < τ_H
```
[**the primer →**](PRIMER.md) · [**the formalism →**](HYPOTHESIS.md)
### Release Tracks
| Track | Install | Docker | Description |
@@ -42,7 +27,7 @@ See [docs/releasing.md](docs/releasing.md) for the full release process.
Turnstone gives LLMs tools — shell, files, search, web, planning — and orchestrates multi-turn conversations where the model investigates, acts, and reports.
- **Local-first & private** — runs entirely on hardware you control, with no telemetry and no phone-home. Point it at local models (vLLM, llama.cpp) or commercial APIs you hold the keys to — your prompts and data never transit a third party you didn't choose.
- **Local-first & private** — runs entirely on hardware you control, with no telemetry and no phone-home. Point it at local models (vLLM, llama.cpp, Ollama) or commercial APIs you hold the keys to — your prompts and data never transit a third party you didn't choose.
- **Bring your own models** — OpenAI-compatible APIs (vLLM, llama.cpp, NIM), the Anthropic Messages API, and Google Gemini, mixed freely per role
- **Interactive sessions** — terminal CLI or browser UI with parallel workstreams
- **Cluster dashboard** — real-time view of every node and workstream, with a rendezvous routing proxy
@@ -101,7 +86,7 @@ LLM; add model backends from the console UI.
For production (released images from ghcr.io, real secrets required), use the
bundled stack: `docker compose -f turnstone/deploy/compose.yaml up`.
See [QUICKSTART.md](QUICKSTART.md) for the install + troubleshooting walkthrough and [docs/docker.md](docs/docker.md) for Docker configuration.
See [QUICKSTART.md](QUICKSTART.md) for the bootstrap wizard and [docs/docker.md](docs/docker.md) for Docker configuration.
### Programmatic (SDK)
@@ -131,9 +116,8 @@ Built-in tools for shell, files, search, web, memory, notifications, and autonom
| `turnstone-console` | Cluster dashboard + routing proxy + admin panel |
| `turnstone-channel` | Channel gateway (Discord and Slack adapters) |
| `turnstone-admin` | User/token management CLI |
| `turnstone-eval` | Headless measurement — scores tool-use against expected actions |
| `turnstone-optimizer` | Prompt/tool optimizer (UCB self-modify loop over the eval substrate) |
| `turnstone-doctor` | LLM-backed cluster diagnostics |
| `turnstone-eval` | Eval harness for prompt/tool optimization |
| `turnstone-bootstrap` | LLM-guided setup wizard |
### Diagrams
@@ -178,14 +162,6 @@ UML diagrams in [`docs/diagrams/`](docs/diagrams/):
- Optional: Discord / Slack channel integrations (`pip install turnstone[discord,slack]`)
- [Git LFS](https://git-lfs.com/) for cloning (diagram PNGs)
## Support
Turnstone is free, Apache-2.0, and self-hosted — no paid tier, no telemetry, no upsell. If it saves you time or you'd like to help keep development moving, you can sponsor the project:
**[❤ Sponsor Turnstone →](https://github.com/sponsors/eous)** · one-off via **[PayPal](https://paypal.me/eousphoros)**
Sponsorship is entirely optional and funds maintenance, new features, and infrastructure. Prefer to contribute in other ways? Filing issues, improving docs, and [pull requests](CONTRIBUTING.md) help just as much.
## Community
Questions, ideas, or want to show what you're building? Join us on Discord:
+16 -53
View File
@@ -29,11 +29,10 @@
# Fewer nodes (lighter machines):
# docker compose up postgres console caddy channel node-1 node-2 node-3
#
# Join a bare-metal host: a turnstone-server running OUTSIDE compose (e.g. to use
# a local GPU) can join this cluster. Postgres, the console's ACME endpoint, and
# SearxNG are published on 127.0.0.1 so a node on THIS machine reaches them via
# localhost. Keep secrets in ~/.config/turnstone/config.toml (chmod 0600 — the
# loader warns otherwise):
# Join a bare-metal host: Postgres is published on 127.0.0.1:5432, so a
# turnstone-server running directly on this machine (e.g. to use a local GPU)
# can join the same cluster. Keep the secret + connection settings in
# ~/.config/turnstone/config.toml (chmod 0600 — the loader warns otherwise):
# [auth]
# jwt_secret = "dev-only-insecure-jwt-secret-change-me-for-real-deployments"
# [database]
@@ -42,20 +41,10 @@
# [api]
# base_url = "http://localhost:8000/v1"
# api_key = "dummy"
# [tls] # only if the cluster runs mTLS
# enabled = true
# then run (node identity isn't a secret, so it stays on the command line):
# TURNSTONE_NODE_ID=host-1 \
# TURNSTONE_ADVERTISE_URL=http://host.docker.internal:8080 \
# TURNSTONE_CONSOLE_URL=http://localhost:8090 \
# TURNSTONE_SEARXNG_URL=http://localhost:8081 \
# TURNSTONE_NODE_ID=host-1 TURNSTONE_ADVERTISE_URL=http://host.docker.internal:8080 \
# turnstone-server --host 0.0.0.0 --port 8080
# The node registers in Postgres, auto-enrolls its mTLS cert from the console's
# ACME endpoint (when the cluster runs mTLS), and the console collector reaches
# it back via host.docker.internal. To join from ANOTHER machine, set
# TURNSTONE_HOST_IP to this host's LAN IP and set TURNSTONE_ACME_EXTERNAL_URL to
# http://<this-host-ip>:8090/acme. Use the same host IP in the node's URLs above
# (and the NODE host's IP in TURNSTONE_ADVERTISE_URL) — see docs/docker.md.
# It registers in Postgres and the console reaches it back via host.docker.internal.
# =============================================================================
name: turnstone
@@ -104,14 +93,13 @@ services:
# INSECURE dev default — override POSTGRES_PASSWORD in .env for real use.
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:-turnstone}
PGDATA: /var/lib/postgresql/data
# Published so a bare-metal turnstone-server can join the cluster (see "Join
# a bare-metal host" in the header). Bound to 127.0.0.1 by default (same-host
# nodes only); set TURNSTONE_HOST_IP to this host's LAN IP to let another
# machine connect — but set a real POSTGRES_PASSWORD first, or you'll expose a
# database with the insecure default password to your network. (The legacy
# POSTGRES_BIND is still honored as a fallback when TURNSTONE_HOST_IP is unset.)
# Published on localhost so a bare-metal turnstone-server running on THIS
# host can join the cluster (see "Join a bare-metal host" in the header).
# Bound to 127.0.0.1 by default; set POSTGRES_BIND=0.0.0.0 to let another
# machine connect — but set a real POSTGRES_PASSWORD first, or you'll expose
# a database with the insecure default password to your network.
ports:
- "${TURNSTONE_HOST_IP:-${POSTGRES_BIND:-127.0.0.1}}:${POSTGRES_PORT:-5432}:5432"
- "${POSTGRES_BIND:-127.0.0.1}:${POSTGRES_PORT:-5432}:5432"
volumes:
- postgres-data:/var/lib/postgresql/data
networks:
@@ -132,12 +120,10 @@ services:
# turnstone-console — cluster dashboard. Reach it ONLY through Caddy at
# https://localhost:8443 (see the caddy service below).
#
# Browsers must reach the dashboard through Caddy (https://localhost:8443): a
# plain HTTP/1.1 origin caps the browser at 6 connections, which starves the
# dashboard's per-pane SSE streams, whereas Caddy serves HTTP/2 (multiplexed)
# and proxies to console:8090 internally. The console's :8090 is published
# below ONLY so bare-metal nodes can reach the plain-HTTP ACME enrollment
# endpoint — don't point a browser at it.
# The console port (8090) is deliberately NOT published to the host: a plain
# HTTP/1.1 origin caps the browser at 6 connections, which starves the
# dashboard's per-pane SSE streams. Caddy serves the browser over HTTP/2
# (multiplexed) and proxies to console:8090 internally, so the cap is gone.
#
# The single `build:` here produces the turnstone:local image every other
# service reuses. extra_hosts lets the console reach a bare-metal server
@@ -152,24 +138,11 @@ services:
- turnstone-console
- --host=0.0.0.0
- --port=8090
# Publishes the console's plain-HTTP listener so a bare-metal node can reach
# the ACME endpoint, fetch the CA, and enroll its cert (the console serves
# HTTP here even under mTLS). Bound to 127.0.0.1 by default; setting
# TURNSTONE_HOST_IP exposes the WHOLE console HTTP API on that interface.
# ACME signing routes require a dedicated short-lived service JWT, but the
# listener and bearer token are still plain HTTP: bind only a trusted LAN
# or VPN interface and restrict it to enrolling nodes. Browsers use Caddy
# :8443, never this port.
ports:
- "${TURNSTONE_HOST_IP:-127.0.0.1}:8090:8090"
environment:
TURNSTONE_JWT_SECRET: *jwt-secret
TURNSTONE_DB_BACKEND: *db-backend
TURNSTONE_DB_URL: *db-url
TURNSTONE_CONSOLE_URL: http://console:8090
# Separate from TURNSTONE_CONSOLE_URL: this is the canonical responder
# base embedded in ACME directory/order URLs for cross-host enrollment.
TURNSTONE_ACME_EXTERNAL_URL: "${TURNSTONE_ACME_EXTERNAL_URL:-}"
extra_hosts:
- "host.docker.internal:host-gateway"
networks:
@@ -246,13 +219,6 @@ services:
# -------------------------------------------------------------------
searxng:
image: searxng/searxng:${SEARXNG_IMAGE_TAG:-latest}
# Published so a bare-metal node's web_search can reach it. SearxNG has NO
# auth, so it is bound to 127.0.0.1 by default; setting TURNSTONE_HOST_IP
# exposes it on that interface — an open search proxy on your LAN, which also
# triggers the SearxNG AGPL-3.0 §13 source-offer obligation (see docs/docker.md).
# In-compose nodes always use the internal http://searxng:8080 and ignore this.
ports:
- "${TURNSTONE_HOST_IP:-127.0.0.1}:${SEARXNG_API_PORT:-8081}:8080"
volumes:
- ./turnstone/deploy/searxng:/etc/searxng:ro
- searxng-cache:/var/cache/searxng # favicon + internal SQLite cache (survives restarts)
@@ -311,9 +277,6 @@ services:
SKIP_PERMISSIONS: ${SKIP_PERMISSIONS:-}
TURNSTONE_NODE_ID: node-1
TURNSTONE_ADVERTISE_URL: http://node-1:8080
# Lets the authenticated ACME client follow the console's canonical LAN
# URLs without trusting destinations learned from the public directory.
TURNSTONE_ACME_EXTERNAL_URL: "${TURNSTONE_ACME_EXTERNAL_URL:-}"
extra_hosts:
- "host.docker.internal:host-gateway"
networks:
+5 -9
View File
@@ -38,18 +38,10 @@ services:
condition: service_completed_successfully
volumes:
- tls-certs:/certs:ro
# The production base keeps :8090 private. The TLS overlay publishes it on
# localhost for same-host enrollment; use a trusted LAN/VPN address for a
# remote node and firewall it to that node.
ports:
- "${TURNSTONE_CONSOLE_HTTP_BIND:-127.0.0.1}:8090:8090"
environment:
TURNSTONE_TLS_ENABLED: "true"
TURNSTONE_TLS_SANS: "console"
TURNSTONE_CONSOLE_URL: "http://console:8090"
# Canonical ACME responder base advertised to enrolling nodes. Set this
# when they reach the console through a different host/address.
TURNSTONE_ACME_EXTERNAL_URL: "${TURNSTONE_ACME_EXTERNAL_URL:-}"
command:
- turnstone-console
- --host=0.0.0.0
@@ -66,7 +58,11 @@ services:
environment:
TURNSTONE_TLS_ENABLED: "true"
TURNSTONE_TLS_SANS: "server"
TURNSTONE_ACME_EXTERNAL_URL: "${TURNSTONE_ACME_EXTERNAL_URL:-}"
# Disable healthcheck — server serves HTTPS with mTLS which the
# stdlib healthcheck script can't satisfy. The base compose
# healthcheck uses plain HTTP which won't work on an HTTPS listener.
healthcheck:
disable: true
# Channel: TLS
channel:
+2 -2
View File
@@ -2,11 +2,11 @@ apiVersion: v2
name: turnstone
description: Multi-node AI orchestration platform with tool use, agent routing, and cluster simulation
type: application
version: 0.2.0
version: 0.1.0
appVersion: "0.3.0"
dependencies:
- name: postgresql
version: ~18.8.0
version: ~18.7.0
repository: https://charts.bitnami.com/bitnami
condition: postgresql.enabled
@@ -110,153 +110,6 @@ Determine the PostgreSQL username.
{{- end }}
{{- end }}
{{/*
The PostgreSQL password when the chart stores it itself, empty when it
does not. Doubles as the predicate for "does <fullname>-secrets need to
carry POSTGRES_PASSWORD", so an inline password is never written
anywhere but <fullname>-secrets, and an operator-supplied Secret is
never duplicated into it.
An operator-supplied existingSecret wins outright: writing the value
into a second Secret nothing reads would only duplicate a credential.
Both branches need "default" because this is reached through include,
which captures rendered text rather than a value: a key that is unset
rather than empty "password:" with nothing after it renders as the
literal "<no value>", and a ten-character string is truthy. Without the
default that lands base64-encoded in POSTGRES_PASSWORD and the workloads
authenticate with it.
*/}}
{{- define "turnstone.db.inlinePassword" -}}
{{- if .Values.postgresql.enabled }}
{{- .Values.postgresql.auth.password | default "" }}
{{- else if not .Values.database.external.existingSecret }}
{{- .Values.database.external.password | default "" }}
{{- end }}
{{- end }}
{{/*
The name of the bundled subchart's own Secret.
Mirrors the subchart's naming rather than calling its helpers, which
expect a context scoped to the subchart that this chart cannot hand
them. Release-derived, so deliberately not turnstone.fullname: a
fullnameOverride here renames this chart's resources and leaves the
subchart's alone, and pointing at "<fullname>-postgresql" would then
name a Secret that does not exist.
The subchart also normalises the release name through a regex before
using it, which is a no-op for the DNS-1123 names Helm accepts, so it is
not reproduced.
*/}}
{{- define "turnstone.postgresql.fullname" -}}
{{- $global := ((.Values.global).postgresql).fullnameOverride }}
{{- if $global }}
{{- $global | trunc 63 | trimSuffix "-" }}
{{- else if .Values.postgresql.fullnameOverride }}
{{- .Values.postgresql.fullnameOverride | trunc 63 | trimSuffix "-" }}
{{- else }}
{{- $name := .Values.postgresql.nameOverride | default "postgresql" }}
{{- if contains $name .Release.Name }}
{{- .Release.Name | trunc 63 | trimSuffix "-" }}
{{- else }}
{{- printf "%s-%s" .Release.Name $name | trunc 63 | trimSuffix "-" }}
{{- end }}
{{- end }}
{{- end }}
{{- define "turnstone.postgresql.secretName" -}}
{{- $existing := coalesce (((.Values.global).postgresql).auth).existingSecret .Values.postgresql.auth.existingSecret }}
{{- if $existing }}
{{- tpl $existing . }}
{{- else }}
{{- include "turnstone.postgresql.fullname" . }}
{{- end }}
{{- end }}
{{/*
The subchart stores the named user's password under "password" and the
superuser's under "postgres-password", and lets an operator rename
either through auth.secretKeys.
*/}}
{{- define "turnstone.postgresql.passwordKey" -}}
{{- $user := .Values.postgresql.auth.username | default "" }}
{{- $keys := .Values.postgresql.auth.secretKeys | default dict }}
{{- if or (empty $user) (eq $user "postgres") }}
{{- $keys.adminPasswordKey | default "postgres-password" }}
{{- else }}
{{- $keys.userPasswordKey | default "password" }}
{{- end }}
{{- end }}
{{/*
Determine the secret holding the PostgreSQL password, and the key within
it. Three sources, and the two helpers agree by construction because
they branch identically:
- an external database pointed at a Secret the chart does not own (a
CloudNativePG-generated secret, an External Secrets target, ...), in
which case the key is rarely "POSTGRES_PASSWORD" hence the
companion existingSecretPasswordKey
- the bundled subchart's own Secret, when it generates the password
- <fullname>-secrets, when the password is supplied inline in values
Note the last is deliberately not turnstone.llm.secretName: that
resolves to llm.existingSecret when the operator supplies one, which
holds LLM API keys and has no reason to carry a database password.
*/}}
{{- define "turnstone.db.secretName" -}}
{{- if not .Values.postgresql.enabled }}
{{- if .Values.database.external.existingSecret }}
{{- .Values.database.external.existingSecret }}
{{- else }}
{{- printf "%s-secrets" (include "turnstone.fullname" .) }}
{{- end }}
{{- else if include "turnstone.db.inlinePassword" . }}
{{- printf "%s-secrets" (include "turnstone.fullname" .) }}
{{- else }}
{{- include "turnstone.postgresql.secretName" . }}
{{- end }}
{{- end }}
{{- define "turnstone.db.passwordKey" -}}
{{- if not .Values.postgresql.enabled }}
{{- if .Values.database.external.existingSecret }}
{{- .Values.database.external.existingSecretPasswordKey | default "password" }}
{{- else }}
{{- printf "POSTGRES_PASSWORD" }}
{{- end }}
{{- else if include "turnstone.db.inlinePassword" . }}
{{- printf "POSTGRES_PASSWORD" }}
{{- else }}
{{- include "turnstone.postgresql.passwordKey" . }}
{{- end }}
{{- end }}
{{/*
Database environment shared by the server, console and migrate Job.
Every value except the password is rendered inline rather than pulled
from the ConfigMap via envFrom, so that one definition serves all three
workloads and the URL is assembled in exactly one place.
POSTGRES_PASSWORD must still precede TURNSTONE_DB_URL: the kubelet
expands $(VAR) only against env entries declared earlier in the list, so
a later definition would leave a literal "$(POSTGRES_PASSWORD)" in the
URL.
*/}}
{{- define "turnstone.db.env" -}}
- name: TURNSTONE_DB_BACKEND
value: {{ .Values.database.backend | quote }}
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: {{ include "turnstone.db.secretName" . }}
key: {{ include "turnstone.db.passwordKey" . }}
- name: TURNSTONE_DB_URL
value: "postgresql+psycopg://{{ include "turnstone.postgresql.username" . }}:$(POSTGRES_PASSWORD)@{{ include "turnstone.postgresql.host" . }}:{{ include "turnstone.postgresql.port" . }}/{{ include "turnstone.postgresql.database" . }}{{ if and (not .Values.postgresql.enabled) .Values.database.external.sslmode }}?sslmode={{ .Values.database.external.sslmode }}{{ end }}"
{{- end }}
{{/*
Determine the secret name for LLM API keys.
*/}}
@@ -7,17 +7,6 @@ metadata:
app.kubernetes.io/component: console
spec:
replicas: {{ .Values.console.replicas }}
{{- if eq (int .Values.console.replicas) 1 }}
# The console registers itself under the fixed service_id "console" and
# deregisters on shutdown. Under RollingUpdate the outgoing pod's
# deregister runs *after* the incoming pod registers and deletes its
# row -- and the heartbeat only touches last_heartbeat, so the row is
# never recreated and the console stays invisible in the registry until
# the next clean start. Recreate orders shutdown strictly before
# startup. Only valid at one replica; see console.replicas.
strategy:
type: Recreate
{{- end }}
selector:
matchLabels:
{{- include "turnstone.selectorLabels" . | nindent 6 }}
@@ -29,18 +18,6 @@ spec:
app.kubernetes.io/component: console
spec:
serviceAccountName: {{ include "turnstone.serviceAccountName" . }}
{{- with .Values.console.nodeSelector }}
nodeSelector:
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.console.affinity }}
affinity:
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.console.tolerations }}
tolerations:
{{- toYaml . | nindent 8 }}
{{- end }}
containers:
- name: console
image: {{ include "turnstone.image" . }}
@@ -59,18 +36,8 @@ spec:
- secretRef:
name: {{ include "turnstone.llm.secretName" . }}
optional: true
env:
{{- include "turnstone.db.env" . | nindent 12 }}
# Self-registration URL for the service registry. Unlike a
# server node the console is one logical endpoint behind its
# Service, so the Service DNS name is correct here. Without
# it the console registers gethostname() (its pod name),
# which no server node can resolve. Stops at ".svc" rather
# than assuming a "cluster.local" DNS domain, which is
# configurable per cluster.
- name: TURNSTONE_CONSOLE_URL
value: "http://{{ include "turnstone.fullname" . }}-console.{{ .Release.Namespace }}.svc:{{ .Values.console.service.port }}"
{{- if or .Values.auth.existingSecret .Values.auth.jwtSecret }}
env:
- name: TURNSTONE_JWT_SECRET
valueFrom:
secretKeyRef:
@@ -18,18 +18,6 @@ spec:
app.kubernetes.io/component: server
spec:
serviceAccountName: {{ include "turnstone.serviceAccountName" . }}
{{- with .Values.server.nodeSelector }}
nodeSelector:
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.server.affinity }}
affinity:
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.server.tolerations }}
tolerations:
{{- toYaml . | nindent 8 }}
{{- end }}
containers:
- name: server
image: {{ include "turnstone.image" . }}
@@ -51,20 +39,8 @@ spec:
name: {{ include "turnstone.llm.secretName" . }}
optional: true
env:
{{- include "turnstone.db.env" . | nindent 12 }}
# Each replica is a distinct node in the rendezvous ring, so it
# must advertise an address that reaches *itself*. The Service
# DNS name would load-balance across every replica, sending
# console traffic routed for node A to an arbitrary pod; the
# default (gethostname(), i.e. the pod name) is not resolvable
# at all. The pod IP is unique, routable in-cluster, and
# re-registered on every start, so churn is self-healing.
- name: POD_IP
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: TURNSTONE_ADVERTISE_URL
value: "http://$(POD_IP):{{ .Values.server.service.port }}"
- name: TURNSTONE_DB_URL
value: "postgresql+psycopg://$(TURNSTONE_DB_USER):$(POSTGRES_PASSWORD)@$(TURNSTONE_DB_HOST):$(TURNSTONE_DB_PORT)/$(TURNSTONE_DB_NAME)"
{{- if or .Values.auth.existingSecret .Values.auth.jwtSecret }}
- name: TURNSTONE_JWT_SECRET
valueFrom:
@@ -6,23 +6,11 @@ metadata:
{{- include "turnstone.labels" . | nindent 4 }}
app.kubernetes.io/component: migrate
annotations:
# post-install, not pre-install: on a first install nothing the
# migration needs exists yet — not the ConfigMap, not the Secret, and
# with the bundled subchart not the database either, since Helm
# creates ordinary resources only once hooks have finished. On an
# upgrade all of it is already running, so pre-upgrade is both safe
# and preferable: migrations land before the new code rolls out
# rather than after.
"helm.sh/hook": post-install,pre-upgrade
"helm.sh/hook": pre-install,pre-upgrade
"helm.sh/hook-weight": "-1"
"helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded
spec:
# Helm does not wait for the database to be ready before running
# post-install hooks, so on a first install this Job is what waits: it
# exits non-zero until PostgreSQL accepts connections, and the retry
# budget has to cover a cold StatefulSet pulling its image and
# initialising.
backoffLimit: 10
backoffLimit: 3
template:
metadata:
labels:
@@ -31,18 +19,6 @@ spec:
spec:
serviceAccountName: {{ include "turnstone.serviceAccountName" . }}
restartPolicy: OnFailure
{{- with .Values.migrate.nodeSelector }}
nodeSelector:
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.migrate.affinity }}
affinity:
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.migrate.tolerations }}
tolerations:
{{- toYaml . | nindent 8 }}
{{- end }}
containers:
- name: migrate
image: {{ include "turnstone.image" . }}
@@ -51,5 +27,12 @@ spec:
- python
- -m
- turnstone.core.storage._migrate
envFrom:
- configMapRef:
name: {{ include "turnstone.fullname" . }}-config
- secretRef:
name: {{ include "turnstone.llm.secretName" . }}
optional: true
env:
{{- include "turnstone.db.env" . | nindent 12 }}
- name: TURNSTONE_DB_URL
value: "postgresql+psycopg://$(TURNSTONE_DB_USER):$(POSTGRES_PASSWORD)@$(TURNSTONE_DB_HOST):$(TURNSTONE_DB_PORT)/$(TURNSTONE_DB_NAME)"
+7 -20
View File
@@ -1,19 +1,4 @@
{{/*
This Secret backs every credential supplied inline in values, so it is
rendered whenever any one of them is set — not, as it once was, only
when llm.existingSecret is empty. Under that older gate an operator who
supplied an LLM Secret lost the unrelated inline values with it: both
POSTGRES_PASSWORD and TURNSTONE_JWT_SECRET silently went unrendered
while the workloads went on referencing them, so every pod stalled in
CreateContainerConfigError.
Each key keeps its own condition, so an operator-supplied Secret still
suppresses the value it replaces and nothing else.
*/}}
{{- $apiKey := and .Values.llm.apiKey (not .Values.llm.existingSecret) }}
{{- $dbPassword := include "turnstone.db.inlinePassword" . }}
{{- $jwtSecret := and .Values.auth.jwtSecret (not .Values.auth.existingSecret) }}
{{- if or $apiKey $dbPassword $jwtSecret }}
{{- if not .Values.llm.existingSecret }}
apiVersion: v1
kind: Secret
metadata:
@@ -22,13 +7,15 @@ metadata:
{{- include "turnstone.labels" . | nindent 4 }}
type: Opaque
data:
{{- if $apiKey }}
{{- if .Values.llm.apiKey }}
OPENAI_API_KEY: {{ .Values.llm.apiKey | b64enc | quote }}
{{- end }}
{{- if $dbPassword }}
POSTGRES_PASSWORD: {{ $dbPassword | b64enc | quote }}
{{- if and .Values.postgresql.enabled .Values.postgresql.auth.password }}
POSTGRES_PASSWORD: {{ .Values.postgresql.auth.password | b64enc | quote }}
{{- else if and (not .Values.postgresql.enabled) .Values.database.external.password }}
POSTGRES_PASSWORD: {{ .Values.database.external.password | b64enc | quote }}
{{- end }}
{{- if $jwtSecret }}
{{- if and .Values.auth.jwtSecret (not .Values.auth.existingSecret) }}
TURNSTONE_JWT_SECRET: {{ .Values.auth.jwtSecret | b64enc | quote }}
{{- end }}
{{- end }}
-21
View File
@@ -14,13 +14,7 @@ database:
port: 5432
database: turnstone
username: turnstone
# Secret holding the password for `username`. Leave empty to supply
# `password` inline below instead.
existingSecret: ""
# Key within existingSecret holding the password. CloudNativePG
# generates "password"; other operators differ.
existingSecretPasswordKey: password
password: ""
sslmode: prefer
# -- Bitnami PostgreSQL subchart
@@ -43,10 +37,6 @@ server:
service:
type: ClusterIP
port: 8080
# -- Node scheduling constraints
nodeSelector: {}
affinity: {}
tolerations: []
# -- Turnstone console (cluster dashboard)
console:
@@ -61,17 +51,6 @@ console:
service:
type: ClusterIP
port: 8090
# -- Node scheduling constraints
nodeSelector: {}
affinity: {}
tolerations: []
# -- Database migration Job (post-install/pre-upgrade hook)
migrate:
# -- Node scheduling constraints
nodeSelector: {}
affinity: {}
tolerations: []
# -- LLM provider configuration
llm:
-87
View File
@@ -1,87 +0,0 @@
# Running a bare-metal turnstone-server under systemd
These units run a `turnstone-server` **outside** Docker (e.g. on a box with a
local GPU) so it joins an existing cluster — typically the docker-compose stack
in [`compose.yaml`](../../compose.yaml). They are the hardened, production-shaped
counterpart to the quick `turnstone-server …` invocation in
[`docs/docker.md`](../../docs/docker.md) ("Join a bare-metal host").
| File | Purpose |
|------|---------|
| `turnstone-server.service` | The hardened server unit (sandboxed; secrets via `config.toml`). |
| `turnstone.slice` | Shared memory/process budget for colocated Turnstone units. |
| `turnstone-server.service.d/node.conf.example` | Per-host identity + cluster URLs drop-in (no secrets). |
## Cluster-side prerequisite
The compose stack must publish Postgres, the console's ACME endpoint, and SearxNG
on an address the bare-metal host can reach. Use a trusted LAN or VPN interface,
firewall it to the joining node, and advertise the same reachable ACME endpoint
(default `127.0.0.1` keeps everything host-local):
```bash
TURNSTONE_HOST_IP=<compose-host-ip> \
TURNSTONE_ACME_EXTERNAL_URL=http://<compose-host-ip>:8090/acme \
docker compose up -d
```
## Install (run as root on the bare-metal host)
```bash
# 1. A dedicated, unprivileged user.
useradd --system --no-create-home --shell /usr/sbin/nologin turnstone
# 2. Install turnstone into a venv at /opt/turnstone-venv (lacme/mTLS is a core dep).
uv venv /opt/turnstone-venv --python 3.12
uv pip install --python /opt/turnstone-venv 'turnstone @ git+https://github.com/turnstonelabs/turnstone'
# …or from a local checkout: uv pip install --python /opt/turnstone-venv /path/to/turnstone
# 3. Secrets — match the cluster's JWT secret + DB credentials (kept out of env).
install -d -m 750 -o turnstone -g turnstone /etc/turnstone
cat > /etc/turnstone/config.toml <<'TOML'
[auth]
jwt_secret = "<same secret as the cluster>"
[database]
backend = "postgresql"
url = "postgresql+psycopg://turnstone:<password>@<compose-host-ip>:5432/turnstone"
[api]
base_url = "http://localhost:8000/v1" # a real model backend is configured in the console UI
api_key = "dummy"
TOML
chown turnstone:turnstone /etc/turnstone/config.toml
chmod 600 /etc/turnstone/config.toml
# 4. Units + per-host drop-in.
cp turnstone-server.service turnstone.slice /etc/systemd/system/
install -d /etc/systemd/system/turnstone-server.service.d
cp turnstone-server.service.d/node.conf.example \
/etc/systemd/system/turnstone-server.service.d/node.conf
$EDITOR /etc/systemd/system/turnstone-server.service.d/node.conf # set the addresses
# 5. Go.
systemctl daemon-reload
systemctl enable --now turnstone-server.service
journalctl -u turnstone-server -f # watch it register + (if the cluster runs mTLS) enroll
```
`tls.enabled` is **not** set here — a joining node inherits it from the cluster's
shared settings (the database). If the cluster runs mTLS, the node auto-enrolls a
cert from the console's ACME endpoint and re-advertises itself over `https://`.
For a node on a different host, `TURNSTONE_ACME_EXTERNAL_URL` is required on the
console and should also be set in the node drop-in. It is the full, externally
reachable responder base
(including `/acme`) that the console embeds in the ACME protocol's follow-up
URLs and that the node trusts as an enrollment-credential destination. The
node's `TURNSTONE_CONSOLE_URL` should point at the same host and port, without
the `/acme` suffix.
For mTLS, `TURNSTONE_ADVERTISE_URL` may use a resolvable DNS hostname or a
literal IP address. Turnstone enrolls literals as IP SANs. Bracket IPv6 literals
inside URLs, for example `http://[2001:db8::10]:8080`; do not use wildcard,
unspecified, or scoped addresses as certificate identities. Restart the node
after changing its advertised identity so it enrolls a matching certificate.
The dedicated service JWT authenticates enrollment but the direct `:8090`
bootstrap is still plain HTTP/TOFU. Use HTTPS through an independently trusted
proxy when the network itself is not trusted.
-85
View File
@@ -1,85 +0,0 @@
# Run a bare-metal turnstone-server as a systemd service so it joins a cluster
# (e.g. the docker-compose stack) from outside Docker — typically to use a local
# GPU. Install steps + the cluster-side prerequisites are in deploy/systemd/README.md
# and docs/docker.md ("Join a bare-metal host"). Per-host identity + the cluster
# URLs go in a drop-in (see node.conf.example); secrets go in config.toml.
[Unit]
Description=Turnstone server (chat workstreams + LLM gateway)
Documentation=https://github.com/turnstonelabs/turnstone
# Postgres is required. After= orders against a colocated postgresql.service
# when present and silently no-ops otherwise (the cluster DB is usually remote).
After=network.target postgresql.service
StartLimitIntervalSec=60
StartLimitBurst=5
[Service]
Type=exec
User=turnstone
Group=turnstone
# Secrets live in config.toml — JWT secret, Postgres URL+password, LLM API key —
# kept out of os.environ so a prompt-injected tool can't dump them via `env`.
Environment=TURNSTONE_CONFIG=/etc/turnstone/config.toml
Environment=TURNSTONE_LOG_LEVEL=info
Slice=turnstone.slice
# Per-host node identity + cluster wiring (TURNSTONE_NODE_ID / _ADVERTISE_URL /
# _CONSOLE_URL / _SEARXNG_URL) go in a drop-in, not here — see node.conf.example.
StateDirectory=turnstone
StateDirectoryMode=0750
LogsDirectory=turnstone
LogsDirectoryMode=0750
WorkingDirectory=/var/lib/turnstone
# --host 0.0.0.0 so the console collector + peer nodes can dial this node back
# at its advertised address. (A single-node, Caddy-fronted install can use
# 127.0.0.1 instead.) Rewrite --port if :8080 is already taken on the host.
ExecStart=/opt/turnstone-venv/bin/turnstone-server --host 0.0.0.0 --port 8080
Restart=on-failure
RestartSec=5s
TimeoutStartSec=120
TimeoutStopSec=30
KillSignal=SIGTERM
KillMode=mixed
# --- Resource limits ---
# SSE keeps an fd per active workstream + outbound LLM stream + MCP stdio pipe.
LimitNOFILE=65535
LimitNPROC=8192
TasksMax=8192
LimitCORE=0
# --- Hardening ---
NoNewPrivileges=true
CapabilityBoundingSet=
AmbientCapabilities=
UMask=0027
PrivateTmp=true
# PrivateDevices=true — disabled: GPU access via /sys/class/drm
ProtectSystem=strict
ProtectHome=true
ProtectKernelTunables=true
ProtectKernelModules=true
ProtectKernelLogs=true
ProtectControlGroups=true
ProtectClock=true
ProtectHostname=true
RestrictNamespaces=true
RestrictRealtime=true
RestrictSUIDSGID=true
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6
LockPersonality=true
MemoryDenyWriteExecute=true
SystemCallArchitectures=native
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @mount
StandardOutput=journal
StandardError=journal
SyslogIdentifier=turnstone-server
[Install]
WantedBy=multi-user.target
@@ -1,33 +0,0 @@
# Per-host node identity + cluster wiring for a bare-metal turnstone-server.
# Copy to /etc/systemd/system/turnstone-server.service.d/node.conf and edit the
# addresses, then `systemctl daemon-reload`. Identity + URLs are NOT secrets, so
# they live here; the JWT secret + DB URL live in /etc/turnstone/config.toml.
#
# Addresses below use RFC 5737 documentation IPs — replace them:
# <this-host> = the bare-metal host's own reachable DNS name or IP address
# (its mTLS identity and the address peers dial)
# <compose-host> = the host running the cluster / docker-compose stack, started
# with TURNSTONE_HOST_IP=<compose-host> and
# TURNSTONE_ACME_EXTERNAL_URL=http://<compose-host>:8090/acme
# so enrollment links and published ports are reachable
# (see docs/docker.md).
[Service]
# Unique node id (defaults to the hostname if unset).
Environment=TURNSTONE_NODE_ID=host-1
# The address peers + the console collector dial back. Auto-upgrades to https://
# once the node enrolls its mTLS cert. IPv6 literals require URL brackets, for
# example http://[2001:db8::10]:8080.
Environment=TURNSTONE_ADVERTISE_URL=http://192.0.2.10:8080
# The cluster console's reachable plain-HTTP ACME/API endpoint. A bare-metal node
# can't resolve the in-cluster name (console:8090), so point it at the published
# port; turnstone-server honors this for cert enrollment.
Environment=TURNSTONE_CONSOLE_URL=http://192.0.2.1:8090
# Trusted canonical responder base. This pins where the node may send its
# enrollment JWT; it must match the console-side value (a literal IP is fine).
Environment=TURNSTONE_ACME_EXTERNAL_URL=http://192.0.2.1:8090/acme
# The cluster's published SearxNG, for the web_search tool.
Environment=TURNSTONE_SEARXNG_URL=http://192.0.2.1:8081
-15
View File
@@ -1,15 +0,0 @@
# Shared resource budget for the colocated Turnstone units. Without a slice each
# unit's MemoryMax= is enforced independently — three units at 85% each can sum
# to 255% of host RAM before any throttles. Under a shared slice the cap is
# hierarchical: the slice ceiling is the real limit. (A bare-metal node that runs
# only turnstone-server still benefits — and keeps the unit's Slice= reference
# valid.) Adjust if the host runs other meaningful workloads alongside Turnstone.
[Unit]
Description=Turnstone services slice (server + console + channel)
Documentation=https://github.com/turnstonelabs/turnstone
Before=slices.target
[Slice]
MemoryHigh=70%
MemoryMax=85%
TasksMax=16384
+160 -718
View File
File diff suppressed because it is too large Load Diff
+261 -858
View File
File diff suppressed because it is too large Load Diff
+25 -12
View File
@@ -12,6 +12,7 @@ Existing bulk endpoints at time of writing:
|---------------------------------------------------------|--------------------------|------------------------------------------|
| `GET /v1/api/cluster/ws/live?ids=a,b,c` | bulk read | `{results, denied, truncated}` |
| model tool `spawn_batch` | bulk create (per-item) | `{results, denied}` |
| `POST /v1/api/workstreams/{ws_id}/stop_cascade` | cascade mutation | `{cancelled, failed, skipped}` |
| `POST /v1/api/workstreams/{ws_id}/close_all_children` | cascade mutation | `{closed, failed, skipped}` |
---
@@ -145,7 +146,7 @@ consistently-typed across the read and create cases.
```
Where `<bucket>` is the endpoint-specific name for "succeeded" —
`closed` for `close_all_children`.
`cancelled` for `stop_cascade`, `closed` for `close_all_children`.
The three buckets partition the input set exactly once:
| Bucket | Meaning |
@@ -160,6 +161,20 @@ be partial. `skipped` is pre-resolved — the target is already in
the terminal state the cascade was aiming at, so it's neither a
win to report nor a fault to fix.
### Example — `stop_cascade`
```json
{
"status": "ok",
"cancelled": ["child-1", "child-3"],
"failed": [],
"skipped": ["child-2"]
}
```
A subsequent retry would target only `failed` ids, not `skipped`
ones — the latter are already done.
### Example — `close_all_children`
```json
@@ -171,12 +186,10 @@ win to report nor a fault to fix.
}
```
Here the success bucket is `closed`. A subsequent retry would
target only `failed` ids, not `skipped` ones — the latter are
already done. When `coord_client` is unavailable (session loaded
but no HTTP client attached — a construction bug) every id goes to
`failed` so the operator notices rather than getting a silent
all-skipped response.
Same partition, different success-bucket name. When `coord_client`
is unavailable (session loaded but no HTTP client attached — a
construction bug) every id goes to `failed` so the operator notices
rather than getting a silent all-skipped response.
---
@@ -219,12 +232,12 @@ all-skipped response.
- **Phase 6** shipped `cluster/ws/live` as the first Shape A endpoint
(`{results, denied, truncated}`).
- **Phase 7** introduced the Shape B cascade-mutation envelope
(`{<bucket>, failed, skipped}`) for the coordinator's
cancel-cascade path.
- **Phase 7** shipped `stop_cascade` as the first Shape B endpoint
(`{cancelled, failed, skipped}`).
- **Phase 8 PR A** shipped `spawn_batch` (Shape A, keyed by idx) and
`close_all_children` (Shape B), which crystallised the
two-shape-per-semantic-category policy codified here.
`close_all_children` (Shape B, twin of `stop_cascade`), which
crystallised the two-shape-per-semantic-category policy codified
here.
Before adding a third shape, read this doc and argue for why the
new surface doesn't fit either A or B. Two idioms in the cluster
+16 -22
View File
@@ -195,17 +195,13 @@ both and the gateway hosts both adapters in one process.
- All subsequent messages in the thread are routed to the same workstream.
- The bot streams responses via message edits, updated approximately every
1.5 seconds.
- If a persisted channel route is no longer active on its owning node, the
router asks the create endpoint to fork the old workstream into a new ID via
`resume_ws`. The saved source can still resolve normally; its
checkpoint-bounded history, configuration, persona, effective project, and
attachment references are cloned before the channel route is repointed. The
old route remains durable until the replacement (and any initial message)
succeeds. If the create endpoint returns the ordinary
source-not-found response *and* a fresh authoritative storage lookup confirms
that the source is gone, the router retries once without `resume_ws` and
starts a fresh conversation. Other access, conflict, routing, and storage
failures remain visible rather than silently discarding history.
- If the workstream is evicted for capacity, the next message in the
thread auto-creates a new workstream and atomically resumes the
previous workstream via the `resume_ws` field on
`CreateWorkstreamMessage`. The server resumes the workstream during
creation (same HTTP request), and the server emits a
`WorkstreamResumedEvent` back to the channel. The thread receives a
*"Resumed: {name} ({count} messages restored)"* confirmation.
### Slash Commands
@@ -288,17 +284,15 @@ See [Security: Database Schema](security.md#database-schema) for the
`channel_routes` table.
2. **Active** — messages are routed bidirectionally. The bot streams
responses via message edits (updated every ~1.5 seconds).
3. **Eviction** — the server evicts an idle workstream for capacity. Its saved
source row and channel route remain durable, and the thread stays open.
4. **Reactivation** — the next message resolves the saved route and probes
whether that workstream is live on its owning node. If it is not, the router
creates a distinct workstream with the old `ws_id` as `resume_ws`. The
create response confirms the fork and message count; there is no separate
resume command or channel-specific resumed event. Only after the replacement
succeeds does the router swap the persisted route. If the source was deleted
or pruned, an exact source-not-found response plus a second authoritative
storage miss triggers one fresh-create retry; other fork failures leave the
old route intact and are surfaced normally.
3. **Eviction** — the server evicts an idle workstream for capacity. The
route is preserved and the thread stays open.
4. **Reactivation** — the next message in the thread detects the stale
route and creates a new workstream with the old `ws_id`
as `resume_ws` on the creation request. The server resumes
the workstream during creation (no separate command or reverse lookup
needed). The channel receives a `WorkstreamResumedEvent`, and
the thread displays *"Resumed: {name} ({count} messages restored)"*.
If the old workstream was pruned, a fresh one starts with no error.
5. **Close**`/close` command closes the workstream via HTTP, deletes the
route, unsubscribes from events, and archives the Discord thread.
+39 -111
View File
@@ -174,32 +174,17 @@ Request:
{
"node_id": "db-west-04",
"name": "perf-analysis",
"model": "gpt-5",
"project_id": "proj_analytics",
"initial_message": "Profile the slow query"
"model": "gpt-5"
}
```
All fields are optional:
- `node_id` — targeting mode:
- **omitted or `"auto"`** — console picks the reachable node with the most available capacity (max_ws - ws_total) and proxies the request to it.
- **`"pool"`** — compatibility alias for automatic placement on the reachable node with the most headroom.
- **`"pool"`** — console picks a reachable node with available capacity using round-robin selection.
- **specific node ID** — proxies the request to that node directly.
- `name` — workstream display name. Auto-generated if omitted.
- `model` — model alias from the target node's registry. Uses the node's default model if omitted.
- `judge_model` — optional judge-model alias for this workstream.
- `initial_message` — first message dispatched after the workstream is published.
- `skill` — enabled profile/skill to snapshot onto a fresh workstream.
- `persona` — enabled persona slug; empty uses the interactive default.
- `project_id` — project to attach, subject to the target node's membership gate.
- `resume_ws` — source ID to **fork** atomically into a new workstream. The
source remains unchanged; its checkpoint-bounded history, configuration,
persona, project, and attachment references are copied transactionally.
The endpoint also accepts the same multipart create shape as a node: one
JSON-encoded `meta` field plus up to ten `file` parts. Files require an
`initial_message` in the dashboard launcher. Files cannot be combined with
`resume_ws`; fork first and upload on the new workstream.
Response:
@@ -211,19 +196,7 @@ Response:
}
```
The response is returned only after the target node has durably published the
workstream. Its hidden `creating` reservation has already crossed to `idle`,
and the node emitted `ws_created` before any initial-message state event. The
cluster SSE event may therefore arrive before or after the HTTP response;
clients should reconcile both by the returned `correlation_id`/workstream ID
rather than treating them as two creates.
For safety, the console masks most target-node failures as the opaque `502`
shape `{"error":"Dispatch to node <node_id> failed"}` instead of reflecting
arbitrary node text or retry-triggering 401/429 responses. The coded
`server.require_project` refusal is the exception and remains a `400` with
actionable wording. Consult the target node's logs for the underlying create
correlation when a reachable node returns a masked 502.
The response confirms the workstream creation request was proxied to the target node. A `ws_created` event on the cluster SSE stream confirms the workstream was actually created.
### `GET /v1/api/cluster/events`
@@ -337,8 +310,8 @@ The auth system uses three scopes instead of the earlier read/full role model:
| Scope | Grants |
|-------|--------|
| `read` | Read-only access: dashboards, workstream lists, SSE streams, health |
| `write` | Non-approval mutations: send, create/open/close/delete, cancel, attachments, rewind, and retry |
| `approve` | Tool-approval and admin HTTP surfaces (with their additional RBAC permission checks) |
| `write` | Send messages, create/close workstreams, approve tool calls |
| `approve` | Admin operations: manage users and API tokens |
Scopes are cumulative — a user with `approve` scope can also perform `write` and `read` operations.
@@ -375,106 +348,64 @@ SSE streams (`/v1/api/workstreams/{ws_id}/events`, `/v1/api/events/global`) are
### Authentication
The proxy mints a short-lived (5-minute) JWT per request carrying the real user's `user_id`, `scopes`, and `permissions` with `aud: turnstone-server`. The user's console JWT (`aud: turnstone-console`) cannot be forwarded directly — it would be rejected by the server's audience validation — so the console re-signs a new server-audience JWT from the validated `AuthResult`. This preserves audit attribution (the upstream server sees the real user, not a service identity) and enforces scope narrowing as defense in depth (a read-only console user's proxied request carries only `read` scope). Ordinary users are re-minted with `src="console-proxy"`; coordinator tokens retain `src="coordinator"` plus `coord_ws_id`, and only the validated console service identity with `service` scope retains `src="console"` for trusted owner forwarding. When no user context is available, the proxy falls back to a `ServiceTokenManager` identity `console-proxy` carrying `src="console"` and `{read, write, approve, service}` scopes. The static `--auth-token` / `proxy_auth_token` is used as a final fallback.
The proxy mints a short-lived (5-minute) JWT per request carrying the real user's `user_id`, `scopes`, and `permissions` with `aud: turnstone-server`. The user's console JWT (`aud: turnstone-console`) cannot be forwarded directly — it would be rejected by the server's audience validation — so the console re-signs a new server-audience JWT from the validated `AuthResult`. This preserves audit attribution (the upstream server sees the real user, not a service identity) and enforces scope narrowing as defense in depth (a read-only console user's proxied request carries only `read` scope). The JWT `src` claim is set to `"console-proxy"` for audit traceability. When no user context is available (auth disabled), the proxy falls back to a `ServiceTokenManager` with service identity `console-proxy`. The static `--auth-token` / `proxy_auth_token` is used as a final fallback.
---
## Browser Dashboard
The console uses an L-shaped application shell: a collapsible navigation rail,
a tab bar, and a pane host. On mobile the rail becomes an off-canvas drawer.
The rail is fed by the cluster SSE snapshot and shows:
The web UI has five views, toggled client-side:
- state/count filters and the live compute-node list, including version drift;
- active coordinator and interactive workstreams, nested under their
coordinator parent and grouped by project when project metadata is visible;
- permission-filtered Manage groups that open the singleton Admin pane.
### 1. Cluster Overview (landing)
Coordinator and interactive conversations open as tabs inside the same shell.
Interactive panes use the owning node's console proxy, so users do not need
direct network access to compute-node ports. Split-right and split-down actions
can display several panes at once. Closing a pane removes only that tab; use the
pane menu's explicit close or delete action to change the workstream lifecycle.
- **State cards** — 5 clickable cards (running, thinking, attention, idle, error) with count and colored top border. Clicking filters to that state.
- **Aggregate bar** — total tokens and tool calls across the cluster.
- **Node table** — columns: NODE, WS, RUN, ATTN, TOKENS, VER, LOAD. Sorted by activity. Clickable rows drill down to node detail. Version column shows per-node version; hidden on mobile.
- **Version drift indicator** — when nodes report different versions, the status bar shows a yellow "DRIFT" warning with a tooltip listing all versions. Node groups show "mixed" with a yellow badge when their members disagree.
- **"+ new" button** — opens the workstream creation modal (see below).
### Dashboard pane
### 2. Node Drill-down
The home view is coordinator-first. It contains the persistent workstream
launcher plus the saved-sessions list. Selecting a state count opens the
filtered workstream table inside the same Dashboard pane; selecting a compute
node opens its proxied node surface. Cluster SSE updates keep rail state,
workstream rows, and tab state glyphs synchronized.
Breadcrumb: `Cluster > db-west-04`. Shows the node's workstreams in a table matching the per-node dashboard layout (STATE, NAME, MODEL, NODE, TASK, TOKENS, CTX) with activity sub-lines. Includes a link to the node's proxied server UI.
### Workstream launcher
**Proxy deep-linking:** Clicking a workstream row opens the node's server UI in a new tab via the proxy at `/node/{node_id}/?ws_id=<id>`, which auto-selects that workstream. Users do not need direct network access to the server node.
The landing-page composer starts a workstream with an optional initial task and
attachments. When the caller can create both kinds, a Coordinator / Interactive
toggle selects the target kind. Its options include:
### 3. Filtered Workstreams
- **Node placement** — "Least loaded" picks the reachable node with the most
headroom, or "Specific node" pins the create to a node from the live list.
- **Persona** — optional dropdown listing the enabled personas for the workstream kind. Sets the system-message composition and capability envelope at creation, snapshotted server-side; empty uses the kind's default. Picking one requires no `persona.*` permission.
- **Skill** — optional dropdown listing enabled skills. Applies the skill's model, auto-approve policy, token budget, and other behavioral settings at creation time.
- **Project** — optional project filing. Private projects require owner/member access. A coordinator child inherits its parent's project unless explicitly routed to another attachable project.
Breadcrumb: `Cluster > Running` or `Cluster > db-west-04`. Server-side paginated workstream table. NODE column values are clickable to filter further. Pagination controls at bottom. Workstream rows use proxy deep-links.
### 4. Workstream Creation Modal
Triggered by the "+ new" header button. A modal dialog with:
- **Node selector** — dropdown with three targeting modes: "Auto (best available)" picks the node with the most headroom, "General pool (any node)" picks a node with available capacity using round-robin, or a specific node from the list (showing capacity).
- **Profile** — optional dropdown listing enabled skills. Applies the skill's model, auto-approve policy, token budget, and other behavioral settings at creation time.
- **Name** — optional text input. Auto-generated if left empty.
- **Model** — optional selector populated from the target model registry.
- **Judge Model** — optional selector for the judge alias (overrides the default
judge model for this workstream).
- **Model** — optional text input for a model alias from the target node's registry.
- **Judge Model** — optional text input for the judge model alias (overrides the default judge model for this workstream).
Interactive launches additionally expose node strategy / node selection.
Submitting uses `POST /v1/api/cluster/workstreams/new`; coordinator launches use
the console's coordinator create surface. A toast confirms the committed
create, while SSE updates the dashboard and opens the resulting pane.
Keyboard shortcuts: Ctrl+Shift+R (refresh title), Ctrl+Shift+E (edit title), Ctrl+Shift+F (fork), Ctrl+Shift+X (delete). Press ? for full shortcut help.
Files require a non-empty initial task so the first turn consumes the staged
attachments. The console shell does not currently expose a fork action; use the
node's standalone workstream UI or the create API's `resume_ws` field.
On submit, `POST /v1/api/cluster/workstreams/new` dispatches the creation request. A toast confirms success; the SSE stream delivers the `ws_created` event to update the dashboard.
### Large pasted text
All five views receive live updates via SSE — state cards update counts, node rows update metrics, workstream rows update state indicators.
Browser composers turn plain text longer than 2,000 Unicode code points into a
`text/plain` attachment named `pasted-text.txt`. A paste exactly at the
threshold stays inline. This applies to the interactive and coordinator send
boxes, the console home launcher, and the node dashboard and new-workstream
composers.
The browser maintains a local `clusterState` object that mirrors the cluster snapshot. It is initialized from the SSE `snapshot` event on connect (or via `GET /v1/api/cluster/snapshot` on initial page load) and updated incrementally by SSE events. View navigation reads from local state — no API round-trips needed after the initial snapshot.
Clipboard files take priority over clipboard text. Text larger than the 512 KiB
attachment ceiling also stays inline, so the browser does not discard it before
a rejected upload. Attachments require a companion message and cannot be sent
as live-turn interjections; a busy composer preserves its message and chips for
an idle retry.
### Saved and filtered sessions
Saved coordinator and interactive sessions share one list with kind and persona
labels, filtering, pagination, and multi-select deletion. Opening a saved
coordinator rehydrates it in the console; opening a saved interactive session
resolves its node, calls `open`, and then connects the node-proxied pane.
The filtered live table carries STATE, NAME, MODEL, NODE, TASK, TOKENS, and CTX
columns. The browser maintains a local `clusterState` initialized from the
cluster snapshot and updated incrementally by SSE; the filtered view normally
renders from that state without another API round trip.
### Admin pane
### 5. Admin Panel
Accessed via the "admin" button in the header (visible when authenticated
with `approve` scope). Provides user, API token, channel link, MCP server,
and skill management with tabs that include Users, API Tokens, Channels,
Schedules, Watches, Personas, Roles, Policies, Prompts, Judge, Skills,
MCP Servers, Usage, Audit, Memories, Models, Nodes, Settings, and TLS. See also
and skill management with 18 tabs (Users, API Tokens, Channels, Schedules,
Watches, Roles, Policies, Prompts, Judge, Skills, MCP Servers, Usage,
Audit, Memories, Models, Nodes, Settings, TLS). See also
[Governance](governance.md) for the Roles, Policies, Skills, Usage, and
Audit tabs, and [Settings](settings.md) for the database-backed
configuration editor.
The **Channels** tab links users to either a Discord or Slack account
via a per-row channel-type selector. The **Models** tab is a CRUD
editor for `model_definitions`, including static and dynamic backend-auth
modes and a per-process **Max concurrent generations** limit for each alias
(`0` means unlimited). The limit is shared by every model-backed role using
that alias and a streaming generation holds its slot through the full decode.
Model edits rebind existing workstreams at their next send while
in-flight requests keep their original definition snapshot; see
[Settings](settings.md#model-definition-reloads) for the full contract. The **Nodes** tab edits per-node
via a per-row channel-type selector. The **Models** tab is a CRUD
editor for `model_definitions`, the **Nodes** tab edits per-node
metadata, and the **TLS** tab manages CA and leaf certificates for the
internal mTLS fabric. The **Settings** tab edits ConfigStore values
live; edits apply without restart.
@@ -572,7 +503,7 @@ Run history is automatically pruned (runs older than 90 days) approximately once
| Mode | Behavior |
|------|----------|
| `auto` | Picks the reachable node with the most available capacity |
| `pool` | Compatibility alias for the reachable node with the most headroom |
| `pool` | Picks a reachable node with available capacity using round-robin |
| `all` | Fan-out to all reachable nodes (capped at `max_fan_out`, default 20) |
| `<node_id>` | Targets a specific node by ID |
@@ -732,7 +663,4 @@ turnstone-server --port 8080
turnstone-console --port 8090
```
Open `http://localhost:8090` for the cluster dashboard. Create workstreams from
the persistent Dashboard launcher. Selecting a workstream opens a coordinator
or node-proxied interactive pane in the console shell — no direct access to
server ports is required.
Open `http://localhost:8090` for the cluster dashboard. Create workstreams via the "+ new" button. Click any workstream to open the proxied server UI — no direct access to server ports required.
+55 -78
View File
@@ -18,7 +18,7 @@ schema changes.
> auth and the `admin.coordinator` permission. A session-scoped JWT
> is minted per login (see [docs/oidc.md](oidc.md) / [docs/security.md](security.md));
> a service token may call the read paths but destructive governance
> paths (`/restrict`, `/close_all_children`) require
> paths (`/restrict`, `/stop_cascade`, `/close_all_children`) require
> the explicit `admin.coordinator` grant — a service-token owner
> match isn't enough.
@@ -37,13 +37,14 @@ schema changes.
| # | Action | Operation |
|---|------------------------------|-------------------------------------------------------------|
| 1 | Create | `POST /v1/api/workstreams/new` |
| 2 | Bootstrap history + subscribe | `GET .../history`, then `GET .../events` (SSE) |
| 2 | Subscribe to events | `GET /v1/api/workstreams/{ws_id}/events` (SSE) |
| 3 | Send a user message | `POST /v1/api/workstreams/{ws_id}/send` |
| 4 | Inspect children | `GET /v1/api/workstreams/{ws_id}/children` |
| 5 | Inspect one workstream | `GET /v1/api/cluster/ws/{ws_id}/detail` |
| 6 | Wait for fan-out | model-side tool `wait_for_workstream` |
| 7 | Govern | `POST /v1/api/workstreams/{ws_id}/trust` |
| | | `POST /v1/api/workstreams/{ws_id}/restrict` |
| | | `POST /v1/api/workstreams/{ws_id}/stop_cascade` |
| | | `POST /v1/api/workstreams/{ws_id}/close_all_children` |
| 8 | Approve / cancel | `POST /v1/api/workstreams/{ws_id}/approve` |
| | | `POST /v1/api/workstreams/{ws_id}/cancel` |
@@ -52,7 +53,7 @@ schema changes.
Refer to `/openapi.json` (Swagger UI at `/docs`) on any
`turnstone-console` process for the authoritative operation ids and
schemas. Coordinator-only verbs (`/children`, `/trust`, `/restrict`,
`/close_all_children`) 404 against `kind=interactive`
`/stop_cascade`, `/close_all_children`) 404 against `kind=interactive`
rows; the shared verbs (`/send`, `/approve`, `/cancel`, `/events`,
`/history`, `/open`, `/close`, etc.) work on both kinds.
@@ -91,34 +92,14 @@ subscribers (step 2) see the session warm up as token traffic starts.
---
## 2. Bootstrap history, then subscribe to the event stream
Read and render history before opening the initial stream:
## 2. Subscribe to the per-coordinator event stream
```http
GET /v1/api/workstreams/{ws_id}/history?limit=100 HTTP/1.1
Authorization: Bearer <token>
```
For a loaded coordinator, `messages` is the requested tail of one total
accepted conversation-row prefix: user, assistant, tool, and system rows,
including projected compaction checkpoints and cancellation-generated markers.
The response's optional `cursor` and `handoff_token` belong to that exact
render. Pass both once on the initial stream URL:
```http
GET /v1/api/workstreams/{ws_id}/events?last_event_id={cursor}&history_token={handoff_token}&user_turn=1&tool_turn=1 HTTP/1.1
GET /v1/api/workstreams/{ws_id}/events HTTP/1.1
Accept: text/event-stream
Authorization: Bearer <token>
```
Omit either query parameter when its history field is `null`. A handoff token
is opaque and process-local: do not parse, persist, or reuse it. Admission of a
later conversation row changes the token; durable acknowledgement does not. If
history returns `503 {"error":"History temporarily unavailable"}`, the response
is not authoritative: retain the current transcript, do not open a tokenless
replacement stream, and retry the read.
One persistent SSE connection per browser tab / SDK caller — the
console fans each event out to every listener queue (cap 500 events
per queue, put_nowait drop on overflow). Events come in flat JSON
@@ -130,10 +111,10 @@ with a `type` field. The recurring shapes a UI has to handle:
| `reasoning` | Reasoning-token stream chunk (when the model exposes it) | `text` |
| `content` | Assistant-content stream chunk | `text` |
| `stream_end` | End of a single provider stream | — |
| `tool_result` | A tool call completed; capable panes also receive the accepted-history replacement | `call_id`, `name`, `output`, `is_error?`, `accepted?`, `_event_id?`, `preview?`, `effect_status?` |
| `tool_result` | A tool call completed (success or error) | `call_id`, `name`, `output`, `is_error?` |
| `tool_output_chunk` | Streaming tool output (e.g. long bash command) | `call_id`, `chunk` |
| `approve_request` | One approval cycle needs operator action; several cycles may coexist | `cycle_id`, `items: [{call_id, header, preview, func_name, approval_label, needs_approval}]` |
| `approval_resolved` | One identified approval cycle was answered | `cycle_id`, `call_ids`, `approved`, `feedback`, `always` |
| `approve_request` | One or more tool calls need operator approval | `items: [{call_id, header, preview, func_name, approval_label, needs_approval}]` |
| `approval_resolved` | Operator answered the approval prompt | `approved`, `feedback` |
| `state_change` | Worker-thread state transition (also re-emitted with the current state on every fresh subscribe so refresh-mid-stream restores composer mode) | `state``running`, `thinking`, `attention`, `idle`, `error` |
| `in_progress_snapshot` | One-shot replay of the in-progress turn's content + reasoning when this client connects mid-stream | `content`, `reasoning` |
| `status` | Token usage + context-window snapshot (fires on every streaming tick) | `prompt_tokens`, `completion_tokens`, `total_tokens`, `context_window`, `pct`, `effort`, `cache_creation_tokens`, `cache_read_tokens` |
@@ -147,11 +128,10 @@ with a `type` field. The recurring shapes a UI has to handle:
| `wait_started` / `wait_progress` / `wait_ended` | `wait_for_workstream` tool lifecycle (see §6) | `call_id`, `ws_ids`, `elapsed`, `results`, `complete` |
| `batch_started` / `batch_ended` | `spawn_batch` / `close_all_children` tool lifecycle | `call_id`, `op`, `total`/`succeeded`/`denied`/`closed`/`failed`/`skipped` |
| `info` / `error` | Operational messages | `message` |
| `history_resync` | The rendered history token no longer names the accepted row prefix | `ws_id`, `reason` |
**Reconnection contract:** a freshly-opened SSE connection receives
one `approve_request` snapshot for every unresolved approval cycle, keyed by
the same stable `cycle_id`, plus any in-flight `wait_*` / `batch_*`
the current snapshot of any pending tool approval (`approve_request`
is re-sent if unresolved), any in-flight `wait_*` / `batch_*`
indicator, the worker's current `state_change`, and an
`in_progress_snapshot` carrying any partial content / reasoning the
model has produced for the in-progress turn — so a tab refresh
@@ -159,11 +139,6 @@ mid-approval, mid-tool-execution, or mid-stream restores both the
correct composer mode and the partial assistant text without waiting
for the response to complete.
`history_resync` is stronger than a numeric replay gap. The server closes that
stream; fetch and render `/history` again, then open a new stream with its new
cursor/token pair. The API and SDK expose these primitives but deliberately do
not choose a reconnect policy for callers.
---
## 3. Send the first user message
@@ -286,10 +261,10 @@ rounds to a 10× token-efficiency win.
---
## 7. Governance — trust, restrict, close_all_children
## 7. Governance — trust, restrict, stop_cascade, close_all_children
These three endpoints let an operator steer a live coordinator session
mid-flight. All three emit an audit event tagged
These four endpoints let an operator steer a live coordinator session
mid-flight. All four emit an audit event tagged
`coordinator.<action>` via the dedicated audit executor so a cascade
burst can't starve audit writes.
@@ -319,6 +294,28 @@ idempotent — calling twice with overlapping lists converges to the
union. Revocations don't survive a session close/reopen; operators
opt in per session. Cap 256 tool names per request, 128 chars each.
### `POST /stop_cascade` — cancel the subtree
```http
POST /v1/api/workstreams/{ws_id}/stop_cascade
{}
```
Cancels the coordinator's in-flight generation AND dispatches
`cancel_workstream` through the routing proxy for every direct
child in the in-memory registry. Returns:
```json
{"status": "ok", "cancelled": ["child-1", "child-3"], "failed": [], "skipped": ["child-2"]}
```
Response uses the [cascade-mutation bulk shape](bulk-endpoints.md):
`cancelled` = accepted, `failed` = dispatch error worth retrying,
`skipped` = upstream 404 (already gone — stale registry entry or
the row was deleted between snapshot and dispatch). Grandchildren
aren't touched directly; they sit behind their parent's cancel and
propagate via the child's SSE stream.
### `POST /close_all_children` — soft-close the direct fan-out
```http
@@ -332,16 +329,16 @@ Response:
{"status": "ok", "closed": ["c-1", "c-2"], "failed": [], "skipped": []}
```
Soft-close cascade bounded by a concurrency semaphore. The `reason`
(up to 512 chars) propagates into each closed child's audit +
`workstream_config` for postmortem. The model-facing tool that
pairs with this endpoint asks for a bounded teardown of the
coordinator's own fan-out. This *soft-closes*; to *cancel* the
fan-out instead, cancel the coordinator (§8) — a coordinator cancel
auto-cascades to its direct children.
Soft-close cascade bounded by the same semaphore as `stop_cascade`.
The `reason` (up to 512 chars) propagates into each closed child's
audit + `workstream_config` for postmortem. Unlike `stop_cascade`
this does NOT recurse into grandchildren — the model-facing tool
that pairs with this endpoint asks for a bounded teardown of the
coordinator's own fan-out. For a full-subtree teardown, use
`stop_cascade`.
See [bulk-endpoints.md](bulk-endpoints.md) for why `close_all_children`
uses the cascade-mutation shape and how it differs from the
See [bulk-endpoints.md](bulk-endpoints.md) for why both endpoints
share the cascade-mutation shape and how it differs from the
`spawn_batch` / `cluster/ws/live` shape.
---
@@ -350,35 +347,21 @@ uses the cascade-mutation shape and how it differs from the
The `approve` endpoint is what resolves an `approve_request` SSE
event. The coordinator's worker thread is blocked inside
`ui.approve_tools` waiting for this POST. Parallel task agents can leave
several approval cycles live at once, so current clients echo the event's
`cycle_id` (or a member `call_id`). A selector-less request resolves the oldest
cycle for compatibility.
`ui.approve_tools` waiting for this POST.
```http
POST /v1/api/workstreams/{ws_id}/approve
{"approved": true, "feedback": null, "always": false, "cycle_id": "cycle_789"}
{"approved": true, "feedback": null, "always": false}
{"approved": false, "feedback": "spawn count looks too high — try 3 not 10"}
{"approved": true, "feedback": null, "always": true} // remember this cycle's tool names
{"approved": true, "feedback": null, "always": true} // always-approve this tool name
```
Success returns `{"status": "ok", "cycle_id": "cycle_789"}`. A stale selector
returns `409` with the currently oldest cycle/call IDs. `always` remembers only
the tool names in the cycle that actually resolved; it does not enable blanket
approval.
`cancel` requests cooperative cancellation of the coordinator's in-flight
generation and auto-cascades to its direct children:
`cancel_workstream` is dispatched through the routing proxy for
every direct child in the registry. The HTTP acknowledgement is immediate;
the worker becomes idle after unwinding. Pass `{"force": true}` only to release
a wedged worker slot immediately. The coordinator itself remains open for a
fresh `send`:
`cancel` drops the in-flight generation but leaves the coordinator
idle and open for a fresh `send`:
```http
POST /v1/api/workstreams/{ws_id}/cancel
{}
{"status": "ok", "dropped": {}}
```
---
@@ -390,30 +373,24 @@ POST /v1/api/workstreams/{ws_id}/close
{}
```
Soft-closes the session — state persists, children keep running
(wind them down first with `close_all_children`, or by cancelling
the coordinator, which cascades the cancel to its direct children),
the worker thread exits, SSE streams send a final `stream_end` and
Soft-closes the session — state persists, children keep running (use
`close_all_children` or `stop_cascade` first to wind them down), the
worker thread exits, SSE streams send a final `stream_end` and
disconnect. The row is reopenable via
`POST /v1/api/workstreams/{ws_id}/open` so long as it hasn't been
deleted.
If any accepted live conversation row still requires persistence
reconciliation, close returns `409 {"error":"workstream has unresolved
persistence"}`. The coordinator remains loaded, its journal is retained, and
no history is discarded; retry after storage recovers.
---
## Further reading
- [coordinator-skills.md](coordinator-skills.md) — writing a skill
that runs on a coordinator session (orchestrator framing,
that runs on a coordinator session (orchestrator persona,
workflow patterns, `SkillKind` classifier).
- [bulk-endpoints.md](bulk-endpoints.md) — the two bulk-shape
idioms (`{results, denied, truncated}` vs
`{<bucket>, failed, skipped}`) used by `cluster/ws/live`,
`spawn_batch`, and `close_all_children`.
`spawn_batch`, `stop_cascade`, and `close_all_children`.
- [architecture.md](architecture.md) — cluster-wide architecture
including how coordinator sessions fit next to node-hosted
interactive workstreams.
+24 -34
View File
@@ -1,13 +1,13 @@
# Writing a coordinator-specific skill
A skill is prompt-level framing that steers a Turnstone session
Skills are prompt-level personas that steer a Turnstone session
toward a narrow task. Most skills target **interactive** sessions —
the single-workstream "do this thing" surface where the model wields
`bash`, `edit_file`, `web_fetch`, and the rest of the maker toolset.
A **coordinator skill** is different. It runs on a session whose job
is to orchestrate other sessions. The toolset is smaller and
narrower, the role is an orchestrator instead of a maker, and the
narrower, the persona is an orchestrator instead of a maker, and the
success metric is "did the plan resolve" instead of "did the code
compile". This doc covers the differences a skill author has to
care about.
@@ -22,8 +22,8 @@ migration 044 added the column). Three values:
| `SkillKind` enum | Stored as | Meaning |
|-------------------------|-----------------|----------------------------------------------------------------------------|
| `SkillKind.INTERACTIVE` | `"interactive"` | Authored for the interactive maker role (single-workstream "do this"). |
| `SkillKind.COORDINATOR` | `"coordinator"` | Authored for the orchestrator role (delegate, monitor, synthesise). |
| `SkillKind.INTERACTIVE` | `"interactive"` | Authored for the interactive maker persona (single-workstream "do this"). |
| `SkillKind.COORDINATOR` | `"coordinator"` | Authored for the orchestrator persona (delegate, monitor, synthesise). |
| `SkillKind.ANY` | `"any"` | Either surface (or audience-neutral). Default on create. |
The `kind` field is a `StrEnum` — drop-in `str` compatible — so DB
@@ -77,7 +77,7 @@ or MCP config can do adds to it. Current members:
| `delete_workstream` | wind-down | Hard-delete one child. Requires approval. |
| `list_nodes` | discover | Enumerate live cluster nodes + capabilities. |
| `skills` (action=find) | discover | Browse the skill catalog; opt-in `kind` filter narrows by audience. |
| `memory` | persist | Durable acting-user orchestration memory (`coordinator`), plus shared memory when attached to a project. |
| `memory` | persist | Orchestration scratchpad keyed by the `coordinator` scope. |
| `notify` | broadcast | Post a status update to a human channel at a narrative beat. |
| `tasks` | plan | Orchestrator-only scratchpad. Children don't see it. |
@@ -96,20 +96,20 @@ for the output. The coordinator stays the orchestrator.
---
## Framing differences
## Persona differences
Interactive skills compose on top of `base_interactive.md` — a
"maker" framing: get the work done, use the tools, edit the code,
"maker" persona: get the work done, use the tools, edit the code,
close the loop.
Coordinator skills compose on top of
[`personas/orchestrator.md`](../turnstone/prompts/personas/orchestrator.md) —
an "orchestrator" framing: decompose, delegate, monitor, synthesise.
[`base_coordinator.md`](../turnstone/prompts/base_coordinator.md) —
an "orchestrator" persona: decompose, delegate, monitor, synthesise.
The base text is short but sets the tone every coordinator skill
inherits:
> You are a coordinator. Your role is to orchestrate work across
> the cluster... You do
> You are a coordinator on a small, focused infrastructure team.
> Your role is to orchestrate work across the cluster... You do
> not edit files, run shell commands, browse the web, or manipulate
> the codebase directly. Children do that.
@@ -126,11 +126,10 @@ the skill should end on.
`tasks` is the coordinator's scratchpad — a persisted, ordered
list of rows with fields `{id, title, status, child_ws_id, created,
updated}`, plus `note` on rows where one has been set (the key is
absent otherwise), that only this coordinator sees. Children don't
see it; the user does via the sidebar. Five actions: `add`,
`update`, `remove`, `reorder`, `list` (only `list` is auto-approved;
the mutators go through the approval flow).
updated}` that only this coordinator sees. Children don't see it;
the user does via the sidebar. Five actions: `add`, `update`,
`remove`, `reorder`, `list` (only `list` is auto-approved; the
mutators go through the approval flow).
The input schema refers to rows by `task_id`; the persisted row
object exposes the same id as `id`. The `child_ws_id` field is a
@@ -143,20 +142,11 @@ A skill's initial prompt can seed the task list by calling
`tasks(action="add", title=...)` as its very first tool calls —
the user gets a visible plan before any child is spawned, and the
coordinator's future self has something concrete to iterate on.
Status transitions (`pending``in_progress``done` / `blocked` /
`needs_user`) are the skill's main feedback loop: mutate the task
when the child covering it finishes, not when the child starts.
`blocked` and `needs_user` are not interchangeable — `blocked` is a
dependency the coordinator may be able to clear itself, while
`needs_user` marks a task that cannot move without a decision,
approval, or grant only the user can give. The distinction is
load-bearing: a coordinator that goes idle holding open tasks gets
nudged to pick them back up — even when children are still running, so
keep the matrix honest rather than expecting the reminder to wait for
an all-clear — and `needs_user` is what tells that nudge the stop was
deliberate. Pair it with `note` to record what is being asked for.
Use `tasks(action="update", task_id=..., child_ws_id=<ws_id>)` to link
a task to the child that owns it once spawn returns.
Status transitions (`pending``in_progress``done` / `blocked`)
are the skill's main feedback loop: mutate the task when the child
covering it finishes, not when the child starts. Use
`tasks(action="update", task_id=..., child_ws_id=<ws_id>)` to
link a task to the child that owns it once spawn returns.
A final gotcha: parallel tool dispatch does NOT serialise reads
after writes in the same batch. If a skill issues an `update` and
@@ -307,11 +297,11 @@ and the coordinator's planning step is itself valuable.
tasks(action='add', title='...') × N # the plan, visible in the sidebar
for task in tasks:
spawn_workstream(skill=..., initial_message=task.brief)
tasks(action='update', task_id=task.id, note='ws=<child_ws_id>')
tasks(action='update', task_id=task.id, notes='ws=<child_ws_id>')
wait_for_workstream(ws_ids=[...], mode='all', timeout=...)
for child in children:
inspect_workstream(ws_id=child)
tasks(action='update', task_id=..., status='done', note='result summary')
tasks(action='update', task_id=..., status='done', notes='result summary')
→ synthesise
```
@@ -349,7 +339,7 @@ For a new coordinator skill:
A full end-to-end test isn't required for every skill; a
prepare-step unit test that asserts "given this initial message, the
first tool call is X with Y args" is usually sufficient to catch
framing drift without a real LLM in the loop.
persona drift without a real LLM in the loop.
---
@@ -361,7 +351,7 @@ framing drift without a real LLM in the loop.
`spawn_batch` and `close_all_children` use, so your skill can
parse results / denied arrays correctly.
- [governance.md](governance.md) — the broader governance surface
(`/trust`, `/restrict`, role-based permissions)
(`/trust`, `/restrict`, `/stop_cascade`, role-based permissions)
that wraps every coord session.
- [settings.md](settings.md) — `coordinator.model_alias` and
`coordinator.reasoning_effort` settings that gate which LLM runs
+9 -11
View File
@@ -13,7 +13,7 @@ cloud "LLM Providers" as llm {
component [OpenAI-compatible API\n(OpenAI, vLLM, llama.cpp)] as llm_openai
component [Anthropic Messages API] as llm_anthropic
}
database "SQLite / PostgreSQL\n(durable state)" as storage
database "SQLite\n(.turnstone.db)" as sqlite
' Turnstone System Boundary
package "Turnstone Platform" {
@@ -33,26 +33,24 @@ eval_user --> eval : Python API
' Internal connections
cli --> llm : LLM Provider API\n(via provider adapters)
cli --> storage : persistence
cli --> sqlite : SQLite
server --> llm : LLM Provider API\n(via provider adapters)
server --> storage : persistence
server --> sqlite : SQLite
eval --> llm : LLM Provider API\n(non-streaming)
eval --> storage : persistence
eval --> sqlite : SQLite
console --> server : HTTP routing/UI proxy + cluster SSE\n(FNV-1a rendezvous placement,\nproxy /node/{id}/* traffic)
console --> server : HTTP proxy\n(hash-ring bucket lookup,\nproxy /node/{id}/* traffic)
channel --> console : multi-node route/create/live/send/approve
channel --> server : direct mode + owning-node SSE\n(POST /v1/api/workstreams/{ws_id}/send,\nGET /v1/api/workstreams/{ws_id}/events)
channel --> server : HTTP + SSE\n(POST /v1/api/workstreams/{ws_id}/send,\nGET /v1/api/workstreams/{ws_id}/events)
' Notes
note right of console
Multi-node router:
- FNV-1a rendezvous placement
- Hash-ring bucket lookup
- Proxies create/send/approve
- Collector aggregates node SSE
- Browser dashboard receives console SSE fanout
- /node/{id} proxies pane HTTP + SSE
- Direct SSE from client to node
- HTTP polling for dashboard
end note
@enduml
+10 -37
View File
@@ -18,7 +18,6 @@ skinparam component {
package "Entry Points" <<Rectangle>> {
component [cli.py\nturnstone] as cli <<entry>>
component [server.py\nturnstone-server] as server <<entry>>
component [console/server.py\nturnstone-console] as consoleentry <<entry>>
component [eval.py\nturnstone-eval] as eval <<entry>>
component [admin.py\nturnstone-admin] as admin <<entry>>
component [bootstrap.py\nturnstone-bootstrap] as bootstrap <<entry>>
@@ -26,16 +25,9 @@ package "Entry Points" <<Rectangle>> {
' Core engine
package "turnstone/core/" <<Rectangle>> {
component [session.py\nChatSession, SessionUI\ngeneration-fenced turn loop] as session <<core>>
component [session_manager.py\nSessionManager\nshared lifecycle invariants] as sessionmanager <<core>>
component [adapters/\ninteractive + coordinator\nconstruction/event policies] as adapters <<core>>
component [model_turn.py\nModelLane, model_turn()\nlower / sample / re-ingest] as modelturn <<core>>
component [trajectory.py\ncanonical Turn IR] as trajectory <<core>>
component [lowering.py\nprovider-wire lowering] as lowering <<core>>
component [state_writer.py\nordered durable state tail] as statewriter <<core>>
component [model_backend_auth.py\nper-call backend credentials] as modelauth <<core>>
component [session.py\nChatSession, SessionUI] as session <<core>>
component [providers/\nLLMProvider, OpenAI, Anthropic, Google] as providers <<core>>
component [workstream.py\nWorkstream types + state] as workstream <<core>>
component [workstream.py\nWorkstreamManager] as workstream <<core>>
component [tools.py\nTool loader] as tools <<core>>
component [memory.py\nPersistence facade] as memory <<core>>
component [storage/\nStorageBackend protocol\nSQLite + PostgreSQL] as storage <<core>>
@@ -87,18 +79,18 @@ package "turnstone/api/" <<Rectangle>> {
package "turnstone/sdk/" <<Rectangle>> {
component [server.py\nTurnstoneServer (sync+async)] as sdkserver <<sdk>>
component [console.py\nTurnstoneConsole (sync+async)] as sdkconsole <<sdk>>
component [events.py\nTyped SSE event stream] as sdkevents <<sdk>>
component [events.py\n27 SSE event types] as sdkevents <<sdk>>
component [_base.py\nhttpx client base] as sdkbase <<sdk>>
}
' Tool schemas
package "turnstone/tools/" <<Rectangle>> {
component [*.json\nBuilt-in tool schemas] as schemas <<artifact>>
component [*.json\n19 tool schemas] as schemas <<artifact>>
}
' Entry point dependencies
cli --> session
cli --> sessionmanager
cli --> workstream
cli --> config
cli --> memory
cli --> colors
@@ -107,8 +99,7 @@ cli --> spinner
cli --> tools
server --> session
server --> sessionmanager
server --> adapters
server --> workstream
server --> config
server --> memory
server --> metrics
@@ -122,26 +113,11 @@ eval --> memory
eval --> config
eval --> tools
consoleentry --> sessionmanager
consoleentry --> adapters
consoleentry --> consoleserver
admin --> auth
bootstrap --> providers
' Core internal deps
sessionmanager --> workstream
sessionmanager --> adapters
sessionmanager --> storage
adapters --> session : constructs
session --> modelturn
session --> trajectory
session --> lowering
session --> statewriter
session --> modelauth
modelturn --> providers
modelturn --> trajectory
modelturn --> lowering
session --> providers
session --> tools
session --> memory
memory --> storage
@@ -153,7 +129,6 @@ session --> mcp : optional
session --> toolsearch : optional
session --> registry : optional
registry --> providers
modelturn --> registry : coherent snapshot
healthcheck --> metrics
mcp --> config
registry --> config
@@ -163,17 +138,15 @@ tools --> schemas
gateway --> discordbot
gateway --> slackbot
gateway --> router
discordbot --> sdkserver : direct HTTP + node SSE
slackbot --> sdkserver : direct HTTP + node SSE
router --> sdkserver : single-node/direct mode
router --> sdkconsole : multi-node route/create/live
discordbot --> sdkserver : HTTP + SSE
slackbot --> sdkserver : HTTP + SSE
router --> storage : channel_routes
' Console dependencies
consoleserver --> collector
consoleserver --> config
consoleserver --> auth
collector --> server : discovery HTTP + cluster SSE aggregation
collector --> server : HTTP polling
' API dependencies
serverspec --> openapi
+31 -146
View File
@@ -32,7 +32,7 @@ class "TerminalUI" as TerminalUI {
class "WorkstreamTerminalUI" as WsTermUI {
- _output_buffer: list[tuple]
- ws_id: str
- manager: SessionManager
- manager: WorkstreamManager
+ flush_buffer()
--
Buffers output when workstream
@@ -41,14 +41,14 @@ class "WorkstreamTerminalUI" as WsTermUI {
class "WebUI" as WebUI {
- _listeners: list[Queue]
- _approval_cycles: dict[str, ApprovalCycle]
- _approval_event: Event
- _ws_prompt_tokens: int
- _ws_tool_calls: dict
+ resolve_approval(approved, feedback, cycle_id?, call_id?)
+ resolve_approval(approved, feedback)
--
Enqueues JSON events for SSE.
Concurrent approval cycles each own
a threading.Event and result slot.
Blocks on threading.Event for
approval.
SSE handlers bridge Queue to
async via run_in_executor().
--
@@ -66,7 +66,8 @@ class "NullUI" as NullUI {
interface "LLMProvider" as LLMProvider <<Protocol>> {
+ provider_name: str {property}
+ get_capabilities(model) → ModelCapabilities
+ create_streaming(client, model, messages, ..., cancel_ref, replay_reasoning_to_model) → Iterator[StreamChunk]
+ create_streaming(client, model, messages, ..., replay_reasoning_to_model) → Iterator[StreamChunk]
+ create_completion(client, model, messages, ..., replay_reasoning_to_model) → CompletionResult
+ convert_tools(tools) → list[dict]
+ extract_reasoning_text(provider_blocks) → str
+ retryable_error_names: frozenset[str] {property}
@@ -126,82 +127,31 @@ class "ModelCapabilities" as ModelCaps <<frozen>> {
+ supports_reasoning_replay: bool
}
class "ModelLane" as ModelLane <<frozen>> {
+ provider: LLMProvider
+ client: Any
+ model: str
+ alias: str
+ capabilities: ModelCapabilities
+ extra_params: dict | None
+ registry: ModelRegistry | None
+ admission: ModelAdmission | None
+ backend_auth_config: ModelConfig | None
+ backend_auth_resolver: Callable | None
}
class "ResolvedModelBinding" as ResolvedBinding <<frozen>> {
+ lane: ModelLane
+ config: ModelConfig | None
+ registry_generation: int
}
class "ModelTurnResult" as ModelTurnResult <<frozen>> {
+ turn: Turn
+ tool_calls: list[dict]
+ finish_reason: str
+ usage: UsageInfo | None
+ wire_msgs: list[dict] | None
+ producer: str
+ serving_model: str
}
class "model_turn()" as ModelTurnFn {
Turn IR → lower → provider stream
→ drain → canonical assistant Turn
--
core/model_turn.py
}
class "Backend auth resolver" as BackendAuth {
+ resolve_model_backend_auth_token(...)
--
Resolves static / Entra OBO /
Entra app / RFC 8693 per call.
Dynamic failure can fail closed.
--
core/model_backend_auth.py
}
' ChatSession
class "ChatSession" as ChatSession {
- _model_binding: ResolvedModelBinding
- _model_binding_lock: Lock
- client: Any
- provider: LLMProvider
- model: str
- ui: SessionUI
- messages: list[Turn]
- messages: list[dict]
- _msg_tokens: list[int]
- _ws_id: str
- _mcp_client: MCPClientManager | None
- _tool_search: ToolSearchManager | None
- _registry: ModelRegistry | None
- _generation: int
- _cancel_event: Event
- _durability_next_ticket: int
+ model_alias: str | None {property}
- _tools: list[dict]
- _task_tools: list[dict]
- _read_files: set[str]
- system_messages: list[dict]
--
+ send(user_input: str, ..., acting_user_id: str | None)
+ cancel()
+ compact_now() → bool
+ fork_from_storage(source_ws_id, principal_id, ...)
+ send(user_input: str)
+ handle_command(command: str)
+ resume(ws_id: str)
- _save_config()
- _stream_response(my_generation) → ModelTurnResult
- _model_turn_with_fallback(consumer, prepare_wire) → ModelTurnResult
- _model_turn_with_retry(lane, tracker, ...) → ModelTurnResult
- _stream_response(stream) → dict
- _create_stream_with_retry(msgs) → Stream (+ fallback)
- _try_stream(client, model, msgs) → Stream
- _execute_tools(tool_calls) → (results, feedback)
- _prepare_tool(tc) → item dict
- _prepare_mcp_tool(call_id, name, args) → item dict
@@ -213,10 +163,8 @@ class "ChatSession" as ChatSession {
- _rebuild_tool_search()
+ close()
- _run_agent(messages, tools, ...) → str
- _compact_messages(auto: bool, my_generation: int)
- _commit_for_generation(generation, commit)
- _publish_for_generation(generation, publish)
- _full_messages() → list[Turn]
- _compact_messages(auto: bool)
- _full_messages() → list[dict]
- _update_token_table(msg)
- _emit_state(state: str)
- _generate_title()
@@ -229,40 +177,19 @@ class "HeadlessSession" as HeadlessSession {
+ send_headless(input, max_turns, ...)
- _override_system_prompt(content)
--
eval.py: drained single-shot turns,
eval.py: non-streaming,
records all tool calls
}
' SessionManager
interface "SessionKindAdapter" as KindAdapter <<Protocol>> {
+ kind: WorkstreamKind
+ build_ui(ws) → SessionUI
+ build_session(ws, ...) → ChatSession
+ cleanup_ui(ws)
}
interface "SessionEventEmitter" as EventEmitter <<Protocol>> {
+ emit_created(ws)
+ emit_rehydrated(ws)
+ emit_state(ws, state)
+ emit_closed(ws_id, reason, name)
}
class "SessionManager" as SessionMgr {
- _adapter: SessionKindAdapter
- _storage: StorageBackend
' WorkstreamManager
class "WorkstreamManager" as WsMgr {
- _session_factory: Callable[[SessionUI], ChatSession]
- _workstreams: dict[str, Workstream]
- _pending_creates: dict[str, Workstream]
- _retiring_ids: set[str]
- _state_writer: StateWriter | None
- _order: list[str]
- _active_id: str
- _on_state_change: Callable
--
+ create(user_id, name, ..., defer_emit_created) → Workstream
+ commit_create(ws) → bool
+ discard(ws, ...) → bool
+ open(ws_id) → Workstream | None
+ delete(ws_id) → bool
+ create(name, ui_factory) → Workstream
+ close(ws_id)
+ get(ws_id) → Workstream
+ get_active() → Workstream
@@ -277,17 +204,11 @@ class "Workstream" as Ws <<dataclass>> {
+ id: str
+ name: str
+ state: WorkstreamState
+ session: ChatSession | None
+ ui: SessionUI | None
+ worker_thread: Thread | None
+ session: ChatSession
+ ui: SessionUI
+ worker_thread: Thread
+ error_message: str
+ last_active: float
+ kind: WorkstreamKind
+ user_id: str
+ parent_ws_id: str | None
+ project_id: str | None
- _fork_reservation_token: str
- _closed: bool
- _lock: Lock
}
@@ -358,13 +279,12 @@ class "ModelRegistry" as ModelReg {
- _models: dict[str, ModelConfig]
- _clients: dict[str, Any]
- _providers: dict[str, LLMProvider]
- _admissions: dict[str, ModelAdmission]
- _client_lock: Lock
+ default: str
+ fallback: list[str]
+ agent_model: str | None
--
+ resolve_binding(alias) → (client, model, config, provider, admission, generation)
+ resolve(alias) → (client, model, config)
+ get_client(alias) → Any
+ get_provider(alias) → LLMProvider
+ has_alias(alias) → bool
@@ -378,22 +298,6 @@ class "ModelRegistry" as ModelReg {
core/model_registry.py
}
class "ModelAdmission" as ModelAdmission {
- alias: str
- _limit: int
- _in_flight: int
- _waiters: deque
+ acquire(cancel_ref) → AdmissionLease
+ set_limit(limit)
+ snapshot() → AdmissionSnapshot
--
Per-process FIFO generation gate.
Stable across alias hot reloads;
queue time is deadline credit.
--
core/admission.py
}
class "ModelConfig" as ModelCfg <<frozen>> {
+ alias: str
+ provider: str
@@ -403,10 +307,6 @@ class "ModelConfig" as ModelCfg <<frozen>> {
+ temperature: float | None
+ max_tokens: int | None
+ reasoning_effort: str | None
+ max_concurrency: int
+ auth_mode: str
+ obo_audience: str
+ obo_scopes: str
}
' Circuit breaker state
@@ -476,35 +376,22 @@ LLMProvider <|.. AnthropicProv
OpenAIProv <|-- GoogleProv
ChatSession --> SessionUI : uses
ChatSession --> ResolvedBinding : owns coherent snapshot
ChatSession --> ModelTurnFn : every model-backed role
ChatSession --> LLMProvider : delegates LLM calls
ChatSession --> MCPMgr : optional
ChatSession --o ToolSearchMgr : _tool_search
ChatSession --> ModelReg : optional
ChatSession <|-- HeadlessSession
SessionMgr --> "*" Ws : manages
SessionMgr --> KindAdapter : delegates construction
SessionMgr --> EventEmitter : lifecycle fan-out
WsMgr --> "*" Ws : manages
Ws --> "1" ChatSession : wraps
Ws --> "1" SessionUI : wraps
Ws --> "1" WsState : has
KindAdapter ..> ChatSession : constructs
WsMgr ..> ChatSession : creates via\nsession_factory(ui, model_alias)
ModelReg --> "*" ModelCfg : holds
ModelReg --> "*" LLMProvider : caches
ModelReg --> "*" ModelAdmission : owns per alias
LLMProvider --> ModelCaps : returns
ModelReg --> ResolvedBinding : resolves atomically
ResolvedBinding --> ModelLane
ModelLane --> LLMProvider
ModelLane --> ModelCaps
ModelLane --> ModelCfg : auth/config snapshot
ModelLane --> ModelAdmission : admission lease
ModelTurnFn --> ModelLane
ModelTurnFn --> ModelTurnResult
ModelTurnFn ..> BackendAuth : per-call resolver
ChatSession --> HealthMon : checks circuit
HealthMon --> "1" CircuitState : has
@@ -517,9 +404,7 @@ note bottom of ChatSession
Provider-agnostic — delegates all LLM
communication to LLMProvider adapters.
Every live/durable publication is fenced by
its generation. Model calls use immutable lanes;
provider-wire mutation stays at lowering.
core/session.py (~2700 lines)
end note
@enduml
+157 -145
View File
@@ -1,171 +1,183 @@
@startuml
!theme plain
title Turnstone — Generation-Fenced Conversation Turn
title Turnstone — Conversation Turn Lifecycle
skinparam sequenceArrowThickness 1.5
skinparam sequenceLifeLineBackgroundColor #F5F5F5
participant "HTTP / CLI\ncaller" as User
participant "SessionManager" as Manager
participant "ChatSession" as Session
participant "SessionUIBase" as UI
participant "Accepted-row handoff\n(total live prefix)" as Handoff
participant "model_turn()\n+ lowering" as Plant
participant "ModelAdmission\n(per alias)" as Admission
participant "LLM provider" as Provider
participant "Tool workers" as Tools
database "StorageBackend\n(SQLite / PostgreSQL)" as Storage
participant "User /\nHTTP Client" as User
participant "ChatSession" as CS
participant "SessionUI" as UI
participant "LLMProvider\n(OpenAI / Anthropic)" as LLM
participant "Tool Executor\n(ThreadPool)" as TP
database "SQLite" as DB
== Admission and generation claim ==
== User Input ==
User -> Manager : dispatch send on one Workstream
Manager -> Session : bind_acting_user(principal)\nsend(text, attachments, send_id)
activate Session
Session -> Session : refresh immutable ResolvedModelBinding
User -> CS : send(user_input)
activate CS
opt token budget exhausted
Session -> UI : approve_tools(__budget_override__)
note right of UI
This gate precedes a generation claim but carries
a monotonic cancellation witness. Stop cannot be
mistaken for a budget-policy denial.
end note
end
CS -> CS : messages.append({role: "user", content: input})
CS -> DB : save_message(ws_id, "user", input)
Session -> Session : _claim_generation() → generation N\ninstall fresh cancel event
Session -> Session : plan memory / participant context
Session -> Handoff : admit USER row\ncommit_key + prefix revision
Session -> Storage : ordered durable batch:\nappend canonical user Turn + metadata
== LLM Call Loop ==
note over Session, Handoff
Every accepted conversation row enters this lane before durability:
USER, ASSISTANT, TOOL, SYSTEM, compaction checkpoints, and cancellation
markers. Admission shares the handoff lock with its live UI transition
or history_resync repair event.
end note
group loop [while tool_calls present]
note over Handoff, Storage
_commit_for_generation(N) admits bounded live mutations under the
generation lock, then executes immutable persistence closures in FIFO
ticket order. A force successor either follows the whole commit or
prevents it. /history projects durable prefix + pending journal suffix;
durable ACK removes the pending copy without changing the prefix revision.
end note
CS -> UI : on_turn_start()
note right of UI
SessionUIBase resets the per-turn inflight
buffers (_ws_inflight_content / reasoning /
seq) that fuel the SSE in_progress_snapshot
event for mid-stream refresh resume.
end note
opt already over the hard context ceiling
Session -> Session : compact before first model call\n(preserve the new user turn)
end
CS -> UI : on_state_change("thinking")
CS -> UI : on_thinking_start()
== Model / tool loop ==
CS -> LLM : provider.create_streaming(\n client, model, messages, tools, ...)\n (normalized StreamChunk iterator)
activate LLM
loop until final answer and no queued input
Session -> UI : on_turn_start()\nreset per-stream replay buffers
Session -> UI : state = thinking\non_thinking_start()
Session -> Session : _stream_response(N)\nretry + fallback policy
Session -> Plant : model_turn(active ModelLane, Turns,\n tools, cancel_ref, on_chunk)
activate Plant
Plant -> Plant : canonical Turns → provider wire\nrestore ids + repair + lane-specific fold
Plant -> Plant : materialize attachment refs\n(nested perception before outer slot)
Plant -> Admission : acquire(cancel_ref)
activate Admission
Plant -> Plant : resolve per-call backend credential\nfrom lane's pinned ModelConfig
Plant -> Provider : create_streaming(...)
activate Provider
note right of CS
Retry up to 3× on transient errors:
RateLimitError, APITimeoutError,
APIConnectionError, InternalServerError,
ServiceUnavailableError, APIError
Backoff: 1s, 2s, 4s
end note
loop normalized stream chunks
Provider --> Plant : StreamChunk
Plant --> Session : on_chunk(StreamChunk)
Session -> Session : check cancel event + generation N
Session -> UI : reasoning / content / info token
end
== Streaming Response ==
Provider --> Plant : finish + usage + native blocks
deactivate Provider
Plant -> Plant : drain + re-ingest assistant Turn\nwith serving-lane provenance
Plant -> Admission : release before retry backoff
deactivate Admission
Plant --> Session : ModelTurnResult
deactivate Plant
Session -> UI : on_stream_end()
Session -> Session : generation-fenced result commit:\nappend assistant Turn + token accounting
Session -> UI : on_turn_committed()
Session -> Handoff : admit ASSISTANT row\ncommit_key + prefix revision
Session -> Storage : ordered durable assistant row\n(content + tool mirror + native lane)
alt no tool calls
opt over soft threshold
Session -> Session : cooperative / end-of-turn compaction
Session -> Handoff : admit SYSTEM/source=compaction\ncheckpoint projection
Session -> Storage : append checkpoint summary marker\nwith source watermark
note right of Storage
Full history remains durable. Resume loads
[summary] + rows after the checkpoint.
end note
opt model stopped for compaction
Session -> Handoff : admit USER/source=compaction_resume row
Session -> Storage : append synthetic compaction_resume Turn
end
end
alt queued messages drained
Session -> Handoff : admit combined queued USER row
Session -> Storage : append combined queued user Turn
else truly complete
Session -> UI : state = idle
end
else tool calls present
Session -> UI : state = running
Session -> Session : prepare items + previews\nattach cancellation witnesses
opt one or more items require a human
Session -> UI : approve_tools(items)\nregister independent ApprovalCycle
note right of UI
Parallel agents may own concurrent cycles.
cycle_id / call_id routes exactly one decision;
Smart Approvals may clear qualifying items.
end note
User -> UI : approve / deny selected cycle
UI --> Session : decision + optional feedback
loop for each chunk in stream
LLM --> CS : delta
note right of CS
on_thinking_stop() called on first
delta token via _stop_spinner_once()
end note
alt reasoning_content present
CS -> UI : on_reasoning_token(text)
else content present
CS -> UI : on_content_token(text)
else tool_call delta
CS -> CS : accumulate in tool_calls_acc
else info_delta present
CS -> UI : on_info(text)\n(e.g. server-side web search status)
end
end
Session -> Tools : execute admitted items in parallel
activate Tools
Tools --> UI : chunks + result card\nwith effect disposition
Tools --> Session : outputs / errors / effect statuses
deactivate Tools
Session -> Session : output-guard evaluation\nthen generation N re-check
note right of CS
**Cancellation checkpoint:**
_check_cancelled() runs per chunk.
If cancel_event is set, raises
GenerationCancelled — preserves
partial content, emits idle state.
end note
opt compaction owed before result sizing
Session -> Session : compact, preserving assistant tool-call Turn
Session -> Handoff : admit SYSTEM/source=compaction checkpoint
Session -> Storage : append checkpoint marker
LLM --> CS : stream complete (usage stats)
deactivate LLM
CS -> UI : on_thinking_stop() (no-op guard: already called by _stop_spinner_once)
CS -> UI : on_stream_end()
CS -> CS : _update_token_table()\ncalibrate chars_per_token ratio
CS -> CS : messages.append(assistant_msg)
CS -> UI : on_turn_committed()
note right of UI
Drops the per-turn inflight buffers — the
assistant message is now in the history
list, so the in_progress_snapshot must
not re-render it during the next tool-
execution window or the next streaming turn.
end note
CS -> DB : save_message(ws_id, "assistant", content)
CS -> DB : save_message(ws_id, "tool_call", ...) ×N
== Tool Dispatch (if tool_calls) ==
alt no tool_calls
CS -> UI : on_status(usage, context_window, effort)
opt prompt_tokens > context_window × auto_compact_pct
CS -> CS : _compact_messages(auto=True)
CS -> LLM : Non-streaming summarization call
CS -> CS : Replace messages with [summary]
end
opt first exchange & no title
CS -> CS : Background thread: _generate_title()
end
CS -> UI : on_state_change("idle")
CS --> User : return
else has tool_calls
CS -> UI : on_state_change("running")
== Phase 1: Prepare ==
CS -> CS : [_prepare_tool(tc) for tc in tool_calls]\nParse JSON args, validate,\nbuild preview + header
== Phase 2: Approve ==
CS -> UI : on_state_change("attention")
CS -> UI : approve_tools(items)
activate UI
note right of UI
TerminalUI: input() prompt
WebUI: _approval_event.wait()
NullUI: returns (True, None)
end note
UI --> CS : (approved: bool, feedback: str?)
deactivate UI
CS -> UI : on_state_change("running")
== Phase 3: Execute ==
CS -> TP : ThreadPoolExecutor(max_workers=4)\nrun_one(item) for each tool
activate TP
note right of TP
Parallel execution:
bash → Popen + line-by-line streaming
read_file → open().read() or base64 image
search → grep subprocess
edit_file → string replace
task → _run_agent() sub-loop
web_fetch → httpx + LLM summarize
web_search → provider-native or SearxNG fallback
memory/recall → SQLite
end note
note right of TP
bash: on_tool_output_chunk(call_id, line)
called per stdout line,
then on_tool_result(call_id, name, output, is_error).
is_error=True when execution failed.
call_id routes chunks/results to correct
tool div during parallel execution.
Other tools: on_tool_result() only.
end note
TP --> CS : [(call_id, output), ...]
deactivate TP
loop for each result
CS -> CS : messages.append({role: "tool", ...})
CS -> DB : save_message(ws_id, "tool_result", ...)
end
opt user_feedback from approval
CS -> CS : messages.append({role: "user", content: feedback})
end
note right of CS : Loop back for next LLM call
else GenerationCancelled
CS -> CS : Preserve partial content\nor roll back incomplete tools
CS -> UI : on_info("[Generation cancelled]")
CS -> UI : on_state_change("idle")
CS --> User : return (no re-raise)
end
Session -> Session : one generation-fenced batch:\nappend all Tool Turns, advisories, feedback
Session -> Handoff : admit FIFO TOOL rows\ncommit keys + prefix revisions
Session -> Storage : FIFO durable tool rows + metadata
end
end
== Stop / force-successor boundary ==
deactivate CS
User -> Session : cancel()
Session -> Session : atomically set generation event; snapshot\nmain stream, child scopes, judges, subprocesses
Session -> Provider : close live stream handle
Session -> Tools : abort child scopes + kill subprocess groups
Session -> UI : resolve only cancelled operation's\napproval cycles
opt cancellation produced accepted conversation rows
Session -> Handoff : admit partial ASSISTANT and/or\nsynthesized TOOL cancellation markers
Session -> Storage : idempotent keyed cancellation rows
end
note over Session, Storage
Every later publish/commit checks generation ownership. An abandoned
worker may unwind, but cannot append Turns, overwrite state, resolve a
successor approval, or repaint the successor UI. Observed tool effects
are preserved as controller-authored cancellation receipts; unreviewed
tool bytes are not laundered into model context.
end note
deactivate Session
@enduml
+103 -86
View File
@@ -1,117 +1,134 @@
@startuml
!theme plain
title Turnstone — Tool Pipeline: Prepare, Approve, Execute, Fold
title Turnstone — Tool Execution Pipeline (Three Phases)
start
partition "Phase 1 Prepare and assess" #E8F5E9 {
:Receive tool calls from one assistant Turn;
:Capture the generation's cancel event\nand acting principal;
partition "Phase 1: Prepare" #E8F5E9 {
:Receive tool_calls list from LLM response;
while (more tool calls?) is (yes)
:Parse arguments and dispatch to\nthe tool-specific preparer;
if (preparation succeeds?) then (yes)
:Build item: call_id, name, header, preview,\nneeds_approval, execute closure;
while (more tool_calls?) is (yes)
:Extract call_id, func_name, raw_args;
if (json.loads(raw_args) succeeds?) then (yes)
:parsed_args = JSON dict;
else (no)
:Build an error item for this call only;\nkeep sibling calls valid;
:Fallback 1: regex extraction;
if (regex found keys?) then (yes)
:parsed_args = extracted dict;
else (no)
:Fallback 2: bare string →\nPRIMARY_KEY_MAP[func_name];
endif
endif
:Attach operation-local cancellation witness\nand pinned principal;
:Dispatch to _prepare_{func_name}();
note right
**Dispatch table (16 built-in + tool_search):**
┌───────────────┬──────────────────┐
│ Tool │ Needs Approval? │
├───────────────┼──────────────────┤
│ bash │ ✓ Yes │
│ read_file │ ✗ Auto-approve │
│ write_file │ ✓ Yes │
│ edit_file │ ✓ Yes │
│ search │ ✗ Auto-approve │
│ diff_file │ ✗ Auto-approve │
│ web_fetch │ ✗ Auto-approve │
│ web_search │ ✗ Auto-approve │
│ tool_search │ ✗ Auto-approve │
│ task_agent │ ✓ Yes │
│ memory │ ✗ Auto-approve │
│ recall │ ✗ Auto-approve │
│ notify │ ✗ Auto-approve │
│ watch │ ✓ create only │
│ skill │ ✓ load only │
│ read_resource │ ✓ Yes │
│ use_prompt │ ✓ Yes │
├───────────────┼──────────────────┤
│ mcp__* │ ✓ Yes (external) │
└───────────────┴──────────────────┘
end note
:Build item dict:
{call_id, func_name, header,
preview, needs_approval,
approval_label, execute: Callable};
endwhile (no)
:Reject only unsafe ordering shapes\n(for example tasks read + write in one batch);
:Run heuristic intent assessment immediately;
:Start generation-pinned LLM judge in background;
:Stamp one immutable Smart Approval\nsettings snapshot on the batch;
note right
Preparation is per-call isolated: one bad preparer
becomes one error Tool Turn rather than orphaning the
assistant's entire tool-call set.
end note
}
partition "Phase 2 Approval cycle" #FFF3E0 {
:Apply explicit bypasses:\nskill / always / policy / blanket;
if (Smart Approvals enabled?) then (yes)
:Wait within the batch's bounded judge deadline;
:Auto-approve only LLM approve verdicts\nat or above the captured threshold;
endif
if (human-gated items remain?) then (yes)
:Acquire approval-publication lease;
:Register independent ApprovalCycle\n(cycle_id, call_ids, event, result);
:Publish approve_request + heuristic verdicts;
partition "Phase 2: Approve" #FFF3E0 {
if (any items need approval?) then (yes)
:_emit_state("attention");
:ui.approve_tools(items);
note right
Parallel task agents can hold several cycles at once.
A decision selects one cycle_id / call_id (or the oldest
cycle for a legacy selector-less client). Double resolve
is a guarded no-op; one cycle cannot wake a sibling.
**auto_approve check is handled
internally by ui.approve_tools()**
**TerminalUI**: Print headers/previews,
prompt [y/n/a, optional message]
If user chose "always":
Add pending tool names to auto_approve_tools
(auto-approve these tool types going forward)
**WebUI**: Enqueue approve_request,
block on _approval_event.wait()
**NullUI**: Return (True, None)
end note
if (operator approves?) then (yes)
:Record decision and optional feedback;
else (denies / policy blocks)
:Mark only pending items denied;\nEffectStatus = none;
if (user approved?) then (yes)
:_emit_state("running");
else (denied)
:Mark all pending items as denied;
:denial_msg = "Denied by user";
:_emit_state("running");
endif
:Publish approval_resolved;\nunregister this cycle;
else (all bypassed / auto-approved)
:Publish tool_info with the exact\nauto-approve reason per item;
endif
if (owning operation cancelled?) then (yes)
:Cancel only cycles carrying that witness;
:Stage every unstarted call as\nEffectStatus = none;
stop
else (all auto-approved)
:ui enqueues tool_info event\n(no blocking);
endif
}
partition "Phase 3 Execute" #E3F2FD {
:Generation + cancellation checkpoint;
if (batch requires serial ordering?) then (yes)
:Execute in provider order;
else (no)
:Execute via bounded ThreadPoolExecutor;
partition "Phase 3: Execute" #E3F2FD {
:_check_cancelled();
note right: Cancellation checkpoint:\nraises GenerationCancelled if\ncancel event is set
if (single tool call?) then (yes)
:Execute sequentially:\nrun_one(items[0]);
else (multiple)
:Execute in parallel:\nThreadPoolExecutor(max_workers=4)\npool.map(run_one, items);
endif
note right
Each worker marks its call started only after the final
generation/cancel check. A missing result after that edge is
conservatively unknown; an unstarted call is definitively none.
**run_one(item):**
if item.error → return error string
if item.denied → return denial message
else → item["execute"](item)
├─ _exec_bash: subprocess.run(["bash", script.sh])
├─ _exec_read_file: open().readlines() or _exec_read_image (base64)
├─ _exec_write_file: makedirs + write
├─ _exec_edit_file: find_occurrences + replace
├─ _exec_search: grep subprocess
├─ _exec_web_fetch: httpx.get + LLM summary
├─ _exec_web_search: SearxNG JSON GET (fallback for local models)
├─ _exec_tool_search: BM25 search + expand_visible()
├─ _exec_task: _run_agent(TASK_AGENT_TOOLS)
├─ _exec_notify: HTTP POST to channel gateway
├─ _exec_memory: structured memory save/search/delete/list
├─ _exec_recall: conversation history FTS5 search
├─ _exec_read_resource: MCPClientManager.read_resource_sync()
├─ _exec_use_prompt: MCPClientManager.get_prompt_sync()
└─ _exec_mcp_tool: MCPClientManager.call_tool_sync()
end note
:Stream tool chunks to the matching call card;
:Capture result / error / preview and effect disposition;
:Collect results: [(call_id, output), ...];
if (Stop interrupts execution?) then (yes)
:Abort child model scopes and subprocess groups;
:Synthesize cancellation receipts;
note right
EffectStatus vocabulary:
committed / none / unknown /
partial / rolled_back.
:_truncate_output() on each result\n(max context_window × chars_per_token × 0.5 chars\ndefault: ~context_window × 2 chars);
Observed but unreviewed bytes are omitted from the
model-facing receipt; effect truth is retained.
end note
endif
:bash: ui.on_tool_output_chunk(call_id, line) per stdout line;
:ui.on_tool_result(call_id, name, output, is_error) for each;
}
partition "Phase 4 — Guard and atomic fold" #F3E5F5 {
if (compaction already owed?) then (yes)
:Compact before sizing/folding results;\npreserve the assistant tool-call Turn;
endif
:Truncate each result against the remaining shared budget;
:Run heuristic + optional LLM output guard;
:Re-check generation after guard work;
:Under one generation commit, append the complete\nTool Turn block + advisories + feedback;
:Persist rows and effect/preview metadata\non the ordered durability lane;
:Return results to the next model turn;
}
:Return (results, user_feedback);
stop
@enduml
+12 -76
View File
@@ -3,7 +3,6 @@
title Turnstone — Workstream State Machine
skinparam state {
BackgroundColor<<lifecycle>> #ECEFF1
BackgroundColor<<idle>> #E8F5E9
BackgroundColor<<thinking>> #E3F2FD
BackgroundColor<<running>> #FFF3E0
@@ -11,18 +10,13 @@ skinparam state {
BackgroundColor<<error>> #FFCDD2
}
state "CREATING (persisted only)" as creating <<lifecycle>> : Hidden durable reservation.\nNot returned by ordinary list/open/history.
state "IDLE" as idle <<idle>> : Waiting for user input.\nNo active LLM call or tool execution.
state "THINKING" as thinking <<thinking>> : LLM streaming response.\nTokens flowing (reasoning + content).
state "RUNNING" as running <<running>> : Tools executing.\nThreadPoolExecutor active.
state "ATTENTION" as attention <<attention>> : Blocked on user action.\nTool approval needed.
state "ERROR" as error <<error>> : Exception occurred.\nRecoverable on next send().
state "CLOSED (persisted only)" as closed <<lifecycle>> : Unloaded, explicitly reopenable row.\nNot a live WorkstreamState member.
[*] --> creating : register exact incarnation\nstate="creating"
creating --> idle : finalize + publish create\nemit ws_created
creating --> [*] : immediate exact-token rollback\n(no lifecycle birth emitted)
creating --> [*] : stale >2h recovery\natomic hard delete; no close event
[*] --> idle : Session created
idle --> thinking : send() called\n_emit_state("thinking")
@@ -44,22 +38,6 @@ running --> error : Exception during\ntool execution
error --> thinking : New send() call\n_emit_state("thinking")
idle --> closed : close / eviction\n[journal reconciled]
error --> closed : close\n[journal reconciled]
thinking --> closed : close\n[journal reconciled]
running --> closed : close\n[journal reconciled]
attention --> closed : close\n[journal reconciled]
closed --> [*] : hard delete
closed --> idle : open / rehydrate
note right of closed
Before every soft-close / eviction transition,
the total accepted conversation-row journal must
be durably reconciled. An unresolved row makes an
explicit close return HTTP 409 (eviction refuses),
and the workstream remains loaded in its live state.
end note
thinking --> idle : cancel() called\nstream aborted\n_emit_state("idle")
running --> idle : cancel() called\n_emit_state("idle")
@@ -67,75 +45,33 @@ running --> idle : cancel() called\n_emit_state("idle")
attention --> idle : cancel() unblocks\napproval wait\n_emit_state("idle")
note left of idle
**Generation-scoped Stop:**
• Sets the active generation event.
• Closes its SDK stream; aborts child model
scopes and judges; kills subprocess groups.
• Sweeps every approval cycle owned by the
cancelled workstream operation.
• Every later send/model live or durable commit
re-checks generation ownership.
**force=true:** also abandons the stuck worker
slot and emits stream_end + IDLE immediately.
An orphaned send/model generation may unwind
but cannot publish into a successor generation.
Quick slash-command workers are a best-effort
escape hatch: without generation checkpoints,
one may finish an in-place mutation concurrently.
**Capacity eviction:** an IDLE candidate is only
a hint. Per-ID + object lifecycle lanes and the
workstream lock revalidate it as worker- and
send-barrier-free,
then install a terminal claim before slot swap.
**Cancel escalation:**
1. **Cooperative**: cancel() sets event + closes
SDK stream → worker exits at next checkpoint
2. **Force**: force=true abandons the worker
thread, emits stream_end immediately.
Orphaned thread still kills subprocesses
but skips message mutations (generation
counter prevents stale writes).
end note
note right of thinking
**Emitted via:**
session._emit_state(state)
→ ui.on_state_change(state)
→ SessionManager state tail
**Propagation:**
• WebUI → global SSE queue (ws_state)
• Console → cluster event / HTTP state
• CLI → SessionManager.set_state()
Non-terminal persistence may use StateWriter;
a per-id tail orders storage + subscribers and
prevents a late state from overwriting CLOSED.
• Console → HTTP polling picks up state
• CLI → WorkstreamManager.set_state()
end note
note left of attention
**Blocking mechanisms:**
• TerminalUI: input() prompt
• WebUI: one Event per ApprovalCycle
• WebUI: threading.Event.wait()
• ChannelBot: SSE event + Discord button
• NullUI: auto-approve (never reaches)
end note
note right of creating
CREATING and CLOSED are storage lifecycle
values, not members of WorkstreamState. The
live enum remains IDLE / THINKING / RUNNING /
ATTENTION / ERROR.
**Crash-abandoned CREATING recovery:**
• Boot pass, then every 5 min even when idle
eviction is disabled.
• Only rows >2h old; manager loaded/pending
IDs and live remote owners are protected.
• The current stable node ID is not a live-owner
exemption, allowing restart recovery.
• Unknown liveness/storage fails closed. Deletion
is atomic across dependents and attachment refs.
• Tokenless legacy/corrupt rows are locked,
reaped, and logged with a warning.
A loaded hard delete closes publication, drains
admitted session durability + state tails, then
conditionally removes the exact durable token.
end note
@enduml
+2 -2
View File
@@ -31,7 +31,7 @@ node "Docker Host" as host {
Command: turnstone-console
--port 8090
Depends: server
FNV-1a rendezvous router for
Hash-ring router for
multi-node clusters
end note
}
@@ -69,7 +69,7 @@ apiclient --> server : HTTP + SSE\nport 8080
' Internal connections
server --> llm_api : OpenAI API\n(HTTPS/HTTP)
console --> server : HTTP proxy\n(FNV-1a rendezvous placement,\nproxy /node/{id}/*)
console --> server : HTTP proxy\n(hash-ring lookup,\nproxy /node/{id}/*)
' Database connections (production/cluster profiles)
server ..> pgbouncer : PostgreSQL\n(pool_size=2)
+2 -18
View File
@@ -32,8 +32,7 @@ package "turnstone/sdk/ (Python)" {
+ approve()
+ command()
+ cancel(ws_id)
+ get_history(ws_id, limit) → WorkstreamHistoryResponse
+ stream_events(ws_id, last_event_id?, history_token?)
+ stream_events(ws_id)
+ stream_global_events()
+ send_and_wait()
+ list_saved_workstreams()
@@ -88,13 +87,6 @@ package "turnstone/sdk/ (Python)" {
+ ok: bool
}
class WorkstreamHistoryResponse <<type>> {
+ ws_id: str
+ messages: list[dict]
+ cursor: int | None
+ handoff_token: str | None
}
class ServerEvent <<event>> {
+ type: str
+ ws_id: str
@@ -113,7 +105,6 @@ package "turnstone/sdk/ (Python)" {
TurnstoneConsole --> AsyncTurnstoneConsole : wraps
TurnstoneConsole --> _SyncRunner : uses
AsyncTurnstoneServer ..> TurnResult : returns
AsyncTurnstoneServer ..> WorkstreamHistoryResponse : renders before SSE
AsyncTurnstoneServer ..> ServerEvent : yields
AsyncTurnstoneConsole ..> ClusterEvent : yields
}
@@ -131,8 +122,7 @@ package "sdk/typescript/ (TypeScript)" {
class "TurnstoneServer" as TSServer <<ts>> {
+ listWorkstreams()
+ send()
+ getHistory() → WorkstreamHistoryResponse
+ streamEvents(cursor?, token?)
+ streamEvents()
+ sendAndWait()
...
}
@@ -164,10 +154,4 @@ note right of AsyncTurnstoneServer
(no type duplication)
end note
note bottom of ServerEvent
history_resync is a typed repair signal.
SDKs expose the REST cursor/token handshake but
never refetch, render, or reconnect automatically.
end note
@enduml
+126 -147
View File
@@ -1,191 +1,170 @@
@startuml
!theme plain
title Turnstone — Storage, Deferred Create, Fork, and Checkpoint Architecture
title Turnstone — Storage Architecture
skinparam class {
BackgroundColor<<protocol>> #E8EAF6
BackgroundColor<<sqlite>> #C8E6C9
BackgroundColor<<postgres>> #B3E5FC
BackgroundColor<<lifecycle>> #FFF9C4
BackgroundColor<<facade>> #FFF9C4
BackgroundColor<<migration>> #FFE0B2
BackgroundColor<<schema>> #F3E5F5
BackgroundColor<<helper>> #FFE0B2
}
interface "StorageBackend" as Storage <<protocol>> {
+ load_message_turns(ws_id, checkpointed=True) → list[Turn]
+ save_message(ws_id, role, content, metadata...)
+ clone_workstream(source, destination, principal, expected_session) → ForkCloneSnapshot
--
+ register_workstream(..., state, reservation_token) → bool
+ ensure_workstream_incarnation_snapshot(ws_id) → row + token
+ finalize_deferred_create(ws_id, token, config...) → bool
+ publish_deferred_create(ws_id, token) → bool
+ delete_workstream_if_fork_reserved(ws_id, token) → bool
+ delete_stale_creating_reservations(...) → list[ws_id]
+ update_workstream_state(ws_id, state)
+ delete_workstream(ws_id) → bool
--
+ attachment / project / memory / auth / governance APIs
' -- Protocol --
interface "StorageBackend" as SB <<protocol>> {
+save_message(ws_id, role, content, ...)
+load_messages(ws_id) → list[dict]
+register_workstream(ws_id, node_id, name, state)
+update_workstream_state(ws_id, state)
+update_workstream_name(ws_id, name)
+set_workstream_alias(ws_id, alias) → bool
+update_workstream_title(ws_id, title)
+resolve_workstream(alias_or_id) → str | None
+delete_workstream(ws_id) → bool
+prune_workstreams(retention_days) → (int, int)
+list_workstreams(node_id, limit, *, parent_ws_id, kind, user_id) → list
+save_workstream_config(ws_id, config)
+load_workstream_config(ws_id) → dict
+kv_get(key) → str | None
+kv_set(key, value) → str | None
+kv_delete(key) → bool
+kv_list() → list[(str, str)]
+kv_search(query) → list[(str, str)]
+search_history(query, limit) → list
+search_history_recent(limit) → list
+create_user(user_id, username, display_name, pw_hash)
+get_user(user_id) / get_user_by_username(username)
+list_users() / delete_user(user_id)
+create_api_token(...) / get_api_token_by_hash(hash)
+list_api_tokens(user_id) / delete_api_token(id)
+close()
}
' -- Backends --
class "SQLiteBackend" as SQLite <<sqlite>> {
- _engine: sa.Engine
- _fts5_available: bool
-_engine: sa.Engine
-_fts5_available: bool
+__init__(path: str)
--
Fork clone: BEGIN IMMEDIATE
FTS5 refresh in same transaction
FTS5 full-text search
Default pool, check_same_thread=False
}
class "PostgreSQLBackend" as PG <<postgres>> {
- _engine: sa.Engine
-_engine: sa.Engine
+__init__(url: str, pool_size: int = 2,\n max_overflow: int = 3)
--
Fork clone: SERIALIZABLE + row locks
Retry SQLSTATE 40001 / 40P01
DML success uses RETURNING rows
tsvector + ILIKE search
Connection pooling (5 max per process)
}
class "_utils.py" as Utils <<helper>> {
+ reconstruct_turns(rows) → list[Turn]
+ recover_trajectory(turns) → list[Turn]
+ reconstruct_turns_checkpointed(...)
+ retain_attachment_refs(conn, ids)
+ release_attachment_refs(conn, ids)
+ clone_workstream_transaction(...) → ForkCloneSnapshot
}
class "ForkCloneExpectation" as Expectation <<lifecycle>> {
+ persona_config
+ project_id / name / writable
+ source_reservation_token
+ destination_reservation_token
}
class "ForkCloneSnapshot" as Snapshot <<lifecycle>> {
+ turns: tuple[Turn, ...]
+ config: dict[str, str]
+ project_id: str | None
}
class "workstreams" as Workstreams <<schema>> {
ws_id PK
state: creating | live state | closed
user_id, node_id, kind, parent_ws_id
project_id, persona, alias, title
}
class "conversations" as Conversations <<schema>> {
canonical persisted Turn rows
provider_data + tool_calls mirror
event_id, source, is_error, meta
attachment-id ref list
' -- Schema --
class "_schema.py" as Schema <<schema>> {
+metadata: MetaData
+memories: Table
+conversations: Table
+workstreams: Table (node_id, alias, title,\n state, skill_id)
+workstream_config: Table
+users: Table (username, password_hash)
+api_tokens: Table (token_hash, scopes)
+channel_users: Table (channel_type)
+scheduled_tasks: Table (..., skill)
--
compaction marker:
source="compaction"
meta.watermark=<folded row id>
SQLAlchemy Core
Single source of truth
}
class "workstream_config" as WorkstreamConfig <<schema>> {
PK (ws_id, key)
stamped persona/session config
private durable incarnation fence:
__fork_destination_reservation
' -- Migration --
class "_migrate.py" as Migrate <<migration>> {
+run_migrations(storage, backend)
-_bootstrap_existing_sqlite()
--
Programmatic Alembic
Auto-bootstrap existing DBs
}
class "workstream_attachments" as Attachments <<schema>> {
content-addressed blob
attachment_id, bytes, kind
refcount
class "migrations/" as Versions <<migration>> {
001_initial_schema.py
002_user_identity.py
}
class "projects + project_members" as Projects <<schema>> {
visibility / owner / membership
active project-memory envelope
' -- Registry --
class "_registry.py" as Registry {
-_storage: StorageBackend | None
+init_storage(backend, path, url) → StorageBackend
+get_storage() → StorageBackend
+reset_storage()
--
Auto-initializes SQLite
if not configured
}
class "SessionManager" as Manager <<lifecycle>> {
+ create(..., defer_emit_created)
+ commit_create(ws)
+ discard(ws)
+ reap_stale_creating_reservations(max_age=2h)
+ open / close / delete
' -- Facade --
class "memory.py" as Facade <<facade>> {
+save_message()
+load_messages()
+register_workstream()
+update_workstream_state()
+save_workstream_config()
+save_memory() / delete_memory()
+search_memories()
+... (all delegated functions)
--
Thin delegation to
get_storage()
Silent failure behavior
}
class "ChatSession" as Session <<lifecycle>> {
+ append canonical Turns
+ compact / resume checkpoint
+ fork_from_storage(...)
' -- Consumers --
class "session.py\nChatSession" as Session {
}
SQLite ..|> Storage
PG ..|> Storage
SQLite --> Utils
PG --> Utils
class "server.py\nWeb UI" as Server {
}
Storage --> Workstreams
Storage --> Conversations
Storage --> WorkstreamConfig
Storage --> Attachments
Storage --> Projects
class "cli.py\nTerminal" as CLI {
}
Manager --> Storage : lifecycle reservation + state
Session --> Storage : turn durability + resume
Session --> Expectation : construction witness
Storage --> Snapshot : atomic clone result
Expectation --> Utils : checked inside transaction
Utils --> Snapshot : builds
' -- Relationships --
SQLite ..|> SB
PG ..|> SB
note right of Manager
**Deferred create publication**
1. INSERT workstream as state="creating" and store a fresh
private token in the same transaction.
2. Construct UI/session and run attachment/fork gates while
ordinary list/open/history reads exclude the row.
3. finalize_deferred_create atomically applies config/alias.
4. publish_deferred_create compare-and-swaps creating → idle.
5. Only then emit ws_created.
SQLite --> Schema : uses
PG --> Schema : uses
Any normal prepublication failure immediately calls exact token-checked
deletion. The token survives publication as the row's incarnation fence:
rollback or later hard delete can never ABA-delete a replacement row.
A legacy row acquires the same private token atomically when rehydrate,
delete, or fork preflight takes its authoritative snapshot. Loaded hard
delete drains admitted session durability before its token-checked delete.
Registry --> SB : creates
Registry --> Migrate : calls
Migrate --> Versions : applies
Migrate --> Schema : references
Facade --> Registry : get_storage()
Session --> Facade : imports
Server --> Facade : imports
CLI --> Facade : imports
' -- Config --
note right of Registry
[database]
backend = "sqlite" | "postgresql"
url = "postgresql+psycopg://..."
path = ".turnstone.db"
pool_size = 2 (+ 3 overflow)
end note
note left of Manager
**Crash-abandoned hidden-create recovery**
• Boot pass; long-lived processes repeat every 5 min,
even when ordinary idle eviction is disabled.
• Candidates remain state="creating", are >2h old,
and are absent from the manager loaded/pending set.
• Live remote owners are protected. The current stable
node ID does not self-protect, enabling restart recovery.
• Unknown liveness or storage failure deletes nothing.
• One transaction rechecks state, age, and token, then
hard-deletes dependents and releases attachment refs.
• Tokenless legacy/corrupt rows use their locked durable
row as the incarnation fence and log a warning.
• Retention pruning excludes creating rows. Recovery never
closes or publishes them as live WorkstreamState values.
note bottom of SQLite
Default backend.
Zero-config for
single-node / dev.
end note
note bottom of Utils
**Atomic fork clone**
• Reject a provisional source; compare the source incarnation captured
by canonical preflight; re-authorize project visibility and compare the
live session envelope inside the transaction.
• Require a same-owner, empty destination still in creating state
with the exact reservation token.
• Copy the checkpoint-bounded canonical trajectory and config;
retain every referenced attachment or roll everything back.
• Preserve/rebase a valid compaction checkpoint watermark and
return the exact snapshot installed into the live destination.
end note
note bottom of Conversations
Full transcript rows are never deleted by compaction. Normal resume
loads the latest valid [summary] + rows after its watermark; audit and
export can request the full marker-free history.
note bottom of PG
Production backend.
Multi-node / Docker default.
Use PgBouncer (transaction mode)
for clusters > 50 nodes.
end note
@enduml
+166 -129
View File
@@ -1,153 +1,190 @@
@startuml
!theme plain
title Turnstone — User Authentication and Model-Backend Credentials
title Turnstone — Authentication Architecture
skinparam class {
BackgroundColor<<core>> #E8EAF6
BackgroundColor<<token>> #C8E6C9
BackgroundColor<<jwt>> #C8E6C9
BackgroundColor<<storage>> #B3E5FC
BackgroundColor<<runtime>> #FFE0B2
BackgroundColor<<model>> #F3E5F5
BackgroundColor<<endpoint>> #FFE0B2
BackgroundColor<<scope>> #F3E5F5
}
package "Request identity" {
class "AuthMiddleware / check_request()" as RequestAuth <<core>> {
Extract bearer or HttpOnly cookie
Validate audience + expiry
Check scope / permission
Publish AuthResult in request state
}
class "AuthResult" as AuthResult <<core>> {
+ user_id: str
+ scopes: frozenset[str]
+ permissions: frozenset[str]
+ token_source: str
}
class "JWT" as JWT <<token>> {
HS256, sub, aud, iat, exp
console proxy mints short-lived
server-audience identity
}
class "API / config token" as ApiToken <<token>> {
ts_* token: SHA-256 DB lookup
config token: constant-time compare
}
class "users / roles / api_tokens" as UserTables <<storage>> {
password hash + token hash
role-derived permissions
}
' -- Core Auth --
class "AuthConfig" as AC <<core>> {
+enabled: bool
+tokens: dict[str, str]
+check(token) → role | None
--
Static config-file tokens
hmac.compare_digest
}
package "Immutable model binding" {
class "ModelRegistry" as Registry <<model>> {
+ resolve_binding(alias)
+ generation: int
--
Atomically resolves client, provider,
model, ModelConfig, generation.
}
class "ModelConfig snapshot" as ModelConfig <<model>> {
+ alias / provider / endpoint / static key
+ auth_mode
+ obo_audience
+ obo_scopes
--
static | entra_obo | entra_app | rfc8693_obo
}
class "ModelLane" as Lane <<model>> {
+ client / provider / model / capabilities
+ backend_auth_config: ModelConfig
+ backend_auth_resolver: Callable
}
class "Model definitions" as ModelTable <<storage>> {
DB + config-file definitions
encrypted protected fields
}
class "AuthResult" as AR <<core>> {
+user_id: str
+scopes: frozenset[str]
+token_source: str
+has_scope(scope) → bool
}
package "Per-call credential resolution" {
class "resolve_model_backend_auth_token()" as Resolver <<runtime>> {
+ alias + pinned ModelConfig
+ initiating principal_id
+ ConfigStore + mint client
→ dynamic token | None | fail closed
}
class "Model mint client" as Mint <<runtime>> {
+ mint_model_obo_token_sync(...)
+ mint_app_token_sync(...)
--
Cached by alias / principal / grant leg;
retains refusal cause for diagnostics.
}
class "OIDC / OBO protected state" as OBOState <<storage>> {
encrypted user refresh credential
deployment Fernet key
configured grant profile
}
class "lane_call_client()" as CallClient <<runtime>> {
cancel check before mint
resolve once per plant call
cancel check after mint
client.with_options(api_key=token)
}
class "Provider SDK request" as ProviderCall <<runtime>> {
Anthropic: x-api-key
OpenAI-style: Authorization Bearer
}
class "check_request()" as CR <<core>> {
auth_config, method, path,
auth_header, cookie_header,
jwt_secret, storage
→ (allowed, status, msg, AuthResult)
--
1. Auth disabled → allow
2. Public path → allow
3. Extract Bearer / cookie
4. Detect token type
5. Validate → AuthResult
6. Check scope vs path
}
RequestAuth --> JWT : validates
RequestAuth --> ApiToken : validates
RequestAuth --> UserTables : lookup + permissions
RequestAuth --> AuthResult : returns
' -- Token Types --
class "JWT (HS256)" as JWT <<jwt>> {
sub: user_id
scopes: "read,write,approve"
src: "password" | "database"
iat, exp (24h default)
--
Detected by: contains "."
Validated locally
No DB call
}
ModelTable --> Registry : load / hot reload
Registry --> ModelConfig : immutable snapshot
Registry --> Lane : coherent binding
class "API Token" as AT <<jwt>> {
Format: ts_ + 64 hex
Stored: SHA-256 hash
--
Detected by: starts with "ts_"
Lookup by hash in DB
Expiry check
}
AuthResult --> Resolver : initiating principal
Lane --> Resolver : callable + pinned config
Resolver --> Mint : dynamic modes only
Mint --> OBOState : decrypt / grant policy
CallClient --> Lane
CallClient --> Resolver
CallClient --> ProviderCall : cloned SDK client
class "Config Token" as CT <<core>> {
Raw value in memory
Role: "read" | "full"
--
Detected by: fallback
hmac.compare_digest
No DB needed
}
note right of Resolver
**Mode policy**
• static: return None; registry client's explicit key remains.
• entra_obo / rfc8693_obo: require an effective principal. HTTP
turns pin the authenticated initiator; single-user internal lanes
may use their session owner. Never borrow another generation's identity.
• entra_app: use deployment app identity, no user required.
• rfc8693_obo alone sends obo_scopes; each dynamic mode is paired
with its required Entra or RFC 8693 grant profile.
' -- Scopes --
class "Scope Hierarchy" as SH <<scope>> {
read: {read}
write: {read, write}
approve: {read, write, approve}
--
GET → read
POST write paths → write
POST /api/workstreams/{ws_id}/approve → approve
/api/admin/* → approve
}
' -- Storage --
class "users" as UT <<storage>> {
user_id (PK)
username (unique)
display_name
password_hash (bcrypt)
created
}
class "api_tokens" as TT <<storage>> {
token_id (PK)
token_hash (SHA-256, unique)
token_prefix
user_id → users
name, scopes
created, expires
}
' -- Endpoints --
class "POST /api/auth/login" as Login <<endpoint>> {
{username, password}
OR {token: "ts_xxx"}
→ {jwt, role, scopes, user_id}
--
Sets HttpOnly cookie
}
class "GET /api/auth/status" as Status <<endpoint>> {
→ {auth_enabled, has_users,
setup_required}
--
Public (no auth)
Drives UI setup wizard
}
class "POST /api/auth/setup" as Setup <<endpoint>> {
{username, display_name, password}
→ {jwt, user_id, scopes}
--
Public (no auth)
Only when zero users exist
Returns 409 if already set up
}
class "Admin API (Console)" as Admin <<endpoint>> {
POST/GET/DELETE users
POST/GET tokens
DELETE tokens/{id}
--
Requires approve scope
}
' -- Relationships --
CR --> AC : config tokens
CR --> JWT : validate
CR --> AT : hash lookup
CR --> CT : hmac check
CR --> AR : returns
CR --> SH : checks
Login --> JWT : issues
Login --> UT : verify password
Login --> TT : verify API token
Setup --> UT : create first user
Setup --> JWT : issues
AT --> TT : lookup by hash
Admin --> UT : CRUD
Admin --> TT : CRUD
AR --> SH : scopes from
JWT ..> AR : produces
AT ..> AR : produces
CT ..> AR : produces
note right of CR
**Middleware Flow**
AuthMiddleware on every request:
1. Extract token from header/cookie
2. Detect type (JWT / ts_ / config)
3. Validate → AuthResult
4. Set ctx_user_id for logging
5. Store auth_result in scope state
end note
note bottom of CallClient
Dynamic credentials are minted at dispatch, not cached in the registry
snapshot. Endpoint, audience, scopes, auth mode, and static-key presence stay
pinned to the same ModelConfig generation as the SDK client. The global
model.auth_fail_closed policy is read live on every mint. A Stop that wins
before or during mint prevents model bytes from being sent afterward.
note bottom of SH
**Console** owns admin endpoints
**Server** validates JWT + config only
Both share JWT signing secret
end note
note bottom of ProviderCall
If minting fails, a configured fail-closed deployment or a keyless alias
raises BackendAuthUnavailableError. A dynamic alias with an explicit static
key may fall back only when policy allows. Authentication refusal is not a
backend-health failure and does not walk to a static fallback model.
note left of JWT
**Console Proxy Token Minting**
When proxying requests to server nodes:
1. Console AuthMiddleware validates user JWT (aud: turnstone-console)
2. Proxy mints new JWT (aud: turnstone-server)
with real user_id, scopes, permissions
3. src: "console-proxy" for audit traceability
4. 5-minute expiry (fresh per request)
5. Fallback: ServiceTokenManager if no user context
end note
@enduml
+14 -30
View File
@@ -82,12 +82,10 @@ class "DiscordBot" as Bot <<service>> {
}
class "ChannelRouter" as Router <<service>> {
+get_or_create_workstream(channel_type, channel_id)
+_is_ws_live(ws_id)
+send_message(ws_id, message)
+send_approval(ws_id, ...)
+lookup_ws_id(channel_type, channel_id)
+resolve_user(channel_type, channel_user_id)
+resolve_route(platform, channel_id)
-> ws_id | None
+register_route(channel_id, ws_id)
+resolve_identity(platform, platform_user_id)
-> user_id | None
--
Maps channels -> workstreams
@@ -95,16 +93,6 @@ class "ChannelRouter" as Router <<service>> {
Caches routes in memory
}
class "turnstone-console router" as ConsoleRouter <<server>> {
POST /v1/api/route/workstreams/new
GET /v1/api/route/workstreams/{ws_id}/live
POST /v1/api/route/workstreams/{ws_id}/send
POST /v1/api/route/workstreams/{ws_id}/approve
GET /v1/api/route?ws_id=...
--
Multi-node rendezvous + durable overrides
}
' -- Server --
class "turnstone-server" as Server <<server>> {
POST /v1/api/workstreams/{ws_id}/send
@@ -160,9 +148,7 @@ Bot --> Router : on_message\non_interaction
Router --> CU : resolve identity
Router --> CR : resolve / register route
Router --> Server : single-node/direct mode\ncreate + send + approve
Router --> ConsoleRouter : multi-node mode\nroute create/live/send/approve/lookup
ConsoleRouter --> Server : routed HTTP to owning node
Router --> Server : POST /v1/api/workstreams/{ws_id}/send\nPOST /v1/api/workstreams/{ws_id}/approve\nPOST /v1/api/workstreams/new
Bot --> Server : GET /v1/api/workstreams/{ws_id}/events\n(SSE via httpx-sse)
Server --> Bot : SSE event stream
@@ -170,7 +156,7 @@ Bot --> Discord : reply / embed\nbutton callback
Slack --> SlackBot : socket-mode\nevents
SlackBot --> Router : on_message / on_action
SlackBot --> Server : owning-node SSE after route lookup
SlackBot --> Server : POST /v1/api/workstreams/{ws_id}/send\nGET /v1/api/workstreams/{ws_id}/events
SlackBot --> Slack : post / update\nBlock Kit button callbacks
Teams .[hidden]. Slack
@@ -189,21 +175,19 @@ note right of Bot
**Inbound Flow**
1. Discord message arrives via gateway
2. Bot.on_message() fires
3. ChannelRouter gets or creates channel -> ws_id
(direct server or multi-node console router)
3. ChannelRouter resolves channel -> ws_id
(or creates new workstream)
4. ChannelRouter resolves platform user -> user_id
via channel_users table
5. Router sends through the configured server/console SDK
5. Router sends POST /v1/api/workstreams/{ws_id}/send to server
**Stale-route recovery (evicted workstreams)**
1. Route health check reports the old ws unavailable
2. Existing ws_id becomes the fork source
**Workstream Resume (evicted workstreams)**
1. Stale route detected (no active SSE listener)
2. Existing ws_id reused directly from route
3. POST /v1/api/workstreams/new with
resume_ws=<ws_id>
4. Server atomically clones source history/config/
persona/project/attachment refs into a new ws_id
5. Router stores the new destination route; source is unchanged
6. If the source was pruned, retry one fresh create
4. Server resumes atomically during creation
5. SSE emits WorkstreamResumedEvent -> thread
end note
note right of Server
+2 -2
View File
@@ -11,7 +11,7 @@ skinparam participant {
participant "ChatSession\n(session.py)" as Session <<session>>
participant "WatchRunner\n(watch.py)" as Runner <<server>>
participant "StorageBackend\n(SQLite / PostgreSQL)" as Storage <<storage>>
participant "StorageBackend\n(SQLite)" as Storage <<storage>>
participant "WebUI / SSE\n(server.py)" as UI <<ui>>
== Create Phase ==
@@ -130,7 +130,7 @@ note right : action="cancel" (auto-approve)
note over Runner, Storage
**Startup:**
1. WatchRunner created in main() with storage + node_id
2. restore_fn closure captures SessionManager
2. restore_fn closure captures WorkstreamManager
3. Initial workstream: session.set_watch_runner(runner)
4. _lifespan(): runner.start() — daemon thread begins
+193 -101
View File
@@ -1,126 +1,218 @@
@startuml
!theme plain
title Turnstone — Intent Judge, Concurrent Approval Cycles, and Output Guard
title Turnstone — Intent Validation (Judge) Architecture
skinparam sequenceArrowThickness 1.5
skinparam sequenceLifeLineBackgroundColor #F5F5F5
skinparam participant {
BackgroundColor<<session>> #C8E6C9
BackgroundColor<<judge>> #FFE0B2
BackgroundColor<<storage>> #B3E5FC
BackgroundColor<<ui>> #E8EAF6
BackgroundColor<<fs>> #F5F5F5
}
participant "ChatSession\ngeneration N" as Session
participant "SessionUIBase" as UI
participant "IntentJudge" as Judge
participant "model_turn()\n(pinned ModelLane)" as Model
participant "Operator / client" as Operator
participant "OutputGuardJudge" as Guard
database "StorageBackend" as Storage
participant "ChatSession\n(session.py)" as Session <<session>>
participant "IntentJudge\n(judge.py)" as Judge <<judge>>
participant "LLM Provider\n(provider)" as LLM <<judge>>
participant "StorageBackend\n(SQLite)" as Storage <<storage>>
participant "WebUI / SSE\n(server.py)" as UI <<ui>>
participant "Filesystem" as FS <<fs>>
== Intent assessment begins during preparation ==
== Tool Call Requires Approval ==
Session -> Session : prepare each tool item independently\nattach principal + cancel witness
Session -> Judge : evaluate(items, callback, cancel_ref)
activate Judge
Judge -> Judge : synchronous heuristic verdict\nfor each call (first matching rule)
Judge --> Session : heuristic verdicts + daemon cancel event
Session -> UI : cache / publish heuristic assessments
Session -> Storage : persist heuristic intent verdicts
note over Judge, Model
The judge owns an immutable resolved binding. Registry/config generations
are freshness watermarks: an effective lane change replaces the judge for
the next batch, while in-flight work keeps the lane it started with.
Dynamic backend auth is resolved for this batch's initiating principal.
parallel_evaluations (1-16) sets per-batch worker width; the model alias's
admission gate remains the process-wide generation ceiling.
Session -> Session : _prepare_tool_calls()
note right
Tool calls parsed from
LLM response. Auto-approved
tools dispatched immediately.
Remaining items need approval.
end note
par LLM judge daemon coordinator
Judge -> Judge : start min(batch size, parallel_evaluations,\npositive alias capacity) workers
loop each worker claims one independent call
Judge -> Model : model_turn(judge lane, canonical Turns,\nread-only evidence tools, cancel_ref)
Model --> Judge : ModelTurnResult
alt evidence tool requested
Judge -> Judge : execute bounded read_file / list_directory
else verdict text
Judge -> Judge : parse + arbitrate against heuristic
Session -> Session : _evaluate_intent(pending_items)
== Tier 1: Heuristic (synchronous, sub-ms) ==
Session -> Judge : evaluate(items, messages, callback)
Judge -> Judge : evaluate_heuristic()\nfor each item
note right
**36 rules (first match wins):**
Critical (0.90, deny): rm /, mkfs,
dd, pipe-to-shell, chmod 777 /,
write/edit /etc/ .ssh/,
download-then-execute chains
High (0.80, review): sudo, kill -9,
destructive git, DROP TABLE,
secrets, HTTP mutations, ssh/scp,
browser+data-export, transitive
install, control-plane mutation
Medium (0.70, review): content
ingestion, interpreter exec,
cloud CLI mutations, pkg install,
write_file, MCP tools, docker ops
Low (0.85, approve): read_file,
list_directory, search, recall,
tool_search, read_resource,
web_search, read-only bash
Default: medium, 0.50, review
end note
Judge --> Session : heuristic_verdicts[]
Session -> Session : attach _heuristic_verdict\nto each pending item
Session -> UI : SSE: approve_request\n{items: [{verdict: ...}],\n judge_pending: true}
note right
Heuristic verdict displayed
immediately as risk badge.
Spinner shown while LLM
judge evaluates.
end note
Session -> Storage : create_intent_verdict()\nfor each heuristic verdict
== Tier 2: LLM Judge (daemon thread, async) ==
Judge -> Judge : spawn daemon thread\n"intent-judge"
note over Judge, LLM
**Context preparation:**
1. FIFO-truncate conversation history
to max_context_ratio of context window
2. Append tool call details as user message
3. System prompt defines judge role + JSON schema
end note
loop up to 3 turns (timeout budget)
Judge -> LLM : create_completion(\nmodel, judge_messages,\ntools=[read_file, list_directory])
LLM --> Judge : CompletionResult
alt tool_calls present (turn < 3)
Judge -> Judge : _exec_read_only_tool()
note right
**Security hardening:**
Blocked: /etc/, /root/,
/proc/, /sys/, /dev/,
.ssh, .gnupg, .aws,
*.pem, *.key, *.p12
File cap: 32KB
Dir cap: 200 entries
end note
Judge -> FS : read_file / list_directory
FS --> Judge : file contents
Judge -> Judge : append tool result\nto judge_messages
else text response (final verdict)
Judge -> Judge : _parse_verdict()
note right
**4-stage JSON parsing:**
1. Direct JSON.loads
2. Markdown code block
3. Brace-counting
4. Regex field extraction
end note
end
Judge --> UI : on_intent_verdict(verdict, judge generation)
UI -> Storage : persist LLM verdict / audit update
end
else approval path continues
Session -> UI : approve_tools(items) with one\nSmart Approval config snapshot
end
== Policy, Smart Approval, and human gate ==
== Tier 3: Arbitration ==
UI -> UI : apply explicit policy / skill / always / blanket bypasses
opt Smart Approvals enabled
UI -> UI : wait within captured deadline for this batch's verdicts
UI -> UI : auto-approve only recommendation=approve\nand confidence >= captured threshold
UI -> Storage : persist auto-approval reason and decision
end
alt human-gated items remain
UI -> UI : acquire publication lease; register ApprovalCycle\n(cycle_id, call_ids, event, result, witnesses)
UI -> Operator : approve_request with cycle_id + item verdicts
Operator -> UI : approve / deny by cycle_id or call_id
UI -> UI : atomically claim exactly one unresolved cycle
UI -> Operator : approval_resolved
UI --> Session : decision + optional feedback
UI -> Storage : stamp tracked verdicts with operator decision
else every item bypassed / auto-approved
UI -> Operator : tool_info with exact auto_approve_reason
UI --> Session : approved
end
note right of UI
Parallel task agents may register several ApprovalCycles. Each cycle owns
its own Event and result slot. A legacy selector-less decision targets the
oldest cycle; double resolution is a no-op. Cached LLM verdicts carry their
judge generation, so reused provider call ids cannot satisfy a new cycle.
Judge -> Judge : compare confidence:\nLLM vs heuristic
note right
Only deliver LLM verdict
if confidence > heuristic.
Otherwise heuristic stands.
end note
== Cancellation boundary ==
opt Stop / close / force-successor
Session -> Judge : abort all judge events owned by the cancelled operation
Session -> UI : resolve_all_approvals(False, "cancelled")
UI -> UI : block new admission leases; wait for admitted bundles;\nclaim only cycles whose cancellation witness is aborted
UI -> Operator : one cancelled resolution per claimed cycle
note over Session, UI
A Stop can win before cycle registration, during publication, or while a
click resolves. The witness + admission sweep makes exactly one terminal
outcome visible; a successor generation's new cycle is not swept.
end note
alt LLM confidence > heuristic confidence
Judge -> Session : callback(llm_verdict)
Session -> UI : SSE: intent_verdict\n{tier: "llm", ...}
note right
UI replaces heuristic badge
with LLM verdict. Spinner
resolves to final assessment.
end note
Session -> Storage : create_intent_verdict()\nfor LLM verdict
end
note over Judge
Normal operator resolution does not necessarily cancel judge inference.
With cancel_on_approval=false, the daemon finishes and late verdicts remain
auditable. With it enabled, the batch event stops remaining judge work.
== User Decision ==
UI -> Session : resolve_approval(\napproved, feedback)
Session -> Storage : update_intent_verdict(\nverdict_id, user_decision)
note right
All tracked verdicts
(heuristic + LLM) updated
with "approved" or "denied".
Swap-and-clear avoids racing
with daemon judge thread.
end note
deactivate Judge
== Tool Execution ==
== Tool output guard ==
Session -> Session : _execute_tools()
note right
Tools execute with
user approval.
end note
Session -> Session : execute admitted tools; truncate each result
Session -> Guard : evaluate(result, tool context, cancel event)
activate Guard
Guard -> Guard : heuristic checks first
opt LLM guard enabled and time remains
Guard -> Model : model_turn(output-guard lane, bounded prompt, cancel_ref)
Model --> Guard : structured verdict
== Output Guard (synchronous, time-budgeted) ==
Session -> Session : _evaluate_output()\nfor each tool result
note right
**Priority-ordered checks (5s budget):**
P1: Prompt injection (role injection,
override phrases, instruction tags)
P2: Credential leakage (API keys,
PEM blocks, connection strings)
P3: Encoded payloads (data URIs,
hex shellcode)
P4: Adversarial URLs (cloud metadata,
credential query params)
P5: System info disclosure (private
IPs, sensitive paths)
Annotates + optionally redacts.
Does NOT gate.
end note
alt output_warning flags detected
Session -> UI : SSE: output_warning\n{call_id, risk_level, flags,\nfunc_name, redacted}
note right
Credential values replaced
with [REDACTED:<type>] before
output enters conversation.
sanitized text excluded from
SSE payload (defense in depth).
end note
UI -> Storage : record_output_assessment()\nfire-and-forget persistence
note right
Stored: flags, risk_level,
annotations, output_length,
redacted (bool). Raw tool
output is never stored.
end note
end
Guard --> Session : assessment / redaction / warning
deactivate Guard
Session -> Session : re-check generation N before folding result
Session -> UI : output warning (no raw secret payload)
Session -> Storage : persist assessment + guarded Tool Turn metadata
note over Guard, Storage
Output-guard objects also pin model/config lanes. Replacement retires the
old object but lets admitted evaluations drain before its private client is
closed. A cancelled or superseded evaluation cannot fold into the successor
trajectory. Raw pre-redaction secrets are never stored in assessment rows.
== Lifecycle ==
note over Session, Judge
**Lazy initialization:**
IntentJudge created on first approval if judge_config.enabled.
Re-uses session's provider/client by default (self-consistency).
Cross-model: separate provider/client from [judge] config.
**Sub-agent exemption:**
Task sub-agents skip intent validation entirely.
**Output guard:**
Runs when judge_config.output_guard is true (default).
Credential redaction when judge_config.redact_secrets is true.
**Storage:**
intent_verdicts table (migration 012), output_assessments table
(migration 022). Both queryable via admin API endpoints
(requires admin.judge permission). Skills store risk_level,
scan_report, scan_version for install-time risk assessment.
end note
@enduml
+39 -45
View File
@@ -13,68 +13,65 @@ skinparam participant {
participant "ChatSession\n(session.py)" as Session <<session>>
participant "MemoryFacade\n(memory.py)" as Facade <<facade>>
participant "MemoryRelevance\n(memory_relevance.py)" as Relevance <<facade>>
participant "StorageBackend\n(SQLite / PostgreSQL)" as Storage <<storage>>
participant "StorageBackend\n(SQLite)" as Storage <<storage>>
participant "Server API\n(server.py)" as API <<api>>
participant "Console Admin\n(console/server.py)" as Admin <<api>>
participant "SDK Client\n(sdk/)" as SDK <<sdk>>
== Phase 1: Tool Path (session.send) ==
Session -> Session : pin acting principal\nparse memory(action=...)
Session -> Session : _prepare_tool_calls()\nparse memory(action=...)
note right
Tool schema: 5 actions
save, get, search, delete, list
Tool schema: 4 actions
save, search, delete, list
Auto-approved (no approval needed)
end note
Session -> Session : resolve live project access\nselect exact/inherited scope
Session -> Session : _exec_memory(item)
alt action = save
Session -> Session : require non-empty description
Session -> Facade : save_structured_memory_strict(\n..., require_active_project)
Session -> Facade : save_structured_memory(\nname, content, description,\nmem_type, scope, scope_id)
Facade -> Facade : normalize_key(name)
Facade -> Storage : guarded atomic upsert\nON CONFLICT ... RETURNING
Storage --> Facade : (saved row, was_update)
Facade --> Session : saved row
Session -> Session : invalidate prefix/cache\naudit acting principal
end
alt action = get
Session -> Facade : get_structured_memory_by_name_strict()
Facade -> Storage : exact scoped-name lookup
Storage --> Session : full row / not found
Facade -> Storage : create_structured_memory()
alt unique constraint violation
Storage --> Facade : IntegrityError
Facade -> Storage : get_structured_memory_by_name()
Storage --> Facade : existing row
Facade -> Storage : update_structured_memory()
end
Storage --> Facade : memory_id
Facade --> Session : (memory_id, old_content)
Session -> Session : _init_system_messages()\nrefresh BM25 context
end
alt action = search
Session -> Storage : search exact scope or\nactor-visible scope union
Session -> Facade : search_structured_memories(\nquery, mem_type, scope,\nscope_id, limit)
Facade -> Storage : search_structured_memories()
Storage --> Session : matched rows
end
alt action = delete
Session -> Facade : delete_structured_memory_returning_strict()
Facade -> Storage : DELETE ... RETURNING
Storage --> Session : deleted row / not found
Session -> Session : invalidate + audit\nmark prefix dirty
Session -> Facade : delete_structured_memory(\nname, scope, scope_id)
Facade -> Storage : delete_structured_memory()
Storage --> Session : bool (existed)
Session -> Session : _init_system_messages()\nrefresh BM25 context
end
== Phase 2: BM25 Relevance Injection ==
Session -> Session : _init_system_messages()\nevery conversation turn
Session -> Session : resolve acting principal\nand live project ACL
Session -> Session : _list_visible_memories(\nlimit=fetch_limit)
note right
**Scope resolution:**
Interactive: global + workstream
+ acting user + readable project
Coordinator: acting user's coordinator
+ readable project
1. global scope (always)
2. workstream scope (ws_id)
3. user scope (user_id, if auth)
Combined and deduplicated.
end note
Session -> Facade : list_visible_structured_memories()
Facade -> Storage : one visibility-union query
Session -> Facade : list_structured_memories()\nper scope
Facade -> Storage : list_structured_memories()
Storage --> Session : up to fetch_limit rows
Session -> Relevance : extract_recent_context(\nmessages, max_messages=3)
@@ -106,31 +103,28 @@ Session -> Session : inject into\nsystem message
== Phase 3: Server API Path ==
SDK -> API : GET /v1/api/memories\n?type=general&limit=20
API -> API : bind scope to caller\ndefault global + caller user
API -> Storage : list visible rows
SDK -> API : GET /v1/api/memories\n?type=project&limit=20
API -> Facade : list_structured_memories()
Facade -> Storage : list_structured_memories()
Storage --> API : rows
API --> SDK : {"memories": [...], "total": N}
SDK -> API : POST /v1/api/memories\n{name, content, description, ...}
API -> API : validate type, scope,\nname/content/description
API -> API : reject internal scopes\nowner-bind workstream scope
API -> Facade : save_structured_memory_strict()
Facade -> Storage : atomic upsert
SDK -> API : POST /v1/api/memories\n{name, content, ...}
API -> API : validate type, scope,\nname length, content length
API -> Facade : save_structured_memory()
Facade -> Storage : create / update
Storage --> API : memory row
API -> API : record_audit(actor)
API --> SDK : 201 (created) / 200 (updated)
SDK -> API : POST /v1/api/memories/search\n{query, type, ...}
API -> API : bind scope to caller
API -> Storage : search visible rows
API -> Facade : search_structured_memories()
Facade -> Storage : search_structured_memories()
Storage --> API : matched rows
API --> SDK : {"memories": [...], "total": N}
SDK -> API : DELETE /v1/api/memories/{name}\n?scope=global
API -> Facade : delete_structured_memory_returning_strict()
Facade -> Storage : DELETE ... RETURNING
API -> API : record_audit(actor)
API -> Facade : delete_structured_memory()
Facade -> Storage : delete row
API --> SDK : {"status": "ok"}
== Phase 4: Console Admin Path ==
@@ -147,7 +141,7 @@ Storage --> Admin : memory row
Admin --> SDK : memory JSON
SDK -> Admin : DELETE /v1/api/admin/memories/{id}
Admin -> Storage : delete_structured_memory_by_id_returning()
Admin -> Storage : delete_structured_memory_by_id()
Admin -> Admin : record_audit(\n"memory.delete")
Admin --> SDK : {"status": "ok"}
+6 -8
View File
@@ -13,7 +13,7 @@ skinparam participant {
participant "Server\n(main)" as Server <<session>>
participant "ConfigStore\n(config_store.py)" as Store <<config>>
participant "SettingsRegistry\n(settings_registry.py)" as Registry <<config>>
participant "StorageBackend\n(SQLite / PostgreSQL)" as Storage <<storage>>
participant "StorageBackend\n(SQLite)" as Storage <<storage>>
participant "Console Admin\n(console/server.py)" as Admin <<api>>
participant "SDK Client\n(sdk/)" as SDK <<sdk>>
participant "ChatSession\n(session.py)" as Session <<session>>
@@ -57,10 +57,9 @@ else key not in cache
Store --> Session : default value
end
note right of Session
Most session settings are captured once
at workstream creation. Documented live readers
(including model.auth_fail_closed per mint)
apply immediately.
Settings are captured once
at workstream creation.
Not re-read on every turn.
end note
== Phase 3: Admin API — List / Schema ==
@@ -129,9 +128,8 @@ Store -> Storage : get_system_settings_bulk(node_id)
Storage --> Store : all settings
Store -> Store : rebuild cache,\nswap atomically,\nincrement _version
note right
Most existing-session settings are unchanged
(frozen at creation time); documented
live readers apply immediately.
Existing sessions: unchanged
(frozen at creation time).
New sessions: pick up
updated values immediately.
end note
+5 -5
View File
@@ -88,7 +88,7 @@
<rect x="300" y="151" width="160" height="2" fill="#161b22"/>
<text x="380" y="178" text-anchor="middle" fill="#e6edf3" font-size="11" font-weight="600">Console</text>
<line x1="318" y1="190" x2="442" y2="190" stroke="#30363d" stroke-width="1"/>
<text x="380" y="208" text-anchor="middle" fill="#8b949e" font-size="9">FNV-1a rendezvous router</text>
<text x="380" y="208" text-anchor="middle" fill="#8b949e" font-size="9">hash-ring router</text>
<text x="380" y="223" text-anchor="middle" fill="#8b949e" font-size="9">cluster dashboard</text>
<text x="380" y="238" text-anchor="middle" fill="#8b949e" font-size="9">reverse proxy</text>
<line x1="318" y1="250" x2="442" y2="250" stroke="#30363d" stroke-width="1"/>
@@ -115,7 +115,7 @@
<rect x="608" y="162" width="184" height="28" rx="3" fill="#1c2128" stroke="#30363d" stroke-width="1"/>
<text x="700" y="180" text-anchor="middle" fill="#8b949e" font-size="9">turnstone-server :8080</text>
<!-- Tools label -->
<text x="700" y="206" text-anchor="middle" fill="#484f58" font-size="8">model lanes + tools / MCP</text>
<text x="700" y="206" text-anchor="middle" fill="#484f58" font-size="8">19 tools + MCP</text>
</g>
<!-- Node B -->
@@ -130,7 +130,7 @@
<rect x="608" y="282" width="184" height="28" rx="3" fill="#1c2128" stroke="#30363d" stroke-width="1"/>
<text x="700" y="300" text-anchor="middle" fill="#8b949e" font-size="9">turnstone-server :8080</text>
<!-- Tools label -->
<text x="700" y="326" text-anchor="middle" fill="#484f58" font-size="8">model lanes + tools / MCP</text>
<text x="700" y="326" text-anchor="middle" fill="#484f58" font-size="8">19 tools + MCP</text>
</g>
<!-- ==================== LLM PROVIDERS ==================== -->
@@ -171,7 +171,7 @@
<rect x="590" y="450" width="220" height="5" fill="#bc8cff"/>
<rect x="590" y="453" width="220" height="2" fill="#161b22"/>
<text x="700" y="476" text-anchor="middle" fill="#e6edf3" font-size="11" font-weight="600">PostgreSQL / SQLite</text>
<text x="700" y="492" text-anchor="middle" fill="#8b949e" font-size="9">workstreams, turns, config, auth</text>
<text x="700" y="492" text-anchor="middle" fill="#8b949e" font-size="9">conversations, memory, auth</text>
</g>
<text x="700" y="444" text-anchor="middle" fill="#bc8cff" font-size="9" font-weight="600" letter-spacing="2">STORAGE</text>
@@ -239,7 +239,7 @@
<!-- Routing rules at bottom, left-aligned -->
<text x="44" y="456" fill="#30363d" font-size="9" font-weight="600" letter-spacing="1">ROUTING</text>
<circle cx="44" cy="474" r="3" fill="#3fb950" opacity="0.6"/>
<text x="54" y="477" fill="#484f58" font-size="9">control plane: client &#x2192; console &#x2192; server node (FNV-1a rendezvous placement)</text>
<text x="54" y="477" fill="#484f58" font-size="9">control plane: client &#x2192; console &#x2192; server node (hash-ring bucket lookup)</text>
<circle cx="44" cy="494" r="3" fill="#58a6ff" opacity="0.6"/>
<text x="54" y="497" fill="#484f58" font-size="9">data plane: client &#x2192; server node (direct SSE, node_url from create response)</text>
<circle cx="44" cy="514" r="3" fill="#f47067" opacity="0.6"/>

Before

Width:  |  Height:  |  Size: 16 KiB

After

Width:  |  Height:  |  Size: 16 KiB

-3
View File
@@ -1,3 +0,0 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e2bfdf968e96f3720ed58674103288e2f57e9c056f5c479a57f37a849f3e69c
size 821878
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b8c1460440784f07e30afea32d4ee17687627df46a24003d761ad79c2676a361
size 169499
oid sha256:881a8b9bce67b5af9a52d5e50deaa72351cd99c76f18aad5caeb2b61131ca1af
size 119798
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:66847ccdf10ef2bd04e93bc0d3924a56ce28462ec9e76a383b53aee4500755e8
size 631799
oid sha256:95dd5ebc899a1261d516686a5aa3319a7f45015d411302825fa28afbfc82e1ce
size 326766
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c74e99c530c3a8af9ab35b1e4d8c4fef0ea35c0c04cc35da7cf3588e71382057
size 661175
oid sha256:9857db23fe3c4316d492073aac69c7e7558b1abe3b95ad7756d4a5933bd0ece7
size 620214
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6261604cc8b75878a8704308929ea64d121cf26547019543fbe1b21cbe700415
size 189791
oid sha256:d9c7769a600c38e6387390e6c42db8152e0f80c31d17b2218f7f636b71c7b868
size 355459
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1b3b7b745f6006ee73d4b31fa598faa0d71ffb74ce349ab08eb3ce09ded506c3
size 266294
oid sha256:23ca090b5656baaf70820cbe4ab6c27f0a3a02e18b4db0695614cf9489c23980
size 281440
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:33dddd8cd8b53fa464cc8a4c899896fee32035d63e0669fe327969e3356349c7
size 329815
oid sha256:04d2069a9b5155ad1e7d842147fd78535ad9106d6856520439c33a9868a47499
size 156694
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:16e3f3bfa0a6af637f7a9fb6765d594eb598428679c88a429c096c3dbae931e4
size 181185
oid sha256:a872556d111185f4531d1b68ee892b4ce5042d7ccf277e2cad08beb6932c9803
size 191144
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:59f14b835665244f3d32981b6c1ac4c4380393a83cc519e271831622aa3f261a
size 197433
oid sha256:e7c3e40c10425d721f833390ae3531c09af501157fd3142531ba4eba86ff719d
size 197112
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1a510eeaabf4ed8dab3b268c8f6bb5b7fef629a664361f7fb5614a6b498db36e
size 294415
oid sha256:b047cdc318c505f0f0895a65e14c5cc7552716053055cca57fa0a77db150e618
size 255458
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ea86c6b6c68ed96f6fd18543d2e7a873f4715332cc3a7d4df167392f668a2de7
size 232403
oid sha256:af5ab3126bf685afe68e24bc4b0ed97371d0ebdb77bf4d76c0331ab120580cc0
size 248809
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:edf02b97e1e1287ebba9e74b9858474dcda42e5542656e505ea133a5b2416f47
size 402992
oid sha256:ae4f79fb22600106f8cb0af4ba5586bb26ea5d57e27ef382fdc59b6549fdbd21
size 415473
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aa9ca9a367c79159a26d1ec544b20fc7118a49082d72f7ee0edbaa85608d49fc
size 238991
oid sha256:96176a09e65e90dadc32d5e9ed778423842be89204d2cf382225f53a90cfaf01
size 258547
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:636a6b2fc1075e4863421e68b99efe7f6f6f62cedbcff36ef0934c055f39fd46
size 281161
oid sha256:79a690c466a5d6f6d4292d78a27b9474e9e9c1373fa80e17dfe37700238c8af8
size 382508
+2 -2
View File
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:137d6c91a34695c820d8b0a33fd753e79165604aa92bf2ac8480d3744b2ef844
size 305199
oid sha256:c89628ed917dfd576c1af75c68fe5fed9beadaaee9dcea7aa7a1643867c4f1b9
size 344323
@@ -1,3 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0455e0dec36ebb8bcfdadcf327a1dd24ddbdcdd8df211c08918180842497727e
size 318681
oid sha256:06fe076f0835a891e00afc804fd1805196ebde9fc0d34998c7873e87287f982b
size 346887
+16 -90
View File
@@ -59,11 +59,9 @@ is gone. Everything goes through `https://localhost:8443`.
## Join a bare-metal host
PostgreSQL, the console's ACME endpoint (`:8090`), and SearxNG (`:8081`) are
published on `127.0.0.1`, so a `turnstone-server` running directly on the same
machine — for example to use a local GPU — can join the same cluster (enrolling
its mTLS cert and running `web_search`) and show up in the console alongside the
containerized nodes.
PostgreSQL is published on `127.0.0.1:5432`, so a `turnstone-server` running
directly on the same machine — for example to use a local GPU — can join the
same cluster and show up in the console alongside the containerized nodes.
Put the secret and connection settings in `~/.config/turnstone/config.toml`
(secrets belong in this file, not the process environment — keep it `0600`,
@@ -87,44 +85,18 @@ command line:
```bash
chmod 600 ~/.config/turnstone/config.toml
TURNSTONE_NODE_ID=host-1 \
TURNSTONE_ADVERTISE_URL=http://host.docker.internal:8080 \
TURNSTONE_CONSOLE_URL=http://localhost:8090 \
TURNSTONE_SEARXNG_URL=http://localhost:8081 \
TURNSTONE_NODE_ID=host-1 TURNSTONE_ADVERTISE_URL=http://host.docker.internal:8080 \
turnstone-server --host 0.0.0.0 --port 8080
```
The host server registers itself in PostgreSQL; the console reaches it back via
`host.docker.internal`. `TURNSTONE_CONSOLE_URL` points the node at the console's
published ACME endpoint so it can enroll its mTLS certificate (needed only when
the cluster runs mTLS; harmless otherwise), and `TURNSTONE_SEARXNG_URL` points
`web_search` at the published SearxNG. The `jwt_secret` and DB credentials above
are the dev-stack defaults — match whatever you set in `.env` if you changed them.
To let a server on a **different** machine join, start the stack with
both `TURNSTONE_HOST_IP=<this host's LAN IP>` and
`TURNSTONE_ACME_EXTERNAL_URL=http://<this host's LAN IP>:8090/acme`. The first
binds PostgreSQL, the console ACME endpoint, and SearxNG to that interface; the
second makes every URL in the ACME directory routable from the remote node (the
full value must include the `/acme` mount). Set the same
`TURNSTONE_ACME_EXTERNAL_URL` on the remote node so its authenticated ACME
client can pin that credential destination. Set `TURNSTONE_CONSOLE_URL` and
`TURNSTONE_SEARXNG_URL` to the compose host's IP, but set
`TURNSTONE_ADVERTISE_URL=http://192.0.2.10:8080` to the **remote** box's own
address. A resolvable DNS name works too. IPv6 literals must be bracketed in
URLs, for example `http://[2001:db8::10]:8080`; Turnstone enrolls literal
addresses as IP SANs rather than numeric DNS SANs.
Use a trusted LAN or VPN address and firewall `:8090` to enrolling nodes. ACME
signing routes require a dedicated short-lived service JWT, but direct bootstrap
is still plain HTTP/TOFU: a bearer token provides authentication, not transport
confidentiality or protection from an active on-path attacker. **Set a strong
`POSTGRES_PASSWORD` first** — `TURNSTONE_HOST_IP` also exposes the database (and
every user account + API-token hash in it), the console API, and the
unauthenticated SearxNG to your network.
To run the bare-metal node as a hardened, persistent service instead of by hand,
use the systemd units in [`deploy/systemd/`](../deploy/systemd/).
`host.docker.internal`. The `jwt_secret` and DB credentials above are the
dev-stack defaults — match whatever you set in `.env` if you changed them. To
let a **different** machine join, start the stack with `POSTGRES_BIND=0.0.0.0`
and use the host's routable IP in the `url` and `TURNSTONE_ADVERTISE_URL`
but **set a strong `POSTGRES_PASSWORD` first**, or you'll expose a database with
the insecure default password (and every user account + API-token hash in it) to
your network.
## Production stack
@@ -138,7 +110,7 @@ docker compose -f turnstone/deploy/compose.yaml up
It's the same shape as the dev stack — Caddy-fronted console, channel, and a
PostgreSQL all share one database so the console discovers the node — but it
pulls released images, runs a single server node, and has **no baked-in
secrets**. Set these in `.env` first (generate with `openssl rand -hex 32`):
secrets**. Set these in `.env` first (`turnstone-bootstrap` generates them):
```bash
TURNSTONE_JWT_SECRET=<python -c "import secrets; print(secrets.token_hex(32))">
@@ -161,11 +133,6 @@ certs via the console's ACME endpoint:
docker compose -f turnstone/deploy/compose.yaml -f deploy/docker-compose.tls.yml up
```
The overlay publishes the console's plain-HTTP bootstrap/API port on
`TURNSTONE_CONSOLE_HTTP_BIND` (default `127.0.0.1`). For a cross-host node, set
that to a trusted LAN/VPN address, set `TURNSTONE_ACME_EXTERNAL_URL` to the same
address plus `/acme`, and firewall the port to enrolling nodes.
See [tls.md](tls.md) for details.
## Configuration
@@ -204,29 +171,18 @@ overrides.
> put [PgBouncer](pgbouncer.md) (transaction pooling) between turnstone and
> PostgreSQL.
> **Lifecycle upgrade:** the release that introduces hidden deferred-create
> reservations must be deployed as a coordinated cohort across every server
> sharing PostgreSQL; older processes do not understand `state='creating'`.
> Drain create traffic until the cohort is upgraded. See
> [PgBouncer: deferred workstream creation](pgbouncer.md#upgrade-note-deferred-workstream-creation).
### Ports
Both stacks publish Caddy (dashboard) and PostgreSQL; the dev stack additionally
publishes the console's ACME endpoint and SearxNG on localhost so a bare-metal
node can enroll its cert and run `web_search`. Everything else is reached through
Caddy or proxied by the console:
publishes the SearxNG UI on localhost. Everything else is reached through Caddy or
proxied by the console:
| Variable | Default | Description |
|----------|---------|-------------|
| `CONSOLE_HTTPS_PORT` | `8443` | Host port for Caddy (dashboard HTTPS) |
| `SEARXNG_HTTPS_PORT` | `8444` | Host port for the SearxNG UI via Caddy (dev: localhost-only; prod: opt-in) |
| `POSTGRES_PORT` | `5432` | Host port for PostgreSQL (for bare-metal joins) |
| `SEARXNG_API_PORT` | `8081` | Host port for the SearxNG API a bare-metal node's `web_search` dials (dev stack) |
| `TURNSTONE_HOST_IP` | `127.0.0.1` | Interface PostgreSQL, the console ACME endpoint, and SearxNG bind on (dev stack). Set to this host's LAN IP so a bare-metal node on **another machine** can reach them — set a strong `POSTGRES_PASSWORD` first (it also exposes the DB and the unauthenticated SearxNG to your network). |
| `TURNSTONE_CONSOLE_HTTP_BIND` | `127.0.0.1` | Production TLS-overlay interface for the console's plain-HTTP bootstrap/API listener. Use only a trusted LAN/VPN address and firewall it to enrolling nodes. |
| `TURNSTONE_ACME_EXTERNAL_URL` | request-derived | Canonical externally reachable ACME responder base, including the final `/acme` mount (for example `http://192.0.2.1:8090/acme`). Set it on the console and clients for cross-host mTLS: the console advertises it, while clients pin it as an allowed enrollment-JWT destination. A reverse-proxy prefix is supported only when the proxy maps it to Turnstone's internal `/acme` mount. |
| `POSTGRES_BIND` | `127.0.0.1` | Production stack (`turnstone/deploy/compose.yaml`) only: interface PostgreSQL binds on; set to the host's LAN IP for remote joins. |
| `POSTGRES_BIND` | `127.0.0.1` | Interface PostgreSQL binds on; set `0.0.0.0` for LAN access |
### Channel gateway
@@ -278,7 +234,6 @@ interface, or anyone who can reach it can search through your instance.
| Variable | Default | Description |
|----------|---------|-------------|
| `WORKSPACE_MOUNT` | empty volume | Host directory bind-mounted at `/workspace` for the model to read/write |
| `TURNSTONE_WORKSPACE` | `/workspace` (image env) | Directory named as the user's workspace in the model's tool descriptions; informational only — see [Working directory](#working-directory) |
| `SKIP_PERMISSIONS` | — | Set to any value to auto-approve all tool calls (dev only) |
| `MCP_CONFIG` | — | Path to an MCP server config file |
| `TURNSTONE_IMAGE_TAG` | `latest` | ghcr.io image tag — production stack |
@@ -287,7 +242,7 @@ interface, or anyone who can reach it can search through your instance.
Both stacks install all entry points into a single image (`turnstone`,
`turnstone-server`, `turnstone-console`, `turnstone-channel`, `turnstone-admin`,
`turnstone-eval`, `turnstone-optimizer`, `turnstone-doctor`):
`turnstone-eval`, `turnstone-bootstrap`):
```bash
docker compose build # build the dev image
@@ -303,35 +258,6 @@ docker compose build --no-cache # rebuild from scratch
| `workspace` | `/workspace` (unless `WORKSPACE_MOUNT` is set) |
| `caddy-data` / `caddy-config` | Caddy's local CA and config (dev stack) |
## Working directory
Node processes run with `/data` as their working directory (the image's
`WORKDIR`), and that is where the model's shell commands execute and
relative file paths resolve — **not** `/workspace`. The shell and file
tool descriptions state both paths (the working directory, and the
workspace named by `TURNSTONE_WORKSPACE`), so the model knows to look in
`/workspace` for your files without being told each session.
To make tools start inside the mount instead, override the working
directory on the node services:
```yaml
services:
turnstone-node:
working_dir: /workspace
```
Two caveats before overriding:
- **SQLite fallback**: when a node runs without PostgreSQL, its fallback
database `.turnstone.db` is created in the process working directory.
Changing `working_dir` on an existing SQLite-fallback deployment makes
the node create a fresh database inside the mount and your prior state
appears lost (it is still in the `turnstone-data` volume under `/data`).
The stock compose stacks use PostgreSQL and are unaffected.
- Migrations (`entrypoint.sh`) run in the same working directory, so the
same SQLite caveat applies to them.
## Cleanup
```bash
+32 -66
View File
@@ -1,19 +1,11 @@
# Evaluation and Prompt Optimization (turnstone-eval, turnstone-optimizer)
# Evaluation and Prompt Optimization (turnstone-eval)
Evaluation for turnstone is split into two commands:
`turnstone-eval` is the evaluation and prompt optimization system for turnstone. It
runs test cases against the LLM, scores tool call sequences against expected
actions, and optionally uses a multi-agent pipeline to optimize the developer
prompt and tool descriptions.
- **`turnstone-eval`** — the measurement substrate. Runs test cases against the LLM
and scores tool call sequences against expected actions. A single measurement pass,
no self-modification.
- **`turnstone-optimizer`** — the prompt/tool optimizer. Loops over the measurement
substrate, using a multi-agent pipeline (analyst, optimizer, observer, diversifier,
tool optimizer) to edit the developer prompt and tool descriptions so more tests pass.
The dependency is strictly one-way: the optimizer consumes the eval substrate; the
substrate never depends on the optimizer.
Source: `turnstone/eval/core.py` (measurement substrate), `turnstone/eval/cli.py`
(the `turnstone-eval` CLI), `turnstone/optimizer.py` (the `turnstone-optimizer` CLI).
Source: `turnstone/eval.py`
---
@@ -35,8 +27,8 @@ This approach (inspired by [Learning to Self-Evolve](https://arxiv.org/abs/2603.
prevents irrecoverable collapse from bad edits — UCB naturally backtracks to
high-scoring ancestors instead of following a linear chain.
The `turnstone-eval` command (or `turnstone-optimizer --no-optimize`) executes only
steps 2-4: a single measurement pass over the root prompt, no optimization.
When optimization is disabled (`--no-optimize`), only steps 2-4 execute
(a single iteration evaluating the root node).
---
@@ -177,8 +169,7 @@ Runs a complete multi-turn conversation:
1. Appends the user message.
2. Checks `_cancelled` event — stops if set (timeout cleanup).
3. Calls the model through the production streaming provider path and drains
the result.
3. Calls the model API (non-streaming).
4. If tool calls are returned, executes them (with stdout suppressed) and
logs each call to `self.tool_call_log`.
5. Repeats up to `max_turns` or until the model responds without tool calls.
@@ -191,15 +182,14 @@ Parallel tool calls are capped at 10 per turn to prevent degenerate repetition.
Each test runs in a `ThreadPoolExecutor(max_workers=1)` with a per-test
timeout (`--test-timeout`). Each attempt gets its own `OpenAI` client with
a matching per-read HTTP transport timeout. Because a trickling stream can
continually reset that read timeout, three layers bound the harness and stop
follow-on work:
a matching httpx read timeout. On timeout, three layers of defense prevent
zombie connections:
1. **Executor wall clock**: The harness stops waiting after `--test-timeout`.
1. **httpx timeout**: Per-request read timeout aborts the HTTP call and
releases the server slot.
2. **`_cancelled` event**: Prevents the orphan thread from starting new turns.
3. **`run_client.close()`**: Retires the connection pool and prevents reuse.
HTTPX2 does not promise that cross-thread client closure immediately aborts
an active body read; that read unwinds on its next wire event or read timeout.
3. **`run_client.close()`**: Closes the connection pool to abort any
in-flight request.
### Retry Logic
@@ -216,8 +206,7 @@ Each test case runs in isolation:
1. A fresh temp directory is created.
2. Setup files are written to the temp directory.
3. The working directory is changed to the temp directory.
4. A per-attempt `OpenAI` client is created with an HTTP transport timeout
matching `--test-timeout`.
4. A per-attempt `OpenAI` client is created with httpx timeout matching `--test-timeout`.
5. A new `HeadlessSession` is created with the current developer prompt.
6. `send_headless()` runs the user prompt through the conversation loop.
7. The tool log is scored against expected actions.
@@ -463,46 +452,30 @@ structure is:
## CLI Usage
Two console scripts (installed as entry points), or the equivalent `python -m`
invocations:
- `turnstone-eval` / `python -m turnstone.eval.cli` — measure only.
- `turnstone-optimizer` / `python -m turnstone.optimizer` — optimize.
### Measure (`turnstone-eval`)
The entry point is `turnstone-eval` (installed as a console script) or
`python -m turnstone.eval`.
```
turnstone-eval tests.json # one measurement pass, print scores
turnstone-eval tests.json --prompt custom.txt # measure a custom prompt
turnstone-eval tests.json --n-runs 5 # more runs per case
turnstone-eval tests.json --parallel 4 # run cases across 4 workers
turnstone-eval tests.json -v # verbose per-turn logging
turnstone-eval tests.json # evaluate + optimize
turnstone-eval tests.json --no-optimize # evaluate only (single iteration)
turnstone-eval tests.json --n-runs 5 --max-iter 10 # more thorough evaluation
turnstone-eval tests.json --prompt custom.txt # start from a custom prompt
turnstone-eval tests.json --optimize-tools # optimize tool descriptions only
turnstone-eval tests.json --diversify 10 # test with prompt variants
turnstone-eval tests.json -v # verbose per-turn logging
```
### Optimize (`turnstone-optimizer`)
### Multi-model setup (local test model, cloud optimizer)
```
turnstone-optimizer tests.json # evaluate + optimize
turnstone-optimizer tests.json --no-optimize # single pass, no optimization
turnstone-optimizer tests.json --n-runs 5 --max-iter 10 # more thorough optimization
turnstone-optimizer tests.json --prompt custom.txt # start from a custom prompt
turnstone-optimizer tests.json --optimize-tools # optimize tool descriptions only
turnstone-optimizer tests.json --diversify 10 # test with prompt variants
```
#### Multi-model setup (local test model, cloud optimizer)
```
turnstone-optimizer tests.json \
turnstone-eval tests.json \
--base-url http://localhost:8000/v1 \
--optimizer-base-url https://api.anthropic.com \
--optimizer-model claude-sonnet-4-6 \
--analyst-model claude-opus-4-6
```
### Measurement Options
Accepted by **both** commands.
### All Options
| Flag | Default | Description |
|-------------------------|----------------------------|-------------|
@@ -511,26 +484,19 @@ Accepted by **both** commands.
| `--model` | auto-detect | Model name. Auto-detected from the API if not specified. |
| `--prompt` | turnstone built-in prompt | Path to initial prompt text file. |
| `--n-runs` | from tests.json or 3 | Number of runs per test case. |
| `--max-iter` | 5 | Maximum optimization iterations. |
| `--no-optimize` | false | Run evaluation only (sets max-iter to 1). |
| `--temperature` | 0.7 | Sampling temperature. |
| `--max-tokens` | 32768 | Max completion tokens. |
| `--reasoning-effort` | `medium` | Reasoning effort: `low`, `medium`, or `high`. |
| `--context-window` | 131072 | Context window size. |
| `--output` | `eval_results.json` | Output results file path. |
| `-v`, `--verbose` | false | Show detailed per-turn logging. |
| `--explore-constant` | 1.414 (sqrt(2)) | UCB exploration constant C. |
| `--test-timeout` | 300 | Per-test timeout in seconds. |
| `--suite-timeout` | 0 (unlimited) | Total suite timeout in seconds. |
| `--no-fast-fail` | false | Disable early termination on all-zero initial runs. |
| `--parallel` | 1 (serial) | Parallel workers (0=auto, N=use N workers). |
### Optimizer Options
Accepted by **`turnstone-optimizer`** only.
| Flag | Default | Description |
|-------------------------|----------------------------|-------------|
| `--max-iter` | 5 | Maximum optimization iterations. |
| `--no-optimize` | false | Run a single measurement pass (sets max-iter to 1). |
| `--explore-constant` | 1.414 (sqrt(2)) | UCB exploration constant C. |
| `--suite-timeout` | 0 (unlimited) | Total suite timeout in seconds. |
| `--optimizer-model` | same as `--model` | Model for prompt optimization. |
| `--optimizer-base-url` | same as `--base-url` | Base URL for optimizer model. |
| `--observer-model` | same as optimizer | Model for meta-optimization (observer). |
+10 -21
View File
@@ -13,25 +13,18 @@ The permission model has two layers:
1. **Scopes** (legacy) — `read`, `write`, `approve`. Checked by `AuthMiddleware`
on every request based on URL path classification.
2. **Permissions** (granular) — named permission strings checked per-endpoint by
2. **Permissions** (granular) — 15 permission strings checked per-endpoint by
`require_permission()`.
**Built-in roles** (seeded by migration 008 and extended by later feature
migrations):
**Built-in roles** (seeded by migration 008):
| Role | Permissions |
|------|-------------|
| admin | Admin-default baseline: ordinary admin, lifecycle, tool-approval, coordinator, project, and persona capabilities. Explicit opt-in capabilities such as `model.skills.write` remain ungranted. |
| operator | read, write, workstreams.create, workstreams.close, conversation.modify |
| admin | read, write, approve, admin.users, admin.roles, admin.orgs, admin.policies, admin.skills, admin.audit, admin.usage, admin.schedules, admin.watches, tools.approve, workstreams.create, workstreams.close |
| operator | read, write, workstreams.create, workstreams.close |
| viewer | read |
Custom roles can be created with any subset of the valid permissions. Built-in
role permission overrides can grant or revoke individual capabilities, so the
admin console is authoritative for the effective set on a deployment.
The `persona.create` / `persona.read` / `persona.write` family gates
persona administration; migration `063` seeds all three onto
`builtin-admin`, and any role can be granted them through the standard
role and permission-override editors.
Custom roles can be created with any subset of the 15 valid permissions.
**Auth flow:**
1. User logs in (password or API token) → `_load_user_permissions()` aggregates
@@ -70,7 +63,7 @@ etc.) since workstream templates were merged into the skills system in v0.8.0.
workstreams, concatenated in alphabetical order by name. Use name prefixes
(e.g. `01-safety`, `02-style`) to control ordering.
- **Explicit selection**: `--skill <name>` CLI flag, `skill` field on
`POST /v1/api/workstreams/new`, console launcher dropdown, scheduled task
`POST /v1/api/workstreams/new`, console creation modal dropdown, scheduled task
config, and channel adapter config. An explicit skill *replaces* defaults.
- **Variables**: Three built-in placeholders resolved at load time:
`{{model}}` (active model name), `{{ws_id}}` (workstream ID),
@@ -134,12 +127,9 @@ Per-LLM-request token and tool call metrics:
LLM response with prompt/completion tokens, cache tokens, tool call count,
model, ws_id
- **Prompt caching**: Anthropic automatic caching (`cache_control: ephemeral`)
and OpenAI caching are enabled by default. Pre-5.6 GPT-5 models request
`prompt_cache_retention: 24h`; GPT-5.6 uses
`prompt_cache_options: {"ttl": "30m"}`. GPT-5.6 cache writes use the
provider's 1.25× input-token rate. `cache_creation_tokens` and
`cache_read_tokens` are tracked per request in `usage_events` and surfaced
in the Usage admin tab
and OpenAI extended retention (`prompt_cache_retention: 24h` for GPT-5.x)
are enabled by default. `cache_creation_tokens` and `cache_read_tokens` are
tracked per request in `usage_events` and surfaced in the Usage admin tab
- **Querying**: `GET /v1/api/admin/usage` with `group_by` (day/hour/model/user)
and time range filtering — includes cache token aggregates
- **Prometheus**: `turnstone_tokens_total{type="cache_creation|cache_read"}`
@@ -187,7 +177,6 @@ All under `/v1/api/admin/` (requires `approve` scope + granular permission).
| Orgs | 3 (list, get, update) | `admin.orgs` |
| Tool Policies | 4 (CRUD) | `admin.policies` |
| Skills | 4 (CRUD) | `admin.skills` |
| Personas | 4 (list, create, get, edit/archive) | `persona.read` / `persona.create` / `persona.write` |
| Schedules | 6 (CRUD + runs) | `admin.schedules` |
| Watches | 3 (list, create, cancel) | `admin.watches` |
| Usage | 1 (aggregated query) | `admin.usage` |
@@ -233,7 +222,7 @@ Both Python and TypeScript console SDKs expose governance methods:
- **Privilege escalation prevented**: `admin_assign_role` blocks self-assignment
and requires caller to hold a superset of the target role's permissions
- **Permission validation**: Role create/update validates permissions against
the permission allowlist (`_VALID_PERMISSIONS`)
a 15-item allowlist (`_VALID_PERMISSIONS`)
- **Self-deletion blocked**: `admin_delete_user` rejects attempts to delete
your own account (matching the self-assignment guard on role endpoints)
- **Field allowlists**: Storage `update_*` methods filter fields against
+70 -116
View File
@@ -17,9 +17,7 @@ evaluation:
read-only tool access. Runs on a daemon thread and delivers its verdict
progressively.
The verdict is advisory by default. The opt-in Smart Approvals mode can use a
completed, high-confidence LLM `approve` verdict to make the decision
automatically under the fail-closed rules below.
The verdict is purely advisory -- the user always makes the final decision.
The heuristic verdict is attached to the `approve_request` SSE event immediately.
The LLM verdict arrives later via an `intent_verdict` SSE event, allowing the
@@ -30,77 +28,55 @@ persisted to the `intent_verdicts` table for audit and future calibration.
## Configuration
### Server and console
### config.toml
Server and console workstreams read database-backed `judge.*` settings from
the settings registry. Edit them at **Admin → Judge** or through the admin
settings API; changes take effect for the next judge batch without a restart.
The principal settings are:
```text
judge.enabled = true
judge.model = "" # empty = same alias as the session
judge.smart_approvals = false # opt-in automatic approval
judge.confidence_threshold = 0.95 # Smart Approvals confidence bar
judge.max_context_ratio = 0.5 # fraction of judge context used for history
judge.timeout = 120.0 # per judge turn and Smart Approvals wait
judge.parallel_evaluations = 1 # concurrent calls within one batch, 1-16
judge.read_only_tools = true # permit read_file/list_directory evidence
judge.cancel_on_approval = false # stop unfinished calls when the gate resolves
```toml
[judge]
enabled = true
model = "" # empty = same as session model
provider = "" # empty = same as session provider
base_url = ""
api_key = ""
smart_approvals = false # auto-approve high-confidence "approve" LLM verdicts (opt-in)
confidence_threshold = 0.95 # Smart Approvals auto-approve bar (LLM recommendation=approve)
max_context_ratio = 0.5 # max % of judge context window for history
timeout = 60.0 # seconds (generous for local models)
read_only_tools = true # judge can use read_file/list_directory
cancel_on_approval = false # stop judging remaining tool calls once user decides
```
`parallel_evaluations = 1` preserves serial evaluation. Raising it reduces the
latency of wide tool-call batches. The selected judge model alias's
`max_concurrency` remains the process-wide generation ceiling, so it can reduce
the actual overlap across judge batches and other roles using that alias.
### Smart Approvals
With `smart_approvals = true` (off by default), a pending batch is approved
automatically — no operator prompt — only when **every** call has a completed
LLM verdict recommending `approve` at or above `confidence_threshold`. The
decision is batch-atomic: one uncertain sibling sends the entire parallel batch
to a human rather than executing the safe-looking subset piecemeal.
With `smart_approvals = true` (off by default) a tool call is approved
automatically — no operator prompt — when the intent judge's **LLM** verdict
recommends `approve` with confidence at or above `confidence_threshold`. Every
other outcome still reaches a human: `review` / `deny` recommendations,
confidence below the threshold, judge errors or timeouts (`llm_fallback`), and
any call the deterministic heuristic rules explicitly flagged `deny` or
`critical`. That heuristic floor blocks only those explicit danger verdicts — it
is **not** a general "never lower the heuristic" rule: the heuristic's default
for an unmatched tool is `review`, and letting a confident LLM `approve` upgrade
a `review` is exactly what Smart Approvals is for. Only `deny` / `critical`
findings are off-limits to auto-approval. Requires the judge to be enabled;
auto-approved calls are tagged `smart_approval` in the dashboard and audit trail.
Smart Approvals applies to the web and coordinator surfaces, not the interactive
CLI.
Every other outcome reaches a human: `review` / `deny` recommendations,
confidence below the threshold, judge errors or timeouts (`llm_fallback`), a
missing/duplicate call ID, an unjudged sibling, and any call the deterministic
heuristic rules explicitly flagged `deny` or `critical`. That heuristic floor
blocks only explicit danger verdicts — it is **not** a general "never lower the
heuristic" rule. The heuristic's default for an unmatched tool is `review`, and
letting a confident LLM upgrade that default is the feature's purpose.
The Smart Approvals enabled flag, threshold, and bounded verdict wait are
captured as one immutable snapshot when each gate batch starts. A settings
reload takes effect on the next batch, while concurrent main-loop and
task-agent gates cannot mix fields from different reload generations. Stop
wakes a batch still waiting for verdicts and is linearized against the final
auto-approval commit: if Stop wins, no `smart_approval` decision or audit row
is recorded for tools that did not cross the gate.
The verdict wait is capped by the snapshot's `judge.timeout`; the judge may
continue evaluating advisory verdicts after that gate falls back to a human.
Requires the judge to be enabled. Auto-approved calls are tagged
`smart_approval` in the dashboard and audit trail. Smart Approvals applies to
the web and coordinator surfaces, not the interactive CLI.
The judge is enabled by default. Disable `judge.enabled` in the admin Judge
settings, or use `--no-judge` in the interactive CLI.
All fields are optional. The judge is enabled by default; use `enabled = false`
(or `--no-judge` on the command line) to disable it.
### CLI flags
```
--judge / --no-judge Enable/disable (default: enabled)
--judge-model ALIAS Registered model alias for judge
--judge-timeout SECONDS LLM judge timeout (default: 120)
--judge-parallel-evaluations N Concurrent evaluations per batch, 1-16 (default: 1)
--judge-model MODEL Model for judge
--judge-provider PROVIDER Provider for judge
--judge-timeout SECONDS LLM judge timeout (default: 60)
--judge-confidence FLOAT Confidence threshold, 0-1 (default: 0.95)
```
The same five values can be placed in the CLI's `config.toml` `[judge]`
section. Smart Approvals is configured through the server/console admin Judge
settings, not a CLI flag—the interactive CLI prompts for approval directly.
(Smart Approvals is configured via `[judge] smart_approvals` / the admin Judge
settings, not a CLI flag — the interactive CLI prompts for approval directly.)
CLI flags override `config.toml` values.
@@ -111,19 +87,20 @@ CLI flags override `config.toml` values.
- **Default (self-consistency)**: When `model` is empty, the session model
evaluates its own tool calls. Research shows self-consistency achieves
comparable accuracy to multi-agent debate at a fraction of the cost.
- **Cross-model**: Register the desired model in the Models tab, then set
`judge.model` to that alias (or pass `--judge-model ALIAS` to the CLI).
- **Cross-provider**: A model alias carries its provider, endpoint, and
credential configuration together, so a judge alias may use a different
provider from the session without separate judge connection settings.
- **Google models**: The judge supports `google` aliases, including read-only
evidence tools. Provider-native reasoning state such as Gemini
`thought_signature` stays attached to the pinned model lane across evidence
turns.
- **Cross-model**: Use a different model for the judge (e.g. local model for
the session, commercial model for the judge). Set `model` and `provider`
in the `[judge]` config section, or use `--judge-model` / `--judge-provider`
CLI flags.
- **Cross-provider**: When both `model` and `provider` are set, the judge
creates its own LLM client. You can optionally specify `base_url` and
`api_key` for non-default endpoints.
- **Google models**: The judge supports `google` as a provider. Note that
read-only tools are disabled for Google models (the Gemini API requires
`thought_signature` in tool call round-trips which the judge's normalized
format does not preserve).
The judge creates one fresh HTTP client per active batch worker and closes each
when that worker finishes, avoiding cross-thread client sharing and stale
connections across runs.
The judge creates a fresh HTTP client for each evaluation run and closes it
when done, avoiding stale connection issues across runs.
If the LLM judge fails or returns no verdict, a fallback verdict with tier
`llm_fallback` is delivered via the callback, ensuring the UI always receives
@@ -216,10 +193,9 @@ Security hardening blocks access to sensitive paths:
### Timeout
The `timeout` setting (default 120 seconds) applies **per turn**, not as a total
budget across turns — each of the up to 5 turns gets a fresh budget, so a slow
earlier turn doesn't starve later ones. If a turn's budget expires, the judge
attempts to parse whatever partial response is available.
The `timeout` setting (default 60 seconds) is a total budget across all judge
turns. Time is decremented after each LLM call. If the budget expires mid-turn,
the judge attempts to parse whatever partial response is available.
---
@@ -255,48 +231,25 @@ calls for approval, it calls `_evaluate_intent()` which:
4. Attaches each heuristic verdict to its item as `_heuristic_verdict`
5. The daemon thread runs the LLM judge and delivers results via `ui.on_intent_verdict()`
The daemon coordinates up to `parallel_evaluations` independent workers for
one batch. Completed verdicts stream to the UI as workers finish, and every call
still receives exactly one LLM or `llm_fallback` verdict. The default of 1 keeps
the historical serial behavior; a higher value collapses a wide batch toward
`ceil(batch size / workers)` judge-call intervals. A smaller positive model
alias capacity also bounds the worker count, avoiding surplus threads queued at
the same admission gate.
With `cancel_on_approval = false` (the default) the daemon runs every item to
completion: verdicts that land after the operator decided still stream to the
UI and persist, stamped with the decision. A newer main-loop batch, session
close, or explicit Stop retires the old generation; unfinished items degrade
to `llm_fallback` verdicts. A judge/model binding or parallelism edit prevents
reuse on the next batch, while already-started calls stay pinned to the binding
and worker count they began with. With `cancel_on_approval = true`, an ordinary
gate decision additionally aborts unfinished work, trading verdict completeness
for inference savings—recommended when the judge shares a single local
inference backend with the session model. Explicit Stop always cancels every
live judge generation, regardless of this preference.
The daemon evaluates items sequentially, so a large parallel batch can outlive
its approval gate. With `cancel_on_approval = false` (the default) the daemon
runs every item to completion: verdicts that land after the operator decided
still stream to the UI and persist, stamped with the decision. The daemon is
aborted only when the next tool batch supersedes it or the session closes —
then each unfinished item degrades to an `llm_fallback` verdict. With
`cancel_on_approval = true` the abort additionally fires the moment the gate
resolves, trading verdict completeness for inference savings — recommended
when the judge shares a single local inference backend with the session model,
where a large batch's remaining judge calls would otherwise compete with the
next turn's completion.
Verdicts that arrive after a *newer batch* has replaced the judge generation
are withheld from the live surfaces (a reused call_id must never ride a stale
`approve` into Smart Approvals) but still persist with
`user_decision = "superseded"` so the audit trail records the judge's answer.
Sub-agent (task agent) tool calls are judge-gated too. Each runs the same
intent pipeline as its own `agent_gate` generation, grounded in that sub-agent's
own trajectory -- its task prompt is the delegation contract the operator
approved, so "does this call serve the task" is the right local question.
Agent-gate generations never occupy the main loop's supersede slot (parallel
siblings would otherwise make each other's verdicts look stale); per-cycle
generation checks enforce staleness instead, and `judge.cancel_on_approval`
fires per gate exactly like the main loop.
Several parallel task agents can therefore leave several approval cycles live
on one workstream. Each cycle owns its event, result, verdict set, and
`cycle_id`; a decision targets exactly one cycle by `cycle_id` or member
`call_id` (selector-less legacy clients resolve the oldest). Workstream Stop or
close performs a workstream-wide denial sweep over all cycles belonging to the
cancelled operation. A force-cancel successor's newly registered cycle carries
a fresh operation witness and is not accidentally denied by the predecessor's
late sweep.
Sub-agents (plan agent, task agent) are exempt from intent validation -- they
always get full tool visibility without judge evaluation.
---
@@ -455,12 +408,13 @@ Redaction types: `api_key`, `private_key`, `password`, `secret`.
### Configuration
```text
judge.output_guard = true # enable output evaluation (default)
judge.redact_secrets = true # auto-redact detected credentials (default)
```toml
[judge]
output_guard = true # enable output evaluation (default)
redact_secrets = true # auto-redact detected credentials (default)
```
Configure both at runtime through the admin Judge settings.
Configurable at runtime via the admin Settings tab.
### Merge semantics (heuristic + LLM judge)
+4 -65
View File
@@ -17,9 +17,8 @@ The MCP server admin form exposes three authorization modes ("Multitenant Author
| `none` | No headers attached. Open MCP server (or one gated by network policy only). | Internal MCP servers on a trusted network. |
| `static` | One static bearer token, configured per server, sent on every request from every user. | Service-to-service MCP servers where per-user attribution doesn't matter, or single-tenant deployments. |
| `oauth_user` *(recommended for user-data servers)* | Each user authorizes separately via OAuth 2.1 + PKCE; Turnstone stores per-user tokens encrypted at rest. | MCP servers that expose user-specific data or that want per-user audit attribution. |
| `oauth_obo` *(sign-in passthrough)* | Each user's Turnstone **org sign-in** (OIDC) mints a per-server access token on demand — no separate per-server consent. One captured credential per user covers every `oauth_obo` server. | Enterprise deployments where the identity provider governs access (Entra, Keycloak) and you want zero per-user connect clicks. See the dedicated section below. |
Switching `auth_type` away from `oauth_user` / `oauth_obo` **deletes** that server's per-user rows (consents / minted cache) — see the transition table below. Switching back later starts clean: users re-consent (or re-mint) on next use. The admin **bulk-revoke** / **flush cache** affordance clears rows without an auth-type change.
Switching `auth_type` away from `oauth_user` orphans existing per-user tokens. Use the admin **bulk-revoke** affordance on the server row (Phase 9) to clear them, or let them expire naturally — they're inert without the matching `auth_type` value.
---
@@ -66,59 +65,6 @@ Keep this in `config.toml` rather than environment variables. An in-process LLM
---
## `auth_type=oauth_obo` — single-credential sign-in passthrough
Where `oauth_user` makes each user complete a **separate** browser consent per MCP server, `oauth_obo` reuses the user's Turnstone **org sign-in** (OIDC). Turnstone captures one refresh credential per user at login and, on each tool call, mints a short-lived access token scoped to that server's audience. There is no per-server connect step, and one credential covers every `oauth_obo` server. This is the right shape when your identity provider already governs who may reach each backend (an Entra tenant with Entra-protected MCP servers; a Keycloak realm with token exchange).
Access is governed **downstream** by the IdP: a user can only mint a token for a server their delegated permissions allow. Removing that grant at the IdP cuts the user off regardless of their Turnstone state.
### Deployment configuration (`[oidc]` in `config.toml`)
`oauth_obo` requires OIDC SSO to be configured (it is the credential source), plus:
```toml
[oidc]
# ... your existing issuer / client_id / client_secret ...
capture_user_credential = true # persist the IdP refresh token at login
obo_grant_profile = "entra" # "entra" | "rfc8693" — how tokens are minted
```
- **`capture_user_credential`** (default `false`): when enabled, Turnstone appends `offline_access` to the login scopes and stores the returned refresh token, encrypted with the same `[security] mcp_token_encryption_key` as `oauth_user` tokens. **The encryption key is required** — Turnstone refuses to start with an `oauth_obo` row (or capture enabled) and no key.
- **`obo_grant_profile`** picks the mint mechanism (the IdP determines which one is valid; this is deployment-wide, not per-server):
- **`entra`** — redeems the user's refresh token directly for a token scoped to `<audience>/.default`. `oauth_scopes` on the server row is **not used** (the admin form rejects it under this profile).
- **`rfc8693`** — a refresh grant for a subject token, then an RFC 8693 token exchange for the server audience. Per-server `oauth_scopes` **are** sent on the exchange (some IdPs require the audience scope explicitly).
### Adding an `oauth_obo` server
In the admin MCP form, choose **Sign-in passthrough** and set **Audience** (required — the downstream resource the token is minted for, e.g. `api://<app-id>` on Entra or the client id on Keycloak). The client-id / secret / registration fields do not apply and are hidden.
`oauth_obo` servers are accepted only when **OIDC sign-in is configured and enabled** and `[oidc] obo_grant_profile` is a valid profile — the write is rejected otherwise, since a row that can never mint would surface to users as a permanent "please retry" that never heals.
### Identity-provider setup
**Entra (`obo_grant_profile = "entra"`):**
1. Turnstone's app registration must hold **delegated permissions** to each MCP server's exposed API, with **admin consent granted** (or the MCP app listed in Turnstone's `preAuthorizedApplications`).
2. Set the server row's Audience to the MCP app's Application ID URI (`api://<guid>`).
3. **Gotcha (verified):** admin-consent issued *immediately* after creating the app/service principal can silently skip a not-yet-propagated resource — the only symptom is `AADSTS65001` at mint time. Verify the delegated grant landed (`az ad app permission list-grants` / the portal's *API permissions* blade shows *Granted*), or grant it explicitly per resource. A missing grant surfaces in Turnstone as a re-login prompt on the affected server (same rail as a revoked credential), and the `mcp_server.oauth.obo_mint_rejected` log line carries the raw `AADSTS…` text.
**Keycloak / RFC 8693 (`obo_grant_profile = "rfc8693"`):**
1. Enable **standard token exchange** on Turnstone's client.
2. Grant the audience: add an audience client scope for each MCP client and attach it to Turnstone's client (optional scopes must be requested — set the server row's Scopes to that scope, or the exchange returns *"Requested audience not available"*).
3. Set the server row's Audience to the downstream client id.
### Revocation & custody
The captured credential is a single per-user secret that can mint for every `oauth_obo` server, so treat it like any long-lived credential:
- **Cut off one user:** unlink their OIDC identity in the admin console (**Users → OIDC identities → delete**). This revokes the captured credential **and** purges their minted cache rows, so future mints fail and cached tokens are dropped. (Warmed in-memory sessions on server nodes self-expire at the access-token TTL; there is no cross-node per-user session-kill.) Removing the user's access at the IdP is the authoritative cut-off.
- The same unlink also purges that user's synthetic `__model_obo__:` gateway-token rows and requests eviction from every registered host's in-process mint memo. Shared `entra_app` model tokens live under the `__app__` pseudo-user and are intentionally not user-deprovisioned; revoking the app credential prevents new mints, while a cached app bearer lasts until `expires_at`.
- **Flush a server's minted tokens** (e.g. after narrowing its audience): the server row's **flush cache** action drops all users' cached tokens for that server. This is **not** a revocation — users re-mint on next use from their still-valid sign-in. It is surfaced honestly (audit `mcp_server.oauth.obo_cache_flushed`, response `effect: cache_flush_remints`) so it is never mistaken for cutting access.
- Per-server revocation in the `oauth_user` sense does not exist for `oauth_obo` — the credential is issuer-scoped and IdP-governed. Revoke at the IdP.
> **Interim for Entra without OBO:** if you don't want host-side minting, admin consent + `preAuthorizedApplications` on each MCP app registration removes the second consent prompt for the plain `oauth_user` flow too (a tenant-config change, no Turnstone code). Tracked in issue #682. It does not remove the per-server connect clicks or per-(user, server) token custody — that is what `oauth_obo` is for.
---
## Lifecycle
1. **First tool call** for a user against an `oauth_user` MCP server: pool dispatch finds no stored token, returns `mcp_consent_required` to the agent. Dashboard renders an inline "Connect" action card.
@@ -129,7 +75,7 @@ The captured credential is a single per-user secret that can mint for every `oau
4. **Step-up scope**: when a tool call hits `403` with `WWW-Authenticate: error="insufficient_scope"`, Turnstone emits `mcp_insufficient_scope` with the parsed scope set; the dashboard offers a "Connect with additional scopes" affordance that opens `/v1/api/mcp/oauth/start?server=<name>&scopes=<extra>` so the union of original + new scopes flows into the AS authorize request.
5. **User revoke** (settings modal): `DELETE /v1/api/mcp/oauth/connections/{server_name}` runs the authoritative local delete + best-effort RFC 7009 upstream revoke (fire-and-forget, capped at 256 concurrent in-flight tasks). `oauth_obo` servers and synthetic model-auth rows are excluded: their rows are mint caches, not consents — deleting one only forces a re-mint — so the connections list hides them and the endpoint refuses them with `409` (revocation for sign-in passthrough happens at the identity layer: unlink the identity or revoke at the IdP).
5. **User revoke** (settings modal): `DELETE /v1/api/mcp/oauth/connections/{server_name}` runs the authoritative local delete + best-effort RFC 7009 upstream revoke (fire-and-forget, capped at 256 concurrent in-flight tasks).
6. **Admin bulk-revoke** (Phase 9): `POST /v1/api/admin/mcp-servers/{name}/bulk-revoke` drops every user's token for the server. Upstream RFC 7009 revoke is intentionally **not** attempted in bulk (avoids N upstream HTTP calls per admin click); tokens at the AS expire naturally. Use the per-user revoke endpoint if you need guaranteed upstream invalidation.
@@ -151,13 +97,10 @@ Additional indicators (circuit-breaker state, encryption-key mismatch) are expos
| From | To | What happens |
|---|---|---|
| `none` / `static``oauth_user` | — | New code path activates for this server. Existing static headers (if any) are no longer sent. Users must authorize on first use. |
| `oauth_user``none` / `static` | — | Existing `mcp_user_tokens` rows are **deleted**: the tokens are bound to the auth model + URL active at consent time, and rows left behind could silently rebind if a row with the old name/URL reappears. Switching back to `oauth_user` later starts clean — users re-consent on next use. This is **not reversible**; the AS-side grants are untouched (revoke upstream via the AS if needed). |
| `oauth_user``none` / `static` | — | Existing `mcp_user_tokens` rows are **orphaned** — inert without a matching `auth_type`. Use admin bulk-revoke to drop them, or let them expire. Switching back to `oauth_user` later re-activates the orphaned rows if they haven't been deleted. |
| OAuth `client_id` or `client_secret` rotated | — | Existing tokens may stop refreshing if the AS treats them as bound to the previous client. Bulk-revoke after rotation. |
| `oauth_user``oauth_obo` | — | The per-user rows are **deleted** on the flip (they mean different things: per-server AS refresh tokens vs. minted cache). `oauth_audience` and `oauth_scopes` mean different things in each model (a resource indicator vs. an IdP app identifier; AS-consent scopes vs. an rfc8693 exchange scope), so on a flip they **never carry** — each is taken from the request for the target model or set NULL. The admin console clears these fields when you change the auth type, so re-enter the correct values for the new mode; via the API, supply them explicitly (a flip into `oauth_obo` with no `oauth_audience` is rejected, and a non-empty `oauth_scopes` under the `entra` profile is rejected since that leg pins `<audience>/.default`). |
| `oauth_obo``none` / `static` | — | Minted cache rows are deleted. |
| `oauth_obo` **audience**, **URL**, or **`oauth_scopes`** changed | — | Minted cache rows are **deleted** (tokens are bound to the audience/URL/scopes at mint time), forcing a fresh mint — so an audience or scope narrowing takes effect immediately, not at token expiry. |
Every transition that changes what a stored row *means* deletes the rows outright — a stale consent or minted token must never be served under new semantics. There is no orphan-and-reactivate path.
The orphan-by-default behavior is chosen so switching back to `oauth_user` is non-destructive. Bulk-revoke is the explicit cleanup path.
---
@@ -170,9 +113,5 @@ Every transition that changes what a stored row *means* deletes the rows outrigh
| `mcp_oauth_url_insecure` | MCP server URL is `http://` (not `https://`) on a non-loopback host | Use `https://`. Per-user bearers must not transit cleartext. |
| Tools fail in scheduled / Discord / Slack runs | OAuth-MCP requires browser-based consent | Users must pre-consent via the web UI. Phase 9 dashboard badge surfaces deferred consents from these runs on next login. |
| Circuit breaker open repeatedly | Transport-level errors on the MCP server (DNS, TLS, 5xx) | Check the per-server error pill; auth errors do not trip the breaker. |
| **`oauth_obo`**: every tool call fails, log shows `obo_misconfigured` | Server row has no Audience, or `obo_grant_profile` is unset/unknown | Set the Audience on the server row; set `[oidc] obo_grant_profile` to `entra` or `rfc8693`. |
| **`oauth_obo`**: `obo_mint_rejected` with `AADSTS65001` | Turnstone's app lacks the (admin-consented) delegated grant to this MCP app — often admin consent that didn't propagate | Grant + admin-consent the delegated permission for this resource; verify it shows *Granted*. See the Entra gotcha above. |
| **`oauth_obo`**: "Sign in to Turnstone again" on one server | Captured credential missing/rejected, or a Conditional Access challenge | User re-logs into Turnstone (re-captures the credential). If it persists, check the IdP grant / CA policy. |
| **`oauth_obo`**: tools don't appear at all for a user | User has not signed in since `capture_user_credential` was enabled (no credential captured) | User logs out and back in via OIDC so the refresh credential is captured. |
See also: `docs/operations/mcp-oauth-headless.md` for the cron / channel-driven run caveat.
+37 -119
View File
@@ -20,79 +20,36 @@ Each memory has three dimensions:
| Type | Purpose |
|-------------|------------------------------------------------------------|
| `user` | User preferences, conventions, working style |
| `general` | General knowledge, architecture, patterns |
| `project` | Project-specific knowledge, architecture, patterns |
| `feedback` | Corrections, lessons learned, things to avoid |
| `reference` | Reference material, documentation, specifications |
### Memory scopes
| Scope | Visibility |
|---------------|-----------------------------------------------------------------|
| `global` | Visible to all workstreams and users |
| `workstream` | Visible only within the originating workstream |
| `user` | Follows the authenticated user across workstreams |
| `coordinator` | Coordinator sessions only; follows the acting user |
| `project` | Shared by workstreams attached to one active project |
| Scope | Visibility |
|--------------|-----------------------------------------------------------|
| `global` | Visible to all workstreams and users |
| `workstream` | Visible only within the originating workstream |
| `user` | Follows the authenticated user across workstreams |
A memory's identity is the tuple `(name, scope, scope_id)`. Saving a memory
with the same identity upserts -- updating content while preserving the ID.
### Inherited target and coordinator scope
Name-based operations use one inherited target when `scope` is omitted:
- An attached active project selects `project` for `save`, `get`, and
`delete`.
- Read-only project access permits `get`, but `save` and `delete` fail. They do
not fall back to a broader namespace.
- Without a project, interactive sessions select `global`; coordinator
sessions select `coordinator`.
A valid explicit scope selects exactly that scope. `search` and `list` are the
only actions that span every visible scope when `scope` is omitted.
Each coordinator's private `coordinator` namespace is keyed by the acting
user's `user_id`. It is durable -- every coordinator session that user runs
(including concurrent ones) shares one orchestration namespace, so procedures
and lessons survive close/reopen.
Isolation is bidirectional and enforced by session kind, not by secrecy of
the scope id:
- A coordinator session sees its acting user's `coordinator` scope and, when
attached, the shared `project` scope. It never sees
`global`/`workstream`/`user` memories.
- Interactive sessions -- including a coordinator's own children, which share
its `user_id` -- are rejected from the `coordinator` scope on every memory
action. Children cannot plant rows the parent coordinator would read.
- The REST memory API (`/v1/api/memories`) does not accept the `coordinator`
scope at all; the scope is written exclusively through a coordinator
session's own memory tool.
Coordinator sessions require an authenticated user identity -- an anonymous
coordinator cannot be constructed, so the scope id is always a real user.
### BM25 relevance injection
On every conversation turn, the system:
1. Resolves the acting principal and their live project access
2. Fetches up to `fetch_limit` memories across that visibility envelope
3. Extracts context from the last 3 user messages
4. Scores memories against that context using a BM25 index
5. Injects the top `relevance_k` memories into the system message as
1. Fetches up to `fetch_limit` memories visible in the current scope
2. Extracts context from the last 3 user messages
3. Scores memories against that context using a BM25 index
4. Injects the top `relevance_k` memories into the system message as
`<memories>` XML tags
6. Appends a hint telling the model how many memories are in scope
5. Appends a hint telling the model how many memories are in scope
This means the model always has its most relevant memories available without
explicit recall -- but can still use `memory(action='search')` for deeper
lookup.
The persona memory lever gates this pathway: a workstream whose persona
turns memory off receives no relevance injection at all -- the steps
above run only when memory is enabled for the session. See
[Personas](personas.md).
### Nudges
The metacognition layer can nudge the model to save memories at appropriate
@@ -120,23 +77,19 @@ All fields are optional. Defaults are shown above.
## Tool Usage
The `memory` tool supports five actions:
The `memory` tool supports four actions:
### save
Store or update a memory.
Every save is a complete write for the relevance summary: `description` must
be supplied and contain non-whitespace text on both creation and update.
Content-only updates are rejected.
```json
{
"action": "save",
"name": "project_architecture",
"content": "The project uses a hexagonal architecture with...",
"description": "Core architecture patterns",
"type": "general",
"type": "project",
"scope": "global"
}
```
@@ -145,26 +98,9 @@ Content-only updates are rejected.
|---------------|----------|-------------|------------------------------------------|
| `name` | yes | -- | Snake_case identifier (max 256 chars) |
| `content` | yes | -- | Memory content (max `max_content` chars) |
| `description` | yes | -- | Non-empty relevance summary, required on create and update |
| `type` | no | `"general"` | One of: user, general, feedback, reference |
| `scope` | no | inherited | Kind-valid scope; see inherited target above |
### get
Retrieve the full content of one memory by name.
```json
{
"action": "get",
"name": "project_architecture",
"scope": "project"
}
```
| Parameter | Required | Default | Description |
|-----------|----------|-----------|----------------------------|
| `name` | yes | -- | Memory name to retrieve |
| `scope` | no | inherited | Exact scope to query |
| `description` | no | `""` | Short description for relevance matching |
| `type` | no | `"project"` | One of: user, project, feedback, reference |
| `scope` | no | `"global"` | One of: global, workstream, user |
### search
@@ -174,7 +110,7 @@ Find memories by query (BM25 full-text search).
{
"action": "search",
"query": "authentication patterns",
"type": "general",
"type": "project",
"limit": 10
}
```
@@ -201,7 +137,7 @@ Remove a memory by name.
| Parameter | Required | Default | Description |
|------------|----------|------------|--------------------------|
| `name` | yes | -- | Memory name to delete |
| `scope` | no | inherited | Exact scope to delete |
| `scope` | no | `"global"` | Scope of the memory |
### list
@@ -231,12 +167,6 @@ Four endpoints on the server for programmatic memory access.
List memories with optional filters.
Without `scope`, the response is restricted to `global` plus the authenticated
caller's `user` namespace. The public API accepts only `global`, `user`, and
`workstream`; internal `project` and `coordinator` namespaces remain available
through the session tool and admin API. Explicit `workstream` access requires
its persisted owner (or a service token).
**Query parameters:**
| Parameter | Type | Required | Default | Description |
@@ -244,10 +174,10 @@ its persisted owner (or a service token).
| `type` | string | no | `""` | Filter by memory type |
| `scope` | string | no | `""` | Filter by scope |
| `scope_id` | string | no | `""` | Filter by scope ID |
| `limit` | int | no | `100` | Max results (1-200) |
| `limit` | int | no | `100` | Max results (capped at 200) |
When `scope=user`, the authenticated user's ID is used automatically and a
different supplied ID is rejected. `scope=workstream` requires `scope_id`.
When `scope=user` and `scope_id` is omitted, the authenticated user's ID is
used automatically.
**Response:** `200`
@@ -258,7 +188,7 @@ different supplied ID is rejected. `scope=workstream` requires `scope_id`.
"memory_id": "a1b2c3d4-e5f6-...",
"name": "project_architecture",
"description": "Core architecture patterns",
"type": "general",
"type": "project",
"scope": "global",
"scope_id": "",
"content": "The project uses a hexagonal architecture...",
@@ -276,9 +206,6 @@ different supplied ID is rejected. `scope=workstream` requires `scope_id`.
Save or upsert a structured memory.
`description` is mandatory for both creates and updates and must contain
non-whitespace text. The API rejects content-only updates.
**Request body:**
```json
@@ -286,7 +213,7 @@ non-whitespace text. The API rejects content-only updates.
"name": "deployment_process",
"content": "Deploy via GitHub Actions. Staging auto-deploys on push to main.",
"description": "CI/CD deployment workflow",
"type": "general",
"type": "project",
"scope": "global",
"scope_id": ""
}
@@ -296,8 +223,8 @@ non-whitespace text. The API rejects content-only updates.
|--------------|--------|----------|-------------|--------------------------------------|
| `name` | string | yes | -- | Memory name (max 256 chars) |
| `content` | string | yes | -- | Memory content (max 65536 chars) |
| `description`| string | yes | -- | Non-empty relevance summary, required on create and update |
| `type` | string | no | unset | user, general, feedback, or reference |
| `description`| string | no | `""` | Short description for search ranking |
| `type` | string | no | `"project"` | One of: user, project, feedback, reference |
| `scope` | string | no | `"global"` | One of: global, workstream, user |
| `scope_id` | string | no | `""` | Scope qualifier (auto-resolved for user scope) |
@@ -308,7 +235,7 @@ non-whitespace text. The API rejects content-only updates.
"memory_id": "a1b2c3d4-e5f6-...",
"name": "deployment_process",
"description": "CI/CD deployment workflow",
"type": "general",
"type": "project",
"scope": "global",
"scope_id": "",
"content": "Deploy via GitHub Actions...",
@@ -324,10 +251,7 @@ same `(name, scope, scope_id)` already existed.
| Status | Condition |
|--------|------------------------------------|
| 400 | Invalid input, scope, scope ID, or limit |
| 403 | Cross-user or non-owner workstream access |
| 404 | Explicit workstream does not exist |
| 500 | Storage mutation failed |
| 400 | Missing name, empty content, invalid type/scope, content too long |
---
@@ -336,15 +260,12 @@ same `(name, scope, scope_id)` already existed.
Search memories by query. Uses POST for the request body but is non-mutating
(requires only `read` scope).
An omitted scope searches the same caller-bound `global` + `user` envelope as
the list endpoint. It never means every row in the table.
**Request body:**
```json
{
"query": "authentication",
"type": "general",
"type": "project",
"scope": "",
"scope_id": "",
"limit": 20
@@ -357,7 +278,7 @@ the list endpoint. It never means every row in the table.
| `type` | string | no | `""` | Filter by type |
| `scope` | string | no | `""` | Filter by scope |
| `scope_id` | string | no | `""` | Filter by scope ID |
| `limit` | int | no | `20` | Max results (1-50) |
| `limit` | int | no | `20` | Max results (capped at 50) |
**Response:** `200`
@@ -368,7 +289,7 @@ the list endpoint. It never means every row in the table.
"memory_id": "a1b2c3d4-e5f6-...",
"name": "auth_patterns",
"description": "Authentication architecture",
"type": "general",
"type": "project",
"scope": "global",
"scope_id": "",
"content": "JWT tokens with HS256...",
@@ -386,9 +307,6 @@ the list endpoint. It never means every row in the table.
Delete a memory by name and scope.
Deletes are atomic: the row used for the success result and audit event is the
row actually removed. A storage failure returns `500`, not a false `404`.
**Path parameters:**
| Parameter | Type | Description |
@@ -443,7 +361,7 @@ List memories across all scopes (no automatic scope resolution).
"memory_id": "a1b2c3d4-e5f6-...",
"name": "project_architecture",
"description": "Core architecture patterns",
"type": "general",
"type": "project",
"scope": "global",
"scope_id": "",
"content": "The project uses...",
@@ -492,7 +410,7 @@ Get a single memory by ID.
"memory_id": "a1b2c3d4-e5f6-...",
"name": "project_architecture",
"description": "Core architecture patterns",
"type": "general",
"type": "project",
"scope": "global",
"scope_id": "",
"content": "The project uses...",
@@ -549,13 +467,13 @@ with TurnstoneServer("http://localhost:8080", token="tok_xxx") as client:
"api_conventions",
"All endpoints use /v1/ prefix. JSON responses.",
description="API design patterns",
mem_type="general",
mem_type="project",
scope="global",
)
print(mem.memory_id)
# Search memories
results = client.search_memories("authentication", mem_type="general", limit=10)
results = client.search_memories("authentication", mem_type="project", limit=10)
for m in results.memories:
print(f"{m['name']}: {m['description']}")
@@ -576,7 +494,7 @@ with TurnstoneConsole("http://localhost:9090", token="tok_xxx") as admin:
result = admin.list_memories(scope="global", limit=100)
# Search
result = admin.search_memories("architecture", mem_type="general")
result = admin.search_memories("architecture", mem_type="project")
# Get by ID
mem = admin.get_memory("a1b2c3d4-e5f6-...")
@@ -600,14 +518,14 @@ const mem = await client.saveMemory({
name: "api_conventions",
content: "All endpoints use /v1/ prefix. JSON responses.",
description: "API design patterns",
type: "general",
type: "project",
scope: "global",
});
// Search memories
const results = await client.searchMemories({
query: "authentication",
type: "general",
type: "project",
limit: 10,
});
+9 -108
View File
@@ -41,7 +41,6 @@ are set.
| `TURNSTONE_OIDC_PASSWORD_ENABLED` | No | `true` | Set to `false` to hide the password form and block all username/password logins (including admin). API tokens continue to work. |
| `TURNSTONE_OIDC_REDIRECT_BASE` | Yes | — | Externally-reachable origin for the OIDC redirect URI (e.g. `https://app.example.com`). Without this, OIDC will refuse to start. The previous Host-header fallback was unsafe under permissive reverse proxies. |
| `TURNSTONE_OIDC_TRUSTED_ENDPOINT_HOSTS` | No | — | Comma-separated list of additional hostnames whose endpoints the IdP discovery document is allowed to reference. See [Cross-host endpoints](#cross-host-endpoints). |
| `TURNSTONE_OIDC_ALLOW_PRIVATE_NETWORK` | No | `false` | Allow the issuer (and its discovered endpoints) to resolve to private/internal addresses — needed for a self-hosted IdP on an internal network. See [Self-hosted and internal IdPs](#self-hosted-and-internal-idps). |
All four required fields — issuer, client ID, client secret, and
`TURNSTONE_OIDC_REDIRECT_BASE` — must be set. If any are missing OIDC
@@ -77,17 +76,17 @@ IdP from redirecting the token-exchange POST (which carries
being aimed at internal services.
A few public IdPs legitimately split endpoints across hostnames. Google
and Microsoft Entra ID are the canonical examples:
is the canonical example:
| IdP | Issuer host | Cross-host endpoint(s) |
|-----|-------------|------------------------|
| Google | `accounts.google.com` | `oauth2.googleapis.com`, `www.googleapis.com`, `openidconnect.googleapis.com` |
| Microsoft Entra | `login.microsoftonline.com` | `graph.microsoft.com` (userinfo) |
| Field | Hostname |
|-------|----------|
| issuer | `accounts.google.com` |
| token_endpoint | `oauth2.googleapis.com` |
| jwks_uri | `www.googleapis.com` |
| userinfo_endpoint | `openidconnect.googleapis.com` |
Both sets are built in — operators using `https://accounts.google.com` or
`https://login.microsoftonline.com/<tenant>/v2.0` need no extra
configuration. (Entra's discovery document advertises `userinfo_endpoint`
on `graph.microsoft.com`, distinct from the issuer host.)
Google's set is built in — operators using `https://accounts.google.com`
need no extra configuration.
For other IdPs whose discovery document references a non-issuer host,
extend the allow-list explicitly:
@@ -100,102 +99,6 @@ The same scheme / no-userinfo / SSRF rules apply to allow-listed hosts —
this knob only relaxes the same-origin check, not the security gates.
Each entry is a hostname (no scheme, no path).
### Self-hosted and internal IdPs
By default Turnstone refuses an issuer whose hostname resolves to a
private or internal address:
```
OIDCError: endpoint URL resolves to non-public address (10.0.0.5): https://auth.example.site
```
This is SSRF hardening, not a licensing or product restriction: the OIDC
flow makes server-side HTTP requests (discovery, JWKS, token exchange),
and refusing non-public destinations keeps a mistyped or maliciously
steered issuer from aiming those fetches at internal services. For a
self-hosted IdP (Keycloak, Authentik, Dex, …) on a private network,
opt in explicitly in `config.toml`:
```toml
[oidc]
allow_private_network = true
```
or via `TURNSTONE_OIDC_ALLOW_PRIVATE_NETWORK=true` (the env var wins
when both are set).
The opt-in admits private-range (RFC 1918), unique-local, site-local,
CGNAT (100.64/10, where overlay VPNs commonly assign hosts), and
loopback addresses. Link-local, multicast, reserved ranges and known
cloud-metadata endpoints stay refused even with the opt-in — no
legitimate IdP lives there. An address is judged by what it actually
reaches, so an IPv6 transition address (NAT64, 6to4, Teredo) wrapping
an internal IPv4 is treated exactly as that IPv4 would be. The HTTPS
requirement and the same-origin endpoint checks are unaffected.
This knob only affects the login-flow IdP configured here. OAuth
endpoints advertised by remote MCP servers are untrusted input and are
always held to the strict public-address rule.
### Model gateway credentials
The same OIDC registration can authenticate model gateways. A model definition
with `auth_mode = "entra_obo"` (Entra grant profile) or `auth_mode =
"rfc8693_obo"` (RFC 8693 token-exchange profile) redeems the driving user's
captured credential for its exact `obo_audience`; `auth_mode = "entra_app"`
uses the registration's client ID and secret with Entra client credentials.
All three bind the result through the provider SDK's native credential option
rather than injecting an override header. The grant mode is never inferred:
missing user context or a failed OBO mint cannot switch a delegated definition
to client credentials.
Each dynamic mode pairs with the grant profile whose dialect it names:
`entra_obo` and `entra_app` require `obo_grant_profile = "entra"`;
`rfc8693_obo` requires `obo_grant_profile = "rfc8693"`. The pairing is
enforced when a write chooses a `(auth_mode, obo_audience)` pair — a same-pair
edit of a row saved before the pairing rule keeps working — and at runtime a
mismatched legacy row refuses to mint with `cause=grant_profile_mismatch` and
no IdP traffic. RFC 8693 client-credentials is not implemented.
The delegated modes need the MCP encryption key, a credential captured for the
driving user, and delegated/admin-consented permission to the audience.
`rfc8693_obo` additionally carries `obo_scopes`, the space-separated scope
list its exchange leg requests: exchange-capable IdPs that gate audiences
behind optional scopes refuse the exchange without it ("Requested audience not
available"), which is why the scope-less Entra-named mode could never mint on
that profile (issue #955). Scopes are stored shape-checked only — whether a
value satisfies the IdP stays the IdP's call at mint time. Turning
`capture_user_credential` off later stops *new* captures but does not
invalidate credentials already stored, so existing users keep minting.
`entra_app` requires a confidential-client secret. Configure the permitted
resource IDs in the runtime setting `model.auth_audience_allowlist` before
saving dynamic model definitions. De-listing an audience later blocks every
write that would arm or re-aim a definition at it, but does not stop aliases
already configured from minting — disabling the row (the `admin.models` disarm
lever) is what stops minting. See
[Settings](settings.md#model-backend-authentication) for permissions, failure
policy, and lane identity rules.
An unrecognised `obo_grant_profile` is warned about at startup and **rejected
at the write choke points**: configuring an `oauth_obo` MCP server or a dynamic
model alias returns a 400 that echoes the configured value, so the typo is the
diagnosis. At runtime an unknown profile never mints — the mint legs resolve by
exact name; the full cause detail is logged once per audience, and every
affected call still logs its per-turn fallback or refusal naming the alias,
the target audience, and the last recorded cause (`cause=` — for example
`unsupported_grant_profile` or `oidc_not_enabled`) — so a pre-existing row
degrades loudly, with the reason visible mid-incident even after the
once-per-process line has rotated out of retained logs, rather than silently
swapping per-user attribution for the shared static key.
The `[security]` token encryption key is deployment-wide, not per-host: rows are
encrypted with `MultiFernet` and carry no key id, so every host that reads them
needs the same keyring. That includes the console, which mints for
coordinator-hosted sessions. A node that needs the key and lacks it refuses to
start; the console starts but withholds its coordinator subsystem and shows
the key requirement as the remediation error instead of failing silently at
call time.
### config.toml alternative
```toml
@@ -208,8 +111,6 @@ provider_name = "Google"
role_claim = "groups"
password_enabled = true
redirect_base = "https://app.example.com"
# Self-hosted IdP on an internal network (see "Self-hosted and internal IdPs")
allow_private_network = false
[oidc.role_map]
admin = "builtin-admin"
-173
View File
@@ -1,173 +0,0 @@
# Personas
A **persona** is a named, reusable bundle attached to a workstream **at
creation** that controls how its system message is composed and what
capability envelope it runs with. Personas answer a recurring operational
complaint: the default composition primes every session for heavy tool use,
and there was no per-workstream dial to launch a "just write prose" or
"evidence-first research" session.
A persona is exactly four levers — no more:
| Lever | What it does |
|---|---|
| **Base prompt** | Replaces the BASE module of the composed system message. *Only* BASE: ENV, CONTEXT, TOOLS, and POLICIES keep composing, so mandatory [prompt policies](governance.md) ride on top of every persona. Built-in personas source their prose from a repo file; operator personas store it inline — see [Where persona prompts live](#where-persona-prompts-live). |
| **Tool visibility** | Which tools the session advertises. Tri-state: *unrestricted* (tracks tool growth and MCP catalogs), *no tools* (the TOOLS prompt block self-suppresses and zero definitions go on the wire), or an *exact set* of names. Including `tool_search` in a set makes it **soft** — tools the model discovers through search join the visible set; omitting it makes the set **hard** (the search pathway is disabled entirely). On commercial providers a soft set costs one prompt-cache re-prime per `tool_search` expansion, since each expansion rewrites the wire tool set and recomposes the prompt. |
| **MCP** | Whether the workstream talks to MCP at all. **Session-wide**: off means no MCP tools for the persona's own hands *or* for in-process task agents, no resource/prompt catalogs, and no listener registrations. This lever expresses infrastructure intent, not behavior shaping. |
| **Memory** | Whether the persona's **own hands** get memory: recalled-memory injection into the prompt, memory-directed metacognitive nudges, and the `memory` tool. Task agents keep their own envelope, and compaction spill/markers are session mechanics that are never persona-gated. An exact tool set that hides `memory` also mutes those nudges, and the compaction-resume pointer follows `recall`'s visibility. |
Visibility is behavior shaping, **not** a security boundary: any tool call
that does reach the wire still clears the same approval, judge, and policy
machinery as always. RBAC and tool policies remain the enforcement layers.
## Snapshot semantics — resolve once, stamp forever
The persona is resolved **once**, at workstream creation, and stamped into
`workstream_config` as five keys (`persona`, `persona_prompt`,
`persona_tools`, `persona_mcp`, `persona_memory`). From then on the session
reads only the stamp:
- **Editing or archiving a persona never changes an existing workstream.**
Rehydrate, resume, and post-compaction resume all run from the stamp.
A mid-session REPL `/resume` adopts the target workstream's stamp for
prompt, tools, and memory; for the MCP lever it can only narrow in
place — adopting an MCP-off stamp drops the live MCP surface, while
adopting an MCP-on stamp into a session whose persona dropped MCP at
construction is refused with an error telling you to reopen the
workstream fresh.
- A workstream outlives its persona — an archived persona keeps labelling
the workstreams stamped with it.
- A partial or unparseable stamp is treated as corruption: session
construction fails loudly rather than silently falling back to a default
envelope the operator never chose.
- Workstreams created before personas existed carry no stamp and keep
legacy behavior, byte-identical to the `engineer` / `orchestrator`
defaults below — with one exception: pre-1.7 workstreams that had
`creative_mode` set are converted by migration `063` into full
`writer` stamps, so they resume as writing sessions rather than as
legacy defaults.
- Forking (`resume_ws` on create) clones the source's stamped persona into the
new workstream; the fork does not re-resolve it.
## Seed personas
Migration `063` seeds six personas. The two per-kind **defaults** carry no
overrides at all, so a zero-touch launch behaves exactly as it did before
personas existed:
| Persona | Kind | Base prompt | Tools | MCP | Memory |
|---|---|---|---|---|---|
| `engineer` *(default)* | interactive | stock | unrestricted | on | on |
| `orchestrator` *(default)* | coordinator | stock | unrestricted | on | on |
| `scribe` | interactive | custom (faithful structuring of given material) | none | off | off |
| `researcher` | interactive | custom (evidence-first) | `read_file`, `search`, `web_fetch`, `web_search`, `recall`, `memory`, `tool_search` (soft) | off | on |
| `writer` | interactive | custom (creative writing partner — replaces the removed `/creative`) | none | off | on |
| `executive` | coordinator | custom (delegate, interrogate plans, judge outcomes) | spawn/inspect/lifecycle tools plus `memory`: `spawn_workstream`, `spawn_batch`, `send_to_workstream`, `wait_for_workstream`, `inspect_workstream`, `list_workstreams`, `list_nodes`, `close_workstream`, `cancel_workstream`, `memory` (hard) | off | on |
Notes:
- `scribe` turns memory off deliberately: recalled memories would
contaminate faithful summarization with unrelated context.
- `researcher`'s set is soft (includes `tool_search`): it starts with
read and evidence tools but can pull in others on demand — e.g. load
`bash` to run a snippet and verify a calculation. It is evidence-first,
not sandboxed; any escalated tool still hits the normal approval path.
- Coordinator sessions do not merge MCP today, so the MCP lever on
coordinator personas is forward-compatible bookkeeping; it bites on
interactive workstreams.
## Where persona prompts live
Prompt source is explicit in the persona row — two nullable columns, never both empty:
| `base_prompt_file` | `base_prompt` | Meaning |
|---|---|---|
| set (e.g. `scribe.md`) | — | **built-in**: prose lives in `prompts/personas/<file>`, code-owned and PR-reviewed |
| set | set | built-in with an **operator override** layered on top (the inline text wins) |
| — | set | **operator** persona, inline prose |
A `CHECK` forbids the both-empty row, so resolution is a plain coalesce —
`base_prompt ?? load(base_prompt_file)` — with no implicit "inherit the default"
branch in application logic. `base_prompt_file` is set only by the migration/code
(the admin API never exposes it): it marks a persona as built-in and blocks
archive, so `engineer` and `orchestrator` can't be removed. To customise a
built-in, set `base_prompt` on it (clear it to revert), or create your own persona.
The resolved prompt is **frozen into the workstream at creation** — later edits to
a built-in's file or an operator's row never change a running workstream; only new
ones pick up the change. "No persona" is not a state: every workstream is stamped,
and an empty `persona=` resolves to the kind's `is_default` (`engineer` /
`orchestrator`).
## Choosing a persona
Every creation surface takes an optional persona; empty always means the
kind's default (or plain legacy behavior on a database with no personas
seeded):
- **Web/console**: the persona select on the console launcher, the server
webui's new-workstream dialog, and the dashboard composer. Selecting a
persona requires **no** `persona.*` permission — the picker feed
(`GET /v1/api/personas`) is authenticated-only and returns display fields.
- **API/SDK**: `CreateWorkstreamRequest.persona` (Python:
`create_workstream(persona=...)`; TypeScript: `{ persona: ... }`).
- **CLI**: `turnstone --persona <name>`. Unknown or disabled names error at
startup. `--resume` ignores `--persona` and adopts the resumed
workstream's stamp.
- **Coordinator spawn**: `spawn_workstream` / `spawn_batch` take a
`persona` argument, validated when the coordinator prepares the spawn
and re-checked by the node that creates the child (children are always
interactive-kind). Omitted means the interactive **default** — a child
never inherits its parent coordinator's persona.
- **Sub-agents**: `task_agent` takes a `persona` argument setting the
sub-agent's identity and capability envelope (resolved against
interactive-kind personas, frozen into the task at prep). Omitted keeps
the default autonomous task-agent identity — never the parent's persona.
## How agents discover personas
Agents are told, not expected to guess: the live persona list (enabled,
interactive-kind — children and sub-agents are always interactive) is
injected into the `persona` parameter description of `task_agent`,
`spawn_workstream`, and `spawn_batch` whenever the session's tool surface
is rendered — session start, MCP catalog change, model-registry reload.
Each entry carries the name, the default marker, and the persona's
one-line description so the model can pick by purpose (descriptions drop
out past 25 personas; the name list always enumerates completely).
A persona created after that render is still reachable — pass its name.
Every resolve failure enumerates the names currently valid for the kind,
so a stale list (or a typo) self-corrects on the next attempt.
Resolution is forgiving on all surfaces (they share one rule):
- names match case-insensitively (`Writer` resolves `writer`);
- an input that uniquely matches a persona's **display name**
(case-insensitive, among the kind's enabled personas — display names are
not unique, and a same-label persona of another kind neither blocks nor
wins) resolves to that persona; an ambiguous match errors, listing the
candidate slugs;
- whatever variant matched, the stamped identity, approval chrome, and
wire always carry the canonical `name` slug.
## Authoring (console)
Personas are managed in the console's **Manage → Governance → Personas**
tab. The admin shelf exposes exactly the four levers plus the kind
list, the default marker, and archive. Rules:
- `name` is an immutable lowercase slug — and the identifier agents and
the CLI launch the persona by (`persona=` on the spawn tools,
`--persona` on the CLI); the create shelf says so under **Name**.
`display_name` is a list label, editable any time, and deliberately
not an identifier (a unique display name happens to resolve, as a
forgiveness fallback — don't design workflows around it).
- Exactly one default per kind, storage-enforced: flipping the flag on a
successor demotes the incumbent atomically, defaults are single-kind,
and a default cannot be archived.
- **Archive only** — there is no delete verb, so every stamped
workstream's provenance stays explicable.
RBAC: `persona.create` / `persona.read` / `persona.write` gate the admin
CRUD (`/v1/api/admin/personas`); all three are granted to `builtin-admin`
by migration `063`, and other roles opt in via role permission overrides.
+8 -44
View File
@@ -15,19 +15,11 @@ down to a small number of real database connections.
## Why PgBouncer works well with turnstone
Most turnstone database operations are short-burst queries: acquire a
connection, execute a small transaction, commit, release. Workstream forks are
the deliberate exception: they clone the source's checkpoint-bounded history
and configuration and retain its attachment references in one transaction.
PostgreSQL runs that clone at `SERIALIZABLE` isolation and retries serialization
or deadlock conflicts as a whole. A large fork can therefore hold its assigned
server connection longer than an ordinary message write.
This still makes **transaction pooling mode** the right fit — no operation
depends on server-session state, and PgBouncer returns the connection as soon
as the transaction finishes. Size and monitor the server pool with concurrent
fork traffic in mind rather than assuming every transaction completes in a few
milliseconds.
All turnstone database operations are short-burst queries: acquire a
connection, execute 13 statements, commit, release. No operation holds
a connection for more than a few milliseconds. This makes **transaction
pooling mode** ideal — PgBouncer assigns a real connection only for the
duration of each transaction, then returns it to the pool.
| Cluster size | Client connections (max) | PgBouncer server connections needed |
|--------------|------------------------|-------------------------------------|
@@ -151,11 +143,9 @@ PgBouncer (which then multiplexes to PostgreSQL):
| `TURNSTONE_DB_URL` | — | Connection URL (point at PgBouncer, not PostgreSQL directly) |
The default pool of 2 + 3 overflow = 5 connections per process is
intentionally small to support large clusters. Most deployments should not
need to increase it. If operators create many large forks concurrently, watch
PgBouncer's `cl_waiting` and PostgreSQL transaction latency before changing
the per-process pool; adding client-side connections cannot help once the
PgBouncer server pool is saturated.
intentionally small to support large clusters. You should not need to
increase this — turnstone's database operations are all short-burst
context-managed queries that hold connections for milliseconds.
SQLAlchemy `pool_pre_ping` is enabled, so stale connections (e.g. after
PgBouncer restarts) are automatically detected and replaced.
@@ -187,32 +177,6 @@ Key metrics to watch:
- **`sv_active`** — active server (PostgreSQL) connections. Should stay
below PostgreSQL `max_connections`.
Short `cl_waiting` spikes during large workstream forks can be normal. Sustained
waiters accompanied by long serializable transactions indicate fork/storage
load, not an SSE or HTTP client-pool problem.
---
## Upgrade note: deferred workstream creation
The workstream lifecycle now uses durable, hidden `state='creating'`
reservations while session construction, upload validation, and optional fork
cloning complete. Older server processes do not understand that private state:
against the same database they may resolve, list, open, or prune a reservation
before its new owner publishes it.
For the upgrade that introduces deferred creation, drain create traffic and
upgrade all server processes sharing the database as one cohort. Do not resume
creates until no older server process remains. The change needs no manual
schema migration, but it is not safe to treat mixed lifecycle implementations
as an ordinary rolling-upgrade state.
A `creating` row should be transient and absent from normal APIs and cluster
events. If one persists after a process crash, inspect the corresponding
`ws.create.*` and `session_mgr.commit_create.*` logs before cleanup. Do not
promote it to `idle` manually: its history, configuration, attachment
references, or lifecycle publication may be incomplete.
---
## Troubleshooting
+29 -138
View File
@@ -50,7 +50,6 @@ with TurnstoneServer("http://localhost:8080") as client:
import asyncio
from turnstone.sdk import AsyncTurnstoneServer
async def main():
async with AsyncTurnstoneServer("http://localhost:8080") as client:
await client.login(username="alice", password="s3cret")
@@ -59,7 +58,6 @@ async def main():
if event.type == "content":
print(event.text, end="", flush=True)
asyncio.run(main())
```
@@ -71,18 +69,17 @@ Both `TurnstoneServer` (sync) and `AsyncTurnstoneServer` (async) expose:
|----------|--------|---------|
| **Workstreams** | `list_workstreams()` | `ListWorkstreamsResponse` |
| | `dashboard()` | `DashboardResponse` |
| | `create_workstream(*, name, model, auto_approve, resume_ws, skill, persona, initial_message, project_id, attachments, ...)` | `CreateWorkstreamResponse` |
| | `create_workstream(*, name, model, auto_approve, skill, initial_message, attachments)` | `CreateWorkstreamResponse` |
| | `close_workstream(ws_id)` | `StatusResponse` |
| **Attachments** | `upload_attachment(ws_id, filename, data, *, mime_type=...)` | `UploadAttachmentResponse` |
| | `list_attachments(ws_id)` | `ListAttachmentsResponse` |
| | `get_attachment_content(ws_id, attachment_id)` | `bytes` |
| | `delete_attachment(ws_id, attachment_id)` | `StatusResponse` |
| **Chat** | `send(message, ws_id, *, attachment_ids=None, client_send_id=None)` | `SendResponse` |
| | `approve(*, ws_id, approved, feedback, always, cycle_id, call_id)` | `ApproveResponse` |
| **Chat** | `send(message, ws_id)` | `SendResponse` |
| | `approve(*, ws_id, approved, feedback, always)` | `StatusResponse` |
| | `command(*, ws_id, command)` | `StatusResponse` |
| | `cancel(ws_id, *, force=False)` | `CancelResponse` |
| **History** | `get_history(ws_id, *, limit=100)` | `WorkstreamHistoryResponse` |
| **Streaming** | `stream_events(ws_id, *, last_event_id=None, history_token=None)` | `Iterator[ServerEvent]` |
| | `cancel(ws_id, *, force=False)` | `StatusResponse` |
| **Streaming** | `stream_events(ws_id)` | `Iterator[ServerEvent]` |
| | `stream_global_events()` | `Iterator[ServerEvent]` |
| **High-level** | `send_and_wait(message, ws_id, *, timeout, on_event)` | `TurnResult` |
| **Saved** | `list_saved_workstreams()` | `ListSavedWorkstreamsResponse` |
@@ -103,7 +100,7 @@ Both `TurnstoneConsole` (sync) and `AsyncTurnstoneConsole` (async) expose:
| | `workstreams(*, state, node, search, sort, page, per_page)` | `ClusterWorkstreamsResponse` |
| | `node_detail(node_id)` | `NodeDetailResponse` |
| | `snapshot()` | `ClusterSnapshotResponse` |
| | `create_workstream(*, node_id, name, model, initial_message, skill, persona, resume_ws)` | `ConsoleCreateWsResponse` |
| | `create_workstream(*, node_id, name, model, initial_message, skill)` | `ConsoleCreateWsResponse` |
| **Schedules** | `list_schedules()` | `ListSchedulesResponse` |
| | `create_schedule(*, name, schedule_type, initial_message, ...)` | `ScheduleInfo` |
| | `get_schedule(task_id)` | `ScheduleInfo` |
@@ -128,12 +125,12 @@ SSE events are deserialized into typed dataclasses. Use `event.type` to discrimi
| Type | Class | Key Fields |
|------|-------|------------|
| `connected` | `ConnectedEvent` | `model`, `model_alias`, `skip_permissions` |
| `user_turn` | `UserTurnEvent` | `ws_id`, `content`, `attachments`, `sender`, `source`, `client_send_ids`, `_event_id` |
| `history` | `HistoryEvent` | `messages` |
| `content` | `ContentEvent` | `text` |
| `reasoning` | `ReasoningEvent` | `text` |
| `tool_info` | `ToolInfoEvent` | `items` |
| `approve_request` | `ApproveRequestEvent` | `cycle_id`, `items` |
| `tool_result` | `ToolResultEvent` | `call_id`, `name`, `output`, `is_error`, `preview`, `accepted`, `effect_status`, `_event_id` |
| `approve_request` | `ApproveRequestEvent` | `items` |
| `tool_result` | `ToolResultEvent` | `call_id`, `name`, `output`, `is_error` |
| `tool_output_chunk` | `ToolOutputChunkEvent` | `call_id`, `chunk` |
| `status` | `StatusEvent` | `prompt_tokens`, `total_tokens`, `pct`, `effort`, `cache_creation_tokens`, `cache_read_tokens` |
| `error` | `ErrorEvent` | `message` |
@@ -141,92 +138,14 @@ SSE events are deserialized into typed dataclasses. Use `event.type` to discrimi
| `stream_end` | `StreamEndEvent` | — |
| `state_change` | `StateChangeEvent` | `state``running`/`thinking`/`attention`/`idle`/`error` |
| `in_progress_snapshot` | `InProgressSnapshotEvent` | `content`, `reasoning` (one-shot mid-stream refresh resume) |
| `approval_resolved` | `ApprovalResolvedEvent` | `cycle_id`, `call_ids`, `approved`, `feedback`, `always` |
| `approval_resolved` | `ApprovalResolvedEvent` | `approved`, `feedback` |
| `cancelled` | `CancelledEvent` | — |
| `history_resync` | `HistoryResyncEvent` | `reason`, optional `ws_id` |
The Python server `send()` and console `coordinator_send()` methods accept an
optional `client_send_id`; TypeScript `send()` accepts the equivalent
`options.clientSendId`. Values match `[A-Za-z0-9_-]{1,128}`. The value is an
opaque optimistic-UI correlation token, not an idempotency key: reusing it
still creates distinct accepted turns and events.
Every upgraded listener on the shared workstream receives `UserTurnEvent`.
Originating panes use `client_send_ids` only to settle the exact optimistic
bubble, while peers render the accepted row once by `_event_id`. A
`message_queued` event carrying the token can establish acceptance even if the
POST acknowledgement is lost. History projects the same correlation alongside
the accepted user row. These tokens are not credentials: when sender and viewer
identities are both known, only a matching sender may settle local optimistic
state; a peer event still renders its canonical row.
The typed projection is negotiated with `?user_turn=1` on the per-workstream
SSE URL. Python `stream_events()` / `send_and_wait()` and TypeScript
`streamEvents()` / `sendAndWait()` set it automatically. Raw consumers that
omit it receive a backward-compatible `replay_truncated` repair signal instead
of the user row and must rebuild from `/history`; its pre-row cursor keeps the
repair retryable if that history request fails.
The browser-only final-tool upsert capability is `?tool_turn=1`. The bundled
Python and TypeScript SDK streaming helpers and channel adapters intentionally
do not negotiate it yet: they retain the executor-receipt `tool_result`
contract and do not own a transcript reducer. `ToolResultEvent` can deserialize
the accepted fields for direct/custom capable clients. Raw capable clients must
deduplicate `_event_id` and replace the newest matching call occurrence; raw
incapable clients receive the pre-row `tool_turn_projection_unsupported` repair
frame and rebuild from history. That staging deliberately prices in two costs
for incapable consumers. A raw client that treats every `replay_truncated`
frame as a rebuild trigger refetches `/history` once per accepted tool row —
one fetch per tool call on a long agentic turn; a client that wants tool
results incrementally should negotiate `tool_turn=1` and reduce, and the
bundled helpers (which ignore the frame rather than rebuild) stay correct
because their receipt-only view never depends on the accepted projection.
Second, only the accepted event carries post-execution output transforms, so a
receipt-rendering consumer (for example, a channel adapter posting the
executor receipt into a thread) keeps the pre-transform text; the accepted
projection is a transcript-consistency mechanism, not a wire confidentiality
boundary — see the API reference note on the preliminary `tool_result`.
Current servers bootstrap conversation history through
`GET /v1/api/workstreams/{ws_id}/history` before the SSE stream; they do not
emit a `history` event. `HistoryEvent` remains deserializable only for
compatibility with older servers. `get_history()` exposes the current REST
bootstrap response, including its optional cursor and one-shot handoff token.
### Caller-managed history handoff
The SDK supplies typed handshake primitives but intentionally does not own a
transcript renderer or reconnect policy. After rendering a successful history
response, pass its cursor and token to exactly one initial stream:
```python
from turnstone.sdk import HistoryResyncEvent
history = client.get_history(ws_id)
render(history.messages)
for event in client.stream_events(
ws_id,
last_event_id=history.cursor,
history_token=history.handoff_token,
):
if isinstance(event, HistoryResyncEvent):
# Stop this stream. The caller chooses when to fetch, render, and
# reconnect with a new history response.
break
apply_live_event(event)
```
`history_resync` means numeric replay cannot prove that the rendered limited
tail came from the same total accepted conversation-row prefix. Stop the
stream, fetch and render history again, and use only the new cursor/token pair.
A 503 history response raises `TurnstoneAPIError`; it is not authoritative, so
retain any existing transcript and do not open a tokenless replacement stream.
**Global events** (from `stream_global_events()`):
| Type | Class | Key Fields |
|------|-------|------------|
| `ws_state` | `WsStateEvent` | `ws_id`, `state`, `tokens`, `activity`, `persistence_state` |
| `ws_state` | `WsStateEvent` | `ws_id`, `state`, `tokens`, `activity` |
| `ws_activity` | `WsActivityEvent` | `ws_id`, `activity`, `activity_state` |
| `ws_rename` | `WsRenameEvent` | `ws_id`, `name` |
| `ws_closed` | `WsClosedEvent` | `ws_id` |
@@ -237,30 +156,24 @@ retain any existing transcript and do not open a tokenless replacement stream.
|------|-------|------------|
| `node_joined` | `NodeJoinedEvent` | `node_id` |
| `node_lost` | `NodeLostEvent` | `node_id` |
| `cluster_state` | `ClusterStateEvent` | `ws_id`, `node_id`, `state`, `tokens`, `persistence_state` |
| `ws_created` | `ClusterWsCreatedEvent` | `ws_id`, `node_id`, `name`, `persistence_state` |
| `cluster_state` | `ClusterStateEvent` | `ws_id`, `node_id`, `state`, `tokens` |
| `ws_created` | `ClusterWsCreatedEvent` | `ws_id`, `node_id`, `name` |
| `ws_closed` | `ClusterWsClosedEvent` | `ws_id` |
| `ws_rename` | `ClusterWsRenameEvent` | `ws_id`, `name` |
| `snapshot` | `ClusterSnapshotEvent` | `nodes`, `overview`, `timestamp` |
Operator-facing workstream rows and rich state events expose only the sanitized
`persistence_state`: `healthy`, `pending`, `retrying`, or `conflict`. SDK types
treat it as optional for compatibility with older nodes; an omitted value means
`healthy`. Retry counts, storage errors, commit keys, and conversation content
are never part of this status surface.
### TurnResult
The `send_and_wait()` method returns a `TurnResult` that aggregates the full response:
```python
result = client.send_and_wait("Hello", ws_id, timeout=60)
result.content # Full text response
result.reasoning # Chain-of-thought (if shown)
result.tool_results # List of (tool_name, output) tuples
result.errors # Any error messages
result.ok # True if no errors and not timed out
result.timed_out # True if timeout expired
result.content # Full text response
result.reasoning # Chain-of-thought (if shown)
result.tool_results # List of (tool_name, output) tuples
result.errors # Any error messages
result.ok # True if no errors and not timed out
result.timed_out # True if timeout expired
```
### Attachments
@@ -270,7 +183,9 @@ Upload files to a workstream and attach them to the next user turn:
```python
# Upload separately, then send a message — attachments auto-attach
with open("screenshot.png", "rb") as f:
att = client.upload_attachment(ws.ws_id, "screenshot.png", f.read(), mime_type="image/png")
att = client.upload_attachment(ws.ws_id, "screenshot.png",
f.read(),
mime_type="image/png")
client.send("What's wrong in this screenshot?", ws.ws_id)
# Or attach at workstream-creation time (multipart upload)
@@ -280,7 +195,9 @@ with open("notes.txt", "rb") as f:
ws = client.create_workstream(
name="triage",
initial_message="Summarize the notes",
attachments=[AttachmentUpload(data=f.read(), filename="notes.txt", mime_type="text/plain")],
attachments=[AttachmentUpload(data=f.read(),
filename="notes.txt",
mime_type="text/plain")],
)
```
@@ -289,26 +206,6 @@ Limits: images ≤ 4 MiB (png/jpeg/gif/webp), text ≤ 512 KiB (UTF-8),
client so cluster-routed callers bind attachments to the owning node
before the request lands.
### Forking a workstream
`resume_ws` is the API's compatibility name for an atomic fork. It creates a
new workstream ID while the source remains unchanged:
```python
fork = client.create_workstream(
resume_ws=ws.ws_id,
name="analysis-branch",
initial_message="Try the alternative plan.",
)
assert fork.resumed
```
The server transaction clones the source's checkpoint-bounded history, saved
session configuration, persona, project, and attachment references. Do not
combine `resume_ws` with `attachments`; fork first, then upload to the new ID.
To rehydrate the original ID rather than branch it, call the server's
`POST /v1/api/workstreams/{ws_id}/open` endpoint.
### Error Handling
Non-2xx responses raise `TurnstoneAPIError`:
@@ -320,7 +217,7 @@ try:
client.send("hi", "bad_ws_id")
except TurnstoneAPIError as e:
print(e.status_code) # 404
print(e.message) # "Unknown workstream"
print(e.message) # "Unknown workstream"
```
---
@@ -345,14 +242,8 @@ const ws = await client.createWorkstream({ name: "demo" });
const result = await client.sendAndWait("Hello!", ws.ws_id);
console.log(result.content);
// Render history, then use its one-shot hints on the initial stream.
const history = await client.getHistory(ws.ws_id);
render(history.messages);
for await (const event of client.streamEvents(ws.ws_id, {
lastEventId: history.cursor ?? undefined,
historyToken: history.handoff_token ?? undefined,
})) {
if (event.type === "history_resync") break; // caller refetches and reconnects
// Stream events
for await (const event of client.streamEvents(ws.ws_id)) {
if (event.type === "content") {
process.stdout.write(event.text);
}
@@ -428,7 +319,7 @@ turnstone/sdk/ Python SDK (sub-package)
_base.py Shared httpx async client, auth, error handling
_sync.py Background event loop for sync wrappers
_types.py TurnResult + TurnstoneAPIError
events.py Typed SSE event dataclasses with type registry
events.py 38 SSE event dataclasses with type registry
server.py AsyncTurnstoneServer + TurnstoneServer
console.py AsyncTurnstoneConsole + TurnstoneConsole
+20 -53
View File
@@ -64,17 +64,15 @@ Scopes are hierarchical — higher scopes imply all lower ones.
### Path-to-scope mapping
| Method | Path pattern | Required scope | Additional RBAC gate |
|--------|-------------|----------------|----------------------|
| GET | Any protected path | `read` | Endpoint-specific where documented |
| POST | `/api/command` | `write` | Project tenancy on the target workstream |
| POST | `/api/workstreams/new`, `/api/cluster/workstreams/new` | `write` | `workstreams.create` or `admin.coordinator` |
| POST | `/api/workstreams/{ws_id}/close` | `write` | `workstreams.close` or `admin.coordinator` |
| POST | `/api/workstreams/{ws_id}/approve` | `approve` | `tools.approve` or `admin.coordinator` |
| POST | `/api/workstreams/{ws_id}/{rewind,retry}` | `write` | `conversation.modify` |
| POST | Other `/api/workstreams/{ws_id}/...` mutation endpoints | `write` | Project tenancy and endpoint-specific gates |
| DELETE | `/api/workstreams/{ws_id}/send` (dequeue), `/api/workstreams/{ws_id}/attachments/{attachment_id}` | `write` | Project tenancy on the target workstream |
| Any | `/api/admin/*` | `approve` | Matching `admin.*` permission |
| Method | Path pattern | Required scope |
|--------|-------------|----------------|
| GET | Any protected path | `read` |
| POST | `/api/command` | `write` |
| POST | `/api/workstreams/new`, `/api/cluster/workstreams/new` | `write` |
| POST | `/api/workstreams/{ws_id}/{send,cancel,close,delete,open,refresh-title,title,attachments}` | `write` |
| DELETE | `/api/workstreams/{ws_id}/send` (dequeue), `/api/workstreams/{ws_id}/attachments/{attachment_id}` | `write` |
| POST | `/api/workstreams/{ws_id}/approve` | `approve` |
| Any | `/api/admin/*` | `approve` |
Public paths bypass authentication entirely: `/`, `/health`, `/metrics`,
`/static/*`, `/shared/*`, `/docs`, `/openapi.json`, `/api/auth/login`,
@@ -86,7 +84,7 @@ Public paths bypass authentication entirely: `/`, `/health`, `/metrics`,
> See also: [Governance documentation](governance.md)
Scopes provide coarse endpoint-level access control. For finer-grained
enforcement, the governance layer adds named permissions checked
enforcement, the governance layer adds 15 named permissions checked
per-endpoint by `require_permission()`. Permissions are bundled into
roles; users are assigned roles via the `user_roles` join table.
@@ -100,8 +98,8 @@ Three built-in roles are seeded by migration 008:
| Role | Permissions |
|------|-------------|
| admin | Admin-default baseline (all ordinary admin and lifecycle permissions; explicitly opt-in capabilities remain ungranted) |
| operator | read, write, workstreams.create, workstreams.close, conversation.modify |
| admin | All 15 permissions |
| operator | read, write, workstreams.create, workstreams.close |
| viewer | read |
Custom roles can be created with any subset of the valid permissions.
@@ -109,34 +107,6 @@ Role creation and update validate permissions against a static allowlist.
Self-assignment is blocked, and assigning a role requires the caller to
hold a superset of the target role's permissions.
### Workstream lifecycle and project boundaries
The remote `/api/command` endpoint is conversation-local. It refuses
`/new`, `/workstreams`, `/resume`, and `/delete` because those local-CLI
helpers enumerate or mutate storage outside the HTTP resource gates. Remote
clients use the dedicated create, open, close, and delete endpoints instead;
`/rewind` and `/retry` have their own path-keyed, `conversation.modify`-gated
endpoints.
Passing `resume_ws` to create is an atomic **fork**, not an in-place resume.
It requires the ordinary create capability and source visibility. A private
project source is visible only to its workstream creator, project owner/member,
or authorized service-to-service cluster plumbing; denials use a not-found
response so guessed IDs do not become an existence oracle. The caller must also
be allowed to attach a new workstream to the source's current project. The
destination always inherits that effective project — a caller-supplied
`project_id` cannot re-file or declassify the conversation.
The canonical preflight atomically captures (and, for a legacy row, installs) a
private source-incarnation fence. The storage transaction compares that source
fence, rejects provisional sources, repeats the ACL/project check, and verifies
the persona/project construction snapshot, destination ownership and
incarnation, emptiness, and every referenced attachment before committing. A
source replacement, membership, project, persona, or destination-incarnation
race aborts the whole fork. Concurrent source-history writes serialize wholly
before or after the snapshot; no mixed or partially authorized history or
attachment reference becomes visible.
---
## Login Flows
@@ -471,15 +441,12 @@ Each proxied request gets a fresh JWT (5-minute expiry). This ensures:
- **Permission forwarding** — granular RBAC permissions from the
console JWT are carried through to the server.
For ordinary users the JWT `src` claim is set to `"console-proxy"`, allowing
servers to distinguish proxied requests from direct logins in audit logs.
Coordinator tokens retain `src="coordinator"` and their signed `coord_ws_id`;
the console service identity retains `src="console"` only when its validated
token also carries the unassignable `service` scope.
The JWT `src` claim is set to `"console-proxy"`, allowing servers to
distinguish proxied requests from direct logins in audit logs.
When no user context is available (auth disabled, or internal requests),
the proxy falls back to a `ServiceTokenManager` with identity `console-proxy`,
`src="console"`, and `{read, write, approve, service}` scopes.
the proxy falls back to a `ServiceTokenManager` with service identity
`console-proxy` and full scopes.
### Service-to-service authentication
@@ -488,8 +455,8 @@ JWTs when communicating with server nodes:
| Service | Identity | Scope | Audience | Purpose |
|---------|----------|-------|----------|---------|
| Console collector | `console-collector` | `read`, `service` | `turnstone-server` | Node health polling and global event collection |
| Console proxy (fallback) | `console-proxy` | `read`, `write`, `approve`, `service` | `turnstone-server` | Proxied API calls when no user context |
| Console collector | `console-collector` | `read` | `turnstone-server` | Node health polling |
| Console proxy (fallback) | `console-proxy` | `approve` | `turnstone-server` | Proxied API calls when no user context |
| Channel notify | `system` | `write` | `turnstone-channel` | Notification delivery to channel gateway |
Service tokens use 1-hour expiry with automatic refresh via
@@ -501,8 +468,8 @@ When the console creates a workstream (the normal path), the
authenticated user's `user_id` is forwarded in the HTTP payload when
calling the server's `POST /v1/api/workstreams/new`. The server
accepts a `user_id` from the request body **only when the caller is a
trusted service** — identified by `token_source="console"` together with the
unassignable `service` scope. `console-proxy`, coordinator, and regular API callers cannot
trusted service** — identified by `token_source` matching
`console-proxy` or `console`. Regular API callers cannot
override `user_id`; the server always uses their JWT identity.
Note that the channel gateway uses a distinct JWT audience
+11 -182
View File
@@ -54,149 +54,6 @@ When a per-model override is `NULL` (empty in the UI), the global default is
used. Switching models via `/model <alias>` re-resolves sampling parameters
from the new model's overrides or global defaults.
### Per-model concurrency
Each model definition may set `max_concurrency` to limit simultaneous model
generations for that alias in one Turnstone process. `0` or an omitted value
means unlimited. The gate is shared by every role using the alias—interactive
turns, coordinators, task agents, judges, output guards, perception, compaction,
and title generation—and a streaming generation holds its slot until the
stream is fully drained or closed.
Admission is strictly per alias. Two aliases remain independent even when they
point to the same URL; Turnstone does not infer shared capacity from endpoint
text. Queue time is excluded from judge/output-guard deadline accounting, and
each retry releases its slot before backoff and reacquires for the next wire
attempt. The cap is local to each process, not cluster-wide; account for the
number of nodes targeting the same inference server. Direct STT/TTS protocol
calls and Cohere/Jina reranking do not currently consume this generation cap.
### Judge batch parallelism
`judge.parallel_evaluations` controls how many independent tool calls from one
approval batch the intent judge evaluates concurrently. It is an integer from
1 through 16 and defaults to 1, preserving serial evaluation until an operator
opts into wider fan-out. Changes are hot-read at the next batch; work already
in flight keeps its captured worker count.
This is a per-batch fan-out setting, not another backend capacity limit. The
judge model alias's `max_concurrency` gate still caps total generations across
all judge batches and every other role using that alias. Actual overlap is
therefore bounded by the batch size, `judge.parallel_evaluations`, and available
alias admission slots. A smaller positive alias cap also narrows the batch's
worker pool so excess judge threads do not queue ahead of later alias traffic.
### Model backend authentication
Model definitions support four backend credential modes:
| `auth_mode` | Identity sent to the model gateway |
|-------------|------------------------------------|
| `static` | The definition's stored `api_key`. |
| `entra_obo` | A caller-delegated Entra access token minted from that user's captured OIDC credential. |
| `entra_app` | A shared app-identity token minted with Turnstone's OIDC client credentials. |
| `rfc8693_obo` | A caller-delegated access token minted from the captured credential via RFC 8693 token exchange, requesting the definition's `obo_scopes`. |
Dynamic modes require an exact `obo_audience` resource identifier. Before an
admin can save one, an operator must add that literal audience to
`model.auth_audience_allowlist` (comma- or newline-separated). Wildcards and
base-URL host matching are intentionally unsupported, and a row whose
effective mode is `static` refuses to store a new non-empty `obo_audience` on
either create or update — an audience cannot be staged for a later flip
(clearing a stale value, or re-saving it unchanged, stays allowed).
`obo_scopes` follows the same staging rule with the mode set inverted: only
`rfc8693_obo` reads it, so every other effective mode refuses to store a new
non-empty value, while clearing or re-saving one unchanged stays open. The
value itself is optional and shape-checked only — whether it satisfies the
IdP is decided at mint time. On a row that is (or becomes) dynamic, every
change except the tuning fields — context window, temperature, max tokens,
reasoning effort, and the two reasoning-persistence toggles — also requires
`admin.mcp`; service tokens do not bypass this capability-escalation gate.
The one exception is de-escalation: a save whose only gated change is
switching `enabled` off is a pure disable, needs only `admin.models`, and
skips validation — a de-listed audience must never block disarming its own
row. The gate is deny-by-default: a field counts as auth-relevant unless it
is provably neutral, so re-enabling a disabled dynamic row, re-pointing its
`base_url`, or swapping its provider or alias all escalate.
Validation runs in two tiers, matching the MCP `oauth_obo` write rules. Row
validity — the audience is allow-listed — applies to every gated write that
touches a dynamic configuration, so a revoked audience can be neither silently
re-pointed at a new `base_url` nor re-armed by an enable flip. Deployment
posture — the token encryption key installed, single sign-on configured, and
the grant profile valid and able to carry the mode — is checked when a write
*chooses* the mode/audience pair and when it re-enables a disabled dynamic
row (arming is the flip that resumes minting, so it must meet what minting
needs); other edits to an existing row stay open if the deployment's posture
changed after it was saved (its mints warn at runtime instead). Refusals name
their cause and echo the configured value.
One asymmetry to be aware of: the write path counts a transient discovery
outage (`enabled=false`, retryable) as configured, but the mints themselves
require discovery to have completed — a config saved during an outage starts
minting only once any authenticated request heals discovery. Until then calls
warn and follow the fail-open/fail-closed policy above.
Every dynamic mode pairs with exactly one grant profile: `entra_obo` and
`entra_app` require `[oidc] obo_grant_profile = "entra"`, and `rfc8693_obo`
requires `"rfc8693"`. The pairing is enforced at the posture tier, so a row
saved before the rule existed keeps accepting same-pair edits; its mints
refuse at runtime with `cause=grant_profile_mismatch` and no IdP traffic.
Judge, output-guard, perception, utility, and sub-agent lanes inherit the
session's effective user for the delegated modes. The perception memo is
partitioned by that principal as well as alias and content hash, so a result
authorized as one user cannot be served to another. Scheduled and wake-driven
work retains the workstream owner even when no user is connected. Eval and
optimizer lanes are registry-less development tools and therefore do not use
dynamic model authentication.
`entra_app` is an explicit model-definition choice; Turnstone never changes a
failed or ownerless delegated call into a client-credentials grant. A
delegated-mode call with no effective user always refuses. A dynamic alias
without a real static key also always refuses instead of issuing its
SDK-construction placeholder. When a real static key is explicitly configured,
mint failures may use it by default; set `model.auth_fail_closed = true` to
prohibit even that fallback. A refusal is not routed through the model
fallback chain.
Dynamic token caches are encrypted in `mcp_user_tokens`, shared across nodes,
and memoized on each host. Unlinking a user's OIDC identity purges their
delegated-mode rows and memo entries. `entra_app` rows belong to the shared
`__app__` identity and are not user-deprovisioned; after client-credential
revocation, an already-minted app bearer remains usable until its recorded
expiry.
Each model call resolves its dynamic credential against the immutable model
definition snapshot that supplied that call's provider, client, endpoint, and
model ID. An admin edit can therefore never pair an old `base_url` with a new
audience, grant mode, or static-key fallback input. The principal and token
remain per-call/live; the connection and model-owned auth configuration move
together as one binding on the next operation. The deployment-wide
`model.auth_fail_closed` switch is intentionally read live on every mint, so an
operator can tighten fallback policy immediately without rebuilding sessions.
`obo_audience` and `obo_scopes` are literal and capped at 2048 characters
each. Environment-variable expansion is deliberately not applied, so the
allow-list decision cannot vary by node or expand beyond the persisted
boundary.
### Responses output controls (per-model)
Models whose capability table declares Responses output controls expose two
additional fields in the Models create/edit shelf:
| Field | Stored capability | Values | Effect |
|-------|-------------------|--------|--------|
| Output verbosity | `verbosity` | `low`, `medium`, `high` | Controls answer length independently of reasoning effort. |
| Reasoning mode | `reasoning_mode` | `standard`, `pro` | Selects standard or higher-compute Pro execution without changing the model ID. |
An empty selection means provider default and omits the capability key. Known
GPT-5.6 models inherit support from the built-in table without persisting
redundant support flags. An OpenAI-compatible model pinned to the Responses API
can opt in with the `supports_verbosity` and `supports_pro_mode` capability
tiles. Chat Completions and non-Responses providers do not surface or submit
these controls.
**Removed settings:** `model.name` and `model.context_window` have been removed
from ConfigStore. Model names and context windows are now configured per-model
in the Models tab. A startup warning is logged if these keys appear in
@@ -250,7 +107,7 @@ initialization:
| Section | Settings |
|---------|----------|
| `model` | default_alias, auth_audience_allowlist, auth_fail_closed, temperature, max_tokens, reasoning_effort, task_alias, task_effort |
| `model` | default_alias, temperature, max_tokens, reasoning_effort, task_alias, task_effort |
| `session` | instructions, retention_days, compact_max_tokens, auto_compact_pct |
| `tools` | timeout, truncation, agent_max_turns, skip_permissions, search, search_threshold, search_max_results |
| `server` | workstream_idle_timeout, max_workstreams |
@@ -258,7 +115,7 @@ initialization:
| `mcp` | config_path, registry_url |
| `ratelimit` | enabled, requests_per_second, burst, trusted_proxies |
| `health` | backend_probe_interval, backend_probe_timeout, circuit_breaker_threshold, circuit_breaker_cooldown |
| `judge` | enabled, model, smart_approvals, confidence_threshold, max_context_ratio, timeout, parallel_evaluations, read_only_tools, output_guard, output_guard_budget_seconds, output_guard_llm, output_guard_model, output_guard_llm_timeout, redact_secrets, cancel_on_approval |
| `judge` | enabled, model, provider, base_url, api_key, confidence_threshold, max_context_ratio, timeout, read_only_tools, output_guard, redact_secrets, cancel_on_approval |
| `interface` | close_tab_action, theme |
| `skills` | discovery_url |
| `memory` | relevance_k, fetch_limit, max_content, nudge_cooldown, nudges |
@@ -435,11 +292,13 @@ Reset a setting to its registry default by removing it from storage.
## Secret Settings
The registry currently defines no production secret system setting. The generic
machinery nevertheless treats any future `is_secret=True` entry as write-only:
list and write responses return `"***"`, and submitting that sentinel preserves
the stored value. Model API keys are fields on model definitions—not
`judge.*` system settings—and use the Models tab's separate write-only flow.
Settings with `is_secret=True` (currently only `judge.api_key`) are blocked
from the write API with a `403` response. This prevents accidental exposure
through the admin UI or audit logs. Secret settings must be configured via
`config.toml` or environment variables.
The list endpoint masks secret values: stored secrets appear as `"***"`
rather than their actual value.
---
@@ -460,40 +319,10 @@ reload.
**Behavior after reload:**
- New workstreams pick up updated values immediately (via `session_factory`)
- Most workstream/session settings remain the snapshot captured at creation or
resume. Component docs call out deliberate live-read exceptions; for
example, Smart Approval settings are snapshotted coherently at the start of
each approval batch.
- Existing sessions keep their frozen configuration (settings are captured at
workstream creation time, not read on every turn)
- Settings marked `restart_required=True` need a server restart to take effect
### Model-definition reloads
The Models tab has a separate live-reload contract from ordinary ConfigStore
settings. Existing sessions remember the concrete registry generation that
supplied their active alias and re-resolve that alias at the start of the next
send. Endpoint, provider, backend model ID, capabilities, extra parameters, and
backend-auth configuration are replaced as one immutable binding. In-flight
turns, judges, and task agents finish or cancel against the binding they
started with; an admin edit never tears one request across two definitions.
The alias's admission gate is retained and resized in place, so a concurrency
edit preserves in-flight accounting and does not reset cached judges or the
output-guard rate limiter.
Sampling and other saved workstream configuration remain workstream state. A
model-definition edit does not silently rewrite a live workstream's chosen
temperature, reasoning effort, max tokens, skill, or persona. Use
`/model <alias>` (or create/fork a workstream) when an explicit session-level
model switch is intended.
If a live workstream's alias is deleted, its next send first attempts the
configured fallback chain. Without a usable fallback, the operator-facing
error names the removed alias and points interactive users to `/model`; adding
the alias back causes the next send to rebind without a process restart. If a
replacement client cannot be constructed, Turnstone logs one
`session.model_refresh_client_construction_failed` warning per registry
generation and retries only after another model reload, avoiding a rebuild
storm on every send.
---
## Migration from config.toml
+24 -106
View File
@@ -1,7 +1,7 @@
---
name: import-conversation-history
description: Use this skill when the user wants to import or migrate conversation history from another LLM chat or coding tool (e.g. ChatGPT, Claude.ai, Cursor, Copilot Chat, Aider, Gemini, a custom JSON export) into Turnstone. The skill teaches Turnstone's destination contracts — workstream identity, the OpenAI-shaped message rows, tool-call/result pairing, provider-fidelity blobs, attachments, and archive-vs-resumable choice — so the agent can map any source format onto them. Trigger phrases: "import my chats", "migrate this transcript into Turnstone", "bring my Claude.ai history over", "load this export as a workstream".
version: 1.1.0
version: 1.0.0
---
# Importing Conversation History into Turnstone
@@ -12,7 +12,7 @@ Source formats vary; the destination does not. Your job is to translate whatever
Two questions to settle with the user before writing anything:
1. **Archive or resumable?** An archive is left closed and is read-only history. A resumable import is also kept closed and unloaded while rows are written, then explicitly opened after validation; this only works cleanly when the source LLM matches a Turnstone-supported provider/model and tool definitions still resolve.
1. **Archive or resumable?** An archive ("saved" workstream — `state="closed"`) is read-only history. A resumable workstream (`state="idle"`) lets the user continue the conversation; this only works cleanly when the source LLM matches a Turnstone-supported provider/model and tool definitions still resolve.
2. **One workstream per source thread, or merge?** Default to one-to-one unless the user explicitly asks to merge.
Default to **archive** when in doubt — resuming a foreign transcript with mismatched tool schemas or stale provider signatures will fail at the next turn.
@@ -25,13 +25,13 @@ Two tables carry the conversation:
| Column | Required | Notes |
|---|---|---|
| `ws_id` | yes | 32-char lowercase hex. Auto-generate with `secrets.token_hex(16)` if you don't already have one. The router hashes the **full ID** — see "Identity & Routing" below. |
| `ws_id` | yes | 32-char lowercase hex. Auto-generate with `secrets.token_hex(16)` if you don't already have one. **First 4 hex chars are the routing bucket** — see "Identity & Routing" below. |
| `name` | yes | Short title. Pull from source thread title; fall back to first ~60 chars of first user message. |
| `state` | yes | Register as `"closed"` while importing. Leave it closed for an archive; explicitly open it after commit for a resumable import. Never set `"running"` or `"creating"` directly. |
| `state` | yes | `"closed"` for archive, `"idle"` for resumable. Never set `"running"` on import. |
| `kind` | yes | `"interactive"` for normal threads. Do NOT use `"coordinator"` for imports — that's reserved for cluster-spawned coordinator workstreams. |
| `parent_ws_id` | no | Leave NULL. Only set if you're importing a coordinator-spawned subtree and re-parenting it; rare. |
| `user_id` | yes | Owner. Must exist in `users`; importer must know which Turnstone user owns the imported history. |
| `node_id` | no | Nullable creation-time service/liveness hint. It is not the routing key or durable owner and may become stale after membership changes. Let a routed create stamp it; a direct shared-storage import may leave it NULL. |
| `node_id` | yes (multi-node) | Denormalized cache of the node that owns this `ws_id`'s bucket. Single-node deployments can leave it NULL or set it to the only node. |
| `alias` | no | Human-typeable short name. Optional; must be unique cluster-wide if set. |
| `title` | no | Auto-titled later by the LLM; safe to leave NULL on import. |
| `skill_id`, `skill_version` | yes | Default `""` and `0` unless the source thread was scoped to a Turnstone skill. |
@@ -55,65 +55,25 @@ The internal format is **OpenAI-shaped**, even when the source was Anthropic or
## Identity & Routing (`ws_id`)
- `ws_id` is **32-char lowercase hex** (i.e. `secrets.token_hex(16)`).
- Ordinary placement is rendezvous (Highest Random Weight, HRW) selection over
the **full `ws_id`** and the current live server set. For each node, Turnstone
computes 32-bit FNV-1a over the node ID, a NUL separator, and the full
workstream ID; it then applies the node weight and selects the highest score.
A live per-workstream override takes precedence.
- The live set comes from recent `services` heartbeats. Placement can therefore
change when nodes join, leave, change weight, or an override changes. There
is no stable prefix-derived placement to pre-compute or persist.
- `workstreams.node_id` is stamped at creation and is not updated as HRW
placement changes. It supports display and liveness-safe cleanup; the console
router does not use it as the ordinary ownership decision.
- For multi-node imports, create through the console routing proxy when the
lifecycle must be published, or write the history once through the cluster's
configured **shared storage backend**. Never partition rows across node-local
databases by ID prefix or by a one-time HRW result: a later membership change
can route the same full ID to another node.
- For single-node imports, HRW placement is degenerate; any valid `ws_id` works.
- The **routing bucket** is `int(ws_id[:4], 16)` — the first 4 hex chars place this workstream on a specific node via the consistent hash ring.
- For multi-node imports: either insert through the console's routing proxy (which forwards to the owning node), or generate `ws_id`s and write directly to each node's database in batches grouped by bucket.
- For single-node imports: bucket math is irrelevant; any `ws_id` works.
- **Do not reuse the source platform's IDs as `ws_id`** unless they happen to be 32-char hex. Generate fresh; if you need the old ID for traceability, store it in `workstream_config` under a key like `import.source_id`.
## Recommended Import Path
Three options, in order of preference:
### 1. Quiesced storage import (recommended for full history)
### 1. Storage protocol (recommended for full history)
Use the current `turnstone.core.storage.StorageBackend` protocol against the
same shared backend as the cluster. The destination must remain absent from all
in-memory session managers while rows are changing: a loaded `ChatSession`
holds its own trajectory and will not observe conversation rows inserted behind
it.
The safe sequence is:
1. Normalize and validate the complete source transcript before writing.
2. Call `register_workstream(..., state="closed")` and require a `True` return;
`False` means the caller-selected ID already exists, so abort rather than
appending to an unrelated workstream.
3. Insert the ordered conversation rows and attachment references.
4. Load the saved rows back and run the validation checklist below.
5. Leave an archive closed. For a resumable import, only now invoke the normal
`POST /v1/api/workstreams/{ws_id}/open` endpoint on the currently routed
node so the session hydrates from the complete transcript.
Do **not** create the destination through the web/SDK create endpoint before a
direct bulk import. Create publishes an empty live session. If that already
happened, close the workstream and confirm the manager-authoritative live probe
returns false before writing, then explicitly open it again after validation.
For attachment-free history, `save_messages_bulk(rows)` is the canonical
single-transaction insert primitive and bypasses the LLM round-trip entirely.
New attachment bytes require the per-row path described under
[Attachments](#attachments).
Use `turnstone.core.storage.Storage.save_messages_bulk(rows)`. This is the canonical bulk-insert primitive and bypasses the LLM round-trip entirely.
```python
from turnstone.core.storage import get_storage # initialized by the host/import entry point
from turnstone.core.storage import get_storage # construct via the same path the server uses
storage = get_storage()
storage = get_storage(...) # see turnstone.core.storage.__init__ for the project's wiring
inserted = storage.register_workstream(
storage.create_workstream( # or whatever the project's exposed creator is — check turnstone/core/storage/_protocol.py
ws_id=ws_id,
user_id=user_id,
name=name,
@@ -121,8 +81,6 @@ inserted = storage.register_workstream(
kind="interactive",
...
)
if not inserted:
raise RuntimeError(f"destination already exists: {ws_id}")
storage.save_messages_bulk([
{"ws_id": ws_id, "role": "user", "content": "Hello"},
@@ -136,19 +94,7 @@ storage.save_messages_bulk([
])
```
`save_messages_bulk` handles `timestamp` and the workstream's `updated` column
internally, so you don't need to compute them per row. Verify the exact
`register_workstream` and message signatures in
`turnstone/core/storage/_protocol.py`; the Storage protocol, not the physical
table layout, is the source of truth.
**Multi-node note:** this path assumes `get_storage()` is connected to the
cluster's shared backend. Do not open a node-local database selected from the
current HRW result, and do not pre-create a live session through the console
routing proxy. After the shared-storage import commits, resolve the current
route and open the closed workstream on that node. Any stored `node_id`
describes creation-time placement, not a permanent shard that should receive a
separate copy.
`save_messages_bulk` handles `timestamp` and the workstream's `updated` column internally, so you don't need to compute them per row. **Verify the exact creator signature** by reading `turnstone/core/storage/_protocol.py` — table layout has shifted across migrations and the Storage protocol is the source of truth.
### 2. SDK `create_workstream(resume_ws=...)` (when the source is already a Turnstone workstream)
@@ -235,48 +181,27 @@ If the source thread had image or file attachments:
- **Size limits**: images ≤ 4 MiB, text documents ≤ 512 KiB. Reject or downsample anything bigger.
- **Allowed types**: server validates magic bytes for images and UTF-8-decodes for text. Binary blobs that aren't images won't pass.
- **Blob identity**: `attachment_id` is the lowercase SHA-256 hex digest of the
bytes. `workstream_attachments` stores that content-addressed blob and its
refcount; it has no workstream or message foreign key.
- **Message link**: the sole message-to-blob link is the ordered JSON ID list in
`conversations.attachments`.
- **No persisted staging lifecycle**: pending upload bytes live only in a
node's in-memory attachment buffer. The old persisted
`pending → reserved → consumed` lifecycle does not apply to storage imports.
- **Lifecycle**: pending → reserved → consumed. For imports, the cleanest path is to upload as pending and immediately consume by attaching to the relevant `conversations.id`.
For new attachment bytes, preserve row order by calling `save_message()` for
each turn. It returns the `conversations.id`; for every attachment referenced by
that turn, call `save_attachment()` with its content hash and bytes, then call
`set_message_attachments(ws_id, message_id, ordered_ids)`. Each
`save_attachment()` call accounts for one reference, while
`set_message_attachments()` records the ordered link.
Two import paths:
`save_messages_bulk(..., attachment_ids=[...])` is appropriate only when those
content-addressed blobs already exist: the bulk transaction retains their
references and writes the ordered lists. Do not first call `save_attachment()`
for a new reference and then pass the same reference to `save_messages_bulk()`;
both paths retain it and would double-count the refcount.
1. **Bulk-insert + post-attach**: insert messages first, get back the assistant/user `conversations.id`, then write `workstream_attachments` rows linking the file to `message_id`.
2. **SDK multipart create**: `create_workstream(attachments=[...], initial_message=...)` for the *first* turn only — the server reserves and consumes them onto that turn. Doesn't help for mid-thread attachments.
SDK multipart create remains useful only for attachments on a new first turn;
it publishes a live session and is not the full-history import path.
For full-history imports with multiple attachments at different turns, path (1) is the only option.
## Validation Checklist
Before declaring success, verify:
- [ ] `ws_id` is 32-char lowercase hex.
- [ ] The workstream remained closed and absent from every live manager while rows were written; archives stay closed and resumable imports are opened only after validation.
- [ ] `workstreams` row exists with the right `user_id` and `kind`.
- [ ] `workstreams` row exists with the right `user_id`, `state`, `kind`.
- [ ] Conversation rows are inserted **in order** (autoincrement `id` will reflect insert order).
- [ ] Every assistant `tool_calls[].id` has a matching `role="tool"` row with the same `tool_call_id`.
- [ ] `tool_calls[].function.arguments` is a JSON-encoded **string**, not a parsed object.
- [ ] First message is typically `role="user"` (not `system`) — Turnstone composes its own system prompt at runtime.
- [ ] No empty assistant rows (`content=NULL` AND `tool_calls=NULL` is invalid).
- [ ] Every attachment ID is the SHA-256 of its stored bytes; each turn's ordered IDs are in `conversations.attachments`, and blob refcounts match message references.
- [ ] If multi-node: the row is in shared storage and the node selected by
`ConsoleRouter.route(ws_id)` from the current live set can load it.
`workstreams.node_id`, when present, is treated as a creation-time hint rather
than asserted equal to the current HRW result.
- [ ] If multi-node: the `ws_id`'s bucket maps to a node that exists; `workstreams.node_id` matches.
- [ ] Round-trip test: run `Storage.load_messages(ws_id)` and confirm the reconstructed list matches what you inserted (modulo timestamps).
## Anti-patterns
@@ -286,20 +211,15 @@ Before declaring success, verify:
- **Don't fabricate `tool_call_id`s without re-pairing.** Mismatched ids silently break the replay chain on the next turn.
- **Don't skip the `tool_name` field on `role="tool"` rows.** Some load paths use it for display and audit; NULL there will render as "unknown tool".
- **Don't write through the LLM (`send()` per turn) for full history.** It's expensive, rewrites assistant turns, and rate-limits will bite long imports.
- **Don't shard imported rows by an ID prefix or a one-time HRW result.** HRW
uses the full ID and live membership; placement may move. In a cluster, write
one copy to shared storage and let request routing select the live node.
## Quick Reference
| Task | Path |
|---|---|
| Generate ws_id | `secrets.token_hex(16)` |
| Multi-node placement | Full-ID 32-bit FNV-1a HRW over live servers; store rows once in shared storage |
| Bulk insert attachment-free messages | `Storage.save_messages_bulk(rows)` |
| Attach new bytes | `save_message()``save_attachment()` per reference → `set_message_attachments()` |
| Bulk insert messages | `Storage.save_messages_bulk(rows)` |
| Archive (read-only) | `state="closed"`, skip `provider_data` |
| Resumable | Register closed, import and validate while unloaded, then explicitly open; populate `provider_data` if same provider |
| Resumable | `state="idle"`, populate `provider_data` if same provider |
| Tool call id | OpenAI shape: `{"id": ..., "type": "function", "function": {"name": ..., "arguments": "<json string>"}}` |
| Tool result row | `role="tool"`, `tool_name`, `tool_call_id`, `content` |
| Source role → Turnstone role | See "Role Mapping" table |
@@ -308,8 +228,6 @@ Before declaring success, verify:
## Files to read before writing the importer
- `turnstone/core/storage/_schema.py` — authoritative table definitions.
- `turnstone/core/storage/_protocol.py``register_workstream`, message, attachment, and load signatures.
- `turnstone/core/rendezvous.py` — authoritative full-ID FNV-1a HRW scoring.
- `turnstone/console/router.py` — live-node discovery, override precedence, and routing behavior.
- `turnstone/core/storage/_protocol.py``save_message`, `save_messages_bulk`, `load_messages` signatures.
- `turnstone/core/session.py` (around the message-save section) — how the runtime constructs in-memory message dicts; mirror this shape on import to round-trip cleanly.
- `turnstone/api/server_schemas.py` — Pydantic shapes for the SDK paths if you go through HTTP.
+19 -71
View File
@@ -54,11 +54,15 @@ docker compose exec caddy \
cat /data/caddy/pki/authorities/local/root.crt # import into your OS/browser
```
**Can Caddy get its cert from the console's internal CA instead?** Not directly.
The console's ACME signing routes require Turnstone's rotating enrollment JWT,
which a standard Caddy ACME issuer does not attach. Keep `tls internal`, or use a
public ACME CA for a publicly trusted certificate. An authenticated gateway or
Caddy plugin would be required to use Turnstone's responder.
**Can Caddy get its cert from the console's internal CA instead?** Technically
yes — the console exposes a real ACME directory (`/acme/directory`) with
auto-approval, so Caddy's `tls { ca http://console:8090/acme/directory }` would
mint a cert for any name. It's not recommended as the default: lacme's ACME
responder is built for turnstone's own client (interop with Caddy's client is
unverified), it couples Caddy startup to the console, and the browser must trust
a private CA either way — so it buys nothing over `tls internal`. For a publicly
trusted cert (no warning), point Caddy at Let's Encrypt with a real domain
instead.
---
@@ -102,15 +106,13 @@ An mTLS listener rejects plain-HTTP probes at the socket, so
it presents the node's own cert as the client cert and pins the cluster
CA, using the PEM files the server writes at boot under
`$TURNSTONE_TLS_PEM_DIR` (default `<tmpdir>/turnstone-tls`). The probe
dials `localhost` for the TLS attempt, which every service certificate carries
as a DNS SAN. Cert renewal rewrites
dials `localhost` for the TLS attempt — the internal CA issues DNS SANs
only, so a literal-IP URL would fail verification. Cert renewal rewrites
the PEM dir alongside the live listener swap, so the probe's client cert
never outlives the served cert. With TLS disabled the plain probe succeeds
and the PEM directory is never consulted. On bare metal with multiple
nodes per host, set `TURNSTONE_TLS_PEM_DIR` per node (each boot clears
stale `lacme-pem-*` dirs under its root).
The production TLS Compose overlay inherits this healthcheck from the base
service; it remains enabled under mTLS.
---
@@ -123,13 +125,6 @@ service; it remains enabled under mTLS.
| `tls.enabled` | `false` | Master switch for internal mTLS |
| `tls.acme_directory` | `""` | External ACME CA URL for console frontend cert |
### ACME topology environment
| Variable | Default | Description |
|----------|---------|-------------|
| `TURNSTONE_ACME_EXTERNAL_URL` | request-derived | Canonical externally reachable responder base, including `/acme` (for example `http://192.0.2.1:8090/acme`). Set it on the console so advertised URLs are routable and on in-cluster clients so their enrollment JWT is allowed only at that configured destination. A public path prefix is valid only when a reverse proxy maps it to Turnstone's internal `/acme` mount. |
| `TURNSTONE_CONSOLE_HTTP_BIND` | `127.0.0.1` | Production TLS-overlay bind for the console's plain-HTTP bootstrap/API port. For cross-host enrollment, use a trusted LAN/VPN interface and firewall it to enrolling nodes. |
### Bootstrap Config (config.toml)
These are needed before storage is available:
@@ -149,7 +144,7 @@ sslkey = "" # path to client key
| CA common name | "Turnstone CA" | |
| CA validity | 10 years | |
| Cert validity | 48 hours | Short-lived, auto-renewed |
| Renewal interval | 12 hours | Leaves retry headroom before expiry |
| Renewal interval | 24 hours | Half of validity |
| ACME auto-approve | true | Internal network, no challenge validation |
---
@@ -182,14 +177,10 @@ turnstone-admin tls-ca-cert --out ca.pem --console-url http://console:8080
# Request a cert for a domain
turnstone-admin tls-issue worker-1.internal --out /certs --console-url http://console:8080
# List managed cluster certs
# List issued certs
turnstone-admin tls-list --console-url http://console:8080
```
`tls-ca-cert` preserves the supplied scheme. An `https://` console URL is
verified with the system trust store; an explicitly supplied `http://` URL is
TOFU and prints a fingerprint that must be checked out of band.
### Console URL Discovery
If `--console-url` is not provided, the CLI discovers it from the `services`
@@ -202,9 +193,7 @@ table in the shared database. The console registers itself on startup.
The **TLS** tab in the console admin panel (System group) shows:
- CA status (common name, certificate count)
- Certificate table (domain, SANs, issued, expires)
- Force-renew for the console-owned internal identity; remote nodes renew and
hot-reload their own keys
- Delete for expired, remotely managed certificate rows
- Force-renew and delete actions per certificate
---
@@ -255,19 +244,9 @@ const client = new TurnstoneServer({
### Node Bootstrap Flow
1. Node starts, connects to shared database (plain connection)
2. Discovers the console URL from the `services` table — or honors an explicit
`TURNSTONE_CONSOLE_URL` (a bare-metal node outside the compose network can't
resolve the in-cluster `console` name, so it points this at the console's
published ACME endpoint)
3. Fetches the CA root from the configured console scheme. Direct deployments
use `http://console/acme/ca.pem` (plain HTTP, TOFU); an explicitly configured
HTTPS proxy is preserved and verified with the system trust store.
4. Requests a service cert via ACME with a dedicated, short-lived Turnstone
service JWT pinned to configured responder origins. lacme emits ACME JWS
messages, but its lightweight responder deliberately does not validate their
signatures or nonces; the service JWT is the enrollment authorization gate.
Direct HTTP bootstrap therefore still requires a trusted LAN/VPN (or an
independently trusted HTTPS proxy). The cert's
2. Discovers console URL from `services` table
3. Fetches CA root cert from `http://console/acme/ca.pem` (plain HTTP, TOFU)
4. Requests a service cert via ACME (plain HTTP, JWS-signed). The cert's
primary domain / SAN is the node's **advertised host** (the host of
`TURNSTONE_ADVERTISE_URL`, e.g. `node-1`) — the name peers actually dial,
not the container hostname. This makes mTLS hostname verification succeed
@@ -282,12 +261,7 @@ const client = new TurnstoneServer({
1. Read `tls.enabled` from ConfigStore
2. Initialize CA (load from DB or generate new root key)
3. Mount ACME responder at `/acme` (serves `/ca.pem` natively). When
`TURNSTONE_ACME_EXTERNAL_URL` is set, use it for every advertised directory,
order, authorization, and certificate URL; otherwise derive URLs from each
request as before. Directory, nonce, and CA bootstrap resources stay public;
account/order/challenge/finalization/certificate routes require the dedicated
enrollment service JWT.
3. Mount ACME responder at `/acme` (serves `/ca.pem` natively)
4. Issue console certs (internal + optional frontend)
5. Start CA-direct auto-renewal (no network, signs directly), scoped to the
console's own cert, plus a periodic GC that reclaims cert rows for
@@ -313,37 +287,11 @@ fronted under a second hostname). Symptom if this is wrong: the console
dashboard shows nodes as unreachable and `openssl s_client` reports the served
cert's SANs don't include the dialed name.
The advertised host and extra SANs may be DNS names or literal IPv4/IPv6
addresses. Turnstone converts IP literals to typed ACME identifiers so the
certificate contains `IPAddress` SANs that normal IP hostname verification can
use; DNS spelling is preserved. Bracket an IPv6 address when it appears in a URL
(for example `TURNSTONE_ADVERTISE_URL=http://[2001:db8::10]:8080`), but use the
bare address in `TURNSTONE_TLS_SANS`. Unspecified bind addresses (`0.0.0.0` and
`::`) and scoped IPv6 addresses such as `fe80::1%eth0` are not certificate
identities. Restart after changing the advertised identity or extra SANs.
### "No console service found"
The console registers itself in the `services` table on startup. If the console
hasn't started or the registration expired (1 hour TTL), nodes can't discover
it. Set `TURNSTONE_CONSOLE_URL` to a reachable console address (this is also how
a bare-metal node that can't resolve the in-cluster `console` name enrolls).
### Cross-host ACME links point at the container
For a node on another host, publish port 8090 on a reachable interface and set
`TURNSTONE_ACME_EXTERNAL_URL` on the console **and in-cluster nodes** to that full
responder base, including `/acme` (for example
`http://192.0.2.1:8090/acme`). The console advertises it; clients use it as a
trusted enrollment-token destination. Keep
`TURNSTONE_CONSOLE_URL=http://console:8090` for in-cluster service discovery.
A remote node whose `TURNSTONE_CONSOLE_URL` already names the public origin can
derive the same `/acme` base, but setting both values explicitly avoids drift.
Bind only a trusted LAN/VPN interface and firewall it to enrolling nodes. The
JWT authenticates the client, but a direct plain-HTTP bootstrap remains TOFU and
does not resist an active on-path attacker. If the network is untrusted, expose
the responder through an independently trusted HTTPS proxy instead.
it. Use `--console-url` explicitly.
### Browser HTTPS to the console
+36 -146
View File
@@ -1,10 +1,9 @@
# Tools Reference
Turnstone exposes a role-specific built-in tool surface plus any configured MCP
tools through provider-native or OpenAI-compatible function calling. Built-in
schemas live under `turnstone/tools/` and are loaded by
`turnstone/core/tools.py`; metadata selects the interactive, coordinator, and
task-agent subsets. MCP tools are discovered from configured servers by
turnstone exposes 16 built-in tools plus any number of external MCP tools to the
LLM via the OpenAI function-calling interface. Built-in tools are defined as JSON
files under `turnstone/tools/` and loaded at startup by `turnstone/core/tools.py`.
MCP tools are discovered from configured MCP servers at startup by
`turnstone/core/mcp_client.py`.
---
@@ -29,19 +28,13 @@ schema plus turnstone-specific metadata keys:
}
```
**Metadata keys** (stripped before sending the schema to the model; the full
set lives in `_META_KEYS` in `turnstone/core/tools.py`):
**Metadata keys** (stripped before sending the schema to the model):
| Key | Type | Meaning |
|------------------|------|---------|
| `task_agent` | bool | Tool is available to task sub-agents. |
| `coordinator` | bool | Tool is available to coordinator sessions. Without `interactive: true` alongside it, this reads as coord-only and the tool is stripped from interactive sessions. |
| `interactive` | bool | Opt a `coordinator: true` tool back into interactive sessions (dual-kind tools like `memory`). |
| `auto_approve` | bool | Tool runs without user confirmation (read-only, safe operations). |
| `primary_key` | str | When the model sends a bare string instead of JSON args, map it to this parameter name. |
| `kind_variants` | dict | Per-kind description / parameter-schema overlays so each session kind sees only the surface it can use (see `memory.json`). |
| `cwd_note` | str | Sentence appended to the description at session build time with `{working_dir}` substituted — declare on tools whose semantics depend on the process working directory (see `bash.json`, `apply_cwd_context`). |
| `workspace_note` | str | Companion sentence naming the operator-configured workspace directory, `{workspace_dir}` substituted; dropped when no workspace is configured. |
| Key | Type | Meaning |
|----------------|------|---------|
| `task_agent` | bool | Tool is available to task sub-agents. |
| `auto_approve` | bool | Tool runs without user confirmation (read-only, safe operations). |
| `primary_key` | str | When the model sends a bare string instead of JSON args, map it to this parameter name. |
---
@@ -51,10 +44,10 @@ set lives in `_META_KEYS` in `turnstone/core/tools.py`):
| Name | Description |
|---------------------|-------------|
| `TOOLS` | The complete loaded built-in union. Sessions send a kind-specific subset (`INTERACTIVE_TOOLS` or `COORDINATOR_TOOLS`). |
| `TOOLS` | All 28 loaded built-in tool definitions (interactive + coordinator union). Sessions send a kind-specific subset (`INTERACTIVE_TOOLS` or `COORDINATOR_TOOLS`). |
| `TASK_AGENT_TOOLS` | Tools with `task_agent: true` -- available to task sub-agents. Includes write operations. |
| `TASK_AUTO_TOOLS` | Set of all tool names with `auto_approve: true` -- used by task-agent sub-sessions to skip confirmation for matching available tools. |
| `BUILTIN_TOOL_NAMES`| Frozenset of the built-in union. Used by tool search to distinguish built-ins from deferrable MCP tools. |
| `BUILTIN_TOOL_NAMES`| Frozenset of all 28 built-in tool names (interactive + coordinator union). Used by tool search to distinguish always-on tools from deferrable MCP tools. |
| `PRIMARY_KEY_MAP` | Dict mapping tool name to its `primary_key` parameter name. |
---
@@ -63,10 +56,7 @@ set lives in `_META_KEYS` in `turnstone/core/tools.py`):
> See also: [Tool Pipeline diagram](diagrams/png/05-tool-pipeline.png)
Tool handling spans a four-phase pipeline. `ChatSession._execute_tools()` owns
prepare, approval, and execution (phases 13); after it returns, the owning
conversation loop guards the observed results and folds them into the
trajectory (phase 4).
Tool execution follows a three-phase pipeline inside `ChatSession._execute_tools()`:
### Phase 1: Prepare
@@ -75,8 +65,9 @@ trajectory (phase 4).
- Parses the JSON arguments (with fallback for malformed JSON).
- If JSON parsing fails entirely, uses `PRIMARY_KEY_MAP` to map a bare string
to the correct parameter.
- Dispatches to the matching `_prepare_{func_name}()` handler, the synthetic
`tool_search` fallback, or the generic `_prepare_mcp_tool()` handler.
- Dispatches to the matching `_prepare_{func_name}()` handler. There are 16
built-in tools plus `tool_search` (synthetic, client-side BM25 fallback) and
the generic `_prepare_mcp_tool()` handler for MCP tools.
- Validates arguments and builds a preview dict containing:
- `call_id`, `func_name`, `header`, `preview` (for display)
- `needs_approval` (bool)
@@ -85,10 +76,7 @@ trajectory (phase 4).
### Phase 2: Approve
Prepared items are sent to the UI via `ui.approve_tools(items)`. Several
parallel task agents may leave independent `ApprovalCycle` objects pending on
one workstream; each round owns a `cycle_id`, event, result, and verdict set.
Remote clients resolve the exact round by `cycle_id` (or a member `call_id`).
All prepared items are sent to the UI via `ui.approve_tools(items)`.
- The UI displays each tool's header and preview to the user.
- Items where `needs_approval` is `False` (auto-approved tools) are shown
@@ -100,10 +88,6 @@ Remote clients resolve the exact round by `cycle_id` (or a member `call_id`).
prompt). This is per-tool, not blanket.
- If `auto_approve` is `True` on the session (via `--skip-permissions` or workstream
template), all tools are approved automatically.
- When Smart Approvals are enabled, one immutable judge/settings snapshot is
stamped onto the whole batch. The batch auto-approves only when every gated
item has a qualifying verdict; partial or mixed qualification fails closed to
the human prompt. Stop is linearized against that terminal decision.
### Phase 3: Execute
@@ -123,28 +107,6 @@ Each item's `execute` callable is invoked:
denials are tracked separately. This removes the need for text-prefix heuristics.
Other tools deliver results atomically via
`ui.on_tool_result(call_id, name, output, is_error=...)` only.
Stop propagates to child model scopes, judges, tracked subprocess groups, and
the approval cycles owned by the cancelled operation. Calls that definitely
never started receive `EffectStatus.none`; an interrupted call whose external
outcome was not observed receives `unknown`, `partial`, or `rolled_back` as
appropriate. These typed receipts preserve effect truth across storage/replay
without exposing unreviewed model output as a tool result.
### Phase 4: Guard and atomic fold
After `_execute_tools()` returns, the main `send()` loop compacts/truncates
completed results to the remaining shared budget and then runs the heuristic
and optional LLM output guard. The task-agent loop deliberately guards the
observed raw output before applying its size cap, so truncation cannot hide a
sensitive result from that check.
After guard work, the owning loop rechecks generation ownership. On the main
conversation path, one generation-fenced commit appends the complete
tool-result block, advisories, feedback, and queued user turns; its durable
records run in FIFO order outside the lifecycle lock. A force-cancelled
predecessor can therefore finish external cleanup, but cannot fold late results
into its successor's trajectory.
---
## Tool Approval Flow
@@ -152,7 +114,7 @@ into its successor's trajectory.
**Auto-approved** (no user confirmation needed at runtime):
- `read_file` -- reads files, no side effects
- `search` -- grep-style search, no side effects
- `memory` -- structured persistent memory (save/get/search/delete/list)
- `memory` -- structured persistent memory (save/search/delete/list)
- `recall` -- searches conversation history
- `notify` -- sends notifications to linked channels (time-sensitive, auto-approved for urgency)
@@ -163,9 +125,6 @@ into its successor's trajectory.
- `web_fetch` -- fetches a URL (SSRF-protected, but makes network requests)
- `web_search` -- web search via self-hosted SearxNG (makes network requests)
- `task_agent` -- spawns an autonomous sub-agent
- `open_preview` -- **URL targets only** (network access, gated like `web_fetch`);
file-path and `attachment:` targets are local reads and run unprompted like
`read_file`
Note: The JSON schema metadata key `auto_approve` controls membership in
`TASK_AUTO_TOOLS` (used for task agent sub-sessions). The actual runtime
@@ -198,7 +157,6 @@ Every tool defines a `primary_key`. The mapping is:
| `search` | `query` |
| `web_fetch` | `url` |
| `web_search` | `query` |
| `open_preview` | `target` |
| `task_agent` | `prompt` |
| `memory` | `name` |
| `recall` | `query` |
@@ -327,7 +285,7 @@ Fetch a URL and extract specific information from it.
| `url` | string | yes | The URL to fetch (must start with `http://` or `https://`). |
| `question` | string | yes | What to extract or answer from the page content. |
- **What it does**: Fetches the URL, strips HTML to plain text, and uses the LLM to extract the answer to the question from the page content. Every redirect hop is SSRF-screened before it is requested. Private/internal addresses are refused by default; enable `tools.allow_private_network` (console Settings → Tools) to make them approvable for self-hosted setups whose services live on the local network — the approval prompt marks such requests, and a public site redirecting into private space is refused regardless. Cloud metadata endpoints and link-local, multicast and reserved addresses are refused even with the opt-in enabled, including as a redirect target from a private address you approved. An address is judged by what it actually reaches, so an IPv6 transition address (NAT64, 6to4, Teredo) wrapping an internal IPv4 is treated exactly as that IPv4 would be.
- **What it does**: Fetches the URL, strips HTML to plain text, and uses the LLM to extract the answer to the question from the page content. Protected against SSRF (blocks private/internal IPs).
- **Auto-approve**: No -- requires user confirmation (makes network requests).
- **Agent availability**: `task_agent`.
@@ -387,39 +345,6 @@ It reports the score scale, whether the endpoint cleanly separates relevant from
---
### open_preview
Show the user rich content in a preview pane beside the conversation.
| Parameter | Type | Required | Description |
|-----------|--------|----------|-------------|
| `target` | string | yes | An http(s) URL, a file path, or `attachment:<id>` for a file attached to the conversation. |
| `kind` | string | no | Rendering override: `web`, `pdf`, `image`, `table`, `text`, or `markdown`. Detected from the content when omitted. |
| `title` | string | no | Pane header title. Defaults to the page title, filename, or URL. |
- **What it does**: Resolves the target to bytes (URLs fetch through the same
SSRF-guarded path as `web_fetch`, screened per redirect hop, honoring the
same `tools.allow_private_network` opt-in), classifies the
content, stores it content-addressed against the workstream, and opens the
frontend preview pane beside the conversation: web pages render in a fully
sandboxed iframe (no scripts, opaque origin), PDFs in the browser viewer,
images inline, CSV/TSV/JSON as a sortable table, text/markdown rendered. A
previewed web page loads none of its remote images or styles by default, so
opening it never reveals the viewer to the page's site; a toggle in the pane
header turns remote content back on for that preview. The
model receives only a one-line confirmation — to reason about content, use
`web_fetch` / `read_file` instead. Preview content is size-capped per kind
(pages 4 MB, PDFs 32 MB, images 4 MB, tables 2 MB, text 512 KB) and GC'd
with the workstream.
- **Auto-approve**: URL targets require confirmation (network access); file
paths and `attachment:` targets run unprompted (local reads).
- **Agent availability**: interactive sessions only (not `task_agent`, not
coordinators).
- **Surfaces**: the pane renders in the web UI (standalone and console). The
CLI prints the confirmation line only — there is no terminal pane.
---
## Agent
The tool name uses the `_agent` suffix — bare `task` collides with
@@ -433,7 +358,7 @@ Delegate a general-purpose task to an autonomous sub-agent.
|-----------|--------|----------|-------------|
| `prompt` | string | yes | Complete task description for the sub-agent. |
- **What it does**: Spawns a sub-agent that inherits the `TASK_AGENT_TOOLS` set (read, write, edit, search, bash, and web tools). The sub-agent runs autonomously to completion. Use for work that requires file modifications or command execution.
- **What it does**: Spawns a sub-agent that inherits the `TASK_AGENT_TOOLS` set (read, write, edit, search, bash, web tools, memory tools). The sub-agent runs autonomously to completion. Use for work that requires file modifications or command execution.
- **Auto-approve**: No -- requires user confirmation.
- **Agent availability**: Top-level only.
@@ -447,26 +372,18 @@ Structured persistent memory across sessions with typed, scoped entries.
| Parameter | Type | Required | Description |
|---------------|---------|----------|-------------|
| `action` | string | yes | `save`, `get`, `search`, `delete`, or `list`. |
| `name` | string | save/get/delete | Short snake_case identifier for the memory. |
| `action` | string | yes | `save`, `search`, `delete`, or `list`. |
| `name` | string | save/delete | Short snake_case identifier for the memory. |
| `content` | string | save | Memory content to store. |
| `description` | string | save | Non-empty description for relevance matching; required on create and update. |
| `type` | string | no | Memory type: `user`, `general`, `feedback`, or `reference`. Default: `general`. |
| `scope` | string | no | Memory scope: `global`, `workstream`, `user`, `coordinator`, or `project`. See defaults below. |
| `description` | string | no | Short description for relevance matching (recommended for `save`). |
| `type` | string | no | Memory type: `user`, `project`, `feedback`, or `reference`. Default: `project`. |
| `scope` | string | no | Memory scope: `global`, `workstream`, or `user`. Default: `global`. |
| `query` | string | search | Search query for finding memories. |
| `limit` | integer | no | Max results for `search` or `list`. Default: 20. |
- **What it does**: Manages structured persistent memories in the database.
Memories persist across sessions, have a type classification, and live in a
role-specific visible scope. Unscoped `save`/`get`/`delete` resolve to one
target: the attached active project, otherwise `global` for an interactive
session or `coordinator` for a coordinator. Read-only project access permits
`get` but makes `save`/`delete` fail without falling back. A valid explicit
scope selects exactly that scope. Unscoped `search`/`list` cover all visible
scopes; use the displayed scope when following a result with `get` or
`delete`.
- **What it does**: Manages structured persistent memories in the database. Memories persist across sessions, have a type classification (user preferences, project knowledge, feedback, reference material) and a scope (global across all workstreams, private to a workstream, or following a user). Relevant memories are included in the system prompt on startup.
- **Auto-approve**: Yes.
- **Agent availability**: Not available to task agents.
- **Agent availability**: Not available to sub-agents (top-level only).
---
@@ -481,7 +398,7 @@ Search conversation history for past messages and tool results.
- **What it does**: Searches conversation history across sessions using FTS5 full-text search. Returns matching messages, tool calls, and tool results with timestamps and workstream context.
- **Auto-approve**: Yes.
- **Agent availability**: Not available to task agents.
- **Agent availability**: Not available to sub-agents (top-level only).
---
@@ -617,11 +534,7 @@ pre-configure skills at workstream creation.
---
## Interactive Tool Summary
This table describes the ordinary interactive surface. Coordinator sessions
receive their delegation/lifecycle tools instead, and task agents receive the
metadata-selected `TASK_AGENT_TOOLS` subset.
## Summary Table
| Tool | Category | Auto-approve | task_agent | primary_key |
|--------------|------------|--------------|------------|-------------|
@@ -632,7 +545,6 @@ metadata-selected `TASK_AGENT_TOOLS` subset.
| `search` | File Ops | Yes | Yes | `query` |
| `web_fetch` | Info | No | Yes | `url` |
| `web_search` | Info | No | Yes | `query` |
| `open_preview`| Info | URL: no; path/attachment: yes | No | `target` |
| `task_agent` | Agent | No | No | `prompt` |
| `memory` | Memory | Yes | No | `name` |
| `recall` | Memory | Yes | No | `query` |
@@ -668,11 +580,6 @@ Tool search uses the best available mechanism for each provider:
`_exec_tool_search()` runs a pure-Python BM25 index over tool names and
descriptions, then expands the matched tools into the visible set.
A persona with a tool-visibility set overrides this selection: any exact
set forces tool search into the client-side BM25 mechanism (tier 3)
regardless of provider, and a **hard** set — one whose visible tools omit
`tool_search` — disables tool search entirely.
### Configuration
Tool search is configured in `config.toml` under the `[tools]` section:
@@ -742,7 +649,7 @@ MCP-compatible service.
3. **Schema conversion**: Each MCP tool's `inputSchema` is converted to OpenAI
function-calling format. The tool name is prefixed: `mcp__{server}__{tool}`.
4. **Merging**: MCP tools are appended after the role's built-in tools via
4. **Merging**: MCP tools are appended after the 16 built-in tools via
`merge_mcp_tools()`. Built-in tools appear first, giving them natural LLM priority.
When dynamic tool search is active, MCP tools are deferred rather than directly
visible -- the model discovers them via search as needed (see
@@ -829,10 +736,7 @@ MCP tool lists stay up-to-date without restart through two mechanisms:
1. **Push notifications** -- MCP servers that declare `tools.listChanged: true` in
their capabilities send `notifications/tools/list_changed` when their tool list
changes. `MCPClientManager` registers a `message_handler` on each `ClientSession`
that triggers an immediate refresh for that server (debounced per server and
notification kind, and run off the receive loop). A refresh that fails while
the connection stays up is retried automatically on the next health-loop tick
until one completes.
that triggers an immediate refresh for that server.
2. **Manual** -- `/mcp refresh` re-fetches tools from all servers immediately.
`/mcp refresh <server>` targets a single server. If a server has disconnected,
@@ -840,10 +744,6 @@ MCP tool lists stay up-to-date without restart through two mechanisms:
same controls (refresh / reconnect buttons per server) for cluster-wide
fan-out.
Reconnects (health-loop, dispatch-driven, or operator-forced) always end in a
full catalog rediscovery, so a server that changed its tools while disconnected
comes back current.
When tools change, `MCPClientManager` rebuilds its merged tool list using copy-on-write
(new list/dict objects assigned atomically) and notifies all active `ChatSession`
instances via registered listener callbacks. Each session rebuilds its `_tools`,
@@ -879,13 +779,6 @@ capabilities for the `resources` capability. For servers that declare it:
2. `list_resource_templates` fetches URI templates (parameterized patterns like
`db://tables/{table}/rows/{id}`).
The protocol advertises both lists through one aggregate `resources`
capability, so a server may implement only one of them. If either request
returns the JSON-RPC `Method not found` code (`-32601`), turnstone treats that
half of the catalog as empty and keeps the other half; authentication,
validation, transport, and all other discovery errors still fail the
connection or refresh.
Both are stored as `{uri, name, description, mimeType, server}` dicts and
merged into a unified catalog.
@@ -921,16 +814,13 @@ catalog.
### Refresh
Resource lists stay current through the same mechanisms as tool lists:
Resource lists stay current through the same three-tier mechanism as tool lists:
1. **Push** -- Servers declaring `resources.listChanged: true` send
`notifications/resources/list_changed`, triggering an immediate refresh
(with the same failed-refresh retry on the health-loop tick).
2. **Manual** -- `/mcp refresh` re-fetches resources alongside tools.
Servers without push support are refreshed whenever they reconnect (every
reconnect ends in full rediscovery) or when an operator refreshes manually;
there is no periodic polling.
`notifications/resources/list_changed`, triggering an immediate refresh.
2. **Periodic** -- Servers without push are polled on the configured refresh
interval (default 4 hours, same timer as tools).
3. **Manual** -- `/mcp refresh` re-fetches resources alongside tools.
---
-56
View File
@@ -1,56 +0,0 @@
{
"defaults": {
"n_runs": 3
},
"cases": [
{
"id": "search-first",
"skill": {
"name": "search-first",
"content": "# Search First\n\nBefore answering ANY question about where something lives in the codebase, you MUST call the `search` tool first. Never answer from memory."
},
"user_prompt": "Where is JWT token validation implemented in this project?",
"expected_actions": [{ "tool": "search" }],
"match_mode": "ordered_subset",
"max_turns": 4
},
{
"id": "test-after-edit",
"skill": {
"name": "test-after-edit",
"content": "# Test After Edit\n\nAfter editing or writing ANY file, you MUST run the test suite with `python -m pytest` via bash before you finish. Do not report done until tests have run."
},
"user_prompt": "Add a function `clamp(x, lo, hi)` that clamps x to [lo, hi] in utils.py.",
"setup": {
"files": {
"utils.py": ""
}
},
"expected_actions": [
{ "tool": "write_file" },
{ "tool": "bash", "args_pattern": { "command": "pytest" } }
],
"match_mode": "ordered_subset",
"max_turns": 8
},
{
"id": "changelog-update",
"skill": {
"name": "changelog-update",
"content": "# Changelog Discipline\n\nWhenever you modify a file, you MUST also append a one-line entry to CHANGELOG.md describing the change in the same task."
},
"user_prompt": "Fix the off-by-one so pager.py shows the last page. Edit pager.py.",
"setup": {
"files": {
"pager.py": "def last_page(total_items, per_page):\n # off-by-one: drops the final partial page\n return total_items // per_page\n",
"CHANGELOG.md": "# Changelog\n"
}
},
"expected_actions": [
{ "tool": "edit_file", "args_pattern": { "path": "CHANGELOG.md" } }
],
"match_mode": "subset",
"max_turns": 8
}
]
}
-9
View File
@@ -1,9 +0,0 @@
__pycache__/
.venv/
*.db
*.db-wal
*.db-shm
.ruff_cache/
.pytest_cache/
.mypy_cache/
uv.lock
-282
View File
@@ -1,282 +0,0 @@
# Understone
A small, multiplayer, BBS-style **ANSI door game** served over the Model
Context Protocol (MCP). It is a text RPG in the spirit of *Legend of the Red
Dragon* — explore an overworld of box-drawing maps, fight wandering monsters,
shop and rest in town, and descend a dungeon — except the "door" is an MCP
server and the player drives it by talking to an AI assistant.
The server is the rules engine and the single source of truth. Players share
**one persistent world**: your assistant calls tools, the server returns
authoritative frames and facts, and the assistant narrates the story around
them.
This is a self-contained reference example. It depends only on `mcp` — there
is no dependency on Turnstone itself — so it runs against any MCP client.
## How to play
There is **no prompt to paste and no persona to configure**. The tool schema
is the whole interface. Once the server is registered with your assistant:
1. Tell your assistant you'd like to play an ANSI door game / text dungeon
RPG (it can discover the tools by name and description).
2. The assistant calls `door_help` to learn how to run the world, then
`door_join` with your adventurer's name.
3. Play unfolds as a conversation: "head east", "fight it", "rest at the inn".
Everything the assistant needs to run the game well is returned by
`door_help`.
## Gameplay
A run is a little RPG loop, played a bit each day:
- **Explore** the overworld of box-drawing maps. Walking is free, but the wild
country has texture — a step may turn up a wandering monster, a purse of
gold, a healing spring, a small trap (which can never kill you), or a scrap
of old Vale lore. Only one such find happens per move, and the non-combat
ones don't interrupt your walk.
- **Fight, shop, and heal** in and around town. Fighting and descending one
rung of the dungeon each spend one of your daily turns; resting, shopping and
moving do not.
- **Delve the deep, a rung at a time.** The dungeon is a ladder of guardians:
each `descend` faces the next one past your deepest and either advances your
depth or bounces you home (your depth persists either way). Carry a few
**potions in your satchel**`quaff` the strongest when you choose, and if a
fight would kill you the satchel saves you automatically, the elixir burning
down your throat at death's edge. Clearing a rung also yields **forge ore**,
which rides the satchel (a won forest fight sometimes turns up a little, too).
- **Forge an edge — with gold AND ore.** At the shop's **forge** you can add a
+1 edge to your equipped weapon or armour, up to a cap, each step dearer than
the last. A step costs gold *and* the ore you won in the deep — so the forge is
fed by descending, not just by a fat purse. Watch, too, for the **rare beasts**
that prowl the forest: felling one is Herald news and always drops a draught.
- **Win the game** by slaying **the Wyrm Below**. Once your hero is seasoned
enough AND has plumbed the deep to its floor, `challenge` it at the dungeon. A
victory frees the Vale, carves your run into the **Hall of Legends**, and — in
the tradition of the classic BBS door games — begins a new life: your
character resets to first-day gear and stats but keeps a permanent ★ for every
Wyrm slain, ready to do it all again.
- **Read the news.** `door_log` is the **Understone Herald**, a shared
broadsheet of notable deeds across the whole world — who joined, who rose a
level, who was dragged home by a goblin, and who freed the Vale.
- **Make it social.** It is a shared world, so you can touch other players.
`ambush` a rival who has not yet acted today — a classic
style player-kill that robs a sleeping foe of some gold, except the surest
defence is simply to take your own turn (an active player is awake and can't
be caught). Lose the ambush and *you* are the one who flees, shamed on the
feed. `post` a private note another player reads on their next visit (it
never reaches the public Herald). Or `gamble` a little gold at the inn's dice
against the house. Ambush spends a turn; mail and dice do not.
- **Bank your coin.** The inn keeps a strongbox: `deposit` gold into the
**vault** and `withdraw` it later (no turn either way). Banked gold is **safe
from ambush** — a sleeping-robber only ever lifts what you carry — and it is
the one thing that **survives a Wyrm-win reset**, carrying wealth across runs.
## Installation
This example uses [`uv`](https://docs.astral.sh/uv/). From the example
directory:
```bash
cd examples/door-game
uv venv
uv pip install -e .
```
That installs the `understone` entry point into the environment.
To run the tests and quality gates:
```bash
uv pip install -e ".[test,dev]"
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run mypy understone/
```
## Running the server
By default the server speaks the **stdio** transport, which is how MCP clients
launch a per-session subprocess:
```bash
understone
```
To host one shared world over HTTP for several clients, run the
**streamable-http** transport as a single long-lived process:
```bash
UNDERSTONE_TRANSPORT=streamable-http understone
```
### Environment variables
| Variable | Default | Description |
|----------|---------|-------------|
| `UNDERSTONE_DB` | `./understone.db` | SQLite database file for the world's state. |
| `UNDERSTONE_WORLD` | _(packaged pack)_ | Directory of a content pack to load instead of the bundled Vale of Understone. |
| `UNDERSTONE_TRANSPORT` | `stdio` | `stdio` or `streamable-http`. |
| `UNDERSTONE_HOST` | `127.0.0.1` | Bind host (streamable-http only). |
| `UNDERSTONE_PORT` | `8077` | Bind port (streamable-http only). |
| `UNDERSTONE_PATH` | `/mcp` | HTTP path for the MCP endpoint (streamable-http only). |
## The Watch — a live spectator view
When the server runs under the **streamable-http** transport, it also serves a
read-only **Watch** page: the lobby TV of the Vale. Point a browser at
```
http://127.0.0.1:8077/watch
```
(the host and port follow `UNDERSTONE_HOST` / `UNDERSTONE_PORT`). It is a
period **CRT spectator console** — a green-and-amber phosphor map of the whole
world with every adventurer's `☻` marker, a live **Understone Herald** feed, the
**Hall of Legends**, and a roster of who is currently abroad. It refreshes every
couple of seconds; if it loses contact it dims and reads `SIGNAL LOST` until the
server returns. The console's palette follows the pack: a world may pick its own
CRT colour with `settings.watch_theme` (`phosphor` green, `amber` gold, `ice`
blue, `ember` red), defaulting to the Vale's green if it says nothing.
The Watch is **strictly read-only**. Input never flows through it — there are no
controls, no forms, nothing that can change the world. It reads the same shared
state the tools do and paints it; that is all. There is no authentication, in
keeping with the rest of this easter-egg server (see the safety note below), so
treat the page as you would the MCP endpoint itself.
> _Screenshot: the Watch console — a phosphor-green overworld map with amber
> `☻` markers, the Herald feed and Hall of Legends down the right-hand rail.
> (Image placeholder; run the server and open the URL to see it live.)_
When the Watch is up, the `door_join` welcome and the `door_help` manual both
print its URL so players (and the assistant narrating for them) know it exists.
If you bind to `0.0.0.0` to share the world across a network, advertise a host
that browsers can actually reach (your machine's LAN address or hostname) rather
than `0.0.0.0` itself — the link is composed from `UNDERSTONE_HOST`.
## Authoring worlds
The Vale of Understone is just the *bundled* world. The whole game — its map,
monsters, economy, and endgame — is a **content pack**: a directory of six JSON
files the server loads at start. Nothing about the Vale is privileged; point
the server at another pack and it runs that world instead. This is the seam
where the game becomes its own authoring target: a pack is plain data, so a
person *or an LLM* can write one, and the same zero-setup philosophy that makes
the game playable with no prompt makes it **authorable with no code**.
The loop has these commands:
```bash
understone newpack mypack # scaffold a pack (copies the Vale as a template)
# ...edit or LLM-generate the JSON in mypack/ to describe your world...
understone validate mypack # check it; prints a report or names what's wrong
understone simulate mypack # play a greedy bot through it and measure the balance
UNDERSTONE_WORLD=mypack understone # serve your world
understone worlds # list the bundled worlds and whether each is sound
```
`newpack` writes a starting template plus an `AUTHORING.md` manual — the
file-by-file schema, the enforced limits, and design guidance — written to be
followed cold by a model. `validate` loads the pack through exactly the same
hardened loader the server uses and either prints a summary ending **"This pack
is sound. The door stands open."** or fails with one precise line naming the
file, the row, and the field at fault.
`simulate` is the **balance instrument**: it drives a deliberately simple,
greedy bot through the *real* game — the same `join`/`move`/`action` calls the
tools make — over a seeded RNG and an injected clock, then prints a report
(final level, gold earned, fights fought, rungs cleared, whether and when the
Wyrm fell). It is a tuning probe, not a player to admire: it answers "is this
world *shaped* right, and is it *winnable*?". Pass `--days N`, `--seed S`, or
`--seeds K` for a multi-seed sweep with means and spreads. `worlds` lists every
bundled world — the default Vale plus any alternate packs shipped under
`understone/world/packs/` — loading each so it can report it as sound or flawed.
**A second bundled world: The Cinder Wastes.** Understone ships a second world
alongside the Vale, in `understone/world/packs/cinder-wastes/` — a volcanic
ash-and-slag map whose Watch page glows ember-red instead of the Vale's green
phosphor. It is the pipeline's own dogfood: it was authored **by an LLM working
only from `AUTHORING.md` and the `validate` loop**, with no engine code touched,
then bundled verbatim. `understone worlds` lists it as sound, and
`understone simulate understone/world/packs/cinder-wastes --days 50 --seeds 3`
shows the greedy bot taking its Magma Wyrm — the end-to-end proof that a world
described purely as data, from the manual alone, is genuinely playable to
victory. Serve it with
`UNDERSTONE_WORLD=understone/world/packs/cinder-wastes understone`.
Packs are validated **hard** at load: every map glyph must render as exactly
one terminal column (no fullwidth runes, no emoji, no combining marks — the
frames are box-drawing rectangles) and may not collide with the frame's
box-drawing lines or the player markers, dimensions and counts are bounded,
display names are length-checked, and every cross-reference (a legend
character, a starting item, the boss monster, a dungeon tier) must resolve. The
loader also pins the rules that keep the endgame coherent: a world has exactly
one boss, and a dungeon tier's lead monster (its fixed rung guardian) may not be
a rare. Because packs are now routinely untrusted, generated output, those error
messages are not a nuisance — they are the **feedback loop**. Iterate against
them until the door stands open.
## Registering with Turnstone
Understone is an ordinary MCP server, so it plugs into Turnstone's MCP client
config two ways.
**Stdio (per-session subprocess).** Turnstone launches the `understone`
command for each session. Each session gets its own subprocess, so for a
truly shared world prefer the HTTP form below; stdio is simplest for solo
play.
```toml
[mcp.servers.understone]
command = "understone"
[mcp.servers.understone.env]
UNDERSTONE_DB = "/var/lib/understone/world.db"
```
**Streamable-HTTP (one shared world).** Run a single Understone process with
`UNDERSTONE_TRANSPORT=streamable-http` and point every client at its URL. This
is the right setup for multiplayer: one process, one database, one world that
all adventurers share.
```toml
[mcp.servers.understone]
url = "http://localhost:8077/mcp"
```
> **Operator note.** For multiplayer, start exactly one shared process —
> `UNDERSTONE_TRANSPORT=streamable-http understone` — and have all clients use
> the url form. The world lives in a single SQLite file written by that one
> process.
## The tools
| Tool | What it does |
|------|--------------|
| `door_help` | The game-master manual. Start here. |
| `door_join` | Create or resume an adventurer; returns the opening map. |
| `door_status` | The character sheet (read-only). |
| `door_look` | Redraw the current view — overworld map or location menu. |
| `door_move` | Walk the overworld (free; no daily turn spent). |
| `door_action` | Context verbs: fight, flee, ambush (a rival), rest, deposit/withdraw (the inn vault), buy, sell, forge (a +1 edge, gold + ore), heal, gamble (inn dice), descend (one rung), challenge (the Wyrm), post (mail another player), quaff (a carried potion), leave. |
| `door_log` | The Understone Herald — the shared feed of notable deeds. |
| `door_rank` | The leaderboard, plus the Hall of Legends (★ marks Wyrm kills). |
| `door_bestow` | Game-master grant of a little gold/healing for a story beat. |
## A note on identity and safety
This example is an **easter egg**, not a hardened service. Identity is
**self-asserted**: a "player" is just a name passed to the tools, and there is
**no authentication** — anyone who can reach the server can act as any name.
That is fine for a shared toy world among people who trust each other, and
deliberately out of scope for a game. Do not store anything sensitive in it,
and if you expose the HTTP transport beyond localhost, put it behind whatever
access control your environment already provides.
The game master's `door_bestow` channel can only grant small, capped amounts
of in-game gold and healing — never items, never turns — and every grant is
written to the public in-world log, so its reach is bounded by design.
-55
View File
@@ -1,55 +0,0 @@
[build-system]
requires = ["hatchling>=1.29"]
build-backend = "hatchling.build"
[project]
name = "understone"
version = "0.10.0"
description = "Understone — a BBS-style ANSI door game served over MCP."
requires-python = ">=3.11"
license = "Apache-2.0"
dependencies = [
"mcp>=1.27,<2",
]
[project.scripts]
understone = "understone.server:main"
[project.optional-dependencies]
test = ["pytest>=9.0"]
dev = ["ruff>=0.9", "mypy>=1.14"]
[tool.hatch.build.targets.wheel]
packages = ["understone"]
[tool.pytest.ini_options]
testpaths = ["tests"]
[tool.ruff]
target-version = "py311"
line-length = 100
[tool.ruff.lint]
select = ["E", "F", "W", "I", "N", "UP", "B", "A", "SIM", "TCH"]
ignore = ["E501"]
[tool.ruff.format]
quote-style = "double"
[tool.mypy]
python_version = "3.11"
strict = true
warn_return_any = true
warn_unused_configs = true
disallow_untyped_defs = true
disallow_incomplete_defs = true
check_untyped_defs = true
no_implicit_optional = true
[[tool.mypy.overrides]]
module = ["mcp", "mcp.*"]
ignore_missing_imports = true
[[tool.mypy.overrides]]
module = "tests.*"
disallow_untyped_defs = false
-256
View File
@@ -1,256 +0,0 @@
"""Shared test fixtures and builders.
These builders construct engine objects directly (no JSON loader) so the
engine tests stay independent of the content pack. Later chunks add
fixtures that load the shipped world and build the game façade.
"""
from __future__ import annotations
from collections import Counter
from datetime import UTC, datetime
from typing import TYPE_CHECKING
import pytest
from understone.engine.models import (
Item,
LocationDef,
Mode,
Monster,
Player,
Settings,
Slot,
TerrainDef,
WorldEvent,
Zone,
)
from understone.engine.world import World
if TYPE_CHECKING:
from collections.abc import Callable
from understone.game import Game
# ---------------------------------------------------------------------------
# Terrain kinds for synthetic test worlds
# ---------------------------------------------------------------------------
GRASS = TerrainDef(key="grass", glyph=".", walkable=True, encounter_rate=0.0, color="floor")
WALL = TerrainDef(key="wall", glyph="", walkable=False, encounter_rate=0.0, color="wall")
WATER = TerrainDef(key="water", glyph="~", walkable=False, encounter_rate=0.0, color="water")
FOREST = TerrainDef(key="forest", glyph="", walkable=True, encounter_rate=1.0, color="tree")
SAFE_FOREST = TerrainDef(key="forest", glyph="", walkable=True, encounter_rate=0.0, color="tree")
DEFAULT_SETTINGS = Settings(
daily_turns=10,
rest_cost=15,
heal_cost_per_hp=2,
starting_gold=20,
starting_weapon="rusty_dagger",
starting_armor="cloth_tunic",
start_hp=20,
start_atk=3,
start_def=0,
xp_base=100,
growth_max_hp=6,
growth_atk=2,
growth_def=1,
bestow_daily_budget=25,
dungeon_tiers=(4, 5),
boss_monster="wyrm_below",
wyrm_min_level=6,
ambush_min_level=3,
ambush_level_band=2,
ambush_gold_pct=25,
post_daily_cap=5,
gamble_max_bet=50,
gamble_daily_cap=5,
satchel_max=3,
forge_base_cost=60,
forge_max_plus=3,
rare_drop_item="minor_potion",
forge_ore_item="iron_ore",
forge_ore_per_plus=1,
ore_dungeon_drop=2,
ore_forest_chance=0.2,
watch_theme="phosphor",
)
def make_settings(**overrides: object) -> Settings:
"""Return DEFAULT_SETTINGS with field overrides for band testing."""
base = {
"daily_turns": DEFAULT_SETTINGS.daily_turns,
"rest_cost": DEFAULT_SETTINGS.rest_cost,
"heal_cost_per_hp": DEFAULT_SETTINGS.heal_cost_per_hp,
"starting_gold": DEFAULT_SETTINGS.starting_gold,
"starting_weapon": DEFAULT_SETTINGS.starting_weapon,
"starting_armor": DEFAULT_SETTINGS.starting_armor,
"start_hp": DEFAULT_SETTINGS.start_hp,
"start_atk": DEFAULT_SETTINGS.start_atk,
"start_def": DEFAULT_SETTINGS.start_def,
"xp_base": DEFAULT_SETTINGS.xp_base,
"growth_max_hp": DEFAULT_SETTINGS.growth_max_hp,
"growth_atk": DEFAULT_SETTINGS.growth_atk,
"growth_def": DEFAULT_SETTINGS.growth_def,
"bestow_daily_budget": DEFAULT_SETTINGS.bestow_daily_budget,
"dungeon_tiers": DEFAULT_SETTINGS.dungeon_tiers,
"boss_monster": DEFAULT_SETTINGS.boss_monster,
"wyrm_min_level": DEFAULT_SETTINGS.wyrm_min_level,
"ambush_min_level": DEFAULT_SETTINGS.ambush_min_level,
"ambush_level_band": DEFAULT_SETTINGS.ambush_level_band,
"ambush_gold_pct": DEFAULT_SETTINGS.ambush_gold_pct,
"post_daily_cap": DEFAULT_SETTINGS.post_daily_cap,
"gamble_max_bet": DEFAULT_SETTINGS.gamble_max_bet,
"gamble_daily_cap": DEFAULT_SETTINGS.gamble_daily_cap,
"satchel_max": DEFAULT_SETTINGS.satchel_max,
"forge_base_cost": DEFAULT_SETTINGS.forge_base_cost,
"forge_max_plus": DEFAULT_SETTINGS.forge_max_plus,
"rare_drop_item": DEFAULT_SETTINGS.rare_drop_item,
"forge_ore_item": DEFAULT_SETTINGS.forge_ore_item,
"forge_ore_per_plus": DEFAULT_SETTINGS.forge_ore_per_plus,
"ore_dungeon_drop": DEFAULT_SETTINGS.ore_dungeon_drop,
"ore_forest_chance": DEFAULT_SETTINGS.ore_forest_chance,
"watch_theme": DEFAULT_SETTINGS.watch_theme,
}
base.update(overrides)
return Settings(**base) # type: ignore[arg-type]
def make_player(**overrides: object) -> Player:
"""Build a Player at sane defaults; override any field by keyword."""
fields = {
"name": "Tester",
"x": 5,
"y": 5,
"hp": 20,
"max_hp": 20,
"level": 1,
"xp": 0,
"gold": 50,
"atk": 5,
"def_": 1,
"weapon_id": "rusty_dagger",
"armor_id": "cloth_tunic",
"turns_left": 10,
"turn_day": 0,
"mode": Mode.TILE,
"at_location": "",
"created_at": "2026-01-01T00:00:00+00:00",
"last_seen": "2026-01-01T00:00:00+00:00",
"log_cursor": 0,
"bestow_spent": 0,
"bestow_day": 0,
"wins": 0,
"posts_sent": 0,
"post_day": 0,
"gambles": 0,
"gamble_day": 0,
}
fields.update(overrides)
return Player(**fields) # type: ignore[arg-type]
def make_monster(**overrides: object) -> Monster:
"""Build a Monster at tier-1 defaults."""
fields = {
"tier": 1,
"name": "Field Rat",
"hp": 6,
"atk": 3,
"def_": 0,
"xp": 8,
"gold": 3,
"monster_id": "",
"boss": False,
}
fields.update(overrides)
return Monster(**fields) # type: ignore[arg-type]
def make_world(
*,
grid: list[list[TerrainDef]] | None = None,
width: int = 11,
height: int = 11,
spawn: tuple[int, int] = (5, 5),
locations: list[LocationDef] | None = None,
zones: list[Zone] | None = None,
monsters: list[Monster] | None = None,
items: list[Item] | None = None,
settings: Settings | None = None,
events: list[WorldEvent] | None = None,
) -> World:
"""Build a small synthetic World (all-grass by default)."""
if grid is None:
grid = [[GRASS for _ in range(width)] for _ in range(height)]
return World(
name="Test Vale",
width=width,
height=height,
spawn=spawn,
terrain=grid,
locations=locations or [],
zones=zones or [],
monsters=monsters or [make_monster()],
items=items or _default_items(),
settings=settings or DEFAULT_SETTINGS,
events=events,
)
def _default_items() -> list[Item]:
return [
Item("rusty_dagger", "Rusty Dagger", Slot.WEAPON, 2, 0, 0, 0),
Item("short_sword", "Short Sword", Slot.WEAPON, 5, 0, 0, 40),
Item("cloth_tunic", "Cloth Tunic", Slot.ARMOR, 0, 1, 0, 0),
Item("leather_armor", "Leather Armor", Slot.ARMOR, 0, 3, 0, 50),
Item("minor_potion", "Minor Potion", Slot.CONSUMABLE, 0, 0, 15, 12),
Item("iron_ore", "Iron Ore", Slot.MATERIAL, 0, 0, 0, 0),
]
def fixed_clock(moment: datetime) -> Callable[[], datetime]:
"""Return a clock callable that always reports *moment*."""
def _clock() -> datetime:
return moment
return _clock
def utc(year: int, month: int, day: int, hour: int = 0, minute: int = 0) -> datetime:
"""Construct a tz-aware UTC datetime."""
return datetime(year, month, day, hour, minute, tzinfo=UTC)
# ---------------------------------------------------------------------------
# Satchel test helpers (the v0.10 stack encoding)
# ---------------------------------------------------------------------------
# The satchel is stack-based ("id:qty"); these wrap the game façade's stack
# helpers so a test can seed/read a bag as a flat id list (duplicate ids
# collapse to one stack), keeping the assertions readable. Shared by the
# descend and Wyrm suites.
def set_satchel(game: Game, player: object, ids: list[str]) -> None:
"""Seed *player*'s satchel from a flat id list (duplicates -> one stack qty)."""
counts = Counter(ids)
stacks = [(item_id, counts[item_id]) for item_id in dict.fromkeys(ids)]
game._satchel_set_stacks(player, stacks) # type: ignore[arg-type]
def satchel_ids(game: Game, player: object) -> list[str]:
"""Return the satchel as a flat id list, each stack expanded by its qty."""
out: list[str] = []
for item_id, qty in game._satchel_stacks(player): # type: ignore[arg-type]
out.extend([item_id] * qty)
return out
@pytest.fixture
def small_world() -> World:
"""An 11x11 all-grass world with the default content tables."""
return make_world()
@@ -1,7 +0,0 @@
┌── The Sleeping Drake ───┐
│ A warm hearth crackles. │
│ A bed costs 15 gold. │
│ │
│ (R)est (L)eave │
└─────────────────────────┘
[ status ]
@@ -1,8 +0,0 @@
┌─ Vale ──┐
│@........│
│.........│
│.........│
│.........│
│.........│
└─────────┘
[ status ]
@@ -1,8 +0,0 @@
┌─ Vale ──┐
│.........│
│.........│
│....@....│
│.........│
│.........│
└─────────┘
[ status ]
-374
View File
@@ -1,374 +0,0 @@
"""Tests for the pack-authoring command surface.
Covers the validate/newpack functions directly (sound and broken packs, the
scaffold round-trip, AUTHORING.md generation from the live loader bands, and
the refuse-non-empty guard), the ``server.main`` argv dispatch (validate routes
through and bare invocation still reaches serve without binding a port), and
one end-to-end subprocess smoke of ``python -m understone validate``.
"""
from __future__ import annotations
import json
import shutil
import subprocess
import sys
from io import StringIO
from pathlib import Path
from typing import TYPE_CHECKING, Any
import pytest
from understone import cli, server
from understone.world import loader
if TYPE_CHECKING:
from collections.abc import Callable
EXAMPLE_DIR = Path(__file__).resolve().parents[1]
SHIPPED = EXAMPLE_DIR / "understone" / "world" / "data"
# The six content files a scaffolded pack must carry, plus the manual.
_PACK_JSONS = {
"terrain.json",
"monsters.json",
"items.json",
"locations.json",
"events.json",
"world.json",
}
# ---------------------------------------------------------------------------
# cli_validate
# ---------------------------------------------------------------------------
def test_cli_validate_sound_pack_reports_and_returns_zero() -> None:
out, err = StringIO(), StringIO()
rc = cli.cli_validate(SHIPPED, out=out, err=err)
assert rc == 0
report = out.getvalue()
assert "This pack is sound. The door stands open." in report
# The report surfaces the headline facts the brief calls for.
assert "The Vale of Understone" in report
assert "96x48" in report
assert "1 boss" in report
assert "% fight" in report
assert err.getvalue() == ""
def test_cli_validate_broken_pack_names_field_and_returns_two(tmp_path: Path) -> None:
# A pack whose daily_turns is out of band: the loader names the field.
pack = _clone_shipped(tmp_path)
_patch_world(pack, _break_daily_turns)
out, err = StringIO(), StringIO()
rc = cli.cli_validate(pack, out=out, err=err)
assert rc == 2
message = err.getvalue()
assert message.startswith("The pack is flawed:")
assert "daily_turns" in message # the offending field is named
assert out.getvalue() == ""
def test_cli_validate_missing_directory_returns_two(tmp_path: Path) -> None:
out, err = StringIO(), StringIO()
rc = cli.cli_validate(tmp_path / "nope", out=out, err=err)
assert rc == 2
assert "The pack is flawed:" in err.getvalue()
# ---------------------------------------------------------------------------
# cli_newpack
# ---------------------------------------------------------------------------
def test_cli_newpack_writes_template_and_manual(tmp_path: Path) -> None:
dest = tmp_path / "mypack"
out, err = StringIO(), StringIO()
rc = cli.cli_newpack(dest, out=out, err=err)
assert rc == 0
present = {p.name for p in dest.iterdir()}
assert present >= _PACK_JSONS # the six content files are all there
assert "AUTHORING.md" in present
# Next-steps guidance points the author at the validate verb.
assert "understone validate" in out.getvalue()
def test_cli_newpack_scaffold_validates(tmp_path: Path) -> None:
"""The load-bearing test: a freshly scaffolded pack loads cleanly.
newpack -> load_world round-trip. If the template the scaffolder copies
ever drifts out of the loader's bands, this fails immediately.
"""
dest = tmp_path / "mypack"
assert cli.cli_newpack(dest, out=StringIO(), err=StringIO()) == 0
world = loader.load_world(dest)
assert world.name == "The Vale of Understone"
assert world.width == 96
def test_cli_newpack_authoring_md_renders_live_band(tmp_path: Path) -> None:
"""AUTHORING.md's bands are generated from the loader, not hand-copied.
The daily_turns band is read straight from the live loader table and must
appear verbatim in the scaffolded manual proving generation from source.
"""
dest = tmp_path / "mypack"
cli.cli_newpack(dest, out=StringIO(), err=StringIO())
manual = (dest / "AUTHORING.md").read_text(encoding="utf-8")
lo, hi = loader.SETTINGS_BANDS["daily_turns"]
assert lo is not None and hi is not None
assert f"`{lo}..{hi}`" in manual
assert "daily_turns" in manual
def test_cli_newpack_authoring_md_has_width_rule_and_live_palette(tmp_path: Path) -> None:
"""AUTHORING.md documents the one-column rule and renders the live palette.
The width section states the Western-monospace assumption, and the safe
palette is generated from ``textwidth.SAFE_PALETTE`` (same can't-drift
pattern as the bands table) every glyph appears, in a backticked cell.
"""
from understone.engine.textwidth import SAFE_PALETTE
dest = tmp_path / "mypack"
cli.cli_newpack(dest, out=StringIO(), err=StringIO())
manual = (dest / "AUTHORING.md").read_text(encoding="utf-8")
assert "## Glyph width" in manual
assert "exactly one terminal column" in manual
assert "Western monospace" in manual # the stated assumption
assert "Safe glyph palette" in manual
for glyph in SAFE_PALETTE:
assert f"`{glyph}`" in manual, f"palette glyph {glyph!r} missing from manual"
def test_cli_newpack_authoring_md_documents_action_sets(tmp_path: Path) -> None:
"""AUTHORING.md documents each building's real verb menu.
The per-building menus are an explicit table: the inn's `gamble` (v0.8) and
the v0.10 vault verbs `deposit`/`withdraw`, the shop's `forge`, and so on.
This pins the table rows and the "quaff anywhere" note so a doc regression
trips.
"""
dest = tmp_path / "mypack"
cli.cli_newpack(dest, out=StringIO(), err=StringIO())
manual = (dest / "AUTHORING.md").read_text(encoding="utf-8")
assert "| `inn` | `rest`, `deposit`, `withdraw`, `gamble`, `leave` |" in manual
assert "| `shop` | `buy`, `sell`, `forge`, `leave` |" in manual
assert "| `healer` | `heal`, `leave` |" in manual
assert "| `dungeon` | `descend`, `challenge`, `leave` |" in manual
assert "`quaff`" in manual and "legal **anywhere**" in manual
# The vault is described where its verbs are listed.
assert "VAULT" in manual and "SAFE from ambush" in manual
def test_cli_newpack_authoring_md_documents_ore_forge(tmp_path: Path) -> None:
"""AUTHORING.md documents the v0.10 ore-gated forge: material slot + settings.
The forge ore is a `material` item earned in combat; the four ore settings
(item, per-plus, dungeon drop, forest chance) are documented, and the band
figures are generated from the live loader so they cannot drift.
"""
dest = tmp_path / "mypack"
cli.cli_newpack(dest, out=StringIO(), err=StringIO())
manual = (dest / "AUTHORING.md").read_text(encoding="utf-8")
assert "`material`" in manual # the new slot
assert "forge_ore_item" in manual
assert "ore_forest_chance" in manual # the float setting (prose, not the band table)
# The two banded ore settings carry their LIVE bands.
lo, hi = loader.SETTINGS_BANDS["ore_dungeon_drop"]
assert f"`{lo}..{hi}`" in manual
assert "earns in combat" in manual or "earned in combat" in manual
def test_cli_newpack_authoring_md_states_color_advisory_and_spawn_walkable(
tmp_path: Path,
) -> None:
"""AUTHORING.md states color is advisory (loader does not validate it) and
that spawn must be on walkable terrain both v0.8 honesty fixes."""
dest = tmp_path / "mypack"
cli.cli_newpack(dest, out=StringIO(), err=StringIO())
manual = (dest / "AUTHORING.md").read_text(encoding="utf-8")
# color is documented as advisory / not validated (it matches loader behaviour).
assert "advisory and not validated" in manual
# spawn's walkability requirement is now stated where spawn is introduced.
assert "must be on walkable terrain" in manual
def test_cli_newpack_authoring_md_color_roles_generated_from_enum(tmp_path: Path) -> None:
"""AUTHORING.md's colour-role vocabulary is generated from the Color enum.
The v0.9 fix: the assignable roles were hand-listed (and went stale road
and the per-building roles were missing). They are now generated from
``Color.assignable()`` the single source for the overlay-vs-assignable
split so the manual lists exactly what the Watch can paint and cannot
drift. This asserts the NEW roles appear, that every assignable enum role
appears, and that the non-assignable roles (overlays + DEFAULT) are NOT
offered as author-assignable.
"""
from understone.screen.palette import Color
dest = tmp_path / "mypack"
cli.cli_newpack(dest, out=StringIO(), err=StringIO())
manual = (dest / "AUTHORING.md").read_text(encoding="utf-8")
# A sampling of the new v0.9 roles is offered in the manual, backticked.
for role in ("road", "forest", "lava", "barren", "inn", "shop", "healer"):
assert f"`{role}`" in manual, f"new colour role {role!r} missing from manual"
# EVERY assignable enum role appears (generated, so the full set is present).
color_section = manual[manual.index("`color` — a palette role string") :].split("###", 1)[0]
for role in Color.assignable():
assert f"`{role.value}`" in manual, f"assignable role {role.value!r} missing from manual"
# The non-assignable roles (runtime overlays + the DEFAULT fallback) are NOT
# offered as terrain/location colours.
non_assignable = {c for c in Color} - set(Color.assignable())
assert Color.DEFAULT in non_assignable # the fallback is not author-pickable
for role in non_assignable:
assert f"`{role.value}`" not in color_section, (
f"non-assignable role {role.value!r} wrongly offered as author-assignable"
)
def test_cli_newpack_authoring_md_has_validate_coverage_split(tmp_path: Path) -> None:
"""AUTHORING.md honestly separates machine-enforced rules from eyeball-only.
The v0.8 subsection lists what `validate` DOES catch (including the two new
enforcements rare-as-guardian and single-boss) and what it does NOT (chief
among them: location menu `actions` contents are unvalidated).
"""
dest = tmp_path / "mypack"
cli.cli_newpack(dest, out=StringIO(), err=StringIO())
manual = (dest / "AUTHORING.md").read_text(encoding="utf-8")
assert "What `validate` checks, and what it cannot" in manual
# The newly-enforced rules are named in the DOES-catch list.
assert "Exactly one boss" in manual
assert "fixed rung guardian) must" in manual # rare-as-guardian enforcement
# The eyeball-only short list names the actions gap and the flavour caveat.
assert "Location menu `actions` contents" in manual
assert "Flavour and narration quality" in manual
def test_cli_newpack_refuses_non_empty_dir(tmp_path: Path) -> None:
dest = tmp_path / "occupied"
dest.mkdir()
(dest / "keep.txt").write_text("mine", encoding="utf-8")
out, err = StringIO(), StringIO()
rc = cli.cli_newpack(dest, out=out, err=err)
assert rc == 2
assert "non-empty" in err.getvalue()
# The pre-existing file is untouched (nothing was scaffolded over it).
assert (dest / "keep.txt").read_text(encoding="utf-8") == "mine"
assert not (dest / "AUTHORING.md").exists()
def test_cli_newpack_into_empty_existing_dir_succeeds(tmp_path: Path) -> None:
"""An existing but empty directory is a fine scaffold target."""
dest = tmp_path / "empty"
dest.mkdir()
assert cli.cli_newpack(dest, out=StringIO(), err=StringIO()) == 0
assert (dest / "AUTHORING.md").exists()
# ---------------------------------------------------------------------------
# server.main argv dispatch
# ---------------------------------------------------------------------------
def test_main_validate_dispatch_returns_status(
tmp_path: Path, capsys: pytest.CaptureFixture
) -> None:
# A broken pack routed through main exits 2; a sound one exits 0.
pack = _clone_shipped(tmp_path)
_patch_world(pack, _break_daily_turns)
with pytest.raises(SystemExit) as broken:
server.main(["validate", str(pack)])
assert broken.value.code == 2
with pytest.raises(SystemExit) as sound:
server.main(["validate", str(SHIPPED)])
assert sound.value.code == 0
assert "The door stands open." in capsys.readouterr().out
def test_main_newpack_dispatch(tmp_path: Path) -> None:
dest = tmp_path / "viamain"
with pytest.raises(SystemExit) as exc:
server.main(["newpack", str(dest)])
assert exc.value.code == 0
assert (dest / "AUTHORING.md").exists()
def test_main_worlds_dispatch(capsys: pytest.CaptureFixture) -> None:
"""`understone worlds` routes through main, exits 0, and lists the Vale."""
with pytest.raises(SystemExit) as exc:
server.main(["worlds"])
assert exc.value.code == 0
out = capsys.readouterr().out
assert "vale" in out
assert "The Vale of Understone" in out
assert "UNDERSTONE_WORLD=" in out
def test_bare_invocation_resolves_to_serve_without_side_effects() -> None:
"""Parsing no argv yields the serve path, and parsing has no side effects.
The transport launch (_serve) is reachable, but argument parsing neither
loads a world nor binds a port so this asserts the resolved command
without ever calling _serve.
"""
args = server._build_parser().parse_args([])
assert args.cmd is None # None => the serve branch in main()
assert callable(server._serve)
def test_subprocess_validate_packaged_world_exits_zero() -> None:
"""End-to-end smoke: `python -m understone validate <packaged dir>` exits 0."""
result = subprocess.run(
[sys.executable, "-m", "understone", "validate", str(SHIPPED)],
cwd=EXAMPLE_DIR,
capture_output=True,
text=True,
timeout=60,
)
assert result.returncode == 0, result.stderr
assert "The door stands open." in result.stdout
# ---------------------------------------------------------------------------
# helpers
# ---------------------------------------------------------------------------
def _clone_shipped(tmp_path: Path) -> Path:
dest = tmp_path / "pack"
shutil.copytree(SHIPPED, dest)
return dest
def _patch_world(pack: Path, mutate: Callable[[dict[str, Any]], None]) -> None:
path = pack / "world.json"
data = json.loads(path.read_text(encoding="utf-8"))
mutate(data)
path.write_text(json.dumps(data), encoding="utf-8")
def _break_daily_turns(data: dict[str, Any]) -> None:
"""Set daily_turns out of its 1..100 band so the pack fails to load."""
data["settings"]["daily_turns"] = 0
-123
View File
@@ -1,123 +0,0 @@
"""Combat resolution tests.
Pins determinism (a fixed seed yields identical results twice, log and
deltas), each outcome (win/lose/flee), xp/gold crediting on victory, and
the defeat contract: the result flags a spawn bounce with no xp/gold and a
zero hp delta (the façade applies hp=1 and the move).
"""
from __future__ import annotations
from tests.conftest import make_monster, make_player
from understone.engine.combat import Outcome, resolve_fight, resolve_flee
from understone.engine.rng import GameRNG
# A strong adventurer vs a Field Rat wins on every probed seed.
_WIN_SEED = 1
# A fragile adventurer vs a Stone Wyrm loses on every probed seed.
_LOSE_SEED = 0
# Flee outcomes (probed): seed 1 escapes clean, seed 0 is caught.
_FLEE_CLEAN_SEED = 1
_FLEE_CAUGHT_SEED = 0
def _strong_player() -> object:
return make_player(hp=20, max_hp=20, atk=5, def_=1, xp=0, gold=50)
def _wyrm() -> object:
return make_monster(tier=5, name="Stone Wyrm", hp=60, atk=18, def_=6, xp=140, gold=60)
def test_fight_is_deterministic_under_fixed_seed() -> None:
r1 = resolve_fight(GameRNG(seed=7), make_player(), make_monster())
r2 = resolve_fight(GameRNG(seed=7), make_player(), make_monster())
assert r1.log == r2.log
assert (r1.outcome, r1.xp_delta, r1.gold_delta, r1.hp_delta) == (
r2.outcome,
r2.xp_delta,
r2.gold_delta,
r2.hp_delta,
)
def test_win_credits_xp_and_gold() -> None:
player = make_player(hp=20, max_hp=20, atk=5, def_=1)
monster = make_monster(hp=6, atk=3, def_=0, xp=8, gold=3)
result = resolve_fight(GameRNG(seed=_WIN_SEED), player, monster)
assert result.outcome is Outcome.WIN
assert result.xp_delta == 8
assert result.gold_delta == 3
# hp_delta is non-positive (you may take a scratch) and never fatal here.
assert result.hp_delta <= 0
assert not result.bounce_to_spawn
def test_win_deltas_are_exact_for_pinned_seed() -> None:
player = make_player(hp=20, max_hp=20, atk=5, def_=1)
monster = make_monster(hp=6, atk=3, def_=0, xp=8, gold=3)
result = resolve_fight(GameRNG(seed=_WIN_SEED), player, monster)
# Pinned from a determinism probe; guards against silent damage drift.
assert result.hp_delta == -1
# The engine no longer emits a "falls + reward" line — that sentence is
# composed by the game façade where the xp/gold are actually banked — so
# the WIN log is one line shorter than before and ends on the kill blow.
assert len(result.log) == 4
assert result.log[-1] == "You strike for 6. (Field Rat: 0 HP)"
def test_win_log_does_not_claim_rewards() -> None:
"""The engine narrates the kill blow only; it never claims xp/gold itself.
Reward ownership lives in the façade (so the Wyrm-win legacy reset, which
keeps no xp/gold, narrates no reward). The deltas are still carried on the
result for the caller to apply.
"""
player = make_player(hp=20, max_hp=20, atk=5, def_=1)
monster = make_monster(hp=6, atk=3, def_=0, xp=8, gold=3)
result = resolve_fight(GameRNG(seed=_WIN_SEED), player, monster)
assert result.outcome is Outcome.WIN
assert result.xp_delta == 8 and result.gold_delta == 3 # deltas still set
joined = "\n".join(result.log)
assert "falls" not in joined # no kill/reward sentence in the engine log
assert "XP" not in joined and "gold" not in joined
def test_loss_flags_bounce_without_rewards() -> None:
result = resolve_fight(GameRNG(seed=_LOSE_SEED), _strong_player_loses(), _wyrm())
assert result.outcome is Outcome.LOSE
assert result.bounce_to_spawn is True
assert result.xp_delta == 0
assert result.gold_delta == 0
# Combat does not set hp to 1 itself — that is the façade's job.
assert result.hp_delta == 0
def _strong_player_loses() -> object:
return make_player(hp=12, max_hp=12, atk=4, def_=0)
def test_flee_can_escape_clean() -> None:
player = make_player(hp=20, max_hp=20, def_=1)
monster = make_monster(atk=8, def_=2)
result = resolve_flee(GameRNG(seed=_FLEE_CLEAN_SEED), player, monster)
assert result.outcome is Outcome.FLED
assert result.hp_delta == 0
def test_flee_caught_costs_hp_but_never_kills() -> None:
player = make_player(hp=20, max_hp=20, def_=1)
monster = make_monster(atk=8, def_=2)
result = resolve_flee(GameRNG(seed=_FLEE_CAUGHT_SEED), player, monster)
assert result.outcome is Outcome.FLED
assert result.hp_delta < 0
# A caught flight cannot drop the player to or below zero.
assert player.hp + result.hp_delta >= 1
def test_flee_caught_never_kills_at_low_hp() -> None:
player = make_player(hp=1, max_hp=20, def_=0)
monster = make_monster(atk=40, def_=0)
result = resolve_flee(GameRNG(seed=_FLEE_CAUGHT_SEED), player, monster)
# At 1 HP the most a failed flee can cost is 0 (cannot go below 1).
assert result.hp_delta == 0
File diff suppressed because it is too large Load Diff
-858
View File
@@ -1,858 +0,0 @@
"""Game façade integration tests over the shipped world.
Drives a full session against a temp store, a frozen clock, and a seeded
RNG: join -> status -> look -> move -> action(buy/rest/fight) -> log ->
rank -> bestow. Persistence is exercised by reopening the store.
Negative-test discipline (turn guard and bestow cap):
Two guards are pinned by assertions here. To confirm each assertion has
teeth, the implementer temporarily reverted the guard line and observed
the matching test FAIL, then restored it:
* Turn guard (engine/turns.py spend_turn): replacing
``if player.turns_left <= 0: return False`` with ``return True``
let fighting continue past the daily budget ``test_turn_budget_blocks``
then failed on the "spent for today" assertion. Restored.
* Bestow cap (game.py bestow): removing the ``if cost > remaining``
refusal let an over-budget bestowal through ``test_bestow_cap_refuses``
then failed on the unchanged-gold assertion. Restored.
* Sanitizer control-char guard (game.py _sanitize): disabling the
``not cleaned.isprintable()`` clause let a newline-injected name create a
player row and a public event ``test_join_rejects_control_char_name``
then failed. Restored. (See the comment block above the hygiene tests.)
"""
from __future__ import annotations
import unicodedata
from pathlib import Path
import pytest
from tests.conftest import fixed_clock, utc
from understone.engine.models import Mode
from understone.engine.rng import GameRNG
from understone.game import Game
from understone.persistence import Store
from understone.world.loader import load_world
PACK = Path(__file__).resolve().parents[1] / "understone" / "world" / "data"
@pytest.fixture
def clock() -> object:
return fixed_clock(utc(2026, 6, 12, 10, 0))
def _game(tmp_path: Path, clock: object, seed: int = 7) -> Game:
world = load_world(PACK)
store = Store(tmp_path / "game.db")
return Game(world, store, clock=clock, rng=GameRNG(seed=seed)) # type: ignore[arg-type]
# ---------------------------------------------------------------------------
# Join / status / look
# ---------------------------------------------------------------------------
def test_join_creates_player_at_spawn(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
out = game.join("Brandr")
player = game.players["Brandr"]
assert (player.x, player.y) == game.world.spawn
assert player.gold == game.world.settings.starting_gold
assert "@" in out
assert game.world.name in out
def test_join_resumes_existing(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
game.players["Brandr"].gold = 123
out = game.join("Brandr")
assert "Welcome back" in out
assert game.players["Brandr"].gold == 123
def test_status_unknown_player_is_friendly(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
out = game.status("Nobody")
assert "has signed the ledger" in out
assert "door_join" in out
def test_look_overworld_has_frame(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
out = game.look("Brandr")
assert "@" in out
assert "" in out and "" in out
assert len(out) < 2048
def test_overworld_frame_textured_borders_intact(tmp_path: Path, clock: object) -> None:
"""The textured overworld frame keeps square borders and a single player marker.
Structural discipline for the v0.6 texture: variants change the GLYPHS but
must never change the geometry. The box rows are uniform width, exactly one
'@' is painted, and the grass field shows more than one variant in a row
(the deterministic stipple, not a flat sheet of '.').
"""
game = _game(tmp_path, clock)
game.join("Brandr")
frame = game.look("Brandr")
lines = frame.split("\n")
# Box rows: top border + VIEW_H grid rows + bottom border, all equal width.
box = [ln for ln in lines if ln and ln[0] in "┌│└"]
widths = {len(ln) for ln in box}
assert len(widths) == 1, f"textured frame rows ragged: {widths}"
# Exactly one player marker, regardless of the surrounding texture.
assert frame.count("@") == 1
# The grass texture varies: a body row carries at least two of . , '
body = [ln for ln in lines if ln.startswith("")]
assert any(len({ch for ch in ln if ch in ".,'"}) >= 2 for ln in body)
def test_look_in_menu_shows_location(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
# Shop is two cells east of spawn along the road.
game.move("Brandr", "", "east", 2)
assert game.players["Brandr"].mode is Mode.MENU
out = game.look("Brandr")
assert "(B)uy" in out and "(L)eave" in out
# ---------------------------------------------------------------------------
# Move
# ---------------------------------------------------------------------------
def test_move_blocked_in_menu(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
game.move("Brandr", "", "east", 2) # into the shop menu
out = game.move("Brandr", "", "east", 2)
assert "inside" in out.lower()
def test_move_enters_location(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
out = game.move("Brandr", "", "west", 2) # inn is two cells west
assert game.players["Brandr"].at_location == "inn"
assert "step inside" in out.lower()
# ---------------------------------------------------------------------------
# Actions: rest, fight, turn budget
# ---------------------------------------------------------------------------
def test_rest_heals_and_charges(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
player.hp = 5
game.move("Brandr", "", "west", 2) # inn
out = game.action("Brandr", "rest", "", "")
assert player.hp == player.max_hp
assert player.gold == game.world.settings.starting_gold - game.world.settings.rest_cost
assert "full health" in out.lower()
def test_rest_when_spent_restores_a_fresh_days_turns(tmp_path: Path, clock: object) -> None:
"""Sleeping at the inn with no turns left rolls into a fresh day's allowance."""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
daily = game.world.settings.daily_turns
player.turns_left = 0 # spent for the day
player.hp = 5
game.move("Brandr", "", "west", 2) # step into the inn
out = game.action("Brandr", "rest", "", "")
assert player.turns_left == daily # a fresh day's turns restored
assert player.hp == player.max_hp # and fully mended
assert f"/{daily} ]" in out # footer reflects the refreshed budget
def test_rest_with_turns_in_hand_never_inflates_the_budget(tmp_path: Path, clock: object) -> None:
"""Resting mid-day mends but adds no turns — the top-up only fires at zero."""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
daily = game.world.settings.daily_turns
player.turns_left = daily - 3 # turns still in hand
player.hp = 5
game.move("Brandr", "", "west", 2) # step into the inn
game.action("Brandr", "rest", "", "")
assert player.turns_left == daily - 3 # unchanged: no farming past the cap
assert player.hp == player.max_hp # but the heal still lands
def test_fight_spends_a_turn_and_credits(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
# Drop into the forest_near zone so an encounter is available.
player.x, player.y = 35, 25
before_turns = player.turns_left
out = game.action("Brandr", "fight", "", "")
assert player.turns_left == before_turns - 1
assert player.xp > 0
assert "XP" in out
def test_turn_budget_blocks(tmp_path: Path, clock: object) -> None:
"""Pins the spend_turn guard: at 0 turns, fighting is refused.
See the module docstring for the revert-and-observe-failure check that
proves this assertion has teeth.
"""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
player.x, player.y = 35, 25
player.turns_left = 0
out = game.action("Brandr", "fight", "", "")
assert "spent for today" in out.lower()
# No turn was consumed past zero, and no XP was gained.
assert player.turns_left == 0
assert player.xp == 0
# ---------------------------------------------------------------------------
# Log / rank
# ---------------------------------------------------------------------------
def test_log_reports_then_advances(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
# A second player acting creates a public event Brandr has not yet seen.
game.join("Sigrun")
first = game.log("Brandr")
assert "Sigrun" in first or "Brandr" in first
assert "The Understone Herald" in first # dressed as the broadsheet
# The cursor advanced; a second read with no new events is quiet.
second = game.log("Brandr")
assert "The Understone Herald" in second # the masthead still prints
assert "still" in second.lower() # the herald-flavoured "all quiet" line
def test_rank_marks_caller(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
game.join("Sigrun")
game.players["Sigrun"].level = 5
out = game.rank("Brandr")
assert "Brandr" in out and "Sigrun" in out
assert "*" in out # the caller's row is marked
assert "" in out # box-drawing table
# ---------------------------------------------------------------------------
# Rank ★ column: stars live in their own column, so a long name keeps them
# ---------------------------------------------------------------------------
def test_win_stars_column_formats() -> None:
"""Zero is blank, 1..5 render as ★ runs, and >5 collapses to ★xN."""
from understone.game import _win_stars
assert _win_stars(0) == ""
assert _win_stars(1) == ""
assert _win_stars(5) == "★★★★★"
assert _win_stars(7) == "★x7"
def test_long_name_with_one_win_keeps_its_star() -> None:
"""A full 24-char name no longer eats its own ★ (the v0.1 truncation bug).
The name occupied the whole 20-wide field before, clipping the star away;
with a separate stars column the survives beside a maximal name.
"""
from understone.engine.rank import RankEntry
from understone.game import _render_rank_table
name = "X" * 24
rows = _render_rank_table([RankEntry(name=name, level=5, xp=100, gold=50, wins=1)], caller="")
body = "\n".join(rows)
assert name in body # the full name is present
assert "" in body # and so is its star
def test_high_win_count_renders_compact_marker() -> None:
"""Seven wins render as the compact ``★x7`` rather than seven glyphs."""
from understone.engine.rank import RankEntry
from understone.game import _render_rank_table
rows = _render_rank_table([RankEntry(name="Champ", level=9, xp=9, gold=9, wins=7)], caller="")
body = "\n".join(rows)
assert "★x7" in body
assert "★★★★★★★" not in body # not seven literal stars
def test_shared_world_other_player_marker(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
game.join("Sigrun")
# Stand Sigrun one cell east of Brandr's spawn so she lands in the view.
sig = game.players["Sigrun"]
brandr = game.players["Brandr"]
sig.x, sig.y = brandr.x + 1, brandr.y
out = game.look("Brandr")
assert "" in out # the other player shows as '☻'
# ---------------------------------------------------------------------------
# Bestow (+ cap negative test)
# ---------------------------------------------------------------------------
def test_bestow_grants_gold(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
before = player.gold
out = game.bestow("Brandr", "a daring rescue", 10, 0)
assert player.gold == before + 10
assert player.bestow_spent == 10
assert "bestowal" in out.lower()
def test_bestow_heal_charges_only_applied(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
player.hp = player.max_hp - 3 # only 3 missing
game.bestow("Brandr", "mercy after a hard fight", 0, 10)
assert player.hp == player.max_hp
# Charged for 3 HP at heal_cost_per_hp, not the requested 10.
assert player.bestow_spent == 3 * game.world.settings.heal_cost_per_hp
def test_bestow_requires_reason(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
out = game.bestow("Brandr", " ", 10, 0)
assert "reason" in out.lower()
assert game.players["Brandr"].gold == game.world.settings.starting_gold
def test_bestow_requires_nonzero(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
out = game.bestow("Brandr", "nothing at all", 0, 0)
assert "at least" in out.lower()
def test_bestow_cap_refuses(tmp_path: Path, clock: object) -> None:
"""Pins the bestow cap: an over-budget grant is refused without mutation.
See the module docstring for the revert-and-observe-failure check that
proves this assertion has teeth.
"""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
budget = game.world.settings.bestow_daily_budget
before_gold = player.gold
out = game.bestow("Brandr", "an absurd windfall", budget + 100, 0)
assert "the fates allow" in out.lower()
# Refused cleanly: no gold moved and no pool spent.
assert player.gold == before_gold
assert player.bestow_spent == 0
def test_bestow_pool_resets_next_day(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
game.bestow("Brandr", "first blessing", 20, 0)
assert player.bestow_spent == 20
# Advance the clock past UTC midnight; the next bestow sees a fresh pool.
game.clock = fixed_clock(utc(2026, 6, 13, 0, 5)) # type: ignore[assignment]
game.bestow("Brandr", "a new day's fortune", 20, 0)
assert player.bestow_spent == 20 # reset to 0 then +20, not 40
# ---------------------------------------------------------------------------
# Persistence round-trip through the façade
# ---------------------------------------------------------------------------
def test_state_survives_store_reopen(tmp_path: Path, clock: object) -> None:
game = _game(tmp_path, clock)
game.join("Brandr")
game.players["Brandr"].x, game.players["Brandr"].y = 35, 25
game.action("Brandr", "fight", "", "")
xp_after = game.players["Brandr"].xp
gold_after = game.players["Brandr"].gold
game.store.close()
world = load_world(PACK)
reopened = Store(tmp_path / "game.db")
revived = Game(world, reopened, clock=clock) # type: ignore[arg-type]
assert revived.players["Brandr"].xp == xp_after
assert revived.players["Brandr"].gold == gold_after
# ---------------------------------------------------------------------------
# Day rollover applies to fight/descend, not just join/bestow
# ---------------------------------------------------------------------------
class _MutableClock:
"""A clock whose reported moment can be advanced between calls."""
def __init__(self, moment: object) -> None:
self.moment = moment
def __call__(self) -> object:
return self.moment
def test_fight_refreshes_budget_across_midnight(tmp_path: Path) -> None:
"""A fight on a new UTC day must reset the budget without re-joining.
Before the fix, _resolve_encounter spent a turn without calling
_ensure_day, so an exhausted player who returned the next day was still
blocked until they happened to re-join.
"""
clk = _MutableClock(utc(2026, 6, 12, 23, 0))
world = load_world(PACK)
store = Store(tmp_path / "game.db")
game = Game(world, store, clock=clk, rng=GameRNG(seed=7)) # type: ignore[arg-type]
game.join("Brandr")
player = game.players["Brandr"]
player.x, player.y = 35, 25 # forest_near zone: an encounter is available
player.turns_left = 0 # spent for the day
daily = game.world.settings.daily_turns
clk.moment = utc(2026, 6, 13, 0, 5) # cross UTC midnight, no re-join
out = game.action("Brandr", "fight", "", "")
assert "spent for today" not in out.lower() # the fresh day let the fight run
assert player.turns_left == daily - 1 # reset to full, then one spent
assert player.xp > 0
assert f"/{daily} ]" in out # footer shows the refreshed budget
def test_descend_refreshes_budget_across_midnight(tmp_path: Path) -> None:
"""Descending on a new UTC day resets the budget without re-joining."""
clk = _MutableClock(utc(2026, 6, 12, 23, 0))
world = load_world(PACK)
store = Store(tmp_path / "game.db")
game = Game(world, store, clock=clk, rng=GameRNG(seed=7)) # type: ignore[arg-type]
game.join("Hero")
player = game.players["Hero"]
# Overwhelming stats so the gauntlet itself never bounces the player.
player.level, player.atk, player.def_ = 20, 200, 100
player.hp = player.max_hp = 500
player.mode = Mode.MENU
player.at_location = "dungeon"
player.turns_left = 0
daily = game.world.settings.daily_turns
clk.moment = utc(2026, 6, 13, 0, 5)
out = game.action("Hero", "descend", "", "")
assert "too weary" not in out.lower()
assert player.turns_left == daily - 1
# ---------------------------------------------------------------------------
# Input hygiene chokepoint (the _sanitize helper)
# ---------------------------------------------------------------------------
#
# Negative-test discipline (security invariant): to prove the control-char
# rejection in Game._sanitize has teeth, the implementer temporarily replaced
# its ``not cleaned.isprintable()`` clause with ``False`` (disabling the
# check) and confirmed test_join_rejects_control_char_name FAILED — the
# injected name created a player row and a public event. The clause was then
# restored. The newline-injection test below is the standing regression for
# that invariant.
def test_join_rejects_control_char_name(tmp_path: Path, clock: object) -> None:
"""A bell/control character in a name is refused with the runes line."""
game = _game(tmp_path, clock)
out = game.join("Bra\x07ndr")
assert "strange runes" in out
assert game.players == {} # no row created
assert game.events == [] # nothing persisted
def test_join_rejects_newline_name_no_persist(tmp_path: Path, clock: object) -> None:
"""An embedded newline (log-injection vector) is refused, nothing written.
The name is kept short so it is the control-char clause not the length
clause that rejects it; this is the standing regression for the
isprintable security invariant documented in the module docstring.
"""
game = _game(tmp_path, clock)
out = game.join("Bra\nndr") # 7 chars: well under the 24 limit
assert "strange runes" in out # the runes (bad-character) refusal, not length
# The security invariant: no player row and no event row escaped the guard.
assert game.players == {}
assert game.events == []
def test_join_rejects_overlong_name(tmp_path: Path, clock: object) -> None:
"""A 25-character name is refused with the narrow-ledger line."""
game = _game(tmp_path, clock)
out = game.join("X" * 25)
assert "ledger is narrow" in out
assert game.players == {}
def test_join_accepts_max_length_name(tmp_path: Path, clock: object) -> None:
"""A 24-character name is exactly at the limit and accepted."""
game = _game(tmp_path, clock)
name = "X" * 24
game.join(name)
assert name in game.players
# ---------------------------------------------------------------------------
# Narrow-ledger width rule (the _sanitize one-column clause, v0.6)
#
# Names/reasons/mail render inside fixed-width frames and tables, so a glyph
# that does not fit a single column would shove a column out of true. The
# sanitizer rejects wide runes and combining marks; a printable-but-wide name
# gets the dedicated narrow-ledger refusal, not the control-char "runes" line.
# ---------------------------------------------------------------------------
def test_join_rejects_wide_cjk_name(tmp_path: Path, clock: object) -> None:
"""A CJK ideograph name is refused with the narrow-ledger line; nothing written."""
game = _game(tmp_path, clock)
out = game.join("")
assert "columns are narrow" in out
assert game.players == {}
assert game.events == []
def test_join_rejects_emoji_name(tmp_path: Path, clock: object) -> None:
"""An emoji in a name (🌲x) is wide and refused with the narrow-ledger line."""
game = _game(tmp_path, clock)
out = game.join("🌲x")
assert "columns are narrow" in out
assert game.players == {}
def test_join_rejects_fullwidth_name(tmp_path: Path, clock: object) -> None:
"""A fullwidth Latin letter () is two columns and refused."""
game = _game(tmp_path, clock)
out = game.join("")
assert "columns are narrow" in out
assert game.players == {}
def test_join_rejects_combining_mark_name(tmp_path: Path, clock: object) -> None:
"""A name with a combining mark (decomposed accent) is refused as wide.
The name is normalised to NFD so the 'o' carries a separate U+0308
combining diaeresis a zero-width code point that desynchronises the
column count. Built explicitly so the source encoding cannot mask it.
"""
game = _game(tmp_path, clock)
decomposed = unicodedata.normalize("NFD", "Bj\u00f6rn")
assert any(unicodedata.combining(ch) for ch in decomposed) # genuinely NFD
out = game.join(decomposed)
assert "columns are narrow" in out
assert game.players == {}
def test_join_accepts_composed_latin_name(tmp_path: Path, clock: object) -> None:
"""A precomposed Latin accent (NFC name) is all single-column and accepted."""
game = _game(tmp_path, clock)
composed = unicodedata.normalize("NFC", "Bj\u00f6rn")
game.join(composed)
assert composed in game.players
def _seed_wide_named_player(db: Path, clock: object, wide_name: str) -> None:
"""Write a stored adventurer whose name is a now-illegal wide rune.
Bypasses ``join`` (which would refuse a wide name at creation) by upserting
a Player row straight through the Store, so the fixture stands in for a save
that predates the narrow-ledger rule. Built by renaming a legitimately-
created hero so every other field stays valid.
"""
from dataclasses import replace
world = load_world(PACK)
seed = Store(db)
game = Game(world, seed, clock=clock, rng=GameRNG(seed=7)) # type: ignore[arg-type]
game.join("Brandr")
base = game.players["Brandr"]
seed.upsert_player(replace(base, name=wide_name))
seed.commit()
seed.close()
def test_join_resumes_stored_wide_name(tmp_path: Path, clock: object) -> None:
"""An existing adventurer with a wide-rune name resumes \u2014 identity is never re-gated.
Resume keys off the exact stored name BEFORE the sanitizer, so a character
whose name predates the narrow-ledger rule is welcomed back rather than
locked out. This is the resume-by-exact-name invariant.
"""
db = tmp_path / "game.db"
wide = "\u9f8d"
_seed_wide_named_player(db, clock, wide)
world = load_world(PACK)
game = Game(world, Store(db), clock=clock, rng=GameRNG(seed=7)) # type: ignore[arg-type]
out = game.join(wide)
assert "Welcome back" in out # resumed, not refused
assert "columns are narrow" not in out
assert wide in game.players
def test_join_still_refuses_new_wide_name(tmp_path: Path, clock: object) -> None:
"""Creation is still gated: a NEW wide name with no stored row is refused.
The resume bypass is exact-name only; a wide name that matches no stored
adventurer falls through to the creation gate and gets the narrow-ledger
refusal, with nothing written.
"""
db = tmp_path / "game.db"
# Seed one wide-named save, then try to CREATE a different wide name.
_seed_wide_named_player(db, clock, "\u9f8d")
world = load_world(PACK)
game = Game(world, Store(db), clock=clock, rng=GameRNG(seed=7)) # type: ignore[arg-type]
out = game.join("\u7363") # a different wide rune \u2014 no stored row for it
assert "columns are narrow" in out
assert "\u7363" not in game.players
def test_bestow_rejects_newline_reason_no_persist(tmp_path: Path, clock: object) -> None:
"""A newline-embedded bestow reason is refused; no event, pool unchanged."""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
events_before = len(game.events)
out = game.bestow("Brandr", "heroics\nand a forged log line", 10, 0)
assert "plainly-spoken" in out
assert len(game.events) == events_before # no bestow event appended
assert player.bestow_spent == 0 # pool untouched
# ---------------------------------------------------------------------------
# Bestow: heal-only at full HP grants nothing (no empty grant persisted)
# ---------------------------------------------------------------------------
def test_bestow_heal_only_at_full_hp_refused(tmp_path: Path, clock: object) -> None:
"""A heal-only bestow at full HP applies nothing and must not persist."""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
assert player.hp == player.max_hp # join starts at full health
events_before = len(game.events)
out = game.bestow("Brandr", "a quiet blessing", 0, 10)
assert "already hale" in out
assert len(game.events) == events_before # no "Fortune favours" line written
assert player.bestow_spent == 0 # nothing charged
# ---------------------------------------------------------------------------
# Descend the deep: one rung per descent (see test_descend.py for the ladder)
# ---------------------------------------------------------------------------
def test_descend_fights_one_rung_and_advances(tmp_path: Path, clock: object) -> None:
"""A strong player clears the next rung: one foe fought, rewards banked, depth +1."""
game = _game(tmp_path, clock)
game.join("Hero")
player = game.players["Hero"]
player.level, player.atk, player.def_ = 20, 200, 100
player.hp = player.max_hp = 500
player.mode = Mode.MENU
player.at_location = "dungeon"
before_turns, before_gold, before_xp = player.turns_left, player.gold, player.xp
out = game.action("Hero", "descend", "", "")
# The first rung is the tier-3 guardian (Forest Wolf); deeper rungs do NOT
# appear in one descent — the deep is fought a rung at a time now.
assert "Forest Wolf" in out
assert "Cave Troll" not in out
assert player.deepest_rung == 1
assert player.turns_left == before_turns - 1
assert player.gold > before_gold
assert player.xp > before_xp
def test_descend_bounces_weak_player_to_spawn(tmp_path: Path, clock: object) -> None:
"""A fresh weak player falls on the first rung and wakes at the spawn.
Depth is NOT advanced by a loss, but it persists at whatever it was (here 0).
"""
game = _game(tmp_path, clock)
game.join("Weakling")
player = game.players["Weakling"]
player.mode = Mode.MENU
player.at_location = "dungeon"
out = game.action("Weakling", "descend", "", "")
assert player.hp == 1
assert player.mode is Mode.TILE
assert player.at_location == ""
assert (player.x, player.y) == game.world.spawn
assert player.deepest_rung == 0 # a loss never advances the deep
# Felled by the first rung (the tier-3 Forest Wolf).
assert "Forest Wolf" in out
# ---------------------------------------------------------------------------
# Shop façade: buy / upgrade / sell / heal stat arithmetic
# ---------------------------------------------------------------------------
def test_shop_buy_upgrade_sell_heal_cycle(tmp_path: Path, clock: object) -> None:
"""Equip deltas apply once on buy/upgrade and unwind cleanly on sell."""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
player.gold = 1000
player.mode = Mode.MENU
player.at_location = "shop"
short_sword = game.world.item_by_id("short_sword")
war_axe = game.world.item_by_id("war_axe")
starter = game.world.item_by_id(game.world.settings.starting_weapon)
assert short_sword is not None and war_axe is not None and starter is not None
starter_atk = player.atk # 3 base + rusty dagger bonus
# Buy the short sword: gold falls by its price, atk rises by the delta.
gold0 = player.gold
game.action("Brandr", "buy", "", "short_sword")
assert player.gold == gold0 - short_sword.price
assert player.atk == starter_atk + (short_sword.atk - starter.atk)
atk_with_sword = player.atk
# Upgrade to the war axe: atk reflects the difference, not a double-add.
gold1 = player.gold
game.action("Brandr", "buy", "", "war_axe")
assert player.gold == gold1 - war_axe.price
assert player.atk == atk_with_sword + (war_axe.atk - short_sword.atk)
# Sell the war axe: half-price refund, atk falls back to the starter bonus.
gold2 = player.gold
game.action("Brandr", "sell", "", "")
assert player.gold == gold2 + war_axe.price // 2
assert player.atk == starter_atk
# Heal at the shrine: HP restored, gold debited per missing point.
player.mode = Mode.MENU
player.at_location = "healer"
player.hp = player.max_hp - 5
per_hp = game.world.settings.heal_cost_per_hp
gold3 = player.gold
game.action("Brandr", "heal", "", "")
assert player.hp == player.max_hp
assert player.gold == gold3 - 5 * per_hp
def test_sell_starter_weapon_refused(tmp_path: Path, clock: object) -> None:
"""The starter blade is unsellable regardless of price (no free-gold loop)."""
game = _game(tmp_path, clock)
game.join("Brandr")
player = game.players["Brandr"]
assert player.weapon_id == game.world.settings.starting_weapon
player.mode = Mode.MENU
player.at_location = "shop"
gold_before = player.gold
out = game.action("Brandr", "sell", "", "")
assert "nothing worth selling" in out.lower()
assert player.gold == gold_before
# ---------------------------------------------------------------------------
# Bounded in-memory event tail (full history stays in SQLite)
# ---------------------------------------------------------------------------
def test_event_tail_is_capped_but_log_still_works(tmp_path: Path, clock: object) -> None:
"""Loading caps the resident tail; door_log still serves recent events."""
from understone.engine.log import since
from understone.game import EVENT_TAIL_KEEP
db = tmp_path / "game.db"
seed_store = Store(db)
last_id = 0
for i in range(EVENT_TAIL_KEEP + 50):
last_id = seed_store.insert_event("t", "sys", "note", f"event {i}")
seed_store.commit()
seed_store.close()
world = load_world(PACK)
game = Game(world, Store(db), clock=clock, rng=GameRNG(seed=7)) # type: ignore[arg-type]
# Only the most recent EVENT_TAIL_KEEP events are resident in memory.
assert len(game.events) == EVENT_TAIL_KEEP
assert game.events[-1].event_id == last_id
# door_log still reports events after a recent cursor.
recent_cursor = game.events[-3].event_id
game.join("Brandr")
game.players["Brandr"].log_cursor = recent_cursor
out = game.log("Brandr")
assert "The Understone Herald" in out # broadsheet masthead
assert "since your last visit" in out
fresh, new_cursor = since(game.events, recent_cursor)
assert fresh # there are events past the cursor
assert new_cursor == game.events[-1].event_id
def test_private_mail_survives_tail_eviction(tmp_path: Path, clock: object) -> None:
"""A private note older than the resident tail is still delivered (durable mail).
Public history that falls off the in-memory tail is gone by design (the
broadsheet does not keep), but mail must not be: a note left while the
recipient was away has to surface however many public events have since
pushed it out of the tail. A third player whose cursor also predates the
note must still never see it, because it was never theirs.
"""
from understone.persistence import EVENT_TAIL_KEEP
db = tmp_path / "game.db"
store = Store(db)
game = Game(load_world(PACK), store, clock=clock, rng=GameRNG(seed=7)) # type: ignore[arg-type]
game.join("Scribe")
game.join("Reader")
game.join("Bystander")
# Scribe leaves Reader a private note; neither Reader nor Bystander reads it.
secret = "the cellar key is under the third barrel"
game.action("Scribe", "post", "Reader", "", secret)
# Flood the feed past the tail bound so the note is evicted from memory.
for i in range(EVENT_TAIL_KEEP + 20):
store.insert_event("t", "sys", "note", f"broadsheet filler {i}")
store.commit()
store.close()
# Reopen: only the newest tail is resident, so the note now lives in the gap.
reopened = Store(db)
revived = Game(load_world(PACK), reopened, clock=clock, rng=GameRNG(seed=7)) # type: ignore[arg-type]
note_id = next(
e.event_id
for e in reopened.targeted_events_since("Reader", 0) # note: from SQLite, not the tail
if secret in e.text
)
assert note_id < revived.events[0].event_id # the note really is past the tail
# The recipient still sees the note, backfilled from SQLite...
reader_log = revived.log("Reader")
assert secret in reader_log
assert "While you were away" in reader_log
# ...but a third player never does, even though their cursor predates it too.
third_log = revived.log("Bystander")
assert secret not in third_log
reopened.close()
-118
View File
@@ -1,118 +0,0 @@
"""XP curve, level-up, and restorative-maths tests.
Pins the threshold edges (at / just below / just above), a multi-level
jump from a single award, the exact growth table, the inn's flat-rate
full heal with affordability gating, and the healer's per-HP cost maths.
"""
from __future__ import annotations
from tests.conftest import DEFAULT_SETTINGS, make_player, make_settings
from understone.engine.leveling import apply_xp, heal, rest, xp_for_level
# Default curve is 100 * (n-1)*n/2 cumulative:
# L2 = 100, L3 = 300, L4 = 600, L5 = 1000.
def test_xp_curve_thresholds() -> None:
assert xp_for_level(1, DEFAULT_SETTINGS) == 0
assert xp_for_level(2, DEFAULT_SETTINGS) == 100
assert xp_for_level(3, DEFAULT_SETTINGS) == 300
assert xp_for_level(4, DEFAULT_SETTINGS) == 600
assert xp_for_level(5, DEFAULT_SETTINGS) == 1000
def test_just_below_threshold_does_not_level() -> None:
player = make_player(level=1, xp=0, hp=20, max_hp=20)
gains = apply_xp(player, 99, DEFAULT_SETTINGS)
assert gains == []
assert player.level == 1
def test_exact_threshold_levels_once() -> None:
player = make_player(level=1, xp=0, hp=10, max_hp=20, atk=5, def_=1)
gains = apply_xp(player, 100, DEFAULT_SETTINGS)
assert len(gains) == 1
assert player.level == 2
# Growth table applied and a full heal granted on level-up.
assert player.max_hp == 26
assert player.atk == 7
assert player.def_ == 2
assert player.hp == player.max_hp
def test_just_above_threshold_levels_once() -> None:
player = make_player(level=1, xp=0)
gains = apply_xp(player, 101, DEFAULT_SETTINGS)
assert len(gains) == 1
assert player.level == 2
assert player.xp == 101
def test_single_award_can_jump_multiple_levels() -> None:
player = make_player(level=1, xp=0, max_hp=20, atk=5, def_=1)
gains = apply_xp(player, 600, DEFAULT_SETTINGS)
# 600 cumulative reaches level 4 (L2=100, L3=300, L4=600).
assert player.level == 4
assert [g.new_level for g in gains] == [2, 3, 4]
# Three levels of growth stacked.
assert player.max_hp == 20 + 3 * 6
assert player.atk == 5 + 3 * 2
assert player.def_ == 1 + 3 * 1
def test_growth_table_respects_settings() -> None:
settings = make_settings(growth_max_hp=10, growth_atk=3, growth_def=2, xp_base=50)
player = make_player(level=1, xp=0, max_hp=20, atk=5, def_=1)
apply_xp(player, 50, settings) # L2 at 50 with xp_base=50
assert player.level == 2
assert player.max_hp == 30
assert player.atk == 8
assert player.def_ == 3
# ---------------------------------------------------------------------------
# rest (inn) and heal (healer)
# ---------------------------------------------------------------------------
def test_rest_full_heals_and_charges() -> None:
player = make_player(hp=5, max_hp=20, gold=50)
assert rest(player, cost=15) is True
assert player.hp == 20
assert player.gold == 35
def test_rest_refused_when_unaffordable() -> None:
player = make_player(hp=5, max_hp=20, gold=10)
assert rest(player, cost=15) is False
assert player.hp == 5
assert player.gold == 10
def test_heal_charges_only_for_hp_restored() -> None:
player = make_player(hp=15, max_hp=20, gold=100)
result = heal(player, amount=10, cost_per_hp=2)
# Only 5 HP were missing.
assert result.healed == 5
assert result.cost == 10
assert player.hp == 20
assert player.gold == 90
def test_heal_bounded_by_affordability() -> None:
player = make_player(hp=2, max_hp=20, gold=7)
result = heal(player, amount=10, cost_per_hp=2)
# 7 gold buys 3 HP at 2/hp.
assert result.healed == 3
assert result.cost == 6
assert player.hp == 5
assert player.gold == 1
def test_heal_noop_when_full() -> None:
player = make_player(hp=20, max_hp=20, gold=100)
result = heal(player, amount=10, cost_per_hp=2)
assert result.healed == 0
assert result.cost == 0
assert player.gold == 100
@@ -1,288 +0,0 @@
"""End-to-end MCP integration test — the only test that touches the network.
Boots the real Understone FastMCP app (backed by a temp DB) in a uvicorn
thread, then drives it over the real streamable-HTTP wire with the real MCP
client: initialize, list_tools (all nine door_* names), join, look. A second
client session joins a second adventurer in the SAME process and world, and
the first player's view then shows the '&' other-player marker — proving the
shared-world, single-process contract over a real wire.
A second test drives the read-only Watch routes that ride inside the same app:
GET /watch (the HTML page), /watch/world.json (the static map), and
/watch/state.json (the live snapshot) confirming the spectator endpoints
serve real world data alongside a working /mcp without breaking either.
"""
from __future__ import annotations
import asyncio
import socket
import threading
import time
from typing import TYPE_CHECKING, Any
import httpx
import pytest
import uvicorn
from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client
from understone import server as understone_server
if TYPE_CHECKING:
from pathlib import Path
PACK = str(understone_server.PACKAGED_WORLD_DIR)
def _find_free_port() -> int:
s = socket.socket()
s.bind(("127.0.0.1", 0))
port = s.getsockname()[1]
s.close()
return int(port)
def _build_server(port: int, db_path: str) -> uvicorn.Server:
app = understone_server.create_app(db_path, PACK)
config = uvicorn.Config(app, host="127.0.0.1", port=port, log_level="warning", access_log=False)
return uvicorn.Server(config)
def _wait_ready(port: int, timeout: float = 5.0) -> None:
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
try:
with socket.create_connection(("127.0.0.1", port), timeout=0.5):
return
except OSError:
time.sleep(0.05)
raise TimeoutError(f"understone server at 127.0.0.1:{port} not ready after {timeout}s")
@pytest.fixture
def live_server(tmp_path: Path) -> Any:
"""Boot the real Understone app in a background uvicorn thread."""
port = _find_free_port()
db_path = str(tmp_path / "wire.db")
server = _build_server(port, db_path)
def _run() -> None:
loop = asyncio.new_event_loop()
asyncio.set_event_loop(loop)
loop.run_until_complete(server.serve())
thread = threading.Thread(target=_run, daemon=True, name="understone-itest")
thread.start()
try:
_wait_ready(port)
yield f"http://127.0.0.1:{port}/mcp"
finally:
server.should_exit = True
thread.join(timeout=5)
# create_app installed a module-level game whose Store holds an open
# SQLite connection; close it and clear the singleton so the next test
# builds its own rather than inheriting this temp DB.
if understone_server._GAME is not None:
understone_server._GAME.store.close()
understone_server._GAME = None
# FastMCP caches a StreamableHTTPSessionManager on the module-level mcp
# singleton and refuses a second lifespan .run() on the same instance.
# Reset it so each fixture instance boots a fresh session manager (the
# production server only ever runs one). Without this, a second
# fixture-using test fails on "run() can only be called once".
understone_server.mcp._session_manager = None
async def _call_text(session: ClientSession, name: str, arguments: dict[str, Any]) -> str:
result = await session.call_tool(name, arguments)
chunks = [block.text for block in result.content if getattr(block, "type", None) == "text"]
return "\n".join(chunks)
async def _drive(url: str) -> dict[str, Any]:
"""Run the full client conversation and return observations."""
observations: dict[str, Any] = {}
async with (
streamable_http_client(url) as (read, write, _get_session_id),
ClientSession(read, write) as session,
):
await session.initialize()
tools = await session.list_tools()
observations["tool_names"] = sorted(t.name for t in tools.tools)
observations["join_one"] = await _call_text(session, "door_join", {"player": "Brandr"})
observations["look_one_before"] = await _call_text(
session, "door_look", {"player": "Brandr"}
)
# A SECOND, independent session joins a second adventurer in the same world.
async with (
streamable_http_client(url) as (read, write, _get_session_id),
ClientSession(read, write) as session,
):
await session.initialize()
# Place player two adjacent to player one so they share the view.
await _call_text(session, "door_join", {"player": "Sigrun"})
await _call_text(
session, "door_move", {"player": "Sigrun", "heading": "east", "distance": 1}
)
# Back as player one: the shared world now shows the other adventurer.
async with (
streamable_http_client(url) as (read, write, _get_session_id),
ClientSession(read, write) as session,
):
await session.initialize()
observations["look_one_after"] = await _call_text(
session, "door_look", {"player": "Brandr"}
)
observations["rank"] = await _call_text(session, "door_rank", {"player": "Brandr"})
return observations
def test_mcp_end_to_end(live_server: str) -> None:
obs = asyncio.run(_drive(live_server))
# All nine tools are advertised over the wire.
expected = {
"door_help",
"door_join",
"door_status",
"door_look",
"door_move",
"door_action",
"door_log",
"door_rank",
"door_bestow",
}
assert set(obs["tool_names"]) == expected
# The join + look frames are real ASCII map frames.
assert "@" in obs["join_one"]
look_before = obs["look_one_before"]
assert "@" in look_before
assert "" in look_before and "" in look_before
# Shared-world proof: after player two joins next door, player one sees '☻'.
assert "" in obs["look_one_after"]
# And the leaderboard lists both adventurers (one process, one world).
assert "Brandr" in obs["rank"]
assert "Sigrun" in obs["rank"]
def _watch_base(mcp_url: str) -> str:
"""Derive the app root (where /watch lives) from the /mcp endpoint URL."""
return mcp_url[: -len("/mcp")] if mcp_url.endswith("/mcp") else mcp_url
async def _join_over_mcp(mcp_url: str, name: str) -> None:
"""Sign one adventurer in over the real MCP wire (so state.json sees them)."""
async with (
streamable_http_client(mcp_url) as (read, write, _get_session_id),
ClientSession(read, write) as session,
):
await session.initialize()
await _call_text(session, "door_join", {"player": name})
def test_watch_routes_serve_world_state(live_server: str) -> None:
base = _watch_base(live_server)
# The MCP join writes the player into the shared world the routes read.
asyncio.run(_join_over_mcp(live_server, "Watcher"))
with httpx.Client(timeout=5.0) as client:
page = client.get(f"{base}/watch")
world = client.get(f"{base}/watch/world.json")
state = client.get(f"{base}/watch/state.json")
# The page is real HTML carrying the static masthead.
assert page.status_code == 200
assert page.headers["content-type"].startswith("text/html")
assert "Understone — Live Watch" in page.text
# The static world payload matches the loaded world.
assert world.status_code == 200
world_body = world.json()
assert world_body["width"] == 96
assert world_body["height"] == 48
assert len(world_body["glyph_rows"]) == world_body["height"]
assert all(len(row) == world_body["width"] for row in world_body["glyph_rows"])
# The live snapshot lists the adventurer who joined over MCP.
assert state.status_code == 200
state_body = state.json()
names = {p["name"] for p in state_body["players"]}
assert "Watcher" in names
def test_watch_routes_coexist_with_mcp(live_server: str) -> None:
"""The custom routes don't shadow /mcp: tool calls still work alongside them."""
base = _watch_base(live_server)
async def _drive_both() -> tuple[str, int]:
async with (
streamable_http_client(live_server) as (read, write, _get_session_id),
ClientSession(read, write) as session,
):
await session.initialize()
joined = await _call_text(session, "door_join", {"player": "Coexist"})
with httpx.Client(timeout=5.0) as client:
status = client.get(f"{base}/watch/state.json").status_code
return joined, status
joined, watch_status = asyncio.run(_drive_both())
assert "@" in joined # the MCP tool still returns a real frame
assert watch_status == 200 # and the watch route still answers
def test_streamable_http_host_gate_off_localhost() -> None:
"""A non-localhost bind must accept remote `Host` headers on /mcp.
REGRESSION: FastMCP freezes DNS-rebinding protection (a localhost-only Host
allowlist) at CONSTRUCTION, and ``server`` builds its FastMCP at import with
the default 127.0.0.1 host. A 0.0.0.0/LAN bind therefore answered TCP and
`/watch` but 421'd `/mcp` for every remote node ("Invalid Host header").
``_serve`` drops the allowlist when bound off localhost; this pins the
mechanism a default instance rejects a foreign Host, a protection-disabled
one accepts it (a 421 in the second case is the bug returning).
Uses fresh FastMCP instances (not the module singleton) so there is no
shared-state or app-cache coupling with the live-server tests above.
"""
from mcp.server.fastmcp import FastMCP
from mcp.server.transport_security import TransportSecuritySettings
from starlette.testclient import TestClient
foreign = {
"Host": "192.168.0.239:8077",
"Accept": "application/json, text/event-stream",
"Content-Type": "application/json",
}
init = {
"jsonrpc": "2.0",
"id": 1,
"method": "initialize",
"params": {
"protocolVersion": "2025-03-26",
"capabilities": {},
"clientInfo": {"name": "probe", "version": "0"},
},
}
# Default (localhost-baked allowlist) — a remote Host is refused.
locked = FastMCP("hostgate-locked")
with TestClient(locked.streamable_http_app()) as client:
assert client.post("/mcp", headers=foreign, json=init).status_code == 421
# Protection disabled (what _serve does off localhost) — remote Host accepted.
opened = FastMCP("hostgate-open")
opened.settings.transport_security = TransportSecuritySettings(
enable_dns_rebinding_protection=False
)
with TestClient(opened.streamable_http_app()) as client:
resp = client.post("/mcp", headers=foreign, json=init)
assert resp.status_code != 421, f"remote Host still rejected: {resp.status_code} {resp.text}"
-336
View File
@@ -1,336 +0,0 @@
"""Movement resolution tests.
Covers edge clipping on all four sides, blocking terrain, the two input
forms (``"NNEE"`` vs heading+distance) and their equivalence, location
entry flipping to MENU, the MAX_STEPS cap, and a stubbed always-encounter
RNG interrupting a walk with a pending fight.
"""
from __future__ import annotations
from tests.conftest import (
FOREST,
GRASS,
WALL,
WATER,
LocationDef,
Zone,
make_player,
make_world,
)
from understone.engine.models import Mode, WorldEvent
from understone.engine.movement import MAX_STEPS, parse_directions, resolve_move
from understone.engine.rng import GameRNG
class _NeverRNG(GameRNG):
"""An RNG whose chance() never fires (no wandering encounters)."""
def __init__(self) -> None:
super().__init__(seed=0)
def chance(self, probability: float) -> bool: # noqa: ARG002
return False
class _AlwaysRNG(GameRNG):
"""An RNG whose chance() always fires (forces an encounter).
The seed still drives ``weighted_index``/``randint``, so different seeds
select different event rows while every encounter roll fires.
"""
def __init__(self, seed: int = 0) -> None:
super().__init__(seed=seed)
def chance(self, probability: float) -> bool: # noqa: ARG002
return True
# ---------------------------------------------------------------------------
# parse_directions
# ---------------------------------------------------------------------------
def test_parse_steps_string() -> None:
assert parse_directions("NNEE", "", 1) == ["N", "N", "E", "E"]
def test_parse_heading_distance() -> None:
assert parse_directions("", "east", 3) == ["E", "E", "E"]
def test_parse_clamps_to_max_steps() -> None:
assert parse_directions("NNNNNNNNNNNN", "", 1) == ["N"] * MAX_STEPS
assert parse_directions("", "north", 99) == ["N"] * MAX_STEPS
def test_parse_rejects_unknown_direction() -> None:
try:
parse_directions("NQ", "", 1)
except ValueError as exc:
assert "Q" in str(exc)
else: # pragma: no cover - failure path
raise AssertionError("expected ValueError")
# ---------------------------------------------------------------------------
# Edge clipping (all four sides)
# ---------------------------------------------------------------------------
def test_clip_north_edge() -> None:
world = make_world()
player = make_player(x=5, y=0)
result = resolve_move(world, player, _NeverRNG(), heading="north", distance=3)
assert player.y == 0
assert result.steps_taken == 0
assert result.blocked
def test_clip_south_edge() -> None:
world = make_world()
player = make_player(x=5, y=10)
result = resolve_move(world, player, _NeverRNG(), heading="south", distance=3)
assert player.y == 10
assert result.blocked
def test_clip_west_edge() -> None:
world = make_world()
player = make_player(x=0, y=5)
result = resolve_move(world, player, _NeverRNG(), heading="west", distance=3)
assert player.x == 0
assert result.blocked
def test_clip_east_edge() -> None:
world = make_world()
player = make_player(x=10, y=5)
result = resolve_move(world, player, _NeverRNG(), heading="east", distance=3)
assert player.x == 10
assert result.blocked
def test_partial_move_then_clip() -> None:
world = make_world()
player = make_player(x=8, y=5)
result = resolve_move(world, player, _NeverRNG(), heading="east", distance=5)
# 8 -> 9 -> 10, then edge.
assert player.x == 10
assert result.steps_taken == 2
assert result.blocked
# ---------------------------------------------------------------------------
# Blocking terrain
# ---------------------------------------------------------------------------
def test_blocked_by_wall() -> None:
grid = [[GRASS for _ in range(11)] for _ in range(11)]
grid[5][6] = WALL
world = make_world(grid=grid)
player = make_player(x=5, y=5)
result = resolve_move(world, player, _NeverRNG(), heading="east", distance=2)
assert player.x == 5
assert result.blocked
assert "wall" in result.blocked_reason
def test_blocked_by_water() -> None:
grid = [[GRASS for _ in range(11)] for _ in range(11)]
grid[4][5] = WATER
world = make_world(grid=grid)
player = make_player(x=5, y=5)
result = resolve_move(world, player, _NeverRNG(), heading="north", distance=2)
assert player.y == 5
assert result.blocked
assert "water" in result.blocked_reason
# ---------------------------------------------------------------------------
# Input-form equivalence and direction correctness
# ---------------------------------------------------------------------------
def test_nnee_lands_at_expected_cell() -> None:
world = make_world()
player = make_player(x=5, y=5)
resolve_move(world, player, _NeverRNG(), steps="NNEE")
# Two north (y-2), two east (x+2).
assert (player.x, player.y) == (7, 3)
def test_heading_equivalent_to_steps() -> None:
world_a = make_world()
player_a = make_player(x=5, y=5)
resolve_move(world_a, player_a, _NeverRNG(), steps="EEE")
world_b = make_world()
player_b = make_player(x=5, y=5)
resolve_move(world_b, player_b, _NeverRNG(), heading="east", distance=3)
assert (player_a.x, player_a.y) == (player_b.x, player_b.y)
def test_max_steps_truncates_long_walk() -> None:
world = make_world(width=40, height=11)
player = make_player(x=0, y=5)
result = resolve_move(world, player, _NeverRNG(), heading="east", distance=99)
assert result.steps_taken == MAX_STEPS
assert player.x == MAX_STEPS
# ---------------------------------------------------------------------------
# Location entry flips to MENU
# ---------------------------------------------------------------------------
def test_entering_location_flips_menu_mode() -> None:
loc = LocationDef(
key="inn",
kind="inn",
name="The Sleeping Drake",
x=7,
y=5,
glyph="I",
color="town",
actions=("rest", "leave"),
)
world = make_world(locations=[loc])
player = make_player(x=5, y=5)
result = resolve_move(world, player, _NeverRNG(), heading="east", distance=4)
assert player.mode is Mode.MENU
assert player.at_location == "inn"
assert result.entered_location == "inn"
# Stopped on the door at x=7 even though distance asked for 4.
assert (player.x, player.y) == (7, 5)
# ---------------------------------------------------------------------------
# Encounter interrupt
# ---------------------------------------------------------------------------
def test_always_encounter_stops_with_pending_fight() -> None:
grid = [[FOREST for _ in range(11)] for _ in range(11)]
zone = Zone(key="wood", x0=0, y0=0, x1=10, y1=10, tier_lo=1, tier_hi=2)
world = make_world(grid=grid, zones=[zone])
player = make_player(x=5, y=5)
result = resolve_move(world, player, _AlwaysRNG(), heading="east", distance=5)
assert result.pending_fight == (1, 2)
# The encounter fires on the first entered cell.
assert result.steps_taken == 1
assert player.x == 6
def test_no_zone_means_no_encounter() -> None:
grid = [[FOREST for _ in range(11)] for _ in range(11)]
world = make_world(grid=grid, zones=[])
player = make_player(x=5, y=5)
result = resolve_move(world, player, _AlwaysRNG(), heading="east", distance=3)
assert result.pending_fight is None
assert result.steps_taken == 3
# ---------------------------------------------------------------------------
# Weighted non-combat overworld events (v0.2)
# ---------------------------------------------------------------------------
def _event_world(*events: WorldEvent) -> object:
"""An all-forest, fully-zoned world carrying a crafted event table."""
grid = [[FOREST for _ in range(11)] for _ in range(11)]
zone = Zone(key="wood", x0=0, y0=0, x1=10, y1=10, tier_lo=1, tier_hi=2)
return make_world(grid=grid, zones=[zone], events=list(events))
def test_event_fight_stops_the_walk() -> None:
"""A fight-kind event sets pending_fight and halts the walk like v0.1."""
world = _event_world(WorldEvent("fight", 1, "", 0, 0))
player = make_player(x=5, y=5)
result = resolve_move(world, player, _AlwaysRNG(), heading="east", distance=5)
assert result.pending_fight == (1, 2)
assert result.event is None
assert result.steps_taken == 1 # stopped on the first triggering cell
def test_event_gold_credits_and_continues() -> None:
"""A gold event credits the rolled amount and does NOT stop the walk."""
world = _event_world(WorldEvent("gold", 1, "a coin-purse", 5, 5))
player = make_player(x=5, y=5, gold=10)
result = resolve_move(world, player, _AlwaysRNG(), heading="east", distance=3)
assert result.event is not None
assert result.event.kind == "gold"
assert result.event.amount == 5 # min == max == 5, so deterministic
assert player.gold == 15
assert result.pending_fight is None
assert result.steps_taken == 3 # the walk ran to completion
def test_event_heal_caps_at_max_hp() -> None:
"""A heal event never overfills: hp is clamped to max_hp."""
world = _event_world(WorldEvent("heal", 1, "a spring", 50, 50))
player = make_player(x=5, y=5, hp=18, max_hp=20)
result = resolve_move(world, player, _AlwaysRNG(), heading="east", distance=1)
assert player.hp == 20 # +50 requested, capped at the 2 missing
assert result.event is not None and result.event.amount == 2
def test_event_trap_floors_hp_at_one_and_spares_gold() -> None:
"""A trap event never kills (floors at 1 HP) and never touches gold."""
world = _event_world(WorldEvent("trap", 1, "old briars", 500, 500))
player = make_player(x=5, y=5, hp=10, max_hp=20, gold=42)
result = resolve_move(world, player, _AlwaysRNG(), heading="east", distance=1)
assert player.hp == 1 # huge trap, but floored
assert player.gold == 42 # gold untouched
assert result.event is not None and result.event.amount == 9 # only 9 could be taken
def test_event_lore_mutates_nothing() -> None:
"""A lore event changes no state and reports a zero amount."""
world = _event_world(WorldEvent("lore", 1, "an old waystone", 0, 0))
player = make_player(x=5, y=5, hp=15, max_hp=20, gold=7)
before = (player.hp, player.gold)
result = resolve_move(world, player, _AlwaysRNG(), heading="east", distance=2)
assert (player.hp, player.gold) == before
assert result.event is not None and result.event.kind == "lore"
assert result.event.amount == 0
assert result.steps_taken == 2
def test_at_most_one_event_per_walk() -> None:
"""Once any event fires, no further cells roll for the rest of the walk.
Two distinct gold rolls would credit 2 gold (1 each); a single fired event
credits exactly 1, proving the walk stops rolling after the first trigger.
"""
world = _event_world(WorldEvent("gold", 1, "a coin", 1, 1))
player = make_player(x=5, y=5, gold=0)
resolve_move(world, player, _AlwaysRNG(), heading="east", distance=5)
assert player.gold == 1 # exactly one event, not five
def test_each_event_kind_reachable_with_crafted_table() -> None:
"""Equal weights make every kind in a crafted table reachable from movement."""
table = [
WorldEvent("fight", 1, "", 0, 0),
WorldEvent("gold", 1, "g", 1, 1),
WorldEvent("heal", 1, "h", 1, 1),
WorldEvent("trap", 1, "t", 1, 1),
WorldEvent("lore", 1, "l", 0, 0),
]
zone = Zone(key="wood", x0=0, y0=0, x1=0, y1=0, tier_lo=1, tier_hi=2)
grid = [[FOREST for _ in range(11)] for _ in range(11)]
world = make_world(grid=grid, zones=[zone], events=table)
seen: set[str] = set()
for seed in range(60):
player = make_player(x=0, y=1, hp=10, max_hp=20) # one step north into the zone cell
result = resolve_move(world, player, _AlwaysRNG(seed), steps="N")
if result.pending_fight is not None:
seen.add("fight")
elif result.event is not None:
seen.add(result.event.kind)
assert seen == {"fight", "gold", "heal", "trap", "lore"}
-9
View File
@@ -1,9 +0,0 @@
"""Smoke test for the packaging skeleton."""
from __future__ import annotations
import understone
def test_version_present() -> None:
assert understone.__version__ == "0.10.0"

Some files were not shown because too many files have changed in this diff Show More