* fix: harden MCP client against misbehaving servers
Misbehaving/failed/misconfigured MCP servers could peg CPU at 100% due
to anyio cancel-scope busy-loops (SDK #2147), uncancelled orphaned
futures, and missing application-layer resilience.
Five fixes:
1. Cancel orphaned futures on timeout — future.cancel() in all sync
bridge methods prevents coroutine accumulation on the event loop
2. Per-server circuit breaker — 3-failure threshold with exponential
cooldown (30s–5min), per-server jitter, auto-reconnect on half-open
probe, McpError excluded (protocol errors from healthy servers)
3. Safe transport stream pre-close — store stream refs and close them
before stack teardown in all error/shutdown paths, preventing the
anyio zero-buffer CPU busy-loop
4. Notification debounce — 5s per-server rate limit on list_changed
refresh storms from buggy servers
5. Periodic refresh backoff with auto-reconnect — disconnected servers
get reconnection attempts with exponential backoff (60s–1hr) instead
of being silently skipped forever
* docs: add MCP resilience section to architecture docs and diagram
Document the circuit breaker, future cancellation, stream pre-close,
notification debounce, and periodic refresh backoff in the architecture
guide and the MCP architecture PlantUML diagram.
* fix: address review — stack leak on transport error, half-open comment
- Widen _connect_one guard to check _per_server_stacks too, not just
_sessions. Transport errors in sync dispatch methods evict the session
but left the stack behind, leaking anyio tasks on reconnect.
- Clarify half-open design: multiple callers are intentionally allowed
through (reconnects serialize on the event loop, first failure re-trips).
* fix: mobile UX for console sidebar drawer and server chat input
Console admin sidebar: add box-shadow elevation, close button with
focus return, 44px touch targets, focus-into-drawer on open, flip
active indicator to left border, cubic-bezier easing, aria-expanded,
fix resize handler state desync, guard toggle injection for panels
without toolbars.
Server chat input: on touch devices Enter inserts newline (tap Send
button to send), hide Shift+Enter hint from placeholder.
* fix: preserve first group label spacing when close header is injected
Add sibling combinator selector so the first sidebar group keeps its
reduced top padding regardless of whether the close header div is
present as first-child.
* feat: render rich media embeds for MCP tool results
Detect structured media JSON (stream_url, results, sessions) in MCP
tool output and render interactive cards instead of plain text.
Web UI: media cards with thumbnail, title, metadata, and click-to-play
video/audio. HLS via lazy-loaded hls.js with direct-stream preference.
Collapsed raw JSON (API keys redacted) for inspection.
Discord: rich embeds with proxied thumbnail images (fetched by the bot
since Discord CDN cannot reach private media servers). Search results
as numbered lists, session state as "Now Playing" cards. Stream URLs
never exposed in embeds — web_url used for safe clickable links.
CI: vendor hls.js 1.6.15 with renovate tracking and update script.
* fix: address PR #292 review — SSRF guards, streaming fetch, tests
- URL validation: reject non-http(s) schemes and userinfo in thumbnail
URLs. Private IPs intentionally allowed (media servers are on LAN).
- Streaming fetch: use http.stream() with aiter_bytes() and a running
byte count to enforce the 2MB cap without buffering the full response.
Validate content-type is image/* before downloading.
- Resilience: wrap try_build_media_embed in try/except in bot.py so a
media embed failure falls through to the code-block path.
- LICENSE: download hls.js LICENSE from npm on update instead of only
copying from old dir.
- Tests: add 19 new tests — try_parse_media (8 cases), _is_safe_image_url
(7 cases), embed builders (4 cases including stream_url exclusion and
string season/episode safety).
* chore: add LICENSE file for vendored hls.js
* fix: remove ANSI escape codes from tool preview fields
Preview text (tool args, URLs, queries) was wrapped in DIM/RESET ANSI
codes at the source in session.py, which leaked into SSE events and
rendered as raw escape sequences in Discord and the web UI.
Move ANSI styling to the CLI consumer (cli.py) where it belongs. Also
escape markdown in Discord tool name titles to prevent __ from being
interpreted as underline formatting.
* fix: drop [MCP: server] prefix from tool descriptions
The prefix made MCP tools look second-class compared to builtins,
causing models to hesitate using them. The server name is already
encoded in the tool name (mcp__server__tool).
* feat: pretty-print JSON tool output, player error state, broader key redaction
- JSON tool results are detected and pretty-printed with 2-space indent
instead of rendering as a wall of text
- API key redaction extended to cover api_key, apiKey, api-key, and
token query params across all tool output (not just media embeds)
- Video/audio player shows styled error message when stream fails to
load instead of leaving a broken player element
- Both appendToolOutput and replayHistory use shared renderToolOutput()
* fix: designer review — player error retry, contrast, tool-cmd cap
- Player error: role="alert" for screen readers, retry button that
reuses existing play handler, includes media title in error message
- Light theme: darken --red from #dc2626 to #b91c1c (5.7:1 contrast
on --code-bg, was 4.3:1 failing WCAG AA at 12px)
- Pretty-print collapsed raw JSON in media embeds (was missed earlier)
- Cap .tool-cmd at 120px to prevent tools with many args from making
approval blocks disproportionately tall in history replay
- Dedicated .media-player-error class instead of reusing .tool-output
* fix: Discord tool info name matching regression, suppress deprecation warning
The escape_markdown call on tool names was stored for matching against
ToolResultEvent.name, but event.name is raw/unescaped. The escaped name
never matched, so the "Running → Done" transition silently failed and
previews disappeared from the status embed.
Fix: store raw name for matching, use escaped name only for display.
Also suppress discord.py's re.sub count deprecation warning (Python
3.13+ issue, fixed upstream).
* fix: update MCP tool description tests to match prefix removal
* fix: address PR #292 review round 2
- Retry button: handle missing span children in click handler so retry
buttons from player error state don't throw
- Footer count: use len(lines) instead of min(len(results), 10) to
reflect actual rendered count after char budget truncation
- Null display: use "null" instead of "None" in JS tool arg preview
- Broader redaction: also redact JSON "api_key": "..." patterns
- SSRF hardening: block loopback and link-local IPs plus cloud metadata
hostnames in thumbnail fetch (private LAN IPs still allowed)
* fix: bundle production compose.yaml for pipx users (#293)
Users who install via pipx don't have a git clone, so there's no
compose.yaml or Dockerfile. Bootstrap now extracts a bundled production
compose file that uses pre-built ghcr.io images instead of local builds.
- Add turnstone/deploy/compose.yaml (ghcr.io images, no build blocks,
single-node production profile only)
- Add write_compose tool to bootstrap wizard
- Update bootstrap system prompt to check for and write compose.yaml
- Remove stale ddgCluster profile references from system prompt
- Include turnstone/deploy/*.yaml in wheel
* fix: use postgresql+psycopg:// DSN scheme in compose fallbacks
The Docker image ships psycopg3, not psycopg2, so the bare
postgresql:// scheme fails. Also clarify PG usage comment in
production compose.
* fix: improve web_fetch reliability — strip scripts, dynamic truncation, more tokens
- strip_html() now removes <script>, <style>, <template>, <noscript>
element content instead of just their tags
- Truncation budget scales with context window (75% in chars, 50k floor)
and takes from the beginning only instead of head+tail splice
- max_tokens bumped from 2000 to 8192 so thinking models don't starve
the visible extraction answer
- reasoning_effort="low" on summarization call to avoid wasting tokens
- Empty responses and empty extractions now report as tool errors
* refactor: extract _utility_completion to fix reasoning_effort duplication
Callers previously had to pass reasoning_effort both as a direct keyword
(for commercial providers) and via _provider_extra_params (for local
model servers). This duplication was easy to get wrong — web_fetch was
already missing the direct keyword.
_utility_completion threads it through both paths from a single call,
used by title generation, compaction, and web_fetch extraction.
* fix: disable thinking when max_tokens too small, cap extraction at 500k
_reasoning_params now returns empty dict when max_tokens can't fit a
thinking budget (e.g. title gen with max_tokens=200). Previously
produced budget_tokens >= max_tokens which is an API error on
manual-thinking Anthropic models.
Also caps web_fetch content truncation at 500k chars — the dynamic
context-window calc was producing 3M chars on 1M-context models.
* fix: clamp utility max_tokens to model output limit, add strip_html tests
_utility_completion now clamps max_tokens to the model's advertised
max_output_tokens so small/local models don't reject 8192-token
requests.
Adds 8 tests for invisible element stripping (script, style, template,
noscript) including multiline, case-insensitive, and attribute cases.
* fix: mock get_capabilities in title retry tests for _utility_completion
_utility_completion calls _get_capabilities to clamp max_tokens. The
existing title tests mocked _provider as a bare MagicMock, so
caps.max_output_tokens was a truthy MagicMock instead of an int. Set
get_capabilities to return a real ModelCapabilities instance.
Build the image once via the profileless console service and reference
it as turnstone:local from server/channel. Prevents stale images when
users run docker compose build without --profile.
* fix: include prompt .md files in wheel, add wheel-completeness CI (#289)
Prompt markdown files were missing from PyPI wheels since the modular
prompts refactor, causing FileNotFoundError on startup for pip-installed
users. Add the missing include pattern and a new CI job that diffs
source-tree data files against wheel contents so omissions are caught
before merge.
* fix: sanitise ALLOW patterns in wheel-completeness check
Strip blank lines and leading whitespace from the allowlist before
passing to grep -vFxf so empty patterns cannot silently match all lines.
* fix: log clean one-liner when PostgreSQL becomes unavailable
Wrap all 174 connection sites in PostgreSQLBackend through a _conn()
context manager that catches OperationalError, emits a single
database.unavailable log line (with connection URL), and suppresses
repeats until the connection is restored (database.connection_restored).
* fix: add StorageUnavailableError and cover all heartbeat loops
Address review feedback:
- Separate connect-phase from execution-phase in _conn() so that
OperationalError during caller code (e.g. BEGIN IMMEDIATE lock
contention) is not misclassified as a connectivity failure.
- Add StorageUnavailableError exception class so callers can
distinguish transient DB outages without redundant tracebacks.
- Apply the same _conn() wrapper to SQLiteBackend for consistency.
- Catch StorageUnavailableError in all 7 periodic loops: watch
runner, server heartbeat, channel heartbeat, console heartbeat,
collector discovery, rebalancer, and scheduler.
- Guard dedup flag with threading.Lock.
- Add tests for dedup logging and PostgreSQL path.
* fix: chunk IN clauses to stay within DB parameter limits
psycopg caps query parameters at 65 535 and SQLite defaults to 999.
assign_buckets, prune_workstreams, and count_skill_resources_bulk were
passing unbounded lists into single IN(...) clauses, causing
OperationalError during rebalancer runs on full-size hash rings.
Chunk sizes: 10 000 (PostgreSQL), 500 (SQLite).
* fix: deduplicate assign_buckets input, add chunking regression tests
Address review feedback: deduplicate bucket list before chunking to
prevent inflated rowcount from cross-chunk duplicates. Add tests that
exercise the multi-chunk path (1200 buckets > SQLite chunk_size of 500)
and verify dedup preserves accurate counts.
Multiple CI completions for the same commit (tag push + branch push)
caused duplicate publish and docker runs. Concurrency group keyed on
head_sha ensures only one publish runs per commit.
* chore: release infrastructure for dual-track stable/experimental
CI/CD changes for the 1.0 release:
- Gate PyPI publish and Docker publish on CI success via workflow_run
- Add docker-publish.yml: builds and pushes to GHCR with smart tagging
(stable gets :X.Y.Z/:X.Y/:stable/:latest, pre-release gets :experimental)
- Add stable/* and v* tags to CI and docker-scan triggers
- Remove stale [mq] extra and types-redis from CI (Redis MQ deleted)
- Remove stale redis from Renovate package rules
Release tooling:
- scripts/release.sh: bump version, uv lock, commit, tag (with --push)
- docs/releasing.md: documents stable/experimental workflow
Docker:
- Add /workspace mount point (WORKSPACE_MOUNT env var, defaults to empty volume)
- Update .env.example: remove stale Redis/auth-token refs, add workspace/model/discord
README:
- Remove beta warning, add hero image and release tracks table
* fix: derive release tag from git instead of workflow_run.head_branch
Use git tag --points-at HEAD after checkout to resolve the release
tag instead of relying on workflow_run.head_branch, which may not
reliably be the tag name for tag-triggered CI runs. Both publish
and docker-publish workflows now skip cleanly when no v* tag exists
at the checked-out commit.
* fix: replay plan review prompt on SSE reconnection
Plan approval prompts were lost when a user navigated to the server
web UI from the console dashboard (triggering a new SSE connection).
Tool approvals stored pending state in _pending_approval and replayed
it on reconnection, but plan reviews used fire-and-forget _enqueue
with no persistent state.
Mirror the _pending_approval pattern: store _pending_plan_review
before blocking, replay it in events_sse for new SSE clients, and
clear it on resolution. Without this fix, plan reviews silently
timed out after 1 hour and were treated as approval.
* test: add plan review SSE replay regression tests
Covers pending state lifecycle: stored during on_plan_review, cleared
on resolve_plan, available for SSE reconnection replay.
The async variants accepted token_factory for auto-rotating JWTs via
ServiceTokenManager, but the sync wrappers did not expose or forward
the parameter. External SDK users calling the sync clients with
token_factory got a TypeError.
Settings rows and admin table rows used align-items: center, which
caused inputs and source badges to drift away from their labels when
descriptions wrapped to multiple lines. Switch to align-items: start
so controls stay next to their label names regardless of row height.
Add 2px top margin on settings toggles to pixel-align with text input
top padding in start-aligned rows.
* fix(examples): rewrite mcp-cluster-ops to use console SDK for cluster routing
The example MCP server was broken after the direct HTTP transport
refactor — it used TurnstoneServer (single-node) for cluster ops that
require TurnstoneConsole (cluster gateway). Rewrites dispatch flow to:
route via console → SSE stream from node → cleanup via console.
- Switch from TurnstoneServer to TurnstoneConsole for node listing and
workstream routing (TURNSTONE_CONSOLE_URL replaces TURNSTONE_SERVER_URL)
- Add proper workstream lifecycle: create via routing proxy, stream from
node, close in finally block with leak-safe ws_id guard
- Catch dispatch exceptions in run_on_node for structured JSON errors
- Extract _extract_node_ids helper, remove dead n.get("id") fallback
- Normalise _console_kwargs to always include token key
- Rewrite tests against Console+Server mocks (36 → 44 tests)
* fix(examples): paginate node listing and clarify auth in README
Address Copilot review feedback on #278:
- _list_nodes_sync now paginates via offset/limit loop so clusters
with >100 nodes are fully discovered
- README step 2 now mentions token passthrough for authenticated clusters
- New test_paginates_large_clusters verifies multi-page fetch (45 tests)
* fix: remove non-auth support from bootstrap wizard
Auth is now mandatory for all deployments. Remove the
TURNSTONE_AUTH_ENABLED toggle and make JWT_SECRET and AUTH_TOKEN
required in the wizard's system prompt.
* fix: remove auth disable support from runtime and infra
Remove AuthConfig.enabled field — auth is always on. Drop
TURNSTONE_AUTH_ENABLED env var, config toggle, and the
check_request bypass. Update compose.yaml, Helm chart,
Terraform, docs, and tests to match.
* feat: deprecate config tokens, require JWT secret, prefer JWT auth
Phase 1 of config-token removal:
- load_jwt_secret() now exits with error if no secret is configured
(was: silently auto-generated ephemeral secret)
- _authenticate_token() logs deprecation warning on config token use
- CLI /cluster commands use ServiceTokenManager when JWT secret is set
- turnstone-admin tls-list uses ServiceTokenManager when JWT secret is set
- Update bootstrap wizard, docker.md, security.md to mark
TURNSTONE_AUTH_TOKEN as deprecated and JWT_SECRET as required
- Console test fixtures use auth token + headers (auth always enforced)
* feat: add service scope for inter-service JWT auth
Add "service" to VALID_SCOPES and SCOPE_HIERARCHY. Service tokens
bypass require_permission() RBAC checks, replacing the old
empty-user-id bypass that config tokens relied on.
All ServiceTokenManager instances that need admin access now include
"service" in their scopes (console proxy, channel gateway, CLI,
admin CLI). Read-only services (collector, notification) unchanged.
* feat: phase 2 config token deprecation
- SDK doc examples now show API tokens (ts_) instead of config tokens
- Remove _get_config_token() from admin CLI (dead code)
- Block config token exchange in handle_auth_login — only password
and API token login allowed
- Update login tests to use password-based auth instead of config
token exchange
* feat: phase 3 — remove config tokens entirely
Complete removal of config-file token authentication:
- Delete AuthConfig.tokens, check(), _ROLE_TO_SCOPES, hmac dispatch
branch, and config token loading from load_auth_config()
- Remove auth_config parameter from _authenticate_token() and
check_request() — callers updated throughout
- Remove TURNSTONE_AUTH_TOKEN from compose.yaml, Helm charts,
Terraform, turnstone.example.toml
- Remove --auth-token CLI flags from turnstone, turnstone-admin,
and turnstone-console
- Simplify console main() — always use ServiceTokenManager
(no fallback to static tokens)
- Delete config-token-specific tests, rewrite check_request and
integration tests to use JWT auth with proper audience claims
- Remove all config token references from docs (security.md,
docker.md, sdk.md, console.md, architecture.md, bootstrap prompt)
* fix: address code review findings
- Fix 33 broken tests: add JWT auth to test_api_versioning,
test_console_routing_proxy, test_tls_admin, test_tls_manager,
test_server_live (jwt_secret + audience-scoped auth headers)
- Add TestRequirePermissionServiceScope: 4 tests covering the
service scope RBAC bypass path
- Remove stale comments referencing config tokens in auth.py and
console/server.py
- Remove dead proxy_auth_token parameter from console create_app()
and static token fallback in _proxy_auth_headers()
- Remove TURNSTONE_AUTH_TOKEN from env.py scrub list
* fix: address Copilot review — JWT audience, compose require secret
- CLI /cluster: add audience=JWT_AUD_CONSOLE to ServiceTokenManager
(console validates audience, JWTs without it were rejected)
- Admin CLI tls-list: same audience fix
- compose.yaml: TURNSTONE_JWT_SECRET now uses :? to fail fast if unset
- SDK console: fix default port from 8081 to 8090
* test: add auth enforcement tests for TLS admin endpoints
5 new tests: unauthenticated requests return 401 (list, renew,
delete), read-only-scoped requests return 403 (renew, delete).
Closes the TLS auth enforcement test gap noted in PROGRESS.md.
* fix: address remaining Copilot review feedback
- Fix token_source="config" → "test" in TLS test fixtures
- Fix AuthResult.token_source docstring to include service origins
- Require TURNSTONE_JWT_SECRET in cluster compose profile (:?)
- Helm: add auth.jwtSecret + auth.existingSecret values, wire
TURNSTONE_JWT_SECRET into secret.yaml and both deployments
- Terraform: replace auth_token with jwt_secret variable + secret,
remove orphaned auth_token resources and IAM reference
- Remove [[auth.tokens]] from security.md config example
* fix: address full code review — 10 findings
Critical:
- Terraform: replace concat(common_env, auth_env) with common_env
(auth_env local was removed but still referenced)
- Channel gateway: remove hmac static token auth from _check_auth(),
use JWT-only validation. Remove --auth-token CLI arg from channel
- Rebalancer: add token_manager support so migration requests carry
JWT auth (was sending unauthenticated POST to /internal/migrate)
Major:
- Guard _permissions_to_scopes() against "service" privilege
escalation from DB role permissions
- Remove dead AuthConfig class, load_auth_config(), and all
auth_config parameters from create_app() signatures
- Helm: inject JWT secret for both inline and existingSecret paths
Minor:
- Remove dead auth_token param from ClusterCollector
- Remove empty TestLoadAuthConfig class
- Short JWT secret now exits instead of warning
- Compose: add generation command comment above JWT_SECRET
- Clean stale config token references from 6 doc files
- Clean stale AUTH_TOKEN reference from bootstrap wizard prompt
* fix: remove remaining stale config token references from docs
- channels.md: remove --auth-token from options table
- oidc.md: remove "config-file tokens still work" claim
- security.md: remove config token section, fix JWT secret docs
(now required/exits, no ephemeral fallback), remove hmac from
ASCII diagram, remove --auth-token reference
* fix: populate model in _last_usage so usage-by-model records correctly
_last_usage was built purely from UsageInfo token counts, never
including a "model" key. server.py's on_status() fell back to
model="" for every record_usage_event call, so GROUP BY model
collapsed all rows into a single empty-key bucket.
* fix: inject model at emission time, preserve dict[str, int] typing
Address Copilot review: keep _last_usage as dict[str, int] for type
safety, inject "model" from self.model when passing to on_status().
This also fixes stale model after /model switch since the value is
read fresh each time.
* fix: MCP tools not surfacing after Sync to Nodes, update Anthropic tool search
Three fixes:
1. session_factory closure captured mcp_client=None when no --mcp-config
was passed at startup. internal_mcp_reload created a new MCPClientManager
on app.state but the factory never saw it. New workstreams got 0 MCP tools.
Fix: mutable _mcp_ref list shared between factory and reload handler.
2. Anthropic dropped the date suffix from tool_search_tool_bm25_20251119
and now requires name == type. Updated constant and tool definition.
3. Add diagnostic logging around API errors (provider, model, base_url,
message counts, full exception chain) and workstream resume (pre/post
provider state, alias resolution warnings).
Also adds Node.js 24 LTS to Dockerfile via multi-stage copy for npx-based
MCP servers.
* fix: address Copilot review — set_storage on reload, sanitize log output
- Call mcp_mgr.set_storage(storage) when internal_mcp_reload creates a
new MCPClientManager so prompt sync works for post-startup servers
- Strip query params from base_url before logging (may contain API keys
in some vLLM deployments)
- Split API error logging: concise warning (type names only) + separate
debug with exc_info=True for full traceback when needed
* chore: remove DDG MCP sidecar, web_search uses built-in ddgs client
The DuckDuckGo MCP server container is redundant — the built-in
DuckDuckGoClient (via ddgs package, included in all extras) auto-detects
when no Tavily key is configured. Removes the ddg-search service,
ddgCluster profile, and mcp-ddg.json config file.
* fix: materialize skill resources to disk for subprocess access
Skill-bundled scripts stored in skill_resources were loaded into memory
but never written to disk, causing FileNotFoundError when the model
tried to execute them. Write resources to a per-workstream temp directory
on skill load, expose via SKILL_RESOURCES_DIR env var and PATH, clean up
on skill change or session close.
* fix: pre-flight validation warns when skill references missing resources
Scan rendered skill content for path references (scripts/foo.py, etc.)
and compare against bundled skill_resources. Warn via on_info if any
referenced paths are not bundled, so operators see the gap at skill
activation rather than at runtime FileNotFoundError.
* fix: address PR #271 review feedback
- Fix trailing colon in PATH when $PATH is empty (cwd-on-PATH risk)
- Move try/except inside per-resource loop so one bad write doesn't
abort all resources
- Explicit encoding="utf-8" for deterministic writes across locales
* fix: address PR #271 review round 2
- Normalize available paths in _validate_skill_resources() to match
referenced paths (both sides use os.path.normpath now)
- Fix flaky traversal test: assert inside base dir, not escaped path
Multiple browser tabs open to the same server could see workstream
names, states, and content mixed up between workstreams when creating,
closing, and switching tabs rapidly.
Root causes and fixes:
- Global SSE ws_created events were never handled — other tabs never
learned about new workstreams, causing blank names and stale tab bars
- SSE reconnection assigned all stale panes to the first workstream
instead of deduplicating; now uses two-pass assignment with tracking
- switchTab left the old EventSource open while reassigning pane.wsId,
creating a window for events to leak; now disconnects SSE first
- Per-workstream events carried no ws_id — server now stamps ws_id on
all events via _enqueue (shallow copy); client handleEvent drops
events with mismatched ws_id as defense-in-depth
- Plan dialog used pane.wsId at resolve time (could drift after tab
switch); now captures ws_id when the dialog opens
- Global ws_closed could reassign panes before per-ws SSE finished
draining; now disconnects per-ws SSE immediately on close
The 5 prompt policy admin endpoints required "admin.prompt_policies"
but the builtin-admin role only grants "admin.policies". Changed to
match the existing permission used by tool policy endpoints.
* fix: harden Discord bot against gateway disconnects and SSE failures
- Isolate Discord API failures from SSE stream — _on_ws_event exceptions
no longer kill the SSE connection and cause missed events
- Fix broken exponential backoff on 4xx/5xx (delay was reset on every
attempt); skip aiter_sse() on error responses
- Add read timeout (90s) to SSE httpx client so half-open TCP
connections are detected and recovered
- Re-resolve node URL on each SSE reconnect attempt
- Add on_resumed handler to recover SSE tasks that died during brief
gateway disconnects (on_ready is not called on session resume)
- Sync slash commands only on first on_ready to avoid Discord rate limits
* fix: SSE backoff on 4xx/5xx and retrieve dead task exceptions
- Replace `continue` with raise+catch so 4xx/5xx errors hit the
exponential backoff path instead of tight-looping
- Retrieve task exceptions in _purge_dead_sse_tasks to suppress
"Task exception was never retrieved" warnings and log the cause
* feat: modular system message composition with admin prompt policies (#267)
Replace the monolithic persona+tools block in _init_system_messages()
with a modular composition harness (turnstone/prompts/). System messages
are now assembled from five typed layers: BASE (persona), ENV (client
surface — web/cli/chat), CONTEXT (datetime, timezone, username), TOOLS
(usage patterns), and POLICIES (behavioral rules with tool gating).
Prompt policies are admin-managed via a new Prompts tab in the Governance
group (CRUD with modal forms, tool gating, priority ordering, enable/disable).
DB policies override file-based defaults by name; file-based policies serve
as deployment defaults. Migration 031 adds the prompt_policies table.
ClientType is threaded end-to-end from channel adapters through the SDK,
HTTP API, WorkstreamManager, and session factory to ChatSession. Discord
sessions now receive chat-optimized formatting (no tables, no Mermaid,
concise output) instead of the web UI's rich markdown instructions.
* fix: address CI failures and Copilot review feedback
- Add client_type param to CLI session_factory (mypy protocol match)
- Add prompt policy CRUD to PostgreSQL backend (test-postgres CI)
- Fix ClientType resolution: compare against enum values, not members
- Fix null client_type coercion (body.get returns None, not "")
- Use local time with astimezone() instead of UTC with local tz name
- Sanitize tool_gate in update endpoint (coerce null to empty string)
* feat: replace console HTTP polling with persistent SSE streams
Console collector now subscribes to each server node's /v1/api/events/global
SSE stream for real-time state updates instead of polling /v1/api/dashboard
and /health every 15 seconds.
Server changes:
- Emit ws_created/ws_closed events on global queue from create/close handlers
- Add node_snapshot on SSE connect (workstreams, health, aggregate)
- Add ?expected_node_id= identity verification (409 on mismatch)
- Add health_changed callback to BackendHealthMonitor circuit breaker
- Add periodic aggregate emitter thread (10s)
Console collector changes:
- Single asyncio event loop on one thread multiplexes all SSE connections
(scales to 1000+ nodes vs thread-per-node)
- Discovery loop spawns/cancels async SSE tasks per node
- Snapshot reconciliation on connect, delta application for live events
- Fix ws_state→cluster_state event type mismatch
- Remove polling code (poll_interval, max_poll_workers, --poll-interval CLI)
SDK changes:
- Add NodeSnapshotEvent, HealthChangedEvent, AggregateEvent dataclasses
- Add stream_node_events() method (async + sync)
* fix: address review feedback on node event streams
- Fix stop() to let SSE manager exit naturally instead of force-stopping
the event loop (ensures finally cleanup runs)
- Guard against empty/invalid SSE data from ping frames
- Treat missing node_id as identity mismatch (409) when expected_node_id
is provided
- Fix stale docstring on _update_metrics
* feat: show thinking indicators, tool calls, and results in Discord threads
Discord threads now surface real-time activity during multi-tool chains
instead of appearing idle. ThinkingStart/Stop events display a transient
italic status message. ToolInfoEvent sends a per-tool "running" embed
that ToolResultEvent edits in-place with the result (FIFO matching by
tool name, fallback to new message). Includes backtick-injection escaping
in tool output. Visibility respects auto-approve config so tool calls
always appear somewhere.
* fix: address review feedback on Discord action visibility
Delete thinking messages in unsubscribe/stale-route cleanup (not just
pop state). Sanitize tool-call previews (escape backticks, strip
mentions). Fix format_tool_result docstring re ellipsis line count.
Add regression test for triple-backtick escaping.
* fix: disable approval buttons on server-side resolution (timeout)
Handle ApprovalResolvedEvent in _on_ws_event to disable buttons and
grey out the approval embed when the server resolves the approval
externally (timeout, auto-approve from another client). Extract
disable_message_buttons helper from views.py so it works on a plain
Message (not just an Interaction).
* fix: reply with guidance when user DMs the bot directly
Non-reply DMs were silently ignored. Now sends a message directing the
user to /ask or @mention in a server channel.
* fix: address round 2 review feedback
- Show error items (policy-denied) in ToolInfoEvent unconditionally
- Match ToolResultEvent to ToolInfoEvent by call_id (deterministic),
fall back to name-based FIFO when call_id is absent
- Escape triple backticks before truncating in format_tool_result so
the 500-char limit holds after expansion
* fix: edit thinking message in-place instead of delete-and-recreate
ThinkingStopEvent now preserves the message for the next event to reuse.
ContentEvent seeds StreamingMessage with the thinking message so the
first flush edits it. ToolInfoEvent edits the thinking message into the
first tool embed. Eliminates the visible delete → gap → new message
flicker during thinking → tool call transitions.
* feat: separate tool call and result into distinct Discord messages
ToolInfoEvent sends a "running" embed (light grey, tool name + preview).
ToolResultEvent marks it "Done"/"Error" (color + title update) and sends
the result as a separate message. This gives clear lifecycle tracing in
chat-style threads where verbosity aids readability.
* fix: show running embed for all tools and remove redundant name prefix
ToolInfoEvent now shows a running embed for every tool regardless of
needs_approval — the running indicator and approval dialog serve
different purposes. Removes the needs_approval/auto_approve filter
that caused missing running embeds when tools were approved via
"Always Approve" or server-side auto-approve.
Also drops the redundant **name** prefix from format_tool_result since
the embed title already carries the tool name.
* fix: concise logging for SSE connection failures
Catch httpx.ConnectError/ConnectTimeout separately from the generic
exception handler. Logs url and error string instead of the full
httpx/httpcore stack trace, which is noise for expected transient
connection failures during node restarts.
* fix: address round 3 review feedback
- Pop _pending_approval_msgs on button click so ApprovalResolvedEvent
doesn't double-update the embed title (e.g. "Approved - Approved")
- Remove unused name/is_error params from format_tool_result — embed
title carries the name, embed color carries the error status
- Fix _disable_buttons docstring to mention title update
- Fix missing /v1 prefix on SSE endpoint URL (caused all SSE connections
to get text/plain 404 responses instead of event streams)
- Stop treating StreamEndEvent as session-terminal (it fires per-segment,
not per-workstream) so multi-turn conversations work in Discord
- Bail on 404 instead of retrying forever for gone workstreams, and clean
up stale routes from storage
- Check response status before iterating SSE events to avoid retrying
non-retryable upstream errors
- Default rebalancer.enabled to True so hash ring routing works without
manual ConfigStore setup
- Add one-shot cache refresh fallback on route endpoints to handle the
startup race between rebalancer and first routed request
- Remove stale "message queues" language from tagline
- Remove duplicated content covered by docs (governance details,
judge config, config.toml reference, health/rate-limit details,
monitoring metrics, tool table, multi-model config)
- Replace tool table with summary + link to docs/tools.md
- Add documentation index table linking to all doc pages
- Add architecture summary (single-node vs multi-node routing)
- Add component table for entry points
- Trim diagram table to most useful subset
- Consolidate quickstart section
README is now a concise landing page that directs to docs for
details, not a duplicated reference manual.
When TURNSTONE_ADVERTISE_URL is set (Docker deployments), the
_advertise_host variable was never assigned. The TLS upgrade path
tried to use it to construct the https:// URL, causing an
UnboundLocalError that made TLS init fail silently.
Fix: derive the TLS URL from _advertise_url (replace http → https)
instead of reconstructing from _advertise_host.
Address Copilot PR feedback:
- Exit with clear error if neither console_url nor server_url is
available after discovery (prevents cryptic failures downstream)
- Fix log field names: console → console_url, server → server_url
for consistency with other channel log events
Add token_factory parameter to SDK clients (_BaseClient, server,
console) — a Callable[[], str] invoked before each request to get
the current auth token. Supports ServiceTokenManager for auto-rotating
JWTs that re-mint transparently before expiry.
Channel gateway creates dual token managers:
- console-audience JWT for routing proxy calls (via AsyncTurnstoneConsole)
- server-audience JWT for direct SSE connections to server nodes
Both _request() and _stream_sse() inject the factory header per-call,
so long-lived connections get fresh tokens on reconnect.
Also adds TURNSTONE_CONSOLE_URL to console compose service for
DNS-resolvable service discovery.