Server SDK create_workstream: add initial_message, auto_approve_tools,
user_id, ws_id params (all optional, omitted when empty).
Console SDK: add auto_approve, auto_approve_tools, user_id to
create_workstream. Add 8 route_* methods for the routing proxy path
(/api/route/*): route_create_workstream, route_send, route_approve,
route_plan_feedback, route_close, route_cancel, route_command,
route_lookup. Sync mirrors for all.
Prepares for channel gateway and scheduler to use SDK clients instead
of raw httpx calls.
Move the consistent hash ring implementation (FNV-1a, virtual nodes,
bisect lookup) from code to docs/design/consistent-hash-ring.md as a
forward-looking reference for future scalability work.
The current rebalancer uses weight-proportional distribution (simpler,
exact splits, no hash variance). The ring algorithm is documented with
test vectors, stability properties, and a comparison table for when
the ring approach becomes advantageous (large clusters, decentralized
routing, cross-language determinism).
hash_ring.py retains: RING_SIZE, bucket_of(), RingNode, NoAvailableNodeError
(all actively used by router and rebalancer).
Replace the full-rehash algorithm (diff ideal vs current across all
65536 buckets) with a donor/recipient algorithm that only moves
buckets from overloaded nodes to underloaded nodes.
Key improvements:
- Adding node C to {A, B} only moves buckets TO C, never between
A and B. Previously the HashRing rehash could shuffle between
existing nodes.
- Seeding uses weight-proportional distribution instead of HashRing
virtual nodes, producing an exact split that doesn't trigger
immediate correction on the next cycle.
- Dead-node buckets are redistributed to the most underloaded
survivors, not rehashed across the whole ring.
- HashRing class is no longer used by the rebalancer (still
available for other uses like the Go rewrite reference).
The threshold check still gates live-to-live moves. Dead-node
recovery remains unconditional.
set_bucket_stat: single-upsert storage method replacing the N-loop
reconciliation in the rebalancer. Reduces DB round-trips from
|ws_delta| per bucket to exactly 1.
Console metrics: /metrics endpoint on the console exposing 6 routing
and ring metrics in Prometheus text format:
- turnstone_router_requests_total (method, status)
- turnstone_router_request_duration_seconds (method)
- turnstone_ring_membership_size
- turnstone_ring_version
- turnstone_ring_rebalance_total (result)
- turnstone_ring_migrations_total
Instrumented in route_create, route_proxy, route_lookup handlers.
Ring gauges updated on collector discovery loop. Rebalance/migration
counters recorded after each rebalancer pass.
When rebalancer.eager_migrate is enabled, the rebalancer POSTs
/_internal/migrate to source nodes after reassigning buckets,
triggering immediate workstream eviction instead of waiting for
lazy resume on the next request.
Only idle workstreams are eagerly migrated — active ones (running,
thinking, attention) are left alone to avoid disrupting in-flight
work. Failed migrations are logged and skipped (the lazy path
handles them eventually).
Rebalancer: daemon thread in the console process that maintains
bucket-to-node assignments in hash_ring_buckets. Seeds the ring on
first run (empty table → 65536 rows via consistent hash). Periodically
checks for membership changes and rebalances: moves cheapest buckets
first (empty > idle > active), respects imbalance threshold, reconciles
bucket_stats against actual workstream counts before each pass.
Uses DB-based leader election (rebalancer_lock in system_settings) for
multi-console deployments. Increments rebalancer_version after writes
so console routers refresh their caches.
Add 6 settings: ring.vnodes_per_unit, rebalancer.enabled/interval/
threshold/eager_migrate, node.weight.
Add /_internal/migrate endpoint on server for eager workstream eviction.
Add routing proxy endpoints to the console server:
- POST /v1/api/route/workstreams/new — hash-ring-routed create with
503 retry, target_node pinning, and node_url injection
- POST /v1/api/route/{send,approve,cancel,command,close} — generic
proxy to workstream owner via O(1) bucket lookup
- GET /v1/api/route?ws_id=X — node URL lookup for direct SSE
Wire ConsoleRouter into console lifespan (cache refresh on startup)
and collector discovery loop (version-based cache invalidation).
Add --console-url to channel gateway CLI for multi-node routing.
ChannelRouter routes control-plane through console when set, SSE
connections go direct to server nodes via node_url from create response.
HashRing: FNV-1a virtual nodes, immutable, computes ideal bucket-to-node
distribution. Used by the rebalancer (next commit) to seed and maintain
the assignment table.
ConsoleRouter: in-memory flat array of 65536 NodeRef entries loaded from
hash_ring_buckets table. O(1) routing via ws_id prefix. Supports
per-workstream overrides, version-based cache refresh, and targeted
ws_id generation.
Both are pure library code with no server integration yet.
Delete the entire turnstone/mq/ package (broker, bridge, protocol,
client) and turnstone/sim/ package. Remove Redis as a dependency.
Channel gateway and console now communicate with server nodes via
direct HTTP (httpx + httpx-sse) instead of Redis pub/sub and queues.
Single-node deployments work with zero infrastructure beyond the
database.
Key changes:
- Channel adapters use httpx POST for create/send/approve/close
and httpx-sse for per-workstream event streaming
- Console collector discovers nodes via services table instead of
Redis SCAN
- Console scheduler dispatches tasks via HTTP POST with DB-based
leader election
- Server registers in services table with 30s heartbeat
- Server accepts optional ws_id in create request (for Phase 2
console-generated routing)
- SDK events gain IntentVerdictEvent and OutputWarningEvent types
- All docs, examples, bootstrap wizard updated
63 files changed, -5968 net lines (Redis transport fully removed)
* fix: sync actual TLS state to ConfigStore on console startup
The console writes tls.enabled to the DB but never clears it when TLS
init fails or isn't configured. Server nodes read the stale DB value
and attempt TLS negotiation with a non-TLS console, producing noisy
SSL errors on every startup.
Console now syncs the actual TLS state after init: if TLS succeeded,
tls.enabled=true; if it failed or wasn't attempted, tls.enabled=false.
Server TLS failure log reduced from full traceback to one-line warning.
* fix: sync TLS state to ConfigStore on console startup
Console now writes the definitive TLS state to ConfigStore so server
nodes don't attempt TLS against a non-TLS console:
- TLS init succeeded → write true
- TLS not configured (DB false/unset) → write false (definitive)
- TLS configured (DB true) but init failed → don't overwrite
(transient failure shouldn't permanently disable)
Server TLS warning reduced to one line with exception type, full
traceback available at debug level.
* feat: auto-detect model changes when LLM backend swaps models
The BackendHealthMonitor already probes /v1/models every 30s but
discarded the response. Now compares the detected model against the
last known one and triggers a registry reload when it changes.
- Extract _extract_context_window() helper for reuse across
detect_model, probe_model_endpoint, and the health monitor
- BackendHealthMonitor: new provider/initial_model/on_model_changed
params; _check_model_change() fires callback on model swap
- Server: wire _handle_model_change callback that updates cli_model_args
and calls registry.reload(); guarded by _user_specified_model flag
so --model overrides are never auto-replaced
- Session: _refresh_model_from_registry() called at top of send();
two string compares when nothing changed, full re-resolve on swap
- 7 new tests for _extract_context_window and model change detection
* fix: address Copilot review on model re-detection
- server: update cli_model_args only after successful reload (not
before), add finally block for new_reg.shutdown(), guard against
cli_model_args not yet initialized
- session: wrap registry lookup in try/except for concurrent reload
race, reset judge on model change, recompute tool_truncation when
context_window changes in auto mode
Replace "You are an expert software engineer" with a grounded
narrative persona: a resident engineer on a focused team with real
tools, real code, and real consequences. Sets expectations about
boundaries, judgment calls, and working within constraints.
* fix: scope-filter memory list/search to current workstream and user
Unscoped memory(action='list') and memory(action='search') returned all
memories across all workstreams. Now applies the same 3-query pattern
(global + current workstream + current user) used by system prompt
injection.
* fix: validate user scope on memory search/list for unauthenticated sessions
Adds _validate_scope guard to search and list prepare paths, matching
save/get/delete. Prevents explicit scope='user' from returning all
user-scoped memories when session is unauthenticated.
* fix: update _get_visible_memories references to _list_visible_memories
* fix: defense-in-depth guard for empty scope_id on search/list
Copilot review: if scope is 'user' or 'workstream' with empty
scope_id, the storage query returns all memories in that scope
across all users/workstreams. The prepare step already validates
via _validate_scope, but add exec-level guard to reject scoped
queries with empty scope_id as defense-in-depth.
* fix: detect context window from vLLM max_model_len field
vLLM exposes the context window as max_model_len on the model object,
not meta.n_ctx_train (llama.cpp format). Both detect_model() and
probe_model_endpoint() now check max_model_len first, falling back
to meta.n_ctx_train for llama.cpp. Fixes 32768 fallback on vLLM
servers that report 262144+ token context windows.
* test: add vLLM max_model_len detection tests
Copilot review: new vLLM context window path had no test coverage.
Add tests for probe_model_endpoint (max_model_len detected, preferred
over meta.n_ctx_train) and detect_model (vLLM model object with
max_model_len).
The first message sent from Discord was silently dropped because the
cog delegated the initial message to the bridge via CreateWorkstream-
Message, but the bridge published response events to the per-workstream
Redis pub/sub channel before the Discord bot had subscribed to it.
Redis pub/sub is fire-and-forget — events with no subscribers are lost.
Fix: create the workstream with initial_message="" (no delegation),
subscribe to the per-workstream event channel, then send the message
through router.send_message() — the same path the second message
already uses successfully.
Applied to both @mention handler and /ask slash command.
* feat: per-workstream status bar above input
Move the global token counter and model name from the header into a
per-pane telemetry strip between messages and the text input. Each
workstream pane now independently shows model name, token usage with
context percentage, tool calls this turn, and turn count.
Backend: add _ws_turn_tool_calls counter (reset per user turn, emitted
in SSE status event alongside turn_count). MQ bridge forwards the new
fields. SDK and TypeScript types updated.
Frontend: build .ws-status-bar DOM in _createDOM, rewrite updateStatus
to target per-pane elements, update SSE connect/disconnect handlers.
Remove #model-name and #status-bar from global header. Restore console
#status-bar CSS in its own stylesheet.
Accessibility: aria-atomic, aria-labels on each field, warning symbols
(▲/⚠) at 80%/95% context for color-blind users, placeholder text
before first status event. Disconnect state uses 2px red border with
dimmed stale fields.
* fix: emit status event on SSE connect so status bar populates on resume
When resuming a workstream, the event_generator only sent connected +
history events. The status bar stayed at placeholder values until the
next LLM response. Now replays session._last_usage as a synthetic
status event right after connected, so token count, tool calls, and
turn count render immediately.
* fix: address Copilot review — remove dead function, clarify locals
Remove updateHeaderForFocusedPane() and its call site (no-op since
status moved per-pane). Rename ambiguous ttc/tc locals to
turn_tool_calls/turn_count in the status replay block.
* feat: add memory get action, reduce search/list preview to 200 chars
search and list truncated memory content to 500 chars with no way to
read the full value. Two changes:
- New 'get' action retrieves a single memory by name with complete
untruncated content. Searches scopes narrowest-first (workstream
→ user → global).
- search/list previews reduced from 500 to 200 chars now that get
exists for full content. Both append a hint:
"Use memory(action='get', name='...') for full content."
Includes get_structured_memory_by_name wrapper in memory.py and
4 tests.
* Update turnstone/tools/memory.json
* fix: include 'get' in _prepare_memory docstring and invalid-action error
PostgreSQL text fields cannot store NUL (0x00) bytes, and SQLite
stores them but they cause downstream issues (API payloads, web UI).
Add sanitize_text() to _utils.py and apply it in both backends'
save_message to content and provider_data fields.
* fix: drop orphaned tool_results with no matching tool_use in _convert_messages
The context window increase from 200K to 1M for Claude 4.6 means
conversations that previously triggered auto-compaction now send their
full history. Older messages with orphaned tool_results (from
pre-fix cancels or compaction boundaries) are now visible to the API,
causing "unexpected tool_use_id in tool_result blocks" errors.
The existing repair code handles orphaned tool_use (synthesizes
missing results), but not the reverse. Now validates each
tool_result against the preceding assistant message's tool_use IDs
and silently drops results with no match.
* fix: filter empty IDs from prev_tool_use_ids, document pass-through
Code review: empty-ID tool_use blocks were added to the filter set,
and the intentional pass-through when prev_tool_use_ids is empty
needed documentation.
* fix: block math sandbox escape via getattr/setattr/type reflection
getattr() with runtime-constructed strings bypassed the AST validator,
allowing full os/subprocess access from the sandboxed math tool via
module.__builtins__['__import__']('os').
Three-layer fix:
- Block getattr, setattr, delattr, type, __import__ in
_MATH_BLOCKED_BUILTINS (prevents direct calls)
- Add AST validation for getattr/setattr/delattr call nodes
(catches them even if builtins dict is bypassed)
- Strip __builtins__ from all pre-imported modules in the execution
namespace (runtime defense — even if AST is somehow bypassed,
module.__builtins__ returns empty dict)
Normal math, sympy, numpy, scipy operations unaffected.
* fix: harden _safe_import to strip __builtins__ from runtime imports
Copilot review: modules imported at runtime via _safe_import still
had their original __builtins__ dict, accessible via
operator.attrgetter('__builtins__'). Now _safe_import strips
__builtins__ from every module it returns. Also blocks
operator.attrgetter/itemgetter at the AST level, and removes the
redundant duplicate getattr check in visit_Call.
* fix: add type ignore for module __builtins__ assignment
* fix: block /proc/*/environ access in bash filter and judge heuristic
/proc/1/environ leaks the full server environment including DB
credentials, API keys, and JWT secrets. Env scrubbing in env.py
only affects subprocess calls, not procfs reads.
- Add /proc/1/environ and /proc/self/environ to BLOCKED_PATTERNS
in safety.py (hard block)
- Add proc-environ-exfil heuristic rule at critical severity with
deny recommendation (catches /proc/<pid>/environ patterns)
* fix: move proc-environ-exfil rule to _CRITICAL_RULES list
Copilot review: rule had risk_level=critical but was placed in
_HIGH_RULES. Move to _CRITICAL_RULES for consistency with the
first-match-wins severity ordering.
The intent judge was receiving up to 50% of the context window in
conversation history (FIFO from end), which grows linearly with
conversation length and causes increasing latency. The judge only
needs the immediate request context to evaluate a tool call's safety.
Now trims to messages from the last user message onward before
applying the FIFO budget cap. Keeps the user's request, the
assistant's response with tool calls, and any recent tool results
while discarding earlier conversation that isn't relevant to the
current intent evaluation.
Claude 4.6 (Opus + Sonnet) unified on 1M token context windows.
Update capabilities table from 200K to 1M for both models. Remove
claude-opus-4 and claude-sonnet-4 entries (end of life). 4.5 models
remain at 200K. Default fallback stays at 200K for unknown models.
* fix: distinguish user cancel from crash in bash tool results
When a user cancels a running bash command, the process is killed
with SIGKILL (exit code -9). Previously this showed as an error,
causing the model to retry. Now checks cancel.is_set() after proc
exit and returns "Cancelled by user." as a non-error result so the
model knows to stop rather than retry.
* fix: use -signal.SIGKILL instead of magic -9
Copilot review: replace hard-coded -9 with -signal.SIGKILL for
clarity. Popen.returncode is negative of signal number when killed.
* feat: add stop_on_error param to bash tool for set -e behavior
New boolean parameter enables 'set -e' in the bash preamble so
multi-step scripts exit on the first command failure instead of
silently continuing. Default false (existing behavior preserved).
pipefail remains always-on.
* fix: strict bool parsing for stop_on_error, treat exit 1 as error with set -e
Copilot review: bool("false") is True — use `is True` for strict
JSON boolean parsing. Also, with stop_on_error enabled, any non-zero
exit code is now treated as an error (set -e means the script halted
on failure), whereas without it exit code 1 remains benign.
* fix: synthesize cancelled tool results instead of stripping turns
When a user cancels during tool execution, the model previously lost
all context about what was attempted (assistant message + tool_calls
stripped entirely). Now synthesizes tool_result messages with
is_error=true and "Cancelled by user." content for any tool_calls
that lack matching results. This keeps the conversation valid for
both providers while preserving the full tool call structure so the
model knows what was tried.
Also applies to KeyboardInterrupt with "Interrupted by user." text.
* fix: persist synthesized cancel results to DB, assert is_error in test
Copilot review: synthesized tool messages were in-memory only,
creating a mismatch with DB that could break rewind/retry. Now
calls save_message() for each synthesized result. Also adds
is_error=True assertion to the cancel test.
* feat: add pagination and longer content to recall tool
- New offset parameter for paginating through recall results
- Content preview increased from 500 to 2000 chars per match with
total length indicator when truncated
- Output passed through _truncate_output for consistency
- OFFSET clause added to SQLite (FTS5 + LIKE) and PostgreSQL
(tsvector + ILIKE) search queries
* fix: defensive int coercion for recall offset/limit
Copilot review: offset/limit could arrive as null, float, or other
non-int types from JSON. Coerce with int() + try/except in prepare,
and int() at the storage layer before binding into SQL OFFSET/LIMIT.
* feat: add diff_file tool for comparing files and content
New read-only tool that shows unified diffs between two files or
between a file and provided content. Useful for verifying edit_file
changes and comparing file versions. Auto-approved (no side effects).
Configurable context lines (default 3). Available to task agents.
* refactor: extract _read_text_lines helper, share across read_file and diff_file
Copilot review: diff_file duplicated file-loading and lacked binary
detection. Extract _read_text_lines() that handles realpath
resolution, null-byte binary detection, and error handling. Used by
both _exec_read_file and _exec_diff for consistent behavior.
* fix: address code review — agent flag, resolved shadowing, read_files
- Add agent: true to diff_file schema so plan agents can use it
- Fix resolved variable shadowing in _exec_read_file (use _ for
unused return from _read_text_lines)
- Register diffed files in _read_files so edit_file read guard
is satisfied after diff_file
- Move difflib import to module level (stdlib, no lazy-load needed)
- Fix description wording ("provided string" not "previous version")
* fix: stream diff with early cutoff, expand paths before header
- Stream difflib output and stop collecting after tool_truncation
chars to avoid large intermediate allocations on big diffs
- Expand paths with expanduser before building the approval header
so display matches actual execution paths
* docs: tool descriptions, bash timeout param, multi-line preview
- task_agent/plan_agent: document the tool subset limitation (no
memory, recall, watch, skill, or further delegation)
- bash: add per-call timeout parameter (1-600s, defaults to 120s),
shown in approval header when specified
- bash: show full command in preview for multi-line scripts so the
approval flow displays the complete command, not just the first line
- bash: document 256KB output cap and stderr prefix in description
* fix: address Copilot review on tool descriptions
- bash: say "truncated" not "256KB" (limit is configurable), document
timeout clamping range (1-600) and global fallback
- bash preview: fix "1 more lines" → "1 more line" singular
- plan_agent: remove bash from listed tools (not in AGENT_TOOLS)
* fix: improve memory save error message, narrow dd command filter
Two minor fixes from harness shakedown:
- memory save: split "both name and content required" into separate
errors for missing name vs empty content
- bash safety: replace blanket "dd if=" block with targeted patterns
for writes to block devices (of=/dev/sd*, /dev/nvme*, /dev/disk/,
etc.) and redirects to the same. Legitimate dd use like generating
test data or benchmarking reads is no longer blocked.
* fix: generalize > /dev/sda redirect pattern to > /dev/sd
Copilot review: only /dev/sda was blocked for redirects while
/dev/sdb, /dev/sdc etc were not. Generalize to match any /dev/sd*
device, consistent with the of= patterns.
* feat: edit_file replace_all, write_file append mode, search match count
Three tool enhancements from harness shakedown feedback:
- edit_file: new replace_all parameter replaces all occurrences of
old_string instead of requiring a unique match. Cannot combine with
near_line or edits array.
- write_file: new mode parameter with "append" option. Appends
content to end of file instead of truncating.
- search: output now includes a summary footer showing total match
count and file count (e.g. "47 matches across 12 files").
* fix: address Copilot review on tool enhancements
- replace_all: skip multi-occurrence rejection in pre-validation so
the feature actually works; show occurrence count in preview
- write_file mode: coerce non-string types safely via str()
- search footer: append before truncation to respect output limits
- edit_file error: mention replace_all as alternative to near_line
read_file silently converted null bytes to spaces, showing corrupted
content with no warning. Now samples the first 8KB for null bytes and
returns a clear error directing the user to bash for binary inspection.
* fix: memory delete searches all scopes when scope not specified
Previously delete defaulted to scope=global, so deleting a
workstream-scoped memory without explicitly passing scope=workstream
silently failed. Now tries narrowest scope first (workstream → user
→ global) and deletes the first match. Explicit scope still honored
when provided.
* fix: reject invalid scope on memory delete instead of silent fallback
Copilot review: invalid scope values were silently treated as
unspecified, which could cause accidental deletion from the wrong
scope. Now returns a clear error listing valid scopes.
* fix: exclude build/vendor/VCS directories from search tool
grep -rn recursed into .git, node_modules, target, __pycache__, etc.
producing hundreds of noise hits from generated content. Add
--exclude-dir flags for common directories that should never appear
in search results.
* fix: glob egg-info pattern and add vendor exclude
Copilot review: .egg-info misses turnstone.egg-info (named dirs),
use *.egg-info glob. Also add vendor to the exclude list.
Agent workflows need git for version control, curl for raw HTTP
requests, jq for JSON processing, and man/info for documentation
lookup. All were missing from the slim base image, leaving the man
tool non-functional and standard dev workflows broken.
* fix: block IPv6 loopback/link-local/private in SSRF filter
check_ssrf used gethostbyname which only resolves IPv4. IPv6 addresses
like ::1, fe80::, fd00:: bypassed the filter entirely. Switch to
getaddrinfo which resolves both address families and check all results.
* fix: handle IPv4-mapped IPv6 and zone IDs in SSRF filter
Copilot review caught two bypasses: ::ffff:127.0.0.1 (IPv4-mapped
IPv6) wasn't normalized before private/loopback checks, and fe80::1%lo0
(zone ID suffix) caused a ValueError that was silently swallowed.
Now normalizes IPv4-mapped addresses and strips zone IDs before parsing.