mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
main
28 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
480a1426b3 |
Fail-closed history-commit handoff (#1005)
* fix(session): fail-closed history-commit handoff (#981) The deleted-workstream discovery is now a terminal, ws_id-keyed latch: keyed conversation commits refuse admission once the durable parent is gone (convergence finalizers and force-abandon are exempt), history handoff refuses to mint a proof token so /history fails closed with a 503 instead of silently wiping the pane, and the SSE stream carries a workstream_gone resync reason. Discarded commits leave a forensic log of commit keys and roles, never content. Conversation rows gain a commit_key (migration 071): keyed saves are idempotent under retry, validated against the full commit identity, and refused when they would cross a workstream deletion. The prune orphan category now requires a NULL alias plus a two-hour updated grace, with cutoffs computed at discovery time and carried into both dialects' rechecks. The mid-turn interjection queue is owner-partitioned with no per-site mode flags: pops take the acting principal's and unowned rows, other participants' rows are structurally retained, and enforcement lives at queue admission plus the shared before_spawn gates. The retraction ledger is bounded by open pop windows: pops open a window atomically with the queue delete, restores close their ids atomically with the ledger consume, every other exit closes through one helper, and misses for unheld ids record nothing. The workstream-gone latch refuses unattended wakes at all three gates (watcher spawn, claim, delivery pre-pop), and the retry dispatcher regained its pre-envelope cancel/error convergence net. Persistence-state reporting derives through the session bound to each UI instead of a registry lookup by id that failed open to healthy during tombstone retention. The dashboard roster no longer re-inserts ghost entries from trailing activity events, the history tool-outcome scan tolerates interleaved non-turn rows, and the shared handoff-deadline handle owns its own retirement. Single-sourced across call sites: keyed-commit row values, attachment save wrappers, tail-truncation and conflict-resolution bodies for both storage dialects; worker-slot lifecycle field sets; the direct-commit admission frame; queued-row layout accessors; the string-aware comment stripper shared by every JS harness suite. Refs #981 #964 * fix(session): sweep handoff fixes to their sibling surfaces The interactive replay loop treated a system row as a tool-batch boundary, so every tool result after an interleaved row vanished from that pane while the coordinator rendered the same history correctly. Only a conversational turn ends the batch window now, matching the shared outcome index. Accepted user turns clear the composer's attachment chips on the same viewer policy that settles optimistic bubbles rather than on having matched a local bubble, so a workstream created with an upload no longer keeps a chip for an attachment the create dispatch already consumed. The coordinator's raced-Stop arm emits the stream-end hook it inherits alongside the idle state, leaving no unfinalized bubble or unflushed tool output. Ending a session surfaces a failure toast when the request never lands or answers with a non-JSON body. The per-second persistence reconcile now probes each session without blocking: a workstream whose generation and handoff locks are held is skipped until the next pass instead of contending the locks every commit needs. The one-shot repair that gates workstream creation at capacity keeps a definite probe — it has no next pass, and the sessions likeliest to be contended are the ones whose unresolved journals emptied its candidate list. Single-sourced: the attachment lane builds its conversation row through the shared commit-identity builder; the ordinary worker exit releases its slot through the lifecycle owner; both operator surfaces snapshot their counters through one non-consuming helper; the replay preamble loses its per-kind wrappers and its config hook; the browser harness suites share one brace walker; and each in-flight history attempt is one record carrying both its abort controller and its deadline. Refs #981 #964 |
||
|
|
7a06f5e8bc |
refactor(session): make ModelLane the provider boundary (#979) (#989)
* refactor(session): make ModelLane the provider boundary (#979) ## Summary This closes the model-lane ownership gap left by #832: `ChatSession` no longer stores raw provider/client handles. `ResolvedModelBinding` now carries the provider, client, model, capabilities, registry generation, and backend-auth configuration as one coherent snapshot. - Atomically rebind existing sessions after model-registry changes while pinning each in-flight send, fallback, judge, output guard, task agent, title, compaction, perception, and voice operation to its initiating principal and binding. - Fence UI publication, canonical trajectory folds, durable writes, streams, retries, child scopes, and judge work by generation. Stop can hand off to a successor without accepting late state; cancelled tools retain typed effect receipts, and concurrent approval batches resolve by exact cycle or call. - Make create, fork, open, close, and delete race-safe with hidden `creating` reservations, incarnation-aware state tails, and an ACL-rechecked transaction that clones checkpoint-bounded history, configuration, project/persona state, and attachment references. - Extend REST/OpenAPI and Python/TypeScript SDK contracts for create/fork inputs, routed-create metadata, live-workstream probes, targeted approvals, and structured cancellation results. - Update architecture, storage, authentication, judge, channel, console, API, and SDK documentation, including regenerated architecture diagrams and OpenAPI artifacts. ## Validation - SQLite suite: 11,188 passed, 9 skipped, 10 deselected - PostgreSQL suite: 11,195 passed, 2 skipped, 10 deselected - Live backend: 3 passed - SSE recovery: 6 passed; browser recovery harness passed all scenarios - Ruff: clean; 595 files correctly formatted - mypy: 243 source files clean - TypeScript: typecheck/build and 35 tests passed - OpenAPI artifacts fresh; all 14 changed diagrams reproduce byte-for-byte - `git diff --check` and Git LFS integrity clean Closes #979. * fix(deps): update nanoid for GHSA-2v37-7h3g-55p8 Refresh the transitive lock entry admitted by PostCSS so the TypeScript security gate no longer resolves the vulnerable custom-generator implementation. Validation: - npm ci - npm audit --audit-level=moderate: 0 vulnerabilities - TypeScript typecheck and build - TypeScript tests: 35 passed * fix(test): assert canonical model registry URLs Replace prefix checks with exact canonical base URL assertions so the tests do not model incomplete URL validation. Validation: tests/test_model_registry.py (185 passed); Ruff check/format; mypy. |
||
|
|
e010124008 |
feat(preview): rich preview pane + open_preview tool
Tool results only ever rendered as plain text in the transcript. This
adds the model-driven rich-preview lane every comparable surface has,
in turnstone's developer-tool idiom: a preview pane that opens BESIDE
the conversation, keyboard-operable, sandboxed, never replacing the
transcript that spawned it.
Backend
- New built-in open_preview(target, kind?, title?): resolves an http(s)
URL, a file path, or attachment:<id> to bytes; classifies into
web/pdf/image/table/text/markdown (magic bytes > MIME hint >
extension > UTF-8 fallback, legacy-charset pages transcoded); caps
size per kind; persists content-addressed with kind="preview" —
refcounted and GC'd with the workstream, skipped by trajectory
reconstruction so preview bytes can never materialize onto the wire.
URL targets gate like web_fetch (network egress); paths/attachments
run unprompted like read_file.
- New core.web.fetch_with_ssrf_guard: manual redirect walk that
SSRF-screens every hop BEFORE requesting it (follow_redirects=True
checked nothing between hops); adopted by both open_preview and
web_fetch. URL userinfo is stripped before the descriptor or the
stored bytes see it; <base href> is injected doctype-safely so
relative assets resolve without quirks mode.
- The preview descriptor rides the tool turn's meta side channel with
ONE shape on every boundary: the live tool_result SSE event, the
conversations.meta column, and the /history projection. Cancelled
batches commit an already-announced preview (blob + meta) instead of
stranding the open pane on a permanent 404.
- New GET {ws}/attachments/{id}/preview (read scope, same ownership
gate as /content) serves the STORED type with per-MIME hardening:
bare CSP sandbox for text/html (renderable, scriptless, opaque
origin), no CSP for application/pdf (Chromium's viewer refuses
sandboxed contexts), full default-src 'none' otherwise; filenames
fold to latin-1-safe ASCII. The console /node proxy now forwards
CSP/nosniff/disposition/cache-control instead of dropping them.
- History loads exclude preview blobs from the bulk content fetch at
the query (they were read and discarded on every load).
Frontend
- New "preview" pane type registered in the shared shell (server +
console): openPaneBeside placement, per-kind renderers — fully
sandboxed iframe for pages, browser PDF viewer, sortable tables
(CSV/TSV/JSON, ragged-file safe, 5k-row cap), rendered markdown,
text — plus back/forward history with arrow keys, reload persistence
via pane meta, and backoff auto-retry (0.9s..7.2s) bridging the gap
between the live descriptor and the batch fold that commits its blob.
- Tool results carrying a descriptor render a credential-redacted
preview chip (the reopen + replay affordance); live results auto-open
the pane only while the originating pane holds focus.
Docs: docs/tools.md + prompts/tools.md. Tests: policy unit tests, tool
prepare/exec (mocked fetch), serving route + proxy header pass-through,
storage exclusion on both backends, cancel-path commit, JS static
guards; a headless-Chrome harness drives the real module graph (32 DOM
assertions).
|
||
|
|
51ed336989 |
fix(storage): survive oversized rows in postgres history search
to_tsvector was computed inline over full row content, so one row whose tsvector exceeds PostgreSQL's 1MB limit aborted every search_history scan. Cap the FTS input at 250K chars (worst-case tsvector expansion stays under the limit; giant rows remain findable by their head). The ILIKE fallback also never ran on postgres: the failed statement leaves the autobegun transaction aborted, so roll it back before falling back. |
||
|
|
7f20b1bc84 |
fix(session): durable shared-workstream state + fork sender persistence
- _known_senders/_shared_workstream are now monotonic: union-only growth, latched shared flag, seeded once per workstream from a full-history distinct-sender read (new StorageBackend.list_message_senders) so compaction narrowing the resumable slice can no longer forget participants (duplicate join notes) or flip the banner back to single-user framing (prompt-prefix cache churn). - Recompute is memoized per turn (invalidated on stamped user-turn append); system-prompt composition no longer pays an O(n) trajectory scan on every recompose. - resume() resets the state: the monotonic guarantees are per workstream, not per session object. - resume(fork=True) bulk-persist now carries the user-turn sender stamp into the fork's meta column (was: _source_meta only, which dropped attribution for every forked user turn on reopen). |
||
|
|
2169559d6e |
feat(projects): governed project containers — memory scope, grouping, manage UI (#724)
* feat(projects): governed project containers — memory scope, grouping, manage UI
A workstream can attach to a project: a first-class, shareable resource
container that owns a `project` memory scope, groups conversations, and is
managed from the console.
Storage / migration 062: projects + project_members tables, workstreams.
project_id, and the memory type default project→general; grants
project.{create,read,write,delete} (admin-default).
Recall + writes: project memory is recalled iff the workstream is attached AND
the user has access (owner ∨ member ∨ public-for-read), resolved once at session
construction; coordinators recall it too. New saves default to the project when
attached + writable; the save and delete paths are write-gated; deleting a
project purges its scoped memory; archived projects aren't recalled.
Access = RBAC capability ∧ per-project ACL (auth.resolve_project_access, a
single-fetch resolver); visibility changes, member management, and delete are
owner-only.
API: project CRUD routes on both the server and console; project_id threaded
through workstream creation, spawn inheritance, the cluster-create proxy, the
dashboard / snapshot / coordinator row builders, and the collector deltas.
UI: a project picker with an inline "+ New project" creator in every creation
box (console launcher + standalone dialog + dashboard); group-by-project in the
rail; a project badge in the composer and on dashboard rows; a console manage
tab (list + create/edit + members shelves). The admin Memories view gains
coordinator/project scope filters and human scope labels (name, not hex). The
memory tool schema documents the project scope and the attach-aware default.
* fix(projects): client refresh hardening, creator race guard, SDK project_id
Addresses PR #724 review feedback plus two bugs found while validating it.
- projects.js refreshProjects: a non-OK status (e.g. 403 when the caller
lacks project.read) or a network/parse error no longer blanks the cache
or masquerades as "no projects" -- the prior cache is preserved, the
failure is recorded (new projectsError()) and warned. Honors the
long-standing "a transient error can't blank the rail" docstring.
- projects.js _fp: the fingerprint separators were raw control bytes,
which made git treat the whole file as binary (no reviewable diff).
Rewritten as escape sequences instead of raw bytes -- behavior is
byte-identical at runtime.
- project_creator.js: createProject() could reject unhandled (authFetch
throws on network/401; r.json() throws on a non-JSON body), leaving the
widget stuck busy/disabled. Added a .catch, plus a generation guard so a
create whose widget was cancelled/reopened mid-flight drops its result
instead of selecting a project the user backed out of.
- types.ts: add project_id to CreateWorkstreamRequest / WorkstreamInfo /
DashboardWorkstream to match the server schemas (was SDK-invisible).
- test_project_api.py: move side-effecting HTTP calls out of asserts so
the requests run even under python -O.
* fix(projects): JSON.stringify the cache fingerprint, drop control-byte separators
_fp joined fields/rows on raw NUL/SOH bytes, which made projects.js read as binary to git. Replace with a collision-proof, escape-free JSON.stringify encoding -- same change-detection semantics, zero embedded control characters.
|
||
|
|
164f74dead |
feat(operator-context): deliver structured per-kind meta to the UI
Operator-context system turns (watch results, output-guard findings, idle children, user interjections) carried their kind (_source) and a flattened text content, but the structured per-kind fields were dropped at every persist/deliver boundary — so the UI rendered every kind as one generic operator bubble and the structured watch-result card was lost. Wire the structured meta through as the single source of truth: - Storage: new conversations.meta JSON column (migration 060); threaded through save_message/save_messages_bulk (facade + protocol + both backends) and rehydrated in reconstruct_turns onto Turn.meta.extra["source_meta"]. - Canonical: make_system_turn carries meta as one _source_meta dict; turn_from_dict/turn_to_dict bridge it to/from Turn.meta.extra. - Live + history: widen on_system_turn(content, source, meta) across all impls + the SSE payload; surface _source_meta -> meta in the /history projection. SDK HistoryEvent docs note the field. - Producers derive both the model-facing content text AND the card from one meta dict, so they cannot drift: render_output_guard_text, build_watch_ reminder carrying output, idle_children and user_interjection metadata. - Frontend: addSystemContext / renderSystemTurn dispatch by source to the watch-result, guard-finding, idle-children, and queued-message cards in both the interactive and coordinator panes; every untrusted field renders via textContent. The meta is a leading-underscore key, stripped before the wire (sanitize_ messages and the native mid-conversation path copy only role+content), so the per-provider wire payloads stay byte-identical. Additive column, no backfill: operator turns predating it reload as plain text bubbles. |
||
|
|
0f1d03755a |
refactor(storage): drop dead ws_id/user_id from workstream_attachments
The blob store is global content-addressed — identical bytes dedupe across workstreams and users, so the per-tenant ws_id/user_id scope columns are dead: nothing reads them, and a committed blob is authorised via the conversations.attachments ref-list (attachment_referenced_in_ws), not a row scope. Drop both columns and idx_ws_attachments_ws_id from the schema and from the save_attachment signature (protocol + both backends + memory wrapper + the caller); fold the column/index drops and their downgrade into the unshipped migration 060. tool_name stays: it is a live denormalised search label (search_history → recall + /history), not trajectory data — "never rehydrated" held only for the wire path, which already ignores it. |
||
|
|
a3f91657c3 |
refactor(storage): reconstruct as row→Turn; extract recover_trajectory
reconstruct_turns is the pure row→Turn deserialize: one positional unpack of
the row tuple, one Turn per row, no wire-validity correction. The scattered
per-role dict-building and the side-channel keys collapse into typed Turn
fields (native ← {producer,blocks}, source ← _source, …); the dead tool_name
column is unpacked but unused. recover_trajectory(turns) is the load-time
trailing-strip policy, lifted out as its own function (one of lowering's three
orphan policies).
reconstruct_messages stays the dict-returning facade for now —
dicts_from_turns(recover_trajectory? · reconstruct_turns) — so every consumer
is unchanged and byte-identical (verified across the storage + reconstruct +
export + wire-payload suites, 7129 green). developer collapses into
Role.SYSTEM (zero writers, wire-identical); a bare-dict provider_data (never a
real native shape — the lane is a block list) no longer round-trips, which the
storage test now reflects.
|
||
|
|
98cee4d20c |
feat(storage): content-addressed refcounted attachments + in-memory upload buffer
Replace the persisted pending/reserved/consumed upload lifecycle (and its orphan-sweep and per-user cap) with a content-addressed, refcounted blob store fronted by the per-node in-memory pending buffer: - Upload stages bytes in the buffer (keyed by sha256); send-commit drains the referenced handles, writes each blob content-addressed (INSERT-OR-IGNORE then refcount += 1, so a stored blob is born referenced and dedupes across messages/workstreams), and records the ordered conversations.attachments ref-list — the sole message->blob link. - reconstruct rebuilds inline image_url/document multipart content from the ref-list, role-agnostically (so tool-produced images via _exec_read_image now persist + rehydrate instead of being flattened to text and lost). Output shape unchanged. - GC is reference counting: delete_messages_after / delete_workstream decrement once per reference and prune a blob at 0; a deduped blob shared with a kept turn (or another ws) survives. - get_content for a committed blob is gated by reference-ownership (the requester owns a turn in the ws whose ref-list names the id), replacing the dropped ws_id/user_id scope. - Migration 060 re-keys legacy consumed attachments to their content hash, dedups, sets refcounts, writes the ref-lists, and drops message_id/reserved_*; pending legacy rows are dropped (pending now lives only in the buffer). Both backends symmetric; the reservation methods, cap, and orphan-sweep are removed across storage/facade/protocol/endpoints/coordinator. Wire harness byte-identical; full suite green. |
||
|
|
c80354880a |
feat(api): enrich saved-workstream list with model/skill/context fields
GET /v1/api/workstreams/saved returned only ws_id/alias/title/created/ updated/message_count — too little to drive the planned saved-list table redesign. Add seven fields, all sourced from already-persisted data (no migration): - state, kind, node_id: columns on the workstreams table - model_alias, launch_skill: from workstream_config via LEFT JOIN - child_count: COUNT of child workstreams via parent_ws_id - context_tokens: most recent usage_events prompt size for the workstream - context_ratio: context-window occupancy (context_tokens / model context window), computed in the handler so the NULL / zero-window cases stay explicit and identical across both storage backends context_window comes from a model_definitions join; aliases defined only in config.toml are absent there, so context_ratio degrades to 0.0 rather than reporting bogus occupancy. The Python SDK reuses the Pydantic model; the TypeScript SDK OpenAPI snapshot and hand-maintained interface are updated. Tests cover the new storage columns (including NULL-when-absent), the handler ratio math + zero-window degradation, and the SDK enriched round-trip. |
||
|
|
d675b237a3 |
feat(mcp): oauth schema + minimum admin form
Adds the data model and admin UI surface required by the OAuth-MCP flow.
Phase 2 of the per-user delegation initiative.
Schema:
- migration 049 creates mcp_user_tokens (PK user_id, server_name) and
mcp_oauth_pending (PK state, indexed by created_at)
- eight new columns on mcp_servers: auth_type ('none' / 'static' /
'oauth_user', NOT NULL DEFAULT 'static') plus six oauth_* config
fields and oauth_as_issuer_cached
- post-upgrade UPDATE normalises auth_type to 'none' for streamable-http
rows whose headers are NULL/empty/'{}'; stdio rows are left at the
'static' default (auth_type is HTTP-auth-only)
- _schema.py kept in lockstep with the migration so metadata.create_all
and alembic upgrade produce identical shapes
- mcp_user_tokens / mcp_oauth_pending TypedDicts in _protocol.py for
Phase 3/4 use (no CRUD methods yet)
Storage / API:
- create_mcp_server gains the eight kwargs across protocol + sqlite +
postgresql
- MCP_SERVER_MUTABLE picks up auth_type and the six text oauth_* fields;
oauth_client_secret_ct is intentionally NOT in the whitelist — Phase 3
will own ciphertext writes via a dedicated method
- McpServerInfo + Create/Update Pydantic schemas extended; oauth_client_secret
accepted as plaintext input but discarded (Phase 3 wires encryption)
Admin handlers:
- _parse_auth_type validates against {'none', 'static', 'oauth_user'} and
rejects empty / unknown values; shared between create and update
- when auth_type changes away from 'oauth_user', the oauth_* config
columns are explicitly nulled in the same UPDATE so the row stays
consistent
- _clean_oauth_text caps text fields at 512 chars (URLs at 2048) to bound
admin write surface
- _mask_mcp_secrets now masks oauth_client_secret_ct to '***' regardless
of reveal=true (write-only field)
- audit detail dict redacts oauth_client_secret if present
Frontend:
- new "Multitenant Authorization" fieldset on the MCP-server modal with
three radio buttons (None / Shared / Per-user OAuth 2.1)
- conditional OAuth subform: AS URL, registration mode (preregistered /
dcr; cimd is future), client ID, client secret, scopes, audience
- secret input is autocomplete=off and never round-trips on edit
- audience auto-populates from the MCP server URL on blur
- headers textarea hidden and submitted as {} when auth_type is 'none' or
'oauth_user' so flipping the radio cleans up server-side state
Tests: storage round-trip for the new columns, oauth_pending table smoke,
migration 049 upgrade/downgrade with stdio-vs-http normalisation, four
admin-API tests for auth_type validation and oauth_*-clear-on-flip-away.
Suite passes 5284 (matched pre-Phase-2 baseline 5267 + 17 new).
Stacks on Phase 0; no behavioural change for existing rows.
|
||
|
|
0e25bad94e |
fix(storage): address PR #457 review feedback
Three issues from the Copilot review on PR #457: 1. SQLite race in bulk_close_stale_orphans (Copilot): the SELECT-then- UPDATE flow doesn't re-apply the eligibility predicates on the UPDATE, so a row that gets touch_workstream-bumped (or set_state- transitioned) between the two statements would still be flipped to closed. Postgres dodges this via UPDATE...RETURNING (one atomic statement); SQLite needs the explicit re-application. Fix: rebuild the WHERE conditions list once, apply on both SELECT and UPDATE, then SELECT-back by ``state='closed' AND updated=now`` to get the accurate closed-id list. A row that became fresh between the two statements skips the UPDATE entirely. 2. SQLite IN-clause bind-parameter limit (Copilot): default 999 cap could be exceeded on a backlog reap (e.g. after a long outage). Chunked the candidate id list at 500 — same chunk size prune_workstreams (line 453) uses for the same reason. 3. Wall-clock-dependent test asserts (Copilot, two locations): the tests asserted ``updated > '2024-01-01T00:00:00'`` which is fragile on systems with skewed clocks or pre-2024 dates. Replaced with ``updated != stale_seed`` — captures the same intent (the value was bumped) without depending on wall-clock date. Two ``...``-as-no-op flags from github-code-quality were false positives — ``...`` is the standard Python idiom for Protocol method bodies and matches every other method in _protocol.py. No code change. |
||
|
|
0debc5d061 |
fix(session_manager): scope orphan reaper by services.last_heartbeat
Replaces the ``node_id == self_node_id`` orphan-scoping heuristic from earlier on this branch with liveness-based scoping using ``services.last_heartbeat``. The heuristic was wrong for the post-#384 world: PR #384 (refactor: replace hash-ring rebalancer with rendezvous hashing) deleted the rebalancer that used to keep workstreams.node_id pointing at a live node. Without it, ``workstreams.node_id`` is now stamped at create time and never updated, so in containerized deployments with dynamic hostnames a dead pod's rows have ``node_id`` matching no surviving service — they'd accumulate forever under the old heuristic. services.last_heartbeat is the same primitive the rendezvous router uses for routing. Reusing it here keeps reap scoping aligned with routing: dead pods' rows fall out of the live set after the heartbeat window and become reapable; alive pods' rows stay protected as long as they heartbeat. Mechanics: - ``bulk_close_stale_orphans`` parameter renamed ``node_id: str | None`` → ``live_node_ids: list[str] | None``. The WHERE clause becomes ``(node_id IS NULL OR node_id NOT IN live_node_ids)``. ``None`` skips the filter entirely (single-process / tests / operator backfill). ``[]`` treats every row as unprotected. - ``SessionManager.close_idle`` pass 2 calls ``storage.list_services(self._service_type)`` to enumerate live peers, passes their service_ids as ``live_node_ids``. ``_service_type`` is derived from ``self.kind`` (INTERACTIVE→"server", COORDINATOR→"console") via a module-level mapping — no constructor param, so production wiring can't miswire the kind/service_type pairing. - list_services failure → pass 2 is skipped this tick (conservative; never reap when liveness state is unknown). Pass 1 still runs. - ``workstreams.node_id`` with NULL value is always eligible — defends against ANSI ``NULL NOT IN (...)`` evaluating to NULL (not TRUE) and silently protecting orphans forever. - Migration 048 simplified to ``(kind, updated)``; the new query's ``NOT IN (small list)`` predicate against an unbounded-cardinality column doesn't index well, so leading ``node_id`` would just add write cost. Tests cover the live-services protection (own/dead/null cases), the empty-peers reap-all case, the list_services-failure conservative fallback, both kind/service_type pairings (interactive→"server", coordinator→"console"), and the combined live_node_ids + exclude_ws_ids filter matrix. |
||
|
|
fff7840de7 |
fix(storage): add bulk_close_stale_orphans + touch_workstream primitives
Two new methods on the StorageBackend Protocol, with implementations on both Postgres (UPDATE ... RETURNING) and SQLite (SELECT-then-UPDATE in one transaction). No callers yet — wiring lands in subsequent commits. bulk_close_stale_orphans(kind, cutoff, exclude_ws_ids, node_id=None) flips rows in BULK_CLOSE_STATE_VALUES (idle/thinking/attention/running) to closed when their updated timestamp is lex-older than cutoff. The node_id filter scopes the reap to a single node's partition — required for multi-node interactive deployments where each node only has authority over its own workstreams.node_id rows. Excludes loaded ids so the in-memory pass owns those. touch_workstream(ws_id) bumps updated without changing state. Used by the open() rehydrate path to defend against the orphan reaper clobbering a freshly-loaded row whose DB updated is older than the cutoff. Pure timestamp write is safe against concurrent close() because close still wins on the state column. BULK_CLOSE_STATE_VALUES is centralized in workstream.py so the two backend implementations and FakeStorage all agree; if a new transient state is added to WorkstreamState, deciding whether it joins this set is part of the change rather than an after-the-fact audit across three files. Storage tests (run against both backends via the conftest fixture) cover the kind/state/cutoff/exclude/node_id matrix plus touch_workstream. |
||
|
|
334edbd580 |
fix(server,console): kind filter on saved-workstreams + closed coords on landing (#380)
* fix(server,console): kind filter on saved-workstreams + closed coords on landing Two independent bugs folded into one hotfix: 1. Coordinators leaking into the interactive UI's "saved workstreams" sidebar. ``list_workstreams_with_history`` (SQLite + postgres) was kind-agnostic — every coordinator row with conversation history came back alongside interactive rows, and ``list_saved_workstreams`` serialized them uniformly with no kind field so the interactive UI rendered coordinators as regular interactive entries. Fix: add optional ``kind: WorkstreamKind | str | None = None`` kwarg on ``list_workstreams_with_history`` (storage protocol + both backends + the ``turnstone.core.memory`` helper). Pass ``kind=WorkstreamKind.INTERACTIVE`` from the /v1/api/workstreams/saved handler so the interactive surface only sees interactive rows. Default ``None`` preserves legacy all-kinds behaviour for any other caller that wants both. 2. Closed coordinators vanish from the console landing page. ``_coordinator_rows`` in console/server.py built dashboard rows exclusively from the in-memory ``CoordinatorManager`` registry, which pops rows on ``close()``. The persisted storage row stays (state='closed') but never reached the landing-page poller at /v1/api/cluster/workstreams?node=console. Fix: two-lane merge in ``_coordinator_rows``. The in-memory lane (manager) stays authoritative for live session state (model / model_alias / current state / tokens). A new persisted lane queries ``storage.list_workstreams(kind=COORDINATOR, user_id=uid, limit=200)`` and appends rows NOT already in the in-memory set — surfacing closed / error / deleted coordinators so the operator can still see them on the landing page. Ownership semantics unchanged — non-admin callers only see their own tenant, admin-bypass via admin.users/admin.roles honored on both lanes, empty-string defense-in-depth matches _check_row_owner_or_404. Tests: - tests/test_storage_sqlite.py — two new tests: kind filter excludes coordinators from the history list; string form of kind accepted (matches the memory.py forwarding shape). - tests/test_coordinator_endpoints.py — four new tests: - closed coordinators from storage surface alongside active ones. - in-memory row wins on ws_id dedup (live state authoritative). - persisted rows respect tenant filter (non-admin, admin bypass). - orphan rows (empty user_id) never leak to empty-sub callers. Gate: ruff + mypy + pytest -m "not live" (4315 passed) all clean. * fix(server,console): address Copilot review on PR #380 Three review comments folded in: 1. Tenancy leak in /v1/api/workstreams/saved — the handler called list_workstreams_with_history without a user_id filter, so any authenticated user could see every other user's saved workstream aliases / titles / names. Fix: - Add ``user_id: str | None = None`` kwarg to list_workstreams_with_history on the protocol + both backends (SQLite + postgres). Pushes the filter into SQL. - memory.py helper forwards the kwarg. - /v1/api/workstreams/saved reads ``_auth_scopes(request)``: a service-scoped caller gets cluster-wide visibility (None), a non-service caller with a blank ``sub`` returns an empty list, otherwise the SQL filter is scoped to the caller's uid. Matches the _visible_workstreams pattern used on /workstreams and /dashboard. 2. Loose type annotation on the memory.py helper — ``kind: Any`` tightened to ``WorkstreamKind | str | None`` so mypy catches invalid callers. WorkstreamKind was already imported in the module. 3. Brittle positional indexing in _coordinator_rows persisted-rows lane — ``row[10]`` for user_id encoded a column offset that would silently corrupt the projection on any future SELECT reorder. Drop the test-double fallback entirely; the storage-protocol contract already requires SQLAlchemy Row with _mapping, and every real caller (SQLite + postgres) provides it. Tests: - test_server_authz.py TestSavedWorkstreamsTenantScoping — four new regression tests covering: non-service caller sees only own rows, service scope sees cluster-wide, blank-sub non-service returns empty, and coordinator rows excluded even for service callers. Gate: ruff + mypy + pytest -m "not live" (4319 passed) all clean. |
||
|
|
553d73109b |
feat(coordinator): phase 5 — harness-test polish + wait_for_workstrea… (#378)
* feat(coordinator): phase 5 — harness-test polish + wait_for_workstream + judge fix
Closes the bug list surfaced by the 2026-04-17 coordinator harness test
plus the post-phase-4 wait_for_workstream ask, and folds in three
adjacent cleanups that landed in the same window. Tightens defense-in-
depth on the model-invoked mutating ops, fixes the LLM judge silent
no-op, kills the inspect-poll token burn, and rounds out a handful of
observability / docstring / spec gaps.
The session-factory pre-resolve at console/session_factory.py and
server.py was rewriting `judge.model` from an alias (e.g. `judge-mini`)
to the resolved underlying id (e.g. `gpt-5-mini`). IntentJudge then
checked `model_registry.has_alias(config.model)`, found nothing, and
fell back to the SESSION's provider/client with that bare model id —
silent `llm_fallback / "did not return a verdict"` whenever the
coordinator and judge alias resolved to different providers.
Pass the alias through unchanged; IntentJudge's existing alias-
resolution path picks up the matching client + provider. Validate
the alias exists so an obvious typo still surfaces, but don't replace
the model field.
Regression: `test_alias_uses_registry_provider_not_session_provider`
constructs an alias whose provider differs from the session's and
asserts the judge picks up the alias's provider/client/model;
`test_coordinator_tool_call_returns_llm_verdict_not_fallback` asserts
the verdict tier is `llm` (not `llm_fallback`) on the happy path.
New `cancel_workstream` tool (approval required, primary_key=ws_id) —
cancels in-flight generation, unblocks any pending approval / plan,
moves the child to idle, leaves the row in storage so a fresh
send_to_workstream lands cleanly. Re-uses the existing
`/v1/api/route/cancel` route + `route.cancel` audit namespace; no
new server endpoint.
`CoordinatorClient.cancel/close_workstream/delete/send` now enforce
a tenant guard inline (`_is_own_subtree`) — only the coordinator
itself or one of its own children is targetable. Foreign ids return
the same 404-shape inspect/wait_for_workstream use, so the model
can't distinguish foreign from missing (no existence oracle).
Defense-in-depth — the upstream node enforcement is the perimeter,
this is the second line.
`list_workstreams` advertised `state="deleted"` and an
`include_closed=true` that surfaced deleted rows. Hard-deletes
cascade the workstream + conversation rows out of storage, so
deleted is unreachable in normal operation. Doc-only fix; the
synthetic-test path that registers `state="deleted"` rows still
works (terminal-state filter still excludes them via
`_terminal_states = {"closed", "deleted"}` in list_children).
Documented that the 120s service-registry heartbeat window means a
node returned by list_nodes can drop out before a follow-up
`spawn_workstream(target_node=…)` lands — the spawn fails with "No
available node for routing" rather than falling back. Two-line
clarification on each tool. No code change (a code fallback is a
bigger discussion deferred to 1.6).
`close_workstream` accepts `reason`; the upstream server handler now
persists it to `workstream_config.close_reason` (capped at 512 BYTES,
sliced on UTF-8 not code points so a CJK / emoji-heavy payload can't
4× the documented budget). `CoordinatorClient.inspect()` reads it
and surfaces as `close_reason` in the result dict — only for
terminal-state children (closed/error/deleted) so the live-child hot
path doesn't pay a per-inspect DB round-trip.
Tests: server-side persistence covers success / no-reason /
length-cap / non-string / storage-failure / multi-byte-utf8 paths;
client-side surface covers terminal vs. live workstreams.
For idle children whose node-dashboard live counter is 0 (the live
block only surfaces in-flight token counters), fall back to
`SUM(prompt_tokens + completion_tokens)` from `usage_events` so the
inspect output reflects cumulative spend.
New `storage.sum_workstream_tokens(ws_id) -> int` on the protocol +
both backends. The fallback is folded INTO `_fetch_cluster_live` so
the merged live block (with persisted total applied) is what gets
cached — back-to-back inspects of an idle child amortize through
the existing 2s LRU cache instead of each firing a fresh aggregation.
`CoordinatorClient.list_skills()` now projects `allowed_tools` per
skill — capped at 20 with a `+N more` sentinel so a skill that
whitelists a wide MCP surface doesn't bloat the per-row payload.
Reads the existing `prompt_templates.allowed_tools` column; no
storage change. Coordinators no longer have to guess what tools a
skill brings.
`route_create` now sets `routing_strategy: "hash_ring" | "target_node"
| "resume"` on the spawn response so the coordinator's spawn
response (and the `spawn_workstream` tool output) carries why a
given node was chosen. 3 lines + 3 covering tests in
test_console_routing_proxy.py.
New coordinator tool `wait_for_workstream(ws_ids, timeout=60,
mode='any'|'all')` that absorbs the wait into a single tool call —
the model sees one call + one result regardless of how long the
children take. Kills the busy-poll inspect loop that burned 20+
turns on a 3-child fan-out.
Storage-poll loop with batched primitives —
`get_workstreams_batch` + `sum_workstream_tokens_batch` issue exactly
two storage calls per tick regardless of N. At the cap (32 ws_ids /
600s / 0.5s tick) that's ~2400 round-trips for a full wait, down
from ~38k under the naive per-id shape.
Validation single-source-of-truth: the client owns mode whitelist,
ws_ids dedup + cap, timeout coerce + clamp. The session preparer
is a thin pass-through that builds the header + dispatches; bad
input surfaces at exec time as a tool error via `result.get("error")`.
Tenant-isolation collapse: missing-row and cross-tenant cases both
return `state="denied"` so wait can't be used as an existence oracle
(matches the 404-mask contract `inspect` uses).
Prompt-side: tools_coordinator.md adds a `wait_for_workstream`
pattern + an explicit "PREFER wait_for_workstream OVER a loop of
inspect_workstream" line in the workflow-shape section.
Replaces the quote-bracketed substring LIKE/ILIKE pattern with proper
JSON-array containment. The previous shape effectively did
`LOWER(tags) LIKE '%"<lower-tag>"%'`, which broke for tag values
containing `"` (the JSON encoder escapes it to `\"` and the literal-
substring search misses), `\` (encoded as `\\`), or non-ASCII
characters that the encoder rendered as `\uXXXX`. Also exposed a
small spoofing surface — `tags=["foo\","bar"]` would have matched a
query for `bar`. Real-world tag values are alphanumeric+dash today
so it hadn't fired in production, but the fix is small.
- SQLite: `EXISTS (SELECT 1 FROM json_each(prompt_templates.tags)
WHERE lower(value) = lower(:tag))` (JSON1 extension; SQLite 3.38+).
- PostgreSQL: `EXISTS (SELECT 1 FROM jsonb_array_elements_text(
prompt_templates.tags::jsonb) AS jat(elem) WHERE lower(jat.elem) =
lower(:tag))`.
Three new tests prove the substring pattern was broken for
quoted / backslash / unicode tag values; the existing case-fold +
wildcard tests continue to pin the contract.
Phase 1 added the coordinator workstream API; phase 2 added only
`/open` to the OpenAPI catalog and missed every other coordinator
endpoint plus phase 3's `/children`, `/tasks`, and the
`/cluster/ws/{ws_id}/detail` aggregator. SDK consumers + operators
browsing `/docs` couldn't discover the surface. Doc-only addition:
12 endpoints + 9 new Pydantic models, all under the `Coordinator`
OpenAPI tag so /docs groups them together.
Sidebar re-fetches `GET /tasks` on every `task_list` `tool_result`
SSE event. A model that runs `add → list` (or any back-to-back
mutation pair) double-fetches the same envelope. Coalesced into
one fetch per 150ms window via a new `loadTasksDebounced` wrapper;
direct UI actions (refresh button, page load) keep calling
`loadTasks` directly so user clicks aren't delayed.
- `ruff check turnstone tests` — clean
- `mypy turnstone` — clean (157 source files)
- `pytest -m "not live"` — 4284 passed, 3 deselected (was 4226 on
main; +58 new tests across coordinator client, tools, judge,
storage, console routing proxy, server close-handler,
storage_skills_filtered, OpenAPI catalog, server close-reason
persistence)
- New tools added: 2 (cancel_workstream, wait_for_workstream) —
TOOLS count 28 → 30; coordinator subset 9 → 11; auto_approve adds
wait_for_workstream; primary_key adds cancel_workstream
- New OpenAPI endpoints: 12 (every phase-1/2/3 coordinator route +
the cluster-inspect aggregator)
- New storage protocol methods: 3 (sum_workstream_tokens,
sum_workstream_tokens_batch, get_workstreams_batch)
All phase 1 / 2 / 3 / 4 invariants preserved: COORDINATOR_TOOLS /
INTERACTIVE_TOOLS disjoint; coordinator sessions have no MCP surface;
list-style tools return {items, truncated}; route-proxy emits
route.<action> audit on 2xx; 404-mask on ownership failures; tenant
filters pushed into SQL; per-coordinator JWT carries scope context.
* fix(coordinator): address Copilot review on PR #378
Three valid Copilot findings on the wait_for_workstream surface:
1. ``wait_for_workstream.json`` description claimed the tool returns a
top-level mapping ``ws_id -> {state, tokens, updated}`` plus
elapsed/complete/mode at the same level, but the actual shape is
``{results: {ws_id: {...}}, elapsed, complete, mode}``. Description
now matches the implementation. Also adds ``deleted`` to the
advertised terminal-state list (it's in ``_WAIT_REAL_TERMINAL_STATES``;
the doc and runtime now agree).
2. ``CoordinatorClient.wait_for_workstream`` docstring listed
``idle / error / closed`` as the real terminal set but the constant
includes ``deleted``. Same fix — list ``deleted`` with a parenthetical
noting it's unreachable in normal operation (hard-delete cascades the
row).
3. Storage protocol docstring math: ``sum_workstream_tokens_batch``
claimed "from ~38k to ~1200" round-trips per wait at the cap, but
``wait_for_workstream`` issues TWO storage calls per tick
(``get_workstreams_batch`` + this one), so 1200 ticks × 2 = ~2400.
Updated to "~2400" with the math spelled out.
Also a clean rebase onto today's main (PR #377 — the rebalancer node_id
snapshot doc — landed since phase 5's last push). Single conflict in
``inspect_workstream.json`` resolved by keeping both notes (rebalancer
node_id binding semantics + the new ``close_reason`` surface from phase
5); ``spawn_workstream.json`` auto-merged.
The github-code-quality bot also flagged three items on
``_protocol.py`` asking to replace ``...`` with ``pass`` in Protocol
method bodies. Refuted: ``...`` is the canonical PEP 544 idiom for
Protocol method bodies and the rest of the file uses it consistently.
The bot's lint rule misfires for ``Protocol`` classes.
Verification:
- ``ruff check turnstone tests`` clean
- ``mypy turnstone`` clean (158 source files)
- ``pytest -m "not live"`` — 4308 passed, 3 deselected (no test count
change; pure doc/comment edits)
|
||
|
|
c397668d21 |
feat(coordinator): tree-view UI, cluster-wide live inspect, dashboard… (#370)
* feat(coordinator): tree-view UI, cluster-wide live inspect, dashboard grouping — phase 3
Closes out the 1.5 coordinator UX surface: a right-sidebar tree view at
/coordinator/{ws_id} showing spawned children + task list, a new
cluster-wide live inspect endpoint that powers the tree's live badges,
and 2-level dashboard tree grouping that nests spawned children under
their coordinator parent.
## Cluster-wide live `inspect_workstream`
New `GET /v1/api/cluster/ws/{ws_id}/detail` on the console, gated by a
new `admin.cluster.inspect` permission (unassigned to any builtin role;
operators opt in). Aggregates `storage.get_workstream` with a
short-timeout (2s) HTTP fetch against the owning node's
`/v1/api/dashboard`. Coordinator-hosted workstreams get their `live`
block from the in-process `CoordinatorManager` instead of a proxy hop.
Response shape `{persisted, live, messages}` — `live: null` on node
unreachability / 5xx / missing-entry with status 200 so the UI can
degrade gracefully without an error state. Correlation-id masks
unexpected exceptions. 404-masks cross-tenant reads (non-admin
callers see only their own workstreams).
`CoordinatorClient.inspect()` best-effort merges the `live` block onto
its storage snapshot so the model-facing `inspect_workstream` tool
gains a `live` key without any schema change. Model-facing tool
schema stays identical.
## Tree-view UI
New right sidebar at `/coordinator/{ws_id}` with a 2-level children
tree + the phase-2 task list.
Backend:
- New `GET /v1/api/coordinator/{ws_id}/children` returns
`{items, truncated}` — identical row shape to the `list_children`
tool — filtered via `storage.list_workstreams(parent_ws_id=..., kind=None)`.
- New `GET /v1/api/coordinator/{ws_id}/tasks` returns the
`{version, tasks}` envelope via the shared module-level
`load_task_envelope` decoder (extracted from `CoordinatorClient`
so both the tool path and the UI read share corruption semantics).
Corrupt envelopes return an empty list for UI resilience — the
`task_list` tool remains the authoritative write + error path.
- `CoordinatorManager` subscribes to the `ClusterCollector`'s
listener channel from the console lifespan and dispatches filtered
`child_ws_created / child_ws_state / child_ws_closed / child_ws_rename`
events onto each coordinator's SSE stream. Filter authoritative
on the server via a per-coordinator child-ws_id registry populated
lazily on `open()` from storage and incrementally on `ws_created`
events; cleared on `close()` / eviction. One SSE connection per
client, no client-side filtering.
Frontend:
- DOM-method-only child-row rendering (no innerHTML of user content).
- State glyph vocabulary (● running / ◐ thinking / ⚠ attention /
✗ error / ○ idle) plus text labels — WCAG 1.4.1 carries info in
both glyph and label.
- Live badges (tokens + pending-approval pip) fetched via
`/cluster/ws/{ws_id}/detail` with a 5s TTL cache and 250ms debounce
per child. One request per state change, not per second.
- SSE child events update in place; renderChildren() re-sorts.
- Mobile (<700px) sidebar collapses to an accordion above the chat
with a toggle button flipping aria-expanded; a `.highlight` flash
marks task→child scroll targets; `prefers-reduced-motion` respected.
- Deep-link child rows to `/node/{node_id}/?ws_id=<child>` via
`<a target="_blank" rel="noopener">` with encodeURIComponent on
regex-validated ids.
## Dashboard tree grouping
Cluster dashboard rows now group by `parent_ws_id`. Coordinator rows
(`kind == "coordinator"` or children present) get an expand/collapse
caret (button with `aria-expanded`); collapsed shows "(N children)".
Expanded renders children indented as sibling rows with a left-border
gutter. Orphaned children (parent missing or closed) render at top
level with a muted "orphan" badge. Expansion state persisted in
`localStorage` keyed per coordinator ws_id so operator preference
survives reloads. Coordinator rows deep-link to `/coordinator/{id}`;
node-backed workstreams keep their existing proxy deep-link.
Per-node `ws_created / ws_state / ws_activity` SSE event payloads
gained `parent_ws_id` + `kind` so the collector can propagate them
through its fan-out to browser clients without a second lookup;
`_build_node_snapshot` and `/v1/api/dashboard` rows include the
same. Coordinators (which don't live on cluster nodes) merge into
`/cluster/workstreams` via a new `_coordinator_rows` helper that
threads them through the collector's `get_workstreams(extra_rows=...)`
parameter — extras share the filter / sort / paginate pipeline with
node-backed rows.
## Tests
- `tests/test_coordinator_endpoints.py` — 19 new cases covering
children (empty / populated / ownership 404 / admin bypass /
invalid ws_id / truncation), tasks (empty / round-trip / corrupt /
ownership), and cluster-inspect (auth gates / 400 / 404 / ownership /
coordinator self-path / unloaded-live-null / message-limit clamp).
- `tests/test_coordinator_manager.py` — 8 new cases covering registry
bootstrap on create + open, dispatch for each event type,
unrelated-parent filtering, shutdown idempotency.
- `tests/test_console.py` — existing `cluster_workstreams` assert
updated for the new `extra_rows` kwarg.
## Verification
- `ruff check turnstone tests` clean.
- `mypy turnstone` clean.
- `pytest -m "not live"` — 4184 passed, 3 deselected.
* fix(coordinator): race in dispatch + ui_factory kwarg filtering — PR #370 review
Addresses feedback from the GitHub Copilot + code-quality bot review
passes on PR #370.
## Race in _dispatch_child_event ws_created branch
Copilot flagged a TOCTOU where the lock-free read of
``self._active_coords`` (line 912) could see the parent coordinator,
then ``close()`` / eviction pops ``_children[parent]`` + drops the
coord from ``_active_coords`` before we acquire ``_children_lock``,
and then ``setdefault(parent, set())`` resurrects the entry —
leaking the registry key forever and fanning events to a closed UI.
Fix: re-check ``parent in self._active_coords`` inside
``_children_lock``. The reference swap is still atomic; holding
``_children_lock`` and re-reading the snapshot catches the race
without serializing back through ``self._lock``.
Regression test: create → close → dispatch a ws_created → assert
neither ``_children`` nor ``_active_coords`` regained the entry.
## ui_factory kwarg filtering via inspect.signature
code-quality bot flagged that the previous ``try ui_factory(…, kind=,
parent_ws_id=) except TypeError`` dance fired on every call with
legacy test factories (``lambda wid: WebUI(ws_id=wid)``) — wasteful
and masks real signature mismatches.
Fix: inspect the factory's signature and only pass kwargs it
actually accepts (explicit param name OR ``**kwargs`` absorber).
Keep a conservative ``except TypeError`` fallback for C-callables
and odd signatures ``inspect`` can't introspect.
Copilot also flagged a comment mismatch (the old comment said
"KeyError on **kwargs" — it's ``TypeError``, which is what the code
caught). The rewritten comment is correct.
## Nit: side-effect in assert
code-quality bot flagged ``assert mgr.close(ws.id)`` in
test_coordinator_manager.py. Split into two statements.
## Verification
- ``ruff check`` clean.
- ``mypy turnstone`` clean.
- ``pytest -m "not live"`` — 4223 passed, 3 deselected, 0 failed.
|
||
|
|
5cbc4bc87c |
feat: bulk message insert for fork performance + endpoint tests (#322)
Add save_messages_bulk() to StorageBackend protocol and both backends. Fork path now inserts all messages in a single transaction instead of N individual save_message() calls — for a 200-message workstream this goes from 200 connection/insert/commit cycles to 1. FTS5 indexing is intentionally skipped for bulk fork data (historical messages indexed on rebuild). Ordering preserved via auto-increment id with a shared timestamp across all rows in the batch. Also adds 22 endpoint tests covering the 6 new workstream management endpoints (delete, open, title, refresh-title, list/update interface settings) and 4 storage-level tests for the bulk insert path. |
||
|
|
c093df274d |
feat: workstream management — fork, rename, delete, open, interface settings
Add workstream forking (resume with fork=True keeps new ws_id), custom naming via aliases, title refresh via LLM, and workstream deletion. New server endpoints: delete, refresh-title, set-title, open-workstream, list/update interface settings. Verdict caching with SSE replay on reconnect, display name fallback (alias→title→name) across all endpoints, judge_model override per workstream, and settings_changed broadcast on config reload. New settings: judge.cancel_on_approval, interface.close_tab_action, interface.theme. Storage backends updated with name in list_workstreams_with_history and new get_workstream_metadata method. |
||
|
|
ab1a71c86c |
feat: add PostgreSQL CI integration tests (#156)
* feat: add PostgreSQL CI integration tests Add --storage-backend pytest option and shared storage_backend fixture in conftest.py that creates SQLiteBackend or PostgreSQLBackend based on the flag. Migrate 13 storage test files to use shared fixture instead of local SQLiteBackend fixtures. Add test-postgres CI job with PostgreSQL 17 service container that runs the full test suite against real PostgreSQL. * fix: use TRUNCATE CASCADE for PG cleanup, wrap in try/finally TRUNCATE is faster than per-table DELETE and resets autoincrement sequences. try/except ensures reset_storage() always runs even if cleanup fails due to a corrupted connection from a failing test. * fix: document _engine coupling in PG cleanup comment |
||
|
|
7e680ee883 |
chore: remove dead code — chat.py, singular touch, unused vars, inline imports (#148)
* chore: remove dead code — chat.py shim, singular touch method, unused vars - Delete turnstone/chat.py (backward-compat re-export shim, zero importers) - Remove touch_structured_memory() singular method from protocol + both backends + 6 tests (only plural batch form is used) - Remove unused _last_err variable in _compact_messages - Remove redundant _AGENT_AUTO_TOOLS / _TASK_AUTO_TOOLS class aliases, use module-level constants directly - Consolidate ~76 inline schema imports to top-level in both storage backends (channel_users, channel_routes, oidc_*, scheduled_tasks, watches, services) Net: -219 lines * fix: address review — remove stale inline timedelta imports in prune_task_runs timedelta is already imported at module scope in both backends. |
||
|
|
bf06102d37 |
fix: memory access tracking and BM25 context caching (#138)
* fix: memory access tracking and BM25 context caching Add touch_structured_memory/touch_structured_memories to storage protocol + SQLite/PostgreSQL backends. Bumps last_accessed and access_count on memory retrieval (BM25 injection + search results). 9 new storage tests. Cache the scored BM25 memory context string on ChatSession, invalidated on memory save/delete. Eliminates ~12 redundant storage queries + index rebuilds per session lifecycle. * fix: address review — deduplicate keys in touch facade, clarify contract Deduplicate keys in the memory.py facade before calling storage so each distinct memory is touched at most once. Update protocol docstring to clarify per-call increment semantics. Add deduplication unit test. * fix: replace unused-import test with real batch duplicate test Replace facade dedup test (which only tested Python set logic) with a real storage-level test that verifies duplicate keys each increment access_count. Fixes ruff F401 lint failure. |
||
|
|
723cad24bb |
feat: structured memory system — typed/scoped memories with BM25 rele… (#53)
* feat: structured memory system — typed/scoped memories with BM25 relevance and metacognitive prompting Replace flat key-value memories table with structured_memories (migration 014). Four memory types (user/project/feedback/reference), three scopes (global/workstream/user). Consolidate remember/recall/forget into two tools: memory (action-based: save/search/delete/list) and recall (conversation history only). BM25 relevance scoring (extracted to turnstone/core/bm25.py) selects top-5 memories for system message injection based on conversation context. Metacognitive prompting injects ephemeral nudges after corrections, tool denials, workstream resume, and completion signals. Scope isolation enforced: system message injection and nudge counts filtered to visible memories only (global + current workstream + authenticated user). User scope requires authentication. Content capped at 32KB. ILIKE/LIKE metacharacters escaped in both backends. 113 new tests (2053 total). * fix: CI failure + copilot review feedback - Fix time.monotonic() cooldown: use None sentinel instead of 0.0 default (monotonic clock starts at boot, not epoch — fresh CI runners have uptime < 300s so cooldown check always triggered) - Catch sa.exc.IntegrityError specifically in upsert instead of broad Exception (copilot review) - Preserve existing description/type on upsert when caller doesn't explicitly set them (copilot review) - Add last_accessed + access_count columns to schema/migration for future LRU/LFU eviction support |
||
|
|
1295919613 |
fix: simplify conversation storage — atomic assistant rows with tool_… (#51)
* fix: simplify conversation storage — atomic assistant rows with tool_calls JSON Replace the denormalized storage model (separate rows for assistant content, tool_call, tool_result) with atomic assistant rows carrying tool_calls as a JSON column. Eliminates the 100-line heuristic reconstruct_messages function and its cross-turn merge bug. Schema: add tool_calls TEXT column to conversations (migration 013). Migration backfills existing data — merges tool_call rows into their parent assistant row as JSON, renames tool_result to tool, deletes consumed tool_call rows. Session save path: assistant content + tool_calls saved in one save_message call before tool execution (crash resilient). Tool results saved as role="tool". Extract shared storage utilities to _utils.py: row_to_dict, mutable field frozensets, reconstruct_messages. Both backends import from _utils — PostgreSQL no longer depends on _sqlite.py. Includes denied/blocked tool call badge fix on resume: _build_history detects denied results and propagates flag to parent assistant entry. Frontend uses flag for correct badge-denied rendering. Denied tools visually muted. role="status" on badges for accessibility. Net -45 lines. 8 new tests for reconstruction, all 1914 tests pass. * fix: migration 013 uses parameterized deletes and ordered downgrade - DELETE of consumed tool_call rows now uses parameterized batches (chunks of 500) instead of string interpolation - Downgrade rebuilds via temp table to preserve chronological id ordering when re-inserting tool_call rows |
||
|
|
fb190f8977 |
Normalize session_id into ws_id as sole persistent identity (#29)
* Normalize session_id into ws_id as sole persistent identity Eliminate the separate session_id concept. The workstream ID (ws_id) is now the single identity used for both real-time routing and conversation persistence, removing a layer of indirection that was 1:1 in practice and buggy on resume (stale pointers, orphaned rows). Schema changes (migration 006): - Drop sessions table; add alias/title columns to workstreams - Rename conversations.session_id → ws_id - Rename session_config table → workstream_config (ws_id column) - Data migration remaps existing conversations to ws_id Storage/API renames: - register_session → register_workstream (already existed, merged) - save_message/load_messages now keyed by ws_id - resolve_session → resolve_workstream - ChatSession.session_id property → ws_id - ChatSession.resume_session() → resume() - resume_session field → resume_ws - SessionResumedEvent → WorkstreamResumedEvent - /api/sessions → /api/workstreams/saved - /sessions slash command → /workstreams - --session-retention-days → --retention-days Channel eviction recovery simplified: reuses old ws_id directly instead of get_session_id_by_ws() reverse lookup. * Fix Copilot review feedback: stale session wording in docs, regenerate OpenAPI spec - docs/channels.md: "resumes the session" → "resumes the workstream", "Session resumed:" → "Resumed:", "old session was pruned" → "old workstream was pruned" - docs/api-reference.md: "Each session object" → "Each saved workstream object", field descriptions updated, removed stale node_id field - sdk/typescript/openapi-server.json: fully regenerated from Python models — removes all stale session_id properties from WorkstreamInfo, DashboardWorkstream, CreateWorkstreamResponse schemas |
||
|
|
a20a058c59 |
Add cluster-scale schema, fix console proxy UX, harden SDK sync runner (#22)
* Add cluster-scale schema, fix console proxy UX, harden SDK sync runner Schema redesign for multi-node deployments: - New `workstreams` table with node_id, state, lifecycle tracking - Add node_id + ws_id columns to sessions table with indexes - Full UUID (32 hex) for session_id and ws_id (was truncated 12/8) - Server generates and owns node_id, bridge retrieves via /health - Bridge retries with exponential backoff, fatal on auth errors - WorkstreamManager persists workstreams and state changes to storage - /health endpoint exposes node_id for bridge discovery Console proxy UX fixes: - Remove duplicate turnstone branding from proxy banner - Same-tab navigation for Open Node UI and workstream deep links SDK _SyncRunner fix: - Sentinel pattern for StopAsyncIteration across thread boundary Remove misplaced PNGs from docs/diagrams/ (correct copies in png/ subdir). * Address PR #22 review feedback - Fix CLI session_factory signature (ws_id param) — CI typecheck failure - First-phase eviction in create() now calls _cleanup_ui + record_eviction - close() persists "closed" state to storage via update_workstream_state - Fix noqa comment in test to pragma: no cover |
||
|
|
2b58c127b1 |
Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging (#20)
* Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging Database abstraction: StorageBackend protocol with 21 methods, SQLAlchemy Core schema, SQLite backend (FTS5), PostgreSQL backend (tsvector/ILIKE), Alembic migrations, singleton registry. memory.py reduced to thin facade. Session.py open_db() calls replaced with generic KV methods. [database] config section with env var support. Deployment: Docker Compose production profile with PostgreSQL, Dockerfile with postgres extras and migration entrypoint, Helm chart with bitnami subcharts, Terraform AWS ECS/Fargate module with RDS + ElastiCache + ALB. 39 new storage tests (934 total). mypy strict clean. Docs and diagrams updated. * Address PR #20 review feedback (16 items) - Backends only call create_all() when Alembic migrations are disabled - Helm configmap uses correct TURNSTONE_DB_BACKEND env var; DB URL constructed via env expansion with secret reference instead of ConfigMap - Migration errors fail fast for PostgreSQL (only non-fatal for SQLite) - save_memory/delete_memory wrapped in exception handling like other facade fns - pool_size passed through from config/env to init_storage() in cli + server - Terraform: DB URL moved to Secrets Manager, auth enabled flag set, optional TLS listeners with certificate_arn, Redis transit encryption on - Docker entrypoint no longer suppresses migration output - Diagram fixes: removed StaticPool claim, removed non-existent migration ref - compose.yaml/README: clarified production profile requires DB env vars |