mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-23 20:34:49 -06:00
main
11 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e75522b0f9 | fix: surface backend errors before stream failure | ||
|
|
bc55210936 |
Add immutable memory index snapshots (#1022)
* feat: add immutable memory index snapshots Capture the visible memory metadata index at first model admission, preserve it as immutable system-prefix context, and emit relevance pointers without rewriting cached history. Align project authorization, MCP actor refresh ordering, storage APIs, SDKs, console surfaces, and regression coverage with the snapshot lifecycle. * fix: stabilize memory index for release candidate * fix(sdk): avoid polynomial description trim * chore: split memory index documentation |
||
|
|
6eae1c3954 | fix(providers): support OpenAI v3 HTTPX2 transport | ||
|
|
480a1426b3 |
Fail-closed history-commit handoff (#1005)
* fix(session): fail-closed history-commit handoff (#981) The deleted-workstream discovery is now a terminal, ws_id-keyed latch: keyed conversation commits refuse admission once the durable parent is gone (convergence finalizers and force-abandon are exempt), history handoff refuses to mint a proof token so /history fails closed with a 503 instead of silently wiping the pane, and the SSE stream carries a workstream_gone resync reason. Discarded commits leave a forensic log of commit keys and roles, never content. Conversation rows gain a commit_key (migration 071): keyed saves are idempotent under retry, validated against the full commit identity, and refused when they would cross a workstream deletion. The prune orphan category now requires a NULL alias plus a two-hour updated grace, with cutoffs computed at discovery time and carried into both dialects' rechecks. The mid-turn interjection queue is owner-partitioned with no per-site mode flags: pops take the acting principal's and unowned rows, other participants' rows are structurally retained, and enforcement lives at queue admission plus the shared before_spawn gates. The retraction ledger is bounded by open pop windows: pops open a window atomically with the queue delete, restores close their ids atomically with the ledger consume, every other exit closes through one helper, and misses for unheld ids record nothing. The workstream-gone latch refuses unattended wakes at all three gates (watcher spawn, claim, delivery pre-pop), and the retry dispatcher regained its pre-envelope cancel/error convergence net. Persistence-state reporting derives through the session bound to each UI instead of a registry lookup by id that failed open to healthy during tombstone retention. The dashboard roster no longer re-inserts ghost entries from trailing activity events, the history tool-outcome scan tolerates interleaved non-turn rows, and the shared handoff-deadline handle owns its own retirement. Single-sourced across call sites: keyed-commit row values, attachment save wrappers, tail-truncation and conflict-resolution bodies for both storage dialects; worker-slot lifecycle field sets; the direct-commit admission frame; queued-row layout accessors; the string-aware comment stripper shared by every JS harness suite. Refs #981 #964 * fix(session): sweep handoff fixes to their sibling surfaces The interactive replay loop treated a system row as a tool-batch boundary, so every tool result after an interleaved row vanished from that pane while the coordinator rendered the same history correctly. Only a conversational turn ends the batch window now, matching the shared outcome index. Accepted user turns clear the composer's attachment chips on the same viewer policy that settles optimistic bubbles rather than on having matched a local bubble, so a workstream created with an upload no longer keeps a chip for an attachment the create dispatch already consumed. The coordinator's raced-Stop arm emits the stream-end hook it inherits alongside the idle state, leaving no unfinalized bubble or unflushed tool output. Ending a session surfaces a failure toast when the request never lands or answers with a non-JSON body. The per-second persistence reconcile now probes each session without blocking: a workstream whose generation and handoff locks are held is skipped until the next pass instead of contending the locks every commit needs. The one-shot repair that gates workstream creation at capacity keeps a definite probe — it has no next pass, and the sessions likeliest to be contended are the ones whose unresolved journals emptied its candidate list. Single-sourced: the attachment lane builds its conversation row through the shared commit-identity builder; the ordinary worker exit releases its slot through the lifecycle owner; both operator surfaces snapshot their counters through one non-consuming helper; the replay preamble loses its per-kind wrappers and its config hook; the browser harness suites share one brace walker; and each in-flight history attempt is one record carrying both its abort controller and its deadline. Refs #981 #964 |
||
|
|
7a06f5e8bc |
refactor(session): make ModelLane the provider boundary (#979) (#989)
* refactor(session): make ModelLane the provider boundary (#979) ## Summary This closes the model-lane ownership gap left by #832: `ChatSession` no longer stores raw provider/client handles. `ResolvedModelBinding` now carries the provider, client, model, capabilities, registry generation, and backend-auth configuration as one coherent snapshot. - Atomically rebind existing sessions after model-registry changes while pinning each in-flight send, fallback, judge, output guard, task agent, title, compaction, perception, and voice operation to its initiating principal and binding. - Fence UI publication, canonical trajectory folds, durable writes, streams, retries, child scopes, and judge work by generation. Stop can hand off to a successor without accepting late state; cancelled tools retain typed effect receipts, and concurrent approval batches resolve by exact cycle or call. - Make create, fork, open, close, and delete race-safe with hidden `creating` reservations, incarnation-aware state tails, and an ACL-rechecked transaction that clones checkpoint-bounded history, configuration, project/persona state, and attachment references. - Extend REST/OpenAPI and Python/TypeScript SDK contracts for create/fork inputs, routed-create metadata, live-workstream probes, targeted approvals, and structured cancellation results. - Update architecture, storage, authentication, judge, channel, console, API, and SDK documentation, including regenerated architecture diagrams and OpenAPI artifacts. ## Validation - SQLite suite: 11,188 passed, 9 skipped, 10 deselected - PostgreSQL suite: 11,195 passed, 2 skipped, 10 deselected - Live backend: 3 passed - SSE recovery: 6 passed; browser recovery harness passed all scenarios - Ruff: clean; 595 files correctly formatted - mypy: 243 source files clean - TypeScript: typecheck/build and 35 tests passed - OpenAPI artifacts fresh; all 14 changed diagrams reproduce byte-for-byte - `git diff --check` and Git LFS integrity clean Closes #979. * fix(deps): update nanoid for GHSA-2v37-7h3g-55p8 Refresh the transitive lock entry admitted by PostCSS so the TypeScript security gate no longer resolves the vulnerable custom-generator implementation. Validation: - npm ci - npm audit --audit-level=moderate: 0 vulnerabilities - TypeScript typecheck and build - TypeScript tests: 35 passed * fix(test): assert canonical model registry URLs Replace prefix checks with exact canonical base URL assertions so the tests do not model incomplete URL validation. Validation: tests/test_model_registry.py (185 passed); Ruff check/format; mypy. |
||
|
|
47524654b3 |
fix(streaming): close the retry window's generation, identity, and masking holes
xhigh review round on the mid-stream retry ladder: 14 verified correctness findings, all fixed, plus the verified-but-capped cleanups mined from the review run. Generation safety — the shared-slot class is removed structurally, not gated per site: a dead attempt's partial now rides the raised exception (thread-private by construction) into a wrapper-local variable, and the _midstream_dead_partial session slot is deleted, so an orphaned superseded generation cannot poison a live generation's preservation. The promotion helper is generation-gated, writes the marker row even for a pre-token death (empty content takes the marker-as-message branch), and backfills a recorded-but-empty partial with the previous attempt's text, so a Stop anywhere in the retry window — backoff, re-create, or TTFT wait — preserves the latest text the user actually saw. _record_cancelled_partial is generation-gated too: a superseded thread touches neither the UI nor the shared slot. Identity — the retry gate and the fatal formatter now consult the provider that actually owns the live stream (recorded at creation, covering the fallback walk by construction), so a fallback stream's provider-specific transient is retryable by ITS OWN contract and failures are labeled with the binding that produced them. The mid-retry rebind check compares the full (client, model, provider) binding — reload() keeps the pooled client on model-only swaps — and a re-prepare also re-exports the wire fold that send()'s token-table calibration counts. Masking — a context overflow raised by the mid-retry re-create surfaces as itself so the compact-and-retry arm can recover the turn, and the overflow arm is split: recovery-machinery failures still surface the original overflow (its wording anticipates them), while post-compaction consumption failures surface as themselves instead of a false overflow diagnosis. Cancellation and terminal paths — a Stop that races the trailing-metadata window is re-checked after the chunk loop, so the turn aborts with the marker instead of committing and running its tool calls; the terminal arm finalizes client-side only, deliberately keeping the in-progress snapshot (the unpersisted partial's only copy) for refresh-replay; KeyboardInterrupt gets the same client-side finalize; the retry arm stops the spinner before restarting it (the CLI's on_thinking_start replaces the spinner without stopping it — a thread leak); and the backoff delay is computed from the pre-increment index, matching the sibling ladders' convention. Mined cleanups: the retry suite wraps the shared session factory instead of duplicating its defaults; the usage projection uses dataclasses.asdict; the partial-content rule lives in one closure serving both preservation paths; the two fatal-log tests are parametrized into one; the test import uses the public providers package. |
||
|
|
5fb27e8f81 |
fix(session): survive mid-stream transport deaths in interactive turns (#937)
A wire death during body streaming (ReadError on a TLS record failure, peer resets) surfaces after the request has already returned its stream handle, so neither the SDK's request retries nor the creation-time retry ladder ever saw it: the interactive turn died with a bare exception string, the partial output was discarded, and no log trace was left. Utility lanes already survived this through drain_stream's normalization; the interactive loop now gets the same treatment. - transport_guarded() in providers/_protocol.py: drain_stream's transport-death conversion made reusable for consumers that keep streaming semantics. Pre-finish deaths raise the retryable IncompleteStreamError (drain's exact message shape); post-finish blips end the stream cleanly, forfeiting only trailing metadata. - The single-pass chunk consumer renames to _stream_attempt; _stream_response is now the resilient wrapper owning ALL stream acquisition plus a bounded mid-stream re-issue ladder (_MID_STREAM_RETRIES, the shared _stop_retrying predicate with a per-loop cap, cancel-aware exponential backoff). Send()'s overflow compact-and-retry arm now wraps the whole turn and passes re-prepared msgs explicitly. - A dead attempt is finalized across every UI consumer before the retry (stream_end then turn_committed then notice then spinner), so retried text never appends onto the dead attempt's in any surface (browser transcript, CLI markdown fences, Slack/Discord streamed messages, SSE replay ring). - Before re-creating, the session re-resolves its registry binding: a concurrent ModelRegistry.reload() closes cached clients, and the retry must not stream into the closed one. A failing re-create logs stream.retry.recreate_failed and re-raises the ORIGINAL stream-death error rather than masking it. - _format_backend_error gains a stream-death branch naming the provider, endpoint, and model, with a short identity-bearing first sentence. _BACKEND_STREAM_EXC_NAMES joins _BACKEND_KNOWN_EXC_NAMES, which also removes those names from _is_ctx_overflow's text-detection eligibility (deliberate: their texts are fixed transport strings that never carry overflow phrases). - _record_fatal_error now logs session.fatal.recorded (INFO for KeyboardInterrupt, ERROR otherwise) so fatal turns leave a journal trace. - _assistant_pending_tokens resets at stream entry so a post-finish blip that loses the trailing usage chunk cannot append the previous turn's completion count as this turn's estimate. Offline SDK boundary pins (openai/anthropic mid-body death identity and no re-request, cross-thread client close surfacing httpx.ReadError) guard the assumptions the retry gate rests on. |
||
|
|
33ace975d2 |
feat(models): default-deny governance and admin UI for per-alias backend auth
Follow-up to the per-alias Entra OBO/app-identity backend auth: the console write path now applies default-deny field classification, the admin shelf gains full backend-auth support, and the session/registry rebind machinery is hardened for config changes landing under live sessions. Console write gate: - Default-deny classification: any non-neutral change to a row that is or becomes dynamic requires admin.mcp plus validation; the provably auth-neutral columns are enumerated (MODEL_AUTH_NEUTRAL_FIELDS) and a live-schema classification test forces every future column to be classified. The derivation is a pure function (_derive_auth_gate) with unit-pinned exclusivity invariants. - Two-tier validation mirroring the MCP oauth_obo validator: the row tier (audience allow-list) runs on every gated write; the posture tier (OIDC configured, token store present) runs on pair changes and on enable-arming. - Pure-disable carve-out: disabling a dynamic row is de-escalation and is never blocked — admin.models suffices and validation is skipped, including for rows with corrupt or skewed stored values. - Capabilities are compared canonically (key order, integral floats), the audience compare normalizes both sides, and staging an audience on a static row is refused on both write twins. - Calibrate writes the capabilities column under an enforced confinement invariant with a compare-and-swap persist. Admin shelf: - Backend-auth section with a per-open constraints fetch (GET /model-definitions/auth-constraints: audience allow-list, grant profile, dynamic modes), datalist audience suggestions, server-defined modes preserved on round-trip, and permission-aware visibility built on cache-skew-safe helpers shared through auth.js. - Refused live-registry swaps surface as an amber registry_warning on the write, delete, reload, and calibrate responses; audit rows carry auth_gated / auth_disarmed markers visible in the audit view. Registry and sessions: - The encryption-key requirement for dynamic auth is enforced inside ModelRegistry.reload() itself — nodes refuse with 503 and the console records coord_registry_error — and reload bumps the generation before the map swap so a racing reader can never pair a stale generation with new maps. - resolve()/resolve_binding() return the generation from inside the registry lock; sessions rebind per send on generation change with atomic client/provider/config commits, fallback-first handling of removed or unconstructable aliases, and judge/limiter resets only when the binding actually changed. - Mint refusals record per-user causes surfaced in the per-turn heartbeat logs; misconfiguration warnings are deduplicated with bounded state. Verification: 10417 tests (99 added on this branch), a 71-scenario browser harness over the real admin shelf, and a live rfc8693 token-exchange e2e run (MCP legs verified end to end; the model-leg scope gap is tracked as #955 under a narrow known-gap signature). Closes #950. |
||
|
|
3af80907c7 |
fix(server): init-message worker exits to error, not idle, on first-turn failure
The initial-message worker (_run_initial) collapsed both cancel and backend-error exits into one `except (Exception, GenerationCancelled)` arm that always stamped state=idle, clobbering the state=error that session.send's _record_fatal_error had persisted+emitted. A spawned child's first-turn backend failure (unreachable model server, exhausted quota, auth error) therefore read as an empty, successful turn — the coordinator's wait/inspect surface reads last_error only for state=='error' — and the real error surfaced only after a manual nudge re-ran the turn synchronously. Split the arm: cancel -> idle, exception -> error. The failed child now settles at state=error and the first wait_for_workstream returns the enriched backend error inline. Also fixes the same latent bug for scheduled tasks, which dispatch through the same endpoint and closure. A failed first turn is deliberately terminal for automated wakes: it settles to a non-ready error terminal, not the idle ready-set that timer/watch wakes recur to, so explicit user/coordinator action reactivates it rather than a silent auto-retry (a self-healing wake-from-error would be a separate wake-gate change). The exception arm routes through a new ChatSession.ensure_error_recorded: a no-op when send already recorded the error in-line (the common backend-boundary path — no duplicate state emit), and the recorder when a pre-try exception (model-registry refresh, user-turn append, system-message recompose) bypassed send's own handler, so state=error always carries a meaningful last_error. Its idempotency guard (_has_persisted_error) is session-lifetime, so ensure_error_recorded is scoped to _run_initial's FRESH first-turn session only; the docstring spells out why a session-reuse caller (retry, /send, coord send, wake) must not route through it until the per-turn error-recorded signal of #865 lands. The other half of making an errored workstream cheap for a model to handle is a stable identifier: the enriched backend error now leads with the model ALIAS the coordinator references everywhere (list_nodes, spawn) and annotates the backend id for the operator — "model=DeepSeek-V4-Flash (id=deepseek-v4-flash)" — so a model routing around a failed model correlates it against those surfaces without a lookup, instead of burning reasoning tokens reconciling the alias against a backend id it never sees anywhere else. Collapses to one token when the alias and id coincide. Tests (TestInitialWorkerFailureState) assert the coordinator-visible manager state and the persisted last_error across the matrix — common- backend and pre-try errors both settle error with a readable last_error; cancel-to-idle settles idle with no error recorded. De-forks the create-app fixture and uses the shared monotonic wait_until helper. The completion-notification honesty surface and the error-recording hygiene of the other send-worker closures (retry / main send / coord send / wake) are deferred to #865. |
||
|
|
c6e5794125 |
fix(compaction): recover from context overflow on resume across providers
A session created under the openai-compatible provider and resumed under the anthropic-compatible provider (same vLLM model) failed with an opaque InternalError instead of recovering. Root cause: vLLM returns a context-window overflow as HTTP 400 BadRequestError on /v1/chat/completions but HTTP 500 InternalServerError on /v1/messages, and the rehydrated resume payload overflowed the window. The 500 was retried four times then surfaced as a bare class name. - Detect overflow by message text, not exception class (_is_ctx_overflow), shared across the fatal-error formatter, both stream-retry gates, the send-loop recovery, the chunker, and the task_agent loop. Overflow is non-retryable (deterministic; no backoff). Phrasing is overflow-specific so a token-quota rate-limit isn't misclassified. - Proactive pre-send compaction (Layer A): when already over the hard ceiling, compact once before the first stream so a resume that arrives over-window (or follows a switch to a smaller-context model, with no prior compaction) doesn't go out blind. Generation-guarded end to end so an orphaned or superseded send can never swap the live generation's history. - Binary-subdivision chunker: an over-window summary batch is split in half and the partials merged (~log2(N) calls, not one per block); a lone over-window block is truncated progressively down to a floor before bailing irreducible. - Cooperative cancellation honored through compaction; send() consumes its own generation's cancel signal on exit, so a stale cancel can't block a later idle /compact and a live cancel is never disarmed. - _format_backend_error surfaces "Context window exceeded ..." instead of an opaque InternalServerError, and only for unrecognized classes. - retry/rewind, the continuation hint, and title generation all exclude the synthetic [Conversation summary] turn so they can't target the label. - task_agent salvages a sub-agent's partial work on any terminal error (not only overflow), re-raising only when there is nothing to salvage. |
||
|
|
ce11f01a80 |
feat(session): enriched backend error messages with provider + URL
A bare ``httpx.ReadTimeout`` previously surfaced as ``ReadTimeout: timed
out`` — no provider, no base URL, no model — leaving the user with no
signal to tell whether a model server hung, the URL was wrong, or the
model isn't loaded on the backend.
``ChatSession._format_backend_error`` now rewrites known boundary
exceptions (httpx ``ReadTimeout`` / ``ConnectError`` / etc. and OpenAI /
Anthropic SDK ``APITimeoutError`` / ``APIConnectionError`` /
``NotFoundError`` / ``AuthenticationError`` / ``RateLimitError``) into
operator-actionable text that names the provider, base URL (query
string stripped before ``sanitize_error_text`` redacts credentials),
and model. Matching is by class name so the helper carries no SDK
imports. Unrecognised exceptions fall through to the legacy
``f"{type(exc).__name__}: {exc}"`` shape, preserving existing grep
targets.
|