mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
33865ca9d2
Multi-stage /review on the full Phase 1+2+3+4 stack surfaced 9 findings (0 critical, 3 major, 5 minor, 1 nit, 1 uncertain). All applied. Major * perf-1 (session_routes.py:2402): make_history_handler ran sync storage.load_workstream_config inside async def history on the cold- workstream path, blocking the event loop on every dashboard /history request for non-resident workstreams. Every other storage call in the same handler correctly used asyncio.to_thread. Wrap the sync call in asyncio.to_thread (preserving the existing try/except so a DB failure still degrades to the conservative-default branch instead of bubbling out). * q-2 (test_reasoning_audit_log_discipline.py): the security-sensitive test (reasoning text never lands at INFO+ severity) only covered the 4 Phase 1 surfaces. Phase 2 added the strip predicate in AnthropicProvider._convert_messages and Phase 3 added 3 more code paths that touch reasoning text — none guarded. Added 4 parallel tests using the existing capture-and-walk infrastructure: OpenAIResponsesProvider.extract_reasoning_text, OpenAIChatCompletionsProvider.extract_reasoning_text, ChatSession._stream_response (drives the synth-block stamp via a fake reasoning-emitting stream), AnthropicProvider._convert_messages with replay_reasoning_to_model=False (drives the Phase 2 strip predicate). * q-1 (model_registry.py:42): the persist_reasoning flag name implied storage-control but actually gates UI rehydration only — operators flipping it could reasonably expect "stop persisting reasoning" but storage of reasoning bytes happens in provider_data regardless. Renamed everywhere to surface_persisted_reasoning: ModelConfig field, migration 052 column (renaming in-place since 052 is not yet on main), schema, MODEL_DEFINITION_MUTABLE allowlist, _postgresql.py + _sqlite.py CRUD impls, _protocol.py create_model_definition signature, 3 console_schemas Pydantic models, console/server.py admin POST + PUT, model_registry row mapper, history_decoration.py helper parameter, server.py _build_history local var, session_routes.py make_history_handler local var, sdk/events.py HistoryEvent docstring, admin.js form id + override pill label, index.html form input id + UI label + tooltip, coordinator.js (none needed), and every test that referenced the old field name. The admin tooltip now reads "Storage of reasoning bytes is unaffected by this flag — they ride in provider_data regardless" so the decoupling stays explicit at the operator surface. Minor * bug-1 (history_decoration.py:336): dispatcher discriminated on provider_content[0]["type"] only. Anthropic's redacted_thinking blocks (sealed by the safety system) can appear before, after, or interleaved with regular thinking blocks per the API docs. When a redacted block lands first, the dispatcher returned "" and the UI silently lost the surrounding thinking text. Registered "redacted_thinking" as a second key in _BLOCK_TYPE_PROVIDER_FACTORY pointing at the same AnthropicProvider factory — the existing extractor's type=="thinking" filter already correctly skips redacted blocks while walking the full list. Regression test added. * q-3 (_protocol.py:155): replay_reasoning_to_model defaults split across 9 sites — operator-side defaults to False (matches DB server_default), provider-API defaults to True (back-compat with direct callers). Original "pick False everywhere" fix would have silently flipped behaviour for any direct provider caller. Instead documented the intentional bifurcation in the Protocol's create_streaming docstring. * q-4+q-5 (_protocol.py:107 + 3 providers): MAX_REASONING_DISPLAY_BYTES was enforced via Python str slicing which counts code points, not UTF-8 bytes — 4-byte CJK/emoji glyphs would blow past the byte ceiling. Renamed to MAX_REASONING_DISPLAY_CHARS to match actual behaviour. Hoisted the 4-line truncation pattern into a shared _join_reasoning_with_cap helper in _protocol.py; each provider's extractor becomes a single line at the tail. * q-6 (tests/_session_helpers.py): _NullUI + _make_session were duplicated verbatim between test_session_replay_reasoning.py and test_session_synth_reasoning_block.py. Hoisted to a shared tests/_session_helpers.py module (importable, leading underscore so pytest doesn't try to collect it). test_model_registry.py's _make_session has a different signature (registry/model_alias args + _FakeUI) and is not a candidate for sharing. Nit * q-7 (history_decoration.py:286): _make_provider_factory used a dict-as-cell workaround for closure read-only scope. Replaced with the more idiomatic nonlocal pattern. Lint + test gate * ruff check + ruff format -- clean. * mypy -- no issues across all 191 source files. * pytest -m 'not live' -- 6115 passed (3 deselected). Net +5 tests (4 audit-log discipline + 1 redacted_thinking dispatcher). Refinements vs the dedupe output (caught during sanity rendering the report) * perf-1 fix preserved the try/except wrapper. The original "wrap in to_thread" one-liner would have let an OperationalError bubble out instead of degrading to the fallback branch. * q-3 fix explicitly documented the bifurcation rather than collapsing both sides to False. "Pick False everywhere" would silently flip back-compat behaviour for direct provider callers. * q-1 fix included the admin.js:5292 fallback site (m.persist_reasoning !== false) that the original threaded-change list missed. * q-6 fix verified the third _make_session in test_model_registry.py is structurally different (different signature + different UI helper) and intentionally NOT a dedupe target.
312 lines
13 KiB
Python
312 lines
13 KiB
Python
"""Tests for ChatSession synthetic ``reasoning_text`` block stamping (Phase 3 path 3).
|
|
|
|
Path 3 covers OpenAI Chat Completions endpoints — vLLM with
|
|
``--reasoning-parser``, llama.cpp with ``reasoning_format``, Gemini's
|
|
``/v1beta/openai/`` endpoint, and any other server that surfaces
|
|
``delta.reasoning_content`` Pydantic extras. These have no native
|
|
provider_blocks shape on the wire, so ``ChatSession._stream_response``
|
|
captures the streamed reasoning text into ``reasoning_parts`` and
|
|
``_maybe_synth_reasoning_block`` stamps it onto ``_provider_content``
|
|
as a synthetic ``{type: "reasoning_text"}`` block at the end of the
|
|
turn.
|
|
|
|
These tests pin:
|
|
1. The synthesizer fires only when no native blocks were emitted AND
|
|
reasoning was captured (Anthropic + OpenAI Responses bypass it).
|
|
2. ``source`` field is tagged with the active model's server_type
|
|
(informational; pulled from ``server_compat.server_type``).
|
|
3. ``OpenAIChatCompletionsProvider.extract_reasoning_text`` round-trips
|
|
the synthetic block on history rehydration.
|
|
4. The synthetic shape is NOT in ``ANTHROPIC_VALID_BLOCK_TYPES`` so
|
|
cross-model resumption (local-model → Anthropic) falls through
|
|
cleanly to the text+tool_calls rebuild path.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from types import SimpleNamespace
|
|
from typing import Any
|
|
|
|
from tests._session_helpers import make_session as _make_session
|
|
from turnstone.core.providers._anthropic import (
|
|
ANTHROPIC_VALID_BLOCK_TYPES,
|
|
AnthropicProvider,
|
|
)
|
|
from turnstone.core.providers._openai_chat import OpenAIChatCompletionsProvider
|
|
|
|
|
|
class TestMaybeSynthReasoningBlock:
|
|
"""Direct unit tests for ``ChatSession._maybe_synth_reasoning_block``."""
|
|
|
|
def test_no_synth_when_provider_blocks_present(self) -> None:
|
|
# Anthropic / OpenAI Responses path — native blocks already
|
|
# carry the reasoning, no synth needed.
|
|
session = _make_session()
|
|
existing = [{"type": "thinking", "thinking": "x"}]
|
|
out = session._maybe_synth_reasoning_block(existing, ["should not be added"])
|
|
assert out is existing
|
|
|
|
def test_no_synth_when_reasoning_parts_empty(self) -> None:
|
|
session = _make_session()
|
|
out = session._maybe_synth_reasoning_block([], [])
|
|
assert out == []
|
|
|
|
def test_no_synth_when_reasoning_parts_only_whitespace(self) -> None:
|
|
session = _make_session()
|
|
out = session._maybe_synth_reasoning_block([], [" ", "\n\t"])
|
|
assert out == []
|
|
|
|
def test_synth_creates_reasoning_text_block(self) -> None:
|
|
session = _make_session()
|
|
out = session._maybe_synth_reasoning_block([], ["thought ", "process"])
|
|
assert len(out) == 1
|
|
assert out[0]["type"] == "reasoning_text"
|
|
assert out[0]["text"] == "thought process"
|
|
|
|
def test_synth_omits_source_when_no_server_type(self) -> None:
|
|
session = _make_session()
|
|
# No registry / no server_compat → source field omitted.
|
|
out = session._maybe_synth_reasoning_block([], ["text"])
|
|
assert "source" not in out[0]
|
|
|
|
def test_synth_includes_source_when_server_type_resolvable(self) -> None:
|
|
session = _make_session()
|
|
session._registry = SimpleNamespace(
|
|
get_config=lambda alias: SimpleNamespace(
|
|
capabilities={"server_compat": {"server_type": "vllm"}},
|
|
)
|
|
)
|
|
session._model_alias = "qwen3-32b"
|
|
out = session._maybe_synth_reasoning_block([], ["text"])
|
|
assert out[0]["source"] == "vllm"
|
|
|
|
def test_synth_handles_registry_exception(self) -> None:
|
|
# _resolve_server_type silently returns "" on any lookup error
|
|
# — synth still fires but omits the source field.
|
|
class BrokenRegistry:
|
|
def get_config(self, alias: str) -> Any:
|
|
raise KeyError(alias)
|
|
|
|
session = _make_session()
|
|
session._registry = BrokenRegistry()
|
|
session._model_alias = "missing"
|
|
out = session._maybe_synth_reasoning_block([], ["text"])
|
|
assert out[0]["text"] == "text"
|
|
assert "source" not in out[0]
|
|
|
|
|
|
class TestSyntheticBlockShapeContract:
|
|
"""The synthetic block shape MUST stay outside Anthropic's valid
|
|
block types so cross-model resumption falls through cleanly."""
|
|
|
|
def test_reasoning_text_not_in_anthropic_valid_types(self) -> None:
|
|
# If this assertion ever fails, the cross-model resumption
|
|
# safety story breaks: a synthetic block from a local-model
|
|
# session would reach Anthropic's wire as a malformed block.
|
|
assert "reasoning_text" not in ANTHROPIC_VALID_BLOCK_TYPES
|
|
|
|
def test_synthetic_block_falls_through_anthropic_shape_filter(self) -> None:
|
|
# Cross-model resumption regression: turn 1 was on a local
|
|
# model (synthetic block stamped), then the operator switched
|
|
# to Anthropic. The shape filter must reject the synthetic
|
|
# block and fall through to text+tool_calls rebuild.
|
|
provider = AnthropicProvider()
|
|
msg = {
|
|
"role": "assistant",
|
|
"content": "spoken answer",
|
|
"_provider_content": [
|
|
{"type": "reasoning_text", "text": "synth thought", "source": "vllm"},
|
|
],
|
|
}
|
|
_, converted = provider._convert_messages([msg])
|
|
assistant = next(m for m in converted if m["role"] == "assistant")
|
|
block_types = [b.get("type") for b in assistant["content"] if isinstance(b, dict)]
|
|
# Foreign block did NOT reach Anthropic's wire. Rebuilt from
|
|
# text only.
|
|
assert "reasoning_text" not in block_types
|
|
assert assistant["content"] == [{"type": "text", "text": "spoken answer"}]
|
|
|
|
|
|
class TestOpenAIChatExtractReasoningText:
|
|
"""``OpenAIChatCompletionsProvider.extract_reasoning_text`` reads
|
|
the synthetic block back out for UI rehydration."""
|
|
|
|
def test_reads_synthetic_reasoning_text_block(self) -> None:
|
|
provider = OpenAIChatCompletionsProvider()
|
|
blocks = [{"type": "reasoning_text", "text": "captured thought"}]
|
|
assert provider.extract_reasoning_text(blocks) == "captured thought"
|
|
|
|
def test_concatenates_multiple_blocks(self) -> None:
|
|
provider = OpenAIChatCompletionsProvider()
|
|
blocks = [
|
|
{"type": "reasoning_text", "text": "first"},
|
|
{"type": "reasoning_text", "text": "second"},
|
|
]
|
|
assert provider.extract_reasoning_text(blocks) == "first\nsecond"
|
|
|
|
def test_skips_other_block_types(self) -> None:
|
|
provider = OpenAIChatCompletionsProvider()
|
|
blocks = [
|
|
{"type": "thinking", "thinking": "anth"},
|
|
{"type": "reasoning", "summary": [{"text": "openai"}]},
|
|
{"type": "reasoning_text", "text": "chat"},
|
|
]
|
|
assert provider.extract_reasoning_text(blocks) == "chat"
|
|
|
|
def test_handles_empty_text_field(self) -> None:
|
|
provider = OpenAIChatCompletionsProvider()
|
|
blocks = [
|
|
{"type": "reasoning_text", "text": ""},
|
|
{"type": "reasoning_text", "text": "kept"},
|
|
]
|
|
assert provider.extract_reasoning_text(blocks) == "kept"
|
|
|
|
def test_handles_missing_text_field(self) -> None:
|
|
provider = OpenAIChatCompletionsProvider()
|
|
blocks = [
|
|
{"type": "reasoning_text"}, # no text
|
|
{"type": "reasoning_text", "text": "kept"},
|
|
]
|
|
assert provider.extract_reasoning_text(blocks) == "kept"
|
|
|
|
def test_returns_empty_for_no_synth_blocks(self) -> None:
|
|
provider = OpenAIChatCompletionsProvider()
|
|
blocks = [{"type": "thinking", "thinking": "x"}]
|
|
assert provider.extract_reasoning_text(blocks) == ""
|
|
|
|
|
|
class TestStreamResponseSynthBlockIntegration:
|
|
"""Integration test: drives a fake reasoning-emitting stream
|
|
through ``ChatSession._stream_response`` and asserts the
|
|
synthesizer wires up correctly. Pins the call site at
|
|
``session.py`` (where ``_maybe_synth_reasoning_block`` is invoked
|
|
on the assembled provider_blocks before stamping ``_provider_content``)
|
|
— without this, a future refactor that drops the synthesizer call
|
|
would silently break path-3 capture (vLLM/llama.cpp/Gemini-compat
|
|
reasoning would be visible live but invisible on history reload).
|
|
"""
|
|
|
|
def _make_stream(self, content: str, reasoning: str) -> Any:
|
|
"""Build an iterator of StreamChunks that mimic a path-3
|
|
capture (reasoning_delta chunks, content chunks, no
|
|
provider_blocks emitted).
|
|
"""
|
|
from turnstone.core.providers._protocol import StreamChunk, UsageInfo
|
|
|
|
chunks = []
|
|
# Reasoning first (matches live SSE order).
|
|
if reasoning:
|
|
chunks.append(StreamChunk(reasoning_delta=reasoning, is_first=True))
|
|
# Content next.
|
|
if content:
|
|
chunks.append(
|
|
StreamChunk(
|
|
content_delta=content,
|
|
is_first=not reasoning,
|
|
)
|
|
)
|
|
# Final chunk with finish_reason + usage.
|
|
chunks.append(
|
|
StreamChunk(
|
|
finish_reason="stop",
|
|
usage=UsageInfo(prompt_tokens=10, completion_tokens=20, total_tokens=30),
|
|
)
|
|
)
|
|
return iter(chunks)
|
|
|
|
def test_stream_response_stamps_synth_block_when_path3_reasoning_captured(
|
|
self,
|
|
) -> None:
|
|
"""Drive a fake stream emitting reasoning_delta chunks (no
|
|
native provider_blocks) through ``_stream_response``; assert
|
|
the resulting assistant_msg carries a synthetic reasoning_text
|
|
block stamped onto ``_provider_content``."""
|
|
session = _make_session()
|
|
# No registry → source field omitted from synth block.
|
|
stream = self._make_stream(content="Final answer.", reasoning="path-3 reasoning")
|
|
msg = session._stream_response(stream)
|
|
assert msg["role"] == "assistant"
|
|
assert msg["content"] == "Final answer."
|
|
# Synthetic block should be stamped onto _provider_content.
|
|
provider_content = msg.get("_provider_content")
|
|
assert isinstance(provider_content, list)
|
|
assert len(provider_content) == 1
|
|
assert provider_content[0]["type"] == "reasoning_text"
|
|
assert provider_content[0]["text"] == "path-3 reasoning"
|
|
|
|
def test_stream_response_no_synth_when_no_reasoning_captured(self) -> None:
|
|
"""Stream emits only content (no reasoning_delta). No synth
|
|
block stamped — _provider_content key absent on assistant_msg."""
|
|
session = _make_session()
|
|
stream = self._make_stream(content="just content", reasoning="")
|
|
msg = session._stream_response(stream)
|
|
assert msg["content"] == "just content"
|
|
# No synth block (and no native blocks either) → key absent.
|
|
assert "_provider_content" not in msg
|
|
|
|
def test_stream_response_synth_block_carries_source_when_server_type_resolvable(
|
|
self,
|
|
) -> None:
|
|
"""When the active model has server_compat.server_type set,
|
|
the synth block carries it as the ``source`` field."""
|
|
session = _make_session()
|
|
session._registry = SimpleNamespace(
|
|
get_config=lambda alias: SimpleNamespace(
|
|
capabilities={"server_compat": {"server_type": "vllm"}},
|
|
)
|
|
)
|
|
session._model_alias = "qwen3-32b"
|
|
stream = self._make_stream(content="answer", reasoning="reasoning text")
|
|
msg = session._stream_response(stream)
|
|
provider_content = msg.get("_provider_content")
|
|
assert isinstance(provider_content, list)
|
|
assert provider_content[0]["source"] == "vllm"
|
|
|
|
|
|
class TestResolveServerType:
|
|
"""Direct unit tests for the helper that pulls server_type from
|
|
the active model's capabilities dict."""
|
|
|
|
def test_returns_empty_when_no_registry(self) -> None:
|
|
session = _make_session()
|
|
session._registry = None
|
|
assert session._resolve_server_type() == ""
|
|
|
|
def test_returns_empty_when_no_alias(self) -> None:
|
|
session = _make_session()
|
|
session._registry = SimpleNamespace(
|
|
get_config=lambda alias: SimpleNamespace(capabilities={})
|
|
)
|
|
session._model_alias = ""
|
|
assert session._resolve_server_type() == ""
|
|
|
|
def test_returns_server_type_when_present(self) -> None:
|
|
session = _make_session()
|
|
session._registry = SimpleNamespace(
|
|
get_config=lambda alias: SimpleNamespace(
|
|
capabilities={"server_compat": {"server_type": "llama.cpp"}}
|
|
)
|
|
)
|
|
session._model_alias = "local-model"
|
|
assert session._resolve_server_type() == "llama.cpp"
|
|
|
|
def test_returns_empty_when_server_compat_missing(self) -> None:
|
|
session = _make_session()
|
|
session._registry = SimpleNamespace(
|
|
get_config=lambda alias: SimpleNamespace(
|
|
capabilities={"context_window": 32768},
|
|
)
|
|
)
|
|
session._model_alias = "local-model"
|
|
assert session._resolve_server_type() == ""
|
|
|
|
def test_returns_empty_on_exception(self) -> None:
|
|
class BrokenRegistry:
|
|
def get_config(self, alias: str) -> Any:
|
|
raise RuntimeError("boom")
|
|
|
|
session = _make_session()
|
|
session._registry = BrokenRegistry()
|
|
session._model_alias = "x"
|
|
assert session._resolve_server_type() == ""
|