fix(providers): segregate inline reasoning at the drain seam

Passthrough servers (parserless vLLM/llama.cpp, LM Studio, bare
gateways) emit reasoning as literal <think>/<reasoning> blocks inside
content, and only three of nine drained lanes stripped them: web_fetch
tool results persisted raw think blocks into every following turn
(#940), judge verdicts parsed through tag noise, and a draft verdict
inside a think block could shadow the real one at the output guard.

One rule at the seam now. drain_stream accumulates content in RUNS
bounded by interleaving signals (provider-parsed reasoning deltas,
tool-call deltas) with the interactive consumer's within-chunk ordering
— reasoning, then content, then the tool-call close — and splits each
run through split_inline_reasoning, the one-shot form of the
interactive lane's ThinkTagSplitter: a pure raw split, exactly
equivalent to the streaming form on every catalog case. One trim policy
exists and the drain owns it: blank edge lines are trimmed once over
the joined runs when a tag was consumed, so tag residue dies at the
edges while genuine inter-run paragraph separators survive. Extracted
text is appended to result.reasoning after any server-parsed reasoning
with a blank-line boundary and rides the native lane as the
reasoning_text synth block. Orphan CLOSE tags deliberately pass through
byte-identical: a close whose open never arrived is indistinguishable
from prose QUOTING the tag, and drained lanes routinely quote
third-party text — reclassifying would let a malicious page containing
the literal tag destroy the extraction that cites it. The title lane
keeps a local rfind peel as display-string formatting. The citations
footer folds only onto non-blank content — sourcing for an answer that
does not exist is dropped rather than handed to emptiness checks as a
footer-only "answer".

Every private strip is deleted: the title lane's strip, the summarizer
strip, _strip_reasoning itself, and the optimizer's five regexes
(_strip_markdown_fence is now the one fence rule, applied to normalized
model output only, never to or-fallback values). Think-only and
whitespace-only responses drain to blank content, and every lane's
no-answer fallback gates on blankness: web_fetch returns an honest
extraction-error card, the intent judge takes the empty-retry ladder,
the task-agent synthesis reports "(no output)", and the optimizer keeps
the current observer system and prompt verbatim on no-answer passes.
Final-say reads (optimizer analyst, eval final_content, the notify
hook) use trajectory.final_assistant_text — the last assistant turn
only, never an earlier narration presented as the conclusion — while
last_assistant_text is the salvage walk (task_agent partial-work
recovery), skipping tool-call-only, all-reasoning, and whitespace-only
turns. Perception memoizes every completed description immediately,
including an empty one — one perceive per key, ever — under a
commit-lock guard so an empty result never overwrites a concurrently
memoized real description; an all-reasoning perception model pins the
placeholder until restart, and the remediation is server-side (a
reasoning parser or the template thinking toggle on the perception
alias). A true double-reasoning shape (inline-extracted text alongside
a native reasoning block) logs chars-only at the drain, where it is
distinguishable from the routine reasoning_delta mirror.

The dialect's semantics are pinned as one table
(tests/_reasoning_dialect.py) driven through shared fixtures
(think_tag_stream, seam_provider): one-shot conformance, the exact
one-shot/streaming equivalence property, the drain seam rules including
quoted-tag safety, run-boundary and separator-preservation pins,
per-lane pins for all nine lanes, and the empty-content assistant wire
shape.

Closes #965. Closes #940.
This commit is contained in:
Patrick Buckley
2026-08-04 08:06:04 -07:00
parent 1d7db73305
commit bc3fa60011
25 changed files with 1363 additions and 173 deletions
+14
View File
@@ -206,6 +206,20 @@ Earlier stable lines (`stable/1.6`, `stable/1.5`) are frozen.
### Fixed ### Fixed
- **Inline `<think>`/`<reasoning>` blocks no longer leak into drained
results (#965, #940).** On servers without a reasoning parser
(parserless vLLM/llama.cpp, LM Studio, bare gateways), reasoning
arrives as literal tags inside content; segregation now happens once
at the drain seam, so web-fetch tool results, sub-agent syntheses,
judge verdicts, titles, summaries, and optimizer prompts receive
tag-free content and the extracted reasoning rides the native lane.
Two behavior notes: a web-fetch extraction whose whole response was
reasoning now returns an explicit `Error: extraction returned no
answer` tool result (previously the raw reasoning text persisted as a
successful result and was replayed every following turn), and a
mismatched-vocabulary close tag (`<think>…</reasoning>`) now closes
the block — matching the interactive lane's long-standing rule —
where the old per-lane strips treated it as unterminated.
- **A transport failure mid-generation no longer kills the interactive - **A transport failure mid-generation no longer kills the interactive
turn (#937).** A wire death during body streaming (TLS record failure, turn (#937).** A wire death during body streaming (TLS record failure,
connection reset — `httpx.ReadError` and kin) surfaces after the connection reset — `httpx.ReadError` and kin) surfaces after the
+167
View File
@@ -0,0 +1,167 @@
"""Inline-reasoning dialect conformance catalog.
Passthrough servers (parserless vLLM/llama.cpp, LM Studio, bare gateways)
emit model reasoning inline as ``<think>``/``<reasoning>`` blocks inside the
content stream a *dialect* of model output. This module is that dialect's
executable specification for ``split_inline_reasoning``: each case maps an
utterance to the exact ``(content, reasoning)`` lanes the one-shot must
produce.
Consumers: the one-shot conformance and one-shotstreaming property suites
in tests/test_think_tag_split.py. Lane suites (session, judge,
output-guard, optimizer, drain-stream) pin their lanes with suite-local
utterances through their own fakes adding a case HERE extends the
semantics spec, not automatically any lane suite.
The split is RAW (residue whitespace stays; ``drain_stream`` owns the one
trim over its joined runs). ``passthrough`` marks cases the split must
return BYTE-IDENTICAL: tag-free text, and text whose only tags are orphan
CLOSE tags. The latter is a
review ruling, not an accident: a close tag whose open never arrived is
indistinguishable from prose QUOTING the tag, and drained lanes routinely
quote third-party text (web-fetch answers citing pages about reasoning
models, guard verdicts echoing judged content) any reclassification
would let quoted text destroy real results. Display lanes wanting
stricter cosmetic peeling (the title) own that locally as formatting.
"""
from dataclasses import dataclass
@dataclass(frozen=True)
class DialectCase:
id: str
utterance: str
content: str
reasoning: str
# The split returns the utterance byte-identical (no tag consumed):
# tag-free text, or orphan-close-only text (quoted-tag safety).
passthrough: bool = False
CASES: tuple[DialectCase, ...] = (
DialectCase(
id="no_tag_byte_identity",
utterance="Just an answer.",
content="Just an answer.",
reasoning="",
passthrough=True,
),
DialectCase(
# The fast path must not strip: unconsumed content is
# byte-identical, whitespace included.
id="no_tag_preserves_whitespace",
utterance=" spaced \n",
content=" spaced \n",
reasoning="",
passthrough=True,
),
DialectCase(
id="leading_block",
utterance="<think>plan</think>Answer",
content="Answer",
reasoning="plan",
),
DialectCase(
# The split is RAW — tag residue stays; drain_stream owns the ONE
# blank-edge-line trim over its joined runs (pinned there), which
# preserves the code block's first-line indentation.
id="indented_code_block_raw_residue",
utterance="<think>plan</think>\n\n print(1)\n more()",
content="\n\n print(1)\n more()",
reasoning="plan",
),
DialectCase(
id="leading_block_raw_residue",
utterance="<think>plan</think>\n\nAnswer\n",
content="\n\nAnswer\n",
reasoning="plan",
),
DialectCase(
id="interleaved_blocks_both_vocabularies",
utterance="Intro <think>a</think>mid <reasoning>b</reasoning>end",
content="Intro mid end",
reasoning="ab",
),
DialectCase(
id="unterminated_open_tail_is_reasoning",
utterance="Answer part<think>never closed",
content="Answer part",
reasoning="never closed",
),
DialectCase(
# QUOTED-CLOSE SAFETY (review ruling): an orphan close is
# indistinguishable from a quoted tag — everything passes through.
# A malicious page embedding the literal string must not be able
# to wipe the extraction that quotes it.
id="orphan_close_passes_through",
utterance="The page says templates emit </think> after the preamble. Answer: 42.",
content="The page says templates emit </think> after the preamble. Answer: 42.",
reasoning="",
passthrough=True,
),
DialectCase(
# Template-pre-injected shape ("reasoning</think>answer"): the seam
# deliberately passes it through — segregating it would require
# treating every quoted close as a boundary. Post-#831 every lane
# streams, and known streaming surfaces strip the orphan close
# server-side; display lanes peel cosmetically on their own.
id="preinject_shape_passes_through",
utterance="plan text</think>\n\nAnswer",
content="plan text</think>\n\nAnswer",
reasoning="",
passthrough=True,
),
DialectCase(
id="immediate_close_passes_through",
utterance="</think>Answer",
content="</think>Answer",
reasoning="",
passthrough=True,
),
DialectCase(
# Any close tag closes any open block (splitter semantics; the old
# pairwise per-caller strip treated this as unterminated).
id="cross_vocabulary_close",
utterance="<think>x</reasoning>Answer",
content="Answer",
reasoning="x",
),
DialectCase(
id="think_only",
utterance="<think>all reasoning</think>",
content="",
reasoning="all reasoning",
),
DialectCase(
id="think_only_unterminated",
utterance="<think>everything",
content="",
reasoning="everything",
),
DialectCase(
# A balanced block followed by a stray close: the block is
# consumed, the stray close stays in content (quoted-tag safety),
# and the consumed-tag strip applies.
id="balanced_block_then_stray_close",
utterance="<think>a</think>b</think>c",
content="b</think>c",
reasoning="a",
),
DialectCase(
id="multiple_blocks_accumulate",
utterance="<think>one</think>mid<think>two</think>tail",
content="midtail",
reasoning="onetwo",
),
DialectCase(
# ACCEPTED RESIDUAL (R2): the split is content-blind, so a literal
# OPEN tag in legitimate prose misroutes the remainder — the same
# false positive the interactive splitter has carried in the
# field. This pin makes any future fix a conscious change.
id="literal_open_tag_false_positive_r2",
utterance="The `<think>` tag opens a block.",
content="The `",
reasoning="` tag opens a block.",
),
)
+38 -1
View File
@@ -19,7 +19,7 @@ from types import SimpleNamespace
from typing import Any from typing import Any
from unittest.mock import MagicMock from unittest.mock import MagicMock
from turnstone.core.providers import StreamChunk, ToolCallDelta from turnstone.core.providers import ModelCapabilities, StreamChunk, ToolCallDelta
from turnstone.core.session import ChatSession from turnstone.core.session import ChatSession
from turnstone.core.session_ui_base import SessionUIBase from turnstone.core.session_ui_base import SessionUIBase
@@ -349,6 +349,43 @@ def as_stream(result: Any) -> list[StreamChunk]:
] ]
def think_tag_stream(utterance: str) -> list[StreamChunk]:
"""``create_streaming`` return value simulating a passthrough server
that emits *utterance* typically think-tag-bearing as plain
streamed content.
The per-lane fixture for inline-reasoning dialect pins: lane tests
supply their own utterances (the dialect's SEMANTICS are specified
once, in ``tests._reasoning_dialect.CASES``, and pinned by the
one-shot suites lane pins assert lane behavior, not tag grammar).
Routes through the real ``drain_stream`` seam exactly like
``as_stream``.
"""
return as_stream(mock_completion_result(content=utterance))
def seam_provider(utterance: str, *, provider_name: str = "openai-compatible") -> MagicMock:
"""Provider fake whose ``create_streaming`` replays *utterance* through
the REAL drain seam (``think_tag_stream``) THE lane-suite seam fake.
One definition so the lane suites cannot drift when the provider
surface ``model_turn`` probes grows: real ``ModelCapabilities`` for
the clamp math, ``provider_name`` overridable per suite.
Assign the RETURNED fake to ``session._provider`` never mutate the
provider a session resolved on its own: with a MagicMock client the
session resolves the process-wide ``create_provider(...)`` singleton,
and writing that shared instance's ``create_streaming`` poisons every
later session in the test run (the SSE-recovery e2e servers resolve
the same instance).
"""
provider = MagicMock()
provider.provider_name = provider_name
provider.get_capabilities.return_value = ModelCapabilities()
provider.create_streaming = MagicMock(return_value=think_tag_stream(utterance))
return provider
class RecordingUI: class RecordingUI:
"""UI adapter recording the ordered event stream ``send()`` emits.""" """UI adapter recording the ordered event stream ``send()`` emits."""
+272 -2
View File
@@ -50,6 +50,144 @@ class TestContentAndReasoning:
assert result.reasoning == "think hard" assert result.reasoning == "think hard"
assert result.content == "answer" assert result.content == "answer"
def test_inline_tags_segregated_at_the_seam(self):
# The one-shot splitter runs on the joined content, so EVERY
# drained consumer receives IR-clean content by construction —
# the tag arrives split across deltas exactly as a passthrough
# server streams it.
result = drain_stream(
iter(
[
StreamChunk(content_delta="<thi"),
StreamChunk(content_delta="nk>plan</think>"),
StreamChunk(content_delta="answer"),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.content == "answer"
assert result.reasoning == "plan"
def test_extracted_reasoning_appends_after_server_parsed(self):
# Segregate, never discard: server-parsed reasoning_delta first,
# inline-extracted after, with a boundary — two distinct passes
# must never read as one run-together sentence.
result = drain_stream(
iter(
[
StreamChunk(reasoning_delta="parsed."),
StreamChunk(content_delta="<think>inline.</think>answer"),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.reasoning == "parsed.\n\ninline."
assert result.content == "answer"
def test_extracted_reasoning_alone_carries_no_separator(self):
result = drain_stream(
iter(
[
StreamChunk(content_delta="<think>inline.</think>answer"),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.reasoning == "inline."
def test_all_reasoning_turn_drops_citations_footer(self):
# An all-reasoning turn's footer is sourcing for an answer that
# does not exist; folding it would hand downstream emptiness
# checks a truthy, footer-only "answer".
result = drain_stream(
iter(
[
StreamChunk(content_delta="<think>all reasoning</think>"),
StreamChunk(finish_reason="stop"),
StreamChunk(info_delta="Sources:\n- x"),
]
)
)
assert result.content == ""
assert result.reasoning == "all reasoning"
def test_double_reasoning_shape_logs_chars_only(self, caplog):
# Inline-extracted text alongside a NATIVE reasoning block has no
# native lane to land in downstream — observable here, chars-only.
import logging
secret = "the plan nobody logs"
with caplog.at_level(logging.DEBUG):
result = drain_stream(
iter(
[
StreamChunk(content_delta=f"<think>{secret}</think>answer"),
StreamChunk(
finish_reason="stop",
provider_blocks=[{"type": "thinking", "thinking": "native"}],
),
]
)
)
assert result.content == "answer"
assert "drain.inline_reasoning_alongside_native" in caplog.text
assert secret not in caplog.text
def test_routine_reasoning_delta_mirror_does_not_log(self, caplog):
# reasoning_delta beside a native block is the NORMAL shape on
# Anthropic/Responses lanes (the delta mirrors the block) — it
# must not drown the anomaly signal.
import logging
with caplog.at_level(logging.DEBUG):
drain_stream(
iter(
[
StreamChunk(reasoning_delta="mirrored"),
StreamChunk(content_delta="answer"),
StreamChunk(
finish_reason="stop",
provider_blocks=[{"type": "thinking", "thinking": "mirrored"}],
),
]
)
)
assert "drain.inline_reasoning_alongside_native" not in caplog.text
def test_trailing_footer_folds_after_split_and_is_never_scanned(self):
# The citations footer is web-controlled text: it folds onto the
# ALREADY-split content, so a tag-shaped citation title cannot
# reclassify the result.
result = drain_stream(
iter(
[
StreamChunk(content_delta="<think>x</think>answer"),
StreamChunk(finish_reason="stop"),
StreamChunk(info_delta="Sources:\n- how </think> works"),
]
)
)
assert result.content == "answer\n\nSources:\n- how </think> works"
assert result.reasoning == "x"
def test_tool_call_turn_with_in_think_tail(self):
# A drained tool-call turn whose trailing content is an
# unterminated think block: the tail is reasoning, the calls ride.
result = drain_stream(
iter(
[
StreamChunk(content_delta="<think>pondering tools"),
StreamChunk(
tool_call_deltas=[ToolCallDelta(index=0, id="c1", name="bash")],
finish_reason="tool_calls",
),
]
)
)
assert result.content == ""
assert result.reasoning == "pondering tools"
assert result.tool_calls is not None and result.tool_calls[0]["id"] == "c1"
def test_stream_without_finish_reason_raises_incomplete(self): def test_stream_without_finish_reason_raises_incomplete(self):
# Complete-or-error: every adapter emits a finish reason on a # Complete-or-error: every adapter emits a finish reason on a
# healthy stream, so its absence means the generation died # healthy stream, so its absence means the generation died
@@ -244,7 +382,12 @@ class TestInfoDelta:
) )
assert result.content == format_citations("body", anns) assert result.content == format_citations("body", anns)
def test_trailing_fold_with_empty_content_matches_too(self): def test_trailing_fold_onto_empty_content_is_dropped(self):
# A turn with NO content (tool-call-only here, all-reasoning in the
# split case) keeps content empty: a footer-only "answer" would be
# truthy and defeat every downstream emptiness check. This pin
# replaced the old byte-match-the-retired-non-streaming-lane rule
# (which appended the footer onto the empty base).
result = drain_stream( result = drain_stream(
iter( iter(
[ [
@@ -253,7 +396,7 @@ class TestInfoDelta:
] ]
) )
) )
assert result.content == "\n\nSources:\n- x" assert result.content == ""
def test_finishless_stream_raises_even_with_trailing_info(self): def test_finishless_stream_raises_even_with_trailing_info(self):
# A stream that dies after a status ping must NOT return the ping # A stream that dies after a status ping must NOT return the ping
@@ -395,3 +538,130 @@ class TestErrorPropagation:
with pytest.raises(RuntimeError, match="upstream broke"): with pytest.raises(RuntimeError, match="upstream broke"):
drain_stream(chunks()) drain_stream(chunks())
def test_whitespace_only_content_does_not_gain_footer():
# Blankness, not truthiness: tag-free whitespace-only content is
# byte-identical at the split (never normalized), and a citations
# footer folded onto it would read as a truthy footer-only "answer".
result = drain_stream(
iter(
[
StreamChunk(content_delta="\n\n"),
StreamChunk(reasoning_delta="all the substance"),
StreamChunk(finish_reason="stop"),
StreamChunk(info_delta="Sources:\n- x"),
]
)
)
assert result.content == "\n\n"
assert result.reasoning == "all the substance"
def test_unterminated_think_before_tool_calls_does_not_swallow_answer():
# Content runs are bounded by interleaving signals, mirroring the
# interactive consumer's flush-and-reset when tool calls begin: an
# unterminated <think> before the calls is reasoning, the post-call
# answer is CONTENT — never swallowed into the open block.
result = drain_stream(
iter(
[
StreamChunk(content_delta="<think>plan"),
StreamChunk(tool_call_deltas=[ToolCallDelta(index=0, id="c1", name="bash")]),
StreamChunk(content_delta="Answer: 42"),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.content == "Answer: 42"
assert result.reasoning == "plan"
assert result.tool_calls is not None and result.tool_calls[0]["id"] == "c1"
def test_unterminated_think_before_reasoning_delta_does_not_swallow_answer():
# The mixed-dialect shape: provider-parsed reasoning interleaving a raw
# content stream also closes the run (the interactive Path-1 reset).
result = drain_stream(
iter(
[
StreamChunk(content_delta="<think>plan"),
StreamChunk(reasoning_delta="parsed."),
StreamChunk(content_delta="answer"),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.content == "answer"
assert result.reasoning == "parsed.\n\nplan"
def test_whitespace_only_extraction_does_not_pollute_reasoning():
# The field-common no-think shape: an empty think body must not append
# blank text (or a separator) into persisted reasoning.
result = drain_stream(
iter(
[
StreamChunk(reasoning_delta="parsed."),
StreamChunk(content_delta="<think>\n\n</think>answer"),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.reasoning == "parsed."
assert result.content == "answer"
def test_combined_content_and_tool_chunk_keeps_content_in_prior_run():
# Within a chunk the interactive ordering holds — reasoning, content,
# THEN the tool-call close: a combined content+tools chunk feeds its
# content into the pre-boundary run, so a close tag arriving in that
# chunk still closes the open block instead of stranding as a literal
# orphan in drained content.
result = drain_stream(
iter(
[
StreamChunk(content_delta="Let me plan. <think>use bash"),
StreamChunk(
content_delta="</think>Running it.",
tool_call_deltas=[ToolCallDelta(index=0, id="c1", name="bash")],
),
StreamChunk(finish_reason="tool_calls"),
]
)
)
assert result.content == "Let me plan. Running it."
assert result.reasoning == "use bash"
assert result.tool_calls is not None and result.tool_calls[0]["id"] == "c1"
def test_inter_run_paragraph_separator_survives_reasoning_boundary():
# The edge trim runs ONCE over the joined whole: a genuine paragraph
# break the model emitted before an interleaving signal is interior
# after the join and must survive.
result = drain_stream(
iter(
[
StreamChunk(content_delta="<think>plan</think>First part.\n\n"),
StreamChunk(reasoning_delta="server-parsed"),
StreamChunk(content_delta="Second part."),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.content == "First part.\n\nSecond part."
assert result.reasoning == "server-parsed\n\nplan"
def test_inter_run_separator_survives_tool_boundary():
result = drain_stream(
iter(
[
StreamChunk(content_delta="Narration.<think>x</think>\n\nIntro:\n\n"),
StreamChunk(tool_call_deltas=[ToolCallDelta(index=0, id="c1", name="bash")]),
StreamChunk(content_delta="Post-call answer."),
StreamChunk(finish_reason="stop"),
]
)
)
assert result.content == "Narration.\n\nIntro:\n\nPost-call answer."
assert result.reasoning == "x"
+38
View File
@@ -1202,3 +1202,41 @@ class TestModelAliasResolution:
assert callback_results[0].tier == "llm" assert callback_results[0].tier == "llm"
assert callback_results[0].tier != "llm_fallback" assert callback_results[0].tier != "llm_fallback"
assert "did not return a verdict" not in callback_results[0].reasoning assert "did not return a verdict" not in callback_results[0].reasoning
class TestInlineReasoningSeam:
"""#965 per-lane pins: judge content arrives IR-clean from the drain."""
def test_think_wrapped_verdict_parses_clean(self):
# Reasoning around the verdict JSON is segregated at the seam, so
# _parse_verdict reads pure JSON — a draft verdict INSIDE the think
# block can no longer shadow the real one.
content = (
'<think>draft: {"recommendation": "block", "risk_level": "critical"}</think>'
+ _good_verdict_json()
)
provider = _make_mock_provider(response_content=content)
judge = _make_judge(provider)
result = judge._evaluate_single(
_make_item(),
[{"role": "user", "content": "test"}],
cancel_event=None,
client=MagicMock(),
)
assert result is not None
assert result.recommendation == "approve"
assert result.risk_level == "low"
def test_think_only_response_takes_empty_ladder(self):
# An all-reasoning judge turn drains to empty content and rides the
# SAME empty-response ladder as a genuinely empty turn — it never
# reaches _parse_verdict with tag text.
provider = _make_mock_provider(response_content="<think>only deliberation</think>")
judge = _make_judge(provider)
result = judge._evaluate_single(
_make_item(),
[{"role": "user", "content": "test"}],
cancel_event=None,
client=MagicMock(),
)
assert result is None
+36
View File
@@ -727,3 +727,39 @@ def test_resolve_capabilities_survives_get_config_raise() -> None:
registry.get_config.side_effect = KeyError("gone") registry.get_config.side_effect = KeyError("gone")
caps = resolve_capabilities(provider, "m", "gone", registry) caps = resolve_capabilities(provider, "m", "gone", registry)
assert caps == ModelCapabilities() assert caps == ModelCapabilities()
def test_inline_tags_segregate_to_native_reasoning_text_and_clean_content() -> None:
# A passthrough server's tagged content, drained through the real seam:
# the turn's text is IR-clean and the extracted reasoning lands in the
# native lane as the path-3 synth block (so it survives reload and the
# operator-gated replay), never in any consumer-visible content.
provider = _FakeProvider([CompletionResult(content="<think>plan</think>answer")])
result = model_turn(_lane(provider), [Turn.user("q")])
assert result.content == "answer"
assert result.turn.text == "answer"
assert result.turn.native is not None
synth = [b for b in result.turn.native.blocks if b.get("type") == "reasoning_text"]
assert len(synth) == 1
assert synth[0]["text"] == "plan"
def test_synth_bail_is_silent_and_leaks_nothing(
caplog: pytest.LogCaptureFixture,
) -> None:
# Bailing on an existing native reasoning block is the ROUTINE no-op on
# Anthropic/Responses lanes (reasoning_delta mirrors the block): no log
# event here, and reasoning text never reaches a log payload. The
# genuinely anomalous shape (inline-EXTRACTED text beside a native
# block) is logged at the drain, where it is distinguishable.
import logging
from turnstone.core.model_turn import synth_reasoning_block
secret_reasoning = "the plan nobody logs"
with caplog.at_level(logging.DEBUG):
blocks = synth_reasoning_block(
[{"type": "thinking", "thinking": "native"}], [secret_reasoning]
)
assert blocks == [{"type": "thinking", "thinking": "native"}]
assert secret_reasoning not in caplog.text
+10 -1
View File
@@ -220,6 +220,8 @@ class TestExtractLastAssistantContent:
assert _extract_last_assistant_content(session) == "world" assert _extract_last_assistant_content(session) == "world"
def test_structured_content(self): def test_structured_content(self):
# Multi-block text flattens via the canonical Turn.text projection
# (the shared final-say read), not a notify-private join.
session = MagicMock() session = MagicMock()
session.messages = turns_from_dicts( session.messages = turns_from_dicts(
[ [
@@ -232,7 +234,14 @@ class TestExtractLastAssistantContent:
}, },
] ]
) )
assert _extract_last_assistant_content(session) == "part one\npart two" assert _extract_last_assistant_content(session) == "part onepart two"
def test_whitespace_only_final_say_reports_empty(self):
# A whitespace-only final say is empty — the notify fallback fires
# instead of sending raw whitespace.
session = MagicMock()
session.messages = turns_from_dicts([{"role": "assistant", "content": " \n"}])
assert _extract_last_assistant_content(session) == ""
def test_empty_messages(self): def test_empty_messages(self):
session = MagicMock() session = MagicMock()
+127
View File
@@ -0,0 +1,127 @@
"""#965 per-lane pins for the optimizer: content arrives IR-clean from the
drain seam, and the emptiness fallbacks are load-bearing.
The optimizer's five private regex strips are gone; these rows pin (a) that
tagged model output still yields clean prompts/systems (the seam does the
work now), and (b) the two ``or``-fallback flips: an all-reasoning pass
keeps the CURRENT observer system / prompt verbatim instead of wiping the
observer system or evaluating an empty-prompt evolution node.
"""
from __future__ import annotations
from tests._session_helpers import seam_provider
from turnstone.optimizer import _observe_and_update_optimizer, _propose_prompt_modification
_OBSERVER_SYSTEM = "You observe optimization runs and refine the optimizer system."
def test_observer_tagged_output_yields_clean_system() -> None:
out = _observe_and_update_optimizer(
client=object(),
model="m",
optimizer_system=_OBSERVER_SYSTEM,
iterations=[],
provider=seam_provider("<think>weighing the history</think>Refined observer text."),
)
assert out == "Refined observer text."
def test_observer_think_only_keeps_current_system_verbatim() -> None:
# Before the seam, the private regex emptied the text AFTER the
# ``or``-fallback check and the observer system was WIPED.
out = _observe_and_update_optimizer(
client=object(),
model="m",
optimizer_system=_OBSERVER_SYSTEM,
iterations=[],
provider=seam_provider("<think>no conclusion reached</think>"),
)
assert out == _OBSERVER_SYSTEM
def test_proposal_tagged_output_yields_clean_prompt() -> None:
out = _propose_prompt_modification(
client=object(),
model="m",
current_prompt="Old prompt.",
test_cases=[],
iteration_result={"cases": {}},
history=[],
provider=seam_provider("<think>rework section two</think>New prompt text."),
)
assert out == "New prompt text."
def test_proposal_think_only_keeps_current_prompt() -> None:
# An all-reasoning pass reads as "no changes" downstream — never an
# empty-prompt evolution-tree node.
out = _propose_prompt_modification(
client=object(),
model="m",
current_prompt="Old prompt.",
test_cases=[],
iteration_result={"cases": {}},
history=[],
provider=seam_provider("<think>hmm</think>"),
)
assert out == "Old prompt."
def test_observer_whitespace_only_content_keeps_current_system() -> None:
# Falsiness is checked AFTER normalization: whitespace-only content
# (tag-free, so seam-byte-identical and truthy) must not strip to ""
# past the fallback and wipe the observer system.
out = _observe_and_update_optimizer(
client=object(),
model="m",
optimizer_system=_OBSERVER_SYSTEM,
iterations=[],
provider=seam_provider("\n\n \n"),
)
assert out == _OBSERVER_SYSTEM
def test_proposal_whitespace_only_content_keeps_current_prompt() -> None:
out = _propose_prompt_modification(
client=object(),
model="m",
current_prompt="Old prompt.",
test_cases=[],
iteration_result={"cases": {}},
history=[],
provider=seam_provider("\n\n"),
)
assert out == "Old prompt."
def test_proposal_fallback_prompt_is_never_fence_stripped() -> None:
# The fence-strip normalizes MODEL output only; a no-answer pass keeps
# a fence-bearing current prompt VERBATIM, never reduced to its fence
# innards.
fenced_prompt = "Do the task.\n```python\nexample()\n```\nBe precise."
out = _propose_prompt_modification(
client=object(),
model="m",
current_prompt=fenced_prompt,
test_cases=[],
iteration_result={"cases": {}},
history=[],
provider=seam_provider("<think>no conclusion</think>"),
)
assert out == fenced_prompt
def test_proposal_model_output_fence_is_unwrapped() -> None:
# The fence rule applies to MODEL output (only): a fenced proposal
# yields its innards, and prose outside the fences is discarded.
out = _propose_prompt_modification(
client=object(),
model="m",
current_prompt="Old prompt.",
test_cases=[],
iteration_result={"cases": {}},
history=[],
provider=seam_provider("Here you go:\n```\nNew prompt text.\n```\nHope that helps!"),
)
assert out == "New prompt text."
+22
View File
@@ -637,3 +637,25 @@ class TestExtractJson:
" (note: not valid JSON, missing braces and quote handling)" " (note: not valid JSON, missing braces and quote handling)"
) )
assert _extract_json(broken) is None assert _extract_json(broken) is None
class TestInlineReasoningSeam:
"""#965 per-lane pins: guard content arrives IR-clean from the drain."""
def test_draft_verdict_inside_think_cannot_shadow_real_verdict(self) -> None:
judge = _make_judge(
content=(
'<think>draft: {"risk_level": "high", "flags": ["exfil"]}</think>'
'{"risk_level": "none", "flags": []}'
)
)
v = judge.evaluate("tool output", func_name="bash", call_id="c1")
assert v.succeeded
assert v.risk_level == "none"
assert v.flags == ()
def test_think_only_response_is_empty_response_error(self) -> None:
judge = _make_judge(content="<think>all deliberation, no verdict</think>")
v = judge.evaluate("tool output", func_name="bash", call_id="c1")
assert not v.succeeded
assert v.error == "empty_response"
+48
View File
@@ -166,3 +166,51 @@ def test_describe_peek_returns_cached_without_recompute() -> None:
is None is None
) )
assert prov.calls == 1 assert prov.calls == 1
def test_describe_cached_memoizes_empty_descriptions() -> None:
# A completed-but-empty description (an all-reasoning pass) memoizes
# like any other result: one perceive per key, ever — bounded cost.
# The pin-until-restart residual is deliberate; the remediation is
# server-side (reasoning parser / template thinking toggle).
prov = _StubProvider(content="")
kw: dict[str, Any] = {
"provider": prov,
"client": object(),
"model": "m",
"principal_id": "user-a",
"alias": "omni",
"content_hash": "h-empty",
"parts": _parts(),
}
assert perception.describe_cached(**kw) == ""
assert perception.describe_cached(**kw) == ""
assert prov.calls == 1 # second call served from the memo
assert (
perception.describe_peek(principal_id="user-a", alias="omni", content_hash="h-empty") == ""
)
def test_racing_empty_result_never_clobbers_memoized_real_description(monkeypatch) -> None:
# The describe call runs unlocked: a racer can memoize a REAL
# description while another call is producing "". The empty commit
# must yield to the existing memo, never overwrite it.
key_kwargs = {"principal_id": "user-a", "alias": "omni", "content_hash": "h-race"}
def _racing_describe(**_kw: Any) -> str:
with perception._cache_lock:
perception._cache[perception._cache_key(**key_kwargs)] = "real from racer"
return ""
monkeypatch.setattr(perception, "describe", _racing_describe)
out = perception.describe_cached(
provider=_StubProvider(content=""),
client=object(),
model="m",
parts=_parts(),
**key_kwargs,
)
assert out == "real from racer"
assert (
perception.describe_peek(**key_kwargs) == "real from racer"
) # the billed real description survived
+21
View File
@@ -6246,3 +6246,24 @@ class TestTransportRetryability:
provider = create_provider(provider_name) provider = create_provider(provider_name)
assert "IncompleteStreamError" in provider.retryable_error_names assert "IncompleteStreamError" in provider.retryable_error_names
def test_sanitize_keeps_empty_content_assistant_turn_with_tool_calls():
# A think-only assistant turn drains to empty content beside its tool
# calls — the exact shape every prose-less tool-call turn already has.
# The chat-lane sanitizer must pass it through unmangled.
msgs = [
{"role": "user", "content": "q"},
{
"role": "assistant",
"content": "",
"tool_calls": [
{"id": "c1", "type": "function", "function": {"name": "bash", "arguments": "{}"}}
],
},
{"role": "tool", "tool_call_id": "c1", "content": "out"},
]
out = sanitize_messages(msgs)
assert out[1]["content"] == ""
assert out[1]["tool_calls"][0]["id"] == "c1"
assert out[2]["tool_call_id"] == "c1"
+114 -3
View File
@@ -18,6 +18,7 @@ from tests._session_helpers import (
mock_completion_result, mock_completion_result,
scripted_anthropic_client, scripted_anthropic_client,
scripted_chat_client, scripted_chat_client,
seam_provider,
) )
from turnstone.core.session import _IMAGE_EXTENSIONS, _IMAGE_SIZE_CAP, ChatSession from turnstone.core.session import _IMAGE_EXTENSIONS, _IMAGE_SIZE_CAP, ChatSession
from turnstone.core.trajectory import ( from turnstone.core.trajectory import (
@@ -1800,9 +1801,10 @@ class TestTitleRetry:
"""A reasoning model's answer can arrive wrapped in an unparsed """A reasoning model's answer can arrive wrapped in an unparsed
``<think>`` span (lanes that don't split it into reasoning_content) ``<think>`` span (lanes that don't split it into reasoning_content)
plus markdown / quotes. There is no portable switch to disable thinking, plus markdown / quotes. There is no portable switch to disable thinking,
so the title pass gives reasoning room (raised max_tokens), reuses so the title pass gives reasoning room (raised max_tokens), relies on
``_strip_reasoning``, and peels wrapping decoration keeping INTERNAL the drain seam's segregation (``split_inline_reasoning`` — this lane
punctuation (the hyphen survives).""" holds no strip of its own), and peels wrapping decoration keeping
INTERNAL punctuation (the hyphen survives)."""
from turnstone.core.providers._protocol import ModelCapabilities from turnstone.core.providers._protocol import ModelCapabilities
from turnstone.core.session import _TITLE_MAX_TOKENS from turnstone.core.session import _TITLE_MAX_TOKENS
@@ -7620,3 +7622,112 @@ def test_record_aux_usage_attributes_explicit_model():
# The explicit agent model wins over the session default. # The explicit agent model wins over the session default.
assert ui.aux_calls[0]["model"] == "plan-model-xyz" assert ui.aux_calls[0]["model"] == "plan-model-xyz"
assert ui.aux_calls[0]["prompt_tokens"] == 900 assert ui.aux_calls[0]["prompt_tokens"] == 900
def _fake_fetched_page() -> MagicMock:
"""Response fake for the monkeypatched web_fetch guard fetch."""
resp = MagicMock()
resp.raise_for_status = MagicMock()
resp.headers = {"content-type": "text/plain"}
resp.text = "Page body."
return resp
class TestInlineReasoningSeamLanes:
"""#965 per-lane pins: web_fetch extraction (the #940 repro) and the
task_agent synthesis path receive IR-clean content from the drain seam."""
def test_web_fetch_extraction_result_is_clean(self, monkeypatch, tmp_db):
# The #940 class: a passthrough server wraps the extraction answer
# in think tags; the tool result (persisted and replayed every
# following turn) must carry ONLY the answer.
session = _make_session()
monkeypatch.setattr(
"turnstone.core.session.fetch_with_ssrf_guard",
lambda url, **kw: _fake_fetched_page(),
)
session._provider = seam_provider(
"<think>scanning the page for the answer</think>HRW hashing weights nodes.",
provider_name="openai",
)
call_id, answer = session._exec_web_fetch(
{"call_id": "wf1", "url": "https://example.com/x", "question": "What is HRW?"}
)
assert answer == "HRW hashing weights nodes."
def test_web_fetch_think_only_extraction_is_honest_error(self, monkeypatch, tmp_db):
# An all-reasoning extraction drains to empty content — the tool
# result flips to an explicit error instead of silently persisting
# think text as a "success".
session = _make_session()
monkeypatch.setattr(
"turnstone.core.session.fetch_with_ssrf_guard",
lambda url, **kw: _fake_fetched_page(),
)
session._provider = seam_provider("<think>hmm, unclear</think>", provider_name="openai")
call_id, answer = session._exec_web_fetch(
{"call_id": "wf2", "url": "https://example.com/x", "question": "What is HRW?"}
)
assert answer == "Error: extraction returned no answer"
def test_task_agent_synthesis_is_clean(self, tmp_db):
# The audit's unverified sibling, scripted: a sub-agent turn wrapped
# in think tags reaches the coordinator-visible synthesis clean.
from turnstone.core.trajectory import Turn
session = _make_session()
session._provider = seam_provider(
"<think>sub-agent deliberation</think>Sub-agent findings.", provider_name="openai"
)
out = session._run_agent(
[Turn.system("You are a test agent."), Turn.user("Report findings.")],
label="task",
tools=[],
auto_tools=set(),
)
assert out == "Sub-agent findings."
def test_task_agent_think_only_turn_reports_no_output(self, tmp_db):
from turnstone.core.trajectory import Turn
session = _make_session()
session._provider = seam_provider(
"<think>nothing but reasoning</think>", provider_name="openai"
)
out = session._run_agent(
[Turn.system("You are a test agent."), Turn.user("Report findings.")],
label="task",
tools=[],
auto_tools=set(),
)
assert out == "(no output)"
class TestWhitespaceOnlyBlanknessGates:
"""Whitespace-only drained content takes the no-answer fallbacks —
blankness, not truthiness, campaign-wide."""
def test_task_agent_whitespace_only_turn_reports_no_output(self, tmp_db):
from turnstone.core.trajectory import Turn
session = _make_session()
session._provider = seam_provider("\n\n", provider_name="openai")
out = session._run_agent(
[Turn.system("You are a test agent."), Turn.user("Report findings.")],
label="task",
tools=[],
auto_tools=set(),
)
assert out == "(no output)"
def test_web_fetch_whitespace_only_extraction_is_honest_error(self, monkeypatch, tmp_db):
session = _make_session()
monkeypatch.setattr(
"turnstone.core.session.fetch_with_ssrf_guard",
lambda url, **kw: _fake_fetched_page(),
)
session._provider = seam_provider("\n\n", provider_name="openai")
call_id, answer = session._exec_web_fetch(
{"call_id": "wf3", "url": "https://example.com/x", "question": "What?"}
)
assert answer == "Error: extraction returned no answer"
+43 -1
View File
@@ -22,11 +22,14 @@ Pinned rules:
text can no longer be a partial tag). text can no longer be a partial tag).
""" """
import random
import pytest import pytest
from tests._reasoning_dialect import CASES as DIALECT_CASES
from tests._session_helpers import make_session from tests._session_helpers import make_session
from turnstone.core.providers import StreamChunk, ToolCallDelta from turnstone.core.providers import StreamChunk, ToolCallDelta
from turnstone.core.streaming_text import ThinkTagSplitter from turnstone.core.streaming_text import ThinkTagSplitter, split_inline_reasoning
class _TokenRecorderUI: class _TokenRecorderUI:
@@ -177,6 +180,45 @@ def test_splitter_standalone_contract():
assert events[-1] == ("tail", True) assert events[-1] == ("tail", True)
@pytest.mark.parametrize("case", DIALECT_CASES, ids=[c.id for c in DIALECT_CASES])
def test_one_shot_dialect_conformance(case):
"""Every dialect-catalog case through the one-shot form, exact lanes."""
content, reasoning = split_inline_reasoning(case.utterance)
assert content == case.content
assert reasoning == case.reasoning
_PASSTHROUGH_CASES = [c for c in DIALECT_CASES if c.passthrough]
@pytest.mark.parametrize("case", _PASSTHROUGH_CASES, ids=[c.id for c in _PASSTHROUGH_CASES])
def test_one_shot_passthrough_byte_identity(case):
"""Unconsumed input (tag-free or orphan-close-only) returns
byte-identical every generated row asserts."""
content, _ = split_inline_reasoning(case.utterance)
assert content == case.utterance
@pytest.mark.parametrize("case", DIALECT_CASES, ids=[c.id for c in DIALECT_CASES])
def test_one_shot_equivalent_to_streaming_over_random_chunkings(case):
"""One-shot ≡ the streaming class fed the same utterance in arbitrary
chunkings, EXACTLY the one-shot is a pure raw split with no rules of
its own for EVERY catalog case."""
rng = random.Random(case.id) # deterministic per case
one_content, one_reasoning = split_inline_reasoning(case.utterance)
for _ in range(25):
spans = []
splitter = ThinkTagSplitter(lambda text, is_r, _s=spans: _s.append((text, is_r)))
i = 0
while i < len(case.utterance):
j = rng.randint(i + 1, len(case.utterance))
splitter.feed(case.utterance[i:j])
i = j
splitter.flush_pending()
assert "".join(t for t, is_r in spans if not is_r) == one_content
assert "".join(t for t, is_r in spans if is_r) == one_reasoning
def test_tool_calls_flush_pending_raw_at_current_state(): def test_tool_calls_flush_pending_raw_at_current_state():
# Once tool calls begin, buffered text cannot be a partial tag: it # Once tool calls begin, buffered text cannot be a partial tag: it
# flushes RAW (no tag scan) at the current in_think state. # flushes RAW (no tag scan) at the current in_think state.
+52
View File
@@ -19,6 +19,8 @@ from turnstone.core.trajectory import (
TextBlock, TextBlock,
ToolCall, ToolCall,
Turn, Turn,
final_assistant_text,
last_assistant_text,
materialize_attachments, materialize_attachments,
resolve_attachment_parts, resolve_attachment_parts,
turn_from_dict, turn_from_dict,
@@ -287,3 +289,53 @@ def test_materialize_attachments_noop_without_resolver_or_placeholders() -> None
messages = [{"role": "user", "content": "hi"}] messages = [{"role": "user", "content": "hi"}]
assert materialize_attachments(messages, None) is messages assert materialize_attachments(messages, None) is messages
assert materialize_attachments(messages, lambda ids: {}) is messages assert materialize_attachments(messages, lambda ids: {}) is messages
def test_last_assistant_text_picks_most_recent_substantive() -> None:
turns = [
Turn.user("q"),
Turn.assistant("early answer"),
Turn.assistant("", tool_calls=(ToolCall(id="c1", name="bash", arguments="{}"),)),
Turn.tool("c1", "out"),
Turn.assistant("final answer"),
]
assert last_assistant_text(turns) == "final answer"
def test_last_assistant_text_skips_textless_finals() -> None:
# Tool-call-only and all-reasoning (empty-text) finals are skipped —
# the walk falls back to the last SUBSTANTIVE assistant text.
turns = [
Turn.assistant("substantive"),
Turn.assistant("", tool_calls=(ToolCall(id="c2", name="bash", arguments="{}"),)),
Turn.assistant(""),
]
assert last_assistant_text(turns) == "substantive"
def test_last_assistant_text_none_when_no_assistant_text() -> None:
assert last_assistant_text([]) is None
assert last_assistant_text([Turn.user("q"), Turn.assistant("")]) is None
def test_last_assistant_text_skips_whitespace_only_turns() -> None:
# "Substantive" means non-blank, not merely truthy: a whitespace-only
# final turn must not be salvaged as partial work.
turns = [Turn.assistant("real work"), Turn.assistant(" \n")]
assert last_assistant_text(turns) == "real work"
assert last_assistant_text([Turn.assistant(" \n")]) is None
def test_final_assistant_text_reads_last_turn_only() -> None:
# The final-say read: no walk-back — an empty final say reports empty,
# never an earlier turn's narration.
substantive_then_empty = [Turn.assistant("mid-loop narration"), Turn.assistant("")]
assert final_assistant_text(substantive_then_empty) == ""
tool_final = [
Turn.assistant("narration"),
Turn.assistant("", tool_calls=(ToolCall(id="c1", name="bash", arguments="{}"),)),
]
assert final_assistant_text(tool_final) == ""
assert final_assistant_text([Turn.assistant(" the answer\n")]) == "the answer"
assert final_assistant_text([Turn.user("q")]) == ""
assert final_assistant_text([]) == ""
+8
View File
@@ -2486,6 +2486,14 @@ def _last_assistant_text(storage: Any, ws_id: str) -> str | None:
"""Walk the conversation tail backward and return the most recent """Walk the conversation tail backward and return the most recent
assistant message's text content. assistant message's text content.
A same-named Turn-based walk with DIFFERENT deliberate semantics
lives in ``core/trajectory.last_assistant_text`` (flattens
multi-block content via ``Turn.text``, skips whitespace-only says,
two-state return); this dict-row walk deliberately skips
list-content rows and keeps its tri-state contract for the storage
try/except. A substantiveness-rule change there does not reach
here, and vice versa; keep both docstrings pointing at each other.
Returns: Returns:
- The content string when the tail contains an assistant message - The content string when the tail contains an assistant message
with a non-empty ``content`` field. with a non-empty ``content`` field.
+15 -11
View File
@@ -54,7 +54,10 @@ from turnstone.core.lowering import (
restore_provider_tool_ids, restore_provider_tool_ids,
sanitize_tool_call_arguments, sanitize_tool_call_arguments,
) )
from turnstone.core.providers._protocol import drain_stream from turnstone.core.providers._protocol import (
drain_stream,
has_reasoning_bearing_block,
)
from turnstone.core.storage._utils import ( from turnstone.core.storage._utils import (
_CLIENT_TOOL_CALL_BLOCK_TYPES, _CLIENT_TOOL_CALL_BLOCK_TYPES,
strip_orphan_client_tool_blocks, strip_orphan_client_tool_blocks,
@@ -75,13 +78,10 @@ _DRAIN_RETRIES = 2
# through. Module-level so tests can zero it. # through. Module-level so tests can zero it.
_DRAIN_RETRY_BASE_DELAY = 0.5 _DRAIN_RETRY_BASE_DELAY = 0.5
# Block types that carry model reasoning natively. Anthropic emits # Native-reasoning block membership lives in providers._protocol
# ``thinking``/``redacted_thinking`` blocks, OpenAI Responses emits # (REASONING_BEARING_BLOCK_TYPES + has_reasoning_bearing_block, beside the
# ``reasoning`` items, and ``reasoning_text`` is our own synthetic # drain's double-reasoning check); this layer consumes the shared
# path-3 block (see :func:`synth_reasoning_block`). # predicate in :func:`synth_reasoning_block`.
REASONING_BEARING_BLOCK_TYPES: frozenset[str] = frozenset(
{"thinking", "redacted_thinking", "reasoning", "reasoning_text"}
)
# --------------------------------------------------------------------------- # # --------------------------------------------------------------------------- #
@@ -520,9 +520,13 @@ def synth_reasoning_block(
text = "".join(reasoning_parts) text = "".join(reasoning_parts)
if not text.strip(): if not text.strip():
return provider_blocks return provider_blocks
for b in provider_blocks: if has_reasoning_bearing_block(provider_blocks):
if isinstance(b, dict) and b.get("type") in REASONING_BEARING_BLOCK_TYPES: # Routine on native-reasoning lanes: reasoning_delta text is the
return provider_blocks # MIRROR of the native block there, so bailing is the correct
# no-op, not an anomaly. The genuinely anomalous shape —
# inline-EXTRACTED text alongside a native block — is logged at
# the drain, where extraction is distinguishable.
return provider_blocks
block: dict[str, Any] = {"type": "reasoning_text", "text": text} block: dict[str, Any] = {"type": "reasoning_text", "text": text}
if cfg is ...: if cfg is ...:
cfg = _get_config_or_none(registry, alias) cfg = _get_config_or_none(registry, alias)
+1 -1
View File
@@ -596,7 +596,7 @@ class OutputGuardJudge:
verdict_id, call_id, start, f"provider_error: {type(e).__name__}" verdict_id, call_id, start, f"provider_error: {type(e).__name__}"
) )
content = (getattr(result, "content", "") or "").strip() content = (result.content or "").strip()
if not content: if not content:
return self._error_verdict(verdict_id, call_id, start, "empty_response") return self._error_verdict(verdict_id, call_id, start, "empty_response")
+14 -1
View File
@@ -178,7 +178,12 @@ def describe_cached(
load-bearing for delegated backend authentication: a description produced load-bearing for delegated backend authentication: a description produced
under one user's OBO grant must never be served to another user without a under one user's OBO grant must never be served to another user without a
call authorized as that user. Returns ``""`` on a backend failure (a call authorized as that user. Returns ``""`` on a backend failure (a
placeholder is rendered upstream) and does *not* cache failures. placeholder is rendered upstream) and does *not* cache failures. A
completed-but-EMPTY description memoizes like any other result one
perceive per key, ever (an all-reasoning pass pins the placeholder;
the remediation is server-side: a reasoning parser or the template
thinking toggle on the perception alias) under one guard: an empty
result NEVER overwrites a concurrently memoized real description.
""" """
key = _cache_key(principal_id=principal_id, alias=alias, content_hash=content_hash) key = _cache_key(principal_id=principal_id, alias=alias, content_hash=content_hash)
with _cache_lock: with _cache_lock:
@@ -201,6 +206,14 @@ def describe_cached(
log.warning("perception fallback failed (alias=%s): %s", alias, exc) log.warning("perception fallback failed (alias=%s): %s", alias, exc)
return "" return ""
with _cache_lock: with _cache_lock:
# Re-check under the lock: the describe call ran unlocked, and a
# concurrent racer may have memoized a REAL description — an empty
# result must never clobber it (the memo has no invalidation
# path, so a clobber would pin the placeholder despite a billed,
# successful perceive).
existing = _cache.get(key)
if existing:
return existing
if key not in _cache and len(_cache) >= _CACHE_MAX: if key not in _cache and len(_cache) >= _CACHE_MAX:
_cache.pop(next(iter(_cache)), None) _cache.pop(next(iter(_cache)), None)
_cache[key] = text _cache[key] = text
+114 -13
View File
@@ -10,6 +10,8 @@ from __future__ import annotations
from dataclasses import dataclass, field, replace from dataclasses import dataclass, field, replace
from typing import TYPE_CHECKING, Any, Protocol, runtime_checkable from typing import TYPE_CHECKING, Any, Protocol, runtime_checkable
from turnstone.core.streaming_text import split_inline_reasoning, strip_blank_edge_lines
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Callable, Iterator from collections.abc import Callable, Iterator
@@ -158,6 +160,39 @@ def accumulate_tool_call_delta(
return tc return tc
def _logger() -> Any:
"""Module logger behind ONE deferred import — structlog stays off the
type-module import path (this module is imported for its dataclasses
by code that must not pay the logging stack's import cost)."""
import structlog # noqa: PLC0415 — deferred off the type-module import path
return structlog.get_logger(__name__)
# Block types that carry model reasoning natively. Anthropic emits
# ``thinking``/``redacted_thinking`` blocks, OpenAI Responses emits
# ``reasoning`` items, and ``reasoning_text`` is the harness's own
# synthetic block (``model_turn.synth_reasoning_block``). Defined here —
# beside the drain that needs it for the double-reasoning observability
# check — and imported by ``model_turn`` (which layers above this module).
REASONING_BEARING_BLOCK_TYPES: frozenset[str] = frozenset(
{"thinking", "redacted_thinking", "reasoning", "reasoning_text"}
)
def has_reasoning_bearing_block(blocks: list[dict[str, Any]]) -> bool:
"""True when any block carries model reasoning natively.
THE membership predicate for :data:`REASONING_BEARING_BLOCK_TYPES`
shared by the drain's double-reasoning check and
``model_turn.synth_reasoning_block``'s bail, so the two can never
disagree about what counts as native reasoning.
"""
return any(
isinstance(b, dict) and b.get("type") in REASONING_BEARING_BLOCK_TYPES for b in blocks
)
def transport_guarded(chunks: Iterator[StreamChunk]) -> Iterator[StreamChunk]: def transport_guarded(chunks: Iterator[StreamChunk]) -> Iterator[StreamChunk]:
"""Normalize mid-body transport deaths on a ``create_streaming`` iterator. """Normalize mid-body transport deaths on a ``create_streaming`` iterator.
@@ -191,14 +226,12 @@ def transport_guarded(chunks: Iterator[StreamChunk]) -> Iterator[StreamChunk]:
return return
except httpx.TransportError as exc: except httpx.TransportError as exc:
if finish_seen: if finish_seen:
import structlog # noqa: PLC0415 — deferred with httpx off the type-module path
# usage_captured distinguishes "completed result kept but # usage_captured distinguishes "completed result kept but
# its spend went missing from usage accounting" (the chat # its spend went missing from usage accounting" (the chat
# lane's usage chunk trails the finish) from a harmless # lane's usage chunk trails the finish) from a harmless
# citation-footer loss — the one log signal that lets a # citation-footer loss — the one log signal that lets a
# missing-spend incident be attributed afterward. # missing-spend incident be attributed afterward.
structlog.get_logger(__name__).warning( _logger().warning(
"stream.post_finish_blip", "stream.post_finish_blip",
error_type=type(exc).__name__, error_type=type(exc).__name__,
usage_captured=usage_seen, usage_captured=usage_seen,
@@ -235,6 +268,16 @@ def drain_stream(chunks: Iterator[StreamChunk]) -> CompletionResult:
metadata on the chat lane the usage chunk trails the finish metadata on the chat lane the usage chunk trails the finish
reason, so that result may report usage=None and the call's spend reason, so that result may report usage=None and the call's spend
goes missing from usage accounting. goes missing from usage accounting.
- Content accumulates in RUNS bounded by interleaving signals
(reasoning_delta, tool-call deltas), each run split independently by
:func:`split_inline_reasoning` the one-shot form of the interactive
lane's tag splitter, with the run boundaries mirroring that
consumer's flush-and-reset at the same signals — so ``content`` is
inline-think-tag-free by construction for EVERY drained consumer, and
non-blank extracted text is APPENDED to ``reasoning`` after any
server-parsed ``reasoning_delta`` (segregated, never discarded).
The split runs BEFORE the citations footer folds back: the footer is
web-controlled text and is never scanned for tags.
- ``usage`` merges via :func:`merge_usage` Anthropic splits prompt - ``usage`` merges via :func:`merge_usage` Anthropic splits prompt
and completion tokens across separate events. and completion tokens across separate events.
- Tool calls accumulate by ``ToolCallDelta.index`` via - Tool calls accumulate by ``ToolCallDelta.index`` via
@@ -248,9 +291,12 @@ def drain_stream(chunks: Iterator[StreamChunk]) -> CompletionResult:
- ``info_delta`` before the finish reason is transient status (server- - ``info_delta`` before the finish reason is transient status (server-
side search pings) that the non-streaming lane never surfaced side search pings) that the non-streaming lane never surfaced
dropped. ``info_delta`` after the finish reason is the citations dropped. ``info_delta`` after the finish reason is the citations
footer (``format_citations("", annotations).strip()``); folding it footer (``format_citations("", annotations).strip()``), folded back
back as ``content + "\\n\\n" + info`` byte-matches the non-streaming as ``content + "\\n\\n" + info`` ONLY when the post-split content is
lane's ``format_citations(content, annotations)`` append. non-blank sourcing for an answer that does not exist is dropped,
never handed to downstream emptiness checks as a footer-only
"answer". (This deliberately replaced the retired non-streaming
lane's unconditional byte-match append.)
Raises whatever the underlying stream raises retry/deadline/fallback Raises whatever the underlying stream raises retry/deadline/fallback
policy stays with the caller, exactly as with the old non-streaming policy stays with the caller, exactly as with the old non-streaming
@@ -262,7 +308,8 @@ def drain_stream(chunks: Iterator[StreamChunk]) -> CompletionResult:
post-finish blip ends the stream cleanly so the completed result is post-finish blip ends the stream cleanly so the completed result is
kept. kept.
""" """
content_parts: list[str] = [] content_segments: list[str] = []
segment_parts: list[str] = []
reasoning_parts: list[str] = [] reasoning_parts: list[str] = []
trailing_info_parts: list[str] = [] trailing_info_parts: list[str] = []
tool_calls_acc: dict[int, dict[str, Any]] = {} tool_calls_acc: dict[int, dict[str, Any]] = {}
@@ -270,11 +317,32 @@ def drain_stream(chunks: Iterator[StreamChunk]) -> CompletionResult:
finish_reason: str | None = None finish_reason: str | None = None
provider_blocks: list[dict[str, Any]] = [] provider_blocks: list[dict[str, Any]] = []
def _close_segment() -> None:
if segment_parts:
content_segments.append("".join(segment_parts))
segment_parts.clear()
for sc in transport_guarded(chunks): for sc in transport_guarded(chunks):
if sc.content_delta: # Content accumulates in RUNS bounded by interleaving signals
content_parts.append(sc.content_delta) # (reasoning_delta, tool-call deltas), and each run is split
# independently below — mirroring the interactive consumer's
# boundary resets (pending flushed and in_think cleared when tool
# calls begin; in_think cleared when content resumes after
# provider-parsed reasoning). Without the boundaries, an
# unterminated ``<think>`` opened before a tool call would swallow
# the post-tool-call answer the interactive lane renders as
# content. Within a chunk the interactive ordering holds:
# reasoning (Path 1), then content (Path 2), then tool calls —
# a combined content+tools chunk feeds its content BEFORE the
# tool-call close, so that content belongs to the pre-boundary
# run exactly as the interactive consumer emits it.
if sc.reasoning_delta: if sc.reasoning_delta:
reasoning_parts.append(sc.reasoning_delta) reasoning_parts.append(sc.reasoning_delta)
_close_segment()
if sc.content_delta:
segment_parts.append(sc.content_delta)
if sc.tool_call_deltas:
_close_segment()
for tcd in sc.tool_call_deltas: for tcd in sc.tool_call_deltas:
accumulate_tool_call_delta(tool_calls_acc, tcd) accumulate_tool_call_delta(tool_calls_acc, tcd)
if sc.usage is not None: if sc.usage is not None:
@@ -287,6 +355,7 @@ def drain_stream(chunks: Iterator[StreamChunk]) -> CompletionResult:
# only the trailing (post-finish) citations footer folds back. # only the trailing (post-finish) citations footer folds back.
if sc.info_delta and finish_reason is not None: if sc.info_delta and finish_reason is not None:
trailing_info_parts.append(sc.info_delta) trailing_info_parts.append(sc.info_delta)
_close_segment()
if finish_reason is None: if finish_reason is None:
raise IncompleteStreamError( raise IncompleteStreamError(
@@ -295,9 +364,41 @@ def drain_stream(chunks: Iterator[StreamChunk]) -> CompletionResult:
"declared in its model capabilities)" "declared in its model capabilities)"
) )
content = "".join(content_parts) # THE one trim: runs split raw, then a single edge trim over the
for info in trailing_info_parts: # joined whole when any run consumed a tag. Per-run trimming cannot
content += "\n\n" + info # distinguish tag residue from a genuine paragraph separator the
# model emitted just before an interleaving signal — trimming each
# run's edges fused sentences across the separator-less join.
split_segments = [split_inline_reasoning(seg) for seg in content_segments]
content = "".join(c for c, _ in split_segments)
extracted = "".join(r for _, r in split_segments)
# The splitter can only REMOVE characters, so a shrunken total is the
# consumed-a-tag signal.
if len(content) != sum(map(len, content_segments)):
content = strip_blank_edge_lines(content)
reasoning = "".join(reasoning_parts)
if extracted.strip():
if reasoning:
# Server-parsed and inline-extracted are distinct passes — keep
# a boundary so they never read as one run-together sentence.
reasoning += "\n\n"
reasoning += extracted
if has_reasoning_bearing_block(provider_blocks):
# True double-reasoning shape (inline tags AND a native
# reasoning block): the extracted text has no native lane to
# land in downstream. Observable here — where extraction is
# distinguishable from the routine reasoning_delta mirror —
# and chars-only: reasoning text is barred from log payloads.
_logger().debug("drain.inline_reasoning_alongside_native", chars=len(extracted))
# A turn with no visible answer — content empty OR whitespace-only —
# must not gain a citations footer: sourcing for an answer that does
# not exist, folded on, would hand every downstream emptiness check a
# truthy footer-only "answer". Blankness, not truthiness; checked
# only when a footer exists (footers ride web-search turns only, and
# the strip scan shouldn't tax every drained completion).
if trailing_info_parts and content.strip():
for info in trailing_info_parts:
content += "\n\n" + info
tool_calls = [tool_calls_acc[i] for i in sorted(tool_calls_acc)] tool_calls = [tool_calls_acc[i] for i in sorted(tool_calls_acc)]
return CompletionResult( return CompletionResult(
@@ -306,7 +407,7 @@ def drain_stream(chunks: Iterator[StreamChunk]) -> CompletionResult:
finish_reason=finish_reason, finish_reason=finish_reason,
usage=usage, usage=usage,
provider_blocks=provider_blocks, provider_blocks=provider_blocks,
reasoning="".join(reasoning_parts), reasoning=reasoning,
) )
+46 -43
View File
@@ -220,6 +220,7 @@ from turnstone.core.trajectory import (
ToolCall, ToolCall,
Turn, Turn,
dicts_from_turns, dicts_from_turns,
last_assistant_text,
turn_from_dict, turn_from_dict,
turn_to_dict, turn_to_dict,
turns_from_dicts, turns_from_dicts,
@@ -946,10 +947,10 @@ _SPEC_ARGUMENTS_LITERAL_RE = re.compile(r"\$ARGUMENTS\b(?!\[)")
# assignment scheme unless the operator set one), so the budget must fit a # assignment scheme unless the operator set one), so the budget must fit a
# full thinking pass at the MODEL'S OWN default — the prompt's hard word cap # full thinking pass at the MODEL'S OWN default — the prompt's hard word cap
# keeps the visible answer trivially cheap, and ``_TITLE_MAX_TOKENS`` carries # keeps the visible answer trivially cheap, and ``_TITLE_MAX_TOKENS`` carries
# the rest. Recover the title from ``content``: reuse # the rest. Recover the title from ``content`` — IR-clean at the drain
# :meth:`ChatSession._strip_reasoning` (the canonical ``<think>``/ # seam (``split_inline_reasoning``; servers that leave reasoning inline
# ``<reasoning>`` remover, for lanes that leave reasoning inline rather than # rather than in ``reasoning_content`` are segregated there, this lane
# in ``reasoning_content``), take the first non-empty line (a model that # holds no strip of its own): take the first non-empty line (a model that
# appends an explanation shouldn't fold prose into the title), then peel a # appends an explanation shouldn't fold prose into the title), then peel a
# ``Title:`` label and wrapping markdown/quote decoration. Internal # ``Title:`` label and wrapping markdown/quote decoration. Internal
# punctuation is preserved so ``.NET``, ``CI/CD``, ``v1.6.0`` survive. # punctuation is preserved so ``.NET``, ``CI/CD``, ``v1.6.0`` survive.
@@ -1196,6 +1197,12 @@ _BACKEND_KNOWN_EXC_NAMES: frozenset[str] = (
) )
def _non_blank_or(text: str | None, fallback: str) -> str:
"""*text* when it has any non-whitespace, else *fallback* — the
campaign-wide blankness doctrine for drained no-answer fallbacks."""
return text if text is not None and text.strip() else fallback
def _is_ctx_overflow(exc: BaseException) -> bool: def _is_ctx_overflow(exc: BaseException) -> bool:
"""True when *exc* looks like a context-window overflow from any backend. """True when *exc* looks like a context-window overflow from any backend.
@@ -3972,18 +3979,22 @@ class ChatSession:
) )
raw = result.content or "" raw = result.content or ""
log.info("ws.title.llm_response", ws_id=ws_id[:8], raw=raw[:200]) log.info("ws.title.llm_response", ws_id=ws_id[:8], raw=raw[:200])
# Take the assistant's answer (``content``), never its reasoning. # ``content`` arrives with balanced/unterminated inline reasoning
# Reuse the canonical reasoning stripper, then drop a leftover close # already segregated at the drain seam (``split_inline_reasoning``).
# tag from lanes that pre-inject the opening ``<think>`` into the # The seam deliberately passes ORPHAN close tags through (a close
# prompt (only ``</think>`` reaches ``content``). See ``_TITLE_*``. # with no open is indistinguishable from quoted prose, and
stripped = self._strip_reasoning(raw) # reclassifying would let quoted text destroy real answers) —
# for a 3-word display string the cheap cosmetic call goes the
# other way, so peel through the LAST stray close tag here.
# This is title formatting like ``_TITLE_WRAP_CHARS``, not
# reasoning segregation. See ``_TITLE_*``.
for _close in ThinkTagSplitter.CLOSE_TAGS: for _close in ThinkTagSplitter.CLOSE_TAGS:
_pos = stripped.rfind(_close) _pos = raw.rfind(_close)
if _pos != -1: if _pos != -1:
stripped = stripped[_pos + len(_close) :] raw = raw[_pos + len(_close) :]
# First non-empty line, with a ``Title:`` label and wrapping # First non-empty line, with a ``Title:`` label and wrapping
# markdown/quote decoration peeled (internal punctuation kept). # markdown/quote decoration peeled (internal punctuation kept).
line = next((ln for ln in stripped.splitlines() if ln.strip()), "") line = next((ln for ln in raw.splitlines() if ln.strip()), "")
line = _TITLE_LABEL_RE.sub("", line.strip(_TITLE_WRAP_CHARS)) line = _TITLE_LABEL_RE.sub("", line.strip(_TITLE_WRAP_CHARS))
title = line.strip(_TITLE_WRAP_CHARS)[:_TITLE_MAX_CHARS] title = line.strip(_TITLE_WRAP_CHARS)[:_TITLE_MAX_CHARS]
if title and self._ws_id == ws_id: if title and self._ws_id == ws_id:
@@ -5579,6 +5590,11 @@ class ChatSession:
"""Run a lightweight internal completion (title gen, compaction, """Run a lightweight internal completion (title gen, compaction,
extraction) through ``model_turn`` on the session's primary lane. extraction) through ``model_turn`` on the session's primary lane.
Inline-reasoning segregation is SEAM-OWNED (``drain_stream`` runs
``split_inline_reasoning``): ``result.content`` is think-tag-free
for every caller of this funnel present and future and no
caller may add a private strip.
``max_tokens`` is clamped to the model's advertised output limit so ``max_tokens`` is clamped to the model's advertised output limit so
small models don't error. small models don't error.
@@ -7784,24 +7800,6 @@ class ChatSession:
self._persist_truncation(removed_count) self._persist_truncation(removed_count)
return content return content
@staticmethod
def _strip_reasoning(text: str) -> str:
"""Remove think-tag blocks and their content.
The tag vocabulary is ThinkTagSplitter's — deriving it keeps this
strip (and the title lane's) from going blind when a variant is
added to the splitter, which would leak raw chain-of-thought into
compaction summaries.
"""
for open_t, close_t in zip(
ThinkTagSplitter.OPEN_TAGS, ThinkTagSplitter.CLOSE_TAGS, strict=True
):
while open_t in text:
start = text.find(open_t)
end = text.find(close_t, start)
text = text[:start] + text[end + len(close_t) :] if end != -1 else text[:start]
return text.strip()
def _ui_stream_discarded(self) -> None: def _ui_stream_discarded(self) -> None:
"""Best-effort dead-segment discard across UI generations. """Best-effort dead-segment discard across UI generations.
@@ -8871,9 +8869,9 @@ class ChatSession:
def _summarize_once(self, system_prompt: str, body: str, my_generation: int = 0) -> str: def _summarize_once(self, system_prompt: str, body: str, my_generation: int = 0) -> str:
"""Run one summary completion over ``body`` and return the cleaned text. """Run one summary completion over ``body`` and return the cleaned text.
Owns the retry loop (transient errors only, exponential backoff) and the Owns the retry loop (transient errors only, exponential backoff).
reasoning-tag strip. Raises on a non-retryable error or retry exhaustion Raises on a non-retryable error or retry exhaustion so the caller
so the caller can abort the whole compaction before any message swap. can abort the whole compaction before any message swap.
""" """
summary_msgs = [ summary_msgs = [
Turn.system(system_prompt), Turn.system(system_prompt),
@@ -8919,8 +8917,11 @@ class ChatSession:
) )
self._backoff_or_cancelled(delay, my_generation) self._backoff_or_cancelled(delay, my_generation)
assert result is not None assert result is not None
# Strip any <think>/<reasoning> tags the summarizer may emit # Inline think tags are already segregated at the drain seam; the
summary = self._strip_reasoning(result.content or "") # trim is summary formatting (tag-free output is deliberately
# byte-identical at the seam, so edge whitespace is trimmed here
# before the summary is persisted and re-joined into merge input).
summary = (result.content or "").strip()
if result.finish_reason == "length": if result.finish_reason == "length":
self._compaction_event( self._compaction_event(
my_generation, {"phase": "progress", "warning": "summary_truncated"} my_generation, {"phase": "progress", "warning": "summary_truncated"}
@@ -17738,10 +17739,10 @@ class ChatSession:
# propagates past this ``except Exception``. # propagates past this ``except Exception``.
overflow = _is_ctx_overflow(e) overflow = _is_ctx_overflow(e)
note = "context limit reached" if overflow else f"error ({type(e).__name__})" note = "context limit reached" if overflow else f"error ({type(e).__name__})"
for t in reversed(agent_turns): salvage = last_assistant_text(agent_turns)
if t.role is Role.ASSISTANT and t.text: if salvage:
self.ui.on_info(f"[{label}] {note}, returning partial work") self.ui.on_info(f"[{label}] {note}, returning partial work")
return self._guard_subagent_synthesis(t.text, label) return self._guard_subagent_synthesis(salvage, label)
# No partial work to salvage: surface overflow as a calm stop message, # No partial work to salvage: surface overflow as a calm stop message,
# but re-raise any other terminal error so the real failure isn't # but re-raise any other terminal error so the real failure isn't
# masked as an empty success. # masked as an empty success.
@@ -17753,7 +17754,9 @@ class ChatSession:
# Handle truncation or content filter — stop agent early # Handle truncation or content filter — stop agent early
if result.finish_reason == "length": if result.finish_reason == "length":
self.ui.on_info(f"[{label}] response truncated, stopping early") self.ui.on_info(f"[{label}] response truncated, stopping early")
return self._guard_subagent_synthesis(result.content or "(truncated)", label) return self._guard_subagent_synthesis(
_non_blank_or(result.content, "(truncated)"), label
)
if result.finish_reason == "content_filter": if result.finish_reason == "content_filter":
self.ui.on_info(f"[{label}] blocked by content filter") self.ui.on_info(f"[{label}] blocked by content filter")
return "(content filter)" return "(content filter)"
@@ -17766,7 +17769,7 @@ class ChatSession:
agent_turns.append(result.turn) agent_turns.append(result.turn)
if not result.tool_calls: if not result.tool_calls:
content = result.content or "(no output)" content = _non_blank_or(result.content, "(no output)")
self.ui.on_info(f"[{label} done] {len(content)} chars") self.ui.on_info(f"[{label} done] {len(content)} chars")
return self._guard_subagent_synthesis(content, label) return self._guard_subagent_synthesis(content, label)
@@ -17922,7 +17925,7 @@ class ChatSession:
) )
) )
result = _api_call(agent_turns, _tools=[]) result = _api_call(agent_turns, _tools=[])
content = result.content or "(no output)" content = _non_blank_or(result.content, "(no output)")
self.ui.on_info(f"[{label} done] {len(content)} chars") self.ui.on_info(f"[{label} done] {len(content)} chars")
return self._guard_subagent_synthesis(content, label) return self._guard_subagent_synthesis(content, label)
@@ -19157,7 +19160,7 @@ class ChatSession:
reasoning_effort=self.reasoning_effort, reasoning_effort=self.reasoning_effort,
) )
answer = result.content or "" answer = result.content or ""
if not answer: if not answer.strip():
answer = "Error: extraction returned no answer" answer = "Error: extraction returned no answer"
except Exception as e: except Exception as e:
answer = f"Extraction failed (page was fetched but summarization errored): {e}" answer = f"Extraction failed (page was fetched but summarization errored): {e}"
+78 -2
View File
@@ -1,4 +1,11 @@
"""Streaming think-tag splitting for interactive content streams.""" """Think-tag splitting: streaming (interactive) and one-shot (drained) forms.
:class:`ThinkTagSplitter` is the ONE tag-semantics engine vocabulary,
earliest-tag-wins selection, any-close-closes-any-open, the partial-tag
carry. :func:`split_inline_reasoning` is its drained form: the same
engine applied to a complete text (feed + flush), plus the residue
whitespace rule. All tag selection happens inside the class; the
one-shot's fast path tests only tag PRESENCE, never position."""
from __future__ import annotations from __future__ import annotations
@@ -33,7 +40,8 @@ class ThinkTagSplitter:
OPEN_TAGS: tuple[str, ...] = ("<think>", "<reasoning>") OPEN_TAGS: tuple[str, ...] = ("<think>", "<reasoning>")
CLOSE_TAGS: tuple[str, ...] = ("</think>", "</reasoning>") CLOSE_TAGS: tuple[str, ...] = ("</think>", "</reasoning>")
MAX_TAG_LEN = max(len(t) for t in OPEN_TAGS + CLOSE_TAGS) ALL_TAGS: tuple[str, ...] = OPEN_TAGS + CLOSE_TAGS
MAX_TAG_LEN = max(len(t) for t in ALL_TAGS)
def __init__(self, emit: Callable[[str, bool], None]) -> None: def __init__(self, emit: Callable[[str, bool], None]) -> None:
self._emit = emit self._emit = emit
@@ -80,3 +88,71 @@ class ThinkTagSplitter:
self._emit(self.pending[:safe], self.in_think) self._emit(self.pending[:safe], self.in_think)
self.pending = self.pending[safe:] self.pending = self.pending[safe:]
break break
def split_inline_reasoning(text: str) -> tuple[str, str]:
"""Split a complete drained text into ``(content, reasoning)``.
The one-shot form of :class:`ThinkTagSplitter` for non-streaming
consumers (``drain_stream``): a plain feed-and-flush of ONE content
run the interactive lane's per-run rule, no more. The caller owns
run boundaries (``drain_stream`` closes a run when tool-call deltas
or provider-parsed reasoning interleave, mirroring the interactive
consumer's flush-and-reset at those signals); a partial tag never
spans an interleaving signal. Balanced
blocks land in the reasoning lane; an unterminated open sends the
tail to reasoning; an orphan CLOSE tag (no prior open) stays in
content untouched. That last case is deliberate: a close tag whose
open never arrived is indistinguishable from prose that merely
QUOTES the tag, and drained lanes routinely quote third-party text
(a web-fetch answer citing a page about reasoning models, a guard
verdict echoing judged content) reclassifying everything before
it would let that text destroy the result. Display-string lanes
that want stricter cosmetic peeling (the title) own it locally as
formatting, not segregation.
The split is RAW: tag residue (the ``"\\n\\n"`` a leading block
leaves behind) stays in the returned content. Exactly ONE trim
policy exists in the tree and ``drain_stream`` owns it it joins
the per-run splits and applies :func:`strip_blank_edge_lines` once
over the whole when any run consumed a tag (a run edge may be
INTERIOR after joining, where a genuine paragraph separator must
survive). With no tag present anywhere the input returns
byte-identical (fast path), so tag-free lanes cannot drift.
"""
if not any(tag in text for tag in ThinkTagSplitter.ALL_TAGS):
return text, ""
content_parts: list[str] = []
reasoning_parts: list[str] = []
def _collect(span: str, is_reasoning: bool) -> None:
(reasoning_parts if is_reasoning else content_parts).append(span)
splitter = ThinkTagSplitter(_collect)
splitter.feed(text)
splitter.flush_pending()
content = "".join(content_parts)
if len(content) == len(text):
# Only orphan close tags were present — nothing was consumed;
# the text passes through byte-identical.
return text, ""
return content, "".join(reasoning_parts)
def strip_blank_edge_lines(text: str) -> str:
"""Remove leading/trailing lines that are entirely whitespace.
The residue-trim unit ``drain_stream`` applies once over its joined
per-run splits: kills the separator lines a consumed tag leaves
behind while preserving the first surviving line's significant
indentation ``.strip()`` would delete it and silently reformat
whitespace-significant output.
"""
lines = text.split("\n")
start, end = 0, len(lines)
while start < end and not lines[start].strip():
start += 1
while end > start and not lines[end - 1].strip():
end -= 1
return "\n".join(lines[start:end])
+38 -1
View File
@@ -23,7 +23,7 @@ from enum import StrEnum
from typing import TYPE_CHECKING, Any from typing import TYPE_CHECKING, Any
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Callable from collections.abc import Callable, Sequence
class Role(StrEnum): class Role(StrEnum):
@@ -381,6 +381,43 @@ def dicts_from_turns(turns: list[Turn]) -> list[dict[str, Any]]:
return [turn_to_dict(t) for t in turns] return [turn_to_dict(t) for t in turns]
def last_assistant_text(turns: Sequence[Turn]) -> str | None:
"""Text of the most recent assistant turn with non-blank text.
The SALVAGE walk (partial-work recovery): assistant turns that said
nothing quotable tool-call-only, all-reasoning, or whitespace-only
are skipped, so the caller recovers the last SUBSTANTIVE assistant
text, or ``None`` when the trajectory has none. A same-named
dict-row walk with DIFFERENT deliberate semantics lives in
``console/coordinator_client._last_assistant_text`` (skips
list-content rows, tri-state return for its storage try/except)
a substantiveness-rule change here does not reach it, and vice
versa; keep both docstrings pointing at each other. Skipping means an
empty final turn falls back to an EARLIER turn's text, which is
exactly what salvage wants and exactly what a final-answer read must
NOT do callers reporting "the final say" (an analyst diagnosis, an
eval's final_content) should read the last assistant turn directly
instead of using this walk.
"""
for t in reversed(turns):
if t.role is Role.ASSISTANT and t.text.strip():
return t.text
return None
def final_assistant_text(turns: Sequence[Turn]) -> str:
"""The LAST assistant turn's text, stripped — ``""`` when the
trajectory has no assistant turn or its final say was empty
(tool-call-only, all-reasoning, or whitespace).
The FINAL-SAY read (optimizer analyst diagnosis, eval final_content):
NO walk-back an empty final say reports empty, never replaced by an
earlier turn's mid-loop narration presented as the conclusion.
Salvage wants the opposite walk: :func:`last_assistant_text`.
"""
return next((t.text for t in reversed(turns) if t.role is Role.ASSISTANT), "").strip()
def resolve_attachment_parts( def resolve_attachment_parts(
messages: list[dict[str, Any]], parts_by_id: dict[str, Any] messages: list[dict[str, Any]], parts_by_id: dict[str, Any]
) -> list[dict[str, Any]]: ) -> list[dict[str, Any]]:
+4 -7
View File
@@ -37,7 +37,7 @@ from turnstone.core.providers import LLMProvider, create_client, create_provider
from turnstone.core.session import ChatSession from turnstone.core.session import ChatSession
from turnstone.core.storage import get_storage, init_storage, reset_storage from turnstone.core.storage import get_storage, init_storage, reset_storage
from turnstone.core.tools import INTERACTIVE_TOOLS, PRIMARY_KEY_MAP from turnstone.core.tools import INTERACTIVE_TOOLS, PRIMARY_KEY_MAP
from turnstone.core.trajectory import Role, Turn, turn_from_dict, turns_from_dicts from turnstone.core.trajectory import Turn, final_assistant_text, turn_from_dict, turns_from_dicts
# Eval evaluates interactive-session agent behaviour — coordinator tools # Eval evaluates interactive-session agent behaviour — coordinator tools
# require a console-hosted session and aren't exercised by the harness. # require a console-hosted session and aren't exercised by the harness.
@@ -779,12 +779,9 @@ def _run_single_test(
return session, _drive return session, _drive
def _finish(session: Any, tool_log: list[dict[str, Any]]) -> dict[str, Any]: def _finish(session: Any, tool_log: list[dict[str, Any]]) -> dict[str, Any]:
# Extract results before releasing session # The run's final content is the final-say read — no walk-back: an
final_content = "" # earlier turn's narration must never be scored as the final answer.
for msg in reversed(session.messages): final_content = final_assistant_text(session.messages)
if msg.role is Role.ASSISTANT and msg.text:
final_content = msg.text
break
return { return {
"tool_log": tool_log, "tool_log": tool_log,
"final_content": final_content, "final_content": final_content,
+35 -69
View File
@@ -30,7 +30,7 @@ from openai import OpenAI
from turnstone.core.model_turn import cap_tool_calls, model_turn, resolve_lane from turnstone.core.model_turn import cap_tool_calls, model_turn, resolve_lane
from turnstone.core.providers import LLMProvider, create_provider from turnstone.core.providers import LLMProvider, create_provider
from turnstone.core.session import ChatSession from turnstone.core.session import ChatSession
from turnstone.core.trajectory import Role, Turn from turnstone.core.trajectory import Turn, final_assistant_text
from turnstone.eval.core import ( from turnstone.eval.core import (
_MCP_ONLY_TOOLS, _MCP_ONLY_TOOLS,
BOLD, BOLD,
@@ -226,6 +226,23 @@ Output the modified optimizer instructions only.\
""" """
def _strip_markdown_fence(text: str) -> str:
"""Unwrap a markdown-fenced block from MODEL output.
THE fence rule, in one place: a complete ``` pair yields its innards
(discarding any prose outside the fences); an unterminated opening
fence drops just the fence line. Callers apply this to normalized
model output ONLY never to an ``or``-fallback value, which must
survive verbatim.
"""
fence_match = re.search(r"```[^\n]*\n(.*?)```", text, re.DOTALL)
if fence_match:
return fence_match.group(1).strip()
if text.startswith("```"):
return "\n".join(text.split("\n")[1:]).strip()
return text
def _diversify_prompts( def _diversify_prompts(
client: Any, client: Any,
model: str, model: str,
@@ -304,17 +321,7 @@ def _diversify_prompts(
max_tokens=8192, max_tokens=8192,
) )
raw = (cr.content or "").strip() raw = (cr.content or "").strip()
# Strip reasoning tags raw = _strip_markdown_fence(raw)
raw = re.sub(
r"<(?:think|reasoning)>.*?</(?:think|reasoning)>",
"",
raw,
flags=re.DOTALL,
).strip()
# Strip markdown fences
fence_match = re.search(r"```[^\n]*\n(.*?)```", raw, re.DOTALL)
if fence_match:
raw = fence_match.group(1).strip()
new_variants = json.loads(raw) new_variants = json.loads(raw)
if isinstance(new_variants, list) and all(isinstance(v, str) for v in new_variants): if isinstance(new_variants, list) and all(isinstance(v, str) for v in new_variants):
@@ -454,18 +461,11 @@ def _observe_and_update_optimizer(
max_tokens=8192, max_tokens=8192,
) )
result = cr.content or optimizer_system # Normalize the MODEL's output first (strip, then unfence), and only
result = re.sub( # then fall back: a no-answer pass — empty, whitespace-only, or
r"<(?:think|reasoning)>.*?</(?:think|reasoning)>", # fence-with-nothing — keeps the current observer system VERBATIM,
"", # never wiped and never itself fence-stripped.
result, result = _strip_markdown_fence((cr.content or "").strip()) or optimizer_system
flags=re.DOTALL,
).strip()
# Strip markdown code fences if wrapped
fence_match = re.search(r"```[^\n]*\n(.*?)```", result, re.DOTALL)
if fence_match:
result = fence_match.group(1).strip()
# Reject degenerate outputs (>200% of input length) # Reject degenerate outputs (>200% of input length)
if len(result) > len(optimizer_system) * 2.0: if len(result) > len(optimizer_system) * 2.0:
@@ -789,22 +789,10 @@ def _run_analyst(
output = _exec_analyst_tool(func_name, tc["function"]["arguments"]) output = _exec_analyst_tool(func_name, tc["function"]["arguments"])
turns.append(Turn.tool(tc["id"], output)) turns.append(Turn.tool(tc["id"], output))
# Extract final text response # The analysis is the analyst's FINAL say only — the no-walk-back read
result = "" # (an all-reasoning or silent final turn yields "" and the analyst
for t in reversed(turns): # section is skipped, never an earlier mid-loop narration).
if t.role is Role.ASSISTANT and t.text: return final_assistant_text(turns)
result = t.text
break
# Strip reasoning tags if present
result = re.sub(
r"<(?:think|reasoning)>.*?</(?:think|reasoning)>",
"",
result,
flags=re.DOTALL,
).strip()
return result
TOOL_OPTIMIZER_SYSTEM = """\ TOOL_OPTIMIZER_SYSTEM = """\
@@ -912,15 +900,7 @@ def _propose_tool_overrides(
max_tokens=8192, max_tokens=8192,
) )
raw = (cr.content or "").strip() raw = _strip_markdown_fence((cr.content or "").strip())
# Strip reasoning tags
raw = re.sub(
r"<(?:think|reasoning)>.*?</(?:think|reasoning)>", "", raw, flags=re.DOTALL
).strip()
# Strip markdown fences
fence_match = re.search(r"```[^\n]*\n(.*?)```", raw, re.DOTALL)
if fence_match:
raw = fence_match.group(1).strip()
try: try:
new_overrides = json.loads(raw) new_overrides = json.loads(raw)
@@ -1053,26 +1033,12 @@ def _propose_prompt_modification(
max_tokens=16384, max_tokens=16384,
) )
new_prompt = cr.content or current_prompt # Normalize the MODEL's output first (strip, then unfence), and only
# then fall back: a no-answer pass — empty, whitespace-only, or
# Strip reasoning tags if present # fence-with-nothing — keeps the current prompt VERBATIM (reads as
new_prompt = re.sub( # "no changes" downstream; never an empty-prompt tree node, and the
r"<(?:think|reasoning)>.*?</(?:think|reasoning)>", # fence-strip must never run on the fallback itself).
"", return _strip_markdown_fence((cr.content or "").strip()) or current_prompt
new_prompt,
flags=re.DOTALL,
).strip()
# Strip markdown code fences if the model wrapped the prompt.
# Also discard any explanation text outside the fences.
fence_match = re.search(r"```[^\n]*\n(.*?)```", new_prompt, re.DOTALL)
if fence_match:
new_prompt = fence_match.group(1).strip()
elif new_prompt.startswith("```"):
# Opening fence without closing — strip just the first line
new_prompt = "\n".join(new_prompt.split("\n")[1:]).strip()
return new_prompt
def _simple_diff(old: str, new: str) -> str: def _simple_diff(old: str, new: str) -> str:
+8 -17
View File
@@ -100,7 +100,7 @@ from turnstone.core.session_ui_base import (
fire_judge_verdict_metric, fire_judge_verdict_metric,
) )
from turnstone.core.tools import TOOLS # noqa: F401 — available for introspection from turnstone.core.tools import TOOLS # noqa: F401 — available for introspection
from turnstone.core.trajectory import turn_to_dict from turnstone.core.trajectory import final_assistant_text
from turnstone.core.web_helpers import version_html as _version_html from turnstone.core.web_helpers import version_html as _version_html
from turnstone.core.workstream import ( from turnstone.core.workstream import (
Workstream, Workstream,
@@ -2225,22 +2225,13 @@ def _validate_notify_targets(raw: Any) -> tuple[str, str]:
def _extract_last_assistant_content(session: Any) -> str: def _extract_last_assistant_content(session: Any) -> str:
"""Return the text content of the last assistant message.""" """Return the text of the session's final assistant say.
for turn in reversed(session.messages):
msg = turn_to_dict(turn) THE final-say read (``trajectory.final_assistant_text``): no
if msg.get("role") == "assistant": walk-back, whitespace-only says report empty so the notify
content = msg.get("content", "") fallback fires instead of sending raw whitespace.
if isinstance(content, str): """
return content return final_assistant_text(session.messages)
if isinstance(content, list):
parts = []
for block in content:
if isinstance(block, dict) and block.get("type") == "text":
text = block.get("text")
if isinstance(text, str) and text:
parts.append(text)
return "\n".join(parts)
return ""
def _fire_notify_targets(ws: Any, content: str) -> None: def _fire_notify_targets(ws: Any, content: str) -> None: