mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-21 11:30:59 -06:00
d100ac92d9
* fix: capacity-aware tool output truncation and context overflow recovery Large tool results (e.g. 593K-char search output) could overflow the context window in a single turn when the conversation was already partially full. The fixed 50%-of-context truncation limit didn't account for current usage. Changes: - _truncate_output() now accepts remaining token budget and uses min(tool_truncation, remaining_budget_chars) as the effective limit - _remaining_token_budget() helper calculates available capacity with reserves for max_tokens response and 5% safety margin - Safety truncation at tool-result append: every string tool result is clamped to remaining budget before entering the message array - _exec_web_search() now calls _truncate_output() (was missing) - Context overflow recovery: catches provider errors indicating context length exceeded (OpenAI + Anthropic patterns), auto-compacts, retries once. Falls back to original error if compact-and-retry fails. * fix: address review — zero-budget floor, nested spinner, Anthropic patterns, tests - Remove 256-char floor from budget truncation — zero budget now returns a placeholder instead of allowing 256 chars through - Stop thinking spinner before compact to avoid nested start/stop - Add Anthropic error patterns (prompt is too long, input tokens) - Wrap compact-and-retry so failures re-raise the original error - Add 15 tests covering budget calculation, capacity-aware truncation, and overflow recovery for both providers * fix: cap response reservation at 25% of context window Reserving the full max_tokens in _remaining_token_budget() zeroed the budget for common configs like max_tokens=32768 on a 32K context, collapsing all tool output to a placeholder. max_tokens is a ceiling, not guaranteed consumption — cap the reserve at context_window // 4. Adds regression test for max_tokens >= context_window.