Commit Graph

4 Commits

Author SHA1 Message Date
Patrick Buckley b6391d1f90 fix(model-turn): one sampling-knob assignment scheme — alias > config > model definition > omit
Round-2 review fixes. The round-1 de-pinning collided with
ConfigStore.get's default-on-miss semantics: the registry defaults
(temperature 1.0, effort "medium") were manufactured onto every
store-backed lane's wire, making the documented "unset -> omit"
terminal unreachable. Unset is now representable end to end, and one
scheme governs every lane: per-model alias value > operator-stored
global setting > in-code model definition (effort only: caps
declaration) > field omitted, inference engine's default rules.

- settings_registry: model.temperature default None, model.reasoning_effort
  default "" — the registered defaults ARE the unset sentinels, so the
  admin UI and the wire agree. Admin webux renders nullable floats blank
  ("(inherit model default)") and maps blank-save to reset; the "" effort
  choice reads "(inherit)".
- model_turn: resolve_temperature_setting/resolve_effort_setting are the
  ONE pair of operator-rung resolvers, shared by resolve_lane, both
  session factories, and the /model switch (the 4th-copy mirror is gone;
  the switch no longer leaks the previous model's override on store-less
  sessions). The caps rung moved out of the lane into model_turn's
  effective computation, below a new request-shaped default_reasoning_effort
  parameter (utility + output guard pass "low": budget coherence with
  their small token caps, not sampling policy — any operator or
  model-definition value beats it). The hidden "medium" terminal is gone.
- providers: Protocol + all adapters take reasoning_effort: str | None =
  None (the Protocol-signature "medium" was the same manufactured pin one
  layer down); ModelCapabilities.default_reasoning_effort defaults "" —
  commercial rows all declare theirs explicitly, so only local lanes and
  Anthropic change, both to match their real serving defaults (Anthropic
  manual-thinking models no longer get implicit thinking-on-medium).
  reasoning_template_kwargs distinguishes unset (inject nothing; template
  default rules) from the explicit "none" off-switch. apply_temperature
  skips temperature unless reasoning is EXPLICITLY off on none-declaring
  models (unset leaves the server default in charge, possibly reasoning-on).
- session: ctor takes temperature: float | None / reasoning_effort:
  str | None = None; _save_config/resume round-trip unset as "" (the
  str(None) era guarded); _run_agent relays session temperature AND
  effort on the same-alias fall-through only (a task alias's configured
  knobs stay reachable in both directions).
- optimizer: the five meta lanes are decoupled from --temperature/
  --reasoning-effort (test-model knobs, per their documented meaning);
  registry-less meta lanes omit both fields.
- cli: --temperature/--reasoning-effort default unset and fall through
  the model config instead of pinning 0.5/"medium" for every CLI session.
- cleanup from the review's below-cap findings: dead resolve_server_type
  deleted (tests re-pointed at _server_type_of), stale ChatSession
  comments in _openai_responses fixed, _store_get_or_none extracted,
  eval system-turn conversion hoisted out of the per-turn loop, dead
  _provider_extra_params patch removed, test_perception uses the shared
  mock_completion_result, effort_ladder uses apply_capability_overrides
  instead of a SimpleNamespace fake config.

Wire goldens regenerated: the only drift is the manufactured "medium"
effort vanishing from unset-effort requests (Responses reasoning.effort,
Chat/Google reasoning_effort, Anthropic output_config.effort) — pure
removals, no additions. Ladder tests now fake ConfigStore with the REAL
get() semantics (registry default on miss) so a forgiving fake can't
mask this class of bug again.
2026-07-13 08:48:27 -07:00
Patrick Buckley 09fd17f2da fix(model-turn): reasoning effort rides the ladder too; round-1 review fixes
Patrick's rulings applied from the round-1 high review:

- reasoning_effort loses every code pin, same as temperature: ModelLane
  resolves the ladder (ModelConfig.reasoning_effort → global
  model.reasoning_effort setting → the lane capabilities'
  default_reasoning_effort), model_turn takes str | None, and the pins
  in both judges, all five optimizer lanes, _utility_completion's
  signature default, and model_turn's own "medium" default are gone.
  Explicit relays of user/operator knobs (session effort on the agent
  seam and web-fetch, harness knobs in eval) stay relays.  Effort's
  terminal is the caps default, not wire omission — it gates thinking
  modes, so unset ≠ the explicit "none" value.
- model.temperature setting default 0.5 → 1.0 (safer for modern models;
  several providers no longer accept temperature at all — those drop it
  via capabilities regardless).  No judge-specific knob: a judge alias
  with a per-model override is the remediation path.
- The agent seam keeps the alias ladder (configured → inherited global
  → none), per ruling; the ModelLane docstring no longer documents the
  removed session-relay convention.
- Optimizer lanes get real operator knobs: the existing --temperature /
  --reasoning-effort CLI flags now relay into all five internal LLM
  steps (previously they reached only the eval sessions, leaving the
  deleted pins with no replacement mechanism).

Round-1 cleanups: create_streaming widened to float | None (the
Protocol's two entry points agree; all callers pass explicitly);
model_turn's provider invocation is a direct keyword call again (strict
mypy re-checks it); perception threads the caller's already-resolved
capabilities (one config generation across gate and wire); redundant
extra_params pre-resolution dropped at utility/agent/eval; the synth
source-tag joins the one-fetch-per-call cfg chain; effort_ladder
delegates its capability merge to resolve_capabilities; stale
_maybe_synth_reasoning_block pointers fixed in the providers package.

Wire goldens regenerated: the only drift is the hidden
"temperature": 0.5 pin vanishing from unset-temperature requests.
2026-07-13 08:48:27 -07:00
Patrick Buckley f4701bf0f9 test(golden): re-baseline Google wire payloads for the effort knob
The Gemini effort fix (ee9e9c1f) adds a flat reasoning_effort to every
Google chat-completions request at the session default knob — the
golden fixtures now carry it. Only the eight google__* goldens change
(one added key each); other providers' goldens are untouched — the
UPDATE_WIRE_GOLDENS pass also wanted to rewrite twelve passing goldens
with escape-format-only churn (raw em-dash vs \u2014), reverted to
keep the diff semantic.
2026-07-04 20:54:20 -07:00
Patrick Buckley c9311e32e4 test(providers): freeze per-provider wire payloads as a refactor baseline
Captures the exact request kwargs each provider hands to its SDK seam (Anthropic
messages.stream, OpenAI chat/responses create, Google OpenAI-compat) for a
representative set of trajectories, asserted against committed goldens. This is the
behavior-equivalence net the canonical-trajectory wire-shape refactor is proven
against. Regenerate the baseline with UPDATE_WIRE_GOLDENS=1.
2026-06-04 11:03:13 -07:00