Commit Graph

3 Commits

Author SHA1 Message Date
Patrick Buckley 3d12798315 feat(audio): voice I/O — speech-to-text + text-to-speech via model roles (#618)
* feat(audio): voice I/O — speech-to-text + text-to-speech via model roles

Browser voice input/output over the OpenAI audio wire protocol, selected
through the existing model-roles system so the same code path serves OpenAI,
vLLM/vLLM-Omni, or any compatible backend — pure registry config, no new
in-process deps. Anthropic has no audio API, so it is capability-gated out of
the audio roles while remaining valid as the agent model.

Backend
- core/audio.py: role resolution + capability gating + transcribe()/synthesize()
  over a registry-resolved client. Typed AudioUnavailableError (503) /
  AudioBackendError (502 — body masked, SDK detail logged). Optional STT prompt.
- Endpoints POST /v1/api/workstreams/{ws_id}/speech-to-text and POST /v1/api/tts,
  registered in v1_routes, write-scoped (direct + proxied), offloaded with
  asyncio.to_thread. Silence -> 422; configured-but-failed backend -> masked 502.
- Model roles: audio.stt_model_alias / audio.tts_model_alias / audio.tts_voice /
  audio.stt_prompt settings; Models -> Roles entries (capability-gated dropdowns,
  "(disabled — voice off)" when unset). /v1/api/models exposes resolved
  stt_default_alias / tts_default_alias + per-model capabilities.
- Capabilities: supports_transcription / supports_speech_synthesis on
  ModelCapabilities; current OpenAI audio lineup (whisper-1, gpt-4o[-mini]-
  transcribe, tts-1[-hd], gpt-4o-mini-tts) registered as known models, with a
  name-inference backstop for local/openai-compatible aliases.

Frontend (interactive UI)
- Mic dictation (record -> transcribe -> fill composer for review) and
  per-message playback, shown only when the role is configured.
- CSS-mask icon set, aria-pressed + live-region announcements, recording timer,
  reduced-motion cue, error-typed toasts + persistent denial, mic disabled while
  busy, code/math stripped before TTS.

Tests: new test_audio.py plus STT/TTS endpoint, settings, openapi, available-
models, and OpenAI-lineup capability coverage. ruff + mypy + node --check clean.

* fix(audio): use const for AUDIO_MODEL_HINTS (var-sweep invariant)
2026-05-30 14:03:39 -07:00
Patrick Buckley 8e32aa09d4 fix(server): advertise registry.default when model.default_alias is foreign
GET /v1/api/models blanked default_alias whenever model.default_alias named
an alias absent from the server's live registry — e.g. when a standalone
turnstone-server shares a ConfigStore with a console whose model.default_alias
points at a console-only / DB alias (or the underlying model id rather than
the alias). The interactive dashboard then showed a bare "Default model"
placeholder even though a new workstream launches on a concrete model.

Mirror session_factory's _effective_default_alias / _effective_routing: fall
back to registry.default (which already incorporates a *valid* model.default_alias
override) when the configured alias is unset or foreign, blanking only if
registry.default is itself unresolvable. The endpoint now reports the model
creation actually uses.
2026-05-29 13:09:18 -07:00
Patrick Buckley 757561db9a feat(ui): backport coordinator selector + pagination UX to interactive dashboard
The interactive dashboard had drifted from the coordinator launcher in two
ways; this backports both for consistency.

Selectors: the Model / Judge Model dropdowns now show the resolved default
model in the placeholder (e.g. "Default — primary (vendor/primary)") instead
of a generic "Default model" / "Default (agent model)". The server's
GET /v1/api/models now returns judge_default_alias (mirroring the console
endpoint), sourced from the judge.model setting. It is intentionally left
blank when judge.model is unset or points at a disabled/removed alias,
because the judge then follows the per-workstream agent model at runtime
(session_factory: judge_config.model or model) — the UI keeps the honest
"Default (agent model)" wording in that case. This also fixes a latent
mislabel: the judge row previously said "agent model" even when an operator
had configured judge.model.

Pagination: Saved Workstreams now paginates at 24/page (Prev · X / Y · Next),
matching Saved Coordinators — page clamp after deletes, hidden on a single
page and in delete mode, Select-All bounded to the visible page. The empty
branch drops out of delete mode (matching the launcher) so the toolbar can't
linger over an empty grid.

The shared .pagination CSS is lifted from console/static/style.css into
shared/cards.css (loaded by both apps) so the two dashboards keep one source
of truth instead of a third copy. The pagination JS render wiring stays
per-app (it binds per-app DOM ids + controller instances) with a
cross-reference comment to its console twin.

Tests: new tests/test_server_available_models.py pins the judge/model
resolution chain (unset / configured / unknown / whitespace / registry-default
fallback).
2026-05-29 12:32:47 -07:00