mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-25 21:34:47 -06:00
498f23c19e3b2c001aabc33f34596e2de30724bf
21 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e057c364b8 |
Fix console proxy regressions and add workstream task field (#21) (#21)
* Fix console proxy regressions and add workstream task field (#21) Bug fixes: - Fix collector polling unversioned /api/dashboard (404 after API versioning PR) — nodes showed red/unreachable, no workstreams - Fix SSE proxy dropping all data events — upstream sends \r\n line endings but proxy split on \n\n only; normalize before parsing - Fix workstream state stuck on idle — on_state_change() only broadcasted via SSE but never updated ws.state on the Workstream object; dashboard polling now sees correct attention/running states - Fix deep-link switchTab early return — when ?ws_id matched the only workstream, switchTab bailed (wsId === currentWsId) before establishing SSE connection; inline init instead of delegating - Fix console banner covering dashboard overlay — inject <style> offsetting .dashboard-overlay below the 32px banner Enhancements: - Add turnstone branding to console proxy banner (turnstone │ Console │ node-id) - Add initial_message field to CreateWorkstreamMessage protocol and console "New Workstream" modal (Task textarea, sent as first message) - Refactor SSE proxy to use shared httpx client with 30s read timeout instead of per-request client creation - Increase approval timeout default from 300s to 3600s (1 hour) Updated: Python SDK, TypeScript SDK, OpenAPI specs, MQ client, API schemas, MQ protocol diagram, SDK docs. * Address PR #21 review feedback (4 items) - Log unknown state strings in on_state_change instead of silently swallowing; remove unnecessary KeyError catch - Wrap initial_message POST in _handle_create_ws with error handling so workstream creation success isn't masked by send failure - Strip all \r from SSE chunks instead of replacing \r\n, fixing chunk-boundary split edge case - Add tests for initial_message wiring in directed and pool targeting * Refactor SSE proxy to use httpx-sse aconnect_sse Replace manual SSE chunk buffering/parsing with httpx_sse.aconnect_sse() which handles line endings, event types, and all SSE spec edge cases. Eliminates the \r\n chunk-boundary bug class entirely. Event types are now always forwarded (sse.event defaults to "message" per spec). |
||
|
|
2b58c127b1 |
Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging (#20)
* Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging Database abstraction: StorageBackend protocol with 21 methods, SQLAlchemy Core schema, SQLite backend (FTS5), PostgreSQL backend (tsvector/ILIKE), Alembic migrations, singleton registry. memory.py reduced to thin facade. Session.py open_db() calls replaced with generic KV methods. [database] config section with env var support. Deployment: Docker Compose production profile with PostgreSQL, Dockerfile with postgres extras and migration entrypoint, Helm chart with bitnami subcharts, Terraform AWS ECS/Fargate module with RDS + ElastiCache + ALB. 39 new storage tests (934 total). mypy strict clean. Docs and diagrams updated. * Address PR #20 review feedback (16 items) - Backends only call create_all() when Alembic migrations are disabled - Helm configmap uses correct TURNSTONE_DB_BACKEND env var; DB URL constructed via env expansion with secret reference instead of ConfigMap - Migration errors fail fast for PostgreSQL (only non-fatal for SQLite) - save_memory/delete_memory wrapped in exception handling like other facade fns - pool_size passed through from config/env to init_storage() in cli + server - Terraform: DB URL moved to Secrets Manager, auth enabled flag set, optional TLS listeners with certificate_arn, Redis transit encryption on - Docker entrypoint no longer suppresses migration output - Diagram fixes: removed StaticPool claim, removed non-existent migration ref - compose.yaml/README: clarified production profile requires DB env vars |
||
|
|
5ee539c983 |
Add Python and TypeScript client SDKs for server and console APIs (#19)
* Add Python and TypeScript client SDKs for server and console APIs Python SDK (turnstone/sdk/) with sync + async clients for both server and console APIs. Returns Pydantic models directly, streams SSE events as typed dataclasses. 27 event types with registry-based deserialization. High-level send_and_wait() for request-response patterns. TypeScript SDK (sdk/typescript/) with zero browser dependencies. Uses fetch + ReadableStream for SSE parsing. Discriminated union event types with type guards. Same API surface as Python SDK. 63 Python tests, 21 TypeScript tests (vitest). Comprehensive docs at docs/sdk.md with SDK architecture diagram. * Address PR #19 review feedback + fix lint - Fix consume_task leak in send_and_wait when send() raises (try/finally) - Fix TS sendAndWait: open SSE before send, plumb AbortSignal for timeout - Add signal param to TS streamSSE for cancellation support - Fix SSE parser: join multi-line data: fields with \n per spec, handle CRLF - Fix generate-types.py sys.path (parents[3] not parents[2]) - Document token ignored when httpx_client provided - Document TS timeout units as milliseconds - Fix stale docstring in test_sdk_sse.py - Fix import sorting (ruff I001) |
||
|
|
62a4ceac96 |
Dev/api versioning openapi (#18)
* Add API versioning under /v1/ prefix with OpenAPI 3.1 spec
All API endpoints move to /v1/api/* (clean break, no unversioned
aliases). Non-API routes (/, /health, /metrics, /static, /shared,
/node proxy) stay unversioned.
New turnstone/api/ package:
- Pydantic v2 models for all request/response schemas (server +
console) used for OpenAPI spec generation
- Programmatic OpenAPI 3.1 spec builder with EndpointSpec catalog
- /openapi.json serves machine-readable spec, /docs serves Swagger UI
Route changes:
- Both servers use Mount("/v1", routes=[...API routes...])
- Auth middleware strips /v1/ prefix before path classification
(PUBLIC_PATHS/WRITE_PATHS stay unversioned internally)
- Console proxy handles /node/{id}/v1/api/ upstream forwarding
- Bridge and CLI HTTP clients updated to /v1/api/ paths
- /openapi.json and /docs added to PUBLIC_PATHS and rate limiter
EXEMPT_PATHS
Security fix from review: required_role() now correctly handles
/node/{id}/v1/api/{path} proxy routes (previously the v1 segment
caused write-path detection to fail, allowing read-only token
escalation).
42 new tests (830 total). All frontend JS, docs, and diagrams updated.
* Fix mypy type errors in turnstone/api/ package
- Add generic type params to dict fields in console_schemas.py
- Add return type annotations to docs.py handler factories
- Move type-only imports (BaseModel, Callable, Awaitable) into
TYPE_CHECKING blocks to satisfy TC002/TC003 ruff rules
* Address PR #18 review feedback + fix mypy errors
Review fixes:
- Add pydantic>=2.0 as explicit dependency in pyproject.toml
(was only transitively available via openai/mcp)
- Auto-detect path parameters from {param} segments in OpenAPI
spec builder (fixes missing required path params)
- Use startswith() with concrete prefix for proxy version
detection instead of fragile substring check
- Make Swagger UI base URL configurable via swagger_ui_base_url
parameter for air-gapped deployments
Mypy fixes:
- Add generic type params to dict fields in console_schemas
- Add return type annotations to docs.py handler factories
- Move type-only imports into TYPE_CHECKING blocks
|
||
|
|
29c00c0cdf |
Extract shared frontend design system into turnstone/shared_static/ (#17)
* Extract shared frontend design system into turnstone/shared_static/ The server UI and console UI had ~60% CSS overlap and significant JS duplication. Extract shared assets into a new turnstone/shared_static/ package mounted at /shared/ in both servers: - base.css: design tokens, reset, typography, login/toast/kb overlays, dashboard table, state dots, health bar, scrollbar, reduced motion - auth.js: authFetch, login overlay with focus trap, logout (hooks for page-specific post-login/logout callbacks) - theme.js: dark/light toggle with system preference detection - toast.js: notification queue with configurable timeout - utils.js: escapeHtml, formatTokens, ctxClass, formatUptime, formatCount - kb.js: keyboard shortcuts overlay with configurable content, focus management, and focus restore on dismiss Console proxy updated: JS shim injection moved from proxy_static (app.js prepend) to proxy_index (inline <script> in HTML) so it runs before any external scripts. New /shared/ path rewriting and proxy_shared_static route added. ~1540 lines removed from page-specific files, 775 lines in shared package. 13 new tests (788 total). * Fix /shared/ auth and remove __init__.py from shared_static Address PR #17 review feedback: 1. Add /shared/ to PUBLIC_PREFIXES in auth.py so shared CSS/JS loads before authentication (required for login overlay to render) 2. Remove turnstone/shared_static/__init__.py to prevent exposing Python package internals (__init__.py, __pycache__) via the StaticFiles mount. Not needed for packaging since pyproject.toml uses explicit glob includes. 3 new auth tests for /shared/ public path access. |
||
|
|
b6e0f0fcca |
Add node version tracking and drift detection to console dashboard (#16)
* Add node version tracking and drift detection to console dashboard Surface the version field from each node's /health endpoint in the console dashboard. Collector extracts version into get_overview() (version_drift + versions fields), promotes it to top-level in get_nodes(), and adds get_version_info() for per-node detail. Console /health endpoint includes drift fields. Frontend adds a VER column to the 7-column node table grid, shows per-node version strings, tracks versions per group with "mixed" + yellow drift badge when nodes disagree, and displays a DRIFT warning or single version in the status bar. Column hidden on mobile (<700px). ARIA labels include version info for accessibility. 10 new tests (745 total). Docs and diagram updated. * Fix drift tooltip text: show 'Versions detected' not 'Nodes running' |
||
|
|
206e37e73e |
Fix circuit breaker, rate limiter, and Anthropic web search correctne… (#15)
* Fix circuit breaker, rate limiter, and Anthropic web search correctness (#15) Three tech debt items addressing correctness and security gaps: Circuit breaker HALF_OPEN single-request permit: - Rename should_allow_request property to acquire_request_permit() method to make the side-effecting, non-idempotent nature explicit - Add _half_open_permit flag: exactly one probe request in HALF_OPEN, subsequent callers blocked until probe completes - Explicitly reset permit on all state transitions (record_success, record_failure) for clean state machine invariants - Session uses BaseException catch to ensure record_failure always fires, preventing permanent circuit deadlock on probe crash Rate limiter X-Forwarded-For support: - Add resolve_client_ip() with rightmost-untrusted XFF parsing - Configurable trusted_proxies via --ratelimit-trusted-proxies CLI flag and [ratelimit] trusted_proxies config (comma-separated CIDRs) - IPv4-mapped IPv6 normalization (::ffff:x.x.x.x → IPv4) for dual-stack - Clientless requests (request.client is None) pass through instead of sharing a single "unknown" bucket - Log warning for invalid CIDR entries in trusted_proxies config - Show trusted proxies in startup log when enabled Anthropic web search multi-turn encrypted content: - Capture raw provider content blocks during streaming via _block_to_dict() using model_dump(exclude_none=True) to avoid Anthropic API rejection - Accumulate thinking_delta into raw_blocks (was silently empty on replay) - Store _provider_content on assistant messages, pass through verbatim in _convert_messages() so encrypted_content/encrypted_index survive turns - Persist to SQLite via new provider_data column (auto-migrated) - Add thinking/signature to _block_to_dict fallback attribute list 23 new tests (735 total), ruff + mypy clean. * Fix Copilot PR #15 review issues: provider data, circuit breaker, IP normalization - Persist assistant message when provider_data exists even if text content is empty — prevents losing Anthropic web search encrypted content needed for multi-turn replay (session.py) - Re-raise KeyboardInterrupt/SystemExit immediately after recording failure instead of attempting fallback models (session.py) - Consume HALF_OPEN permit for the transition caller — prevents two concurrent probe requests when only one should be allowed (healthcheck.py) - Normalize IPv4-mapped IPv6 addresses consistently in resolve_client_ip() — prevents duplicate rate-limit buckets for ::ffff:x.x.x.x vs x.x.x.x (ratelimit.py) |
||
|
|
6c5441435b |
Add console workstream creation + server reverse proxy (#14)
* Add console workstream creation + server reverse proxy (#14) Enable the console dashboard to create workstreams and proxy server UIs, so users only need network access to the console port. Workstream creation via MQ: - POST /api/cluster/workstreams/new with three targeting modes: specific node (directed queue), auto (best node by capacity), or general pool (shared queue, any bridge picks up) - Console pushes CreateWorkstreamMessage to Redis; bridge handles the rest (server creation, ownership registration, SSE events) Reverse proxy for server UIs: - /node/{node_id}/ serves the server's HTML with static path rewriting and a console-return banner injected after <body> - JS proxy shim prepended to app.js overrides fetch() and EventSource() to route root-relative URLs through /node/{id}/api/... - SSE streams proxied via httpx.AsyncClient(timeout=None) with per- connection clients for long-lived streams - GET/POST API requests forwarded with body and auth token Security: - Proxy write paths checked against WRITE_PATHS to prevent read-token escalation (read tokens cannot POST /api/send through proxy) - html.escape() on node_id in banner HTML to prevent XSS - String length limits on name/model inputs Frontend: - "+ new" button in header opens creation modal with node dropdown (Auto / General pool / specific nodes with capacity display) - Modal has focus trap, backdrop dismiss, scroll lock, keyboard handling - Workstream rows and node links deep-link via proxy paths - Custom select arrow, Instrument Panel modal styling Documentation: - docs/console.md rewritten with proxy and creation API docs - docs/architecture.md console section updated - PlantUML diagrams 01, 11, 12 updated + PNGs re-rendered - README.md updated 28 new tests (741 total), ruff + mypy clean. * Fix Copilot PR #14 review issues: auth bypass, XSS, proxy robustness - Normalize trailing slashes in required_role() to prevent write-role bypass via /api/send/ or /node/{id}/api/send/ (auth.py) - Validate node_id format in proxy handlers (alphanumeric, dot, dash, underscore only) to prevent injection vectors - Use json.dumps() for JS proxy shim prefix to prevent script injection - URL-quote node_id in HTML attribute contexts (proxy_index, proxy_static) - Check upstream status in _proxy_sse() — emit error event on non-200 instead of keeping a dead SSE connection open - Check upstream status in proxy_index() — propagate non-2xx errors - Forward query string in _proxy_post() (consistency with _proxy_get) - Handle JSON null values in create_workstream() — treat null as empty, reject non-string types with 400 - Fix docs/diagram LPUSH → RPUSH to match actual broker implementation |
||
|
|
f02972c11d |
Add provider-native web search with Tavily fallback (#13)
* Add provider-native web search with Tavily fallback Replace client-side Tavily web search with provider-native implementations: - Anthropic: inject web_search_20250305 server-side tool, handle server_tool_use / web_search_tool_result streaming blocks, emit info_delta for search status display - OpenAI: inject web_search_options for gpt-5-search-api, format url_citation annotations as footnote sources - Local/vLLM: preserve existing Tavily-based web_search tool as fallback Add supports_web_search to ModelCapabilities and info_delta to StreamChunk. Remove end-of-life GPT-4o model entries from capability tables. Update docs, diagrams, and README. 88 provider tests (32 new). * Fix Copilot PR #13 review: capture streaming url_citation annotations Accumulate url_citation annotations during OpenAI streaming and emit formatted citations as a final info_delta chunk after the stream ends. Previously annotations were only captured in non-streaming mode, so search model users in the interactive path never saw citation sources. |
||
|
|
d28879208f |
Add multi-provider LLM adapter with model capability flags (#12)
* Add multi-provider LLM adapter with model capability flags Introduce a provider abstraction layer between ChatSession and LLM SDK clients, enabling native support for Anthropic alongside OpenAI-compatible APIs. Each provider translates at the API boundary while the internal message format remains OpenAI-like throughout session history and persistence. - LLMProvider protocol with StreamChunk/CompletionResult normalized types - OpenAIProvider: GPT-4o, GPT-5.x, O-series capability tables with conditional temperature, reasoning_effort, and token param handling - AnthropicProvider: native streaming, message/tool format conversion, adaptive vs manual thinking modes, effort parameter for 4.6 models - ModelCapabilities per-model flags: temperature support, token param name, thinking mode, effort levels, context window, max output - Smart auto-detect: latest Opus for Anthropic, latest base GPT for OpenAI - --provider CLI flag for both turnstone and turnstone-server - anthropic SDK as optional dependency (pip install turnstone[anthropic]) - 56 new provider tests, 672 total passing - Updated architecture docs and 4 PlantUML diagrams * Fix Copilot PR #12 review: reasoning_effort gating, Anthropic thinking, provider factory - Default reasoning_effort_values to () so unknown/local models don't receive unsupported top-level reasoning_effort param. Models that need it (GPT-5.x, search models) have explicit capability declarations. - Fix Anthropic _reasoning_params: "none" and "" effort now return {} instead of enabling thinking with 4096 budget. - Use create_provider("openai") singleton instead of OpenAIProvider() in ChatSession fallback for consistency with registry path. - Add 8 parameter gating tests: unknown model no reasoning_effort, GPT-5 no temperature, GPT-5.1 conditional temperature, O-series no temperature, Anthropic none/empty/low effort. |
||
|
|
a1f00092f5 |
Migrate HTTP servers from stdlib to Starlette/ASGI + uvicorn (#11)
* Migrate HTTP servers from stdlib to Starlette/ASGI + uvicorn Replace Python stdlib http.server (ThreadedHTTPServer, BaseHTTPRequestHandler) with Starlette ASGI applications served by uvicorn across all three HTTP entry points. SSE endpoints use sse-starlette EventSourceResponse with async generators that bridge sync queue.Queue via run_in_executor(). Bridge SSE parser replaced with httpx-sse EventSource. - turnstone/server.py: Starlette app factory with create_app(), pure ASGI middleware (auth, rate limit, metrics, CORS), async route handlers, lifespan context manager for startup/shutdown. WebUI and ChatSession remain fully synchronous — worker threads unchanged. - turnstone/console/server.py: Same pattern, simpler (no ChatSession). Path params replace manual string slicing for node detail route. - turnstone/mq/bridge.py: _iter_sse_data() uses httpx_sse.EventSource instead of hand-rolled line parser. - Tests: All ThreadedHTTPServer fixtures replaced with starlette.testclient.TestClient via create_app() factories. - Docs: Updated architecture.md, api-reference.md, README.md, and PlantUML diagrams (03, 11) + regenerated PNGs. * Fix Copilot PR #11 review: TestClient cleanup, JSON error handling, SSE timeout - Close TestClient in teardown for TestConsoleAuth and TestConsoleLogin to avoid lifespan/resource leaks - Close TestClient via yield/finally in TestConsoleHTTPEndpoints fixture - Add _read_json() helper for safe JSON body parsing (returns {} on invalid JSON instead of 500, matching old stdlib handler behavior) - Apply same try/except pattern to console auth_login endpoint - Increase SSE queue.get timeout from 1s to 5s to align with sse-starlette ping interval, reducing executor task churn |
||
|
|
c006be25de | Bump version to 0.3.0 | ||
|
|
167d63b385 |
Add operational features: health degradation, rate limiting, workstre… (#10)
* Add operational features: health degradation, rate limiting, workstream eviction Backend health monitor with circuit breaker (CLOSED/OPEN/HALF_OPEN) probes LLM backend periodically; /health returns "degraded" when unreachable. Token-bucket per-IP rate limiter with 429 + Retry-After responses; /health and /metrics exempt. Workstream auto-eviction of oldest idle when at configurable max_workstreams capacity. New modules: healthcheck.py (BackendHealthMonitor, CircuitState), ratelimit.py (TokenBucket, RateLimiter). 5 new Prometheus metrics. Both UIs: health indicator, 429 retry with toast, eviction notifications, node degradation badges (console), circuit state in dashboard footer. Config: [health] and [ratelimit] TOML sections, max_workstreams in [server]. Docs: README, architecture, API reference, PlantUML diagrams updated. 616 tests pass (35 new), mypy clean, ruff clean. * Fix Copilot PR #10 review: version import, capacity check order, validations, docs - Use turnstone.__version__ instead of hard-coded "0.2.1" in /health and /metrics endpoints - Move capacity check/eviction before session creation in WorkstreamManager.create() to avoid wasted work when at capacity - Validate rate > 0 and burst >= 1 in RateLimiter when enabled - Validate max_workstreams >= 1 in WorkstreamManager.__init__ - Parse do_POST path with urlparse for consistent rate limit exemptions and metrics labeling - Fix should_allow_request docstring: HALF_OPEN allows requests through (not just one probe) - Fix /health docstring: degraded when circuit is not CLOSED (includes HALF_OPEN) - Add class="health-ok" to health indicator HTML to prevent visible empty pill before first poll - Update PlantUML: remove stale MAX_WORKSTREAMS constant, fix RateLimiter.check and TokenBucket signatures; regenerate PNG |
||
|
|
2c48f694db |
Add multi-model support with ModelRegistry, fallback routing, and per… (#9)
* Add multi-model support with ModelRegistry, fallback routing, and per-workstream selection Introduces a ModelRegistry that holds named model configurations loaded from [models.*] sections in config.toml. Each workstream can select its model at creation time or switch mid-session via /model <alias>. When the primary model is unreachable, a configurable fallback chain tries alternative models. Sub-agents (plan/task) can optionally use a cheaper model via the agent_model setting. Core changes: - New turnstone/core/model_registry.py: ModelConfig (frozen, api_key redacted from repr), ModelRegistry (thread-safe lazy client creation, resolve, fallback chain), load_model_registry() with backwards-compatible config loading - session.py: registry/model_alias params, /model show+switch command, fallback in _create_stream_with_retry (extracted _try_stream), agent model override in _run_agent - workstream.py: factory signature accepts optional model_alias, create() gains model param - cli.py + server.py: build registry, updated session factories, banner, shutdown - protocol.py: model field on CreateWorkstreamMessage - bridge.py: pass model through workstream creation chain Frontend: - MODEL column added to dashboard tables in both server and console UIs - Responsive: hidden alongside NODE at narrow viewports - ARIA labels include model info, title attributes for truncated text - SSE connected event includes model_alias Documentation: - README: architecture tree, Multi-Model Support section, config keys - docs/architecture.md: module map, Multi-Model Registry subsection - docs/api-reference.md: model field in workstream creation, model_alias in SSE - PlantUML diagrams 02 + 03 updated with ModelRegistry Tests: 43 new tests (576 total), mypy clean, ruff clean. * Fix Copilot PR #9 review: model_alias property, preserve manual tool_truncation - Expose model_alias as a public @property on ChatSession instead of accessing the private _model_alias from server.py and tests - Track _manual_tool_truncation flag so /model switch only recomputes tool_truncation when it was auto-derived, preserving --tool-truncation overrides - Update PlantUML diagram to reflect the public property |
||
|
|
5118808f24 |
Add MCP client support for external tool servers (#8)
* Add MCP client support for external tool servers
MCPClientManager connects to stdio and HTTP MCP servers via a background
asyncio event loop, discovers tools at startup, and converts schemas to
OpenAI function-calling format with mcp__{server}__{tool} prefixing.
- New turnstone/core/mcp_client.py: async-sync bridge, config loader
(TOML [mcp.servers.*] + standard mcpServers JSON), tool discovery
- session.py: mcp_client param, self._tools/_task_tools/_agent_tools,
_prepare_mcp_tool/_exec_mcp_tool, /mcp introspection command
- tools.py: merge_mcp_tools() helper
- cli.py + server.py: --mcp-config arg, client lifecycle, banner info
- pyproject.toml: mcp>=1.6 required dependency, mypy override
- 30 new tests (config, schema conversion, session integration, errors)
- Docs: README MCP section, tools.md MCP reference, architecture.md
MCP subsection, 3 updated PlantUML diagrams + PNGs
* Fix Copilot PR #8 review: hermetic MCP config tests, approval docs wording
- Patch load_config in test_json_file_not_found and test_invalid_json so
a developer's local config.toml doesn't leak into test results
- Clarify MCP approval docs: tools require approval by default, but
--skip-permissions and UI auto-approve override this
|
||
|
|
b6c1c3a676 |
Add console-to-server deep linking via ?ws_id= query parameter (#7)
* Add console-to-server deep linking via ?ws_id= query parameter Server UI parses ?ws_id= on load (both direct init and post-login) and auto-selects the matching workstream instead of showing the dashboard. URL is cleaned from the address bar via history.replaceState after navigation. Defers initial SSE connection to avoid redundant connect when deep-linking switches tabs immediately. Console workstream rows are now clickable — opens the node's server UI in a new tab with ?ws_id= targeting that workstream. Uses URL constructor for safe URL building. External-link indicator (↗) appears on hover. Rows without server_url have role/tabindex removed to avoid broken affordance. currentServerUrl reset on showOverview() to prevent stale fallback across views. Collector injects server_url into workstream dicts in both the poll path and ws_created event path so deep links work immediately. * Fix Copilot PR #7 review: deep-link duplicate history entry, server_url test coverage - Suppress history.pushState in switchTab() during deep-link navigation by setting _historyNavigation=true around both call sites (post-login and direct init). Fixes Back button appearing to do nothing on first press. - Add server_url assertions to poll and ws_created collector tests to prevent regressions of deep-link functionality. |
||
|
|
7fbcb70ec1 |
Add call_id routing for streaming tool output during parallel execution (#6)
* Add call_id routing for streaming tool output during parallel execution
Thread call_id through tool_info, approve_request, and tool_result SSE
events so the browser can route streaming output chunks and final results
to the correct tool div when multiple bash tools run in parallel.
Server: include call_id in serialized approval items and tool_result events.
Protocol: add call_id to on_tool_result signature (session, cli, eval, server)
and ToolResultEvent dataclass; pass through MQ bridge.
Client: set data-call-id on tool divs, match by call_id in appendToolOutputChunk
and appendToolOutput with func_name fallback; extract makeCollapsible
helper; use CSS.escape for querySelector safety; fix replayHistory
\\n typo and missing keyboard accessibility on collapsed output.
Bridge: fix pre-existing bug using "name" instead of "func_name" for
auto-approval matching; include call_id in _build_history for replay.
Also adds on_tool_result calls to write_file and edit_file exec methods.
* Update docs/tools.md
|
||
|
|
5c452a1239 |
Harden session persistence: config storage, interrupted repair, prune… (#4)
* Harden session persistence: config storage, interrupted repair, prune tests Session config persistence: - Add session_config table to SQLite schema for persisting LLM-affecting parameters (temperature, reasoning_effort, max_tokens, instructions, creative_mode) across resume - Add save_session_config() and load_session_config() to memory.py - ChatSession._save_config() called on init and when /instructions, /effort, /creative slash commands change config - resume_session() restores persisted config and rebuilds system messages Interrupted session repair: - load_session_messages() now strips trailing incomplete tool call turns where tool_calls exist but fewer tool results than expected (session was interrupted mid-execution via Ctrl+C or crash) Cleanup: - delete_session() now also removes session_config rows - 14 new tests: interrupted repair (4), config persistence (5), prune_sessions (5) Docs & diagrams: - Document config persistence and interrupted repair in architecture.md - Add _save_config() to ChatSession in 03-core-engine-classes.puml * Fix Copilot PR #4 review: prune config cleanup, /new config persist, resume instructions - prune_sessions() now deletes session_config rows for orphaned/stale sessions - /new command calls _save_config() so config persists immediately - resume_session() uses key presence check for instructions, fixing cross-session leak - Add test_prune_removes_session_config covering both orphan and stale paths |
||
|
|
14a9ff9513 |
Stream bash tool output incrementally via SSE
Replace subprocess.run() with Popen for bash tool execution, streaming stdout line-by-line through a new on_tool_output_chunk callback. Web UI renders chunks incrementally with a pulsing amber border indicator. Core: - Add on_tool_output_chunk(call_id, chunk) to SessionUI protocol - Rewrite _exec_bash() with Popen, process-group kill via start_new_session + os.killpg, background stderr drain thread, threading.Event-based timeout detection - Guard UI callback with contextlib.suppress so errors don't interrupt output collection Server/CLI/eval: - Add tool_output_chunk SSE event type in WebUI - No-op implementations in TerminalUI, BackgroundTerminalUI, SilentUI MQ: - Add ToolOutputChunkEvent to mq/protocol.py and _OUTBOUND_REGISTRY - Handle tool_output_chunk in bridge._handle_ws_event Web UI: - Add appendToolOutputChunk() with call_id-keyed DOM elements, inner auto-scroll, ARIA attributes, and empty chunk guards - Fix appendToolOutput() streaming cleanup using adjacency matching - Make collapsed output keyboard-accessible (tabindex, role, keydown) - Improve stripAnsi() to handle CSI, OSC, and two-byte escapes; use it consistently in replayHistory, addInfoMessage, addErrorMessage - Add .tool-output-stream CSS with soft pulse animation, mobile max-height cap, and consolidated prefers-reduced-motion support Docs & diagrams: - Document tool_output_chunk SSE event in api-reference.md - Update SessionUI protocol (14 methods) in architecture.md - Update Phase 3 execution flow in tools.md - Add on_tool_output_chunk to 03-core-engine-classes.puml - Update 04-conversation-turn.puml, 05-tool-pipeline.puml - Add ToolOutputChunkEvent to 06-mq-protocol.puml - Add to event list in 07-message-routing.puml - Regenerate all 5 affected PNG diagrams |
||
|
|
9be155b97a |
Quality overhaul: code tooling, CI/CD, architecture diagrams, UI rede… (#1)
* Quality overhaul: code tooling, CI/CD, architecture diagrams, UI redesign, and legacy cleanup - Add ruff (lint+format) and mypy (strict) with zero errors across 37 source files - Add GitHub Actions CI (lint, typecheck, test matrix 3.11/3.12/3.13) and PyPI publish workflow - Create 12 PlantUML architecture diagrams with PNG renders covering all subsystems - Refresh README and docs with badges, diagram links, and current descriptions - Refactor test_server_live.py with mock streaming helpers for deterministic CI testing - Update dependencies to current versions (openai>=2.24, httpx>=0.28, redis>=7.2) Console dashboard: - Move state indicators from top cards to fixed bottom status bar with cluster metrics - Replace flat 50-node list with hostname-prefix grouped nodes (expand/collapse, up to 1000) - Apply "Instrument Panel" visual redesign: IBM Plex Mono + Outfit fonts, warm amber accent, LED glow state indicators, deep charcoal surfaces, WCAG AA contrast compliance - Add render cache, stale indicator, active filter highlight, loading states Server web UI: - Apply matching Instrument Panel aesthetic for visual consistency with console - Fix branding (pcode → turnstone), extract inline styles to CSS classes - Rename pcode localStorage keys and history state to turnstone Legacy cleanup: - Remove persona-model-specific --persona flag and /persona slash command - Remove model_identity from chat_template_kwargs (vLLM-specific mechanism) - Refactor plan agent to use standard developer message instead of model_identity - Remove dead code (unused date/has_tools variables, noqa suppressions) * Fix CI typecheck: add mypy overrides for optional sympy/numpy imports The math sandbox optionally imports sympy and numpy at runtime (try/except ImportError). In CI these packages are not installed, so mypy raises import-not-found rather than import-untyped. Add mypy overrides to ignore missing imports for these optional dependencies. * Fix Copilot review findings: ARIA role, status bar cache, and pulse opacity - Change #node-table from role="tree" to role="list" and group elements from role="treeitem" to role="listitem" (proper ARIA semantics) - Include currentView and currentFilter.state in renderStatusBar cache key so active pill highlight updates when switching views - Align pulse animation to 0.35 opacity (already applied in CSS) |
||
|
|
0d6252dd7d | Initial commit — turnstone multi-node AI orchestration platform. |