mirror of
https://github.com/turnstonelabs/turnstone.git
synced 2026-08-12 23:12:23 -06:00
main
23 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
29f1f34cf3 |
feat(helm): add node scheduling properties (#977)
Signed-off-by: Dennis Witt <dennis@derwitt.de> |
||
|
|
9334cf0cef |
fix(helm): make the bundled-PostgreSQL default installable (#949)
* fix(helm): render the chart Secret for every inline credential Setting llm.existingSecret suppressed the chart's whole Secret, not just the LLM API key it replaces. POSTGRES_PASSWORD and TURNSTONE_JWT_SECRET went unrendered with it while server, console and the migrate Job went on referencing them, so every pod stalled in CreateContainerConfigError. Supplying an LLM Secret is a supported, documented configuration, and it took the install down on both the bundled and external database paths. turnstone.db.secretName compounded it by falling back to turnstone.llm.secretName, pointing the password lookup at the operator's LLM Secret — which has no reason to carry a database password. Both now derive from one predicate. turnstone.db.inlinePassword returns the password when the chart stores it itself and empty when an operator supplies it, so secret.yaml renders on exactly the condition under which turnstone.db.secretName resolves to <fullname>-secrets. The two cannot disagree about where the password lives, which is what the earlier llm.secretName fallback was working around. Each key keeps its own condition, so an existingSecret still suppresses the value it replaces and nothing else. Verified by rendering nine values permutations against both this and the previous templates and diffing every secretKeyRef against the Secrets each tree creates: three permutations fixed, six byte-identical, none regressed. helm lint passes on all nine. The bundled-PostgreSQL default is unaffected and still broken: the subchart generates its password into <fullname>-postgresql, which the chart never reads. It is separately blocked by the migrate hook running before the database exists, so it needs the design decision called for in #932 rather than a secret-name change. * fix(helm): default the inline password so an unset key cannot become one turnstone.db.inlinePassword is reached through include, which captures rendered text rather than a value. A key that is unset rather than empty — "password:" with nothing after it, or --set database.external.password=null — renders as the literal "<no value>", and a ten-character string is truthy, so it satisfied the gate in templates/secret.yaml and landed base64-encoded in POSTGRES_PASSWORD. Workloads then authenticated with the string "<no value>". Reaching the values through default "" keeps unset and empty equivalent, which is what the previous templates got for free by testing the value directly instead of the rendered text. Introduced by the commit before this one; caught in review. The two null spellings are now permanent cases in the render matrix. Across eleven permutations, three are fixed relative to main, eight are byte-identical, none regress, and the inline password still round-trips byte-exact. helm lint passes on all eleven. * docs(helm): narrow the inlinePassword guarantee to what it holds The comment claimed secret.yaml and turnstone.db.secretName cannot disagree about where the password lives. That holds wherever the chart or the operator supplies the password, but not where the bundled subchart generates its own — that lands in the subchart's Secret, which neither helper reads. State the two guarantees that do hold instead. * fix(helm): make the bundled-PostgreSQL default installable The default values have never produced a working install. Two faults, and the first is why the second could not be fixed on its own. The migrate Job ran as a pre-install hook, and Helm creates ordinary resources only once hooks have finished. On a first install that means none of what the migration needs exists yet: not the ConfigMap, not the Secret, and — because the subchart is an ordinary resource — not the database either. #932 worked around the first two by dropping the Job's ServiceAccount reference and inlining its environment, but nothing can work around the third: no reference to the subchart's Secret, however derived, is readable by a hook that runs before the subchart exists. So the Job moves to post-install, and to pre-upgrade rather than post-upgrade: on an upgrade everything is already running, and migrations belong before the new code rolls out rather than after. Helm does not wait for readiness before post-install hooks, so the Job's own retry is what waits for a cold database, and backoffLimit rises to cover an image pull and cluster initialisation. That in turn unwinds the workarounds. The Job takes the chart's ServiceAccount back, and templates/secret.yaml drops the hook annotations it was given so the pre-install Job could read it — those made it a hook resource, untracked by the release, so the credentials survived helm uninstall and were skipped by helm rollback. With ordering fixed the password resolves properly. When the subchart generates its own, turnstone.db.secretName now points at the subchart's Secret instead of at <fullname>-secrets, which never carried the key. The naming is mirrored rather than delegated, since the subchart's helpers expect a context this chart cannot hand them, and it is derived from the release name: a fullnameOverride here renames this chart's resources and leaves the subchart's alone, so "<fullname>-postgresql" would name a Secret that does not exist. Verified across fifteen values permutations against origin/main: nine fixed, six byte-identical, none regressed, helm lint clean on all fifteen. The permutations cover both fullnameOverride spellings, a subchart existingSecret with a renamed key, and the superuser key rule. An external database with no password and no existingSecret is unchanged and still fails at pod start. Passwordless authentication is not something the chart models — the URL always references a password — so that stays as it was rather than becoming a template-time error. |
||
|
|
989f51edc5 |
fix(helm): render the chart Secret for every inline credential (#948)
* fix(helm): render the chart Secret for every inline credential Setting llm.existingSecret suppressed the chart's whole Secret, not just the LLM API key it replaces. POSTGRES_PASSWORD and TURNSTONE_JWT_SECRET went unrendered with it while server, console and the migrate Job went on referencing them, so every pod stalled in CreateContainerConfigError. Supplying an LLM Secret is a supported, documented configuration, and it took the install down on both the bundled and external database paths. turnstone.db.secretName compounded it by falling back to turnstone.llm.secretName, pointing the password lookup at the operator's LLM Secret — which has no reason to carry a database password. Both now derive from one predicate. turnstone.db.inlinePassword returns the password when the chart stores it itself and empty when an operator supplies it, so secret.yaml renders on exactly the condition under which turnstone.db.secretName resolves to <fullname>-secrets. The two cannot disagree about where the password lives, which is what the earlier llm.secretName fallback was working around. Each key keeps its own condition, so an existingSecret still suppresses the value it replaces and nothing else. Verified by rendering nine values permutations against both this and the previous templates and diffing every secretKeyRef against the Secrets each tree creates: three permutations fixed, six byte-identical, none regressed. helm lint passes on all nine. The bundled-PostgreSQL default is unaffected and still broken: the subchart generates its password into <fullname>-postgresql, which the chart never reads. It is separately blocked by the migrate hook running before the database exists, so it needs the design decision called for in #932 rather than a secret-name change. * fix(helm): default the inline password so an unset key cannot become one turnstone.db.inlinePassword is reached through include, which captures rendered text rather than a value. A key that is unset rather than empty — "password:" with nothing after it, or --set database.external.password=null — renders as the literal "<no value>", and a ten-character string is truthy, so it satisfied the gate in templates/secret.yaml and landed base64-encoded in POSTGRES_PASSWORD. Workloads then authenticated with the string "<no value>". Reaching the values through default "" keeps unset and empty equivalent, which is what the previous templates got for free by testing the value directly instead of the rendered text. Introduced by the commit before this one; caught in review. The two null spellings are now permanent cases in the render matrix. Across eleven permutations, three are fixed relative to main, eight are byte-identical, none regress, and the inline password still round-trips byte-exact. helm lint passes on all eleven. * docs(helm): narrow the inlinePassword guarantee to what it holds The comment claimed secret.yaml and turnstone.db.secretName cannot disagree about where the password lives. That holds wherever the chart or the operator supplies the password, but not where the bundled subchart generates its own — that lands in the subchart's Secret, which neither helper reads. State the two guarantees that do hold instead. |
||
|
|
73fb84b459 |
fix(helm): make the Kubernetes chart installable and multi-node capable (#932)
* fix(helm): repair install-blocking template bugs
The chart could not complete `helm install` in any cluster. Three
independent faults, each hit in sequence on a clean namespace:
1. The console Deployment never set TURNSTONE_DB_URL. The console
requires it (console/server.py exits with "Storage backend is
required for the console") so the pod could never start. Only the
server Deployment defined it.
2. The migrate Job is a pre-install hook but referenced the chart's
ServiceAccount. Helm creates ordinary resources only after hooks
complete, so the Job could never be scheduled:
Error creating: pods "turnstone-migrate-" is forbidden: error
looking up service account <ns>/turnstone: serviceaccount
"turnstone" not found
The migration talks to PostgreSQL and never to the Kubernetes API,
so it now runs under the namespace default ServiceAccount.
3. The same Job took its config via `envFrom` on the chart's ConfigMap
and Secret -- also ordinary resources -- so once (2) was fixed it
failed with:
Error: configmap "turnstone-config" not found
The Job is now self-contained. Where it still needs the chart's own
Secret for POSTGRES_PASSWORD, that Secret carries matching
pre-install/pre-upgrade hook annotations at a lower weight (-3 against
the Job's -1) so it exists by the time the hook runs.
Also wires up two values that were documented but referenced by no
template: database.external.existingSecret and database.external.sslmode.
An external database frequently keeps its password in a secret the chart
does not own (CloudNativePG, External Secrets, ...), where the key is
rarely named POSTGRES_PASSWORD, so existingSecretPasswordKey is added
alongside. sslmode is appended to the URL only on the external path.
The shared turnstone.db.env helper renders every connection value inline
rather than relying on envFrom expansion, which is what lets the hook
stand alone; the server, console and Job now cannot drift apart. Its
secret-name fallback resolves through turnstone.llm.secretName rather
than hardcoding "<fullname>-secrets", because templates/secret.yaml is
skipped entirely when llm.existingSecret is set -- hardcoding it would
point every workload at a Secret that is never created.
Verified against an external CloudNativePG cluster: `helm install`
completes, the migration creates all 45 tables, and both workloads reach
PostgreSQL over TLS. `helm lint` passes, and every referenced Secret is
either chart-created or operator-supplied, across the bundled,
bundled+llm.existingSecret, external+inline-password and
external+existingSecret paths.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(helm): advertise per-pod URLs so multi-node routing works
Neither workload advertised an address peers could reach, so the console
could not talk to server nodes at all and server.replicas > 1 was
unusable.
Server nodes register in the `services` table and the console routes to
them with rendezvous (HRW) hashing: route(ws_id) picks exactly one node
and proxies to that node's advertised URL. The chart set nothing, so a
node fell back to gethostname() -- the pod name -- which nothing in the
cluster can resolve, and the console's SSE collector could never attach.
The fix cannot be the Service DNS name: that load-balances across every
replica, so traffic the router computed for node A lands on an arbitrary
pod. With three replicas that produces a steady stream of 404s through
the router's retry path. Each pod now advertises its own pod IP via the
downward API, which is unique, routable in-cluster on any CNI, and
re-registered on every start.
The console is the opposite case -- one logical endpoint behind its
Service -- so it advertises the Service DNS name via TURNSTONE_CONSOLE_URL.
That name stops at ".svc" rather than assuming a "cluster.local" DNS
domain, which is configurable per cluster.
Verified at server.replicas=3: all three nodes register distinct
addresses, and six workstreams created through
/v1/api/route/workstreams/new distribute across the ring and complete
real inference turns.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(helm): use Recreate for the single-replica console
Workaround for a service-registry race, kept as its own commit so it can
be dropped if the underlying bug is fixed in the application instead.
The console registers itself under the fixed service_id "console" and
deregisters on shutdown. Under RollingUpdate the incoming pod registers
first and the outgoing pod's deregister then deletes that row. The
console's heartbeat only updates last_heartbeat -- heartbeat_service()
returns False when the row is missing and the caller discards it -- so
the registration is never recreated and the console stays invisible in
the registry for the life of the process.
Recreate orders shutdown strictly before startup. It is gated on
console.replicas == 1, since Recreate is meaningless above that and the
fixed service_id makes multiple console replicas overwrite each other
regardless.
The better fix is arguably in the application: have heartbeat_service()
re-register when its row has gone, which would make this unnecessary.
Happy to drop this commit in favour of that.
Note for existing deployments: switching strategy on a live Deployment
fails with `spec.strategy.rollingUpdate: Forbidden: may not be specified
when strategy type is 'Recreate'` because the stored object still
carries the defaulted rollingUpdate block. It needs a one-off
`kubectl patch` to remove that field. Fresh installs are unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8240d2c00d | chore(deps): update helm release postgresql to ~18.8.0 | ||
|
|
04b3a3abe4 |
feat(deploy): systemd units for a bare-metal turnstone-server node
Hardened service + slice + node-identity drop-in template + a README for running a turnstone-server outside Docker that joins the compose cluster — the production-shaped counterpart to the one-liner in docs/docker.md. Secrets stay in config.toml; per-host identity + cluster URLs go in the drop-in. The README notes the cross-host mTLS caveat (turnstonelabs/lacme#22). |
||
|
|
fb4c6eefe0 | chore(deps): update helm release postgresql to ~18.7.0 | ||
|
|
1728a4c0af |
feat(web-search): replace Tavily/DuckDuckGo backends with self-hosted SearxNG
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single self-hosted SearxNG service bundled into the docker-compose stacks. Core: - New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to (backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard unchanged. _resolve_search_client follows storage -> toml -> env -> default precedence (explicit "" disables, via ConfigStore.stored_keys()). - Drop the Tavily-era topic=finance (no SearxNG category); topic is now general/news. Settings/config: - Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key. - Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines, with get_searxng_url/get_searxng_engines. Compose + bundled config: - Internal-only searxng service (no published API port, :ro config, /healthz healthcheck, persistent searxng-cache volume) in both stacks; bundle turnstone/deploy/searxng/settings.yml (JSON output on, limiter off). - Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in). - bootstrap extractor + wheel packaging updated. Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the lxml/h2/brotli transitives). Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG; docs/docker.md carries the AGPL-3.0 §13 operator note. BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg"; tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance. Closes #545 |
||
|
|
631f1b0021 |
feat(compose): cluster-by-default Caddy-fronted stack with bare-metal join
`docker compose up` from a clone builds one image and brings up the whole stack — PostgreSQL, console, Caddy, channel, and 10 server nodes — sharing one Postgres so the console discovers every node. The dashboard is reachable only through Caddy (HTTP/2 avoids the browser's 6-connection cap on the dashboard's SSE streams); the console's plain-HTTP port is no longer published. Postgres binds 127.0.0.1 so a bare-metal turnstone-server can join the cluster — the bare-metal overlay is folded in and removed. Insecure dev defaults keep it zero-config; the bundled production stack mirrors the shape but pulls ghcr images and requires real secrets. Move the Caddyfile under turnstone/deploy so it ships in the wheel; update docs, QUICKSTART, and .env.example to match. |
||
|
|
d820168f3f |
fix(tls): repair cluster mTLS — cert identity, renewal scoping, hot-reload
Enabling mTLS broke the cluster in three layered ways: - Service certs were keyed on socket.gethostname() (the container ID) and never carried the advertised service name as a SAN, so every collector and routing-proxy handshake failed the hostname check. build_cert_hostnames() now puts the advertised host first: it becomes the cert's primary domain (hence a SAN) and a stable store key that survives container recreation. - lacme's RenewalManager renews everything in the store; with the store shared cluster-wide, every node renewed every other node's (and every dead container's) cert — an N×M renewal storm. _SingleDomainStore scopes each node's sweep to its own cert, and the console adds a periodic GC for the certs of long-departed nodes. - uvicorn loads its cert once at boot and never reloads, so renewed certs never reached the listener and the served cert expired mid-process. swap_context_cert() hot-swaps renewed material into the live SSL context (server listener and console client context) via load_cert_chain. Observability and browser access: - The collector logged connection/TLS failures at DEBUG, so a persistent mTLS-verify failure was invisible. It now logs the first failure per node (reachable->unreachable) at WARNING and stays at DEBUG on retries. - The console serves plain HTTP (it is the ACME bootstrap endpoint) and no longer rewrites its advertised URL to https://. Browser->console TLS is terminated by a reverse proxy: the cluster profile gains a caddy service (browser h2/HTTPS -> caddy -> console h1.1/HTTP) plus browser-TLS docs. Tests: tests/test_tls_san_renewal.py, tests/test_collector_reachability.py. |
||
|
|
bf36461187 |
chore(deps): update helm release postgresql to ~18.6.0 (#405)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
217688547e |
fix: expose channel gateway port for bare-metal deploys
The channel gateway registers with its Docker-internal hostname (e.g. http://channel:8091) which is unreachable from a host-side server. Publish port 8091 and set TURNSTONE_CHANNEL_ADVERTISE_URL to localhost so the server can reach it for schedule notifications. |
||
|
|
62d2a0fe6a |
fix: remove non-auth support from bootstrap wizard (#274)
* fix: remove non-auth support from bootstrap wizard Auth is now mandatory for all deployments. Remove the TURNSTONE_AUTH_ENABLED toggle and make JWT_SECRET and AUTH_TOKEN required in the wizard's system prompt. * fix: remove auth disable support from runtime and infra Remove AuthConfig.enabled field — auth is always on. Drop TURNSTONE_AUTH_ENABLED env var, config toggle, and the check_request bypass. Update compose.yaml, Helm chart, Terraform, docs, and tests to match. * feat: deprecate config tokens, require JWT secret, prefer JWT auth Phase 1 of config-token removal: - load_jwt_secret() now exits with error if no secret is configured (was: silently auto-generated ephemeral secret) - _authenticate_token() logs deprecation warning on config token use - CLI /cluster commands use ServiceTokenManager when JWT secret is set - turnstone-admin tls-list uses ServiceTokenManager when JWT secret is set - Update bootstrap wizard, docker.md, security.md to mark TURNSTONE_AUTH_TOKEN as deprecated and JWT_SECRET as required - Console test fixtures use auth token + headers (auth always enforced) * feat: add service scope for inter-service JWT auth Add "service" to VALID_SCOPES and SCOPE_HIERARCHY. Service tokens bypass require_permission() RBAC checks, replacing the old empty-user-id bypass that config tokens relied on. All ServiceTokenManager instances that need admin access now include "service" in their scopes (console proxy, channel gateway, CLI, admin CLI). Read-only services (collector, notification) unchanged. * feat: phase 2 config token deprecation - SDK doc examples now show API tokens (ts_) instead of config tokens - Remove _get_config_token() from admin CLI (dead code) - Block config token exchange in handle_auth_login — only password and API token login allowed - Update login tests to use password-based auth instead of config token exchange * feat: phase 3 — remove config tokens entirely Complete removal of config-file token authentication: - Delete AuthConfig.tokens, check(), _ROLE_TO_SCOPES, hmac dispatch branch, and config token loading from load_auth_config() - Remove auth_config parameter from _authenticate_token() and check_request() — callers updated throughout - Remove TURNSTONE_AUTH_TOKEN from compose.yaml, Helm charts, Terraform, turnstone.example.toml - Remove --auth-token CLI flags from turnstone, turnstone-admin, and turnstone-console - Simplify console main() — always use ServiceTokenManager (no fallback to static tokens) - Delete config-token-specific tests, rewrite check_request and integration tests to use JWT auth with proper audience claims - Remove all config token references from docs (security.md, docker.md, sdk.md, console.md, architecture.md, bootstrap prompt) * fix: address code review findings - Fix 33 broken tests: add JWT auth to test_api_versioning, test_console_routing_proxy, test_tls_admin, test_tls_manager, test_server_live (jwt_secret + audience-scoped auth headers) - Add TestRequirePermissionServiceScope: 4 tests covering the service scope RBAC bypass path - Remove stale comments referencing config tokens in auth.py and console/server.py - Remove dead proxy_auth_token parameter from console create_app() and static token fallback in _proxy_auth_headers() - Remove TURNSTONE_AUTH_TOKEN from env.py scrub list * fix: address Copilot review — JWT audience, compose require secret - CLI /cluster: add audience=JWT_AUD_CONSOLE to ServiceTokenManager (console validates audience, JWTs without it were rejected) - Admin CLI tls-list: same audience fix - compose.yaml: TURNSTONE_JWT_SECRET now uses :? to fail fast if unset - SDK console: fix default port from 8081 to 8090 * test: add auth enforcement tests for TLS admin endpoints 5 new tests: unauthenticated requests return 401 (list, renew, delete), read-only-scoped requests return 403 (renew, delete). Closes the TLS auth enforcement test gap noted in PROGRESS.md. * fix: address remaining Copilot review feedback - Fix token_source="config" → "test" in TLS test fixtures - Fix AuthResult.token_source docstring to include service origins - Require TURNSTONE_JWT_SECRET in cluster compose profile (:?) - Helm: add auth.jwtSecret + auth.existingSecret values, wire TURNSTONE_JWT_SECRET into secret.yaml and both deployments - Terraform: replace auth_token with jwt_secret variable + secret, remove orphaned auth_token resources and IAM reference - Remove [[auth.tokens]] from security.md config example * fix: address full code review — 10 findings Critical: - Terraform: replace concat(common_env, auth_env) with common_env (auth_env local was removed but still referenced) - Channel gateway: remove hmac static token auth from _check_auth(), use JWT-only validation. Remove --auth-token CLI arg from channel - Rebalancer: add token_manager support so migration requests carry JWT auth (was sending unauthenticated POST to /internal/migrate) Major: - Guard _permissions_to_scopes() against "service" privilege escalation from DB role permissions - Remove dead AuthConfig class, load_auth_config(), and all auth_config parameters from create_app() signatures - Helm: inject JWT secret for both inline and existingSecret paths Minor: - Remove dead auth_token param from ClusterCollector - Remove empty TestLoadAuthConfig class - Short JWT secret now exits instead of warning - Compose: add generation command comment above JWT_SECRET - Clean stale config token references from 6 doc files - Clean stale AUTH_TOKEN reference from bootstrap wizard prompt * fix: remove remaining stale config token references from docs - channels.md: remove --auth-token from options table - oidc.md: remove "config-file tokens still work" claim - security.md: remove config token section, fix JWT secret docs (now required/exits, no ephemeral fallback), remove hmac from ASCII diagram, remove --auth-token reference |
||
|
|
a315cabe71 |
chore: polish — remove dead code, update diagrams and docs
Remove stale Redis/Bridge/MQ references found via vulture scan and manual grep: - bot.py docstring: remove Redis MQ reference - server.py trusted_sources: remove "bridge" - tls.py docstring: remove "bridge" from service list Delete 4 obsolete diagram pairs (puml + png): - 06-mq-protocol, 07-message-routing, 08-redis-key-schema, 10-simulator-architecture Update 7 diagrams to reflect direct HTTP architecture: - system-context, package-structure, workstream-states, console-data-flow, deployment, channel-architecture, settings-architecture Redraw architecture-overview.svg: Console router replaces Redis MQ, direct SSE data plane, hash ring routing. |
||
|
|
62ff3217d0 |
fix: TLS Docker end-to-end testing fixes (#185)
* fix: TLS Docker end-to-end testing fixes Fixes discovered during Docker Compose TLS integration testing: - Dockerfile: use --extra all (prevents missing optional deps) - lacme 1.0.3: fixes CACertificateIssued event logging crash - chmod PermissionError: guard for Docker volume mounts - socket import: moved to top of main() (was inside TLS conditional, caused NameError in _default_node_id) - redis.SSLConnection: ConnectionPool needs explicit connection_class, not ssl=True (which only works on Redis() directly) - Empty redis password: pass None instead of "" to avoid AUTH error - TURNSTONE_CONSOLE_URL: env var for Docker service discovery (0.0.0.0 bind address isn't reachable from other containers) - HTTP01Handler: ACME client needs a challenge handler even when server auto-approves - Docker overlay: tls-init as root with chmod, Redis conditional password, console Redis TLS flags, TURNSTONE_CONSOLE_URL * feat: full mTLS end-to-end with lacme 1.0.4 Completes the mTLS chain across all services: lacme 1.0.4: - Dual EKU certs (serverAuth + clientAuth) — fixes mTLS rejection - Configurable CA name (name="turnstone") — consistent store key Bootstrap CA import: - Console imports bootstrap CA from /certs volume on first boot - Single trust root: bootstrap CA → console → all service certs Bridge mTLS: - TLSClient init when TURNSTONE_TLS_ENABLED set - Auto-upgrades server URL from http:// to https:// - SSLContext passed to all 3 httpx clients via verify= Console collector mTLS: - upgrade_tls() method replaces httpx client with mTLS context - Called in lifespan after cert issuance alongside proxy upgrade - Fixes "Failed to poll node" when server serves HTTPS Docker overlay: - TURNSTONE_TLS_SANS on all services (Docker service names as SANs) - TURNSTONE_TLS_ENABLED on bridge - Channel service with Redis TLS flags - TURNSTONE_CONSOLE_URL for service discovery - Server healthcheck disabled (mTLS healthcheck deferred) - Redis conditional password from env Verified end-to-end: bootstrap → console CA → server HTTPS → bridge mTLS → Redis TLS → channel Redis TLS → console collector polls server over mTLS → workstream creation works through bridge * fix: lint + copilot feedback on TLS Docker e2e - SIM105: contextlib.suppress(PermissionError) for chmod - F401: remove unused get_storage import in bridge - Redis healthcheck: pass password when REDIS_PASSWORD is set * fix: sort imports in admin.py and bridge.py * fix: tls-init key permissions, healthcheck env, collector race - tls-init: add set -e, chown to turnstone:turnstone with restrictive perms (keys 0600, certs 0640, dirs 0750) instead of world-readable - Redis healthcheck: use container runtime $$REDIS_PASSWORD instead of Compose-time interpolation for consistency with --requirepass block - collector upgrade_tls(): don't close old httpx client while concurrent poll threads may still be using it — let GC handle cleanup |
||
|
|
b086390558 |
fix: TLS deferred work — wiring, security, tests, Docker overlay (#184)
* fix: TLS deferred work — wiring, security, tests, Docker overlay
Security fixes:
- PostgreSQL SSL: validate sslmode against known values, urlencode
all params to prevent URL injection
- ConfigStore env seeding: removed redundant type coercion, delegate
to validate_value() which handles all coercion correctly
Functional wiring:
- Database SSL: init_storage() passes SSL params to PostgreSQL URL
- Server: env var fallbacks for DB SSL (TURNSTONE_DB_SSLMODE etc.)
- Proxy mTLS: re-create proxy clients after TLS cert issuance
- Channel gateway: --ssl-certfile/keyfile/ca-certs CLI args, HTTPS
advertise URL when SSL configured
- ConfigStore env seeding: TURNSTONE_{SECTION}_{KEY} seeds on first boot
- Console deregistration on shutdown (with debug logging)
Specs, tests, Docker:
- OpenAPI: 5 TLS admin endpoints in console_spec.py
- Auth enforcement test (401 without auth)
- SDK ValueError test (mismatched cert/key)
- Docker overlay: TURNSTONE_TLS_ENABLED, bridge --redis-tls, Redis
healthcheck with client cert
- Removed stale type:ignore comments (lacme 1.0.2 type stubs)
* review: address copilot feedback on TLS deferred work
- ConfigStore env seeding: use config_store.set() instead of
storage.set_system_setting() (correct API, updates cache)
- Remove unused defn variable (iterate SETTINGS keys only)
- Fix structlog call-arg error (positional args, not kwargs)
- Channel gateway: validate cert+key provided together
- Restore type:ignore[no-any-return] for CI mypy (lacme 1.0.2
type stubs not in CI's mypy overrides yet)
* fix: rename _VALID_SSLMODES to lowercase (N806)
|
||
|
|
d08a57dfc2 |
feat: SDK TLS support, Docker Compose overlay, TLS docs (#183)
Python SDK: - ca_cert, client_cert, client_key on all 4 client classes - ValueError if only one of client_cert/client_key provided - Passed to httpx verify=/cert= TypeScript SDK: - TlsOptions type exported (zero runtime code) - Fix picomatch vulnerability (npm audit fix) Docker Compose: - deploy/docker-compose.tls.yml overlay with tls-init bootstrap - Notes it's an overlay requiring a base compose file Documentation: - docs/tls.md: architecture, config, CLI, SDK examples, troubleshooting - Fixed package name (@turnstone/sdk), Node.js 18+ note |
||
|
|
10165bb8a1 |
feat: add OpenShell sandbox policy for turnstone-server (#128)
* feat: add OpenShell sandbox policy for turnstone-server Curated policy for running turnstone-server inside an OpenShell sandbox with kernel-enforced security boundaries (Landlock, netns, seccomp). - Filesystem: workdir read-write, /usr+/etc read-only, /tmp+/dev/null read-write, Landlock best_effort compatibility - Network: default-deny with allowlisted LLM APIs (OpenAI, Anthropic), Tavily, skills.sh, GitHub (read-only L7), MCP registry (read-only L7), Redis localhost, package registries, curated web_fetch domains - Git: L7-enforced read-only (info/refs + git-upload-pack only) - Process: privilege drop to sandbox:sandbox - Inference routing template for credential isolation (real API keys never enter the sandbox, resolved at proxy layer) * fix: address PR #128 review feedback + add integration guide Review fixes: - Use python3 (not python) in usage examples to match binary allowlist - Fix network_policy → network_policies in comment - Remove /usr/bin/git from github_api (git uses github.com not api.github.com; already covered by git_operations policy) - Remove pip/uv from bash_network_tools (package_registries already covers their PyPI access; no need for StackOverflow/Wikipedia reach) - Restructure routes.yaml so commented blocks are indented under routes: key (uncomment without restructuring YAML) New: docs/openshell.md covering policy customization, inference routing, domain allowlisting, MCP subprocess inheritance, and the dual-layer security model. |
||
|
|
f494553020 |
chore(deps): update helm release redis to v25 (#93)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
c766c81f25 |
chore(deps): update helm release postgresql to v18 (#92)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
387ef06da7 |
chore(deps): update helm release redis to ~20.13.0 (#89)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
e4e2200c33 |
chore(deps): update helm release postgresql to ~16.7.0 (#88)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> |
||
|
|
2b58c127b1 |
Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging (#20)
* Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging Database abstraction: StorageBackend protocol with 21 methods, SQLAlchemy Core schema, SQLite backend (FTS5), PostgreSQL backend (tsvector/ILIKE), Alembic migrations, singleton registry. memory.py reduced to thin facade. Session.py open_db() calls replaced with generic KV methods. [database] config section with env var support. Deployment: Docker Compose production profile with PostgreSQL, Dockerfile with postgres extras and migration entrypoint, Helm chart with bitnami subcharts, Terraform AWS ECS/Fargate module with RDS + ElastiCache + ALB. 39 new storage tests (934 total). mypy strict clean. Docs and diagrams updated. * Address PR #20 review feedback (16 items) - Backends only call create_all() when Alembic migrations are disabled - Helm configmap uses correct TURNSTONE_DB_BACKEND env var; DB URL constructed via env expansion with secret reference instead of ConfigMap - Migration errors fail fast for PostgreSQL (only non-fatal for SQLite) - save_memory/delete_memory wrapped in exception handling like other facade fns - pool_size passed through from config/env to init_storage() in cli + server - Terraform: DB URL moved to Secrets Manager, auth enabled flag set, optional TLS listeners with certificate_arn, Redis transit encryption on - Docker entrypoint no longer suppresses migration output - Diagram fixes: removed StaticPool claim, removed non-existent migration ref - compose.yaml/README: clarified production profile requires DB env vars |