23 Commits

Author SHA1 Message Date
Dennis Witt 29f1f34cf3 feat(helm): add node scheduling properties (#977)
Signed-off-by: Dennis Witt <dennis@derwitt.de>
2026-08-05 13:52:45 -07:00
Patrick Buckley 9334cf0cef fix(helm): make the bundled-PostgreSQL default installable (#949)
* fix(helm): render the chart Secret for every inline credential

Setting llm.existingSecret suppressed the chart's whole Secret, not just
the LLM API key it replaces. POSTGRES_PASSWORD and TURNSTONE_JWT_SECRET
went unrendered with it while server, console and the migrate Job went on
referencing them, so every pod stalled in CreateContainerConfigError.
Supplying an LLM Secret is a supported, documented configuration, and it
took the install down on both the bundled and external database paths.

turnstone.db.secretName compounded it by falling back to
turnstone.llm.secretName, pointing the password lookup at the operator's
LLM Secret — which has no reason to carry a database password.

Both now derive from one predicate. turnstone.db.inlinePassword returns
the password when the chart stores it itself and empty when an operator
supplies it, so secret.yaml renders on exactly the condition under which
turnstone.db.secretName resolves to <fullname>-secrets. The two cannot
disagree about where the password lives, which is what the earlier
llm.secretName fallback was working around. Each key keeps its own
condition, so an existingSecret still suppresses the value it replaces
and nothing else.

Verified by rendering nine values permutations against both this and the
previous templates and diffing every secretKeyRef against the Secrets
each tree creates: three permutations fixed, six byte-identical, none
regressed. helm lint passes on all nine.

The bundled-PostgreSQL default is unaffected and still broken: the
subchart generates its password into <fullname>-postgresql, which the
chart never reads. It is separately blocked by the migrate hook running
before the database exists, so it needs the design decision called for
in #932 rather than a secret-name change.

* fix(helm): default the inline password so an unset key cannot become one

turnstone.db.inlinePassword is reached through include, which captures
rendered text rather than a value. A key that is unset rather than empty
— "password:" with nothing after it, or --set database.external.password=null
— renders as the literal "<no value>", and a ten-character string is
truthy, so it satisfied the gate in templates/secret.yaml and landed
base64-encoded in POSTGRES_PASSWORD. Workloads then authenticated with
the string "<no value>".

Reaching the values through default "" keeps unset and empty equivalent,
which is what the previous templates got for free by testing the value
directly instead of the rendered text. Introduced by the commit before
this one; caught in review.

The two null spellings are now permanent cases in the render matrix.
Across eleven permutations, three are fixed relative to main, eight are
byte-identical, none regress, and the inline password still round-trips
byte-exact. helm lint passes on all eleven.

* docs(helm): narrow the inlinePassword guarantee to what it holds

The comment claimed secret.yaml and turnstone.db.secretName cannot
disagree about where the password lives. That holds wherever the chart
or the operator supplies the password, but not where the bundled
subchart generates its own — that lands in the subchart's Secret, which
neither helper reads. State the two guarantees that do hold instead.

* fix(helm): make the bundled-PostgreSQL default installable

The default values have never produced a working install. Two faults,
and the first is why the second could not be fixed on its own.

The migrate Job ran as a pre-install hook, and Helm creates ordinary
resources only once hooks have finished. On a first install that means
none of what the migration needs exists yet: not the ConfigMap, not the
Secret, and — because the subchart is an ordinary resource — not the
database either. #932 worked around the first two by dropping the Job's
ServiceAccount reference and inlining its environment, but nothing can
work around the third: no reference to the subchart's Secret, however
derived, is readable by a hook that runs before the subchart exists.

So the Job moves to post-install, and to pre-upgrade rather than
post-upgrade: on an upgrade everything is already running, and
migrations belong before the new code rolls out rather than after. Helm
does not wait for readiness before post-install hooks, so the Job's own
retry is what waits for a cold database, and backoffLimit rises to cover
an image pull and cluster initialisation.

That in turn unwinds the workarounds. The Job takes the chart's
ServiceAccount back, and templates/secret.yaml drops the hook
annotations it was given so the pre-install Job could read it — those
made it a hook resource, untracked by the release, so the credentials
survived helm uninstall and were skipped by helm rollback.

With ordering fixed the password resolves properly. When the subchart
generates its own, turnstone.db.secretName now points at the subchart's
Secret instead of at <fullname>-secrets, which never carried the key.
The naming is mirrored rather than delegated, since the subchart's
helpers expect a context this chart cannot hand them, and it is derived
from the release name: a fullnameOverride here renames this chart's
resources and leaves the subchart's alone, so "<fullname>-postgresql"
would name a Secret that does not exist.

Verified across fifteen values permutations against origin/main: nine
fixed, six byte-identical, none regressed, helm lint clean on all
fifteen. The permutations cover both fullnameOverride spellings, a
subchart existingSecret with a renamed key, and the superuser key rule.

An external database with no password and no existingSecret is unchanged
and still fails at pod start. Passwordless authentication is not
something the chart models — the URL always references a password — so
that stays as it was rather than becoming a template-time error.
2026-08-02 14:52:31 -07:00
Patrick Buckley 989f51edc5 fix(helm): render the chart Secret for every inline credential (#948)
* fix(helm): render the chart Secret for every inline credential

Setting llm.existingSecret suppressed the chart's whole Secret, not just
the LLM API key it replaces. POSTGRES_PASSWORD and TURNSTONE_JWT_SECRET
went unrendered with it while server, console and the migrate Job went on
referencing them, so every pod stalled in CreateContainerConfigError.
Supplying an LLM Secret is a supported, documented configuration, and it
took the install down on both the bundled and external database paths.

turnstone.db.secretName compounded it by falling back to
turnstone.llm.secretName, pointing the password lookup at the operator's
LLM Secret — which has no reason to carry a database password.

Both now derive from one predicate. turnstone.db.inlinePassword returns
the password when the chart stores it itself and empty when an operator
supplies it, so secret.yaml renders on exactly the condition under which
turnstone.db.secretName resolves to <fullname>-secrets. The two cannot
disagree about where the password lives, which is what the earlier
llm.secretName fallback was working around. Each key keeps its own
condition, so an existingSecret still suppresses the value it replaces
and nothing else.

Verified by rendering nine values permutations against both this and the
previous templates and diffing every secretKeyRef against the Secrets
each tree creates: three permutations fixed, six byte-identical, none
regressed. helm lint passes on all nine.

The bundled-PostgreSQL default is unaffected and still broken: the
subchart generates its password into <fullname>-postgresql, which the
chart never reads. It is separately blocked by the migrate hook running
before the database exists, so it needs the design decision called for
in #932 rather than a secret-name change.

* fix(helm): default the inline password so an unset key cannot become one

turnstone.db.inlinePassword is reached through include, which captures
rendered text rather than a value. A key that is unset rather than empty
— "password:" with nothing after it, or --set database.external.password=null
— renders as the literal "<no value>", and a ten-character string is
truthy, so it satisfied the gate in templates/secret.yaml and landed
base64-encoded in POSTGRES_PASSWORD. Workloads then authenticated with
the string "<no value>".

Reaching the values through default "" keeps unset and empty equivalent,
which is what the previous templates got for free by testing the value
directly instead of the rendered text. Introduced by the commit before
this one; caught in review.

The two null spellings are now permanent cases in the render matrix.
Across eleven permutations, three are fixed relative to main, eight are
byte-identical, none regress, and the inline password still round-trips
byte-exact. helm lint passes on all eleven.

* docs(helm): narrow the inlinePassword guarantee to what it holds

The comment claimed secret.yaml and turnstone.db.secretName cannot
disagree about where the password lives. That holds wherever the chart
or the operator supplies the password, but not where the bundled
subchart generates its own — that lands in the subchart's Secret, which
neither helper reads. State the two guarantees that do hold instead.
2026-08-02 14:42:33 -07:00
posixpositive 73fb84b459 fix(helm): make the Kubernetes chart installable and multi-node capable (#932)
* fix(helm): repair install-blocking template bugs

The chart could not complete `helm install` in any cluster. Three
independent faults, each hit in sequence on a clean namespace:

1. The console Deployment never set TURNSTONE_DB_URL. The console
   requires it (console/server.py exits with "Storage backend is
   required for the console") so the pod could never start. Only the
   server Deployment defined it.

2. The migrate Job is a pre-install hook but referenced the chart's
   ServiceAccount. Helm creates ordinary resources only after hooks
   complete, so the Job could never be scheduled:

     Error creating: pods "turnstone-migrate-" is forbidden: error
     looking up service account <ns>/turnstone: serviceaccount
     "turnstone" not found

   The migration talks to PostgreSQL and never to the Kubernetes API,
   so it now runs under the namespace default ServiceAccount.

3. The same Job took its config via `envFrom` on the chart's ConfigMap
   and Secret -- also ordinary resources -- so once (2) was fixed it
   failed with:

     Error: configmap "turnstone-config" not found

   The Job is now self-contained. Where it still needs the chart's own
   Secret for POSTGRES_PASSWORD, that Secret carries matching
   pre-install/pre-upgrade hook annotations at a lower weight (-3 against
   the Job's -1) so it exists by the time the hook runs.

Also wires up two values that were documented but referenced by no
template: database.external.existingSecret and database.external.sslmode.
An external database frequently keeps its password in a secret the chart
does not own (CloudNativePG, External Secrets, ...), where the key is
rarely named POSTGRES_PASSWORD, so existingSecretPasswordKey is added
alongside. sslmode is appended to the URL only on the external path.

The shared turnstone.db.env helper renders every connection value inline
rather than relying on envFrom expansion, which is what lets the hook
stand alone; the server, console and Job now cannot drift apart. Its
secret-name fallback resolves through turnstone.llm.secretName rather
than hardcoding "<fullname>-secrets", because templates/secret.yaml is
skipped entirely when llm.existingSecret is set -- hardcoding it would
point every workload at a Secret that is never created.

Verified against an external CloudNativePG cluster: `helm install`
completes, the migration creates all 45 tables, and both workloads reach
PostgreSQL over TLS. `helm lint` passes, and every referenced Secret is
either chart-created or operator-supplied, across the bundled,
bundled+llm.existingSecret, external+inline-password and
external+existingSecret paths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(helm): advertise per-pod URLs so multi-node routing works

Neither workload advertised an address peers could reach, so the console
could not talk to server nodes at all and server.replicas > 1 was
unusable.

Server nodes register in the `services` table and the console routes to
them with rendezvous (HRW) hashing: route(ws_id) picks exactly one node
and proxies to that node's advertised URL. The chart set nothing, so a
node fell back to gethostname() -- the pod name -- which nothing in the
cluster can resolve, and the console's SSE collector could never attach.

The fix cannot be the Service DNS name: that load-balances across every
replica, so traffic the router computed for node A lands on an arbitrary
pod. With three replicas that produces a steady stream of 404s through
the router's retry path. Each pod now advertises its own pod IP via the
downward API, which is unique, routable in-cluster on any CNI, and
re-registered on every start.

The console is the opposite case -- one logical endpoint behind its
Service -- so it advertises the Service DNS name via TURNSTONE_CONSOLE_URL.
That name stops at ".svc" rather than assuming a "cluster.local" DNS
domain, which is configurable per cluster.

Verified at server.replicas=3: all three nodes register distinct
addresses, and six workstreams created through
/v1/api/route/workstreams/new distribute across the ring and complete
real inference turns.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(helm): use Recreate for the single-replica console

Workaround for a service-registry race, kept as its own commit so it can
be dropped if the underlying bug is fixed in the application instead.

The console registers itself under the fixed service_id "console" and
deregisters on shutdown. Under RollingUpdate the incoming pod registers
first and the outgoing pod's deregister then deletes that row. The
console's heartbeat only updates last_heartbeat -- heartbeat_service()
returns False when the row is missing and the caller discards it -- so
the registration is never recreated and the console stays invisible in
the registry for the life of the process.

Recreate orders shutdown strictly before startup. It is gated on
console.replicas == 1, since Recreate is meaningless above that and the
fixed service_id makes multiple console replicas overwrite each other
regardless.

The better fix is arguably in the application: have heartbeat_service()
re-register when its row has gone, which would make this unnecessary.
Happy to drop this commit in favour of that.

Note for existing deployments: switching strategy on a live Deployment
fails with `spec.strategy.rollingUpdate: Forbidden: may not be specified
when strategy type is 'Recreate'` because the stored object still
carries the defaulted rollingUpdate block. It needs a one-off
`kubectl patch` to remove that field. Fresh installs are unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 13:42:22 -07:00
renovate[bot] 8240d2c00d chore(deps): update helm release postgresql to ~18.8.0 2026-07-16 10:51:18 -07:00
Patrick Buckley 04b3a3abe4 feat(deploy): systemd units for a bare-metal turnstone-server node
Hardened service + slice + node-identity drop-in template + a README for
running a turnstone-server outside Docker that joins the compose cluster —
the production-shaped counterpart to the one-liner in docs/docker.md. Secrets
stay in config.toml; per-host identity + cluster URLs go in the drop-in. The
README notes the cross-host mTLS caveat (turnstonelabs/lacme#22).
2026-06-15 03:41:24 -07:00
renovate[bot] fb4c6eefe0 chore(deps): update helm release postgresql to ~18.7.0 2026-06-03 18:55:55 -07:00
Patrick Buckley 1728a4c0af feat(web-search): replace Tavily/DuckDuckGo backends with self-hosted SearxNG
Drop the Tavily and DuckDuckGo (ddgs) web_search backends for a single
self-hosted SearxNG service bundled into the docker-compose stacks.

Core:
- New SearXNGClient + _format_searxng; rewrite resolve_web_search_client to
  (backend, searxng_url, searxng_engines, ...). MCP backend + oauth_user guard
  unchanged. _resolve_search_client follows storage -> toml -> env -> default
  precedence (explicit "" disables, via ConfigStore.stored_keys()).
- Drop the Tavily-era topic=finance (no SearxNG category); topic is now
  general/news.

Settings/config:
- Remove tools.tavily_api_key, get_tavily_key, $TAVILY_API_KEY, [api].tavily_key.
- Add tools.searxng_url (default http://searxng:8080) + tools.searxng_engines,
  with get_searxng_url/get_searxng_engines.

Compose + bundled config:
- Internal-only searxng service (no published API port, :ro config, /healthz
  healthcheck, persistent searxng-cache volume) in both stacks; bundle
  turnstone/deploy/searxng/settings.yml (JSON output on, limiter off).
- Caddy serves the SearxNG web UI on :8444 (dev: localhost-only; prod: opt-in).
- bootstrap extractor + wheel packaging updated.

Deps: drop the ddg extra + ddgs mypy override (regenerates uv.lock, removing the
lxml/h2/brotli transitives).

Docs: tools/docker/architecture/openshell + diagrams + config example + CHANGELOG;
docs/docker.md carries the AGPL-3.0 §13 operator note.

BREAKING: tools.web_search_backend no longer accepts "tavily"/"ddg";
tools.tavily_api_key and the ddg extra are removed. Run the bundled SearxNG (ships
in the compose stacks) or set TURNSTONE_SEARXNG_URL to an external instance.

Closes #545
2026-05-31 19:03:16 -07:00
Patrick Buckley 631f1b0021 feat(compose): cluster-by-default Caddy-fronted stack with bare-metal join
`docker compose up` from a clone builds one image and brings up the whole stack
— PostgreSQL, console, Caddy, channel, and 10 server nodes — sharing one
Postgres so the console discovers every node. The dashboard is reachable only
through Caddy (HTTP/2 avoids the browser's 6-connection cap on the dashboard's
SSE streams); the console's plain-HTTP port is no longer published. Postgres
binds 127.0.0.1 so a bare-metal turnstone-server can join the cluster — the
bare-metal overlay is folded in and removed. Insecure dev defaults keep it
zero-config; the bundled production stack mirrors the shape but pulls ghcr
images and requires real secrets.

Move the Caddyfile under turnstone/deploy so it ships in the wheel; update docs,
QUICKSTART, and .env.example to match.
2026-05-31 16:05:21 -07:00
Patrick Buckley d820168f3f fix(tls): repair cluster mTLS — cert identity, renewal scoping, hot-reload
Enabling mTLS broke the cluster in three layered ways:

- Service certs were keyed on socket.gethostname() (the container ID) and
  never carried the advertised service name as a SAN, so every collector and
  routing-proxy handshake failed the hostname check. build_cert_hostnames()
  now puts the advertised host first: it becomes the cert's primary domain
  (hence a SAN) and a stable store key that survives container recreation.

- lacme's RenewalManager renews everything in the store; with the store shared
  cluster-wide, every node renewed every other node's (and every dead
  container's) cert — an N×M renewal storm. _SingleDomainStore scopes each
  node's sweep to its own cert, and the console adds a periodic GC for the
  certs of long-departed nodes.

- uvicorn loads its cert once at boot and never reloads, so renewed certs
  never reached the listener and the served cert expired mid-process.
  swap_context_cert() hot-swaps renewed material into the live SSL context
  (server listener and console client context) via load_cert_chain.

Observability and browser access:

- The collector logged connection/TLS failures at DEBUG, so a persistent
  mTLS-verify failure was invisible. It now logs the first failure per node
  (reachable->unreachable) at WARNING and stays at DEBUG on retries.

- The console serves plain HTTP (it is the ACME bootstrap endpoint) and no
  longer rewrites its advertised URL to https://. Browser->console TLS is
  terminated by a reverse proxy: the cluster profile gains a caddy service
  (browser h2/HTTPS -> caddy -> console h1.1/HTTP) plus browser-TLS docs.

Tests: tests/test_tls_san_renewal.py, tests/test_collector_reachability.py.
2026-05-30 16:27:15 -07:00
renovate[bot] bf36461187 chore(deps): update helm release postgresql to ~18.6.0 (#405)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-04-23 21:15:38 -07:00
Patrick Buckley 217688547e fix: expose channel gateway port for bare-metal deploys
The channel gateway registers with its Docker-internal hostname
(e.g. http://channel:8091) which is unreachable from a host-side
server. Publish port 8091 and set TURNSTONE_CHANNEL_ADVERTISE_URL
to localhost so the server can reach it for schedule notifications.
2026-04-06 10:47:50 -07:00
Patrick Buckley 62d2a0fe6a fix: remove non-auth support from bootstrap wizard (#274)
* fix: remove non-auth support from bootstrap wizard

Auth is now mandatory for all deployments. Remove the
TURNSTONE_AUTH_ENABLED toggle and make JWT_SECRET and AUTH_TOKEN
required in the wizard's system prompt.

* fix: remove auth disable support from runtime and infra

Remove AuthConfig.enabled field — auth is always on. Drop
TURNSTONE_AUTH_ENABLED env var, config toggle, and the
check_request bypass. Update compose.yaml, Helm chart,
Terraform, docs, and tests to match.

* feat: deprecate config tokens, require JWT secret, prefer JWT auth

Phase 1 of config-token removal:

- load_jwt_secret() now exits with error if no secret is configured
  (was: silently auto-generated ephemeral secret)
- _authenticate_token() logs deprecation warning on config token use
- CLI /cluster commands use ServiceTokenManager when JWT secret is set
- turnstone-admin tls-list uses ServiceTokenManager when JWT secret is set
- Update bootstrap wizard, docker.md, security.md to mark
  TURNSTONE_AUTH_TOKEN as deprecated and JWT_SECRET as required
- Console test fixtures use auth token + headers (auth always enforced)

* feat: add service scope for inter-service JWT auth

Add "service" to VALID_SCOPES and SCOPE_HIERARCHY. Service tokens
bypass require_permission() RBAC checks, replacing the old
empty-user-id bypass that config tokens relied on.

All ServiceTokenManager instances that need admin access now include
"service" in their scopes (console proxy, channel gateway, CLI,
admin CLI). Read-only services (collector, notification) unchanged.

* feat: phase 2 config token deprecation

- SDK doc examples now show API tokens (ts_) instead of config tokens
- Remove _get_config_token() from admin CLI (dead code)
- Block config token exchange in handle_auth_login — only password
  and API token login allowed
- Update login tests to use password-based auth instead of config
  token exchange

* feat: phase 3 — remove config tokens entirely

Complete removal of config-file token authentication:

- Delete AuthConfig.tokens, check(), _ROLE_TO_SCOPES, hmac dispatch
  branch, and config token loading from load_auth_config()
- Remove auth_config parameter from _authenticate_token() and
  check_request() — callers updated throughout
- Remove TURNSTONE_AUTH_TOKEN from compose.yaml, Helm charts,
  Terraform, turnstone.example.toml
- Remove --auth-token CLI flags from turnstone, turnstone-admin,
  and turnstone-console
- Simplify console main() — always use ServiceTokenManager
  (no fallback to static tokens)
- Delete config-token-specific tests, rewrite check_request and
  integration tests to use JWT auth with proper audience claims
- Remove all config token references from docs (security.md,
  docker.md, sdk.md, console.md, architecture.md, bootstrap prompt)

* fix: address code review findings

- Fix 33 broken tests: add JWT auth to test_api_versioning,
  test_console_routing_proxy, test_tls_admin, test_tls_manager,
  test_server_live (jwt_secret + audience-scoped auth headers)
- Add TestRequirePermissionServiceScope: 4 tests covering the
  service scope RBAC bypass path
- Remove stale comments referencing config tokens in auth.py and
  console/server.py
- Remove dead proxy_auth_token parameter from console create_app()
  and static token fallback in _proxy_auth_headers()
- Remove TURNSTONE_AUTH_TOKEN from env.py scrub list

* fix: address Copilot review — JWT audience, compose require secret

- CLI /cluster: add audience=JWT_AUD_CONSOLE to ServiceTokenManager
  (console validates audience, JWTs without it were rejected)
- Admin CLI tls-list: same audience fix
- compose.yaml: TURNSTONE_JWT_SECRET now uses :? to fail fast if unset
- SDK console: fix default port from 8081 to 8090

* test: add auth enforcement tests for TLS admin endpoints

5 new tests: unauthenticated requests return 401 (list, renew,
delete), read-only-scoped requests return 403 (renew, delete).
Closes the TLS auth enforcement test gap noted in PROGRESS.md.

* fix: address remaining Copilot review feedback

- Fix token_source="config" → "test" in TLS test fixtures
- Fix AuthResult.token_source docstring to include service origins
- Require TURNSTONE_JWT_SECRET in cluster compose profile (:?)
- Helm: add auth.jwtSecret + auth.existingSecret values, wire
  TURNSTONE_JWT_SECRET into secret.yaml and both deployments
- Terraform: replace auth_token with jwt_secret variable + secret,
  remove orphaned auth_token resources and IAM reference
- Remove [[auth.tokens]] from security.md config example

* fix: address full code review — 10 findings

Critical:
- Terraform: replace concat(common_env, auth_env) with common_env
  (auth_env local was removed but still referenced)
- Channel gateway: remove hmac static token auth from _check_auth(),
  use JWT-only validation. Remove --auth-token CLI arg from channel
- Rebalancer: add token_manager support so migration requests carry
  JWT auth (was sending unauthenticated POST to /internal/migrate)

Major:
- Guard _permissions_to_scopes() against "service" privilege
  escalation from DB role permissions
- Remove dead AuthConfig class, load_auth_config(), and all
  auth_config parameters from create_app() signatures
- Helm: inject JWT secret for both inline and existingSecret paths

Minor:
- Remove dead auth_token param from ClusterCollector
- Remove empty TestLoadAuthConfig class
- Short JWT secret now exits instead of warning
- Compose: add generation command comment above JWT_SECRET
- Clean stale config token references from 6 doc files
- Clean stale AUTH_TOKEN reference from bootstrap wizard prompt

* fix: remove remaining stale config token references from docs

- channels.md: remove --auth-token from options table
- oidc.md: remove "config-file tokens still work" claim
- security.md: remove config token section, fix JWT secret docs
  (now required/exits, no ephemeral fallback), remove hmac from
  ASCII diagram, remove --auth-token reference
2026-04-01 19:38:24 -07:00
Patrick Buckley a315cabe71 chore: polish — remove dead code, update diagrams and docs
Remove stale Redis/Bridge/MQ references found via vulture scan and
manual grep:
- bot.py docstring: remove Redis MQ reference
- server.py trusted_sources: remove "bridge"
- tls.py docstring: remove "bridge" from service list

Delete 4 obsolete diagram pairs (puml + png):
- 06-mq-protocol, 07-message-routing, 08-redis-key-schema,
  10-simulator-architecture

Update 7 diagrams to reflect direct HTTP architecture:
- system-context, package-structure, workstream-states,
  console-data-flow, deployment, channel-architecture,
  settings-architecture

Redraw architecture-overview.svg: Console router replaces Redis MQ,
direct SSE data plane, hash ring routing.
2026-03-30 20:30:05 -07:00
Patrick Buckley 62ff3217d0 fix: TLS Docker end-to-end testing fixes (#185)
* fix: TLS Docker end-to-end testing fixes

Fixes discovered during Docker Compose TLS integration testing:

- Dockerfile: use --extra all (prevents missing optional deps)
- lacme 1.0.3: fixes CACertificateIssued event logging crash
- chmod PermissionError: guard for Docker volume mounts
- socket import: moved to top of main() (was inside TLS conditional,
  caused NameError in _default_node_id)
- redis.SSLConnection: ConnectionPool needs explicit connection_class,
  not ssl=True (which only works on Redis() directly)
- Empty redis password: pass None instead of "" to avoid AUTH error
- TURNSTONE_CONSOLE_URL: env var for Docker service discovery
  (0.0.0.0 bind address isn't reachable from other containers)
- HTTP01Handler: ACME client needs a challenge handler even when
  server auto-approves
- Docker overlay: tls-init as root with chmod, Redis conditional
  password, console Redis TLS flags, TURNSTONE_CONSOLE_URL

* feat: full mTLS end-to-end with lacme 1.0.4

Completes the mTLS chain across all services:

lacme 1.0.4:
- Dual EKU certs (serverAuth + clientAuth) — fixes mTLS rejection
- Configurable CA name (name="turnstone") — consistent store key

Bootstrap CA import:
- Console imports bootstrap CA from /certs volume on first boot
- Single trust root: bootstrap CA → console → all service certs

Bridge mTLS:
- TLSClient init when TURNSTONE_TLS_ENABLED set
- Auto-upgrades server URL from http:// to https://
- SSLContext passed to all 3 httpx clients via verify=

Console collector mTLS:
- upgrade_tls() method replaces httpx client with mTLS context
- Called in lifespan after cert issuance alongside proxy upgrade
- Fixes "Failed to poll node" when server serves HTTPS

Docker overlay:
- TURNSTONE_TLS_SANS on all services (Docker service names as SANs)
- TURNSTONE_TLS_ENABLED on bridge
- Channel service with Redis TLS flags
- TURNSTONE_CONSOLE_URL for service discovery
- Server healthcheck disabled (mTLS healthcheck deferred)
- Redis conditional password from env

Verified end-to-end: bootstrap → console CA → server HTTPS →
bridge mTLS → Redis TLS → channel Redis TLS → console collector
polls server over mTLS → workstream creation works through bridge

* fix: lint + copilot feedback on TLS Docker e2e

- SIM105: contextlib.suppress(PermissionError) for chmod
- F401: remove unused get_storage import in bridge
- Redis healthcheck: pass password when REDIS_PASSWORD is set

* fix: sort imports in admin.py and bridge.py

* fix: tls-init key permissions, healthcheck env, collector race

- tls-init: add set -e, chown to turnstone:turnstone with restrictive
  perms (keys 0600, certs 0640, dirs 0750) instead of world-readable
- Redis healthcheck: use container runtime $$REDIS_PASSWORD instead of
  Compose-time interpolation for consistency with --requirepass block
- collector upgrade_tls(): don't close old httpx client while concurrent
  poll threads may still be using it — let GC handle cleanup
2026-03-26 14:24:44 -07:00
Patrick Buckley b086390558 fix: TLS deferred work — wiring, security, tests, Docker overlay (#184)
* fix: TLS deferred work — wiring, security, tests, Docker overlay

Security fixes:
- PostgreSQL SSL: validate sslmode against known values, urlencode
  all params to prevent URL injection
- ConfigStore env seeding: removed redundant type coercion, delegate
  to validate_value() which handles all coercion correctly

Functional wiring:
- Database SSL: init_storage() passes SSL params to PostgreSQL URL
- Server: env var fallbacks for DB SSL (TURNSTONE_DB_SSLMODE etc.)
- Proxy mTLS: re-create proxy clients after TLS cert issuance
- Channel gateway: --ssl-certfile/keyfile/ca-certs CLI args, HTTPS
  advertise URL when SSL configured
- ConfigStore env seeding: TURNSTONE_{SECTION}_{KEY} seeds on first boot
- Console deregistration on shutdown (with debug logging)

Specs, tests, Docker:
- OpenAPI: 5 TLS admin endpoints in console_spec.py
- Auth enforcement test (401 without auth)
- SDK ValueError test (mismatched cert/key)
- Docker overlay: TURNSTONE_TLS_ENABLED, bridge --redis-tls, Redis
  healthcheck with client cert
- Removed stale type:ignore comments (lacme 1.0.2 type stubs)

* review: address copilot feedback on TLS deferred work

- ConfigStore env seeding: use config_store.set() instead of
  storage.set_system_setting() (correct API, updates cache)
- Remove unused defn variable (iterate SETTINGS keys only)
- Fix structlog call-arg error (positional args, not kwargs)
- Channel gateway: validate cert+key provided together
- Restore type:ignore[no-any-return] for CI mypy (lacme 1.0.2
  type stubs not in CI's mypy overrides yet)

* fix: rename _VALID_SSLMODES to lowercase (N806)
2026-03-25 22:43:35 -07:00
Patrick Buckley d08a57dfc2 feat: SDK TLS support, Docker Compose overlay, TLS docs (#183)
Python SDK:
- ca_cert, client_cert, client_key on all 4 client classes
- ValueError if only one of client_cert/client_key provided
- Passed to httpx verify=/cert=

TypeScript SDK:
- TlsOptions type exported (zero runtime code)
- Fix picomatch vulnerability (npm audit fix)

Docker Compose:
- deploy/docker-compose.tls.yml overlay with tls-init bootstrap
- Notes it's an overlay requiring a base compose file

Documentation:
- docs/tls.md: architecture, config, CLI, SDK examples, troubleshooting
- Fixed package name (@turnstone/sdk), Node.js 18+ note
2026-03-25 20:43:34 -07:00
Patrick Buckley 10165bb8a1 feat: add OpenShell sandbox policy for turnstone-server (#128)
* feat: add OpenShell sandbox policy for turnstone-server

Curated policy for running turnstone-server inside an OpenShell sandbox
with kernel-enforced security boundaries (Landlock, netns, seccomp).

- Filesystem: workdir read-write, /usr+/etc read-only, /tmp+/dev/null
  read-write, Landlock best_effort compatibility
- Network: default-deny with allowlisted LLM APIs (OpenAI, Anthropic),
  Tavily, skills.sh, GitHub (read-only L7), MCP registry (read-only L7),
  Redis localhost, package registries, curated web_fetch domains
- Git: L7-enforced read-only (info/refs + git-upload-pack only)
- Process: privilege drop to sandbox:sandbox
- Inference routing template for credential isolation (real API keys
  never enter the sandbox, resolved at proxy layer)

* fix: address PR #128 review feedback + add integration guide

Review fixes:
- Use python3 (not python) in usage examples to match binary allowlist
- Fix network_policy → network_policies in comment
- Remove /usr/bin/git from github_api (git uses github.com not
  api.github.com; already covered by git_operations policy)
- Remove pip/uv from bash_network_tools (package_registries already
  covers their PyPI access; no need for StackOverflow/Wikipedia reach)
- Restructure routes.yaml so commented blocks are indented under
  routes: key (uncomment without restructuring YAML)

New: docs/openshell.md covering policy customization, inference routing,
domain allowlisting, MCP subprocess inheritance, and the dual-layer
security model.
2026-03-18 18:09:47 -07:00
renovate[bot] f494553020 chore(deps): update helm release redis to v25 (#93)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-03-15 15:58:41 -07:00
renovate[bot] c766c81f25 chore(deps): update helm release postgresql to v18 (#92)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-03-15 15:58:38 -07:00
renovate[bot] 387ef06da7 chore(deps): update helm release redis to ~20.13.0 (#89)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-03-15 15:56:13 -07:00
renovate[bot] e4e2200c33 chore(deps): update helm release postgresql to ~16.7.0 (#88)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-03-15 15:56:11 -07:00
Patrick Buckley 2b58c127b1 Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging (#20)
* Add pluggable storage backend (SQLite + PostgreSQL) and deployment packaging

Database abstraction: StorageBackend protocol with 21 methods, SQLAlchemy Core
schema, SQLite backend (FTS5), PostgreSQL backend (tsvector/ILIKE), Alembic
migrations, singleton registry. memory.py reduced to thin facade. Session.py
open_db() calls replaced with generic KV methods. [database] config section
with env var support.

Deployment: Docker Compose production profile with PostgreSQL, Dockerfile with
postgres extras and migration entrypoint, Helm chart with bitnami subcharts,
Terraform AWS ECS/Fargate module with RDS + ElastiCache + ALB.

39 new storage tests (934 total). mypy strict clean. Docs and diagrams updated.

* Address PR #20 review feedback (16 items)

- Backends only call create_all() when Alembic migrations are disabled
- Helm configmap uses correct TURNSTONE_DB_BACKEND env var; DB URL
  constructed via env expansion with secret reference instead of ConfigMap
- Migration errors fail fast for PostgreSQL (only non-fatal for SQLite)
- save_memory/delete_memory wrapped in exception handling like other facade fns
- pool_size passed through from config/env to init_storage() in cli + server
- Terraform: DB URL moved to Secrets Manager, auth enabled flag set,
  optional TLS listeners with certificate_arn, Redis transit encryption on
- Docker entrypoint no longer suppresses migration output
- Diagram fixes: removed StaticPool claim, removed non-existent migration ref
- compose.yaml/README: clarified production profile requires DB env vars
2026-03-03 22:57:34 -08:00