The openclaw-mlx-tts voice helper pulls in the full mlx-swift Metal shader
stack, which some beta Xcode toolchains (e.g. Xcode 27 / macOS 27 SDK) cannot
compile: the metal compiler dies non-deterministically (a different .metal file
each run, 'Could not read serialized diagnostics file'). The main app builds
fine, so an unrelated dev/proof build should not be blocked by the helper.
Add OPENCLAW_SKIP_MLX_TTS=1 (matching the sibling SKIP_TSC/SKIP_UI_BUILD
toggles) to package the app without the voice helper, gating both the per-arch
build and the bundle copy. Refuse the flag for release builds, which must ship
the helper (notarization verifies it), so a skipped build can never become a
silently incomplete release.
* fix(snapshot): survive cold PowerShell starts in Windows staging gates
CI run 31775262530, checks-windows-node-test-1 attempt 1, showed the fail-closed ACL probe timing out during PowerShell first-use module preparation. Centralize encoded one-shot spawning, budget 60 seconds for cold starts, and preserve the underlying probe failure as the error cause.
* fix(snapshot): sanitize PowerShell failure causes in Windows staging gates
* fix(secrets): explain the sanitized plan-file failure cause suppression
check-lint-core-2 flagged preserve-caught-error at the private plan file
catch; retaining the raw error would re-leak the -EncodedCommand argv the
sanitization contract strips, so the suppression is intentional (same
idiom as setup-inference-activate.ts).
* test(lint): register the private-plan-file suppression in the inventory
* test(infra): give the LAN-host real PowerShell spawn a cold-start budget
checks-windows-node-test-2 (run 31804325922) hit the same cold-start flake
class this PR fixes: the codepage-proof test spawns real powershell.exe
bounded at 3s, which a cold runner cannot meet. Production keeps its
fail-open 3s route-hint probe; only the test's real-spawn verification
uses the shared cold-spawn budget.
Hosted CI runners restored the boundary-artifact cache and rebuilt it anyway: fresh checkouts re-stamp every input mtime, so mtime freshness never passed. Stamp files now record the input content digest and byte-identical inputs skip the rebuild (~60s saved per hosted lint/boundary job, 0.17s verify). Telegram CI shards pack ten files per job instead of five now that per-file import cost is back to seconds (#123607), halving the ~42-job fanout.
Five-file Telegram jobs finished the first file, then isolate re-imported the next graph in silence until the 300s watchdog killed the worker. Recycle the Vitest process after each file and keep five files per CI job.
Preserve externally scoped Telegram and Matrix test plans instead of expanding each CI shard back into the full extension suite. Keep broad runs bounded and retain external include ownership through directory run specs.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Bound Telegram extension tests to five files per Vitest process across explicit config, directory, and full-suite routes while preserving serial isolated execution.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(test): keep shared jsdom window in step with the per-file module reset
The non-isolated runner already resets the module graph after every test file,
so each file evaluates its own component classes. The jsdom window it shares
across the whole worker was never reset with it: every Control UI component
registers with `if (!customElements.get(tag))`, so the first file to import a
component owned that tag for the rest of the run and `document.createElement`
kept building elements closed over that file's module instances. Later files'
singletons, module mocks, and spies were never the ones production reached, so
assertions failed as "expected ... to be called once, but got 0 times" in
whichever files the size-based sequencer happened to place after a warming one.
Drop repo-owned tags with the graph they came from. Dependency packages are
externalized and register once per worker through native ESM, so their
definitions are attributed by define call site and kept.
The same window also carried mounted DOM forward, so helpers reading
`document.body.querySelector(...)` answered an earlier file's leaked dialog and
focus assertions read its stale activeElement. Clear the body with it;
`document.head` stays, since dependency styles cannot be replayed either.
* fix(test): keep the jsdom definition shape local to its module
* feat(gateway): transfer node worker workspaces
* fix(gateway): harden node workspace transfer
* fix(gateway): isolate transfer HTTP contract
* fix(gateway): trim transfer HTTP exports
PR #123406's tooling-tail split re-claimed the file (twice) while #123383
had made attempt-light its sole owner; the vitest-projects coverage
invariant fails on main. attempt-light keeps ownership.
* refactor(i18n): re-key native i18n artifacts to content-hash identity (v2)
The native inventory stored a write-only 'line' field per entry, so any
unrelated edit above a string rewrote apps/.i18n/native-source.json
(~half of all commits touching it were pure line-number churn). Identity
was (surface, path, source), duplicating the same string per file
(5385 entries for 4187 unique pairs) and churning IDs on file moves.
Locale artifacts were positional arrays repeating full English source
text, so one inserted string rewrote diff spans in all 21 files.
v2 artifacts: inventory entries keyed by (surface, source) with merged
per-site {path, kind} lists and pure sha256 content-hash IDs; locale
files become id-keyed sorted translation maps. Existing translations
carry over by source match with a deterministic duplicate pick; the
sticky-ID reuse machinery and positional validation are deleted.
Everything under apps/.i18n plus generated platform locale artifacts is
marked linguist-generated. ci-changed-scope gains a one-time
owner-complete migration escape mirroring the control-ui precedent.
CLI surface (baseline/check/sync/verify) and the locale-refresh
workflow are unchanged.
* ci: register run-attempt-state test in its Vitest lane
Commit e04dfd26e2 added extensions/codex/src/app-server/run-attempt-state.test.ts
without a lane owner, so the full-suite ownership audit
(test/vitest-projects-config.test.ts) fails on main. Register it in the
attempt-light shard alongside its run-attempt siblings.
* feat(docker): schedule image refreshes
* docs(docker): explain weekly image refreshes
* test(scripts): gate workflow-step execution on bash 4 mapfile support
Stock macOS bash 3.2 lacks mapfile; CI truth is Linux bash 5.
* test: cover docker-release suffix threading and sanctioned second caller
* test(codex): wire run-attempt-state into the attempt-extra project
#123345 added the file without a project owner; the full-suite coverage
guard fails for any PR that runs it.
* fix(ci): run build-artifacts PR validation on hosted runners
ci-build-artifacts-testbox.yml pinned PR runs to blacksmith-16vcpu and
ran Testbox lifecycle steps unconditionally, so the prepare-run landing
gate starved for every PR during a Blacksmith outage even with
OPENCLAW_CI_RUNNER_BACKEND=github. PR events now build on ubuntu-24.04
with dispatch-only Testbox steps, mirroring ci-check-testbox.yml.
* test(ci): align build-artifacts dispatch guard
Move llama.cpp chat and local embeddings onto a verified externally managed llama-server runtime. Remove the in-process native runtime, forked embedding workers, and node-llama-cpp dependency while preserving guided setup, local GGUF models, tool-capable agent runs, diagnostics, and operator docs.