GitHub-hosted images carry Node 24.19.0 in /opt/hostedtoolcache, so hosted jobs
resolved from there and never populated the cached root. They still ran a
restore that could only miss and then a save whose path did not exist, logging
"Path Validation Error: Path(s) specified in the action for caching do(es) not
exist" on every hosted job. No entry was ever written, so nothing was poisoned,
but the warning is noise and the steps are pure waste on ~30 jobs per run.
Gate both steps on runner.environment != 'github-hosted', the signal this
workflow already uses for Blacksmith-only behavior, and gate the save on the
setup step reporting that a download actually populated the root. That second
guard also covers a future self-hosted image whose toolcache clears the floor.
Verified on Blacksmith: a satisfying image toolcache leaves the root absent so
the save is skipped, and a download populates it so the save runs.
Blacksmith's image tracks an older runner-images snapshot. Measured on a leased
box 2026-08-16, its toolcache holds Node 20.20.0, 22.22.0 and 24.13.0 while this
repo's engines floor is >=22.22.3 and >=24.15.0 -- short by three and two
patches. Every candidate is rejected, so all 306 of 306 sampled jobs fell
through to a nodejs.org download. GitHub-hosted runners carry 24.19.0 and
resolve from the toolcache in about a second, which is why only Blacksmith pays.
Normally that download is 2.6s (p99 3.3s), but ~46 jobs fetch the same 50 MB
simultaneously and the mirror throttles: three of 53 sampled runs had setup-node
medians of 44-93s with maxes to 139s, and because every job pays at once it
lands whole on the wall -- those runs went ~210s to 325s.
Keep the payload in the Actions cache, which Blacksmith serves from its
colocated backend. Measured on Blacksmith: cold 1605ms, warm 77ms.
Restores are prefix-keyed and the save carries the resolved patch. An exact key
would be worse than nothing: cache entries are immutable and an exact hit
suppresses the post-job save, so a floating `24.x` key would pin the first Node
it ever saw and, once the floor advanced past it, every job would restore the
rejected payload and re-download forever. Keying the save on the installed
version lets a newer resolve publish a new entry that later prefix restores
pick up.
A rejected payload is pruned before the replacement installs, because the entry
is saved wholesale and a leftover would ride along in every future save.
Windows keeps its existing path. Proven on Blacksmith across cold, warm, stale
and truncated-binary cases; both guards are mutation-checked.
This stays useful even if Blacksmith refreshes their image: the floor moves
independently of the snapshot, so the gap recurs. The image refresh is still
the better fix and is worth asking them for.
The manifest planner closure and the protocol coverage script import only
node builtins and relative files (verified importing the full closure with
an empty node_modules under native type stripping), so push/PR preflight
drops the pnpm store restore and install (~30s off the barrier every lane
waits behind). Manual dispatches keep the tsx path for frozen targets, and
the coverage script inlines the record guard under the documented
dependency-free exception.
The store archive accretes every prior lockfile generation through
prefix-key restores (measured 2.05 GiB, ~36s restore in every hosted job);
the warmup writer now prunes to the current lockfile closure before saving.
* build(deps): remove npm shrinkwrap; mirror pnpm lock into transient package locks
npm 12 removed shrinkwrap (command + tarball/root loading). Delete all 82
committed npm-shrinkwrap.json files and stop publishing lockfiles; keep
pnpm-lock.yaml as the single reviewed dependency boundary. The generator
becomes scripts/generate-npm-package-lock.mjs and feeds plugin bundling via
a transient package-lock.json + npm ci (works on npm 11 and 12). Tarball
validation treats the published 2026.7.2 beta train as a shrinkwrap
transition; self-update npm detection now uses install topology instead of
the shipped shrinkwrap.
* fix(deps): repair lint, deadcode, and test-type lanes for the npm 12 migration
- sort integrity comparisons with an explicit comparator (oxlint)
- keep resolveBunGlobalNodeModules module-local (knip unused-export gate)
- model npm pack --json as npm<=11 array / npm 12 name-keyed object
- default calver destructuring in the tarball test fixture
The installation hit Blacksmith's backing-disk cap because sticky keys
embedded PR numbers and content hashes, minting a new disk per PR and
per input change until every mount 429-failed fleet-wide. node-deps
(#109752) and ext-boundary (#109804) were already re-keyed; this
converts the last two minters and audits the third suspect.
Keys changed:
- Vitest fs transform cache: `vitest-fs-v2-<pr-N|protected>-...` ->
the one existing `vitest-fs-v2-protected-<os>-<arch>-node-<ver>`
disk. Entries are content-hash keyed (vitest sha1 over
id+content+NODE_ENV+version+env config), so cross-PR sharing is safe
by construction: PRs now mount the protected snapshot directly
(read-only, commit false) and the separate PR-only seed mount is
deleted. Writers (non-PR shard writer + scheduled warm run) keep
explicit commit true. Disk count: 1 + O(open PRs) -> 1.
- Gradle: `gradle-v1-<task>-<pr-N|protected>-<hashFiles(deps)>` ->
`gradle-v2-<task>`. Task scope stays (light ktlint must not seed
heavy build lanes); PR number and dependency hash leave the key. The
dependency hash moved to an in-job fingerprint marker that makes the
non-PR writer rebuild its snapshot cold when inputs change, bounding
disk growth; PR mounts are read-only and safely reuse stale
snapshots because Gradle caches are content-addressed. Disk count:
O(tasks x PRs x dependency bumps) -> O(tasks) (7 today).
- Docker builder (useblacksmith/setup-docker-builder): audited, out of
scope. The action owns its key internally and always uses the repo
name only (src/setup_builder.ts getStickyDisk), so it is already one
disk per installation repo and cannot mint per-PR disks.
PR warm-path behavior changes:
- Vitest: PRs lose only PR-local warm entries for files the PR itself
changed (previously persisted on the pr-N disk between pushes);
changed files re-transform once per push, unchanged files still hit
the protected snapshot. PRs touching transform inputs
(lockfile/tsconfig/package.json) run cold per push since the
generation wipe is now local to the discarded clone.
- Gradle: dependency-bump PRs get warmer (stale-but-valid protected
snapshot instead of a cold fresh key); other PRs are unchanged.
Deleted `.github/workflows/pr-cache-cleanup.yml`: it existed for the
per-PR cache layer (#109425) and only deleted GitHub actions/cache
archives, which GitHub's own LRU/TTL eviction already handles for the
remaining fork/Windows per-PR archive paths.
Guards: pinned the O(1) vitest key + non-PR commit gate, added a
Gradle per-task key/writer/fingerprint guard, added a repo-wide scan
asserting no useblacksmith/stickydisk key ever contains
github.event.pull_request.number or hashFiles(), and replaced the
cleanup-workflow assertions with a stays-deleted check.
* perf(ci): serve node_modules snapshots from O(1) protected sticky disks
Every Blacksmith mount of the dependency sticky disk currently 429s:
the v2 key minted one backing disk per PR and per manifest hash, which
saturated Blacksmith's installation-wide 1000-sticky-disk budget (run
29559333389, job checks-node-compact-small-15: "Sticky disk limit
exceeded"). Same-repo shards then fall back to a cold, storeless pnpm
install (~40s) on every run - worse than the ~22s actions/cache path
the sticky rollout replaced.
Re-key the snapshot to one stable disk per node-version and move the
install inputs into a runtime fingerprint marker:
- Key is now `<repo>-node-deps-bind-v3-<node-version>`; dependency
changes refresh the disk in place instead of allocating a new one.
- The fingerprint (manifest hashFiles set + node version + lockfile
mode) is evaluated on the bind step, before the mount lands, so the
'**/package.json' glob cannot sweep snapshot-internal manifests.
- Consumers mount read-only (commit: false); on fingerprint match they
restore importer archives and skip pnpm install entirely, on mismatch
they install against the clone's on-disk store (warm store, no
actions/cache download) without capturing.
- Writers are trusted non-PR jobs only: build-artifacts on canonical
pushes plus the scheduled vitest-cache-warm run as a deadline writer
(main pushes cancel each other under merge traffic and would starve
the snapshot). commit: on-change keeps warm no-op runs from
committing.
- Roll the consumer flag out to every Linux Blacksmith lane that
installs dependencies (build-artifacts, check-shard,
check-additional-shard, check-docs, checks-ui, control-ui-i18n,
native-i18n, qa-smoke-ci-profile, checks-fast-core, both contract
shards) with the same fork/dispatch gates as the nondist shard.
Guards now enumerate the consumer set, pin the O(1) key shape, enforce
single-writer commit expressions, and cover the fingerprint marker in
the importer capture/restore helper.
Verified locally: importer archive capture 0.96s/228K and restore
0.07s against the real 2.0GB hoisted tree; a snapshot untarred into a
fresh workspace resolves modules, runs bin shims, `pnpm exec`, and a
Vitest suite (one-time pnpm reconcile only when the absolute path
changes, which Blacksmith's fixed /home/runner/_work layout avoids).
* fix(ci): commit dependency snapshots explicitly, not via on-change heuristic
stickydisk's on-change mode compares allocated disk bytes with a 4KB
threshold, so a fingerprint refresh whose reinstall keeps usage stable
(metadata-only manifest edits, same-sized dependency swaps) could be
silently discarded, stranding every consumer on a stale marker and a
permanent reinstall path. Writers now commit explicitly, mirroring the
Vitest transform disk's rationale; warm no-op writer runs re-commit an
identical snapshot, which is cheap and safe. Also document that the
non-PR commit gate binds cooperating code only - the enforced trust
boundary stays the fork/dispatch runner gate, matching the protected
node-compile disk's posture.
* test(ci): guard sticky consumers against writerless node-version key splits
Reviewer follow-up: the snapshot key is partitioned by node-version and
both writers (build-artifacts, vitest-cache-warm) rely on the action
default, so a consumer pinning any other version would split onto a
key nobody seeds and silently regress to permanently cold installs.
Pin the action default to 24.x and assert every sticky consumer
resolves to that same key segment.
* test: narrow sticky-consumer step.with before indexing
The filter's runtime narrowing did not carry into the map callback, so
step.with indexing failed check-test-types (TS18048). Collect a narrowed
stepWith instead.
One-time maintainer-authorized bootstrap merge for the release-gate verifier policy. Exact hosted CI and all supporting workflow gates passed on 66133de419.
Recreated from #85108 because the original branch could not be updated by maintainers.
Preserves current-main pnpm install hardening while switching workflow pnpm setup to packageManager, and adds exact version-scoped release-age exclusions for already-locked packages that pnpm 11.2.2 audits during install.
Co-authored-by: Altay <altay@hey.com>