You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The self-hosted buildx local cache on the homeserver keeps ending up with an index.json that references blobs which do not exist under blobs/sha256/. BuildKit then fails the whole image build with an error that points nowhere near the cause:
ERROR: NotFound: rpc error: code = NotFound desc = content sha256:<digest>: not found
ERROR: failed to build: failed to solve: lease "<id>": not found
Measured occurrences
Both observed on 2026-09-25, in different image dirs, on different runs:
/var/buildx-cache/dicompot — 9 blobs referenced but absent (64 MB dir, only 18 blob files on disk)
/var/buildx-cache/backend-service — 5 blobs referenced but absent (645 MB)
In both cases the set of digests CI errored on matched the set of missing blobs exactly, so the diagnosis is not a guess. An audit of all 21 cache dirs found only these two corrupt; the other 19 were clean.
Why this matters
A corrupt cache dir is a hard build failure, not a slow build, and the error names a lease and a content digest rather than the cache. It reads like a buildx/runner fault, which sends you looking at the builder (I did — that was wrong) or at the base image digest (also wrong; the golang:1.27-alpine digest resolves fine on a fresh registry token). Meanwhile the whole Containers matrix for that image is red on every open PR.
Manual recovery is mv /var/buildx-cache/<image> /var/buildx-cache/.quarantine-<image>-<ts> then rerun. That is a manual step on a box that is otherwise unattended, and it recurred on a second image within the hour.
Known contributing factor, not yet the cause
Two things are known about how the cache dir is written:
containers.yml sets umask 002 before mkdir, but that governs only the workflow shell. BuildKit writes index.json / oci-layout / blobs from inside its buildkitd container under that container's own umask. Fixed in fix(ci): keep buildx cache group-writable across runner users #3304 for the group-permission half.
scripts/prune-buildx-cache.sh deletes blobs older than PRUNE_DAYS (14) individually, and only resets the whole directory when the total exceeds MAX_BYTES (2 GiB). Its own comment notes that a single missing blob voids the entire import:
WARNING: local cache import at <dir> skipped: digest sha256:... unavailable
The script's reasoning is that this "cannot HARD-fail a build". That appears to be wrong, or at least incomplete: a partially-pruned directory is exactly the state observed here, and it did hard-fail. Worth checking whether the age-based pass is the thing creating the corruption.
Concrete questions
Does the age-based blob deletion in prune-buildx-cache.sh produce a directory that BuildKit still exports an index.json for? If so, that is the corruption source and the age pass should be dropped in favour of the existing whole-directory reset, at the cost of a cold cache every PRUNE_DAYS.
Do concurrent matrix rows for the same image ever overlap on the same cache dir? Two runs building dicompot at once was observed to produce a permission denied on oci-layout. If cache-to is not concurrency-safe per image, rows need serialising or a per-run dir.
Does docker builder prune -af (run during this investigation) or an aborted build mid-export leave a partial index.json behind?
Notes
#3304 (fix(ci): keep buildx cache group-writable across runner users) fixes the permission half of this — the group-writable oci-layout that locked out all but the last-writing runner. It does not address corruption. Do not close this as fixed by #3304.
The self-hosted buildx local cache on the homeserver keeps ending up with an
index.jsonthat references blobs which do not exist underblobs/sha256/. BuildKit then fails the whole image build with an error that points nowhere near the cause:Measured occurrences
Both observed on 2026-09-25, in different image dirs, on different runs:
/var/buildx-cache/dicompot— 9 blobs referenced but absent (64 MB dir, only 18 blob files on disk)/var/buildx-cache/backend-service— 5 blobs referenced but absent (645 MB)In both cases the set of digests CI errored on matched the set of missing blobs exactly, so the diagnosis is not a guess. An audit of all 21 cache dirs found only these two corrupt; the other 19 were clean.
Why this matters
A corrupt cache dir is a hard build failure, not a slow build, and the error names a
leaseand acontent digestrather than the cache. It reads like a buildx/runner fault, which sends you looking at the builder (I did — that was wrong) or at the base image digest (also wrong; thegolang:1.27-alpinedigest resolves fine on a fresh registry token). Meanwhile the whole Containers matrix for that image is red on every open PR.Manual recovery is
mv /var/buildx-cache/<image> /var/buildx-cache/.quarantine-<image>-<ts>then rerun. That is a manual step on a box that is otherwise unattended, and it recurred on a second image within the hour.Known contributing factor, not yet the cause
Two things are known about how the cache dir is written:
containers.ymlsetsumask 002beforemkdir, but that governs only the workflow shell. BuildKit writesindex.json/oci-layout/ blobs from inside its buildkitd container under that container's own umask. Fixed in fix(ci): keep buildx cache group-writable across runner users #3304 for the group-permission half.scripts/prune-buildx-cache.shdeletes blobs older thanPRUNE_DAYS(14) individually, and only resets the whole directory when the total exceedsMAX_BYTES(2 GiB). Its own comment notes that a single missing blob voids the entire import:The script's reasoning is that this "cannot HARD-fail a build". That appears to be wrong, or at least incomplete: a partially-pruned directory is exactly the state observed here, and it did hard-fail. Worth checking whether the age-based pass is the thing creating the corruption.
Concrete questions
prune-buildx-cache.shproduce a directory that BuildKit still exports anindex.jsonfor? If so, that is the corruption source and the age pass should be dropped in favour of the existing whole-directory reset, at the cost of a cold cache everyPRUNE_DAYS.dicompotat once was observed to produce apermission deniedonoci-layout. Ifcache-tois not concurrency-safe per image, rows need serialising or a per-run dir.docker builder prune -af(run during this investigation) or an aborted build mid-export leave a partialindex.jsonbehind?Notes
#3304 (
fix(ci): keep buildx cache group-writable across runner users) fixes the permission half of this — the group-writableoci-layoutthat locked out all but the last-writing runner. It does not address corruption. Do not close this as fixed by #3304.