Skip to content

self-hosted buildx cache corrupts: index.json references absent blobs #3305

Description

@Xore

The self-hosted buildx local cache on the homeserver keeps ending up with an index.json that references blobs which do not exist under blobs/sha256/. BuildKit then fails the whole image build with an error that points nowhere near the cause:

ERROR: NotFound: rpc error: code = NotFound desc = content sha256:<digest>: not found
ERROR: failed to build: failed to solve: lease "<id>": not found

Measured occurrences

Both observed on 2026-09-25, in different image dirs, on different runs:

  • /var/buildx-cache/dicompot — 9 blobs referenced but absent (64 MB dir, only 18 blob files on disk)
  • /var/buildx-cache/backend-service — 5 blobs referenced but absent (645 MB)

In both cases the set of digests CI errored on matched the set of missing blobs exactly, so the diagnosis is not a guess. An audit of all 21 cache dirs found only these two corrupt; the other 19 were clean.

Why this matters

A corrupt cache dir is a hard build failure, not a slow build, and the error names a lease and a content digest rather than the cache. It reads like a buildx/runner fault, which sends you looking at the builder (I did — that was wrong) or at the base image digest (also wrong; the golang:1.27-alpine digest resolves fine on a fresh registry token). Meanwhile the whole Containers matrix for that image is red on every open PR.

Manual recovery is mv /var/buildx-cache/<image> /var/buildx-cache/.quarantine-<image>-<ts> then rerun. That is a manual step on a box that is otherwise unattended, and it recurred on a second image within the hour.

Known contributing factor, not yet the cause

Two things are known about how the cache dir is written:

  1. containers.yml sets umask 002 before mkdir, but that governs only the workflow shell. BuildKit writes index.json / oci-layout / blobs from inside its buildkitd container under that container's own umask. Fixed in fix(ci): keep buildx cache group-writable across runner users #3304 for the group-permission half.

  2. scripts/prune-buildx-cache.sh deletes blobs older than PRUNE_DAYS (14) individually, and only resets the whole directory when the total exceeds MAX_BYTES (2 GiB). Its own comment notes that a single missing blob voids the entire import:

    WARNING: local cache import at <dir> skipped: digest sha256:... unavailable
    

    The script's reasoning is that this "cannot HARD-fail a build". That appears to be wrong, or at least incomplete: a partially-pruned directory is exactly the state observed here, and it did hard-fail. Worth checking whether the age-based pass is the thing creating the corruption.

Concrete questions

  • Does the age-based blob deletion in prune-buildx-cache.sh produce a directory that BuildKit still exports an index.json for? If so, that is the corruption source and the age pass should be dropped in favour of the existing whole-directory reset, at the cost of a cold cache every PRUNE_DAYS.
  • Do concurrent matrix rows for the same image ever overlap on the same cache dir? Two runs building dicompot at once was observed to produce a permission denied on oci-layout. If cache-to is not concurrency-safe per image, rows need serialising or a per-run dir.
  • Does docker builder prune -af (run during this investigation) or an aborted build mid-export leave a partial index.json behind?

Notes

#3304 (fix(ci): keep buildx cache group-writable across runner users) fixes the permission half of this — the group-writable oci-layout that locked out all but the last-writing runner. It does not address corruption. Do not close this as fixed by #3304.

Activity

  1. added
    opsDeployment, runners, observability, host access
    ci-queue-stallCI queue stall alarm (scripts/ci-queue-watch.py)
    on Sep 25, 2026
  2. github-actions commented on Sep 25, 2026

    @github-actions

    Queue recovered as of 2026-09-25T17:07:45Z: no runs parked beyond 60m. Closing the alarm; the next stalled sweep reopens.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ci-queue-stallCI queue stall alarm (scripts/ci-queue-watch.py)opsDeployment, runners, observability, host access

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions