Skip to content

fix(codebuild): sweep only build containers whose owning process is gone - #2526

Merged
vieiralucas merged 2 commits into
mainfrom
fix/codebuild-sweep-live-owners
Sep 14, 2026
Merged

vieiralucas merged 2 commits into
mainfrom
fix/codebuild-sweep-live-owners

Conversation

@vieiralucas

@vieiralucas vieiralucas commented Sep 14, 2026 •

Copy link
Copy Markdown
Member

Summary

With #2520's retried-failure output, #2523's E2E run captured why codebuild_real_execution::cross_phase_shell_state_persists still failed intermittently after #2519:

Provisioning: FAILED  CLIENT_ERROR  "failed to create /codebuild/build: "

That is the docker exec mkdir into the build container right after docker run -d -- the container had already been removed.

Cause

At startup, CodeBuild swept leaked build containers: every fakecloud-codebuild container whose fakecloud-instance label was not the current process. That includes containers of other fakecloud processes that are still running on the same daemon. When a second server starts -- parallel e2e test servers, or two installs side by side -- it removes the first server's in-flight build container.

The shared startup reaper (fakecloud-server/src/reaper.rs) already does this correctly: it removes an object only when its owning PID is dead. CodeBuild's own sweep skipped that check.

Reproduced with a new e2e: start a build on one server, start a second server while it runs. Old sweep: the first build ends FAILED. New sweep: SUCCEEDED.

Fix

  • pid_alive moves from the server reaper to fakecloud_core::container_net, plus owned_by_dead_process(label, is_alive): an object is orphaned only when its fakecloud-<pid> owner is neither this process nor alive, and an unparseable label is never treated as orphaned.
  • The server reaper and CodeBuild's sweep both use it, so the two cannot diverge again.
  • libc becomes a cfg(unix) dependency of fakecloud-core (already in the lockfile).

Also, test-only: ec2_instance_runtime's container should be running assertion now reports the container's status, exit code, OOM flag, error, start/finish time and logs. That test also failed once in #2523 (container present but stopped, not a pull failure) and the cause isn't identified yet; the next occurrence will name it.

Surfaces

No API, SDK, docs, conformance or count change: startup cleanup behavior only.

Test plan

  • owned_by_dead_process unit test (live other owner kept, dead owner swept, self kept, unparseable labels kept) and a pid_alive probe test in fakecloud-core; the reaper's existing tests still pass against the moved function.
  • New e2e another_server_starting_does_not_kill_a_running_build: passes with the fix; with the old sweep restored the build ends FAILED.
  • cargo clippy -p fakecloud-core -p fakecloud-codebuild -p fakecloud --all-targets -D warnings, e2e test clippy, fmt clean; fakecloud-codebuild lib 67/67.

Summary by cubic

Fixes CodeBuild's startup sweep to only remove build containers whose owning fakecloud process is dead. Previously, any container not owned by the current process was removed, so a second server starting on the same daemon would delete another server's in-flight build and fail it in PROVISIONING.

Bug Fixes

  • Moves the ownership check (pid_alive + owned_by_dead_process) into fakecloud_core::container_net and reuses it in both the server reaper and CodeBuild's sweep.
  • Improves the ec2_instance_runtime "container should be running" assertion to include container state and logs on failure.

Written for commit 774b459. Summary will update on new commits.

Review in cubic

At startup CodeBuild removed every `fakecloud-codebuild` container whose
`fakecloud-instance` label was not the current process -- including those of
other fakecloud processes that are still running and sharing the daemon.
Any second server starting (parallel e2e test servers, side-by-side
installs) killed the first server's in-flight build container, failing the
build in PROVISIONING with "failed to create /codebuild/build: " (the
`docker exec` into a container that had just been removed). The
codebuild_real_execution e2e tests hit exactly that when run in parallel.

The shared startup reaper already got this right by checking whether the
owning PID is alive. Move that check to fakecloud_core::container_net
(`pid_alive` plus `owned_by_dead_process`) and use it from both, so a
container is swept only when its owner is neither this process nor alive.

Also: ec2_instance_runtime's "container should be running" assertion now
reports the container's status, exit code, OOM flag, error, start/finish
time and logs, so its remaining intermittent failure names its cause.

Tests: owned_by_dead_process unit test; a new e2e starts a second server
while a build runs and asserts the build still succeeds (fails with the
build FAILED under the old sweep).
@vieiralucas
vieiralucas merged commit 1252fc2 into main Sep 14, 2026
157 checks passed
@vieiralucas
vieiralucas deleted the fix/codebuild-sweep-live-owners branch September 14, 2026 07:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant