Repository navigation
fix(runner): refuse a wave when the process table is nearly full - #109
Merged
Merged
Conversation
…ly full omnigent host's orphan reaper stalled on 16 September and ~35,900 zombies filled the runner's pids limit; every wave from then on died at its first command with a Go thread-creation trace from gh that named neither cause nor fix. preflight_processes now runs first, reads pids.current/pids.max with builtins only (a full table cannot start cat), warns at 50%, refuses at 90% and names the zombie count and the restart.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why
On 2026-09-28 every wave from 11:26Z failed at
gh issue listwith a Go "failed to create new OS thread (errno=11)" trace. The runner container had hit its process limit: 38,076 of 38,079, of which about 35,900 were zombies. omnigent host (PID 12) is a child subreaper, and its orphan reaper had stopped collecting on 2026-09-16. The trace named neither the cause nor the fix. A container restart cleared it.preflight_processesnow runs first in every wave, before the kernel preflight and before any model call. It readspids.currentandpids.maxwith shell builtins only, because a full table cannot startcateither. It warns at 50%, and at 90% refuses the wave with the zombie count and the restart command. Where there is no cgroup limit (macOS, no cgroup v2, ormax) it skips.Effect on the merge gate
None.
Testing
bash batch/run-queue.sh --self-testpasses (macOS; Linux via CI)tests/execution/test_preflight_processes.pycovers refuse, warn, healthy, no limit, no cgroup files, and that it runs beforepreflight_kernelinmainWhole suite: 2027 passed, 3 skipped.
Anything a reviewer should look at twice
The underlying stall is omnigent's. We are upgrading to v0.15.0 before deciding whether to report it upstream. This guard only makes the failure legible and early.