What
Most CI runs on the single homeserver runner (honeypot-ci), which also runs the whole honeypot fleet and GPU benchmark/training work (#3135). #3304 (buildx cache permissions) and #3303 (lease errors) were both runner-state problems. ci-queue-watch.yml and ci-heartbeat.yml watch the queue, but there's no written contract for how much CPU, RAM and disk a CI job may take, which jobs are "heavy", or how to drain the runner before host maintenance.
Proposal
docs/CI-RUNNER.md (or a section in docs/CI-CD.md):
- The runner's labels, and which workflows target it versus hosted runners (the
ci-router.yml logic in plain words).
- Measured peak RSS and disk for the heavy lanes (Rust build, Playwright, container builds).
- Where temp and cache live, and their cleanup (
cache-cleanup.yml, scripts/prune-buildx-cache.sh).
- Drain steps: pause the runner service, wait for in-flight jobs, and verify no orphaned build containers.
Sources
Found by the #3194 OmniRoute ops/CI deep-check (pinned 18bbb101, APIARY main 3dca4457).
What
Most CI runs on the single homeserver runner (
honeypot-ci), which also runs the whole honeypot fleet and GPU benchmark/training work (#3135). #3304 (buildx cache permissions) and #3303 (lease errors) were both runner-state problems.ci-queue-watch.ymlandci-heartbeat.ymlwatch the queue, but there's no written contract for how much CPU, RAM and disk a CI job may take, which jobs are "heavy", or how to drain the runner before host maintenance.Proposal
docs/CI-RUNNER.md(or a section indocs/CI-CD.md):ci-router.ymllogic in plain words).cache-cleanup.yml,scripts/prune-buildx-cache.sh).Sources
docs/ops/RUNNER_BOX.mdFound by the #3194 OmniRoute ops/CI deep-check (pinned
18bbb101, APIARYmain3dca4457).