Skip to content

docs: self-hosted runner resource classes, capacity budget and drain-before-maintenance procedure #3327

Description

@Xore

What

Most CI runs on the single homeserver runner (honeypot-ci), which also runs the whole honeypot fleet and GPU benchmark/training work (#3135). #3304 (buildx cache permissions) and #3303 (lease errors) were both runner-state problems. ci-queue-watch.yml and ci-heartbeat.yml watch the queue, but there's no written contract for how much CPU, RAM and disk a CI job may take, which jobs are "heavy", or how to drain the runner before host maintenance.

Proposal

docs/CI-RUNNER.md (or a section in docs/CI-CD.md):

  • The runner's labels, and which workflows target it versus hosted runners (the ci-router.yml logic in plain words).
  • Measured peak RSS and disk for the heavy lanes (Rust build, Playwright, container builds).
  • Where temp and cache live, and their cleanup (cache-cleanup.yml, scripts/prune-buildx-cache.sh).
  • Drain steps: pause the runner service, wait for in-flight jobs, and verify no orphaned build containers.

Sources

Found by the #3194 OmniRoute ops/CI deep-check (pinned 18bbb101, APIARY main 3dca4457).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

documentationImprovements or additions to documentationenhancementNew feature or requestin-progressActively being worked onopsDeployment, runners, observability, host access

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions