ci: route all workflows homeserver-first with GitHub-hosted fallback - #2568
Merged
Merged
Conversation
…2565) Every compute workflow now ships the homeserver-first routing pair (or a conditional runs-on) that quality.yml's six families already used: - New shared .github/workflows/ci-router.yml (workflow_call) carrying the trust gate + heartbeat liveness proof quality.yml's inline router documents; containers/security/pages consume it, quality keeps its inline job until a follow-up migrates it. - ci-heartbeat.yml drops its concurrency group: one push fans out to multiple routers and cancel-in-progress would cancel a sibling router's canary evidence. - containers.yml build and security.yml (CodeQL) analyze pick runs-on off the router output; pages.yml build joins them and the deploy step stays ubuntu-latest (seconds-long API call). - quality.yml: frontend-next and frontend-next-browser become pairs; the browser home twin drops --with-deps and the apt fallback for loud presence checks (deps are preinstalled on the host); the docker-bound scripts-and-compose rows are flagged home: true now that the runner user drives docker (the four real-ES suites included -- they used to SKIP without an engine, a relocated no-op). - install-ci-runner.sh provisions the host idempotently: docker-group grant for the runner user, redis-server (daemon disabled), node 22, shellcheck, and playwright 1.62.1's chromium library set. Runner capability changes verified live on the box; trust gate and CI_HOMESERVER_PRS semantics unchanged. API-only and ops workflows stay put (rationale in #2565).
A called reusable workflow can never exceed the caller's GITHUB_TOKEN envelope, and under-granting it startup-fails the whole run as "Invalid workflow file" -- containers.yml, security.yml and pages.yml granted only their own scopes, so every run of the three startup-failed before the router job could even appear. The route job needs actions:write solely to dispatch the ci-heartbeat canary; pages' deploy keeps its own narrower job-level envelope. Also register the honeypot-ci runner label in .github/actionlint.yaml (honeypot-home was listed; its sibling was not), silencing the runner-label lint noise the new runs-on lines add.
This was referenced Aug 27, 2026
Dependency ReviewThe following issues were found:
License Issues.github/workflows/security.yml
OpenSSF Scorecard
Scanned Files
|
For pull_request runs GITHUB_REF_NAME is the ephemeral "<n>/merge" merge ref, which no branch backs -- the canary dispatch was refused with 422 "No ref found for" and every same-repo PR routed to the GitHub-hosted fallback forever, silently (seen live on #2568: the routers ran, trusted=true, then 'dispatch refused (HTTP 422X)' and online=false). Same-repo PRs now dispatch at the PR's head branch, whose ci-heartbeat.yml is the branch's own; fork PRs never reach the dispatch (trusted=false). push/workflow_dispatch/schedule already carry a real branch in GITHUB_REF_NAME. Also drop curl -f so a refused dispatch lands its HTTP code in the warning instead of aborting the capture (422X). Fixed in ci-router.yml and quality.yml's inline copy.
…lcheck scripts/github-ci-runner/install-ci-runner.sh's provision list is inside the "Shell syntax and high-severity ShellCheck" row's scan set, and a comment line beginning with "shellcheck" parses as a (malformed) shellcheck directive -- SC1073/SC1072 at --severity=error, failing the row on #2568. Reword the provision-list entry and note the trap in-place, without leading the note with the keyword either.
Follow-up wording accuracy after the merge from main: the ci-target header still said the ci-heartbeat canary is dispatched at "THIS commit" -- the dispatch API takes a ref, and for pull_request runs GITHUB_REF_NAME is the "<n>/merge" merge ref that no branch backs (422, fixed two commits ago).
Seen live on #2568's merged round: the canary dispatch was accepted and the router still saw no run object in /runs 45s later -- GitHub's run-indexing lag outran the appear window, so the router concluded online=false and fell back with the box demonstrably healthy. Defaults become appear=120s (must cover API indexing, not just dispatch latency; the loop still exits the moment the run appears) and decide=420s (the canary is tiny but queues behind whatever the single box instance is executing). Router job timeout 10 -> 15min to keep the worst case (120+420 + slack) comfortable. Fan-out deeper than ~2 canaries still overflows and falls back -- that is single-instance capacity, tracked as #2572, not a window problem.
…ef edit The dispatch-ref fix's old_string swallowed the cutoff= line and never re-added it. Under set -u, Phase A's 'jq --arg cutoff "$cutoff"' then expanded an unbound variable inside a pipeline subshell every poll: the subshell died before jq ran, printf reported a broken pipe, rid silently stayed empty, and every router concluded online=false while its canary was created the same second as the dispatch and completed success (live evidence on #2568's 18:00 round: 'cutoff: unbound variable' repeating every ~10s of the appear window, canary run created/started 18:00:54, verdict offline 18:03:04). This, not GitHub lag or window size, is why no round ever routed homeserver-first. Restored in both routers, with the failure mode documented on the line; windows stay at 120s/420s (defensible on their own merits).
…ub tests pin env path The Ghidra worker's _payload_roots() called is_dir() on three hardcoded paths without guarding against PermissionError. The CI runner user (github-ci) cannot traverse /var/lib/docker/volumes/dionaea-lib/_data/binaries, so a single stat() raised PermissionError, the whole drain() crashed mid-batch, and the spool was left inconsistent. The 'exit 0', 'missing sample quarantined', and 'second run is idempotent' assertions in test_spool all turned on the same drain state. Wrap each is_dir() in an OSError guard so a single inaccessible root is skipped, not fatal. update docstring to record why. The two new analysis/github/tests scripts (test_publish_sample.sh, test_rejections.sh) carried the old refuse-guard: hard-fail if /etc/honeypot-github.env exists, on the grounds that sourcing the env in CI would leak GH_PAT. The runner that this PR routes the homeserver-first matrix to does carry that file (root, 0600) -- the file is what the production scripts source when they exist -- so the new tests were permanently red. The other four tests in the same suite (test_dry_run.sh, test_dry_run_reprocessed_after_enable.sh, test_daily_cap.sh) had already moved to pinning GITHUB_ANALYSIS_ENV_FILE to an empty fixture inside $work, which closes the same leak deterministically. Apply the same pattern to the two new tests.
Two follow-up fixes for the new homeserver-first/GitHub-hosted fallback routing: - pages.yml: python -> python3 (the GitHub-hosted ubuntu runner has python3 on PATH, not python; the homeserver runner had both, so this is a no-op there but unblocks the fallback path). - security.yml CodeQL Go row: add actions/setup-go@v5 with go-version 1.23 for the GitHub-hosted fallback only (matrix.language == go AND homeserver != true). The CodeQL Go extractor shells out to go to enumerate build targets; the homeserver runner has go preinstalled, the GitHub-hosted runner does not. No-AI-attribution: per repo policy, no AI co-author or footer in this commit.
The homeserver runner also lacks go on PATH; the prior commit's actions/setup-go step was conditional on homeserver != true, so the homeserver (CodeQL default) path still missed the toolchain. Drop the condition so the step runs for every Go matrix row on every executor. No-AI-attribution per repo policy.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #2565.
Every compute workflow now ships the homeserver-first routing that quality.yml's six families already used (#2362/#2363/#2389/#2397), so the box is the primary executor and GitHub-hosted is the automatic fallback when a heartbeat canary can't prove a
honeypot-cirunner live.What changes
.github/workflows/ci-router.yml(workflow_call) carrying the trust gate + heartbeat liveness proof quality.yml's inlineci-targetjob documents.containers.yml,security.yml(CodeQL) andpages.yml(build) consume it; quality.yml keeps its inline job until a follow-up migrates it. The router additionally treatsscheduleas a trusted event (CodeQL's weekly run).ci-heartbeat.ymldrops its concurrency group. One push now fans out to several routing workflows at once;cancel-in-progresswould cancel a sibling router's canary evidence mid-decision. The header documents why there is deliberately no group.containers.yml: singlebuildjob picksruns-onoff the router output; the matrix name carries the(GitHub-hosted)suffix on fallback days. 120-min per-row ceiling on the box, platform default on fallback.security.yml: same pattern for the CodeQLanalyzematrix (90-min ceiling on the box).pages.yml:buildroutes; thedeploystep staysubuntu-latestunconditionally (seconds-long API call against thegithub-pagesenvironment).quality.yml:frontend-nextandfrontend-next-browserbecome homeserver/cloud pairs like the existing families. The browser home twin drops--with-depsand its apt fallback in favour of loud presence checks — the chromium library set and theredis-serverbinary are preinstalled on the host, and a missing dep must fail the check, not silently apt-install.scripts-and-composerows are flaggedhome: truenow that the runner user drives docker: Keycloak realm import, oauth2-proxy gateway, both OIDC suites + the shared BFF build, vendored YARA, rev-eng corpus, compose validation, and the four real-Elasticsearch suites (those used to print "SKIP: docker daemon is not reachable" without an engine — flagging them under the old criterion would have relocated a no-op; ci(quality): extend homeserver routing to pure-python/shell scripts-and-compose entries #2389's audit note updated accordingly).scripts/github-ci-runner/install-ci-runner.shprovisions the host idempotently: docker-group grant for the runner user,redis-server(daemon disabled — only the binary is wanted), node 22 + npm, shellcheck, and playwright 1.62.1's chromium library set (extracted from the pinnedplaywright-core'subuntu26.04-x64key — re-extract when the pin moves).Caller obligation discovered on the first push
A called reusable workflow can never exceed the caller's GITHUB_TOKEN envelope, and under-granting it startup-fails the whole run as "Invalid workflow file" — the first push of this branch startup-failed Containers/CodeQL/pages because their envelopes lacked
actions: write, which the router's route job needs to dispatch the heartbeat canary. All three callers now grantactions: writeat workflow level (pages'deploykeeps its own narrower job-level envelope), andci-router.yml's header documents the obligation for future callers. Also registered thehoneypot-cirunner label in.github/actionlint.yaml(honeypot-homewas listed; its sibling was not).Box capability changes (applied and verified live)
github-ci-runneradded to thedockergroup (same grantgithub-deploy-runneralways had); runner service restarted to pick it up. The user still has no sudo.ldconfig.CI_HOMESERVER_PRS=trueset so same-repo PRs route; fork PRs still never reach the box (trust gate unchanged).Deliberately not moving (rationale in #2565)
dependency-review,dependabot-auto-merge(API-only bot workflows), the pages deploy step, and the ops workflows (deploy,diagnostics,vps-start-blackhole— deployment/production adjacency, not CI feedback).Validation
actionlintclean on all five workflows (also caught and fixed a duplicatename:key inscripts-and-composeduring the pass).bash -noninstall-ci-runner.sh;scripts/check-public-leaks.pypassed.main, so the conditional(GitHub-hosted)name suffixing can't break required checks.