Skip to content

ci: route all workflows homeserver-first with GitHub-hosted fallback - #2568

Merged
Xore merged 12 commits into
mainfrom
worktree-ci-all-homeserver-first
Aug 28, 2026
Merged

Xore merged 12 commits into
mainfrom
worktree-ci-all-homeserver-first

Conversation

@Xore

@Xore Xore commented Aug 27, 2026 •

Copy link
Copy Markdown
Owner

Closes #2565.

Every compute workflow now ships the homeserver-first routing that quality.yml's six families already used (#2362/#2363/#2389/#2397), so the box is the primary executor and GitHub-hosted is the automatic fallback when a heartbeat canary can't prove a honeypot-ci runner live.

What changes

  • New shared router .github/workflows/ci-router.yml (workflow_call) carrying the trust gate + heartbeat liveness proof quality.yml's inline ci-target job documents. containers.yml, security.yml (CodeQL) and pages.yml (build) consume it; quality.yml keeps its inline job until a follow-up migrates it. The router additionally treats schedule as a trusted event (CodeQL's weekly run).
  • ci-heartbeat.yml drops its concurrency group. One push now fans out to several routing workflows at once; cancel-in-progress would cancel a sibling router's canary evidence mid-decision. The header documents why there is deliberately no group.
  • containers.yml: single build job picks runs-on off the router output; the matrix name carries the (GitHub-hosted) suffix on fallback days. 120-min per-row ceiling on the box, platform default on fallback.
  • security.yml: same pattern for the CodeQL analyze matrix (90-min ceiling on the box).
  • pages.yml: build routes; the deploy step stays ubuntu-latest unconditionally (seconds-long API call against the github-pages environment).
  • quality.yml:
    • frontend-next and frontend-next-browser become homeserver/cloud pairs like the existing families. The browser home twin drops --with-deps and its apt fallback in favour of loud presence checks — the chromium library set and the redis-server binary are preinstalled on the host, and a missing dep must fail the check, not silently apt-install.
    • The docker-bound scripts-and-compose rows are flagged home: true now that the runner user drives docker: Keycloak realm import, oauth2-proxy gateway, both OIDC suites + the shared BFF build, vendored YARA, rev-eng corpus, compose validation, and the four real-Elasticsearch suites (those used to print "SKIP: docker daemon is not reachable" without an engine — flagging them under the old criterion would have relocated a no-op; ci(quality): extend homeserver routing to pure-python/shell scripts-and-compose entries #2389's audit note updated accordingly).
  • scripts/github-ci-runner/install-ci-runner.sh provisions the host idempotently: docker-group grant for the runner user, redis-server (daemon disabled — only the binary is wanted), node 22 + npm, shellcheck, and playwright 1.62.1's chromium library set (extracted from the pinned playwright-core's ubuntu26.04-x64 key — re-extract when the pin moves).

Caller obligation discovered on the first push

A called reusable workflow can never exceed the caller's GITHUB_TOKEN envelope, and under-granting it startup-fails the whole run as "Invalid workflow file" — the first push of this branch startup-failed Containers/CodeQL/pages because their envelopes lacked actions: write, which the router's route job needs to dispatch the heartbeat canary. All three callers now grant actions: write at workflow level (pages' deploy keeps its own narrower job-level envelope), and ci-router.yml's header documents the obligation for future callers. Also registered the honeypot-ci runner label in .github/actionlint.yaml (honeypot-home was listed; its sibling was not).

Box capability changes (applied and verified live)

  • github-ci-runner added to the docker group (same grant github-deploy-runner always had); runner service restarted to pick it up. The user still has no sudo.
  • Host provision installed per the script: redis-server 8.0.5 (daemon disabled), node v22.22.1/npm 9.2.0, shellcheck, chromium libs verified via ldconfig.
  • Repo variable CI_HOMESERVER_PRS=true set so same-repo PRs route; fork PRs still never reach the box (trust gate unchanged).

Deliberately not moving (rationale in #2565)

dependency-review, dependabot-auto-merge (API-only bot workflows), the pages deploy step, and the ops workflows (deploy, diagnostics, vps-start-blackhole — deployment/production adjacency, not CI feedback).

Validation

  • actionlint clean on all five workflows (also caught and fixed a duplicate name: key in scripts-and-compose during the pass).
  • YAML parse check: quality 20 jobs, containers 2, security 2, pages 3.
  • bash -n on install-ci-runner.sh; scripts/check-public-leaks.py passed.
  • No branch protection on main, so the conditional (GitHub-hosted) name suffixing can't break required checks.

Xore added 2 commits August 27, 2026 17:19
…2565)

Every compute workflow now ships the homeserver-first routing pair (or a
conditional runs-on) that quality.yml's six families already used:

- New shared .github/workflows/ci-router.yml (workflow_call) carrying the
  trust gate + heartbeat liveness proof quality.yml's inline router
  documents; containers/security/pages consume it, quality keeps its
  inline job until a follow-up migrates it.
- ci-heartbeat.yml drops its concurrency group: one push fans out to
  multiple routers and cancel-in-progress would cancel a sibling
  router's canary evidence.
- containers.yml build and security.yml (CodeQL) analyze pick runs-on
  off the router output; pages.yml build joins them and the deploy step
  stays ubuntu-latest (seconds-long API call).
- quality.yml: frontend-next and frontend-next-browser become pairs; the
  browser home twin drops --with-deps and the apt fallback for loud
  presence checks (deps are preinstalled on the host); the docker-bound
  scripts-and-compose rows are flagged home: true now that the runner
  user drives docker (the four real-ES suites included -- they used to
  SKIP without an engine, a relocated no-op).
- install-ci-runner.sh provisions the host idempotently: docker-group
  grant for the runner user, redis-server (daemon disabled), node 22,
  shellcheck, and playwright 1.62.1's chromium library set.

Runner capability changes verified live on the box; trust gate and
CI_HOMESERVER_PRS semantics unchanged. API-only and ops workflows stay
put (rationale in #2565).
A called reusable workflow can never exceed the caller's GITHUB_TOKEN
envelope, and under-granting it startup-fails the whole run as "Invalid
workflow file" -- containers.yml, security.yml and pages.yml granted
only their own scopes, so every run of the three startup-failed before
the router job could even appear. The route job needs actions:write
solely to dispatch the ci-heartbeat canary; pages' deploy keeps its own
narrower job-level envelope. Also register the honeypot-ci runner label
in .github/actionlint.yaml (honeypot-home was listed; its sibling was
not), silencing the runner-label lint noise the new runs-on lines add.
@github-actions

github-actions Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Dependency Review

The following issues were found:
  • ✅ 0 vulnerable package(s)
  • ✅ 0 package(s) with incompatible licenses
  • ✅ 0 package(s) with invalid SPDX license definitions
  • ⚠️ 1 package(s) with unknown licenses.
See the Details below.

License Issues

.github/workflows/security.yml

PackageVersionLicenseIssue Type
actions/setup-go5.*.*NullUnknown License

OpenSSF Scorecard

PackageVersionScoreDetails
actions/actions/setup-go 5.*.* 🟢 6.1
Details
CheckScoreReason
Binary-Artifacts🟢 10no binaries found in the repo
Dangerous-Workflow🟢 10no dangerous workflow patterns detected
Packaging⚠️ -1packaging workflow not detected
Code-Review🟢 10all changesets reviewed
Maintained🟢 56 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 5
CII-Best-Practices⚠️ 0no effort to earn an OpenSSF best practices badge detected
Token-Permissions⚠️ 0detected GitHub workflow tokens with excessive permissions
Pinned-Dependencies🟢 7dependency not pinned by hash detected -- score normalized to 7
Fuzzing⚠️ 0project is not fuzzed
License🟢 10license file detected
Signed-Releases⚠️ -1no releases found
Security-Policy🟢 9security policy file detected
Branch-Protection⚠️ 0branch protection not enabled on development/release branches
SAST🟢 10SAST tool is run on all commits

Scanned Files

  • .github/workflows/security.yml

Xore and others added 10 commits August 27, 2026 18:24
For pull_request runs GITHUB_REF_NAME is the ephemeral "<n>/merge"
merge ref, which no branch backs -- the canary dispatch was refused
with 422 "No ref found for" and every same-repo PR routed to the
GitHub-hosted fallback forever, silently (seen live on #2568: the
routers ran, trusted=true, then 'dispatch refused (HTTP 422X)' and
online=false). Same-repo PRs now dispatch at the PR's head branch,
whose ci-heartbeat.yml is the branch's own; fork PRs never reach the
dispatch (trusted=false). push/workflow_dispatch/schedule already
carry a real branch in GITHUB_REF_NAME. Also drop curl -f so a refused
dispatch lands its HTTP code in the warning instead of aborting the
capture (422X). Fixed in ci-router.yml and quality.yml's inline copy.
…lcheck

scripts/github-ci-runner/install-ci-runner.sh's provision list is inside
the "Shell syntax and high-severity ShellCheck" row's scan set, and a
comment line beginning with "shellcheck" parses as a (malformed)
shellcheck directive -- SC1073/SC1072 at --severity=error, failing the
row on #2568. Reword the provision-list entry and note the trap
in-place, without leading the note with the keyword either.
Follow-up wording accuracy after the merge from main: the ci-target
header still said the ci-heartbeat canary is dispatched at "THIS
commit" -- the dispatch API takes a ref, and for pull_request runs
GITHUB_REF_NAME is the "<n>/merge" merge ref that no branch backs
(422, fixed two commits ago).
Seen live on #2568's merged round: the canary dispatch was accepted
and the router still saw no run object in /runs 45s later -- GitHub's
run-indexing lag outran the appear window, so the router concluded
online=false and fell back with the box demonstrably healthy. Defaults
become appear=120s (must cover API indexing, not just dispatch
latency; the loop still exits the moment the run appears) and
decide=420s (the canary is tiny but queues behind whatever the single
box instance is executing). Router job timeout 10 -> 15min to keep the
worst case (120+420 + slack) comfortable. Fan-out deeper than ~2
canaries still overflows and falls back -- that is single-instance
capacity, tracked as #2572, not a window problem.
…ef edit

The dispatch-ref fix's old_string swallowed the cutoff= line and never
re-added it. Under set -u, Phase A's 'jq --arg cutoff "$cutoff"' then
expanded an unbound variable inside a pipeline subshell every poll:
the subshell died before jq ran, printf reported a broken pipe, rid
silently stayed empty, and every router concluded online=false while
its canary was created the same second as the dispatch and completed
success (live evidence on #2568's 18:00 round: 'cutoff: unbound
variable' repeating every ~10s of the appear window, canary run
created/started 18:00:54, verdict offline 18:03:04). This, not GitHub
lag or window size, is why no round ever routed homeserver-first.
Restored in both routers, with the failure mode documented on the
line; windows stay at 120s/420s (defensible on their own merits).
…ub tests pin env path

The Ghidra worker's _payload_roots() called is_dir() on three hardcoded paths without guarding against PermissionError. The CI runner user (github-ci) cannot traverse /var/lib/docker/volumes/dionaea-lib/_data/binaries, so a single stat() raised PermissionError, the whole drain() crashed mid-batch, and the spool was left inconsistent. The 'exit 0', 'missing sample quarantined', and 'second run is idempotent' assertions in test_spool all turned on the same drain state. Wrap each is_dir() in an OSError guard so a single inaccessible root is skipped, not fatal. update docstring to record why.

The two new analysis/github/tests scripts (test_publish_sample.sh, test_rejections.sh) carried the old refuse-guard: hard-fail if /etc/honeypot-github.env exists, on the grounds that sourcing the env in CI would leak GH_PAT. The runner that this PR routes the homeserver-first matrix to does carry that file (root, 0600) -- the file is what the production scripts source when they exist -- so the new tests were permanently red. The other four tests in the same suite (test_dry_run.sh, test_dry_run_reprocessed_after_enable.sh, test_daily_cap.sh) had already moved to pinning GITHUB_ANALYSIS_ENV_FILE to an empty fixture inside $work, which closes the same leak deterministically. Apply the same pattern to the two new tests.
Two follow-up fixes for the new homeserver-first/GitHub-hosted fallback routing:
- pages.yml: python -> python3 (the GitHub-hosted ubuntu runner has python3
  on PATH, not python; the homeserver runner had both, so this is a no-op
  there but unblocks the fallback path).
- security.yml CodeQL Go row: add actions/setup-go@v5 with go-version 1.23
  for the GitHub-hosted fallback only (matrix.language == go AND
  homeserver != true). The CodeQL Go extractor shells out to go to
  enumerate build targets; the homeserver runner has go preinstalled,
  the GitHub-hosted runner does not.

No-AI-attribution: per repo policy, no AI co-author or footer in this
  commit.
The homeserver runner also lacks go on PATH; the prior commit's
actions/setup-go step was conditional on homeserver != true, so the
homeserver (CodeQL default) path still missed the toolchain. Drop the
condition so the step runs for every Go matrix row on every executor.

No-AI-attribution per repo policy.
@Xore
Xore merged commit 8a3c35a into main Aug 28, 2026
122 checks passed
@Xore
Xore deleted the worktree-ci-all-homeserver-first branch August 28, 2026 10:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Route all CI homeserver-first with automatic GitHub-hosted fallback

1 participant