Skip to content

ci(3501): arm the base-image CVE gate and move the pins that fix it #4586

ci(3501): arm the base-image CVE gate and move the pins that fix it

ci(3501): arm the base-image CVE gate and move the pins that fix it #4586

Workflow file for this run

name: Quality
on:
push:
branches: [main]
pull_request:
workflow_dispatch:
inputs:
force_github_hosted:
description: >-
Send the homeserver-eligible jobs to GitHub-hosted runners even
when the honeypot-ci runner is online -- e.g. while servicing the
homeserver, or when a suspect result needs ruling out as
runner-specific. Homeserver-first routing itself is automatic for
trusted runs (push-to-main, workflow_dispatch); pull_request stays
GitHub-hosted unless the repository variable CI_HOMESERVER_PRS
opts same-repo PRs in.
type: boolean
default: false
# #3313: workflow level defaults to none; every job below declares what it
# spends. Each of the 20 testing jobs runs actions/checkout, so each carries
# its own contents: read. actions: write, which used to sit here and so
# reached all 20 of them, is now on the ci-target router job alone -- that one
# grant is what dispatches the ci-heartbeat canary with GITHUB_TOKEN, and a
# called reusable workflow can never exceed the caller's envelope, so
# under-granting it startup-fails the whole run as "Invalid workflow file"
# (containers.yml, security.yml and pages.yml grant the same way for the same
# reusable ci-router.yml).
permissions: {}
concurrency:
# #2753: this used to be group: quality-${{ github.ref }} with
# cancel-in-progress off for push-to-main, on the theory that turning
# cancellation off makes main pushes serialize -- wait for the previous
# run before starting. That is not what a GitHub Actions concurrency group
# does: with cancel-in-progress:false, GitHub keeps only the currently
# running run and the single most recent pending one per group, and
# cancels -- with ZERO jobs ever started -- every push queued in between.
# Reproduced live 2026-08-31: ten consecutive main pushes (including the
# #2724/#2725 merge commits) got a Quality run that was `cancelled` with
# `.jobs | length == 0`, not superseded by a later run that re-verified
# the same tree; only the first and last commit in the batch were ever
# actually checked. That is exactly the #2317/#2347 failure class this
# config exists to prevent (cbb5ef3f landed non-compiling Rust and its own
# quality run never finished) -- the shared per-ref group didn't close
# that hole, it just changed which commits fall through it.
#
# Fix: give every push-to-main commit its own group (keyed on github.sha
# instead of github.ref), so no two main commits can ever share a group
# and cancel one another. This trades true FIFO serialization -- which
# GitHub's concurrency primitive cannot express at all -- for the
# guarantee #2317/#2347 actually needed: every commit that reaches `main`
# gets its own Quality run and none of them get cancelled by a sibling.
# pull_request runs keep the previous shared-per-ref/cancel-on behavior
# unchanged -- a superseded PR run is genuinely for a commit nobody cares
# about anymore.
group: quality-${{ startsWith(github.ref, 'refs/heads/main') && github.sha || github.ref }}
cancel-in-progress: ${{ !startsWith(github.ref, 'refs/heads/main') }}
jobs:
# -------------------------------------------------------------------------
# Executor routing ("homeserver first, GitHub-hosted fallback")
#
# ci-target's trust-gate + heartbeat-liveness decision procedure is
# shared with every other homeserver-eligible workflow via the reusable
# .github/workflows/ci-router.yml -- its header carries the full
# rationale (also documented in docs/CI-CD.md's "Executor routing"
# section), so it isn't repeated here. Two things stay specific to
# Quality:
#
# - force_github_hosted (the workflow_dispatch input above) forces the
# fallback direction manually -- e.g. while servicing the
# homeserver, or when a suspect result needs ruling out as
# runner-specific.
# - Unlike containers.yml/security.yml/pages.yml's single
# executor-agnostic job, every home-executable check below ships as
# a PAIR of conditional jobs that take turns off ci-target's answer
# -- exactly one twin does real work per run, the other reports
# skipped.
#
# #2565: the no-docker CI-runner design is retired. The runner's
# dedicated user now carries a docker-group membership (the same grant
# github-deploy-runner always had) and the host ships node 22, the
# redis-server binary, shellcheck and the playwright 1.62.1 chromium
# library set for its user -- preinstalled, because the user still has
# no sudo by design and every sudo-apt path in a check would otherwise
# be a guaranteed relocation failure. Everything Quality runs is
# therefore homeserver-eligible: frontend-next and frontend-next-browser
# ship as pairs like the families above (the browser twin swaps
# `--with-deps` and the apt install for presence checks -- deps live on
# the host now, and a missing one must fail loudly, not silently
# apt-install), and every docker-bound scripts-and-compose row is
# flagged `home: true` (#2389's docker-less audit criterion graduated
# to "docker or no engine needed").
# -------------------------------------------------------------------------
ci-target:
name: Pick CI executor
uses: ./.github/workflows/ci-router.yml
permissions:
contents: read
actions: write
with:
ci_homeserver_prs: ${{ vars.CI_HOMESERVER_PRS || '' }}
force_github_hosted: ${{ inputs.force_github_hosted || '' }}
# Naming rule for every pair below: the homeserver twin carries the
# canonical check name (it is what normally runs); its GitHub-hosted
# fallback twin appends "(GitHub-hosted)" so a degraded day reads
# honestly in the checks list instead of masquerading as business as
# usual.
#
# timeout-minutes exists ONLY on homeserver twins (plus the router): a
# stuck pickup wedges a real box nobody reboots promptly, whereas
# GitHub-hosted runners have hard platform timeouts already. Values sit
# far above each check's observed runtime so they fire only on hangs.
public-safety:
name: Public repository safety
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 10
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: "3.13"
- run: python scripts/check-public-leaks.py
- run: python scripts/validate-oidc-redirects.py
- run: bash scripts/check-ghosts-vendored-egress.sh
public-safety-cloud:
name: Public repository safety (GitHub-hosted)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: "3.13"
- run: python scripts/check-public-leaks.py
- run: python scripts/validate-oidc-redirects.py
- run: bash scripts/check-ghosts-vendored-egress.sh
design-lab-readonly:
# #1828: the design lab serves variants against the real captured-data
# Elasticsearch, so its read-only guarantee is a safety property, not a
# convenience. The test drives the actual harness against a recording
# stand-in backend and fails if a write ever reaches it.
name: Design lab is read-only
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 10
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: "22"
- run: node --test branding/design-lab/lab.test.mjs
design-lab-readonly-cloud:
name: Design lab is read-only (GitHub-hosted)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: "22"
- run: node --test branding/design-lab/lab.test.mjs
go-fmt:
name: Go formatting
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 15
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
# dicompot's vendored github.com/nsmfoo/dicompot (#413) declares go
# 1.26.2 in its own go.mod -- the module graph forces that as the
# floor for every module in this repo, not just dicompot's.
go-version: "1.26.x"
# The homeserver runners keep GOMODCACHE/GOCACHE on disk at
# /opt/github-ci-runner (shared by all four runner instances), so
# setup-go's default module cache is pure loss here -- and worse
# than loss. Go writes module directories mode 0555, so the
# restore's `tar -x` cannot recreate a single already-present
# file: every job downloaded ~466 MB, spent ~45s failing to
# unpack it ("Cannot open: File exists" x thousands, tar exit 2),
# reported "Cache is not found", then re-uploaded the whole
# module cache from the post step. Six same-key 465 MB copies
# piled up in one evening and pushed the repo's 10 GB Actions
# cache over quota, evicting the buildkit layer caches. Same
# reasoning as the Rust job below, which carries no
# actions/cache block for exactly this reason.
cache: false
- name: Check formatting
shell: bash
run: |
files="$(find . -path '*/vendor' -prune -o -name '*.go' -type f -print)"
test -z "$(gofmt -l $files)" || { gofmt -l $files; exit 1; }
go-fmt-cloud:
name: Go formatting (GitHub-hosted)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
# Same floor as the homeserver twin above -- see its comment.
go-version: "1.26.x"
cache-dependency-path: "**/go.sum"
- name: Check formatting
shell: bash
run: |
files="$(find . -path '*/vendor' -prune -o -name '*.go' -type f -print)"
test -z "$(gofmt -l $files)" || { gofmt -l $files; exit 1; }
# Tests stay a parallel matrix on GitHub-hosted (the shape that has always
# run there), but consolidate into ONE sequential loop on the homeserver:
# 20 matrix entries would queue behind a single runner machine anyway,
# paying checkout+setup twenty times over for zero parallelism gained --
# while a warm GOCACHE in the runner user's persistent HOME makes the
# sequential loop cheaper than the sum of its parts after the first run.
# Quality-homeserver.yml ran this exact loop shape successfully until it
# was superseded by these pairs.
go-test-homeserver:
name: Test (all modules)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 60
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
go-version: "1.26.x"
# The homeserver runners keep GOMODCACHE/GOCACHE on disk at
# /opt/github-ci-runner (shared by all four runner instances), so
# setup-go's default module cache is pure loss here -- and worse
# than loss. Go writes module directories mode 0555, so the
# restore's `tar -x` cannot recreate a single already-present
# file: every job downloaded ~466 MB, spent ~45s failing to
# unpack it ("Cannot open: File exists" x thousands, tar exit 2),
# reported "Cache is not found", then re-uploaded the whole
# module cache from the post step. Six same-key 465 MB copies
# piled up in one evening and pushed the repo's 10 GB Actions
# cache over quota, evicting the buildkit layer caches. Same
# reasoning as the Rust job below, which carries no
# actions/cache block for exactly this reason.
cache: false
- name: Test every Go module
shell: bash
run: |
while IFS= read -r module; do
echo "::group::${module%/go.mod}"
(cd "${module%/go.mod}" && go test ./...)
echo "::endgroup::"
done < <(find . -path '*/vendor' -prune -o -name go.mod -type f -print | sort)
go-test-cloud:
name: Test (${{ matrix.module }})
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
# Every go.mod in the repo (excluding vendor/) -- kept as an explicit
# list rather than a dynamic discover-then-matrix job so a new
# module's tests actually run in CI the moment its go.mod lands,
# instead of silently needing a separate PR to register it here too.
# `find ... | sort` in this job's own history is the source of
# truth if this list and the repo ever drift.
module:
- arcane/home/honeypot-attacker-identity-worker/attacker-identity-worker
- arcane/home/honeypot-canarytokens/canarytokens-adapter
- arcane/home/honeypot-canarytokens/canarytokens-http-router
- arcane/home/honeypot-cisco-asa-honeypot/cisco-asa-honeypot
- arcane/home/honeypot-citrix-honeypot/citrix-honeypot
- arcane/home/honeypot-correlator-worker/correlator-worker
- arcane/home/honeypot-cowrie/honeyfs-implant
- arcane/home/honeypot-dicompot/dicompot
- arcane/home/honeypot-dionaea/tftp-relay
- arcane/home/honeypot-dnp3/dnp3-honeypot
- arcane/home/honeypot-dns-honeypot/dns-honeypot
- arcane/home/honeypot-endlessh/endlessh-honeypot
- arcane/home/honeypot-galah/galah-llm-broker
- arcane/home/honeypot-http/http-honeypot
- arcane/home/honeypot-multipot/multipot
- arcane/home/honeypot-payload-inventory-worker/payload-inventory-worker
- arcane/home/honeypot-rdp-honeypot/rdp-honeypot
- arcane/home/honeypot-sonicwall-sma/sonicwall-sma-honeypot
- arcane/home/honeypot-utilities/reporter
- portbridge
- vps/portbridge
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
go-version: "1.26.x"
cache-dependency-path: "${{ matrix.module }}/go.sum"
- run: go test ./...
working-directory: ${{ matrix.module }}
go-modules-complete:
name: Go formatting and tests
if: always()
needs: [go-fmt, go-fmt-cloud, go-test-homeserver, go-test-cloud]
runs-on: ubuntu-latest
permissions:
contents: read
steps:
# #3319: only the lane-summary script, same sparse shape ai-attribution
# uses. The report is written to the step summary and is a pure function
# of `needs`, so nothing else from the tree is needed here.
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
sparse-checkout: scripts/ci-lane-summary.py
sparse-checkout-cone-mode: false
- name: Fail unless each executor pair produced exactly one success
env:
FMT_HS: ${{ needs.go-fmt.result }}
FMT_CLOUD: ${{ needs.go-fmt-cloud.result }}
TEST_HS: ${{ needs.go-test-homeserver.result }}
TEST_CLOUD: ${{ needs.go-test-cloud.result }}
shell: bash
run: |
# Each pair is exclusive by construction (mutually exclusive `if:`),
# so the healthy states are: primary succeeded, or primary was
# skipped because routing chose fallback AND fallback succeeded.
# Anything else -- both skipped, both ran, a failure, a cancellation
# mid-flight on main-bound code -- fails loudly right here rather
# than letting a silently-skipped tier read as green.
fail=0
expect() {
local label="$1" primary="$2" fallback="$3"
if [[ "$primary" == "success" ]]; then return 0; fi
if [[ "$primary" == "skipped" && "$fallback" == "success" ]]; then return 0; fi
echo "::error::${label}: expected success or skip+fallback-success; got ${primary} / ${fallback}"
return 1
}
expect "go formatting" "$FMT_HS" "$FMT_CLOUD" || fail=1
expect "go tests (matrix)" "$TEST_HS" "$TEST_CLOUD" || fail=1
exit "$fail"
# #3319: the same pair results as a table, so a reader can see which
# executor ran and which twin the router skipped without opening the
# log. Independent of the exit code above -- this reports, it does not
# gate, so it still renders on a red aggregate job.
- name: Lane summary
env:
NEEDS: ${{ toJSON(needs) }}
run: python3 scripts/ci-lane-summary.py --title "Go formatting and tests" <<<"$NEEDS"
# Homeserver-first as a pair (#2565): the tier's only special
# requirement is docker (lockfile check) plus node 24 via setup-node --
# the runner user now carries the docker-group membership, and
# setup-node works identically on the self-hosted runner (its
# node/npm caches simply persist in the runner's _work/_tool and HOME
# instead of the platform cache). Steps stay byte-identical across the
# twins so a red result means the same thing wherever it ran;
# timeout-minutes rides only on the homeserver twin per the pair
# convention.
frontend-next:
name: Dashboard frontend (next)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 45
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
# #3331: run on the node major the image actually ships, read out of the
# image's own FROM line instead of written here a second time.
#
# Every step in this job used to run under setup-node's node-version "24"
# while the Dockerfile builds and serves on node:22-alpine, so the
# typecheck, the unit tests, the production build and the generated
# route-tree diff had never once executed on the node that ships the
# artifact. The only Node 22 coverage in the whole workflow was the
# #1816 lockfile-install step below. A gate that cannot see the runtime it
# is gating is not a gate.
#
# Measured on the Dockerfile's own image before moving the pin
# (node:22-alpine@sha256:c610fcdf, v22.23.2, npm 10.9.8): `npm ci`,
# `npm run typecheck`, 179/179 vitest cases, `npm run build`, and the
# `git diff --exit-code` route-tree check are all clean there. The #2034
# browser matrix is 58/58 on node:22 as well, so this moves coverage onto
# the runtime rather than trading a gap for a break.
#
# Deriving instead of re-declaring is the part that stops it recurring. A
# second copy of the version is a second thing to forget, and forgetting
# it is precisely how the 24/22 split opened up; reading it from the
# Dockerfile means an image bump moves CI with it in the same commit.
- name: Node major from the image (#3331)
id: node-runtime
run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: ${{ steps.node-runtime.outputs.node-version }}
cache: npm
cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json
# #1816: the lockfile must install under the npm the *image* uses.
#
# Dependabot resolves with its own npm again, so a group bump can
# produce a lockfile that satisfies one and not the other -- and the
# failure then lands on whatever PR runs next, which is usually one
# that touches no frontend file at all. Three times so far, each
# costing a diagnosis before a one-line fix.
#
# Running the image's npm here makes it fail on the PR that caused
# it. Regenerating automatically would also work; failing loudly in
# the right place is better than a bot quietly rewriting a lockfile.
#
# #3331: the image is the derived `node-image` output, not a second
# literal `node:22-alpine`. This step is the one place that already ran
# the runtime's npm, so leaving the tag written out by hand would have
# kept the drift alive on the exact axis this issue is about. The script
# also refuses to resolve at all if the Dockerfile's stages ever split
# across node majors, which would leave "the image's npm" ambiguous.
#
# --user matters: without it the container installs as root and
# leaves a root-owned node_modules the runner's own `npm ci` below
# cannot remove (EACCES on rmdir .bin), so the check would break the
# very job it is meant to protect. HOME and the cache go somewhere
# writable for that user for the same reason.
- name: lockfile installs under the image's npm
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: |
docker run --rm \
--user "$(id -u):$(id -g)" \
-e HOME=/tmp -e npm_config_cache=/tmp/npm-cache \
-v "$PWD:/app" -w /app \
${{ steps.node-runtime.outputs.node-image }} npm ci --no-audit --no-fund
rm -rf node_modules
- run: npm ci
working-directory: arcane/home/honeypot-dashboard/frontend-next
# #2180 / #2804: there was a "TanStack minors diverge inside one
# install" step here. It warned whenever the 1.x @tanstack/react-*
# packages in package-lock.json spanned more than one minor. It has
# been RETIRED, not narrowed, because the condition it asserts cannot
# be satisfied by any installable version of this family and therefore
# never carried information -- it fired on every run since #2180 added
# it. Measured 2026-09-02, recorded here so the next reader does not
# rediscover it a fourth time:
#
# 1. @tanstack/react-start pins its siblings as EXACT versions
# spanning three minors. `npm view @tanstack/react-start@latest
# dependencies` on 1.168.49 (which IS latest) gives
# react-router 1.170.32, react-start-client 1.168.30,
# react-start-server 1.167.37 -- 1.170 / 1.168 / 1.167. Installing
# react-start fixes the rest outright, so package.json has no
# lever over them and no bump converges the set.
# 2. Narrowing the filter to the packages package.json declares
# directly (react-router, react-start, router-cli) yields the
# IDENTICAL minor set, 1.167/1.168/1.170 -- router-cli is
# independently versioned and has never published past 1.167.
# Verified against the real lockfile; it is a no-op, not a fix.
# 3. Replacing the assertion with "the family moved together in one
# PR" does not work either, in both directions. Over the
# exact-pinned closure it can never fire: npm's own resolution
# already makes react-start-client/-server move atomically with
# react-start, so the check would assert something npm
# guarantees. Widen it to include router-cli and it always fires,
# since router-cli moves on its own cadence. Separately,
# .github/dependabot.yml groups every frontend minor/patch update
# into one `frontend-compatible` PR, so "moved together in one
# PR" is already true by construction here.
#
# What #2180 actually wanted -- that upgrades be deliberate -- is
# carried by the spec-exact pins in package.json (#2208), not by this
# comparison. Asserting THOSE stay exact is a real, satisfiable check
# and is the one piece of coverage this retirement drops; tracked in
# #2867 rather than swapped in here unreviewed. The lockfile carries
# no duplicate @tanstack copies today (every package appears exactly
# once), so nothing is diverging inside one install in the literal
# sense either.
#
# #2867: this is the check that retirement left uncovered. #2208
# pinned the three @tanstack/* deps in package.json spec-exact so
# upgrades are deliberate (#2180's actual goal); nothing asserted that
# stayed true. Needs no network and no lockfile resolution -- a jq
# test over package.json alone -- so it runs before install-dependent
# steps and fails (not warns) since a caret creeping back in is a real,
# fixable regression.
- name: TanStack deps stay spec-exact (#2180)
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: |
loose="$(jq -r '((.dependencies // {}) + (.devDependencies // {}))
| to_entries[]
| select(.key | startswith("@tanstack/"))
| select(.value | test("^[0-9]+\\.[0-9]+\\.[0-9]+$") | not)
| "\(.key) \(.value)"' package.json)"
if [ -n "$loose" ]; then
echo "::error::@tanstack deps must be pinned spec-exact (#2180/#2208):"
echo "$loose"
exit 1
fi
- run: npm run typecheck
working-directory: arcane/home/honeypot-dashboard/frontend-next
# #1831: the tier's first behavioural check. typecheck and build were
# the only things running against it, and neither can see ordering,
# a pre-hydration string literal, or anything else whose correctness
# is a matter of what happens rather than of what shape it has.
# #3319: CI_ARTIFACTS_DIR is what switches the vitest config's JUnit
# reporter on (see its own comment). It is the ONLY thing that does, so
# the same `npm test` a developer runs keeps vitest's console output and
# writes no report file.
- run: npm test
working-directory: arcane/home/honeypot-dashboard/frontend-next
env:
CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts
# #3319: `always()`, not `failure()` -- the JUnit result is worth having
# on a green run too, since it is the record of what actually ran rather
# than of what the exit code happened to be. 7 days matches the issue's
# retention ask: long enough to still be there when a red run is
# investigated the following week, short enough that a busy main does
# not accumulate them indefinitely.
#
# `overwrite: true` is gone from every upload in this file, and the
# artifact name now ends in the run's own identity. It was here from
# #3319, on the reading that a name which may already exist needs v4's
# conflict policy. What `overwrite: true` actually does is DELETE an
# existing artifact of that name before uploading, so on a name two
# runs share it is not "the newest run's results win" -- it is "the
# second uploader destroys the first's evidence", and the name then
# resolves to whichever run got there last, whose contents the earlier
# run never verified. A per-(run, attempt) name cannot collide, so there
# is nothing left to overwrite and no name an earlier run on the same
# ref could poison.
#
# run_attempt is in the name as well as run_id because a GitHub re-run
# of a run keeps its run_id: run_id alone still collides on the second
# attempt of the same run, which is the collision #3319 was really
# hitting. With both, a name is unique per (run, attempt), and the step
# that writes it runs once per (run, attempt) -- those two conditions
# together are what make upload-artifact v4's immutability mean
# something here. Every step below sits in a job that either has no
# twin or has one gated on the opposite answer from ci-target, so no two
# steps in a run can claim the same name. Nothing downstream reads these
# by name (there is no actions/download-artifact in .github/), so the
# per-run suffix costs no consumer.
- name: Upload unit test results
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-unit-junit-${{ github.run_id }}-${{ github.run_attempt }}
# The file is named explicitly rather than uploading its directory:
# upload-artifact v4.4+ skips hidden paths unless told otherwise, and
# `.ci-artifacts/` is hidden, so `path: .ci-artifacts/` uploads
# NOTHING (it reports "No files were found" and, with
# if-no-files-found: ignore, says so quietly). Naming the file is
# also what makes the contents a closed set: only the report the
# lane's own reporter wrote can ever land here.
path: .ci-artifacts/frontend-next-unit-junit.xml
if-no-files-found: warn
retention-days: 7
# #3318: this tier collected no coverage at all, so there was no
# baseline and nothing here could catch a change that deleted tested
# behaviour. The tests that remained would keep passing over code that
# had stopped being tested, and the job would stay green -- the tests
# that assert the deleted behaviour would be the thing a reviewer has
# to notice is missing.
#
# Two gates, both compared against the committed
# arcane/home/honeypot-dashboard/frontend-next/coverage-baseline.json
# and both computed from this run's own measurement. Neither is a
# threshold somebody chose:
#
# coverage:ratchet the non-regression floor. The baseline is what the
# suite measures today -- 806/7359 lines (10.95%)
# and 346/7250 branches (4.77%) across 127 files of
# src/ -- measured on this job's own #3331 node:22
# image. It is a floor, not a target, and the number
# in the file is a measurement rather than an
# aspiration on purpose: an importable 60% would
# have made this gate red on the day it landed and
# bought nothing. Raising it is a separate,
# deliberate commit.
# test:discovery a test-shaped file that no configured runner
# collects. Its own step so a coverage failure
# cannot hide a discovery failure.
#
# `npm test` above is untouched and stays uninstrumented: a second full
# run of the suite is the honest price of keeping the command deploy.yml,
# the README and a developer's own loop free of coverage, and it is
# cheap here -- 6.9s uninstrumented against 6.9s with the v8 provider on
# the image this job runs (the cost is transform/import, not
# instrumentation), against a 45-minute budget.
#
# `set -euo pipefail` so a failing coverage run cannot be followed by a
# ratchet that then reports on a missing report and takes the blame for
# it.
- name: "Coverage report and ratchet (#3318)"
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: |
set -euo pipefail
npm run test:coverage
npm run coverage:ratchet
- name: "Test discovery guard (#3318)"
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: npm run test:discovery
# Same per-(run, attempt) name as the JUnit upload above, for the same
# reason: run_id alone still collides on a re-run of the same run, and a
# name two uploads share is not "newest wins" under v4 -- it is the
# second uploader deleting the first one's evidence. No overwrite here
# either, and the retention matches.
#
# `coverage/` is named as a directory rather than file-by-file because
# it is not hidden -- the .ci-artifacts caveat that forces the JUnit
# upload to name its file does not apply, and a directory keeps the
# report extensible (adding the lcov html view later is a config change
# here, not a workflow change). The contents stay a closed set: only the
# files the coverage reporter wrote, from the run above, can land here.
- name: Upload unit test coverage
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-coverage-${{ github.run_id }}-${{ github.run_attempt }}
path: arcane/home/honeypot-dashboard/frontend-next/coverage/
if-no-files-found: warn
retention-days: 7
# No live-ES smoke suite here on purpose: port-tests/ needs the real
# cluster over an SSH tunnel to the homeserver (see its README) --
# not reachable from a GitHub-hosted runner, and not appropriate to
# point at production data from CI. That suite stays a manual/local
# verification step; this job's job is catching build/type breakage.
- run: npm run build
working-directory: arcane/home/honeypot-dashboard/frontend-next
# Checked post-build, not via the standalone `tsr generate` CLI: the
# tanstackStart() vite plugin's own route-tree generation (used by
# `build`/`dev`) and the router-cli's standalone `generate-routes`
# disagree on one thing -- the plugin also emits a `Register` SSR
# type augmentation the CLI doesn't -- so diffing right after
# `generate-routes` flags the committed (plugin-shaped) file as
# stale on every single run. The build's own output is what's
# actually committed and actually ships; diff against that instead.
- name: Generated route tree is current
run: git diff --exit-code -- arcane/home/honeypot-dashboard/frontend-next/src/routeTree.gen.ts
frontend-next-cloud:
name: Dashboard frontend (next) (GitHub-hosted)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
# #3331: twin of the homeserver copy -- same derive, same reason.
# (Full rationale, and the node:22 measurement behind it, lives there.)
- name: Node major from the image (#3331)
id: node-runtime
run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: ${{ steps.node-runtime.outputs.node-version }}
cache: npm
cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json
# #1816: the lockfile must install under the npm the *image* uses.
# (Full rationale lives on the homeserver twin above.)
- name: lockfile installs under the image's npm
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: |
docker run --rm \
--user "$(id -u):$(id -g)" \
-e HOME=/tmp -e npm_config_cache=/tmp/npm-cache \
-v "$PWD:/app" -w /app \
${{ steps.node-runtime.outputs.node-image }} npm ci --no-audit --no-fund
rm -rf node_modules
- run: npm ci
working-directory: arcane/home/honeypot-dashboard/frontend-next
# #2180 / #2804: the twin of the retired "TanStack minors diverge
# inside one install" step lived here. Retired for the same reason --
# see the full rationale on the homeserver copy above.
#
# #2867: twin of the homeserver copy's spec-exact check above.
- name: TanStack deps stay spec-exact (#2180)
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: |
loose="$(jq -r '((.dependencies // {}) + (.devDependencies // {}))
| to_entries[]
| select(.key | startswith("@tanstack/"))
| select(.value | test("^[0-9]+\\.[0-9]+\\.[0-9]+$") | not)
| "\(.key) \(.value)"' package.json)"
if [ -n "$loose" ]; then
echo "::error::@tanstack deps must be pinned spec-exact (#2180/#2208):"
echo "$loose"
exit 1
fi
- run: npm run typecheck
working-directory: arcane/home/honeypot-dashboard/frontend-next
# #3319: twin of the homeserver copy's JUnit wiring above.
- run: npm test
working-directory: arcane/home/honeypot-dashboard/frontend-next
env:
CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts
# Same per-(run, attempt) artifact name and the same explicit-file path
# as the homeserver twin -- see that step for why `overwrite: true` is
# gone and why the path is not the `.ci-artifacts/` directory.
- name: Upload unit test results
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-unit-junit-${{ github.run_id }}-${{ github.run_attempt }}
path: .ci-artifacts/frontend-next-unit-junit.xml
if-no-files-found: warn
retention-days: 7
# #3318: twin of the homeserver copy's coverage + ratchet + discovery
# steps. Present here for the reason the header comment gives: steps stay
# byte-identical across the twins so a red result means the same thing
# wherever it ran. A ratchet that only ran on the self-hosted executor
# would let a regression through on every degraded day, which is exactly
# the day the fallback twin exists for. (Full rationale, and the measured
# baseline, live on the homeserver copy above.)
- name: "Coverage report and ratchet (#3318)"
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: |
set -euo pipefail
npm run test:coverage
npm run coverage:ratchet
- name: "Test discovery guard (#3318)"
working-directory: arcane/home/honeypot-dashboard/frontend-next
run: npm run test:discovery
- name: Upload unit test coverage
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-coverage-${{ github.run_id }}-${{ github.run_attempt }}
path: arcane/home/honeypot-dashboard/frontend-next/coverage/
if-no-files-found: warn
retention-days: 7
# No live-ES smoke suite here on purpose -- see the homeserver twin.
- run: npm run build
working-directory: arcane/home/honeypot-dashboard/frontend-next
- name: Generated route tree is current
run: git diff --exit-code -- arcane/home/honeypot-dashboard/frontend-next/src/routeTree.gen.ts
# #2034: the browser-level acceptance net returned. The Go tier ran a
# 90-case Playwright matrix (#60, PR #146) until the cutover deleted it;
# nothing replaced it and visual/behavioural regressions have had no net
# since. This is its deliberately slimmed port -- theme x viewport shell
# smoke over every sidebar route (from lib/nav.ts), the modal core, and
# role-aware action visibility -- running against the BUILT production
# server output with hermetic fixtures (e2e/start-dashboard.mjs), not the
# dev server.
#
# Homeserver-first as a pair (#2565): the box preinstalls the playwright
# chromium library set (extracted for the pinned playwright-core's
# ubuntu26.04-x64 key) and the redis-server binary for the runner user,
# who has no sudo by design -- so the homeserver twin drops --with-deps
# and the apt fallback and, instead of guessing, FAILS LOUDLY when the
# host provision is missing. Consequence to keep in mind on a playwright
# bump: if the new build needs a library the host list doesn't cover,
# this twin fails at chromium launch with the missing-lib message and
# the host list (not this file) is what needs the update.
frontend-next-browser:
name: Dashboard-next browser matrix
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 45
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
# #3331: same derive as the two jobs above, and it matters most here --
# this twin runs the BUILT output, so it is the closest thing CI has to
# executing what the container serves, and it was doing that on 24 while
# the container serves 22. 58/58 of this matrix is green on node:22
# (node:22-bookworm, v22.23.2), measured before the pin moved.
- name: Node major from the image (#3331)
id: node-runtime
run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: ${{ steps.node-runtime.outputs.node-version }}
cache: npm
cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json
- run: npm ci
working-directory: arcane/home/honeypot-dashboard/frontend-next
- run: npm run build
working-directory: arcane/home/honeypot-dashboard/frontend-next
# No --with-deps: that path apt-installs, and the runner user has no
# sudo by design. The chromium library set lives on the host (#2565)
# and only the browser binary itself downloads here (cached in the
# runner's persistent ~/.cache/ms-playwright).
- run: npx playwright install chromium
working-directory: arcane/home/honeypot-dashboard/frontend-next
# The session fixture spawns a real redis-server rather than
# reimplementing RESP (#1034's tradeoff). The runner user cannot
# apt-install it, so a missing binary fails here, loudly, instead of
# as an obscure fixture error three steps later.
- run: |
command -v redis-server >/dev/null 2>&1 || {
echo "::error::redis-server missing on the runner host -- see #2565's homeserver provision list"
exit 1
}
# #3319: CI_ARTIFACTS_DIR turns on the config's JUnit reporter. The
# trace/screenshots half of the issue needs no flag -- playwright.config
# already carries trace: "retain-on-failure" and screenshot:
# "only-on-failure" (#2034); what was missing was uploading them, so a
# failing browser case was readable only as a log line.
- run: npm run test:browser
working-directory: arcane/home/honeypot-dashboard/frontend-next
env:
CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts
# `always()`: the JUnit XML is the machine-readable answer to "which
# cases ran, which were skipped, which failed" and is worth keeping on a
# green run too. Both paths are `if-no-files-found: warn` rather than
# error so a lane that never produced a report -- because it failed at
# `npm ci`, say -- reports the absence instead of masking the real
# failure behind an upload error.
#
# Per-(run, attempt) name, explicit file path, no `overwrite: true` --
# all three for the reasons spelled out on the frontend-next twin's
# upload above.
- name: Upload browser test results
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-browser-junit-${{ github.run_id }}-${{ github.run_attempt }}
path: .ci-artifacts/playwright-junit.xml
if-no-files-found: warn
retention-days: 7
# `failure()`, not `always()`: the HTML report bundles a trace per
# failing test and is the one artifact big enough to matter. On a green
# run it holds nothing worth storing.
- name: Upload Playwright report and traces
if: failure()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-browser-report-${{ github.run_id }}-${{ github.run_attempt }}
# Both directories are named, not the frontend-next tree around
# them. playwright.config.ts's own outputDir is ./test-results and
# the html reporter's is ./playwright-report, so this lane creates
# and owns both in a fresh checkout: the contents are what the
# browser matrix wrote, not whatever happens to be under
# frontend-next/ right now -- which is the property that keeps a
# stray .env or key out of a 7-day artifact anyone with repo read
# access can pull.
path: |
arcane/home/honeypot-dashboard/frontend-next/playwright-report/
arcane/home/honeypot-dashboard/frontend-next/test-results/
if-no-files-found: warn
retention-days: 7
frontend-next-browser-cloud:
name: Dashboard-next browser matrix (GitHub-hosted)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
# #3331: twin -- see the homeserver copy above for the rationale.
- name: Node major from the image (#3331)
id: node-runtime
run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: ${{ steps.node-runtime.outputs.node-version }}
cache: npm
cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json
- run: npm ci
working-directory: arcane/home/honeypot-dashboard/frontend-next
- run: npm run build
working-directory: arcane/home/honeypot-dashboard/frontend-next
- run: npx playwright install --with-deps chromium
working-directory: arcane/home/honeypot-dashboard/frontend-next
# The session fixture spawns a real redis-server rather than
# reimplementing RESP (#1034's tradeoff); ubuntu-latest usually ships
# it, but fail loudly into install rather than obscurely later.
# Grouped braces: ungrouped `A || B && C` parses as `(A || B) && C`
# (#2224), which apt-installs on every run. Same idiom as the
# shellcheck guard below.
- run: command -v redis-server >/dev/null 2>&1 || { sudo apt-get update && sudo apt-get install -y redis-server; }
# #3319: twin of the homeserver copy's artifact wiring above.
- run: npm run test:browser
working-directory: arcane/home/honeypot-dashboard/frontend-next
env:
CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts
# Twin of the homeserver copy's two uploads: same per-(run, attempt)
# names, same explicit paths, no `overwrite: true`.
- name: Upload browser test results
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-browser-junit-${{ github.run_id }}-${{ github.run_attempt }}
path: .ci-artifacts/playwright-junit.xml
if-no-files-found: warn
retention-days: 7
- name: Upload Playwright report and traces
if: failure()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: frontend-next-browser-report-${{ github.run_id }}-${{ github.run_attempt }}
path: |
arcane/home/honeypot-dashboard/frontend-next/playwright-report/
arcane/home/honeypot-dashboard/frontend-next/test-results/
if-no-files-found: warn
retention-days: 7
# First CI coverage for this crate (#1608's Rust service tier had none
# until now -- a broken build or a real regression wouldn't have
# surfaced until Arcane tried to build it live on the homeserver).
#
# Homeserver-first, unlike when #1608 landed: back then the self-hosted
# runner had no Rust toolchain to fall back on. Instead of relying on
# host state (the old assumption that rustup happened to be preinstalled
# on GH images), the homeserver twin bootstraps rustup itself
# into the runner user's persistent HOME -- rustup then honours the
# crate's own rust-toolchain.toml (#1720 pin) exactly like every other
# execution path does.
#
# This comment used to also claim `~/.cargo` + `target/` "simply stay warm
# on disk between runs, which is why no actions/cache block exists here".
# Only the first half was ever true, and the second half was the reason the
# 159s job rebuilt its whole crate graph on every single run:
#
# * `~/.cargo` (registry/, git/) really does persist -- it is in $HOME,
# outside the checkout.
# * `target/` does NOT persist. It lives inside the workspace, and
# actions/checkout defaults to `clean: true`, i.e. `git clean -ffdx`.
# The `-x` is what matters: it deletes *ignored* files, and `target/` is
# ignored (.gitignore line 7). Verified locally, not inferred: a
# `sub/target/debug/libfoo.rlib` under a `target/` ignore rule does not
# survive `git clean -ffdx`. So every run recompiled all 287 packages in
# Cargo.lock from an empty target/.
#
# The cloud twin already carries an actions/cache block over target/ and
# pays this correctly (each ubuntu-latest runner is a fresh VM, so the
# cache service is the only disk that survives). This twin has the opposite
# situation -- a persistent disk that the checkout was destroying -- so the
# fix is to move the one directory that is being wiped to a path that is
# not. Deliberately NOT an actions/cache block: see the "Reuse a persistent,
# ref-scoped target/" step below and the go-fmt job's comment for what
# actions/cache did on this runner (six same-key 465 MB copies put the repo
# over its 10 GB quota).
#
# `cargo fmt --check` deliberately isn't part of this gate: the crate
# predates any formatting pass and has never been run through rustfmt
# (532 diff hunks against its default profile) -- landing that gate
# now would force an unrelated repo-wide reformat into this PR instead
# of a follow-up that can be reviewed on its own.
backend-service:
name: Dashboard backend-service (Rust)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 90
defaults:
run:
working-directory: arcane/home/honeypot-dashboard/backend-service
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Bootstrap rustup into the persistent HOME (no-op after first)
shell: bash
run: |
set -euo pipefail
if ! command -v rustup >/dev/null 2>&1; then
# Pinned version + hardcoded checksum instead of piping
# sh.rustup.rs straight into sh (SAST-flagged, #3115). The
# unversioned dist/ path is rolling-latest, and its .sha256
# sidecar is fetched from the same host over the same TLS
# session moments later -- that detects transport corruption,
# which TLS already covers, not a compromised upstream. Same
# discipline as the trivy block below: fetch the versioned
# archive and verify against a digest hardcoded here, taken
# from that archive's own .sha256 once at authoring time.
rustup_version=1.28.2
rustup_sha256=20a06e644b0d9bd2fbdbfd52d42540bdde820ea7df86e92e533c073da0cdd43c
target=x86_64-unknown-linux-gnu
tmp="$(mktemp -d)"
curl --proto '=https' --tlsv1.2 -sSf -o "$tmp/rustup-init" \
"https://static.rust-lang.org/rustup/archive/${rustup_version}/${target}/rustup-init"
echo "${rustup_sha256} $tmp/rustup-init" | sha256sum -c -
chmod +x "$tmp/rustup-init"
"$tmp/rustup-init" -y --profile minimal --no-modify-path
rm -rf "$tmp"
fi
echo "$HOME/.cargo/bin" >>"$GITHUB_PATH"
# `rustup show` installs whatever rust-toolchain.toml pins (including
# its declared clippy component); explicit component add stays as a
# cheap idempotent safety net for toolchains whose config predates
# components support.
- name: Install pinned toolchain (rust-toolchain.toml)
shell: bash
run: |
export PATH="$HOME/.cargo/bin:$PATH"
rustup show active-toolchain
rustup component add clippy
cargo --version && rustc --version
# #3405: put target/ somewhere `git clean -ffdx` cannot reach, so the
# 159s of "recompile all 287 packages" stops happening every run.
#
# This is the self-hosted counterpart of the cloud twin's actions/cache
# block, and it is NOT an actions/cache block. On this runner the
# persistent disk already *is* the cache: the go-fmt job's comment
# records what happens when actions/cache is used anyway -- Go's 0555
# module dirs make the restore's `tar -x` fail, the job then re-uploads,
# and six same-key 465 MB copies push the repo past its 10 GB Actions
# quota and evict unrelated caches. ~/.cargo/registry and ~/.cargo/git
# need nothing from us for the same reason: they are in $HOME and have
# been warm all along. target/ is the only one checkout destroys.
#
# Keyed per ref, never shared across refs. The honeypot-ci label is
# served by several runner instances that between them run every open
# branch, so a single shared target/ would be written concurrently by
# unrelated checkouts. One directory per ref makes that structurally
# impossible: two refs never touch the same path.
#
# NOT keyed per commit, and that is the whole trick. Cargo fingerprints
# each unit against its own inputs, so reusing a branch's directory
# across that branch's commits is exactly what it is designed for -- and
# it is the case that pays. A commit-keyed directory would be cold on
# every push (a push *is* a new commit), which is a slower spelling of
# the 159s job we are trying to remove.
#
# The root is resolved best-effort and this step can never fail the job:
# if no persistent root is writable, CARGO_TARGET_DIR is left unset,
# cargo builds into the in-tree target/ it uses today, and the run costs
# what it always cost. A cold cache is therefore a normal outcome here,
# not an error -- and it cannot turn a red crate green, because cargo
# rebuilds anything whose source, toolchain or deps changed.
- name: Reuse a persistent, ref-scoped target/ (best-effort)
shell: bash
run: |
set -uo pipefail
warn() { echo "::warning::cargo-target: $*"; }
# Precedence: ops override, then a cross-instance shared root, then
# this runner instance's own HOME. The middle option needs the box
# provisioned (install-ci-runner.sh already sets UMask=0002 and a
# setgid shared group, so a group-writable root is all it takes) and
# raises the hit rate from "the same instance happened to take this
# job again" to "any instance". The HOME option needs no
# provisioning at all, so the cache is live on an un-provisioned box
# rather than being dead code waiting for an ops ticket.
root=""
for candidate in \
"${CI_CARGO_TARGET_ROOT:-}" \
/var/cargo-target-cache \
"$HOME/.cache/cargo-target"; do
[ -n "$candidate" ] || continue
if mkdir -p "$candidate" 2>/dev/null && [ -w "$candidate" ]; then
root="$candidate"
break
fi
done
if [ -z "$root" ]; then
warn "no writable persistent root; building into the in-tree target/ as before"
exit 0
fi
# GITHUB_REF is fully qualified -- refs/heads/main on push,
# refs/pull/<n>/merge on pull_request -- and for a given PR that
# merge ref is stable across every run of that PR, so one directory
# is reused for a PR's whole review loop and then abandoned when it
# closes. The slug is for humans; the digest is what actually
# guarantees uniqueness, since GITHUB_REF_NAME mangles ("a/b" and
# "a-b" both slug to "a-b").
slug="$(printf '%s' "${GITHUB_REF_NAME:-noref}" \
| tr -c 'A-Za-z0-9._-' '-' | cut -c1-40)"
ref_digest="$(printf '%s' "${GITHUB_REF:-noref}" | sha256sum | cut -c1-8)"
channel="$(sed -n 's/^channel[[:space:]]*=[[:space:]]*"\(.*\)"/\1/p' \
rust-toolchain.toml 2>/dev/null | head -1)"
[ -n "$channel" ] || channel="unpinned"
# Toolchain in the key so a pin bump starts clean rather than
# leaning on cargo noticing the compiler changed underneath a
# long-lived shared directory.
toolchain_digest="$(printf '%s' "$channel" | sha256sum | cut -c1-8)"
dir="$root/${slug}-${ref_digest}-${toolchain_digest}"
if ! mkdir -p "$dir" 2>/dev/null; then
warn "could not create $dir; building into the in-tree target/ as before"
exit 0
fi
# Heartbeat, not the directory's own mtime, is what eviction reads.
# A warm rebuild writes into target/debug/ and does not necessarily
# touch the target/ directory entry itself, so an mtime rule could
# delete a directory a running job is actively using. The job below
# has a 90-minute timeout against a 7-day floor, so a heartbeat
# touched now cannot be mistaken for an abandoned one.
touch "$dir/.ci-target-heartbeat" 2>/dev/null || true
if [ -d "$dir/debug" ] || [ -d "$dir/.fingerprint" ]; then
warm=yes
else
warm=no
fi
if ! echo "CARGO_TARGET_DIR=$dir" >>"$GITHUB_ENV"; then
warn "could not export CARGO_TARGET_DIR; building into the in-tree target/ as before"
exit 0
fi
{
echo "### Rust target/ reuse"
echo ""
echo "- root: \`$root\`"
echo "- directory: \`$dir\`"
echo "- ref: \`${GITHUB_REF:-noref}\` (slug \`$slug\`, digest \`$ref_digest\`)"
echo "- toolchain: \`$channel\` (digest \`$toolchain_digest\`)"
echo "- state on arrival: **$warm** (\"warm\" = a previous run left a target tree here)"
} >>"$GITHUB_STEP_SUMMARY"
echo "cargo-target: using $dir (warm=$warm)"
# Eviction: drop directories whose job has not run in
# CI_CARGO_TARGET_PRUNE_DAYS days, keyed on the heartbeat. Without
# this the shared box accumulates one full target/ per ref ever
# opened. Never fatal, and only ever removes a directory that has
# provably been idle for a week.
prune_days="${CI_CARGO_TARGET_PRUNE_DAYS:-7}"
find "$root" -maxdepth 2 -name '.ci-target-heartbeat' -type f \
-mtime "+$prune_days" -print0 2>/dev/null \
| while IFS= read -r -d '' heartbeat; do
victim="$(dirname "$heartbeat")"
# Defence in depth: never hand rm -rf anything but a direct
# child of the root we resolved above.
case "$victim" in
"$root"/*) rm -rf "$victim" && echo "cargo-target: pruned $victim" ;;
*) warn "refusing to prune $victim (not under $root)" ;;
esac
done || true
# Teed so the "did the cache actually hit" question is answered by cargo's
# own output rather than by inference. pipefail keeps a build failure a
# build failure; nothing here asserts anything new.
- name: Build
run: |
set -o pipefail
PATH="$HOME/.cargo/bin:$PATH" cargo build --all-targets 2>&1 \
| tee "${RUNNER_TEMP}/cargo-build.log"
# #3319: tees the test output into .ci-artifacts/ and uploads it on
# failure. Not a JUnit file: cargo has no built-in JUnit reporter, and
# its per-test result stream is only available on the unstable
# `--format json`, so the honest artifact here is the full log rather
# than a hand-rolled conversion that could misrepresent a result. The
# log is what the issue said was the only evidence a red run had; it is
# now retained rather than scrolled away.
#
# `mkdir -p` plus a named file, never a directory upload: the log is
# the only thing this job writes under .ci-artifacts/, and naming it
# keeps the artifact's contents a closed set (see the frontend-next
# twin's upload for why the directory form uploads nothing at all).
- name: Test (log retained on failure)
shell: bash
run: |
set -o pipefail
mkdir -p "$CI_ARTIFACTS_DIR"
PATH="$HOME/.cargo/bin:$PATH" cargo test 2>&1 | tee "$CI_ARTIFACTS_DIR/cargo-test.log"
env:
CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts
# Per-(run, attempt) name, no `overwrite: true` -- see the frontend-next
# twin's upload for the full argument.
- name: Upload test log
if: failure()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: backend-service-cargo-test-log-${{ github.run_id }}-${{ github.run_attempt }}
path: .ci-artifacts/cargo-test.log
if-no-files-found: warn
retention-days: 7
- run: PATH="$HOME/.cargo/bin:$PATH" cargo clippy --all-targets -- -D warnings
# #3325: the committed contract has to be what the generator
# renders. `cargo test` already fails on a stale openapi.json, but
# only the *tests* fail, and only once someone reads which of the
# four assertions went red -- `diff` states the actual change and
# is the step a reviewer reads when they wonder why the file moved.
# Regenerate locally with `cargo run --bin openapi > openapi.json`.
- name: OpenAPI contract is not stale (#3325)
run: PATH="$HOME/.cargo/bin:$PATH" cargo run --quiet --bin openapi | diff -u openapi.json -
# The measurement, not an assertion. Cargo prints one "Compiling <pkg>"
# line per unit it actually rebuilt, so this counts the work the cache
# did NOT save: ~287 on a cold run, 0 when every unit was still fresh.
# That number is the proof the cache hits, and it is read off a real
# build rather than inferred from a cache-hit log line.
- name: Report target/ reuse outcome
if: always()
shell: bash
run: |
set -uo pipefail
log="${RUNNER_TEMP}/cargo-build.log"
# `grep -c` prints the count AND exits 1 when the count is 0, so a
# bare `|| echo 0` would yield "0\n0" and break the arithmetic
# below. Swallow the status instead and keep only the count.
count() { grep -c "$1" "$2" 2>/dev/null || true; }
total="$(count '^[[:space:]]*\[\[package\]\]' Cargo.lock)"
[ -n "$total" ] || total=0
if [ -r "$log" ]; then
compiled="$(count '^[[:space:]]*Compiling ' "$log")"
[ -n "$compiled" ] || compiled=0
reused="$(( total - compiled ))"
else
compiled="n/a"
reused="n/a"
fi
size="n/a"
if [ -n "${CARGO_TARGET_DIR:-}" ] && [ -d "$CARGO_TARGET_DIR" ]; then
size="$(du -sh "$CARGO_TARGET_DIR" 2>/dev/null | cut -f1)"
[ -n "$size" ] || size="n/a"
fi
{
echo "### Rust build work actually done"
echo ""
echo "- packages compiled this run: **$compiled** of $total in Cargo.lock"
echo "- units served from the reused target/: **$reused**"
echo "- target/ size after the build: \`$size\`"
echo ""
echo "A cold cache is a valid outcome: the step above falls back to the"
echo "in-tree target/ and this run costs what it always cost."
} >>"$GITHUB_STEP_SUMMARY"
echo "cargo-target: compiled=$compiled reused=$reused size=$size"
backend-service-cloud:
name: Dashboard backend-service (Rust) (GitHub-hosted)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
timeout-minutes: 60
defaults:
run:
working-directory: arcane/home/honeypot-dashboard/backend-service
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- run: rustup component add clippy
- uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: |
~/.cargo/registry
~/.cargo/git
arcane/home/honeypot-dashboard/backend-service/target
key: cargo-${{ runner.os }}-${{ hashFiles('arcane/home/honeypot-dashboard/backend-service/Cargo.lock') }}
restore-keys: cargo-${{ runner.os }}-
- run: cargo build --all-targets
# #3319: twin of the homeserver copy's log retention above.
- name: Test (log retained on failure)
shell: bash
run: |
set -o pipefail
mkdir -p "$CI_ARTIFACTS_DIR"
cargo test 2>&1 | tee "$CI_ARTIFACTS_DIR/cargo-test.log"
env:
CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts
# Twin of the homeserver copy's upload: same per-(run, attempt) name,
# same named-file path, no `overwrite: true`.
- name: Upload test log
if: failure()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: backend-service-cargo-test-log-${{ github.run_id }}-${{ github.run_attempt }}
path: .ci-artifacts/cargo-test.log
if-no-files-found: warn
retention-days: 7
- run: cargo clippy --all-targets -- -D warnings
# #3325: same drift gate as the homeserver lane -- see the note
# there. The two lanes both build this contract, so the contract
# cannot pass on one toolchain and fail on the other.
- name: OpenAPI contract is not stale (#3325)
run: cargo run --quiet --bin openapi | diff -u openapi.json -
vendored-theme:
name: Vendored Xore/theme is in sync
needs: [ci-target]
if: needs.ci-target.outputs.homeserver == 'true'
runs-on: [self-hosted, linux, x64, honeypot-ci]
timeout-minutes: 10
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: theme.css matches the pinned commit
run: scripts/check-vendored-theme.sh
# python3, not python: this job has no setup-python step, so on the
# self-hosted lane it runs on the runner's own PATH and the homeserver
# runner ships no bare `python` -- `python branding/...` died with
# exit 127 there (#2461) while the vendored sheet itself was fine.
- name: Portable APIARY tokens match Xore/theme
run: python3 branding/scripts/check_theme_sync.py
# Canvas surfaces resolve tokens to pixels and carry dark-theme
# fallbacks, so a rename upstream degrades silently instead of failing.
# #1825 renamed eight tokens and 31 call sites kept rendering, wrong.
- name: Frontend theme tokens all exist
run: scripts/check-theme-tokens.sh
# The catalogue lives in the frontend, the tokens live in the vendored
# stylesheet. A theme in one and not the other fails silently either
# way -- default styling, or unpickable. #1758.
- name: Theme catalogue matches theme.css
run: scripts/check-theme-catalogue.sh
vendored-theme-cloud:
name: Vendored Xore/theme is in sync (GitHub-hosted)
needs: [ci-target]
if: needs.ci-target.outputs.homeserver != 'true'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: theme.css matches the pinned commit
run: scripts/check-vendored-theme.sh
# Same python3 choice as the self-hosted twin above: one spelling
# across both lanes, valid on the runner's own PATH either way.
- name: Portable APIARY tokens match Xore/theme
run: python3 branding/scripts/check_theme_sync.py
- name: Frontend theme tokens all exist
run: scripts/check-theme-tokens.sh
- name: Theme catalogue matches theme.css
run: scripts/check-theme-catalogue.sh
scripts-and-compose:
# Per-row executor routing (#2389, widened by #2565): `home: true`
# marks matrix rows whose whole run needs nothing beyond checkout
# files, this workflow's own setup interpreters, pip installs, and --
# since #2565 retired the runner's no-docker design -- the docker
# daemon the runner user now drives via its docker-group membership.
# Eligibility under the ORIGINAL (docker-less) criterion was audited
# by executing every candidate in a docker-less, sudo-less
# environment with the pip-flavored rows re-run in clean venvs -- the
# audit ledger lives on the #2389 PR. The four
# analysis/tests/test_{honeypot_ilm_rollover,geoip_pipeline,
# dionaea_incidents_index,conpot_persona_pipeline}.sh rows were left
# unflagged then because they print "SKIP: docker daemon is not
# reachable" without a container engine -- on the box they now run
# their real Elasticsearch containers, so they are flagged.
#
# A flagged row lands on [self-hosted, linux, x64, honeypot-ci] whenever
# ci-target proves the box live -- the same heartbeat gate every paired
# family above uses -- else falls back to ubuntu-latest, gaining the
# "(GitHub-hosted)" name suffix that day per the pair-naming rule above.
# timeout-minutes rides along only while the leg actually sits on the
# metal (a wedged pickup blocks a real box nobody reboots promptly);
# fallback legs keep the platform ceiling (360 = effectively unset).
#
# Unflagged rows stay on ubuntu-latest: the inert "Set up Node for the
# dashboard OIDC suites" row (a matrix can't inject a use-step -- its
# node arrives from the runner image either way), and the
# "Sandbox shell tests (#2268, #2253)" row -- that one is eligible
# under the criterion above (toolchain-only: bash, coreutils, a
# mktemp workdir) but is deliberately kept off the metal anyway,
# because it drives guest-runner.sh's tail through a PATH-stubbed
# `systemctl poweroff` and run_pending.sh's flock path. Both are
# neutralised by the stubs and by WINDOWS_SANDBOX_* redirection into
# the temp dir, but a disposable GitHub-hosted VM is the right blast
# radius for a suite whose production form powers a machine off and
# takes the shared KVM detonation lock. Nothing else -- every
# remaining row is flagged.
name: ${{ matrix.name }}${{ matrix.home == true && needs.ci-target.outputs.homeserver != 'true' && ' (GitHub-hosted)' || '' }}
needs: [ci-target]
runs-on: ${{ matrix.home == true && needs.ci-target.outputs.homeserver == 'true' && fromJSON('["self-hosted", "linux", "x64", "honeypot-ci"]') || fromJSON('["ubuntu-latest"]') }}
timeout-minutes: ${{ (matrix.home == true && needs.ci-target.outputs.homeserver == 'true') && 45 || 360 }}
strategy:
fail-fast: false
matrix:
include:
- name: Python syntax
home: true
run: find . -path '*/vendor' -prune -o -name '*.py' -type f -print0 | xargs -0 -r python -m py_compile
- name: Dionaea log rotation patch (#1389)
home: true
run: |
# Build-time patch applied to dionaea's vendored log_json.py/
# log_incident.py -- exercises apply_patch() against fixtures built
# from the patch's own OLD_CLASS/OLD_IMPORT text, and the actual
# patched FileHandler class body (exec'd, not reimplemented) for
# rotation/data-loss behavior.
python arcane/home/honeypot-dionaea/dionaea/tests/test_log_rotation_patch.py
- name: Conpot JSON log rotation patch (#2892)
home: true
run: |
# Same shape as the dionaea row above, plus a behavioural
# check the dionaea suite doesn't have: the patched JsonLogger
# is imported (not reimplemented) against a fixture package
# and run past CONPOT_JSON_LOG_MAX_BYTES to confirm the
# close/rename/reopen rotation actually fires.
python arcane/home/honeypot-conpot/conpot/tests/test_json_log_rotation_patch.py
- name: Galah JSON log rotation patch (#2892)
home: true
run: |
# Same shape as the conpot row above, but the target is Go
# rather than Python: the suite extracts the patch's added
# rotatingWriter type into a standalone Go program and
# `go run`s it, confirming the close/rename/reopen rotation
# actually fires -- skips (not fails) if `go` isn't on PATH.
python arcane/home/honeypot-galah/galah/tests/test_json_log_rotation_patch.py
- name: Beelzebub JSON log rotation patch (#2892)
home: true
run: |
# Same shape as the Galah row above (Go target, `go run`s the
# extracted rotatingWriter, skips without `go`) -- the other
# vendored Go sensor, patched in internal/builder/builder.go.
python arcane/home/honeypot-beelzebub/beelzebub/tests/test_json_log_rotation_patch.py
- name: HellPot X-Forwarded-For trust patch (#1876)
home: true
run: |
# Two build-time patches run in sequence against the same file
# and the second matches on the first's output. That coupling
# is invisible at review time and breaks the image build with a
# string mismatch several layers down, so it is asserted here --
# along with the trust rule that makes reintroducing header
# trust safe after #1419 removed it as spoofable.
python arcane/home/honeypot-hellpot/hellpot/tests/test_xff_trust_patch.py
- name: "Mailoney JSON-line sink patches (#1422, #2197)"
home: true
run: |
# Exact-text build-time patches over vendored core.py. The suite
# execs the injected helper block under the container's own
# TZ=Europe/Berlin pin to prove emitted Z-suffixed stamps are
# true UTC (#2197: Berlin wall clock used to wear a Z, shifting
# ip-enrichment-worker's portbridge join by the DST offset),
# that every match target still fits upstream verbatim, and
# that AUTH PLAIN capture survives unpadded or hostile base64.
python arcane/home/honeypot-mailoney/mailoney/tests/test_json_log_patch.py
- name: Sensor timestamps carry explicit UTC sources (#2197)
home: true
run: |
# Repo-wide grep-level guard for the class behind #2197: any
# emitter of a %SZ stamp must pass an explicit UTC source
# (gmtime/timezone.utc/date -u) on the same line, or carry a
# `utc-verified:` waiver with its reason inline. Parse-side
# formats are exempt by construction; Rust .format() calls are
# out of scope by design -- the zone rides on the value's type,
# not the template.
python scripts/check-timestamp-utc.py
- name: "JSON sink retention parity (#120, #2196, #2216)"
home: true
run: |
# Every self-rotating JSON directory under /logs must have BOTH
# halves of the #120 contract present -- a writer that actually
# rotates (proven by its rotation token in source) and a pruner
# glob line in log-maintenance.sh, credited to that directory's
# own find line -- plus every find line there must belong to a
# ledgered sink. #2196 shipped because mailoney lacked both
# halves at once and nothing structural connected compose
# volumes to vendored writers to this script; adding a sensor
# now means adding one reviewable ledger row instead.
#
# #2216: the checks above only see halves an author already
# touched, so the script also enumerates every
# /opt/stacks/apiary/logs/<dir> bind mount in arcane/home/*/
# compose.yml -- a new stack's mount is the one thing it cannot
# omit and still write to disk. Each must be a ledger row, an
# EXEMPT entry with a written reason, or a KNOWN_UNCOVERED entry
# naming the issue tracking the gap.
#
# #2826: and a written reason is not evidence. An EXEMPT entry
# that asserts a mechanism ("log-maintenance.sh rotates it",
# "the writer self-bounds") now carries that mechanism as a
# checked key -- deleting the rotate line, or the ported
# rotateIfOversized, fails here instead of leaving the ledger
# asserting coverage that no longer exists.
python scripts/check-json-sink-retention-parity.py
#
# #2921: the writer proof above is a grep, and a grep can't
# tell source from a test asserting on the same token or a
# stale __pycache__/*.pyc left over from a previous run --
# this exercises tree_contains() against a synthetic fixture
# to prove both false-positive paths stay closed.
python3 -B scripts/tests/test_check_json_sink_retention_parity.py
# #2926: the persona overlay's bare-name entries (free, uptime,
# id, ...) were shadowed by cowrie's own builtins and silently
# unreachable -- nproc disagreeing with lscpu in the same session
# was the cheapest tell. Asserts the priority patch still matches
# upstream's getCommand() exactly once and is idempotent, plus the
# persona-consistency fixtures it unblocks (nproc == lscpu's CPU
# count, free's static seed is persona-scale).
#
# The behavioural half is what earns the row: it execs the patched
# getCommand() and pins the two regressions an unscoped version
# causes -- a path-form invocation (/bin/echo, and 66 other real
# binaries in the runtime image) reading the container's own ELF
# instead of being emulated, and a canned overlay outranking
# uname's cowrie.cfg-driven persona or dd's operand parsing and
# input_data capture.
- name: "Cowrie txtcmds overlay priority patch (#2926)"
home: true
run: python arcane/home/honeypot-cowrie/cowrie/tests/test_txtcmds_priority_patch.py
- name: "cowrie auth-phase service-request patch (#3307)"
home: true
run: python arcane/home/honeypot-cowrie/cowrie/tests/test_service_request_log_patch.py
# #1982: compose.yml is the only place a stack's knobs are exercised,
# .env.example is what a deployer copies when standing the stack up
# manually -- drift between them means a variable exists at runtime
# but can't be discovered at setup time (that's how ZEEK_PROXY_IFACE
# shipped as refuse-to-start-by-default and CANARY_PUBLIC_HOSTNAME as
# silent placeholder-domain). Two rules: every ${VAR} interpolated in
# arcane/home/<stack>/compose.yml must appear as a KEY= line in that
# stack's own .env.example (per-stack, pointers to a sibling don't
# count -- each file is copied on its own), and no key may be listed
# twice within one environment: block (the ZEEK_PROXY_ATTRIBUTION_
# INTERVAL_SECONDS dup meant tuning line one silently changed nothing).
- name: Compose interpolation documented per stack (#1982)
home: true
run: python scripts/check-compose-env-docs.py
# #119/#2051: a healthcheck (compose-level or baked into the
# service's own Dockerfile) with no autoheal=true label just
# detects a wedge and does nothing about it. #119 fixed the fleet
# once; the gap came back on every stack that didn't exist yet
# when it closed, so this is the standing guard #119 itself asked
# for, scanning both layers -- a compose-only matcher reproduces
# the exact blind spot that let four image-level-HEALTHCHECK
# services hide from the first audit.
- name: Healthcheck/autoheal label pairing (#2051)
home: true
run: |
python -m pip install -q pyyaml
python scripts/check-autoheal-labels.py
# #2320: every WORKER_LOOPS consumer boots through the same
# apiary-backend entrypoint whose gate (#2183) refuses an empty
# SERVICE_TOKEN unless APIARY_ALLOW_UNAUTH_DEV=1 -- both arrive as
# per-container env, so a loop tier forwarding only the token half
# crash-looped on the dev-tier deployment (#2320) while its serving
# siblings booted. Mechanical form of checking every consumer
# against the gate's inputs: if backend ever grows a third gating
# variable, AUTH_VARS in the script and this gate move together.
- name: WORKER_LOOPS consumers carry the apiary-backend auth pair (#2320)
home: true
run: python scripts/check-backend-boot-contract.py
# #2352: #1502 moved the YARA scanner under
# honeypot-payload-analysis/analysis/yara/ and updated the checker
# script, but the surrounding prose never got the same pass -- six
# documents kept pointing readers at the dead repository-root path,
# including the operator README's own sync command, which failed
# verbatim from a fresh checkout. Grep-level guard for that class:
# outside arcane/**, a .md may cite analysis/yara/ only on a line
# that also names its real home (or carries an inline waiver).
- name: Docs cite moved trees by their real home (#2352)
run: python scripts/check-doc-stale-paths.py
# #2962: the orphaned-script sweep (#1609 Phase 6) found these two
# regression tests describe real coverage in their own headers
# (guest-runner.sh's post-detonation poweroff-on-artifact-failure
# path, #2268; run_pending.sh's stale-claim reconciliation, #2253)
# but were never referenced by name, glob, or doc anywhere in the
# tree -- unlike analysis/github/tests/*.sh, nothing ran them.
# Neither needs a hypervisor or network (both stub the
# orchestrator/guest side per their own headers), so plain
# ubuntu-latest is enough -- not the self-hosted KVM host
# guest-runner.sh and run_pending.sh actually run on in
# production. Fixed here as part of wiring them in:
# run_pending.sh's ${VAR:-default} treated the stale-claims test's
# explicit `export WINDOWS_SANDBOX_KVM_SHARED_LOCK=""` (meant to
# disable the shared lock, per the script's own "Empty disables
# it" comment) the same as unset, so the never-run test would
# have flocked the real production lock path the first time it
# ran anywhere with permission to.
- name: "Sandbox shell tests (#2268, #2253)"
run: |
for t in sandbox/windows/tests/*.sh; do
echo "=== $t ==="; bash "$t" || exit 1
done
# #2456: vps/zeek and dev/sensing-lab build the same parser set on
# purpose -- the lab's measurements are only evidence for production
# while production runs the same parsers. A component pinned in one
# file but floating (or pinned to a different commit) in the other
# is exactly the lab/prod drift the #1821 hassh pin pattern exists
# to prevent; this makes it a red build instead of a silent rebuild
# difference. Offline and stdlib-only, so it rides the docker-less
# homeserver rows.
- name: Zeek lab/prod build pin parity (#2456)
home: true
run: python scripts/check-zeek-pin-parity.py
# #2458: the pointer-rot family (#2352, #2353, #2356/#2453, #2455,
# #2357, five docs' worth invalidated by #1659) kept being fixed by
# hand after PR #2367's one-off sweep script proved the shape. This
# makes that sweep standing: every repo-path-like citation in
# docs/** and the root README must resolve to a git-tracked path.
# Offline, stdlib-only; deliberate era-record and host-side
# references live in scripts/doc-path-lint-allowlist.txt with
# per-entry reasons instead of passing silently.
- name: Docs cite existing repo paths (#2458)
home: true
run: python scripts/check-doc-paths-exist.py
# #3332: the inverse of #2458 -- every docs/**/*.md must be reachable
# from docs/README.md through relative links (transitively), except
# dated record trees listed in the script. 54 of 120 docs, runbooks
# among them, were unreachable from the map when this landed.
- name: Docs reachable from the documentation map (#3332)
home: true
run: python scripts/check-docs-reachable.py
# #3395: #2458 scans only README.md and docs/**, so the component
# trees under arcane/**, analysis/**, sandbox/** and branding/** were
# never link-checked -- the agent-intrusion-corpus README's dead
# `../../docs/...` hop (two levels short of a five-deep path) lived
# there unnoticed. Whole-tree relative-link resolution. Fenced blocks
# are skipped: their paths belong to the reader's project, not ours.
- name: Every tracked doc's relative links resolve (#3395)
home: true
run: python scripts/check-doc-links.py
# #3395: two of the forty mermaid diagrams in the tree did not parse
# at all -- AI_TRIAGE.md named a node `call` (a reserved mermaid
# token) and ghidra/README.md left a colon unquoted inside an edge
# label. Both rendered as an error box on GitHub and nothing noticed,
# so mermaid parsing is now a gate rather than a review habit. Renders
# every block in one headless browser; mermaid is resolved from the
# npx cache, so the step is self-bootstrapping (no pinned version to
# drift out of date with the docs). The warm-up below is what makes
# that true: a fresh runner's npx cache is empty, so the lookup inside
# check-mermaid.mjs finds nothing and the gate cannot run at all.
- name: Mermaid diagrams parse (#3395)
home: true
run: |
npx -y @mermaid-js/mermaid-cli --version
node scripts/check-mermaid.mjs
# #2576: the CAPE implementation-plan doc claimed the detail page
# (/cape/{sha256}) was admin-gated. Nothing in backend-service
# enforces that -- require_service_token is the actual gate on
# /api/v1/cape/{sha}, the same middleware every other /api/v1
# detail route carries (main.rs:183, wired at main.rs:468). A real
# admin-role check does exist, but on a different route entirely
# (the Workbench "cape" analyzer entry that triggers a fresh
# detonation) -- this regression-tests the doc text, not the code,
# so the wrong claim can't silently drift back in on a rewrap or a
# synonym swap. pytest installed inline, same as the other
# pytest-based rows above -- it isn't a production dependency.
- name: Docs regression tests (#2576)
home: true
run: |
python -m pip install pytest
python -m pytest tests/docs/ -v
- name: "Dionaea SMB exploit-signature incidents (#622, #1775)"
home: true
run: |
# Executes the code this patch injects into smb.py, against a
# stand-in connection, rather than asserting on the substituted
# text -- the property that matters is one incident per
# connection, not one per SMB packet. Reporting per packet made
# a single DoublePulsar connection produce ~643 documents and
# 47.5% of everything the fleet ingested.
python arcane/home/honeypot-dionaea/dionaea/tests/test_smb_exploit_patch.py
- name: Ghidra worker, spool discipline and AI triage
home: true
run: |
# Stdlib only and stubbed on both sides, so it runs here in seconds.
# Worth running: this worker decides what reaches an analyst's screen
# from a live malware sample, and the triage half will not send
# sample-derived text to a non-local model endpoint — a rule that is
# only worth anything if something checks it on every change.
python analysis/ghidra/worker/tests/test_ghidra_worker.py
# GPU-queue drain crash recovery (#2075): a drainer death
# mid-generation used to strand its job as an eternally-running
# zombie. Stubbed queue, no ES/docker/nvidia-smi -- the sweep
# must requeue-once-then-fail past the staleness bound, never
# touch a legitimately-running job, and run before the queued
# early return that used to hide the problem.
python analysis/ghidra/worker/tests/test_gpu_queue_drain.py
# Replacement Ghidra REST service (#245): fake analyzeHeadless, no
# real Ghidra/JVM in CI -- verifies server.py's HTTP/queue layer
# only, same "stub both sides" reasoning as the worker test above.
python analysis/ghidra/service/tests/test_server.py
# Governance tests use synthetic snapshots/reports only. CI never
# pulls or loads multi-GB model weights.
python analysis/ghidra/models/tests/test_model_governance.py
python analysis/ghidra/models/tests/test_model_status_adapter.py
# #3334: the host-side wiring that actually runs the injection
# corpus weekly and on a pin change. The corpus itself needs the
# real model, so what is provable without a GPU is the part that
# made it dead weight before: that the units exist, that the
# .path unit watches files the installer deploys, and that the
# runner fails closed and skips only when nothing changed. Driven
# against a stub `docker` on PATH -- no model, no Ollama.
python analysis/ghidra/models/tests/test_llm_injection_suite.py
# Benchmark transcript records (#1805). No model involved -- this
# asserts the prompt is stored as sent, that a refusal or timeout
# is stored rather than dropped, that a stored transcript is never
# rewritten, and that captured real-data transcripts cannot be
# written inside the repo. That last one is a data-handling rule,
# so it is worth a gate rather than reviewer vigilance.
python analysis/ghidra/benchmarks/tests/test_transcripts.py
# Tier B evidence cache (#1805). No Ghidra and no service in CI --
# covers the cache key (every component that shapes the evidence
# must change it, or a Ghidra upgrade silently reuses output from
# a different decompiler) and the injection assertion, which must
# report honestly that the payload never reached the evidence.
python analysis/ghidra/benchmarks/tests/test_ghidra_cache.py
# Cross-path fidelity for the ghidra slot (#1805). No model, no
# Ghidra, no service. The other Tier B tests all synthesise their
# cache entries, so they agree with whatever the loader happens to
# expect; this one drives the same path over real recorded Ghidra
# output, which is what makes a drift in the expected response
# shape visible. Also measures the standing gap between the
# qualification gate's hand-written fixtures and production's
# _evidence() output, so a change to either renderer is a
# measured delta rather than a silent one.
python analysis/ghidra/benchmarks/tests/test_ghidra_fidelity.py
# Corpus manifest validator (#2038). Pure structural checks over
# fixture manifests -- no compiler, no corpus rebuild. Covers the
# three gaps that let a broken corpus pass: a required alternative
# that trips its own case's forbidden list, a rubric case whose
# build has disappeared, and a partial toolchain/opt-level grid.
# The overlap check reuses polarity.forbidden_hit(), so the guard
# that keeps it from regressing to #2517's naive-substring false
# positives lives here too and has to run on every change.
python analysis/ghidra/benchmarks/tests/test_validate_manifest.py
# Claim-pool scoring (#1805). Embedder is stubbed, so no Ollama.
# Covers the properties that decide whether a claim-pool score can
# be believed: the adjudicator cannot be a contestant, rephrasing
# earns nothing, unadjudicated claims never count as correct, and
# the pool version moves when a verdict does.
python analysis/ghidra/benchmarks/tests/test_claims.py
# Corpus scorer (#1952). An empty answer used to collect the
# injection-resistance point, giving a model that returned nothing
# a floor of 14/69 -- which is how a thinking model scoring its
# empty-answer floor read as a fifth of a pass instead of a zero.
python analysis/ghidra/benchmarks/tests/test_record_baseline.py
# Session-slot critical gate (#2232). The rubric-vocabulary leg
# no longer gates: only MITRE correctness, forbidden content,
# and severity can fail critical_ok, since governance treats
# that boolean as disqualifying (#1947 rule 4) and five of
# eleven round models proved the old AND tripped on wording.
python analysis/ghidra/benchmarks/tests/test_session_scoring.py
# Harmony-family serving adaptation (#2233). gpt-oss tags through
# Ollama's /api/chat need a different wire shape or `content`
# comes back empty; pins the dispatch so the calibrated Qwen
# request shape is untouched.
python analysis/ghidra/benchmarks/tests/test_harmony_chat.py
# Offline transcript re-scoring (#2266). No GPU and no Ollama --
# the fixtures are written through the real TranscriptWriter and
# replayed, so this pins the property the tool exists for: a
# rescore of a stored answer scores identically to what a live
# evaluate_slot() would have scored for it. Also pins
# scorer_git_sha's dirty marker, without which the documented
# "run at two commits and diff the reports" workflow silently
# attributes both sides to the same commit.
python analysis/ghidra/benchmarks/tests/test_rescore_from.py
# #2980: the four files below existed with no workflow running
# them. This block is hand-enumerated -- there is no pytest
# discovery over analysis/ghidra/benchmarks/tests/ -- so a test
# file lands here only if someone remembers, and four had drifted
# off. test_injection_gate.py is the sharpest case: its 48 tests
# are the whole regression guarantee the #2694 Tier B gate rests
# on, and they were executing only by hand. All four are stubbed
# the same way as the rows above (no model, no Ghidra, no
# network) and run standalone.
python analysis/ghidra/benchmarks/tests/test_injection_gate.py
python analysis/ghidra/benchmarks/tests/test_corpus_eval.py
python analysis/ghidra/benchmarks/tests/test_run_real_corpus_eval.py
python analysis/ghidra/benchmarks/tests/test_regenerate_pre_2393.py
# full_capabilities.py (#800): inventories env vars straight out of
# the real pipeline source, so its own test runs against the real
# tree rather than fixtures -- proving that stays true.
python analysis/ghidra/tests/test_full_capabilities.py
# CAPE worker's status.json discipline (#319 follow-up). No live
# CAPE involved -- see the test file's own docstring for why that
# stays with --selftest --round-trip (#318) instead.
python sandbox/cape/worker/tests/test_cape_worker.py
- name: Real-data session probe metadata honesty (#2387)
home: true
run: |
# contracts is stubbed, so /app, pydantic, ES, docker and Ollama
# stay out of CI. Pins both halves of the metadata honesty fix:
# stage 0's correlated auth/duration evidence reaches
# session_prompt() unchanged (the old code pinned
# auth_success=False / duration=0.0 under an EXACT-prompt claim),
# and sessions whose evidence missed the lookback window are
# reported and skipped instead of being prompted with defaults.
# Also drives the extracted stage-0 jq program (#2426) against
# committed fixtures -- latest-close-wins, string-duration
# coercion, failed-only auth, empty second-response gap -- on
# plain jq; skips itself gracefully where jq is absent so this
# row stays homeserver-first per #2389 without dropping coverage.
python analysis/ghidra/benchmarks/tests/test_probe_real_session.py
- name: Statictools lief contract (#2072)
home: true
run: |
# lief_parse() is the structural read on PE/Mach-O/ELF samples,
# and lief was the one unpinned install in an image whose whole
# point is that nothing drifts underneath it silently (#2072).
# CI installs the version the Dockerfile pins -- read back out
# of that Dockerfile, so the tested pin and the shipped pin
# cannot diverge.
LIEF_PIN=$(grep -om1 'lief==[0-9.]*' analysis/ghidra/statictools/Dockerfile | cut -d '=' -f3)
python -m pip install -q "lief==${LIEF_PIN}"
python analysis/ghidra/statictools/tests/test_lief_parse.py
- name: Guarded LLM worker contracts
home: true
run: |
python -m pip install -r llm-worker/requirements.txt
python -m unittest discover -s llm-worker/tests -v
python llm-worker/worker.py --selftest
# ml-worker/, analysis/es-results-importer/, personas/, and
# services-adapter/ are all deployed by deploy.yml but had zero test
# coverage wired into any workflow -- 189 tests across 11 files
# (found while auditing #982's own CI gap; this is the same class of
# issue, just for four whole modules instead of two scripts).
# ml-worker's suite uses pytest specifically (parametrized fixtures),
# unlike every other Python test dir in this workflow, which is
# plain unittest -- pytest isn't a production dependency so it's
# installed here rather than added to ml-worker/requirements.txt.
# #3319: --junitxml is pytest's own machine-readable reporter, so
# this needs no new dependency. Both pytest rows carry it and both
# write into the one .ci-artifacts/ dir the upload step below
# collects; the other ~20 rows are plain unittest/shell/go and
# simply produce no file, which `if-no-files-found: ignore` treats
# as the normal case rather than a warning on every one of them.
- name: ml-worker anomaly pipeline contracts
home: true
run: |
python -m pip install -r ml-worker/requirements.txt pytest
mkdir -p .ci-artifacts
python -m pytest ml-worker/tests/ -v --junitxml=.ci-artifacts/ml-worker-junit.xml
# #2219: Keycloak caps every admin-events page silently, and the
# single-page fetch permanently dropped everything past row 1,000
# on exactly the credential-spray days this telemetry exists for.
# The pagination contracts (stubbed HTTP layer, no live service)
# pin the drained/undrained pager and its checkpoint-hold rule.
# pytest again only installed here -- not a production dependency.
- name: auth-events-worker Keycloak pagination contracts
home: true
run: |
python -m pip install -r auth-events-worker/requirements.txt pytest
mkdir -p .ci-artifacts
python -m pytest auth-events-worker/tests/ -v \
--junitxml=.ci-artifacts/auth-events-worker-junit.xml
- name: es-results-importer chunked-upload and dedup contracts
home: true
run: |
python -m pip install -r arcane/home/honeypot-dashboard/analysis/es-results-importer/requirements.txt
python -m unittest discover -s arcane/home/honeypot-dashboard/analysis/es-results-importer/tests -v
- name: Persona application is idempotent
home: true
run: python -m unittest discover -s personas/tests -v
- name: services-adapter allowlist enforcement
home: true
run: python -m unittest discover -s arcane/home/honeypot-dashboard/services-adapter/tests -v
- name: Sandbox guest string-cleaning filter (#530)
home: true
run: python3 sandbox/test_guest_clean_strings.py -v
# #2878: these two rows used to name test_extract_iocs.py (#482) and
# test_run_sample_golden_check.py (#100, #2023) individually, so the
# two delivery-check suites #2252 added next to them (Windows and
# GHOSTS) ran nowhere at all -- a regression test nothing runs is a
# file, not a check. Discovery instead, the way the worker suites
# are already run, so the next test_*.py dropped in these
# directories is picked up without a workflow edit. #2775's
# legacy-WinRM result test is the first to arrive that way.
- name: "Windows sandbox orchestrator suites (#482, #2023, #2252, #2775)"
home: true
run: python3 -m unittest discover -s sandbox/windows/orchestrate -p 'test_*.py' -v
- name: GHOSTS sandbox orchestrator suites (#2252)
home: true
run: python3 -m unittest discover -s sandbox/ghosts/orchestrate -p 'test_*.py' -v
# #3314: .github/actionlint.yaml existed but nothing ran actionlint.
# Blocking. SHELLCHECK_OPTS matches the repo's "high-severity
# ShellCheck" policy: info-level notes in run: blocks do not fail.
# Archive checksum is upstream's (actionlint_1.7.7_checksums.txt).
- name: "Workflow lint (actionlint, #3314)"
home: true
run: |
set -euo pipefail
tgz="$RUNNER_TEMP/actionlint.tgz"
curl -fsSL -o "$tgz" https://github.com/rhysd/actionlint/releases/download/v1.7.7/actionlint_1.7.7_linux_amd64.tar.gz
echo "023070a287cd8cccd71515fedc843f1985bf96c436b7effaecce67290e7e0757 $tgz" | sha256sum -c -
tar -xzf "$tgz" -C "$RUNNER_TEMP" actionlint
SHELLCHECK_OPTS="-S warning" "$RUNNER_TEMP/actionlint" -color
# A plain-scalar `name: Foo (bar, #123)` is valid YAML that silently
# truncates at " #" (a comment) -- eight names were shipping as
# "Foo (bar," before #3314. actionlint cannot see it; this can.
if grep -nE '^[[:space:]]*(- )?name: [^"'"'"'].* #' .github/workflows/*.yml; then
echo "::error::unquoted name: containing ' #' -- quote it, YAML reads the rest as a comment"
exit 1
fi
# #3313: every SHA-pinned third-party `uses:` must carry a full
# `# vX.Y.Z` comment. zizmor's unpinned-uses audit reads the SHA
# and ignores the comment entirely, so a pin can be perfectly
# well-formed and still say the wrong thing: two pins here were
# labelled `# v7` where the SHA is v7.0.0, and one was labelled
# `# v6` for a v4.6.2 SHA. A reviewer reading the comment is
# misled about what a job actually runs, which defeats the point
# of recording it. Offline and deterministic -- no network, so it
# cannot flake. Dependabot rewrites both halves of a pin on a
# bump, so this stays satisfied as versions move.
if grep -nEo 'uses: +[A-Za-z0-9._-]+/[A-Za-z0-9._/-]+@[0-9a-f]{40}( *# *[^ ]+)?' \
.github/workflows/*.yml \
| grep -vE '@[0-9a-f]{40} *# *v[0-9]+\.[0-9]+\.[0-9]+$'; then
echo "::error::action pin without a full '# vX.Y.Z' comment -- SHA is pinned but its version is not stated, or is stated as a bare major"
exit 1
fi
# #3314: zizmor security audit of the workflows. Fails on every
# medium+ finding EXCEPT the two rules left in ADVISORY below.
# #3313 landed the other three (unpinned-uses, excessive-permissions
# and artipacked) across the whole tree, so they are now dropped
# from ADVISORY and a regression in any of them is blocking. zizmor
# publishes no checksum file; the pin below was recorded on first
# download (2026-09-26).
- name: "Workflow security audit (zizmor, #3314)"
home: true
run: |
set -euo pipefail
tgz="$RUNNER_TEMP/zizmor.tgz"
curl -fsSL -o "$tgz" https://github.com/zizmorcore/zizmor/releases/download/v1.30.1/zizmor-x86_64-unknown-linux-gnu.tar.gz
echo "e65324f4430c2717591937edcec90ccbefaf14c174f8ec9415e03ca875b46e1a $tgz" | sha256sum -c -
tar -xzf "$tgz" -C "$RUNNER_TEMP" zizmor
"$RUNNER_TEMP/zizmor" --offline --min-severity medium --format json .github/workflows > "$RUNNER_TEMP/zizmor.json" || true
python3 - "$RUNNER_TEMP/zizmor.json" <<'PY'
import collections, json, sys
# Both remaining advisory rules are Low/High-by-rule but not
# exploitable as written, and neither is a checkout/pin/permission
# property -- so #3313 does not own them:
#
# self-repository: the five `uses: ./.github/workflows/ci-router.yml`
# calls. zizmor wants the explicit owner/repo/path@ref form; the
# local-path form is GitHub's own first-party-reusable-workflow
# syntax and is what makes the router pick up edits to the
# shared router without a second SHA to bump. Every one of the
# five is triggered by pull_request/push/schedule, never by
# pull_request_target, so the called workflow is always the
# default branch's own file.
#
# dangerous-triggers: main-health-watch.yml's `workflow_run`. It
# runs the default branch's copy of the script and never reads
# the triggering event: every value it reports comes from a
# main-scoped `gh api`/`gh run list` read, and the only
# untrusted-looking text it echoes is main's own commit
# subjects. Changing the trigger would blind the #3324 alarm
# that exists precisely to catch runs nobody started.
ADVISORY = {"self-repository", "dangerous-triggers"}
findings = json.load(open(sys.argv[1]))
by_rule = collections.Counter(f["ident"] for f in findings)
for rule, n in sorted(by_rule.items()):
print(f"{'advisory' if rule in ADVISORY else 'BLOCKING'} {n:3d} {rule}")
blocking = [f for f in findings if f["ident"] not in ADVISORY]
for f in blocking:
loc = f["locations"][0]["symbolic"]
print(f"::error::zizmor {f['ident']}: {f['desc']} ({loc['key']})")
sys.exit(1 if blocking else 0)
PY
# #3320: Dockerfile lint, separate from image-security-scan's CVE
# check. Policy (and why package-version pins are exempt) lives in
# .hadolint.yaml; fails on warnings and errors. Pinned binary,
# checksum pinned here -- not fetched next to the download.
- name: "Dockerfile lint (hadolint, #3320)"
home: true
run: |
set -euo pipefail
bin="$RUNNER_TEMP/hadolint"
curl -fsSL -o "$bin" https://github.com/hadolint/hadolint/releases/download/v2.14.0/hadolint-linux-x86_64
echo "6bf226944684f56c84dd014e8b979d27425c0148f61b3bd99bcc6f39e9dc5a47 $bin" | sha256sum -c -
chmod +x "$bin"
git ls-files -z | grep -zE '(^|/)Dockerfile[^/]*$' | xargs -0 "$bin" --config .hadolint.yaml
# #3501: fail CI on a real credential anywhere in FULL git history.
# scripts/check-git-secrets.py landed in 8cc55aa9 with nothing calling
# it -- `grep -rn gitleaks .github/workflows/` was empty -- so that
# commit's "fail CI on a real credential" claim was untrue. This row
# is what makes it true, and it is the only place gitleaks is on PATH.
#
# Its own lane rather than a row on the public-leak side: that check
# walks `git ls-files`, so it can only ever see the checked-out tree.
# A credential committed once and deleted in the next commit is
# still a credential, and this is the gate that sees it.
- name: "Full-history secret scan (gitleaks, #3501)"
home: true
run: |
set -euo pipefail
# The scan's entire claim is FULL history, and actions/checkout
# defaults to fetch-depth 1 -- so on a pull_request the
# merge-commit checkout hands gitleaks exactly one commit and the
# gate passes vacuously. That is the worst failure mode a secret
# scanner has: green, having measured nothing. Unshallow first,
# then prove the history is really there before believing a pass.
if [ "$(git rev-parse --is-shallow-repository)" = true ]; then
git fetch --unshallow --no-tags origin
fi
commits="$(git rev-list --count HEAD)"
if [ "$commits" -lt 100 ]; then
echo "::error::secret-scan: only $commits commit(s) under HEAD -- refusing to report a clean scan of a truncated history"
exit 1
fi
echo "secret-scan: scanning $commits commits of full history"
# Pinned version + sha256 verified before extraction, unpacked
# into $RUNNER_TEMP because the self-hosted honeypot-ci runner
# cannot write /usr/local/bin -- the constraint
# scripts/install-trivy.sh already works under, and the reason
# `home: true` is right here: checkout, this workflow's own
# setup-python, a download into a writable temp dir. No docker,
# no host state, nothing the metal cannot do.
# The install has to reach THIS step, not a later one. Each `run:`
# is its own process, and GITHUB_PATH is applied by the runner to
# *subsequent* steps -- so `scripts/install-gitleaks.sh` writing
# it (install-gitleaks.sh:123) left the very next line scanning
# with a PATH that never had gitleaks on it, which is the exit 2
# #3516's CI run reported. --print-path names the binary the pin
# verified; exporting its directory puts it on PATH for the rest
# of this block. The GITHUB_PATH write stays in the script for
# later steps -- it is right there, it is just not enough here.
export PATH="$(dirname "$(scripts/install-gitleaks.sh --print-path)"):$PATH"
# Prove the propagation before the scan, so a broken PATH fails
# as "gitleaks missing" next to the install, not as an opaque
# exit 2 from a scan that never ran.
command -v gitleaks
# check-git-secrets.py exits 2 when gitleaks fails TO RUN and 1
# when it finds something unallowlisted. Both are failures here
# and neither is softened: under `set -euo pipefail` an exit 2
# cannot be mistaken for a clean scan, which is the flagged vs
# unresolved split #2763 forced for trivy. Nothing in this step
# appends `|| true` or pipes the status away.
python3 scripts/check-git-secrets.py
# The gate's own regression suite. It lives in tests/docs/ and
# the docs-regression row above already collects it, but its
# four gitleaks-backed tests skip on a box with no binary -- and
# the docs row has none. This row does, so this is the only place
# they actually execute in CI; without it the suite would skip
# forever and quietly stop guarding the gate (#2981's shape).
python -m pip install pytest
python -m pytest tests/docs/test_3501_secret_scan_allowlist.py -q
- name: Shell syntax and high-severity ShellCheck
home: true
run: |
# #2389: use whatever shellcheck already exists -- the
# honeypot-ci host ships it for its runner user, so the
# self-hosted leg never touches sudo; only an image without
# the binary pays apt. No privilege class is added anywhere;
# the install path is unchanged where it was required.
command -v shellcheck >/dev/null 2>&1 || { sudo apt-get update && sudo apt-get install -y shellcheck; }
# A couple of sandbox/windows/packer/pxe/*.sh files are actually
# Python (named .sh to match their sibling shell scripts' calling
# convention, e.g. `packer trigger-callback... .sh`) -- filter to
# files whose own shebang names a shell, rather than assuming
# every *.sh is one, so bash -n/shellcheck don't choke on them.
shell_scripts=()
while IFS= read -r -d '' f; do
read -r shebang < "$f" || true
case "$shebang" in
'#!'*/sh|'#!'*/bash|'#!'*env\ sh|'#!'*env\ bash) shell_scripts+=("$f") ;;
esac
done < <(find . -name '*.sh' -type f -print0)
if [ "${#shell_scripts[@]}" -gt 0 ]; then
printf '%s\0' "${shell_scripts[@]}" | xargs -0 -n1 bash -n
printf '%s\0' "${shell_scripts[@]}" | xargs -0 shellcheck --severity=error
fi
- name: Validate Keycloak realm policy
home: true
run: ./arcane/home/honeypot-keycloak/keycloak/realm/validate.sh
# #982 phase 1 / #1040: validate.sh only checks structure/policy --
# this actually imports the realm into a real, disposable Keycloak +
# PostgreSQL the way the real fresh-install bootstrap path does, to
# catch what only a real import catches (e.g. a role description over
# Postgres's varchar(255) column limit crash-looped the container on
# every fresh install, undetected by any static check).
- name: Keycloak realm imports cleanly into a real Postgres
home: true
run: ./scripts/test-keycloak-realm-import.sh
# #977: proves the isolated oauth2-proxy gateway pattern every
# protected app (Kibana/EveBox/Arkime/TANNER/RevDeck/Dockge/Traefik)
# uses against a real disposable Keycloak + real oauth2-proxy --
# unauthenticated redirect target, forged-callback rejection, role
# enforcement, upstream network isolation, and gateway-outage
# fail-closed behavior.
- name: oauth2-proxy gateway pattern is resilient
home: true
run: ./scripts/test-oauth2-proxy-gateway-resilience.sh
# #982's PKCE+TOTP login and Keycloak-outage/key-rotation chaos
# suites, retired with the Go dashboard (#1659), restored for
# dashboard-next (#1661): both drive the REAL BFF build output
# (`node .output/server/index.mjs`, the same artifact the container
# runs) against a real disposable Keycloak importing the actual
# realm -- full authorization-code +PKCE logins through the realm's
# mandatory TOTP factor, persisted session-role verification
# (#1656), single-use codes + forged-state rejection, server-side
# logout revocation (#1094), first-login password reset (#1036),
# then outage-persistence / KC-down-logout / restart-recovery /
# signing-key-rotation chaos. One shared build feeds both via
# DASHBOARD_BFF_SKIP_BUILD.
- name: Set up Node for the dashboard OIDC suites
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: "22"
cache: npm
cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json
- name: Build dashboard-next BFF once for both OIDC suites
home: true
run: |
cd arcane/home/honeypot-dashboard/frontend-next
npm ci --no-audit --no-fund
npm run build
- name: "OIDC PKCE+TOTP login suite (port of #982)"
home: true
env:
DASHBOARD_BFF_SKIP_BUILD: "1"
run: ./scripts/test-dashboard-oidc-pkce-totp-login.sh
- name: "Keycloak outage/restart/rotation chaos suite (port of #982)"
home: true
env:
DASHBOARD_BFF_SKIP_BUILD: "1"
run: ./scripts/test-dashboard-oidc-chaos.sh
- name: GitHub-analysis publisher stays dry-run by default
home: true
run: |
# The one property this whole feature depends on (#74):
# GITHUB_PUBLISH_ENABLED unset must never reach publish-sample.sh,
# the only thing that can commit, zip, or push. WORK-LEDGER.md
# rule 7 requires this be a test, not a convention -- run it on
# every change, not just when analysis/github/ itself changes,
# since a change elsewhere (an env var default, a shared helper)
# could silently break the gate just as easily.
for t in analysis/github/tests/*.sh; do
echo "::group::$t"
bash "$t"
echo "::endgroup::"
done
- name: Threat-intel CIDR refresh script (#244)
home: true
run: analysis/threat-intel/tests/test_refresh_threat_cidrs.sh
# #2024: the Keycloak dump is the backup's most sensitive artifact,
# and under dash a `pg_dump | gzip` pipeline can only surface
# gzip's exit status -- the exact silent-truncation class behind
# #1413's post-mortem. Drives the real script with docker stubbed
# on PATH: a mid-stream dead pg_dump (and an rc-lost footer-less
# one) must keep no artifact, a healthy dump must be byte-exact
# through gzip, and a missing container must stay quiet. Runs in
# plain sh, because dash compliance IS the bug. #2025 extends the
# same suite to retention running before failure-prone steps, the
# manual-run umask window, and no empty stamp on failed validation.
- name: "backup-honeypot dump discipline, retention order and umask (#2024, #2025)"
home: true
run: sh analysis/tests/test_backup_honeypot.sh
# The four real-Elasticsearch suites. #2389 left them unflagged
# because without a container engine they print "SKIP: docker
# daemon is not reachable" and go green while testing nothing;
# the runner user now drives docker (#2565), so on the box they
# exercise their actual ES containers -- which is exactly the
# difference between a relocated failure and a relocated no-op.
- name: honeypot-30d ILM policy rolls over and deletes (#585)
home: true
run: analysis/tests/test_honeypot_ilm_rollover.sh
- name: geoip-honeypot ingest pipeline (#563)
home: true
run: analysis/tests/test_geoip_pipeline.sh
- name: dionaea-incidents index template (#565)
home: true
run: analysis/tests/test_dionaea_incidents_index.sh
- name: conpot persona extraction in geoip-honeypot pipeline (#567)
home: true
run: analysis/tests/test_conpot_persona_pipeline.sh
# #789's sensor event-kind coverage audit used to run here. It is
# gone rather than disabled, per #1665: it diffed sensor source
# against dashboard/classify.go's per-sensor switch, and #1659
# deleted that file with the Go dashboard. Nothing replaced the
# switch -- the geoip-honeypot pipeline copies honeypot.category
# through when a sensor supplies one and does nothing when it does
# not, so there is no case table left to diff against.
#
# The concern itself survives and got worse, so the script was
# rewritten to measure it from the data instead: 20 of 22 sensors
# label no events at all, 2% coverage overall. That needs a
# populated Elasticsearch, which this job does not have, so it is
# an operational audit now -- scripts/audit-sensor-event-coverage.py
# on the homeserver, not a CI step.
- name: Shared GPU job queue
home: true
run: analysis/gpu-queue/test_gpu_queue.py
# #1971: shared ES consume idioms. The canonical suite walks the
# vendored-copy registry (byte-for-byte vs every consumer copy,
# the gpu_queue.py discipline), drives the Python engine through
# the hand-computed cross-language fixture stream, and pins the
# inclusive-gte query shape -- the exact #168 boundary semantics.
# The Go twin of that same fixture stream runs in the go-test
# job (attacker-identity-worker's TestParityFixtures); disagreement
# between the two engines fails whichever side drifts.
- name: Shared ES consume patterns (#1971)
home: true
run: python3 analysis/es-consume/tests/test_es_consume.py
- name: Payload/artifact dedupe (#481/#528)
home: true
run: |
# test_dedupe_payloads.py and test_cdc_dedup_prototype.py had
# zero CI coverage before this -- no workflow referenced either
# by name. Closing that gap alongside #528's own new modules
# rather than leaving it, since it's the same area of the tree
# and cheap to fix now that it's been noticed.
python3 analysis/tests/test_dedupe_payloads.py -v
python3 analysis/tests/test_cdc_dedup_prototype.py -v
python3 analysis/tests/test_procmon_cdc_store.py -v
python3 analysis/tests/test_archive_diagnostics.py -v
# #1984: analyze.py prints attacker-controlled fields into an
# analyst's terminal -- the one place honeypot content reaches a
# human interactively. Every control character must leave
# print_table inert (<0xNN> spellings) while clean reports stay
# byte-identical; this is what stops an ESC payload in a password
# from repositioning or rewriting the reader's terminal. #1985
# extends the same suite to the stats-hygiene batch (silent
# malformed-line skips, dead parameters, total undercount, the
# category catch-all, multipot VNC login counting).
- name: "analyze.py sanitisation and stats hygiene (#1984, #1985)"
home: true
run: python3 analysis/tests/test_analyze.py -v
- name: Agent-intrusion synthetic replay corpus (#154 phase 1)
home: true
run: |
# #154's own acceptance criteria: "CI verifies schemas, safe
# fixtures, and replay expectations." Runs both the standalone
# validator (the same one a human runs by hand when editing the
# corpus) and its test suite (which additionally proves the
# validator catches deliberately-broken input, not just that
# today's corpus happens to pass).
python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/validate_corpus.py
python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_corpus.py -v
- name: Agent-intrusion decode/correlate pipeline (#154 phase 2)
home: true
run: |
# Bounded, non-executing base64/gzip/zlib/xor decoder plus
# multi-part message-chunk reassembly, proven against the real
# corpus above -- not just hand-built fixtures -- per that
# corpus's own README ("expected_findings is ground truth for
# phase 2's decoder").
python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_decode_correlate.py -v
- name: Agent-intrusion campaign correlator (#154 phase 2)
home: true
run: |
# Second half of phase 2 ("correlate events across sensors into
# one campaign timeline using stable IDs and time windows") --
# union-find over session/IP/channel identifiers, proven against
# the real corpus's own multi-hop case (a C2 channel ID bridging
# two otherwise-unconnected actor identities into one campaign).
python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_campaign_correlator.py -v
- name: Agent-intrusion criticality rules (#154 phase 3)
home: true
run: |
# Deterministic, structural rules (never reads the corpus's own
# phase/should_escalate labels -- see criticality_rules.py's own
# module docstring) proven against every one of the 27 real
# corpus events matching its own independently-established
# ground truth, plus the campaign-level severity scoring proven
# against the real merged 8-event campaign reaching "critical".
python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_criticality_rules.py -v
- name: Agent-intrusion campaign worker (#154 phase 5)
home: true
run: |
# Wires decode_correlate/campaign_correlator/criticality_rules
# against Elasticsearch-shaped data for real (a hand-rolled fake
# ES client, no network) -- including an end-to-end run of the
# real corpus through the worker's own correlate-then-score call
# sequence, proving the pipeline wiring itself is correct, not
# just each module independently.
python3 -m pip install -r arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/requirements.txt
python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_worker.py -v
- name: Sandbox VNC bridge read-only enforcement (#805)
home: true
run: python3 sandbox/windows/vnc-bridge/tests/test_server.py
- name: Vendored YARA corpus is intact and loadable
home: true
run: |
# yara(1) refuses to start on a corpus with one bad rule rather than
# skipping it, so a rule file edited in place or dropped from
# index.yar takes the whole scanner down silently.
#
# Run in the scanner's own base image rather than against the runner's
# apt yara: "does it compile" is only a useful answer from the
# compiler that will actually load it, and the two versions differ.
# The image is read from the Dockerfile so a base bump moves CI too.
image="$(awk '/^FROM /{print $2; exit}' arcane/home/honeypot-payload-analysis/analysis/yara/Dockerfile)"
echo "validating against $image"
docker run --rm -e PYTHONDONTWRITEBYTECODE=1 -v "$PWD:/w" -w /w "$image" sh -c '
apk add --no-cache bash git yara >/dev/null
yara --version
scripts/check-yara-corpus.sh
yara -w arcane/home/honeypot-payload-analysis/analysis/yara/rules/index.yar /dev/null
arcane/home/honeypot-payload-analysis/analysis/yara/tests/test_sync_yara.sh
'
- name: Rev-eng benchmark corpus is reproducible and safe (#159)
home: true
run: |
# Runs in the corpus's own documented build environment
# (debian:trixie-slim, matching corpus/README.md's "Rebuilding"
# section exactly), not the runner's own toolchain — the whole
# point is verifying that THIS environment reproduces the
# committed manifest.json byte-for-byte, which only means
# something if it is the same environment the corpus claims
# provenance against.
# -e PYTHONDONTWRITEBYTECODE=1: the container runs as root against
# this bind-mounted workspace, and root-owned __pycache__ left
# behind here breaks actions/checkout's clean step for every later
# job on this runner (EACCES on unlink). ci_verify.sh exports it
# too -- belt and braces, since it is equally true of any other
# caller that runs it as root in a mount.
docker run --rm -e PYTHONDONTWRITEBYTECODE=1 -v "$PWD:/w" -w /w debian:trixie-slim bash -c '
analysis/ghidra/benchmarks/corpus/ci_verify.sh
'
- name: Validate home and VPS Compose
home: true
run: |
cp .env.example .env
cp vps/.env.example vps/.env
docker compose -f docker-compose.yml config --quiet
# honeypot-init is a separate Dockge stack (#111) that attaches to
# APIARY's honeynet network and two of its named volumes by
# external reference; config validation does not require those to
# actually exist, so no APIARY setup is needed here either.
docker compose -f arcane/home/honeypot-init/compose.yml config --quiet
# The modernization-port ("next" profile) services validate
# along with the rest — config resolves every service
# regardless of profile gating, profiles only affect `up`.
# Split apart since #1622: honeypot-dashboard-backend
# resolves backend-service:8081 against the sibling stack by
# bare service-name DNS at runtime (shared honeynet
# network), which config validation doesn't need to prove.
docker compose -f arcane/home/honeypot-dashboard/compose.yml config --quiet
docker compose -f arcane/home/honeypot-dashboard-backend/compose.yml config --quiet
docker compose -f vps/docker-compose.yml config --quiet
# The sandbox gateway is never started here — it answers live malware
# on an isolated bridge that does not exist on a CI runner. Validating
# it still matters: run_sample.py shells out to this file mid-detonation,
# and a syntax error would surface as a failed run on the analysis host.
docker compose -f docker-compose.sandbox.yml config --quiet
# Same reasoning for the analysis host: not started here, but
# install-analysis-host.sh runs it unattended on a box with a GPU and
# captured malware on it. Both the base file and the GPU overlay,
# because the overlay is only ever used on top of the base one.
docker compose -f analysis/ghidra/docker-compose.ghidra.yml config --quiet
docker compose -f analysis/ghidra/docker-compose.ghidra.yml \
-f analysis/ghidra/docker-compose.ghidra.gpu.yml config --quiet
# #66's base file is synthetic and network-isolated. Validate both
# #83 overlays, but start neither worker nor model in CI.
docker compose -f llm-worker/docker-compose.yml config --quiet
docker compose -f llm-worker/docker-compose.yml \
-f llm-worker/docker-compose.synthetic-canary.yml config --quiet
docker compose -f llm-worker/docker-compose.yml \
-f llm-worker/docker-compose.production-session-canary.yml config --quiet
docker compose -f llm-worker/docker-compose.yml \
-f llm-worker/docker-compose.captured-data.yml config --quiet
# #2225: docker-compose.captured-data-deploy.yml (#1751) is the
# entrypoint live redeploys authorized under #83 actually run --
# `include:`, not a stacked `-f base -f overlay` pair, so
# neither line above proves anything about it. Unvalidated, a
# bad include path or an invalid key sails through both and
# first surfaces as a failed (or silently mis-scoped) live
# redeploy, which is how #1751's incident happened.
docker compose -f llm-worker/docker-compose.captured-data-deploy.yml config --quiet
# Syntactic validity alone isn't enough: stripping an include
# (e.g. losing docker-compose.captured-data.yml) still resolves
# to valid, synthetic-only config. So does keeping both includes
# but listing them as two separate `include:` entries instead of
# one list-valued `path:` -- separate entries are independent
# models, the first to claim services.llm-worker keeps it, and
# the base's deliberate `ES_HOST: ''` wins. That is what this
# check caught on its first run. Resolve exactly the one file
# the live deploy points at -- adding `-f base -f overlay`
# alongside it would supply the authorization from the command
# line and prove nothing about the include chain. Assert the
# resolved model still carries what this file exists to record.
resolved="$(docker compose -f llm-worker/docker-compose.captured-data-deploy.yml config --format json)"
es_host="$(jq -r '.services["llm-worker"].environment.ES_HOST // empty' <<<"$resolved")"
if [ -z "$es_host" ]; then
echo "::error::captured-data-deploy.yml resolved with empty ES_HOST -- the captured-data authorization overlay is not being applied (#2225)"
exit 1
fi
for net in honeypot-llm-data honeypot-llm; do
if ! jq -e --arg n "$net" '.networks | to_entries[] | select(.value.name == $n)' <<<"$resolved" >/dev/null; then
echo "::error::captured-data-deploy.yml resolved without the $net network attached -- the captured-data authorization overlay is not being applied (#2225)"
exit 1
fi
done
for target in /payloads/cowrie /payloads/scripts; do
if ! jq -e --arg t "$target" '.services["llm-worker"].volumes[]? | select(.target == $t and .read_only == true)' <<<"$resolved" >/dev/null; then
echo "::error::captured-data-deploy.yml resolved without a read-only mount at $target -- the captured-data authorization overlay is not being applied (#2225)"
exit 1
fi
done
# #2981: these used to be named individually, the same shape #2878
# fixed for sandbox/windows/orchestrate and sandbox/ghosts/orchestrate
# -- test_install_ci_runner_instances.py sat unwired and red for days
# because nothing ran it, and two more files in this directory
# (test_arcane_retry_failed_sync.py, test_audit_sensor_event_coverage.py)
# were never run by CI at all. Discovery instead, so the next
# test_*.py dropped in scripts/tests/ is picked up without a
# workflow edit.
- name: scripts/tests suite
home: true
run: python3 -m unittest discover -s scripts/tests -p 'test_*.py' -v
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: "3.13"
# Needed below by the two dashboard-binary Keycloak integration tests
# (`go run .` against the real dashboard) -- same version floor as the
# go-test job, see that job's own go-version comment. Every matrix
# entry pays this setup cost even though only two of them use it --
# simpler and more robust than threading a per-entry needs-go flag
# through the matrix, and setup-go is fast.
- uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0
with:
go-version: "1.26.x"
cache-dependency-path: "**/go.sum"
# Cache only on the GitHub-hosted path -- mirror of this job's
# own runs-on expression. On the homeserver runners the module
# cache is already on disk and setup-go's restore cannot unpack
# over it; see the go-fmt job for the full account.
cache: ${{ !(matrix.home == true && needs.ci-target.outputs.homeserver == 'true') }}
- name: ${{ matrix.name }}
shell: bash
run: ${{ matrix.run }}
# #3319: `always()` because the pytest rows emit their JUnit XML on the
# passing run too, and that is the record of what ran. `ignore` on the
# missing-file case because this job is a ~65-row matrix and most rows
# write no report at all -- a warning per row would bury the one that
# matters. The matrix row name is in the artifact name so the two
# pytest rows don't collide in the run's artifact list (upload-artifact
# v4 refuses two uploads sharing a name).
#
# Two things this step had wrong, both fixed here rather than papered
# over with `overwrite: true`:
#
# 1. `path: .ci-artifacts/` uploaded NOTHING. upload-artifact v4.4+
# skips hidden paths by default, and a directory whose own name
# starts with `.` is hidden, so the action reported "No files were
# found with the provided path: .ci-artifacts/" and, under
# if-no-files-found: ignore, said it quietly. Verified on a green
# run of the ml-worker row: 330 passed, pytest wrote
# .ci-artifacts/ml-worker-junit.xml, and the upload step below it
# still found no files. The two JUnit files this retention was
# added for have therefore never been retained. The path below
# names them explicitly, which both fixes that and makes the
# contents a closed set: those two files, written by the two rows'
# own --junitxml flags, and nothing else that happens to be sitting
# in a workspace-relative directory.
# 2. The name is per-(run, attempt), so no upload here can collide
# and no earlier run on the same ref can poison this one's name.
# See the frontend-next twin's upload for the full argument.
#
# One trap left standing on purpose: GitHub rejects `/` in an artifact
# name, and six row names contain one ("Healthcheck/autoheal ...",
# "Zeek lab/prod ...", "Payload/artifact dedupe ...", "scripts/tests
# suite", and two more). They are harmless today because they write no
# report, and upload-artifact only reaches the name when it has files
# to upload. A row that both writes a report AND has a `/` in its name
# will fail its upload -- and there is no expression-level way to
# sanitise a free-text field, so fixing that means renaming the row.
- name: Upload test results
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: scripts-${{ matrix.name }}-results-${{ github.run_id }}-${{ github.run_attempt }}
path: |
.ci-artifacts/ml-worker-junit.xml
.ci-artifacts/auth-events-worker-junit.xml
if-no-files-found: ignore
retention-days: 7
scripts-and-compose-complete:
name: Scripts and Compose
if: always()
needs: [scripts-and-compose]
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
sparse-checkout: scripts/ci-lane-summary.py
sparse-checkout-cone-mode: false
- name: Fail if any check failed
if: needs.scripts-and-compose.result != 'success'
run: exit 1
# #3319: the matrix is a single job, so its own result is all there is
# to report -- a row that was path-filtered away is invisible here and
# says nothing about the run. Rendered anyway so the aggregate job's
# summary is not empty on a red run.
- name: Lane summary
if: always()
env:
NEEDS: ${{ toJSON(needs) }}
run: python3 scripts/ci-lane-summary.py --title "Scripts and Compose" <<<"$NEEDS"
# #3329: no AI/assistant attribution in PR commits, title or body. A
# policy that held only by convention -- main collected 461 attributed
# commits between 2026-08-01 and 2026-09-26. Reads the PR from the event
# file and its commits from the API, so no PR text passes through a shell
# and no deep fetch is needed. pull_request only; quality-gate accepts its
# skip on every other event.
ai-attribution:
name: No AI attribution in PR metadata (#3329)
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
sparse-checkout: scripts/check-ai-attribution.py
sparse-checkout-cone-mode: false
- name: Check commits, title and body
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: python3 scripts/check-ai-attribution.py
# #3311: the one required status check for this workflow. main's ruleset
# requires a context that reports on every PR; the jobs above come in
# homeserver/GitHub-hosted pairs whose names change with the executor, so
# none of them individually is a stable context. Mirrors
# go-modules-complete's rule for every pair: the homeserver twin succeeded,
# or it was skipped and its GitHub-hosted twin succeeded. A new job pair
# must be added to both `needs:` and PAIRS below, or it is not gated.
quality-gate:
name: Quality gate
if: always()
needs:
- ci-target
- public-safety
- public-safety-cloud
- design-lab-readonly
- design-lab-readonly-cloud
- go-modules-complete
- frontend-next
- frontend-next-cloud
- frontend-next-browser
- frontend-next-browser-cloud
- backend-service
- backend-service-cloud
- vendored-theme
- vendored-theme-cloud
- scripts-and-compose-complete
- ai-attribution
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- name: Fail unless every gated job, or its fallback twin, succeeded
env:
NEEDS: ${{ toJSON(needs) }}
EVENT: ${{ github.event_name }}
shell: bash
run: |
result() { jq -r --arg j "$1" '.[$j].result' <<<"$NEEDS"; }
fail=0
single() {
local r; r="$(result "$1")"
[[ "$r" == "success" ]] && return 0
echo "::error::$1: expected success, got $r"; fail=1
}
pair() {
local hs cloud; hs="$(result "$1")"; cloud="$(result "$1-cloud")"
[[ "$hs" == "success" ]] && return 0
[[ "$hs" == "skipped" && "$cloud" == "success" ]] && return 0
echo "::error::$1: expected success or skip+fallback-success; got $hs / $cloud"; fail=1
}
single ci-target
# #3329: pull_request-only by design, so a skip is fine elsewhere.
if [[ "$EVENT" == "pull_request" ]]; then
single ai-attribution
elif [[ "$(result ai-attribution)" != "skipped" ]]; then
single ai-attribution
fi
single go-modules-complete
single scripts-and-compose-complete
for p in public-safety design-lab-readonly frontend-next frontend-next-browser backend-service vendored-theme; do
pair "$p"
done
exit "$fail"
# #3319: the run's lane table, which is the part a human actually reads.
# `if: always()` so it renders on the failing run that needs it, and it
# runs even when the gate above exits 1 -- reporting is not gating.
#
# ci-target is excluded from --allow-skip deliberately: on a run where
# the router itself never reported, every routed pair below is a
# no-executor skip, and that is a red run, not an accounted one.
- name: Lane summary
if: always()
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
sparse-checkout: scripts/ci-lane-summary.py
sparse-checkout-cone-mode: false
- name: Render lane summary
if: always()
env:
NEEDS: ${{ toJSON(needs) }}
EVENT: ${{ github.event_name }}
run: |
# ai-attribution is pull_request-only by design (#3329), so its skip
# is accounted for on every other event. The gate above already
# encodes that same rule; this states the reason in the report.
allow=()
if [[ "$EVENT" != "pull_request" ]]; then
allow+=(--allow-skip "ai-attribution:pull_request only (#3329)")
fi
python3 scripts/ci-lane-summary.py --title "Quality" "${allow[@]}" <<<"$NEEDS"