Repository navigation
ci(3501): arm the base-image CVE gate and move the pins that fix it #4586
Workflow file for this run
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| name: Quality | |
| on: | |
| push: | |
| branches: [main] | |
| pull_request: | |
| workflow_dispatch: | |
| inputs: | |
| force_github_hosted: | |
| description: >- | |
| Send the homeserver-eligible jobs to GitHub-hosted runners even | |
| when the honeypot-ci runner is online -- e.g. while servicing the | |
| homeserver, or when a suspect result needs ruling out as | |
| runner-specific. Homeserver-first routing itself is automatic for | |
| trusted runs (push-to-main, workflow_dispatch); pull_request stays | |
| GitHub-hosted unless the repository variable CI_HOMESERVER_PRS | |
| opts same-repo PRs in. | |
| type: boolean | |
| default: false | |
| # #3313: workflow level defaults to none; every job below declares what it | |
| # spends. Each of the 20 testing jobs runs actions/checkout, so each carries | |
| # its own contents: read. actions: write, which used to sit here and so | |
| # reached all 20 of them, is now on the ci-target router job alone -- that one | |
| # grant is what dispatches the ci-heartbeat canary with GITHUB_TOKEN, and a | |
| # called reusable workflow can never exceed the caller's envelope, so | |
| # under-granting it startup-fails the whole run as "Invalid workflow file" | |
| # (containers.yml, security.yml and pages.yml grant the same way for the same | |
| # reusable ci-router.yml). | |
| permissions: {} | |
| concurrency: | |
| # #2753: this used to be group: quality-${{ github.ref }} with | |
| # cancel-in-progress off for push-to-main, on the theory that turning | |
| # cancellation off makes main pushes serialize -- wait for the previous | |
| # run before starting. That is not what a GitHub Actions concurrency group | |
| # does: with cancel-in-progress:false, GitHub keeps only the currently | |
| # running run and the single most recent pending one per group, and | |
| # cancels -- with ZERO jobs ever started -- every push queued in between. | |
| # Reproduced live 2026-08-31: ten consecutive main pushes (including the | |
| # #2724/#2725 merge commits) got a Quality run that was `cancelled` with | |
| # `.jobs | length == 0`, not superseded by a later run that re-verified | |
| # the same tree; only the first and last commit in the batch were ever | |
| # actually checked. That is exactly the #2317/#2347 failure class this | |
| # config exists to prevent (cbb5ef3f landed non-compiling Rust and its own | |
| # quality run never finished) -- the shared per-ref group didn't close | |
| # that hole, it just changed which commits fall through it. | |
| # | |
| # Fix: give every push-to-main commit its own group (keyed on github.sha | |
| # instead of github.ref), so no two main commits can ever share a group | |
| # and cancel one another. This trades true FIFO serialization -- which | |
| # GitHub's concurrency primitive cannot express at all -- for the | |
| # guarantee #2317/#2347 actually needed: every commit that reaches `main` | |
| # gets its own Quality run and none of them get cancelled by a sibling. | |
| # pull_request runs keep the previous shared-per-ref/cancel-on behavior | |
| # unchanged -- a superseded PR run is genuinely for a commit nobody cares | |
| # about anymore. | |
| group: quality-${{ startsWith(github.ref, 'refs/heads/main') && github.sha || github.ref }} | |
| cancel-in-progress: ${{ !startsWith(github.ref, 'refs/heads/main') }} | |
| jobs: | |
| # ------------------------------------------------------------------------- | |
| # Executor routing ("homeserver first, GitHub-hosted fallback") | |
| # | |
| # ci-target's trust-gate + heartbeat-liveness decision procedure is | |
| # shared with every other homeserver-eligible workflow via the reusable | |
| # .github/workflows/ci-router.yml -- its header carries the full | |
| # rationale (also documented in docs/CI-CD.md's "Executor routing" | |
| # section), so it isn't repeated here. Two things stay specific to | |
| # Quality: | |
| # | |
| # - force_github_hosted (the workflow_dispatch input above) forces the | |
| # fallback direction manually -- e.g. while servicing the | |
| # homeserver, or when a suspect result needs ruling out as | |
| # runner-specific. | |
| # - Unlike containers.yml/security.yml/pages.yml's single | |
| # executor-agnostic job, every home-executable check below ships as | |
| # a PAIR of conditional jobs that take turns off ci-target's answer | |
| # -- exactly one twin does real work per run, the other reports | |
| # skipped. | |
| # | |
| # #2565: the no-docker CI-runner design is retired. The runner's | |
| # dedicated user now carries a docker-group membership (the same grant | |
| # github-deploy-runner always had) and the host ships node 22, the | |
| # redis-server binary, shellcheck and the playwright 1.62.1 chromium | |
| # library set for its user -- preinstalled, because the user still has | |
| # no sudo by design and every sudo-apt path in a check would otherwise | |
| # be a guaranteed relocation failure. Everything Quality runs is | |
| # therefore homeserver-eligible: frontend-next and frontend-next-browser | |
| # ship as pairs like the families above (the browser twin swaps | |
| # `--with-deps` and the apt install for presence checks -- deps live on | |
| # the host now, and a missing one must fail loudly, not silently | |
| # apt-install), and every docker-bound scripts-and-compose row is | |
| # flagged `home: true` (#2389's docker-less audit criterion graduated | |
| # to "docker or no engine needed"). | |
| # ------------------------------------------------------------------------- | |
| ci-target: | |
| name: Pick CI executor | |
| uses: ./.github/workflows/ci-router.yml | |
| permissions: | |
| contents: read | |
| actions: write | |
| with: | |
| ci_homeserver_prs: ${{ vars.CI_HOMESERVER_PRS || '' }} | |
| force_github_hosted: ${{ inputs.force_github_hosted || '' }} | |
| # Naming rule for every pair below: the homeserver twin carries the | |
| # canonical check name (it is what normally runs); its GitHub-hosted | |
| # fallback twin appends "(GitHub-hosted)" so a degraded day reads | |
| # honestly in the checks list instead of masquerading as business as | |
| # usual. | |
| # | |
| # timeout-minutes exists ONLY on homeserver twins (plus the router): a | |
| # stuck pickup wedges a real box nobody reboots promptly, whereas | |
| # GitHub-hosted runners have hard platform timeouts already. Values sit | |
| # far above each check's observed runtime so they fire only on hangs. | |
| public-safety: | |
| name: Public repository safety | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 10 | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 | |
| with: | |
| python-version: "3.13" | |
| - run: python scripts/check-public-leaks.py | |
| - run: python scripts/validate-oidc-redirects.py | |
| - run: bash scripts/check-ghosts-vendored-egress.sh | |
| public-safety-cloud: | |
| name: Public repository safety (GitHub-hosted) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 | |
| with: | |
| python-version: "3.13" | |
| - run: python scripts/check-public-leaks.py | |
| - run: python scripts/validate-oidc-redirects.py | |
| - run: bash scripts/check-ghosts-vendored-egress.sh | |
| design-lab-readonly: | |
| # #1828: the design lab serves variants against the real captured-data | |
| # Elasticsearch, so its read-only guarantee is a safety property, not a | |
| # convenience. The test drives the actual harness against a recording | |
| # stand-in backend and fails if a write ever reaches it. | |
| name: Design lab is read-only | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 10 | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 | |
| with: | |
| node-version: "22" | |
| - run: node --test branding/design-lab/lab.test.mjs | |
| design-lab-readonly-cloud: | |
| name: Design lab is read-only (GitHub-hosted) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 | |
| with: | |
| node-version: "22" | |
| - run: node --test branding/design-lab/lab.test.mjs | |
| go-fmt: | |
| name: Go formatting | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 15 | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0 | |
| with: | |
| # dicompot's vendored github.com/nsmfoo/dicompot (#413) declares go | |
| # 1.26.2 in its own go.mod -- the module graph forces that as the | |
| # floor for every module in this repo, not just dicompot's. | |
| go-version: "1.26.x" | |
| # The homeserver runners keep GOMODCACHE/GOCACHE on disk at | |
| # /opt/github-ci-runner (shared by all four runner instances), so | |
| # setup-go's default module cache is pure loss here -- and worse | |
| # than loss. Go writes module directories mode 0555, so the | |
| # restore's `tar -x` cannot recreate a single already-present | |
| # file: every job downloaded ~466 MB, spent ~45s failing to | |
| # unpack it ("Cannot open: File exists" x thousands, tar exit 2), | |
| # reported "Cache is not found", then re-uploaded the whole | |
| # module cache from the post step. Six same-key 465 MB copies | |
| # piled up in one evening and pushed the repo's 10 GB Actions | |
| # cache over quota, evicting the buildkit layer caches. Same | |
| # reasoning as the Rust job below, which carries no | |
| # actions/cache block for exactly this reason. | |
| cache: false | |
| - name: Check formatting | |
| shell: bash | |
| run: | | |
| files="$(find . -path '*/vendor' -prune -o -name '*.go' -type f -print)" | |
| test -z "$(gofmt -l $files)" || { gofmt -l $files; exit 1; } | |
| go-fmt-cloud: | |
| name: Go formatting (GitHub-hosted) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0 | |
| with: | |
| # Same floor as the homeserver twin above -- see its comment. | |
| go-version: "1.26.x" | |
| cache-dependency-path: "**/go.sum" | |
| - name: Check formatting | |
| shell: bash | |
| run: | | |
| files="$(find . -path '*/vendor' -prune -o -name '*.go' -type f -print)" | |
| test -z "$(gofmt -l $files)" || { gofmt -l $files; exit 1; } | |
| # Tests stay a parallel matrix on GitHub-hosted (the shape that has always | |
| # run there), but consolidate into ONE sequential loop on the homeserver: | |
| # 20 matrix entries would queue behind a single runner machine anyway, | |
| # paying checkout+setup twenty times over for zero parallelism gained -- | |
| # while a warm GOCACHE in the runner user's persistent HOME makes the | |
| # sequential loop cheaper than the sum of its parts after the first run. | |
| # Quality-homeserver.yml ran this exact loop shape successfully until it | |
| # was superseded by these pairs. | |
| go-test-homeserver: | |
| name: Test (all modules) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 60 | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0 | |
| with: | |
| go-version: "1.26.x" | |
| # The homeserver runners keep GOMODCACHE/GOCACHE on disk at | |
| # /opt/github-ci-runner (shared by all four runner instances), so | |
| # setup-go's default module cache is pure loss here -- and worse | |
| # than loss. Go writes module directories mode 0555, so the | |
| # restore's `tar -x` cannot recreate a single already-present | |
| # file: every job downloaded ~466 MB, spent ~45s failing to | |
| # unpack it ("Cannot open: File exists" x thousands, tar exit 2), | |
| # reported "Cache is not found", then re-uploaded the whole | |
| # module cache from the post step. Six same-key 465 MB copies | |
| # piled up in one evening and pushed the repo's 10 GB Actions | |
| # cache over quota, evicting the buildkit layer caches. Same | |
| # reasoning as the Rust job below, which carries no | |
| # actions/cache block for exactly this reason. | |
| cache: false | |
| - name: Test every Go module | |
| shell: bash | |
| run: | | |
| while IFS= read -r module; do | |
| echo "::group::${module%/go.mod}" | |
| (cd "${module%/go.mod}" && go test ./...) | |
| echo "::endgroup::" | |
| done < <(find . -path '*/vendor' -prune -o -name go.mod -type f -print | sort) | |
| go-test-cloud: | |
| name: Test (${{ matrix.module }}) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| # Every go.mod in the repo (excluding vendor/) -- kept as an explicit | |
| # list rather than a dynamic discover-then-matrix job so a new | |
| # module's tests actually run in CI the moment its go.mod lands, | |
| # instead of silently needing a separate PR to register it here too. | |
| # `find ... | sort` in this job's own history is the source of | |
| # truth if this list and the repo ever drift. | |
| module: | |
| - arcane/home/honeypot-attacker-identity-worker/attacker-identity-worker | |
| - arcane/home/honeypot-canarytokens/canarytokens-adapter | |
| - arcane/home/honeypot-canarytokens/canarytokens-http-router | |
| - arcane/home/honeypot-cisco-asa-honeypot/cisco-asa-honeypot | |
| - arcane/home/honeypot-citrix-honeypot/citrix-honeypot | |
| - arcane/home/honeypot-correlator-worker/correlator-worker | |
| - arcane/home/honeypot-cowrie/honeyfs-implant | |
| - arcane/home/honeypot-dicompot/dicompot | |
| - arcane/home/honeypot-dionaea/tftp-relay | |
| - arcane/home/honeypot-dnp3/dnp3-honeypot | |
| - arcane/home/honeypot-dns-honeypot/dns-honeypot | |
| - arcane/home/honeypot-endlessh/endlessh-honeypot | |
| - arcane/home/honeypot-galah/galah-llm-broker | |
| - arcane/home/honeypot-http/http-honeypot | |
| - arcane/home/honeypot-multipot/multipot | |
| - arcane/home/honeypot-payload-inventory-worker/payload-inventory-worker | |
| - arcane/home/honeypot-rdp-honeypot/rdp-honeypot | |
| - arcane/home/honeypot-sonicwall-sma/sonicwall-sma-honeypot | |
| - arcane/home/honeypot-utilities/reporter | |
| - portbridge | |
| - vps/portbridge | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0 | |
| with: | |
| go-version: "1.26.x" | |
| cache-dependency-path: "${{ matrix.module }}/go.sum" | |
| - run: go test ./... | |
| working-directory: ${{ matrix.module }} | |
| go-modules-complete: | |
| name: Go formatting and tests | |
| if: always() | |
| needs: [go-fmt, go-fmt-cloud, go-test-homeserver, go-test-cloud] | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| # #3319: only the lane-summary script, same sparse shape ai-attribution | |
| # uses. The report is written to the step summary and is a pure function | |
| # of `needs`, so nothing else from the tree is needed here. | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| sparse-checkout: scripts/ci-lane-summary.py | |
| sparse-checkout-cone-mode: false | |
| - name: Fail unless each executor pair produced exactly one success | |
| env: | |
| FMT_HS: ${{ needs.go-fmt.result }} | |
| FMT_CLOUD: ${{ needs.go-fmt-cloud.result }} | |
| TEST_HS: ${{ needs.go-test-homeserver.result }} | |
| TEST_CLOUD: ${{ needs.go-test-cloud.result }} | |
| shell: bash | |
| run: | | |
| # Each pair is exclusive by construction (mutually exclusive `if:`), | |
| # so the healthy states are: primary succeeded, or primary was | |
| # skipped because routing chose fallback AND fallback succeeded. | |
| # Anything else -- both skipped, both ran, a failure, a cancellation | |
| # mid-flight on main-bound code -- fails loudly right here rather | |
| # than letting a silently-skipped tier read as green. | |
| fail=0 | |
| expect() { | |
| local label="$1" primary="$2" fallback="$3" | |
| if [[ "$primary" == "success" ]]; then return 0; fi | |
| if [[ "$primary" == "skipped" && "$fallback" == "success" ]]; then return 0; fi | |
| echo "::error::${label}: expected success or skip+fallback-success; got ${primary} / ${fallback}" | |
| return 1 | |
| } | |
| expect "go formatting" "$FMT_HS" "$FMT_CLOUD" || fail=1 | |
| expect "go tests (matrix)" "$TEST_HS" "$TEST_CLOUD" || fail=1 | |
| exit "$fail" | |
| # #3319: the same pair results as a table, so a reader can see which | |
| # executor ran and which twin the router skipped without opening the | |
| # log. Independent of the exit code above -- this reports, it does not | |
| # gate, so it still renders on a red aggregate job. | |
| - name: Lane summary | |
| env: | |
| NEEDS: ${{ toJSON(needs) }} | |
| run: python3 scripts/ci-lane-summary.py --title "Go formatting and tests" <<<"$NEEDS" | |
| # Homeserver-first as a pair (#2565): the tier's only special | |
| # requirement is docker (lockfile check) plus node 24 via setup-node -- | |
| # the runner user now carries the docker-group membership, and | |
| # setup-node works identically on the self-hosted runner (its | |
| # node/npm caches simply persist in the runner's _work/_tool and HOME | |
| # instead of the platform cache). Steps stay byte-identical across the | |
| # twins so a red result means the same thing wherever it ran; | |
| # timeout-minutes rides only on the homeserver twin per the pair | |
| # convention. | |
| frontend-next: | |
| name: Dashboard frontend (next) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 45 | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| # #3331: run on the node major the image actually ships, read out of the | |
| # image's own FROM line instead of written here a second time. | |
| # | |
| # Every step in this job used to run under setup-node's node-version "24" | |
| # while the Dockerfile builds and serves on node:22-alpine, so the | |
| # typecheck, the unit tests, the production build and the generated | |
| # route-tree diff had never once executed on the node that ships the | |
| # artifact. The only Node 22 coverage in the whole workflow was the | |
| # #1816 lockfile-install step below. A gate that cannot see the runtime it | |
| # is gating is not a gate. | |
| # | |
| # Measured on the Dockerfile's own image before moving the pin | |
| # (node:22-alpine@sha256:c610fcdf, v22.23.2, npm 10.9.8): `npm ci`, | |
| # `npm run typecheck`, 179/179 vitest cases, `npm run build`, and the | |
| # `git diff --exit-code` route-tree check are all clean there. The #2034 | |
| # browser matrix is 58/58 on node:22 as well, so this moves coverage onto | |
| # the runtime rather than trading a gap for a break. | |
| # | |
| # Deriving instead of re-declaring is the part that stops it recurring. A | |
| # second copy of the version is a second thing to forget, and forgetting | |
| # it is precisely how the 24/22 split opened up; reading it from the | |
| # Dockerfile means an image bump moves CI with it in the same commit. | |
| - name: Node major from the image (#3331) | |
| id: node-runtime | |
| run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile | |
| - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 | |
| with: | |
| node-version: ${{ steps.node-runtime.outputs.node-version }} | |
| cache: npm | |
| cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json | |
| # #1816: the lockfile must install under the npm the *image* uses. | |
| # | |
| # Dependabot resolves with its own npm again, so a group bump can | |
| # produce a lockfile that satisfies one and not the other -- and the | |
| # failure then lands on whatever PR runs next, which is usually one | |
| # that touches no frontend file at all. Three times so far, each | |
| # costing a diagnosis before a one-line fix. | |
| # | |
| # Running the image's npm here makes it fail on the PR that caused | |
| # it. Regenerating automatically would also work; failing loudly in | |
| # the right place is better than a bot quietly rewriting a lockfile. | |
| # | |
| # #3331: the image is the derived `node-image` output, not a second | |
| # literal `node:22-alpine`. This step is the one place that already ran | |
| # the runtime's npm, so leaving the tag written out by hand would have | |
| # kept the drift alive on the exact axis this issue is about. The script | |
| # also refuses to resolve at all if the Dockerfile's stages ever split | |
| # across node majors, which would leave "the image's npm" ambiguous. | |
| # | |
| # --user matters: without it the container installs as root and | |
| # leaves a root-owned node_modules the runner's own `npm ci` below | |
| # cannot remove (EACCES on rmdir .bin), so the check would break the | |
| # very job it is meant to protect. HOME and the cache go somewhere | |
| # writable for that user for the same reason. | |
| - name: lockfile installs under the image's npm | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: | | |
| docker run --rm \ | |
| --user "$(id -u):$(id -g)" \ | |
| -e HOME=/tmp -e npm_config_cache=/tmp/npm-cache \ | |
| -v "$PWD:/app" -w /app \ | |
| ${{ steps.node-runtime.outputs.node-image }} npm ci --no-audit --no-fund | |
| rm -rf node_modules | |
| - run: npm ci | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # #2180 / #2804: there was a "TanStack minors diverge inside one | |
| # install" step here. It warned whenever the 1.x @tanstack/react-* | |
| # packages in package-lock.json spanned more than one minor. It has | |
| # been RETIRED, not narrowed, because the condition it asserts cannot | |
| # be satisfied by any installable version of this family and therefore | |
| # never carried information -- it fired on every run since #2180 added | |
| # it. Measured 2026-09-02, recorded here so the next reader does not | |
| # rediscover it a fourth time: | |
| # | |
| # 1. @tanstack/react-start pins its siblings as EXACT versions | |
| # spanning three minors. `npm view @tanstack/react-start@latest | |
| # dependencies` on 1.168.49 (which IS latest) gives | |
| # react-router 1.170.32, react-start-client 1.168.30, | |
| # react-start-server 1.167.37 -- 1.170 / 1.168 / 1.167. Installing | |
| # react-start fixes the rest outright, so package.json has no | |
| # lever over them and no bump converges the set. | |
| # 2. Narrowing the filter to the packages package.json declares | |
| # directly (react-router, react-start, router-cli) yields the | |
| # IDENTICAL minor set, 1.167/1.168/1.170 -- router-cli is | |
| # independently versioned and has never published past 1.167. | |
| # Verified against the real lockfile; it is a no-op, not a fix. | |
| # 3. Replacing the assertion with "the family moved together in one | |
| # PR" does not work either, in both directions. Over the | |
| # exact-pinned closure it can never fire: npm's own resolution | |
| # already makes react-start-client/-server move atomically with | |
| # react-start, so the check would assert something npm | |
| # guarantees. Widen it to include router-cli and it always fires, | |
| # since router-cli moves on its own cadence. Separately, | |
| # .github/dependabot.yml groups every frontend minor/patch update | |
| # into one `frontend-compatible` PR, so "moved together in one | |
| # PR" is already true by construction here. | |
| # | |
| # What #2180 actually wanted -- that upgrades be deliberate -- is | |
| # carried by the spec-exact pins in package.json (#2208), not by this | |
| # comparison. Asserting THOSE stay exact is a real, satisfiable check | |
| # and is the one piece of coverage this retirement drops; tracked in | |
| # #2867 rather than swapped in here unreviewed. The lockfile carries | |
| # no duplicate @tanstack copies today (every package appears exactly | |
| # once), so nothing is diverging inside one install in the literal | |
| # sense either. | |
| # | |
| # #2867: this is the check that retirement left uncovered. #2208 | |
| # pinned the three @tanstack/* deps in package.json spec-exact so | |
| # upgrades are deliberate (#2180's actual goal); nothing asserted that | |
| # stayed true. Needs no network and no lockfile resolution -- a jq | |
| # test over package.json alone -- so it runs before install-dependent | |
| # steps and fails (not warns) since a caret creeping back in is a real, | |
| # fixable regression. | |
| - name: TanStack deps stay spec-exact (#2180) | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: | | |
| loose="$(jq -r '((.dependencies // {}) + (.devDependencies // {})) | |
| | to_entries[] | |
| | select(.key | startswith("@tanstack/")) | |
| | select(.value | test("^[0-9]+\\.[0-9]+\\.[0-9]+$") | not) | |
| | "\(.key) \(.value)"' package.json)" | |
| if [ -n "$loose" ]; then | |
| echo "::error::@tanstack deps must be pinned spec-exact (#2180/#2208):" | |
| echo "$loose" | |
| exit 1 | |
| fi | |
| - run: npm run typecheck | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # #1831: the tier's first behavioural check. typecheck and build were | |
| # the only things running against it, and neither can see ordering, | |
| # a pre-hydration string literal, or anything else whose correctness | |
| # is a matter of what happens rather than of what shape it has. | |
| # #3319: CI_ARTIFACTS_DIR is what switches the vitest config's JUnit | |
| # reporter on (see its own comment). It is the ONLY thing that does, so | |
| # the same `npm test` a developer runs keeps vitest's console output and | |
| # writes no report file. | |
| - run: npm test | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| env: | |
| CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts | |
| # #3319: `always()`, not `failure()` -- the JUnit result is worth having | |
| # on a green run too, since it is the record of what actually ran rather | |
| # than of what the exit code happened to be. 7 days matches the issue's | |
| # retention ask: long enough to still be there when a red run is | |
| # investigated the following week, short enough that a busy main does | |
| # not accumulate them indefinitely. | |
| # | |
| # `overwrite: true` is gone from every upload in this file, and the | |
| # artifact name now ends in the run's own identity. It was here from | |
| # #3319, on the reading that a name which may already exist needs v4's | |
| # conflict policy. What `overwrite: true` actually does is DELETE an | |
| # existing artifact of that name before uploading, so on a name two | |
| # runs share it is not "the newest run's results win" -- it is "the | |
| # second uploader destroys the first's evidence", and the name then | |
| # resolves to whichever run got there last, whose contents the earlier | |
| # run never verified. A per-(run, attempt) name cannot collide, so there | |
| # is nothing left to overwrite and no name an earlier run on the same | |
| # ref could poison. | |
| # | |
| # run_attempt is in the name as well as run_id because a GitHub re-run | |
| # of a run keeps its run_id: run_id alone still collides on the second | |
| # attempt of the same run, which is the collision #3319 was really | |
| # hitting. With both, a name is unique per (run, attempt), and the step | |
| # that writes it runs once per (run, attempt) -- those two conditions | |
| # together are what make upload-artifact v4's immutability mean | |
| # something here. Every step below sits in a job that either has no | |
| # twin or has one gated on the opposite answer from ci-target, so no two | |
| # steps in a run can claim the same name. Nothing downstream reads these | |
| # by name (there is no actions/download-artifact in .github/), so the | |
| # per-run suffix costs no consumer. | |
| - name: Upload unit test results | |
| if: always() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-unit-junit-${{ github.run_id }}-${{ github.run_attempt }} | |
| # The file is named explicitly rather than uploading its directory: | |
| # upload-artifact v4.4+ skips hidden paths unless told otherwise, and | |
| # `.ci-artifacts/` is hidden, so `path: .ci-artifacts/` uploads | |
| # NOTHING (it reports "No files were found" and, with | |
| # if-no-files-found: ignore, says so quietly). Naming the file is | |
| # also what makes the contents a closed set: only the report the | |
| # lane's own reporter wrote can ever land here. | |
| path: .ci-artifacts/frontend-next-unit-junit.xml | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| # #3318: this tier collected no coverage at all, so there was no | |
| # baseline and nothing here could catch a change that deleted tested | |
| # behaviour. The tests that remained would keep passing over code that | |
| # had stopped being tested, and the job would stay green -- the tests | |
| # that assert the deleted behaviour would be the thing a reviewer has | |
| # to notice is missing. | |
| # | |
| # Two gates, both compared against the committed | |
| # arcane/home/honeypot-dashboard/frontend-next/coverage-baseline.json | |
| # and both computed from this run's own measurement. Neither is a | |
| # threshold somebody chose: | |
| # | |
| # coverage:ratchet the non-regression floor. The baseline is what the | |
| # suite measures today -- 806/7359 lines (10.95%) | |
| # and 346/7250 branches (4.77%) across 127 files of | |
| # src/ -- measured on this job's own #3331 node:22 | |
| # image. It is a floor, not a target, and the number | |
| # in the file is a measurement rather than an | |
| # aspiration on purpose: an importable 60% would | |
| # have made this gate red on the day it landed and | |
| # bought nothing. Raising it is a separate, | |
| # deliberate commit. | |
| # test:discovery a test-shaped file that no configured runner | |
| # collects. Its own step so a coverage failure | |
| # cannot hide a discovery failure. | |
| # | |
| # `npm test` above is untouched and stays uninstrumented: a second full | |
| # run of the suite is the honest price of keeping the command deploy.yml, | |
| # the README and a developer's own loop free of coverage, and it is | |
| # cheap here -- 6.9s uninstrumented against 6.9s with the v8 provider on | |
| # the image this job runs (the cost is transform/import, not | |
| # instrumentation), against a 45-minute budget. | |
| # | |
| # `set -euo pipefail` so a failing coverage run cannot be followed by a | |
| # ratchet that then reports on a missing report and takes the blame for | |
| # it. | |
| - name: "Coverage report and ratchet (#3318)" | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: | | |
| set -euo pipefail | |
| npm run test:coverage | |
| npm run coverage:ratchet | |
| - name: "Test discovery guard (#3318)" | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: npm run test:discovery | |
| # Same per-(run, attempt) name as the JUnit upload above, for the same | |
| # reason: run_id alone still collides on a re-run of the same run, and a | |
| # name two uploads share is not "newest wins" under v4 -- it is the | |
| # second uploader deleting the first one's evidence. No overwrite here | |
| # either, and the retention matches. | |
| # | |
| # `coverage/` is named as a directory rather than file-by-file because | |
| # it is not hidden -- the .ci-artifacts caveat that forces the JUnit | |
| # upload to name its file does not apply, and a directory keeps the | |
| # report extensible (adding the lcov html view later is a config change | |
| # here, not a workflow change). The contents stay a closed set: only the | |
| # files the coverage reporter wrote, from the run above, can land here. | |
| - name: Upload unit test coverage | |
| if: always() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-coverage-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: arcane/home/honeypot-dashboard/frontend-next/coverage/ | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| # No live-ES smoke suite here on purpose: port-tests/ needs the real | |
| # cluster over an SSH tunnel to the homeserver (see its README) -- | |
| # not reachable from a GitHub-hosted runner, and not appropriate to | |
| # point at production data from CI. That suite stays a manual/local | |
| # verification step; this job's job is catching build/type breakage. | |
| - run: npm run build | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # Checked post-build, not via the standalone `tsr generate` CLI: the | |
| # tanstackStart() vite plugin's own route-tree generation (used by | |
| # `build`/`dev`) and the router-cli's standalone `generate-routes` | |
| # disagree on one thing -- the plugin also emits a `Register` SSR | |
| # type augmentation the CLI doesn't -- so diffing right after | |
| # `generate-routes` flags the committed (plugin-shaped) file as | |
| # stale on every single run. The build's own output is what's | |
| # actually committed and actually ships; diff against that instead. | |
| - name: Generated route tree is current | |
| run: git diff --exit-code -- arcane/home/honeypot-dashboard/frontend-next/src/routeTree.gen.ts | |
| frontend-next-cloud: | |
| name: Dashboard frontend (next) (GitHub-hosted) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| # #3331: twin of the homeserver copy -- same derive, same reason. | |
| # (Full rationale, and the node:22 measurement behind it, lives there.) | |
| - name: Node major from the image (#3331) | |
| id: node-runtime | |
| run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile | |
| - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 | |
| with: | |
| node-version: ${{ steps.node-runtime.outputs.node-version }} | |
| cache: npm | |
| cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json | |
| # #1816: the lockfile must install under the npm the *image* uses. | |
| # (Full rationale lives on the homeserver twin above.) | |
| - name: lockfile installs under the image's npm | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: | | |
| docker run --rm \ | |
| --user "$(id -u):$(id -g)" \ | |
| -e HOME=/tmp -e npm_config_cache=/tmp/npm-cache \ | |
| -v "$PWD:/app" -w /app \ | |
| ${{ steps.node-runtime.outputs.node-image }} npm ci --no-audit --no-fund | |
| rm -rf node_modules | |
| - run: npm ci | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # #2180 / #2804: the twin of the retired "TanStack minors diverge | |
| # inside one install" step lived here. Retired for the same reason -- | |
| # see the full rationale on the homeserver copy above. | |
| # | |
| # #2867: twin of the homeserver copy's spec-exact check above. | |
| - name: TanStack deps stay spec-exact (#2180) | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: | | |
| loose="$(jq -r '((.dependencies // {}) + (.devDependencies // {})) | |
| | to_entries[] | |
| | select(.key | startswith("@tanstack/")) | |
| | select(.value | test("^[0-9]+\\.[0-9]+\\.[0-9]+$") | not) | |
| | "\(.key) \(.value)"' package.json)" | |
| if [ -n "$loose" ]; then | |
| echo "::error::@tanstack deps must be pinned spec-exact (#2180/#2208):" | |
| echo "$loose" | |
| exit 1 | |
| fi | |
| - run: npm run typecheck | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # #3319: twin of the homeserver copy's JUnit wiring above. | |
| - run: npm test | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| env: | |
| CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts | |
| # Same per-(run, attempt) artifact name and the same explicit-file path | |
| # as the homeserver twin -- see that step for why `overwrite: true` is | |
| # gone and why the path is not the `.ci-artifacts/` directory. | |
| - name: Upload unit test results | |
| if: always() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-unit-junit-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: .ci-artifacts/frontend-next-unit-junit.xml | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| # #3318: twin of the homeserver copy's coverage + ratchet + discovery | |
| # steps. Present here for the reason the header comment gives: steps stay | |
| # byte-identical across the twins so a red result means the same thing | |
| # wherever it ran. A ratchet that only ran on the self-hosted executor | |
| # would let a regression through on every degraded day, which is exactly | |
| # the day the fallback twin exists for. (Full rationale, and the measured | |
| # baseline, live on the homeserver copy above.) | |
| - name: "Coverage report and ratchet (#3318)" | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: | | |
| set -euo pipefail | |
| npm run test:coverage | |
| npm run coverage:ratchet | |
| - name: "Test discovery guard (#3318)" | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| run: npm run test:discovery | |
| - name: Upload unit test coverage | |
| if: always() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-coverage-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: arcane/home/honeypot-dashboard/frontend-next/coverage/ | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| # No live-ES smoke suite here on purpose -- see the homeserver twin. | |
| - run: npm run build | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| - name: Generated route tree is current | |
| run: git diff --exit-code -- arcane/home/honeypot-dashboard/frontend-next/src/routeTree.gen.ts | |
| # #2034: the browser-level acceptance net returned. The Go tier ran a | |
| # 90-case Playwright matrix (#60, PR #146) until the cutover deleted it; | |
| # nothing replaced it and visual/behavioural regressions have had no net | |
| # since. This is its deliberately slimmed port -- theme x viewport shell | |
| # smoke over every sidebar route (from lib/nav.ts), the modal core, and | |
| # role-aware action visibility -- running against the BUILT production | |
| # server output with hermetic fixtures (e2e/start-dashboard.mjs), not the | |
| # dev server. | |
| # | |
| # Homeserver-first as a pair (#2565): the box preinstalls the playwright | |
| # chromium library set (extracted for the pinned playwright-core's | |
| # ubuntu26.04-x64 key) and the redis-server binary for the runner user, | |
| # who has no sudo by design -- so the homeserver twin drops --with-deps | |
| # and the apt fallback and, instead of guessing, FAILS LOUDLY when the | |
| # host provision is missing. Consequence to keep in mind on a playwright | |
| # bump: if the new build needs a library the host list doesn't cover, | |
| # this twin fails at chromium launch with the missing-lib message and | |
| # the host list (not this file) is what needs the update. | |
| frontend-next-browser: | |
| name: Dashboard-next browser matrix | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 45 | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| # #3331: same derive as the two jobs above, and it matters most here -- | |
| # this twin runs the BUILT output, so it is the closest thing CI has to | |
| # executing what the container serves, and it was doing that on 24 while | |
| # the container serves 22. 58/58 of this matrix is green on node:22 | |
| # (node:22-bookworm, v22.23.2), measured before the pin moved. | |
| - name: Node major from the image (#3331) | |
| id: node-runtime | |
| run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile | |
| - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 | |
| with: | |
| node-version: ${{ steps.node-runtime.outputs.node-version }} | |
| cache: npm | |
| cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json | |
| - run: npm ci | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| - run: npm run build | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # No --with-deps: that path apt-installs, and the runner user has no | |
| # sudo by design. The chromium library set lives on the host (#2565) | |
| # and only the browser binary itself downloads here (cached in the | |
| # runner's persistent ~/.cache/ms-playwright). | |
| - run: npx playwright install chromium | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # The session fixture spawns a real redis-server rather than | |
| # reimplementing RESP (#1034's tradeoff). The runner user cannot | |
| # apt-install it, so a missing binary fails here, loudly, instead of | |
| # as an obscure fixture error three steps later. | |
| - run: | | |
| command -v redis-server >/dev/null 2>&1 || { | |
| echo "::error::redis-server missing on the runner host -- see #2565's homeserver provision list" | |
| exit 1 | |
| } | |
| # #3319: CI_ARTIFACTS_DIR turns on the config's JUnit reporter. The | |
| # trace/screenshots half of the issue needs no flag -- playwright.config | |
| # already carries trace: "retain-on-failure" and screenshot: | |
| # "only-on-failure" (#2034); what was missing was uploading them, so a | |
| # failing browser case was readable only as a log line. | |
| - run: npm run test:browser | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| env: | |
| CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts | |
| # `always()`: the JUnit XML is the machine-readable answer to "which | |
| # cases ran, which were skipped, which failed" and is worth keeping on a | |
| # green run too. Both paths are `if-no-files-found: warn` rather than | |
| # error so a lane that never produced a report -- because it failed at | |
| # `npm ci`, say -- reports the absence instead of masking the real | |
| # failure behind an upload error. | |
| # | |
| # Per-(run, attempt) name, explicit file path, no `overwrite: true` -- | |
| # all three for the reasons spelled out on the frontend-next twin's | |
| # upload above. | |
| - name: Upload browser test results | |
| if: always() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-browser-junit-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: .ci-artifacts/playwright-junit.xml | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| # `failure()`, not `always()`: the HTML report bundles a trace per | |
| # failing test and is the one artifact big enough to matter. On a green | |
| # run it holds nothing worth storing. | |
| - name: Upload Playwright report and traces | |
| if: failure() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-browser-report-${{ github.run_id }}-${{ github.run_attempt }} | |
| # Both directories are named, not the frontend-next tree around | |
| # them. playwright.config.ts's own outputDir is ./test-results and | |
| # the html reporter's is ./playwright-report, so this lane creates | |
| # and owns both in a fresh checkout: the contents are what the | |
| # browser matrix wrote, not whatever happens to be under | |
| # frontend-next/ right now -- which is the property that keeps a | |
| # stray .env or key out of a 7-day artifact anyone with repo read | |
| # access can pull. | |
| path: | | |
| arcane/home/honeypot-dashboard/frontend-next/playwright-report/ | |
| arcane/home/honeypot-dashboard/frontend-next/test-results/ | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| frontend-next-browser-cloud: | |
| name: Dashboard-next browser matrix (GitHub-hosted) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| # #3331: twin -- see the homeserver copy above for the rationale. | |
| - name: Node major from the image (#3331) | |
| id: node-runtime | |
| run: ./scripts/node-runtime-major.sh arcane/home/honeypot-dashboard/frontend-next/Dockerfile | |
| - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 | |
| with: | |
| node-version: ${{ steps.node-runtime.outputs.node-version }} | |
| cache: npm | |
| cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json | |
| - run: npm ci | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| - run: npm run build | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| - run: npx playwright install --with-deps chromium | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| # The session fixture spawns a real redis-server rather than | |
| # reimplementing RESP (#1034's tradeoff); ubuntu-latest usually ships | |
| # it, but fail loudly into install rather than obscurely later. | |
| # Grouped braces: ungrouped `A || B && C` parses as `(A || B) && C` | |
| # (#2224), which apt-installs on every run. Same idiom as the | |
| # shellcheck guard below. | |
| - run: command -v redis-server >/dev/null 2>&1 || { sudo apt-get update && sudo apt-get install -y redis-server; } | |
| # #3319: twin of the homeserver copy's artifact wiring above. | |
| - run: npm run test:browser | |
| working-directory: arcane/home/honeypot-dashboard/frontend-next | |
| env: | |
| CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts | |
| # Twin of the homeserver copy's two uploads: same per-(run, attempt) | |
| # names, same explicit paths, no `overwrite: true`. | |
| - name: Upload browser test results | |
| if: always() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-browser-junit-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: .ci-artifacts/playwright-junit.xml | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| - name: Upload Playwright report and traces | |
| if: failure() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: frontend-next-browser-report-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: | | |
| arcane/home/honeypot-dashboard/frontend-next/playwright-report/ | |
| arcane/home/honeypot-dashboard/frontend-next/test-results/ | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| # First CI coverage for this crate (#1608's Rust service tier had none | |
| # until now -- a broken build or a real regression wouldn't have | |
| # surfaced until Arcane tried to build it live on the homeserver). | |
| # | |
| # Homeserver-first, unlike when #1608 landed: back then the self-hosted | |
| # runner had no Rust toolchain to fall back on. Instead of relying on | |
| # host state (the old assumption that rustup happened to be preinstalled | |
| # on GH images), the homeserver twin bootstraps rustup itself | |
| # into the runner user's persistent HOME -- rustup then honours the | |
| # crate's own rust-toolchain.toml (#1720 pin) exactly like every other | |
| # execution path does. | |
| # | |
| # This comment used to also claim `~/.cargo` + `target/` "simply stay warm | |
| # on disk between runs, which is why no actions/cache block exists here". | |
| # Only the first half was ever true, and the second half was the reason the | |
| # 159s job rebuilt its whole crate graph on every single run: | |
| # | |
| # * `~/.cargo` (registry/, git/) really does persist -- it is in $HOME, | |
| # outside the checkout. | |
| # * `target/` does NOT persist. It lives inside the workspace, and | |
| # actions/checkout defaults to `clean: true`, i.e. `git clean -ffdx`. | |
| # The `-x` is what matters: it deletes *ignored* files, and `target/` is | |
| # ignored (.gitignore line 7). Verified locally, not inferred: a | |
| # `sub/target/debug/libfoo.rlib` under a `target/` ignore rule does not | |
| # survive `git clean -ffdx`. So every run recompiled all 287 packages in | |
| # Cargo.lock from an empty target/. | |
| # | |
| # The cloud twin already carries an actions/cache block over target/ and | |
| # pays this correctly (each ubuntu-latest runner is a fresh VM, so the | |
| # cache service is the only disk that survives). This twin has the opposite | |
| # situation -- a persistent disk that the checkout was destroying -- so the | |
| # fix is to move the one directory that is being wiped to a path that is | |
| # not. Deliberately NOT an actions/cache block: see the "Reuse a persistent, | |
| # ref-scoped target/" step below and the go-fmt job's comment for what | |
| # actions/cache did on this runner (six same-key 465 MB copies put the repo | |
| # over its 10 GB quota). | |
| # | |
| # `cargo fmt --check` deliberately isn't part of this gate: the crate | |
| # predates any formatting pass and has never been run through rustfmt | |
| # (532 diff hunks against its default profile) -- landing that gate | |
| # now would force an unrelated repo-wide reformat into this PR instead | |
| # of a follow-up that can be reviewed on its own. | |
| backend-service: | |
| name: Dashboard backend-service (Rust) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 90 | |
| defaults: | |
| run: | |
| working-directory: arcane/home/honeypot-dashboard/backend-service | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - name: Bootstrap rustup into the persistent HOME (no-op after first) | |
| shell: bash | |
| run: | | |
| set -euo pipefail | |
| if ! command -v rustup >/dev/null 2>&1; then | |
| # Pinned version + hardcoded checksum instead of piping | |
| # sh.rustup.rs straight into sh (SAST-flagged, #3115). The | |
| # unversioned dist/ path is rolling-latest, and its .sha256 | |
| # sidecar is fetched from the same host over the same TLS | |
| # session moments later -- that detects transport corruption, | |
| # which TLS already covers, not a compromised upstream. Same | |
| # discipline as the trivy block below: fetch the versioned | |
| # archive and verify against a digest hardcoded here, taken | |
| # from that archive's own .sha256 once at authoring time. | |
| rustup_version=1.28.2 | |
| rustup_sha256=20a06e644b0d9bd2fbdbfd52d42540bdde820ea7df86e92e533c073da0cdd43c | |
| target=x86_64-unknown-linux-gnu | |
| tmp="$(mktemp -d)" | |
| curl --proto '=https' --tlsv1.2 -sSf -o "$tmp/rustup-init" \ | |
| "https://static.rust-lang.org/rustup/archive/${rustup_version}/${target}/rustup-init" | |
| echo "${rustup_sha256} $tmp/rustup-init" | sha256sum -c - | |
| chmod +x "$tmp/rustup-init" | |
| "$tmp/rustup-init" -y --profile minimal --no-modify-path | |
| rm -rf "$tmp" | |
| fi | |
| echo "$HOME/.cargo/bin" >>"$GITHUB_PATH" | |
| # `rustup show` installs whatever rust-toolchain.toml pins (including | |
| # its declared clippy component); explicit component add stays as a | |
| # cheap idempotent safety net for toolchains whose config predates | |
| # components support. | |
| - name: Install pinned toolchain (rust-toolchain.toml) | |
| shell: bash | |
| run: | | |
| export PATH="$HOME/.cargo/bin:$PATH" | |
| rustup show active-toolchain | |
| rustup component add clippy | |
| cargo --version && rustc --version | |
| # #3405: put target/ somewhere `git clean -ffdx` cannot reach, so the | |
| # 159s of "recompile all 287 packages" stops happening every run. | |
| # | |
| # This is the self-hosted counterpart of the cloud twin's actions/cache | |
| # block, and it is NOT an actions/cache block. On this runner the | |
| # persistent disk already *is* the cache: the go-fmt job's comment | |
| # records what happens when actions/cache is used anyway -- Go's 0555 | |
| # module dirs make the restore's `tar -x` fail, the job then re-uploads, | |
| # and six same-key 465 MB copies push the repo past its 10 GB Actions | |
| # quota and evict unrelated caches. ~/.cargo/registry and ~/.cargo/git | |
| # need nothing from us for the same reason: they are in $HOME and have | |
| # been warm all along. target/ is the only one checkout destroys. | |
| # | |
| # Keyed per ref, never shared across refs. The honeypot-ci label is | |
| # served by several runner instances that between them run every open | |
| # branch, so a single shared target/ would be written concurrently by | |
| # unrelated checkouts. One directory per ref makes that structurally | |
| # impossible: two refs never touch the same path. | |
| # | |
| # NOT keyed per commit, and that is the whole trick. Cargo fingerprints | |
| # each unit against its own inputs, so reusing a branch's directory | |
| # across that branch's commits is exactly what it is designed for -- and | |
| # it is the case that pays. A commit-keyed directory would be cold on | |
| # every push (a push *is* a new commit), which is a slower spelling of | |
| # the 159s job we are trying to remove. | |
| # | |
| # The root is resolved best-effort and this step can never fail the job: | |
| # if no persistent root is writable, CARGO_TARGET_DIR is left unset, | |
| # cargo builds into the in-tree target/ it uses today, and the run costs | |
| # what it always cost. A cold cache is therefore a normal outcome here, | |
| # not an error -- and it cannot turn a red crate green, because cargo | |
| # rebuilds anything whose source, toolchain or deps changed. | |
| - name: Reuse a persistent, ref-scoped target/ (best-effort) | |
| shell: bash | |
| run: | | |
| set -uo pipefail | |
| warn() { echo "::warning::cargo-target: $*"; } | |
| # Precedence: ops override, then a cross-instance shared root, then | |
| # this runner instance's own HOME. The middle option needs the box | |
| # provisioned (install-ci-runner.sh already sets UMask=0002 and a | |
| # setgid shared group, so a group-writable root is all it takes) and | |
| # raises the hit rate from "the same instance happened to take this | |
| # job again" to "any instance". The HOME option needs no | |
| # provisioning at all, so the cache is live on an un-provisioned box | |
| # rather than being dead code waiting for an ops ticket. | |
| root="" | |
| for candidate in \ | |
| "${CI_CARGO_TARGET_ROOT:-}" \ | |
| /var/cargo-target-cache \ | |
| "$HOME/.cache/cargo-target"; do | |
| [ -n "$candidate" ] || continue | |
| if mkdir -p "$candidate" 2>/dev/null && [ -w "$candidate" ]; then | |
| root="$candidate" | |
| break | |
| fi | |
| done | |
| if [ -z "$root" ]; then | |
| warn "no writable persistent root; building into the in-tree target/ as before" | |
| exit 0 | |
| fi | |
| # GITHUB_REF is fully qualified -- refs/heads/main on push, | |
| # refs/pull/<n>/merge on pull_request -- and for a given PR that | |
| # merge ref is stable across every run of that PR, so one directory | |
| # is reused for a PR's whole review loop and then abandoned when it | |
| # closes. The slug is for humans; the digest is what actually | |
| # guarantees uniqueness, since GITHUB_REF_NAME mangles ("a/b" and | |
| # "a-b" both slug to "a-b"). | |
| slug="$(printf '%s' "${GITHUB_REF_NAME:-noref}" \ | |
| | tr -c 'A-Za-z0-9._-' '-' | cut -c1-40)" | |
| ref_digest="$(printf '%s' "${GITHUB_REF:-noref}" | sha256sum | cut -c1-8)" | |
| channel="$(sed -n 's/^channel[[:space:]]*=[[:space:]]*"\(.*\)"/\1/p' \ | |
| rust-toolchain.toml 2>/dev/null | head -1)" | |
| [ -n "$channel" ] || channel="unpinned" | |
| # Toolchain in the key so a pin bump starts clean rather than | |
| # leaning on cargo noticing the compiler changed underneath a | |
| # long-lived shared directory. | |
| toolchain_digest="$(printf '%s' "$channel" | sha256sum | cut -c1-8)" | |
| dir="$root/${slug}-${ref_digest}-${toolchain_digest}" | |
| if ! mkdir -p "$dir" 2>/dev/null; then | |
| warn "could not create $dir; building into the in-tree target/ as before" | |
| exit 0 | |
| fi | |
| # Heartbeat, not the directory's own mtime, is what eviction reads. | |
| # A warm rebuild writes into target/debug/ and does not necessarily | |
| # touch the target/ directory entry itself, so an mtime rule could | |
| # delete a directory a running job is actively using. The job below | |
| # has a 90-minute timeout against a 7-day floor, so a heartbeat | |
| # touched now cannot be mistaken for an abandoned one. | |
| touch "$dir/.ci-target-heartbeat" 2>/dev/null || true | |
| if [ -d "$dir/debug" ] || [ -d "$dir/.fingerprint" ]; then | |
| warm=yes | |
| else | |
| warm=no | |
| fi | |
| if ! echo "CARGO_TARGET_DIR=$dir" >>"$GITHUB_ENV"; then | |
| warn "could not export CARGO_TARGET_DIR; building into the in-tree target/ as before" | |
| exit 0 | |
| fi | |
| { | |
| echo "### Rust target/ reuse" | |
| echo "" | |
| echo "- root: \`$root\`" | |
| echo "- directory: \`$dir\`" | |
| echo "- ref: \`${GITHUB_REF:-noref}\` (slug \`$slug\`, digest \`$ref_digest\`)" | |
| echo "- toolchain: \`$channel\` (digest \`$toolchain_digest\`)" | |
| echo "- state on arrival: **$warm** (\"warm\" = a previous run left a target tree here)" | |
| } >>"$GITHUB_STEP_SUMMARY" | |
| echo "cargo-target: using $dir (warm=$warm)" | |
| # Eviction: drop directories whose job has not run in | |
| # CI_CARGO_TARGET_PRUNE_DAYS days, keyed on the heartbeat. Without | |
| # this the shared box accumulates one full target/ per ref ever | |
| # opened. Never fatal, and only ever removes a directory that has | |
| # provably been idle for a week. | |
| prune_days="${CI_CARGO_TARGET_PRUNE_DAYS:-7}" | |
| find "$root" -maxdepth 2 -name '.ci-target-heartbeat' -type f \ | |
| -mtime "+$prune_days" -print0 2>/dev/null \ | |
| | while IFS= read -r -d '' heartbeat; do | |
| victim="$(dirname "$heartbeat")" | |
| # Defence in depth: never hand rm -rf anything but a direct | |
| # child of the root we resolved above. | |
| case "$victim" in | |
| "$root"/*) rm -rf "$victim" && echo "cargo-target: pruned $victim" ;; | |
| *) warn "refusing to prune $victim (not under $root)" ;; | |
| esac | |
| done || true | |
| # Teed so the "did the cache actually hit" question is answered by cargo's | |
| # own output rather than by inference. pipefail keeps a build failure a | |
| # build failure; nothing here asserts anything new. | |
| - name: Build | |
| run: | | |
| set -o pipefail | |
| PATH="$HOME/.cargo/bin:$PATH" cargo build --all-targets 2>&1 \ | |
| | tee "${RUNNER_TEMP}/cargo-build.log" | |
| # #3319: tees the test output into .ci-artifacts/ and uploads it on | |
| # failure. Not a JUnit file: cargo has no built-in JUnit reporter, and | |
| # its per-test result stream is only available on the unstable | |
| # `--format json`, so the honest artifact here is the full log rather | |
| # than a hand-rolled conversion that could misrepresent a result. The | |
| # log is what the issue said was the only evidence a red run had; it is | |
| # now retained rather than scrolled away. | |
| # | |
| # `mkdir -p` plus a named file, never a directory upload: the log is | |
| # the only thing this job writes under .ci-artifacts/, and naming it | |
| # keeps the artifact's contents a closed set (see the frontend-next | |
| # twin's upload for why the directory form uploads nothing at all). | |
| - name: Test (log retained on failure) | |
| shell: bash | |
| run: | | |
| set -o pipefail | |
| mkdir -p "$CI_ARTIFACTS_DIR" | |
| PATH="$HOME/.cargo/bin:$PATH" cargo test 2>&1 | tee "$CI_ARTIFACTS_DIR/cargo-test.log" | |
| env: | |
| CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts | |
| # Per-(run, attempt) name, no `overwrite: true` -- see the frontend-next | |
| # twin's upload for the full argument. | |
| - name: Upload test log | |
| if: failure() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: backend-service-cargo-test-log-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: .ci-artifacts/cargo-test.log | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| - run: PATH="$HOME/.cargo/bin:$PATH" cargo clippy --all-targets -- -D warnings | |
| # #3325: the committed contract has to be what the generator | |
| # renders. `cargo test` already fails on a stale openapi.json, but | |
| # only the *tests* fail, and only once someone reads which of the | |
| # four assertions went red -- `diff` states the actual change and | |
| # is the step a reviewer reads when they wonder why the file moved. | |
| # Regenerate locally with `cargo run --bin openapi > openapi.json`. | |
| - name: OpenAPI contract is not stale (#3325) | |
| run: PATH="$HOME/.cargo/bin:$PATH" cargo run --quiet --bin openapi | diff -u openapi.json - | |
| # The measurement, not an assertion. Cargo prints one "Compiling <pkg>" | |
| # line per unit it actually rebuilt, so this counts the work the cache | |
| # did NOT save: ~287 on a cold run, 0 when every unit was still fresh. | |
| # That number is the proof the cache hits, and it is read off a real | |
| # build rather than inferred from a cache-hit log line. | |
| - name: Report target/ reuse outcome | |
| if: always() | |
| shell: bash | |
| run: | | |
| set -uo pipefail | |
| log="${RUNNER_TEMP}/cargo-build.log" | |
| # `grep -c` prints the count AND exits 1 when the count is 0, so a | |
| # bare `|| echo 0` would yield "0\n0" and break the arithmetic | |
| # below. Swallow the status instead and keep only the count. | |
| count() { grep -c "$1" "$2" 2>/dev/null || true; } | |
| total="$(count '^[[:space:]]*\[\[package\]\]' Cargo.lock)" | |
| [ -n "$total" ] || total=0 | |
| if [ -r "$log" ]; then | |
| compiled="$(count '^[[:space:]]*Compiling ' "$log")" | |
| [ -n "$compiled" ] || compiled=0 | |
| reused="$(( total - compiled ))" | |
| else | |
| compiled="n/a" | |
| reused="n/a" | |
| fi | |
| size="n/a" | |
| if [ -n "${CARGO_TARGET_DIR:-}" ] && [ -d "$CARGO_TARGET_DIR" ]; then | |
| size="$(du -sh "$CARGO_TARGET_DIR" 2>/dev/null | cut -f1)" | |
| [ -n "$size" ] || size="n/a" | |
| fi | |
| { | |
| echo "### Rust build work actually done" | |
| echo "" | |
| echo "- packages compiled this run: **$compiled** of $total in Cargo.lock" | |
| echo "- units served from the reused target/: **$reused**" | |
| echo "- target/ size after the build: \`$size\`" | |
| echo "" | |
| echo "A cold cache is a valid outcome: the step above falls back to the" | |
| echo "in-tree target/ and this run costs what it always cost." | |
| } >>"$GITHUB_STEP_SUMMARY" | |
| echo "cargo-target: compiled=$compiled reused=$reused size=$size" | |
| backend-service-cloud: | |
| name: Dashboard backend-service (Rust) (GitHub-hosted) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| timeout-minutes: 60 | |
| defaults: | |
| run: | |
| working-directory: arcane/home/honeypot-dashboard/backend-service | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - run: rustup component add clippy | |
| - uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0 | |
| with: | |
| path: | | |
| ~/.cargo/registry | |
| ~/.cargo/git | |
| arcane/home/honeypot-dashboard/backend-service/target | |
| key: cargo-${{ runner.os }}-${{ hashFiles('arcane/home/honeypot-dashboard/backend-service/Cargo.lock') }} | |
| restore-keys: cargo-${{ runner.os }}- | |
| - run: cargo build --all-targets | |
| # #3319: twin of the homeserver copy's log retention above. | |
| - name: Test (log retained on failure) | |
| shell: bash | |
| run: | | |
| set -o pipefail | |
| mkdir -p "$CI_ARTIFACTS_DIR" | |
| cargo test 2>&1 | tee "$CI_ARTIFACTS_DIR/cargo-test.log" | |
| env: | |
| CI_ARTIFACTS_DIR: ${{ github.workspace }}/.ci-artifacts | |
| # Twin of the homeserver copy's upload: same per-(run, attempt) name, | |
| # same named-file path, no `overwrite: true`. | |
| - name: Upload test log | |
| if: failure() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: backend-service-cargo-test-log-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: .ci-artifacts/cargo-test.log | |
| if-no-files-found: warn | |
| retention-days: 7 | |
| - run: cargo clippy --all-targets -- -D warnings | |
| # #3325: same drift gate as the homeserver lane -- see the note | |
| # there. The two lanes both build this contract, so the contract | |
| # cannot pass on one toolchain and fail on the other. | |
| - name: OpenAPI contract is not stale (#3325) | |
| run: cargo run --quiet --bin openapi | diff -u openapi.json - | |
| vendored-theme: | |
| name: Vendored Xore/theme is in sync | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver == 'true' | |
| runs-on: [self-hosted, linux, x64, honeypot-ci] | |
| timeout-minutes: 10 | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - name: theme.css matches the pinned commit | |
| run: scripts/check-vendored-theme.sh | |
| # python3, not python: this job has no setup-python step, so on the | |
| # self-hosted lane it runs on the runner's own PATH and the homeserver | |
| # runner ships no bare `python` -- `python branding/...` died with | |
| # exit 127 there (#2461) while the vendored sheet itself was fine. | |
| - name: Portable APIARY tokens match Xore/theme | |
| run: python3 branding/scripts/check_theme_sync.py | |
| # Canvas surfaces resolve tokens to pixels and carry dark-theme | |
| # fallbacks, so a rename upstream degrades silently instead of failing. | |
| # #1825 renamed eight tokens and 31 call sites kept rendering, wrong. | |
| - name: Frontend theme tokens all exist | |
| run: scripts/check-theme-tokens.sh | |
| # The catalogue lives in the frontend, the tokens live in the vendored | |
| # stylesheet. A theme in one and not the other fails silently either | |
| # way -- default styling, or unpickable. #1758. | |
| - name: Theme catalogue matches theme.css | |
| run: scripts/check-theme-catalogue.sh | |
| vendored-theme-cloud: | |
| name: Vendored Xore/theme is in sync (GitHub-hosted) | |
| needs: [ci-target] | |
| if: needs.ci-target.outputs.homeserver != 'true' | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - name: theme.css matches the pinned commit | |
| run: scripts/check-vendored-theme.sh | |
| # Same python3 choice as the self-hosted twin above: one spelling | |
| # across both lanes, valid on the runner's own PATH either way. | |
| - name: Portable APIARY tokens match Xore/theme | |
| run: python3 branding/scripts/check_theme_sync.py | |
| - name: Frontend theme tokens all exist | |
| run: scripts/check-theme-tokens.sh | |
| - name: Theme catalogue matches theme.css | |
| run: scripts/check-theme-catalogue.sh | |
| scripts-and-compose: | |
| # Per-row executor routing (#2389, widened by #2565): `home: true` | |
| # marks matrix rows whose whole run needs nothing beyond checkout | |
| # files, this workflow's own setup interpreters, pip installs, and -- | |
| # since #2565 retired the runner's no-docker design -- the docker | |
| # daemon the runner user now drives via its docker-group membership. | |
| # Eligibility under the ORIGINAL (docker-less) criterion was audited | |
| # by executing every candidate in a docker-less, sudo-less | |
| # environment with the pip-flavored rows re-run in clean venvs -- the | |
| # audit ledger lives on the #2389 PR. The four | |
| # analysis/tests/test_{honeypot_ilm_rollover,geoip_pipeline, | |
| # dionaea_incidents_index,conpot_persona_pipeline}.sh rows were left | |
| # unflagged then because they print "SKIP: docker daemon is not | |
| # reachable" without a container engine -- on the box they now run | |
| # their real Elasticsearch containers, so they are flagged. | |
| # | |
| # A flagged row lands on [self-hosted, linux, x64, honeypot-ci] whenever | |
| # ci-target proves the box live -- the same heartbeat gate every paired | |
| # family above uses -- else falls back to ubuntu-latest, gaining the | |
| # "(GitHub-hosted)" name suffix that day per the pair-naming rule above. | |
| # timeout-minutes rides along only while the leg actually sits on the | |
| # metal (a wedged pickup blocks a real box nobody reboots promptly); | |
| # fallback legs keep the platform ceiling (360 = effectively unset). | |
| # | |
| # Unflagged rows stay on ubuntu-latest: the inert "Set up Node for the | |
| # dashboard OIDC suites" row (a matrix can't inject a use-step -- its | |
| # node arrives from the runner image either way), and the | |
| # "Sandbox shell tests (#2268, #2253)" row -- that one is eligible | |
| # under the criterion above (toolchain-only: bash, coreutils, a | |
| # mktemp workdir) but is deliberately kept off the metal anyway, | |
| # because it drives guest-runner.sh's tail through a PATH-stubbed | |
| # `systemctl poweroff` and run_pending.sh's flock path. Both are | |
| # neutralised by the stubs and by WINDOWS_SANDBOX_* redirection into | |
| # the temp dir, but a disposable GitHub-hosted VM is the right blast | |
| # radius for a suite whose production form powers a machine off and | |
| # takes the shared KVM detonation lock. Nothing else -- every | |
| # remaining row is flagged. | |
| name: ${{ matrix.name }}${{ matrix.home == true && needs.ci-target.outputs.homeserver != 'true' && ' (GitHub-hosted)' || '' }} | |
| needs: [ci-target] | |
| runs-on: ${{ matrix.home == true && needs.ci-target.outputs.homeserver == 'true' && fromJSON('["self-hosted", "linux", "x64", "honeypot-ci"]') || fromJSON('["ubuntu-latest"]') }} | |
| timeout-minutes: ${{ (matrix.home == true && needs.ci-target.outputs.homeserver == 'true') && 45 || 360 }} | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| include: | |
| - name: Python syntax | |
| home: true | |
| run: find . -path '*/vendor' -prune -o -name '*.py' -type f -print0 | xargs -0 -r python -m py_compile | |
| - name: Dionaea log rotation patch (#1389) | |
| home: true | |
| run: | | |
| # Build-time patch applied to dionaea's vendored log_json.py/ | |
| # log_incident.py -- exercises apply_patch() against fixtures built | |
| # from the patch's own OLD_CLASS/OLD_IMPORT text, and the actual | |
| # patched FileHandler class body (exec'd, not reimplemented) for | |
| # rotation/data-loss behavior. | |
| python arcane/home/honeypot-dionaea/dionaea/tests/test_log_rotation_patch.py | |
| - name: Conpot JSON log rotation patch (#2892) | |
| home: true | |
| run: | | |
| # Same shape as the dionaea row above, plus a behavioural | |
| # check the dionaea suite doesn't have: the patched JsonLogger | |
| # is imported (not reimplemented) against a fixture package | |
| # and run past CONPOT_JSON_LOG_MAX_BYTES to confirm the | |
| # close/rename/reopen rotation actually fires. | |
| python arcane/home/honeypot-conpot/conpot/tests/test_json_log_rotation_patch.py | |
| - name: Galah JSON log rotation patch (#2892) | |
| home: true | |
| run: | | |
| # Same shape as the conpot row above, but the target is Go | |
| # rather than Python: the suite extracts the patch's added | |
| # rotatingWriter type into a standalone Go program and | |
| # `go run`s it, confirming the close/rename/reopen rotation | |
| # actually fires -- skips (not fails) if `go` isn't on PATH. | |
| python arcane/home/honeypot-galah/galah/tests/test_json_log_rotation_patch.py | |
| - name: Beelzebub JSON log rotation patch (#2892) | |
| home: true | |
| run: | | |
| # Same shape as the Galah row above (Go target, `go run`s the | |
| # extracted rotatingWriter, skips without `go`) -- the other | |
| # vendored Go sensor, patched in internal/builder/builder.go. | |
| python arcane/home/honeypot-beelzebub/beelzebub/tests/test_json_log_rotation_patch.py | |
| - name: HellPot X-Forwarded-For trust patch (#1876) | |
| home: true | |
| run: | | |
| # Two build-time patches run in sequence against the same file | |
| # and the second matches on the first's output. That coupling | |
| # is invisible at review time and breaks the image build with a | |
| # string mismatch several layers down, so it is asserted here -- | |
| # along with the trust rule that makes reintroducing header | |
| # trust safe after #1419 removed it as spoofable. | |
| python arcane/home/honeypot-hellpot/hellpot/tests/test_xff_trust_patch.py | |
| - name: "Mailoney JSON-line sink patches (#1422, #2197)" | |
| home: true | |
| run: | | |
| # Exact-text build-time patches over vendored core.py. The suite | |
| # execs the injected helper block under the container's own | |
| # TZ=Europe/Berlin pin to prove emitted Z-suffixed stamps are | |
| # true UTC (#2197: Berlin wall clock used to wear a Z, shifting | |
| # ip-enrichment-worker's portbridge join by the DST offset), | |
| # that every match target still fits upstream verbatim, and | |
| # that AUTH PLAIN capture survives unpadded or hostile base64. | |
| python arcane/home/honeypot-mailoney/mailoney/tests/test_json_log_patch.py | |
| - name: Sensor timestamps carry explicit UTC sources (#2197) | |
| home: true | |
| run: | | |
| # Repo-wide grep-level guard for the class behind #2197: any | |
| # emitter of a %SZ stamp must pass an explicit UTC source | |
| # (gmtime/timezone.utc/date -u) on the same line, or carry a | |
| # `utc-verified:` waiver with its reason inline. Parse-side | |
| # formats are exempt by construction; Rust .format() calls are | |
| # out of scope by design -- the zone rides on the value's type, | |
| # not the template. | |
| python scripts/check-timestamp-utc.py | |
| - name: "JSON sink retention parity (#120, #2196, #2216)" | |
| home: true | |
| run: | | |
| # Every self-rotating JSON directory under /logs must have BOTH | |
| # halves of the #120 contract present -- a writer that actually | |
| # rotates (proven by its rotation token in source) and a pruner | |
| # glob line in log-maintenance.sh, credited to that directory's | |
| # own find line -- plus every find line there must belong to a | |
| # ledgered sink. #2196 shipped because mailoney lacked both | |
| # halves at once and nothing structural connected compose | |
| # volumes to vendored writers to this script; adding a sensor | |
| # now means adding one reviewable ledger row instead. | |
| # | |
| # #2216: the checks above only see halves an author already | |
| # touched, so the script also enumerates every | |
| # /opt/stacks/apiary/logs/<dir> bind mount in arcane/home/*/ | |
| # compose.yml -- a new stack's mount is the one thing it cannot | |
| # omit and still write to disk. Each must be a ledger row, an | |
| # EXEMPT entry with a written reason, or a KNOWN_UNCOVERED entry | |
| # naming the issue tracking the gap. | |
| # | |
| # #2826: and a written reason is not evidence. An EXEMPT entry | |
| # that asserts a mechanism ("log-maintenance.sh rotates it", | |
| # "the writer self-bounds") now carries that mechanism as a | |
| # checked key -- deleting the rotate line, or the ported | |
| # rotateIfOversized, fails here instead of leaving the ledger | |
| # asserting coverage that no longer exists. | |
| python scripts/check-json-sink-retention-parity.py | |
| # | |
| # #2921: the writer proof above is a grep, and a grep can't | |
| # tell source from a test asserting on the same token or a | |
| # stale __pycache__/*.pyc left over from a previous run -- | |
| # this exercises tree_contains() against a synthetic fixture | |
| # to prove both false-positive paths stay closed. | |
| python3 -B scripts/tests/test_check_json_sink_retention_parity.py | |
| # #2926: the persona overlay's bare-name entries (free, uptime, | |
| # id, ...) were shadowed by cowrie's own builtins and silently | |
| # unreachable -- nproc disagreeing with lscpu in the same session | |
| # was the cheapest tell. Asserts the priority patch still matches | |
| # upstream's getCommand() exactly once and is idempotent, plus the | |
| # persona-consistency fixtures it unblocks (nproc == lscpu's CPU | |
| # count, free's static seed is persona-scale). | |
| # | |
| # The behavioural half is what earns the row: it execs the patched | |
| # getCommand() and pins the two regressions an unscoped version | |
| # causes -- a path-form invocation (/bin/echo, and 66 other real | |
| # binaries in the runtime image) reading the container's own ELF | |
| # instead of being emulated, and a canned overlay outranking | |
| # uname's cowrie.cfg-driven persona or dd's operand parsing and | |
| # input_data capture. | |
| - name: "Cowrie txtcmds overlay priority patch (#2926)" | |
| home: true | |
| run: python arcane/home/honeypot-cowrie/cowrie/tests/test_txtcmds_priority_patch.py | |
| - name: "cowrie auth-phase service-request patch (#3307)" | |
| home: true | |
| run: python arcane/home/honeypot-cowrie/cowrie/tests/test_service_request_log_patch.py | |
| # #1982: compose.yml is the only place a stack's knobs are exercised, | |
| # .env.example is what a deployer copies when standing the stack up | |
| # manually -- drift between them means a variable exists at runtime | |
| # but can't be discovered at setup time (that's how ZEEK_PROXY_IFACE | |
| # shipped as refuse-to-start-by-default and CANARY_PUBLIC_HOSTNAME as | |
| # silent placeholder-domain). Two rules: every ${VAR} interpolated in | |
| # arcane/home/<stack>/compose.yml must appear as a KEY= line in that | |
| # stack's own .env.example (per-stack, pointers to a sibling don't | |
| # count -- each file is copied on its own), and no key may be listed | |
| # twice within one environment: block (the ZEEK_PROXY_ATTRIBUTION_ | |
| # INTERVAL_SECONDS dup meant tuning line one silently changed nothing). | |
| - name: Compose interpolation documented per stack (#1982) | |
| home: true | |
| run: python scripts/check-compose-env-docs.py | |
| # #119/#2051: a healthcheck (compose-level or baked into the | |
| # service's own Dockerfile) with no autoheal=true label just | |
| # detects a wedge and does nothing about it. #119 fixed the fleet | |
| # once; the gap came back on every stack that didn't exist yet | |
| # when it closed, so this is the standing guard #119 itself asked | |
| # for, scanning both layers -- a compose-only matcher reproduces | |
| # the exact blind spot that let four image-level-HEALTHCHECK | |
| # services hide from the first audit. | |
| - name: Healthcheck/autoheal label pairing (#2051) | |
| home: true | |
| run: | | |
| python -m pip install -q pyyaml | |
| python scripts/check-autoheal-labels.py | |
| # #2320: every WORKER_LOOPS consumer boots through the same | |
| # apiary-backend entrypoint whose gate (#2183) refuses an empty | |
| # SERVICE_TOKEN unless APIARY_ALLOW_UNAUTH_DEV=1 -- both arrive as | |
| # per-container env, so a loop tier forwarding only the token half | |
| # crash-looped on the dev-tier deployment (#2320) while its serving | |
| # siblings booted. Mechanical form of checking every consumer | |
| # against the gate's inputs: if backend ever grows a third gating | |
| # variable, AUTH_VARS in the script and this gate move together. | |
| - name: WORKER_LOOPS consumers carry the apiary-backend auth pair (#2320) | |
| home: true | |
| run: python scripts/check-backend-boot-contract.py | |
| # #2352: #1502 moved the YARA scanner under | |
| # honeypot-payload-analysis/analysis/yara/ and updated the checker | |
| # script, but the surrounding prose never got the same pass -- six | |
| # documents kept pointing readers at the dead repository-root path, | |
| # including the operator README's own sync command, which failed | |
| # verbatim from a fresh checkout. Grep-level guard for that class: | |
| # outside arcane/**, a .md may cite analysis/yara/ only on a line | |
| # that also names its real home (or carries an inline waiver). | |
| - name: Docs cite moved trees by their real home (#2352) | |
| run: python scripts/check-doc-stale-paths.py | |
| # #2962: the orphaned-script sweep (#1609 Phase 6) found these two | |
| # regression tests describe real coverage in their own headers | |
| # (guest-runner.sh's post-detonation poweroff-on-artifact-failure | |
| # path, #2268; run_pending.sh's stale-claim reconciliation, #2253) | |
| # but were never referenced by name, glob, or doc anywhere in the | |
| # tree -- unlike analysis/github/tests/*.sh, nothing ran them. | |
| # Neither needs a hypervisor or network (both stub the | |
| # orchestrator/guest side per their own headers), so plain | |
| # ubuntu-latest is enough -- not the self-hosted KVM host | |
| # guest-runner.sh and run_pending.sh actually run on in | |
| # production. Fixed here as part of wiring them in: | |
| # run_pending.sh's ${VAR:-default} treated the stale-claims test's | |
| # explicit `export WINDOWS_SANDBOX_KVM_SHARED_LOCK=""` (meant to | |
| # disable the shared lock, per the script's own "Empty disables | |
| # it" comment) the same as unset, so the never-run test would | |
| # have flocked the real production lock path the first time it | |
| # ran anywhere with permission to. | |
| - name: "Sandbox shell tests (#2268, #2253)" | |
| run: | | |
| for t in sandbox/windows/tests/*.sh; do | |
| echo "=== $t ==="; bash "$t" || exit 1 | |
| done | |
| # #2456: vps/zeek and dev/sensing-lab build the same parser set on | |
| # purpose -- the lab's measurements are only evidence for production | |
| # while production runs the same parsers. A component pinned in one | |
| # file but floating (or pinned to a different commit) in the other | |
| # is exactly the lab/prod drift the #1821 hassh pin pattern exists | |
| # to prevent; this makes it a red build instead of a silent rebuild | |
| # difference. Offline and stdlib-only, so it rides the docker-less | |
| # homeserver rows. | |
| - name: Zeek lab/prod build pin parity (#2456) | |
| home: true | |
| run: python scripts/check-zeek-pin-parity.py | |
| # #2458: the pointer-rot family (#2352, #2353, #2356/#2453, #2455, | |
| # #2357, five docs' worth invalidated by #1659) kept being fixed by | |
| # hand after PR #2367's one-off sweep script proved the shape. This | |
| # makes that sweep standing: every repo-path-like citation in | |
| # docs/** and the root README must resolve to a git-tracked path. | |
| # Offline, stdlib-only; deliberate era-record and host-side | |
| # references live in scripts/doc-path-lint-allowlist.txt with | |
| # per-entry reasons instead of passing silently. | |
| - name: Docs cite existing repo paths (#2458) | |
| home: true | |
| run: python scripts/check-doc-paths-exist.py | |
| # #3332: the inverse of #2458 -- every docs/**/*.md must be reachable | |
| # from docs/README.md through relative links (transitively), except | |
| # dated record trees listed in the script. 54 of 120 docs, runbooks | |
| # among them, were unreachable from the map when this landed. | |
| - name: Docs reachable from the documentation map (#3332) | |
| home: true | |
| run: python scripts/check-docs-reachable.py | |
| # #3395: #2458 scans only README.md and docs/**, so the component | |
| # trees under arcane/**, analysis/**, sandbox/** and branding/** were | |
| # never link-checked -- the agent-intrusion-corpus README's dead | |
| # `../../docs/...` hop (two levels short of a five-deep path) lived | |
| # there unnoticed. Whole-tree relative-link resolution. Fenced blocks | |
| # are skipped: their paths belong to the reader's project, not ours. | |
| - name: Every tracked doc's relative links resolve (#3395) | |
| home: true | |
| run: python scripts/check-doc-links.py | |
| # #3395: two of the forty mermaid diagrams in the tree did not parse | |
| # at all -- AI_TRIAGE.md named a node `call` (a reserved mermaid | |
| # token) and ghidra/README.md left a colon unquoted inside an edge | |
| # label. Both rendered as an error box on GitHub and nothing noticed, | |
| # so mermaid parsing is now a gate rather than a review habit. Renders | |
| # every block in one headless browser; mermaid is resolved from the | |
| # npx cache, so the step is self-bootstrapping (no pinned version to | |
| # drift out of date with the docs). The warm-up below is what makes | |
| # that true: a fresh runner's npx cache is empty, so the lookup inside | |
| # check-mermaid.mjs finds nothing and the gate cannot run at all. | |
| - name: Mermaid diagrams parse (#3395) | |
| home: true | |
| run: | | |
| npx -y @mermaid-js/mermaid-cli --version | |
| node scripts/check-mermaid.mjs | |
| # #2576: the CAPE implementation-plan doc claimed the detail page | |
| # (/cape/{sha256}) was admin-gated. Nothing in backend-service | |
| # enforces that -- require_service_token is the actual gate on | |
| # /api/v1/cape/{sha}, the same middleware every other /api/v1 | |
| # detail route carries (main.rs:183, wired at main.rs:468). A real | |
| # admin-role check does exist, but on a different route entirely | |
| # (the Workbench "cape" analyzer entry that triggers a fresh | |
| # detonation) -- this regression-tests the doc text, not the code, | |
| # so the wrong claim can't silently drift back in on a rewrap or a | |
| # synonym swap. pytest installed inline, same as the other | |
| # pytest-based rows above -- it isn't a production dependency. | |
| - name: Docs regression tests (#2576) | |
| home: true | |
| run: | | |
| python -m pip install pytest | |
| python -m pytest tests/docs/ -v | |
| - name: "Dionaea SMB exploit-signature incidents (#622, #1775)" | |
| home: true | |
| run: | | |
| # Executes the code this patch injects into smb.py, against a | |
| # stand-in connection, rather than asserting on the substituted | |
| # text -- the property that matters is one incident per | |
| # connection, not one per SMB packet. Reporting per packet made | |
| # a single DoublePulsar connection produce ~643 documents and | |
| # 47.5% of everything the fleet ingested. | |
| python arcane/home/honeypot-dionaea/dionaea/tests/test_smb_exploit_patch.py | |
| - name: Ghidra worker, spool discipline and AI triage | |
| home: true | |
| run: | | |
| # Stdlib only and stubbed on both sides, so it runs here in seconds. | |
| # Worth running: this worker decides what reaches an analyst's screen | |
| # from a live malware sample, and the triage half will not send | |
| # sample-derived text to a non-local model endpoint — a rule that is | |
| # only worth anything if something checks it on every change. | |
| python analysis/ghidra/worker/tests/test_ghidra_worker.py | |
| # GPU-queue drain crash recovery (#2075): a drainer death | |
| # mid-generation used to strand its job as an eternally-running | |
| # zombie. Stubbed queue, no ES/docker/nvidia-smi -- the sweep | |
| # must requeue-once-then-fail past the staleness bound, never | |
| # touch a legitimately-running job, and run before the queued | |
| # early return that used to hide the problem. | |
| python analysis/ghidra/worker/tests/test_gpu_queue_drain.py | |
| # Replacement Ghidra REST service (#245): fake analyzeHeadless, no | |
| # real Ghidra/JVM in CI -- verifies server.py's HTTP/queue layer | |
| # only, same "stub both sides" reasoning as the worker test above. | |
| python analysis/ghidra/service/tests/test_server.py | |
| # Governance tests use synthetic snapshots/reports only. CI never | |
| # pulls or loads multi-GB model weights. | |
| python analysis/ghidra/models/tests/test_model_governance.py | |
| python analysis/ghidra/models/tests/test_model_status_adapter.py | |
| # #3334: the host-side wiring that actually runs the injection | |
| # corpus weekly and on a pin change. The corpus itself needs the | |
| # real model, so what is provable without a GPU is the part that | |
| # made it dead weight before: that the units exist, that the | |
| # .path unit watches files the installer deploys, and that the | |
| # runner fails closed and skips only when nothing changed. Driven | |
| # against a stub `docker` on PATH -- no model, no Ollama. | |
| python analysis/ghidra/models/tests/test_llm_injection_suite.py | |
| # Benchmark transcript records (#1805). No model involved -- this | |
| # asserts the prompt is stored as sent, that a refusal or timeout | |
| # is stored rather than dropped, that a stored transcript is never | |
| # rewritten, and that captured real-data transcripts cannot be | |
| # written inside the repo. That last one is a data-handling rule, | |
| # so it is worth a gate rather than reviewer vigilance. | |
| python analysis/ghidra/benchmarks/tests/test_transcripts.py | |
| # Tier B evidence cache (#1805). No Ghidra and no service in CI -- | |
| # covers the cache key (every component that shapes the evidence | |
| # must change it, or a Ghidra upgrade silently reuses output from | |
| # a different decompiler) and the injection assertion, which must | |
| # report honestly that the payload never reached the evidence. | |
| python analysis/ghidra/benchmarks/tests/test_ghidra_cache.py | |
| # Cross-path fidelity for the ghidra slot (#1805). No model, no | |
| # Ghidra, no service. The other Tier B tests all synthesise their | |
| # cache entries, so they agree with whatever the loader happens to | |
| # expect; this one drives the same path over real recorded Ghidra | |
| # output, which is what makes a drift in the expected response | |
| # shape visible. Also measures the standing gap between the | |
| # qualification gate's hand-written fixtures and production's | |
| # _evidence() output, so a change to either renderer is a | |
| # measured delta rather than a silent one. | |
| python analysis/ghidra/benchmarks/tests/test_ghidra_fidelity.py | |
| # Corpus manifest validator (#2038). Pure structural checks over | |
| # fixture manifests -- no compiler, no corpus rebuild. Covers the | |
| # three gaps that let a broken corpus pass: a required alternative | |
| # that trips its own case's forbidden list, a rubric case whose | |
| # build has disappeared, and a partial toolchain/opt-level grid. | |
| # The overlap check reuses polarity.forbidden_hit(), so the guard | |
| # that keeps it from regressing to #2517's naive-substring false | |
| # positives lives here too and has to run on every change. | |
| python analysis/ghidra/benchmarks/tests/test_validate_manifest.py | |
| # Claim-pool scoring (#1805). Embedder is stubbed, so no Ollama. | |
| # Covers the properties that decide whether a claim-pool score can | |
| # be believed: the adjudicator cannot be a contestant, rephrasing | |
| # earns nothing, unadjudicated claims never count as correct, and | |
| # the pool version moves when a verdict does. | |
| python analysis/ghidra/benchmarks/tests/test_claims.py | |
| # Corpus scorer (#1952). An empty answer used to collect the | |
| # injection-resistance point, giving a model that returned nothing | |
| # a floor of 14/69 -- which is how a thinking model scoring its | |
| # empty-answer floor read as a fifth of a pass instead of a zero. | |
| python analysis/ghidra/benchmarks/tests/test_record_baseline.py | |
| # Session-slot critical gate (#2232). The rubric-vocabulary leg | |
| # no longer gates: only MITRE correctness, forbidden content, | |
| # and severity can fail critical_ok, since governance treats | |
| # that boolean as disqualifying (#1947 rule 4) and five of | |
| # eleven round models proved the old AND tripped on wording. | |
| python analysis/ghidra/benchmarks/tests/test_session_scoring.py | |
| # Harmony-family serving adaptation (#2233). gpt-oss tags through | |
| # Ollama's /api/chat need a different wire shape or `content` | |
| # comes back empty; pins the dispatch so the calibrated Qwen | |
| # request shape is untouched. | |
| python analysis/ghidra/benchmarks/tests/test_harmony_chat.py | |
| # Offline transcript re-scoring (#2266). No GPU and no Ollama -- | |
| # the fixtures are written through the real TranscriptWriter and | |
| # replayed, so this pins the property the tool exists for: a | |
| # rescore of a stored answer scores identically to what a live | |
| # evaluate_slot() would have scored for it. Also pins | |
| # scorer_git_sha's dirty marker, without which the documented | |
| # "run at two commits and diff the reports" workflow silently | |
| # attributes both sides to the same commit. | |
| python analysis/ghidra/benchmarks/tests/test_rescore_from.py | |
| # #2980: the four files below existed with no workflow running | |
| # them. This block is hand-enumerated -- there is no pytest | |
| # discovery over analysis/ghidra/benchmarks/tests/ -- so a test | |
| # file lands here only if someone remembers, and four had drifted | |
| # off. test_injection_gate.py is the sharpest case: its 48 tests | |
| # are the whole regression guarantee the #2694 Tier B gate rests | |
| # on, and they were executing only by hand. All four are stubbed | |
| # the same way as the rows above (no model, no Ghidra, no | |
| # network) and run standalone. | |
| python analysis/ghidra/benchmarks/tests/test_injection_gate.py | |
| python analysis/ghidra/benchmarks/tests/test_corpus_eval.py | |
| python analysis/ghidra/benchmarks/tests/test_run_real_corpus_eval.py | |
| python analysis/ghidra/benchmarks/tests/test_regenerate_pre_2393.py | |
| # full_capabilities.py (#800): inventories env vars straight out of | |
| # the real pipeline source, so its own test runs against the real | |
| # tree rather than fixtures -- proving that stays true. | |
| python analysis/ghidra/tests/test_full_capabilities.py | |
| # CAPE worker's status.json discipline (#319 follow-up). No live | |
| # CAPE involved -- see the test file's own docstring for why that | |
| # stays with --selftest --round-trip (#318) instead. | |
| python sandbox/cape/worker/tests/test_cape_worker.py | |
| - name: Real-data session probe metadata honesty (#2387) | |
| home: true | |
| run: | | |
| # contracts is stubbed, so /app, pydantic, ES, docker and Ollama | |
| # stay out of CI. Pins both halves of the metadata honesty fix: | |
| # stage 0's correlated auth/duration evidence reaches | |
| # session_prompt() unchanged (the old code pinned | |
| # auth_success=False / duration=0.0 under an EXACT-prompt claim), | |
| # and sessions whose evidence missed the lookback window are | |
| # reported and skipped instead of being prompted with defaults. | |
| # Also drives the extracted stage-0 jq program (#2426) against | |
| # committed fixtures -- latest-close-wins, string-duration | |
| # coercion, failed-only auth, empty second-response gap -- on | |
| # plain jq; skips itself gracefully where jq is absent so this | |
| # row stays homeserver-first per #2389 without dropping coverage. | |
| python analysis/ghidra/benchmarks/tests/test_probe_real_session.py | |
| - name: Statictools lief contract (#2072) | |
| home: true | |
| run: | | |
| # lief_parse() is the structural read on PE/Mach-O/ELF samples, | |
| # and lief was the one unpinned install in an image whose whole | |
| # point is that nothing drifts underneath it silently (#2072). | |
| # CI installs the version the Dockerfile pins -- read back out | |
| # of that Dockerfile, so the tested pin and the shipped pin | |
| # cannot diverge. | |
| LIEF_PIN=$(grep -om1 'lief==[0-9.]*' analysis/ghidra/statictools/Dockerfile | cut -d '=' -f3) | |
| python -m pip install -q "lief==${LIEF_PIN}" | |
| python analysis/ghidra/statictools/tests/test_lief_parse.py | |
| - name: Guarded LLM worker contracts | |
| home: true | |
| run: | | |
| python -m pip install -r llm-worker/requirements.txt | |
| python -m unittest discover -s llm-worker/tests -v | |
| python llm-worker/worker.py --selftest | |
| # ml-worker/, analysis/es-results-importer/, personas/, and | |
| # services-adapter/ are all deployed by deploy.yml but had zero test | |
| # coverage wired into any workflow -- 189 tests across 11 files | |
| # (found while auditing #982's own CI gap; this is the same class of | |
| # issue, just for four whole modules instead of two scripts). | |
| # ml-worker's suite uses pytest specifically (parametrized fixtures), | |
| # unlike every other Python test dir in this workflow, which is | |
| # plain unittest -- pytest isn't a production dependency so it's | |
| # installed here rather than added to ml-worker/requirements.txt. | |
| # #3319: --junitxml is pytest's own machine-readable reporter, so | |
| # this needs no new dependency. Both pytest rows carry it and both | |
| # write into the one .ci-artifacts/ dir the upload step below | |
| # collects; the other ~20 rows are plain unittest/shell/go and | |
| # simply produce no file, which `if-no-files-found: ignore` treats | |
| # as the normal case rather than a warning on every one of them. | |
| - name: ml-worker anomaly pipeline contracts | |
| home: true | |
| run: | | |
| python -m pip install -r ml-worker/requirements.txt pytest | |
| mkdir -p .ci-artifacts | |
| python -m pytest ml-worker/tests/ -v --junitxml=.ci-artifacts/ml-worker-junit.xml | |
| # #2219: Keycloak caps every admin-events page silently, and the | |
| # single-page fetch permanently dropped everything past row 1,000 | |
| # on exactly the credential-spray days this telemetry exists for. | |
| # The pagination contracts (stubbed HTTP layer, no live service) | |
| # pin the drained/undrained pager and its checkpoint-hold rule. | |
| # pytest again only installed here -- not a production dependency. | |
| - name: auth-events-worker Keycloak pagination contracts | |
| home: true | |
| run: | | |
| python -m pip install -r auth-events-worker/requirements.txt pytest | |
| mkdir -p .ci-artifacts | |
| python -m pytest auth-events-worker/tests/ -v \ | |
| --junitxml=.ci-artifacts/auth-events-worker-junit.xml | |
| - name: es-results-importer chunked-upload and dedup contracts | |
| home: true | |
| run: | | |
| python -m pip install -r arcane/home/honeypot-dashboard/analysis/es-results-importer/requirements.txt | |
| python -m unittest discover -s arcane/home/honeypot-dashboard/analysis/es-results-importer/tests -v | |
| - name: Persona application is idempotent | |
| home: true | |
| run: python -m unittest discover -s personas/tests -v | |
| - name: services-adapter allowlist enforcement | |
| home: true | |
| run: python -m unittest discover -s arcane/home/honeypot-dashboard/services-adapter/tests -v | |
| - name: Sandbox guest string-cleaning filter (#530) | |
| home: true | |
| run: python3 sandbox/test_guest_clean_strings.py -v | |
| # #2878: these two rows used to name test_extract_iocs.py (#482) and | |
| # test_run_sample_golden_check.py (#100, #2023) individually, so the | |
| # two delivery-check suites #2252 added next to them (Windows and | |
| # GHOSTS) ran nowhere at all -- a regression test nothing runs is a | |
| # file, not a check. Discovery instead, the way the worker suites | |
| # are already run, so the next test_*.py dropped in these | |
| # directories is picked up without a workflow edit. #2775's | |
| # legacy-WinRM result test is the first to arrive that way. | |
| - name: "Windows sandbox orchestrator suites (#482, #2023, #2252, #2775)" | |
| home: true | |
| run: python3 -m unittest discover -s sandbox/windows/orchestrate -p 'test_*.py' -v | |
| - name: GHOSTS sandbox orchestrator suites (#2252) | |
| home: true | |
| run: python3 -m unittest discover -s sandbox/ghosts/orchestrate -p 'test_*.py' -v | |
| # #3314: .github/actionlint.yaml existed but nothing ran actionlint. | |
| # Blocking. SHELLCHECK_OPTS matches the repo's "high-severity | |
| # ShellCheck" policy: info-level notes in run: blocks do not fail. | |
| # Archive checksum is upstream's (actionlint_1.7.7_checksums.txt). | |
| - name: "Workflow lint (actionlint, #3314)" | |
| home: true | |
| run: | | |
| set -euo pipefail | |
| tgz="$RUNNER_TEMP/actionlint.tgz" | |
| curl -fsSL -o "$tgz" https://github.com/rhysd/actionlint/releases/download/v1.7.7/actionlint_1.7.7_linux_amd64.tar.gz | |
| echo "023070a287cd8cccd71515fedc843f1985bf96c436b7effaecce67290e7e0757 $tgz" | sha256sum -c - | |
| tar -xzf "$tgz" -C "$RUNNER_TEMP" actionlint | |
| SHELLCHECK_OPTS="-S warning" "$RUNNER_TEMP/actionlint" -color | |
| # A plain-scalar `name: Foo (bar, #123)` is valid YAML that silently | |
| # truncates at " #" (a comment) -- eight names were shipping as | |
| # "Foo (bar," before #3314. actionlint cannot see it; this can. | |
| if grep -nE '^[[:space:]]*(- )?name: [^"'"'"'].* #' .github/workflows/*.yml; then | |
| echo "::error::unquoted name: containing ' #' -- quote it, YAML reads the rest as a comment" | |
| exit 1 | |
| fi | |
| # #3313: every SHA-pinned third-party `uses:` must carry a full | |
| # `# vX.Y.Z` comment. zizmor's unpinned-uses audit reads the SHA | |
| # and ignores the comment entirely, so a pin can be perfectly | |
| # well-formed and still say the wrong thing: two pins here were | |
| # labelled `# v7` where the SHA is v7.0.0, and one was labelled | |
| # `# v6` for a v4.6.2 SHA. A reviewer reading the comment is | |
| # misled about what a job actually runs, which defeats the point | |
| # of recording it. Offline and deterministic -- no network, so it | |
| # cannot flake. Dependabot rewrites both halves of a pin on a | |
| # bump, so this stays satisfied as versions move. | |
| if grep -nEo 'uses: +[A-Za-z0-9._-]+/[A-Za-z0-9._/-]+@[0-9a-f]{40}( *# *[^ ]+)?' \ | |
| .github/workflows/*.yml \ | |
| | grep -vE '@[0-9a-f]{40} *# *v[0-9]+\.[0-9]+\.[0-9]+$'; then | |
| echo "::error::action pin without a full '# vX.Y.Z' comment -- SHA is pinned but its version is not stated, or is stated as a bare major" | |
| exit 1 | |
| fi | |
| # #3314: zizmor security audit of the workflows. Fails on every | |
| # medium+ finding EXCEPT the two rules left in ADVISORY below. | |
| # #3313 landed the other three (unpinned-uses, excessive-permissions | |
| # and artipacked) across the whole tree, so they are now dropped | |
| # from ADVISORY and a regression in any of them is blocking. zizmor | |
| # publishes no checksum file; the pin below was recorded on first | |
| # download (2026-09-26). | |
| - name: "Workflow security audit (zizmor, #3314)" | |
| home: true | |
| run: | | |
| set -euo pipefail | |
| tgz="$RUNNER_TEMP/zizmor.tgz" | |
| curl -fsSL -o "$tgz" https://github.com/zizmorcore/zizmor/releases/download/v1.30.1/zizmor-x86_64-unknown-linux-gnu.tar.gz | |
| echo "e65324f4430c2717591937edcec90ccbefaf14c174f8ec9415e03ca875b46e1a $tgz" | sha256sum -c - | |
| tar -xzf "$tgz" -C "$RUNNER_TEMP" zizmor | |
| "$RUNNER_TEMP/zizmor" --offline --min-severity medium --format json .github/workflows > "$RUNNER_TEMP/zizmor.json" || true | |
| python3 - "$RUNNER_TEMP/zizmor.json" <<'PY' | |
| import collections, json, sys | |
| # Both remaining advisory rules are Low/High-by-rule but not | |
| # exploitable as written, and neither is a checkout/pin/permission | |
| # property -- so #3313 does not own them: | |
| # | |
| # self-repository: the five `uses: ./.github/workflows/ci-router.yml` | |
| # calls. zizmor wants the explicit owner/repo/path@ref form; the | |
| # local-path form is GitHub's own first-party-reusable-workflow | |
| # syntax and is what makes the router pick up edits to the | |
| # shared router without a second SHA to bump. Every one of the | |
| # five is triggered by pull_request/push/schedule, never by | |
| # pull_request_target, so the called workflow is always the | |
| # default branch's own file. | |
| # | |
| # dangerous-triggers: main-health-watch.yml's `workflow_run`. It | |
| # runs the default branch's copy of the script and never reads | |
| # the triggering event: every value it reports comes from a | |
| # main-scoped `gh api`/`gh run list` read, and the only | |
| # untrusted-looking text it echoes is main's own commit | |
| # subjects. Changing the trigger would blind the #3324 alarm | |
| # that exists precisely to catch runs nobody started. | |
| ADVISORY = {"self-repository", "dangerous-triggers"} | |
| findings = json.load(open(sys.argv[1])) | |
| by_rule = collections.Counter(f["ident"] for f in findings) | |
| for rule, n in sorted(by_rule.items()): | |
| print(f"{'advisory' if rule in ADVISORY else 'BLOCKING'} {n:3d} {rule}") | |
| blocking = [f for f in findings if f["ident"] not in ADVISORY] | |
| for f in blocking: | |
| loc = f["locations"][0]["symbolic"] | |
| print(f"::error::zizmor {f['ident']}: {f['desc']} ({loc['key']})") | |
| sys.exit(1 if blocking else 0) | |
| PY | |
| # #3320: Dockerfile lint, separate from image-security-scan's CVE | |
| # check. Policy (and why package-version pins are exempt) lives in | |
| # .hadolint.yaml; fails on warnings and errors. Pinned binary, | |
| # checksum pinned here -- not fetched next to the download. | |
| - name: "Dockerfile lint (hadolint, #3320)" | |
| home: true | |
| run: | | |
| set -euo pipefail | |
| bin="$RUNNER_TEMP/hadolint" | |
| curl -fsSL -o "$bin" https://github.com/hadolint/hadolint/releases/download/v2.14.0/hadolint-linux-x86_64 | |
| echo "6bf226944684f56c84dd014e8b979d27425c0148f61b3bd99bcc6f39e9dc5a47 $bin" | sha256sum -c - | |
| chmod +x "$bin" | |
| git ls-files -z | grep -zE '(^|/)Dockerfile[^/]*$' | xargs -0 "$bin" --config .hadolint.yaml | |
| # #3501: fail CI on a real credential anywhere in FULL git history. | |
| # scripts/check-git-secrets.py landed in 8cc55aa9 with nothing calling | |
| # it -- `grep -rn gitleaks .github/workflows/` was empty -- so that | |
| # commit's "fail CI on a real credential" claim was untrue. This row | |
| # is what makes it true, and it is the only place gitleaks is on PATH. | |
| # | |
| # Its own lane rather than a row on the public-leak side: that check | |
| # walks `git ls-files`, so it can only ever see the checked-out tree. | |
| # A credential committed once and deleted in the next commit is | |
| # still a credential, and this is the gate that sees it. | |
| - name: "Full-history secret scan (gitleaks, #3501)" | |
| home: true | |
| run: | | |
| set -euo pipefail | |
| # The scan's entire claim is FULL history, and actions/checkout | |
| # defaults to fetch-depth 1 -- so on a pull_request the | |
| # merge-commit checkout hands gitleaks exactly one commit and the | |
| # gate passes vacuously. That is the worst failure mode a secret | |
| # scanner has: green, having measured nothing. Unshallow first, | |
| # then prove the history is really there before believing a pass. | |
| if [ "$(git rev-parse --is-shallow-repository)" = true ]; then | |
| git fetch --unshallow --no-tags origin | |
| fi | |
| commits="$(git rev-list --count HEAD)" | |
| if [ "$commits" -lt 100 ]; then | |
| echo "::error::secret-scan: only $commits commit(s) under HEAD -- refusing to report a clean scan of a truncated history" | |
| exit 1 | |
| fi | |
| echo "secret-scan: scanning $commits commits of full history" | |
| # Pinned version + sha256 verified before extraction, unpacked | |
| # into $RUNNER_TEMP because the self-hosted honeypot-ci runner | |
| # cannot write /usr/local/bin -- the constraint | |
| # scripts/install-trivy.sh already works under, and the reason | |
| # `home: true` is right here: checkout, this workflow's own | |
| # setup-python, a download into a writable temp dir. No docker, | |
| # no host state, nothing the metal cannot do. | |
| # The install has to reach THIS step, not a later one. Each `run:` | |
| # is its own process, and GITHUB_PATH is applied by the runner to | |
| # *subsequent* steps -- so `scripts/install-gitleaks.sh` writing | |
| # it (install-gitleaks.sh:123) left the very next line scanning | |
| # with a PATH that never had gitleaks on it, which is the exit 2 | |
| # #3516's CI run reported. --print-path names the binary the pin | |
| # verified; exporting its directory puts it on PATH for the rest | |
| # of this block. The GITHUB_PATH write stays in the script for | |
| # later steps -- it is right there, it is just not enough here. | |
| export PATH="$(dirname "$(scripts/install-gitleaks.sh --print-path)"):$PATH" | |
| # Prove the propagation before the scan, so a broken PATH fails | |
| # as "gitleaks missing" next to the install, not as an opaque | |
| # exit 2 from a scan that never ran. | |
| command -v gitleaks | |
| # check-git-secrets.py exits 2 when gitleaks fails TO RUN and 1 | |
| # when it finds something unallowlisted. Both are failures here | |
| # and neither is softened: under `set -euo pipefail` an exit 2 | |
| # cannot be mistaken for a clean scan, which is the flagged vs | |
| # unresolved split #2763 forced for trivy. Nothing in this step | |
| # appends `|| true` or pipes the status away. | |
| python3 scripts/check-git-secrets.py | |
| # The gate's own regression suite. It lives in tests/docs/ and | |
| # the docs-regression row above already collects it, but its | |
| # four gitleaks-backed tests skip on a box with no binary -- and | |
| # the docs row has none. This row does, so this is the only place | |
| # they actually execute in CI; without it the suite would skip | |
| # forever and quietly stop guarding the gate (#2981's shape). | |
| python -m pip install pytest | |
| python -m pytest tests/docs/test_3501_secret_scan_allowlist.py -q | |
| - name: Shell syntax and high-severity ShellCheck | |
| home: true | |
| run: | | |
| # #2389: use whatever shellcheck already exists -- the | |
| # honeypot-ci host ships it for its runner user, so the | |
| # self-hosted leg never touches sudo; only an image without | |
| # the binary pays apt. No privilege class is added anywhere; | |
| # the install path is unchanged where it was required. | |
| command -v shellcheck >/dev/null 2>&1 || { sudo apt-get update && sudo apt-get install -y shellcheck; } | |
| # A couple of sandbox/windows/packer/pxe/*.sh files are actually | |
| # Python (named .sh to match their sibling shell scripts' calling | |
| # convention, e.g. `packer trigger-callback... .sh`) -- filter to | |
| # files whose own shebang names a shell, rather than assuming | |
| # every *.sh is one, so bash -n/shellcheck don't choke on them. | |
| shell_scripts=() | |
| while IFS= read -r -d '' f; do | |
| read -r shebang < "$f" || true | |
| case "$shebang" in | |
| '#!'*/sh|'#!'*/bash|'#!'*env\ sh|'#!'*env\ bash) shell_scripts+=("$f") ;; | |
| esac | |
| done < <(find . -name '*.sh' -type f -print0) | |
| if [ "${#shell_scripts[@]}" -gt 0 ]; then | |
| printf '%s\0' "${shell_scripts[@]}" | xargs -0 -n1 bash -n | |
| printf '%s\0' "${shell_scripts[@]}" | xargs -0 shellcheck --severity=error | |
| fi | |
| - name: Validate Keycloak realm policy | |
| home: true | |
| run: ./arcane/home/honeypot-keycloak/keycloak/realm/validate.sh | |
| # #982 phase 1 / #1040: validate.sh only checks structure/policy -- | |
| # this actually imports the realm into a real, disposable Keycloak + | |
| # PostgreSQL the way the real fresh-install bootstrap path does, to | |
| # catch what only a real import catches (e.g. a role description over | |
| # Postgres's varchar(255) column limit crash-looped the container on | |
| # every fresh install, undetected by any static check). | |
| - name: Keycloak realm imports cleanly into a real Postgres | |
| home: true | |
| run: ./scripts/test-keycloak-realm-import.sh | |
| # #977: proves the isolated oauth2-proxy gateway pattern every | |
| # protected app (Kibana/EveBox/Arkime/TANNER/RevDeck/Dockge/Traefik) | |
| # uses against a real disposable Keycloak + real oauth2-proxy -- | |
| # unauthenticated redirect target, forged-callback rejection, role | |
| # enforcement, upstream network isolation, and gateway-outage | |
| # fail-closed behavior. | |
| - name: oauth2-proxy gateway pattern is resilient | |
| home: true | |
| run: ./scripts/test-oauth2-proxy-gateway-resilience.sh | |
| # #982's PKCE+TOTP login and Keycloak-outage/key-rotation chaos | |
| # suites, retired with the Go dashboard (#1659), restored for | |
| # dashboard-next (#1661): both drive the REAL BFF build output | |
| # (`node .output/server/index.mjs`, the same artifact the container | |
| # runs) against a real disposable Keycloak importing the actual | |
| # realm -- full authorization-code +PKCE logins through the realm's | |
| # mandatory TOTP factor, persisted session-role verification | |
| # (#1656), single-use codes + forged-state rejection, server-side | |
| # logout revocation (#1094), first-login password reset (#1036), | |
| # then outage-persistence / KC-down-logout / restart-recovery / | |
| # signing-key-rotation chaos. One shared build feeds both via | |
| # DASHBOARD_BFF_SKIP_BUILD. | |
| - name: Set up Node for the dashboard OIDC suites | |
| uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 | |
| with: | |
| node-version: "22" | |
| cache: npm | |
| cache-dependency-path: arcane/home/honeypot-dashboard/frontend-next/package-lock.json | |
| - name: Build dashboard-next BFF once for both OIDC suites | |
| home: true | |
| run: | | |
| cd arcane/home/honeypot-dashboard/frontend-next | |
| npm ci --no-audit --no-fund | |
| npm run build | |
| - name: "OIDC PKCE+TOTP login suite (port of #982)" | |
| home: true | |
| env: | |
| DASHBOARD_BFF_SKIP_BUILD: "1" | |
| run: ./scripts/test-dashboard-oidc-pkce-totp-login.sh | |
| - name: "Keycloak outage/restart/rotation chaos suite (port of #982)" | |
| home: true | |
| env: | |
| DASHBOARD_BFF_SKIP_BUILD: "1" | |
| run: ./scripts/test-dashboard-oidc-chaos.sh | |
| - name: GitHub-analysis publisher stays dry-run by default | |
| home: true | |
| run: | | |
| # The one property this whole feature depends on (#74): | |
| # GITHUB_PUBLISH_ENABLED unset must never reach publish-sample.sh, | |
| # the only thing that can commit, zip, or push. WORK-LEDGER.md | |
| # rule 7 requires this be a test, not a convention -- run it on | |
| # every change, not just when analysis/github/ itself changes, | |
| # since a change elsewhere (an env var default, a shared helper) | |
| # could silently break the gate just as easily. | |
| for t in analysis/github/tests/*.sh; do | |
| echo "::group::$t" | |
| bash "$t" | |
| echo "::endgroup::" | |
| done | |
| - name: Threat-intel CIDR refresh script (#244) | |
| home: true | |
| run: analysis/threat-intel/tests/test_refresh_threat_cidrs.sh | |
| # #2024: the Keycloak dump is the backup's most sensitive artifact, | |
| # and under dash a `pg_dump | gzip` pipeline can only surface | |
| # gzip's exit status -- the exact silent-truncation class behind | |
| # #1413's post-mortem. Drives the real script with docker stubbed | |
| # on PATH: a mid-stream dead pg_dump (and an rc-lost footer-less | |
| # one) must keep no artifact, a healthy dump must be byte-exact | |
| # through gzip, and a missing container must stay quiet. Runs in | |
| # plain sh, because dash compliance IS the bug. #2025 extends the | |
| # same suite to retention running before failure-prone steps, the | |
| # manual-run umask window, and no empty stamp on failed validation. | |
| - name: "backup-honeypot dump discipline, retention order and umask (#2024, #2025)" | |
| home: true | |
| run: sh analysis/tests/test_backup_honeypot.sh | |
| # The four real-Elasticsearch suites. #2389 left them unflagged | |
| # because without a container engine they print "SKIP: docker | |
| # daemon is not reachable" and go green while testing nothing; | |
| # the runner user now drives docker (#2565), so on the box they | |
| # exercise their actual ES containers -- which is exactly the | |
| # difference between a relocated failure and a relocated no-op. | |
| - name: honeypot-30d ILM policy rolls over and deletes (#585) | |
| home: true | |
| run: analysis/tests/test_honeypot_ilm_rollover.sh | |
| - name: geoip-honeypot ingest pipeline (#563) | |
| home: true | |
| run: analysis/tests/test_geoip_pipeline.sh | |
| - name: dionaea-incidents index template (#565) | |
| home: true | |
| run: analysis/tests/test_dionaea_incidents_index.sh | |
| - name: conpot persona extraction in geoip-honeypot pipeline (#567) | |
| home: true | |
| run: analysis/tests/test_conpot_persona_pipeline.sh | |
| # #789's sensor event-kind coverage audit used to run here. It is | |
| # gone rather than disabled, per #1665: it diffed sensor source | |
| # against dashboard/classify.go's per-sensor switch, and #1659 | |
| # deleted that file with the Go dashboard. Nothing replaced the | |
| # switch -- the geoip-honeypot pipeline copies honeypot.category | |
| # through when a sensor supplies one and does nothing when it does | |
| # not, so there is no case table left to diff against. | |
| # | |
| # The concern itself survives and got worse, so the script was | |
| # rewritten to measure it from the data instead: 20 of 22 sensors | |
| # label no events at all, 2% coverage overall. That needs a | |
| # populated Elasticsearch, which this job does not have, so it is | |
| # an operational audit now -- scripts/audit-sensor-event-coverage.py | |
| # on the homeserver, not a CI step. | |
| - name: Shared GPU job queue | |
| home: true | |
| run: analysis/gpu-queue/test_gpu_queue.py | |
| # #1971: shared ES consume idioms. The canonical suite walks the | |
| # vendored-copy registry (byte-for-byte vs every consumer copy, | |
| # the gpu_queue.py discipline), drives the Python engine through | |
| # the hand-computed cross-language fixture stream, and pins the | |
| # inclusive-gte query shape -- the exact #168 boundary semantics. | |
| # The Go twin of that same fixture stream runs in the go-test | |
| # job (attacker-identity-worker's TestParityFixtures); disagreement | |
| # between the two engines fails whichever side drifts. | |
| - name: Shared ES consume patterns (#1971) | |
| home: true | |
| run: python3 analysis/es-consume/tests/test_es_consume.py | |
| - name: Payload/artifact dedupe (#481/#528) | |
| home: true | |
| run: | | |
| # test_dedupe_payloads.py and test_cdc_dedup_prototype.py had | |
| # zero CI coverage before this -- no workflow referenced either | |
| # by name. Closing that gap alongside #528's own new modules | |
| # rather than leaving it, since it's the same area of the tree | |
| # and cheap to fix now that it's been noticed. | |
| python3 analysis/tests/test_dedupe_payloads.py -v | |
| python3 analysis/tests/test_cdc_dedup_prototype.py -v | |
| python3 analysis/tests/test_procmon_cdc_store.py -v | |
| python3 analysis/tests/test_archive_diagnostics.py -v | |
| # #1984: analyze.py prints attacker-controlled fields into an | |
| # analyst's terminal -- the one place honeypot content reaches a | |
| # human interactively. Every control character must leave | |
| # print_table inert (<0xNN> spellings) while clean reports stay | |
| # byte-identical; this is what stops an ESC payload in a password | |
| # from repositioning or rewriting the reader's terminal. #1985 | |
| # extends the same suite to the stats-hygiene batch (silent | |
| # malformed-line skips, dead parameters, total undercount, the | |
| # category catch-all, multipot VNC login counting). | |
| - name: "analyze.py sanitisation and stats hygiene (#1984, #1985)" | |
| home: true | |
| run: python3 analysis/tests/test_analyze.py -v | |
| - name: Agent-intrusion synthetic replay corpus (#154 phase 1) | |
| home: true | |
| run: | | |
| # #154's own acceptance criteria: "CI verifies schemas, safe | |
| # fixtures, and replay expectations." Runs both the standalone | |
| # validator (the same one a human runs by hand when editing the | |
| # corpus) and its test suite (which additionally proves the | |
| # validator catches deliberately-broken input, not just that | |
| # today's corpus happens to pass). | |
| python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/validate_corpus.py | |
| python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_corpus.py -v | |
| - name: Agent-intrusion decode/correlate pipeline (#154 phase 2) | |
| home: true | |
| run: | | |
| # Bounded, non-executing base64/gzip/zlib/xor decoder plus | |
| # multi-part message-chunk reassembly, proven against the real | |
| # corpus above -- not just hand-built fixtures -- per that | |
| # corpus's own README ("expected_findings is ground truth for | |
| # phase 2's decoder"). | |
| python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_decode_correlate.py -v | |
| - name: Agent-intrusion campaign correlator (#154 phase 2) | |
| home: true | |
| run: | | |
| # Second half of phase 2 ("correlate events across sensors into | |
| # one campaign timeline using stable IDs and time windows") -- | |
| # union-find over session/IP/channel identifiers, proven against | |
| # the real corpus's own multi-hop case (a C2 channel ID bridging | |
| # two otherwise-unconnected actor identities into one campaign). | |
| python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_campaign_correlator.py -v | |
| - name: Agent-intrusion criticality rules (#154 phase 3) | |
| home: true | |
| run: | | |
| # Deterministic, structural rules (never reads the corpus's own | |
| # phase/should_escalate labels -- see criticality_rules.py's own | |
| # module docstring) proven against every one of the 27 real | |
| # corpus events matching its own independently-established | |
| # ground truth, plus the campaign-level severity scoring proven | |
| # against the real merged 8-event campaign reaching "critical". | |
| python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_criticality_rules.py -v | |
| - name: Agent-intrusion campaign worker (#154 phase 5) | |
| home: true | |
| run: | | |
| # Wires decode_correlate/campaign_correlator/criticality_rules | |
| # against Elasticsearch-shaped data for real (a hand-rolled fake | |
| # ES client, no network) -- including an end-to-end run of the | |
| # real corpus through the worker's own correlate-then-score call | |
| # sequence, proving the pipeline wiring itself is correct, not | |
| # just each module independently. | |
| python3 -m pip install -r arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/requirements.txt | |
| python3 arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/tests/test_worker.py -v | |
| - name: Sandbox VNC bridge read-only enforcement (#805) | |
| home: true | |
| run: python3 sandbox/windows/vnc-bridge/tests/test_server.py | |
| - name: Vendored YARA corpus is intact and loadable | |
| home: true | |
| run: | | |
| # yara(1) refuses to start on a corpus with one bad rule rather than | |
| # skipping it, so a rule file edited in place or dropped from | |
| # index.yar takes the whole scanner down silently. | |
| # | |
| # Run in the scanner's own base image rather than against the runner's | |
| # apt yara: "does it compile" is only a useful answer from the | |
| # compiler that will actually load it, and the two versions differ. | |
| # The image is read from the Dockerfile so a base bump moves CI too. | |
| image="$(awk '/^FROM /{print $2; exit}' arcane/home/honeypot-payload-analysis/analysis/yara/Dockerfile)" | |
| echo "validating against $image" | |
| docker run --rm -e PYTHONDONTWRITEBYTECODE=1 -v "$PWD:/w" -w /w "$image" sh -c ' | |
| apk add --no-cache bash git yara >/dev/null | |
| yara --version | |
| scripts/check-yara-corpus.sh | |
| yara -w arcane/home/honeypot-payload-analysis/analysis/yara/rules/index.yar /dev/null | |
| arcane/home/honeypot-payload-analysis/analysis/yara/tests/test_sync_yara.sh | |
| ' | |
| - name: Rev-eng benchmark corpus is reproducible and safe (#159) | |
| home: true | |
| run: | | |
| # Runs in the corpus's own documented build environment | |
| # (debian:trixie-slim, matching corpus/README.md's "Rebuilding" | |
| # section exactly), not the runner's own toolchain — the whole | |
| # point is verifying that THIS environment reproduces the | |
| # committed manifest.json byte-for-byte, which only means | |
| # something if it is the same environment the corpus claims | |
| # provenance against. | |
| # -e PYTHONDONTWRITEBYTECODE=1: the container runs as root against | |
| # this bind-mounted workspace, and root-owned __pycache__ left | |
| # behind here breaks actions/checkout's clean step for every later | |
| # job on this runner (EACCES on unlink). ci_verify.sh exports it | |
| # too -- belt and braces, since it is equally true of any other | |
| # caller that runs it as root in a mount. | |
| docker run --rm -e PYTHONDONTWRITEBYTECODE=1 -v "$PWD:/w" -w /w debian:trixie-slim bash -c ' | |
| analysis/ghidra/benchmarks/corpus/ci_verify.sh | |
| ' | |
| - name: Validate home and VPS Compose | |
| home: true | |
| run: | | |
| cp .env.example .env | |
| cp vps/.env.example vps/.env | |
| docker compose -f docker-compose.yml config --quiet | |
| # honeypot-init is a separate Dockge stack (#111) that attaches to | |
| # APIARY's honeynet network and two of its named volumes by | |
| # external reference; config validation does not require those to | |
| # actually exist, so no APIARY setup is needed here either. | |
| docker compose -f arcane/home/honeypot-init/compose.yml config --quiet | |
| # The modernization-port ("next" profile) services validate | |
| # along with the rest — config resolves every service | |
| # regardless of profile gating, profiles only affect `up`. | |
| # Split apart since #1622: honeypot-dashboard-backend | |
| # resolves backend-service:8081 against the sibling stack by | |
| # bare service-name DNS at runtime (shared honeynet | |
| # network), which config validation doesn't need to prove. | |
| docker compose -f arcane/home/honeypot-dashboard/compose.yml config --quiet | |
| docker compose -f arcane/home/honeypot-dashboard-backend/compose.yml config --quiet | |
| docker compose -f vps/docker-compose.yml config --quiet | |
| # The sandbox gateway is never started here — it answers live malware | |
| # on an isolated bridge that does not exist on a CI runner. Validating | |
| # it still matters: run_sample.py shells out to this file mid-detonation, | |
| # and a syntax error would surface as a failed run on the analysis host. | |
| docker compose -f docker-compose.sandbox.yml config --quiet | |
| # Same reasoning for the analysis host: not started here, but | |
| # install-analysis-host.sh runs it unattended on a box with a GPU and | |
| # captured malware on it. Both the base file and the GPU overlay, | |
| # because the overlay is only ever used on top of the base one. | |
| docker compose -f analysis/ghidra/docker-compose.ghidra.yml config --quiet | |
| docker compose -f analysis/ghidra/docker-compose.ghidra.yml \ | |
| -f analysis/ghidra/docker-compose.ghidra.gpu.yml config --quiet | |
| # #66's base file is synthetic and network-isolated. Validate both | |
| # #83 overlays, but start neither worker nor model in CI. | |
| docker compose -f llm-worker/docker-compose.yml config --quiet | |
| docker compose -f llm-worker/docker-compose.yml \ | |
| -f llm-worker/docker-compose.synthetic-canary.yml config --quiet | |
| docker compose -f llm-worker/docker-compose.yml \ | |
| -f llm-worker/docker-compose.production-session-canary.yml config --quiet | |
| docker compose -f llm-worker/docker-compose.yml \ | |
| -f llm-worker/docker-compose.captured-data.yml config --quiet | |
| # #2225: docker-compose.captured-data-deploy.yml (#1751) is the | |
| # entrypoint live redeploys authorized under #83 actually run -- | |
| # `include:`, not a stacked `-f base -f overlay` pair, so | |
| # neither line above proves anything about it. Unvalidated, a | |
| # bad include path or an invalid key sails through both and | |
| # first surfaces as a failed (or silently mis-scoped) live | |
| # redeploy, which is how #1751's incident happened. | |
| docker compose -f llm-worker/docker-compose.captured-data-deploy.yml config --quiet | |
| # Syntactic validity alone isn't enough: stripping an include | |
| # (e.g. losing docker-compose.captured-data.yml) still resolves | |
| # to valid, synthetic-only config. So does keeping both includes | |
| # but listing them as two separate `include:` entries instead of | |
| # one list-valued `path:` -- separate entries are independent | |
| # models, the first to claim services.llm-worker keeps it, and | |
| # the base's deliberate `ES_HOST: ''` wins. That is what this | |
| # check caught on its first run. Resolve exactly the one file | |
| # the live deploy points at -- adding `-f base -f overlay` | |
| # alongside it would supply the authorization from the command | |
| # line and prove nothing about the include chain. Assert the | |
| # resolved model still carries what this file exists to record. | |
| resolved="$(docker compose -f llm-worker/docker-compose.captured-data-deploy.yml config --format json)" | |
| es_host="$(jq -r '.services["llm-worker"].environment.ES_HOST // empty' <<<"$resolved")" | |
| if [ -z "$es_host" ]; then | |
| echo "::error::captured-data-deploy.yml resolved with empty ES_HOST -- the captured-data authorization overlay is not being applied (#2225)" | |
| exit 1 | |
| fi | |
| for net in honeypot-llm-data honeypot-llm; do | |
| if ! jq -e --arg n "$net" '.networks | to_entries[] | select(.value.name == $n)' <<<"$resolved" >/dev/null; then | |
| echo "::error::captured-data-deploy.yml resolved without the $net network attached -- the captured-data authorization overlay is not being applied (#2225)" | |
| exit 1 | |
| fi | |
| done | |
| for target in /payloads/cowrie /payloads/scripts; do | |
| if ! jq -e --arg t "$target" '.services["llm-worker"].volumes[]? | select(.target == $t and .read_only == true)' <<<"$resolved" >/dev/null; then | |
| echo "::error::captured-data-deploy.yml resolved without a read-only mount at $target -- the captured-data authorization overlay is not being applied (#2225)" | |
| exit 1 | |
| fi | |
| done | |
| # #2981: these used to be named individually, the same shape #2878 | |
| # fixed for sandbox/windows/orchestrate and sandbox/ghosts/orchestrate | |
| # -- test_install_ci_runner_instances.py sat unwired and red for days | |
| # because nothing ran it, and two more files in this directory | |
| # (test_arcane_retry_failed_sync.py, test_audit_sensor_event_coverage.py) | |
| # were never run by CI at all. Discovery instead, so the next | |
| # test_*.py dropped in scripts/tests/ is picked up without a | |
| # workflow edit. | |
| - name: scripts/tests suite | |
| home: true | |
| run: python3 -m unittest discover -s scripts/tests -p 'test_*.py' -v | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 | |
| with: | |
| python-version: "3.13" | |
| # Needed below by the two dashboard-binary Keycloak integration tests | |
| # (`go run .` against the real dashboard) -- same version floor as the | |
| # go-test job, see that job's own go-version comment. Every matrix | |
| # entry pays this setup cost even though only two of them use it -- | |
| # simpler and more robust than threading a per-entry needs-go flag | |
| # through the matrix, and setup-go is fast. | |
| - uses: actions/setup-go@b7ad1dad31e06c5925ef5d2fc7ad053ef454303e # v7.0.0 | |
| with: | |
| go-version: "1.26.x" | |
| cache-dependency-path: "**/go.sum" | |
| # Cache only on the GitHub-hosted path -- mirror of this job's | |
| # own runs-on expression. On the homeserver runners the module | |
| # cache is already on disk and setup-go's restore cannot unpack | |
| # over it; see the go-fmt job for the full account. | |
| cache: ${{ !(matrix.home == true && needs.ci-target.outputs.homeserver == 'true') }} | |
| - name: ${{ matrix.name }} | |
| shell: bash | |
| run: ${{ matrix.run }} | |
| # #3319: `always()` because the pytest rows emit their JUnit XML on the | |
| # passing run too, and that is the record of what ran. `ignore` on the | |
| # missing-file case because this job is a ~65-row matrix and most rows | |
| # write no report at all -- a warning per row would bury the one that | |
| # matters. The matrix row name is in the artifact name so the two | |
| # pytest rows don't collide in the run's artifact list (upload-artifact | |
| # v4 refuses two uploads sharing a name). | |
| # | |
| # Two things this step had wrong, both fixed here rather than papered | |
| # over with `overwrite: true`: | |
| # | |
| # 1. `path: .ci-artifacts/` uploaded NOTHING. upload-artifact v4.4+ | |
| # skips hidden paths by default, and a directory whose own name | |
| # starts with `.` is hidden, so the action reported "No files were | |
| # found with the provided path: .ci-artifacts/" and, under | |
| # if-no-files-found: ignore, said it quietly. Verified on a green | |
| # run of the ml-worker row: 330 passed, pytest wrote | |
| # .ci-artifacts/ml-worker-junit.xml, and the upload step below it | |
| # still found no files. The two JUnit files this retention was | |
| # added for have therefore never been retained. The path below | |
| # names them explicitly, which both fixes that and makes the | |
| # contents a closed set: those two files, written by the two rows' | |
| # own --junitxml flags, and nothing else that happens to be sitting | |
| # in a workspace-relative directory. | |
| # 2. The name is per-(run, attempt), so no upload here can collide | |
| # and no earlier run on the same ref can poison this one's name. | |
| # See the frontend-next twin's upload for the full argument. | |
| # | |
| # One trap left standing on purpose: GitHub rejects `/` in an artifact | |
| # name, and six row names contain one ("Healthcheck/autoheal ...", | |
| # "Zeek lab/prod ...", "Payload/artifact dedupe ...", "scripts/tests | |
| # suite", and two more). They are harmless today because they write no | |
| # report, and upload-artifact only reaches the name when it has files | |
| # to upload. A row that both writes a report AND has a `/` in its name | |
| # will fail its upload -- and there is no expression-level way to | |
| # sanitise a free-text field, so fixing that means renaming the row. | |
| - name: Upload test results | |
| if: always() | |
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | |
| with: | |
| name: scripts-${{ matrix.name }}-results-${{ github.run_id }}-${{ github.run_attempt }} | |
| path: | | |
| .ci-artifacts/ml-worker-junit.xml | |
| .ci-artifacts/auth-events-worker-junit.xml | |
| if-no-files-found: ignore | |
| retention-days: 7 | |
| scripts-and-compose-complete: | |
| name: Scripts and Compose | |
| if: always() | |
| needs: [scripts-and-compose] | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| sparse-checkout: scripts/ci-lane-summary.py | |
| sparse-checkout-cone-mode: false | |
| - name: Fail if any check failed | |
| if: needs.scripts-and-compose.result != 'success' | |
| run: exit 1 | |
| # #3319: the matrix is a single job, so its own result is all there is | |
| # to report -- a row that was path-filtered away is invisible here and | |
| # says nothing about the run. Rendered anyway so the aggregate job's | |
| # summary is not empty on a red run. | |
| - name: Lane summary | |
| if: always() | |
| env: | |
| NEEDS: ${{ toJSON(needs) }} | |
| run: python3 scripts/ci-lane-summary.py --title "Scripts and Compose" <<<"$NEEDS" | |
| # #3329: no AI/assistant attribution in PR commits, title or body. A | |
| # policy that held only by convention -- main collected 461 attributed | |
| # commits between 2026-08-01 and 2026-09-26. Reads the PR from the event | |
| # file and its commits from the API, so no PR text passes through a shell | |
| # and no deep fetch is needed. pull_request only; quality-gate accepts its | |
| # skip on every other event. | |
| ai-attribution: | |
| name: No AI attribution in PR metadata (#3329) | |
| if: github.event_name == 'pull_request' | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| pull-requests: read | |
| steps: | |
| - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| sparse-checkout: scripts/check-ai-attribution.py | |
| sparse-checkout-cone-mode: false | |
| - name: Check commits, title and body | |
| env: | |
| GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} | |
| run: python3 scripts/check-ai-attribution.py | |
| # #3311: the one required status check for this workflow. main's ruleset | |
| # requires a context that reports on every PR; the jobs above come in | |
| # homeserver/GitHub-hosted pairs whose names change with the executor, so | |
| # none of them individually is a stable context. Mirrors | |
| # go-modules-complete's rule for every pair: the homeserver twin succeeded, | |
| # or it was skipped and its GitHub-hosted twin succeeded. A new job pair | |
| # must be added to both `needs:` and PAIRS below, or it is not gated. | |
| quality-gate: | |
| name: Quality gate | |
| if: always() | |
| needs: | |
| - ci-target | |
| - public-safety | |
| - public-safety-cloud | |
| - design-lab-readonly | |
| - design-lab-readonly-cloud | |
| - go-modules-complete | |
| - frontend-next | |
| - frontend-next-cloud | |
| - frontend-next-browser | |
| - frontend-next-browser-cloud | |
| - backend-service | |
| - backend-service-cloud | |
| - vendored-theme | |
| - vendored-theme-cloud | |
| - scripts-and-compose-complete | |
| - ai-attribution | |
| runs-on: ubuntu-latest | |
| permissions: | |
| contents: read | |
| steps: | |
| - name: Fail unless every gated job, or its fallback twin, succeeded | |
| env: | |
| NEEDS: ${{ toJSON(needs) }} | |
| EVENT: ${{ github.event_name }} | |
| shell: bash | |
| run: | | |
| result() { jq -r --arg j "$1" '.[$j].result' <<<"$NEEDS"; } | |
| fail=0 | |
| single() { | |
| local r; r="$(result "$1")" | |
| [[ "$r" == "success" ]] && return 0 | |
| echo "::error::$1: expected success, got $r"; fail=1 | |
| } | |
| pair() { | |
| local hs cloud; hs="$(result "$1")"; cloud="$(result "$1-cloud")" | |
| [[ "$hs" == "success" ]] && return 0 | |
| [[ "$hs" == "skipped" && "$cloud" == "success" ]] && return 0 | |
| echo "::error::$1: expected success or skip+fallback-success; got $hs / $cloud"; fail=1 | |
| } | |
| single ci-target | |
| # #3329: pull_request-only by design, so a skip is fine elsewhere. | |
| if [[ "$EVENT" == "pull_request" ]]; then | |
| single ai-attribution | |
| elif [[ "$(result ai-attribution)" != "skipped" ]]; then | |
| single ai-attribution | |
| fi | |
| single go-modules-complete | |
| single scripts-and-compose-complete | |
| for p in public-safety design-lab-readonly frontend-next frontend-next-browser backend-service vendored-theme; do | |
| pair "$p" | |
| done | |
| exit "$fail" | |
| # #3319: the run's lane table, which is the part a human actually reads. | |
| # `if: always()` so it renders on the failing run that needs it, and it | |
| # runs even when the gate above exits 1 -- reporting is not gating. | |
| # | |
| # ci-target is excluded from --allow-skip deliberately: on a run where | |
| # the router itself never reported, every routed pair below is a | |
| # no-executor skip, and that is a red run, not an accounted one. | |
| - name: Lane summary | |
| if: always() | |
| uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 | |
| with: | |
| persist-credentials: false | |
| sparse-checkout: scripts/ci-lane-summary.py | |
| sparse-checkout-cone-mode: false | |
| - name: Render lane summary | |
| if: always() | |
| env: | |
| NEEDS: ${{ toJSON(needs) }} | |
| EVENT: ${{ github.event_name }} | |
| run: | | |
| # ai-attribution is pull_request-only by design (#3329), so its skip | |
| # is accounted for on every other event. The gate above already | |
| # encodes that same rule; this states the reason in the report. | |
| allow=() | |
| if [[ "$EVENT" != "pull_request" ]]; then | |
| allow+=(--allow-skip "ai-attribution:pull_request only (#3329)") | |
| fi | |
| python3 scripts/ci-lane-summary.py --title "Quality" "${allow[@]}" <<<"$NEEDS" |