Repository navigation
ci(containers): boot-smoke the two dashboard images we build but never run (#3316) - #3390
Merged
Merged
Conversation
…r run (#3316) containers.yml builds all nineteen rows and image-security-scan.yml scans them, but nothing ever started the result. A green compile gate and a clean vulnerability scan are both silent about the four things that decide whether a container comes up at all: the entrypoint, the runtime user, the file layout, and the env contract. The #2183/#2299 boot-refusal class was found by hand on a deployed host for exactly that reason. scripts/boot-smoke-image.sh starts one image and gates on the image's own HEALTHCHECK going healthy, then adds per-tier HTTP assertions on top: backend-service /healthz 200, /readyz reports ready with a green cluster dashboard-next / bounces unauthenticated traffic to /auth/login, and the vendored /static/theme.css the image's own healthcheck fetches is still served It runs the image by ID rather than by tag, because a tag can move underneath the step; on a pull_request the build loads the image locally and the row is found through a per-run label (the executor's docker daemon is shared by seven runner users, so "newest dangling image" would be a race), while on a push the build's own digest names it. backend-service needs a reachable Elasticsearch before it reports ready, so the script starts a stub bound to the docker bridge gateway -- not 0.0.0.0, this executor has a public interface -- and passes it in under ELASTICSEARCH_URL. It runs with a real throwaway SERVICE_TOKEN rather than APIARY_ALLOW_UNAUTH_DEV so the #2183 production boot path is the one exercised. Only the two long-lived HTTP service rows carry boot_smoke: true. The capture daemons keep building exactly as before: `load` and the extra label are both inert without the flag, and nothing is started for them. The smokes are steps inside the existing build job, so containers-gate already covers them with no change to the required checks. 44 tests cover the script's failure paths against stub docker/curl binaries and read the committed Dockerfiles and containers.yml directly, so a smoke that drifts from the image it smokes -- an EXPOSE, a renamed route, a rotated healthcheck path -- fails the test rather than passing against the substitute.
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3316.
containers.yml builds all nineteen rows and image-security-scan.yml scans them, but nothing ever started the result. A green compile gate and a clean vulnerability scan are both silent about the four things that decide whether a container comes up at all — the entrypoint, the runtime user, the file layout, and the env contract. The #2183/#2299 boot-refusal class was found by hand on a deployed host for exactly that reason.
What this adds
scripts/boot-smoke-image.shstarts one built image and gates on the image's own HEALTHCHECK going healthy, then adds per-tier HTTP assertions on top:backend-service/healthz200,/readyzreportsreadywith a green clusterdashboard-next/bounces unauthenticated traffic to/auth/login; the vendored/static/theme.cssthe image's own healthcheck fetches is still servedThe healthcheck is part of the artifact, so if it is wrong (wrong port, wrong path, a
curlthat is not in the image) the smoke fails on the real defect instead of passing against a substitute. The HTTP assertions check the contract — what a response says — not merely that something answered.Decisions worth reviewing
...:pr-1234would be a statement about whatever that tag resolved to at that instant. On apull_requestthe build loads the image locally and the row is found via a per-run, per-attempt label — the executor's docker daemon is shared by seven runner users, so "newest dangling image" would be a race that boots a neighbour's build. On a push the build's own digest names it (three pull attempts, for the ghcr read-after-write lag containers: generate a digest-bound SBOM for backend-service and dashboard-next builds #3321's syft step already had to handle).backend-serviceneeds a reachable ES before it reports ready. The stub answers exactly whates.rsasks — the product document withX-Elastic-Product: Elasticsearch(the 8.x client refuses without it),_cluster/healthgreen,*/_settings→{}— and anything unmodelled gets a 404 rather than hanging, so a future readiness dependency fails the assertion instead of the timeout. It binds the bridge gateway, not0.0.0.0, because this executor is a honeypot host with a public interface.SERVICE_TOKEN, notAPIARY_ALLOW_UNAUTH_DEV=1, so the Hardening: unset SERVICE_TOKEN silently disables both auth tiers — refuse to boot outside an explicit dev override #2183 production boot path is the one exercised and the log is not full of the override's warning.OIDC_DISABLEDandWEB_CONCURRENCYleft unset on purpose. The multi-worker fork loop inserver/cluster.mjsis part of the entrypoint; a smoke that ran it single-process would skip the part most likely to break.BACKEND_URL/LISTEN_ADDRstay at their image defaults so the smoke proves the image's own ENV works.buildjob, not a new job. A second buildx invocation would double the build cost, and the gate question is settled for free:containers-gatealreadyneeds: [ci-target, build], so the smokes are on the required Containers gate with no change to the required checks.boot_smoke: true.load:and the extra label are both inert without the flag, so the capture daemons build byte-identically to before and nothing is started for them.Tests
44 tests, auto-discovered by quality.yml's
scripts/tests suite. The script's failure paths run against stubdocker/curlbinaries — no image builds in CI:docker runfailure, failing status/redirect/json assertions, non-JSON bodyHEALTHCHECKdegrades to the assertions alone rather than spinning the timeout--keep-image) on every pathThe
ParityWithTheDockerfilesclass reads the committed Dockerfiles andcontainers.ymldirectly, so drift fails the test instead of quietly weakening the smoke. Verified by deliberately breaking each:EXPOSE 8081→9090,--health-timeout 120→60, a/readyzroute rename, backend HEALTHCHECK path rot, deletion of the theme.css assertion, deletion of aboot_smoke: trueflag, and acontainers-gateneeds:that dropsbuild.Checks run
actionlintclean ·shellcheck --severity=errorclean ·bash -nclean ·zizmor --min-severity mediumreports no new findings (only the pre-existingunpinned-uses/excessive-permissionsin quality.yml's ADVISORY set, and notemplate-injection) ·check-backend-boot-contract.pypasses ·orphaned-script-sweep.shpasses ·check-ai-attribution.pyclean on the commit message and PR title.Full
scripts/testssuite: 239 tests, 2 pre-existing failures intest_compose_drift_watch_sweep.pythat also fail on a cleanmaincheckout (unrelated — verified by stashing).The script was also driven end-to-end against purpose-built fixture images on real Docker: pass, boot-refusal, health timeout,
docker runfailure, the assertion failures above, missing-HEALTHCHECK degradation, stub-ES/readyz200, and the negative (ES unreachable →/readyz503, assertion correctly fails). Three real bugs turned up that way and are fixed: assertion flags shifted 3 instead of 4, stored$1|$2|$3instead of$2|$3|$4, and the no-HEALTHCHECK case spun the full timeout instead of degrading.Not in this PR
A post-deploy readiness check in
deploy.yml. The brief says this contract can back one, which is a follow-up rather than part of #3316.