ci(audit): make a failing scheduled security audit reach someone - #786
Merged
Merged
Conversation
The three audit jobs have been doing their work. Nothing was listening. The scheduled run on `main` failed on 2026-08-31, 2026-09-07 and 2026-09-14 — three consecutive weeks, every one a genuine red — and no issue was opened for any of them. `main` sat with a failing advisory gate and all 17 open PRs blocked behind it until someone went looking by hand three weeks later (#785). The detection was never the problem; the alerting did not exist. A scheduled failure is the one case with no audience. A PR failure is on the PR. A push failure is in front of whoever pushed. A cron failure on `main` is a red square on a page nobody opens. So the new `report` job runs only for `schedule`, and turns that square into an issue. It also closes the issue when a later scheduled run comes back green, so an open `security-audit-failure` issue means "main's advisory gate is red right now", not "was red once" — a stale alert nobody trusts is how this gets ignored a second time. The issue body says the part that cost the most time to work out: this class of failure is usually not caused by a code change. cargo-deny and cargo-audit read an advisory DB that updates daily, so a lockfile that was green when it merged goes red on its own, every open PR is blocked because these are required checks, and reading an older run's log UNDERSTATES the problem — that log lists only the advisories that existed when it ran. The 2026-08-30 run named one yanked crate; by 2026-09-17 there were five. Deliberately not a required check, and it never fails the run: every step is best-effort and a failure to file is logged rather than masking the advisory result that prompted it. This reports on the audit; it does not add a second way for the audit to break. It is skipped on `merge_group` (schedule-only), so it cannot hang the queue. Runs on a GitHub-hosted runner — two `gh` calls with no Rust in them have no business taking a slot in the self-hosted CI pool. Verified: all four paths exercised against a mocked `gh` (fail/no issue -> create; fail/existing -> comment; green/existing -> comment + close; green/none -> no-op), exit 0 in every case; multi-job failure lists each failing job. `python3 scripts/check_workflow_policy.py` -> PASS: 22 workflows; its 15 unit tests OK; typos clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The finding
The three audit jobs have been doing their work. Nothing was listening.
The scheduled run on
mainfailed three consecutive weeks — and no issue was opened for any of them:failurefailurefailuresuccessmainsat with a failing advisory gate and all 17 open PRs blocked behind it until someone went looking by hand three weeks later (#785). The detection was never the problem. The alerting did not exist.Why a scheduled failure specifically
It is the one case with no audience:
mainis a red square on a page nobody opens.So the new
reportjob runs only forschedule, and turns that square into an issue.It also closes the issue when a later scheduled run comes back green, so an open
security-audit-failureissue means "main's advisory gate is red right now" rather than "was red once". A stale alert nobody trusts is how this gets ignored a second time.What the issue says
The body carries the part that cost the most time to work out this week — that this class of failure is usually not caused by a code change.
cargo denyandcargo auditread an advisory DB that updates daily, so a lockfile that was green when it merged goes red on its own. Two consequences worth stating up front: every open PR is blocked because these are required checks, and reading an older run's log understates the problem — that log lists only the advisories that existed when it ran. The 2026-08-30 run named one yanked crate; by 2026-09-17 there were five.It points at
cargo deny check/cargo auditagainst today's DB, and recommends a lockfile-only--precisebump over a manifest change, with majors kept as their own reviewed PR.Safety
merge_group(schedule-only), so it cannot hang the merge queue.permissions: issues: writescoped to this job alone; the workflow default stayscontents: read.ghcalls with no Rust in them shouldn't take a slot in the self-hosted pool.Verification
All four paths exercised against a mocked
gh, exit 0 in every case:Multi-job failure lists each failing job with its result. Also:
python3 scripts/check_workflow_policy.py→PASS: 22 workflows, its 15 unit tests OK,typosclean.🤖 Generated with Claude Code