Skip to content

ci(audit): make a failing scheduled security audit reach someone - #786

Merged
CueCrux-Myles merged 1 commit into
mainfrom
fix/alert-on-scheduled-audit-failure
Sep 17, 2026
Merged

CueCrux-Myles merged 1 commit into
mainfrom
fix/alert-on-scheduled-audit-failure

Conversation

@CueCrux-Myles

Copy link
Copy Markdown
Contributor

The finding

The three audit jobs have been doing their work. Nothing was listening.

The scheduled run on main failed three consecutive weeks — and no issue was opened for any of them:

Date Result
2026-08-31 failure
2026-09-07 failure
2026-09-14 failure
2026-08-24 and earlier success

main sat with a failing advisory gate and all 17 open PRs blocked behind it until someone went looking by hand three weeks later (#785). The detection was never the problem. The alerting did not exist.

Why a scheduled failure specifically

It is the one case with no audience:

  • A PR failure is on the PR.
  • A push failure is in front of whoever pushed.
  • A cron failure on main is a red square on a page nobody opens.

So the new report job runs only for schedule, and turns that square into an issue.

It also closes the issue when a later scheduled run comes back green, so an open security-audit-failure issue means "main's advisory gate is red right now" rather than "was red once". A stale alert nobody trusts is how this gets ignored a second time.

What the issue says

The body carries the part that cost the most time to work out this week — that this class of failure is usually not caused by a code change. cargo deny and cargo audit read an advisory DB that updates daily, so a lockfile that was green when it merged goes red on its own. Two consequences worth stating up front: every open PR is blocked because these are required checks, and reading an older run's log understates the problem — that log lists only the advisories that existed when it ran. The 2026-08-30 run named one yanked crate; by 2026-09-17 there were five.

It points at cargo deny check / cargo audit against today's DB, and recommends a lockfile-only --precise bump over a manifest change, with majors kept as their own reviewed PR.

Safety

  • Not a required check, and it never fails the run. Every step is best-effort; a failure to file is logged rather than masking the advisory result that prompted it. This reports on the audit — it does not add a second way for the audit to break.
  • Skipped on merge_group (schedule-only), so it cannot hang the merge queue.
  • permissions: issues: write scoped to this job alone; the workflow default stays contents: read.
  • Runs on a GitHub-hosted runner — two gh calls with no Rust in them shouldn't take a slot in the self-hosted pool.

Verification

All four paths exercised against a mocked gh, exit 0 in every case:

Scenario Behaviour
fail, no open issue creates
fail, issue open comments
green, issue open comments + closes
green, none open no-op

Multi-job failure lists each failing job with its result. Also: python3 scripts/check_workflow_policy.py → PASS: 22 workflows, its 15 unit tests OK, typos clean.

🤖 Generated with Claude Code

The three audit jobs have been doing their work. Nothing was listening.

The scheduled run on `main` failed on 2026-08-31, 2026-09-07 and
2026-09-14 — three consecutive weeks, every one a genuine red — and no
issue was opened for any of them. `main` sat with a failing advisory gate
and all 17 open PRs blocked behind it until someone went looking by hand
three weeks later (#785). The detection was never the problem; the
alerting did not exist.

A scheduled failure is the one case with no audience. A PR failure is on
the PR. A push failure is in front of whoever pushed. A cron failure on
`main` is a red square on a page nobody opens. So the new `report` job
runs only for `schedule`, and turns that square into an issue.

It also closes the issue when a later scheduled run comes back green, so
an open `security-audit-failure` issue means "main's advisory gate is red
right now", not "was red once" — a stale alert nobody trusts is how this
gets ignored a second time.

The issue body says the part that cost the most time to work out: this
class of failure is usually not caused by a code change. cargo-deny and
cargo-audit read an advisory DB that updates daily, so a lockfile that was
green when it merged goes red on its own, every open PR is blocked because
these are required checks, and reading an older run's log UNDERSTATES the
problem — that log lists only the advisories that existed when it ran. The
2026-08-30 run named one yanked crate; by 2026-09-17 there were five.

Deliberately not a required check, and it never fails the run: every step
is best-effort and a failure to file is logged rather than masking the
advisory result that prompted it. This reports on the audit; it does not
add a second way for the audit to break. It is skipped on `merge_group`
(schedule-only), so it cannot hang the queue.

Runs on a GitHub-hosted runner — two `gh` calls with no Rust in them have
no business taking a slot in the self-hosted CI pool.

Verified: all four paths exercised against a mocked `gh` (fail/no issue ->
create; fail/existing -> comment; green/existing -> comment + close;
green/none -> no-op), exit 0 in every case; multi-job failure lists each
failing job. `python3 scripts/check_workflow_policy.py` -> PASS: 22
workflows; its 15 unit tests OK; typos clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@CueCrux-Myles
CueCrux-Myles added this pull request to the merge queue Sep 17, 2026
Merged via the queue into main with commit 1175216 Sep 17, 2026
23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant