Skip to content

fix(ci): a retry budget sized on the outage that was measured, not on a guess - #794

Merged
stephrobert merged 1 commit into
mainfrom
fix/a-retry-budget-that-crosses-an-outage
Sep 25, 2026
Merged

stephrobert merged 1 commit into
mainfrom
fix/a-retry-budget-that-crosses-an-outage

Conversation

@stephrobert

Copy link
Copy Markdown
Owner

The nightly conformance run of 2026-09-25 died installing the Scaleway CLI:

curl: (22) The requested URL returned error: 504

from a step that already said --retry 3 --retry-connrefused. Three jobs died
the same way on 2026-09-14.

curl does retry on a 504. The budget is seven seconds.

The reflex reading is that curl ignores HTTP errors. Measured against a local
server answering 504 forever, --retry 3 sends four requests — 1s, 2s, 4s —
and gives up after seven seconds
. --retry-all-errors changes nothing: 504 is
already in curl's own list of transient errors.

Retrying was never the gap. How long it retried was.

How long an outage actually lasts here

One real measurement, taken 2026-09-14 while this family of failure was killing
three jobs at once: terraform-provider-outscale v1.8.0 answered 200 to five
requests out of fifteen for roughly twenty minutes
, then fifteen out of
fifteen. A 67% failure rate, not a total one.

Against that rate, taking attempts as independent — which they are not when a
nearby cache is what broke, so this is a floor rather than a promise:

attempts chance of failing anyway
4 (what every step had) 19.8%
8 3.9%
10 1.7%

19.8% is exactly what was observed: passing most nights, failing some.

A fixed interval, not a backoff

Exponential backoff spares a service struggling under load. What breaks here is
a CDN answering 504 to everyone, and waiting longer does not help it. At an
equal time budget a fixed interval buys far more attempts: ten cost 135s
spaced 15s apart, against 511s doubling.

What is not retried

A 404 is a pin naming an asset that does not exist. Retrying it for two minutes
turns a clear error into a slow one, and the log then says the same thing for two
different problems.

What falsify found in this work

Every test drove the script with the budget lowered through the environment so
it would run in seconds — and none of them touched the default. Setting
attempts=4, the exact value that failed, left all of them green. The defaults
now have a test of their own, asserting the calculation rather than the file.

Proven

  • Four tests driving the real script: a server that fails then recovers, one
    that never recovers, one that answers 404, and one that must leave no file
    behind.
  • A guard refusing a workflow that downloads with curl instead of the helper,
    so the next tool added does not inherit the seven seconds. 33 call sites moved.
  • Four mutations, each compiling and each biting.
  • The five falsify specs mutating the seven workflows this touches: green, and
    falsify:lint finds all 1309 fragments across 209 specs — no guard was
    disarmed by moving the call sites.
  • mise run prepush green.

🤖 Generated with Claude Code

… a guess

The nightly conformance run of 2026-09-25 died installing the Scaleway CLI:

    curl: (22) The requested URL returned error: 504

from a step that already said `--retry 3 --retry-connrefused`. Three jobs had
died the same way on 2026-09-14, and the Outscale half of the functional leg
with them.

CURL DOES RETRY ON A 504. THE BUDGET IS SEVEN SECONDS.

The reflex reading is that curl ignores HTTP errors. Measured against a local
server answering 504 forever, `--retry 3` sends four requests, one second apart
then two then four — and gives up after seven seconds. `--retry-all-errors`
changes nothing: 504 is already in curl's own list of transient errors.

So retrying was never the gap. How long it retried was.

HOW LONG AN OUTAGE ACTUALLY LASTS HERE

One real measurement, taken 2026-09-14 while this family of failure was killing
three jobs at once: the asset `terraform-provider-outscale v1.8.0` answered 200
to five requests out of fifteen for roughly twenty minutes, then fifteen out of
fifteen. A 67% failure rate, not a total one.

Against that rate, taking attempts as independent — which they are not when a
nearby cache is what broke, so this is a floor rather than a promise:

     4 attempts   19.8% chance of failing anyway   <- what every step had
     8 attempts    3.9%
    10 attempts    1.7%

19.8% is exactly what was observed: passing most nights, failing some.

A FIXED INTERVAL, NOT A BACKOFF

Exponential backoff spares a service struggling under load. What breaks here is
a CDN answering 504 to everyone, and waiting longer does not help it. At an
equal time budget a fixed interval buys far more attempts: ten of them cost 135s
spaced 15s apart, against 511s doubling.

WHAT IS NOT RETRIED

A 404 is a pin naming an asset that does not exist. Retrying it for two minutes
turns a clear error into a slow one, and the log then says the same thing for
two different problems. curl's own notion of transient is the right one and
excludes 404.

WHAT FALSIFY FOUND IN THIS WORK

Every test drove the script with the budget lowered through the environment so
it would run in seconds — and none of them touched the default. Setting
`attempts=4`, the exact value that failed, left all of them green. The defaults
now have a test of their own, and it asserts the calculation rather than the
file: eight attempts is where the measured outage stops being likely to win.

PROVEN

- Four tests driving the real script against a server that fails then recovers,
  one that never recovers, and one that answers 404.
- A guard refusing a workflow that downloads with curl instead of the helper,
  so the next tool added does not inherit the seven seconds.
- Four mutations, each compiling and each biting.
- The five falsify specs that mutate the seven workflows this touches: green,
  and `falsify:lint` finds all 1309 fragments across 209 specs.
- `mise run prepush` green.

Assisted-by: Claude Code (claude-opus-5)
@stephrobert
stephrobert merged commit cd8c1f4 into main Sep 25, 2026
31 checks passed
@stephrobert
stephrobert deleted the fix/a-retry-budget-that-crosses-an-outage branch September 25, 2026 09:28
stephrobert added a commit that referenced this pull request Sep 25, 2026
…e tag (#794) (#797)

`runtime-proof.yml`'s stacks job runs twice: once on main, once with the working
tree moved to the last release tag, as a positive control — the thing that tells
"our night is red" from "the world is red".

#794 moved every workflow download onto `tools/ci/fetch.sh`. The control checks
out v0.13.0 before it installs anything, so the steps that install Incus and
Terraform looked for that helper inside a tag cut before it existed:

  tools/ci/fetch.sh: No such file or directory
  Process completed with exit code 127

Thirty-five seconds in, before a single stack was applied, on 36150807823. No
pull request could have caught it: the stacks job only runs on the nightly
schedule and on dispatch, and the control leg carries `continue-on-error`, so it
does not even redden the job it belongs to. It was found by reading a run.

THE FIX IS THE ORDER, and it is also the better control. Everything that
provisions the RUNNER now happens before the move — Go, Incus, OVN, Terraform —
and everything that is the PRODUCT happens after it: the binary, the images, the
suites. A witness should differ from its subject in one thing, the code under
test, not in how its runner was built. Running the tag's own installation steps
mixed the runtime under test with the harness that fetched it, and this is what
made that mixing visible.

`Install Terraform` moves up with the other tools for the same reason; nothing
else changes order, and the step that carries the gate is untouched.

What holds it: TestTheControlNeverRunsARepositoryScriptInsideTheTag asserts that
no `tools/ci/` script runs after the detach, and — the accepting half — that one
does run before it, so a job that simply lost its control cannot pass by being
empty. tools/falsify/specs/the-control-runs-on-its-own-tools.json carries two
mutations: a CI script placed after the detach, and the detach moved back before
the tools. Both compile and both redden the test.

`mise run prepush` green.

Assisted-by: Claude Code (claude-opus-5)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant