Skip to content

20260904 - Retry deployments that failed before the artifact arrived - #9

Merged
Purple10101 merged 2 commits into
mainfrom
20260904-retry-failed-deployments
Sep 24, 2026
Merged

Purple10101 merged 2 commits into
mainfrom
20260904-retry-failed-deployments

Conversation

@Purple10101

Copy link
Copy Markdown
Contributor

Adds mender-auto-accept/deploy_retry.py, a timer-driven pass that re-issues Mender deployments which failed because the artifact never finished downloading. Hosted Mender's own Retries field is a paid-plan feature (403 on our os plan).

Ships observing only. The systemd unit runs without --apply, so it logs a verdict per failure and creates nothing. Turning it on is a separate, deliberate step.

Safety

  • Allowlist, default deny: only Unexpected status code while fetching artifact and Giving up on resuming the download are retried. Disk-full, failed install steps, rollbacks and unreadable logs are refused.
  • Only the final attempt in Mender's cumulative device log is classified.
  • 2 attempts per device and artifact, 5 per pass, 30 min backoff. Skips busy devices, devices where the artifact has since landed, and devices no longer accepted.
  • Refuses to create anything if its attempt-count state cannot be written.

Evidence

  • 34 unit tests pass.
  • Replayed against all 45 historical failure logs (2025-10 to 2026-09): 2 retry verdicts (nightcrawler1 v0.4.4.0, nightcrawler2 v0.16.1, both genuine download failures), 43 refusals, no wrong retries. It misses two real download failures (a 3 h Update Module timeout and a stream-next client error), which is the conservative direction.

Tracked in ClickUp: OTA deployments miss part of the fleet on every release (123zgec4tmx), Design item 6.

🤖 Generated with Claude Code

Purple10101 and others added 2 commits September 4, 2026 11:37
nightcrawler2 lost owl-os-pi5-v0.16.1 on 2026-09-03 to a broken download:
two truncated streams, then a resumed range request that R2 answered 400,
and the deployment was over three minutes after it started with eight of
its ten client retries unused. Recovering it meant creating a deployment
by hand, which does not scale with the fleet.

Hosted Mender's own Retries field is a paid-plan feature. On this tenant's
`os` plan, creating a deployment with one is rejected outright with
403 "Feature not available in your Plan.", so the retry lives here instead.

Only one failure is treated as fixable: the artifact never finished
arriving. Nothing was written to the node and nothing about the node caused
it, so an identical request has an independent chance of succeeding.
Everything else is refused, including failures the log cannot explain. The
costs are asymmetric. A missed retry costs one hand deployment; a wrong one
reboots a live node on a timer for a reason that will not change.

Two properties of Mender's device log make the obvious implementation
wrong, and both were found against real fleet logs rather than reasoned
about:

  * "Installing artifact..." is printed about a second after the deployment
    starts, before the download, so it does not mark the install phase. A
    phase test built on it would have refused every genuine download
    failure. Wilderness A shows the line one second in; nightcrawler1 shows
    it five hours before its download gave up.

  * The log is cumulative across attempts. nightcrawler1's carries three
    attempts at one deployment plus lines dated four months earlier from a
    clock skew, so matching over the whole log would let an old disk-full
    line veto a retry forever and an old transport error authorise one.
    Only the final attempt is classified.

Blast radius is bounded deliberately: it can only re-send an artifact a
human-created deployment had already targeted at that device, two attempts
per pair, five per pass, and it aborts before creating anything if the
attempt counts cannot be persisted. Deployments it cannot count are how a
retry becomes a storm, and ProtectSystem=strict leaves the checkout
read-only under the packaged unit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The classifier is validated against every failure log currently in the
fleet, but the path that creates a deployment has never run against real
hardware, and the classification depends on log wording that only Mender
controls. So the packaged unit reports its verdict on each failure and
creates nothing.

The journal is the evidence for whether to trust it. Once its calls match
the ones a person would have made, --apply is a one-line change and a
daemon-reload. Until then a failure it would have retried still needs a
deployment by hand, which is what the README now says.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Purple10101
Purple10101 merged commit f83f614 into main Sep 24, 2026
2 checks passed
@Purple10101
Purple10101 deleted the 20260904-retry-failed-deployments branch September 24, 2026 14:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant