20260904 - Retry deployments that failed before the artifact arrived - #9
Merged
Merged
Conversation
nightcrawler2 lost owl-os-pi5-v0.16.1 on 2026-09-03 to a broken download:
two truncated streams, then a resumed range request that R2 answered 400,
and the deployment was over three minutes after it started with eight of
its ten client retries unused. Recovering it meant creating a deployment
by hand, which does not scale with the fleet.
Hosted Mender's own Retries field is a paid-plan feature. On this tenant's
`os` plan, creating a deployment with one is rejected outright with
403 "Feature not available in your Plan.", so the retry lives here instead.
Only one failure is treated as fixable: the artifact never finished
arriving. Nothing was written to the node and nothing about the node caused
it, so an identical request has an independent chance of succeeding.
Everything else is refused, including failures the log cannot explain. The
costs are asymmetric. A missed retry costs one hand deployment; a wrong one
reboots a live node on a timer for a reason that will not change.
Two properties of Mender's device log make the obvious implementation
wrong, and both were found against real fleet logs rather than reasoned
about:
* "Installing artifact..." is printed about a second after the deployment
starts, before the download, so it does not mark the install phase. A
phase test built on it would have refused every genuine download
failure. Wilderness A shows the line one second in; nightcrawler1 shows
it five hours before its download gave up.
* The log is cumulative across attempts. nightcrawler1's carries three
attempts at one deployment plus lines dated four months earlier from a
clock skew, so matching over the whole log would let an old disk-full
line veto a retry forever and an old transport error authorise one.
Only the final attempt is classified.
Blast radius is bounded deliberately: it can only re-send an artifact a
human-created deployment had already targeted at that device, two attempts
per pair, five per pass, and it aborts before creating anything if the
attempt counts cannot be persisted. Deployments it cannot count are how a
retry becomes a storm, and ProtectSystem=strict leaves the checkout
read-only under the packaged unit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The classifier is validated against every failure log currently in the fleet, but the path that creates a deployment has never run against real hardware, and the classification depends on log wording that only Mender controls. So the packaged unit reports its verdict on each failure and creates nothing. The journal is the evidence for whether to trust it. Once its calls match the ones a person would have made, --apply is a one-line change and a daemon-reload. Until then a failure it would have retried still needs a deployment by hand, which is what the README now says. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
mender-auto-accept/deploy_retry.py, a timer-driven pass that re-issues Mender deployments which failed because the artifact never finished downloading. Hosted Mender's own Retries field is a paid-plan feature (403 on ourosplan).Ships observing only. The systemd unit runs without
--apply, so it logs a verdict per failure and creates nothing. Turning it on is a separate, deliberate step.Safety
Unexpected status code while fetching artifactandGiving up on resuming the downloadare retried. Disk-full, failed install steps, rollbacks and unreadable logs are refused.Evidence
stream-nextclient error), which is the conservative direction.Tracked in ClickUp: OTA deployments miss part of the fleet on every release (123zgec4tmx), Design item 6.
🤖 Generated with Claude Code