Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -9,3 +9,4 @@ __pycache__/

# Runtime state
.pending-deploys.json
.deploy-retries.json
86 changes: 86 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,3 +60,89 @@ Edit `.env`:
2. Script fetches pending devices from Mender API
3. Filters by `node_id` prefix (if configured)
4. Accepts matching devices


## Failed-deployment retry

Re-issues Mender deployments that failed **before the artifact reached the
node**, so an update lost to a broken download recovers without anyone
deploying it by hand.

Hosted Mender has a "Retries" field on a deployment, but it is a paid-plan
feature: on the `os` plan, creating a deployment with one is rejected with
`403 Feature not available in your Plan.` The client's own retry does not cover
this either. It resumes a broken download, but once the CDN answers with an
unexpected HTTP status it treats that as fatal and fails the deployment with
most of its ten retries unused.

### What it will and will not retry

Only one failure is considered fixable: the artifact never finished arriving.
Nothing was written to the node and nothing about the node caused it, so an
identical request has an independent chance of succeeding.

Everything else is refused, **including failures it cannot read**. The costs are
asymmetric: a missed retry costs one hand deployment, while a wrong retry
reboots a live node on a timer for a reason that will not change. So the test is
an allowlist and the default is to do nothing.

| In the final attempt of the device log | Retried |
|---|---|
| `Unexpected status code while fetching artifact` | yes |
| `Giving up on resuming the download` | yes |
| `No space left on device`, too big, incompatible, bad signature | no |
| `Process returned non-zero exit status`, `ArtifactRollback` | no |
| Anything else, an unreadable log, no log at all | no |

Two properties of Mender's device log make the naive version of this wrong, and
both were measured against real fleet logs:

* **The log is cumulative across attempts.** One node's carries three attempts
at the same deployment, plus lines dated four months earlier from a clock
skew. Only the text after the final `Deployment with ID ... started` is
classified, so an old disk-full line cannot veto a retry forever, and an old
transport error cannot authorise one.
* **`Installing artifact...` does not mark the install phase.** It is printed
about a second after the deployment starts, before the download. On one node
it appears five hours before the download gives up. Phase cannot be inferred
from it.

### Other limits

* A retry is skipped while the device is already running a deployment, if the
artifact has landed since the failure, or if the device is no longer accepted.
* `DEPLOY_RETRY_MAX_ATTEMPTS` (2) per device+artifact, then it stops and leaves
it for a person.
* `DEPLOY_RETRY_MAX_PER_PASS` (5) caps one pass, so a fleet-wide outage cannot
become a fleet-wide burst.
* Attempt counts are written before the next deployment is created, and the pass
aborts up front if they cannot be written at all. Deployments that cannot be
counted are how a retry becomes a storm.

### Install

```bash
sudo cp systemd/retina-deploy-retry.* /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now retina-deploy-retry.timer
```

**The unit runs without `--apply`, so it creates nothing.** It reports a verdict
on every real failure and leaves them alone. That is deliberate: the create path
has never run against the fleet, and the classifier reads log wording that only
Mender controls, so the journal earns the trust first.

```bash
journalctl -u retina-deploy-retry
```

If its verdicts match the calls you would have made, add `--apply` to
`ExecStart` and `systemctl daemon-reload`. Until then a failure it would have
retried still needs a deployment by hand.

### Check what it would do, by hand

```bash
cd ~/retina/node-infra/mender-auto-accept
MENDER_PAT=your-token .venv/bin/python deploy_retry.py
```
26 changes: 26 additions & 0 deletions mender-auto-accept/.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -37,3 +37,29 @@ NODE_ID_PREFIX=ret

# Where tunnel_sync records what it has provisioned (optional).
# TUNNEL_STATE_FILE=


# --- Failed-deployment retry (deploy_retry.py) ---

# Hosted Mender's own "Retries" field is a paid-plan feature and this tenant is
# on `os`, so a deployment created with one is rejected 403. These drive the
# replacement. Only downloads that never completed are retried; see the README.

# How many times one device+artifact pair may be retried (optional, default 2).
# DEPLOY_RETRY_MAX_ATTEMPTS=2

# Ceiling on deployments one pass may create, so a fleet-wide outage cannot
# become a fleet-wide burst. The remainder is picked up next pass.
# DEPLOY_RETRY_MAX_PER_PASS=5

# Minimum wait between attempts on the same pair (optional, default 1800).
# DEPLOY_RETRY_BACKOFF_SECONDS=1800

# How far back to look for failures (optional, default 7 days). Also bounds how
# long an attempt count is remembered.
# DEPLOY_RETRY_WINDOW_DAYS=7

# Where attempt counts are recorded. Leave unset under systemd: the unit's
# StateDirectory= is picked up automatically. Set it only for a hand run that
# needs the counts somewhere specific.
# DEPLOY_RETRY_STATE_FILE=
Loading
Loading