Skip to content

20260924 - Boot guard, stack-health inventory, watchdog fix - #61

Merged
Purple10101 merged 2 commits into
mainfrom
20260924-boot-guard-stack-health
Sep 24, 2026
Merged

Purple10101 merged 2 commits into
mainfrom
20260924-boot-guard-stack-health

Conversation

@Purple10101

Copy link
Copy Markdown
Contributor

Stacked on #60. Phase 2 items 2b and 5a of ClickUp OTA deployments miss part of the fleet on every release (123zgec4tmx), plus a watchdog fix found during #60's testing. Draft until proven on the os-v0.17.2-dev image.

feaf2cc: boot guard and stack health

  • retina-env-guard, ExecStartPre=- in retina-node.service. Moves a NUL or unparseable manifests/.env aside at every boot, same rule as retina-node's install preflight, so the config-merger can regenerate it from user.yml. Without it a zeroed .env keeps the stack down through every reboot: compose cannot load the project, so it cannot run the config-merger that would repair it (ret9573ecda, 2026-08-27 to 2026-09-24). The - means it can never stop the stack starting.
  • mender-inventory-retina-stack, picked up by the existing inventory glob. Reports retina_stack (up, degraded, down, absent), retina_stack_running, retina_stack_restarting, retina_blah2, retina_env_ok, retina_compose_ok, retina_mode. blah2 stopped on purpose in spectrum or sdrconnect mode is not degraded. Hand-run on the test nodes: Josh Test Node up, Josh Test Node 2 degraded (no RSPduo), 116-160 ms.
  • 13 tests under dash with a stub docker; the Tests workflow now runs every test directory (28 tests) and checks all three scripts are executable.

eb4ca91: watchdog fix

blah2_rspduo_restart.bash errored with operand expected whenever blah2-api was down, because TIMESTAMP was empty. A missing or non-numeric timestamp now reads as stale. Behaviour is unchanged: the FIRST_CHAR test already triggered the restart.

Expected results on os-v0.17.2-dev

  • Josh Test Node, .env zeroed then hard reboot: guard sets it aside (journal of retina-node.service), stack and radar come up by themselves.
  • Mender inventory shows retina_stack=up for Josh Test Node and degraded for Josh Test Node 2 within 10 minutes of boot.
  • Watchdog on Josh Test Node 2 runs without the operand expected error.
  • 20260924 - Install a fork of the docker-compose Update Module #60's fork still behaves identically (regression install).

🤖 Generated with Claude Code

Purple10101 and others added 2 commits September 24, 2026 13:55
retina-env-guard runs as ExecStartPre=- in retina-node.service. A power cut
during the config-merger's write left ret9573ecda's compose .env entirely
NUL on 2026-08-27, and compose then refused to load the project at every
boot, including the run of the config-merger that would have repaired it:
four weeks without radar. The guard moves a NUL or unparseable .env aside,
by the same rule as retina-node's install preflight, and the config-merger
regenerates it from user.yml. The "-" means it can never stop the stack
starting.

mender-inventory-retina-stack reports retina_stack (up, degraded, down,
absent), running and restarting counts, blah2's state, whether .env and the
compose project load, and retina-gui's mode. Telemetry runs inside the same
stack, so a node whose stack is broken went silent instead of unhealthy;
mender-updated runs this every 600 s regardless. blah2 stopped on purpose in
spectrum or sdrconnect mode is not counted as degraded. On the test nodes it
read Josh Test Node up and Josh Test Node 2 degraded (no RSPduo, blah2
restarting), in 116-160 ms.

13 tests run both scripts under dash with a stub docker; the Tests workflow
now runs every test directory and checks all three scripts are executable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With no answer from /api/map, TIMESTAMP was empty and
DIFF_TIMESTAMP=$(($CURR_TIMESTAMP-$TIMESTAMP)) failed with "operand
expected" on every check. The restart still happened through the FIRST_CHAR
test, so this was noise, but noise in exactly the situation where the
watchdog's output gets read. A missing or non-numeric timestamp now reads
as stale.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Purple10101

Copy link
Copy Markdown
Contributor Author

Evidence on os-v0.17.2-dev, 2026-09-24 (built from eb4ca91; Tests workflow green)

# Test Node Result
1 OS install both committed; /usr/local/sbin/retina-env-guard 0755; ExecStartPre=-/usr/local/sbin/retina-env-guard in retina-node.service; inventory script 0755; watchdog fix present; fork 8abfe66d with its diversion; stack active, 8 containers; guard silent on a clean boot
2 Health in Mender both Josh Test Node retina_stack=up, 8 running, blah2 running, .env/compose ok; Josh Test Node 2 retina_stack=degraded, 1 restarting, blah2 restarting (no RSPduo)
3 Boot guard end to end Josh Test Node .env zeroed (208 NUL), sysrq-b at 13:51:24. 13:51:58 retina-env-guard: set aside .../.env (contains NUL bytes) as .env.corrupt-20260924T135158Z; 13:52:02 config-merger wrote a new .env (0 NUL, identical to the original); 13:52:03 service finished; blah2 running, /api/map 1 s old, tar1090 on the node's location. No hands on the node.
4 Watchdog fix Josh Test Node 2 blah2-api stopped (/api/map unreachable): no operand expected; same decision (crash-looping ... bypassing grace period, Successfully restarted blah2); blah2-api back in 6 s
5 Fork regression Josh Test Node dev2 to v0.4.6.0: .env seeded, stop 13:57:43.3 to start 13:57:54.1 (10.8 s), blah2 0 restarts, map 1 s old

Every expected result in the PR description met.

Base automatically changed from 20260924-fork-compose-module to main September 24, 2026 14:01
@Purple10101
Purple10101 marked this pull request as ready for review September 24, 2026 14:02
@Purple10101
Purple10101 merged commit 1e20412 into main Sep 24, 2026
2 checks passed
@Purple10101
Purple10101 deleted the 20260924-boot-guard-stack-health branch September 24, 2026 14:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant