Skip to content

Wait for the modem, and supervise startup instead of dying silently - #20

Merged
bazauto merged 1 commit into
mainfrom
fix/startup-supervision
Aug 27, 2026
Merged

bazauto merged 1 commit into
mainfrom
fix/startup-supervision

Conversation

@bazauto

@bazauto bazauto commented Aug 27, 2026

Copy link
Copy Markdown
Owner

The IO node did not always come up after a power cycle, and when it failed it failed
completely and silently: the onboard LED never lit, nothing reached the broker, and the
board sat at a REPL until someone power-cycled it again. The orchestrator did the right
thing — the Goods Shed sensors went stale and the block degraded to unknown, which the
fail-safe rule treats as occupied — but nothing anywhere pointed at the node.

Closes #19.

The race

The Pico and the ESP-AT modem leave a power cycle together, and the Pico is much faster.
Measured on the bench:

Event Time from power-on
Pico reaches its first AT command ~300 ms (MicroPython boot, plus 19 ms of I2C and expander setup)
Modem reaches ready 1.2 s – 3.5 s

Both ends of the modem's spread came from the same session — one reset reached ready at
3471 ms, another at 1232 ms — and the variance is Ethernet link negotiation. setup()
opened by sending AT+RST on the constructor's 2000 ms default, so the deadline for a
reply fell about 2.3 s after power-on. Whether the node started at all therefore depended
on which side of that the link happened to land, which is exactly the reported "sometimes
it doesn't come up".

A command sent into that window is not refused, it is swallowed whole: no echo, no reply.

What changed

setup() asks before it commands. wait_for_modem() polls a bare AT until the
modem answers, for up to 15 s, and only then sends the reset. The startup stages now raise
distinct exception types — ModemNotResponding, NetworkNotReady, BrokerUnreachable —
so a supervisor can tell which stage failed without matching on message text. They
subclass RuntimeError, which is what callers already caught.

Startup is supervised. main() was called bare at module scope, so any exception on
the way up simply ended main.py. node_startup.start_supervised() retries three times,
flashing for 2 s / 5 s / 10 s between attempts, then resets the board and starts over. It
never returns without a running node. A crash in the main loop gets the same treatment,
for the same reason — it is equally a node that stopped reporting without saying so.

The LED carries a code. Solid on is running, dark is now only no power or a hard
crash, and N flashes means a failure at stage N:

Code Stage Where to look
2 Sensors Ribbon to boards 1/2, config.py addresses
3 Modem UART GP8/GP9, modem power
4 Network Ethernet cable, DHCP, switch port
5 Broker Mosquitto on the bench box, MQTT_HOST
6 Runtime crash USB console — a bug, not wiring
7 Unknown USB console

Codes start at 2 because a single flash is too easily confused with a board merely
blinking on boot, and the 1400 ms gap between cycles is deliberately much longer than the
250 ms between flashes so a 2 cannot read as a 4. A cycle is always completed once
started, so a code is never cut off half way through and left uncountable.

Also closed: wait_for_line() only inspects lines recorded after it is called, so an
+ETH_GOT_IP arriving in the same read as the reset's own OK would have been missed and
setup() would have sat out its full 30 s before raising as though the network were down.
It does not bite today only because the real gap is seconds.

Decisions worth knowing

Retrying a wiring fault forever is deliberate. An expander missing at boot is usually a
loose ribbon, not a permanent fact, and a board that keeps trying and keeps flashing is
easier to find than one that is simply dark. The cost is that a genuinely broken node now
reboots in a loop instead of sitting still — but it presents identically at the
orchestrator either way, so nothing is lost by the node being loud about it.

No machine import in either new module. The pin and the reset function are injected,
which is what lets the retry policy, the code mapping and the flash pattern all be proven
on the host rather than by standing in front of the baseboard counting.

Verification

Host suite, run on this branch:

128 passed in 0.08s

The regression tests were mutation-checked rather than trusted: reverting the probe and
restoring the 2 s AT+RST timeout fails 7 of the new tests with exactly the original
error, ModemNotResponding: AT+RST failed.

/mqtt-check run against the diff — it adds no contract surface, and the 30 s re-assert
survives the restructure (layout.tick() still runs every loop pass).

On the bench, after deploying:

12:44:41  sensor/cs---goods-shed/reading {"state": "occupied"}
12:45:04  sensor/cs---goods-shed/reading {"state": "occupied"}   <- re-assert
block/1d89a49e.../state  "occupancy":"occupied"

The LED path was run on the real board — the pin reads back 1 and 0, three attempts each
flashed code 3, StartupFailed was raised and the reset hook called once.

What the host suite cannot prove is that the modem's real boot stays inside the 15 s
budget. Measured at 1.2–3.5 s, so about 4x headroom, but that is a bench measurement
rather than a test.

Docs

docs/startup-and-status-led.md is new and carries the race, the timings, the policy and
the code table. CLAUDE.md gains the doc in its authoritative table, the two new modules
in the src/lib listing, and a trap entry — the old one would otherwise have left the LED
being dark undocumented as a symptom.

)

The IO node did not always come up after a power cycle, and when it failed it
failed completely: the LED never lit, nothing reached the broker, and the board
sat at a REPL until someone power-cycled it again.

The Pico reaches its first AT command about 300 ms after power-on. The modem
needs 1.2-3.5 s to reach `ready`, the spread being Ethernet link negotiation.
`setup()` opened with `AT+RST` on a 2000 ms deadline, so whether the node
started at all depended on which side of ~2.3 s the link landed on. A command
sent into that window is swallowed whole — no echo, no reply.

`setup()` now polls a bare `AT` until the modem answers before commanding it,
and the startup stages raise distinct exception types so a failure can be told
apart without matching on message text.

Startup is also supervised. `main()` was called bare at module scope, so any
exception on the way up ended `main.py`. It now retries, flashes a code on the
onboard LED, and resets the board rather than exiting — the LED staying dark is
how this was noticed, so it is the right channel to report on. A crash in the
main loop is treated the same way.

Also closes a related blind spot: `wait_for_line()` only inspects lines recorded
after it is called, so an `+ETH_GOT_IP` arriving in the same read as the reset's
own `OK` would have been missed and `setup()` would have sat out its full 30 s
before raising as though the network were down.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@bazauto
bazauto merged commit b9a4c23 into main Aug 27, 2026
1 check passed
@bazauto
bazauto deleted the fix/startup-supervision branch August 27, 2026 11:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

IO node dies silently when the modem is still booting after a power cycle

1 participant