Skip to content

fix(observer): make worker, config, and diagnostics task-owned - #60

Open
agessaman wants to merge 13 commits into
observer-firmware-devfrom
feat/observer-reliability
Open

agessaman wants to merge 13 commits into
observer-firmware-devfrom
feat/observer-reliability

Conversation

@agessaman

Copy link
Copy Markdown
Owner

Summary

  • Give the MQTT worker sole ownership of runtime state: pointer-free snapshots, incarnation-checked callback events, and a revisioned config mailbox.
  • Keep diagnostic work off the mesh loop: queued NTP/OTA jobs, completion-based NTP scheduling, and ota update that waits for the check then flashes itself.
  • Fail safely when a transport, portal mutex, or MQTT event sink cannot allocate, and require native/sanitizer tests plus preset parity before observer publication.

Test plan

  • pio test -e native -e native_sanitized observer handoff/init/NTP suites
  • Heltec V3 (no PSRAM) flash on /dev/cu.usbserial-32: two TLS slots connected; after connect memory reported Free 57128 / Min 45120 / IntMax 46068
  • GitHub run-unit-tests + PR smoke on this SHA
  • Repeat portal polling + reconfigure on V3 while watching loop gaps

Host-testable building blocks for the single-owner runtime model: a
value mailbox, a coalescing configuration mailbox, an async job handle,
an SDK event channel with its slot transition, a completion-driven NTP
schedule, a durable-preference commit and slot stats JSON emission.
esp_mqtt_client_init() succeeded without a registered event sink when
registration failed, and the WebSocket wrap returned a transport with
the unpadded SDK buffer when the padded allocation failed. Both paths
now fail cleanly instead, the IDF 4 WebSocket layout is pinned to the
validated patch level, and the connection flags esp-mqtt tasks write
become atomic.
Slot objects, counters and NTP state were mutated by the worker and
read live from the loop task, CLI and SDK callbacks through volatile
flags and the diagnostic singleton. The worker now owns them outright:

- SDK callbacks push event records; the worker applies the transitions.
- Configuration reaches the worker as a coalesced mailbox revision
  carrying the prefs and radio metadata that caused it.
- Diagnostics read a published snapshot instead of live slots, and
  radio stats are sampled on the loop task that owns them.
- Forced NTP sync and the NTP diagnostic become queued jobs, so no CLI
  caller blocks the mesh loop; `get mqtt.runtime` reports sample age,
  configuration revisions and lost SDK events.
The alert reporter sampled slot monitoring, outage start and preset
name in three separate live reads, so a slot could change underneath
them. It now takes one outage snapshot per slot. The role stats JSON
uses the shared slot emitter, and both roles call bridge->loop() so
the loop task can publish its owned inputs to the worker.
Per-field writes from the mesh task could be read mid-update by the
SNMP worker. Producers now publish a complete record through a mailbox
and the worker copies it into OID storage once per loop.
AsyncTCP handlers read NodePrefs and MQTTPrefs while the loop task
could be editing them, including the password compared at login. The
server now publishes a snapshot between commands and serves every
response from it under the existing mutex, and refuses to start when
that mutex is unavailable.
Observer setters mutated live prefs and rolled back on save failure,
so a concurrent reader could see a value that never reached flash.
Edits now land on a candidate copy that the serializer writes; durable
RAM is replaced only after a verified save.
The WiFi reconnect backoff counter and the configured-SSID test were
read from tasks other than the one writing them.
`ota check` spun a worker task and blocked the caller in a delay loop
until it finished. It now queues the check and returns immediately;
repeating the command reads the cached result with its ID and age. The
flash path keeps its blocking barrier and its done flag becomes atomic.
Both commands now return immediately and are polled by repeating them.
Documents the cached result windows and the new `get mqtt.runtime`.
Release builds published without running the unit tests or the preset
parity check. Both channels now call those workflows and depend on
them, parity pins the channel under test to the candidate SHA rather
than a branch tip that may have moved, and check_observer_ci.py fails
the run if that graph is weakened. Adds a sanitized native environment
and extends the PR smoke matrix to the observer channels.
MQTT_OWNERSHIP.md described hazards the worker-ownership work has
resolved and deferred items it implemented. The phase plan moves to
docs/observer-ownership-history.md, and the restore notes flag the
Phase 2 list as historical so it is not acted on again.
Sample snapshot heap at 1 Hz, finish ota update after the queued
check, and use a task mutex for mailbox copies so multi-kilobyte
preference images no longer disable interrupts.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant