Skip to content

feat: email alert when all mesh radios are down - #171

Merged
zvx-echo6 merged 1 commit into
mainfrom
feat/radio-outage-alert
Sep 16, 2026
Merged

zvx-echo6 merged 1 commit into
mainfrom
feat/radio-outage-alert

Conversation

@zvx-echo6

Copy link
Copy Markdown
Owner

Summary

Email alert when every mesh radio is down, since mesh alerts cannot be delivered then. On 2026-09-15 both radios were down for 1-3.5 h unnoticed; a Buckhorn fire update only went out after they came back.

  • Detection: Meshtastic via the new connector.link_up (in sync with the watchdog's link-status write); MeshCore via connected.
  • Alert after radio_outage.threshold_seconds (300), then a recovery email. Short outages send nothing.
  • Restart-safe: outage state lives in mesh_health_events, and there are no transitions during the startup grace.
  • Delivery: existing EmailChannel via a notifications.destinations entry. Now sets Message-ID/Date.
  • Secrets: notifications.destinations.*.smtp_password is added to the protected fields.

Notes

  • A Meshtastic-only deployment still crashes at boot if the radio is unreachable (MeshtasticTransport.connect() raises), so the monitor never starts there. Composite (Meshtastic + MeshCore) deployments tolerate it. Not changed here.
  • Disabled until radio_outage.destination names a configured email destination.

Tests

2263 passed, 0 failed (was 2241). New: test_radio_outage.py (20), plus 2 destination secret-preservation tests.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CfJYSn4wcmPKhVb6MfSQnr

Adds RadioOutageMonitor, which emails an ops destination when every
configured radio (Meshtastic and MeshCore) has been down for
radio_outage.threshold_seconds (default 300), and sends a recovery email
when one reconnects. On 2026-09-15 both radios were down for over an hour
and nobody noticed.

- Meshtastic state comes from a new connector.link_up, kept in sync with
  the supervisor's link-status write (_connected is never cleared on link
  loss). MeshCore uses its event-driven connected flag.
- Outage start/alert/end rows are persisted in the previously unused
  mesh_health_events table, so an open outage survives restarts (the
  host watchdog restarts meshai every 2 min while the Meshtastic port is
  unreachable). No state transitions happen during startup_grace_seconds,
  which absorbs the stale "up" before the supervisor's first check.
- Email goes through the existing EmailChannel, using a named
  notifications.destinations entry (radio_outage.destination). A failed
  send retries every email_retry_seconds; the alert is only recorded once
  sent; recovery gives up after 3 failures.
- EmailChannel now sets Message-ID and Date (Postfix does not add them)
  and passes a bare envelope sender.
- notifications.destinations.*.smtp_password is now a protected secret
  field (SMTP_PASSWORD), so dashboard saves keep the ${VAR} reference.
- Hot settings: radio_outage.enabled, threshold_seconds,
  check_interval_seconds, startup_grace_seconds, destination,
  email_retry_seconds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CfJYSn4wcmPKhVb6MfSQnr
@zvx-echo6
zvx-echo6 merged commit ebf7a64 into main Sep 16, 2026
2 checks passed
@zvx-echo6
zvx-echo6 deleted the feat/radio-outage-alert branch September 16, 2026 22:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant