Repository navigation
feat: email alert when all mesh radios are down - #171
Merged
Merged
Conversation
Adds RadioOutageMonitor, which emails an ops destination when every
configured radio (Meshtastic and MeshCore) has been down for
radio_outage.threshold_seconds (default 300), and sends a recovery email
when one reconnects. On 2026-09-15 both radios were down for over an hour
and nobody noticed.
- Meshtastic state comes from a new connector.link_up, kept in sync with
the supervisor's link-status write (_connected is never cleared on link
loss). MeshCore uses its event-driven connected flag.
- Outage start/alert/end rows are persisted in the previously unused
mesh_health_events table, so an open outage survives restarts (the
host watchdog restarts meshai every 2 min while the Meshtastic port is
unreachable). No state transitions happen during startup_grace_seconds,
which absorbs the stale "up" before the supervisor's first check.
- Email goes through the existing EmailChannel, using a named
notifications.destinations entry (radio_outage.destination). A failed
send retries every email_retry_seconds; the alert is only recorded once
sent; recovery gives up after 3 failures.
- EmailChannel now sets Message-ID and Date (Postfix does not add them)
and passes a bare envelope sender.
- notifications.destinations.*.smtp_password is now a protected secret
field (SMTP_PASSWORD), so dashboard saves keep the ${VAR} reference.
- Hot settings: radio_outage.enabled, threshold_seconds,
check_interval_seconds, startup_grace_seconds, destination,
email_retry_seconds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CfJYSn4wcmPKhVb6MfSQnr
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Email alert when every mesh radio is down, since mesh alerts cannot be delivered then. On 2026-09-15 both radios were down for 1-3.5 h unnoticed; a Buckhorn fire update only went out after they came back.
connector.link_up(in sync with the watchdog's link-status write); MeshCore viaconnected.radio_outage.threshold_seconds(300), then a recovery email. Short outages send nothing.mesh_health_events, and there are no transitions during the startup grace.notifications.destinationsentry. Now sets Message-ID/Date.notifications.destinations.*.smtp_passwordis added to the protected fields.Notes
MeshtasticTransport.connect()raises), so the monitor never starts there. Composite (Meshtastic + MeshCore) deployments tolerate it. Not changed here.radio_outage.destinationnames a configured email destination.Tests
2263 passed, 0 failed (was 2241). New:
test_radio_outage.py(20), plus 2 destination secret-preservation tests.🤖 Generated with Claude Code
https://claude.ai/code/session_01CfJYSn4wcmPKhVb6MfSQnr