Skip to content

fix(ops): a message is worth sending only if a human must act on it - #433

Merged
github-actions[bot] merged 2 commits into
mainfrom
fix/alert-noise-and-env-ownership
Aug 28, 2026
Merged

fix(ops): a message is worth sending only if a human must act on it#433
github-actions[bot] merged 2 commits into
mainfrom
fix/alert-noise-and-env-ownership

Conversation

@catomean

Copy link
Copy Markdown
Collaborator

Thirteen Telegrams reached George on the evening of 2026-08-28. Two were real — and one of those two, botsmann being down, was buried in the noise of the other eleven. This makes every category of the eleven impossible, and fixes the outage class that produced the two.

What each of last night's messages was, and what happens to it now

Message Cause Now
UNIT DOWN: vitareba-app in 60s crash loop, no cooldown already fixed in #432; now via alert_once
UNIT DOWN: restic-check failed on a 25-day-old stale lock from the laptop, retried 46s later and finished clean oneshot grace: a retry in flight or already succeeded does not page. Root cause also fixed — ExecStartPre=-restic unlock on the box
FAILED UNIT: ● / RECOVERED: unit____ systemd's bullet read as the unit name already fixed earlier that evening
FAILED UNIT: cloud-init-hotplugd failed since 2026-07-22, nothing to do, never recovers masked — a unit that can only produce noise should not be in the failed set at all
zz-audit-probe failed + recovered a planted test probe ALERT_DRY_RUN=1: probes journal, they do not deliver
fleet register check failed on TEST a deliberate test send to the real chat same
2× identical register check failed … substrata port 4022 read a working tree ~15 agent sessions share, mid-edit; and had its own Telegram call with no state judges the committed register; one message per finding per day
UNIT DOWN: botsmann-app real: /opt/botsmann/shared/.env was root-owned, app runs as ubuntu repaired live, and the class is now self-healing

The changes

  • lib-alert.sh — one delivery point. alert_once/alert_clear give every sender a per-subject cooldown. An identical-text floor sits under the unkeyed path only, so a caller with no state cannot repeat itself; keyed senders bypass it, because two mechanisms arguing about one message means the quieter one silently wins. Suppressed is never invisible — the journal gets every alert, including suppressed ones.
  • notify-failure.sh — a oneshot a retry already fixed is not an incident. The unit template gains TimeoutStartSec=300: DefaultTimeoutStartSec is 90s, so the notifier would otherwise be SIGTERMed inside its own grace window and the page never sent — an anti-noise guard that becomes an anti-alert bug.
  • host-check.sh — an app whose .env its own unit User= cannot read is repaired and reported, not asked about. This took two apps down on 2026-08-28 (vitareba 18:38, botsmann 20:16) and both times the remedy was one deterministic command. The owner is not a guess: it is the unit's User=. deploy.sh asserts the same invariant where the file is written.
  • scripts/local/fleet-register-check moves into the repo — it was untracked, which is why neither of its bugs was ever reviewed.

Verification

  • npm run test:ops: 52 passed, 0 failed in the host-alert suite (33 before), whole bundle green.
  • Six mutations, each reintroducing one behaviour, each turning the suite red — floor removed, dry-run removed, oneshot grace removed, env repair downgraded to a report, per-subject cooldown removed, env stamp-clear removed. 0 inert. Mutations ran against a throwaway copy; the pushed artifact was grepped before and after.
  • Register check proven both ways live: a dirty working tree with a duplicate port → passes (reads the commit); a scratch commit that really has one → fails, would send once, second run says not repeating.

🤖 Generated with Claude Code

https://claude.ai/code/session_01AqzRcMP1uJzd7Tav5fNxQz

catomean and others added 2 commits August 28, 2026 22:40
Thirteen Telegrams reached George on the evening of 2026-08-28. Two of them
were real, and one of those two was an outage nobody had told him about in a
way he could act on. This makes every category of the other eleven impossible.

- lib-alert: one delivery point. alert_once/alert_clear give every sender a
  per-subject cooldown; an identical-text floor sits under the UNKEYED path so
  a caller with no state (the register check, which sent the same failure twice
  in 60s) still cannot repeat itself. Keyed senders bypass the floor — two
  mechanisms arguing about one message means the quieter one silently wins.
  ALERT_DRY_RUN=1 journals without delivering, because at 22:14 a test of a new
  check rang a real phone with a message saying it was a test. Suppressed is
  never invisible: the journal gets every alert, including the suppressed ones.

- notify-failure: a oneshot that a retry has already fixed is not an incident.
  restic-check failed at 18:41:42 on a 25-day-old stale lock, was re-run at
  18:42:28, finished clean at 18:44:46 — and paged, because OnFailure fires on
  the failing run and nothing looked again. It looks again. The unit template
  gains TimeoutStartSec: DefaultTimeoutStartSec is 90s, so the notifier would
  otherwise be killed inside its own grace window and the page never sent.

- host-check: an app whose .env its own unit User= cannot read is REPAIRED and
  reported, not asked about. This took two apps down today (vitareba 18:38,
  botsmann 20:16, both root-owned .env, both crash-looping) and both times the
  remedy was one deterministic command — the owner is not a guess, it is the
  unit's User=. deploy.sh asserts the same invariant where the file is written.

- cloud-init-hotplugd is masked. Failed since 2026-07-22 with nothing to do, it
  was the member that pinned the old aggregate latch at `bad` for six weeks;
  per-unit keying turned that silence into a recurring page. A unit that can
  only produce noise should not be in the failed set at all, so that
  `systemctl --failed` keeps meaning "something is wrong".

- fleet-register-check moves into the repo (it was untracked, which is why
  neither of its bugs was ever reviewed) and now judges the register as
  COMMITTED. It read a working tree that ~15 agent sessions share and paged
  twice about a duplicate port that had never been committed. CI still judges
  the tree, which is right there — in CI the tree IS the commit under review.

Six mutations, each reintroducing one of the above, each turning the suite red.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AqzRcMP1uJzd7Tav5fNxQz
watch.sh carried a byte-identical alert(), so every property added to the
shared one — the duplicate floor, ALERT_DRY_RUN, the journal-always guarantee
— silently did not apply to the watchdog's DOWN/RECOVERED messages. A second
copy of a send path is exactly how the fleet register check ended up able to
page twice in sixty seconds with no state at all: nothing is wrong with the
copy on the day it is made, and nothing updates it afterwards.

Falls back to journal-only if lib-alert.sh is somehow absent, rather than
going silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AqzRcMP1uJzd7Tav5fNxQz
@github-actions
github-actions Bot merged commit 644de63 into main Aug 28, 2026
3 checks passed
@github-actions
github-actions Bot deleted the fix/alert-noise-and-env-ownership branch August 28, 2026 20:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant