Skip to content

Measure node availability over a trailing week, not time since registration - #36

Merged
Babissimo merged 1 commit into
mainfrom
node-availability
Sep 24, 2026
Merged

Babissimo merged 1 commit into
mainfrom
node-availability

Conversation

@Babissimo

Copy link
Copy Markdown
Contributor

Replaces a node's uptime_s with its availability: the share of the minutes the server was up in which the node delivered a frame, over the trailing 7 days and 24 hours.

Why

uptime_s was the time since the node last registered with this process. Every reconnect, config change and startup reset it, and nothing persisted it, so after a deploy every node on the retina-server leaderboard read the same few minutes (all online nodes showed "0h 4m" on production on 2026-09-23).

What changes

  • availability.MinuteRing: one bit per minute over the trailing week (1,260 bytes), indexed by epoch minute, so two rings compare slot for slot.
  • NodeMetrics.record_frame marks the minute. first_seen replaces connected_at and is kept across re-registration.
  • NodeAnalyticsManager.server_minutes holds the server's own up minutes, marked through mark_server_up(). Minutes the server was down count for no one.
  • The summary's uptime_s becomes availability_7d, availability_24h (shares from 0 to 1, or None until a whole minute has been measured) and availability_measured_s.
  • A node is measured from its first whole minute after first_seen, not over the full week, so a new node is not scored against days before it existed.
  • A blocked node's frames still count as delivered. Availability says whether a node is sending; whether it is believed is reputation's business.
  • NodeMetrics.to_state/from_state and MinuteRing.to_state/from_state let the caller persist the ring, first_seen and the running counts. The counts travel together because reputation reads detections per frame. A count missing from saved state starts at zero, and a ring saved over another window is dropped, so neither can stop a restore.

Tests

tests/test_availability.py, 15 tests. Each guard was checked by a control mutation that fails a test: the slot clearing as the week advances, the intersection with server minutes, the first-seen start, the minute-in-progress exclusion, re-registration keeping first_seen, a persisted count, the late-frame branch, the blocked path, the missing-count default and the window-length check.

Full suite: 533 passed. pre-commit run --all-files is clean.

Consumer

retina-server bumps to this commit in its own PR, which persists the state in its snapshot, runs the server clock and replaces uptime on the dashboard. ClickUp: https://app.clickup.com/t/123zgec4k7v

🤖 Generated with Claude Code

…ration

uptime_s was the time since the node last registered with this process.
Registration happens on every reconnect, config change and startup
priming, and the clock was never persisted, so after a deploy every node on
the leaderboard read the same few minutes. It said nothing about how
reliably a node delivers.

NodeMetrics now marks, in a one-bit-per-minute ring over the trailing
week, each minute in which the node delivered a frame. The manager keeps
the same ring for the minutes the server itself was up (the caller marks
them through mark_server_up), and availability is the share of those up
minutes in which the node delivered, over 7 days and over 24 hours. Minutes
the server was down count for no one, so a deploy is not held against a
node. A node is measured from its first whole minute after it was first
seen, so a new node is not scored against a week it was not part of.

A blocked node's frames still count as delivered: availability says
whether a node is sending, and whether it is believed is reputation's
business.

NodeMetrics.to_state/from_state carry the ring, first_seen and the running
counts across a restart for the caller to persist. The counts travel
together because reputation reads detections per frame: detections restored
without frames would read as a flood. A count missing from saved state
starts at zero and a ring saved over another window is dropped, so a later
change to either cannot stop a restore.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Babissimo
Babissimo merged commit 1724418 into main Sep 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant