Skip to content

20260925 - Serve each frame's confirmed tracks by timestamp - #36

Merged
Purple10101 merged 2 commits into
mainfrom
20260925-serve-each-frames-tracks
Sep 26, 2026
Merged

Purple10101 merged 2 commits into
mainfrom
20260925-serve-each-frames-tracks

Conversation

@Purple10101

@Purple10101 Purple10101 commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Node ingest 1.6.0 lets a detection frame carry the node tracker's confirmed tracks, each active one naming the detection it took by index into the frame's arrays. Nothing the tracker exposed could fill that: events.jsonl writes nothing for a coasting track and nothing at all when one dies, and names detections by timestamp rather than index. This adds a read-only route that can.

What it adds

GET /frame?timestamp=<ms> on the existing control port (127.0.0.1:30101) returns the tracker's result for that frame: a per-process run id and every confirmed track, each with its state, the detection it took (hit) and its counters. Without a timestamp it returns the latest.

  • hit indexes the arrays as received over TCP, before any rejection or SNR gate. process_streaming_frame stamps each detection with frame_index for this; the key is never serialised anywhere, and a test checks every writer.
  • State follows the hit, not the internal TrackState. A track that took a detection this frame is active; a confirmed one that did not is coasting; one removed this frame is deleted, once. So an M-of-N promotion on a missed frame first appears as coasting.
  • The last 32 frames are held. A frame not held is a 404 carrying run and latest, which tells a caller whether waiting is worth it: latest behind the asked timestamp means not processed yet, ahead means evicted, null means never fed.
  • run separates id namespaces. Track ids repeat after a same-day restart; run changes only on a process restart, and survives /reset, since the id counter does too.

Nothing existing changes: events.jsonl, /events, /reset and the ingest socket behave as before.

Evidence

  • ruff check, ruff format --check, pytest (390 passed, 1 expected xfail) and pre-commit run --all-files.
  • Against production, on jonathan-node-1 on 2026-09-25: this branch mounted into the node's own tracker container and fed by blah2-api's real forwarding, with retina-telemetry sending its tracks to api.retina.fm. Over 2,200 frames, every one had a /frame answer for its own timestamp, and all 279 active hits checked named the detection this tracker's own events.jsonl recorded for that track. One track was bound to aircraft a92361. See 20260925 - Adopt node ingest v1.6.1: ADS-B position tags and the node's tracks retina-telemetry#23 for the full run, including the node's unexplained reboot at the end.

Worth knowing

  • avg_snr reads high on a young track: total_snr includes the detection that started it and n_associated does not count it, so a steady 15 dB reads 20.0 at three associations. It is sent as Track.to_dict has always computed it. A separate fix, if wanted.
  • latest is the most recently recorded frame, not the largest timestamp held. The two differ only after the node's clock steps backwards.

Release

retina-telemetry reads this route. It is safe in either order: against a tracker without /frame it sends frames without tracks and logs one line. Once released, retina-node's RETINA_TRACKER_V default should move to the new tag.

🤖 Generated with Claude Code

Purple10101 and others added 2 commits September 25, 2026 13:23
retina-telemetry now has to send, with every detection frame it uplinks, the
node's confirmed tracks as they stood after that frame: which are alive, which
detection each one took, and which died. Nothing in the tracker could say
that. events.jsonl is written per association, so a coasting track emits
nothing there and a deletion is silent, and the /events stream is built from
the same writes. A consumer could infer liveness from gaps, but only by
guessing at the tracker's own deletion rules, and it could never learn which
detection in a frame a track had claimed.

GET /frame on the control server is that answer. The tracker keeps a record
of the last 32 frames, keyed by the frame's timestamp exactly as received,
and the endpoint serves one by timestamp or the latest when none is given. The
timestamp is the key because it is the one identifier both sides already
share: telemetry reads the frame from blah2-api, which forwards the same frame
here. A 404 is an ordinary answer (not arrived yet, aged out, or cleared by a
reset) and an unparseable timestamp is a 400.

The record is built at the end of process_frame, under the lock the frame
path already holds, and lists confirmed tracks only. A track's wire state is
decided by whether it associated in this frame, not by its internal
TrackState: M-of-N promotion fires on the frame count, so a track can be
promoted on a frame it missed and read ACTIVE internally with no detection
behind it. Deciding from association means an active track always carries a
hit and a coasting one never does. A track promoted by tracklet initiation
associated in that frame by construction, so it is active with its hit.
Deleted tracks are captured at the deletion step, before the archive merge,
because the merge can fold a later track into one that died this frame and
the record would otherwise report the sum.

The hit has to index the arrays as blah2-api sent them, because that is the
only indexing a consumer holding the same frame can resolve. The tracker
rejects impossible detections and partitions out those below the SNR floor
before association, so its own list indices are useless outside it.
process_streaming_frame now stamps each detection with frame_index, its
position in the arrays as received, and the record reads it back. Every writer
that serialises a detection picks its keys explicitly (the events file, the
innovation log, the in-memory history and Track.to_dict), so the key reaches
none of them, and a test holds that. A detection without the key, as in file
mode, gives a null hit and the track is still reported.

Each Tracker also carries a run id, generated at construction from the UTC
start time and a few random hex characters, and served on /frame and /health.
Track ids come from a per-process daily counter, so a same-day restart reuses
them, and the run id is how a consumer tells the namespaces apart. A reset
clears the frame records but keeps the run id: the id counter is class-level
and survives a reset, so ids do not repeat across one.

avg_snr uses the same formula as Track.to_dict, total_snr over n_associated,
and inherits its quirk: total_snr includes the birth detection while
n_associated does not count it, so the figure reads slightly high on a young
track. Kept as is so the two never disagree. Values stay in the tracker's own
units, milliseconds included; the telemetry side converts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…wait

A 404 from GET /frame said only "frame not held", which leaves retina-telemetry
guessing. It polls blah2-api and then asks the tracker for that frame's
timestamp, and blah2-api forwards the same frame here at about the moment
telemetry reads it. So a miss usually means the frame has not been processed
yet and a short retry will find it. But it can also mean the frame has already
aged out of the 32 held, or that this tracker is not being fed at all, and in
both of those waiting is wasted time on every frame.

The 404 now carries the newest held record's timestamp as `latest`, or null
when nothing is held. Older than the timestamp asked for means not yet
processed, so retry briefly. Newer means evicted, so give up now. Null means
never fed or just reset, so never wait. The consumer compares, the tracker
does not guess on its behalf.

It also carries the run id. The contract has telemetry send the tracker's run
alongside a null track list to say "the tracker produced nothing for this
frame", and that must be sayable even when the frame is not held, so the miss
cannot be the one response that omits it.

`latest` is the most recently recorded frame rather than the numerically
largest timestamp. The two only differ across a backwards clock step, where
the most recent record is the one that says what the tracker is doing now.
The fields are read under the tracker lock, in the same critical section as
the failed lookup, so they describe the same moment. The 400 is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@claude

claude Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Claude finished @Purple10101's task in 3m 48s —— View job


Claude's review

  • Gather context (read changed files, diff against origin/main)
  • Understand the request (code review)
  • Review retina_tracker/control.py changes
  • Review retina_tracker/server.py changes
  • Review retina_tracker/tracker.py changes
  • Review tests/test_frame_record.py
  • Post review feedback

I traced the whole feature end to end — the frame_index stamping in server.py, the took/deleted bookkeeping and frame_records eviction in tracker.py, and the /frame route in control.py — plus the concurrency story and the test suite. This is careful, well-tested work. I didn't find any correctness bugs. Notes below are things worth confirming or minor observations, not blockers.

Correctness checks that held up:

  • Snapshot-before-merge ordering (tracker.py:398-400): deleted is materialized via _frame_track before _merge_tracks() runs on self.all_tracks, so a merge folding a later track into one that died this frame can't corrupt the frozen "deleted" record. The comment explains this and the code matches it.
  • took decided at association time, not from state_status afterward (tracker.py:284-289): correctly avoids the trap where a track promoted this frame reads ACTIVE internally regardless of whether it associated. test_a_track_promoted_on_a_frame_it_missed_is_coasting nails this case.
  • Locking: process_frame (called from the TCP ingest path) and the /frame HTTP handler both take tracker_lock, so frame_records mutation and reads are properly serialized (server.py grep confirms _handle_line wraps process_frame in the same lock passed to start_control_server). The lock is released before the socket write in _frame(), which is right.
  • frame_index isolation: confirmed via grep that no other code assumes a fixed detection key set, so adding frame_index in server.py:40 is safe. test_frame_index_reaches_no_output independently backs this by checking events.jsonl, innovations.jsonl, history.snapshot(), and tracker.to_dict().
  • Eviction / replacement: frame_records.pop(timestamp, None) before re-insert correctly handles a repeated timestamp (moves it to the newest position rather than leaving a stale entry), and the while len(...) > MAX_FRAME_RECORDS: popitem(last=False) correctly evicts oldest-first. Both are exercised by tests.
  • track.id truthiness (tracker.py:396, 469): Track._generate_id returns a non-empty hex string, so if not track.id / if track.id can't misfire on a falsy-but-valid id (e.g. 0).

Worth confirming / minor notes:

  • The avg_snr inflation on young tracks (total_snr includes the seeding detection, n_associated doesn't) is already flagged explicitly in the PR description as a known, separate issue — not something this PR needs to fix, just noting I saw the same thing in _frame_track (tracker.py, near the bottom) and agree it's out of scope here.
  • _frame() in control.py parses timestamp with plain int(raw), which accepts negative numbers and leading + (e.g. timestamp=-5 won't 400, it'll just miss and 404 with latest). That's harmless since it's just a dict-key lookup, but worth confirming that's intentional rather than an oversight, since every other value in that space is a real epoch-ms timestamp.
  • Ordering in the returned tracks list is: live tracks in self.tracks order, then deleted ones appended at the end. Nothing in the contract (or tests) implies an ordering guarantee, so this is a non-issue unless a consumer starts relying on order.

Verification note: I was not able to execute ruff check, ruff format --check, or pytest in this sandbox (the commands required approval that wasn't available non-interactively), so I relied on static reading of the diff, the surrounding tracker.py/track.py/server.py code, and the new test file rather than running the suite myself. The PR body states pytest (390 passed, 1 xfail), ruff check, ruff format --check, and pre-commit run --all-files all pass, and the production validation against jonathan-node-1 (2,200+ frames, 279 active hits cross-checked against events.jsonl) is strong independent evidence this behaves as designed.

Overall: clean, well-reasoned design (especially the hit-indexes-raw-arrays and run-id-survives-reset decisions), thorough test coverage including the HTTP-level 404/400 cases, and no regressions to existing routes. Nothing here should block merge.

@Purple10101
Purple10101 merged commit dec791e into main Sep 26, 2026
2 checks passed
@Purple10101
Purple10101 deleted the 20260925-serve-each-frames-tracks branch September 26, 2026 11:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant