20260925 - Serve each frame's confirmed tracks by timestamp - #36
Conversation
retina-telemetry now has to send, with every detection frame it uplinks, the node's confirmed tracks as they stood after that frame: which are alive, which detection each one took, and which died. Nothing in the tracker could say that. events.jsonl is written per association, so a coasting track emits nothing there and a deletion is silent, and the /events stream is built from the same writes. A consumer could infer liveness from gaps, but only by guessing at the tracker's own deletion rules, and it could never learn which detection in a frame a track had claimed. GET /frame on the control server is that answer. The tracker keeps a record of the last 32 frames, keyed by the frame's timestamp exactly as received, and the endpoint serves one by timestamp or the latest when none is given. The timestamp is the key because it is the one identifier both sides already share: telemetry reads the frame from blah2-api, which forwards the same frame here. A 404 is an ordinary answer (not arrived yet, aged out, or cleared by a reset) and an unparseable timestamp is a 400. The record is built at the end of process_frame, under the lock the frame path already holds, and lists confirmed tracks only. A track's wire state is decided by whether it associated in this frame, not by its internal TrackState: M-of-N promotion fires on the frame count, so a track can be promoted on a frame it missed and read ACTIVE internally with no detection behind it. Deciding from association means an active track always carries a hit and a coasting one never does. A track promoted by tracklet initiation associated in that frame by construction, so it is active with its hit. Deleted tracks are captured at the deletion step, before the archive merge, because the merge can fold a later track into one that died this frame and the record would otherwise report the sum. The hit has to index the arrays as blah2-api sent them, because that is the only indexing a consumer holding the same frame can resolve. The tracker rejects impossible detections and partitions out those below the SNR floor before association, so its own list indices are useless outside it. process_streaming_frame now stamps each detection with frame_index, its position in the arrays as received, and the record reads it back. Every writer that serialises a detection picks its keys explicitly (the events file, the innovation log, the in-memory history and Track.to_dict), so the key reaches none of them, and a test holds that. A detection without the key, as in file mode, gives a null hit and the track is still reported. Each Tracker also carries a run id, generated at construction from the UTC start time and a few random hex characters, and served on /frame and /health. Track ids come from a per-process daily counter, so a same-day restart reuses them, and the run id is how a consumer tells the namespaces apart. A reset clears the frame records but keeps the run id: the id counter is class-level and survives a reset, so ids do not repeat across one. avg_snr uses the same formula as Track.to_dict, total_snr over n_associated, and inherits its quirk: total_snr includes the birth detection while n_associated does not count it, so the figure reads slightly high on a young track. Kept as is so the two never disagree. Values stay in the tracker's own units, milliseconds included; the telemetry side converts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…wait A 404 from GET /frame said only "frame not held", which leaves retina-telemetry guessing. It polls blah2-api and then asks the tracker for that frame's timestamp, and blah2-api forwards the same frame here at about the moment telemetry reads it. So a miss usually means the frame has not been processed yet and a short retry will find it. But it can also mean the frame has already aged out of the 32 held, or that this tracker is not being fed at all, and in both of those waiting is wasted time on every frame. The 404 now carries the newest held record's timestamp as `latest`, or null when nothing is held. Older than the timestamp asked for means not yet processed, so retry briefly. Newer means evicted, so give up now. Null means never fed or just reset, so never wait. The consumer compares, the tracker does not guess on its behalf. It also carries the run id. The contract has telemetry send the tracker's run alongside a null track list to say "the tracker produced nothing for this frame", and that must be sayable even when the frame is not held, so the miss cannot be the one response that omits it. `latest` is the most recently recorded frame rather than the numerically largest timestamp. The two only differ across a backwards clock step, where the most recent record is the one that says what the tracker is doing now. The fields are read under the tracker lock, in the same critical section as the failed lookup, so they describe the same moment. The 400 is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Claude finished @Purple10101's task in 3m 48s —— View job Claude's review
I traced the whole feature end to end — the Correctness checks that held up:
Worth confirming / minor notes:
Verification note: I was not able to execute Overall: clean, well-reasoned design (especially the |
Node ingest 1.6.0 lets a detection frame carry the node tracker's confirmed tracks, each active one naming the detection it took by index into the frame's arrays. Nothing the tracker exposed could fill that:
events.jsonlwrites nothing for a coasting track and nothing at all when one dies, and names detections by timestamp rather than index. This adds a read-only route that can.What it adds
GET /frame?timestamp=<ms>on the existing control port (127.0.0.1:30101) returns the tracker's result for that frame: a per-processrunid and every confirmed track, each with its state, the detection it took (hit) and its counters. Without a timestamp it returns the latest.hitindexes the arrays as received over TCP, before any rejection or SNR gate.process_streaming_framestamps each detection withframe_indexfor this; the key is never serialised anywhere, and a test checks every writer.TrackState. A track that took a detection this frame isactive; a confirmed one that did not iscoasting; one removed this frame isdeleted, once. So an M-of-N promotion on a missed frame first appears ascoasting.404carryingrunandlatest, which tells a caller whether waiting is worth it:latestbehind the asked timestamp means not processed yet, ahead means evicted,nullmeans never fed.runseparates id namespaces. Track ids repeat after a same-day restart;runchanges only on a process restart, and survives/reset, since the id counter does too.Nothing existing changes:
events.jsonl,/events,/resetand the ingest socket behave as before.Evidence
ruff check,ruff format --check,pytest(390 passed, 1 expected xfail) andpre-commit run --all-files.api.retina.fm. Over 2,200 frames, every one had a/frameanswer for its own timestamp, and all 279 active hits checked named the detection this tracker's ownevents.jsonlrecorded for that track. One track was bound to aircrafta92361. See 20260925 - Adopt node ingest v1.6.1: ADS-B position tags and the node's tracks retina-telemetry#23 for the full run, including the node's unexplained reboot at the end.Worth knowing
avg_snrreads high on a young track:total_snrincludes the detection that started it andn_associateddoes not count it, so a steady 15 dB reads 20.0 at three associations. It is sent asTrack.to_dicthas always computed it. A separate fix, if wanted.latestis the most recently recorded frame, not the largest timestamp held. The two differ only after the node's clock steps backwards.Release
retina-telemetry reads this route. It is safe in either order: against a tracker without
/frameit sends frames without tracks and logs one line. Once released, retina-node'sRETINA_TRACKER_Vdefault should move to the new tag.🤖 Generated with Claude Code