20260922 - Adopt node ingest v1.4.0, and claim a node by email - #21
Merged
Merged
Conversation
The contract moved twice while we were on 1.2.2, both additively. 1.3.0 added `GET`/`PUT /v1/nodes/claim` and `POST /v1/nodes/claim/resend` with the `NodeClaimRequest` and `ClaimResponse` schemas, an `invalid_claim` slug on 400, and a 409 that answers with `ClaimResponse` rather than `Error`. 1.4.0 added `claim_state`, `claim_email` and `claim_undeliverable`, all required, to `HeartbeatResponse` and `ContactResponse`. Both landed on production on 2026-09-18. Nothing was broken meanwhile, which is why this is adoption rather than a fix: `comms/client.py` parses responses into plain dicts and never validates them against the generated response models, so the three new keys have been arriving and being ignored since that deploy. Claim codes were removed server-side in the same stack and cost us nothing, because we never implemented them. The spec file is byte-identical to `contracts/nodes-v1.openapi.yaml` in retina-server, verified by blob hash rather than by eye. `NodeConfig.beam_width_deg` was checked on the way in, as the working agreement requires: still nullable, a third revision running. The reason this is not a pure regeneration is `ClaimResponse.state`. It carries `title: State`, exactly like `HeartbeatRequest.state`, and datamodel-codegen names generated enum classes after that title. Faced with two it keeps the first and renames the second `State1`, and in 1.4.0 the loser is the node's own six-value state. That breaks quietly rather than loudly: `comms/lifecycle.py` and `wire/heartbeat.py` both do `from ...wire.models import State as WireState`, so they would go on importing a name that still exists and now means "where a claim stands". The first symptom would be `WireState.streaming` raising AttributeError while building a heartbeat, at runtime, on a node. Chasing the number would be worse than useless, since which enum gets demoted depends on declaration order in someone else's file. So `tools/normalise_spec.py` grows a second rewrite alongside the nullable one, on the same terms: applied to a temporary copy, never to the checked-in contract. Each enum in `NAMED_ENUMS` is hoisted into a component of that name and every inline occurrence replaced with a `$ref` to it. The claim enum becomes `ClaimState`, `State` keeps its meaning, and the three claim-state fields share one class instead of generating two identical ones under different names. The RootModel count is unchanged at 3. `tests/test_normalise_spec.py` is new, and covers both rewrites. The test worth keeping is `test_no_two_inline_enums_share_a_title`: it is the general form of this fault, so the next collision fails a test rather than silently renaming something. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
v1.4.0 restates `claim_state`, `claim_email` and `claim_undeliverable` on every
heartbeat and contact response. This reads them and writes them to the status
document, which is the whole point of the revision: the node binds no ports, so
an owner waiting on a claim link, or one whose address bounced, learns it there
or nowhere. retina-gui already reads that file.
None of it gates anything. An unclaimed node registers, streams and beats
exactly as an owned one does; ownership decides who sees the data at the far
end, not whether it is sent. `Claim` is carried on `Snapshot` and applied by
`State.apply_levels`, which is where it belongs: the server restates the block
on every beat, so it is a level rather than an edge, and a node that missed one
is told again a minute later. That is also how a claim performed months after
setup reaches a node that was never watching for it.
It is one argument rather than three, and read as a unit gated on `claim_state`.
That is the bug this shape exists to prevent. `claim_state` is required and
non-nullable wherever it appears, so its presence says the block is there;
`claim_email` is required and *nullable*, so its null is a value, meaning nobody
has offered an address. Read field by field, a detection ack carries neither, so
its absent `claim_email` would be indistinguishable from a cleared one and the
ack arriving seconds after a heartbeat would blank the address. Both directions
are pinned by tests.
`Claim.state` is a plain string rather than the generated `ClaimState`. A fourth
value from a later server should reach an operator as itself, not take the
address and the bounce flag down with it on the way to being dropped.
The log line says whether an address is on file and never what it is, matching
what `collect/contact.py` already does when it logs the names of unrecognised
fields rather than their values. The owner sees their own address in the status
document; a container log is not the place for it. It logs only on a change,
because an unchanged level arrives once a minute forever.
Three notes for whoever picks up the rest of the claim work:
- Nothing here calls a claim endpoint, and this deliberately stops short of
that. `PUT /v1/nodes/claim` needs the address that owns the node, which
nothing on a node knows. It is not the contact email: that answers "whom do
we ring", and reusing it would mail a link that hands the node to whoever
happened to be listed. The server has how an address relates to an account
open as its own question.
- `pending` is not durable. A claim link declined while a second is
outstanding can strand a node there for about fifteen minutes before it
falls back to `unclaimed`. That is a known server-side race, tracked there.
Nothing should wait on `pending` or read reaching it as progress.
- The mock holds the three values rather than deriving them, and exposes them
through `/_control/levels`. The claim endpoints are deliberately not
implemented in it: a mock that answers calls nobody makes would be asserting
a shape the node has never had to parse.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both docs still said v1.2.2, and neither mentioned a claim. The fact worth recording is the one that will otherwise be re-derived wrongly, probably by someone being helpful: `telemetry-contact.json` already holds an `email`, and `PUT /v1/nodes/claim` wants an email, so the two look like the same field. They are different questions. The contact address answers "whom do we ring about this node" and is optional throughout. The claim address answers "who owns this node", and getting it wrong mails a stranger a link that hands them the node. An owner may well give the same address twice; that is their answer to two questions, not a licence to infer the second from the first. The same discipline the consent records are held to, and for the same reason: it reaches a person. `docs/data-sources.md` gains that comparison in §4, next to the contact document it will be confused with, plus what a node *is* told since 1.4.0 and the fact that `pending` is not durable. `CLAUDE.md` moves to 1.4.0, records the beam-field check on this adoption, and gains a row in the retina-gui table for the address nobody collects yet, marked "no, but" rather than blocking: a node that is never claimed still registers, streams and beats. The required-and-nullable count stays fourteen. The three new nullables are response-side, and `tests/wire/test_serialise.py` counts payload schemas, which is the only place `exclude_none` could do damage. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`apply_response` turned every 409 into `request_config_resend()`. That was right while the only 409 in the contract was a frame naming a `config_version` the server never issued, and v1.3.0 added a second one that means something entirely different: the node already has an owner. Nothing about it concerns our configuration. Left alone, the first `PUT /nodes/claim` against an already-owned node would have queued a `PUT /nodes/config` as a side effect, and every retry after it another. The node would have looked, from the server, like one whose geometry kept going stale for no reason anybody could trace back to a claim. A status code cannot tell the two apart, so `Outcome` now carries the path it answered. It is the last field and defaults to the empty string, so the positional constructions in the tests still read as they did. `_classify` already had the path in hand; it was being spent on the error string and discarded. On a claim path the answer is handled separately and never asks for a resend. It also adopts the body, because these endpoints speak `ClaimResponse`, which names the three fields without the `claim_` prefix the heartbeat and contact responses use, and they speak it on a refusal too: the 409 is the one refusal in the contract that does not wear `Error`. It carries the address that won precisely so a node can reconcile from it rather than retry, which is what `_apply_claim_answer` does. A refusal that does wear `Error` (`invalid_claim`, a 429, a 5xx) carries no `state`, applies nothing, and leaves the last known claim alone. `test_a_409_anywhere_else_still_asks_for_a_config_resend` guards the rule this is an exception to, rather than a replacement for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A harness, not a feature. `PUT /v1/nodes/claim` still has no caller in the service and this does not give it one: nothing on a node knows the address that owns it, and the contact email is a different question with a worse failure. What is missing is the trigger and the source of the address, so `tools/probe_claim.py` supplies exactly those two by hand and nothing else. Every request is built and sent by the service's own code: `Client`, the generated `NodeClaimRequest`, `to_wire` and `apply_response`. Unlike every other live-*.sh, `tools/live-claim.sh` does not point the node at the mock. It talks to whichever server minted the node's token, which for a node with no `RETINA_API_URL` override is production, so `offer` makes that server mail a real person and a click on the link binds the node to the account behind the address. Releasing it is the owner's to do from the dashboard. That is why the phases are separate invocations rather than one script that runs start to finish: `read` and `watch` send no mail and are safe to repeat, while `offer` and `resend` both refuse to run without `--confirm-send`, and `offer` additionally makes the caller repeat the address in `--confirm-address`. A mistyped address here mails a stranger. It deliberately does not run the service. The service registers on startup, and the spec answers a node already holding a valid token with the opaque 403, so it would never get a token; the one path where registration does succeed revokes the token the node's live container is using, which the server alerts on. A second process with the same `node_id` posting frames would also interleave `seq` and `boot_id` with the real container's and make the node look like it was flapping. So only the claim endpoints are driven, with the token the node already holds, and /data is mounted read-only so that token cannot be written even by accident. One thing the probe reports needs explaining, because it cost a confusing rehearsal: loading a token always queues a configuration resend, since a restored token never comes with a `config_version`. That is `State._load` doing its job and has nothing to do with the claim, so it is cleared on startup. Past that point a queued resend means a response asked for one, which is exactly what the claim 409 must not do. The mock grows the three endpoints so the whole sequence can be rehearsed locally before it is pointed at a real server, which is how the two bugs above were found. It models what the contract describes: offering an address the node already holds changes nothing and mails nothing, an address is trimmed and lower cased before it is judged so the 255 bound describes the trimmed form, and a claim on an owned node answers 409 with a `ClaimResponse` rather than an `Error`, having written nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…does
Found by running the claim end to end against production on jn1, which is the
argument for the mock being written rather than generated: a generated one
would have agreed with itself.
Two things were wrong, both in the direction of the mock being stricter than
the server.
**The 409 turns on the address, not on ownership.** The mock refused any offer
on an owned node. The spec's wording is "a nomination names an address that is
not theirs", and production accepts the owner's own address with a 200 on the
idempotent path. Re-offering what the node already holds is not a conflict,
because there is nothing to refuse.
**Offering an address the node already holds really does change nothing**,
including the state, which the mock was quietly moving to `pending`. The case
that proves it is a declined claim, and it is worth writing down because it is
a trap for whoever wires this into retina-gui:
- Declining returns the node to `unclaimed` immediately rather than
stranding it on `pending` until the challenge expires. That settles the
open question in the server's own ticket about what a decline covers.
- The address survives the decline. The node reports `unclaimed` with an
address still on file, which the very first read of an untouched node does
not: that returns a null address.
- So offering that same address again is "an address the node already
holds". It answers 200, changes nothing, mails nothing, and leaves the
node `unclaimed`. An owner who declined by accident and asked to claim
again would get silence.
- `POST /nodes/claim/resend` is the only call that produces another link,
and it is what moved jn1 from `unclaimed` back to `pending`.
The mock now models all of it, and the tests name production and the date they
were checked against it, so a later revision that changes any of this shows up
as a test to revisit rather than a comment nobody trusts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The other half of retina-gui's Node claim section. It writes /data/retina-gui/telemetry-claim.json; this reads it, decides which call achieves what the owner asked for, and makes it. Until now nothing here called a claim endpoint at all, because nothing on a node knew an address to offer. `collect/claim.py` returns a `Nomination`, the spec's own word for offering an address, and deliberately not `Claim`: `state.Claim` is where the claim actually stands, which is the server's answer rather than the owner's ask, and the two disagree for as long as it takes us to notice the file. ## Two keys, two kinds of thing `email` is state, so a change in it is a local change and goes out as `PUT /nodes/claim`, which is the call that mails a link. `send_requested_at` is an event, and it has to exist separately because of something the live run against production found rather than anything the spec says outright: offering an address the node already holds is accepted, writes nothing and mails nothing, and a declined link leaves the node `unclaimed` with the address still on file. So a node can sit unclaimed with an address against it and no number of offers will ever move it. `POST /nodes/claim/resend` is the only way out. Choosing between the two lives here rather than in retina-gui, which records what the owner wants and nothing about the wire. This is the only side that knows where the claim stands and what the server does with each call, and a rule implemented on both sides is one that will eventually disagree with itself. ## The bound on acting twice Nothing durable records that an ask was acted on, and the ask stays in the file, so a restart would re-read whatever it last held and mail the owner another link, every time this container came up. `CLAIM_ASK_FRESH_FOR_S` bounds that: five minutes, because somebody is looking at a page when they press that button, so an ask this service was not running to see is one they will simply make again. The cost is at most one duplicate from a restart inside the window, against a duplicate on every restart for ever. An offer adopts the stamp stored beside it for the same reason. An owner who fills in the box and presses send again in one go has the link sent by the offer; the stamp has been answered by it and must not produce a second. The resend is sent through the client directly rather than through `send_until_delivered`. The endpoint takes no body and that wrapper cannot make a request without one, and passing an empty object instead would be a shape this has never been checked against. The retry is the owner's anyway: the spec says a resend happens "when somebody asks for the mail again, and never automatically", so a failure is recorded and left for the person who is already looking at the page. Nothing here gates anything, so every failure goes to `errors[]` and never to the status document's `detail`, which is for what stops a node working. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Half of one feature, across two repos. The other half is
offworldlabs/retina-gui#94, which collects the address. This is safe on
its own: with no claim file present, nothing here calls a claim endpoint and
the node behaves exactly as it does today.
What moved
The contract went 1.2.2 to 1.4.0 while we were away, in two additive steps
that were merged and deployed to production on 2026-09-18. 1.3.0 added
GET/PUT /v1/nodes/claimandPOST /v1/nodes/claim/resend; 1.4.0 putclaim_state,claim_emailandclaim_undeliverableon the heartbeat andcontact responses.
docs/node-ingest-v1.ymlis byte-identical to the serverauthor's copy, verified by blob hash rather than by eye.
Nothing was broken meanwhile, which is why this is adoption rather than a fix:
responses are parsed as plain dicts and never validated, so the three new keys
have been arriving and being ignored since that deploy.
Two things here are not mechanical
The generator wanted to rename our state enum.
ClaimResponse.stateistitled
State, exactly likeHeartbeatRequest.state, and datamodel-codegennames enum classes after that title, keeping the first and renaming the second
State1. In 1.4.0 the loser is the node's own six-value state, and it breaksquietly:
comms/lifecycle.pyandwire/heartbeat.pybothimport State as WireState, so they would go on importing a name that still exists and nowmeans "where a claim stands".
tools/normalise_spec.pygrows a second rewriteto stop the collision, on the same terms as the first: applied to a temporary
copy, never to the checked-in contract.
A claim conflict was asking for a configuration resend.
apply_responseturned every
409intorequest_config_resend(), which was right while theonly
409meant an unrecognisedconfig_version. The claim409means thenode already has an owner. Left alone, every offer to an already-owned node
would have put a spurious
PUT /nodes/configon the wire.Outcomenowcarries the path it answered, because a status code cannot tell the two apart.
Evidence
All gates green, including
tools/check.sh --tracked.Beyond that, the claim endpoints were driven against production on a real
node, and the server disagreed with the spec's plain reading twice. Both are
now modelled in the mock, with tests naming production and the date:
unclaimedimmediately, and leavesthe address on file. Offering that same address again is accepted, writes
nothing and mails nothing, so the node sits unclaimed and no
PUTwill evermove it.
resendis the only way out. This is why the resend trigger exists.409turns on the address, not on ownership: re-offering the owner'sown address on an owned node returns
200.The whole chain was then run on a second node against the mock, so nothing was
mailed: the GUI wrote the file, this read it across the read-only mount, offered
it, and the answer came back through
status.jsonto the page.What has no evidence
proved by hand with
tools/probe_claim.py; the send path in__main__hasonly talked to the mock.
undeliverablehas only come from the mock's knob, never a real bounce.One call for the reviewer
tools/probe_claim.pyandtools/live-claim.shland on main with this. Theytalk to whichever server minted the node's token, which is production, and
offermakes it mail a real person. They are guarded (--confirm-send, plus--confirm-addressrepeated back) and documented as a harness rather than afeature, and they are what found both mock bugs and the
409defect. Worthmerging on purpose rather than by accident; say the word and I will split them
onto their own branch.
🤖 Generated with Claude Code