Skip to content

20260922 - Adopt node ingest v1.4.0, and claim a node by email - #21

Merged
Purple10101 merged 7 commits into
mainfrom
20260922-adopt-spec-v140
Sep 22, 2026
Merged

Purple10101 merged 7 commits into
mainfrom
20260922-adopt-spec-v140

Conversation

@Purple10101

@Purple10101 Purple10101 commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Half of one feature, across two repos. The other half is
offworldlabs/retina-gui#94, which collects the address. This is safe on
its own
: with no claim file present, nothing here calls a claim endpoint and
the node behaves exactly as it does today.

What moved

The contract went 1.2.2 to 1.4.0 while we were away, in two additive steps
that were merged and deployed to production on 2026-09-18. 1.3.0 added
GET/PUT /v1/nodes/claim and POST /v1/nodes/claim/resend; 1.4.0 put
claim_state, claim_email and claim_undeliverable on the heartbeat and
contact responses. docs/node-ingest-v1.yml is byte-identical to the server
author's copy, verified by blob hash rather than by eye.

Nothing was broken meanwhile, which is why this is adoption rather than a fix:
responses are parsed as plain dicts and never validated, so the three new keys
have been arriving and being ignored since that deploy.

Two things here are not mechanical

The generator wanted to rename our state enum. ClaimResponse.state is
titled State, exactly like HeartbeatRequest.state, and datamodel-codegen
names enum classes after that title, keeping the first and renaming the second
State1. In 1.4.0 the loser is the node's own six-value state, and it breaks
quietly: comms/lifecycle.py and wire/heartbeat.py both import State as WireState, so they would go on importing a name that still exists and now
means "where a claim stands". tools/normalise_spec.py grows a second rewrite
to stop the collision, on the same terms as the first: applied to a temporary
copy, never to the checked-in contract.

A claim conflict was asking for a configuration resend. apply_response
turned every 409 into request_config_resend(), which was right while the
only 409 meant an unrecognised config_version. The claim 409 means the
node already has an owner. Left alone, every offer to an already-owned node
would have put a spurious PUT /nodes/config on the wire. Outcome now
carries the path it answered, because a status code cannot tell the two apart.

Evidence

All gates green, including tools/check.sh --tracked.

Beyond that, the claim endpoints were driven against production on a real
node, and the server disagreed with the spec's plain reading twice. Both are
now modelled in the mock, with tests naming production and the date:

  • Declining a link returns the node to unclaimed immediately, and leaves
    the address on file. Offering that same address again is accepted, writes
    nothing and mails nothing, so the node sits unclaimed and no PUT will ever
    move it. resend is the only way out. This is why the resend trigger exists.
  • The 409 turns on the address, not on ownership: re-offering the owner's
    own address on an owned node returns 200.

The whole chain was then run on a second node against the mock, so nothing was
mailed: the GUI wrote the file, this read it across the read-only mount, offered
it, and the answer came back through status.json to the page.

What has no evidence

  • The service has never claimed against production. The endpoints were
    proved by hand with tools/probe_claim.py; the send path in __main__ has
    only talked to the mock.
  • undeliverable has only come from the mock's knob, never a real bounce.
  • Nothing has run as a released image.

One call for the reviewer

tools/probe_claim.py and tools/live-claim.sh land on main with this. They
talk to whichever server minted the node's token, which is production, and
offer makes it mail a real person. They are guarded (--confirm-send, plus
--confirm-address repeated back) and documented as a harness rather than a
feature, and they are what found both mock bugs and the 409 defect. Worth
merging on purpose rather than by accident; say the word and I will split them
onto their own branch.

🤖 Generated with Claude Code

Purple10101 and others added 7 commits September 22, 2026 11:36
The contract moved twice while we were on 1.2.2, both additively. 1.3.0 added
`GET`/`PUT /v1/nodes/claim` and `POST /v1/nodes/claim/resend` with the
`NodeClaimRequest` and `ClaimResponse` schemas, an `invalid_claim` slug on 400,
and a 409 that answers with `ClaimResponse` rather than `Error`. 1.4.0 added
`claim_state`, `claim_email` and `claim_undeliverable`, all required, to
`HeartbeatResponse` and `ContactResponse`. Both landed on production on
2026-09-18.

Nothing was broken meanwhile, which is why this is adoption rather than a fix:
`comms/client.py` parses responses into plain dicts and never validates them
against the generated response models, so the three new keys have been arriving
and being ignored since that deploy. Claim codes were removed server-side in the
same stack and cost us nothing, because we never implemented them.

The spec file is byte-identical to `contracts/nodes-v1.openapi.yaml` in
retina-server, verified by blob hash rather than by eye. `NodeConfig.beam_width_deg`
was checked on the way in, as the working agreement requires: still nullable, a
third revision running.

The reason this is not a pure regeneration is `ClaimResponse.state`. It carries
`title: State`, exactly like `HeartbeatRequest.state`, and datamodel-codegen
names generated enum classes after that title. Faced with two it keeps the first
and renames the second `State1`, and in 1.4.0 the loser is the node's own
six-value state. That breaks quietly rather than loudly: `comms/lifecycle.py`
and `wire/heartbeat.py` both do `from ...wire.models import State as WireState`,
so they would go on importing a name that still exists and now means "where a
claim stands". The first symptom would be `WireState.streaming` raising
AttributeError while building a heartbeat, at runtime, on a node.

Chasing the number would be worse than useless, since which enum gets demoted
depends on declaration order in someone else's file. So `tools/normalise_spec.py`
grows a second rewrite alongside the nullable one, on the same terms: applied to
a temporary copy, never to the checked-in contract. Each enum in `NAMED_ENUMS`
is hoisted into a component of that name and every inline occurrence replaced
with a `$ref` to it. The claim enum becomes `ClaimState`, `State` keeps its
meaning, and the three claim-state fields share one class instead of generating
two identical ones under different names. The RootModel count is unchanged at 3.

`tests/test_normalise_spec.py` is new, and covers both rewrites. The test worth
keeping is `test_no_two_inline_enums_share_a_title`: it is the general form of
this fault, so the next collision fails a test rather than silently renaming
something.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
v1.4.0 restates `claim_state`, `claim_email` and `claim_undeliverable` on every
heartbeat and contact response. This reads them and writes them to the status
document, which is the whole point of the revision: the node binds no ports, so
an owner waiting on a claim link, or one whose address bounced, learns it there
or nowhere. retina-gui already reads that file.

None of it gates anything. An unclaimed node registers, streams and beats
exactly as an owned one does; ownership decides who sees the data at the far
end, not whether it is sent. `Claim` is carried on `Snapshot` and applied by
`State.apply_levels`, which is where it belongs: the server restates the block
on every beat, so it is a level rather than an edge, and a node that missed one
is told again a minute later. That is also how a claim performed months after
setup reaches a node that was never watching for it.

It is one argument rather than three, and read as a unit gated on `claim_state`.
That is the bug this shape exists to prevent. `claim_state` is required and
non-nullable wherever it appears, so its presence says the block is there;
`claim_email` is required and *nullable*, so its null is a value, meaning nobody
has offered an address. Read field by field, a detection ack carries neither, so
its absent `claim_email` would be indistinguishable from a cleared one and the
ack arriving seconds after a heartbeat would blank the address. Both directions
are pinned by tests.

`Claim.state` is a plain string rather than the generated `ClaimState`. A fourth
value from a later server should reach an operator as itself, not take the
address and the bounce flag down with it on the way to being dropped.

The log line says whether an address is on file and never what it is, matching
what `collect/contact.py` already does when it logs the names of unrecognised
fields rather than their values. The owner sees their own address in the status
document; a container log is not the place for it. It logs only on a change,
because an unchanged level arrives once a minute forever.

Three notes for whoever picks up the rest of the claim work:

  - Nothing here calls a claim endpoint, and this deliberately stops short of
    that. `PUT /v1/nodes/claim` needs the address that owns the node, which
    nothing on a node knows. It is not the contact email: that answers "whom do
    we ring", and reusing it would mail a link that hands the node to whoever
    happened to be listed. The server has how an address relates to an account
    open as its own question.
  - `pending` is not durable. A claim link declined while a second is
    outstanding can strand a node there for about fifteen minutes before it
    falls back to `unclaimed`. That is a known server-side race, tracked there.
    Nothing should wait on `pending` or read reaching it as progress.
  - The mock holds the three values rather than deriving them, and exposes them
    through `/_control/levels`. The claim endpoints are deliberately not
    implemented in it: a mock that answers calls nobody makes would be asserting
    a shape the node has never had to parse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both docs still said v1.2.2, and neither mentioned a claim. The fact worth
recording is the one that will otherwise be re-derived wrongly, probably by
someone being helpful: `telemetry-contact.json` already holds an `email`, and
`PUT /v1/nodes/claim` wants an email, so the two look like the same field.

They are different questions. The contact address answers "whom do we ring
about this node" and is optional throughout. The claim address answers "who
owns this node", and getting it wrong mails a stranger a link that hands them
the node. An owner may well give the same address twice; that is their answer to
two questions, not a licence to infer the second from the first. The same
discipline the consent records are held to, and for the same reason: it reaches
a person.

`docs/data-sources.md` gains that comparison in §4, next to the contact document
it will be confused with, plus what a node *is* told since 1.4.0 and the fact
that `pending` is not durable. `CLAUDE.md` moves to 1.4.0, records the beam-field
check on this adoption, and gains a row in the retina-gui table for the address
nobody collects yet, marked "no, but" rather than blocking: a node that is never
claimed still registers, streams and beats.

The required-and-nullable count stays fourteen. The three new nullables are
response-side, and `tests/wire/test_serialise.py` counts payload schemas, which
is the only place `exclude_none` could do damage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`apply_response` turned every 409 into `request_config_resend()`. That was
right while the only 409 in the contract was a frame naming a `config_version`
the server never issued, and v1.3.0 added a second one that means something
entirely different: the node already has an owner. Nothing about it concerns
our configuration.

Left alone, the first `PUT /nodes/claim` against an already-owned node would
have queued a `PUT /nodes/config` as a side effect, and every retry after it
another. The node would have looked, from the server, like one whose geometry
kept going stale for no reason anybody could trace back to a claim.

A status code cannot tell the two apart, so `Outcome` now carries the path it
answered. It is the last field and defaults to the empty string, so the
positional constructions in the tests still read as they did. `_classify`
already had the path in hand; it was being spent on the error string and
discarded.

On a claim path the answer is handled separately and never asks for a resend.
It also adopts the body, because these endpoints speak `ClaimResponse`, which
names the three fields without the `claim_` prefix the heartbeat and contact
responses use, and they speak it on a refusal too: the 409 is the one refusal
in the contract that does not wear `Error`. It carries the address that won
precisely so a node can reconcile from it rather than retry, which is what
`_apply_claim_answer` does. A refusal that does wear `Error` (`invalid_claim`,
a 429, a 5xx) carries no `state`, applies nothing, and leaves the last known
claim alone.

`test_a_409_anywhere_else_still_asks_for_a_config_resend` guards the rule this
is an exception to, rather than a replacement for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A harness, not a feature. `PUT /v1/nodes/claim` still has no caller in the
service and this does not give it one: nothing on a node knows the address that
owns it, and the contact email is a different question with a worse failure.
What is missing is the trigger and the source of the address, so
`tools/probe_claim.py` supplies exactly those two by hand and nothing else.
Every request is built and sent by the service's own code: `Client`, the
generated `NodeClaimRequest`, `to_wire` and `apply_response`.

Unlike every other live-*.sh, `tools/live-claim.sh` does not point the node at
the mock. It talks to whichever server minted the node's token, which for a
node with no `RETINA_API_URL` override is production, so `offer` makes that
server mail a real person and a click on the link binds the node to the account
behind the address. Releasing it is the owner's to do from the dashboard.

That is why the phases are separate invocations rather than one script that
runs start to finish: `read` and `watch` send no mail and are safe to repeat,
while `offer` and `resend` both refuse to run without `--confirm-send`, and
`offer` additionally makes the caller repeat the address in
`--confirm-address`. A mistyped address here mails a stranger.

It deliberately does not run the service. The service registers on startup, and
the spec answers a node already holding a valid token with the opaque 403, so
it would never get a token; the one path where registration does succeed
revokes the token the node's live container is using, which the server alerts
on. A second process with the same `node_id` posting frames would also
interleave `seq` and `boot_id` with the real container's and make the node look
like it was flapping. So only the claim endpoints are driven, with the token
the node already holds, and /data is mounted read-only so that token cannot be
written even by accident.

One thing the probe reports needs explaining, because it cost a confusing
rehearsal: loading a token always queues a configuration resend, since a
restored token never comes with a `config_version`. That is `State._load`
doing its job and has nothing to do with the claim, so it is cleared on
startup. Past that point a queued resend means a response asked for one, which
is exactly what the claim 409 must not do.

The mock grows the three endpoints so the whole sequence can be rehearsed
locally before it is pointed at a real server, which is how the two bugs above
were found. It models what the contract describes: offering an address the node
already holds changes nothing and mails nothing, an address is trimmed and
lower cased before it is judged so the 255 bound describes the trimmed form,
and a claim on an owned node answers 409 with a `ClaimResponse` rather than an
`Error`, having written nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…does

Found by running the claim end to end against production on jn1, which is the
argument for the mock being written rather than generated: a generated one
would have agreed with itself.

Two things were wrong, both in the direction of the mock being stricter than
the server.

**The 409 turns on the address, not on ownership.** The mock refused any offer
on an owned node. The spec's wording is "a nomination names an address that is
not theirs", and production accepts the owner's own address with a 200 on the
idempotent path. Re-offering what the node already holds is not a conflict,
because there is nothing to refuse.

**Offering an address the node already holds really does change nothing**,
including the state, which the mock was quietly moving to `pending`. The case
that proves it is a declined claim, and it is worth writing down because it is
a trap for whoever wires this into retina-gui:

  - Declining returns the node to `unclaimed` immediately rather than
    stranding it on `pending` until the challenge expires. That settles the
    open question in the server's own ticket about what a decline covers.
  - The address survives the decline. The node reports `unclaimed` with an
    address still on file, which the very first read of an untouched node does
    not: that returns a null address.
  - So offering that same address again is "an address the node already
    holds". It answers 200, changes nothing, mails nothing, and leaves the
    node `unclaimed`. An owner who declined by accident and asked to claim
    again would get silence.
  - `POST /nodes/claim/resend` is the only call that produces another link,
    and it is what moved jn1 from `unclaimed` back to `pending`.

The mock now models all of it, and the tests name production and the date they
were checked against it, so a later revision that changes any of this shows up
as a test to revisit rather than a comment nobody trusts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The other half of retina-gui's Node claim section. It writes
/data/retina-gui/telemetry-claim.json; this reads it, decides which call
achieves what the owner asked for, and makes it. Until now nothing here called
a claim endpoint at all, because nothing on a node knew an address to offer.

`collect/claim.py` returns a `Nomination`, the spec's own word for offering an
address, and deliberately not `Claim`: `state.Claim` is where the claim
actually stands, which is the server's answer rather than the owner's ask, and
the two disagree for as long as it takes us to notice the file.

## Two keys, two kinds of thing

`email` is state, so a change in it is a local change and goes out as
`PUT /nodes/claim`, which is the call that mails a link.
`send_requested_at` is an event, and it has to exist separately because of
something the live run against production found rather than anything the spec
says outright: offering an address the node already holds is accepted, writes
nothing and mails nothing, and a declined link leaves the node `unclaimed` with
the address still on file. So a node can sit unclaimed with an address against
it and no number of offers will ever move it. `POST /nodes/claim/resend` is the
only way out.

Choosing between the two lives here rather than in retina-gui, which records
what the owner wants and nothing about the wire. This is the only side that
knows where the claim stands and what the server does with each call, and a
rule implemented on both sides is one that will eventually disagree with
itself.

## The bound on acting twice

Nothing durable records that an ask was acted on, and the ask stays in the
file, so a restart would re-read whatever it last held and mail the owner
another link, every time this container came up. `CLAIM_ASK_FRESH_FOR_S`
bounds that: five minutes, because somebody is looking at a page when they
press that button, so an ask this service was not running to see is one they
will simply make again. The cost is at most one duplicate from a restart inside
the window, against a duplicate on every restart for ever.

An offer adopts the stamp stored beside it for the same reason. An owner who
fills in the box and presses send again in one go has the link sent by the
offer; the stamp has been answered by it and must not produce a second.

The resend is sent through the client directly rather than through
`send_until_delivered`. The endpoint takes no body and that wrapper cannot make
a request without one, and passing an empty object instead would be a shape
this has never been checked against. The retry is the owner's anyway: the spec
says a resend happens "when somebody asks for the mail again, and never
automatically", so a failure is recorded and left for the person who is already
looking at the page.

Nothing here gates anything, so every failure goes to `errors[]` and never to
the status document's `detail`, which is for what stops a node working.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Purple10101
Purple10101 merged commit 2a6f3a1 into main Sep 22, 2026
2 checks passed
@Purple10101
Purple10101 deleted the 20260922-adopt-spec-v140 branch September 22, 2026 12:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant