Skip to content

Add just verify-deploy, which asks an environment whether its deploy worked - #613

Merged
jehanazad merged 2 commits into
mainfrom
tools/verify-deploy
Oct 1, 2026
Merged

jehanazad merged 2 commits into
mainfrom
tools/verify-deploy

Conversation

@Babissimo

@Babissimo Babissimo commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

ClickUp: 2.5 Add a verify-deploy command for each environment

Stacked on #617, in the stack main ← #595 ← #602 ← #620 ← #607 ← #617 ← #613 ← #596, which lands in one deploy.

Summary

Checking a deploy was a manual procedure spread over a CLAUDE.md paragraph and six memory notes. just verify-deploy <prod|staging|test> asks the environment itself, over its public hostnames only, and exits non-zero with a reason:

  • Where it looks. The hostnames, RETINA_ENV and the synthetic-fleet and radar-polling flags come from the environment's compose overlay. The origin marker comes from deploy/origin-marker.sh. The server's own answers are then read back against both.
  • Which box answered. /api/health's synthetic_fleet and polled_radar_polling are read against the overlay. Between them they tell prod, staging and test apart without a key.
  • Waiting out a restart. Health, the public map's feed, the public node list and, on the fleet's box, the unfiltered feed are polled together for up to --wait minutes. That rides out a restart's 5xx answers, unpublished feed and empty map. A 4xx, or a certificate it cannot check, ends the wait. A failed read on the last poll alone, after a good one, is forgiven as a blip.
  • Health reads degraded as the baseline it is, so it warns rather than fails.
  • The public map. It is judged on aircraft-live.json, the real-node feed the map draws on every deployed host, not on the unfiltered aircraft.json.
    • An empty public map with real nodes connected warns: real nodes can solve nothing for hours.
    • A feed that errors, times out, stops being written or has the wrong shape fails.
    • So does a box with no real node connected, except the fleet's.
  • The fleet. On the fleet's box, aircraft.json must carry aircraft that a connected synthetic node saw, so real nodes' aircraft cannot stand in for the fleet. A fleet that is not connected fails.
  • Keyed checks. Since Answer the test router's reads only for an administrator or the radar key #608, /api/test/dashboard answers only an administrator or the radar key. RETINA_ENV and task health are read from it only when the key is in VERIFY_DEPLOY_RADAR_KEY. Without the key they are reported as SKIP, and the summary line says they were skipped, not passed. The key goes in one header to that one route, and a redirect is never followed with it. It is redacted from the answer's body before anything quotes it, and again from every report line.
  • The front end. Every script and stylesheet the page names, and every chunk and lazy stylesheet those reach, must come back as what it is. A 404 is named as "the deploy does not carry it", and a 200 that is not JavaScript as the SPA fallback. --expect-string and --expect-absent search all of them and the page itself.
  • Real nodes. The connected real nodes on the public list are listed, so a run before the deploy can be compared with one after.

CLAUDE.md's deploy paragraph and ONBOARDING now point at it. No public endpoint reports the deployed commit, so the script cannot say which commit shipped. 123zgec4n48 carries that, and it is best built on 1.6's deploy/deploy.sh.

Testing

  • backend/tests/test_verify_deploy.py has 86 tests against a fake environment. Every branch was checked by deliberately breaking it: 70 such changes, and each one turns at least one test red. Among them:
    • Each map outcome has a test and a control: an empty public map over real nodes warns, a feed that errors, times out or has the wrong shape fails, and the fleet's box with no fleet aircraft fails.
    • The key stays out of stdout and stderr, including when the dashboard echoes it in an error, across the 200-character cut, or in a form repr would escape. It travels only in the dashboard request's header, and that request does not follow redirects.
    • Skipped keyed checks print a SKIP row and "skipped, not passed", and the dashboard is not asked.
  • One test pins the script's accepted content types to deploy/page-asset.sh.
  • Ran it without a key against all three environments today, and all three passed with the keyed checks skipped. Each named its box correctly: prod polled_radar_polling=True, staging neither flag, test synthetic_fleet=True with 23 fleet aircraft. On prod, --expect-string ps-sim-anomalous was found in a lazy chunk (PhysicsPage-*.js). The keyed path has run only against the fake, since nobody here holds the key.
  • Pre-commit is clean.
  • The full backend suite ran with CI's flags (-m "not external" -n 4 --dist worksteal) on this branch rebased onto main: 5225 passed, 2 skipped. The last amend touched only the script, its tests and CLAUDE.md, and the test file passes under the repo's conftest.

Review notes

  • It touches no workflow, and it only reads. It makes no SSH connection and names no address.
  • Rebased onto main after Answer the test router's reads only for an administrator or the radar key #608. The CLAUDE.md paragraph keeps main's tst and User-Agent sentences and adds this command beside them.
  • deploy/page-asset.sh and deploy/origin-marker.sh change only their header comments. They no longer claim to be the single implementation.
  • The recipe sits at the end of the justfile, in its own section.

🤖 Generated with Claude Code

@claude

This comment has been minimized.

@claude

This comment has been minimized.

@claude

This comment has been minimized.

…worked

Checking a deploy was a procedure spread across a CLAUDE.md paragraph and six memory notes, each recording a trap: Cloudflare answers 1010 to a script without a browser User-Agent, localhost curls on the droplets answer 400, /api/health reads degraded hourly and after every restart, the server answers 5xx and publishes an empty feed while it restarts and then takes minutes to refill the map, a status code cannot tell a bundle from the SPA fallback, and a lazily loaded page's change is not in the index chunk. Worktree sessions load none of that memory, so each trap was found again.

deploy/verify-deploy.py encodes them. It reads the environment's hostnames, and the flags and RETINA_ENV it should report, from that environment's compose overlay, and the origin marker from deploy/origin-marker.sh, so none can drift from its source. It polls health, the public map's feed, the public node list and, on the fleet's box, the unfiltered feed together for up to --wait minutes, failing only what is still wrong at the end; reads a degraded health as a warning; checks the marker on every /api/ answer; walks every script and stylesheet the page reaches; and with --expect-string finds a string in any of them. It only reads, and only public hostnames.

Its verdict rests on public reads because /api/test/* answers only an administrator or the radar key since #608:
- The map is judged on aircraft-live.json, the real-node feed the public map draws on every deployed host, not the unfiltered aircraft.json. An empty public map with real nodes connected warns, since real nodes can solve nothing for hours; a feed that errors, stops or has the wrong shape fails, and so does a box with no real node connected.
- On the fleet's box the unfiltered feed must carry aircraft a synthetic node saw, so real nodes' aircraft cannot stand in for the fleet.
- /api/health's synthetic_fleet and polled_radar_polling, read against the overlay, tell prod, staging and test apart without a key.
- RETINA_ENV and task health need the dashboard, so they run only when the radar key is in VERIFY_DEPLOY_RADAR_KEY. Otherwise they are reported as skipped, and the summary says they were not passed. The key goes in one header to that one route, is never followed through a redirect, and is redacted from an answer's body before anything quotes it.

No public endpoint reports the deployed commit, so it cannot say what shipped beyond --expect-string; ClickUp 123zgec4n48 carries exposing one. The connected real nodes are listed for comparison with a run before the deploy, and the in-browser look at the map stays a human check.

ClickUp: 123zgec4mxu

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Babissimo
Babissimo changed the base branch from main to chore/env-example-parity September 28, 2026 11:33
@claude

This comment has been minimized.

Base automatically changed from chore/env-example-parity to main October 1, 2026 01:33
@claude

claude Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Claude finished @jehanazad's task in 13m 1s —— View job


Claude's review of #613

Todo list

  • Gather context (PR body, changed files, prior review)
  • Full read of deploy/verify-deploy.py and the backend routes it depends on
  • Verify each candidate finding against the actual code (line numbers, backend contract)
  • Post findings

Summary

This is a synchronize re-run. I did a fresh deep pass over deploy/verify-deploy.py plus the backend endpoints it reads (/api/health, /api/radar/nodes, the aircraft feeds, /api/test/dashboard) rather than reusing the prior review's conclusions. Backend contract assumptions (health flags, feed JSON shape incl. node_ref/contributing_node_refs renaming, dashboard auth/shape, and that status has exactly two possible values on the unqualified /api/health call) all check out against backend/routes/health.py, backend/services/health.py, backend/routes/sim_ingest.py, backend/services/blah2_poller.py, backend/routes/radar.py, backend/services/tasks/aircraft_flush.py, backend/services/node_refs.py, and backend/routes/test.py. No bugs there.

Five findings from this pass, none blocking but worth a look:

  1. _asset's 4xx handling doesn't forgive transient codes (deploy/verify-deploy.py:719). if resp.status < 500 or attempt == 2: return None, why fails immediately on any 4xx, including 429 (rate limit) or 408/425, with no retry — unlike every other check in the script, which explicitly tolerates TRANSIENT_4XX (line 113). _walk() fires up to 8 concurrent requests across up to MAX_ASSETS=400 scripts/stylesheets against the production CDN; tripping rate limiting on even one asset fails the whole bundle check (and any --expect-string/--expect-absent search depending on it) on an otherwise-healthy deploy.

  2. A genuinely broken subsystem can't fail the keyed check, only warn (deploy/verify-deploy.py:579-581, backend/routes/test.py:215). backend/routes/test.py:215 sets radar_pipeline = "error" when the pipeline/tracker is truly broken (not just starting up), while other subsystems use "waiting" for an expected transient state. _keyed() bins both into one waiting dict and only ever emits a WARN — never a FAIL — so just verify-deploy prod --radar-key ... against a deploy with a dead radar pipeline still reports "PASS with warnings" and exits 0.

  3. _overlay_setting hand-parses compose YAML with a line regex instead of reusing the repo's real YAML parser (deploy/verify-deploy.py:191-194). backend/tests/test_env_example.py:65-81 already has a correct _ComposeLoader/_server_environment() for this exact file set, handling mapping-form environment: blocks, inline comments, and anchors/merge keys. The regex ^\s*-\s*{key}=(\S+)\s*$ only matches the - KEY=VALUE list form on its own line; any overlay written with a mapping form or a <<: *defaults merge key would make _overlay_setting silently return None, and environment() would raise a bare SystemExit claiming the overlay "sets no HOST_API" even though the file is fine.

  4. The fleet's unfiltered feed has no freshness check (deploy/verify-deploy.py:511, vs. the public feed's feed_age/FEED_STALE_S at lines 405/431). _fleet() only checks reading.seen_by(synthetic) > 0 from a single snapshot — there's no age field on reading.fleet analogous to feed_age. Today this is masked because aircraft_flush.py writes aircraft.json and aircraft-live.json together, so a dead flush already fails via the public feed's staleness check — but nothing in verify-deploy.py enforces that coupling, so a future change decoupling the two writes (or a bug freezing only the unfiltered one) would let the fleet check pass forever on stale data.

  5. Minor/non-blocking: _keyed() and _bundle() run sequentially in run() (deploy/verify-deploy.py:338,340) even though neither depends on the other's result. On a slow front end, _bundle()'s walk of up to 400 assets dominates wall-clock; running the two concurrently would shave time off every invocation.

Nothing here blocks the PR — #1 and #2 are the two I'd actually want addressed (both are plausible false-negative/false-positive paths in a deploy gate), #3 and #4 are latent/masked today, and #5 is pure polish.

Fix this →
· Branch: tools/verify-deploy

@jehanazad
jehanazad merged commit 28eaf3b into main Oct 1, 2026
18 checks passed
@jehanazad
jehanazad deleted the tools/verify-deploy branch October 1, 2026 01:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants