fix(#499): the delivery that passed every check and served nothing — report the posterior, capture P8 anchors, probe conda for capability - #529
Conversation
… the last step's hard dependency Third gap found in this preflight, and the third of the same shape: a thing the LAST step needs that nothing checks before the FIRST step spends money. Found during the live rehearsal with four of eight models already trained. The un_fao postprocessor is fatal without the Appwrite coordinate registry — platform_env_require_registry() states it outright, "the registry is the ONLY source of coordinates" (#308). Its default path is a relative hop to a SIBLING checkout of views-appwrite: $REPO/../views-appwrite/docs/ADRs/platform/coordinate_registry.toml A pod that cloned only views-models does not have one. The delivery would have reached step 4, the last, and died there after every GPU hour was spent. WHY THE EXISTING CHECKS DID NOT COVER IT, which is the part worth remembering: preflight already resolves un_fao's REGION and that passed. But REGION resolves a COVERAGE declaration — a different registry entirely. Two registries, two meanings, and one passing says nothing about the other. A check that looks adjacent is not a check. Also recorded in the message, because it cost a real decision during the rehearsal: exporting APPWRITE_REGISTRY after a run has started does NOT help. The delivery sources its environment once at launch and every child inherits that copy, so the only remedy mid-run is to place the file at the default path. Preflight now says this rather than leaving the next operator to work it out at 2am. The registry holds non-secret identifiers only — verified before copying it to rented hardware, not assumed: zero lines match a token pattern and every "secret" hit is commentary about slots. Its own header says "Coordinates are NON-SECRET identifiers. Secrets appear only as SLOTS." Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…re P8 anchors from the runner The 2026-09-30 rehearsal delivered successfully and served nothing. Every structural check passed — valid manifest, coverage exactly land_gaul's 64,742 cells, findability guard resolving all 111 objects — while tower_point, the estimator faoapi serves, returned ZERO for all 2,333,448 cells. 98.2% of cells had no signal in any of 128 draws. It was found by hand at 4am, by someone computing anchor values for an unrelated probe. That is the defect: not that an undertrained model produced an empty posterior, which is what 40 lessons means, but that NOTHING IN THE CHAIN SAID SO. THREE CHANGES, all converting "remember to look" into "cannot miss". 1. tools/prereg/posterior_health.py — runs on every delivery, between the publish and the postprocessor, because that is the first moment the pooled frame exists on disk. Reports the fraction of cells with no signal in any draw, the largest draw, and how many cells carry a non-zero served point estimate. The same numbers mean opposite things for a rehearsal and a production run, so it is TOLD which it is. A degenerate rehearsal exits 0 and says "expected"; a degenerate production run exits non-zero and says the delivery is structurally valid and serves nothing. Failing the rehearsal it was designed to produce would teach the operator to ignore the exit code, which costs more than it saves. The DEGENERATE threshold is the starkest available fact — the served estimate is zero EVERYWHERE — rather than a tuned fraction, because a threshold that needs calibration is one that gets argued with. 2. tools/prereg/capture_anchors.py — P8's anchors, moved off the pod. Until now this existed only at /root/capture_anchors.py on a machine that will be destroyed, which is precisely the defect this whole effort is about. Run FROM THE RUNNER, which is what makes P8's protection structural: the rule is that anchors must not be chosen after seeing what the API returns, and running it by hand afterwards preserves that only if nobody looked first. It records the raw draws as well as the point estimate. On the run that motivated it every point estimate was 0.0, so comparing them would have passed against an unrelated empty dataset. A cell with 3 non-zero draws of 128 and a specific 301,274 spike is a fingerprint; a probe that can only DETECT a disagreement is worth less than one that can also DIAGNOSE it. It also records the views_frames version, because P8 claimed to remove "compared the wrong thing" BY CONSTRUCTION on the basis that both sides run the same version — they float independently inside >=1.10.2,<2, and the SERVING version is not observable from outside. 3. conda is probed for CAPABILITY, not presence. Preflight checked `command -v conda` and the delivery then failed at step 4 — the last, after every GPU hour — because miniconda 26.7.1 will not create an environment until its channel Terms of Service are accepted. A --dry-run create exercises resolution, channel access and the ToS gate together in seconds, and the refusal now names the two `conda tos accept` commands rather than leaving an operator to find them at 4am. Both tools read the pooled frame as the ensemble WRITES it — y_pred.npy plus identifiers.npz, not PredictionFrame.load's values.npy. That mismatch cost twenty minutes and is stated in the code rather than left to be rediscovered. Ten new tests EXECUTE the tools against synthetic degenerate, sparse and healthy posteriors, including a control that a healthy one is not flagged — without which a tool that printed DEGENERATE unconditionally would pass everything else. They assert on behaviour, not source text: the previous round of guards here went 22/22 green with every fix reverted. Neither tool can PREVENT anything. Pooling and publishing are one invocation, so by the time a posterior can be inspected it is already on the shelf. The refusal belongs where the publish decision is made (#523). 8017 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ucceed while it is dark
On 2026-09-30 the delivery reported success at every stage and faoapi served nothing. The
artefacts were perfect: that seat downloaded all 108 shards and independently re-assembled
2,330,712 rows. One of 108 downloads stalled, there is no retry on that path, and the run was
refused with reason="ingest_failed". NOTHING ON THIS SIDE COULD HAVE DETECTED IT — only asking
the API distinguishes "delivered" from "served", and "served" is what the definition of done
says.
The delivery now queries `forecast_serving_state` after the postprocessor, and if the run was
refused it triggers an ingest and re-checks.
THREE THINGS THAT COST REAL TIME TO ESTABLISH AND ARE NOT GUESSABLE. All three are in the
script, because the next person to hit this will be reading it at 3am:
* THE ROUTE THAT RECOVERS A REFUSED RUN IS NOT THE ROUTE FAO USES. A subset query
re-evaluates the newest run and performs the full ingest (~200s, verified twice).
`/pg/data/forecast/bulk` short-circuits on the missing grid artefact and 503s in 0.2s
WITHOUT ATTEMPTING ANYTHING — polling it reports failure forever and never retries. It is
the obvious thing to try and it is the wrong thing.
* `/health`'s `status` and `forecast_freshness` are NOT the authority. During the refusal
they read "healthy, age 0.02d, not stale" while nothing was served at all. They read the
newest record in the STORE. `forecast_serving_state` reads what is served.
* Without a consumer API key the step says CANNOT VERIFY rather than passing quietly. The
failure being fixed here is a stage reporting a success it had not established; a check
that silently skips would reproduce it.
The upstream fix is views-faoapi's and is written, mutation-verified and committed — three
bounded attempts, retrying exceptions, failed results AND successful results carrying no bytes.
It does not help the run in flight, because nothing is deployed: live is v1.7.1 and v1.7.2,
v1.7.3 and 1.7.4 are all undeployed. Deploying is the operator's. So the interim procedure is
the real remedy today, and it is now in the script instead of in a message.
ALSO FIXED, and it is the helper defeating the guard: `_code_only` strips from the first `#`,
which eats the `###` inside echo strings. Two assertions about OPERATOR-VISIBLE OUTPUT now read
raw source, with the reason stated — the helper exists to stop a COMMENT standing in for code,
and an echo string is the code's behaviour, not a comment.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…dependence diagnostic
Simon's question: sum the served tower-map surface over a conflict-affected country and you get
far less than history records. Is that bad calibration?
Mostly it is the wrong arithmetic, and the tool exists to separate the two readings.
WHY SUMMING MAPS MUST UNDERSHOOT. tower_point is a MODE-like summary — its own docstring says it
returns 0 "when the zero atom dominates, by density". For a zero-inflated heavy-tailed variable
the mode is ~0 almost everywhere while the mean is not. Expectation is linear and modes are not:
E[sum] = sum E[x] always holds, mode(sum) != sum mode(x) for skewed data. So summing MAP surfaces
undershoots even for a perfectly calibrated model and tells you nothing about calibration.
The right object is the distribution of the TOTAL, from summing the DRAWS.
THE FAILURE MODE THAT IS WORTH LOOKING FOR, and the reason this is more than a rebuttal. Every
cell's marginal can be perfectly calibrated while the TOTAL is badly calibrated, because the
total's spread depends on the DEPENDENCE between cells. If each draw is a coherent joint
scenario, totals have realistic spread. If draws are independent per-cell marginals stitched
together, summing ~64,818 of them washes out the tail — variances add instead of covariances
accumulating — and aggregate intervals become far too narrow. Measured as:
dependence_ratio = Var_d(total) / sum_cells Var_d(cell)
~1 independent (aggregate tail understated) · >1 joint structure · <1 would itself want explaining
FALSIFIERS PRE-REGISTERED IN THE DOCSTRING, stated before the numbers were looked at: F1
independence; F2 PIT piling up near 1 (timid under-prediction); F3 90% coverage far below 0.90;
F4 MCR far below 1. F2 and F4 are not speculative here — this platform already records that MSLE
rewards timid under-prediction, which is why MCR exists as a magnitude guardrail and why ranking
on MSLE alone is refused.
HONEST ABOUT WHAT IT CANNOT DO: a FORECASTING run predicts months with no observations, so
F2-F4 need calibration or validation. Against a forecast it reports predictive totals and the
dependence ratio and SAYS so, rather than implying a comparison it did not make.
Eight tests against synthetic posteriors whose answer is known by construction: independent
cells (ratio must land near 1), joint scenarios built from a shared per-draw multiplier (ratio
must exceed 2), observed totals set to a known multiple of the predictive mean so MCR is exactly
predictable, and a CONTROL that a well-centred model trips neither F2 nor F4 — without which a
tool printing the flags unconditionally would pass everything else.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Code reviewFound 7 issues:
views-models/tools/podrun/pod_run_fao_delivery.sh Lines 300 to 322 in a77f1b0
views-models/tools/prereg/posterior_health.py Lines 30 to 45 in a77f1b0
views-models/tools/podrun/pod_run_fao_delivery.sh Lines 81 to 83 in a77f1b0
views-models/tools/podrun/pod_run_fao_delivery.sh Lines 233 to 235 in a77f1b0
views-models/tools/prereg/posterior_health.py Lines 70 to 85 in a77f1b0
views-models/tools/prereg/__init__.py Lines 1 to 31 in a77f1b0 Issues 1, 3, 4 and 5 were reproduced locally rather than inferred. 🤖 Generated with Claude Code - If this code review was useful, please react with 👍. Otherwise, react with 👎. |
The 2026-09-30 rehearsal delivered successfully and served nothing.
Every structural check passed: valid manifest, coverage exactly
land_gaul's 64,742 cells, the findability guard resolving all 111 objects. Andtower_point— the estimator views-faoapi serves — returned zero for all 2,333,448 cells. 98.2% of cells had no signal in any of 128 draws.It was found by hand at 4am, by someone computing anchor values for an unrelated probe.
That is the defect. Not that an undertrained model produced an empty posterior — that is what 40 lessons means — but that nothing in the chain said so.
1.
tools/prereg/posterior_health.pyRuns on every delivery, between the publish and the postprocessor, because that is the first moment the pooled frame exists on disk. Reports cells with no signal in any draw, the largest draw, and how many cells carry a non-zero served point estimate.
The same numbers mean opposite things for a rehearsal and a production run, so it is told which it is. A degenerate rehearsal exits 0 and says expected; a degenerate production run exits non-zero and says the delivery is structurally valid and serves nothing. Failing the rehearsal it was designed to produce would teach the operator to ignore the exit code, which costs more than it saves.
The
DEGENERATEthreshold is the starkest available fact — the served estimate is zero everywhere — rather than a tuned fraction, because a threshold that needs calibration is one that gets argued with.2.
tools/prereg/capture_anchors.pyP8's anchors, moved off the pod. Until now this existed only at
/root/capture_anchors.pyon a machine that will be destroyed — precisely the defect this whole effort is about.Run from the runner, which is what makes P8's protection structural rather than procedural: the rule is that anchors must not be chosen after seeing what the API returns, and running it by hand afterwards preserves that only if nobody looked first.
It records the raw draws as well as the point estimate. On the run that motivated it every point estimate was
0.0, so comparing them would have passed against an unrelated empty dataset. A cell with 3 non-zero draws of 128 and a specific 301,274 spike is a fingerprint. A probe that can only detect a disagreement is worth less than one that can also diagnose it.It also records the
views_framesversion — P8 claimed to remove "compared the wrong thing" by construction on the basis that both sides run the same version. They float independently inside>=1.10.2,<2, and the serving version is not observable from outside (faoapi's/versionreports the app, not its dependencies).3. conda is probed for capability, not presence
Preflight checked
command -v conda. The delivery then failed at step 4 — the last, after every GPU hour — because miniconda 26.7.1 will not create an environment until its channel Terms of Service are accepted.A
--dry-runcreate exercises resolution, channel access and the ToS gate together in seconds, and the refusal now names the twoconda tos acceptcommands rather than leaving an operator to find them at 4am.Also in this branch
The Appwrite coordinate registry check (earlier commit):
platform_env_require_registry()states "the registry is the ONLY source of coordinates" (#308) and defaults to a relative hop to a sibling checkout of views-appwrite, which a pod that cloned only views-models does not have. Found mid-rehearsal with four of eight models trained.The pattern across all four: a thing the last step needs that nothing checks before the first step spends money. And one trap worth stating — preflight already resolved un_fao's
REGIONand that passed, butREGIONresolves a coverage declaration, a different registry entirely. A check that looks adjacent is not a check.Verification
Ten new tests execute the tools against synthetic degenerate, sparse and healthy posteriors — including a control that a healthy posterior is not flagged, without which a tool printing
DEGENERATEunconditionally would pass everything else. They assert on behaviour, not source text: the previous round of guards in this repository went 22/22 green with every fix reverted.Both tools read the pooled frame as the ensemble writes it —
y_pred.npyplusidentifiers.npz, notPredictionFrame.load'svalues.npy. That mismatch cost twenty minutes and is stated in the code rather than left to be rediscovered.ruffclean. 8017 passed, 234 skipped, 5 xfailed.What these cannot do
Neither tool can prevent anything. Pooling and publishing are a single invocation, so by the time a posterior can be inspected it is already on the shelf. They make it impossible to miss, not impossible to happen. The refusal belongs where the publish decision is made — #523, which is an ADR-013 contract change and a maintainer's call.
🤖 Generated with Claude Code
https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K