Skip to content

fix(#499): the delivery that passed every check and served nothing — report the posterior, capture P8 anchors, probe conda for capability - #529

Open
Polichinel wants to merge 5 commits into
developmentfrom
fix/499-preflight-appwrite-registry
Open

Polichinel wants to merge 5 commits into
developmentfrom
fix/499-preflight-appwrite-registry

Conversation

@Polichinel

Copy link
Copy Markdown
Collaborator

The 2026-09-30 rehearsal delivered successfully and served nothing.

Every structural check passed: valid manifest, coverage exactly land_gaul's 64,742 cells, the findability guard resolving all 111 objects. And tower_point — the estimator views-faoapi serves — returned zero for all 2,333,448 cells. 98.2% of cells had no signal in any of 128 draws.

It was found by hand at 4am, by someone computing anchor values for an unrelated probe.

That is the defect. Not that an undertrained model produced an empty posterior — that is what 40 lessons means — but that nothing in the chain said so.

1. tools/prereg/posterior_health.py

Runs on every delivery, between the publish and the postprocessor, because that is the first moment the pooled frame exists on disk. Reports cells with no signal in any draw, the largest draw, and how many cells carry a non-zero served point estimate.

The same numbers mean opposite things for a rehearsal and a production run, so it is told which it is. A degenerate rehearsal exits 0 and says expected; a degenerate production run exits non-zero and says the delivery is structurally valid and serves nothing. Failing the rehearsal it was designed to produce would teach the operator to ignore the exit code, which costs more than it saves.

The DEGENERATE threshold is the starkest available fact — the served estimate is zero everywhere — rather than a tuned fraction, because a threshold that needs calibration is one that gets argued with.

2. tools/prereg/capture_anchors.py

P8's anchors, moved off the pod. Until now this existed only at /root/capture_anchors.py on a machine that will be destroyed — precisely the defect this whole effort is about.

Run from the runner, which is what makes P8's protection structural rather than procedural: the rule is that anchors must not be chosen after seeing what the API returns, and running it by hand afterwards preserves that only if nobody looked first.

It records the raw draws as well as the point estimate. On the run that motivated it every point estimate was 0.0, so comparing them would have passed against an unrelated empty dataset. A cell with 3 non-zero draws of 128 and a specific 301,274 spike is a fingerprint. A probe that can only detect a disagreement is worth less than one that can also diagnose it.

It also records the views_frames version — P8 claimed to remove "compared the wrong thing" by construction on the basis that both sides run the same version. They float independently inside >=1.10.2,<2, and the serving version is not observable from outside (faoapi's /version reports the app, not its dependencies).

3. conda is probed for capability, not presence

Preflight checked command -v conda. The delivery then failed at step 4 — the last, after every GPU hour — because miniconda 26.7.1 will not create an environment until its channel Terms of Service are accepted.

A --dry-run create exercises resolution, channel access and the ToS gate together in seconds, and the refusal now names the two conda tos accept commands rather than leaving an operator to find them at 4am.

Also in this branch

The Appwrite coordinate registry check (earlier commit): platform_env_require_registry() states "the registry is the ONLY source of coordinates" (#308) and defaults to a relative hop to a sibling checkout of views-appwrite, which a pod that cloned only views-models does not have. Found mid-rehearsal with four of eight models trained.

The pattern across all four: a thing the last step needs that nothing checks before the first step spends money. And one trap worth stating — preflight already resolved un_fao's REGION and that passed, but REGION resolves a coverage declaration, a different registry entirely. A check that looks adjacent is not a check.

Verification

Ten new tests execute the tools against synthetic degenerate, sparse and healthy posteriors — including a control that a healthy posterior is not flagged, without which a tool printing DEGENERATE unconditionally would pass everything else. They assert on behaviour, not source text: the previous round of guards in this repository went 22/22 green with every fix reverted.

Both tools read the pooled frame as the ensemble writes it — y_pred.npy plus identifiers.npz, not PredictionFrame.load's values.npy. That mismatch cost twenty minutes and is stated in the code rather than left to be rediscovered.

ruff clean. 8017 passed, 234 skipped, 5 xfailed.

What these cannot do

Neither tool can prevent anything. Pooling and publishing are a single invocation, so by the time a posterior can be inspected it is already on the shelf. They make it impossible to miss, not impossible to happen. The refusal belongs where the publish decision is made — #523, which is an ADR-013 contract change and a maintainer's call.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

Polichinel and others added 5 commits September 30, 2026 02:53
… the last step's hard dependency

Third gap found in this preflight, and the third of the same shape: a thing the LAST step needs
that nothing checks before the FIRST step spends money. Found during the live rehearsal with
four of eight models already trained.

The un_fao postprocessor is fatal without the Appwrite coordinate registry —
platform_env_require_registry() states it outright, "the registry is the ONLY source of
coordinates" (#308). Its default path is a relative hop to a SIBLING checkout of
views-appwrite:

    $REPO/../views-appwrite/docs/ADRs/platform/coordinate_registry.toml

A pod that cloned only views-models does not have one. The delivery would have reached step 4,
the last, and died there after every GPU hour was spent.

WHY THE EXISTING CHECKS DID NOT COVER IT, which is the part worth remembering: preflight
already resolves un_fao's REGION and that passed. But REGION resolves a COVERAGE declaration —
a different registry entirely. Two registries, two meanings, and one passing says nothing about
the other. A check that looks adjacent is not a check.

Also recorded in the message, because it cost a real decision during the rehearsal: exporting
APPWRITE_REGISTRY after a run has started does NOT help. The delivery sources its environment
once at launch and every child inherits that copy, so the only remedy mid-run is to place the
file at the default path. Preflight now says this rather than leaving the next operator to
work it out at 2am.

The registry holds non-secret identifiers only — verified before copying it to rented
hardware, not assumed: zero lines match a token pattern and every "secret" hit is commentary
about slots. Its own header says "Coordinates are NON-SECRET identifiers. Secrets appear only
as SLOTS."

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…re P8 anchors from the runner

The 2026-09-30 rehearsal delivered successfully and served nothing. Every structural check
passed — valid manifest, coverage exactly land_gaul's 64,742 cells, findability guard
resolving all 111 objects — while tower_point, the estimator faoapi serves, returned ZERO for
all 2,333,448 cells. 98.2% of cells had no signal in any of 128 draws.

It was found by hand at 4am, by someone computing anchor values for an unrelated probe. That
is the defect: not that an undertrained model produced an empty posterior, which is what 40
lessons means, but that NOTHING IN THE CHAIN SAID SO.

THREE CHANGES, all converting "remember to look" into "cannot miss".

1. tools/prereg/posterior_health.py — runs on every delivery, between the publish and the
postprocessor, because that is the first moment the pooled frame exists on disk. Reports the
fraction of cells with no signal in any draw, the largest draw, and how many cells carry a
non-zero served point estimate.

The same numbers mean opposite things for a rehearsal and a production run, so it is TOLD which
it is. A degenerate rehearsal exits 0 and says "expected"; a degenerate production run exits
non-zero and says the delivery is structurally valid and serves nothing. Failing the rehearsal
it was designed to produce would teach the operator to ignore the exit code, which costs more
than it saves.

The DEGENERATE threshold is the starkest available fact — the served estimate is zero
EVERYWHERE — rather than a tuned fraction, because a threshold that needs calibration is one
that gets argued with.

2. tools/prereg/capture_anchors.py — P8's anchors, moved off the pod. Until now this existed
only at /root/capture_anchors.py on a machine that will be destroyed, which is precisely the
defect this whole effort is about.

Run FROM THE RUNNER, which is what makes P8's protection structural: the rule is that anchors
must not be chosen after seeing what the API returns, and running it by hand afterwards
preserves that only if nobody looked first.

It records the raw draws as well as the point estimate. On the run that motivated it every
point estimate was 0.0, so comparing them would have passed against an unrelated empty dataset.
A cell with 3 non-zero draws of 128 and a specific 301,274 spike is a fingerprint; a probe that
can only DETECT a disagreement is worth less than one that can also DIAGNOSE it. It also
records the views_frames version, because P8 claimed to remove "compared the wrong thing" BY
CONSTRUCTION on the basis that both sides run the same version — they float independently
inside >=1.10.2,<2, and the SERVING version is not observable from outside.

3. conda is probed for CAPABILITY, not presence. Preflight checked `command -v conda` and the
delivery then failed at step 4 — the last, after every GPU hour — because miniconda 26.7.1
will not create an environment until its channel Terms of Service are accepted. A --dry-run
create exercises resolution, channel access and the ToS gate together in seconds, and the
refusal now names the two `conda tos accept` commands rather than leaving an operator to find
them at 4am.

Both tools read the pooled frame as the ensemble WRITES it — y_pred.npy plus identifiers.npz,
not PredictionFrame.load's values.npy. That mismatch cost twenty minutes and is stated in the
code rather than left to be rediscovered.

Ten new tests EXECUTE the tools against synthetic degenerate, sparse and healthy posteriors,
including a control that a healthy one is not flagged — without which a tool that printed
DEGENERATE unconditionally would pass everything else. They assert on behaviour, not source
text: the previous round of guards here went 22/22 green with every fix reverted.

Neither tool can PREVENT anything. Pooling and publishing are one invocation, so by the time a
posterior can be inspected it is already on the shelf. The refusal belongs where the publish
decision is made (#523).

8017 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ucceed while it is dark

On 2026-09-30 the delivery reported success at every stage and faoapi served nothing. The
artefacts were perfect: that seat downloaded all 108 shards and independently re-assembled
2,330,712 rows. One of 108 downloads stalled, there is no retry on that path, and the run was
refused with reason="ingest_failed". NOTHING ON THIS SIDE COULD HAVE DETECTED IT — only asking
the API distinguishes "delivered" from "served", and "served" is what the definition of done
says.

The delivery now queries `forecast_serving_state` after the postprocessor, and if the run was
refused it triggers an ingest and re-checks.

THREE THINGS THAT COST REAL TIME TO ESTABLISH AND ARE NOT GUESSABLE. All three are in the
script, because the next person to hit this will be reading it at 3am:

  * THE ROUTE THAT RECOVERS A REFUSED RUN IS NOT THE ROUTE FAO USES. A subset query
    re-evaluates the newest run and performs the full ingest (~200s, verified twice).
    `/pg/data/forecast/bulk` short-circuits on the missing grid artefact and 503s in 0.2s
    WITHOUT ATTEMPTING ANYTHING — polling it reports failure forever and never retries. It is
    the obvious thing to try and it is the wrong thing.

  * `/health`'s `status` and `forecast_freshness` are NOT the authority. During the refusal
    they read "healthy, age 0.02d, not stale" while nothing was served at all. They read the
    newest record in the STORE. `forecast_serving_state` reads what is served.

  * Without a consumer API key the step says CANNOT VERIFY rather than passing quietly. The
    failure being fixed here is a stage reporting a success it had not established; a check
    that silently skips would reproduce it.

The upstream fix is views-faoapi's and is written, mutation-verified and committed — three
bounded attempts, retrying exceptions, failed results AND successful results carrying no bytes.
It does not help the run in flight, because nothing is deployed: live is v1.7.1 and v1.7.2,
v1.7.3 and 1.7.4 are all undeployed. Deploying is the operator's. So the interim procedure is
the real remedy today, and it is now in the script instead of in a message.

ALSO FIXED, and it is the helper defeating the guard: `_code_only` strips from the first `#`,
which eats the `###` inside echo strings. Two assertions about OPERATOR-VISIBLE OUTPUT now read
raw source, with the reason stated — the helper exists to stop a COMMENT standing in for code,
and an echo string is the code's behaviour, not a comment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…dependence diagnostic

Simon's question: sum the served tower-map surface over a conflict-affected country and you get
far less than history records. Is that bad calibration?

Mostly it is the wrong arithmetic, and the tool exists to separate the two readings.

WHY SUMMING MAPS MUST UNDERSHOOT. tower_point is a MODE-like summary — its own docstring says it
returns 0 "when the zero atom dominates, by density". For a zero-inflated heavy-tailed variable
the mode is ~0 almost everywhere while the mean is not. Expectation is linear and modes are not:
E[sum] = sum E[x] always holds, mode(sum) != sum mode(x) for skewed data. So summing MAP surfaces
undershoots even for a perfectly calibrated model and tells you nothing about calibration.

The right object is the distribution of the TOTAL, from summing the DRAWS.

THE FAILURE MODE THAT IS WORTH LOOKING FOR, and the reason this is more than a rebuttal. Every
cell's marginal can be perfectly calibrated while the TOTAL is badly calibrated, because the
total's spread depends on the DEPENDENCE between cells. If each draw is a coherent joint
scenario, totals have realistic spread. If draws are independent per-cell marginals stitched
together, summing ~64,818 of them washes out the tail — variances add instead of covariances
accumulating — and aggregate intervals become far too narrow. Measured as:

    dependence_ratio = Var_d(total) / sum_cells Var_d(cell)
    ~1 independent (aggregate tail understated) · >1 joint structure · <1 would itself want explaining

FALSIFIERS PRE-REGISTERED IN THE DOCSTRING, stated before the numbers were looked at: F1
independence; F2 PIT piling up near 1 (timid under-prediction); F3 90% coverage far below 0.90;
F4 MCR far below 1. F2 and F4 are not speculative here — this platform already records that MSLE
rewards timid under-prediction, which is why MCR exists as a magnitude guardrail and why ranking
on MSLE alone is refused.

HONEST ABOUT WHAT IT CANNOT DO: a FORECASTING run predicts months with no observations, so
F2-F4 need calibration or validation. Against a forecast it reports predictive totals and the
dependence ratio and SAYS so, rather than implying a comparison it did not make.

Eight tests against synthetic posteriors whose answer is known by construction: independent
cells (ratio must land near 1), joint scenarios built from a shared per-draw multiplier (ratio
must exceed 2), observed totals set to a known multiple of the predictive mean so MCR is exactly
predictable, and a CONTROL that a well-centred model trips neither F2 nor F4 — without which a
tool printing the flags unconditionally would pass everything else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
@Polichinel

Copy link
Copy Markdown
Collaborator Author

Code review

Found 7 issues:

  1. BLOCKER — both new tools fail to import on every real run. Section 3 leaves the shell in $REPO/ensembles/$ENSEMBLE and there is no cd "$REPO" before python -m tools.prereg.*, which resolves tools from the current directory (nothing is installed). Reproduced: ModuleNotFoundError: No module named 'tools.prereg' from the ensemble dir, exit 0 from the repo root. The || echo fallback then prints "posterior health reported a problem above", so a check that never ran reports its subject is broken — and P8_ANCHORS.json is never written, which is the artefact the PR body says must stop living only on a pod.

stage pool_and_publish
cd "$REPO/ensembles/$ENSEMBLE" || die "cannot enter $ENSEMBLE"
export WANDB_MODE=offline WANDB_SILENT=true
# -sa/--saved pools the member forecasts just written. Without it the ensemble refetches and
# the eight runs above are wasted. -p publishes the wire shards (ADR-013).
"$VENV/bin/python" main.py -r forecasting -f -sa -p || die "$ENSEMBLE pool-and-publish exited non-zero"
# ── 3b. is there anything IN the posterior? ───────────────────────────────────────────
# Runs between the publish and the postprocessor because that is the first moment the pooled
# frame exists on disk. It cannot prevent anything — pooling and publishing are one invocation,
# so the numbers are already on the shelf — but on 2026-09-30 a delivery passed every structural
# check while `tower_point` returned zero for all 2,333,448 cells, and that was found by hand at
# 4am. A readout every run turns "remember to look" into "cannot miss".
stage posterior_health
POOLED=$(ls -d "$REPO/ensembles/$ENSEMBLE/data/generated/predictions_forecasting_"* 2>/dev/null | tail -1)
if [ -z "$POOLED" ]; then
echo "!!! no pooled output found under $ENSEMBLE — cannot report posterior health"
else
MODE_ARG=production
[ -n "$REHEARSAL_LESSONS" ] && MODE_ARG=rehearsal
"$VENV/bin/python" -m tools.prereg.posterior_health "$POOLED" \
--mode "$MODE_ARG" --json-out "$OUT/POSTERIOR_HEALTH.json" \
|| echo "### posterior health reported a problem above — the delivery is already published"

  1. Three copies of the pooled-frame loader land in one PR, and all three drop the validation the existing reader has. tools/collapse/collapse_predictions.py raises on a missing file, wrong ndim, <2 draws, non-finite, negative, and an identifier length that does not match the row count — the last guarded by a test written because a truncated identifiers.npz would silently label the wrong cells. None of that survives into the two tools that now run on every delivery.

"""Load a pooled PredictionFrame written as y_pred.npy + identifiers.npz.
Deliberately NOT `PredictionFrame.load`: that expects `values.npy`, and the ensemble writes
`y_pred.npy`. The mismatch cost twenty minutes on 2026-09-30 and is worth stating here rather
than rediscovering.
"""
import views_frames as vf
y = np.load(target_dir / "y_pred.npy")
ids = np.load(target_dir / "identifiers.npz")
index = vf.SpatioTemporalIndex(
time=ids["time"], unit=ids["unit"], level=vf.SpatialLevel.PGM
)
return vf.PredictionFrame(y_pred=y, index=index), y, ids

  1. STATUS is written OK unconditionally, after the run has possibly just determined the posterior is DEGENERATE or that faoapi is not serving it. Both verdicts reach stdout only. The artefact bundle that comes home says OK / mode: production for a run the script itself judged bad — the shape this PR exists to fix.

https://github.com/views-platform/views-models/blob/a77f1b09a7af7d0db8acbb04ebee6c7b0c013c8e/tools/podrun/pod_run_fao_delivery.sh#L447-L450

  1. The two new $OUT artefacts are not in the start-of-run clear. A rehearsal writes P8_ANCHORS.json; a later production run that dies early leaves it in place, and the guide's rsync brings it home as that run's. P8's protection is that the anchors belong to the run being probed.

}
rm -f "$OUT/STATUS" "$OUT/REHEARSAL" "$OUT/PUBLISHED"
rm -rf "$OUT/.conda_probe"

  1. The registry path is re-derived instead of reusing the canonical helper. tools/credentials/platform_env.sh:59 already holds this expression. If that default moves, this preflight passes and the postprocessor still dies at step 4 — the failure the check was added to prevent.

# the file, or set the variable, before launching.
REGISTRY_PATH="${APPWRITE_REGISTRY:-$REPO/../views-appwrite/docs/ADRs/platform/coordinate_registry.toml}"
if [ -f "$REGISTRY_PATH" ]; then

  1. The SPARSE threshold cannot fire on the incident that motivated the tool. The 2026-09-30 posterior was 98.2% all-zero; the branch requires > 0.999. Only the DEGENERATE branch would report it, and one surviving non-zero cell out of 2.3M flips the verdict to OK. draws_nonzero_fraction is computed and used in no branch.

calibration is a threshold that gets argued with; this one cannot be.
"""
if stats["point_nonzero"] == 0:
return "DEGENERATE", (
"the served point estimate is zero for EVERY one of "
f"{stats['rows']:,} cells. Whatever is delivered, a consumer reads nothing from it."
)
if stats["rows_all_zero_fraction"] > 0.999:
return "SPARSE", (
f"{stats['rows_all_zero_fraction']:.3%} of cells have no signal in any draw; "
f"only {stats['rows_with_any_signal']:,} cells carry anything."
)
return "OK", (
f"{stats['point_nonzero']:,} cells carry a non-zero point estimate "
f"(max {stats['point_max']:.6g})."
)

  1. tools/prereg/aggregate_calibration.py (209 lines) and its 139 lines of tests appear nowhere in the PR description, are invoked by nothing, and the package docstring says the package "produces the two things" while shipping three. CLAUDE.md: "Record what you chose and why in the PR description" and "A new developer should understand the responsibilities from the package layout".

"""Evidence about a delivery that its structural checks cannot produce.
A delivery can pass every check in the chain and still be worthless. On 2026-09-30 one did: the
manifest was valid, coverage was exactly ``land_gaul``'s 64,742 cells, the findability guard
resolved all 111 objects — and ``tower_point``, the estimator views-faoapi serves, returned zero
for every one of 2,333,448 cells. The posterior was empty. Nothing in the chain said so, and it
was found by hand at 4am by someone computing anchor values for an unrelated probe.
This package produces the two things that would have said so, on every run:
python -m tools.prereg.posterior_health <pooled_dir> --mode rehearsal|production
python -m tools.prereg.capture_anchors <pooled_dir> --mode ... --out P8_ANCHORS.json
``posterior_health`` answers *is there anything in it?* — the fraction of cells with no signal in
any draw, the largest draw, and how many cells carry a non-zero served point estimate. The same
numbers mean opposite things for a rehearsal and a production run, so it is told which it is; it
cannot be inferred from the data, and that is the whole reason a human kept having to.
``capture_anchors`` implements pre-registration v3's **P8** — *the values served ARE the values
we produced*. It records both the point estimate and the raw draws, because the point estimate
alone could not discriminate on the run that motivated this: comparing 0.0 to 0.0 would have
passed against an unrelated empty dataset.
House rules. Both read the pooled frame as the ensemble writes it — ``y_pred.npy`` plus
``identifiers.npz``, not ``PredictionFrame.load``'s ``values.npy``. Anchor selection is seeded, so
it is reproducible and provably not steered by what the API returned. Neither tool can PREVENT
anything: pooling and publishing are a single invocation, so by the time a posterior can be
inspected it is already on the shelf. They exist to make it impossible to miss, not impossible
to happen — the refusal, if one is ever wanted, belongs where the publish decision is made
(views-models#523).
"""

Issues 1, 3, 4 and 5 were reproduced locally rather than inferred.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant