Skip to content

fix(#516,#517,#509,#518,#523): a fresh pod can reproduce the run that delivered — and a cheap rehearsal that cannot be mistaken for one - #524

Merged
Polichinel merged 11 commits into
developmentfrom
fix/516-517-fresh-pod-reproducibility
Sep 29, 2026
Merged

Polichinel merged 11 commits into
developmentfrom
fix/516-517-fresh-pod-reproducibility

Conversation

@Polichinel

@Polichinel Polichinel commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #516, #517, #509, #518. Implements #523 (its deferred half stays open).

A /falsify audit of "we are ready to start a new full 40-lesson FAO delivery run — nothing remains but creating a pod" returned FALSIFIED: three hard falsifications, all with one cause.

The cause

Every fix that made the 2026-09-29 run work was applied by hand to the pod, and the pod was destroyed. The repository never learned any of them. The pod was the artefact and nothing in git described it, so a fresh pod reproduces the original broken state.

That is the finding worth keeping. The three defects below are its instances.

H1 — the run could not be launched at all

pod_run_model.sh refuses total_lessons < 300; all eight configs declare 300; main.py has no hyperparameter override. A cheap end-to-end test — the ordinary thing to want after a run that failed late — was reachable only by deleting the guard.

The floor is right: it stops a full GPU budget being spent by accident. But it was written as though wasted money were the risk, and it is not the main one. The risk is an undertrained model reaching a partner as a real forecast, and on that axis the guard made things worse — the only way past it produced output indistinguishable from a production run.

--rehearsal <lessons> takes the count on the command line, patches only the pod's ephemeral clone, re-imports to confirm the patch took, and marks the output (REHEARSAL file; mode: and config: PATCHED after checkout in MANIFEST). The production floor is untouched when the flag is absent.

H2 — a fresh environment built successfully and WRONG (#516)

Measured, not reasoned about:

declared requirements resolve to the pod that delivered
numpy 1.26.4 1.26.4
xarray 2025.12.0 2024.3.0
pandas 3.0.6 1.5.3

#516 called this env unbuildable. It is not — it builds, on a different pandas major, silently, which is the more dangerous of the two shapes. views-datafactory requires xarray outright and pandas only in an optional extra that is not installed, so xarray is the sole carrier (its own pyproject says so) and pinning the carrier fixes it.

The obvious cap would have been wrong. xarray<2025 looks right; xarray 2024.11.0 already requires pandas>=2.1. Measured boundary: 2024.3.0 is the last release accepting pandas 1.x, 2024.5.0 moved to >=2.0, 2024.9.0 to >=2.1. The cliff is inside the 2024 line.

H3 — the pod could not publish (#517)

The string appwrite appeared nowhere in pod_run_model.sh. views-pipeline-core[appwrite] was never installed, so _build_datastore raised at publish — after the full training run. That is exactly how 2026-09-29 failed. Now installed, and imported in preflight: the whole script is written to fail early, and the one dependency that failed late was the one not checked there.

Also fixed

The second audit, and the part of this PR that matters most

This branch's first set of tests was 22 green assertions that protected nothing. An independent /falsify guard-mode pass reverted every fix in the commit, kept the comments, and got 22/22 green. 20 of 24 mutations survived; 10 of 14 guards were DECORATIVE, 4 WEAK, none held.

The mechanism is the lesson. Every assertion read the script as text, and pod_run_model.sh is unusually well commented — each fix carries a paragraph naming the incident and quoting the exact strings. So the better the comment, the weaker the guard: deleting the code left the comment, and the comment satisfied the assertion.

Worst survivor: os.environ.get("REHEARSAL_LESSONS") or "" → or "40". One word. Every production run then patches itself to 40 lessons, the floor is dead, the leftover-patch detector is dead, and because the shell variable stays empty the MANIFEST says mode: production and no marker is written. A 40-lesson model labelled a production delivery, suite fully green.

And one guard certified a safety property that did not exist. It claimed the guide documented the /workspace chmod trap. The guide did not — no "world-readable", no "network filesystem", no C-154, no #518. It passed because /workspace is on line 104 and chmod on line 180, joined by .* under re.S — the identical defect the same commit message boasted of having found and deleted elsewhere. That is worse than no guard, because it stopped the next reader looking. The warning is now written.

pod_run_model.sh warns against exactly this in three places, citing #501 "the guard that was not one". The guards did it anyway. That is the argument for the independence rule: the author's own model named the trap three times in one commit and stepped into it.

What the guards do now

  • the config-check program is extracted from its heredoc and executed against fixture models in real git repos — that block holds all the rehearsal/production logic and was wholly unguarded. Includes an adversarial fixture whose literal is patchable but whose get_hp_config() returns 300 regardless, so a verification comparing target against itself is caught.
  • the shell argument parser is executed — deleting the --rehearsal) case arm removes the escape hatch and left every text guard green.
  • remaining text assertions read comment-stripped source.
  • requirements guards assert specifier semantics via packaging.SpecifierSet — is 1.5.3 admitted, is 3.0.6 refused — not pin presence. pandas>=3.0 passed the old guard, and marker-gated pins (; python_version < "3.10", inert on 3.11) defeated the identical-pins check while keeping the captured text identical.
  • the roster guard calls get_hp_config() instead of matching the literal.

Mutation replay — all 18 now caught

X7 production self-patch · X8 bare flag · X22 case arm deleted · X1 extra dropped · X2 imports commented out · X5 2026-09-29 install order · X6 floor prints instead of exits · X9 detector on wrong branch · X10 tautological verification · X11 marker block deleted · X12 marker cleared at end · X13 MANIFEST branches swapped · X14 pinned to the broken versions · X15 upper bounds dropped · X17 marker-gated inert pins · X18 get_hp_config() returns 40 · X20b toolz override as comment · control: guide warning deleted.

X16 (xarray == 2024.3.0 — stricter and correct) previously failed claiming no pin existed; it now passes, so the guard no longer false-alarms on a legal PEP 508 spelling.

Two honest notes: X18 was recorded SURVIVED by the audit but had not applied — the config returns a dict literal so the mutation's sed matched nothing; re-applied correctly, it is caught. And X19 now passes because the warning it was meant to defeat genuinely exists; the control proves the guard fires when that warning is removed.

Verification

  • Clean-checkout simulation (C-110 discipline — a green working tree proves nothing): cloned this branch fresh and resolved both environments from its own files. Postprocessor env → numpy 1.26.4, pandas 1.5.3, xarray 2024.3.0, views-datafactory 1.13.0. Pod training venv → appwrite 13.6.1, views-hydranet 0.1.2, views-pipeline-core 3.3.4, torch 2.14.0, pandas 1.5.3. Both match the pod that delivered.
  • The chain's fixes are reachable: hydranet 0.1.2 (sniffer), pipeline-core 3.3.4 (dedup), postprocessing 1.4.0 (pinned in both launchers by fix(#439): bump VIEWS_POSTPROCESSING_PIN 1.1.1 → 1.4.0 in both launchers #522).
  • ruff check . clean. 8077 passed, 209 skipped, 5 xfailed.
  • The 8 untracked calibration_delivery_* directories are generated model output and stay excluded.

A correction to an earlier claim in this description. Two full-suite runs on an unchanged tree reported 8058 passed/209 skipped and 7758 passed/470 skipped/3 failed, and I attributed that to pytest-randomly. That was wrong — no such plugin is installed (pytest 9.0.3, nothing matching random in the environment), so the -p no:randomly I added to subsequent runs was a no-op and proved nothing.

The three failures (test_falsification_catalog_enhancements, test_update_readme_survives_missing_data_client) passed in isolation, pass in a clean clone of development, and have not recurred in any of the six later full-suite runs, the last of which was plain pytest with no flags and returned the same 8097 passed/209 skipped/5 xfailed. The 261 extra skips alongside 3 failures is the shape of a shared fixture erroring and its dependents skipping, which fits a transient condition during that one run — the guard-audit subagent was mutating and reverting files in this tree at the time. I am recording this as unexplained rather than explained, because a confident wrong diagnosis is what this PR is largely about.

Deliberately not done

A rehearsal is MARKED, not REFUSED. The publish step is not in this runner — its header is accurate that it uploads nothing — so nothing stops a rehearsal's forecasts being published. views-pipeline-core reviewed the shape and argued convincingly against detecting a rehearsal: that repo cannot know what a lesson is, and inferring undertraining would be guessing at a property it has no access to. The agreed design is the inverse — refuse to publish unless a run positively declares itself complete, absence of a declaration being a refusal — homed in the manifest's existing provenance block and applied at both doors (sampled_forecast_publisher.py::publish_sampled_forecast and prediction/io.py::_upload_to_prediction_store).

That is an ADR-013 contract change across four repos on a partner-visible manifest, so it is a maintainer decision, not a peer request. Recorded on #523, which stays open. Until it lands, do not publish a directory containing a REHEARSAL file.

Register entries for this audit are deferred behind a named trigger: #520 merging, since it already consumes C-154/C-155 and both branches would otherwise fight over the header counts.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

Polichinel and others added 11 commits September 29, 2026 23:14
…nd WRONG

#516 said a fresh un_fao environment is unbuildable. Measured on 2026-09-29, it is
not: it builds, on a different pandas MAJOR, in silence.

    declared:  views-datafactory>=1.9.0,<2.0.0 + numpy>=1.26.4,<2.0.0
    resolved:  numpy 1.26.4, xarray 2025.12.0, pandas 3.0.6
    the pod that produced the first FAO delivery: xarray 2024.3.0, pandas 1.5.3

Neither xarray nor pandas was pinned anywhere. The working versions existed only as
hand-pins on a pod that has since been destroyed, so a fresh pod silently picks
different ones. A build failure is loud; this is not, which makes it the more
dangerous of the two shapes the issue conflated.

views-datafactory requires xarray outright and pandas only in an optional extra that
is not installed, so pandas arrives THROUGH xarray. xarray is the sole carrier —
datafactory's own pyproject comment says so — and pinning the carrier fixes it.

The bound is measured, not guessed. xarray 2024.3.0 is the LAST release that accepts
pandas 1.x; 2024.5.0 moved to pandas>=2.0 and 2024.9.0 to pandas>=2.1. A tidy-looking
`<2025` cap would therefore have been WRONG: the cliff is inside the 2024 line, not at
the year boundary. Verified by resolving the file, not by reading version numbers.

Declared identically in both postprocessors because they share one prefix (C-116);
pinning only one lets whichever runs last decide pandas for both, which is the same
trap the existing numpy pin already documents one line above.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…d --rehearsal

Three findings from one /falsify audit of "we are ready for a new 40-lesson run"
(FALSIFIED). All three had one cause: every fix that made the 2026-09-29 run work was
applied BY HAND to the pod, and the pod was destroyed. The pod was the artefact and
nothing in git described it.

#517 — the publish path was never installed. The string "appwrite" appeared nowhere in
pod_run_model.sh, so `_build_datastore` raised at publish time, AFTER the full training
run. That is exactly how the first FAO delivery attempt failed. The extra is now
requested explicitly with no version, so views-hydranet's own range still decides which
pipeline-core is installed, and the client is IMPORTED in preflight — the whole script
is written to fail early, and the one dependency that failed late was the one not
checked there. Verified by a real resolve: appwrite 13.6.1 alongside views-hydranet
0.1.2 and views-pipeline-core 3.3.4, with pandas 1.5.3 and numpy 1.26.4 intact. The
same resolve confirms the C-151 toolz override is necessary rather than folklore —
without it the resolver lands on toolz 0.11.2.

#523 — a deliberate cheap run was reachable only by deleting the guard against an
accidental one. The >=300 floor stays; `--rehearsal <lessons>` is the escape hatch. It
takes the count on the command line, because main.py has no hyperparameter override and
a flag without a count still forces an edit to a tracked config. It patches the POD's
clone only, then re-imports to CONFIRM the patch took — a substitution that silently
missed would give a 300-lesson run wearing a rehearsal label, or the reverse.

The marking matters more than the permission. A rehearsal's parquets are structurally
identical to a production run's and pass every check including the runner's own, so a
REHEARSAL file is written beside them and MANIFEST gains `mode:` and `config: PATCHED
after checkout` — the git sha alone no longer describes the run.

Found while reviewing this change, not in the audit: section 2 does not re-clone when
.git exists, so a production run on a pod that had rehearsed would read the LEFTOVER
patch. The floor still refused it, but blamed the committed config and advised
--rehearsal — a correct refusal for the wrong reason, sending the operator to edit the
wrong file. It now detects the dirty config and names it.

Deferred, and it is the weaker half: this runner MARKS a rehearsal but cannot REFUSE to
publish one, because the publish step is not in this script. That guard belongs in
views-pipeline-core. An xfail(strict) stub for it was written and DELETED — its regex
spanned the whole file under re.S and XPASSed by accident, i.e. the exact class of
guard-that-cannot-fire this audit exists to find.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
… the missing warning

An independent /falsify guard-mode audit of 53bfba9 reverted EVERY fix in that commit,
kept the comments, and got 22/22 green. 20 of 24 mutations survived. 10 of 14 guards
were DECORATIVE, 4 WEAK, none held.

The mechanism is the lesson. Every assertion read pod_run_model.sh as TEXT, and that
script is unusually well commented — each fix carries a paragraph naming the incident and
quoting the exact strings. So THE BETTER THE COMMENT, THE WEAKER THE GUARD: deleting the
code left the comment, and the comment satisfied the assertion. Deleting the appwrite
extra from the install line was invisible because the comment above it names
views-pipeline-core[appwrite].

Worst survivor: `os.environ.get("REHEARSAL_LESSONS") or ""` -> `or "40"`. One word. Every
production run then patches itself to 40 lessons, the floor is dead, the leftover-patch
detector is dead, and because the SHELL variable stays empty the MANIFEST says
`mode: production` and no REHEARSAL marker is written. A 40-lesson model labelled a
production delivery, suite fully green.

AND ONE GUARD CERTIFIED A SAFETY PROPERTY THAT DID NOT EXIST. It claimed the RunPod guide
documents the /workspace chmod trap. The guide did not: no "world-readable", no "network
filesystem", no C-154, no #518. It passed because `/workspace` is on line 104 and `chmod`
on line 180, joined by `.*` under re.S — the identical defect the previous commit message
boasted of having found and deleted elsewhere. That is worse than no guard, because it
stopped the next reader looking. The warning is now written (#518, C-154): /workspace is a
network filesystem where chmod 600 returns success and does nothing, leaving a credential
at mode 666 with no error to notice.

What changed in the tests:
  - the config-check program is EXTRACTED FROM ITS HEREDOC AND EXECUTED against fixture
    models in real git repos. That block holds all the rehearsal/production logic and was
    wholly unguarded. Includes an adversarial fixture whose literal is patchable but whose
    get_hp_config() returns 300 regardless, so a "verification" that compared target
    against itself is caught.
  - remaining text assertions read COMMENT-STRIPPED source, so a comment can never stand
    in for code.
  - requirements guards assert SPECIFIER SEMANTICS via packaging.SpecifierSet — is 1.5.3
    admitted, is 3.0.6 refused — instead of pin presence. `pandas>=3.0` passed the old one,
    and marker-gated pins (`; python_version < "3.10"`, inert on 3.11) defeated the
    identical-pins check while keeping the captured text identical.
  - the roster guard CALLS get_hp_config() instead of matching the literal, which is what
    pod_run_model.sh itself insists on, citing #501 "the guard that was not one".

pod_run_model.sh warned against this exact failure in three places and the guards did it
anyway. That is the argument for the independence rule: the author's own model named the
trap three times in one commit and stepped into it.

35 tests, up from 22.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
… parsing

Two gaps found by replaying the guard-mode audit's mutations against the rewritten
guards — the replay is the point, since the previous version passed 22/22 with every fix
reverted.

#509 — the pod installed views-datafactory>=1.9.0. The credential-handling fixes landed
in 1.13.0: before it the client could carry a netrc credential across a redirect to
another host and embed it in error messages. This script only ever runs on hardware we do
not own, carrying exactly that credential, and the guide already prescribes >=1.13.0
explicitly while the script did not. A resolver picks 1.13.0 anyway today, but "the
resolver will probably do the right thing" is precisely the reasoning that put xarray
2025.12.0 and pandas 3.0.6 into a fresh environment (#516). Verified: the floor still
resolves to 1.13.0 with hydranet 0.1.2, pipeline-core 3.3.4, appwrite 13.6.1, pandas
1.5.3.

The shell argument parser was unguarded. Executing the config-check program does not
cover it — the flag never reaches Python if the shell refuses it first. Deleting the
`--rehearsal)` case arm removes H1's escape hatch entirely and left every guard green,
because the string survives in the header comment and in USAGE. Now executed against the
real script: the flag and its count must be CONSUMED (distinguished from "unknown option"
by which error appears), a missing count must be refused, a non-integer must be refused,
and — the control, without which the first assertion proves nothing — a genuinely unknown
option must still be rejected. Argument parsing precedes `mkdir -p "$OUT"`, so these cases
exit before touching the filesystem.

MUTATION REPLAY — every surviving mutation from the audit is now caught:

  X7  `or ""` -> `or "40"` (production runs patch themselves)        3 failed
  X8  --rehearsal becomes a bare flag                               3 failed
  X22 the --rehearsal case arm deleted                              3 failed
  X1  extra dropped from the install line, comment kept             1 failed
  X2  both appwrite imports commented out, prose comment kept       1 failed
  X5  2026-09-29 install order restored + reassuring comment        1 failed
  X6  the floor prints instead of exiting                           1 failed
  X9  leftover-patch detector flipped to the wrong branch           1 failed
  X10 patch verification becomes `if target != target`              1 failed
  X11 the marker-writing block deleted                              1 failed
  X12 stale-marker clear moved to the end of the run                1 failed
  X13 MANIFEST mode branches swapped                                1 failed
  X14 both files pinned to xarray 2025.12 / pandas 3.0              8 failed
  X15 upper bounds dropped                                          8 failed
  X17 pins marker-gated inert on 3.11                               6 failed
  X18 get_hp_config() returns 40, the 300 literal untouched         1 failed
  X20b toolz override replaced by a comment                         1 failed
  control: the guide's chmod warning deleted                        1 failed

X16 (`xarray == 2024.3.0`, a stricter and CORRECT pin) previously failed with a message
claiming no pin existed; it now passes, so the guard no longer raises a false alarm
against a legal PEP 508 spelling.

Two honest notes. X18 was recorded SURVIVED in the audit but had not applied — the config
returns a dict literal, so the mutation's sed matched nothing. Re-applied correctly, it is
caught. And X19 (replacing the credential chmod with an unrelated one) now passes, because
the /workspace warning it was meant to defeat genuinely exists; the control above proves
the guard fires when that warning is removed.

40 tests, up from 35.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ested for

The previous commit shipped the guard and not the fix. The sequence: edit the install
line, write the test, verify the test fires by reverting the line with sed, then
`git checkout -- tools/podrun/pod_run_model.sh` to undo the sed — which reverted to the
last COMMIT, discarding the uncommitted 1.13.0 edit along with the sed. The commit then
captured the test plus a 1.9.0 install line.

Caught by that test on the next full-suite run, which is the whole argument for writing
guards that execute: a text-matching guard for this would have been satisfied by the
comment above the line, which names 1.13.0 and explains why.

Also worth recording, because it made the failure hard to see for two runs: this suite
runs pytest-randomly, so ordering varies between runs. Two earlier runs reported 8058
passed/209 skipped and 7758 passed/470 skipped/3 failed with an unchanged tree, and the
three failures were ordering interference in pre-existing tests, not in anything here.
With `-p no:randomly` the counts are stable at 8076 passed/209 skipped/5 xfailed and the
only failure was this one. Verify against a fixed order before concluding a change is
responsible for a red suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ds typed on a pod

The fourth instance of the pattern the /falsify audit found three of, and the one that
mattered most. On 2026-09-29 the whole FAO delivery — eight forecasting runs, the
rusty_bucket pool, the publish, the un_fao postprocessor — was typed by hand on a rented
pod. It worked, the pod was destroyed, and nothing recorded what had been done.

The audit missed it because it audited the script that exists (pod_run_model.sh) instead of
asking what the delivery needs. pod_run_model.sh runs `-r calibration -t -e` and its header
is accurate that it uploads nothing: it is Track A, the research frames. Nothing in tools/
drove Track B at all — `grep -rln "prediction_store\|r forecasting" tools/` returns one
unrelated file.

This is not a stopgap. fimbulthul, which #499 names for every Track B step, has been
unreachable since 2026-09-25, so a pod is the only hardware there is.

pod_run_model.sh gains `--forecast`: `-r forecasting -t -f`, and it SKIPS sections 4 and 5.
Those are calibration deliverables — a forecast has one origin, so the 13-parquet count
check would refuse it, and what the FAO chain consumes is the pooled ensemble output. Kept
in this script rather than a copy because everything before section 3 is identical and
already exercised; a second copy would be a second place for C-151 to rot.

pod_run_fao_delivery.sh drives the chain. Two details that are not obvious and that cost
real time to establish:

  - `-sa/--saved` on the ensemble is REQUIRED. Without it rusty_bucket refetches instead of
    pooling the member forecasts just written, and the eight runs are wasted.
  - the two legs need DIFFERENT INTERPRETERS. This pod builds a uv venv;
    tools/launcher/postprocessor.sh uses `conda shell.bash hook` / `conda create --prefix`.
    A pod can satisfy one and not the other, and step 4 is last — so a missing conda kills
    the delivery after every GPU hour is spent. Preflight now checks for it.

`--preflight` is the mode that protects the money. At 300 lessons step 1 alone is ~16 GPU
hours, and everything it checks is knowable in seconds: checkout, venv, netrc, the three
publish secrets, conda, that un_fao's REGION resolves to land_gaul rather than a disarmed
value (C-110), GPU, disk. It accumulates and reports every problem rather than dying on the
first, because a round trip per problem is billed by the second.

The publish secrets are checked HERE and not at first publish because
views-pipeline-core's PredictionStoreConfig claims to read them "once at startup and fail
loud … preventing silent failures after hours of training" and does not — it is called from
_build_datastore, after training (views-pipeline-core#557). Until that moves, this preflight
is the only check that happens before the money is spent.

A rehearsal is marked at every level, and the marker says the uncomfortable part out loud:
nothing downstream refuses a marked rehearsal (#523), so its undertrained forecasts DO reach
the FAO shelf and ARE servable. That is deliberate — it is the only way to test the chain —
and it is why the final stage reads back what landed with `tools.liveness`, by name, rather
than inferring success from an exit code. On 2026-09-29 the publish reported success having
written nothing (C-155).

19 guards, written after the lesson that 10 of 14 of the previous set were decorative. The
parser and `--preflight` are EXECUTED; the remaining text assertions read comment-stripped
source and cover only ordering facts that cannot be executed without a GPU, a store and a
partner bucket. PODRUN_ROOT exists so preflight is exercisable off a pod; it relocates the
workspace and relaxes no check.

Every preflight check was proven able to fire, including by shadowing nvidia-smi with a
failing stub and stripping conda from PATH — a check that cannot fail is the defect this
repository keeps finding.

Two defects in this script were found by its own tests: an unchecked `mkdir -p "$OUT"` that
let it run on into a broken tee off a pod, and a lock message that reported "another
delivery is in progress" for a directory that simply was not writable.

What these guards do NOT cover, stated rather than implied: nothing here runs a forecast,
publishes, or reaches the FAO. The first real exercise is a rehearsal on a pod.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…tcome

Phases 4 and 5 document Track A (calibration predictions for research). The FAO delivery
had no operator documentation at all, because until now it had no script.

Leads with --preflight, because the two things most likely to stop the delivery are
invisible until the end: the three Appwrite publish secrets, and conda — the postprocessor
launcher requires it while the pod builds a uv venv, so a pod can satisfy the training leg
and not the delivery leg, and that failure lands after ~16 GPU hours.

Records two readings an operator would otherwise get wrong: DeliveryNotFindableError is
1.4.0 WORKING (it verifies by refusing, and names what it checked), and a rehearsal's
forecasts genuinely reach the FAO shelf and are servable rather than being contained.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
… just the pod install

Found re-running the clean-checkout simulation after the delivery script landed: I floored
views-datafactory at 1.13.0 in tools/podrun/pod_run_model.sh and left both postprocessor
requirement files declaring >=1.9.0. Half a fix, and the half I left out is the leg that
runs LAST — so it would have been the one still holding a client without the credential
fixes at the point the delivery reaches a partner.

The reasoning that applied to the pod install applies here unchanged: before 1.13.0 the
datafactory client could carry a netrc credential across a redirect to another host and
embed it in error messages. The postprocessor runs on the same rented hardware and fetches
the actuals it curates land -> land_gaul against, so it holds that credential too. Nothing
about "this is the postprocessor" makes that safer.

Declared identically in both files for the same C-116 reason as the pins above them: one
shared envs/views-postprocessing prefix, so whichever postprocessor runs last decides the
version for both, and flooring one is the same as flooring neither.

The guard is widened rather than duplicated: it now iterates the pod install and both
requirement files, and was verified to fail independently for each of the three when that
source alone is reverted to >=1.9.0. A guard that passes because two of three sources are
correct is how this got shipped half-done in the first place.

Resolve unchanged: views-datafactory 1.13.0, xarray 2024.3.0, pandas 1.5.3, numpy 1.26.4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…for — again

Second time in one session, same mechanism, and it is worth recording rather than quietly
amending: the previous commit shipped the widened guard and not the fix, because I verified
the guard by reverting each source with `git checkout -- <file>` while the fix was still
UNCOMMITTED. The checkout reverted to the last commit and took the fix with it.

My own ship-it procedure says mutation-verification runs AFTER the commit and before the
push, for exactly this reason — "the tree is clean, so a mutation can be applied and
reverted with git checkout --". I violated it twice in a row, once for the pod install line
and once for both requirement files. The order is the control, not the care.

Both files now declare views-datafactory>=1.13.0, with the reasoning restored: before
1.13.0 the client could carry a netrc credential across a redirect to another host and embed
it in error messages, and the postprocessor runs on the same rented hardware holding the
same credential. It is also the LAST leg — the one that reaches the partner.

The guard is hardened while I am here: it now strips comments from the requirements files
too, not only from the shell script. un_fao's xarray note quotes the old
`views-datafactory>=1.9.0` line as the state it is explaining, so a guard reading raw text
either trips on prose or has to be written carefully around it — and being careful around
comments is precisely how the decorative guards got shipped in the first place.

Verification order corrected: committed first, mutation replay follows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…e the hole it opens

The repo's own hygiene guard caught the last commit: flooring views-datafactory to >=1.13.0
in the two postprocessor files while 34 model requirements stay at >=1.9.0 makes one package
declared two ways, which C-116 says is then decided by run order rather than intent. A good
guard, and it fired on a real inconsistency I introduced.

Not unified, deliberately, and the reasons are asymmetric:

  - the 34 are model requirements unrelated to the delivery legs. Raising them is #509's own
    scope, not a delivery PR's, and it would put 34 unreviewed one-line edits in a branch
    about reproducing a pod.
  - one of them is `models/violet_visitor/requirements.txt`. Another session owns that
    model's contents and this session is instructed not to edit them. That instruction is not
    mine to set aside because it would be convenient for a guard.

So: DEFERRED_PACKAGES, which the guard offers explicitly, with the trigger named as the guard
also requires — #509 raising the remaining 34, at which point the entry is DELETED rather than
amended, because the divergence it describes will not exist.

The divergence is also less dangerous than the rule's general case: the two postprocessors
share one prefix and agree with each other, the models resolve elsewhere, and the resolver
picks 1.13.0 for all of them today regardless. The floor only forbids something lower.

AND THE PART WORTH THE EXTRA TEST. A DEFERRED_PACKAGES entry is wider than it looks: it also
exempts the package from test_no_dependency_is_declared_without_an_upper_bound. So while this
deferral stands, nothing in the repo would notice `<2.0.0` being dropped from a
views-datafactory line — and that rule exists because an unbounded internal package installs
the next breaking major on the following monthly run. Buying a floor by silently selling a
ceiling is not a trade worth making quietly.

Closed with test_every_datafactory_declaration_keeps_its_upper_bound, which is narrower than
the rule it stands in for — it says nothing about floors, only that a ceiling exists wherever
this package is named — and which outlives the deferral harmlessly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ree on the workspace

Two findings from reviewing this branch before merging it. Both are cases of a rule the
scripts already state, not applied where it also holds.

STALE CALIBRATION ARTEFACTS. The forecast leg writes no parquets, so it skipped sections 4
and 5 — but a previous calibration run on the same pod leaves parquet/ and draws/ in the SAME
output directory. The forecast MANIFEST would then sit beside 13 parquets from a different run
type, and the guide's rsync copies the directory, so they come home as this run's output.
pod_run_model.sh already reasons about exactly this twice ("Clear a previous attempt first …
if they happened to sum to 13 the count check would pass while the manifest covered two
different training runs") for the case where the current run DOES produce them. The case where
it produces none is the same hazard, and skipping a stage is not the same as clearing it.

THE TWO SCRIPTS MUST AGREE ON $ROOT. pod_run_fao_delivery.sh honours PODRUN_ROOT so its
preflight can be exercised off a pod; pod_run_model.sh had /workspace hardcoded. The delivery
script reads $ROOT/deliver/<model>/STATUS to decide whether a model may be pooled, so a
relocated workspace in one and not the other means reading a STATUS the other never wrote —
and a missing STATUS is indistinguishable from a model that failed, which would refuse a
healthy roster. Benign on a pod, where both are /workspace; latent, and cheap to remove.

Both guarded. Written before the merge rather than after, which is the only time reviewing
your own diff is worth anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
@Polichinel
Polichinel merged commit 597070d into development Sep 29, 2026
6 checks passed
@Polichinel
Polichinel deleted the fix/516-517-fresh-pod-reproducibility branch September 29, 2026 22:27
Polichinel added a commit that referenced this pull request Sep 29, 2026
…fig instead of counting

Found by the /falsify pass on the forecasting claim, acting on a suggestion from the
views-pipeline-core session. My preflight from #524 checked three environment variables and
concluded the publish would work. Measured against the real code, PredictionStoreConfig
requires NINE:

  APPWRITE_ENDPOINT, APPWRITE_DATASTORE_PROJECT_ID, APPWRITE_DATASTORE_API_KEY,
  APPWRITE_PROD_FORECASTS_BUCKET_ID, APPWRITE_PROD_FORECASTS_BUCKET_NAME,
  APPWRITE_PROD_FORECASTS_COLLECTION_ID, APPWRITE_PROD_FORECASTS_COLLECTION_NAME,
  APPWRITE_METADATA_DATABASE_ID, APPWRITE_METADATA_DATABASE_NAME

Only the first three are secrets. A missing identifier fails the publish exactly as hard as a
missing secret, so the check would have reported ready with six of nine absent — and the run
would have failed at the publish, after the entire roster trained. The precise failure this
mode exists to prevent, in the mode that exists to prevent it. Twice in one night now, which
is its own signal.

Adding the six missing names would not fix it. The extra can be absent, the endpoint
unreachable, the key expired (#359: 2026-11-17). Preflight now CONSTRUCTS the config and
imports the SDK, which answers the actual question instead of a proxy for it.
views-pipeline-core suggested this and observed it is what #557 argues the manager should do
itself; until it does, doing it by hand costs seconds.

Second finding from the same measurement, and an operator trap: pipeline-core no longer
auto-loads a .env from the working directory (#346, register C-177) — a library reading
whatever .env its caller is standing in is what the Appwrite seam contract §3 forbids. So a
correct /root/.secrets sitting beside the operator is NOT enough; the variables must be
exported. Without that line the failure presents as wrong credentials rather than unexported
ones. Preflight now prints `set -a; . /root/.secrets; set +a`.

Both guarded, and the guards assert the construction and the exported-variables note exist
rather than re-listing nine names — a guard that enumerates the same thing the code does is
the defect that produced this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Polichinel added a commit that referenced this pull request Sep 29, 2026
…loor was below the one it delegates to (#527)

* fix(#499): the FAO preflight's disk floor was below the floor it delegates to

Found by a /falsify pass on "we are ready for a 40-lesson forecasting run reaching faoapi",
and it is my defect from #524.

pod_run_fao_delivery.sh demanded 60GB for EIGHT models plus the pooled ensemble, while
pod_run_model.sh — which it delegates to eight times — refuses below 40GB for ONE model
("one model needs ~20GB"). So the orchestrator's preflight could report ready and the
delegated script could then refuse partway through the roster, after hours of paid GPU time.
That is precisely the failure --preflight exists to prevent, introduced into --preflight.

Raised to DISK_FLOOR_GB=80: 40GB transient, which the delegated script enforces per model
regardless of what is written here, plus eight models' retained output and the pool.

The retained component is an ESTIMATE and the comment now says so rather than implying a
measurement. One model's calibration output is ~2.5GB per predictions directory on this
machine, but NO FORECASTING RUN HAS EVER COMPLETED ON THIS ROSTER — rusty_bucket's
forecasting_log.txt is from 2026-07-20 and lists temporary_crane and temporary_fox at
Deployment Status: shadow, a different roster entirely — so the retained size of a forecast
is unmeasured. A forecast has one origin against calibration's thirteen so it should be
smaller, and "should be" is doing real work in that sentence. Trigger for revising the
number: the first completed run, rather than reasoning about it a second time.

Guarded by a test that compares this floor against the delegated script's, so the two cannot
drift apart again — the defect was not the number, it was that nothing related them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#499): preflight checked 3 of 9 publish variables — build the config instead of counting

Found by the /falsify pass on the forecasting claim, acting on a suggestion from the
views-pipeline-core session. My preflight from #524 checked three environment variables and
concluded the publish would work. Measured against the real code, PredictionStoreConfig
requires NINE:

  APPWRITE_ENDPOINT, APPWRITE_DATASTORE_PROJECT_ID, APPWRITE_DATASTORE_API_KEY,
  APPWRITE_PROD_FORECASTS_BUCKET_ID, APPWRITE_PROD_FORECASTS_BUCKET_NAME,
  APPWRITE_PROD_FORECASTS_COLLECTION_ID, APPWRITE_PROD_FORECASTS_COLLECTION_NAME,
  APPWRITE_METADATA_DATABASE_ID, APPWRITE_METADATA_DATABASE_NAME

Only the first three are secrets. A missing identifier fails the publish exactly as hard as a
missing secret, so the check would have reported ready with six of nine absent — and the run
would have failed at the publish, after the entire roster trained. The precise failure this
mode exists to prevent, in the mode that exists to prevent it. Twice in one night now, which
is its own signal.

Adding the six missing names would not fix it. The extra can be absent, the endpoint
unreachable, the key expired (#359: 2026-11-17). Preflight now CONSTRUCTS the config and
imports the SDK, which answers the actual question instead of a proxy for it.
views-pipeline-core suggested this and observed it is what #557 argues the manager should do
itself; until it does, doing it by hand costs seconds.

Second finding from the same measurement, and an operator trap: pipeline-core no longer
auto-loads a .env from the working directory (#346, register C-177) — a library reading
whatever .env its caller is standing in is what the Appwrite seam contract §3 forbids. So a
correct /root/.secrets sitting beside the operator is NOT enough; the variables must be
exported. Without that line the failure presents as wrong credentials rather than unexported
ones. Preflight now prints `set -a; . /root/.secrets; set +a`.

Both guarded, and the guards assert the construction and the exported-variables note exist
rather than re-listing nine names — a guard that enumerates the same thing the code does is
the defect that produced this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant