Repository navigation
fix(#516,#517,#509,#518,#523): a fresh pod can reproduce the run that delivered — and a cheap rehearsal that cannot be mistaken for one - #524
Merged
Conversation
…nd WRONG #516 said a fresh un_fao environment is unbuildable. Measured on 2026-09-29, it is not: it builds, on a different pandas MAJOR, in silence. declared: views-datafactory>=1.9.0,<2.0.0 + numpy>=1.26.4,<2.0.0 resolved: numpy 1.26.4, xarray 2025.12.0, pandas 3.0.6 the pod that produced the first FAO delivery: xarray 2024.3.0, pandas 1.5.3 Neither xarray nor pandas was pinned anywhere. The working versions existed only as hand-pins on a pod that has since been destroyed, so a fresh pod silently picks different ones. A build failure is loud; this is not, which makes it the more dangerous of the two shapes the issue conflated. views-datafactory requires xarray outright and pandas only in an optional extra that is not installed, so pandas arrives THROUGH xarray. xarray is the sole carrier — datafactory's own pyproject comment says so — and pinning the carrier fixes it. The bound is measured, not guessed. xarray 2024.3.0 is the LAST release that accepts pandas 1.x; 2024.5.0 moved to pandas>=2.0 and 2024.9.0 to pandas>=2.1. A tidy-looking `<2025` cap would therefore have been WRONG: the cliff is inside the 2024 line, not at the year boundary. Verified by resolving the file, not by reading version numbers. Declared identically in both postprocessors because they share one prefix (C-116); pinning only one lets whichever runs last decide pandas for both, which is the same trap the existing numpy pin already documents one line above. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…d --rehearsal Three findings from one /falsify audit of "we are ready for a new 40-lesson run" (FALSIFIED). All three had one cause: every fix that made the 2026-09-29 run work was applied BY HAND to the pod, and the pod was destroyed. The pod was the artefact and nothing in git described it. #517 — the publish path was never installed. The string "appwrite" appeared nowhere in pod_run_model.sh, so `_build_datastore` raised at publish time, AFTER the full training run. That is exactly how the first FAO delivery attempt failed. The extra is now requested explicitly with no version, so views-hydranet's own range still decides which pipeline-core is installed, and the client is IMPORTED in preflight — the whole script is written to fail early, and the one dependency that failed late was the one not checked there. Verified by a real resolve: appwrite 13.6.1 alongside views-hydranet 0.1.2 and views-pipeline-core 3.3.4, with pandas 1.5.3 and numpy 1.26.4 intact. The same resolve confirms the C-151 toolz override is necessary rather than folklore — without it the resolver lands on toolz 0.11.2. #523 — a deliberate cheap run was reachable only by deleting the guard against an accidental one. The >=300 floor stays; `--rehearsal <lessons>` is the escape hatch. It takes the count on the command line, because main.py has no hyperparameter override and a flag without a count still forces an edit to a tracked config. It patches the POD's clone only, then re-imports to CONFIRM the patch took — a substitution that silently missed would give a 300-lesson run wearing a rehearsal label, or the reverse. The marking matters more than the permission. A rehearsal's parquets are structurally identical to a production run's and pass every check including the runner's own, so a REHEARSAL file is written beside them and MANIFEST gains `mode:` and `config: PATCHED after checkout` — the git sha alone no longer describes the run. Found while reviewing this change, not in the audit: section 2 does not re-clone when .git exists, so a production run on a pod that had rehearsed would read the LEFTOVER patch. The floor still refused it, but blamed the committed config and advised --rehearsal — a correct refusal for the wrong reason, sending the operator to edit the wrong file. It now detects the dirty config and names it. Deferred, and it is the weaker half: this runner MARKS a rehearsal but cannot REFUSE to publish one, because the publish step is not in this script. That guard belongs in views-pipeline-core. An xfail(strict) stub for it was written and DELETED — its regex spanned the whole file under re.S and XPASSed by accident, i.e. the exact class of guard-that-cannot-fire this audit exists to find. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
… the missing warning An independent /falsify guard-mode audit of 53bfba9 reverted EVERY fix in that commit, kept the comments, and got 22/22 green. 20 of 24 mutations survived. 10 of 14 guards were DECORATIVE, 4 WEAK, none held. The mechanism is the lesson. Every assertion read pod_run_model.sh as TEXT, and that script is unusually well commented — each fix carries a paragraph naming the incident and quoting the exact strings. So THE BETTER THE COMMENT, THE WEAKER THE GUARD: deleting the code left the comment, and the comment satisfied the assertion. Deleting the appwrite extra from the install line was invisible because the comment above it names views-pipeline-core[appwrite]. Worst survivor: `os.environ.get("REHEARSAL_LESSONS") or ""` -> `or "40"`. One word. Every production run then patches itself to 40 lessons, the floor is dead, the leftover-patch detector is dead, and because the SHELL variable stays empty the MANIFEST says `mode: production` and no REHEARSAL marker is written. A 40-lesson model labelled a production delivery, suite fully green. AND ONE GUARD CERTIFIED A SAFETY PROPERTY THAT DID NOT EXIST. It claimed the RunPod guide documents the /workspace chmod trap. The guide did not: no "world-readable", no "network filesystem", no C-154, no #518. It passed because `/workspace` is on line 104 and `chmod` on line 180, joined by `.*` under re.S — the identical defect the previous commit message boasted of having found and deleted elsewhere. That is worse than no guard, because it stopped the next reader looking. The warning is now written (#518, C-154): /workspace is a network filesystem where chmod 600 returns success and does nothing, leaving a credential at mode 666 with no error to notice. What changed in the tests: - the config-check program is EXTRACTED FROM ITS HEREDOC AND EXECUTED against fixture models in real git repos. That block holds all the rehearsal/production logic and was wholly unguarded. Includes an adversarial fixture whose literal is patchable but whose get_hp_config() returns 300 regardless, so a "verification" that compared target against itself is caught. - remaining text assertions read COMMENT-STRIPPED source, so a comment can never stand in for code. - requirements guards assert SPECIFIER SEMANTICS via packaging.SpecifierSet — is 1.5.3 admitted, is 3.0.6 refused — instead of pin presence. `pandas>=3.0` passed the old one, and marker-gated pins (`; python_version < "3.10"`, inert on 3.11) defeated the identical-pins check while keeping the captured text identical. - the roster guard CALLS get_hp_config() instead of matching the literal, which is what pod_run_model.sh itself insists on, citing #501 "the guard that was not one". pod_run_model.sh warned against this exact failure in three places and the guards did it anyway. That is the argument for the independence rule: the author's own model named the trap three times in one commit and stepped into it. 35 tests, up from 22. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
… parsing Two gaps found by replaying the guard-mode audit's mutations against the rewritten guards — the replay is the point, since the previous version passed 22/22 with every fix reverted. #509 — the pod installed views-datafactory>=1.9.0. The credential-handling fixes landed in 1.13.0: before it the client could carry a netrc credential across a redirect to another host and embed it in error messages. This script only ever runs on hardware we do not own, carrying exactly that credential, and the guide already prescribes >=1.13.0 explicitly while the script did not. A resolver picks 1.13.0 anyway today, but "the resolver will probably do the right thing" is precisely the reasoning that put xarray 2025.12.0 and pandas 3.0.6 into a fresh environment (#516). Verified: the floor still resolves to 1.13.0 with hydranet 0.1.2, pipeline-core 3.3.4, appwrite 13.6.1, pandas 1.5.3. The shell argument parser was unguarded. Executing the config-check program does not cover it — the flag never reaches Python if the shell refuses it first. Deleting the `--rehearsal)` case arm removes H1's escape hatch entirely and left every guard green, because the string survives in the header comment and in USAGE. Now executed against the real script: the flag and its count must be CONSUMED (distinguished from "unknown option" by which error appears), a missing count must be refused, a non-integer must be refused, and — the control, without which the first assertion proves nothing — a genuinely unknown option must still be rejected. Argument parsing precedes `mkdir -p "$OUT"`, so these cases exit before touching the filesystem. MUTATION REPLAY — every surviving mutation from the audit is now caught: X7 `or ""` -> `or "40"` (production runs patch themselves) 3 failed X8 --rehearsal becomes a bare flag 3 failed X22 the --rehearsal case arm deleted 3 failed X1 extra dropped from the install line, comment kept 1 failed X2 both appwrite imports commented out, prose comment kept 1 failed X5 2026-09-29 install order restored + reassuring comment 1 failed X6 the floor prints instead of exiting 1 failed X9 leftover-patch detector flipped to the wrong branch 1 failed X10 patch verification becomes `if target != target` 1 failed X11 the marker-writing block deleted 1 failed X12 stale-marker clear moved to the end of the run 1 failed X13 MANIFEST mode branches swapped 1 failed X14 both files pinned to xarray 2025.12 / pandas 3.0 8 failed X15 upper bounds dropped 8 failed X17 pins marker-gated inert on 3.11 6 failed X18 get_hp_config() returns 40, the 300 literal untouched 1 failed X20b toolz override replaced by a comment 1 failed control: the guide's chmod warning deleted 1 failed X16 (`xarray == 2024.3.0`, a stricter and CORRECT pin) previously failed with a message claiming no pin existed; it now passes, so the guard no longer raises a false alarm against a legal PEP 508 spelling. Two honest notes. X18 was recorded SURVIVED in the audit but had not applied — the config returns a dict literal, so the mutation's sed matched nothing. Re-applied correctly, it is caught. And X19 (replacing the credential chmod with an unrelated one) now passes, because the /workspace warning it was meant to defeat genuinely exists; the control above proves the guard fires when that warning is removed. 40 tests, up from 35. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ested for The previous commit shipped the guard and not the fix. The sequence: edit the install line, write the test, verify the test fires by reverting the line with sed, then `git checkout -- tools/podrun/pod_run_model.sh` to undo the sed — which reverted to the last COMMIT, discarding the uncommitted 1.13.0 edit along with the sed. The commit then captured the test plus a 1.9.0 install line. Caught by that test on the next full-suite run, which is the whole argument for writing guards that execute: a text-matching guard for this would have been satisfied by the comment above the line, which names 1.13.0 and explains why. Also worth recording, because it made the failure hard to see for two runs: this suite runs pytest-randomly, so ordering varies between runs. Two earlier runs reported 8058 passed/209 skipped and 7758 passed/470 skipped/3 failed with an unchanged tree, and the three failures were ordering interference in pre-existing tests, not in anything here. With `-p no:randomly` the counts are stable at 8076 passed/209 skipped/5 xfailed and the only failure was this one. Verify against a fixed order before concluding a change is responsible for a red suite. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ds typed on a pod The fourth instance of the pattern the /falsify audit found three of, and the one that mattered most. On 2026-09-29 the whole FAO delivery — eight forecasting runs, the rusty_bucket pool, the publish, the un_fao postprocessor — was typed by hand on a rented pod. It worked, the pod was destroyed, and nothing recorded what had been done. The audit missed it because it audited the script that exists (pod_run_model.sh) instead of asking what the delivery needs. pod_run_model.sh runs `-r calibration -t -e` and its header is accurate that it uploads nothing: it is Track A, the research frames. Nothing in tools/ drove Track B at all — `grep -rln "prediction_store\|r forecasting" tools/` returns one unrelated file. This is not a stopgap. fimbulthul, which #499 names for every Track B step, has been unreachable since 2026-09-25, so a pod is the only hardware there is. pod_run_model.sh gains `--forecast`: `-r forecasting -t -f`, and it SKIPS sections 4 and 5. Those are calibration deliverables — a forecast has one origin, so the 13-parquet count check would refuse it, and what the FAO chain consumes is the pooled ensemble output. Kept in this script rather than a copy because everything before section 3 is identical and already exercised; a second copy would be a second place for C-151 to rot. pod_run_fao_delivery.sh drives the chain. Two details that are not obvious and that cost real time to establish: - `-sa/--saved` on the ensemble is REQUIRED. Without it rusty_bucket refetches instead of pooling the member forecasts just written, and the eight runs are wasted. - the two legs need DIFFERENT INTERPRETERS. This pod builds a uv venv; tools/launcher/postprocessor.sh uses `conda shell.bash hook` / `conda create --prefix`. A pod can satisfy one and not the other, and step 4 is last — so a missing conda kills the delivery after every GPU hour is spent. Preflight now checks for it. `--preflight` is the mode that protects the money. At 300 lessons step 1 alone is ~16 GPU hours, and everything it checks is knowable in seconds: checkout, venv, netrc, the three publish secrets, conda, that un_fao's REGION resolves to land_gaul rather than a disarmed value (C-110), GPU, disk. It accumulates and reports every problem rather than dying on the first, because a round trip per problem is billed by the second. The publish secrets are checked HERE and not at first publish because views-pipeline-core's PredictionStoreConfig claims to read them "once at startup and fail loud … preventing silent failures after hours of training" and does not — it is called from _build_datastore, after training (views-pipeline-core#557). Until that moves, this preflight is the only check that happens before the money is spent. A rehearsal is marked at every level, and the marker says the uncomfortable part out loud: nothing downstream refuses a marked rehearsal (#523), so its undertrained forecasts DO reach the FAO shelf and ARE servable. That is deliberate — it is the only way to test the chain — and it is why the final stage reads back what landed with `tools.liveness`, by name, rather than inferring success from an exit code. On 2026-09-29 the publish reported success having written nothing (C-155). 19 guards, written after the lesson that 10 of 14 of the previous set were decorative. The parser and `--preflight` are EXECUTED; the remaining text assertions read comment-stripped source and cover only ordering facts that cannot be executed without a GPU, a store and a partner bucket. PODRUN_ROOT exists so preflight is exercisable off a pod; it relocates the workspace and relaxes no check. Every preflight check was proven able to fire, including by shadowing nvidia-smi with a failing stub and stripping conda from PATH — a check that cannot fail is the defect this repository keeps finding. Two defects in this script were found by its own tests: an unchecked `mkdir -p "$OUT"` that let it run on into a broken tee off a pod, and a lock message that reported "another delivery is in progress" for a directory that simply was not writable. What these guards do NOT cover, stated rather than implied: nothing here runs a forecast, publishes, or reaches the FAO. The first real exercise is a rehearsal on a pod. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…tcome Phases 4 and 5 document Track A (calibration predictions for research). The FAO delivery had no operator documentation at all, because until now it had no script. Leads with --preflight, because the two things most likely to stop the delivery are invisible until the end: the three Appwrite publish secrets, and conda — the postprocessor launcher requires it while the pod builds a uv venv, so a pod can satisfy the training leg and not the delivery leg, and that failure lands after ~16 GPU hours. Records two readings an operator would otherwise get wrong: DeliveryNotFindableError is 1.4.0 WORKING (it verifies by refusing, and names what it checked), and a rehearsal's forecasts genuinely reach the FAO shelf and are servable rather than being contained. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
… just the pod install Found re-running the clean-checkout simulation after the delivery script landed: I floored views-datafactory at 1.13.0 in tools/podrun/pod_run_model.sh and left both postprocessor requirement files declaring >=1.9.0. Half a fix, and the half I left out is the leg that runs LAST — so it would have been the one still holding a client without the credential fixes at the point the delivery reaches a partner. The reasoning that applied to the pod install applies here unchanged: before 1.13.0 the datafactory client could carry a netrc credential across a redirect to another host and embed it in error messages. The postprocessor runs on the same rented hardware and fetches the actuals it curates land -> land_gaul against, so it holds that credential too. Nothing about "this is the postprocessor" makes that safer. Declared identically in both files for the same C-116 reason as the pins above them: one shared envs/views-postprocessing prefix, so whichever postprocessor runs last decides the version for both, and flooring one is the same as flooring neither. The guard is widened rather than duplicated: it now iterates the pod install and both requirement files, and was verified to fail independently for each of the three when that source alone is reverted to >=1.9.0. A guard that passes because two of three sources are correct is how this got shipped half-done in the first place. Resolve unchanged: views-datafactory 1.13.0, xarray 2024.3.0, pandas 1.5.3, numpy 1.26.4. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…for — again Second time in one session, same mechanism, and it is worth recording rather than quietly amending: the previous commit shipped the widened guard and not the fix, because I verified the guard by reverting each source with `git checkout -- <file>` while the fix was still UNCOMMITTED. The checkout reverted to the last commit and took the fix with it. My own ship-it procedure says mutation-verification runs AFTER the commit and before the push, for exactly this reason — "the tree is clean, so a mutation can be applied and reverted with git checkout --". I violated it twice in a row, once for the pod install line and once for both requirement files. The order is the control, not the care. Both files now declare views-datafactory>=1.13.0, with the reasoning restored: before 1.13.0 the client could carry a netrc credential across a redirect to another host and embed it in error messages, and the postprocessor runs on the same rented hardware holding the same credential. It is also the LAST leg — the one that reaches the partner. The guard is hardened while I am here: it now strips comments from the requirements files too, not only from the shell script. un_fao's xarray note quotes the old `views-datafactory>=1.9.0` line as the state it is explaining, so a guard reading raw text either trips on prose or has to be written carefully around it — and being careful around comments is precisely how the decorative guards got shipped in the first place. Verification order corrected: committed first, mutation replay follows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…e the hole it opens The repo's own hygiene guard caught the last commit: flooring views-datafactory to >=1.13.0 in the two postprocessor files while 34 model requirements stay at >=1.9.0 makes one package declared two ways, which C-116 says is then decided by run order rather than intent. A good guard, and it fired on a real inconsistency I introduced. Not unified, deliberately, and the reasons are asymmetric: - the 34 are model requirements unrelated to the delivery legs. Raising them is #509's own scope, not a delivery PR's, and it would put 34 unreviewed one-line edits in a branch about reproducing a pod. - one of them is `models/violet_visitor/requirements.txt`. Another session owns that model's contents and this session is instructed not to edit them. That instruction is not mine to set aside because it would be convenient for a guard. So: DEFERRED_PACKAGES, which the guard offers explicitly, with the trigger named as the guard also requires — #509 raising the remaining 34, at which point the entry is DELETED rather than amended, because the divergence it describes will not exist. The divergence is also less dangerous than the rule's general case: the two postprocessors share one prefix and agree with each other, the models resolve elsewhere, and the resolver picks 1.13.0 for all of them today regardless. The floor only forbids something lower. AND THE PART WORTH THE EXTRA TEST. A DEFERRED_PACKAGES entry is wider than it looks: it also exempts the package from test_no_dependency_is_declared_without_an_upper_bound. So while this deferral stands, nothing in the repo would notice `<2.0.0` being dropped from a views-datafactory line — and that rule exists because an unbounded internal package installs the next breaking major on the following monthly run. Buying a floor by silently selling a ceiling is not a trade worth making quietly. Closed with test_every_datafactory_declaration_keeps_its_upper_bound, which is narrower than the rule it stands in for — it says nothing about floors, only that a ceiling exists wherever this package is named — and which outlives the deferral harmlessly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…ree on the workspace
Two findings from reviewing this branch before merging it. Both are cases of a rule the
scripts already state, not applied where it also holds.
STALE CALIBRATION ARTEFACTS. The forecast leg writes no parquets, so it skipped sections 4
and 5 — but a previous calibration run on the same pod leaves parquet/ and draws/ in the SAME
output directory. The forecast MANIFEST would then sit beside 13 parquets from a different run
type, and the guide's rsync copies the directory, so they come home as this run's output.
pod_run_model.sh already reasons about exactly this twice ("Clear a previous attempt first …
if they happened to sum to 13 the count check would pass while the manifest covered two
different training runs") for the case where the current run DOES produce them. The case where
it produces none is the same hazard, and skipping a stage is not the same as clearing it.
THE TWO SCRIPTS MUST AGREE ON $ROOT. pod_run_fao_delivery.sh honours PODRUN_ROOT so its
preflight can be exercised off a pod; pod_run_model.sh had /workspace hardcoded. The delivery
script reads $ROOT/deliver/<model>/STATUS to decide whether a model may be pooled, so a
relocated workspace in one and not the other means reading a STATUS the other never wrote —
and a missing STATUS is indistinguishable from a model that failed, which would refuse a
healthy roster. Benign on a pod, where both are /workspace; latent, and cheap to remove.
Both guarded. Written before the merge rather than after, which is the only time reviewing
your own diff is worth anything.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Polichinel
added a commit
that referenced
this pull request
Sep 29, 2026
…fig instead of counting Found by the /falsify pass on the forecasting claim, acting on a suggestion from the views-pipeline-core session. My preflight from #524 checked three environment variables and concluded the publish would work. Measured against the real code, PredictionStoreConfig requires NINE: APPWRITE_ENDPOINT, APPWRITE_DATASTORE_PROJECT_ID, APPWRITE_DATASTORE_API_KEY, APPWRITE_PROD_FORECASTS_BUCKET_ID, APPWRITE_PROD_FORECASTS_BUCKET_NAME, APPWRITE_PROD_FORECASTS_COLLECTION_ID, APPWRITE_PROD_FORECASTS_COLLECTION_NAME, APPWRITE_METADATA_DATABASE_ID, APPWRITE_METADATA_DATABASE_NAME Only the first three are secrets. A missing identifier fails the publish exactly as hard as a missing secret, so the check would have reported ready with six of nine absent — and the run would have failed at the publish, after the entire roster trained. The precise failure this mode exists to prevent, in the mode that exists to prevent it. Twice in one night now, which is its own signal. Adding the six missing names would not fix it. The extra can be absent, the endpoint unreachable, the key expired (#359: 2026-11-17). Preflight now CONSTRUCTS the config and imports the SDK, which answers the actual question instead of a proxy for it. views-pipeline-core suggested this and observed it is what #557 argues the manager should do itself; until it does, doing it by hand costs seconds. Second finding from the same measurement, and an operator trap: pipeline-core no longer auto-loads a .env from the working directory (#346, register C-177) — a library reading whatever .env its caller is standing in is what the Appwrite seam contract §3 forbids. So a correct /root/.secrets sitting beside the operator is NOT enough; the variables must be exported. Without that line the failure presents as wrong credentials rather than unexported ones. Preflight now prints `set -a; . /root/.secrets; set +a`. Both guarded, and the guards assert the construction and the exported-variables note exist rather than re-listing nine names — a guard that enumerates the same thing the code does is the defect that produced this commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Polichinel
added a commit
that referenced
this pull request
Sep 29, 2026
…loor was below the one it delegates to (#527) * fix(#499): the FAO preflight's disk floor was below the floor it delegates to Found by a /falsify pass on "we are ready for a 40-lesson forecasting run reaching faoapi", and it is my defect from #524. pod_run_fao_delivery.sh demanded 60GB for EIGHT models plus the pooled ensemble, while pod_run_model.sh — which it delegates to eight times — refuses below 40GB for ONE model ("one model needs ~20GB"). So the orchestrator's preflight could report ready and the delegated script could then refuse partway through the roster, after hours of paid GPU time. That is precisely the failure --preflight exists to prevent, introduced into --preflight. Raised to DISK_FLOOR_GB=80: 40GB transient, which the delegated script enforces per model regardless of what is written here, plus eight models' retained output and the pool. The retained component is an ESTIMATE and the comment now says so rather than implying a measurement. One model's calibration output is ~2.5GB per predictions directory on this machine, but NO FORECASTING RUN HAS EVER COMPLETED ON THIS ROSTER — rusty_bucket's forecasting_log.txt is from 2026-07-20 and lists temporary_crane and temporary_fox at Deployment Status: shadow, a different roster entirely — so the retained size of a forecast is unmeasured. A forecast has one origin against calibration's thirteen so it should be smaller, and "should be" is doing real work in that sentence. Trigger for revising the number: the first completed run, rather than reasoning about it a second time. Guarded by a test that compares this floor against the delegated script's, so the two cannot drift apart again — the defect was not the number, it was that nothing related them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#499): preflight checked 3 of 9 publish variables — build the config instead of counting Found by the /falsify pass on the forecasting claim, acting on a suggestion from the views-pipeline-core session. My preflight from #524 checked three environment variables and concluded the publish would work. Measured against the real code, PredictionStoreConfig requires NINE: APPWRITE_ENDPOINT, APPWRITE_DATASTORE_PROJECT_ID, APPWRITE_DATASTORE_API_KEY, APPWRITE_PROD_FORECASTS_BUCKET_ID, APPWRITE_PROD_FORECASTS_BUCKET_NAME, APPWRITE_PROD_FORECASTS_COLLECTION_ID, APPWRITE_PROD_FORECASTS_COLLECTION_NAME, APPWRITE_METADATA_DATABASE_ID, APPWRITE_METADATA_DATABASE_NAME Only the first three are secrets. A missing identifier fails the publish exactly as hard as a missing secret, so the check would have reported ready with six of nine absent — and the run would have failed at the publish, after the entire roster trained. The precise failure this mode exists to prevent, in the mode that exists to prevent it. Twice in one night now, which is its own signal. Adding the six missing names would not fix it. The extra can be absent, the endpoint unreachable, the key expired (#359: 2026-11-17). Preflight now CONSTRUCTS the config and imports the SDK, which answers the actual question instead of a proxy for it. views-pipeline-core suggested this and observed it is what #557 argues the manager should do itself; until it does, doing it by hand costs seconds. Second finding from the same measurement, and an operator trap: pipeline-core no longer auto-loads a .env from the working directory (#346, register C-177) — a library reading whatever .env its caller is standing in is what the Appwrite seam contract §3 forbids. So a correct /root/.secrets sitting beside the operator is NOT enough; the variables must be exported. Without that line the failure presents as wrong credentials rather than unexported ones. Preflight now prints `set -a; . /root/.secrets; set +a`. Both guarded, and the guards assert the construction and the exported-variables note exist rather than re-listing nine names — a guard that enumerates the same thing the code does is the defect that produced this commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #516, #517, #509, #518. Implements #523 (its deferred half stays open).
A
/falsifyaudit of "we are ready to start a new full 40-lesson FAO delivery run — nothing remains but creating a pod" returned FALSIFIED: three hard falsifications, all with one cause.The cause
Every fix that made the 2026-09-29 run work was applied by hand to the pod, and the pod was destroyed. The repository never learned any of them. The pod was the artefact and nothing in git described it, so a fresh pod reproduces the original broken state.
That is the finding worth keeping. The three defects below are its instances.
H1 — the run could not be launched at all
pod_run_model.shrefusestotal_lessons < 300; all eight configs declare 300;main.pyhas no hyperparameter override. A cheap end-to-end test — the ordinary thing to want after a run that failed late — was reachable only by deleting the guard.The floor is right: it stops a full GPU budget being spent by accident. But it was written as though wasted money were the risk, and it is not the main one. The risk is an undertrained model reaching a partner as a real forecast, and on that axis the guard made things worse — the only way past it produced output indistinguishable from a production run.
--rehearsal <lessons>takes the count on the command line, patches only the pod's ephemeral clone, re-imports to confirm the patch took, and marks the output (REHEARSALfile;mode:andconfig: PATCHED after checkoutinMANIFEST). The production floor is untouched when the flag is absent.H2 — a fresh environment built successfully and WRONG (#516)
Measured, not reasoned about:
#516 called this env unbuildable. It is not — it builds, on a different pandas major, silently, which is the more dangerous of the two shapes.
views-datafactoryrequires xarray outright and pandas only in an optional extra that is not installed, so xarray is the sole carrier (its own pyproject says so) and pinning the carrier fixes it.The obvious cap would have been wrong.
xarray<2025looks right; xarray 2024.11.0 already requirespandas>=2.1. Measured boundary: 2024.3.0 is the last release accepting pandas 1.x, 2024.5.0 moved to>=2.0, 2024.9.0 to>=2.1. The cliff is inside the 2024 line.H3 — the pod could not publish (#517)
The string
appwriteappeared nowhere inpod_run_model.sh.views-pipeline-core[appwrite]was never installed, so_build_datastoreraised at publish — after the full training run. That is exactly how 2026-09-29 failed. Now installed, and imported in preflight: the whole script is written to fail early, and the one dependency that failed late was the one not checked there.Also fixed
views-datafactory>=1.9.0. The credential-handling fixes landed in 1.13.0: before it the client could carry a netrc credential across a redirect to another host and embed it in error messages. This script only ever runs on hardware we do not own, carrying exactly that credential, and the guide already prescribed>=1.13.0while the script did not./workspacesilently ignoreschmod. It is a network filesystem;chmod 600returns success and does nothing, leaving a credential at mode666with no error to notice. See below for why this was missing.The second audit, and the part of this PR that matters most
This branch's first set of tests was 22 green assertions that protected nothing. An independent
/falsifyguard-mode pass reverted every fix in the commit, kept the comments, and got 22/22 green. 20 of 24 mutations survived; 10 of 14 guards were DECORATIVE, 4 WEAK, none held.The mechanism is the lesson. Every assertion read the script as text, and
pod_run_model.shis unusually well commented — each fix carries a paragraph naming the incident and quoting the exact strings. So the better the comment, the weaker the guard: deleting the code left the comment, and the comment satisfied the assertion.Worst survivor:
os.environ.get("REHEARSAL_LESSONS") or ""→or "40". One word. Every production run then patches itself to 40 lessons, the floor is dead, the leftover-patch detector is dead, and because the shell variable stays empty theMANIFESTsaysmode: productionand no marker is written. A 40-lesson model labelled a production delivery, suite fully green.And one guard certified a safety property that did not exist. It claimed the guide documented the
/workspacechmod trap. The guide did not — no "world-readable", no "network filesystem", no C-154, no #518. It passed because/workspaceis on line 104 andchmodon line 180, joined by.*underre.S— the identical defect the same commit message boasted of having found and deleted elsewhere. That is worse than no guard, because it stopped the next reader looking. The warning is now written.pod_run_model.shwarns against exactly this in three places, citing #501 "the guard that was not one". The guards did it anyway. That is the argument for the independence rule: the author's own model named the trap three times in one commit and stepped into it.What the guards do now
get_hp_config()returns 300 regardless, so a verification comparingtargetagainst itself is caught.--rehearsal)case arm removes the escape hatch and left every text guard green.packaging.SpecifierSet— is 1.5.3 admitted, is 3.0.6 refused — not pin presence.pandas>=3.0passed the old guard, and marker-gated pins (; python_version < "3.10", inert on 3.11) defeated the identical-pins check while keeping the captured text identical.get_hp_config()instead of matching the literal.Mutation replay — all 18 now caught
X7production self-patch ·X8bare flag ·X22case arm deleted ·X1extra dropped ·X2imports commented out ·X52026-09-29 install order ·X6floor prints instead of exits ·X9detector on wrong branch ·X10tautological verification ·X11marker block deleted ·X12marker cleared at end ·X13MANIFEST branches swapped ·X14pinned to the broken versions ·X15upper bounds dropped ·X17marker-gated inert pins ·X18get_hp_config()returns 40 ·X20btoolz override as comment · control: guide warning deleted.X16(xarray == 2024.3.0— stricter and correct) previously failed claiming no pin existed; it now passes, so the guard no longer false-alarms on a legal PEP 508 spelling.Two honest notes: X18 was recorded SURVIVED by the audit but had not applied — the config returns a dict literal so the mutation's
sedmatched nothing; re-applied correctly, it is caught. And X19 now passes because the warning it was meant to defeat genuinely exists; the control proves the guard fires when that warning is removed.Verification
numpy 1.26.4, pandas 1.5.3, xarray 2024.3.0, views-datafactory 1.13.0. Pod training venv →appwrite 13.6.1, views-hydranet 0.1.2, views-pipeline-core 3.3.4, torch 2.14.0, pandas 1.5.3. Both match the pod that delivered.ruff check .clean. 8077 passed, 209 skipped, 5 xfailed.calibration_delivery_*directories are generated model output and stay excluded.A correction to an earlier claim in this description. Two full-suite runs on an unchanged tree reported
8058 passed/209 skippedand7758 passed/470 skipped/3 failed, and I attributed that topytest-randomly. That was wrong — no such plugin is installed (pytest 9.0.3, nothing matchingrandomin the environment), so the-p no:randomlyI added to subsequent runs was a no-op and proved nothing.The three failures (
test_falsification_catalog_enhancements,test_update_readme_survives_missing_data_client) passed in isolation, pass in a clean clone ofdevelopment, and have not recurred in any of the six later full-suite runs, the last of which was plainpytestwith no flags and returned the same8097 passed/209 skipped/5 xfailed. The 261 extra skips alongside 3 failures is the shape of a shared fixture erroring and its dependents skipping, which fits a transient condition during that one run — the guard-audit subagent was mutating and reverting files in this tree at the time. I am recording this as unexplained rather than explained, because a confident wrong diagnosis is what this PR is largely about.Deliberately not done
A rehearsal is MARKED, not REFUSED. The publish step is not in this runner — its header is accurate that it uploads nothing — so nothing stops a rehearsal's forecasts being published.
views-pipeline-corereviewed the shape and argued convincingly against detecting a rehearsal: that repo cannot know what a lesson is, and inferring undertraining would be guessing at a property it has no access to. The agreed design is the inverse — refuse to publish unless a run positively declares itself complete, absence of a declaration being a refusal — homed in the manifest's existingprovenanceblock and applied at both doors (sampled_forecast_publisher.py::publish_sampled_forecastandprediction/io.py::_upload_to_prediction_store).That is an ADR-013 contract change across four repos on a partner-visible manifest, so it is a maintainer decision, not a peer request. Recorded on #523, which stays open. Until it lands, do not publish a directory containing a
REHEARSALfile.Register entries for this audit are deferred behind a named trigger: #520 merging, since it already consumes C-154/C-155 and both branches would otherwise fight over the header counts.
🤖 Generated with Claude Code
https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K