You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
RunPod guide: everything the 2026-09-29/30 FAO delivery effort learned that the guide does not yet carry #526
A holding issue. docs/runpod_run_guide.md was written for the calibration campaign — eight HydraNets on global land, predictions rsynced home for the research team. The 2026-09-29/30 effort was a different thing: the FAO delivery leg, which the guide's own post-mortem correctly flagged as untested (reports/postmortem_runpod_first_deployment_2026-09.md:458 — "Nothing here has been tested on the forecasting partition, which is what the FAO delivery [needs]").
#524 added the two corrections that could not wait, because guards depend on them: Phase 4b (the delivery chain) and the /workspace chmod warning. Everything below is the rest — recorded here rather than written into the guide now, so the guide is edited deliberately and in one pass rather than accreting.
New post-mortems and reports for this effort will be written separately; this issue is only the guide's share.
1. The two tracks are not distinguished, and conflating them cost real confusion
The guide's Phases 4–6 are Track A: main.py -r calibration -t -e, 13-origin parquets, rsync home, teardown. The runner's header is accurate that it uploads nothing.
Track B — the FAO delivery — is a different chain and was never in the guide:
1. the 8 models main.py -r forecasting -t -f (~2h each at 300 lessons)
2. rusty_bucket main.py -r forecasting -f -sa pooled from the SAVED member forecasts
3. publish -p wire shards to production_forecasts (ADR-013)
4. un_fao postprocessors/un_fao/run.sh curates land (64,818) -> land_gaul (64,742)
5. faoapi ingest -> /data endpoints
A reader who follows Phase 4 expecting an FAO delivery gets calibration predictions and no upload. That is the single most valuable correction here.
2. -sa/--saved at step 2 is load-bearing
Without it rusty_bucketrefetches instead of pooling the member forecasts step 1 just wrote, and the eight runs are wasted. The guide should state this as a consequence, not a flag.
3. run.sh does not work on a pod
ensembles/rusty_bucket/run.sh (and every generated model run.sh) does eval "$(conda shell.bash hook)" and expects a conda prefix. A pod builds a uv venv, so the invocation is "$VENV/bin/python" main.py ... directly. Anyone following the runbook's bash run.sh -r forecasting ... on a pod will be confused, and reports/fao_delivery_runbook.md says exactly that.
4. The runbook says fimbulthul; that is no longer available
#499 Track B assigns B3, B4 and B8 to "You, on fimbulthul."fimbulthul has been unreachable since 2026-09-25. A pod is not a stopgap for this work — it is the only hardware. The guide should stop implying a choice exists.
5. What the guide should say about reading the outcome
Two readings an operator will otherwise get wrong, both learned the hard way:
DeliveryNotFindableError is views-postprocessing 1.4.0 WORKING. That build verifies a delivery by what it refuses, and the message names every object it checked. Read as a failure, it invites exactly the wrong remedy.
A re-publish is not a safe retry. Content-hash dedup returns the pre-existing file's id and reports success having written nothing (C-155, views-pipeline-core#551, fixed in 3.3.4). The documented remedy for an invisible delivery was the trigger for it. Worst case was 0 of 110 objects servable.
6. Credentials — beyond the chmod trap now in #524
Nine env vars, three of them real secrets (APPWRITE_ENDPOINT, APPWRITE_DATASTORE_PROJECT_ID, APPWRITE_DATASTORE_API_KEY); the rest are identifiers. Canonical home is views-faoapi/.env.
views-pipeline-core's PredictionStoreConfig docstring claims it reads these "once at startup … preventing silent failures after hours of training."It does not — it is called from _build_datastore, i.e. after training (views-pipeline-core#557). Until that moves, pod_run_fao_delivery.sh --preflight is the only early check, and the guide should say why running it is not optional.
7. Environment facts a fresh pod needs and the guide does not state
The obvious cap would have been wrong.xarray<2025 looks right; 2024.11.0 already requires pandas>=2.1. The cliff is inside the 2024 line: 2024.3.0 is the last release accepting pandas 1.x, 2024.5.0 moved to >=2.0, 2024.9.0 to >=2.1.
views-datafactory>=1.13.0, not >=1.9.0, wherever a leg runs on rented hardware: before 1.13.0 the client could carry a netrc credential across a redirect to another host and embed it in error messages (34 models pin views-datafactory>=1.9.0; the credential-handling fixes landed in 1.13.0 #509). Currently deferred for the 34 model declarations — see tests/test_requirements_hygiene.DEFERRED_PACKAGES.
WANDB_MODE=offline is not optional. Without it main.py calls wandb.login() and blocks on a prompt nobody is watching; the first forecasting run died there after the environment was already built.
--rehearsal <lessons> now exists on both runners. Two things the guide must be blunt about:
A rehearsal's forecasts reach the FAO shelf and are servable. Nothing downstream refuses a marked rehearsal (A deliberate cheap end-to-end run is only reachable by deleting the guard that forbids an accidental one #523), and they are structurally indistinguishable from real forecasts — same columns, same coverage, all finite and non-negative, every check green. /workspace/deliver/_fao/REHEARSAL says so. They must be superseded, not left.
A rehearsal proves the chain, not the capacity. 40 lessons has a different duration and memory profile from 300, and memory is where this platform has failed before.
9. Operational details worth a line each
setsid matters — already in the guide, and it earned its place.
Watch with cat /workspace/deliver/_fao/STAGE, or tail -f .../run.log.
The console shows ~0% GPU and looks idle; it genuinely is idle between sampling steps. Already noted for Track A, equally true here.
Do not take PRO 6000 MIG 24GB — 25× slower on real work. Already in the guide; worth keeping loud.
109 orphan objects from run 20260929_172325 are inert. If the guide grows a cleanup section, they belong in it.
A holding issue.
docs/runpod_run_guide.mdwas written for the calibration campaign — eight HydraNets on global land, predictions rsynced home for the research team. The 2026-09-29/30 effort was a different thing: the FAO delivery leg, which the guide's own post-mortem correctly flagged as untested (reports/postmortem_runpod_first_deployment_2026-09.md:458— "Nothing here has been tested on the forecasting partition, which is what the FAO delivery [needs]").#524 added the two corrections that could not wait, because guards depend on them: Phase 4b (the delivery chain) and the
/workspacechmod warning. Everything below is the rest — recorded here rather than written into the guide now, so the guide is edited deliberately and in one pass rather than accreting.New post-mortems and reports for this effort will be written separately; this issue is only the guide's share.
1. The two tracks are not distinguished, and conflating them cost real confusion
The guide's Phases 4–6 are Track A:
main.py -r calibration -t -e, 13-origin parquets, rsync home, teardown. The runner's header is accurate that it uploads nothing.Track B — the FAO delivery — is a different chain and was never in the guide:
A reader who follows Phase 4 expecting an FAO delivery gets calibration predictions and no upload. That is the single most valuable correction here.
2.
-sa/--savedat step 2 is load-bearingWithout it
rusty_bucketrefetches instead of pooling the member forecasts step 1 just wrote, and the eight runs are wasted. The guide should state this as a consequence, not a flag.3.
run.shdoes not work on a podensembles/rusty_bucket/run.sh(and every generated modelrun.sh) doeseval "$(conda shell.bash hook)"and expects a conda prefix. A pod builds a uv venv, so the invocation is"$VENV/bin/python" main.py ...directly. Anyone following the runbook'sbash run.sh -r forecasting ...on a pod will be confused, andreports/fao_delivery_runbook.mdsays exactly that.4. The runbook says
fimbulthul; that is no longer available#499 Track B assigns B3, B4 and B8 to "You, on fimbulthul." fimbulthul has been unreachable since 2026-09-25. A pod is not a stopgap for this work — it is the only hardware. The guide should stop implying a choice exists.
5. What the guide should say about reading the outcome
Two readings an operator will otherwise get wrong, both learned the hard way:
DeliveryNotFindableErroris views-postprocessing 1.4.0 WORKING. That build verifies a delivery by what it refuses, and the message names every object it checked. Read as a failure, it invites exactly the wrong remedy.6. Credentials — beyond the chmod trap now in #524
APPWRITE_ENDPOINT,APPWRITE_DATASTORE_PROJECT_ID,APPWRITE_DATASTORE_API_KEY); the rest are identifiers. Canonical home isviews-faoapi/.env.views-pipeline-core'sPredictionStoreConfigdocstring claims it reads these "once at startup … preventing silent failures after hours of training." It does not — it is called from_build_datastore, i.e. after training (views-pipeline-core#557). Until that moves,pod_run_fao_delivery.sh --preflightis the only early check, and the guide should say why running it is not optional.7. Environment facts a fresh pod needs and the guide does not state
xarray 2025.12.0/pandas 3.0.6where every successful run used2024.3.0/1.5.3, and neither was pinned (A fresh un_fao environment is unbuildable: datafactory's xarray cap permits a release needing pandas>=2.1, the viewser chain caps pandas below 2 #516). Not a build failure — a silent major-version change.xarray<2025looks right; 2024.11.0 already requirespandas>=2.1. The cliff is inside the 2024 line: 2024.3.0 is the last release accepting pandas 1.x, 2024.5.0 moved to>=2.0, 2024.9.0 to>=2.1.views-datafactory>=1.13.0, not>=1.9.0, wherever a leg runs on rented hardware: before 1.13.0 the client could carry a netrc credential across a redirect to another host and embed it in error messages (34 models pin views-datafactory>=1.9.0; the credential-handling fixes landed in 1.13.0 #509). Currently deferred for the 34 model declarations — seetests/test_requirements_hygiene.DEFERRED_PACKAGES.WANDB_MODE=offlineis not optional. Without itmain.pycallswandb.login()and blocks on a prompt nobody is watching; the first forecasting run died there after the environment was already built.8. Rehearsals
--rehearsal <lessons>now exists on both runners. Two things the guide must be blunt about:/workspace/deliver/_fao/REHEARSALsays so. They must be superseded, not left.9. Operational details worth a line each
setsidmatters — already in the guide, and it earned its place.cat /workspace/deliver/_fao/STAGE, ortail -f .../run.log.PRO 6000 MIG 24GB— 25× slower on real work. Already in the guide; worth keeping loud.20260929_172325are inert. If the guide grows a cleanup section, they belong in it.Related
#516 · #517 · #518 · #509 · #523 · #525 · #499 · views-pipeline-core#551 · #555 · #557 · views-postprocessing#313 · #319