You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
EPIC: the 11 pgm darts models — calibration point predictions on rented hardware, delivered to the research team #532
The 11 pgm datafactory r2darts2 models have been on development since #491 and not one of them has ever produced a prediction. This epic makes them runnable on rented hardware and turns their calibration output into the point predictions the research team consumes. It is the r2darts2 sibling of #505 (the HydraNet half), and it closes the last open line of epic #488.
The models: blue_oceanbrave_heartdancing_monkeydark_necessitiesdark_rivergolden_eaglelittle_talksmister_blueskyold_rulesred_hawksilent_fox. NBEATS / TSMixer / TiDE, n_epochs: 300.
Why this matters
Researchers are waiting on calibration-partition point predictions to run views-platform/ensemble-updater and other analyses. The HydraNet half of that delivery happened in September. The darts half never started, because there is no way to run a darts model on a pod and no way to turn its output into what the tool reads.
Three things make this more than "run the models":
fimbulthul no longer resolves. This is rented hardware or nothing, and budget has been genuinely tight ($20–58 on the last campaign).
The output is a trap. The dataframe path writes predictions_calibration_<ts>_{00..12}.parquet — almost the deliverable. But darts_bridge.prediction_frames_to_dataframe puts a list in every cell, and ADR-023:148-149 records that ensemble-updater's _as_float_prediction_array"takes float(x[0]) on a list cell — one draw, silently, which is the exact failure this ADR is meant to prevent." Handing the parquets over unconverted is a silent wrong answer, not a crash.
Desired end state
A pod runner that takes one darts model from a bare pod to verified delivery parquets, with a free --preflight that refuses before any GPU time is bought.
A converter that turns list cells into one scalar per cell, in count space, with the validations the HydraNet path already earned.
Documentation good enough that someone who was not here can repeat the run — tools/podrun/__init__.py's own promotion criterion.
In:tools/podrun/, tools/collapse/, tests/, docs/runpod_run_guide.md, ADR-023 amendment, one new ADR, and three config lines each in little_talks / mister_bluesky.
Out:
run.sh files — standing rule, they are production infrastructure and generated from a pipeline-core template.
The nine other models' configs — already correct (REGION = "land", loa: priogrid_month, prediction_format: "dataframe", num_samples: 1, mc_dropout: False, point metrics active, maturity candidate).
views-r2darts2 code. The one upstream finding is filed as an issue, not a patch.
Publishing views-r2darts2 0.2.4 to PyPI. The runner installs from the git tag; a release is irreversible and other repos resolve against it.
Each of these was measured, not recalled, and each one shapes a story.
finding
consequence
Calibration is train (121, 456), test (457, 504); max(steps) = 36 → 13 rolling origins, 2 333 448 rows each. Identical across all 11 (config_partitions.py is one blob).
The deliverable shape matches the HydraNet delivery exactly.
PyPI's latest r2darts2 is 0.2.3, which never frees the prediction scratch directory (views-r2darts2#54: ~4 000 dirs at ~53 GB filled a 2 TB disk). 0.2.4 fixes it and is tagged but unpublished.
Install from the git tag 0.2.4 and assert the installed version.
PredictionScratch() takes no base_dir, so scratch honours TMPDIR and otherwise lands in /tmp — the container disk — while the runner's disk floor measures /workspace, the volume.
A floor that measures a filesystem the workload does not use cannot fire. Set TMPDIR.
little_talks and mister_bluesky are the only two at num_samples: 100 / mc_dropout: True with point metrics commented out. The list-cell conversion materialises all 13 origins at once — measured 23.3 GB per origin → ~303 GB, against a pod rule of RAM ≥ 50 GB. The nine measure ~11 GB.
Those two cannot run as configured. Align them to the other nine (S4) and file the measurement upstream.
pred_type is derived from the data, not the config (native_evaluator.py:258). Dropping to 1 sample without re-activating regression_point_metrics makes a run train fully, then raise "No metrics configured for (regression, point)", writing nothing.
S4 is three lines, not one.
A calibration run uploads nothing, and the datafactory authenticates by HTTP Basic from ~/.netrc mode 600 — VIEWS_DATAFACTORY is a phantom no code reads (postmortem §2.8).
No [appwrite] extra, no publish variables on rented hardware. New ADR.
blue_ocean, brave_heart, golden_eagle declare ~71 covariates; the other eight declare 3.
A second cost class. Measure before fanning out.
run_integration_tests.sh sets no WANDB_MODE (so main.py blocks on wandb.login()), wants a pre-existing named conda env while a pod builds a uv venv, and defaults to a 1800 s timeout.
It cannot be the pod runner.
Stories
Ordered checklist: #539. Tick the boxes there as stories land.
story
depends on
S1
#533 — The darts collapse: list cells to one scalar per cell
#533 is the hinge — #534's final stage invokes the converter. #536 is independent of the three code
stories. #537 is the money gate: nothing after it is planned until its numbers exist.
Epic acceptance criteria
bash tools/podrun/pod_run_darts_calibration.sh --preflight <model> on a machine with nothing configured names every missing precondition at once and reaches no run stage.
The converter refuses, with a test proving it, to emit x[0] from a multi-sample list cell.
Delivery parquets validate: 2 333 448 rows, month_id/priogrid_id as flat int64 columns, one float scalar per target per cell, finite, non-negative, no duplicate (month_id, priogrid_id).
Runtime, peak RAM and peak disk are recorded for one model before any second pod is rented.
All 11 models delivered, or the exceptions named with the reason.
tools/podrun/__init__.py records that a second model family has run through it.
The 11 pgm datafactory r2darts2 models have been on
developmentsince #491 and not one of them has ever produced a prediction. This epic makes them runnable on rented hardware and turns their calibration output into the point predictions the research team consumes. It is the r2darts2 sibling of #505 (the HydraNet half), and it closes the last open line of epic #488.The models:
blue_oceanbrave_heartdancing_monkeydark_necessitiesdark_rivergolden_eaglelittle_talksmister_blueskyold_rulesred_hawksilent_fox. NBEATS / TSMixer / TiDE,n_epochs: 300.Why this matters
Researchers are waiting on calibration-partition point predictions to run
views-platform/ensemble-updaterand other analyses. The HydraNet half of that delivery happened in September. The darts half never started, because there is no way to run a darts model on a pod and no way to turn its output into what the tool reads.Three things make this more than "run the models":
fimbulthulwrote zero predictions — the engine labelled the indexcountry_idinstead ofpriogrid_id,CorePredictionSniffercorrectly refused all 13 origins, andconcurrent.futures.wait(futures)discarded every refusal so the run reportedPASS(fixed in fix(#499): the eleven pgm darts models declare entity_id: priogrid_id — their predictions were labelled country_id and refused #504 /b2cd4a07; never re-run). Whatever this epic produces is the first pgm r2darts2 output in existence.fimbulthulno longer resolves. This is rented hardware or nothing, and budget has been genuinely tight ($20–58 on the last campaign).dataframepath writespredictions_calibration_<ts>_{00..12}.parquet— almost the deliverable. Butdarts_bridge.prediction_frames_to_dataframeputs a list in every cell, and ADR-023:148-149 records thatensemble-updater's_as_float_prediction_array"takesfloat(x[0])on a list cell — one draw, silently, which is the exact failure this ADR is meant to prevent." Handing the parquets over unconverted is a silent wrong answer, not a crash.Desired end state
--preflightthat refuses before any GPU time is bought.tools/podrun/__init__.py's own promotion criterion.Scope boundaries
In:
tools/podrun/,tools/collapse/,tests/,docs/runpod_run_guide.md, ADR-023 amendment, one new ADR, and three config lines each inlittle_talks/mister_bluesky.Out:
run.shfiles — standing rule, they are production infrastructure and generated from a pipeline-core template.REGION = "land",loa: priogrid_month,prediction_format: "dataframe",num_samples: 1,mc_dropout: False, point metrics active, maturitycandidate).pod_run_fao_delivery.sh, and theprediction_framemigration (Migrate the r2darts2 family to prediction_format: prediction_frame — deferred from #489 #492).What the investigation established
Each of these was measured, not recalled, and each one shapes a story.
train (121, 456),test (457, 504);max(steps) = 36→ 13 rolling origins, 2 333 448 rows each. Identical across all 11 (config_partitions.pyis one blob).0.2.4and assert the installed version.PredictionScratch()takes nobase_dir, so scratch honoursTMPDIRand otherwise lands in/tmp— the container disk — while the runner's disk floor measures/workspace, the volume.TMPDIR.little_talksandmister_blueskyare the only two atnum_samples: 100/mc_dropout: Truewith point metrics commented out. The list-cell conversion materialises all 13 origins at once — measured 23.3 GB per origin → ~303 GB, against a pod rule of RAM ≥ 50 GB. The nine measure ~11 GB.pred_typeis derived from the data, not the config (native_evaluator.py:258). Dropping to 1 sample without re-activatingregression_point_metricsmakes a run train fully, then raise "No metrics configured for (regression, point)", writing nothing.~/.netrcmode 600 —VIEWS_DATAFACTORYis a phantom no code reads (postmortem §2.8).[appwrite]extra, no publish variables on rented hardware. New ADR.blue_ocean,brave_heart,golden_eagledeclare ~71 covariates; the other eight declare 3.run_integration_tests.shsets noWANDB_MODE(somain.pyblocks onwandb.login()), wants a pre-existing named conda env while a pod builds a uv venv, and defaults to a 1800 s timeout.Stories
Ordered checklist: #539. Tick the boxes there as stories land.
little_talksandmister_blueskyto the point deliverydark_river#533 is the hinge — #534's final stage invokes the converter. #536 is independent of the three code
stories. #537 is the money gate: nothing after it is planned until its numbers exist.
Epic acceptance criteria
bash tools/podrun/pod_run_darts_calibration.sh --preflight <model>on a machine with nothing configured names every missing precondition at once and reaches no run stage.x[0]from a multi-sample list cell.month_id/priogrid_idas flatint64columns, onefloatscalar per target per cell, finite, non-negative, no duplicate(month_id, priogrid_id).tools/podrun/__init__.pyrecords that a second model family has run through it.Related: #488 · #489 · #492 · #494 · #499 · #504 · #505 · views-r2darts2#54 · views-r2darts2#55