Skip to content

EPIC: the 11 pgm darts models — calibration point predictions on rented hardware, delivered to the research team #532

Description

@Polichinel

The 11 pgm datafactory r2darts2 models have been on development since #491 and not one of them has ever produced a prediction. This epic makes them runnable on rented hardware and turns their calibration output into the point predictions the research team consumes. It is the r2darts2 sibling of #505 (the HydraNet half), and it closes the last open line of epic #488.

The models: blue_ocean brave_heart dancing_monkey dark_necessities dark_river golden_eagle little_talks mister_bluesky old_rules red_hawk silent_fox. NBEATS / TSMixer / TiDE, n_epochs: 300.


Why this matters

Researchers are waiting on calibration-partition point predictions to run views-platform/ensemble-updater and other analyses. The HydraNet half of that delivery happened in September. The darts half never started, because there is no way to run a darts model on a pod and no way to turn its output into what the tool reads.

Three things make this more than "run the models":

  1. No r2darts2 model has ever predicted at pgm. The laptop's 31 GB ran out twice (S2 (#488): bring the 11 pgm datafactory r2darts2 models from staging_202608 onto development #489). The one run that completed on fimbulthul wrote zero predictions — the engine labelled the index country_id instead of priogrid_id, CorePredictionSniffer correctly refused all 13 origins, and concurrent.futures.wait(futures) discarded every refusal so the run reported PASS (fixed in fix(#499): the eleven pgm darts models declare entity_id: priogrid_id — their predictions were labelled country_id and refused #504 / b2cd4a07; never re-run). Whatever this epic produces is the first pgm r2darts2 output in existence.
  2. fimbulthul no longer resolves. This is rented hardware or nothing, and budget has been genuinely tight ($20–58 on the last campaign).
  3. The output is a trap. The dataframe path writes predictions_calibration_<ts>_{00..12}.parquet — almost the deliverable. But darts_bridge.prediction_frames_to_dataframe puts a list in every cell, and ADR-023:148-149 records that ensemble-updater's _as_float_prediction_array "takes float(x[0]) on a list cell — one draw, silently, which is the exact failure this ADR is meant to prevent." Handing the parquets over unconverted is a silent wrong answer, not a crash.

Desired end state

  • A pod runner that takes one darts model from a bare pod to verified delivery parquets, with a free --preflight that refuses before any GPU time is bought.
  • A converter that turns list cells into one scalar per cell, in count space, with the validations the HydraNet path already earned.
  • Documentation good enough that someone who was not here can repeat the run — tools/podrun/__init__.py's own promotion criterion.
  • 11 models × 13 parquets of 2 333 448 rows, delivered.
  • Epic EPIC: the 19 target models — 11 pgm datafactory r2darts2 + 8 HydraNets — run on the latest platform #488 definition-of-done line 4 ticked.

Scope boundaries

In: tools/podrun/, tools/collapse/, tests/, docs/runpod_run_guide.md, ADR-023 amendment, one new ADR, and three config lines each in little_talks / mister_bluesky.

Out:

What the investigation established

Each of these was measured, not recalled, and each one shapes a story.

finding consequence
Calibration is train (121, 456), test (457, 504); max(steps) = 36 → 13 rolling origins, 2 333 448 rows each. Identical across all 11 (config_partitions.py is one blob). The deliverable shape matches the HydraNet delivery exactly.
PyPI's latest r2darts2 is 0.2.3, which never frees the prediction scratch directory (views-r2darts2#54: ~4 000 dirs at ~53 GB filled a 2 TB disk). 0.2.4 fixes it and is tagged but unpublished. Install from the git tag 0.2.4 and assert the installed version.
PredictionScratch() takes no base_dir, so scratch honours TMPDIR and otherwise lands in /tmp — the container disk — while the runner's disk floor measures /workspace, the volume. A floor that measures a filesystem the workload does not use cannot fire. Set TMPDIR.
little_talks and mister_bluesky are the only two at num_samples: 100 / mc_dropout: True with point metrics commented out. The list-cell conversion materialises all 13 origins at once — measured 23.3 GB per origin → ~303 GB, against a pod rule of RAM ≥ 50 GB. The nine measure ~11 GB. Those two cannot run as configured. Align them to the other nine (S4) and file the measurement upstream.
pred_type is derived from the data, not the config (native_evaluator.py:258). Dropping to 1 sample without re-activating regression_point_metrics makes a run train fully, then raise "No metrics configured for (regression, point)", writing nothing. S4 is three lines, not one.
A calibration run uploads nothing, and the datafactory authenticates by HTTP Basic from ~/.netrc mode 600 — VIEWS_DATAFACTORY is a phantom no code reads (postmortem §2.8). No [appwrite] extra, no publish variables on rented hardware. New ADR.
blue_ocean, brave_heart, golden_eagle declare ~71 covariates; the other eight declare 3. A second cost class. Measure before fanning out.
run_integration_tests.sh sets no WANDB_MODE (so main.py blocks on wandb.login()), wants a pre-existing named conda env while a pod builds a uv venv, and defaults to a 1800 s timeout. It cannot be the pod runner.

Stories

Ordered checklist: #539. Tick the boxes there as stories land.

story depends on
S1 #533 — The darts collapse: list cells to one scalar per cell —
S2 #534 — The darts pod runner #533
S3 #535 — Make the run repeatable by someone who was not here #534
S4 #536 — Align little_talks and mister_bluesky to the point delivery — (needs a decision)
S5 #537 — Measure one model on a pod: dark_river #534, #535
S6 #538 — Run the remaining ten and deliver #537, #536
#533 ──> #534 ──> #535 ──> #537 ──> #538
                            ▲         ▲
#536 ───────────────────────┘─────────┘

#533 is the hinge — #534's final stage invokes the converter. #536 is independent of the three code
stories. #537 is the money gate: nothing after it is planned until its numbers exist.

Epic acceptance criteria

  • bash tools/podrun/pod_run_darts_calibration.sh --preflight <model> on a machine with nothing configured names every missing precondition at once and reaches no run stage.
  • The converter refuses, with a test proving it, to emit x[0] from a multi-sample list cell.
  • One darts model completes calibration on a GPU and writes 13 parquets — epic EPIC: the 19 target models — 11 pgm datafactory r2darts2 + 8 HydraNets — run on the latest platform #488 DoD line 4.
  • Delivery parquets validate: 2 333 448 rows, month_id/priogrid_id as flat int64 columns, one float scalar per target per cell, finite, non-negative, no duplicate (month_id, priogrid_id).
  • Runtime, peak RAM and peak disk are recorded for one model before any second pod is rented.
  • All 11 models delivered, or the exceptions named with the reason.
  • tools/podrun/__init__.py records that a second model family has run through it.

Related: #488 · #489 · #492 · #494 · #499 · #504 · #505 · views-r2darts2#54 · views-r2darts2#55

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions