Skip to content

Release: development → main (442 commits) — G5 of the runbook, HOLD at G4b - #410

Open
Polichinel wants to merge 497 commits into
mainfrom
development
Open

Polichinel wants to merge 497 commits into
mainfrom
development

Conversation

@Polichinel

@Polichinel Polichinel commented Aug 22, 2026 •

Copy link
Copy Markdown
Collaborator

⚠️ HOLD — one gate is half-verified, not red

Merging this is G5 of the release runbook (#282), the final gate. CORRECTED 2026-08-22: this PR first claimed G4 was red on the strength of tools.liveness, which reports forecast_verdict: NEVER_DELIVERED. That reading was wrong — the forecast is delivered. The detector matches on a filename prefix the delivery does not use (filed separately). What remains genuinely unverified is the servable half of G4. Opened so the release is visible, reviewable and CI-checked.

This is the views-models release: development → main, 442 commits, the first since 2026-06-26. Per #282 §1, a release of this repo is exactly this merge plus a tag — nothing is published, nothing installs it, and no sibling repo depends on it. Its artifact is the forecast, which is why the gate is a real run rather than an install check.

1346 files changed, 47250 insertions(+), 3777 deletions(-)

Release gate — measured today, not inherited from the runbook

# gate status
G0 all five workflows green on development ✅ now green — Bootstrap, Run Tests, Runtime Smoke, Secret Scan, Update Model Catalogs
G1 main ahead of development = 0; no partition numbers moved ✅ re-verified below
G2 ensembles/pink_ponyclub/run.sh -m gets past import on a real machine ⬜ unverified
G3 monthly_run.sh end to end, env snapshots committed ⬜ unverified
G4a a real FAO forecast in unfao_bucket ✅ met — 110 shards, verified by listing the bucket
G4b that forecast servable by faoapi ⬜ unverified — needs FAO's caller key
G5 this merge, tag 0.3.0, branch protection on main ⬜ pending G2, G3, G4b

G0 — the blocker is gone

The runbook recorded G0 at 4/5, blocked by "the one item nobody in this repo can fix" — violet_visitor's committed loss_reg. #297 and #336 are both closed and all five workflows pass on development.

G1 — re-verified, and one claim corrected

main ahead of development is 0. On partitions, a naive number-scrape of the 132 changed config_partitions.py files reports 87 "moved", and that reading is wrong. Two things account for all of it:

  1. ViewsMonth.now().id (an ingester3 dependency) replaced by a local _current_month_id() using the January-1980 epoch — the 1980 is the epoch constant, not a boundary. Equivalence checked: (2026−1980)×12+8 = 560, matching the liveness surface's now_month_id: 560.
  2. A dropped comment line: # ViewsMonth reference: 121 = Jan 1990, 444 = Dec 2016, ….

The partition tuples themselves are byte-identical on both sides:

origin/main         (121, 456) (457, 504) (121, 504) (505, 552)
origin/development  (121, 456) (457, 504) (121, 504) (505, 552)

0 partition boundaries moved. The forecast windows this release ships are the ones main already has.

G4 — corrected

First reading, which was wrong. python -m tools.liveness reports:

forecast_verdict:   NEVER_DELIVERED
historical_verdict: DELIVERING
other_files: 110

That was taken at face value and this PR was opened claiming the release was blocked.

What is actually in the bucket, listed directly:

110 x rusty_bucket_forecasting_20260727_095355__lr_ged_sb__m000559 ... m000668  (0.92 MB each)
  1 x historical_dataset_20260813_080043.parquet                                (171.8 MB)

The forecast is there — 110 shards, delivered 2026-08-13. G4's delivery half is met.

The detector cannot see it. tools/liveness/unfao_delivery.py:40-41:

FORECAST_PREFIX   = "forecast_dataset_"      # nothing is named this
HISTORICAL_PREFIX = "historical_dataset_"    # this one matches

The delivery writes rusty_bucket_forecasting_*, so the forecast half has never matched a real file, while the historical half matches and reports correctly. Those 110 shards are counted as other_files. Filed separately; it matters beyond this PR because #320 is "FAO forecast delivery has been stalled for 145 days and nothing detected it" — this detector exists for that, and is blind in the stream it was built to watch.

What is still unverified is G4b: that faoapi can serve the delivered forecast. That is the half this repo cannot check alone.

Deletions — 2, both intended

models/bright_starship/scripts/audit_data_parity.py
scripts/update_partitions.py

Superseded by tools/{scaffold,catalogs,partitions,audit}/ under C-60. The runbook anticipated 9; 7 of those already flowed to main.

Still outstanding at G5, from the runbook

  • Tag 0.3.0 — bare MAJOR.MINOR.PATCH, no v. The v prefix is an active defect: git tag --list | sort -V | tail -1 returns v0.0.1 as "newest" because the prefix sorts after digits.
  • Branch protection on main, requiring the five workflows. main currently has none, which is how it took four direct hotfix commits in June 2026 that never flowed back.
  • ADR-019, the release procedure — does not exist. No document defines what main means. CLAUDE.md's own test applies: a rule a contributor could violate without knowing it existed.

What this release explicitly does not do

Publish anything · restructure the 131-requirements mismatch (C-116). Two items that were on this list are done and no longer gate anything here: the views-r2darts2 three-way split (C-115) was closed by #461 — all 31 declare one spec; and both views-hydranet (0.1.0, 2026-09-15, views-hydranet#340; the 8 models pin ~=0.1.0, #462) and views-postprocessing (1.2.0) are on PyPI. The FAO launcher still installs views-postprocessing from git+…@main — a moving branch pointer, tracked with views-pipeline-core#222 — but PyPI availability is no longer the reason. (Updated 2026-09-15, #468.)

Polichinel and others added 30 commits July 19, 2026 15:43
…tion

docs(liveness): naming convention — assumption upgraded to cited fact (views_api wiki)
…ndows (#240)

python -m tools.liveness.datafactory_input answers: does the input store's
observed coverage (live last_valid_month_id) reach what meta/partitions.json
requires? Requirement DERIVED at runtime (max test-window end — re-arms on
every partition bump automatically; automates the C-96 tripwire). TDD (14
tests first, injected reader/netrc seams); netrc presence reported as a fact,
values never read; exit 0/1/2. First live run: INPUT_FRESH (558 vs 552,
margin 6 months).

Also: S1 review-finding remediation — parse_run_name now rejects legacy
month-00 names (fatalities001_2022_00_t01) with a pinned test, keeping the
month math clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tory-input

feat(liveness): S2 — datafactory input check (coverage vs partition windows) (#240)
…ion IDs encoded (#241)

python -m tools.liveness.appwrite_store answers: is the new shelf reachable,
when did the newest forecast land, and do the REAL metadata IDs hold? The
IDs are now discovered + encoded with receipts (db 'file_metadata',
collection 'production_forecasts'); the June 2026 failure's phantom
'forecasts_metadata' is documented as the historical wrong value — closing
C-100's founding mystery. TDD (19 tests first): injected fetch/credentials/
clock; creds resolved env-first then known .env (ancestor-walk discovery —
robust from worktrees); secrets never rendered (pinned redaction test);
truthful SKIP without creds. Verdicts STORE_ACTIVE/IDLE/UNREACHABLE/SKIP,
exit 0/1/2.

First live run: STORE_IDLE — newest file 2025-11-27 (234 days), 318 files,
real_collection_present=True. The check reporting a true problem on day one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e-store

feat(liveness): S3 — Appwrite production_forecasts check + real collection IDs (#241)
…_bucket (#242)

python -m tools.liveness.unfao_delivery answers "when did FAO last receive
anything?" per delivery stream (forecast_dataset_* / historical_dataset_*),
with verdicts DELIVERING/STALLED/NEVER_DELIVERED per stream and an overall
DELIVERING/DELIVERY_STALLED/UNREACHABLE/SKIP. TDD (12 tests first, fixtures =
the real bucket listing); credentials reused from appwrite_store (same
project/.env — S7 homes it in a shared module, noted); secrets never
rendered; truthful SKIP without creds; exit 0/1/2.

First live run: DELIVERY_STALLED — forecast 131 days (2026-03-10),
historical 111 days (2026-03-30). The 4-month FAO stall (C-99-class silent
lapse) is now machine-visible.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…elivery

feat(liveness): S4 — FAO unfao_bucket delivery check (per-stream freshness) (#242)
…d a false idle verdict (#241, #242)

Appwrite lists 25 files/page by default; sorting one page of the 318-file
production_forecasts bucket yielded a FALSE 'newest = 2025-11-27' (truth:
2026-06-29 — the June monthly runs DID upload to the shelf). The shipped S3
check inherited the flaw; S4 was correct only because its bucket fits one
page. Both checks now request server-side orderDesc($createdAt)+limit
(S4: per-stream startsWith+orderDesc+limit, with per-stream totals);
regression test pins the ordering requirement. Ground-truth docstrings
corrected.

Corrected live verdicts: production_forecasts STORE_ACTIVE (newest
2026-06-29, 19 days — July cycle NOT uploaded, a real finding); unfao
DELIVERY_STALLED unchanged (131/111 days).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ination

fix(liveness): Appwrite server-side ordering — pagination bug gave a false idle verdict
…cycle? (#243)

python -m tools.liveness.wandb_execution answers execution recency per
monthly ensemble (roster mirrors monthly_run.sh, hand-encoded with a
keep-in-sync note): latest finished forecasting run, created_at, state,
days_since, and the train-window end month_id — the data-cutoff receipt
that resolved the run-naming ambiguity. Verdicts COMPUTED/NOT_COMPUTED/
NEVER_RUN per ensemble; overall EXECUTION_CURRENT/STALE/UNREACHABLE/SKIP;
exit 0/1/2. TDD (13 tests first, injected client/netrc/clock); wandb
imported lazily (already installed, zero new deps); truthful SKIP without
~/.netrc api.wandb.ai.

First live run: EXECUTION_CURRENT — pink_ponyclub/skinny_love finished
2026-06-29 (19d, cutoff 557), rude_boy/first_love 2026-07-15 (3d, cutoff
558). The 2026-07-19 trust-crisis question is now a one-command fact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…xecution

feat(liveness): S5 — wandb execution check (did the team compute this cycle?) (#243)
…244)

python -m tools.liveness.vpn_store answers: does the legacy store hold a
fresh fatalities run (computed uploads awaiting promotion)? Off-VPN it
reports VPN_REQUIRED (exit 0) — the 2026-07-19 observation-boundary trap
is now a named verdict, never a false RED. On-VPN: lists runs via
views_forecasts.db_ops.ViewsMetadata, judges the newest fatalities run with
the S1 parser + freshness budget (one parser, one convention). TDD (12
tests first); host-resolution failures classified VPN_REQUIRED, missing
packages SKIP_NO_PACKAGE, else UNREACHABLE.

Receipt encoded: the store's Postgres schema is literally
'forecasts_metadata' — the origin of the phantom Appwrite collection ID
that killed the June 2026 run (legacy schema name copied into new-store
config). The C-100 incident now has its full causal chain documented in
code.

On-VPN behavior pending live confirmation at the maintainer's next VPN
session (noted in #244).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
feat(liveness): S6 — VPN store (gjoll) check with truthful VPN_REQUIRED (#244)
…#245)

python -m tools.liveness runs all six surface checks, prints one raw-facts
block per surface, and exits with the worst verdict (0 healthy/skips, 1
attention, 2 unreachable); a crashing check is contained and reported, never
hiding the others.

Extraction strictly limited to demonstrated duplication (WET-before-DRY,
maintainer doctrine): report.py (fact renderer + the merged verdict->exit
map, unknown verdicts fail loud) and appwrite_api.py (credentials dataclass/
.env parsing/resolution order, the pagination-cure query builders, the
default stdlib fetch — previously duplicated across the two Appwrite
checks). All six modules now delegate; appwrite_store re-exports for
backward compatibility. Behavior preservation proven: all 92 pre-existing
liveness tests pass UNMODIFIED; +10 runner/report tests (102 total).

First full dashboard run: API fresh, input fresh (margin 6), shelf active
(19d), execution current (4 ensembles), FAO stalled (131/111d — the one
true red), vpn store VPN_REQUIRED. worst_exit: 1. The maintainer's stated
bare-minimum MVP — "you must be able to check whether your forecasts are
live" — exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
feat(liveness): S7 — DRY consolidation + python -m tools.liveness (the dashboard) (#245)
tools/liveness/README.md: one command + per-surface usage, the 0/1/2 exit
contract (truthful skips are 0), every verdict per surface, and the three
encoded conventions with receipts (data-cutoff run naming, the real Appwrite
metadata IDs vs the phantom forecasts_metadata, server-side orderDesc
listing). Validated against live output and the suite (102 passed); every
command copy-paste correct.

Register: C-100 -> Mitigated (the config-vs-reality exit exists as
tools/liveness; residual = scheduling, which is C-99's exit); C-96 note
(tripwire automated by datafactory_input); C-98 note (both stores now
observable; authority still undecided); C-99 note (the instrument is the
dead-man's-switch primitive; the heartbeat remains). Header 44/16.

Closes the liveness epic #238.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs(liveness): S8 — README + register closure, C-100 → Mitigated (#246)
… charter coverage gap (T3) — falsify audit 2026-07-19

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 2026-07-19 falsification audit (claim: "tools.liveness is air and water
tight") found 3 hard + 3 soft defects. Fixes:

- report.one_line(): fact values with embedded newlines (live sqlalchemy
  errors) no longer break the one-fact-per-line contract (P1)
- datafactory_input: missing datafactory_query is now a truthful
  SKIP_NO_PACKAGE (exit 0), mirroring vpn_store, not a false UNREACHABLE
  alarm (P2)
- wandb_execution: _judge moved inside the per-ensemble try — a malformed
  created_at is a failure fact, never an uncaught crash (P4)
- all six main()s classify the verdict BEFORE printing, so an unregistered
  verdict fails loud without a half-block the runner would contradict (P7)
- roster tripwire: MONTHLY_ENSEMBLES is now asserted against monthly_run.sh
  itself, not a literal copy (P5)
- README documents the C-102 non-goals (viewser, website, content sanity)
  so all-green cannot be over-read (P8, xfail-pinned)

All findings enforced by tests/test_liveness_falsifications.py
(107 liveness tests green, 1 xfail = the open C-102 gap). Register:
C-101 -> Resolved, C-102 Open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-fixes

fix(liveness): falsification-audit fixes — C-101 Resolved, C-102 registered
…p (T3) — falsify audit

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…coverage — C-103 Resolved

Falsify audit (claim: "100% covered, green/beige/red all around") found the
suite at 95% coverage, zero beige tests, and red misused for live network
probes (ADR-005 red = adversarial/error-path — the maintainer's reading was
correct; the suites' was not).

- ADR-005 amended: fourth category `live` (real-external-service probes,
  truthful skip offline) + pyproject marker
- 6 live probes relabeled red -> live; 22 genuine error-path tests red-marked
- 8 beige structural tests added (module conventions, verdict-registration
  in the exit map, SURFACES registry, roster tripwire)
- Coverage 95% -> 100% branch: real offline tests for the default clients
  (fake wandb/views_forecasts modules, monkeypatched urllib, credential
  resolution paths, netrc-probe failure paths); pragma only on __main__
  guards, with reasons
- tests/test_liveness_taxonomy.py enforces the taxonomy from now on

130 liveness tests green (was 107), 1 xfail (C-102 scope marker), ruff clean.
Register: C-103 -> Resolved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
test(liveness): ADR-005 'live' marker + taxonomy fix + 100% branch coverage (C-103)
…ing check

The serving probe sampled only /{run}/cm — a run serving country-month but
empty at grid level read LIVE_FRESH (the "one endpoint" gap, 2026-07-19).
Both public levels are now sampled at the first forecast month; a run empty
at EITHER level is LIVE_NOT_SERVING. Facts: serving_rows_cm +
serving_rows_pgm (replaces serving_rows_sampled); per-level errors prefixed.
Live receipt encoded: fatalities003_2026_05_t01 serves 2 rows at both
levels (cm: country_id keys, pgm: pg_id keys).

TDD (pgm-empty and both-levels-rendered tests written first); 132 liveness
tests green; branch coverage stays 100%.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
feat(liveness): old_api serving probe covers both cm and pgm levels
…d-model corrections + live blocker status for #230

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs(reports): ADR-013 wire-contract review (views-models seat)
…gnment, liveness instrument (still Proposed; decision unchanged)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs(adr-017): 2026-07-19 amendment — production evidence + ADR-013 alignment (stays Proposed)
…ion) + strengthen C-85 to Tier 2

falsify audit of the 2026-07-20 rusty_bucket delivery: four failed ensemble
runs traced to two compounding silent defects.

- C-85 raised Tier 3 -> 2: ensemble --saved reloads a stale y_pred.npy keyed
  on the sample-agnostic model-artifact timestamp; editing n_samples never
  changes the artifact, so the config change is silently discarded (no config
  fingerprint on the cache). Live evidence: mixed S=128/S=32 npy on disk, zero
  at the configured S=16; parent OOM ~18 GB every run regardless of config.
- C-104 (new, Tier 2): posterior sample count is 4 different config keys
  (baseline n_samples, hydranet n_posterior_samples, r2darts num_samples,
  stepshifter pred_samples); the CI/parity contract reads only
  n_posterior_samples, a decoy for baseline — CI green while the runtime
  produces a different count. C-52's readiness fix seeded the decoy.

The silence = three layers missing on one axis: no guardrail (cache/parity
blind to config), fragmented knowledge (4 names + decoy), no test of the
runtime pf.sample_count. Header Open 45 -> 46.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…nfig-load guard runs in CI, six tests off hand-patched paths (#470)

* docs(#468): C-50 and C-42 Resolved — the packages they said were unpublished have been published

C-50 said views-baseline was not on PyPI. It has been since 1.0.0 (2026-07-28); 1.0.2 (2026-09-09)
is the first version that runs, all 37 models pin it (#460), and runtime_smoke.yml installs it from
the index. C-42 said the synthetic models depended on an unreleased pipeline-core branch. That branch
merged and released: PredictionFrameEnsembleManager imports from published 3.2.0 (verified), CI pins
3.0.1, the ensembles declare >=3.0.0 (#372). Both had been stale since July.

Header: Open 70 -> 68, Resolved 43 -> 45. Tiers unchanged.

reports/conda_to_uv_migration_investigation.md gets one dated line under its metadata: the three
packages it proposes git+ dependencies for are on PyPI now. Body kept as the 2026-05-23 record it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* ci(#469): run the roster config-load guard in CI, in its own job

tests/test_roster_configs_load.py constructs each of the 8 HydraNet roster configs through
HydraNetConfig. It is the guard for C-259 — two production ensemble members sat unloadable from
August to September while every test passed, because test_roster_conformance.py compares values
and never builds the object. The load test importorskips views_hydranet, and run_tests.yml
installs only pipeline-core, so it has skipped in CI since the day it was written. views-hydranet
0.1.0 reached PyPI on 2026-09-15 (views-hydranet#340); this job is where the guard now runs.

Separate job, modelled on runtime_smoke.yml: run_tests.yml's install is pinned on purpose, and a
red here means exactly one thing.

#469 assumed torch was avoidable. It is not — importing config_initializer is torch-free, but
get_config -> validate_loss_reg -> utils/utils.py:7 imports it; all 8 configs failed with
ModuleNotFoundError when the job was first run locally. Installed from the CPU-only index, which is
all a validator needs. pipeline-core and views-frames are not on this path and not installed.

Released pin, not a git ref (C-106); ==0.1.0 moves with the models' ~=0.1.0 (#462).

Verified locally in a venv built with exactly this job's install lines: 9 passed. Mutation —
delete ss_feedback from heavy_freighter, where scheduled sampling is on — 2 failed, 7 passed:
the load test on heavy_freighter and test_scheduled_sampling_declares_its_feedback. Reverted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* test(#467): the six directory-creation tests use the _root redirect too

Four ModelScaffoldBuilder tests and two EnsembleScaffoldBuilder tests set
builder._model.model_dir and a hand-written builder._subdirs list after construction — the
same incomplete-redirect shape #467 fixed for the script-building tests, and the module guard's
own docstring named them as the second instance. They did not leak, because
build_model_directory happens to read only the two attributes they patched. That is luck, not a
property.

Now: monkeypatch ModelPathManager._root to tmp_path before constructing the builder, and delete
the hand-patching. The builder derives model_dir and _subdirs from the redirected root itself.
EnsemblePathManager inherits _root, so the ensemble tests use the identical line.

The assertions got stronger, not weaker. creates_subdirs and creates_gitkeep used to assert on a
list the test invented (two paths, or one); they now assert on every subdirectory the path
manager actually declares, and that each is under tmp_path. test_build_model_scripts_without_
directory_raises asserts the redirected directory does not exist before expecting the raise.
test_build_model_directory_creates_dir took a monkeypatch fixture it never used; now it does.

25 passed; nothing written into the repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@Polichinel

Copy link
Copy Markdown
Collaborator Author

G4b status, 2026-09-15 — verified in substance; one key-scope re-run outstanding, and it is Simon's

The gate says "unverified — needs FAO's caller key". That overstates what is missing.

It was verified on 24 August. views-faoapi's own record (reports/post_mortems/2026-08-25_the_week_the_delivery_started_working.md, 08-24 07:17): v1.7.1 deployed, smoke test ALL PASS, forecast grid served for the first time — 32.3 MiB, 2,330,712 rows, 36 months, 27/27 value columns. The live service reports deployed_tag v1.7.1 today, so that check ran against the code that is running now.

Confirmed today by the faoapi session, read-only: /version 200 v1.7.1 · /health 200 healthy, forecast age 33.5 d (SLA 45) · /pg/data/forecast/bulk 200, 33,817,005 bytes · /data/forecast/latest 413 with the byte estimate, as designed.

Why it is not ticked yet: those calls and the 08-24 smoke ran on the write-scoped datastore key. faoapi's per-key cache (ADR-027) means a fresh read-scoped caller — which is what FAO is — hits a cold partition and can see a different state; that is the 2026-08-13 post-mortem exactly. So G4b should be ticked on a read-scoped run. No such key exists in the faoapi session's environment, and it correctly declined to look for one.

What closes it: a read-scoped Appwrite key (the one FAO was issued, or one minted in the console), then from views-faoapi:

APPWRITE_DATASTORE_API_KEY=<read-scoped key> python -m views_faoapi.smoke --expect-tag v1.7.1

Exit 0 → paste here → G4b met. Not an external-party dependency; a console action.

Side finding: v1.7.2 and v1.7.3 are tagged and were never deployed. The served-code delta is one dict entry in the GET / index (forecast_pg_bulk); nothing FAO reads. The /latest 413 behaviour described to FAO on 08-25 is in v1.6.0 and has been live throughout.

Polichinel and others added 27 commits September 17, 2026 10:41
…-core 3.2.0, one file per source guard (#476)

* test(#455): the guard #444 lacked — CI on 3.2.0, the sniffer test selects the file like the loader, no source may carry both

PR #444 added config_maturity.py to 14 models and broke all 14 while 3,684 tests passed. Two
reasons, both fixed here before a single config moves.

CI pinned pipeline-core 3.0.1. Every 3.x before 3.2.0 crashes a config_maturity.py source at the
sniffer — 3.2.0 (2026-09-08) is the first where the loader prefers the new file, deployment_status
is no longer a mandatory key, and the five read sites accept either name (#495, #497). CI now runs
3.2.0, which the ensembles' >=3.0.0,<4.0.0 already permits. The pin's comment records the next
bump's named trigger: from 3.3.0, get_queryset raises instead of returning None, and the catalogs
job calls it for every model with no isolation.

tests/test_core_config_sniffer_contract.py hardcoded config_deployment.py. It fed the sniffer the
dict the real loader would have IGNORED whenever config_maturity.py existed, and stayed green.
File selection is now delegated to pipeline-core's own load_maturity_config (3.2.0) — the same
rule the managers use, so this test cannot drift from them again. _is_deprecated becomes
_is_retired: 3.2.0 refuses to run a retired source by design, in either vocabulary.

The #455 guard: tests/test_config_completeness.py::TestMaturityConfig asserts every source carries
EXACTLY ONE of config_maturity.py / config_deployment.py, with the value valid in that file's
vocabulary. Same for ensembles in test_ensemble_configs.py. ADR-017 Phase 2 is a rename; two files
is the #444 state, where pipeline-core reads one and ignores the other with nothing noticing.
config_deployment.py leaves the fixed required-file lists for the same reason: the requirement is
"one of the two", not a filename.

Mutation-tested on black_ranger, each reverted:
  both files present              -> test_exactly_one_maturity_file FAILS, naming the source
  renamed, value "deployed"       -> sniffer contract FAILS: "maturity='deployed' is not valid"
  renamed, value "retired"        -> green; excluded from sniffer subjects, as 3.2.0 refuses it
  renamed, value "candidate"      -> green — the dry run of the rename itself

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* feat(#449): every reader accepts both maturity vocabularies — the transition window, made real

ADR-017 Phase 2 renames config_deployment.py -> config_maturity.py one source at a time, gated on
the source's engine running pipeline-core >= 3.2.0. 69 of 119 sources (38 stepshifter, 31
r2darts2) run on 2.3.0, which requires the legacy file and knows nothing of maturity, so they keep
it until their engine moves (views-stepshifter#103, views-r2darts2#24). During that window a
source carries exactly one file, and it is the READERS that understand both. No config file moves
in this commit; every reader is ready for the ones that will.

deliveries/coherence.py — maturity_of() reads config_maturity.py first and returns the declared
value after checking it is one of candidate/graduate/retired; falls through to the existing §3
translation for a legacy file; refuses a source carrying both. Its module docstring said
"maturity does not exist yet". A latent bug in the error path is fixed on the way: it called
_source_dir(source).relative_to(...), which is None for anything outside the repo.

tools/catalogs/create_catalogs.py and update_readme.py — load either file; the catalog column is
now "Maturity", in one vocabulary, translated for legacy sources. tests/test_tooling_scripts.py's
"exact copy" of the table generator moves with it.

run_integration_tests.sh — pre-flight classifies by maturity from whichever file exists; skips
retired (or legacy deprecated) models as RETIRED. pipeline-core >= 3.2.0 refuses them anyway.

tools/scaffold/build_{model,ensemble}_scaffold.py — new sources are born with config_maturity.py
(candidate) via template_config_maturity. Verified: the builder writes the new file, not the old.

tests/test_roster_configs_load.py _assemble, tests/test_delivery_coherence.py — prefer the new
file; three new tests cover the declared path, the closed set, and the two-file refusal.

Four CICs move with their scripts (cic-sync): CatalogExtractor, IntegrationTestRunner,
ModelScaffoldBuilder, EnsembleScaffoldBuilder.

WET, on purpose: the §3 table is now in coherence.py, create_catalogs.py and update_readme.py —
three copies of four entries. deliveries/ runs standalone; the tools run under pipeline-core in
CI; sharing an import couples them (coherence.py:71-83 says why). Revisit at a fourth reader.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUBUwUyR5V82K

* fix(#449): one reader of maturity, not three — catalog tools call coherence.maturity_of; ADR-017 §11 dated; docs follow the script

Review findings on #476 (five agents + review-diff), all addressed:

- tools/catalogs/create_catalogs.py, update_readme.py: the copied _LEGACY_TO_MATURITY dict
  mapped deployed -> graduate unconditionally, so the catalog showed ensembles/white_mustang as
  graduate while deliveries/coherence.py (R2: composite is graduate only if every member is)
  said candidate. Two readers of the same fact disagreed. Both tools now import
  deliveries.coherence.maturity_of; coherence.py is stdlib-only, so this couples nothing.
  The "WET, deliberately" justification in the plan was wrong and is withdrawn.
- docs/ADRs/017_source_composition_delivery.md §11 Phase 2: the "cannot ride 3.0" blocker
  expired with pipeline-core 3.0.1/3.2.0; a dated status paragraph records that the rename is
  per source, gated on the engine's floor >=3.2.0, and that the window closes on views-models'
  say-so. Status line extended.
- docs/run_integration_tests.md, README.md: the summary class the script now prints is RETIRED,
  not DEPRECATED.
- run_tests.yml: the floor to bump with is the ensembles', not "the launchers'" (launchers do
  not pin pipeline-core).
- two test docstrings: pipeline-core does not "silently" ignore the second file — it logs a
  warning and runs anyway.

Full suite on 3.2.0: 7111 passed. ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(register): C-147 two maturity readers disagreed (mitigated), C-148 inert guard blind to dict readers (open); C-130 noted

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…rity.py — nothing graduates (#479)

* feat(#449): rename the 50 sources on pipeline-core 3.x to config_maturity.py — nothing graduates

ADR-017 Phase 2, the rename. 29 baseline, 8 hydranet and 13 ensembles — every source whose
engine runs on pipeline-core 3.x — now carry configs/config_maturity.py -> get_maturity_config()
-> {'maturity': ...}, generated from pipeline-core 3.2.0's own template so the fleet has one shape.
The 69 stepshifter and r2darts2 sources keep config_deployment.py: their engines pin pipeline-core
<3.0.0, and 2.3.0 requires the old file (views-stepshifter#103, views-r2darts2#24).

Values, by ADR-017 §3: 40 shadow + 6 baseline -> candidate; 3 deprecated -> retired
(bashful/dopey/happy_dwarf); the one deployed, ensembles/white_mustang, -> candidate under R2
(its members lavender_haze and blank_space are shadow). Nothing becomes graduate: a script is
not an author, and maturity is the author's sign-off.

With it:
- tests/test_ensemble_maturity_rules.py replaces test_falsify_deployment_status_convention.py.
  That file's two xfail stubs said they would "flip to a hard gate when the rule is decided +
  the violation resolved"; ADR-017 §5 decided it and this rename resolved it. R1 and R2 are now
  fleet-wide hard gates in both vocabularies via deliveries.coherence.maturity_of (coherence.py's
  own R2 check only sees ensembles a delivery names, C-144). Mutation-tested: white_mustang ->
  graduate trips R2 naming both members; lavender_haze -> deprecated trips R1 for two ensembles.
  The second stub's subject ('production' literal in pipeline-core's ensemble check) was removed
  in pipeline-core 3.2.0 and no longer exists in any release CI installs.
- tests/test_delivery_map_truth.py counts both files and both keys; the map states the figures
  per file (47 candidate / 3 retired across 50; 68 shadow / 1 deprecated across 69).
  Mutation-tested: a wrong figure in the map turns it red.
- ADR-001, -003, -004, -009: the lines that stated the config_deployment.py contract, each with
  a dated status note (#453). README.md's two config descriptions rewritten.
- KNOWN_REJECTED in test_core_config_sniffer_contract.py is unchanged: _is_retired already
  excludes the three retired sources from the sniffer's subjects, as the docstring says.
- The root catalog and per-model READMEs regenerate from update_catalogs.yml on merge (it
  triggers on models/*/configs/config_*.py and commits); not hand-edited.

Full suite on pipeline-core 3.2.0: 7112 passed, 0 failed, 0 xpass. ruff clean. validate_docs passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(ADR-017): §11 status names #479 as the rename PR

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(ADR-004): the vocabulary row is the maturity set; "production gating depends on it" was the false claim #453 named

Review of #479 (agents 1, 3, 4 independently): the PR amended the required-keys row and left
the next row asserting shadow/deployed/baseline/deprecated as the Tier-1 vocabulary with
"production gating depends on it" — the exact line #453 called false. Now the closed maturity
set, with the legacy four as the translated remainder on pipeline-core 2.x sources.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(register): C-130 resolved by #479; C-144 narrowed — R1/R2 have a fleet-wide edit-time gate

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…d by the sniffer since 2026-06-27 (#482)

#220 set evaluation_mode='point' on zero_/locf_/average_{cm,pgm}baseline and the three _dream
fixtures without aggregate_method, which pipeline-core's CoreConfigSniffer has required for
point mode since 2026-03-13. Every run of these nine died at config load, before any
views-baseline code ran. One line each: "aggregate_method": "arithmetic_mean", the convention
purple_alien and the hydranets already follow.

Scope checked, not assumed: pipeline-core 3.2.0's sniffer over every non-retired model and
ensemble refuses exactly these nine (tests/test_core_config_sniffer_contract.py); pipeline-core
2.3.0's sniffer — what the 68 non-retired stepshifter/r2darts2 sources actually run on — refuses
none of them. KNOWN_REJECTED is therefore empty, as its comment always said it should become;
removing one of the nine lines turns the contract test red naming the model.

Proof of life on the versions requirements.txt installs (views-baseline 1.0.2, pipeline-core
3.2.0): diagonal_dream calibration --train --evaluate runs to "Done" — sniffer audited, model
saved, metrics computed. Independent of #445/#459 (the targets outage), which is also closed.


Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…odel whose data client is not installed (#483)

* chore(#478): CI on pipeline-core 3.3.0; update_readme.py survives a model whose data client is not installed

pipeline-core 3.3.0 (PyPI, 2026-09-19). The named trigger from #476's pin comment fired as
written: the queryset loader now raises ImportError with an install hint instead of returning
None (their #514), and tools/catalogs/update_readme.py called it for every model with no
isolation — under 3.3.0 the catalogs job died on purple_alien, the first datafactory model,
because the job installs no data client. Measured, not assumed: create_catalogs.py exit 0,
update_readme.py exit 1 before; both exit 0 after, 23 models take the isolated branch, and
the regenerated READMEs are byte-identical to what 3.2.0 produced (the None branch: "No
description provided"). Rendering those querysets for real means installing views-datafactory
in the job — that belongs with #474, not a pin bump.

Guard: tests/test_update_readme_survives_missing_data_client.py runs the real script (it is
monolithic, C-81/C-93) against a temporary repo holding one model whose config_queryset.py
imports a client that cannot exist. Green on 3.2.0 and 3.3.0; red on 3.3.0 with the isolation
removed.

Both workflow pins 3.2.0 -> 3.3.0; the pin comment rewritten with what 3.3.0 changes for this
repo (ADR-064 entity refusal, wandb <1.0, views-frames <3 — the CI venv now resolves wandb
0.30.0 and views-frames 2.0.0, the same resolution the ensembles get) and the next named
trigger (4.0: viewser becomes an optional extra, their ADR-063).

Full suite on 3.3.0 / wandb 0.30.0 / views-frames 2.0.0: 7257 passed, 0 failed. ruff clean.
A real ensemble run on that resolution — synthetic_chant, calibration, --saved, after its three
_dream members trained on views-baseline 1.0.2 — finishes with metrics.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* test(docstring): say what the missing-client guard covers and what C-81 still does not

Review: the fixture's unknown client name takes the loader's bare re-raise path, not the
install-hint path — deliberate, and now said; and the guard is narrower than C-81's
validate=True crash at ModelPathManager construction, which stays open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…<0.3.0 (#485)

* build(deps): move all 31 r2darts2 models to views-r2darts2[manager]>=0.2.3,<0.3.0

r2darts2 0.2.3 (released 2026-09-19) is the first 0.2.x that installs beside the published
platform: darts 0.40 / pandas <2 pins, wandb >=0.28.2 (pipeline-core 3.3.0 allows <1.0), plus
fixes for three runtime bugs present on every 0.2.x before it. The `[manager]` extra is now
required: on 0.2.x pipeline-core is optional in r2darts2, and every model's main.py imports it.

Resolver-verified from PyPI alone (uv, py3.11, one model's requirements.txt):
views-r2darts2 0.2.3, views-pipeline-core 3.3.0, viewser 6.6.4, darts 0.40.0, pandas 1.5.3,
wandb 0.30.0. tests/test_requirements_hygiene.py and test_environment_sharing.py: 10 passed.

11 of these 31 models were run end to end on 0.2.3's code (editable) in the views_pipeline env
before release; a fresh-env run on the published wheel is the check this PR still needs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016d7duESygAoK9Dm9Vzwzs7

* test: the requirements-name parser stops at an extra — views-r2darts2[manager] is views-r2darts2

test_algorithm_coherence split the requirement on version operators only, so the [manager]
extra every r2darts2 model now declares (#485) read as a different package from the one
main.py imports. 318 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* test(comments): the two blocks that said 0.2.x is not adoptable now say when and why it became adoptable

test_requirements_hygiene.py and test_environment_sharing.py explained the <0.2.0 cap and
the four conflicts behind it (#317). Both walls fell 2026-09-19 — pipeline-core 3.3.0 widened
wandb, r2darts2 0.2.3 dropped darts 0.46 / pandas 2 — and #485 moves the spec. Dated, both
states kept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… one point forecast of 31 that did not (#487)

* fix(#486): bad_romance declares regression_point_metrics — the one point forecast of 31 that did not

views-evaluation 2.0 refuses a point forecast with no regression_point_metrics ("No metrics
configured for (regression, point)"), after training: bad_romance (num_samples 1, mc_dropout
False) had the line commented out and cost an 816 s run on fimbulthul. Census of the 31
r2darts2 sources: it was the only one. Uncommented, in dancing_queen's shape.

Guard: tests/test_point_forecast_declares_point_metrics.py — every source declaring
num_samples <= 1 must declare regression_point_metrics. 29 point sources checked; red on
bad_romance with the line re-commented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#486): bad_romance's point baselines come back too — 8114034 commented out both lines, not one

Review (git history): the commit that disabled regression_point_metrics disabled
regression_point_baselines in the same hunk. dancing_queen — the sibling this fix says it
matches — has both. Now it does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#485 is not the pandas lock lifting

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…202608 onto development (#491)

* feat(#489): the eleven pgm datafactory r2darts2 models, from staging_202608 onto development

Story S2 of epic #488. blue_ocean brave_heart dancing_monkey dark_necessities dark_river
golden_eagle little_talks mister_bluesky old_rules red_hawk silent_fox — NBEATS / TSMixer /
TiDE, pgm, all three targets, datafactory (loa priogrid_month). Copied directory by
directory from origin/staging_202608 (no branch merge; the four staging ensembles over them
stay on #402), then four corrections each, applied by a one-shot script that is not committed:

- requirements.txt: `views-r2darts2>=0.1.0` (unbounded, no client) -> the fleet spec
  `views-r2darts2[manager]>=0.2.3,<0.3.0` plus `views-datafactory>=1.9.0,<2.0.0`, which the
  queryset imports.
- run.sh: env_path pointed at envs/views-hydranet — a copy-paste from the hydranet template
  — now envs/views_r2darts2 like the 31 siblings.
- config_meta.py: prediction_format "prediction_frame" -> "dataframe". Every other r2darts2
  model runs on "dataframe" (13 of them passed on fimbulthul today); staging hid these 11
  from test_pfe_production_readiness.py with a conftest skip-list that development does not
  have, and they declare neither n_posterior_samples nor hp regression_targets. The
  prediction_frame migration is a separate matter.
- config_deployment.py (shadow) -> config_maturity.py (candidate), from pipeline-core's
  template: on 3.x from day one, one file per source.

New guard, tests/test_datafactory_client_is_declared.py: a config_queryset.py that imports
datafactory_query must have views-datafactory in requirements.txt — the state the eleven
arrived in, which no existing hygiene test judged. 34 importers, 34 declare; red when one
is dropped.

Counts that moved: test_environment_sharing views_r2darts2 22 -> 33; forecast_delivery_map
61 config_maturity.py (58 candidate / 3 retired), 130 files, 131 source directories.

Full suite on pipeline-core 3.3.0: 7682 passed. The sniffer contract accepts all eleven
(126 subjects, none refused). ruff and validate_docs clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(#489): the eleven READMEs regenerated on 3.3.0 — maturity, tree and targets as the configs say; comment count 31 -> 42

Review: the copied READMEs still said 'shadow' and listed config_deployment.py; dancing_monkey's
listed one target of three. Regenerated with tools/catalogs/update_readme.py under the CI mirror
— the bytes the merge job would have written — so nothing on development is stale for a minute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(register): C-149 — run.sh's env prefix is a prompt answer nothing checks against requirements.txt

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…: True on views-hydranet 0.1.1 (#493)

* feat(#484): the eight HydraNets refuse to train on CPU — require_cuda: True on views-hydranet 0.1.1

Story S3 of epic #488. views-hydranet 0.1.1 (2026-09-19) adds `require_cuda`: True turns a
CPU device into a RuntimeError at the top of training, evaluation and forecasting instead of
a warning banner. The reason is views-hydranet#377: a fresh install resolved a torch whose
CUDA build the machine's driver could not run, torch fell back to CPU, and violet_visitor
trained for 6 h 46 m at 1/50th speed under a banner. On a server, that is how a week is lost.

- 8 x configs/config_hyperparameters.py: 'require_cuda': True, next to freeze_recurrent.
- 8 x requirements.txt: views-hydranet~=0.1.0 -> ~=0.1.1. Not optional: HydraNetConfig has
  extra="allow", so on 0.1.0 the key is accepted and does nothing — the floor and the key
  land together or the key is decoration.
- .github/workflows/roster_configs_load.yml: 0.1.1.
- tests/test_roster_conformance.py FOUNDATION: require_cuda True, pinned per model.

Measured:
- roster tests (test_roster_configs_load constructs all 8 HydraNetConfigs, conformance)
  on a venv built as the CI job builds it — pipeline-core 3.3.0, CPU torch, views-hydranet
  0.1.1 --no-deps: 201 passed. Full suite on 3.3.0: 7726 passed.
- The point of the story, end-to-end: a fresh venv from violet_visitor's own
  requirements.txt with torch swapped for the CPU build (the #377 shape), `main.py -r
  calibration -t`: config audited, pgm data fetched and audited, then "require_cuda=True but
  the device for training is 'cpu' ... Refusing to run on CPU (#377)" — exit 1 after 17 s,
  zero epochs.
- Mutation: remove the key from bold_comet -> test_foundation_value red, naming it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* ci(comment): the roster job's header names the pin it must match by date, not by a literal that just went stale

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(register): C-150 — a fresh env's torch can outrun the driver; loud since #493, fixed by nobody yet (#494)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… red on the header count (#495)

The merge-readiness test counts '| Status | Open |' exactly; 'Open — made loud, not fixed' was
not counted and development went red on the #493 squash. Qualifier lives in Notes, where it was
already said.


Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…he 38 stepshifter sources remain on the legacy vocabulary (#496)

* feat(#490): config_maturity.py for the 31 cm r2darts2 models — 38 stepshifter sources are all that is left on the legacy vocabulary

Story S4 of epic #488; ADR-017 Phase 2, the r2darts2 family. #479 could not rename these 31:
their engine pinned pipeline-core <3.0.0, and 2.3.0 requires config_deployment.py. #485 moved
the engine (views-r2darts2 0.2.3, pipeline-core >=3.0), so the #479 procedure applies: each
file generated from pipeline-core 3.3.0's template_config_maturity, values by ADR-017 §3 —
all 31 were `shadow`, all 31 are `candidate`. Nothing graduates; a script is not an author.
One shape across the fleet (the 92 new-vocabulary files hash identically modulo the value).

With it: docs/forecast_delivery_map.md restated — 92 config_maturity.py (89 candidate /
3 retired), 38 config_deployment.py (37 shadow / 1 deprecated), 130 files, 131 source
directories; ADR-017 §11 status: the window now closes when views-stepshifter#103 ships.
test_delivery_map_truth follows the figures; KNOWN_REJECTED stays empty; the catalogs job
regenerates the 31 READMEs on merge (it did so for #479 and #491).

Full suite on pipeline-core 3.3.0: 7725 passed — the sniffer contract accepts all 31 through
load_maturity_config, and R1/R2 (test_ensemble_maturity_rules) hold. ruff and validate_docs
clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs: README and ADR-017's status line no longer call r2darts2 a legacy-vocabulary engine

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs: the maturity gate is the resolved pipeline-core, not the declared floor; README names only the engine still on 2.x

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…iewser | datafactory | synthetic (#497)

* feat(#474): the catalog says which data source each model reaches — viewser | datafactory | synthetic

Story S5 of epic #488. #473 measured 77 models on viewser (pandas 1, Python <=3.11) and 29 on
the datafactory or synthetic data, and nothing showed it: the queryset file knew, the catalog
and READMEs did not. Now:

- tools/catalogs/data_source.py — one reader, by AST, never a guess: `viewser` (imports
  viewser), `datafactory` (imports datafactory_query), `synthetic` (generate() returns
  {"source": "synthetic"}), `none` (no file), `unknown` (both clients, or neither). Reads the
  source; imports nothing from it, so it answers where the client is not installed (#483).
  The same shape as deliveries.coherence.maturity_of (C-147).
- create_catalogs.py: a `Data Source` column in the root README's model table, after Input
  Features. update_readme.py: a `Data Source` row in every model README. Both import the one
  reader.
- models/README_scaffold.md, ensembles/README_ensemble_scaffold.md: the row that has held the
  maturity value since #476 is now labelled Maturity, not Deployment Status; models get the
  Data Source row.
- READMEs regenerated here (root + 111 models + 10 ensembles): the catalogs job triggers on
  models/*/configs/config_*.py, which this change does not touch, so the output moves with the
  script in the same PR.
- tests/test_data_source_catalog.py: each classifier branch on a synthetic file (viewser,
  datafactory, synthetic, both -> unknown, neither -> unknown, relative import -> unknown,
  no file -> none); no model in the fleet is `unknown`; and the pin — 77 viewser / 34
  datafactory / 6 synthetic of 117 — changed on purpose when a model migrates.
  test_tooling_scripts's characterisation copy follows the generator. CatalogExtractor CIC
  lists the key.

Full suite on pipeline-core 3.3.0: 7735 passed. ruff and validate_docs clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* review(#474): the synthetic check reads generate()'s return only; the fleet test imports conftest's discovery; docstring and CIC name the key

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ple_alien's exclusion was an env-isolation workaround (#500)

* chore(#499): the integration runner excludes nothing by default — purple_alien's exclusion dates from when the HydraNets could not run

The default was set while purple_alien (a HydraNet) was unrunnable; all eight run now (#488). The
banner and summary print 'none' for an empty list; the CIC and the user doc say when and why the
default went. --exclude still replaces rather than appends.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(#499): the purple_alien default was an env-isolation workaround (5a2fd2e, 2026-03-15), not a broken model — say so; README's flag table follows

Review: git history says the runner moved to one shared conda env and purple_alien alone needed
views-hydranet, which that env lacked; it was the only HydraNet in the tree then. The 'could not
run' wording was mine and wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…r the first integration pass (#501)

* feat(#499): the eight HydraNets read global land; total_lessons 40 for the first integration pass

Runbook #499 Step 1. REGION "africa_me_legacy" (13,110 cells) -> "land" (64,818 — the set the
FAO delivery cuts land_gaul from) in all eight config_queryset.py. total_lessons 300 -> 40 in
all eight hyperparameter files, commented as what it is: a run-time budget for the pass that
measures whether a HydraNet at global land fits fimbulthul, not a model choice; 300 (#463) is
restored by a follow-up PR once the pass has run. total_lessons is deliberately unpinned by the
roster tests for exactly this reason.

The roster test that pinned the region asserted `"africa_me_legacy" in <file text>` — a
substring, which the comment recording the history now satisfies, so it passed vacuously on
the new region. Replaced by test_global_land_region, which reads the REGION assignment line
and pins "land". Mutation: one model back to africa_me_legacy -> red, naming it.

Not changed: the FAO postprocessor's own queryset (still africa_me_legacy in git; #499 B5,
C-110), and every other config value on the eight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* review(#499): the lessons comment names its trigger and owner; violet_visitor's docstring stops saying africa_me_legacy

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(register): C-146 — a third occurrence: the region guard matched a substring, not the assignment (#501)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…factory fetch on current Python 3.11 (#502)

Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…, offsets 1/1 — global land does not fit the Africa+ME crop (#503)

* feat(#499): the eight HydraNets' volume is the whole PRIO-GRID raster — global land does not fit the Africa+ME crop

Runbook #499, found by the first pass on fimbulthul (2026-09-19 21:42): with REGION "land"
(#501) views-hydranet's DataSniffer refused every model in 9 s — "Spatial Span Violation! Data
spans 278.0x719.0, but volume resolution is 180x180". The four crop values row_offset 87,
col_offset 310, height 180, width 180 are the Africa+ME window. The sniffer's own docstring
names the global shape ("region=land in a 360x720 volume ... leading rows/cols zero, expected"):
row_offset 0, col_offset 0, height 360, width 720, on all eight — the roster test requires the
eight to share one grid topology, and they do.

Not measured here: memory. The volume is 8x the crop's area; training samples 32x32 windows
(window_dim), so the per-step cost need not scale with it, and views-hydranet's disk_guard
refuses an oversize posterior cube before allocating. The pass measures it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#499): row/col offsets are 1, not 0 — pipeline-core numbers the grid from 1 (the April fix, 9540a7b, applied to all eight)

Review: pipeline-core 3.3.0 dataloaders.py:540 builds row = (pgid-1)//720 + 1, so rows run
1..360 and cols 1..720; views-hydranet's VolumeHandler indexes row - row_offset and refuses
r_idx >= height. With 0/0 the last row/column of land overflows the 360x720 volume — or, if
no land cell sits on the edge, every prediction lands one cell off. 9540a7b set 1/1 for
heavy_freighter on this exact combination in April; f0b4436 later returned it to the crop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… — their predictions were labelled country_id and refused (#504)

* fix(#499): the eleven pgm darts models declare entity_id: priogrid_id — their predictions were labelled country_id and refused

Runbook #499. dark_river's calibration run on fimbulthul (2026-09-22) trained on 64,818
priogrid entities, then had all 13 evaluation sequences refused by pipeline-core's
CorePredictionSniffer:

  MultiIndex names ('month_id', 'country_id') do not match expected layout
  for level='pgm': ('priogrid_id', 'month_id')

views-r2darts2 reads the index name from the combined config and defaults it to "country_id"
whatever the level (darts_forecasting_model_manager.py:140, published 0.2.3) — right for the
31 cm models, wrong for these 11. The prediction VALUES were always priogrid cells; only the
label was wrong. Filed upstream as views-r2darts2#55 (derive it from level, as
dataset/base.py:934 already does); this declares it explicitly until that ships.

Nobody saw the refusal because pipeline-core evaluates in threads and discards their
exceptions: zero predictions written, run reported PASS (views-pipeline-core#529). That is
how three days of GPU produced metric frames and no data.

Guard: tests/test_darts_entity_id_matches_level.py — every r2darts2 model's effective
entity_id (declared, else the engine's country_id default) must match its level. 43 checked;
red when one of the eleven drops the key.

Full suite on pipeline-core 3.3.0: 7773 passed. ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* review(#499): the level lookup fails readably, and the third 'is it r2darts2' discovery names its trigger

Review of #504: EXPECTED[level] would have raised a bare KeyError if the platform ever gained a
third level (today impossible — test_config_completeness pins cm|pgm); now it says so. And this
is the third way the suite asks whether a model is r2darts2 — left as three, with the fourth
caller named as the trigger to extract one helper into conftest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…hers consume (#506)

* feat(#505): collapse posterior draws to the point predictions researchers consume (#506)

The eight HydraNets emit D x K draws per cell, not a number. `ensemble-updater`
takes one number per cell. Nothing currently bridges that.

The platform already defines the bridge: `inference_orchestrator.py:175` inverts
the log1p scaling, `:179` collapses the draw axis by the model's declared
`aggregate_method` (views-hydranet's ADR 021 / ADR 039 stage 5). Stage 5 is gated
on `evaluation_mode == "point"` and all eight run `stochastic`, so the draws reach
disk uncollapsed and nothing downstream folds them. `tools/collapse` performs that
stage, on the operator's machine, from the preserved numpy.

It honours the declared method rather than hard-coding one — the first draft
hard-coded the mean, which is the "two places to state one fact" pattern vmo_021
exists to stop. `test_collapse_declaration_matches_the_converter` pins the
converter's default against all eight models' `aggregate_method`, and against
their `evaluation_mode`, so the two cannot drift apart silently.

Verification:
- 28 tests: contract, input mutation, and real output already on this machine.
- 21 mutations applied to the converter, 21 caught, 0 survived. The first pass
  caught 19; the two survivors (lexicographic origin ordering, and a dropped
  identifier-length check) were real coverage holes, which is why the fixture
  now carries 13 origins rather than 3.
- Cross-checked against a pure-Python per-row recomputation over all 13 origins
  of all eight models' validation output: zero mismatches.
- Testing found two defects in the converter: a guard that raised ValueError
  while building its own error message, and a float32 accumulator.

Not validated: global land, the calibration partition, D x K = 32. Everything
measured here is Africa+ME validation at 16 draws, because that is what exists
locally. ADR-023 says so in Validation & Monitoring.

ADR-023 records the decision; C-152 records the coupling it accepts (a third
reader of an on-disk layout no repo publishes). Both 021s are cited, so the
README's cross-repo prefix rule now covers vhy_021 / vmo_021.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#505): the roster test imports tools.collapse lazily — roster-configs-load has no pandas

`roster-configs-load` installs views-hydranet and nothing else, then imports
tests/test_roster_conformance.py through tests/test_roster_configs_load.py.
A module-level `from tools.collapse.collapse_predictions import ...` pulled
pandas into that job's collection and failed it at import, before any test ran.

The import now sits inside the one test that uses it, with a comment saying why
it must stay there. Verified by importing the module with pandas blocked from
sys.meta_path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#505): the scale guard could not fire on a single corrupted target

Code review (5 agents) found the one defect that mattered, and it was a guard
that could not fire.

`_check_scale` took the maximum of the three targets FLATTENED TOGETHER. A
scaler-registry mismatch upstream does not have to hit all three targets, and
when it hits one, a healthy sibling carries the combined maximum over the
threshold. The corrupted target then ships as log1p(count) — a plausible-looking
parquet with a silently wrong column, which is the exact failure this guard was
written to prevent.

Its test could not detect the difference either: it set all three targets to
log-space values at once, so it passed under the broken implementation and the
correct one alike. A test that cannot distinguish the two designs is not
evidence about either.

- the check is now per target, and names the offending column
- new test puts ONE target in log space with healthy siblings; a mutation that
  restores the flattened form is now caught (M22)
- added a duplicate (month_id, priogrid_id) guard: ensemble-updater joins on that
  pair, and every target agreeing on a duplicated identifier is still a duplicate,
  so cross-target alignment is blind to it (M23)

Also from review:
- module docstring named `--max-plausible`; the flag is `--min-plausible-max`
- bare "ADR-069" cross-repo citation qualified to views-hydranet `vhy_069`, per the
  README's own prefix rule
- the docstring described only the mean; the module honours `median` too
- docs/ADRs/README.md listed `[ADR-021]` while the paragraph directly below it
  documents that 021 collides across repos — now `[vmo_021]`, applying the
  convention the PR itself introduces
- `test_collapse_declaration_matches_the_converter` moved out of TestGridAndTargets
  (documented as "Region grid and target channels") into its own class

Verified: 8037 passed, ruff clean, 23/23 mutations caught with 0 survivors and
0 unrun, tree byte-identical after the campaign.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ation run (#507)

* feat(#499): the eight HydraNets train 300 lessons for the real calibration run

40 was the integration-pass value set by #501 to make the first global-land pass
cheap. It is not a training length: measured on a rented RTX PRO 4500 SE, one
lesson is 84 seconds, and the bias against observed fatalities is roughly twice
as bad at 160 lessons as at 300 (0.05 vs 0.13 predicted/observed total, measured
on the validation output already on the operator's laptop).

Cost of this line, measured rather than guessed: 300 lessons is ~7 h of training
plus ~1 h of evaluation per model, so ~8 h and ~$6 per model at $0.75/hr, and
~$48 for all eight.

`n_head_samples` deliberately NOT raised. The plan had been D x K = 4 x 8, but
`ensembles/rusty_bucket/configs/config_hyperparameters.py:16` pins
`expected_samples_per_model: 16`, and its own comment records that the pool was
thinned from 1024 to 128 because the full-S run peaks at ~28.6 GB. Raising K to 8
would break that pin, double each model's prediction volume from 13 GB to 26 GB,
and lengthen evaluation — to buy finer quantile resolution that the estimator
analysis did not need at 16 draws. If the ensemble's memory behaviour is fixed
later, K and `expected_samples_per_model` move together, in one PR, not this one.

Verified: full suite green (7997 passed), roster conformance, ensemble configs
and the ADR-015 sample-count report all unaffected, since only the lesson count
moves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#499): the comment above total_lessons said the restoring PR was still owed

Code review (5 agents) on #507. Two findings, both about text that would mislead
the next reader rather than about the value itself.

1. The comment directly above the changed line, identical in all eight files,
   read: "40 for the eleven-model calibration pass ... Production is 300 (#463);
   the PR that restores it is owed by whoever ticks that pass on #499." After
   this change it lied twice over — it named a value the line no longer holds,
   and it described as owed the very work it sits on top of. Rewritten to say
   what is true now, and to record the measured cost (84 s per lesson at global
   land, so ~7 h per model) so the number is not a mystery to whoever reads it
   next. The stale "eleven-model" phrasing, which belonged to the darts group and
   not to these eight, goes with it.

2. `run_integration_tests.sh`'s 1800 s default was sized when these models
   trained 40 lessons. At 300 they need ~7 h and WILL report TIMEOUT. That is a
   budget, not a regression, but a person reading a wall of TIMEOUT rows would
   reasonably conclude otherwise. The header and --help now say so and give the
   flag. Comments only; no behaviour changed.

Two corrections to my own claims in the PR description, both raised by review:

- I called the quality evidence "measured, not guessed". It is weaker than that.
  The 0.05-vs-0.13 predicted/observed figures compare 160-lesson and 300-lesson
  models on Africa+ME VALIDATION output, not the 40-vs-300 before/after this diff
  actually makes, and not at global land. No 40-vs-300 ablation exists. It is a
  reasonable proxy — 40 was never claimed to be adequate — but the framing
  overstated how directly it was measured. The 84 s per lesson figure IS directly
  measured, on the pod, and is the only number here I would defend as such.

- The gate the code comment named was the eleven-model fimbulthul pass, which
  never formally closed. The evidence offered instead is a fresh end-to-end pod
  run. Same intent — prove it fits, measure the cost — but different evidence
  than the gate specified, and that substitution should be visible rather than
  glossed.

Also merges origin/development (#506, the converter) so the branch is current
and the suite runs against the merged state.

Verified: 8037 passed, ruff clean, bash -n on the runner.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(#499): the runner's CIC records that a HydraNet TIMEOUT is now the budget

`cic-sync` was right to fail. `run_integration_tests.sh` is CIC-governed by
docs/CICs/IntegrationTestRunner.md (ADR-006), and my change there was comments
only — so `[cic-skip]` in the PR title would have passed the gate.

That would have been the wrong way through it. The contract document is where
someone looks to find out what the runner guarantees, and "the default timeout no
longer fits eight of the models in the roster" is exactly that kind of fact. Put
in the title, it would have been invisible a week later.

So the CIC now carries it: why 1800 s was right at 40 lessons, why 300 lessons
means ~7 h of training plus ~1 h of evaluation per model, the flag to use, and
why the default is deliberately NOT raised — it suits every other library, and
raising it globally would turn a genuine hang in a cheap model into a half-day
wait.

The two tables that a reader hits first — the `--timeout` row and the "Model
exceeds timeout" failure-mode row — now both point at that note, because a wall
of TIMEOUT rows is the exact shape a real failure takes and nobody should have to
guess which one they are looking at.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…nsit, and we put it on rented hardware (#510)

The datafactory speaks plain HTTP. HTTP Basic sends the credential base64-encoded
on every chunk request, and base64 is encoding, not encryption. views-datafactory
accepted that as their C-318 when the audience was a trusted circle on trusted
networks, which on fimbulthul was reasonable.

On 2026-09-28 we changed the audience without changing the mechanism: the
credential went onto five rented machines in datacentres we do not control. The
operator chose knowingly to use his personal login rather than provision a
throwaway, having been told what it meant. That is recorded, not second-guessed —
the work was owed and the server was gone.

What makes the residual risk outlive the run is two properties of the credential
itself: no expiry and no per-host registration. It stays valid until a person
rotates it by hand, and it authenticates from anywhere. So a pod image, a
snapshot or a detached volume that survives a campaign carries a live permanent
credential, and only housekeeping closes that.

Tier 3: a real but unquantified interception risk rather than a demonstrated
compromise, with cheap known mitigations. The trigger is deliberately the general
case — any run on hardware we do not own — not RunPod specifically, because the
next one may be a collaborator's machine or a CI runner.

Mitigations in preference order: a throwaway login retired at campaign end (about
three commands for whoever administers the data server); TLS on the data server,
which removes the class entirely; or deleting rented volumes and images at the
end and rotating afterwards.

Raised by the views-datafactory session reviewing the RunPod post-mortem (#508).
Cross-refs C-151, views-datafactory C-318, and #509 — the client floor still
permits a version that carried this credential across redirects to other hosts.

Header: 160 -> 161 entries, 151 -> 152 concerns, Accepted 4 -> 5, T3 64 -> 65.
Verified: 8036 passed, register header tests green.


Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…rator guide (#508)

* docs(#499): the first RunPod deployment — post-mortem, cost note, operator guide

Three documents from one effort, split by audience because they were fighting
each other in a single file. Plus `tools/podrun` v0.1.0, and a correction to a
number that #507 merged.

- reports/postmortem_runpod_first_deployment_2026-09.md — what happened and what
  must change. Owns the narrative, the misdiagnoses, and the failures.
- reports/runpod_cost_and_time_note_2026-09.md — what it cost. Written for a
  manager, cost and time only, one clearly-labelled extrapolation.
- docs/runpod_run_guide.md — how to do it. Owns the procedure and the gates.
- tools/podrun/ — the runner, marked v0.1.0 PROVISIONAL in four places, with the
  conditions for promoting it out of 0.1.0 stated.

Each fact has one owner; the others cross-reference. The guide is the how, the
post-mortem is the why, per fao_delivery_runbook.md's precedent.

THE CORRECTION. #507 merged "one lesson is 84 s, so 300 lessons is ~7 h" into the
eight config comments, run_integration_tests.sh and the runner's CIC, stated as
measured fact. It is wrong. The 84 s came from the FIRST lesson of a cold
two-lesson smoke run, which carries warm-up and is not representative of the
other 299. Three completed models give 202, 253 and 272 minutes end to end —
40-54 s per lesson, ~4 h per model. Corrected in all three places, with the
error named rather than quietly overwritten. No operational harm: the
recommended --timeout 30000 was over-provisioned against the wrong number and
remains safe against the right one.

This repeats, within hours, the lesson the post-mortem's own §2.6 records —
never characterise against a throwaway model's output. It is now §7 item 8.

REVIEW. Three peer sessions reviewed independently and each found something real.

views-hydranet: the guide stated two different runtimes (the above); the
disk_guard problem is worse than described — even taught to read the cgroup it
counts only the posterior cube, missing the ~2.3 GB input volume and the torch
context, so it would still under-report peak by 2-3x; the 2 steps/s threshold is
pod-relative, not a hardware expectation (a 4070 laptop does ~15); both q95
figures were measured at S=16, so the claim is stability across datasets, NOT
across sample counts; and C-151 must be cited as views-models C-151 because
views-hydranet's C-151 is an unrelated entry.

views-faoapi: "the upload can run from a machine Simon controls" was a procedure
claim I was not entitled to — the architecture supports it (UPLOAD_ENABLED
defaults False, no store client constructed when disarmed) but there is one
entrypoint, no --no-upload, and disarming edits a committed delivery
declaration. Restated as architecture + open question. Also: the ensemble and
delivery steps are per-run and memory-bound, so the per-model extrapolation is
silent about them; and the two documents disagreed on how many models were
complete.

views-datafactory: the preflight check I wrote passes only because urllib follows
a 308 and carries the Authorization header across it — the exact pattern their
#388 removed. Verified on a live pod: bare URL 308, /.zmetadata 200, and
Authorization sits in req.headers not unredirected_hdrs. Now points at
.zmetadata. Also raised the client floor to >=1.13.0, where the
credential-handling fixes landed.

Verified: 8036 passed 0 failed, ruff clean, validate_docs.sh passes,
test_tools_layout passes with the new group. Downloaded predictions deliberately
not committed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(podrun): the draws archive could be empty and still report success

Bug review of tools/podrun/pod_run_model.sh before merging #508. Eight
successful runs had exercised the happy path; this is about the paths they did
not.

CRITICAL, and verified empirically rather than reasoned:

    find <no matches> -print0 | tar --null -T -   ->  exit 0, valid 22-byte archive

So compress_draws could write an empty posterior and the script would proceed to
STATUS=OK. The only other signal was a small number in MANIFEST that a human had
to notice. On the one model family this has run the find pattern matches; on the
next one it might not, and `__init__.py` already warns that no other family has
been through it. Now counted before and verified after:

  - refuse if find matches zero lr_* draw files
  - after tar, list the archive and require the entry count to equal the count found

HIGH: two config guards checked file TEXT, not the parsed value. A regex took the
first `'total_lessons': <digits>` anywhere in the file and a grep matched
`REGION = "land"` anywhere including comments. That is the defect this repo
shipped once already — #501's "the guard that was not one", where a substring
assertion was satisfied by a comment recording the region's history. Both now
import the config module and read the value the interpreter sees. Verified: reads
300 and "land" from violet_visitor, and the file does contain "300 lessons" in a
comment, which is the surface a text check exposes.

HIGH: `SRC=$(ls -d ...predictions_calibration_*)` had no existence guard. An empty
result meant `cd "$SRC/.."` -> `cd /..`, i.e. the filesystem root. It happened to
fail loud because find errors on an empty path argument, which is an accident to
depend on. Now refuses explicitly.

MEDIUM: STATUS was never cleared at start, so a stale FAILED from an earlier
attempt shadowed a fresh run for its whole multi-hour duration -- and the script's
own header advertises `cat STATUS` as the way to watch from outside. Cleared now.

MEDIUM: output directories were not cleared between attempts. Parquets are named
from the SOURCE run's timestamp, so an old set and a new set could coexist; if
they summed to 13 the count check would pass while the manifest covered two
different training runs. Both output dirs are now cleared per attempt.

MEDIUM: the model and converter existence checks had no stage of their own, so
die() reported whatever stage ran last -- FAILED:verify_env for "no such model".

MEDIUM: die()'s explanation went only through the tee subshell, which can lose
its last buffered lines on a hard kill, at exactly the moment the reason is
wanted. The reason is now also written directly to $OUT/FAILURE, as STATUS and
STAGE already were.

LOW-MEDIUM: no lock, so two invocations for the same model on one pod raced on
the clone, the log and STATUS. I nearly caused this myself today when a chain and
a manual assignment could both have taken bright_starship. Now takes $OUT/.lock
and releases it on exit.

Known and NOT fixed, recorded in __init__.py and the post-mortem's open items:
the script has no automated tests, and MANIFEST's git sha and lesson count are
unchecked, so they would ship blank rather than refuse. Both are accepted at
v0.1.0.

Verified: all three guards fire on constructed inputs (empty archive refused;
counts match on real files; empty SRC refused), the config check reads real
values through the interpreter, bash -n clean, ruff clean, test_tools_layout
passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ers (#522)

* fix(#439): bump VIEWS_POSTPROCESSING_PIN 1.1.1 -> 1.4.0 in both launchers

Closes #439. Two lines, one value, and it is the only thing standing between the FAO
findability fix and a pod.

WHY THIS IS NOT A NO-OP. The launcher does NOT install views-postprocessing from PyPI. It
installs from a git ref:

    tools/launcher/postprocessor.sh:132
    pip install "git+https://${GITHUB_TOKEN}@.../views-postprocessing.git@${VIEWS_POSTPROCESSING_PIN}"

and refuses the run at :164 if the installed ref does not match the pin. So
views-postprocessing#314 being merged does nothing for a pod, and releasing it to PyPI
(1.4.0 is there) does nothing either. Verified empirically rather than reasoned: the
2026-09-29 pod ended up with views-postprocessing 1.1.1 while PyPI was at 1.3.0.

This was invisible because the two halves of the same incident reach a pod by DIFFERENT
mechanisms — views-pipeline-core transitively from PyPI through views-hydranet's dependency
range, views-postprocessing from a git tag by exact pin. "Is the fix released?" therefore
has two different correct answers depending on which half is meant, and nothing states
that anywhere. views-postprocessing#310 is the same family.

WHAT 1.4.0 CARRIES: #314, the C-94 findability guard that verifies every artefact by name
rather than trusting a returned file id — the defect that let the 2026-09-29 delivery
report "Postprocessor Run Completed" while being unservable; and #315, a guard deriving
the port's documented datastore contract from source so the prose cannot drift from the
code.

VERIFIED BEFORE COMMITTING, not assumed:
  - tag 1.4.0 resolves on the remote the launcher actually pulls from (8db8c9fb)
  - no other reference to 1.1.1 remains in the repo outside historical reports
  - the ref-vs-pin guard at postprocessor.sh:164 still reads the same variable
  - `bash -n` clean on both files

BOTH LAUNCHERS, deliberately. un_crafd carries the identical pin and the identical
exposure; leaving CRAF'd on 1.1.1 would fix one partner delivery and silently leave the
other on the version with the defect.

NOT pinned to `development`, which would have been the lazy choice and is worse than a
moving target here: the :164 guard compares the installed ref against the pin STRING, so a
branch name satisfies the check designed to catch staleness while installing whatever
landed most recently. It would pass the guard and still be unreproducible for a
partner-visible delivery.

Expect a behaviour change: a delivery that previously completed can now stop. 1.4.0 adds
two refusal situations and no new exception types. A DeliveryNotFindableError names every
object that does not resolve and states explicitly when nothing in the run is servable.

Edited under an explicit lift of the standing "do not modify run.sh" rule, granted for this
change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(#439): record WHY the pin moved to 1.4.0, in both launchers

/review-diff on the pin bump, and the finding is one the value hid: the rationale blocks in
both launchers document 1.1.0 and then 1.1.1 in detail, and stopped there. The value said
1.4.0 with no recorded reason.

These blocks are not commentary. `tools/launcher/postprocessor.sh:26` says "EVERY STEP
BELOW IS A SCAR. Read the comment before reordering anything." A future reader would have
found a version two moves ahead of its own history and no way to learn what it carries.

Which is the exact pattern the views-postprocessing session named against itself twice
today — fixing the visible half of a finding and leaving the half that bites. The value is
the visible half. The scar record is the half that explains it.

un_fao now records the 2026-09-29 incident as the reason: the findability guard in 1.1.1
checked TWO artefacts out of the 110 a run uploads, and by returned file id rather than by
name, so the first-ever FAO delivery uploaded 109 of 110, logged "Postprocessor Run
Completed", fired a success alert and was refused by views-faoapi. It also records that the
producer half of that defect is views-pipeline-core#552, shipped in 3.3.4 and reached from
PyPI, while this pin is the other half and is reached from a git tag — which is precisely
why "is the fix released?" had two different correct answers and this line was missed for
hours.

un_crafd records the same, framed by the shared-prefix argument its own block already makes
twice: both launchers install into one conda prefix, so moving only the leg that failed
downgrades the environment out from under the armed FAO delivery (C-139). CRAF'd runs the
identical `_assert_delivery_is_findable`, so the defect is byte-identical there and has
simply not fired on that leg yet — the same sentence the 1.1.1 note had to write about #268.

Also replaced the verification recipe. The old one greps for `success is not True`, which
verifies the 1.1.0-era C-79 fix and is still correct but says nothing about what 1.4.0 adds.
The new one greps `DeliveryNotFindableError` in `delivery/findability.py` — verifying the
build by what it REFUSES rather than by what it claims, which is the lesson of the whole
incident.

No behaviour change in this commit; the pin value is unchanged from the previous one.
ruff clean, 8037 passed, `bash -n` clean on both files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…at delivered, and the FAO chain is a script (#524)

* fix(#516): pin xarray and pandas — the fresh env built successfully and WRONG

#516 said a fresh un_fao environment is unbuildable. Measured on 2026-09-29, it is
not: it builds, on a different pandas MAJOR, in silence.

    declared:  views-datafactory>=1.9.0,<2.0.0 + numpy>=1.26.4,<2.0.0
    resolved:  numpy 1.26.4, xarray 2025.12.0, pandas 3.0.6
    the pod that produced the first FAO delivery: xarray 2024.3.0, pandas 1.5.3

Neither xarray nor pandas was pinned anywhere. The working versions existed only as
hand-pins on a pod that has since been destroyed, so a fresh pod silently picks
different ones. A build failure is loud; this is not, which makes it the more
dangerous of the two shapes the issue conflated.

views-datafactory requires xarray outright and pandas only in an optional extra that
is not installed, so pandas arrives THROUGH xarray. xarray is the sole carrier —
datafactory's own pyproject comment says so — and pinning the carrier fixes it.

The bound is measured, not guessed. xarray 2024.3.0 is the LAST release that accepts
pandas 1.x; 2024.5.0 moved to pandas>=2.0 and 2024.9.0 to pandas>=2.1. A tidy-looking
`<2025` cap would therefore have been WRONG: the cliff is inside the 2024 line, not at
the year boundary. Verified by resolving the file, not by reading version numbers.

Declared identically in both postprocessors because they share one prefix (C-116);
pinning only one lets whichever runs last decide pandas for both, which is the same
trap the existing numpy pin already documents one line above.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#517,#523): install the appwrite extra, prove it at preflight, add --rehearsal

Three findings from one /falsify audit of "we are ready for a new 40-lesson run"
(FALSIFIED). All three had one cause: every fix that made the 2026-09-29 run work was
applied BY HAND to the pod, and the pod was destroyed. The pod was the artefact and
nothing in git described it.

#517 — the publish path was never installed. The string "appwrite" appeared nowhere in
pod_run_model.sh, so `_build_datastore` raised at publish time, AFTER the full training
run. That is exactly how the first FAO delivery attempt failed. The extra is now
requested explicitly with no version, so views-hydranet's own range still decides which
pipeline-core is installed, and the client is IMPORTED in preflight — the whole script
is written to fail early, and the one dependency that failed late was the one not
checked there. Verified by a real resolve: appwrite 13.6.1 alongside views-hydranet
0.1.2 and views-pipeline-core 3.3.4, with pandas 1.5.3 and numpy 1.26.4 intact. The
same resolve confirms the C-151 toolz override is necessary rather than folklore —
without it the resolver lands on toolz 0.11.2.

#523 — a deliberate cheap run was reachable only by deleting the guard against an
accidental one. The >=300 floor stays; `--rehearsal <lessons>` is the escape hatch. It
takes the count on the command line, because main.py has no hyperparameter override and
a flag without a count still forces an edit to a tracked config. It patches the POD's
clone only, then re-imports to CONFIRM the patch took — a substitution that silently
missed would give a 300-lesson run wearing a rehearsal label, or the reverse.

The marking matters more than the permission. A rehearsal's parquets are structurally
identical to a production run's and pass every check including the runner's own, so a
REHEARSAL file is written beside them and MANIFEST gains `mode:` and `config: PATCHED
after checkout` — the git sha alone no longer describes the run.

Found while reviewing this change, not in the audit: section 2 does not re-clone when
.git exists, so a production run on a pod that had rehearsed would read the LEFTOVER
patch. The floor still refused it, but blamed the committed config and advised
--rehearsal — a correct refusal for the wrong reason, sending the operator to edit the
wrong file. It now detects the dirty config and names it.

Deferred, and it is the weaker half: this runner MARKS a rehearsal but cannot REFUSE to
publish one, because the publish step is not in this script. That guard belongs in
views-pipeline-core. An xfail(strict) stub for it was written and DELETED — its regex
spanned the whole file under re.S and XPASSed by accident, i.e. the exact class of
guard-that-cannot-fire this audit exists to find.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#518): the guards were decorative — execute the script, and write the missing warning

An independent /falsify guard-mode audit of 53bfba9 reverted EVERY fix in that commit,
kept the comments, and got 22/22 green. 20 of 24 mutations survived. 10 of 14 guards
were DECORATIVE, 4 WEAK, none held.

The mechanism is the lesson. Every assertion read pod_run_model.sh as TEXT, and that
script is unusually well commented — each fix carries a paragraph naming the incident and
quoting the exact strings. So THE BETTER THE COMMENT, THE WEAKER THE GUARD: deleting the
code left the comment, and the comment satisfied the assertion. Deleting the appwrite
extra from the install line was invisible because the comment above it names
views-pipeline-core[appwrite].

Worst survivor: `os.environ.get("REHEARSAL_LESSONS") or ""` -> `or "40"`. One word. Every
production run then patches itself to 40 lessons, the floor is dead, the leftover-patch
detector is dead, and because the SHELL variable stays empty the MANIFEST says
`mode: production` and no REHEARSAL marker is written. A 40-lesson model labelled a
production delivery, suite fully green.

AND ONE GUARD CERTIFIED A SAFETY PROPERTY THAT DID NOT EXIST. It claimed the RunPod guide
documents the /workspace chmod trap. The guide did not: no "world-readable", no "network
filesystem", no C-154, no #518. It passed because `/workspace` is on line 104 and `chmod`
on line 180, joined by `.*` under re.S — the identical defect the previous commit message
boasted of having found and deleted elsewhere. That is worse than no guard, because it
stopped the next reader looking. The warning is now written (#518, C-154): /workspace is a
network filesystem where chmod 600 returns success and does nothing, leaving a credential
at mode 666 with no error to notice.

What changed in the tests:
  - the config-check program is EXTRACTED FROM ITS HEREDOC AND EXECUTED against fixture
    models in real git repos. That block holds all the rehearsal/production logic and was
    wholly unguarded. Includes an adversarial fixture whose literal is patchable but whose
    get_hp_config() returns 300 regardless, so a "verification" that compared target
    against itself is caught.
  - remaining text assertions read COMMENT-STRIPPED source, so a comment can never stand
    in for code.
  - requirements guards assert SPECIFIER SEMANTICS via packaging.SpecifierSet — is 1.5.3
    admitted, is 3.0.6 refused — instead of pin presence. `pandas>=3.0` passed the old one,
    and marker-gated pins (`; python_version < "3.10"`, inert on 3.11) defeated the
    identical-pins check while keeping the captured text identical.
  - the roster guard CALLS get_hp_config() instead of matching the literal, which is what
    pod_run_model.sh itself insists on, citing #501 "the guard that was not one".

pod_run_model.sh warned against this exact failure in three places and the guards did it
anyway. That is the argument for the independence rule: the author's own model named the
trap three times in one commit and stepped into it.

35 tests, up from 22.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#509): floor views-datafactory at 1.13.0, and guard the shell arg parsing

Two gaps found by replaying the guard-mode audit's mutations against the rewritten
guards — the replay is the point, since the previous version passed 22/22 with every fix
reverted.

#509 — the pod installed views-datafactory>=1.9.0. The credential-handling fixes landed
in 1.13.0: before it the client could carry a netrc credential across a redirect to
another host and embed it in error messages. This script only ever runs on hardware we do
not own, carrying exactly that credential, and the guide already prescribes >=1.13.0
explicitly while the script did not. A resolver picks 1.13.0 anyway today, but "the
resolver will probably do the right thing" is precisely the reasoning that put xarray
2025.12.0 and pandas 3.0.6 into a fresh environment (#516). Verified: the floor still
resolves to 1.13.0 with hydranet 0.1.2, pipeline-core 3.3.4, appwrite 13.6.1, pandas
1.5.3.

The shell argument parser was unguarded. Executing the config-check program does not
cover it — the flag never reaches Python if the shell refuses it first. Deleting the
`--rehearsal)` case arm removes H1's escape hatch entirely and left every guard green,
because the string survives in the header comment and in USAGE. Now executed against the
real script: the flag and its count must be CONSUMED (distinguished from "unknown option"
by which error appears), a missing count must be refused, a non-integer must be refused,
and — the control, without which the first assertion proves nothing — a genuinely unknown
option must still be rejected. Argument parsing precedes `mkdir -p "$OUT"`, so these cases
exit before touching the filesystem.

MUTATION REPLAY — every surviving mutation from the audit is now caught:

  X7  `or ""` -> `or "40"` (production runs patch themselves)        3 failed
  X8  --rehearsal becomes a bare flag                               3 failed
  X22 the --rehearsal case arm deleted                              3 failed
  X1  extra dropped from the install line, comment kept             1 failed
  X2  both appwrite imports commented out, prose comment kept       1 failed
  X5  2026-09-29 install order restored + reassuring comment        1 failed
  X6  the floor prints instead of exiting                           1 failed
  X9  leftover-patch detector flipped to the wrong branch           1 failed
  X10 patch verification becomes `if target != target`              1 failed
  X11 the marker-writing block deleted                              1 failed
  X12 stale-marker clear moved to the end of the run                1 failed
  X13 MANIFEST mode branches swapped                                1 failed
  X14 both files pinned to xarray 2025.12 / pandas 3.0              8 failed
  X15 upper bounds dropped                                          8 failed
  X17 pins marker-gated inert on 3.11                               6 failed
  X18 get_hp_config() returns 40, the 300 literal untouched         1 failed
  X20b toolz override replaced by a comment                         1 failed
  control: the guide's chmod warning deleted                        1 failed

X16 (`xarray == 2024.3.0`, a stricter and CORRECT pin) previously failed with a message
claiming no pin existed; it now passes, so the guard no longer raises a false alarm
against a legal PEP 508 spelling.

Two honest notes. X18 was recorded SURVIVED in the audit but had not applied — the config
returns a dict literal, so the mutation's sed matched nothing. Re-applied correctly, it is
caught. And X19 (replacing the credential chmod with an unrelated one) now passes, because
the /workspace warning it was meant to defeat genuinely exists; the control above proves
the guard fires when that warning is removed.

40 tests, up from 35.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#509): actually apply the 1.13.0 floor the previous commit only tested for

The previous commit shipped the guard and not the fix. The sequence: edit the install
line, write the test, verify the test fires by reverting the line with sed, then
`git checkout -- tools/podrun/pod_run_model.sh` to undo the sed — which reverted to the
last COMMIT, discarding the uncommitted 1.13.0 edit along with the sed. The commit then
captured the test plus a 1.9.0 install line.

Caught by that test on the next full-suite run, which is the whole argument for writing
guards that execute: a text-matching guard for this would have been satisfied by the
comment above the line, which names 1.13.0 and explains why.

Also worth recording, because it made the failure hard to see for two runs: this suite
runs pytest-randomly, so ordering varies between runs. Two earlier runs reported 8058
passed/209 skipped and 7758 passed/470 skipped/3 failed with an unchanged tree, and the
three failures were ordering interference in pre-existing tests, not in anything here.
With `-p no:randomly` the counts are stable at 8076 passed/209 skipped/5 xfailed and the
only failure was this one. Verify against a fixed order before concluding a change is
responsible for a red suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* feat(#499): script the FAO delivery chain — it existed only as commands typed on a pod

The fourth instance of the pattern the /falsify audit found three of, and the one that
mattered most. On 2026-09-29 the whole FAO delivery — eight forecasting runs, the
rusty_bucket pool, the publish, the un_fao postprocessor — was typed by hand on a rented
pod. It worked, the pod was destroyed, and nothing recorded what had been done.

The audit missed it because it audited the script that exists (pod_run_model.sh) instead of
asking what the delivery needs. pod_run_model.sh runs `-r calibration -t -e` and its header
is accurate that it uploads nothing: it is Track A, the research frames. Nothing in tools/
drove Track B at all — `grep -rln "prediction_store\|r forecasting" tools/` returns one
unrelated file.

This is not a stopgap. fimbulthul, which #499 names for every Track B step, has been
unreachable since 2026-09-25, so a pod is the only hardware there is.

pod_run_model.sh gains `--forecast`: `-r forecasting -t -f`, and it SKIPS sections 4 and 5.
Those are calibration deliverables — a forecast has one origin, so the 13-parquet count
check would refuse it, and what the FAO chain consumes is the pooled ensemble output. Kept
in this script rather than a copy because everything before section 3 is identical and
already exercised; a second copy would be a second place for C-151 to rot.

pod_run_fao_delivery.sh drives the chain. Two details that are not obvious and that cost
real time to establish:

  - `-sa/--saved` on the ensemble is REQUIRED. Without it rusty_bucket refetches instead of
    pooling the member forecasts just written, and the eight runs are wasted.
  - the two legs need DIFFERENT INTERPRETERS. This pod builds a uv venv;
    tools/launcher/postprocessor.sh uses `conda shell.bash hook` / `conda create --prefix`.
    A pod can satisfy one and not the other, and step 4 is last — so a missing conda kills
    the delivery after every GPU hour is spent. Preflight now checks for it.

`--preflight` is the mode that protects the money. At 300 lessons step 1 alone is ~16 GPU
hours, and everything it checks is knowable in seconds: checkout, venv, netrc, the three
publish secrets, conda, that un_fao's REGION resolves to land_gaul rather than a disarmed
value (C-110), GPU, disk. It accumulates and reports every problem rather than dying on the
first, because a round trip per problem is billed by the second.

The publish secrets are checked HERE and not at first publish because
views-pipeline-core's PredictionStoreConfig claims to read them "once at startup and fail
loud … preventing silent failures after hours of training" and does not — it is called from
_build_datastore, after training (views-pipeline-core#557). Until that moves, this preflight
is the only check that happens before the money is spent.

A rehearsal is marked at every level, and the marker says the uncomfortable part out loud:
nothing downstream refuses a marked rehearsal (#523), so its undertrained forecasts DO reach
the FAO shelf and ARE servable. That is deliberate — it is the only way to test the chain —
and it is why the final stage reads back what landed with `tools.liveness`, by name, rather
than inferring success from an exit code. On 2026-09-29 the publish reported success having
written nothing (C-155).

19 guards, written after the lesson that 10 of 14 of the previous set were decorative. The
parser and `--preflight` are EXECUTED; the remaining text assertions read comment-stripped
source and cover only ordering facts that cannot be executed without a GPU, a store and a
partner bucket. PODRUN_ROOT exists so preflight is exercisable off a pod; it relocates the
workspace and relaxes no check.

Every preflight check was proven able to fire, including by shadowing nvidia-smi with a
failing stub and stripping conda from PATH — a check that cannot fail is the defect this
repository keeps finding.

Two defects in this script were found by its own tests: an unchecked `mkdir -p "$OUT"` that
let it run on into a broken tee off a pod, and a lock message that reported "another
delivery is in progress" for a directory that simply was not writable.

What these guards do NOT cover, stated rather than implied: nothing here runs a forecast,
publishes, or reaches the FAO. The first real exercise is a rehearsal on a pod.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* docs(#499): Phase 4b — the FAO delivery chain, and how to read its outcome

Phases 4 and 5 document Track A (calibration predictions for research). The FAO delivery
had no operator documentation at all, because until now it had no script.

Leads with --preflight, because the two things most likely to stop the delivery are
invisible until the end: the three Appwrite publish secrets, and conda — the postprocessor
launcher requires it while the pod builds a uv venv, so a pod can satisfy the training leg
and not the delivery leg, and that failure lands after ~16 GPU hours.

Records two readings an operator would otherwise get wrong: DeliveryNotFindableError is
1.4.0 WORKING (it verifies by refusing, and names what it checked), and a rehearsal's
forecasts genuinely reach the FAO shelf and are servable rather than being contained.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#509): floor views-datafactory in the postprocessor legs too, not just the pod install

Found re-running the clean-checkout simulation after the delivery script landed: I floored
views-datafactory at 1.13.0 in tools/podrun/pod_run_model.sh and left both postprocessor
requirement files declaring >=1.9.0. Half a fix, and the half I left out is the leg that
runs LAST — so it would have been the one still holding a client without the credential
fixes at the point the delivery reaches a partner.

The reasoning that applied to the pod install applies here unchanged: before 1.13.0 the
datafactory client could carry a netrc credential across a redirect to another host and
embed it in error messages. The postprocessor runs on the same rented hardware and fetches
the actuals it curates land -> land_gaul against, so it holds that credential too. Nothing
about "this is the postprocessor" makes that safer.

Declared identically in both files for the same C-116 reason as the pins above them: one
shared envs/views-postprocessing prefix, so whichever postprocessor runs last decides the
version for both, and flooring one is the same as flooring neither.

The guard is widened rather than duplicated: it now iterates the pod install and both
requirement files, and was verified to fail independently for each of the three when that
source alone is reverted to >=1.9.0. A guard that passes because two of three sources are
correct is how this got shipped half-done in the first place.

Resolve unchanged: views-datafactory 1.13.0, xarray 2024.3.0, pandas 1.5.3, numpy 1.26.4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#509): apply the postprocessor floor the last commit only tested for — again

Second time in one session, same mechanism, and it is worth recording rather than quietly
amending: the previous commit shipped the widened guard and not the fix, because I verified
the guard by reverting each source with `git checkout -- <file>` while the fix was still
UNCOMMITTED. The checkout reverted to the last commit and took the fix with it.

My own ship-it procedure says mutation-verification runs AFTER the commit and before the
push, for exactly this reason — "the tree is clean, so a mutation can be applied and
reverted with git checkout --". I violated it twice in a row, once for the pod install line
and once for both requirement files. The order is the control, not the care.

Both files now declare views-datafactory>=1.13.0, with the reasoning restored: before
1.13.0 the client could carry a netrc credential across a redirect to another host and embed
it in error messages, and the postprocessor runs on the same rented hardware holding the
same credential. It is also the LAST leg — the one that reaches the partner.

The guard is hardened while I am here: it now strips comments from the requirements files
too, not only from the shell script. un_fao's xarray note quotes the old
`views-datafactory>=1.9.0` line as the state it is explaining, so a guard reading raw text
either trips on prose or has to be written carefully around it — and being careful around
comments is precisely how the decorative guards got shipped in the first place.

Verification order corrected: committed first, mutation replay follows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#509): record the views-datafactory split as a deferral, and close the hole it opens

The repo's own hygiene guard caught the last commit: flooring views-datafactory to >=1.13.0
in the two postprocessor files while 34 model requirements stay at >=1.9.0 makes one package
declared two ways, which C-116 says is then decided by run order rather than intent. A good
guard, and it fired on a real inconsistency I introduced.

Not unified, deliberately, and the reasons are asymmetric:

  - the 34 are model requirements unrelated to the delivery legs. Raising them is #509's own
    scope, not a delivery PR's, and it would put 34 unreviewed one-line edits in a branch
    about reproducing a pod.
  - one of them is `models/violet_visitor/requirements.txt`. Another session owns that
    model's contents and this session is instructed not to edit them. That instruction is not
    mine to set aside because it would be convenient for a guard.

So: DEFERRED_PACKAGES, which the guard offers explicitly, with the trigger named as the guard
also requires — #509 raising the remaining 34, at which point the entry is DELETED rather than
amended, because the divergence it describes will not exist.

The divergence is also less dangerous than the rule's general case: the two postprocessors
share one prefix and agree with each other, the models resolve elsewhere, and the resolver
picks 1.13.0 for all of them today regardless. The floor only forbids something lower.

AND THE PART WORTH THE EXTRA TEST. A DEFERRED_PACKAGES entry is wider than it looks: it also
exempts the package from test_no_dependency_is_declared_without_an_upper_bound. So while this
deferral stands, nothing in the repo would notice `<2.0.0` being dropped from a
views-datafactory line — and that rule exists because an unbounded internal package installs
the next breaking major on the following monthly run. Buying a floor by silently selling a
ceiling is not a trade worth making quietly.

Closed with test_every_datafactory_declaration_keeps_its_upper_bound, which is narrower than
the rule it stands in for — it says nothing about floors, only that a ceiling exists wherever
this package is named — and which outlives the deferral harmlessly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#499): clear stale calibration output on the forecast leg, and agree on the workspace

Two findings from reviewing this branch before merging it. Both are cases of a rule the
scripts already state, not applied where it also holds.

STALE CALIBRATION ARTEFACTS. The forecast leg writes no parquets, so it skipped sections 4
and 5 — but a previous calibration run on the same pod leaves parquet/ and draws/ in the SAME
output directory. The forecast MANIFEST would then sit beside 13 parquets from a different run
type, and the guide's rsync copies the directory, so they come home as this run's output.
pod_run_model.sh already reasons about exactly this twice ("Clear a previous attempt first …
if they happened to sum to 13 the count check would pass while the manifest covered two
different training runs") for the case where the current run DOES produce them. The case where
it produces none is the same hazard, and skipping a stage is not the same as clearing it.

THE TWO SCRIPTS MUST AGREE ON $ROOT. pod_run_fao_delivery.sh honours PODRUN_ROOT so its
preflight can be exercised off a pod; pod_run_model.sh had /workspace hardcoded. The delivery
script reads $ROOT/deliver/<model>/STATUS to decide whether a model may be pooled, so a
relocated workspace in one and not the other means reading a STATUS the other never wrote —
and a missing STATUS is indistinguishable from a model that failed, which would refuse a
healthy roster. Benign on a pod, where both are /workspace; latent, and cheap to remove.

Both guarded. Written before the merge rather than after, which is the only time reviewing
your own diff is worth anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…loor was below the one it delegates to (#527)

* fix(#499): the FAO preflight's disk floor was below the floor it delegates to

Found by a /falsify pass on "we are ready for a 40-lesson forecasting run reaching faoapi",
and it is my defect from #524.

pod_run_fao_delivery.sh demanded 60GB for EIGHT models plus the pooled ensemble, while
pod_run_model.sh — which it delegates to eight times — refuses below 40GB for ONE model
("one model needs ~20GB"). So the orchestrator's preflight could report ready and the
delegated script could then refuse partway through the roster, after hours of paid GPU time.
That is precisely the failure --preflight exists to prevent, introduced into --preflight.

Raised to DISK_FLOOR_GB=80: 40GB transient, which the delegated script enforces per model
regardless of what is written here, plus eight models' retained output and the pool.

The retained component is an ESTIMATE and the comment now says so rather than implying a
measurement. One model's calibration output is ~2.5GB per predictions directory on this
machine, but NO FORECASTING RUN HAS EVER COMPLETED ON THIS ROSTER — rusty_bucket's
forecasting_log.txt is from 2026-07-20 and lists temporary_crane and temporary_fox at
Deployment Status: shadow, a different roster entirely — so the retained size of a forecast
is unmeasured. A forecast has one origin against calibration's thirteen so it should be
smaller, and "should be" is doing real work in that sentence. Trigger for revising the
number: the first completed run, rather than reasoning about it a second time.

Guarded by a test that compares this floor against the delegated script's, so the two cannot
drift apart again — the defect was not the number, it was that nothing related them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

* fix(#499): preflight checked 3 of 9 publish variables — build the config instead of counting

Found by the /falsify pass on the forecasting claim, acting on a suggestion from the
views-pipeline-core session. My preflight from #524 checked three environment variables and
concluded the publish would work. Measured against the real code, PredictionStoreConfig
requires NINE:

  APPWRITE_ENDPOINT, APPWRITE_DATASTORE_PROJECT_ID, APPWRITE_DATASTORE_API_KEY,
  APPWRITE_PROD_FORECASTS_BUCKET_ID, APPWRITE_PROD_FORECASTS_BUCKET_NAME,
  APPWRITE_PROD_FORECASTS_COLLECTION_ID, APPWRITE_PROD_FORECASTS_COLLECTION_NAME,
  APPWRITE_METADATA_DATABASE_ID, APPWRITE_METADATA_DATABASE_NAME

Only the first three are secrets. A missing identifier fails the publish exactly as hard as a
missing secret, so the check would have reported ready with six of nine absent — and the run
would have failed at the publish, after the entire roster trained. The precise failure this
mode exists to prevent, in the mode that exists to prevent it. Twice in one night now, which
is its own signal.

Adding the six missing names would not fix it. The extra can be absent, the endpoint
unreachable, the key expired (#359: 2026-11-17). Preflight now CONSTRUCTS the config and
imports the SDK, which answers the actual question instead of a proxy for it.
views-pipeline-core suggested this and observed it is what #557 argues the manager should do
itself; until it does, doing it by hand costs seconds.

Second finding from the same measurement, and an operator trap: pipeline-core no longer
auto-loads a .env from the working directory (#346, register C-177) — a library reading
whatever .env its caller is standing in is what the Appwrite seam contract §3 forbids. So a
correct /root/.secrets sitting beside the operator is NOT enough; the variables must be
exported. Without that line the failure presents as wrong credentials rather than unexported
ones. Preflight now prints `set -a; . /root/.secrets; set +a`.

Both guarded, and the guards assert the construction and the exported-variables note exist
rather than re-listing nine names — a guard that enumerates the same thing the code does is
the defect that produced this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants